Systems and methods for determining nucleic acids

By using multiple nucleic acid probes and reading sequences, combined with error correction technology, the problem of low throughput in smFISH technology is solved, and efficient detection and high-resolution imaging of mRNAs in multiple genes in cells is achieved.

CN120099143APending Publication Date: 2025-06-06PRESIDENT & FELLOWS OF HARVARD COLLEGE
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510214028.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2015-04-03
Filing Date
2015-07-29
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing single-molecule fluorescence in situ hybridization (smFISH) technology is difficult to efficiently detect mRNA molecules of multiple genes in cells due to low throughput.

Method used

Using multiple nucleic acid probes, the probe contains a target sequence and a multiple reading sequence. By measuring the binding of the probe in the sample, code words are generated, and effective code words are formed through error correction to improve the resolution and efficiency of detection.

Benefits of technology

It realizes efficient detection of mRNAs of multiple genes in cells, improves the resolution and efficiency of detection, and reduces the rate of misidentification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120099143A_ABST
    Figure CN120099143A_ABST
Patent Text Reader

Abstract

The present invention generally relates to systems and methods for imaging or determining nucleic acids, e.g., within cells. In some embodiments, a transcriptome of a cell may be determined. Certain embodiments relate to the determination of nucleic acids, such as mRNA, within a cell at relatively high resolution. In some embodiments, multiple nucleic acid probes may be applied to a sample and their binding within the sample, for example using fluorescence, to determine the location of the nucleic acid probes within the sample. In some embodiments, a codeword may be based on binding of multiple nucleic acid probes, and in some cases, the codeword may define an error correction code to reduce or prevent misidentification of nucleic acids. In some cases, a relatively large number of different targets may be identified using relatively small numbers of markers, such as by using various combination methods.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of CN202010422932.0.

[0002] Related Applications

[0003] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 031,062, entitled “Systems and Methods for Determining Nucleic Acids,” filed by Zhuang et al. on July 30, 2014, U.S. Provisional Patent Application Serial No. 62 / 142,653, entitled “Systems and Methods for Determining Nucleic Acids,” filed by Zhuang et al. on April 3, 2015, and U.S. Provisional Patent Application Serial No. 62 / 050,636, entitled “Probe Library Construction,” filed by Zhuang et al. on September 15, 2014. Each of the above is incorporated herein by reference.

[0004] Government funding

[0005] This invention was made with government support under Grant No. GM096450 awarded by the National Institutes of Health. The government has certain rights in this invention.

[0006] field

[0007] The present invention generally relates to systems and methods for imaging or measuring nucleic acids, for example, within a cell. In some embodiments, the transcriptome of a cell can be measured. background

[0008] Single molecule fluorescence in situ hybridization (smFISH) is a powerful method for detecting individual mRNA molecules in cells. The high detection efficiency and large dynamic range of the method provide fine details to the expression state, spatial distribution and intercellular differences of individual mRNA in cells and intact tissues. Such methods are necessary for understanding many recent insights of gene regulation and expression. However, the basic limitation of smFISH is its low throughput, usually only a few genes at a time. This low throughput is due to the lack of distinguishable probes for labeling cells and the cost of a large number of labeled probes required for efficient staining. Therefore, it is necessary to improve the detection of mRNA molecules.

[0009] Overview

[0010] The present invention generally relates to systems and methods for imaging or determining nucleic acids in, for example, cells. In some embodiments, the transcriptome of the cell can be determined. In some cases, the subject matter of the present invention relates to related products, alternative solutions to specific problems and / or multiple different uses of one or more systems and / or articles.

[0011] In one aspect, the present invention generally relates to compositions. According to one set of embodiments, the composition comprises a plurality of nucleic acid probes, at least some of which comprise a first portion comprising a target sequence and a plurality of reading sequences. In some cases, each nucleic acid probe comprises a first portion comprising a target sequence and a plurality of reading sequences. In some embodiments, the plurality of reading sequences are distributed over the plurality of nucleic acid probes to define an error correction code.

[0012] In another aspect, the invention generally relates to a method. In one set of embodiments, the method comprises the following operations: exposing a sample to a plurality of nucleic acid probes; for each of the nucleic acid probes, determining the binding of the nucleic acid probe within the sample; generating a codeword based on the binding of the nucleic acid probes; and for at least some of the codewords, matching the codeword with a valid codeword, wherein if no match is found, applying error correction to the codeword to form a valid codeword.

[0013] In another set of embodiments, the method comprises the following operations: exposing a sample to a plurality of nucleic acid probes, wherein the nucleic acid probes comprise a first portion comprising a target sequence and a second portion comprising one or more reading sequences, and wherein at least some of the plurality of nucleic acid probes comprise distinguishable nucleic acid probes formed from a combinatorial combination of one or more reading sequences taken from the plurality of reading sequences; and for each nucleic acid probe, determining the binding of the target sequence of the nucleic acid probe within the sample.

[0014] In another set of embodiments, the method includes the following operations: exposing the sample to a plurality of primary nucleic acid probes (also called encoding probes); exposing the plurality of primary nucleic acid probes to a sequence of secondary nucleic acid probes (also called readout probes) and measuring the fluorescence of each secondary nucleic acid probe within the sample; generating code words based on the fluorescence of the secondary nucleic acid probes; and for at least some of the code words, matching the code words with valid code words, wherein if no match is found, applying error correction to the code words to form valid code words.

[0015] In one set of embodiments, the method comprises the following operations: exposing a plurality of primary nucleic acid probes to a sample; and exposing the plurality of nucleic acid probes to a sequence of secondary nucleic acid probes, and determining the fluorescence of each secondary probe within the sample. In some embodiments, at least some of the plurality of secondary nucleic acid probes comprise distinguishable secondary nucleic acid probes formed from a combinatorial combination of one or more read sequences (or readout probe sequences) taken from a plurality of read sequences (or readout probe sequences).

[0016] In another set of embodiments, the method comprises the following operations: exposing cells to multiple nucleic acid probes, exposing the multiple nucleic acid probes to a first secondary probe comprising a first signaling entity, measuring the first signaling entity with an accuracy of better than 500 nm, inactivating the first signaling entity, exposing the multiple nucleic acid probes to a second secondary probe comprising a second signaling entity, and measuring the second signaling entity with an accuracy of better than 500 nm.

[0017] In another set of embodiments, the method includes the following operations: exposing cells to multiple nucleic acid probes, exposing the multiple nucleic acid probes to a first secondary probe comprising a first signaling entity, measuring the first signaling entity with a resolution better than 100 nm, inactivating the first signaling entity, exposing the multiple nucleic acid probes to a second secondary probe comprising a second signaling entity, and measuring the second signaling entity with a resolution better than 100 nm.

[0018] In another set of embodiments, the method includes the following operations: exposing cells to multiple nucleic acid probes, exposing the multiple nucleic acid probes to a first secondary probe comprising a first signaling entity, measuring the first signaling entity using super-resolution imaging technology, inactivating the first signaling entity, exposing the multiple nucleic acid probes to a second secondary probe comprising a second signaling entity, and measuring the second signaling entity using super-resolution imaging technology.

[0019] In certain embodiments, the method comprises the following operations: associating a plurality of targets with a plurality of target sequences and a plurality of code words, wherein the code words comprise a plurality of positions and a value for each position, and the code words form an error checking and / or error correction code space; associating a plurality of distinguishable read sequences with the plurality of code words such that each distinguishable read sequence represents a value for a position within the code word; and forming a plurality of nucleic acid probes, each nucleic acid probe comprising a target sequence and one or more read sequences.

[0020] In addition, in one set of embodiments, the method includes the following operations: associating multiple targets with multiple target sequences and multiple code words, wherein the code words include multiple positions and values ​​for each position, and the code words form an error checking and / or error correction code space; forming multiple nucleic acid probes, each nucleic acid probe comprising a target sequence; and forming a group comprising the multiple nucleic acid probes so that each group of nucleic acid probes corresponds to at least one common value of a position within the code words.

[0021] In another set of embodiments, the method includes the following operations: associating multiple targets with multiple target sequences and multiple code words, wherein the code words contain multiple positions less than the number of targets and wherein each code word is associated with a single target, associating multiple distinguishable read sequences with the multiple code words so that each distinguishable read sequence represents the value of a position within the code word, and forming multiple nucleic acid probes, each nucleic acid probe comprising a target sequence and one or more read sequences.

[0022] In another set of embodiments, the method comprises the following operations: exposing a plurality of nucleic acid probes to a cell, exposing the plurality of nucleic acid probes to a sequence of secondary probes, and measuring the fluorescence of each of the secondary probes in the cell, and based on the fluorescence sequence of each secondary probe, measuring the nucleic acid in the cell.

[0023] In another set of embodiments, the method includes the following operations: associating multiple targets with multiple target sequences and multiple code words, wherein the code words include multiple positions and values ​​for each position, and the code words form an error checking and / or error correction code space; forming multiple nucleic acid probes, each nucleic acid probe comprising a target sequence; and forming a group comprising the multiple nucleic acid probes so that each group of nucleic acid probes corresponds to at least one common value for a position within the code words.

[0024] In another set of embodiments, the method includes the following operations: exposing cells to multiple nucleic acid probes, exposing the multiple nucleic acid probes to a first secondary probe comprising a first signaling entity, determining the first signaling entity using super-resolution imaging technology, inactivating the first signaling entity, exposing the multiple nucleic acid probes to a second secondary probe comprising a second signaling entity, and determining the second signaling entity using super-resolution imaging technology.

[0025] In another set of embodiments, the method comprises the following operations: exposing cells to multiple nucleic acid probes, exposing the multiple nucleic acid probes to a first secondary probe comprising a first signaling entity, measuring the first signaling entity with an accuracy better than 500 nm, inactivating the first signaling entity, exposing the multiple nucleic acid probes to a second secondary probe comprising a second signaling entity, and measuring the second signaling entity with an accuracy better than 500 nm.

[0026] In another set of embodiments, the method comprises the following operations: exposing cells to multiple nucleic acid probes, exposing the multiple nucleic acid probes to a first secondary probe comprising a first signaling entity, measuring the first signaling entity with a resolution better than 100 nm, inactivating the first signaling entity, exposing the multiple nucleic acid probes to a second secondary probe comprising a second signaling entity, and measuring the second signaling entity using super-resolution imaging technology.

[0027] In another set of embodiments, the method includes the following operations: associating multiple nucleic acid targets with multiple target sequences and multiple code words, wherein the code words include multiple positions and values ​​for each position, and the code words form an error checking and / or error correction code; associating a unique reading sequence with each possible value for each position in the code word, wherein the reading sequence is taken from a set of orthogonal sequences that have limited homology with each other and with nucleic acid species in the sample; forming multiple primary nucleic acid probes, each primary nucleic acid probe comprising a target sequence that uniquely binds to a nucleic acid target and one or more reading sequences; forming multiple secondary nucleic acid probes comprising a signaling entity and a sequence complementary to one of the reading sequences; exposing the sample to the primary nucleic acid probes so that the nucleic acid probes hybridize with nucleic acid targets in the sample; exposing the primary nucleic acid probes in the sample to secondary nucleic acid probes so that the secondary nucleic acid probes hybridize with the reading sequences on at least some of the primary nucleic acid probes; imaging the sample; and repeating the exposing and imaging steps one or more times, using different secondary nucleic acid probes for at least some of the repetitions.

[0028] According to another set of embodiments, the method includes the following operations: associating multiple nucleic acid targets with multiple target sequences and multiple code words, wherein the code words include multiple positions and values ​​for each position, and the code words form an error checking and / or error correction code space; forming multiple nucleic acid probes comprising a signaling entity and a target sequence that specifically binds to one of the nucleic acid targets; grouping the nucleic acid probes into multiple probe libraries, wherein each probe library corresponds to a specific value of a unique position within the code word; exposing a sample to one of the probe libraries; imaging the sample; and repeating the exposing and imaging steps one or more times, using different probe libraries for at least some of the repetitions.

[0029] In another aspect, the invention includes methods of making one or more of the embodiments described herein. In another aspect, the invention includes methods of using one or more of the embodiments described herein.

[0030] When considered in conjunction with the accompanying drawings, other advantages and novel features of the present invention will become apparent from the following detailed description of various non-limiting embodiments of the present invention. In the case where this specification and the documents incorporated by reference include conflicting and / or inconsistent disclosures, this specification shall prevail. If two or more documents incorporated by reference include conflicting and / or inconsistent disclosures, the document with the later effective date shall prevail. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Non-limiting embodiments of the present invention will be described by way of example with reference to the accompanying drawings, which are schematic and are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown is generally represented by a single numeral. For the sake of clarity, where illustration is not necessary for an ordinary technician in the field to understand the present invention, not every component is marked in each figure, nor is every component of each embodiment of the present invention shown. In the drawings:

[0032] Figures 1A-1C The coding scheme of the nucleic acid probe in certain embodiments of the present invention is illustrated;

[0033] Figure 2A-2G Illustrate the determination of mRNA in cells in some embodiments of the present invention;

[0034] Figure 3A-3B The determination of nucleic acids according to various embodiments of the present invention is illustrated;

[0035] Figures 4A-4B is a non-limiting example of a plurality of read sequences distributed among a population of different nucleic acid probes according to certain embodiments of the present invention;

[0036] Figures 5A-5E The determination of nucleic acid according to another embodiment of the present invention is illustrated;

[0037] Figures 6A-6H illustrates the simultaneous determination of multiple nucleic acid species in a cell in certain embodiments of the present invention;

[0038] Figures 7A-7E showing expression noise of genes and co-variation of expression between different genes determined according to some embodiments of the present invention;

[0039] Figures 8A-8E Illustrate the spatial distribution of RNA in cells determined according to one embodiment of the present invention;

[0040] Figures 9A-9C Illustrate the simultaneous determination of multiple nucleic acid species in a cell in another embodiment of the present invention;

[0041] Figures 10A-10B showing the expression between different genes determined according to another embodiment of the present invention;

[0042] Fig.11 is a schematic depiction of a combination-mark according to another embodiment of the present invention;

[0043] FIG. 12 shows a schematic description of the Hamming distance in another embodiment of the present invention.

[0044] Fig.13 The generation of a probe library in another embodiment of the present invention is illustrated.

[0045] Figures 14A-14B An example of a fluorescent spot assay in another embodiment of the present invention is described;

[0046] Figures 15A-15B By way of example, in another embodiment of the present invention, error correction facilitates RNA detection;

[0047] Figures 16A-16B A representation showing misidentification rates and calling rates in one embodiment of the present invention; Fig. 16C and 16D Displays the 1-->0 error for each hybridization round ( Fig. 16C ) and 0-->1 error ( Fig.16D )’s average error rate; Fig.16E The call rate of each RNA species estimated from the 1-->0 and 0-->1 error rates is shown;

[0048] Figures 17A-17DShowing the representation of the false positive rate and the call rate in another embodiment of the present invention;

[0049] Figures 18A-18C showing a comparison of experiments according to one embodiment of the present invention; and

[0050] Figures 19A-19D Decoding and error estimation in another embodiment of the present invention is illustrated.

[0051] Sequence Description

[0052] SEQ ID NO: 1 is: GTTGGCGACGAAAGCACTGCGATTGGAACCGTCCCAAGCGTTGCGCTTAATGGATCATCAATTTTGTCTCACTACGACGGTCAATCGCGCTGCATACTTGCGTCGGTCGGACAAACGAGG;

[0053] SEQ ID NO:2 is CGCAACGCTTGGACGGTTCCAATCGGATC;

[0054] SEQ ID NO: 3 is CGAATGCCTGGCCTCGAACGAACGATAGC;

[0055] SEQ ID NO: 4 is ACAAATCCGACCAGATCGGACGATCATGGG;

[0056] SEQ ID NO:5 is CAAGTATGCAGCGCGATTGACCGTCTCGTT;

[0057] SEQ ID NO: 6 is TGCGTCGTCTGGCTAGCACGGCACGCAAAT;

[0058] SEQ ID NO:7 is AAGTCGTACGCCGATGCGCAGCAATTCACT;

[0059] SEQ ID NO:8 is CGAAACATCGGCCACGGTCCCGTTGAACTT;

[0060] SEQ ID NO:9 is ACGAATCCACCGTCCAGCGCGTCAAACAGA;

[0061] SEQ ID NO: 10 is CGCGAAATCCCCGTAACGAGCGTCCCTTGC;

[0062] SEQ ID NO:11 is GCATGAGTTGCCTGGCGTTGCGACGACTAA;

[0063] SEQ ID NO:12 is CCGTCGTCTCCGGTCCACCGTTGCGCTTAC;

[0064] SEQ ID NO:13 is GGCCAATGGCCCAGGTCCGTCACGCAATTT;

[0065] SEQ ID NO:14 is TTGATCGAATCGGAGCGTAGCGGAATCTGC;

[0066] SEQ ID NO:15 is CGCGCGGATCCGCTTGTCGGGAACGGATAC;

[0067] SEQ ID NO:16 is GCCTCGATTACGACGGATGTAATTCGGCCG;

[0068] SEQ ID NO:17 is GCCCGTATTCCCGCTTGCGAGTAGGGCAAT

[0069] SEQ ID NO:18 is GTTGGTCGGCACTTGGGTGC;

[0070] SEQ ID NO:19 is CGATGCGCCAATTCCGGTTC;

[0071] SEQ ID NO:20 is CGCGGGCTATATGCGAACCG;

[0072] SEQ ID NO:21 is TAATACGACTCACTATAGGGAAAGCCGGTTCATCCGGTGG;

[0073] SEQ ID NO:22 is TAATACGACTCACTATAGGGTGATCATCGCTCGCGGGTTG;

[0074] SEQ ID NO:23 is TAATACGACTCACTATAGGGCGTGGAGGGCATACAACGC;

[0075] SEQ ID NO:24 is CGCAACGCTTGGGACGGTTCCAATCGGATC / 3Cy5Sp / ;

[0076] SEQ ID NO:25 is CGAATGCTCTGGCCTCGAACGAACGATAGC / 3Cy5Sp / ;

[0077] SEQ ID NO:26 is ACAAATCCGACCAGATCGGACGATCATGGG / 3Cy5Sp / ;

[0078] SEQ ID NO:27 is CAAGTATGCAGCGCGATTGACCGTCTCGTT / 3Cy5Sp / ;

[0079] SEQ ID NO:28 is GCGGGAAGCACGTGGATTAGGGCATCGACC / 3Cy5Sp / ;

[0080] SEQ ID NO:29 is AAGTCGTACGCCGATGCGCAGCAATTCACT / 3Cy5Sp / ;

[0081] SEQ ID NO:30 is CGAAACATCGGCCACGGTCCCGTTGAACTT / 3Cy5Sp / ;

[0082] SEQ ID NO:31 is ACGAATCCACCGTCCAGCGCGTCAAACAGA / 3Cy5Sp / ;

[0083] SEQ ID NO:32 is CGCGAAATCCCCGTAACGAGCGTCCCTTGC / 3Cy5Sp / ;

[0084] SEQ ID NO:33 is GCATGAGTTGCCTGGCGTTGCGACGACTAA / 3Cy5Sp / ;

[0085] SEQ ID NO:34 is CCGTCGTCTCCGGTCCACCGTTGCGCTTAC / 3Cy5Sp / ;

[0086] SEQ ID NO:35 is GGCCAATGGCCCAGGTCCGTCACGCAATTT / 3Cy5Sp / ;

[0087] SEQ ID NO:36 is TTGATCGAATCGGAGCGTAGCGGAATCTGC / 3Cy5Sp / ;

[0088] SEQ ID NO:37 is CGCGCGGATCCGCTTGTCGGGAACGGATAC / 3Cy5Sp / ;

[0089] SEQ ID NO:38 is GCCTCGATTACGACGGATGTAATTCGGCCG / 3Cy5Sp / ; and

[0090] SEQ ID NO: 39 is GCCCGTATTCCCGCTTGCGAGTAGGGCAAT / 3Cy5Sp / .

[0091] Figure 20-1 to Figure 20-8 Two different codebooks for the 140-gene experiment are shown. The specific codewords of the 16-bit MHD4 code assigned to each RNA species were studied in two shuffles of the 140-gene experiment. The "Gene" column contains the name of the gene. The "Codeword" column contains the specific binary word assigned to each gene.

[0092] Details

[0093] The present invention generally relates to systems and methods for imaging or measuring nucleic acids in, for example, cells. In some embodiments, the transcriptome of cells can be measured. Certain embodiments relate to measuring nucleic acids in cells, such as mRNA, with relatively high resolution. In some embodiments, multiple nucleic acid probes can be applied to samples, and for example, their combination in samples is measured using fluorescence to measure the position of nucleic acid probes in samples. In some embodiments, code words can be based on the combination of multiple nucleic acid probes, and in some cases, the code words can define error correction codes to reduce or prevent the misidentification of nucleic acids. In some cases, relatively small amounts of labels can be used, for example, by using various combination methods to identify relatively large amounts of different targets.

[0094] Two exemplary methods are now discussed. However, it should be understood that these methods are presented in an illustrative and non-limiting manner; other aspects and embodiments are further discussed in detail herein. In one exemplary method, a primary probe (also referred to as a coding probe) and a secondary probe (also referred to as a readout probe) are used, wherein the primary probe encodes a "code word" and binds to a target nucleic acid in a sample, and the secondary probe is used to read out the code word of the primary probe. In another exemplary method, a plurality of different primary probes containing a code word are divided into as many separate libraries as there are positions in the code word, so that each primary probe library corresponds to a certain value in a certain position of the code word (e.g., a "1" in the first position as in "1001").

[0095] Now refer to Figure 3ADescribe the first example.As will be discussed in more detail below, in other embodiments, other configurations may also be used.In this first example, a series of nucleic acid probes are used to for example qualitatively or quantitatively measure the nucleic acid in a cell or other sample.For example, nucleic acid may be identified as existing or not existing, and / or may measure the number or concentration of some nucleic acid in a cell or other sample.In some cases, may be with relatively high resolution, and in some cases, may be better than the resolution of visible wavelength to measure the position of probe in a cell or other sample.

[0096] This example generally relates to, for example, detecting nucleic acid in a cell or other sample spatially with a relatively high resolution. For example, the nucleic acid can be mRNA or other nucleic acid as described herein. In one group of embodiments, the nucleic acid in the cell can be measured by delivering or applying a nucleic acid probe to the cell. In some cases, by using a combined approach, a relatively large amount of nucleic acid can be measured using a relatively small amount of different labels on the nucleic acid probe. Therefore, for example, due to the simultaneous combination of nucleic acid probes and different nucleic acids in the sample, a relatively large amount of nucleic acid in the sample can be measured using a relatively small amount of experiments.

[0097] In one group of embodiments, a primary nucleic acid probe colony capable of binding to a nucleic acid suspected to be present in a cell is applied to a cell (or other sample). Subsequently, secondary nucleic acid probes that can bind to some primary nucleic acids or otherwise interact with some primary nucleic acids are sequentially added and, for example, measured using imaging techniques such as fluorescence microscopy (e.g., conventional fluorescence microscopy), STORM (random optical reconstruction microscopy) or other imaging techniques. After imaging, the secondary nucleic acid probes are inactivated or removed, and different secondary nucleic acid probes are added to the sample. This can be repeated multiple times with a variety of different secondary nucleic acid probes. The binding patterns of various secondary nucleic acid probes can be used to measure the primary nucleic acid probes at the position in a cell or other sample, which can be used to measure the mRNA or other nucleic acids present.

[0098] For example, Figure 3A As shown, a nucleic acid population 10 (here represented by nucleic acids 11, 12, and 13) within a cell can be exposed to a primary nucleic acid probe population 20 (including probes 21 and 22). The primary nucleic acid probes can include, for example, a target sequence that can recognize a nucleic acid (e.g., a sequence within nucleic acid 11). Probes 21 and 22 can contain the same or different targeting sequences, which can bind to or hybridize with the same or different nucleic acids. As an example, Figure 3AAs shown, probe 21 comprises a first targeting sequence 25 that targets the probe to nucleic acid 11, while probe 22 comprises a second targeting sequence 26 that is different from the first targeting sequence 25 and that targets the probe to nucleic acid 12. The target sequence may be substantially complementary to at least a portion of the target nucleic acid, and sufficient target sequence may be present so that specific binding of the nucleic acid probe to the target nucleic acid can occur.

[0099] The primary nucleic acid probe 20 may also contain one or more "reading" sequences. In this example, two such reading sequences are used, but in other embodiments, there may be 1, 3, 4 or more reading sequences in the primary nucleic acid probe. The reading sequences may all be independently the same or different. In addition, in one set of embodiments, different nucleic acid probes may use one or more common reading sequences. For example, more than one reading sequence may be combinatorially combined on different nucleic acid probes, thereby generating a relatively large number of different nucleic acid probes that can be individually identified even if only a relatively small number of reading sequences are used. Thus, for example, in Figure 3A In the example, probe 21 contains reading sequences 27 and 29, and probe 22 contains reading sequences 27 and 28, wherein the two reading sequences 27 are identical and different from the reading sequences 28 and 29.

[0100] After the primary nucleic acid probe 20 has been introduced into the sample and allowed to interact with nucleic acids 11, 12 and 13, one or more secondary nucleic acid probes 30 can be applied to the sample to determine the primary nucleic acid probe. The secondary nucleic acid probe may contain a recognition sequence that can identify one of the reading sequences present in the primary nucleic acid probe population. For example, the recognition sequence may be substantially complementary to at least a portion of the reading sequence so that the secondary nucleic acid probe can be combined or hybridized with the corresponding primary nucleic acid probe. For example, in this example, the recognition sequence 35 can recognize the reading sequence 27. In addition, the secondary nucleic acid probe may contain one or more signal transduction entities 33. For example, the signal transduction entity can be a fluorescent entity attached to the probe, or certain nucleic acid sequences that can be determined in some way. More than one secondary sequence can be used, for example, in sequence. For example, as shown in the figure, the initial secondary probe 30 can be removed (for example, as described below), and a new secondary probe 31 can be added, which contains a recognition sequence 36 that can recognize the reading sequence 28 and one or more signal transduction entities. This can also be repeated multiple times, for example, to determine the reading sequence 29 or other reading sequences that may exist.

[0101] The position of the secondary nucleic acid probes 30, 31, etc. can be determined by measuring the signaling entity 33. For example, if the signaling entity is fluorescent, fluorescence microscopy can be used to measure the signaling entity. In some embodiments, the signaling entity can be measured using imaging of the sample at a relatively high resolution, and in some cases, super-resolution imaging techniques (e.g., resolution better than the wavelength of visible light or the diffraction limit of light) can be used. Examples of super-resolution imaging techniques include STORM or other techniques as discussed herein. In some cases, for example, when using certain super-resolution imaging techniques such as STORM, more than one sample image can be acquired.

[0102] More than one type of secondary nucleic acid probe can be applied to cells or other samples. For example, a first secondary nucleic acid probe that can recognize a first reading sequence can be applied, which can then be inactivated or removed by itself or its attached signaling entity, and a second secondary nucleic acid probe that can recognize a second reading sequence can be applied. This process can be repeated multiple times, using different secondary nucleic acid probes each time, for example, to determine the reading sequences present in various primary nucleic acid probes. Therefore, the primary nucleic acid in the sample can be determined based on the binding pattern of the secondary nucleic acid probe.

[0103] For example, a first location within a cell or other sample may show binding of a first secondary probe and a third secondary probe, but not binding of a second or fourth secondary probe, while a second location may show different binding patterns for the various secondary probes. The primary nucleic acid probe to which the secondary probes are capable of binding or hybridizing can be determined by considering the binding patterns of the various secondary probes. For example, referring to Figure 3A , if the first secondary probe is capable of determining the reading sequence 27, the second secondary probe is capable of determining the reading sequence 28, and the third secondary probe is capable of determining the reading sequence 29, then the primary nucleic acid 25 can be determined by the binding of the first and third secondary probes (but not the second secondary probe), and the primary nucleic acid 26 can be determined by the binding of the first and second secondary probes (but not the third secondary probe). Similarly, if it is known that the first probe 21 contains the target sequence 25 and the second probe 22 contains the target sequence 26, the nucleic acids 11 and 12 in the sample can also be determined spatially based on the binding patterns of the various secondary nucleic acid probes. In addition, it should be noted that since there is more than one reading sequence on the primary nucleic acid probe, even if the first probe 21 and the second probe 22 contain a common reading sequence (reading sequence 27), these probes in the sample can be distinguished due to the different binding patterns of the various secondary nucleic acid probes.

[0104] In certain embodiments, this pattern of binding or hybridization of the secondary nucleic acid probes can be converted into a "code word". In this example, for example, for the first probe 21 and the second probe 22, the code words are "101" and "110", respectively, where a value of 1 indicates binding and a value of 0 indicates no binding. In other embodiments, the code words may also have a longer length; only three probes are shown here for clarity. The code words can be directly related to the specific target nucleic acid sequence of the primary nucleic acid probe. Therefore, different primary nucleic acid probes can match certain code words, which can then be used to identify different targets of the primary nucleic acid probe based on the binding pattern of the secondary probe, even if in some cases there is overlap in the reading sequences of different secondary probes, for example, Figure 3A However, if there is no obvious binding (e.g., for nucleic acid 13), the code word would be "000" in this example.

[0105] In some embodiments, the values ​​in each code word can also be assigned in different ways. For example, a value of 0 can represent binding, while a value of 1 represents no binding. Similarly, a value of 1 can represent the binding of a secondary nucleic acid probe to a type of signaling entity, while a value of 0 can represent the binding of a secondary nucleic acid probe to another type of distinguishable signaling entity. These signaling entities can be distinguished, for example, by fluorescence of different colors. In some cases, the values ​​in the code word do not have to be limited to 0 and 1. The values ​​can also be derived from larger alphabets such as ternary (e.g., 0, 1, and 2) or quaternary (e.g., 0, 1, 2, and 3) systems. Each different value can be represented, for example, by a different distinguishable signaling entity, including (in some cases) a value that can be represented by the absence of a signal.

[0106] The codewords for each target may be assigned sequentially, or may be assigned randomly. Figure 3A As shown, the first nucleic acid target can be designated as 101, and the second nucleic acid target can be designated as 110. In addition, in some embodiments, an error detection system or an error correction system, such as a Hamming system, a Golay code, or an extended Hamming system (or a SECDED system, i.e., single error correction, double error detection) can be used to assign code words. In general, such systems can be used to identify the position where an error has occurred, and in some cases, such systems can also be used for error correction and determine what the correct code word should be. For example, a code word such as 001 can be detected as invalid and corrected to 101 using such a system, for example, if 001 has not been assigned to a different target sequence before. A variety of different error correction codes can be used, many of which have been previously developed for use in the computer industry; however, such error correction systems are not typically used in biological systems. Other examples of such error correction codes are discussed in more detail below.

[0107] It should also be understood that in some cases it is not necessary to use all possible code words in the code. For example, in some embodiments, unused code words can be used as negative controls. Similarly, in some embodiments, some code words can be omitted because they are more prone to error in measurement than other code words. For example, in some implementations, reading a code word with a value of more "1" may be more prone to error than reading a code word with a value of less "1".

[0108] It should be understood that the above description is an example of one embodiment of the present invention, and primary and secondary nucleic acid probes are not necessary in all embodiments. For example, in some embodiments, a series of nucleic acid probes containing signaling entities are used to measure nucleic acids in cells or other samples, without necessarily requiring secondary probes.

[0109] For example, now go to Figure 3B In this example, nucleic acids 11, 12, and 13 are exposed to different rounds of probes 21, 22, 23, 24, etc. These probes can each contain a target sequence that can recognize a nucleic acid (e.g., a sequence within nucleic acid 11 or 12). These probes can each target the same nucleic acid, but different regions of the nucleic acid. In addition, some or all of the probes can contain one or more signaling entities, such as signaling entity 29 on probe 21. For example, the signaling entity can be a fluorescent entity attached to the probe, or a nucleic acid sequence that can be measured in some way.

[0110] The first round of probes (e.g., probe 21 and probe 22) can be applied to cells or other samples. Probe 21 can be bound to nucleic acid 11 through target sequence 25. This binding can be determined by measuring signaling entity 29. For example, if the signaling entity is fluorescent, fluorescence microscopy can be used to, for example, spatially measure the signaling entity within a cell or other sample. In some but not all embodiments, imaging of the sample can be used at a relatively high resolution to measure the signaling entity, and in some cases, super-resolution imaging techniques can be used. In addition, different probes may also be present; for example, probe 22 containing target sequence 26 can bind to nucleic acid 12 and be measured by signaling entity 29 within probe 22. These can occur, for example, sequentially or simultaneously. Optionally, probes 21 and 22 may also be removed or inactivated, for example, between the application of different rounds of probes.

[0111] Next, a second round of probes (e.g., probe 23) is applied to the sample. In this example, probe 23 is able to bind to nucleic acid 11 through the targeted region, although there is no probe that can bind to nucleic acid 12 in the second round. As discussed above, the binding of the probes is allowed to occur, and the determination of the binding can be performed by signaling entities. These signaling entities can be the same or different from the probes of the first round. This process can be repeated any number of times with different probes. For example, Figure 3B As shown, round 2 contains a probe capable of binding nucleic acid 11, while round 3 contains a probe capable of binding nucleic acid 12.

[0112] In certain embodiments, the combination or hybridization of each round of nucleic acid probes can be converted into a "code word". In this example, by using probes 21, 22, 23 and 24, code words 101 or 110 can be formed, wherein 1 represents combination, 0 represents no combination, and the first position corresponds to the combination of probes 21 or 22, while the second position corresponds to the combination of probe 22, and the third position corresponds to the combination of probe 24. The code word of 000 will represent no combination, for example, as shown for nucleic acid 13 in this example. By designing suitable nucleic acid probes, code words can be directly related to the specific target nucleic acid sequence of nucleic acid probes. Therefore, for example, 110 can correspond to the first target nucleic acid 12 (for example, the first and second round nucleic acid probes containing probes capable of targeting nucleic acid 11, and these probes can target the same or different regions of nucleic acid 11), and 101 can correspond to the second target nucleic acid (for example, the first and third round nucleic acid probes containing probes capable of targeting nucleic acid 12, and these probes can target the same or different regions of nucleic acid 12). In addition, it should be noted that each round of detection can contain the same or different signaling entities as other probes in the same round and / or other probes in different rounds. For example, in one set of embodiments, only one signaling entity is used in all rounds of probes.

[0113] Similar to the above, the codewords for each target may be assigned sequentially, or may be assigned randomly. In some embodiments, an error detection or error correction system, such as a Hamming system, Golay code, or an extended Hamming system or a SECDED system (single error correction, double error detection), may be used to assign codewords within the code space. Generally, such error correction systems may be used to identify where errors occur, and in some cases, such systems may also be used to correct errors and determine what the correct codeword should be.

[0114] Similar to above, in certain embodiments, a value at each position in a codeword can be arbitrarily assigned to binding or non-binding of a probe comprising more than one distinguishable signaling entity.

[0115] In some cases, nucleic acid probes can be formed into "pools" or groups of nucleic acids that share common characteristics. For example, probes for all targets having a codeword containing a 1 in the first position (e.g., 110 and 101 but not 011) can comprise one pool, while probes for all targets containing a 1 in the second position (e.g., 110 and 011 but not 101) can comprise another pool. See also Figure 1C In some cases, nucleic acid probes can be members of more than one group or library. In addition to target sequences, reading sequences and / or signaling entities, members of a nucleic acid library may also contain features that allow them to be distinguished from other groups. These features may be short nucleic acid sequences that are used for amplification, generation or separation of these sequences. The nucleic acid probes of each group may be applied to a sample in sequence, for example as described herein.

[0116] Thus, in some aspects, the present invention generally relates to systems and methods for determining nucleic acids in cells or other samples. Samples can include cell cultures, cell suspensions, biological tissues, biopsies, organisms, etc. Samples can also be cell-free, but still contain nucleic acids. If the sample contains cells, the cells can be human cells or any other suitable cells, such as mammalian cells, fish cells, insect cells, plant cells, etc. In some cases, more than one cell can be present.

[0117] The nucleic acid to be determined can be, for example, DNA, RNA, or other nucleic acids present in a cell (or other sample). The nucleic acid can be endogenous to the cell, or added to the cell. For example, the nucleic acid can be viral or artificially produced. In some cases, the nucleic acid to be determined can be expressed by the cell. In some embodiments, the nucleic acid is RNA. RNA can be coding and / or non-coding RNA. Non-limiting examples of RNA that can be studied in cells include mRNA, siRNA, rRNA, miRNA, tRNA, lncRNA, snoRNA, snRNA, exRNA, piRNA, etc. In some embodiments, for example, at least some of the multiple oligonucleotides are complementary to a portion of a specific chromosome sequence (e.g., a human chromosome).

[0118] In some cases, a significant portion of the nucleic acids in a cell can be studied. For example, in some cases, sufficient RNA present in a cell can be assayed to generate a partial or complete transcriptome of the cell. In some cases, at least 4 types of mRNA are assayed in a cell, and in some cases, at least 3, at least 4, at least 7, at least 8, at least 12, at least 14, at least 15, at least 16, at least 22, at least 30, at least 31, at least 32, at least 50, at least 63, at least 64, at least 72, at least 75, at least 100, at least 127, at least 128, at least 140, at least 255, at least 256, at least 500, at least 630, at least 640, at least 720, at least 750, at least 1000, at least 1270, at least 1280, at least 1400, at least 2550, at least 2560, at least 50 ... , at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 4,000, at least 5,000, at least 7,500, at least 10,000, at least 12,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000 or at least 100,000 types of mRNA.

[0119] In some cases, the transcriptome of the cell can be determined. It should be understood that the transcriptome generally includes all RNA molecules produced in the cell, not just mRNA. Therefore, for example, the transcriptome can also include rRNA, tRNA, siRNA, etc. In some embodiments, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90% or 100% of the transcriptome of the cell can be determined.

[0120] The determination of one or more nucleic acids in a cell or other sample can be qualitative and / or quantitative. In addition, the determination can also be spatial, for example, the position of the nucleic acid in a cell or other sample can be determined in two or three dimensions. In some embodiments, the position, quantity and / or concentration of the nucleic acid in a cell (or other sample) can be determined.

[0121] In some cases, a significant portion of the genome of the cell can be assayed. The genomic segments assayed can be contiguous or interspersed across the genome. For example, in some cases, at least 4 genomic segments are assayed within a cell, and in some cases, at least 3, at least 4, at least 7, at least 8, at least 12, at least 14, at least 15, at least 16, at least 22, at least 30, at least 31, at least 32, at least 50, at least 63, at least 64, at least 72, at least 75, at least 100, at least 127, at least 128, at least 140, at least 255, at least 256, at least 257, at least 261, at least 262, at least 263, at least 264, at least 265, at least 266, at least 267, at least 268, at least 269, at least 270, at least 271, at least 272, at least 273, at least 274, at least 275, at least 276, at least 277, at least 278, at least 279, at least 280, at least 281, at least 282, at least 283, at least 284, at least 285, at least 286, at least 287, at least 288, at least 289, at least 300, at least 300, at least 301, At least 500, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 4,000, at least 5,000, at least 7,500, at least 10,000, at least 12,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000, or at least 100,000 genomic fragments.

[0122] In some cases, the entire genome of the cell can be determined. It should be understood that the genome generally includes all DNA molecules produced in the cell, not just the chromosomal DNA. Therefore, for example, in some cases, the genome can also include mitochondrial DNA, chloroplast DNA, plasmid DNA, etc. In some embodiments, at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90% or 100% of the cell genome can be determined.

[0123] As discussed herein, a variety of nucleic acid probes can be used to measure one or more nucleic acids in cells or other samples. The probe may include nucleic acids such as DNA, RNA, LNA (locked nucleic acid), PNA (peptide nucleic acid) or a combination thereof (or an entity that can, for example, specifically hybridize with nucleic acids). In some cases, additional components may also be present in the nucleic acid probe, for example, as described below. Any suitable method may be used to introduce the nucleic acid probe into a cell.

[0124] For example, in some embodiments, cells are fixed before introducing nucleic acid probes, for example to maintain the position of nucleic acid in the cell. The technology used to fix cells is known to those of ordinary skill in the art. As a non-limiting example, cells can be fixed using chemicals such as formaldehyde, paraformaldehyde, glutaraldehyde, ethanol, methanol, acetone, acetic acid, etc. In one embodiment, cells can be fixed using an organic solvent (HOPE) mediated by Hepes-glutamic acid buffer.

[0125] Any suitable method can be used to introduce nucleic acid probes into cells (or other samples). In some cases, the cells can be permeabilized sufficiently so that the nucleic acid probes can be introduced into the cells by flowing a fluid containing the nucleic acid probes around the cells. In some cases, the cells can be permeabilized sufficiently as part of a fixation process; in other embodiments, the cells can be permeabilized by exposing the cells to certain chemicals such as ethanol, methanol, Triton, etc. In addition, in some embodiments, techniques such as electroporation or microinjection can be used to introduce nucleic acid probes into cells or other samples.

[0126] Certain aspects of the present invention generally relate to nucleic acid probes introduced into cells (or other samples). The probe may include any of a variety of entities that can be hybridized with nucleic acids (e.g., DNA, RNA, LNA, PNA, etc., usually by Watson-Crick base pairing). Nucleic acid probes generally contain a target sequence that can bind (in some cases, specifically) at least a portion of a target nucleic acid. When introduced into a cell or other system, the target system may be able to bind to a specific target nucleic acid (e.g., mRNA or other nucleic acids discussed herein). In some cases, a signaling entity (e.g., as discussed below) and / or by using a secondary nucleic acid probe that can bind to a nucleic acid probe (i.e., a primary nucleic acid probe) is used to determine the nucleic acid probe. The determination of such nucleic acid probes is discussed in detail below.

[0127] In some cases, more than one type of (primary) nucleic acid probe may be applied to a sample, e.g., simultaneously. For example, there may be at least 2, at least 5, at least 10, at least 25, at least 50, at least 75, at least 100, at least 300, at least 1,000, at least 3,000, at least 10,000, or at least 30,000 distinguishable nucleic acid probes, which may be applied to a sample, e.g., simultaneously or sequentially.

[0128] The target sequence can be located anywhere within the nucleic acid probe (or primary nucleic acid probe or encoding nucleic acid probe). The target sequence may contain a region that is substantially complementary to a portion of the target nucleic acid. In some cases, the portion may be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% complementary. In some cases, the length of the target sequence may be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400 or at least 450 nucleotides. In some cases, the length of the target sequence can be no more than 500, no more than 450, no more than 400, no more than 350, no more than 300, no more than 250, no more than 200, no more than 175, no more than 150, no more than 125, no more than 100, no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, no more than 35, no more than 30, no more than 20, or no more than 10 nucleotides. Combinations of any of these are also possible, for example, the target sequence can have a length of 10 to 30 nucleotides, 20 to 40 nucleotides, 5 to 50 nucleotides, 10 to 200 nucleotides, or 25 to 35 nucleotides, 10 to 300 nucleotides, etc. Typically, complementarity is determined based on Watson-Crick nucleotide base pairing.

[0129] The target sequence of a (primary) nucleic acid probe can be determined with reference to a target nucleic acid suspected to be present in a cell or other sample. For example, the sequence of a protein can be used to determine the target nucleic acid for the protein by measuring the nucleic acid expressed to form the protein. In some cases, only a portion of the nucleic acid encoding the protein is used, for example, a portion having a length as described above. In addition, in some cases, more than one target sequence that can be used to identify a specific target can be used. For example, a variety of probes that can be combined or hybridized with different regions of the same target can be used sequentially and / or simultaneously. Hybridization generally refers to an annealing process, by which complementary single-stranded nucleic acids are associated to form double-stranded nucleic acids by Watson-Crick nucleotide base pairing (e.g., hydrogen bonding, guanine-cytosine and adenine-thymine).

[0130] In some embodiments, nucleic acid probes, such as elementary nucleic acid probes, may also comprise one or more "reading" sequences. However, it should be appreciated that reading sequences are not essential in all cases. In some embodiments, nucleic acid probes may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more, 20 or more, 32 or more, 40 or more, 50 or more, 64 or more, 75 or more, 100 or more, 128 or more reading sequences. The reading sequence may be located at any position in the nucleic acid probe. If there is more than one reading sequence, the reading sequence may be located adjacent to each other, and / or dispersed with other sequences.

[0131] The read sequence, if present, can have any length. If more than one read sequence is used, the read sequences can independently have the same or different lengths. For example, the read sequence can be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. In some cases, the length of the read sequence can be no more than 500, no more than 450, no more than 400, no more than 350, no more than 300, no more than 250, no more than 200, no more than 175, no more than 150, no more than 125, no more than 100, no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, no more than 35, no more than 30, no more than 20, or no more than 10 nucleotides. Combinations of any of these are also possible, for example, the read sequence can have a length of 10 to 30 nucleotides, 20 to 40 nucleotides, 5 to 50 nucleotides, 10 to 200 nucleotides, or 25 to 35 nucleotides, 10 to 300 nucleotides, etc.

[0132] In some embodiments, the reading sequence can be arbitrary or random. In some cases, the reading sequence is selected to reduce or minimize the homology with other components of the cell or other samples, for example, so that the reading sequence itself is not combined or hybridized with other nucleic acids suspected to be present in the cell or other samples. In some cases, the homology can be less than 10%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2% or less than 1%. In some cases, there can be less than 20 base pairs, less than 18 base pairs, less than 15 base pairs, less than 14 base pairs, less than 13 base pairs, less than 12 base pairs, less than 11 base pairs or less than 10 base pairs of homology. In some cases, the base pairs are sequential.

[0133] In one set of embodiments, a population of nucleic acid probes may contain a number of read sequences, which in some cases may be less than the number of targets of the nucleic acid probes. One of ordinary skill in the art will recognize that if there is one signaling entity and n read sequences, then typically 2 can be uniquely identified. n -1 different nucleic acid target. However, not all possible combinations need to be used. For example, a nucleic acid probe population can target 12 different nucleic acid sequences, but contain no more than 8 reading sequences. As another example, a nucleic acid population can target 140 different nucleic acid species, but contain no more than 16 reading sequences. Different nucleic acid sequence targets can be identified individually by using different combinations of reading sequences within each probe. For example, each probe can contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more reading sequences. In some cases, nucleic acid probe populations can each contain the same number of reading sequences, but in other cases, different numbers of reading sequences can be present on various probes.

[0134] As a non-limiting example, a first nucleic acid probe may contain a first target sequence, a first reading sequence, and a second reading sequence, while a second, different nucleic acid probe may contain a second target sequence, the same first reading sequence, and a third reading sequence instead of the second reading sequence. Such probes can thus be distinguished by determining the various reading sequences present or associated with a given probe or position, as discussed herein.

[0135] In addition, in certain embodiments, nucleic acid probes (and their corresponding complementary sites on coding probes) can be prepared using only 2 or 3 of the 4 bases (such as omitting all "G" or all "C" in the probe). Sequences lacking "G" or "C" can form very few secondary structures in certain embodiments and can contribute to more uniform and faster hybridization.

[0136] In some embodiments, the nucleic acid probe may contain a signaling entity. However, it should be understood that a signaling entity is not required in all cases; for example, in some embodiments, a secondary nucleic acid probe may be used to assay the nucleic acid probe, as discussed in further detail below. Examples of signaling entities that may be used will also be discussed in more detail below.

[0137] Other components may also be present in the nucleic acid probe. For example, in one group of embodiments, one or more primer sequences may be present, for example to allow enzymatic amplification of the probe. One of ordinary skill in the art will know primer sequences suitable for applications such as amplification (e.g., using PCR or other suitable techniques). Many such primer sequences are commercially available. Other examples of sequences that may be present in the primary nucleic acid probe include, but are not limited to, promoter sequences, operators, identification sequences, nonsense sequences, etc.

[0138] In some embodiments, primers are single-stranded or partially double-stranded nucleic acids (e.g., DNA) used as the starting point of nucleic acid synthesis, thereby allowing polymerases such as nucleic acid polymerases to extend primers and replicate complementary strands. Primers are (e.g., designed to) complementary to target nucleic acids and hybridize with them. In some embodiments, primers are synthetic primers. In some embodiments, primers are non-naturally occurring primers. Primers generally have a length of 10 to 50 nucleotides. For example, primers can have a length of 10 to 40, 10 to 30, 10 to 20, 25 to 50, 15 to 40, 15 to 30, 20 to 50, 20 to 40 or 20 to 30 nucleotides. In some embodiments, primers have a length of 18 to 24 nucleotides.

[0139] In addition, the components of the nucleic acid probe can be arranged in any suitable order. For example, in one embodiment, the components can be arranged in the nucleic acid probe as: primer-reading sequence-targeting sequence-reading sequence-reverse primer. The "reading sequence" in this structure can each contain any number (including 0) of reading sequences, as long as there is at least one reading sequence in the probe. Non-limiting exemplary structures include primer-targeting sequence-reading sequence-reverse primer, primer-reading sequence-targeting sequence-reverse primer, targeting sequence-primer-targeting sequence-reading sequence-reverse primer, targeting sequence-primer-reading sequence-targeting sequence-reverse primer, primer-targeting sequence-reading sequence--targeting sequence-reverse primer, targeting sequence-primer-reading sequence-reverse primer, targeting sequence-primer-reading sequence-reverse primer, targeting sequence-reading sequence-primer, reading sequence-targeting sequence-primer, reading sequence-primer-targeting sequence-reverse primer, etc. In addition, in some embodiments (including in all the above examples), the reverse primer is optional.

[0140] According to certain aspects of the invention, after the nucleic acid probe is introduced into a cell or other sample, the nucleic acid probe can be directly measured by measuring the signaling entity (if present), and / or the nucleic acid probe can be measured by using one or more secondary nucleic acid probes. As mentioned, in some cases, the determination can be spatial (e.g., in two or three dimensions). In addition, in some cases, the determination can be quantitative, for example, the amount or concentration of the primary nucleic acid probe (and the target nucleic acid) can be determined. In addition, depending on the application, the secondary probe can include any of a variety of entities that can hybridize with nucleic acids (e.g., DNA, RNA, LNA and / or PNA, etc.). The signaling entity is discussed in more detail below.

[0141] The secondary nucleic acid probe may comprise a recognition sequence that can bind or hybridize to the reading sequence of the primary nucleic acid probe. In some cases, the binding is specific, or the binding may allow the recognition sequence to preferentially bind or hybridize to only one sequence in the reading sequence present. The secondary nucleic acid probe may also comprise one or more signaling entities. If more than one secondary nucleic acid probe is used, the signaling entities may be the same or different.

[0142] Recognition sequence can have any length, and multiple recognition sequences can have identical or different lengths.If use more than one recognition sequence, recognition sequence can have identical or different lengths independently.For example, the length of recognition sequence can be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40 or at least 50 nucleotides.In some cases, the length of recognition sequence can be no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, more than 35, no more than 30, no more than 20 or no more than 10 nucleotides.Any combination among these is also possible, for example, recognition sequence can have the length of 10 to 30, 20 to 40 or 25 to 35 nucleotides etc.In one embodiment, recognition sequence has the length identical with reading sequence. In addition, in some cases, the recognition sequence can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 100% complementary to the reading sequence of the primary nucleic acid probe.

[0143] As mentioned, in some cases, the secondary nucleic acid probe can include one or more signaling entities. Examples of signaling entities are discussed in more detail below.

[0144] As discussed, in certain aspects of the invention, nucleic acid probes comprising various "reading sequences" are used. For example, a population of primary nucleic acid probes can contain certain "reading sequences" that can bind to certain secondary nucleic acid probes, and the location of the primary probes within a sample is determined using, for example, secondary nucleic acid probes that contain signaling entities. As mentioned, in some cases, populations of reading sequences can be combined in various combinations to generate different nucleic acid probes, for example so that a relatively small number of reading sequences can be used to generate a relatively large number of different nucleic acid probes.

[0145] Therefore, in some cases, primary nucleic acid probe (or other nucleic acid probe) colonies can each contain a certain number of reading sequences, some of which are shared between different primary nucleic acid probes, so that the total population probe of the primary nucleic acid probe can include a certain number of reading sequences. The nucleic acid probe colony can have any suitable number of reading sequences. For example, the primary nucleic acid probe colony can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, etc. reading sequences. In some embodiments, more than 20 are also possible. Additionally, in some cases, a population of nucleic acid probes can have a total of 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 20 or more, 24 or more, 32 or more, 40 or more, 50 or more, 60 or more, 64 or more, 100 or more, 128 or more, etc. possible read sequences present, although some or all of the probes may each contain more than one read sequence, as discussed herein. In addition, in some embodiments, the nucleic acid probe population can have no more than 100, no more than 80, no more than 64, no more than 60, no more than 50, no more than 40, no more than 32, no more than 24, no more than 20, no more than 16, no more than 15, no more than 14, no more than 13, no more than 12, no more than 11, no more than 10, no more than 9, no more than 8, no more than 7, no more than 6, no more than 5, no more than 4, no more than 3, or no more than 2 read sequences present. Any combination of these read sequences is also possible, for example, the nucleic acid probe population can include a total of 10 to 15 read sequences.

[0146] As a non-limiting example of a method for combinatorially generating a relatively large number of nucleic acid probes from a relatively small number of reading sequences, in a population of 6 different types of nucleic acid probes, each nucleic acid probe comprising one or more reading sequences, the total number of reading sequences within the population may be no greater than 4. It should be understood that although 4 reading sequences are used in this example for ease of explanation, in other embodiments, depending on the application, a larger number of nucleic acid probes may be achieved using, for example, 5, 8, 10, 16, 32, or more reading sequences, or any other suitable number of reading sequences described herein. Figure 4A , if each primary nucleic acid probe contains two different reading sequences, then by using 4 such reading sequences (A, B, C, and D), up to 6 probes can be identified separately. It should be noted that in this example, the ordering of the reading sequences on the nucleic acid probes is not required, that is, "AB" and "BA" can be considered synonymous (although in other embodiments, the ordering of the reading sequences may be required, and "AB" and "BA" may not necessarily be synonymous). Similarly, if 5 reading sequences (A, B, C, D, and E) are used in the primary nucleic acid probe population, up to 10 probes can be identified separately, such as Figure 4B For example, one of ordinary skill in the art will appreciate that, assuming that ordering of the reads is not necessary, up to k reads can be generated for a population with n reads at each probe. Because not all probes need to have the same number of read sequences, and not all combinations of read sequences need to be used in every embodiment, more or less than this number of different probes may be used in certain embodiments. In addition, it should be understood that in some embodiments, the number of read sequences on each probe need not be the same. For example, some probes may include 2 read sequences, while other probes may include 3 read sequences.

[0147] In some aspects, the binding pattern and / or reading sequence of the nucleic acid probe within the sample can be used to define error detection and / or error correction codes, for example, to reduce or prevent misidentification or errors of nucleic acids, for example, as discussed with reference to FIG. 3. Thus, for example, if binding is indicated (e.g., as determined using a signaling entity), the position can be identified with "1"; conversely, if binding is not indicated, the position can be identified with "0" (or vice versa in some cases). Multiple rounds of binding assays (e.g., using different nucleic acid probes) can then be used to generate, for example, a "code word" for the spatial position. In some embodiments, the code word can be error detected and / or corrected. For example, the code word can be organized so that if a match is not found for a given set of reading sequences or binding patterns of nucleic acid probes, the match can be identified as an error, and optionally, the sequence can be error corrected to determine the correct target of the nucleic acid probe. In some cases, a code word may have fewer "letters" or positions than the total number of nucleic acids encoded by the code word, for example, when each code word encodes a different nucleic acid.

[0148] Such error detection and / or error correction codes can take many forms. Previously, multiple such codes have been developed in other contexts such as the telecommunications industry, such as Golay codes or Hamming codes. In one group of embodiments, the reading sequence or binding pattern of the nucleic acid probe is assigned so that not every possible combination is assigned.

[0149] For example, if 4 read sequences are possible and the primary nucleic acid probe contains 2 read sequences, up to 6 primary nucleic acid probes can be identified; however, the number of primary nucleic acid probes used can be less than 6. Similarly, for k read sequences in a population with n read sequences on each primary nucleic acid probe, 6 read sequences can be generated. different probes, but the number of primary nucleic acid probes used may be greater or less than Additionally, these may be randomly assigned or assigned in a specific manner to enhance the ability to detect and / or correct errors.

[0150] As another example, if multiple rounds of nucleic acid probes are used, the number of rounds can be chosen arbitrarily. If in each round, each target can give two possible results, such as being detected or not detected, for n rounds of probes, up to 2 n Different targets may be possible, but the number of nucleic acid targets actually used may be less than 2. n For example, if in each round, each target can give more than two possible outcomes, such as being detected in different color channels, then for n rounds of probes, more than 2 n Species (e.g. 3 n , 4n ...) different targets may be possible. In some cases, the number of nucleic acid targets actually used may be any number less than this number. In addition, these may be randomly assigned or assigned in a specific manner to enhance the ability to detect and / or correct errors.

[0151] For example, in one set of embodiments, code words or nucleic acid probes can be assigned within the code space so that the assignments are separated by a Hamming distance, which measures the number of incorrect "reads" in a given pattern that cause the nucleic acid probe to be misinterpreted as a different valid nucleic acid probe. In some cases, the Hamming distance can be at least 2, at least 3, at least 4, at least 5, at least 6, etc. In addition, in one set of embodiments, the assignments can be formed into Hamming codes, such as Hamming (7, 4) codes, Hamming (15, 11) codes, Hamming (31, 26) codes, Hamming (63, 57) codes, Hamming (127, 120) codes, etc. In another set of embodiments, the assignments can be formed into SECDED codes, such as SECDED (8, 4) codes, SECDED (16, 4) codes, SCEDED (16, 11) codes, SCEDED (22, 16) codes, SCEDED (39, 32) codes, SCEDED (72, 64) codes, etc. In another set of embodiments, the allocation may form an extended binary Golay code, a perfect binary Golay code, or a ternary Golay code. In another set of embodiments, the allocation may represent a subset of possible values ​​taken from any of the above codes.

[0152] For example, a code having the same error correction properties of a SECDED code may be formed by using only binary words containing a fixed number (such as 4) of '1' bits to encode the target. In another set of embodiments, the allocation may represent a subset of possible values ​​obtained from the above code for the purpose of addressing asymmetric readout errors. For example, in some cases, when the ratio of a '0' bit being measured as a '1' or a '1' bit being measured as a '0' is different, where the number of '1' bits may be fixed for all binary words used, the code may eliminate biased measurements of words with different numbers of '1's.

[0153] Therefore, in some embodiments, once a code word is determined (e.g., as discussed herein), the code word can be compared with a known nucleic acid code word. If a match is found, the nucleic acid target can be identified or determined. If a match is not found, the error in the reading of the code word can be identified. In some cases, error correction can also be applied to determine the correct code word, and thus the correct identity of the nucleic acid target is caused. In some cases, the code word can be selected so that when assuming that there is only one error, only one possible correct code word is available, and therefore, the nucleic acid target can only have one correct identity. In some cases, this can also be generalized to a larger code word interval or Hamming distance; for example, a code word can be selected so that if there are two, three or four errors (or more errors in some cases), only one possible correct code word is available, and therefore, the nucleic acid target can only have one correct identity.

[0154] The error correction code may be a binary error correction code, or it may be based on other numbering systems, such as a ternary or quaternary error correction code. For example, in one set of embodiments, more than one type of signaling entity may be used and assigned to different numbers within the error correction code. Thus, as a non-limiting example, a first signaling entity (or in some cases, more than one signaling entity) may be assigned a '1', and a second signaling entity (or in some cases, more than one signaling entity) may be assigned a '2' (where '0' indicates that no signaling entity is present), and a code word is assigned to define a ternary error correction code. Similarly, a third signaling entity may additionally be assigned a '3' to produce a quaternary error correction code, and so on.

[0155] As described above, in certain aspects, a signaling entity is determined, e.g., to determine a nucleic acid probe and / or generate a codeword. In some cases, various techniques can be used, e.g., to spatially determine a signaling entity within a sample. In some embodiments, the signaling entity can be fluorescent, and techniques for determining fluorescence within a sample (e.g., fluorescence microscopy or confocal microscopy) can be used to spatially identify the location of the signaling entity within the cell. In some cases, the location of the entity within the sample can be determined in two or even three dimensions. In addition, in some embodiments, more than one signaling entity can be determined at once (e.g., with signaling entities of different colors or emissions) and / or sequentially.

[0156] In addition, in some embodiments, the confidence level of the nucleic acid target identified can be measured.For example, the ratio of the number of exact matches and the number of matches with one or more 1-bit errors can be used to measure the confidence level.In some cases, only the matching with a confidence ratio greater than a certain value can be used.For example, in certain embodiments, only when the confidence ratio of matching is greater than about 0.01, greater than about 0.03, greater than about 0.05, greater than about 0.1, greater than about 0.3, greater than about 0.5, greater than about 1, greater than about 3, greater than about 5, greater than about 10, greater than about 30, greater than about 50, greater than about 100, greater than about 300, greater than about 500, greater than about 1000 or any other suitable value, matching can be accepted. Additionally, in some embodiments, a match is accepted only if the confidence ratio of the identified nucleic acid target is greater than about 0.01, about 0.03, about 0.05, about 0.1, about 0.3, about 0.5, about 1, about 3, about 5, about 10, about 30, about 50, about 100, about 300, about 500, about 1000, or any other suitable value over an internal standard or false positive control.

[0157] In some embodiments, the spatial position of an entity (and, thus, a nucleic acid probe with which the entity may be associated) can be determined with a relatively high resolution. For example, the position can be determined with a spatial resolution of better than about 100 microns, better than about 30 microns, better than about 10 microns, better than about 3 microns, better than about 1 micron, better than about 800 nm, better than about 600 nm, better than about 500 nm, better than about 400 nm, better than about 300 nm, better than about 200 nm, better than about 100 nm, better than about 90 nm, better than about 80 nm, better than about 70 nm, better than about 60 nm, better than about 50 nm, better than about 40 nm, better than about 30 nm, better than about 20 nm, or better than about 10 nm, etc.

[0158] There are multiple technologies that can, for example, optically measure the spatial position of an entity using a fluorescence microscope or image it. In some cases, the spatial position can be measured with super-resolution or with a resolution superior to the wavelength or diffraction limit of light. Non-limiting examples include STORM (stochastic optical reconstruction microscopy), STED (stimulated emission depletion microscopy), NSOM (near-field scanning optical microscopy), 4Pi microscopy, SIM (structured illumination microscopy), SMI (spatial modulation illumination) microscopy, RESOLFT (reversible saturable optical linear fluorescence conversion microscopy), GSD (ground state depletion microscopy), SSIM (saturated structured illumination microscopy), SPDM (spectral precise distance microscopy), photoactivated localization microscopy (PALM), fluorescence photoactivated localization microscopy (FPALM), LIMON (3D optical microscopy nanomicroscopy), super-resolution optical wave imaging (SOFI), etc. See, for example, U.S. Pat. No. 7,838,302, entitled “Sub-Diffraction Limit Image Resolution and Other Imaging Techniques,” issued on November 23, 2010 to Zhuang et al.; U.S. Pat. No. 8,564,792, entitled “Sub-diffraction Limit Image Resolution in Three Dimensions,” issued on October 22, 2013 to Zhuang et al.; or International Patent Application Publication No. WO 2013 / 090360, entitled “High Resolution Dual-Objective Microscopy,” issued on June 20, 2013 to Zhuang et al., each of which is incorporated herein by reference in its entirety.

[0159] As an illustrative, non-limiting example, in one set of embodiments, a sample may be imaged with a large numerical aperture, an oil immersion objective having a 100x magnification, and light collected on an electron multiplying CCD camera. In another example, a sample may be imaged with a large numerical aperture, an oil immersion lens having a 40x magnification, and light collected with a widefield scientific CMOS camera. In various non-limiting embodiments, by utilizing different combinations of objective lenses and cameras, a single field of view may correspond to no less than 40×40 microns, 80×80 microns, 120×120 microns, 240×240 microns, 340×340 microns, or 500×500 microns, etc. Similarly, in some embodiments, a single camera pixel may correspond to a sample area of ​​no less than 80×80 nm, 120×120 nm, 160×160 nm, 240×240 nm, or 300×300 nm, etc. In another example, a sample may be imaged with a low numerical aperture, an air lens having a 10x magnification, and light collected with a sCMOS camera. In other embodiments, the sample can be optically sectioned by illuminating the sample through a single or multiple scanned diffraction-limited focus produced by a scanning mirror or a spinning disk, and the collected sample passes through a single or multiple pinholes. In another embodiment, the sample can also be irradiated by a thin layer of light produced by any of a variety of methods known to those skilled in the art.

[0160] In one embodiment, the sample can be illuminated by a single Gaussian mode laser line. In some embodiments, the profiled illumination can be flattened by passing these laser lines through a multimode optical fiber that is vibrated by a piezoelectric or other mechanical device. In some embodiments, the illumination profile can be flattened by passing a single mode Gaussian beam through various refractive beam shapers (such as a piShaper or a series of stacked Powell lenses). In another set of embodiments, the Gaussian beam can be passed through a variety of different diffusing elements, such as frosted glass or an engineered diffuser, which can be rotated at high speed in some cases to remove residual laser speckle. In another embodiment, the laser illumination can be passed through a series of lenslet arrays to produce overlapping images of the illumination that approximate a planar illumination field.

[0161] In some embodiments, the centroid of the spatial location of the entity can be determined. For example, the centroid of the signaling entity can be determined within an image or a series of images using image analysis algorithms known to those of ordinary skill in the art. In some cases, the algorithm can be selected to determine non-overlapping single emitters and / or partially overlapping single emitters in a sample. Non-limiting examples of suitable techniques include maximum likelihood algorithms, least squares algorithms, Bayesian algorithms, compressed sensing algorithms, and the like. Combinations of these techniques may also be used in some cases.

[0162] Additionally, in some cases, the signaling entity can be inactivated. For example, in some embodiments, a first secondary nucleic acid probe containing a signaling entity can be applied to a sample that can recognize a first reading sequence, and then the first secondary nucleic acid probe can be inactivated before a second secondary nucleic acid probe is applied to the sample. If multiple signaling entities are used, the same or different techniques can be used to inactivate the signaling entities, and some or all of the multiple signaling entities can be inactivated, for example, sequentially or simultaneously.

[0163] Inactivation can be caused by removing the signaling entity (e.g., from the sample or from the nucleic acid probe, etc.) and / or by chemically altering the signaling entity in some manner (e.g., by photobleaching the signaling entity, bleaching or chemically altering the structure of the signaling entity (e.g., by reduction), etc.). For example, in one set of embodiments, the fluorescent signaling entity can be inactivated by chemical or optical techniques such as oxidation, photobleaching, chemical bleaching, stringent washing or enzymatic digestion, or by exposure to an enzyme reaction, dissociation of the signaling entity from other components (e.g., probes), chemical reaction of the signaling entity (e.g., with reactants that can alter the structure of the signaling entity), etc. For example, bleaching can occur by exposure to oxygen, a reducing agent, or the signaling entity can be chemically cleaved from the nucleic acid probe and washed away by a fluid flow.

[0164] In some embodiments, various nucleic acid probes (including primary and / or secondary nucleic acid probes) may include one or more signaling entities. If more than one nucleic acid probe is used, the signaling entities may be the same or different from each other. In certain embodiments, the signaling entity is any entity capable of emitting light. For example, in one embodiment, the signaling entity is fluorescent. In other embodiments, the signaling entity may be phosphorescent, radioactive, absorptive, etc. In some cases, the signaling entity is any entity that can be measured in a sample with a relatively high resolution (e.g., with a resolution better than the wavelength of visible light or the diffraction limit). The signaling entity may be, for example, a dye, a small molecule, a peptide, or a protein, etc. In some cases, the signaling entity may be a single molecule. If multiple secondary nucleic acid probes are used, the nucleic acid probe may contain the same or different signaling entities.

[0165] Non-limiting examples of signaling entities include fluorescent entities (fluorophores) or phosphorescent entities, such as cyanine dyes (e.g., Cy2, Cy3, Cy3B, Cy5, Cy5.5, Cy7, etc.), Alexa Fluor dyes, Atto dyes, photoswitchable dyes, photoactivatable dyes, fluorescent dyes, metal nanoparticles, semiconductor nanoparticles or "quantum dots", fluorescent proteins such as GFP (green fluorescent protein) or photoactivatable fluorescent proteins such as PAGFP, PSCFP, PSCFP2, Dendra, Dendra2, EosFP, tdEos, mEos2, mEos3, PAmCherry, PAtagRFP, mMaple, mMaple2 and mMaple3. Other suitable signaling entities are known to those of ordinary skill in the art. See, for example, U.S. Patent No. 7,838,302 or U.S. Patent Application Serial No. 61 / 979,436, each of which is incorporated herein by reference in its entirety.

[0166] In one set of embodiments, the signaling entity can be attached to the oligonucleotide sequence by a bond that can be cleaved to release the signaling entity. In one set of embodiments, the fluorophore can be conjugated to the oligonucleotide by a cleavable bond (such as a photocleavable bond). Non-limiting examples of photocleavable bonds include, but are not limited to, 1-(2-nitrophenyl)ethyl, 2-nitrobenzyl, biotin phosphoramidite, acrylic acid phosphoramidite, diethylaminocoumarin, 1-(4,5-dimethoxy-2-nitrophenyl)ethyl, cyclododecyl(dimethoxy-2-nitrophenyl)ethyl, 4-aminomethyl-3-nitrobenzyl, (4-nitro-3-(1-chlorocarbonyloxyethyl)phenyl)methyl-S-acetylthioate, (4-nitro-3-(1-chlorocarbonyloxyethyl)phenyl)methyl-3-(2-pyridyldithiopropionate), 3-(4,4'-dimethoxytrityl)-1-(2-nitrophenyl)-propane-1,3 ... ]-phosphoramidite, 1-[2-nitro-5-(6-(4,4'-dimethoxytrityloxy)butyramidomethyl)phenyl]-ethyl-[2-cyanoethyl-(N,N-diisopropyl)]-phosphoramidite, 1-[2-nitro-5-(6-(4,4'-dimethoxytrityloxy)butyramidomethyl)phenyl]-ethyl-[2-cyanoethyl-(N,N-diisopropyl)]-phosphoramidite, 1-[2-nitro-5-(6-(N-(4,4'-dimethoxytrityl))-biotinamidohexamethyleneiminomethyl)phenyl]-ethyl-[2-cyanoethyl-(N,N-diisopropyl)]-phosphoramidite or similar linkers. In another set of embodiments, the fluorophore can be conjugated to the oligonucleotide via a disulfide bond. Disulfide bonds can be cleaved by a variety of reducing agents, such as, but not limited to, dithiothreitol, dithioerythritol, β-mercaptoethanol, sodium borohydride, thioredoxin, glutaredoxin, trypsinogen, hydrazine, diisobutylaluminum hydride, oxalic acid, formic acid, ascorbic acid, phosphorous acid, tin chloride, glutathione, thioglycolate, 2,3-dimercaptopropanol, 2-mercaptoethylamine, 2-aminoethanol, tris(2-carboxyethyl)phosphine, bis(2-mercaptoethyl)sulfone, N,N'-dimethyl-N,N'-di(mercaptoacetyl)hydrazine, 3-mercaptohexanoate, dimethylformamide, thiopropyl-agarose, tri-n-butylphosphine, cysteine, ferrous sulfate, sodium sulfite, phosphites, hypophosphites, thiophosphates, and the like, and / or combinations of any of these reducing agents. In another embodiment, the fluorophore may be conjugated to the oligonucleotide via one or more phosphorothioate-modified nucleotides in which the sulfur modification replaces the bridging and / or non-bridging oxygen. In certain embodiments, the fluorophore may be cleaved from the oligonucleotide by the addition of compounds such as, but not limited to, iodoethanol, iodine mixed in ethanol, silver nitrate, or mercuric chloride. In another set of embodiments, the signaling entity may be chemically inactivated by reduction or oxidation.For example, in one embodiment, sodium borohydride can be used to reduce chromophores such as Cy5 or Cy7 to a stable non-fluorescent state. In another set of embodiments, the fluorophore can be conjugated to the oligonucleotide via an azo bond, and the azo bond can be cleaved with 2-[(2-N-arylamino)phenylazo]pyridine. In another set of embodiments, the fluorophore can be conjugated to the oligonucleotide via a suitable nucleic acid segment, which can be cleaved when appropriately exposed to a DNA enzyme such as an exodeoxyribonuclease or an endodeoxyribonuclease. Examples include, but are not limited to, deoxyribonuclease I or deoxyribonuclease II. In one set of embodiments, cleavage can be performed by a restriction endonuclease. Non-limiting examples of potentially suitable restriction endonucleases include BamHI, BsrI, NotI, XmaI, PspAI, DpnI, MboI, MnII, Eco57I, Ksp632I, DraIII, AhaII, SmaI, MluI, HpaI, ApaI, BclI, BstEII, TaqI, EcoRI, SacI, HindII, HaeII, DraII, Tsp509I, Sau3AI, PacI, etc. More than 3000 restriction enzymes have been studied in detail, and more than 600 are commercially available. In another set of embodiments, the fluorophore can be conjugated to biotin, and the oligonucleotide can be conjugated to avidin or streptavidin. The interaction between biotin and avidin or streptavidin allows the fluorophore to be conjugated to the oligonucleotide, while sufficient exposure to excess added free biotin can "compete out" the linkage, causing cleavage to occur. In addition, in another set of embodiments, corresponding "toe-hold-probes" can be used to remove the probe, which contains the same sequence as the probe and an additional number of bases (e.g., 1-20 additional bases, such as 5 additional bases) with homology to the encoding probe. These probes can remove the labeled readout probe through strand displacement interactions.

[0167] As used herein, the term "light" generally refers to electromagnetic radiation having any suitable wavelength (or equivalently, frequency). For example, in some embodiments, light can include wavelengths in the optical or visible range (e.g., having a wavelength between about 400nm and about 700nm, i.e., "visible light"), infrared wavelengths (e.g., having a wavelength between about 300 microns and 700nm), ultraviolet wavelengths (e.g., having a wavelength between about 400nm and about 10nm), etc. In some cases, as discussed in detail below, more than one entity can be used, i.e., chemically different or, for example, structurally different entities. However, in other cases, the entities can be chemically identical or at least substantially chemically identical.

[0168] In one set of embodiments, the signaling entity is "switchable", that is, the entity can be switched between two or more states, at least one of which emits light of a desired wavelength. In other states, the entity may not emit light, or emit light of different wavelengths. For example, the entity can be "activated" to a first state capable of producing light of a desired wavelength, and "deactivated" to a second state in which light of the same wavelength cannot be emitted. If the entity can be activated by incident light of a suitable wavelength, the entity is "photoactivated". As a non-limiting example, Cy5 can be switched between fluorescence and dark states in a controlled and reversible manner by light of different wavelengths, i.e., 633nm (or 642nm, 647nm, 656nm) red light can switch or deactivate Cy5 to a stable dark state, while 405nm green light can switch or activate Cy5 back to a fluorescent state. In some cases, an entity can be reversibly switched between two or more states, for example, when exposed to an appropriate stimulus. For example, a first stimulus (e.g., light of a first wavelength) can be used to activate a switchable entity, and a second stimulus (e.g., light of a second wavelength) can be used to deactivate a switchable entity, e.g., to a non-luminescent state. Any suitable method can be used to activate the entity. For example, in one embodiment, incident light of a suitable wavelength can be used to activate the entity to emit light, i.e., the entity is "photo-switchable". Therefore, a photo-switchable entity can be switched between different luminescent or non-luminescent states by, for example, incident light of different wavelengths. The light can be monochromatic (e.g., generated using a laser) or polychromatic. In another embodiment, the entity can be activated when stimulated by an electric field and / or a magnetic field. In other embodiments, the entity can be activated when exposed to a suitable chemical environment (e.g., by adjusting pH, or inducing a reversible chemical reaction involving the entity, etc.). Similarly, any suitable method can be used to inactivate the entity, and the methods for activating and deactivating the entity do not need to be the same. For example, the entity can be deactivated when exposed to incident light of a suitable wavelength, or the entity can be deactivated by waiting for a sufficient time.

[0169] Typically, one of ordinary skill in the art can identify a "switchable" entity by determining conditions under which the entity in the first state can emit light when exposed to an excitation wavelength, switching the entity from the first state to the second state (e.g., when exposed to light at a switching wavelength), and then showing that the entity in the second state is no longer able to emit light (or emits light at a greatly reduced intensity) when exposed to the excitation wavelength.

[0170] In one set of embodiments, as discussed, the switchable entity can be switched when exposed to light. In some cases, the light used to activate the switchable entity can come from an external source, for example, a light source such as a laser light source, another light emitting entity close to the switchable entity, etc. In some cases, the second light emitting entity can be a fluorescent entity, and in certain embodiments, the second light emitting entity itself can also be a switchable entity.

[0171] In some embodiments, the switchable entity includes a first luminescent moiety (e.g., a fluorophore) and a second moiety that activates or "switches" the first moiety. For example, upon exposure to light, the second moiety of the switchable entity can activate the first moiety, causing the first moiety to emit light. Examples of activator moieties include, but are not limited to, Alexa Fluor 405 (Invitrogen), Alexa Fluor 488 (Invitrogen), Cy2 (GE Healthcare), Cy3 (GE Healthcare), Cy3B (GE Healthcare), Cy3.5 (GE Healthcare), or other suitable dyes. Examples of luminescent moieties include, but are not limited to, Cy5, Cy5.5 (GE Healthcare), Cy7 (GE Healthcare), Alexa Fluor 647 (Invitrogen), Alexa Fluor 680 (Invitrogen), Alexa Fluor 700 (Invitrogen), Alexa Fluor 750 (Invitrogen), Alexa Fluor 790 (Invitrogen), DiD, DiR, YOYO-3 (Invitrogen), YO-PRO-3 (Invitrogen), TOT-3 (Invitrogen), TO-PRO-3 (Invitrogen), or other suitable dyes.These can be linked together, for example, covalently, e.g., directly or via a linker, to form compounds such as, but not limited to, Cy5-Alexa Fluor 405, Cy5-Alexa Fluor 488, Cy5-Cy2, Cy5-Cy3, Cy5-Cy3.5, Cy5.5-AlexaFluor 405, Cy5.5-Alexa Fluor 488, Cy5.5-Cy2, Cy5.5-Cy3, Cy5.5-Cy3.5, Cy7-AlexaFluor 405, Cy7-Alexa Fluor 488, Cy7-Cy2, Cy7-Cy3, Cy7-Cy3.5, Alexa Fluor 647-AlexaFluor 405, Alexa Fluor 647-Alexa Fluor 488, Alexa Fluor 647-Cy2, Alexa Fluor 647-Cy3, Alexa Fluor 647-Cy3.5, Alexa Fluor 750-Alexa Fluor 405, Alexa Fluor 750-Alexa Fluor 488, Alexa Fluor 750-Cy2, Alexa Fluor 750-Cy3, or Alexa Fluor 750-Cy3.5. Those of ordinary skill in the art know the structures of these and other compounds (many of which are commercially available). The moieties may be connected by covalent bonds or by joints such as those described in detail below. Other luminescent or activator moieties may include moieties having two quaternized nitrogen atoms connected by a polymethylene chain, wherein each nitrogen is independently a part of a heteroaromatic moiety such as pyrrole, imidazole, thiazole, pyridine, quinoline, indole, benzothiazole, etc., or a part of a non-aromatic amine. In some cases, there may be 5, 6, 7, 8, 9 or more carbon atoms between the two nitrogen atoms.

[0172] In some cases, when separated from each other, the luminescent portion and the activator portion can each be a fluorophore, i.e., an entity that can emit light of a certain emission wavelength when exposed to a stimulus (e.g., an excitation wavelength). However, when a switchable entity comprising a first fluorophore and a second fluorophore is formed, the first fluorophore forms a first luminescent portion, and the second fluorophore forms an activator portion that activates or "switches" the first portion in response to a stimulus. For example, the switchable entity may comprise a first fluorophore directly bonded to a second fluorophore, or the first and second entities may be connected via a joint or a common entity. Whether a pair of luminescent portions and activator portions produce a suitable switchable entity can be tested by methods known to those of ordinary skill in the art. For example, light of various wavelengths can be used to excite the pair of luminescent portions and activator portions, and the emitted light from the luminescent portions can be measured to determine whether the pair forms a suitable switch.

[0173] As a non-limiting example, Cy3 and Cy5 can be linked together to form such an entity. In this example, Cy3 is an activator moiety that can activate Cy5 (the luminescent moiety). Thus, light at or near the absorption maximum of the activation or second moiety of the entity (e.g., light near 532 nm for Cy3) can cause the moiety to activate the first luminescent moiety, thereby causing the first moiety to emit light (e.g., near 647 nm for Cy5). See, e.g., U.S. Patent No. 7,838,302, which is incorporated herein by reference in its entirety. In some cases, the first luminescent moiety can then be deactivated by any suitable technique (e.g., by directing 647 nm red light to the Cy5 moiety of the molecule).

[0174] Other non-limiting examples of potentially suitable activator moieties include 1,5IAEDANS, 1,8-ANS, 4-methylumbelliferone, 5-carboxy-2,7-dichlorofluorescein, 5-carboxyfluorescein (5-FAM), 5-carboxynaphthofluorescein, 5-carboxytetramethylrhodamine (5-TAMRA), 5-FAM (5-carboxyfluorescein), 5-HAT (hydroxytryptamine), 5-hydroxytryptamine (HAT), 5-ROX (carboxy-X-rhodamine), 5-TAMRA (5-carboxytetramethylrhodamine), 6-carboxyrhodamine 6G, 6-CR 6G, 6-JOE, 7-amino-4-methylcoumarin, 7-aminoactinomycin D (7-AAD), 7-hydroxy-4-methylcoumarin, 9-amino-6-chloro-2-methoxyacridine, ABQ, acid fuchsin, ACMA (9-amino-6-chloro-2-methoxyacridine), acridine orange, acridine red, acridine yellow, acridine flavin, acridine flavin Feulgen SITSA, Alexa Fluor 350, Alexa Fluor 405, Alexa Fluor 430, Alexa Fluor 488, Alexa Fluor 500, Alexa Fluor 514, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 610, Alexa Fluor 633, Alexa Fluor 635, Alizarin complex, Alizarin red, AMC, AMCA-S, AMCA (aminomethylcoumarin), AMCA-X, aminoactinomycin D, aminocoumarin, aminomethylcoumarin (AMCA), aniline blue, stearic acid anthracene, APTRA-BTC, APTS, astrazone brilliant red 4G, astrazone orange R, astrazone red 6B, astrazone yellow 7GLL, mepaline, ATTO390, ATTO 425, ATTO 465, ATTO 488, ATTO 495, ATTO 520, ATTO 532, ATTO 550, ATTO 565, ATTO 590, ATTO 594, ATTO 610, ATTO 611X, ATTO 620, ATTO 633, ATTO 635, ATTO 647, ATTO647N, ATTO 655, ATTO 680, ATTO 700, ATTO 725, ATTO 740, ATTO-TAG CBQCA, ATTO-TAG FQ, Auramine, Aurophosphine G, Aurophosphine, BAO 9 (diaminophenyloxadiazole), BCECF (high pH), BCECF (low pH), berberine sulfate, Bimane, Bis-benzamide, Bis-benzimide (Hoechst), Bis-BTC, BlancophorFFG,Blancophor SV, BOBO-1, BOBO-3, Bodipy 492 / 515, Bodipy 493 / 503, Bodipy 500 / 510, Bodipy 505 / 515, Bodipy 530 / 550, Bodipy 542 / 563, Bodipy 558 / 568, Bodipy 564 / 570, Bodipy 576 / 589, Bodipy 581 / 591, Bodipy 630 / 650-X, Bodipy 650 / 665-X, Bodipy 665 / 676, Bodipy Fl, Bodipy FL ATP, Bodipy Fl-ceramide, Bodipy R6G, Bodipy TMR, BodipyTMR-X conjugate, Bodipy TMR-X, SE, Bodipy TR, Bodipy TR ATP, Bodipy TR-XSE, BO-PRO-1, BO-PRO-3, Leucoflavin FF, BTC, BTC-5N, Calcein, Calcein Blue, Calcein Red, Calcein Green, Calcein Green-1Ca, 2+ Dye, Calcium Green-2Ca 2+ , Calcium Green-5N Ca 2+ , Calcium Green-C18 Ca 2+, calcium orange, fluorescent brightener, carboxy-X-rhodamine (5-ROX), Cascade Blue, Cascade Yellow, catecholamine, CCF 2 (GeneBlazer), CFDA, chromomycin A, chromomycin A, CL-NERF, CMFDA, coumarin phalloidin, CPM methylcoumarin, CTC, CTC methyl, Cy2, Cy3.1 8, Cy3.5, Cy3, Cy5.1 8, cyclic AMP fluoride sensor (FiCRhR), benzoyl, dansyl, dansylamide, dansylcadaverine, dansyl chloride, dansyl DHPE, dansyl fluoride, DAPI, Dapoxyl, Dapoxyl 2. Dapoxyl3'DCFDA, DCFH (dichlorodihydrofluorescein diacetate), DDAO, DHR (dihydrorhodamine 123), Di-4-ANEPPS, Di-8-ANEPPS (non-ratio), DiA (4-Di-16-ASP), dichlorodihydrofluorescein diacetate (DCFH)), DiD-lipophilic tracer, DiD (DiIC18 (5)), DIDS, dihydrorhodamine 123 (DHR), DiI (DiIC18 (3)), dinitrophenol, DiO (DiOC18 (3)), DiR, DiR (DiIC18 (7)), DM-NERF (high pH), DNP, dopamine, DTAF, DY-630-NHS, DY-635-NHS, DyLight405, DyLight 488, DyLight 549, DyLight633, DyLight 649, DyLight680, DyLight 800, ELF 97, Eosin, Erythrosine, Erythrosine ITC, Ethidium Bromide, Ethidium Homodimer-1 (EthD-1), Euchrysin, EukoLight, Europium (III) Chloride, Fast Blue, FDA, Feulgen (p-nitroaniline), FIF (Formaldehyde Induced Fluorescence), FITC, Flazo Orange, Fluo-3, Fluo-4, Fluorescein (FITC), Fluorescein diacetate, Fluoro-Emerald, Fluoro-Gold (Hydroxystilbene), Fluor-Ruby, FluorX, FM1-43, FM 4-46, Fura Red (High pH), Fura Red / Fluo-3, Fura-2, Fura-2 / BCECF, Genacryl Brilliant Red B, Genacryl Brilliant Yellow 10GF, Genacryl Pink 3G, Genacryl Yellow 5GF, GeneBlazer (CCF2), Gloxalic Acid, Granular Blue, Hematoporphyrin, Hoechst 33258, Hoechst 33342, Hoechst34580, HPTS, Hydroxycoumarin, Hydroxystilbene (FluoroGold), Serotonin, Indo-1, High Calcium, Indo-1, Low Calcium, Indodicarbocyanine (DiD), Indotricarbocyanine (DiR), Intrawhite Cf, JC-1, JO-JO-1, JO-PRO-1, LaserPro, Laurodan, LDS 751 (DNA), LDS 751 (RNA), Leucophor PAF, Leucophor SF, Leucophor WS, Lissamine Rhodamine, Lissamine Rhodamine B, Calcein / Ethidium Homodimer, LOLO-1, LO-PRO-1, Lucifer Yellow, Lyso Tracer Blue, Lyso Tracer Blue-White, Lyso Tracer Green, Lyso Tracer Red, Lyso Tracer Yellow, LysoSensor Blue, LysoSensor Green, LysoSensor Yellow / Blue, Mag Green, Naphthalene Red (Root Bark Red B), Mag-Fura Red, Mag-Fura-2, Mag-Fura-5, Mag-Indo-1, Magnesium Green, Magnesium Orange, Peacock Green, Sea Blue, Maxilon Brilliant Yellow 10GFF, Maxilon Brilliant Yellow 8GFF, Merocyanin, Methoxycoumarin, Mitotracker Green FM, Mitotracker Orange, Mitotracker Red, Dinitromycin, Monobromomethane, Monobromomethane (mBBr-GSH), Monochlorobimane, MPS (Methyl Green Pyronitrile), NBD, NBD Amine, Nile Red, Nitrobenzoxadiazole, Norepinephrine, Nuclear Fast Red, Nuclear Yellow, Nylosan BrilliantIavin E8G, Oregon Green, Oregon Green 488-X, Oregon Green, Oregon Green 488, Oregon Green 500, Oregon Green 514, Pacific Blue, Feulgen, PBFI, Phloxin B (Magdala Red), Phorwite AR, Phorwite BKL, Phorwite Rev, Phorwite RPA, Phosphine 3R, PKH26 (Sigma), PKH67, PMIA, Pontochrome Blue Black, POPO-1, POPO-3, PO-PRO-1, PO-PRO-3, Primulin, Prosin Yellow, Propidium Iodide (PI), PyMPO, Pyrene, Pyronine, Pyronine B, Pyzel Brilliant Yellow 7GF, QSY 7. Resorufin, RH414, Rhod-2, Rhodamine, Rhodamine 110, Rhodamine 123, Rhodamine 5GLD, Rhodamine 6G, Rhodamine B, Rhodamine B200, Rhodamine B Extra, Rhodamine BB, Rhodamine BG, Rhodamine Green, Rhodamine Phallicidine, Rhodamine Phalloidin, Rhodamine Red, Rhodamine WT, Rose Bengal, S65A, S65C, S65L, S65T, SBFI, 5-HT, Swinney Brilliant Red 2B, Swinney Brilliant Red 4G, Swinney Brilliant Red B, Swinney Orange, Swinney Yellow L, SITS, SITS (Primrose Lin), SITS (Stilbene Isothiosulfonic Acid), SNAFL Calcein, SNAFL-1, SNAFL-2, SNARF Calcein, SNARF1, Sodium Green, SpectrumAqua, SpectrumGreen, SpectrumOrange, SpectrumRed, SPQ (6-methoxy-N-(3-sulfopropyl)quinoline), Stilbene, Sulforhodamine B can C, Sulforhodamine Extra, SYTO 11, SYTO 12. SYTO 13, SYTO14, SYTO 15, SYTO 16, SYTO 17, SYTO 18, SYTO 20, SYTO 21, SYTO 22, SYTO 23, SYTO 24, SYTO 25, SYTO 40, SYTO 41, SYTO 42, SYTO 43, SYTO 44, SYTO 45, SYTO 59. SYTO 60, SYTO 61, SYTO 62, SYTO 63, SYTO 64, SYTO 80, SYTO 81, SYTO 82, SYTO 83, SYTO 84, SYTO 85, SYTOX Blue, SYTOX Green, SYTOX Orange, Tetracycline, Tetramethylrhodamine (TAMRA), Texas Red, Texas Red-X-conjugate, Thiadiazole cyanine (DiSC3), Thiazide Red R, Thiazole Orange, Thioflavin 5, Thioflavin S, Thioflavin TCN, Thiolyte, Thiazole Orange, Tinopol CBS (fluorescent brightener), TMR, TO-PRO-1, TO-PRO-3, TO-PRO-5, TOTO-1, TOTO-3, TRITC (tetramethylisocyanate), True Blue, TruRed, Ultralite, Fluorescein sodium B, Uvitex SFC, WW 781, X-rhodamine, XRITC, xylene orange, Y66F, Y66H, Y66W, YO-PRO-1, YO-PRO-3, YOYO-1, YOYO-3, SYBR Green, thiazole orange (inter-chelating dyes), or a combination thereof.

[0175] Another aspect of the present invention relates to computer-implemented methods. For example, a computer and / or automated system capable of automatically and / or repeatedly performing any method described herein may be provided. As used herein, an "automated" device refers to a device capable of operating without human guidance, i.e., the automated device may perform the function during a period of time after anyone has completed taking any action to initiate the function (e.g., by inputting instructions into a computer to initiate the process). Typically, the automated device may perform repeated functions after that point in time. In some cases, the processing steps may also be recorded on a machine-readable medium.

[0176] For example, in some cases, a computer can be used to control the imaging of a sample, such as using fluorescence microscopy, STORM, or super-resolution techniques such as those described herein. In some cases, a computer can also control operations such as drift correction in image analysis, physical registration, hybridization and cluster alignment, cluster decoding (e.g., fluorescence cluster decoding), error detection or correction (e.g., as discussed herein), noise reduction, identification of foreground features from background features (such as noise or debris in an image), etc. As an example, a computer can be used to control the activation and / or excitation of a signaling entity within a sample, and / or the acquisition of an image of a signaling entity. In one set of embodiments, a sample can be excited using light having various wavelengths and / or intensities, and a computer can be used to associate the sequence of wavelengths of light used to excite the sample with an acquired image of the sample containing the signaling entity. For example, a computer can apply light having various wavelengths and / or intensities to a sample to produce a different average number of signaling entities in each target region (e.g., one activated entity per position, two activated entities per position, etc.). In some cases, as described above, this information can be used to construct (in some cases at high resolution) an image and / or determine the location of a signaling entity, as noted above.

[0177] In some aspects, the sample is placed on a microscope. In some cases, the microscope may include one or more channels, such as microfluidic channels, to guide or control fluid to enter or leave the sample. For example, in one embodiment, by allowing fluid to flow into or out of the sample through one or more channels, nucleic acid probes, such as the nucleic acid probes discussed herein, can be introduced and / or removed from the sample. In some cases, one or more chambers or reservoirs may also be provided for holding the fluid, such as being communicated with the channel and / or with the sample fluid. One of ordinary skill in the art will be familiar with channels, including microfluidic channels, for moving fluids into or out of the sample.

[0178] As used herein, "microfluidic", "microscopic", "micro", "micro-" prefixes (e.g., such as "microchannel"), etc. generally refer to elements or articles having a width or diameter less than about 1 mm, in some cases less than about 100 microns (micrometers). In some embodiments, for any embodiment discussed herein, a larger channel can be used instead of a microfluidic channel or in combination with a microfluidic channel. For example, in some cases, a channel having a width or diameter less than about 10 mm, less than about 9 mm, less than about 8 mm, less than about 7 mm, less than about 6 mm, less than about 5 mm, less than about 4 mm, less than about 3 mm, or less than about 2 mm can be used. In some cases, an element or article includes a channel through which a fluid can flow. In all embodiments, the specified width can be a minimum width (i.e., a specified width, wherein, at this position, the article can have a larger width in different dimensions), or a maximum width (i.e., wherein, at this position, the article has a width that is not wider than the specified width, but can have a larger length). Thus, for example, the microfluidic channel can have an average cross-sectional dimension (e.g., perpendicular to the direction of flow of the fluid in the microfluidic channel) of less than about 1 mm, less than about 500 microns, less than about 300 microns, or less than about 100 microns. In some cases, the microfluidic channel can have an average diameter of less than about 60 microns, less than about 50 microns, less than about 40 microns, less than about 30 microns, less than about 25 microns, less than about 10 microns, less than about 5 microns, less than about 3 microns, or less than about 1 micron.

[0179] As used herein, "channel" means a feature on or in an article (e.g., a substrate) that at least partially directs the flow of a fluid. In some cases, a channel may be formed at least in part by a single component, e.g., an etched substrate or a molded unit. A channel may have any cross-sectional shape, e.g., circular, elliptical, triangular, irregular, square, or rectangular (with any aspect ratio), etc., and may be covered or uncovered (i.e., open to the external environment surrounding the channel). In embodiments where the channel is completely covered, at least a portion of the channel may have a completely enclosed cross-section, and / or the entire channel may be completely enclosed along its entire length except for its inlet and outlet.

[0180] The channel can have any aspect ratio, for example, at least about 2: 1, more generally at least about 3: 1, at least about 5: 1, at least about 10: 1, etc. aspect ratio (length to average cross-sectional dimension). As used herein, the "cross-sectional dimension" of a fluid or microfluidic channel is measured in a direction generally perpendicular to the fluid flow in the channel. The channel will generally include characteristics that facilitate control of fluid transmission, for example, structural characteristics and / or physical or chemical characteristics (hydrophobicity versus hydrophilicity) and / or other characteristics that can apply force (e.g., including force) to the fluid. The fluid in the channel can partially or completely fill the channel. In some cases, the fluid can be maintained or confined in a channel or a portion of a channel in some manner, for example, using surface tension (e.g., so that the fluid remains in a channel in a meniscus (such as a concave or convex meniscus)). In an article or substrate, some (or all) channels may have a particular size or less, for example, in some cases having a maximum dimension perpendicular to fluid flow of less than about 5 mm, less than about 2 mm, less than about 12 mm, less than about 500 microns, less than about 200 microns, less than about 100 microns, less than about 60 microns, less than about 50 microns, less than about 40 microns, less than about 30 microns, less than about 25 microns, less than about 10 microns, less than about 3 microns, less than about 1 micron, less than about 300 nm, less than about 100 nm, less than about 30 nm, or less than about 10 nm, or less. In one embodiment, the channel is a capillary.

[0181] According to certain aspects of the present invention, a variety of materials and methods can be used to form devices or components containing microfluidic channels, chambers, etc. For example, various devices or components can be formed from solid materials, where channels can be formed by micromachining, film deposition processes such as spin coating and chemical vapor deposition, physical vapor deposition, laser fabrication, photolithography, etching methods (including wet chemical or plasma processes), electrodeposition, etc. See, for example, Scientific American, 248:44-55, 1983 (Angell et al.).

[0182] In one set of embodiments, the various structures or components can be made of polymers, for example, elastomeric polymers such as polydimethylsiloxane ("PDMS"), polytetrafluoroethylene ("PTFE") or ) etc. For example, according to one embodiment, channels such as microfluidic channels can be realized by separately manufacturing a fluid system using PDMS or other soft lithography techniques (details of soft lithography techniques suitable for this embodiment are discussed in Younan Xia and George M. Whitesides, entitled “Soft Lithography” published in Annual Review of Material Science, 1998, Vol. 28, pp. 153-184, and in references entitled “Soft Lithography in Biology and Biochemistry” published in Annual Review of Biomedical Engineering, 2001, Vol. 3, pp. 335-373 by George M. Whitesides, Emanuele Ostuni, Shuichi Takayama, Xingyu Jiang, and Donald E. Ingber, each of which is incorporated herein by reference) to realize channels such as microfluidic channels.

[0183] Other examples of potentially suitable polymers include, but are not limited to, polyethylene terephthalate (PET), polyacrylates, polymethacrylates, polycarbonates, polystyrene, polyethylene, polypropylene, polyvinyl chloride, cyclic olefin copolymers (COC), polytetrafluoroethylene, fluorinated polymers, silicones such as polydimethylsiloxane, polyvinylidene chloride, bis-benzocyclobutene ("BCB"), polyimides, fluorinated derivatives of polyimides, and the like. Combinations, copolymers, or mixtures comprising the foregoing polymers are also contemplated. The device may also be formed of a composite material (e.g., a composite material of a polymer and a semiconductor material).

[0184] In some embodiments, the various microfluidic structures or parts of the device are made of polymers and / or flexible and / or elastomeric materials, and can be conveniently formed by hardenable fluids, which are convenient for manufacturing by molding (e.g., replication molding, injection molding, casting molding, etc.). Hardenable fluids can be substantially any fluid that can be induced to solidify or spontaneously solidify into a solid that can accommodate and / or convey a fluid that is envisioned to be used in a fluid network and used with a fluid network. In one embodiment, the hardenable fluid comprises a polymer liquid or a liquid polymer precursor (i.e., a "prepolymer"). Suitable polymer liquids may include, for example, thermoplastic polymers, thermosetting polymers, waxes, metals, or mixtures or composites thereof that are heated to above their melting points. As another example, suitable polymer liquids may include a solution of one or more polymers in a suitable solvent that forms a solid polymer material after removing the solvent (e.g., by evaporation). Such polymer materials that can be solidified from, for example, a molten state or by solvent evaporation are well known to those of ordinary skill in the art. For embodiments in which one or two mold masters are composed of elastomeric materials, a variety of polymeric materials (many of which are elastomeric) are suitable, and are also suitable for forming molds or mold masters. A non-limiting list of examples of such polymers includes polymers of the general classes of silicone polymers, epoxy polymers, and acrylate polymers. Epoxy polymers are characterized by the presence of three-membered cyclic ether groups commonly referred to as epoxy, 1,2-epoxide, or oxirane. For example, in addition to compounds based on aromatic amines, triazines, and alicyclic backbones, diglycidyl ethers of bisphenol A may also be used. Another example includes the well-known novolac polymers. Non-limiting examples of silicone elastomers suitable for use in the present invention include those formed from precursors including chlorosilanes such as methylchlorosilane, ethylchlorosilane, phenylchlorosilane, and the like.

[0185] In certain embodiments, silicone polymers are used, for example, silicone elastomer polydimethylsiloxane. Non-limiting examples of PDMS polymers include those sold by Dow Chemical Co., Midland, MI under the trademark Sylgard, particularly Sylgard 182, Sylgard 184 and Sylgard 186. Silicone polymers including PDMS have several beneficial properties that simplify the manufacture of various structures of the present invention. For example, such materials are cheap, easily available, and can be cured from prepolymer liquids by thermal curing. For example, PDMS can generally be cured by exposing the prepolymer liquid to a temperature of, for example, about 65°C to about 75°C (for an exposure time of, for example, at least about 1 hour). Similarly, silicone polymers such as PDMS can be elastomeric, and therefore can be used to form very small features with a relatively high aspect ratio, which is necessary in certain embodiments of the present invention. In this regard, a flexible (e.g., elastomer) mold or master mold can be advantageous.

[0186] One advantage of forming structures such as microfluidic structures or channels from silicone polymers (such as PDMS) is the ability of such polymers to be oxidized, for example by exposure to oxygen-containing plasmas such as air plasma, so that the oxidized structure contains chemical groups on its surface that can crosslink with other oxidized silicone polymer surfaces or with the oxidized surfaces of a variety of other polymers and non-polymer materials. Therefore, structures can be manufactured, subsequently oxidized and substantially irreversibly sealed to other silicone polymer surfaces, or to the surfaces of other substrates that react with the oxidized silicone polymer surfaces, without the need for separate adhesives or other sealing devices. In most cases, sealing can be accomplished simply by contacting the oxidized silicone surface with another surface, without the need to apply auxiliary pressure to form a seal. That is, pre-oxidized silicone surfaces are used as contact adhesives for suitable mating surfaces. Specifically, in addition to being irreversibly sealed itself, oxidized silicones such as oxidized PDMS can also be irreversibly sealed to a series of oxidized materials other than itself, including, for example, glass, silicon, silicon oxide, quartz, silicon nitride, polyethylene, polystyrene, glassy carbon and epoxy polymers, which are oxidized (for example, by exposure to oxygen-containing plasma) in a manner similar to the PDMS surface. Oxidation and sealing methods and integral molding techniques that can be used in the context of the present invention are described in the art, for example, in a paper entitled "Rapid Prototyping of Microfluidic Systems and Polydimethylsiloxane," Anal. Chem., 70:474-480, 1998 (Duffy et al.), incorporated herein by reference.

[0187] The following documents are each incorporated herein by reference in their entirety: U.S. Patent No. 7,838,302, entitled “Sub-Diffraction Limit Image Resolution and Other Imaging Techniques,” issued by Zhuang et al. on November 23, 2010; U.S. Patent No. 8,564,792, entitled “Sub-diffraction Limit Image Resolution in Three Dimensions,” issued by Zhuang et al. on October 22, 2013; and International Patent Application Publication No. WO 2013 / 090360, entitled “High Resolution Dual-Objective Microscopy,” published by Zhuang et al. on June 20, 2013.

[0188] In addition, incorporated by reference in their entirety are U.S. Provisional Patent Application Serial No. 62 / 031,062, entitled “Systems and Methods for Determining Nucleic Acids,” filed by Zhuang et al. on July 30, 2014; U.S. Provisional Patent Application Serial No. 62 / 050,636, entitled “Probe Library Construction,” filed by Zhuang et al. on September 15, 2014; U.S. Provisional Patent Application Serial No. 62 / 142,653, entitled “Systems and Methods for Determining Nucleic Acids,” filed by Zhuang et al. on April 3, 2015; and a PCT application entitled “Probe Library Construction” filed by Zhuang et al. on the same day as this application.

[0189] The following examples are intended to illustrate certain embodiments of the invention, but are not intended to illustrate the full scope of the invention.

[0190] Example 1

[0191] This embodiment provides a platform that enables the use of a novel form of highly multiplexed fluorescence in situ hybridization (FISH) to simultaneously detect the number and spatial organization of thousands of different mRNAs in a single cell with high efficiency and low error rate. This embodiment accomplishes these measurements by integrating and innovating methods for massively parallel probe synthesis, super-resolution imaging, and self-correcting error-checking codes.

[0192] Here, these examples provide some or all of the methods for detecting thousands of unique RNAs expressed in cells simultaneously. The method is not only expected to innovate the flux of effective single molecule FISH (smFISH) method, but also allows researchers to benefit from the hypothesis free discovery method that has made other full genome system methods so effective for biology. For example, the full genome method can allow researchers to find the RNA that its expression level and / or subcellular localization pattern change under certain target conditions (such as disease states), without a priori knowing which mRNA will change in abundance or positioning. Hundreds of genes are measured simultaneously in a single cell and also allow to identify the correlation between the gene and the localization pattern expressed in some cases.

[0193] This can be realized using the method for highly multiplexed smFISH for continuous hybridization and super-resolution imaging by orthogonal detection probes, thereby reducing the cost of developing highly automated systems for synthesizing probes and minimizing the requirements of users, as discussed herein. This provides an integrated platform for bioinformatics, mathematical operations of error correction codes, image registration and analysis of processing probe design, and for carrying out cumbersome fluid handling by a simple user-friendly interface. This integration allows easy operation and is convenient for rapid data collection under limited user training.

[0194] This example illustrates: (1) the computational design of "codewords" attached to all RNA targets in a cell that will allow unique identification of each RNA with a degree of experimental error tolerance, (2) the translation of these codewords into nucleotide sequences and synthesis of the desired single-stranded (ss) oligonucleotide (e.g., ssDNA) probes, (3) sample fixation and in situ hybridization of these probes to RNA targets, (4) the readout of these codewords by sequential rounds of hybridization of different fluorescent probes imaged using conventional fluorescence microscopy or super-resolution fluorescence microscopy, and (5) automatic decoding of the measured codewords combined with computational error correction to uniquely and robustly identify individual mRNAs.

[0195] In the first step, a "code word" is assigned to each RNA to be labeled. In a typical design, these can be strings of N binary letters or positions. Code words can be selected from the same wide range of existing fault-tolerant or error-correcting coding schemes developed for digital storage and communication, such as using Hamming codes, etc. For example, actin-RNA can be assigned the binary code word 11001010. Each code word can be unique and separated from other code words by a Hamming distance h, which measures the number of letters or positions that must be read incorrectly for a code word to be misunderstood as a different code word. A Hamming distance greater than 1 between all code words allows detection of some measurement errors-because simple errors will produce code words that are not used to encode RNA. For Hamming distances greater than 2, it is also possible to correct some errors because a code word with one error will be closest to a single unique code word in Hamming distance. The total number of different RNAs to be detected from the transcriptome and the amount of error correction required determine the length of the code word. Information theory provides several efficient algorithms for assembling error-corrected binary code books.

[0196] In the second step, the coding scheme is translated into a set of oligonucleotide (e.g., DNA) probe sequences, which may be referred to as primary probes or coding probes, wherein each sequence not only targets the probe to the target RNA, but also encodes a unique binary codeword (Fig. 1) within a set of secondary binding sites. For example, the primary binding sequence of each targeted mRNA may be designed first. These sequences are "target sequences", which comprise complementary nucleotide sequences of their target RNAs that are selected by calculation to meet a set of stringent hybridization conditions (including uniqueness in the target genome). In order to improve the efficiency of hybridization with individual mRNA, multiple primary target sequences are designed for each individual RNA. Subsequently, each position within the group of codewords is assigned a unique oligonucleotide (e.g., DNA) sequence, which is referred to as a reading sequence. These tags are designed not to interact with endogenous mRNA sequences or to interact with each other. For example, for all values ​​"1" in the codewords of individual mRNAs, the corresponding reading sequence is attached to the primary targeting sequence for the mRNA. Typically, each probe will contain a target sequence and one or more reading sequences. If the total length of the necessary reading sequence and the primary target sequence exceeds the synthetic capacity, a subset of the reading sequence may be attached to different target sequences. For example, consider the potential codeword for actin 11001010. The probe sequence for this RNA may contain read sequences corresponding to positions 1, 2, 5, and 7 in the codeword attached to a variety of actin-specific target sequences. After all sequences have been designed, the resulting complex set of unique custom oligonucleotide (e.g., DNA) sequences is manufactured and amplified using the methods described below.

[0197] In the third step, the resulting DNA pool is hybridized to, for example, fixed permeabilized cells. In this process, individual probes can be attached to each RNA in the cell through hybridization of their corresponding target sequence to the RNA, while the reading sequence remains free to bind to the appropriate secondary probe as discussed below.

[0198] In the fourth step - the readout step - the fluorescently labeled secondary nucleic acid probe (also referred to as the readout probe) is sequentially hybridized with the reading sequence attached to the target sequence bound to the mRNA target in the above step. When a large number of different RNA species are imaged simultaneously in a cell, the density of the labeled RNA can exceed the density at which each RNA can be resolved by conventional imaging methods. Therefore, this can be performed using super-resolution imaging methods, such as STORM (stochastic optical reconstruction microscopy), to resolve the labeled molecules. After each round of hybridization and imaging with a secondary probe, the fluorophore is quenched or otherwise inactivated by chemical or optical techniques such as oxidation, chemical bleaching, photobleaching, stringent washing or enzymatic digestion. The sample is then stained with the next secondary probe, and the cycle is continued until all positions of the codeword have been read out. In the simplest form, there will be one hybridization step for each position within the codeword, for example, there are 8 hybridization steps for an 8-letter codeword (Fig. 1).

[0199] FIG1 shows a schematic diagram of this embodiment. Figure 1A Each position showing a code word is assigned a unique oligonucleotide sequence when that position has a value of "1." All mRNA code words are then translated into combinations of read sequences attached to the targeting sequence. Figure 1B The various steps of the labeling scheme of this embodiment are shown. In the first step, all mRNAs (I-III) are labeled with a plurality of oligonucleotide (e.g., ssDNA) probes containing a primary targeting sequence that hybridizes with the target RNA and a "tail" (i.e., containing a reading sequence) with a translated code word that does not interact with the endogenous nucleotide sequence. In the next step, a first secondary probe is added that can bind to all probes whose tails have a reading sequence corresponding to a value "1" in the first position. The dyes on these secondary probes are imaged and bleached, and then the next secondary probe is added to bind to the probe attached to the mRNA, which has a value "1" at the second position of its assigned code word, and so on.

[0200] In a final step, the microscopic images from each staining and imaging round are, for example, computationally aligned (e.g., using feducial beads or other markers tracked during image acquisition), and localization clusters resolved by conventional fluorescence microscopy or super-resolution imaging (e.g., STORM) from different rounds are identified. These localization clusters are generated from individual target mRNA molecules, and the hybridization round in which a spot is detected in a given cluster corresponds to a "1" in the code word for that mRNA. If there are no missed detection events or false positive signals in the image, the code word will perfectly match one of the expected code words. Figure 1 describes an example in which the code word has three letters, i.e., three positions, and the three target mRNAs have the code words 110, 101, and 011 assigned to them. In a real experimental example, the code words may contain more numbers. For example, the mRNA for actin may be assigned the code word 11001010. In this case, detected clusters containing overlapping localization signals in the 1st, 2nd, 5th, and 7th hybridization steps (meaning the 1st, 2nd, 5th, and 7th secondary probes bound to the site) can be identified as individual actin mRNA molecules because the pattern of positive binding matches the code word for actin (11001010). In addition, if there are missed detection events or false positive signals in the image data, these deviations can be corrected by the error correction scheme implemented. For example, localization clusters with detected code words that have only one digit deviation from 11001010 (such as 11000010 or 11101010) can also be identified as actin mRNA because all other valid code words in this embodiment are different from the detected pattern in two or more positions.

[0201] Example 2

[0202] This example describes another alternative method that differs in several of the above steps. This method starts with the first step of constructing the codeword into the desired mRNA target, as described above.

[0203] In the second step of this method, nucleic acid probes are designed that uniquely bind to the target mRNA target as described above. However, instead of appending unique reading sequences to these target sequences, unique probe libraries or probe sets are constructed from these target sequences. Each library contains all or a subset of sequences targeting all mRNAs that contain the same value at a given position in their codewords. For example, the first library will have all or a subset of target sequences designed for all mRNAs that contain 1 (e.g., 110 and 101 instead of 011) at the first position of its codeword; the second library may have all or a subset of target sequences designed for all mRNAs that contain 1 (e.g., 110 and 011 instead of 101) at the second position of its codeword; the third library may have all or a subset of target sequences designed for all mRNAs that contain 1 (e.g., 011 and 101 instead of 110) in the third position of its codeword. Figure 1C ). As another example, consider the potential codeword 11001010 for actin. Probes targeting this mRNA would be included in libraries 1, 2, 5, and 7, but not in libraries 3, 4, 6, and 8. The same target for a given mRNA may or may not be included in a library. For example, probes targeting the same region of actin may be included in libraries 1, 2, 5, and 7, or any subset of these libraries. After all libraries have been designed, each complex set of unique custom oligonucleotide sequences is made and amplified using the methods described below.

[0204] In the third step of this method, the first probe library is hybridized with, for example, fixed permeabilized cells. In the process, the fluorophore attached to each probe in the library binds to each target in the library. The binding of these probes is then determined by fluorescence microscopy. As described above, these images can be collected via a series of methods including conventional fluorescence imaging or super-resolution imaging methods (such as STORM). After one round of imaging, the probes from the first library are inactivated or removed from the sample by the above method. The process is then repeated for each consecutive probe library until some or all libraries have been applied to the sample and imaged so that all positions in the code word have been read out. In the simplest form, there will be one hybridization and imaging step for each position in the code word, for example, for a code word with 3 positions ( Figure 1C ) There are 3 rounds of hybridization and imaging, or 8 rounds of hybridization and imaging for a codeword with 8 positions.

[0205] The final step of this method is the same as described above.

[0206] Example 3

[0207] In this example, a set of (8,4) SECDED codes is used to encode 14 genes (PGK1, H3F3B, PKM, ENO1, GPI, EEF2, GNAS, HSPA8, GAPDH, CALM1, RHOA, PPIA, UBA52, and VCP) ( Figure 2A-2E To determine the precision of these measurements, the measured abundance of these 14 mRNAs was compared to abundance measured from bulk RNA-seq of A549 cells (published ENCODE data). Strikingly, excellent agreement was found between the two measurements, as transcript counts measured using the sequential hybridization approach correlated with gene expression measured using RNA-seq with a Pearson correlation coefficient r of 0.75 ( Figure 2F ). Gene expression from three other cells was also measured, and it was found that gene expression of these 14 genes was highly correlated between cells, with an r of 0.96 ( Figure 2G ).

[0208] Codebook design. Single error correction double error detection (SECDED) codes are used to assign binary code words to each mRNA in the target group. SECDED is an extended Hamming code book with additional parity bits. In short, Matlab's communication system toolbox is used to generate SECDED codes of 8 or 16 letters or positions. In both cases, only those code words containing four 1s are used. These words are randomly assigned to the mRNAs in the target group. [0 1 0 1 1 1 0 0] is an example of the 8-letter code words used (i.e., these code words each contain 4 1s and 4 0s). [0 1 0 1 1 1 0 0 0 0 0 0 0 00 0] is an example of the 16-letter code words used (i.e., each code word contains 4 1s and 12 0s). Not every code word must be assigned to an mRNA.

[0209] Computational assembly of ssDNA primary probe sequences. Depending on the experiment, the number of primary nucleic acid probes used for hybridization with mRNA targets ranges from 200 to 2000 unique oligonucleotides. For example, to label 14 mRNAs with 28 oligonucleotides targeting each gene, 392 unique sequences are used. Large quantities of oligonucleotides with unique sequences are purchased from libraries from LC Sciences or CustomArray. However, the oligonucleotides synthesized by the array are trace and insufficient for in situ hybridization. Their amplification protocols are described below.

[0210] Each primary probe contains three components: flanking primer sequences that allow enzymatic amplification of the probe, a targeting sequence for in situ hybridization with mRNA, and a secondary tag sequence containing one or more read sequences for sequential readout of the codewords.

[0211] The following are examples of primary probes:

[0212]

[0213] (SEQ ID NO:1)

[0214] Components are arranged in the following order: forward primer (not underlined), secondary reading sequence 1 (underlined), mRNA targeting sequence (not underlined), secondary reading sequence 2 (underlined) and reverse primer (not underlined). The secondary reading sequence is the reverse complementary sequence of the corresponding secondary probe. Since only the code words containing four '1' are used, the primary probe of every kind of mRNA in this example needs to contain 4 different secondary reading sequences. However, in order to reduce the total length of the primary probe, the target sequence library of every kind of mRNA target is randomly divided into two libraries. Two secondary reading sequences are connected to each probe in one of the two libraries, and the other two secondary reading sequences are connected to the probe in the other library. The design criteria of each component are as described below.

[0215] Primer design. Specific index primers were generated by the collection of 240,000 published sequences of orthogonal 25-bp long sequences. These sequences were trimmed to 20bp, and the sequences were selected for narrow melting temperatures of 70 to 80°C, the absence of continuous repetitions of 3 or more bases, and the presence of a GC clamp (i.e., one of the two 3' terminal bases must be G or C). In order to further improve specificity, these sequences were subsequently screened for the human genome using BLAST+ (Camacho et al. 2009), and primers with 14 or more continuous base homologies were eliminated. In the screening by BLAST+ subsequently, primers with 11 or more continuous bases or more than 5 bases on the 3' ends of any other primer or T7 promoter were also removed.

[0216] Secondary probe design. A 30-bp long secondary probe sequence was generated by concatenating the fragments of the above-mentioned orthogonal primer set. These secondary molecules were then screened for orthogonality to other secondary molecules (no more than 11 base pairs of homology) and potential off-target binding sites in the human genome (no more than 14 base pairs of homology). The secondary sequences used in this example are provided in Table 1.

[0217] Table 1

[0218]

[0219]

[0220]

[0221] mRNA targeting sequence design. In order to determine the relative abundance of all isoforms of all genes expressed in these cell lines, transcriptome analysis data from the ENCODE project for total RNA from A549 and IMR90 cells were processed using the publicly available software cufflinks, and human genome annotations from gencode v18. Gene models corresponding to the most highly expressed isoforms were used to construct a FASTA format sequence library that records the dominant isoforms of each gene. Target genes were selected from this library. These genes were divided into 1 kb segments, and then the software OligoArray2.1 was used to generate primary probe sequences for the human transcriptome, which had the following restrictions: 30-bp or 40-bp in length, depending on the experiment; probe-target melting temperature greater than 70°C (variable parameter); no cross-hybridization targets with a melting temperature greater than 72°C (variable parameter); no predicted internal secondary structures with a melting temperature greater than 76°C (variable parameter); and no continuous repeats of single nucleotides of 6 or more bases. After OligoArray probe selection, all potential probes that were mapped to different genes were rejected, yet all potential probes with multiple alignments to the same gene were retained. BLAST databases were assembled from the FASTA libraries of all expressed genes to screen for the uniqueness of the probes. For each gene, 14 to 28 targeting sequences generated during the OligoArray treatment were selected.

[0222] Probe synthesis-index PCR. The template of the specific probe group is selected from the complex oligonucleotide library by limited cycle PCR. In brief, 0.5 to 1ng of the complex oligonucleotide library is combined with 0.5 micromoles of each primer. The forward primer matches the priming sequence of the desired subset, and the reverse primer is the 5' series of the sequence and the T7 promoter. In order to avoid the generation of G-quadruplexes that may be difficult to synthesize, the terminal G required in the T7 promoter is generated from the G located at the 5' of the appropriate priming region. All primers are synthesized by IDT. A 50 microliter reaction volume is amplified using a KAPA real-time library amplification kit (KAPA Biosystems; KK2701) or by a homemade qPCR mixture including 0.8X EvaGreen (Biotum; 31000-T) and a hot start Phusion polymerase (New England Biolabs; M0535S). Amplification is performed in real time using Agilent's MX300P or Biorad's CFXConnect. Individual samples are removed immediately before the amplification plateau to minimize the distortion of template abundance caused by over-amplification. Individual templates were purified using columns according to the manufacturer's instructions (Zymo DNA Clean and Concentrator; D4003) and eluted in RNase-free deionized water.

[0223] Amplification by in vitro transcription. Subsequently, template was amplified by in vitro transcription. In brief, 0.5 to 1 microgram of template DNA was amplified to 100-200 micrograms of RNA in a single 20-30 microliter reaction with high-yield RNA polymerase (New England Biolabs; E2040S). The reaction was supplemented with 1X RNAse inhibitor (Promega RNasin; N2611). Amplification was usually performed at 37°C for 4 to 16 hours to maximize yield. RNA was not purified after the reaction, and was stored at -80°C or converted into DNA immediately as described below.

[0224] Reverse transcription. Use reverse transcriptase Maxima H- (Thermo Scientific; EP0751) to produce 1 to 2 nmol of fluorescently labeled ssDNA probe from the above in vitro transcription reaction. This enzyme is used because it has high continuous synthesis ability and temperature resistance, which allows a large amount of RNA to be converted into DNA in a small volume at a temperature that is not conducive to the formation of secondary structures. Use 1.6mM of each dNTP, 1-2nmol of fluorescently labeled forward primer, 300 units of Maxima H-, 60 units of RNasin and a final 1X concentration of Maxima RT buffer to supplement the unpurified RNA generated above. The final 75 microliter volume is incubated at 50°C for 60 minutes.

[0225] Chain selection and purification. The template RNA in the above reaction was subsequently removed from the DNA by alkaline hydrolysis. 75 microliters of 0.25M EDTA and 0.5NNaOH were added to each reverse transcription reaction, and the samples were incubated at 95°C for 10 minutes. The reaction was immediately neutralized by purifying the ssDNA probe with a modified version of the Zymo Oligo Clean and Concentrator protocol. Specifically, the 5 microgram capacity column was replaced with a 25 microgram or 100 microgram capacity DNA column, as appropriate. The remainder of the protocol was performed according to the manufacturer's instructions. The probe was eluted in 100 microliters of RNase-free deionized water and then evaporated in a vacuum concentrator. The final pellet was resuspended in 10 microliters of RNase-free water and stored at -20°C. Denaturing polyacrylimide gel electrophoresis and absorption spectroscopy showed that the protocol typically produced 90-100% incorporation of fluorescent primers into full-length probes and 75-90% recovery of total fluorescent probes. Therefore, this protocol can be used to generate ∼2 nmol of fluorescent probe in a reaction volume of no more than 150 μl.

[0226] Cell culture and fixation. A549 and IMR90 cells (American Type Culture Collection) were cultured in Dulbecco's modified Eagle's medium and Eagle's minimum essential medium, respectively. Cells were incubated at 37°C and 5% CO2 Cells were fixed in 3% paraformaldehyde (Electron Microscopy Sciences) in PBS for 15 min, washed with PBS, and permeabilized in 70°C ethanol at 4°C overnight.

[0227] Fluorescence in situ hybridization (FISH) - primary (encoding) probes. Cells were hydrated in wash buffer (2xSSC, 50% formamide) for 10 minutes, labeled with primary oligonucleotides (0.5 nM / sequence) in hybridization buffer (2xSSC, 50% formamide, 1 mg / mL yeast tRNA and 10% dextran sulfate) at 37°C overnight, washed twice with wash buffer at 47°C for 10 minutes, and washed twice with 2xSSC. Prior to imaging, fluorescent reference beads (Molecular Probes, F-8809) were added at a dilution of 1:10,000 in 2xSSC.

[0228] Secondary probe. The secondary (readout) probe (10 nM) was hybridized to its primary target in secondary hybridization buffer (2xSSC, 20% formamide and 10% dextran sulfate) at 37°C for 30 minutes. During hybridization, the cells remained on the microscope stage. The temperature was maintained at 37°C using an objective heater. The cells were washed with secondary wash buffer (2xSSC, 20% formamide).

[0229] Fluidics and STORM imaging. Multiple rounds of continuous labeling, washing, imaging, and bleaching were performed on an automated platform consisting of a fluidics setup and a STORM (Stochastic Optical Reconstruction Microscopy) microscope. The fluidics setup included a flow chamber (Bioptech FCS2), a peristaltic pump (Rainin Dynamax RP-1), and three computer-controlled 8-way valves (Hamilton MVP and Hamilton HVXM 8-5). The system allows automated integration of secondary hybridization and STORM image acquisition.

[0230] The imaging buffer included 50 mM Tris (pH 8), 10% (w / v) glucose, 1% βME (2-mercaptoethanol) or 25 mM MEA (with or without 2 mM 1,5-cyclooctadiene) and an oxygen scavenging system (0.5 mg / ml glucose oxidase (Sigma-Aldrich) and 40 micrograms / ml catalase (Sigma-Aldrich). A layer of mineral oil was used to seal the imaging buffer to prevent acidification during multiple hybridizations.

[0231] The STORM setup includes an Olympus IX-71 inverted microscope configured for oblique incidence excitation. The sample is continuously illuminated with a 642nm diode-pumped solid-state laser (VFL-P500-642; MPB Communications). A 405nm solid-state laser (Cube405-100C; Coherent) is used for dye activation. Fluorescence is collected using an Olympus (UPlanSApo 100x, 1.4NA) objective lens and passed through a custom dichroic and four-view beam splitter. All images are recorded using an EMCCD camera (Andor iAxon897) and imaged at 60Hz. Before saving, the camera's 512x256 field of view is split into separate 256x256 pixel images. The left half of the field of view contains the STORM data, and the right half contains an image of the fluorescent reference beads. These latter images are downsampled to 1Hz before saving. During data acquisition, a homemade focus lock is used to maintain a constant focal plane. STORM images consisted of 20,000 to 30,000 frames in STORM buffer, while bleach images consisted of 10,000 frames in wash buffer.

[0232] Image Analysis - Analysis of Single-molecule Localization Images of single-molecule localization and fluorescent fiducial beads were processed separately using previously published single emitter localization software.

[0233] Image Registration. The starting positions of the beads from each round of hybridization are used to align the images from each round. 2D autocorrelation between consecutive hybridized bead images is used, followed by nearest neighbor matching to match beads between images. Pairs of beads with the most similar displacement vectors are used to compute rigid translation-rotation bending to align the beads. This alignment method is robust to samples in which multiple primitives are displaced or detached and reattached during imaging.

[0234] Drift correction. Drift during image acquisition was corrected using the trajectory of the reference beads (recorded at 1 Hz). The bead positions were concatenated in each frame. The trajectory of the two beads that moved in the most correlated manner was taken as the drift trajectory.

[0235] mRNA cluster calling.Location is first screened as photons above a threshold number (usually 2000), and needs to be within 32nm (adjustable parameters) of 5 other locations.The remaining molecules are located in a 2D histogram of 10×10nm boxes (box size is a variable parameter) and binned.All connected boxes are considered as part of a cluster (diagonal contacts are classified as connected).Clusters need to have more than 80 global locations (variable parameters) in all hybridizations to be called mRNA clusters.The weighted centroids of these clusters from the 2D histogram are recorded as mRNA positions.

[0236] A given cluster was recorded as representative in a single hybridization round if more than 9 localizations were found within a 48 nm radius (variable parameter) of the centroid of that mRNA in each hybridization round.

[0237] Cluster decoding. For each mRNA cluster, a codeword is read out, which includes "0" (for all hybridization rounds in which less than a threshold number of localizations are found near the centroid) and "1" (for rounds in which the above threshold number of localizations are counted). The SECDED codebook decodes these as exact matches to the target mRNA codeword, correctable errors that can be unambiguously mapped back to the target mRNA, or uncorrectable errors that differ from a word in the codebook by two or more letters.

[0238] Figure 2A STORM images of cells are shown. Figure 2B show Figure 2A Magnification of the boxed area in . Each spot represents localization. Localizations from different rounds of imaging are differentially shown. Figure 2C Show from Figure 2B Representative localization cluster of the boxed region in . The cluster shows localization signals from 4 different hybridizations. The cluster is a putative mRNA encoded with the codeword [0 1 0 1 1 1 0 0]. Figure 2D Reconstructed cell images of 14 genes after decoding and error correction are shown. Each gene is displayed differentially. Figure 2E Measured gene expression of 14 genes from cells is shown. Figure 2F Comparison of transcript counts to the entire RNA-sequencing data is shown. Figure 2G The correlation of transcript expression levels between two cells detected using the method is shown.

[0239] Example 4

[0240] The following examples generally relate to multiple single-molecule imaging with error robustness coding, which allows thousands of RNA species to be measured simultaneously in a single cell. Generally, the knowledge of the expression profile and spatial landscape of RNA in individual cells is necessary for understanding the rich repertoire of cellular behavior. The following examples report various techniques related to single-molecule imaging methods, which allow the copy number and spatial localization of thousands of RNA species to be determined in a single cell. Some of these techniques are referred to as multiplexed error-robust fluorescence in situ hybridization or "MERFISH".

[0241] By using an error-robust encoding scheme to combat single-molecule labeling and detection errors, these examples show the imaging of hundreds of unique RNA species in hundreds of individual cells. 4 to 10 6Correlation analysis of genes has allowed the constraint of gene regulatory networks, the prediction of novel functions of many unannotated genes, and the identification of distinct spatial distribution patterns of RNA that are associated with the properties of the encoded proteins.

[0242] Systematic analysis of the abundance and spatial organization of RNA in single cells is expected to transform the understanding of many areas of cell and developmental biology, such as the mechanisms of gene regulation, the heterogeneous behavior of cells, and the development and maintenance of cell fate. Single-molecule fluorescence in situ hybridization (smFISH) has become a powerful tool for studying the copy number and spatial organization of RNA in single cells isolated or in their natural tissue environment. By utilizing its ability to map the spatial distribution of specific RNAs with high resolution, smFISH has revealed the importance of subcellular RNA localization in different processes such as cell migration, development, and polarization. At the same time, the ability of smFISH to accurately measure the copy number of specific RNAs without amplification bias allows quantitative measurement of natural fluctuations in gene expression, which in turn clarifies the regulatory mechanisms that form such fluctuations and their role in a variety of biological processes.

[0243] However, the application of smFISH methods to many systems-level questions is still limited by the number of RNA species that can be measured simultaneously in a single cell. State-of-the-art efforts using combinatorial labeling via color-based barcoding or sequential hybridization have been able to measure 10-30 different RNA species simultaneously in individual cells, but many interesting biological questions will benefit from the measurement of hundreds to thousands of RNAs within a single cell, which is not achievable using such techniques. For example, analysis of how the expression profiles of such a large number of RNAs vary from cell to cell and how these variations are correlated between different genes can be used to systematically identify co-regulated genes and map regulatory networks; knowledge of the subcellular organization of many RNAs and their correlations can help elucidate the molecular mechanisms behind the establishment and maintenance of many local cellular structures; and characterization of the RNA profiles of individual cells in native tissues can allow in situ identification of cell types.

[0244] The following examples generally discuss certain techniques called MERFISH, which is a highly multiplexed smFISH imaging method that significantly increases the number of RNA species that can be imaged simultaneously in individual cells by using a combination of labeling and sequential imaging that utilizes an error robustness encoding scheme. These examples simultaneously measure 140 kinds of RNA species using an encoding scheme that can detect and correct errors, and simultaneously measure 1001 kinds of RNA species using an encoding scheme that can detect but not correct errors to demonstrate the multiplexed imaging method. It should be understood that these numbers are only exemplary, not restrictive. The copy number variation of these genes and the correlation analysis of spatial distribution allow us to identify groups of genes that are co-regulated and groups of genes that share similar spatial distribution patterns in cells.

[0245] Combinatorial labeling using error-robust encoding schemes. Combinatorial labeling that identifies each RNA species by multiple (N) different signals provides a way to rapidly increase the number of RNA species that can be simultaneously detected in individual cells ( Figure 5A However, scaling up the throughput of smFISH to a system-scale approach faces significant challenges, as not only does the number of addressable RNA species increase exponentially with N, but the detection error rate also increases exponentially with N ( Figures 5B-5D ). Imagine a conceptually simple scheme to implement combinatorial labeling, where each RNA species is encoded with an N-bit binary word, and the sample is probed with N corresponding hybridization rounds, each round targeting only the subset of RNAs that should read '1' in the corresponding bit ( Fig.11 ). N rounds of hybridization will allow detection of 2 N -1 RNA. Using only 16 hybridizations, more than 64,000 RNA species can be identified, which should cover the entire human transcriptome including messenger RNA (mRNA) and non-coding RNA ( Figure 5B ; upper symbols). However, as N increases, the fraction of RNAs that are properly detected (the call rate) will decrease rapidly, and more troublesomely, the fraction of RNAs that are identified as incorrect species (the misidentification rate) will increase rapidly ( Figure 5C , the following symbols; Figure 5D , symbols above). By using the actual error rate for each hybridization (measured below), most RNA molecules will be misidentified after 16 rounds of hybridization!

[0246] To address this challenge, an error-robust coding scheme is designed in which only two pairs of N -1 words to encode RNA. In a codebook where the minimum Hamming distance is 4 (HD4 code), at least 4 bits must be read incorrectly to change one codeword to another ( Fig. 12A As a result, each single-bit error produces a word that is uniquely close to a single codeword, which allows detection and correction of such errors ( Fig. 12B ). Double-bit errors produce words from multiple code words with equal Hamming distance of 2 and thus can be detected but not corrected ( Fig. 12C Such a code should significantly increase the recall rate and reduce the false positive rate ( Figure 5C and 5D, middle symbol). To further account for the fact that hybridization events are more likely to be missed (1-->0 errors) than to misidentify background spots as RNA (0-->1 errors) in smFISH measurements, a modified HD4 (MHD4) code was designed in which the number of '1' bits is kept constant and relatively low (only 4 per word) to reduce errors and avoid biased detection. This MHD4 code should further increase the call rate and reduce the false identification rate ( Figure 5C , the symbol above; Figure 5D , the symbol below).

[0247] In addition to error considerations, several practical challenges make it difficult to detect a large number of RNA species, such as the high cost of the large number of fluorescently labeled FISH probes required and the long time required to complete many rounds of hybridization. To overcome these challenges, in this example, a two-step labeling scheme was designed to encode and read out cellular RNA ( Figure 5E ). First, cellular RNA is labeled with a set of encoding probes (also called primary probes), each of which contains an RNA targeting sequence and two flanking readout sequences. Four of N unique readout sequences are assigned to each RNA species based on the RNA-based MHD4 codeword. Second, these N readout sequences are identified through N rounds of hybridization and imaging (each round using a unique readout probe) using complementary FISH probes (readout probes (also called secondary probes)). To increase the signal-to-background ratio, each cellular RNA is labeled with ~192 encoding probes. Because each encoding probe contains two of the four readout sequences associated with that RNA ( Figure 5E ), so that up to ~96 readout probes can bind to each cellular RNA in each hybridization round. To generate the large number of encoding probes required, they are amplified from array-derived oligonucleotide pools containing tens of thousands of custom sequences using an enzymatic amplification method involving in vitro transcription followed by reverse transcription ( Fig.13 , see below for probe synthesis). This two-step labeling approach significantly reduced the total hybridization time for experiments: efficient hybridization to the readout sequence was found to take only 15 minutes, whereas efficient direct hybridization to cellular RNA required more than 10 hours.

[0248] FIG5 depicts MERFISH, a highly multiplexed smFISH method using combined labeling and error robust encoding. Figure 5A A schematic depiction of identifying multiple RNA species in N rounds of imaging is shown. Each RNA species is encoded with an N-bit binary word, and during each round of imaging, only a subset of RNAs that should read '1' in the corresponding bit emit a signal. Figures 5B-5D The number of addressable RNA species is shown ( Figure 5B ), the ratio of these RNAs to be correctly identified (call rate) ( Figure 5C) and the ratio of RNA being incorrectly identified as different RNA species (misidentification rate) ( Figure 5D ) (plotted as a function of the number of bits (N) in the binary word encoding the RNA). Figure 5B and 5D The points above include all 2 N -1 possible binary words; the middle point is the HD4 code where the Hamming distance separating the words is 4; and the lower point is the modified HD4 (MHD4) code where the number of '1' bits remains at 4. These are Figure 5C It's the other way around.

[0249] The call and misidentification rates were calculated at a bit error rate of 10% (for 1-->0 errors) and 4% (for 0-->1 errors). Figure 5E is a schematic diagram of the implementation of the MHD4 code for RNA identification. Each RNA species is first labeled with ~192 encoding probes that convert the RNA into a unique combination of readout sequences (encoding hybs). Each of these encoding probes contains a central RNA targeting region flanked by two readout sequences extracted from a library of N different sequences, each of which is associated with a specific hybridization round. The encoding probe for a particular RNA species contains a unique combination of 4 of the N readout sequences, which corresponds to the 4 hybridization rounds in which the RNA should read as '1'. The readout sequences (hyb 1, hyb 2, ..., hyb N) are detected using N subsequent rounds of hybridization using fluorescent readout probes. The bound probes are inactivated by photobleaching between consecutive rounds of hybridization. For clarity, only one possible pairing of readout sequences is described here for the encoding probes; however, all possible pairs of the 4 readout sequences are used with the same frequency and are randomly distributed along each cellular RNA in actual experiments.

[0250] Fig.11 A schematic depiction of a combinatorial marking method based on a simple binary code is shown. In a conceptually simple marking method, 2 N -1 different RNA species can be uniquely encoded with all N bit binary words (excluding words with all '0'). In each hybridization round, FISH probes targeted to all RNA species with '1' in the corresponding bit are included. To increase the ability to distinguish RNA spots from background, each RNA is addressed with multiple FISH probes per hybridization round. The signal from the bound probe disappears before the next round of hybridization. This process is continued for all N hybridization rounds (hyb 1, hyb 2, ...), and all 2 can be identified by the unique on-off pattern of the fluorescent signal in each hybridization round. N -1 RNA species.

[0251] FIG. 12 shows a schematic depiction of the Hamming distance and its use in the identification and correction of errors. Fig. 12A This is a schematic diagram of a Hamming distance of 4. Fig. 12B and 12C is a diagram showing that the coding scheme using a Hamming distance of 4 corrects the single-bit error ( Fig. 12B ) or detect but not correct double-bit errors ( Fig. 12C ) is a schematic diagram of the ability of the measured word to be used in conjunction with a human. The arrows highlight the bits where the words indicated thereon differ. If one of the words must flip 4 bits from '1' to '0' or from '0' to '1' to convert into the other word, the two code words are separated by a Hamming distance of 4. Single-bit error correction is possible because if the measured word differs from the correct code word by only one bit, it is likely an error due to misreading the code word because the code words of all other RNA species will differ from the measured word by at least 3 bits. In this case, the measured word can be corrected to a code word that differs by only 1 bit. If the measured word differs from the correct code word by 2 bits, the measured word can still be identified as an error, but correction is no longer possible because more than one correct code word differs from the measured word by 2 bits.

[0252] Fig.13 The generation of coding probe library is shown. A complex oligonucleotide library synthesized by array containing ~100k sequences is used as a template for enzymatic amplification of coding probes (for different experiments). Each template sequence in the oligonucleotide library contains a central target region that can bind to cellular RNA, two flanking readout sequences and two flanking index primers. In the first step, the desired template molecule for a specific experiment is selected and amplified with an index PCR reaction. In order to allow amplification by in vitro transcription, a T7 promoter is added to the PCR product during this step. In the second step, RNA is amplified from these template molecules by in vitro transcription. In the third step, the RNA is reversely transcribed back to DNA. In the last step, the template RNA is removed by alkaline hydrolysis, leaving only the desired ssDNA probe. This scheme produces a complex library of ~2nmol coding probes, which contains ~20,000 different sequences for 140-gene experiments or ~100,000 different sequences for 1001-gene experiments.

[0253] Example 5

[0254] This example illustrates the use of a 16-bit MHD4 code to measure 140 genes using MERFISH. To test the feasibility of this error-robust multiplexed imaging method, this example uses a 140-gene measurement on human fibroblasts (IMR90), using a 16-bit MHD4 code to encode 130 RNA species, while leaving 10 code words as misidentification controls (Figure 20). After each round of hybridization with the fluorescent readout probe, the cells were imaged by conventional widefield imaging using oblique-incidence illumination geometry. Fluorescent spots corresponding to individual RNAs were clearly detected and then effectively extinguished by a brief photobleaching step ( Fig. 6A The samples were stable throughout the 16 rounds of iterative labeling and imaging. No changes in the number of fluorescent spots between rounds were observed that matched the changes predicted based on the relative abundance of the targeted RNA species in each round derived from bulk sequencing, and no systematic trend of decrease with increasing number of hybridization rounds was observed ( Fig.14A The average brightness of the spots varied between rounds with a standard deviation of 40%, which may be attributed to the different binding efficiencies of the readout probes to different readout sequences on the encoding probes ( Fig. 14B Only a small systematic decrease in spot brightness with increasing number of hybridization rounds was observed, which was an average of 4% per round ( Fig. 14B ).

[0255] Next, binary words were constructed from the observed fluorescent spots based on their on-off patterns over 16 hybridization rounds ( Figure 6B-6D If the word matches one of the 140 MHD4 code words exactly (exact match) or differs by only 1 bit (error-correctable match), it is assigned to the corresponding RNA species ( Fig.6D ).exist Fig. 6A and 6B In a single cell depicted in Figure 2, more than 1500 RNA molecules corresponding to 87% of the 130 encoded RNA species were detected after error correction ( Fig. 6E ). Similar observations were made in ~400 cells from 7 independent experiments. On average, about 4 times more RNA molecules and about 2 times more RNA species were detected per cell after error correction compared to the values ​​obtained before error correction ( FIG. 15 ).

[0256] Two types of errors can occur in the copy number measurement of each RNA species: 1) some molecules of that RNA species are not detected, resulting in a reduced call rate, and 2) some molecules from other RNA species are misidentified as that RNA species. To assess the extent of misidentification, 10 misidentified control words, i.e., code words that are not associated with any cellular RNA, were used. Although matches to these control words were observed, they occurred much less frequently than the real RNA code words: 95% of the 130 RNA code words were counted more frequently than the median count of these control words. In addition, it was generally found that for the actual RNA code words, the ratio of the number of exact matches to the number of matches with 1-bit errors was significantly higher than the same ratio observed for the misidentified controls, as expected ( Fig.16A and 16B By using this ratio as a measure of the confidence of the RNA identification, it was found that 91% of the 130 RNA species had confidence ratios greater than the maximum confidence ratio observed for the misidentification control ( Fig. 6F ), indicating high accuracy of RNA identification. Subsequent analysis was performed only on these 91% genes.

[0257] To evaluate the call rate, the error correction capability of the MHD4 code was used to determine the 1-->0 error rate (average 10%) and 0-->1 error rate (average 4%) for each hybridization round ( Fig. 16C and 16D ). By using these error rates, a call rate of about 80% for individual RNA species after error correction was estimated, i.e., about 80% of the fluorescent spots corresponding to the RNA species were correctly decoded ( Fig.16E ). Note that although the remaining 20% ​​of spots result in a loss in detection efficiency, most of them do not cause misidentification of species because they are decoded as double-bit error words and discarded.

[0258] To test for potential technical bias in these measurements, the same 130 RNA species were probed with a different MHD4 codebook by shuffling the codewords between different RNA species ( FIG. 20 ) and varying the coding probe sequences. Measurements using this alternative codebook gave similar misidentification and call rates ( FIG. 17 ). The copy numbers of individual RNA species per cell measured with these two codebooks showed excellent agreement, with a Pearson correlation coefficient of 0.94 ( Figure 6G ), which indicates that the choice of coding scheme does not bias the measured counts.

[0259] To validate the copy numbers obtained from the MERFISH experiments, conventional smFISH measurements were performed for 15 of the 130 genes (selected from the full measured abundance range of 3 orders of magnitude). For each of these genes, the mean copy number and copy number distribution across many cells were quantitatively consistent between MERFISH and conventional smFISH measurements ( Fig.18A and 18B The ratio of copy numbers determined by these two methods was 0.82 + / - 0.06 (mean + / - SEM across 15 measured RNA species, Fig.18B ), which is consistent with the estimated 80% call rate for the multiplex imaging method. The quantitative match between this ratio and the estimated call rate over the full range of measured abundance additionally supports the assessment that the misidentification error is low. It is assumed that the agreement between MERFISH and conventional smFISH results extends to genes with the lowest measured abundance (<1 copy per cell, Fig.18B ), with an estimated measurement sensitivity of at least 1 copy per cell.

[0260] As a final validation, the abundance of each RNA species averaged across hundreds of cells was compared to those obtained from bulk RNA sequencing measurements of the same cell line. The imaging results correlated significantly with the bulk sequencing results, with a Pearson correlation coefficient of 0.89 ( Figure 6H ).

[0261] FIG6 shows the simultaneous measurement of 140 RNAs in individual cells using MERFISH utilizing a 16-bit MHD4 code. Fig. 6A Images of RNA molecules in IMR90 cells after each hybridization round (hyb 1-hyb 16) are shown. Images after photobleaching (bleach 1) demonstrate the efficient removal of fluorescent signal between hybridizations. Figure 6B The localization of all detected single molecules in this stained cell based on their measured binary words is shown. Inset: Composite fluorescence image of 16 hybridization rounds, boxed sub-regions with numbered circles indicate potential RNA molecules. Circles indicate unidentifiable molecules whose binary words do not match any of the 16-bit MHD4 code words even after error correction. Figure 6C show Figure 6B The boxed subregions in the fluorescence images are from each round of hybridization, where circles indicate potential RNA molecules. Fig.6D show Figure 6C The corresponding words of the spots identified in . The crosses indicate the corrected bits. Fig. 6E Shown are the RNA copy numbers for each gene observed in this cell without (lower) or with (higher) error correction. Fig. 6FShown are confidence ratios measured for 130 RNA species (left) and 10 misidentified control characters (right), normalized to the maximum value observed from the misidentified control (dashed line). Figure 6G is a scatter plot of the average copy number of each RNA species per cell measured using two shuffled codebooks of the MHD4 code. The Pearson correlation coefficient is 0.94 and the p-value is 1×10 -53 The dashed line corresponds to the y=x line. Figure 6H is a scatter plot of the average copy number of each RNA species per cell versus its abundance in fragments per kilobase per million reads (FPKM) as determined by bulk sequencing. The Pearson correlation coefficient between the log abundances of the two measurements was 0.89 with a p-value of 3x10 -39 .

[0262] FIG. 14 shows the number and average brightness of fluorescent spots detected in 16 rounds of hybridization before and after bleaching. Fig.14A The number of fluorescent spots observed per cell before (higher) and after (lower) photobleaching is shown as a function of hybridization rounds averaged across all measurements using the first 16-bit MHD4 code. Photobleaching reduces the number of fluorescent spots by two or more orders of magnitude. Hybridization rounds without a lower bar represent rounds where no molecules were observed after bleaching. Also depicted is the expected variation in the number of fluorescent spots between rounds (cycles) predicted based on the relative abundance of the targeted RNA species in each hybridization round derived from bulk RNA sequencing. The average difference between the observed and predicted number of spots per hybridization is only 15% of the average number of spots. This difference does not increase systematically with increasing number of hybridization rounds. Fig. 14B The average brightness of the fluorescent spots identified in each hybridization round, averaged across all measurements using the first 16-bit MHD4 code, is shown before (top) and after (bottom) photobleaching. The brightness variation between different hybridization rounds is 40% (standard deviation). The variation pattern is reproducible between experiments with the same code, which may be attributed to differences in the binding efficiency of the readout probes to different readout sequences. There is a small systematic trend of decreasing brightness with increasing hybridization rounds, which is an average of 4% per round. Photobleaching extinguishes the fluorescence to a level similar to the autofluorescence of the cells.

[0263] FIG. 15 shows that error correction significantly increases the number of RNA molecules and RNA species detected in individual cells. Fig.15A Histogram showing the ratio of the total number of molecules detected per cell with error correction to the number measured without error correction. Fig. 15Bis a histogram of the ratio of the total number of RNA species detected per cell with error correction to the total number of RNA species detected without error correction. Both ratios were determined for ~200 cells and the histograms were constructed from these ratios.

[0264] FIG. 16 shows a characterization of the misidentification and call rates of RNA species for a 140-gene experiment using a specific 16-bit MHD4 code. Fig.16A The number of measured words that exactly match the codeword corresponding to FLNA is shown, which is represented by the bar in the center of the circle, and the number of measured words that have a 1-bit error compared to the codeword of FLNA, which is represented by the 16 bars on the circle. In addition to the codewords that are not assigned to any RNA (i.e., the misidentified control words), Fig. 16B and Fig.16A The solid line connects the exact match to the 1-bit error word generated by the 1->0 error. Based on the observation that the ratio of the number of exact matches to the number of error-correctable matches for the actual RNA code words is generally significantly higher than the same ratio observed for the misidentified controls, this ratio is defined as the confidence ratio for the RNA identification. The confidence ratios measured using the 16-bit MHD4 code for all 130 RNA species (center bar) and 10 misidentified control words not assigned to any RNA (outer bars) are shown in Fig. 6F middle. Fig. 16C and 16D Displays the 1-->0 error for each hybridization round ( Fig. 16C ) and 0-->1 error ( Fig.16D )’s average error rate. Fig.16E The call rate of each RNA species estimated from the 1-->0 and 0-->1 error rates is shown. Genes are ranked from left to right based on the measured abundance (which spans 3 orders of magnitude). The call rate is largely independent of the abundance of the gene.

[0265] Figure 17 shows a characterization of the false identification rate and the call rate for the second 16-bit MHD4 code. In this second encoding scheme, the 140 code words were shuffled between different RNA species and the encoding probe sequences were varied. Fig.17A Shown are normalized confidence ratios measured for 130 RNA species (left) and 10 misidentified control words not assigned to any RNA (right). Fig. 6F The normalized confidence ratio was determined in the same manner as in . Fig. 17B and 17C The error for 1-->0 for each hybridization round is shown ( Fig. 17B ) and 0-->1 error ( Fig. 17C ) was used to determine the average error rate. Fig.17DThe call rates determined for each RNA species estimated from the 1-->0 and 0-->1 error rates are shown. Genes are ordered from left to right based on measured abundance.

[0266] FIG. 18 shows MERFISH measurements of a subset of genes compared to conventional smFISH results. Fig.18A RNA copy number distributions in single cells are shown for 3 exemplary genes KIAA1199, DYNC1H1 and LMTK2 in the high, medium and low abundance ranges, respectively. Lighter bars: distribution constructed from ˜400 cells in a 140-gene measurement using the MHD4 code. Darker bars: distribution constructed from ˜100 cells in a conventional smFISH measurement. Fig.18B Shown is a comparison of the average RNA copy number per cell measured in a 140-gene experiment using the MHD4 code with that determined by conventional smFISH for 15 genes. The average ratio of copy number measured using the MHD4 measurement to copy number measured using conventional smFISH is 0.82 + / - 0.06 (mean + / - SEM for 15 genes). The dashed line corresponds to the y = x line.

[0267] Figure 20 shows two different codebooks used for the 140-gene experiment. The specific codewords of the 16-bit MHD4 code assigned to each RNA species were studied in two shuffles of the 140-gene experiment. The "Gene" column contains the name of the gene. The "Codeword" column contains the specific binary word assigned to each gene.

[0268] Example 6

[0269] This example generally relates to high-throughput analysis of cell-to-cell variation in gene expression. The MERFISH method allows for parallelization of measurements of many individual RNA species and analysis of covariation between different RNA species. In this example, the parallelization aspect is first illustrated by examining the cell-to-cell variation in expression levels of each measured gene ( Fig. 7A To quantify the variation in the measurements, the Fano factor (defined as the ratio of the variance to the mean RNA copy number) was calculated for all measured RNA species. For many genes, the Fano factor deviated significantly from 1 (the value expected for a simple Poisson approach) and showed a trend to increase with mean RNA abundance ( Figure 7B ). This trend of increasing Fano factors with average RNA abundance could be explained by changes in transcription rate and / or promoter-off switching rate rather than changes in promoter-on rate.

[0270] In addition, several RNA species were identified as having Fano factors much larger than this average trend. For example, SLC5A3, CENPF, MKI67, TNC, and KIAA1199 were found to show Fano factor values ​​significantly higher than the Fano factor values ​​of other genes expressed at similar abundance levels. The high variability of some of these genes can be explained by their association with the cell cycle. For example, two of these particularly "noisy" genes, MKI67 and CENPF, were annotated as cell cycle-related genes, and based on their bimodal expression ( Figure 7C ), whose transcription is believed to be strongly regulated by the cell cycle. Other highly variable genes did not show the same bimodal expression pattern and were not known to be cell cycle related.

[0271] Analyzing covariations in the expression levels of different genes can reveal which genes are co-regulated and elucidate gene regulatory pathways. At the population level, such analyses usually require the application of external stimuli to drive changes in gene expression; thus, correlated expression changes can be observed in genes that share common regulatory elements affected by the stimulus. At the single-cell level, natural random fluctuations in gene expression can be exploited to perform such analyses, and multiple regulatory networks can therefore be studied without having to stimulate each of them individually. Such covariation analyses can constrain regulatory networks, propose new regulatory pathways, and predict the functions of unannotated genes based on their association with co-varying genes.

[0272] This method was applied to a 140-gene measurement and ~10,000 pairwise correlation coefficients, which describe how to examine the expression levels of each pair of genes that co-vary between cells. Many highly variable genes showed tightly correlated or anti-correlated changes ( Figure 7C To better understand the correlation of all gene pairs, a hierarchical clustering approach was used to organize these genes based on their correlation coefficients ( Fig.7D From the cluster tree structure, seven groups of genes with significantly correlated expression patterns were identified ( Fig.7D ). In each of the 7 groups, each gene showed a significantly stronger average correlation with other members of the group than genes outside the group. To further validate and understand these groups, we identified the gene ontology (GO) terms enriched in each of the 7 groups. Notably, the GO terms enriched within each group shared similar functions and were largely unique to each group ( Fig. 7E ), thus validating the notion that the observed covariations in expression reflect some commonality in the regulation of these genes.

[0273] This example describes two of these groups as illustrative examples. The main GO terms associated with group 1 are those related to the extracellular matrix (ECM) ( Fig.7D and 7E). Notable members of this group include ECM components such as FBN1, FBN2, COL5A, COL7A and TNC, and glycoproteins that connect ECM to cell membranes such as VCAN and THBS1. This group also includes the unannotated gene KIAA1199, which can be predicted to play a role in ECM metabolism based on its association with this cluster. In fact, this gene has recently been identified as an enzyme involved in the regulation of hyaluronic acid, the main sugar component of ECM.

[0274] Group 6 contains many genes encoding vesicular transporters and proteins associated with cell motility ( Fig.7D and 7E ). Vesicle transport genes include microtubule motors and related genes DYNC1H, CKAP1, and factors related to vesicle formation and transport such as DNAJC13 and RAB3B. Once again, the unannotated gene KIAA1462 was found within this cluster. Based on its strong correlation with DYNC1H1 and DNAJC13, this gene is predicted to be involved in vesicle transport. Cell motility genes in this group include actin binding proteins such as AFAP1, SPTAN1, SPTBN1, and MYH10, as well as genes involved in adhesion complex formation, such as FLNA and FLNC. Several GTPase-related factors involved in the regulation of cell motility, attachment, and contraction also fall into this group, including DOCK7, ROCK2, IQGAP1, PRKCA, and AMOTL1. The observation that some cell motility genes are associated with vesicle transport genes is consistent with a role for vesicle transport in cell migration. Another interesting feature of group 6 is that a subset of these genes (particularly those associated with cell motility) are anti-correlated with members of the ECM group discussed above ( Fig.7D ). This anti-correlation may reflect regulatory interactions that mediate the transition of cells between adhesion and migration states.

[0275] FIG. 7 shows the cell-to-cell variation and pairwise correlations of RNA species determined from the 140-gene measurement. Fig. 7A Comparison of gene expression levels in cells from two individuals is shown. Figure 7B Fano factors for individual genes are shown. Error bars represent standard error of the mean determined from 7 independent data sets. Figure 7C Z-scores for expression changes of 4 pairs of exemplary genes are shown, showing correlated (upper two) or anti-correlated (lower two) changes in 100 randomly selected cells. Z-scores are defined as the difference from the mean normalized by the standard deviation. Fig.7DIt is a matrix of the pairwise correlation coefficients of the intercellular variation of the expression of the measured genes displayed together with the hierarchical clustering tree. The 7 groups identified by the specific threshold (dashed line) on the cluster tree are represented by the black box in the matrix and the line on the tree, and the gray line on the tree indicates the ungrouped genes. Different threshold selections can be carried out on the cluster tree to select smaller subgroups with closer correlation or larger supergroups containing more weakly coupled subgroups. Two groups in the 7 groups are enlarged on the right. Fig. 7E Shows enrichment of 30 selected statistically significantly enriched GO terms in 7 groups. Enrichment refers to the ratio of the score of genes within a group with a specific GO term to the score of all measured genes with that term. Not all GO terms provided here are in the top 10 lists.

[0276] Example 7

[0277] This example illustrates mapping the spatial distribution of RNA. As an imaging-based method, MERFISH also allows the spatial distribution of many RNA species to be studied simultaneously. Several patterns emerge from visual inspection of individual genes, with some RNA transcripts enriched in the perinuclear region, some enriched at the cell periphery, and some scattered throughout the cell ( Fig. 8A ). To identify genes with similar spatial distributions, the correlation coefficients of the spatial density distributions of all pairs of RNA species were determined, and hierarchical clustering was used again to organize these RNAs based on pairwise correlations. The correlation coefficient matrix shows groups of genes with correlated spatial organization, and the two most significant groups with the strongest correlations are shown in Figure 8B The RNA of group I was enriched in the perinuclear region, while the RNA of group II was enriched near the cell periphery ( Figure 8C Quantitative analysis of the distance between each RNA molecule and the nucleus or cell periphery indeed confirmed this visual impression ( Fig.8D ).

[0278] Group I contains genes encoding extracellular proteins such as FBN1, FBN2, and THSB1, secreted proteins such as PAPPA, and integral membrane proteins such as LRP1 and GPR107. These proteins have no obvious commonalities in function. In contrast, GO analysis showed significant enrichment for location terms such as extracellular region, basement membrane, or perivitelline space ( Fig. 8E To reach these locations, proteins must pass through the secretory pathway, which generally requires translation of mRNAs on the endoplasmic reticulum (ER). Therefore, it is believed that the observed spatial pattern of these mRNAs reflects their co-translational enrichment on the ER. These mRNAs are expressed in the perinuclear region ( Figure 8C and 8D , lighter shade) supports this conclusion.

[0279] Group II contains genes encoding actin-binding proteins, including filament proteins FLNA and FLNC, talin TLN1, and spectrin SPTAN1 and SPTBN1; microtubule-binding protein CKAP5; and dyneins MYH10 and DYNC1H1. This group is enriched for GO terms such as cortical actin cytoskeleton, actin filament binding, and cell-cell adhesion junctions ( Fig. 8E β-actin mRNA can be enriched near the cell periphery in fibroblasts, and mRNAs encoding members of the actin-binding Arp2 / 3 complex are also enriched near the cell periphery in fibroblasts. Enrichment of Group II mRNA in the cell periphery region ( Figure 8C and 8D ) suggest that the spatial distribution of group II genes may be related to the distribution of actin cytoskeleton mRNA.

[0280] FIG8 shows the different spatial distributions of RNA observed in the 140-gene survey. Fig. 8A Examples of the spatial distribution observed for four different RNA species in cells are shown. Figure 8B is a matrix of pairwise correlation coefficients describing the degree of correlation in the spatial distribution of each gene pair, displayed together with the hierarchical clustering tree. Two strongly correlated groups are indicated by black boxes on the matrix and shading on the tree. Figure 8C The spatial distribution of all RNAs in the two groups in two exemplary cells is shown. Lighter symbols: Group I genes; darker symbols: Group II genes. Fig.8D The mean distances of genes in group I and genes in group II to the cell edge or nucleus normalized to the mean distance of all genes are shown. Error bars represent SEM across 7 data sets. Fig. 8E Enrichment of GO terms in each of the two groups is shown.

[0281] Example 8

[0282] This example illustrates the measurement of 1001 genes using a 14-bit MHD2 code. This example further increases the throughput of MERFISH measurements by imaging ~1000 RNA species simultaneously. This increase can be achieved using MHD4 codes by increasing the number of bits per code word to 32 while keeping the number of "1" bits per word at 4 ( Figure 5B). Although the stability of the sample over many hybridization rounds ( FIG. 14 ) suggests that such an extension is potentially feasible, an alternative approach is shown here that does not require increasing the number of hybridizations by relaxing the error correction requirements, but rather maintains the error detection capability. For example, by reducing the Hamming distance from 4 to 2, 1001 genes can be encoded using all 14-bit words containing four '1' bits, and these RNAs are probed using only 14 rounds of hybridization. However, because a single error can produce words that are equally close to two different code words, error correction is no longer possible for this modified Hamming distance-2 (MHD2) code. Therefore, it is expected that the call rate will be lower and the misidentification rate will be higher using this encoding scheme.

[0283] To evaluate the performance of this 14-bit MHD2 code, 16 of the 1001 possible code words were retained as misidentification controls, and the remaining 985 words were used to encode cellular RNA. Included in these 985 RNAs were 107 RNA species probed in the 140-gene experiment as additional controls. The 1001-gene experiment was performed in IMR90 cells using a method similar to that described above. To allow all encoding probes to be synthesized from a single 100,000-member oligonucleotide pool, the number of encoding probes for each RNA species was reduced to ~94. Fluorescent spots corresponding to individual RNA molecules were detected again in each round of hybridization with the readout probe, and based on their on-off patterns, these spots were decoded as RNA ( Fig. 9A , 19A 9B). 430 RNA species were detected in the cells shown in FIG9, and similar results were obtained in ˜200 imaged cells in 3 independent experiments.

[0284] As expected, the misidentification rate for this scheme was higher than that for the MHD4 code. 77% of all true RNA words were detected with a frequency higher than the median count of the misidentified control (rather than the 95% value observed in the MHD4 measurement). Using the same confidence ratio analysis as above, it was found that 73% of the 985 RNA species were measured with a confidence ratio greater than the maximum value observed for the misidentified control (rather than the 91% measured for MHD4) ( Fig.19C ). RNA copy numbers measured from these 73% RNA species showed excellent correlation with bulk RNA sequencing results (Pearson correlation coefficient r = 0.76; Fig. 9B , black). Notably, the remaining 27% of genes still showed good (although low) correlation with bulk RNA sequencing data (r = 0.65; Fig. 9B , red), but with the conservative measure of excluding them from further analysis.

[0285] The lack of error correction also reduced the call rate of each RNA species: when comparing 107 RNA species that were common to the 1001-gene and 140-gene measurements, the copy numbers per cell of these RNA species were found to be lower in the 1001-gene measurement ( Fig. 9C and 19D ). The total counts of these RNAs per cell were ∼1 / 3 of the total observed in the 140-gene measurement. Thus, the lack of error correction in the MHD2 code produces a ∼3-fold reduction in call rate, which is consistent with the ∼4-fold reduction in call rate observed for the MHD4 code when error correction was not applied. As expected from the quantitative agreement between the 140-gene measurement and conventional smFISH results, comparison of the 1001-gene measurement with conventional smFISH results for 10 RNA species also showed a drop in call rate to approximately 1 / 3 ( Fig. 18C Despite the expected decrease in call rate, the good correlation found between the copy number observed in the 1001-gene measurement and that observed in the 140-gene measurement as well as in conventional smFISH and bulk RNA sequencing measurements suggests that the relative abundance of these RNAs can be quantified using the MHD2 encoding scheme.

[0286] Simultaneous imaging of ~1000 genes in individual cells significantly expands the ability to detect co-regulated genes. Fig. 10A A matrix showing pairwise correlation coefficients determined from cell-to-cell variation in expression levels of these genes. Using the same hierarchical cluster analysis as above, ~100 groups of genes with correlated expression were identified. Surprisingly, almost all of these ~100 groups showed statistically significant enrichment of functionally related GO terms ( Fig. 10B These include some groups identified in the 140-gene measurement, such as a group associated with cell replication genes and a group associated with cell motility genes ( Fig. 10A and 10B , groups 7 and 102), as well as many new groups. The groups identified here include 46 RNA species that lack any previous GO annotations, for which functions based on their group associations can be hypothesized. For example, KIAA1462 is part of the cell motility group, also shown in the 140-gene experiment, indicating a potential role for this gene in cell motility ( Fig. 10A , group 102). Similarly, KIAA0355 was part of a novel group enriched in genes associated with cardiac development ( Fig. 10A , group 79), C17orf70 is part of a group associated with ribosomal RNA processing ( Fig. 10A, group 22). By using these groupings, the cellular functions of 61 transcription factors and other partially annotated proteins of unknown function can be hypothesized. For example, transcription factors Z3CH13 and CHD8 are members of the cell motility group, indicating their potential role in the transcriptional regulation of cell motility genes.

[0287] FIG. 9 shows the simultaneous measurement of 1001 genes in a single cell using MERFISH utilizing a 14-bit MHD2 code. Fig. 9A The localization of all detected single molecules based on their measured binary words in the stained cell is shown. Inset: Composite false-color fluorescence image of 14 hybridization rounds for the boxed sub-region with numbered circles indicating potential RNA molecules. Circles represent unidentifiable molecules whose binary words do not match any of the 14-bit MHD2 code words. Images of individual hybridization rounds are shown in Fig.19A middle. Fig. 9B is a scatter plot of the average copy number per cell measured in the 1001-gene experiment versus the abundance measured by bulk sequencing. The upper symbol is for 73% of the genes detected with a confidence ratio above the maximum ratio observed for the misidentified control. The Pearson correlation coefficient is 0.76 and the p-value is 3 × 10 -133 The symbols below are for the remaining 27% of genes. The Pearson correlation coefficient is 0.65 and the p-value is 3x10 -33 . Fig. 9C Figure 3 is a scatter plot of the average copy number of 107 genes shared between the 1001-gene measurement using the MHD2 code and the 140-gene measurement using the MHD4 code. The Pearson correlation coefficient is 0.89 and the p value is 9 × 10 -30 The dashed line corresponds to the y=x line.

[0288] FIG. 10 shows covariation analysis of RNA species measured in the 1001-gene measurement. Fig. 10A is a matrix of all pairwise correlation coefficients for cell-to-cell variation in expression of measured genes displayed as a hierarchical clustering tree. The ~100 identified groups of correlated genes are indicated by shading on the tree. Magnifications of 4 of the groups described in the text are shown on the right. Fig. 10B is the enrichment of 20 selected statistically significantly enriched GO terms in the four groups.

[0289] FIG. 19 shows the decoding and error estimation of the 1001-gene experiment. Fig.19A For each of the 14 hybridization rounds, Fig. 9A The final image is a composite image of these 14 rounds. The circles represent fluorescent spots that have been identified as potential RNA molecules. Some circles in the composite image represent unidentifiable molecules whose binary words do not match any of the 14-bit MHD2 code words. Fig. 9B show Fig. 9A The corresponding binary word for each spot identified in and the RNA species it was decoded into. "Unidentified" means that the measured binary word does not match any of the 1001 code words. Fig.19C Normalized confidence ratios measured for 985 RNA species (left) and 16 misidentified control words that were not targeted to any RNA (right) are shown. The normalized confidence ratio is Fig. 6F Defined in . Fig.19D Histogram showing the decrease in detected abundance of 107 genes present in the 1001-gene experiment and the 140-gene experiment. "Fold reduction in copy number" is defined as the mean number of RNA molecules per cell of each species measured in the 140-gene experiment divided by the corresponding mean number measured in the 1001-gene experiment.

[0290] Fig. 18C Comparison of the average RNA copy number per cell measured in a 1001-gene experiment using the MHD2 code with the average RNA copy number of 10 genes determined by conventional smFISH. The average ratio of the copy number measured using the MHD2 measurement method to the copy number measured using conventional smFISH is 0.30 + / - 0.05 (mean + / - SEM of 10 genes). The dashed line corresponds to the y = x line and the dotted line corresponds to the y = 0.30 x line.

[0291] Example 9

[0292] The above examples illustrate highly multiplexed detection schemes for system-level RNA imaging in single cells. By using combined labeling, sequential hybridization and imaging and two different error robustness encoding schemes, 140 or 1001 genes in hundreds of individual human fibroblasts can be imaged simultaneously. Of the two encoding schemes presented here, the MHD4 code is capable of both error detection and error correction, and thus can provide a higher call rate and a lower misidentification rate than the MHD2 code, which can only detect but not correct errors. On the other hand, MHD2 provides faster scaling of the degree of multiplexing using the number of bits than MHD4. Other error robustness encoding schemes can also be used for such multiplexed imaging, and the experimenter can set the balance between detection accuracy and ease of multiplexing based on the specific requirements of the experiment.

[0293] By increasing the number of bits in the codeword, it should be possible to use MERFISH using, for example, MHD4 or MHD2 codes to further increase the number of detectable RNA species. For example, using an MHD4 code with 32 total bits and 4 or 6 '1' bits can increase the number of addressable RNA species to 1,240 or 27,776, respectively. The latter is the approximate size of the human transcriptome. The predicted false positive and call rates are still reasonable for a 32-bit MHD4 code (for an MHD4 code with 4 '1' bits, e.g. Figure 5C and 5D , and similar ratios are calculated for the MHD4 code with 6 "1" bits). If more precise measurements are required, the additional increase in the number of bits will allow the use of coding schemes with Hamming distances greater than 4, further enhancing error detection and correction capabilities. While increasing the number of bits by adding more hybridization rounds will increase data collection time and may cause sample degradation, these problems can be mitigated by utilizing multiple colors to read out multiple bits in each hybridization round.

[0294] As multiplexing increases, it is important to consider the potential increase in the density of RNA that needs to be resolved in each round of imaging. Based on the imaging and sequencing results, it can be estimated that including the entire transcriptome of IMR90 cells will result in ∼200 molecules / μm 3 Using current imaging and analysis methods, 2-3 molecules / μm can be resolved per hybridization round. 3 , which will reach ∼20 molecules / μm after 32 rounds of hybridization 3 This density should allow all but the top 10% most expressed genes to be imaged simultaneously, or include subsets of genes with even higher expression levels. By utilizing more advanced image analysis algorithms to better resolve overlapping images of individual molecules, such as compressed sensing, it may be possible to expand the resolvable density by ~4-fold, allowing all but the top 2% most expressed genes to be imaged together.

[0295] These examples have illustrated the power of data derived from highly multiplexed RNA imaging by using covariation and correlation analysis to reveal distinct subcellular distribution patterns of RNA, constrain gene regulatory networks, and predict the functions of many previously unannotated or partially annotated genes of unknown function. Given its ability to quantify RNA over a wide range of abundances without amplification bias while preserving the native context, systems and methods such as MERFISH will allow many applications for in situ transcriptome analysis of individual cells in culture or complex tissues.

[0296] Example 10

[0297] The following are various materials and methods used in the above examples.

[0298] Probe Design. Each RNA species in the target set was randomly assigned a binary codeword from all 140 possible codewords of the 16-bit MHD4 code or all 1001 possible codewords of the 14-bit MHD2 code.

[0299] The array-synthesized oligonucleotide pool is used as a template to prepare the encoding probes. The template molecule of each encoding probe contains three components: i) a central targeting sequence for in situ hybridization with the target RNA, ii) two flanking readout sequences designed to hybridize with each of the two different readout probes, and iii) two flanking primer sequences that allow enzymatic amplification of the probe ( Fig.13 ). The readout sequence is taken from 16 possible readout sequences, each corresponding to one hybridization round. The readout sequences are assigned to the coding probes so that for any RNA species, each of the 4 readout sequences is evenly distributed along the length of the target RNA and appears at the same frequency. The template molecule of the 140-gene library also includes a common 20 nucleotide (nt) priming region between the first PCR primer and the first readout sequence. The priming sequence is used in the reverse transcription step described below.

[0300] Multiple experiments are embedded in the oligonucleotide library synthesized by single array, and PCR is used to selectively amplify the oligonucleotides required for specific experiments only. The primer sequences for the index PCR reaction are generated from a set of orthogonal 25-nt sequences. These sequences are trimmed to 20nt and selected for the following: i) narrow melting temperature range (70°C to 80°C), ii) there is no continuous repetition of 3 or more identical nucleotides, and iii) there is a GC clamp, i.e. one of the two 3' terminal bases must be G or C. In order to further improve specificity, BLAST+ is used to screen these sequences for the human transcriptome, and primers with 14 or more continuous base homologies are eliminated. Finally, BLAST+ is used again to identify and exclude primers with 11-nt homology regions at the 3' end of any other primer or with 5-nt homology regions at the 3' end of the T7 promoter. The forward primer sequence (primer 1) is measured as described above, and the reverse primer each contains the 20-nt sequence as described above plus the 20-nt T7 promoter sequence to promote the amplification (primer 2) by in vitro transcription. The primer sequences used in the 140-gene and 1001-gene experiments are listed below.

[0301] Table 2

[0302]

[0303] 30-nt long read sequences were generated by concatenating fragments of the same orthogonal primer set generated above (by combining a 20-nt primer with a 10-nt fragment of another). These read sequences were then screened for orthogonality to the index primer sequence and other read sequences (no more than 11 nt homology) and potential off-target binding sites in the human genome (no more than 14 nt homology) using BLAST+. These read sequences were probed using fluorescently labeled read probes with sequences complementary to the read sequences, with one read sequence being probed in each hybridization round. All read probe sequences used are listed below.

[0304] Table 3

[0305]

[0306]

[0307] The readout probes for the 140-gene library were probes 1 to 16. The readout probes for the 1001-gene experiment were probes 1 to 14. / 3Cy5Sp / indicates 3'Cy5 modification.

[0308] To design the central targeting sequence of the coding probe, the abundance of different transcripts in IMR90 cells using Cufflinks v2.1, total RNA data from the ENCODE project, and human genome annotations from Gencode v18 were followed. Probes were designed from gene models corresponding to the most abundant isoforms using OligoArray2.1 with the following constraints: the length of the target sequence region was 30-nt; the melting temperature of the hybridization region of the probe and the cellular RNA target was greater than 70°C; there was no cross-hybridization target at a melting temperature greater than 72°C; there was no predicted internal secondary structure at a melting temperature greater than 76°C; and there were no consecutive repeats of 6 or more identical nucleotides. The melting temperature was adjusted to optimize the specificity of these probes and minimize the secondary structure, while still generating a sufficient number of probes for the library. In order to reduce computational costs, isoforms were divided into 1-kb regions for probe design. All potential probes mapped to more than one cellular RNA species were rejected by using BLAST+. Probes with multiple targets on the same RNA were retained.

[0309] For each gene in the 140-gene experiment, 198 putative coding probe sequences were generated by concatenating the appropriate index primer, read sequence, and target region, as Fig.13As shown. To address the possibility that concatenation of these sequences introduced new regions with homology to off-target RNAs, these putative sequences were screened against all human rRNA and tRNA sequences and highly expressed genes (genes with FPKM>10,000) using BLAST+. Probes with greater than 14 nt homology to rRNA or tRNA or greater than 17 nt homology to highly expressed genes were removed. After these excisions, there were ∼192 (standard deviation 2) probes per gene for the two MHD4 codebooks used in the 140-gene experiment. The same protocol used for the 1001-gene experiment was used as follows: starting with 96 putative targeting sequences per gene, after these additional homology excisions, ∼94 (standard deviation 6) encoding probes per gene were obtained. For the 1001-gene experiment, the number of encoding probes for each RNA was reduced so that these probes could be synthesized from a single 100,000-member oligonucleotide pool rather than two separate pools. Each coding probe is designed to contain two of the four readout sequences associated with each code word, so in any given hybridization round, only half of the bound coding probes can bind to the readout probe. Use ~192 or ~94 coding probes per RNA to obtain a high signal-to-background ratio for individual RNA molecules. The number of coding probes per RNA can be significantly reduced, but still allows identification of single RNA molecules. In addition, increasing the number of readout sequences per coding probe or using optical sectioning methods to reduce fluorescent background can allow further reduction in the number of coding probes per RNA.

[0310] Two types of misidentification controls were designed. The first control (blank word) was not represented by a coding probe. The second type of control (no target word) had a coding probe that did not target any cellular RNA. The target region of these probes included a random nucleotide sequence that was subjected to the same constraint for designing the above-mentioned RNA targeting sequence. In addition, these random sequences were screened for the human transcriptome to ensure that they did not have significant homology (>14-nt) to any human RNA. The 140-gene measurement included 5 blank words and 5 no target words. The 1001-gene measurement included 11 blank words and 5 no target words.

[0311] Probe Synthesis. Encoded probes were synthesized using the following steps and Fig.13 The synthesis scheme is illustrated in FIG.

[0312] Step 1: A template oligonucleotide pool (CustomArray) was amplified by limited cycle PCR on a Bio-Rad CFX96 using primer sequences specific to the desired probe set. To facilitate subsequent amplification by in vitro transcription, the reverse primer contained a T7 promoter. All primers were synthesized by IDT. The reaction was column purified (Zymo DNA Clean and Concentrator; D4003).

[0313] Step 2: According to the manufacturer's instructions (New England Biolabs, E2040S), the purified PCR product is further amplified ~ 200 times and converted into RNA by high-yield in vitro transcription. Each 20 microliters of reactants contain ~ 1 microgram from the above-mentioned template DNA, each NTP of 10mM, 1x reaction buffer, 1x RNAse inhibitor (PromegaRNasin, N2611) and 2 microliters of T7 polymerase. This reactant is incubated at 37 ℃ for 4 hours to maximize productivity. This reactant is not purified before the following steps.

[0314] Step 3: The RNA product from the above in vitro transcription reaction was then converted back to DNA by reverse transcription. Each 50 microliter reaction contained the unpurified RNA product from step 2, which was supplemented with 1.6 mM of each dNTP, 2 nmol of reverse transcription primer, 300 units of Maxima H reverse transcriptase (Thermo Scientific, EP0751), 60 units of RNasin, and a final 1x concentration of Maxima RT buffer. The reaction was incubated at 50°C for 45 minutes, and the reverse transcriptase was inactivated at 85°C for 5 minutes. The template of the 140-gene library contained a common priming region for this reverse transcription step; therefore, when generating these probes, a single primer was used for this step. Its sequence is CGGGTTTAGCGCCGGAAATG (SEQ ID NO: 40). For the 1001-gene library, a common priming region was not included; therefore, reverse transcription was performed with a forward primer: CGCGGGCTATATGCGAACCG (SEQ ID NO: 20).

[0315] Step 4: In order to remove the template RNA, 20 microliters of 0.25M EDTA and 0.5NNaOH are added to the above reaction to selectively hydrolyze RNA, and the sample is incubated at 95°C for 10 minutes. Immediately thereafter, a 100 microgram capacity column (ZymoResearch, D4030) and Zymo Oligo Clean and Concentrator scheme are used to purify the reactant by column purification. The final probe is eluted in 100 microliters of deionized water without RNase, evaporated in a vacuum concentrator, and then resuspended in 10 μL encoding hybridization buffer (see below). The probe is stored at -20°C. Denaturing polyacrylamide gel electrophoresis and absorption spectra are used to confirm the quality of the probe, and reveal that the probe synthesis scheme converts 90-100% of the reverse transcription primer into a full-length probe, and in the probe constructed, 70-80% is recovered during the purification step.

[0316] Fluorescently labeled readout probes have a sequence complementary to the readout sequences described above and a Cy5 dye attached to the 3' end. These probes were obtained from IDT and purified by HPLC.

[0317] Sample preparation and labeling with coded probes. Human primary fibroblasts (American Type Culture Collection, IMR90) were used for this work. These cells are relatively large and flat, facilitating wide-field imaging without the need for optical sectioning. Cells were cultured with Eagle's Minimum Essential Medium. Cells were seeded at 350,000 cells / coverslip on 22-mm, #1.5 coverslips (Bioptechs, 0420-0323-2) and incubated at 37°C with 5% CO 2 Incubate in a Petri dish for 48-96 hours. The cells were fixed in 4% paraformaldehyde (Electron Microscopy Sciences, 15714) in 1x phosphate buffered saline (PBS; Ambion, AM9625) for 20 minutes at room temperature, reduced with 0.1% w / v sodium borohydride (Sigma, 480886) aqueous solution for 5 minutes to reduce background fluorescence, washed three times with ice-cold 1x PBS, permeabilized with 0.5% v / v Triton (Sigma, T8787) in 1x PBS for 2 minutes at room temperature, and washed three times with ice-cold 1x PBS.

[0318] Cells are incubated in the coding wash buffer containing 2x saline-sodium citrate buffer (SSC) (Ambion, AM9763), 30% v / v formamide (Ambion, AM9342) and 2mM vanadyl ribonucleoside complex (NEB, S1402S) for 5 minutes. 100 micromoles (140-gene experiment) or 200 micromoles (1001-gene experiment) coding probes in 10 microliters of coding hybridization buffer are added to the coverslip containing cells, and it is evenly spread by placing another coverslip on the top of the sample. Samples are then incubated in a wet chamber in 37°C hybridization ovens for 18-36 hours. Coding hybridization buffer includes the coding wash buffer supplemented with 1mg / mL yeast tRNA (Life Technologies, 15401-011) and 10% w / v dextran sulfate (Sigma, D8906-50G).

[0319] The cells were then washed with primary coding wash buffer, incubated at 47°C for 10 minutes, and the wash was repeated three times. A 1:1000 dilution of 0.2 micron diameter carboxylate-modified orange fluorescent beads (Life Technologies, F-8809) in 2xSSC was sonicated for 3 minutes, followed by incubation with the sample for 5 minutes. The beads were used as reference markers to align images obtained from multiple consecutive rounds of hybridization, as described below. The sample was washed once with 2xSSC, followed by post-fixation with 4% v / v paraformaldehyde in 2×SSC for 30 minutes at room temperature. The sample was then washed three times with 2xSSC, and imaged immediately or stored at 4°C for no longer than 12 hours before imaging. All solutions were prepared to be RNase-free.

[0320] MERFISH imaging using multiple continuous rounds of hybridization. The sample coverslip is assembled into the FCS2 flow chamber of Bioptech, and the flow through the chamber is controlled by a homemade fluid system consisting of three computer-controlled 8-way valves (Hamilton, MVP and HVXM 8-5) and a computer-controlled peristaltic pump (Rainin, Dynamax RP-1). The sample is imaged on a homemade microscope constructed around an Olympus IX-71 body and a 1.45NA, 100x oil-immersion objective lens and configured for oblique incident excitation. The objective lens is heated to 37 ° C using Bioptechs objective lens heater. A homemade autofocus system is used to maintain constant focus throughout the imaging process. Solid-state lasers (MPB Communications, VFL-P500-642; Coherent, 561-200CWCDRH; and Coherent, 1069413 / AT) were used to provide illumination at 641 nm, 561 nm, and 405 nm for excitation of the Cy5-labeled readout probe, reference beads, and nuclear counterstain, respectively. These lines were combined with a custom dichroic (Chroma, zy405 / 488 / 561 / 647 / 752RP-UF1) and the emission was filtered with a custom dichroic (Chroma, ZET405 / 488 / 561 / 647-656 / 752m). Use dichroic T560lpxr, T650lpxr, 750dcxxr (Chroma) and emission filter ET525 / 50m, WT59550m-2f, ET700 / 75m, HQ770lp (Chroma), utilize QuadView (Photometrics) to separate fluorescence, and image with EMCCD camera (Andor, iXon-897).Configure the camera so that the pixel corresponds to 167nm in the sample plane.The whole system is fully automated, so that the whole experiment can be imaged and fluid handled without user intervention.

[0321] Sequential hybridization, imaging and bleaching are performed as follows. 10nM of the appropriate fluorescently labeled read probe in 1mL of read hybridization buffer (2xSSC; 10% v / v formamide; 10% w / v dextran sulfate and 2mM vanadyl ribonucleoside complex) is flowed through the sample, the flow is stopped, and the sample is incubated for 15 minutes. Subsequently, 2mL of read wash buffer (2xSSC, 20% v / v formamide; and 2mM vanadyl ribonucleoside complex) is flowed through the sample, the flow is stopped, and the sample is incubated for 3 minutes. 2mL of imaging buffer containing 2xSSC, 50mM TrisHCl pH 8, 10% w / v glucose, 2mM Trolox (Sigma-Aldrich, 238813), 0.5mg / mL glucose oxidase (Sigma-Aldrich, G2133) and 40 micrograms / mL catalase (Sigma-Aldrich, C30) is flowed through the sample. The flow was then stopped and approximately 75 to 100 regions were subsequently exposed to 25 mW 642-nm and 1 mW 561-nm light and imaged. Each region was 40 microns by 40 microns. The laser power was measured at the rear port of the microscope. Because the imaging buffer was sensitive to oxygen, 50 mL of imaging buffer for a single experiment was freshly prepared at the beginning of the experiment and then stored under a layer of mineral oil throughout the measurement. The buffer stored in this manner was stable for more than 24 hours.

[0322] After imaging, the fluorescence of the readout probe was eliminated by photobleaching. The sample was washed with 2 mL of photobleaching buffer (2xSSC and 2 mM vanadyl ribonucleoside complex) and each imaging area of ​​the sample was exposed to 200 mW of 641-nm light for 3 seconds. To confirm the efficacy of this photobleaching treatment, the imaging buffer was introduced again and the sample was imaged as described above.

[0323] The above hybridization, imaging, and photobleaching process was repeated 16 times for a 140-gene measurement using the MHD4 code, or 14 times for a 1001-gene measurement using the MHD2 code. The entire experiment was typically completed in ~20 hours.

[0324] After imaging was completed, 2 mL of a 1:1000 dilution of Hoescht (ENZ-52401) in 2×SSC was flowed through the chamber to label the nuclei of the cells. The sample was then immediately washed with 2 mL of 2×SSC followed by 2 mL of imaging buffer. Each area of ​​the sample was then imaged again with ~1 mW of 405 nm light.

[0325] Because cells are imaged using widefield imaging utilizing oblique incident illumination without optical sectioning and z-scanning, conventional smFISH is used to quantify the fraction of individual RNA species outside the axial range of the geometric structure of imaging for 6 different RNA species. For this purpose, by collecting the stack of images at different focal depths through the entire depth of the cell, these cells are optically sectioned. The images are aligned in continuous focal planes, and the fraction of RNA detected in the three-dimensional stacking but not detected in the basic focal plane is calculated for each cell subsequently. It is found that only a small part, 15%+ / -1% (mean value + / -SEM across 6 different RNA species) of RNA molecules are outside the imaging range of the fixed focal plane without z-scanning. These measurements also confirm that the excitation geometry irradiates the complete depth of the cell. Any optical sectioning technique can be adopted in MERFISH to allow RNA to be imaged in thicker cells or tissues.

[0326] Construction of measured words. Use a multi-Gaussian fitting algorithm (presuming a Gaussian with a uniform width of 167nm) to identify and locate fluorescent spots in each image. This algorithm is used to allow the distinction and fitting of partially overlapping spots. By setting the intensity threshold required for fitting spots using the software, RNA spots are distinguished from background signals (i.e., signals produced by probes of non-specific binding). Due to the change in spot brightness between hybridization rounds, the threshold is appropriately adjusted for each hybridization round to minimize the combined mean value of 1-->0 and 0-->1 error rates in all hybridization rounds (140-gene measurement), or to maximize the ratio of the number of measured words with 4 "1" positions to the number of measured words with 3 or 5 "1" positions (1001-gene measurement). Use a faster single Gaussian fitting algorithm to identify the position of the reference beads in each frame.

[0327] Images of the same sample area in different rounds of hybridization were recorded by rotating and translating the image to align the two reference beads within the same image that were most similar in position after a rough initial alignment by image correlation. All images were aligned to the coordinate system established by the images collected in the first round of hybridization. The quality of this alignment was determined from the residual distances between 5 additional reference beads, with alignment errors typically being ~20 nm.

[0328] If the distance between spots is less than 1 pixel (167nm), the fluorescent spots in different hybridization rounds are connected into a single string corresponding to the potential RNA molecule. For each spot string, the on-off sequence of the fluorescent signal in all hybridization rounds is used to assign a binary word to the potential RNA molecule, where '1' is assigned to the hybridization round containing a fluorescent signal above the threshold and '0' is assigned to the other hybridization rounds. The measured words are then decoded into RNA species using the above-mentioned 16-bit MHD4 code or 14-bit MHD2 code. In the case of the 16-bit MHD4 code, if the measured binary word completely matches the code word of a specific RNA or differs from the code word by a single bit, it is assigned to the RNA. In the case of the 14-bit MHD2 code, it is assigned to the RNA only if the measured binary word completely matches the code word of a specific RNA. To determine the copy number per cell, the number of each RNA species is counted in individual cells within each 40 micron × 40 micron imaging area. Note that this number accounts for most but not all RNA molecules within the cell, because part of the cell can be outside the imaging area or depth of focus. Tiled images of adjacent areas and adjacent focal planes can be used to improve counting accuracy.

[0329] In the 140-gene experiment, some areas of the nucleus occasionally contain too much fluorescent signal to correctly identify individual RNA spots. In the 1001-gene experiment, the nucleus usually contains too much fluorescent signal to allow identification of individual RNA molecules. These bright areas are excluded from all subsequent analyses. This work focuses on the mRNA enriched in the cytoplasm. In order to estimate the fraction of mRNA lost by excluding the nuclear region, conventional smFISH is used to quantitatively find the molecular fractions of 6 different mRNA species found inside the nucleus. It is found that only 5%+ / -2% (mean value + / -SEM across 6 RNA species) of these RNA molecules are found in the nucleus. The use of super-resolution imaging and / or optical sectioning may allow identification of individual molecules in these dense nuclear regions, which is particularly useful for detecting those non-coding RNAs enriched in the nucleus.

[0330] smFISH measurement of individual genes.The library of 48 fluorescently labeled (Quasar670) oligonucleotide probes of each RNA is purchased from Biosearch Technologies.The probe sequence of 30-nt is directly obtained from the random subset of the target area for multiplexing measurement.Fix and permeabilize cells as described above.The 250nM oligonucleotide probe in 10 microliters of encoding hybridization buffer (as described above) is added to the cover glass containing cells, and is evenly spread by placing another cover glass on the top of the sample.The sample is then incubated in a wet chamber in 37°C hybridization ovens for 18 hours.Then the cells are washed for 10 minutes at 37°C with encoding wash buffer (as described above), and the washing is repeated three times in total.The sample is then washed three times with 2xSSC and imaged in imaging buffer using the same imaging geometry as described above for MERFISH.

[0331] Bulk RNA sequencing. Total RNA was extracted from IMR90 cells cultured as above using the Zymo Quick RNA MiniPrep Kit (R1054) according to the manufacturer's instructions. PolyA RNA was then selected (NEB; E7490), and sequencing libraries were constructed using the NEBNext Ultra RNA Library Preparation Kit (NEB; E7530), amplified with custom oligonucleotides, and 150-bp reads were obtained from MiSeq. These sequences were aligned to the human genome (Gencode v18), and isoform abundance was calculated using cufflinks.

[0332] Computation of prediction scaling and error properties for different coding schemes. Analytical expressions are derived for the dependence of the number of possible codewords, the call rate, and the misidentification rate on N. The call rate is defined as the fraction of RNA molecules that are correctly identified. The misidentification rate is defined as the fraction of RNA molecules that are misidentified as the wrong RNA species. For coding schemes with error detection capability, the call rate and the misidentification rate do not add up to 1 because a fraction of molecules that are not called correctly can be detected as errors and discarded, and therefore are not misidentified as the wrong species. These calculations assume that the probability of a misread bit is constant for all hybridization rounds, but is different for 1-->0 and 0-->1 errors. The experimentally measured average 1-->0 and 0-->1 error rates (10% and 4%, respectively) are used for Figures 5B-5D The estimates shown. For simplicity, words corresponding to all '0's are not removed from the calculation.

[0333] For a simple binary encoding scheme in which all possible N-bit binary words are assigned to unique RNA species, the number of possible code words is 2 N The number of words available to encode RNA is actually 2N = -1, since the code word '00...0' contains no detectable fluorescence in any hybridization round, but for simplicity, the words corresponding to all '0's are not removed from subsequent calculations. The error introduced by this approximation is negligible. For any given word with m '1's and Nm '0's, the probability of measuring the word without error (the fraction of RNAs called correctly) is:

[0334] (1-p 1 ) m (1-p 0 ) N-m ,(1)

[0335] where p 1 is the error rate of 1-->0 per bit, p 0 is the error rate for each bit 0->1. Because different words in this simple binary encoding scheme can have different numbers of "1" bits, if p 1 ≠p 0 , then the call rates of different words will be different. The average call rate is determined by the weighted average of the values ​​of equation (1) for all words (e.g. Figure 5C The weighted average is:

[0336]

[0337] in is the binomial coefficient and corresponds to the number of words with m '1' bits in this coding scheme. Since in this coding scheme, each error produces a binary word encoding a different RNA, Figure 5D The average misidentification rate for this coding scheme reported in comes directly from (2):

[0338]

[0339] To calculate the scaling and error properties of an extended Hamming distance 4 (HD4) code, a generator matrix for the expected number of data bits using standard methods is first created. The generator matrix determines the specific words present in a given coding scheme and is used to directly determine the number of encoded words as a function of the number of bits. In this coding scheme, the call rate corresponds to the fraction of measured words with no errors and the fraction of measured words with a single bit error. For a code word with m '1' bits, this fraction is determined by the following expression:

[0340] (1-p 1 ) m (1-p 0 ) N-m +mp 1 1 (1-p 1 )m-1 (1-p 0 ) N-m +(Nm)p 0 1 (1-p 1 ) m (1-p 0 ) N-m-1 (4)

[0341] The first term is the probability of not making any error, the second term corresponds to the total probability of making a 1->0 error at any of the m '1' bits without causing any other 0->1 errors, and the last term corresponds to the total probability of making a 0->1 error at any of the Nm '0' bits without causing any 1->0 errors. Because the number of '1' bits can vary between words in this coding scheme, Figure 5C The average call rates reported in are calculated from the weighted average of equation (4) for different values ​​of m. The weight of each term is determined from the number of words containing m '1' bits determined from the generator matrix above.

[0342] Because RNA code words are separated by a minimum Hamming distance of 4, at least 4 errors are required to convert one word into another. If error correction is applied, 3 or 5 errors can also convert one RNA into another. Therefore, the misidentification rate from all possible combinations of 3, 4, and 5 errors is estimated for a code word with m '1' bits. Technically, >5 errors can also convert one RNA into another, but since the error rate per bit is small, the probability of such errors is negligible. The expression is approximately:

[0343]

[0344] The first sum corresponds to all ways in which exactly 4 errors can be produced. Similarly, the second and third sums correspond to all ways in which exactly 3 or 5 errors can be produced. Equation (5) provides an upper limit on the misidentification rate because not all 3, 4, or 5 bit errors produce a word that matches or will be corrected to another correct word. Again, because the number of '1' bits can vary between words, Figure 5D The average misidentification rate reported in is calculated as the weighted average of equation (5) for the number of words with m '1' bits.

[0345] In order to generate an MHD4 code in which the number of "1" bits of each code word is set to 4, an HD4 code is first generated as described above, and then all code words not containing 4 "1"s are removed. Figure 5CThe recall rate for this code reported in is calculated directly from equation (4), but with m=4 since all codewords in this code have 4 '1' bits. It is calculated by modifying equation (5) with the following considerations Figure 5D The false positive rate of the code reported in : (i) the number of '1' bits m is set to 4, and (ii) errors that produce words without 4, 3 or 5 '1' bits are excluded. Therefore, equation (5) simplifies to

[0346]

[0347] Again, this expression is an upper bound on the actual misidentification rate, since not all words with 4 "1's" are valid code words.

[0348] Estimation of the error rate of 1->0 and 0->1 for each hybridization round. To calculate the probability of misreading a bit in a given hybridization round, the error correction property of the MHD4 code is used. In brief, the probability of a 1->0 or 0->1 error is derived in the following way. Assume that the probability of an error in the i-th bit (i.e., the i-th hybridization round) is p i , and the actual number of RNA molecules of a given species is A, then the number of exact matches for that RNA will be And the number of 1-bit error-corrected matches for this RNA corresponding to an error at the ith position will be

[0349] p i It can be directly derived from the ratio: This ratio assumes that the counts for 1-bit error correction arise only from single-bit errors of the correct word, and that multi-error contamination from other RNA words is negligible. Given that the error rate per hybridization round is small, and that at least 3 errors are required to convert an RNA codeword into a word that can be misidentified as another RNA, the above approximation should be a good approximation.

[0350] To calculate the average 1-->0 or 0-->1 error probability for each of the 16 hybridization rounds, the per-bit error rate for each bit of each gene was calculated using the method described above, and the errors were ranked based on whether they corresponded to 1-->0 or 0-->1 errors, and the average of these errors for each bit weighted by the number of counts observed for the corresponding gene was taken.

[0351] The call rates of individual RNA species were estimated from the actual imaging data. By estimating the 1-->0 or 0-->1 error probability for each round of hybridization determined as above, it is possible to estimate the call rate of each RNA based on the specific word used to encode it. Specifically, the fraction of RNA species called correctly is determined by

[0352]

[0353] where the first term represents the probability of observing an exact match of the codeword and the second term represents the probability of observing an error-corrected match (i.e., with a 1-bit error). The bit error rate p for each RNA species is i The value of p is determined by the specific codeword for the RNA and the measured 1->0 or 0->1 error rate for each hybridization round. If the codeword for an RNA contains a '1' at the i-th position, then p is determined from the 1->0 error rate for the i-th hybridization round. i ; If the word contains a '0' at position i, then p is determined from the 0-->1 error rate of the i-th hybridization round i .

[0354] The hierarchical cluster analysis of covariation of RNA abundance.The hierarchical cluster of covariation of gene expression of 140-gene and 1001-gene experiment is carried out as follows.First, the distance between every pair of genes is determined as 1 minus the Pearson correlation coefficient of the cell-to-cell variation of the copy number (both of which are standardized by the total RNA counted in the cell) of the measurement of these two kinds of RNA species.Therefore, highly correlated genes are "closer" to each other, and highly anti-correlated genes are "further" separated.Use unweighted paired group arithmetic mean method (UPGMA) to build a cohesive hierarchical cluster tree from these distances subsequently.Specifically, starting from individual genes, hierarchical clustering is built by identifying the two closest clusters (or individual genes) to each other according to the arithmetic mean of the distance between all inter-cluster gene pairs.The cluster (or individual gene) with the minimum distance is then grouped together, and the process is repeated.The matrix of two-to-two correlations is subsequently sorted based on the order of the genes in these trees.

[0355] By selecting the threshold on the hierarchical clustering tree (given by Fig.7D and 10A The hierarchical clustering tree generates about 10 groups of genes (where each group contains at least 4 members) for the 140-gene experiment, or about 100 groups of genes (where each group contains at least 3 members) for the 1001-gene experiment. Note that the threshold can be changed to identify smaller groups that are more tightly coupled or larger groups with relatively loose coupling.

[0356] The probability value of the confidence that a gene belongs to a specific group is determined by calculating the difference between the average correlation coefficient between the gene and all other members of the group and the average correlation coefficient between the gene and all other measured genes outside the group. The significance (p-value) of the difference is determined using the Student's t-test.

[0357] Because hierarchical clustering is inherently a one-dimensional analysis, i.e., any given gene can only be a member of a single group, this analysis does not allow identification of all related gene groups. Higher-dimensional analyses, such as principal component analysis or k-means clustering, can be used to identify more clusters of co-varying genes.

[0358] Analysis of RNA spatial distribution.In order to identify genes with similar spatial distribution, each measured cell is subdivided into 2 × 2 regions, and the score of each RNA species present in each of these boxes is calculated.In order to control the fact that some regions of the cell naturally contain more RNA than other regions, the enrichment of each gene is calculated, i.e., the ratio of the score observed in a given region for a given RNA species to the average score observed for all genes in the same region.For each pair of RNA species, the Pearson correlation coefficient of the regional difference of the enrichment of these two RNA species of each cell is measured, and the mean value of the correlation coefficient is taken in the cells exceeding~400 imaging in 7 independent data sets.The above-mentioned identical hierarchical clustering algorithm is used subsequently to cluster RNA species based on these average correlation coefficients.Due to a large number of cells for analysis, it is found that coarse spatial binning (2 × 2 regions for each cell) is sufficient to capture the spatial correlation between genes, and more detailed binning does not produce more significantly related groups.

[0359] In order to measure the distance between the gene and the nucleus and the cell edge, the brightness threshold on the cell image is first used to segment the identified nucleus and the cell edge. The distance from the nearest part of each RNA molecule to the nucleus and the nearest part of the cell edge is then measured. For each data set, the average distance of each RNA species averaged for all cells measured is then calculated. The mean value of these distances is taken for group I genes, group II genes or all genes. In this analysis, only those RNA species with at least 10 counts per cell are used to minimize the statistical error of the distance value.

[0360] Gene Ontology (GO) analysis.As mentioned above, select genome from hierarchical tree.Use the GO term of annotation and the term that is immediately upstream or downstream of the annotation found, measure the set of GO terms of the RNA species of all measurements and the RNA species associated with each group from the nearest people's GO annotation.The enrichment of these annotations is calculated from the ratio of the score of the gene in each group with this term and the score of the gene of all measurements with this term, and the p-value of this enrichment is calculated by hypergeometric function.Only consider the GO term of statistically significant enrichment less than 0.05 of p value.

[0361] Although several embodiments of the present invention have been described and illustrated herein, a person of ordinary skill in the art will easily envision various other methods and / or structures for performing the functions and / or obtaining the results and / or one or more advantages described herein, and each of such changes and / or modifications is considered to be within the scope of the embodiments of the present invention described herein. More generally, it will be readily understood by those skilled in the art that all parameters, dimensions, materials and configurations described herein are intended to be exemplary and that actual parameters, dimensions, materials and / or configurations will depend on the specific application for which the teachings of the present invention are used. Those skilled in the art will recognize or be able to determine many equivalents of the specific embodiments of the present invention described herein (by using no more than routine experiments). Therefore, it should be understood that the aforementioned embodiments are presented only by way of example, and within the scope of the appended claims and their equivalents, embodiments of the present invention may be implemented in a manner different from that explicitly described and claimed. The present invention relates to various individual characteristics, systems, articles, materials, kits and / or methods described herein. In addition, if such characteristics, systems, articles, materials, kits and / or methods are not mutually contradictory, any combination of two or more such characteristics, systems, articles, materials, kits and / or methods is included within the scope of the invention of the present disclosure.

[0362] As defined and used herein, all definitions should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0363] The indefinite articles "a" and "an," as used herein in the specification and claims, should be understood to mean "at least one," unless explicitly indicated to the contrary.

[0364] The phrase "and / or", as used herein in the specification and claims, should be understood to mean "either or both" of the connected elements, i.e., the elements are present in combination in some cases and separately in other cases. Multiple elements listed with "and / or" should be interpreted in the same manner, i.e., "one or more" of the connected elements. In addition to the elements explicitly identified by the "and / or" clause, other elements may optionally be present, whether related or unrelated to those elements explicitly identified. Thus, as a non-limiting example, a reference to "A and / or B", when used in conjunction with an open-ended phrase such as "comprising", may refer to A only (optionally including elements other than B) in one embodiment; may refer to B only (optionally including elements other than A) in another embodiment; may refer to A and B (optionally including other elements) in another embodiment; and so on.

[0365] As used herein in the specification and claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" should be understood to be inclusive, i.e., including at least one, but also including more than one of a plurality of elements or a list of elements, and optionally, including additional unlisted items. Only terms that clearly indicate the opposite, such as "only one of..." or "exactly one of..." or "consisting of..." when used in the claims will refer to the inclusion of exactly one element of a plurality of elements or a list of elements. Generally, the term "or" as used herein should be interpreted as indicating an exclusive selection (i.e., "one or the other but not two") only when it is crowned with an exclusive term such as "either", "one of...", "only one of..." or "exactly one of...". "Substantially consisting of...", when used in the claims, should have its ordinary meaning used in the field of patent law.

[0366] As used herein in the specification and claims, the phrase "at least one" with respect to a list of one or more elements should be understood to mean at least one element selected from any one or more elements in the list of elements, but not necessarily including at least one of each element explicitly listed in the list of elements and not excluding any combination of elements in the list of elements. This definition also allows for the optional presence of elements other than the elements explicitly identified in the list of elements to which the phrase "at least one" refers, whether related or unrelated to those explicitly identified elements. Thus, as a non-limiting example, "at least one of A and B" (or, equivalently, "at least one of A or B" or, equivalently, "at least one of A and / or B") may refer in one embodiment to at least one (optionally including more than one) A without B present (and optionally including elements other than B); in another embodiment to at least one (optionally including more than one) B without A present (and optionally including elements other than A); in another embodiment to at least one (optionally including more than one) A and at least one (optionally including more than one) B (and optionally including other elements); and so on.

[0367] It should also be understood that, unless explicitly stated to the contrary, in any method claimed herein that includes more than one step or act, the order of the steps or acts of the method are not necessarily limited to the order in which the steps or acts of the method are recited.

[0368] In the claims, as well as in the foregoing specification, all transitional phrases such as "comprising," "including," "carrying," "having," "containing," "possessing," "involving," "holding," "consisting of," and the like are to be construed as open-ended, i.e., meaning including, but not limited to, "consisting of." Only the transitional phrases "consisting of" and "consisting essentially of" shall be closed or semi-closed transitional phrases, respectively, as set forth in Section 2111.03 of the U.S. Patent Office Manual of Patent Examining Procedures.

[0369] This application involves the following technical solutions:

[0370] 1. A method comprising:

[0371] exposing the sample to a plurality of nucleic acid probes;

[0372] For each of the nucleic acid probes, determining binding of the nucleic acid probe within the sample;

[0373] generating a codeword based on the binding of the nucleic acid probes; and

[0374] For at least some of the codewords, the codewords are matched to valid codewords, wherein if no match is found, error correction is applied to the codewords to form valid codewords.

[0375] 2. The method of technical solution 1 comprises exposing the sample to at least 5 different nucleic acid probes.

[0376] 3. The method of any one of technical solutions 1 or 2, which comprises exposing the sample to at least 10 different nucleic acid probes.

[0377] 4. The method of any one of technical solutions 1-3, which comprises exposing the sample to at least 100 different nucleic acid probes.

[0378] 5. The method of any one of technical solutions 1-4, which comprises exposing the sample to multiple nucleic acid probes simultaneously.

[0379] 6. The method according to any one of technical solutions 1 to 5, comprising sequentially exposing the sample to a plurality of nucleic acid probes.

[0380] 7. The method of technical solutions 1-6, wherein the multiple nucleic acid probes comprise a combinatorial combination of nucleic acid probes with different sequences.

[0381] 8. The method of technical solution 7, wherein the combinatorial combination of nucleic acid probes targets the combinatorial combination of RNA species in the sample.

[0382] 9. The method of technical solution 7 or 8, wherein the combinatorial combination of nucleic acid probes targets the combinatorial combination of DNA sequences in the sample.

[0383] 10. The method of any one of technical solutions 1 to 9, wherein at least some of the plurality of nucleic acid probes comprise a first portion comprising a target sequence and a second portion comprising one or more reading sequences.

[0384] 11. The method of technical solution 10, wherein the plurality of nucleic acid probes comprise distinguishable nucleic acid probes formed by a combinatorial combination of one or more reading sequences taken from the one or more reading sequences.

[0385] 12. The method of any one of technical solutions 10 or 11, wherein the target sequence is substantially complementary to a nucleic acid sequence encoding a protein.

[0386] 13. The method of any one of technical solutions 10-12, wherein the target sequence is substantially complementary to messenger RNA (mRNA).

[0387] 14. The method of any one of technical solutions 10-13, wherein the multiple nucleic acid probes comprise at least 8 possible reading sequences.

[0388] 15. The method of any one of technical solutions 10-14, wherein the multiple nucleic acid probes contain no more than 32 possible reading sequences.

[0389] 16. The method of any one of technical solutions 10-15, wherein the multiple nucleic acid probes contain no more than 16 possible reading sequences.

[0390] 17. The method of any one of technical solutions 10-16, wherein the multiple nucleic acid probes contain no more than 8 possible reading sequences.

[0391] 18. The method of any one of technical solutions 10-17, wherein the multiple reading sequences are distributed on the multiple nucleic acid probes so as to define an error correction code.

[0392] 19. The method of any one of technical solutions 10-18, wherein the target sequences of the plurality of nucleic acid probes have an average length of 10 to 200 nucleotides.

[0393] 20. The method of any one of technical solutions 10-19, wherein the multiple reading sequences have an average length of 5 nucleotides to 50 nucleotides.

[0394] 21. The method of any one of technical solutions 10-20, wherein at least some of the multiple nucleic acid probes contain no more than 10 reading sequences.

[0395] 22. The method of any one of technical solutions 10-21, wherein at least some of the multiple nucleic acid probes contain no more than 6 reading sequences.

[0396] 23. The method of any one of technical solutions 10-22, wherein at least some of the multiple nucleic acid probes contain no more than 4 reading sequences.

[0397] 24. The method of any one of technical solutions 10-23, wherein at least some of the multiple nucleic acid probes contain no more than 3 reading sequences.

[0398] 25. The method of any one of technical solutions 10-24, wherein at least some of the multiple nucleic acid probes contain no more than 2 reading sequences.

[0399] 26. The method of technical solutions 1-25 also includes exposing the sample to a first secondary probe comprising a first signal transduction entity, wherein the first secondary probe is capable of binding to some reading sequences of the nucleic acid probe, and determining the binding of the nucleic acid probe by determining the first signal transduction entity in the sample.

[0400] 27. The method of technical solution 26, wherein the first signaling entity is fluorescent.

[0401] 28. The method of any one of technical solutions 26 or 27, wherein the first signaling entity is a protein.

[0402] 29. The method of any one of technical solutions 26-28, wherein the first signal transduction entity is a dye.

[0403] 30. The method of any one of technical solutions 26-29, wherein the first signal transduction entity is a nanoparticle.

[0404] 31. The method of any one of technical solutions 26-30 further comprises exposing the sample to a second secondary probe comprising a second signaling entity, wherein the second secondary probe is capable of binding to some reading sequences of the nucleic acid probe, and determining the binding of the nucleic acid probe by determining the second signaling entity within the sample.

[0405] 32. The method of technical solution 31, wherein the first signal conduction entity is the same as the second signal conduction entity.

[0406] 33. The method of technical solution 31, wherein the first signal conduction entity is different from the second signal conduction entity.

[0407] 34. The method of any one of technical solutions 31-33 further includes inactivating the first signaling entity before exposing the sample to the second secondary probe.

[0408] 35. The method of technical solution 34 comprises inactivating the first signaling entity by photobleaching at least some of the first signaling entities.

[0409] 36. The method of any one of technical solutions 34 or 35, comprising inactivating the first signaling entity by chemically bleaching at least some of the first signaling entities.

[0410] 37. The method of any one of technical solutions 34-36 comprises inactivating the first signaling entity by exposing the first signaling entity to a reactant capable of changing the structure of the signaling entity.

[0411] 38. The method of any one of technical solutions 34-37, comprising inactivating the first signaling entity by removing at least some of the first signaling entity.

[0412] 39. The method of any one of technical solutions 34-38, comprising inactivating the first signaling entity by dissociating the first signaling entity from the first secondary probe.

[0413] 40. The method of any one of technical solutions 34-39 comprises inactivating the first signaling entity by dissociating the first secondary probe containing the first signaling entity from the sample.

[0414] 41. The method of any one of technical solutions 34-40, comprising inactivating the first signaling entity by chemically cleaving it from the first secondary probe.

[0415] 42. The method of any one of technical solutions 34-41, comprising inactivating the first signaling entity by enzymatic cleavage from the first secondary probe.

[0416] 43. The method of any one of technical solutions 34-42 comprises inactivating the first signaling entity by exposing the signaling entity or the first secondary probe to an enzyme.

[0417] 44. The method of any one of technical solutions 1-43, comprising determining the center of mass of the first signaling entity using an algorithm for determining non-overlapping single emitters.

[0418] 45. The method of any one of technical solutions 1-44, comprising determining the center of mass of the first signaling entity using an algorithm for determining partially overlapping single emitters.

[0419] 46. ​​The method of any one of technical solutions 44 or 45, which includes using a maximum likelihood algorithm to determine the center of mass.

[0420] 47. The method of any one of technical solutions 44-46 comprises using a least squares algorithm to determine the center of mass.

[0421] 48. The method of any one of technical solutions 44-47 comprises using a Bayesian algorithm to determine the center of mass.

[0422] 49. The method of any one of technical solutions 44-48 comprises using a compressed sensing algorithm to determine the center of mass.

[0423] 50. The method of any one of technical solutions 1-49 further comprises determining the confidence level of the identified nucleic acid target.

[0424] 51. The method of technical solution 51 includes determining the confidence level using the ratio of the number of exact matches to the number of matches with one or more 1-bit errors to the code word.

[0425] 52. The method of technical solution 51 includes using the ratio of the number of exact matches to the number of matches with exactly one 1-bit error to the code word to determine the confidence level.

[0426] 53. The method of technical solutions 1-52, wherein at least some of the multiple nucleic acid probes contain DNA.

[0427] 54. The method of any one of technical solutions 1-53, wherein at least some of the multiple nucleic acid probes contain RNA.

[0428] 55. The method of any one of technical solutions 1-54, wherein at least some of the multiple nucleic acid probes contain PNA.

[0429] 56. The method of any one of technical solutions 1-55, wherein at least some of the multiple nucleic acid probes contain LNA.

[0430] 57. The method of any one of technical solutions 1-56, wherein the plurality of nucleic acid probes have an average length of 10 to 300 nucleotides.

[0431] 58. The method of any one of technical solutions 1-57, wherein at least some of the multiple nucleic acid probes are constructed to bind to nucleic acids within the sample.

[0432] 59. The method of any one of technical solutions 1-58, wherein at least some of the binding of the nucleic acid probe to the target within the sample is specific binding.

[0433] 60. The method of any one of technical solutions 1-59, wherein at least some of the binding of the nucleic acid probe to the target within the sample is through Watson-Crick base pairing.

[0434] 61. The method of any one of technical solutions 1-60, wherein at least some of the multiple nucleic acid probes are constructed to bind RNA.

[0435] 62. The method of any one of technical solutions 1-61, wherein at least some of the multiple nucleic acid probes are constructed to bind to non-coding RNA.

[0436] 63. The method of any one of technical solutions 1-62, wherein at least some of the multiple nucleic acid probes are constructed to bind to mRNA.

[0437] 64. The method of any one of technical solutions 1-63, wherein at least some of the multiple nucleic acid probes are constructed to bind to transfer RNA (tRNA).

[0438] 65. The method of any one of technical solutions 1-64, wherein at least some of the multiple nucleic acid probes are constructed to bind to ribosomal RNA (rRNA).

[0439] 66. The method of any one of technical solutions 1-65, wherein at least some of the multiple nucleic acid probes are constructed to bind to lncRNA.

[0440] 67. The method of any one of technical solutions 1-66, wherein at least some of the multiple nucleic acid probes are constructed to bind to snoRNA.

[0441] 68. The method of any one of technical solutions 1-67, wherein at least some of the multiple nucleic acid probes are constructed to bind to non-coding RNA.

[0442] 69. The method of any one of technical solutions 1-68, wherein the multiple nucleic acid probes are constructed to bind to DNA.

[0443] 70. The method of any one of technical solutions 1-69, wherein at least some of the multiple nucleic acid probes are constructed to bind to genomic DNA.

[0444] 71. The method of any one of technical solutions 1-70 comprises determining the binding of the nucleic acid probe within the sample with a resolution better than 300 nm.

[0445] 72. The method of any one of technical solutions 1-71 comprises determining the binding of the nucleic acid probe within the sample with a resolution better than 100 nm.

[0446] 73. The method of any one of technical solutions 1-72 comprises determining the binding of the nucleic acid probe within the sample with a resolution better than 80 nm.

[0447] 74. The method of any one of technical solutions 1-73 comprises determining the binding of the nucleic acid probe within the sample with a resolution better than 50 nm.

[0448] 75. The method of any one of technical solutions 1-74, wherein the sample comprises cells.

[0449] 76. The method of technical solution 75, wherein the cells are human cells.

[0450] 77. The method of any one of technical solutions 76 or 77, wherein the cells are fixed.

[0451] 78. The method of any one of technical solutions 1-77 comprises determining the binding of the nucleic acid probe by imaging at least a portion of the sample.

[0452] 79. The method of any one of technical solutions 1-78 comprises using optical imaging technology to determine the binding of the nucleic acid probe.

[0453] 80. The method of any one of technical solutions 1-79 comprises using fluorescence imaging technology to determine the binding of the nucleic acid probe.

[0454] 81. The method of any one of technical solutions 1-80 comprises using multi-color fluorescence imaging technology to determine the binding of the nucleic acid probe.

[0455] 82. The method of any one of technical solutions 1-81, which comprises using super-resolution fluorescence imaging technology to determine the binding of the nucleic acid probe.

[0456] 83. The method of technical solution 82 comprises using stochastic optical reconstruction microscopy (STORM) to determine the binding of the nucleic acid probe.

[0457] 84. The method of any one of technical solutions 82 or 83, which comprises using photo-activated localization microscopy (PALM) or fluorescence photo-activated localization microscopy (FPALM) to determine the binding of the nucleic acid probe.

[0458] 85. The method of any one of technical solutions 82-84, comprising using stimulated emission depletion microscopy (STED) to determine the binding of the nucleic acid probe.

[0459] 86. The method of any one of technical solutions 82-85, which comprises using structured illumination microscopy (SIM) to determine the binding of the nucleic acid probe.

[0460] 87. The method of any one of technical solutions 82-86 comprises using reversible saturated optical linear fluorescence transition (RESOLFT) microscopy to determine the binding of the nucleic acid probe.

[0461] 88. A method comprising:

[0462] exposing the sample to a plurality of nucleic acid probes, wherein the nucleic acid probes comprise a first portion comprising a target sequence and a second portion comprising one or more reading sequences, and wherein at least some of the plurality of nucleic acid probes comprise distinguishable nucleic acid probes formed from a combinatorial combination of one or more reading sequences taken from the plurality of reading sequences; and

[0463] For each of the nucleic acid probes, binding of the target sequence of the nucleic acid probe in the sample is determined.

[0464] 89. The method of technical solution 88 further comprises:

[0465] Codewords are generated based on determination of read sequences within the sample.

[0466] 90. The method of technical solution 89 further includes:

[0467] For at least some of the codewords, the codewords are matched to valid codewords, wherein if no match is found, error correction is applied to the codewords to form valid codewords.

[0468] 91. A method comprising:

[0469] exposing the sample to a plurality of primary nucleic acid probes;

[0470] exposing the plurality of primary nucleic acid probes to a sequence of secondary nucleic acid probes and measuring the fluorescence of each secondary nucleic acid probe within the sample;

[0471] generating a codeword based on the fluorescence of the secondary nucleic acid probe; and

[0472] For at least some of the codewords, the codewords are matched to valid codewords, wherein if no match is found, error correction is applied to the codewords to form valid codewords.

[0473] 92. A method comprising:

[0474] exposing a plurality of primary nucleic acid probes to the sample; and

[0475] The plurality of nucleic acid probes are exposed to secondary nucleic acid probe sequences and the fluorescence of each secondary probe within the sample is measured, wherein at least some of the plurality of secondary nucleic acid probes comprise distinguishable secondary nucleic acid probes formed by combinatorial combinations of one or more read sequences obtained from the plurality of read sequences.

[0476] 93. The method of technical solution 92 further comprises:

[0477] A codeword is generated based on the fluorescence of the secondary nucleic acid probe.

[0478] 94. The method of technical solution 93 further comprises:

[0479] For at least some of the codewords, the codewords are matched to valid codewords, wherein if no match is found, error correction is applied to the codewords to form valid codewords.

[0480] 95. A composition comprising:

[0481] A plurality of nucleic acid probes, each nucleic acid probe comprising a first portion comprising a target sequence and a plurality of reading sequences, wherein the plurality of reading sequences are distributed over the plurality of nucleic acid probes so as to define an error correction code.

[0482] 96. A composition of technical solution 95, wherein the multiple nucleic acid probes define a code space having a Hamming distance of at least 2.

[0483] 97. A composition of technical solution 96, wherein the multiple nucleic acid probes define a code space having a Hamming distance of at least 3.

[0484] 98. A composition of any one of technical solutions 96 or 97, wherein the code is a Hamming (7,4) code, a Hamming (15,11) code, a Hamming (31,26) code, a Hamming (63,57) code or a Hamming (127,120) code.

[0485] 99. A composition according to any one of technical solutions 95-98, wherein the code is a SECDED code.

[0486] 100. A composition of technical solution 99, wherein the code is a SECDED (8,4) code.

[0487] 101. A composition of technical solution 99, wherein the code is a SECDED (16,4) code.

[0488] 102. A composition of technical solution 99, wherein the code is a SECDED (16, 11) code, a SECDED (22, 16) code, a SECDED (39, 32) code or a SECDED (72, 64) code.

[0489] 103. A combination of any one of technical solutions 99-102, wherein only code words with a constant number of 1s are used.

[0490] 104. A composition of technical solution 99, wherein the code is an MHD4 code.

[0491] 105. A composition of technical solution 99, wherein the code is an MHD2 code.

[0492] 106. The composition of any one of technical solutions 95-105, wherein the multiple nucleic acid probes contain no more than 100 possible reading sequences.

[0493] 107. The composition of any one of technical solutions 95-106, wherein the multiple nucleic acid probes contain no more than 64 possible reading sequences.

[0494] 108. The composition of any one of technical solutions 95-107, wherein the multiple nucleic acid probes contain no more than 32 possible reading sequences.

[0495] 109. The composition of any one of technical solutions 95-108, wherein the multiple nucleic acid probes contain no more than 16 possible reading sequences.

[0496] 110. The composition of any one of technical solutions 95-109, wherein the multiple nucleic acid probes contain no more than 8 possible reading sequences.

[0497] 111. A composition according to any one of technical solutions 95-110, wherein the target sequences of the plurality of nucleic acid probes have an average length of 10 to 200 nucleotides.

[0498] 112. A composition according to any one of technical solutions 95-111, wherein the plurality of reading sequences have an average length of 5 to 50 nucleotides.

[0499] 113. A composition of any one of technical solutions 95-112, wherein at least some of the multiple nucleic acid probes contain no more than 10 reading sequences.

[0500] 114. A composition of any one of technical solutions 95-113, wherein at least some of the multiple nucleic acid probes contain no more than 6 read sequences.

[0501] 115. A composition according to any one of technical solutions 95-114, wherein at least some of the multiple nucleic acid probes contain no more than 4 read sequences.

[0502] 116. A composition according to any one of technical solutions 95-115, wherein at least some of the multiple nucleic acid probes contain no more than 2 reading sequences.

[0503] 117. A composition of any one of technical solutions 95-116, wherein at least some of the multiple nucleic acid probes contain DNA.

[0504] 118. A composition according to any one of technical solutions 95-117, wherein at least some of the multiple nucleic acid probes contain RNA.

[0505] 119. A composition according to any one of technical solutions 95-118, wherein at least some of the multiple nucleic acid probes comprise PNA.

[0506] 120. A composition according to any one of technical solutions 95-119, wherein at least some of the multiple nucleic acid probes contain LNA.

[0507] 121. A composition according to any one of technical solutions 95-120, wherein the multiple nucleic acid probes have an average length of 10 to 300 nucleotides.

[0508] 122. A method comprising:

[0509] associating a plurality of targets with a plurality of target sequences and a plurality of codewords, wherein the codewords comprise a plurality of positions and a value for each position, and the codewords form an error checking and / or error correction code space;

[0510] associating a plurality of distinguishable read sequences with the plurality of code words such that each distinguishable read sequence represents a value of a position within the code word; and

[0511] A plurality of nucleic acid probes are formed, each nucleic acid probe comprising a target sequence and one or more reading sequences.

[0512] 123. The method of technical solution 122, wherein the code space has a Hamming distance of at least 2.

[0513] 124. The method of any one of technical solutions 122 or 123, wherein the multiple targets include RNA.

[0514] 125. The method of any one of technical solutions 122-124, wherein the multiple targets include DNA.

[0515] 126. A method comprising:

[0516] Exposing the nucleic acid probe formed by the method of any one of technical solutions 122-125 to a sample;

[0517] sequentially determining the binding of each of the plurality of secondary nucleic acid probes to the nucleic acid probe in the sample;

[0518] determining a sample codeword comprising a binding pattern of the plurality of secondary nucleic acid probes to the nucleic acid probe; and

[0519] The sample codeword is matched to the codeword to determine the target within the sample.

[0520] 127. The method of technical solution 126, wherein at least some of the multiple secondary nucleic acid probes are capable of binding to the reading sequence of the nucleic acid probe.

[0521] 128. The method of technical solution 127, wherein at least some of the multiple secondary nucleic acid probes are capable of specifically binding to only one reading sequence.

[0522] 129. The method of technical solutions 125-128, wherein if no match is found between the valid code word and the sample code word, error correction is applied to the sample code word to identify the matching code word.

[0523] 130. The method of technical solutions 125-129, wherein at least some of the secondary nucleic acid probes each contain a signal transduction entity.

[0524] 131. The method of technical solution 130 comprises imaging the signal transduction entity to determine the binding of the multiple secondary nucleic acid probes within the sample.

[0525] 132. The method of technical solution 131 comprises imaging the signal transduction entity using super-resolution imaging technology.

[0526] 133. The method of any one of technical solutions 131 or 132 also includes applying drift correction to the image.

[0527] 134. The method of any one of technical solutions 131-133 also includes applying physical alignment to the image.

[0528] 135. The method of technical solutions 125-134 comprises determining the binding of the multiple secondary nucleic acid probes in the sample with a resolution better than 100 nm.

[0529] 136. A method comprising:

[0530] associating a plurality of targets with a plurality of target sequences and a plurality of codewords, wherein the codewords comprise a plurality of positions and a value for each position, and the codewords form an error checking and / or error correction code space;

[0531] forming a plurality of nucleic acid probes, each nucleic acid probe comprising a target sequence; and

[0532] Groups comprising the plurality of nucleic acid probes are formed such that each group of nucleic acid probes corresponds to at least one common value for a position within the codeword.

[0533] 137. The method of technical solution 136, wherein the codeword includes a plurality of positions less than the number of targets.

[0534] 138. A method comprising:

[0535] associating a plurality of targets with a plurality of target sequences and a plurality of codewords, wherein the codewords comprise a plurality of positions less than the number of targets, and wherein each codeword is associated with a single target;

[0536] associating a plurality of distinguishable read sequences with the plurality of code words such that each distinguishable read sequence represents a value of a position within the code word; and

[0537] A plurality of nucleic acid probes are formed, each nucleic acid probe comprising a target sequence and one or more reading sequences.

[0538] 139. The method of technical solution 138, wherein the code words form an error checking space.

[0539] 140. A method according to any one of technical solutions 138 or 139, wherein the code words form an error correction code space.

[0540] 141. A method comprising:

[0541] associating a plurality of targets with a plurality of target sequences and a plurality of codewords, wherein the codewords comprise a plurality of positions and a value for each position, and the codewords form an error checking and / or error correction code space;

[0542] forming a plurality of nucleic acid probes, each nucleic acid probe comprising a target sequence; and

[0543] Groups comprising the plurality of nucleic acid probes are formed such that each group of nucleic acid probes corresponds to at least one common value for a position within the codeword.

[0544] 142. The method of any one of technical solutions 138-141, wherein the groups are applied to the sample sequentially using a fluid device.

[0545] 143. The method of any one of technical solutions 138-142, wherein each group of nucleic acid probes contains a group identification sequence that can be distinguished from the group identification sequences of other groups.

[0546] 144. A method comprising:

[0547] associating a plurality of nucleic acid targets with a plurality of target sequences and a plurality of code words, wherein the code words comprise a plurality of positions and a value for each position, and wherein the code words form an error checking and / or error correction code;

[0548] associating a unique read sequence with each possible value for each position in the codeword, wherein the read sequence is taken from a set of orthogonal sequences having limited homology to each other and to nucleic acid species in the sample;

[0549] forming a plurality of primary nucleic acid probes, each primary nucleic acid probe comprising a target sequence that uniquely binds to a nucleic acid target and one or more reading sequences;

[0550] forming a plurality of secondary nucleic acid probes comprising a signaling entity and a sequence complementary to one of the reading sequences;

[0551] exposing the sample to the primary nucleic acid probe so that the nucleic acid probe hybridizes to a nucleic acid target in the sample;

[0552] exposing the primary nucleic acid probes in the sample to secondary nucleic acid probes so that the secondary nucleic acid probes hybridize to the reading sequences on at least some of the primary nucleic acid probes;

[0553] imaging the sample; and

[0554] The exposing and imaging steps are repeated one or more times, with at least some of the repetitions being performed using different secondary nucleic acid probes.

[0555] 145. The method of technical solution 144, wherein the reading sequence is taken from a set of orthogonal sequences, and the orthogonal sequences have less than 15 base pairs of homology with each other or with the nucleic acid types in the sample.

[0556] 146. The method of any one of technical solutions 144 or 145 comprises combining the images of the samples into a combined image.

[0557] 147. The method of technical solution 146 also includes identifying prospect features corresponding to putative nucleic acid targets in the sample.

[0558] 148. The method of any one of technical solutions 144-147, comprising determining the center of mass of the first signaling entity using an algorithm for determining non-overlapping single emitters.

[0559] 149. The method of any one of technical solutions 144-148, comprising determining the center of mass of the first signaling entity using an algorithm for determining partially overlapping single emitters.

[0560] 150. The method of any one of technical solutions 144-149, which comprises using a maximum likelihood algorithm to determine the center of mass.

[0561] 151. The method of any one of technical solutions 144-150 comprises using a least squares algorithm to determine the center of mass.

[0562] 152. The method of any one of technical solutions 144-151, which includes using a Bayesian algorithm to determine the center of mass.

[0563] 153. The method of any one of technical solutions 144-152 comprises using a compressed sensing algorithm to determine the center of mass.

[0564] 154. The method of any one of technical solutions 144-153 comprises determining a code word corresponding to each foreground feature by determining a binding pattern between a nucleic acid probe library and the sample.

[0565] 155. The method of technical solution 154 includes matching the code word with a valid code word, wherein if no match is found, error correction is applied to the code word to form a valid code word corresponding to the nucleic acid target.

[0566] 156. The method of any one of technical solutions 144-155 comprises imaging the sample using a fluorescence microscope.

[0567] 157. The method of any one of technical solutions 144-156 comprises imaging the sample with a resolution better than 500 nm.

[0568] 158. The method of any one of technical solutions 144-157 comprises imaging the sample with a resolution better than 300 nm.

[0569] 159. The method of any one of technical solutions 144-158 comprises imaging the sample with a resolution better than 100 nm.

[0570] 160. The method of any one of technical solutions 144-159 comprises imaging the sample with a resolution better than 80 nm.

[0571] 161. The method of any one of technical solutions 144-160 comprises imaging the sample with a resolution better than 50 nm.

[0572] 162. The method of any one of technical solutions 144-161 comprises imaging the sample with a resolution better than 20 nm.

[0573] 163. The method of any one of technical solutions 144-162 comprises imaging the sample using super-resolution fluorescence imaging technology.

[0574] 164. The method of any one of technical solutions 144-163 comprises imaging the sample using a technique selected from STORM, PALM, FPALM, STED, SIM, RESOLFT, SOFI or SPDM.

[0575] 165. The method of any one of technical solutions 144-164, wherein the codeword comprises a plurality of positions less than the number of targets.

[0576] 166. A method comprising:

[0577] associating a plurality of nucleic acid targets with a plurality of target sequences and a plurality of code words, wherein the code words comprise a plurality of positions and a value for each position, and the code words form an error checking and / or error correction code space;

[0578] forming a plurality of nucleic acid probes comprising a signaling entity and a target sequence that specifically binds to one of the nucleic acid targets;

[0579] grouping the nucleic acid probes into a plurality of probe pools, wherein each probe pool corresponds to a specific value of a unique position within the codeword;

[0580] exposing the sample to one of the probe libraries;

[0581] imaging the sample; and

[0582] The exposing and imaging steps are repeated one or more times, using different probe pools for at least some of the repeats.

[0583] 167. The method of technical solution 166 comprises combining the images of the sample into a combined image.

[0584] 168. The method of technical solution 167 also includes identifying prospect features corresponding to putative nucleic acid targets in the sample.

[0585] 169. The method of technical solution 168 comprises determining a code word corresponding to each foreground feature by determining a binding pattern between the nucleic acid probe library and the sample.

[0586] 170. The method of technical solution 169 includes matching the determined code word with a valid code word, wherein if no match is found, error correction is applied to the determined code word to form a valid code word corresponding to the nucleic acid target.

[0587] 171. The method of any one of technical solutions 166-170, which comprises imaging the sample using a fluorescence microscope.

[0588] 172. The method of any one of technical solutions 166-171 comprises imaging the sample with a resolution better than 300 nm.

[0589] 173. The method of any one of technical solutions 166-172 comprises imaging the sample with a resolution better than 100 nm.

[0590] 174. The method of any one of technical solutions 166-173 comprises imaging the sample using super-resolution fluorescence imaging technology.

[0591] 175. The method of any one of technical solutions 166-174 comprises imaging the sample using a technique selected from STORM, PALM, FPALM, STED, SIM, RESOLFT, SOFI or SPDM.

[0592] 176. The method of any one of technical solutions 166-175, wherein the codeword comprises a plurality of positions less than the number of targets.

Claims

1. A method wherein include: associating a plurality of nucleic acid targets in a sample with a plurality of codewords, wherein the codewords include a plurality of positions and a value for each position; exposing the sample to a plurality of nucleic acid probes, wherein at least some of the plurality of nucleic acid probes comprise a target sequence; For each of the nucleic acid probes, determining binding of the nucleic acid probe within the sample; generating a codeword based on binding of the plurality of nucleic acid probes within the sample, wherein a value of a position in the codeword is based on the presence or absence of binding of different subsets of the plurality of nucleic acid probes; and for at least some of the codewords, matching the codewords to valid codewords, wherein if no match is found, discarding the codeword or applying error correction to the codeword to form a valid codeword, the valid codeword being a plurality of codewords assigned to the plurality of nucleic acid targets; and The abundance and / or spatial distribution of nucleic acids within the sample is determined using valid code words corresponding to the binding of the plurality of nucleic acid probes within the sample.

2. The method of claim 1, wherein the valid code words form an error checking and / or error correction code space.

3. The method of any one of claims 1-2, comprising exposing the sample to at least 5 different nucleic acid probes.

4. The method of any one of claims 1 or 2, comprising exposing the sample to at least 10 different nucleic acid probes.

5. The method of any one of claims 1-2, comprising exposing the sample to at least 100 different nucleic acid probes.

6. The method of any one of claims 1-2, comprising exposing the sample to a plurality of nucleic acid probes simultaneously.

7. The method of any one of claims 1-2, comprising sequentially exposing the sample to a plurality of nucleic acid probes.

8. The method of claims 1-2, wherein the plurality of nucleic acid probes comprises a combinatorial combination of nucleic acid probes having different sequences.

9. The method of claim 8, wherein the combinatorial combination of nucleic acid probes targets a combinatorial combination of RNA species in the sample.

10. The method of claim 7, wherein the combinatorial combination of nucleic acid probes targets a combinatorial combination of DNA sequences in the sample.

Citation Information

Patent Citations

  • Systems and methods for detecting nucleic acids

    CN112029826B

  • Sub-diffraction limit image resolution and other imaging techniques

    US7838302B2

  • Sub-diffraction limit image resolution in three dimensions

    US8564792B2

  • High resolution dual-objective microscopy

    WO2013090360A2