Imaging-based pooled crispr screening
The method improves imaging-based pooled CRISPR screening by introducing DNA with guide and identification portions and using co-localization to create codewords, enabling high-throughput determination of genotype-phenotype correspondences and complex cellular properties.
Patent Information
- Application Number
- JP2025167369
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-05-01
- Filing Date
- 2025-10-03
- Publication Date
- 2026-02-10
AI Technical Summary
Existing imaging-based pooled CRISPR screening methods struggle to determine genotype-phenotype correspondences in situ for individual cells due to challenges in genotyping phenotypically imaged cells, limiting the measurement of cellular structures and intracellular molecular organization.
A method involving the introduction of DNA with guide, reporter, and identification portions into cells, followed by imaging and co-localization of readout probes to create a codeword for genotype-phenotype determination, using techniques like MERFISH to improve decoding accuracy.
Enables high-throughput screening for genotype-phenotype correspondences, allowing for the analysis of complex cellular properties such as cell shape and intracellular molecular organization with reduced misidentification of background signals.
Smart Images

Figure 2026021332000001_ABST
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 836,578, filed April 19, 2019, by Zhuang et al., entitled "Imaging-Based Pooled CRISPR Screening"; and U.S. Provisional Patent Application No. 62 / 841,715, filed May 1, 2019, by Zhuang et al., entitled "Imaging-Based Pooled CRISPR Screening," each of which is incorporated herein by reference in its entirety.
[0002] Government funding This invention was made with government support under grant number MH113094 awarded by the National Institutes of Health. The United States Federal Government has certain rights in this invention.
[0003] The present invention relates generally to imaging cells, e.g., determining phenotypes and / or genotypes within a population of cells. In some cases, the cells may be engineered, e.g., using CRISPR or other techniques. [Background technology]
[0004] The development of CRISPR-based gene editing systems has significantly advanced our ability to manipulate genes and probe the molecular mechanisms underlying cellular function through gene perturbations. Driven by the ability to generate highly diverse nucleic acid libraries, CRISPR-based pooled library screening can substantially accelerate the discovery of genes involved in cellular processes. However, the accessible phenotypes in pooled library screening are primarily limited to cell viability and marker expression. Recently, single-cell RNA sequencing and mass cytometry have been combined with CRISPR screening to expand the accessible phenotype space of pooled library screening and enable genetic screening based on single-cell profiles of RNA and protein expression.
[0005] However, many important cellular phenotypes remain beyond the reach of high-throughput pooled library screening. These include the shape of cellular structures and intracellular molecular organization, as well as their dynamics, which can only be measured with techniques such as high-resolution imaging. High-content imaging further enables these properties to be measured simultaneously for many molecular species in a parallelized manner (e.g., the recent development of single-cell transcriptome imaging has increased the number of molecular phenotypes that can be imaged in individual cells in a single experiment to the genome scale). Despite the power of imaging in assessing cellular phenotypes, imaging-based pooled library screening remains challenging, primarily due to the difficulties associated with genotyping individual, phenotypically imaged cells in pooled library screening. Methods have been developed that allow cells with a particular phenotype to be genotyped by sequencing after physical separation. However, to determine the complete genotype-phenotype correspondence, all imaging-based pooled library screening methods are required, in which both genotype and phenotype are imaged in situ for individual cells. Summary of the Invention [Problem to be solved by the invention]
[0006] The present invention generally relates to cell imaging, for example, determining the phenotype and / or genotype in a cell population.In some cases, cells can be manipulated, for example, by using CRISPR or other techniques.The subject matter of the present disclosure may involve interrelated products, alternative solutions to specific problems, and / or multiple different uses of one or more systems and / or items. [Means for solving the problem]
[0007] In one aspect, the present invention is generally directed to a method. According to one set of embodiments, the method includes: (a) introducing into a plurality of cells DNA comprising a guide portion comprising a recognition sequence, a reporter portion, and an identification portion comprising a lead sequence; (b) determining the location of an RNA molecule expressed from the reporter portion of the introduced DNA into the plurality of cells by determining the reporter portion; (c) determining the lead sequence on the RNA molecule expressed from the DNA comprising the reporter portion and the identification portion introduced into the plurality of cells by exposing the cells to a readout probe capable of binding to the lead sequence; (d) co-localizing the binding of the readout probe with the location of the RNA molecule expressed from the reporter portion of the introduced DNA; (e) repeating (b), (c), and (d) multiple times using different read sequences; and (f) creating a codeword corresponding to the binding of the co-localized readout probe, the numerical value of the codeword being based on the binding of the readout probe to the lead sequence.
[0008] According to another set of embodiments, a method includes introducing into a plurality of cells DNA comprising a guide portion comprising a recognition sequence, a reporter portion, and an identification portion comprising a lead sequence; determining the location of an RNA molecule expressed from the reporter portion of the introduced DNA in the plurality of cells by determining the reporter portion; determining the lead sequence in the plurality of cells by exposing the cells to a plurality of readout probes, each capable of binding to the lead sequence; co-localizing the binding of the readout probes with the location of the RNA molecule expressed from the reporter portion of the introduced DNA; and creating a code word corresponding to the binding of the co-localized readout probes, the numerical value of the code word being based on the binding of the readout probe to the lead sequence.
[0009] In yet another set of embodiments, a method includes introducing a nucleic acid into a plurality of cells, wherein the nucleic acid comprises a guide portion comprising a recognition sequence, a reporter portion, and an identifying portion comprising a lead sequence; and imaging the plurality of cells, wherein the cells exhibit differences in an imageable phenotype due to expression of the guide portion; and collecting a plurality of images of the plurality of cells, wherein the images of the cells exhibit differences due to differences in the identifying portions of the nucleic acid within the cells.
[0010] In yet another set of embodiments, a method includes using a lentivirus to introduce DNA into a plurality of cells, wherein the DNA comprises a guide portion comprising a recognition sequence, a reporter portion, and an identification portion comprising a lead sequence; determining a phenotype of the plurality of cells; and determining a genotype of the plurality of cells; and determining a correspondence between the genotype and the phenotype.
[0011] According to yet another set of embodiments, a method includes using a lentivirus to introduce DNA into a plurality of cells, wherein the DNA comprises a guide portion comprising a recognition sequence and an identification portion comprising a lead sequence; determining a phenotype of the plurality of cells; determining a genotype of the plurality of cells; and determining a correspondence between the genotype and the phenotype.
[0012] In another aspect, the invention includes methods of performing one or more of the embodiments described herein. In yet another aspect, the invention includes methods of using one or more of the embodiments described herein.
[0013] Other advantages and novel features of the present invention will become apparent from the following detailed description of various non-limiting embodiments of the invention when considered in conjunction with the accompanying drawings.
[0014] Non-limiting embodiments of the present invention are described, by way of example, with reference to the accompanying drawings, which are schematic and are not intended to be drawn to scale. In the figures, each identical or nearly identical component illustrated is typically represented by a single numeral. For purposes of clarity, unless illustration is necessary to enable those skilled in the art to understand the invention, not every component will be shown in every figure, nor will every component of every embodiment of the invention be shown. [Brief explanation of the drawings]
[0015] [Figure 1-1] 1A-1F illustrate imaging-based barcode detection for genotyping in accordance with one embodiment of the present invention. [Figure 1-2] Same as above. [Figure 1-3] Same as above. [Figure 1-4] Same as above. [Figure 2-1] 2A to 2D illustrate the barcode misidentification rate in another embodiment of the present invention. [Figure 2-2] Same as above. [Figure 3-1] 3A-3E illustrate lentiviral designs in yet another embodiment of the present invention. [Figure 3-2] Same as above. [Figure 3-3] Same as above. [Figure 4-1] 4A-4D illustrate imaging-based pooled CRISPR screening in yet another embodiment of the present invention. [Figure 4-2] Same as above. [Figure 5-1] 5A-5C illustrate genetic elements involved in regulation according to one embodiment of the present invention. [Figure 5-2] Same as above. [Figure 6-1] 6A-6B illustrate certain genes used for transcription inhibition in another embodiment of the present invention. [Figure 6-2] Same as above. [Figure 7] FIG. 7 illustrates a cloning strategy for a library in one embodiment of the present invention. [Figure 8-1] FIG. 8 illustrates a colocalization rate analysis in another embodiment of the present invention. [Figure 8-2] Same as above. [Figure 9] FIG. 9 illustrates a cloning strategy for a library in yet another embodiment of the present invention. [Figure 10-1] 10A-10D illustrate the knockdown of a specific gene in one embodiment of the present invention. [Figure 10-2] Same as above. [Figure 11] FIG. 11 illustrates enrichment changes in nuclear speckles for MALAT1 in another embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] The present invention generally relates to imaging cells, e.g., determining phenotypes and / or genotypes within a cell population, e.g., establishing genotype-phenotype correspondences for high-throughput screening. In some cases, cells can be engineered using, e.g., CRISPR or other techniques. In certain embodiments, nucleic acids can be introduced into cells using, e.g., lentivirus. The nucleic acids can contain a guide portion including a DNA or RNA recognition sequence, a reporter portion, and an identifying portion including one or more lead sequences. The guide portion can be used to alter the phenotype of a cell, e.g., using a sequence that can be targeted using CRISPR or other techniques, e.g., an sgRNA sequence, and in some cases, the phenotype of the cell can be determined using various imaging methods. The identifying portion can be determined using MERFISH or other suitable techniques. In addition, in some cases, the correlation or co-localization of reporter sequence determination and lead sequence determination can substantially improve decoding accuracy, e.g., due to reduced misidentification of background signals. Other aspects are generally directed to compositions or devices for use in such methods, kits for use in such methods, and the like.
[0017] One exemplary aspect of the present invention is generally directed to systems and methods for manipulating the genetic material of a cell, for example, using CRISPR or other techniques, and determining the phenotype of the cell resulting from this manipulation. The genotype of the cell can also be determined using code words encoding the read sequences, for example, code words used in MERFISH or similar techniques. By determining both the genotype of a cell and how the phenotype of the cell is altered, certain embodiments discussed herein can be useful for understanding complex questions, such as, for example, understanding cell shape, intracellular molecular organization, and the like, spatially within a cell, e.g., a mammalian cell.
[0018] An exemplary embodiment of the present invention will now be described with reference to Figure 1A. In this figure, members of a library of nucleic acids can be introduced into a cell, such as a mammalian cell. In one set of embodiments, the nucleic acids include a guide portion (e.g., containing an sgRNA or another recognition sequence that can be used to recognize a target site), a reporter portion (e.g., that can directly or indirectly generate a signal, such as a fluorescent signal or an immunoprecipitation signal), and an identifying portion or "barcode" portion (e.g., containing a lead sequence that can be used to distinguish various nucleic acids containing different guide portions from one another).
[0019] Various methods can be used to introduce nucleic acids into cells. These include, for example, viral delivery (e.g., using lentivirus, retrovirus, adenovirus, adeno-associated virus, etc.), electroporation, ballistic delivery, etc. In some cases, lentivirus can be useful because it allows stable integration of nucleic acids into the genome of a cell. In addition, in certain embodiments, the rate of nucleic acid introduction into cells can be controlled so that the majority of cells contain only one such nucleic acid. For example, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the cells can have only one such nucleic acid introduced therein.
[0020] In this example, for lentiviruses generated from a pooled library, each lentivirus may contain two members of the library. During the lentiviral infection process, the guide portion and the identifying portion may be recombined. Such recombination may result in misidentification of the guide portion based on the measurement of the identifying portion. Thus, in one set of embodiments, the guide portion and the identifying portion may be positioned adjacent to each other within the 3'LTR region of the lentivirus, i.e., after the PPT (polypurine tract) sequence, so that the distance between the guide portion and the identifying portion is minimal, for example, 100 bases or less for the non-variable region of the sgRNA for Cas9. In this way, the recombination rate may be reduced to improve accurate association between the guide portion and the identifying portion. In some cases, the guide portion is duplicated, for example, within the 5' region of the lentiviral proviral DNA. This may allow the guide portion to be integrated into the host cell genome, resulting in expression of the guide portion.
[0021] After introduction, the cells can be studied to determine the phenotype and genotype of the cells (e.g., using an identifying moiety). For example, the phenotype can be measured using imaging methods that detect proteins, RNA, or DNA within cells or subcompartments of cells. In certain embodiments, the phenotype can also relate to cell growth, shape, or cell-cell interactions. In some cases, the phenotype can be a change in a cell's characteristics over time, dynamics, etc. In some cases, the phenotype can include multiplexed features, i.e., multidimensional readouts.
[0022] The identifying portion can be determined, for example, using MERFISH (multiplexed error-robust fluorescence in situ hybridization) or other techniques. Those skilled in the art will be familiar with MERFISH and related techniques. See, for example, International Patent Application Publication Nos. WO2016 / 018960, WO2016 / 018963, WO2018 / 089445, WO2018 / 218150, and WO2018 / 089438. In some embodiments, the identifying portion can contain multiple "read sequences" or nucleic acid sequences that can be specifically identified using corresponding nucleic acid probes (e.g., "readout probes") sequentially. In techniques such as MERFISH, the presence or absence of a lead sequence can be coded as a numerical value, and thus the sequence of the readout probe can be coded as a code word. Additionally, various error detection and / or correction methods may optionally be applied to the codewords, such as Hamming codes or Golay codes.
[0023] In some cases, determinations of reporter moieties may be interspersed with determinations of various portions of identifying moieties (e.g., using one or more readout probes). Additionally, in some embodiments, association or co-localization of reporter moiety positions with determinations of identifying moieties may be used to substantially improve decoding accuracy. For example, binding events or code words that do not sufficiently correspond to positions where reporter moieties are present may be ignored as background noise, non-specific labeling, etc. Association or co-localization of reporter moieties with identifying moieties may substantially improve detection accuracy.
[0024] The above discussion is a non-limiting example of one embodiment of the present invention, which can be used to image cells and, for example, determine phenotypes and / or genotypes in a cell population.However, other embodiments are also possible.Therefore, more generally, various aspects of the present invention are directed to various systems and methods for determining phenotypes and / or genotypes in a cell population, for example, via imaging, and / or manipulating cells using CRISPR or other techniques.
[0025] According to one aspect, the present invention generally relates to a system and method for determining the phenotype and / or genotype of a cell population using imaging.In addition, the genome of a cell can be manipulated, for example, by using CRISPR or other techniques.In some cases, using suitable imaging methods, for example, the techniques described herein, a relatively large number of cells can be studied to determine their phenotype and genotype, for example, after manipulation.In some embodiments, by using such imaging methods as discussed herein, which allow for relatively large-scale or high-throughput screening, a relatively large number of cells can be determined.For example, a plurality of cells can be determined for a specific phenotype (for example, after editing by CRISPR), and the cells with a certain phenotype or desired phenotype can also be determined genotypically.
[0026] In some cases, a relatively large number of cells can be determined. For example, depending on the magnification, a single field of view can contain a relatively large number of cells (for example, at least 10, at least 100, at least 1,000, at least 10,000, at least 100,000, etc.). In addition, the sample can be larger than a single field of view (for example, particularly in the case of a relatively high magnification), and multiple images of different parts of the sample can be collected, for example, manually or automatically (for example, using computer control). This can allow the use of more than one field of view to study even larger numbers of cells, for example, at least 10, at least 100, at least 1,000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, etc. For example, an overall image of a sample may be assembled using multiple fields of view (e.g., imaged simultaneously or nearly simultaneously) to provide an image; e.g., at least two, at least three, at least five, at least seven, at least 10, at least 15, at least 20, at least 30, at least 50, at least 75, or at least 100 images may be collected in different fields of view (e.g., corresponding to different portions of the sample) to provide the overall image. Thus, in some cases, the sample may be substantially larger than a single field of view. For example, the sample may have an area of at least about 0.01 cm, at least about 0.03 cm, at least about 0.1 cm, at least about 0.3 cm, at least about 1 cm, at least about 3 cm, or at least about 10 cm, etc.
[0027] Additionally, in some embodiments, multiple images may be taken of the same field of view, for example, at least 2, at least 3, at least 5, at least 7, at least 10, at least 15, at least 20, at least 30, at least 50, at least 75, or at least 100 images may be collected of the same field of view.
[0028] In some cases, in one set of embodiments, multiple images may be taken of each field of view being imaged in the sample. In some embodiments, different wavelengths may be used. For example, in some cases, images may be collected, for example, with different illumination sources and captured using different optical filters to produce different colors of the image that probe the presence of different fluorescent compounds. Thus, in some embodiments, multiple images may be taken at different wavelengths, for example, to view the images in different colors (e.g., red-green-blue, red-yellow-blue, cyan-magenta-yellow, etc.).
[0029] In some embodiments, these images can be collected at defined time intervals to create time-lapse images of the sample. This can be useful, for example, to determine characteristics that change over time, such as cell proliferation. For example, an image (or multiple images) can be collected at different time points, for example, at a periodicity of about 5 seconds, about 10 seconds, about 15 seconds, about 30 seconds, about 1 minute, about 2 minutes, about 3 minutes, about 5 minutes, about 10 minutes, about 15 minutes, about 20 minutes, about 30 minutes, about 1 hour, about 2 hours, about 3 hours, about 4 hours, about 1 day, etc. Similarly, in some embodiments, images can be collected after different treatments of the same sample.
[0030] In some embodiments, multiple images may also be collected by different imaging modalities, including techniques described herein, such as super-resolution optical microscopy, conventional epifluorescence microscopy, confocal microscopy, etc. Such images may optionally be combined to create high-content optical measurements of cellular properties.
[0031] The cell can be any suitable cell, for example, a mammalian cell (e.g., a human cell or a non-human cell), a bacterial cell (e.g., E. coli), a eukaryotic cell, a prokaryotic cell, a yeast cell, or other type of cell. The cell can originate from any suitable source, for example, a cell culture. In some cases, the cell can be obtained from a tissue sample, for example, from a biopsy, or can be artificially grown or cultured, etc. In some cases, the cell is genetically engineered. In some cases, the tissue sample can be analyzed. In certain embodiments, a plurality of cells can be transfected as discussed herein, and the phenotype of the resulting cells is determined.
[0032] In certain embodiments, the nucleic acid that can be used to modify the genetic material of cell, for example, its genome, is introduced into cell.Technology such as CRISPR or other related techniques can be used to modify the genetic material of cell, for example, as guided by nucleic acid.In some embodiments, this can allow the genetic manipulation of cell and the accurate identification of its corresponding phenotype, using identifying part, so as to identify the genotype that causes observed phenotype.
[0033] For example, in one set of embodiments, the nucleic acid delivered to the cell may include a guide moiety, and / or a reporter moiety, and / or an identifying moiety. The guide moiety may contain, for example, an sgRNA or another recognition sequence that can be used to recognize a target site within the genome of the cell. The reporter moiety may be capable of directly or indirectly generating a signal, such as a fluorescent signal. For example, the reporter moiety may encode a fluorescent protein (e.g., GFP), an enzyme (e.g., luciferase) that can be used to make another molecule fluorogenic, an enzyme that produces a detectable chemical reaction, etc. The identifying moiety may include a sequence that can be used to distinguish various nucleic acids containing different guide moieties from one another. For example, the identifying moiety may include one or more sequences (e.g., "read sequences") that can be read using a corresponding nucleic acid probe (e.g., "readout probe").
[0034] When present, the guide portion, and / or reporter portion, and / or identifying portion can be arranged in any suitable order on the nucleic acid to be introduced into the cell. In some cases, these portions can be relatively close to each other (e.g., less than 5,000 bases, less than 3,000 bases, less than 1,000 bases, less than 500 bases, less than 300 bases, less than 100 bases, less than 50 bases, less than 30 bases, or less than 10 bases apart from each other within the nucleic acid). In addition, in some cases, one or more of these portions can at least partially overlap, for example, within the nucleic acid. Furthermore, in some embodiments, other portions or sequences can also be present within the nucleic acid. For example, one or more of these portions can contain a promoter sequence, such as the promoter sequence discussed herein.
[0035] In one set of embodiments, the nucleic acid comprises an expression portion or a guide portion. The guide portion may comprise any suitable nucleic acid sequence suspected of being capable of altering the phenotype of a cell and / or that can be used to intentionally alter or manipulate the genome of a cell, e.g., resulting in an observable change in the phenotype of a cell. For example, the guide portion may encode a sequence encoding a gene, a protein, a regulatory sequence (e.g., an operon, a promoter such as a CMV promoter, a repressor, a transcription factor binding site, etc.), a non-coding RNA (e.g., miRNA, siRNA, rRNA, tRNA, lncRNA, snoRNA, snRNA, exRNA, piRNA, tsRNA, rsRNA, shRNA, Cas9 guide RNA, sgRNA, etc.), etc. In some cases, the guide portion may be part of the same nucleic acid as the identifying portion; in other cases, the expression portion may be part of a different nucleic acid.
[0036] Thus, for example, the guide portion may include a sequence, such as an RNA sequence, that recognizes a target region of interest on DNA (e.g., on the genome of a cell). In some cases, the guide portion may also include a binding sequence, such as a Cas binding sequence, that can be recognized by a Cas nuclease or another nuclease. For example, in certain cases, the guide portion may be a guide portion suitable for enabling CRISPR editing of the genome to occur. For example, the guide portion may include a gRNA (guide RNA) or an sgRNA (single-stranded guide RNA). In some embodiments, the sgRNA may include a crisprRNA portion (crRNA) that is sequence-complementary to the target sequence (e.g., target DNA) and a tracrRNA portion that can be recognized by a Cas nuclease or another nuclease. In some cases, the crRNA portion may have 17, 18, 19, or 20 nucleotides. A variety of different Cas nucleases can be used, such as Cas9 (derived from Streptococcus pyogenes), Cas14, CasX, CasY, Cas12a, Cas13a, Cas13b, Cas13d, Cas14a, etc. Mutant forms of these Cas nucleases are also envisioned, such as high-fidelity Cas9, eSpCas9, SpCas9-HF1, HypaCas9, FokI-fused dCas9, xCas9, dCas9, etc. Non-limiting examples of binding sequences suitable for Cas are provided below. In addition, those skilled in the art are familiar with CRISPR and related techniques, and kits useful for carrying out CRISPR experiments are readily available commercially.
[0037] In certain embodiments, there may be more than one possibility for a guide moiety. For example, a library of nucleic acids may be prepared, e.g., having different crRNA moieties, e.g., for binding to different target sequences within a genome. In certain cases, there may be at least 10, e.g., at least 10, e.g., for guide moieties capable of binding to different target sites within a genome and / or causing different changes or manipulations of the genome. 2 , at least 10 3, at least 10 4 , at least 10 5 There may be multiple possibilities, such as 100,000,000. Thus, in certain embodiments, multiple distinguishable nucleic acids may be prepared using one or more identifying moieties (such as the techniques described herein) and one or more guide moieties. However, it should be understood that the number of possible identifying moieties need not equal the number of possible guide moieties, i.e., there may be some degree of redundancy, as discussed below, for example.
[0038] In some embodiments, nucleic acids may contain reporter moieties that can be determined, for example, by fluorescence or other detection methods. For example, reporter moieties may include genes encoding fluorescent proteins such as GFP (green fluorescent protein), red fluorescent protein, derived from dsRed, PAGFP, PSCFP, PSCFP2, Dendra, Dendra2, EosFP, tdEos, mEos2, mEos3, PAmCherry, PAtagRFP, mMaple, mMaple2, and mMaple3. Other suitable fluorescent proteins are known to those skilled in the art. For example, see U.S. Patent No. 7,838,302; or U.S. Patent Application No. 61 / 979,436, each of which is incorporated herein by reference in its entirety.
[0039] In another set of embodiments, the reporter moiety may encode an enzyme (e.g., luciferase) that can be used to make another molecule fluorogenic. When expressed within a cell, an appropriate substrate (e.g., luciferin) can be added that can be converted to a fluorogenic form upon exposure to the enzyme. However, in areas where no nucleic acid is present, no such fluorescence occurs. In this way, nucleic acids can be localized or located within a cell (or in portions of a cell).
[0040] It should be understood that the reporter moiety need not be determinable solely via fluorescence. Other reporter moieties may be used in other embodiments. For example, in one embodiment, an enzyme or the like that produces a detectable chemical reaction may be encoded within the reporter moiety. Further examples of reporters that may be used include, but are not limited to, proteins that are detectable by immunoprecipitation, immunofluorescence, and the like. Non-limiting examples of suitable proteins include Myc tags or HA tags.
[0041] Any suitable technique may be used to determine the reporter moiety, and the exact method may depend on the type of reporter. Examples include, but are not limited to, in situ hybridization methods such as single-molecule fluorescent in situ hybridization (smFISH), multiplexed FISH, CASFISH, or other techniques known to those skilled in the art. In one embodiment, smFISH is used to localize the reporter moiety, for example, within a cell.
[0042] Furthermore, as discussed herein, the location of the identifying moiety may also be determined, and it may be associated with or co-localized with a reporter moiety, which may be useful, for example, to reduce background noise and / or improve decoding accuracy. For example, the reporter moiety of a nucleic acid may emit a first signal (e.g., a first fluorescence), and the identifying moiety may emit a second signal (e.g., a second fluorescence, which may be of the same wavelength as the first fluorescence or a different wavelength), which may be associated with or co-localized with each other.
[0043] In some embodiments, the nucleic acids may include an identifying portion or "barcode" of nucleotides that can be used to distinguish nucleic acids from one another. The identifying portion may be located at any suitable position on the nucleic acid. For example, in one embodiment, the identifying portion may be located within the 3'UTR of the reporter gene.
[0044] In some cases, other sequences may be present within the identifying portion. For example, in some cases, the identifying portion may include a promoter or another regulatory sequence (e.g., an operon, a promoter such as a CMV promoter, a repressor, a transcription factor binding site, etc.). A promoter may drive transcription. In some embodiments, the promoter of the identifying portion may be the same as or different from the promoter of the guide portion.
[0045] In certain embodiments, e.g., at least 10, at least 10 2 , at least 10 3 , at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 Libraries of identifying portions can be used, containing unique sequences, such as 10, 20, 30, 40, or 50. In some cases, the identifying portion can be defined as a number of variable portions (or "bits"), e.g., within a sequence, although the unique sequences can all be determined individually (e.g., randomly). For example, the identifying portion can include at least 2, at least 3, at least 5, at least 7, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, or at least 50 variable portions. Each of the variable portions can include at least 2, at least 3, at least 4, at least 5, at least 7, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, or more possibilities.
[0046] Thus, for example, for each variable region, an identifying portion defined by 22 variable regions and 2 unique possibilities would be 2 22 As another non-limiting example, the identifying moiety would define a library of identifying moieties with 7 = 4,194,304 members. 10A library of identifying portions can be defined by 10 variable regions and 7 unique possibilities per variable region, so that each member defines a library of identifying portions. It should be understood that the variable portions can contain any suitable number of nucleotides, and different variable portions within an identifying portion can independently have the same or different number of nucleotides. Different variable regions can also have the same number of unique possibilities or different numbers of unique possibilities.
[0047] For example, the variable portion may be defined as having a length of at least 2, at least 3, at least 4, at least 5, at least 7, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50 or more nucleotides, and / or a length of no more than 50, no more than 40, no more than 30, no more than 25, no more than 20, no more than 15, no more than 10, no more than 7, no more than 5, no more than 4, no more than 3, or no more than 2 nucleotides. Combinations of these are also possible; for example, the variable portion may be between 5 and 50 nt, or between 15 and 25 nt, etc.
[0048] Each readout array position may be thought of as a "bit" (e.g., a 1 or a 0 in this example), although it should be understood that the number of possibilities for each "bit" is not necessarily limited to just two, as is the case in a computer. In other embodiments, rather than just two possibilities, there may be three possibilities (i.e., a "trit"), four possibilities (i.e., a "quad bit"), five possibilities, etc. For example, in the example below, a variety of trits are used. However, the use of bits (with any number of possibilities) to form the identifying portion may in some, but not all, embodiments enable the use of code words, error detection codes, error correction codes, etc., within the identifying portion, e.g., as discussed in detail herein.
[0049] In some cases, variable portions of an identifying portion can be linked together to form the identifying portion. However, in other cases, one or more variable portions can be separated, for example, by non-variable portions of nucleotides, to form the identifying portion. In addition, in some cases, some or all of the possible variable portions in the library can be unique, for example, to minimize errors. Any method can be used for linking. For example, portions can be linked together using ligation, overlap PCR, oligonucleotide pool synthesis, or other techniques known to those skilled in the art for connecting or linking nucleic acids together.
[0050] In certain embodiments, all members of a library are generated and / or used. However, in other embodiments, not all members of a library are necessarily generated and / or used. For example, in some embodiments, a smaller subset of a library may be used, e.g., to reduce or eliminate ambiguity or accidental reuse, e.g., less than 75%, 50%, 40%, 30%, 20%, 10%, 5%, 3%, 2%, 1%, 0.5%, 0.3%, or 0.1% of all possible members of a library are generated and / or used.
[0051] In some embodiments, the genotype of a cell can be determined, for example, using an identification moiety. To determine the genotype of a cell, various different techniques can be used, such as FISH, smFISH, MERFISH, in situ hybridization, multiplexed FISH, CASFISH, or other techniques known to those skilled in the art. In some embodiments, these techniques can involve direct hybridization with the identification moiety or the molecule generated from this moiety via cells. In certain cases, these techniques can also involve the binding of a separate adapter entity, which directly binds to the identification moiety or the molecule generated from it. Further non-limiting examples of techniques include those disclosed in U.S. Patent Application No. 15 / 329,683; or International Patent Application Publication No. WO2016 / 018960, each of which is incorporated herein by reference in its entirety.
[0052] In one set of embodiments, determining the genotype of a cell can be facilitated by determining the identifying portion of the nucleic acid in the cell. For example, a nucleic acid containing an identifying portion and a guide portion may have been introduced into the cell; the guide portion, as discussed above, may have caused a different phenotype, for example, by allowing editing of the target sequence to occur, for example, on the genome. However, it may also be important to know which nucleic acid has been introduced into which cell, thereby enabling understanding between the observed phenotype (e.g., altered phenotype) and the genotype that causes these phenotypes. As discussed herein, determining the identifying portion in the cell may determine the identity of the nucleic acid contained in each cell; thus, for example, if the nucleic acid contains an identifying portion and a guide portion on the same individual nucleic acid, the specific guide portion may also be determined.
[0053] By way of non-limiting example, in one set of embodiments, cells may be sequentially exposed to nucleic acid probes capable of binding to different portions of the identifying portion or a molecule, such as RNA, expressed by the cell from the identifying portion, e.g., a nucleic acid probe comprising a target sequence (e.g., capable of optionally specifically binding to at least a portion of the identifying portion), and a lead sequence (e.g., which may be partially "read" to determine binding), and binding of the nucleic acid probes within the cell may be determined. For example, cells may be exposed to a secondary nucleic acid probe that may contain a recognition sequence capable of binding to or hybridizing with the lead sequence and may contain a signal-generating entity. By determining the signal-generating entity in the images (and optionally inactivating the signal-generating entity between images and exposure to a different nucleic acid probe), the identifying portion of the cell may be determined.
[0054] As discussed herein, various nucleic acid probes can be used to determine one or more nucleic acids in a cell. The probe can include nucleic acids (or entities that can specifically hybridize with nucleic acids, for example) such as DNA, RNA, LNA (locked nucleic acid), PNA (peptide nucleic acid), or a combination thereof. In some cases, for example, as discussed below, additional components can also be present in the nucleic acid probe. In some embodiments, the nucleic acid probe can be made from other components, for example, proteins or other small molecules, and can represent a combination of these components with nucleic acids such as DNA, RNA, LNA, PNA, etc.
[0055] Nucleic acid probes can be introduced into cells using any suitable method. In some cases, cells can be sufficiently permeabilized so that the nucleic acid probe can be introduced into the cells by a fluid containing the nucleic acid probe in the vicinity of the cells. In some cases, cells can be sufficiently permeabilized as part of a fixation process, and in other embodiments, cells can be permeabilized by exposure to certain chemicals, such as ethanol, methanol, Triton, etc. Additionally, in some embodiments, techniques such as electroporation or microinjection can be used to introduce the nucleic acid probe into cells.
[0056] The determination of nucleic acid in a cell can be qualitative and / or quantitative. In addition, the determination can be spatial, for example, the location of nucleic acid in a cell can be determined in two or three dimensions. In some embodiments, the location, number, and / or concentration of nucleic acid in a cell can be determined.
[0057] As mentioned, in certain embodiments, the correlation or co-localization between the location of the reporter gene and the detection of the read sequence when reading the code word can substantially improve decoding accuracy, for example, due to reduced misidentification of background signals introduced by non-specific labeling. In some cases, for example, the readout portion of the code word can contain only one sequence for the readout, so that the readout signal is more difficult to identify, for example, compared to the background. For example, in certain cases, the reporter moiety can be determined, for example, locally or spatially, as discussed herein, and the portion of the identifying sequence can be determined as discussed herein. In some cases, the portion of the apparent identifying sequence that is not co-localized with the reporter moiety can be eliminated from further consideration. For example, the apparent identifying sequence can be an incorrect signal, background noise, or the like. In addition, in some cases, the reporter moiety can be determined between different determinations of the identifying sequence. Such an approach can improve accuracy by reducing errors due to sample movement, for example, stage drift. Thus, the association or co-localization between the location of the reporter gene and the detection of the lead sequence can be used to determine whether a tentative signal in the lead sequence is the lead sequence or background noise (and therefore not worthy of further consideration), etc.
[0058] It should be understood that the number of guide moieties and / or identifying moieties can have a relatively large number of possibilities (e.g., millions), which can be readily achieved by one of ordinary skill in the art using techniques such as computers and automated nucleic acid synthesizers (many of which are commercially available), as well as solid-phase synthesis, and / or isothermal assembly, and / or error-prone PCR, and / or ligating or otherwise assembling, for example, multiple variable regions in combinatorial, overlapping PCR. Correspondingly, a relatively large number of unique identifying moieties can be correlated with such a large number of possibilities for guide moieties, for example, through the use of a relatively small number of suitable variable regions and unique "bits" that can be created for each variable region. Thus, for example, at least 10, at least 10 2 , at least 10 3 , at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 A library of nucleic acids (e.g., each containing an identifying portion and a guide portion) can be prepared containing unique members such as
[0059] In certain embodiments, nucleic acids from a library of nucleic acids can be introduced into cells.Any suitable technique can be used to introduce nucleic acids.For example, in one set of embodiments, nucleic acids can be delivered into cells using viruses such as lentivirus, retrovirus, adenovirus, or adeno-associated virus.In some cases, the virus can transfect nucleic acids into cells or deliver them into the genome of cells, and in some cases, deliver them stably into the genome.
[0060] For example, in one set of embodiments, a lentiviral delivery system can be used to introduce nucleic acids into cells. Lentiviral systems can allow the number of nucleic acids introduced into cells to be controlled. For example, by controlling the titer of the lentivirus used for transduction, the number of library members delivered to each cell can be controlled to be one or more. In some embodiments, the guide portion and the identifying portion can be positioned adjacent to each other within the 3'LTR region of the lentivirus, i.e., after the lentiviral PPT (polypurine tract) sequence, so that the distance between the guide portion and the identifying portion is minimal, e.g., 100 bases or less for the non-variable region of the sgRNA for Cas9. In certain embodiments, the distance can also be less than 500 bases, less than 300 bases, less than 200 bases, less than 100 bases, less than 50 bases, less than 30 bases, or less than 10 bases. Such lentiviral constructs can reduce the genomic distance between the guide portion and the identifying portion. This may result in a reduction in recombination efficiency, which may allow for more accurate identification of the guide moiety by measuring the identified moiety. Those skilled in the art will be familiar with lentivirus and other virus-based delivery systems for introducing nucleic acids into cells. Many kits are readily available commercially that allow for such delivery of nucleic acids into cells using viruses.
[0061] In addition, in some embodiments, other techniques can be used to introduce nucleic acids into cells. For example, nucleic acids can be incorporated into a plasmid that can be taken up by cells. Other methods for introducing nucleic acids into cells include, but are not limited to, calcium phosphate (e.g., tricalcium phosphate), electroporation, cell squeezing, mixing cationic lipids with materials to create liposomes that fuse with the cell membrane, and the like. Further non-limiting examples of suitable methods include dendrimers, cationic polymers, lipofection, FuGENE, sonoporation, optical transfection, protoplast fusion, impale infection, gene guns, magnetofection, particle bombardment, viral infection, and the like.
[0062] In certain embodiments, nucleic acids may be introduced into cells, or cells may be transfected with nucleic acids, such that at least 50% of the cells have zero or only one nucleic acid introduced therein. In some cases, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 99%, etc. of the cells may have zero and / or only one nucleic acid introduced therein. This may be achieved, for example, using a lentivirus, such as the lentivirus discussed above, an appropriate dilution technique, a cell sorting technique, or through the use of other techniques, such as microfluidic dropping. In other cases, the percentage of transfected cells may be small, such as less than 50%, less than 20%, less than 10%, or less than 1%. In some embodiments, cells into which such nucleic acids have not been introduced may be removed. Non-limiting examples of cell removal include, for example, treatment with a chemical (such as an antibiotic) that kills or prevents non-transfected cells from dividing. In another example, some or all of the cells that do not contain the introduced nucleic acid can be removed from the sample using, for example, fluorescence activated cell sorting and / or other suitable cell sorting or microfluidics methods.
[0063] In certain embodiments, the identifying portion and the guide portion may be combined within a single source, e.g., nucleic acid contained within a single virus. In other embodiments, these portions may be delivered to the cell by separate sources, e.g., two different viral delivery vehicles. Other examples of introducing nucleic acids into cells are disclosed herein, and may involve the same or different methods of introduction.
[0064] The combination of identifying portion and guide portion, whether it is on the same medium, for example, a virus, or on different mediums, for example, viruses, can be, for example, randomly determined or deterministically determined.For example, a given CRISPR compilation can be assigned to a given barcode and expressed in cells.In some embodiments, the specific association of identifying portion and guide portion can be measured by any of various techniques.For example, PCR can be used to amplify the portion of nucleic acid that contains both identifying portion and guide portion, and then sequencing methods, including next-generation sequencing, can be used to identify which identifying region occurs with which guide portion through direct sequencing of this PCR product.Those skilled in the art will be familiar with other techniques that can be used to sequence nucleic acids that contain identifying portion and guide portion, for example.For example, any technique can be used for sequencing, such as Sanger sequencing, high-throughput sequencing, next-generation sequencing, nanopore sequencing, sequencing by ligation, sequencing by synthesis, etc. Those skilled in the art will be aware of different techniques for sequencing nucleic acids.
[0065] In certain embodiments, cells can be analyzed to determine their phenotype. Optionally, in some embodiments, the phenotype can be altered, for example, through the use of CRISPR or other techniques that may interact with the cell's genome, as discussed herein. Phenotypes can be determined using any suitable technique, such as optical techniques through analysis of cell behavior. Specific examples include, but are not limited to, microscopy or other optical methods, such as light microscopy, fluorescence microscopy, confocal microscopy, near-field microscopy, two-photon microscopy, or phase contrast microscopy, or other techniques described herein. Optionally, super-resolution methods, including any of the techniques described herein, can be used. In some embodiments, phenotypes can also be probed by other techniques, such as atomic force microscopy or patch clamping. Additionally, in some embodiments, phenotypes can be determined using proteins. For example, proteins can be determined using fluorescence, immunofluorescence, etc. Specific, non-limiting examples include fluorescent labeling, such as fluorescent proteins or organic dyes. Optionally, both microscopy and other techniques can be used in combination to determine phenotypes.
[0066] Examples of phenotypes that can be determined include, but are not limited to, cell shape (e.g., shape, size, visual appearance, organelles, subcompartments, state (e.g., in the cell cycle), etc.), certain characteristics of cell movement (e.g., speed, persistence, chemotactic behavior, etc.), certain characteristics of intercellular interactions (e.g., cell-to-cell adhesion, cell-to-cell avoidance, cell-to-cell interactions, etc.), or certain subcellular characteristics (e.g., protein or nucleic acid location, protein or nucleic acid diffusion, binding of two or more proteins and / or nucleic acids, etc.). Shape can include whole cell shape or subcompartment shape. In one embodiment, smFISH is used to determine the phenotype of a cell.
[0067] In some cases, the phenotype can be determined dynamically, for example, as a change in a cell over time.
[0068] In certain embodiments, the cells are present on a substrate suitable for cell culture and / or imaging. For example, the substrate can be glass, silicon, plastic (e.g., polystyrene, polypropylene, polycarbonate, etc.), etc. In some cases, at least a portion of the substrate can be at least partially optically transparent. The substrate can also be untreated or treated in a manner that facilitates cell attachment.
[0069] In some embodiments, the phenotype that can be determined includes all or at least a part of the transcriptome of a cell.Various techniques can be used to determine the transcriptome, including but not limited to smFISH, MERFISH, or other techniques such as those described herein.See also U.S. Patent Application No. 15 / 329,683; or International Patent Application Publication No. WO2016 / 018960, each of which is incorporated herein by reference in its entirety.In some cases, the transcriptome can be determined spatially within one or more cells.
[0070] In addition, in some cases, the phenotype that can be determined includes all or at least a portion of the chromosomes of the cell and / or agents, such as proteins or RNAs, that are bound to or otherwise associated with the chromosomes of the cell. For example, the concentration, spatial location, activity, association, etc., of the chromosomes and / or other associated agents can be determined according to certain embodiments of the present invention. In some cases, the chromosomes can be spatially determined within one or more cells. Non-limiting examples of techniques that can be used to determine the chromosomes include multiplexed DNA FISH or CASFISH. As yet another example, epigenetic modifications of the cell can also be determined.
[0071] Additionally, in some cases, the phenotype that can be determined includes all or at least a portion of the cell's proteome. Various techniques can be used to determine the proteome, including antibody labeling, sequential antibody labeling, multiplexed antibody imaging, or other multiplexed protein imaging methods. For example, the concentration, spatial location, activity, association, etc. of proteins and / or other associated agents can be determined.
[0072] In certain embodiments, one or more markers may be determined within a cell to determine the phenotype. For example, the marker may indicate a particular cellular protein, nucleic acid, morphological feature, etc., or the marker may indicate cellular behavior. In addition, the marker may optionally be a marker that can be determined visually. For example, the marker may be a fluorescent marker that alters the fluorescence of another fluorescent entity within the cell (e.g., via enhancement or quenching). In some embodiments, the marker may also be a dye that changes color. Thus, differences in intensity, wavelength, frequency, location, distribution, etc., between cells within an image may be determined to determine the phenotype of the cell. In some cases, other methods of determining the marker may also be used; for example, the marker may be a radioactive marker. Many such markers may be commercially available.
[0073] Furthermore, it should be understood that these measurements are not mutually exclusive. Any combination of these measurements can be performed in a single sample. Furthermore, in some embodiments, such measurements can be repeated, for example, on the same sample. For example, measurements can be repeated to ensure validity or reduce potential errors (e.g., measurement errors), and measurements can be repeated after exposure to various stimuli or conditions, such as treatment with different nutrient sources, small molecules, or other suitable agents that may interact with the cells.
[0074] In some cases, the phenotype of a cell can be changed, for example, by applying the guide moiety discussed above, which can be expressed by the cell to change its phenotype in some forms. For example, the guide moiety can be used to induce changes to the genome of a cell, for example, through CRISPR or other suitable techniques, including those described herein. As another example, a guide moiety encoding a protein can be added to a cell, and the cell can express the protein. When different proteins are encoded in different cells, the cells can exhibit different phenotypes, which can be determined as described above. Thus, for example, multiple cells can be transfected with multiple different guide moieties, or multiple different guide moieties can be otherwise introduced into multiple cells, and the cells can then be studied to determine the effects that the different guide moieties have on their phenotypes.
[0075] Thus, certain embodiments are generally directed to nucleic acid probes introduced into cells (or other samples). Depending on the application, the probe may comprise any of a variety of entities capable of hybridizing to a nucleic acid, e.g., a target site, such as DNA, RNA, LNA, PNA, etc., typically by Watson-Crick base pairing. The nucleic acid probe typically contains a target sequence capable of binding to at least a portion of the target, e.g., the target site. In some cases, the binding may be specific (e.g., via complementary binding). Once introduced into a cell or other system, the target sequence may be capable of binding to a specific target (e.g., mRNA or other nucleic acid discussed herein). The nucleic acid probe may also contain one or more lead sequences, discussed below.
[0076] In some cases, more than one type of nucleic acid probe may be applied to a sample, for example, sequentially or simultaneously. For example, at least 2, at least 5, at least 10, at least 25, at least 50, at least 75, at least 100, at least 300, at least 1,000, at least 3,000, at least 10,000, or at least 30,000 distinguishable nucleic acid probes may be applied to a sample. In some cases, nucleic acid probes may be added sequentially. However, in some cases, more than one nucleic acid probe may be added simultaneously.
[0077] Nucleic acid probes can include one or more target sequences, which can be located anywhere within the nucleic acid probe.Target sequences can contain a region that is substantially complementary to a target, for example, a portion of a target nucleic acid.For example, in some cases, a portion can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary, for example, to achieve specific binding.Typically, complementarity is determined based on Watson-Crick nucleotide base pairing.
[0078] In some cases, the target sequence can be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. In some cases, the target sequence may be no more than 500, no more than 450, no more than 400, no more than 350, no more than 300, no more than 250, no more than 200, no more than 175, no more than 150, no more than 125, no more than 100, no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, no more than 35, no more than 30, no more than 20, or no more than 10 nucleotides in length. Combinations of any of these may also be possible; for example, the target sequence may have a length of between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, between 10 and 300 nucleotides, etc.
[0079] For targets that are likely to exist in cells or other samples, the target sequence of a nucleic acid probe can be determined.For example, the target nucleic acid for a protein can be determined using the sequence of the protein, for example, by determining the nucleic acid that is expressed to form the protein.In some cases, only a portion of the nucleic acid that encodes the protein, for example, having the length discussed above, can be used.In addition, in some cases, more than one target sequence can be used to identify a specific target.For example, multiple probes can be used that can sequentially and / or simultaneously bind to the same or different regions of the same target, or sequentially and / or simultaneously hybridize with it.Hybridization typically refers to the annealing process in which complementary single-stranded nucleic acids associate through Watson-Crick nucleotide base pairing (for example, hydrogen bonds between guanine and cytosine and between adenine and thymine) to form double-stranded nucleic acid.
[0080] In some embodiments, the nucleic acid probe may also include one or more "lead" sequences, as discussed above. The lead sequence may be used to identify the nucleic acid probe, for example, through association with a signal-generating entity, as discussed below. In some embodiments, the nucleic acid probe may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more, 20 or more, 24 or more, 32 or more, 40 or more, 48 or more, 50 or more, 64 or more, 75 or more, 100 or more, or 128 or more lead sequences. The lead sequences may be located anywhere within the nucleic acid probe. When more than one lead sequence is present, the lead sequences may be located adjacent to each other and / or may be interrupted by other sequences.
[0081] The lead sequence can be any length. When more than one lead sequence is used, the lead sequences can be independently the same or different. For example, the lead sequence can be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. In some cases, the lead sequence may be no more than 500, no more than 450, no more than 400, no more than 350, no more than 300, no more than 250, no more than 200, no more than 175, no more than 150, no more than 125, no more than 100, no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, no more than 35, no more than 30, no more than 20, or no more than 10 nucleotides in length. Combinations of any of these may also be possible, for example, the lead sequence may have a length between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, between 10 and 300 nucleotides, etc.
[0082] In some embodiments, the lead sequence may be arbitrary or random. In certain cases, the lead sequence is selected to reduce or minimize homology with other components of a cell or other sample, for example, so that the lead sequence itself does not bind or hybridize with other nucleic acids likely to be present in the cell or other sample. In some cases, the homology may be less than 10%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, or less than 1%. In some cases, there may be less than 20 base pairs, less than 18 base pairs, less than 15 base pairs, less than 14 base pairs, less than 13 base pairs, less than 12 base pairs, less than 11 base pairs, or less than 10 base pairs of homology. In some cases, such base pairs are consecutive.
[0083] In one set of embodiments, the nucleic acid probe population may contain a certain number of lead sequences, which may in some cases be less than the number of nucleic acid probe targets. Those skilled in the art will appreciate that when there is one signal-generating entity and n lead sequences, generally there are 2 n It will be appreciated that up to 1 different nucleic acid target can be uniquely identified. However, not all possible combinations need to be used. For example, a nucleic acid probe population can target 12 different nucleic acid sequences but contain no more than 8 lead sequences. As another example, a nucleic acid population can target 140 different nucleic acid sequences but contain no more than 16 lead sequences. By using different combinations of lead sequences within each probe, different nucleic acid sequence targets can be individually identified. For example, each probe can contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, etc., or more lead sequences. In some cases, nucleic acid probe populations can each contain the same number of lead sequences, but in other cases, different numbers of lead sequences can be present on various probes.
[0084] By way of non-limiting example, a first nucleic acid probe may contain a first target sequence, a first lead sequence, and a second lead sequence, while a second, different nucleic acid probe may contain a second target sequence, the same first lead sequence, but not the second lead sequence, but a third lead sequence. Such probes may thus be distinguished by determining the various lead sequences present in or associated with a given probe or location, as discussed herein. For example, probes may be sequentially identified and encoded using "code words," as discussed below. Code words may also be subject to error detection and / or correction.
[0085] Additionally, in certain embodiments, nucleic acid probe populations (and their corresponding, complimentary sites on coded probes) can be made using only two or three of the four naturally occurring nucleotide bases, such as excluding all "G"s or excluding all "C"s within the probe population. In certain embodiments, sequences lacking "G"s or "C"s may form fewer secondary structures and contribute to more homogeneous and rapid hybridization. Thus, in some cases, nucleic acid probes can contain only A, T, and G; only A, T, and C; only A, C, and G; or only T, C, and G.
[0086] In one embodiment, the lead sequence on the nucleic acid probe may be capable of binding (e.g., specifically) to a recognition sequence on the corresponding primary amplified nucleic acid. Thus, when the nucleic acid probe recognizes a target in a biological sample, such as a DNA or RNA target, the primary amplified nucleic acid may also associate with the target via the nucleic acid probe through interaction, e.g., complementary binding, between the lead sequence of the nucleic acid probe and the corresponding recognition sequence on the primary amplified nucleic acid. For example, the recognition sequence may be capable of recognizing the target lead sequence but not substantially recognizing or substantially binding to other non-target lead sequences. The primary amplified nucleic acid may also include any of a variety of entities capable of hybridizing with nucleic acids, such as DNA, RNA, LNA, and / or PNA, depending on the application. For example, such entities may form part or all of the recognition sequence. Thus, the recognition sequence may recognize a nucleic acid sequence, such as DNA or RNA.
[0087] In some cases, the recognition sequence can be substantially complementary to the target lead sequence.In some cases, the sequence can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary.Typically, complementarity is determined based on Watson-Crick nucleotide base pairing.The structure of the target lead sequence can include the structures previously described.
[0088] In some cases, the recognition sequence can be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. In some cases, the recognition sequence may be no more than 500, no more than 450, no more than 400, no more than 350, no more than 300, no more than 250, no more than 200, no more than 175, no more than 150, no more than 125, no more than 100, no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, no more than 35, no more than 30, no more than 20, or no more than 10 nucleotides in length. Combinations of any of these may also be possible; for example, the recognition sequence may have a length of between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, between 10 and 300 nucleotides, etc.
[0089] In some embodiments, the primary amplified nucleic acid may also include one or more lead sequences capable of binding to the secondary amplified nucleic acid, as discussed below. For example, the primary amplified nucleic acid may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more, 20 or more, 32 or more, 40 or more, 50 or more, 64 or more, 75 or more, 100 or more, or 128 or more lead sequences. The lead sequences may be located anywhere within the primary amplified nucleic acid. When more than one lead sequence is present, the lead sequences may be located adjacent to each other and / or may be interrupted by other sequences. In one embodiment, the primary amplified nucleic acid includes a recognition sequence at a first end and multiple lead sequences at a second end.
[0090] In some cases, the lead sequence in the primary amplified nucleic acid can be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. In some cases, the lead sequence may be no more than 500, no more than 450, no more than 400, no more than 350, no more than 300, no more than 250, no more than 200, no more than 175, no more than 150, no more than 125, no more than 100, no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, no more than 35, no more than 30, no more than 20, or no more than 10 nucleotides in length. Combinations of any of these may also be possible, for example, the lead sequence may have a length of between 10 and 20 nucleotides, between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, between 10 and 300 nucleotides, etc.
[0091] Any number of lead sequences can be present in the primary amplified nucleic acid. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more lead sequences can be present in the primary amplified nucleic acid. When more than one lead sequence is present in the primary amplified nucleic acid, the lead sequences can be the same or different. In some cases, for example, the lead sequences can all be identical.
[0092] In some embodiments, the population of primary amplified nucleic acids may be made using only two or three of the four naturally occurring nucleotide bases, such as excluding all "G"s or excluding all "C"s in the nucleic acid population. In certain embodiments, sequences lacking "G"s or "C"s may form fewer secondary structures and contribute to more homogeneous and rapid hybridization. Thus, in some cases, the primary amplified nucleic acids may contain only A, T, and G; only A, T, and C; only A, C, and G; or only T, C, and G.
[0093] In some cases, more than one type of primary amplification nucleic acid may be applied to sample, for example, sequentially or simultaneously.For example, at least 2, at least 5, at least 10, at least 25, at least 50, at least 75, at least 100, at least 300, at least 1,000, at least 3,000, at least 10,000, or at least 30,000 distinct primary amplification nucleic acids may be applied to sample.In some cases, primary amplification nucleic acids may be added sequentially.However, in some cases, more than one primary amplification nucleic acid may be added simultaneously.
[0094] In one set of embodiments, the lead sequence on the primary amplified nucleic acid may be capable of binding (e.g., specifically) to a recognition sequence on the corresponding secondary amplified nucleic acid. Thus, if the nucleic acid probe recognizes a target in a biological sample, e.g., a DNA or RNA target, the secondary amplified nucleic acid may also associate with the target via the primary amplified nucleic acid through interaction, e.g., complementary binding, between the lead sequence of the primary amplified nucleic acid and the recognition sequence on the corresponding secondary amplified nucleic acid. For example, the recognition sequence on the secondary amplified nucleic acid may be capable of recognizing the lead sequence on the primary amplified nucleic acid, but may not substantially recognize or substantially bind to other non-target lead sequences. The secondary amplified nucleic acid may also include any of a variety of entities capable of hybridizing to nucleic acids, such as DNA, RNA, LNA, and / or PNA, depending on the application. For example, such entities may form part or all of the recognition sequence.
[0095] In some cases, the recognition sequence on the secondary amplified nucleic acid can be substantially complementary to the lead sequence on the primary amplified nucleic acid. In some cases, the sequence can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% complementary.
[0096] In some cases, the recognition sequence on the secondary amplified nucleic acid can be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. In some cases, the recognition sequence may be no more than 500, no more than 450, no more than 400, no more than 350, no more than 300, no more than 250, no more than 200, no more than 175, no more than 150, no more than 125, no more than 100, no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, no more than 35, no more than 30, no more than 20, or no more than 10 nucleotides in length. Combinations of any of these may also be possible; for example, the recognition sequence may have a length of between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, between 10 and 300 nucleotides, etc.
[0097] In some embodiments, the secondary amplified nucleic acid may also contain one or more lead sequences capable of binding to a signal-generating entity, as discussed herein. For example, the secondary amplified nucleic acid may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 or more, 20 or more, 32 or more, 40 or more, 50 or more, 64 or more, 75 or more, 100 or more, or 128 or more lead sequences capable of binding to a signal-generating entity. The lead sequences may be located anywhere within the secondary amplified nucleic acid. When more than one lead sequence is present, the lead sequences may be located adjacent to each other and / or may be interrupted by other sequences. In one embodiment, the secondary amplified nucleic acid contains a recognition sequence at a first end and multiple lead sequences at a second end. This structure may also be the same as or different from the structure of the primary amplified nucleic acid.
[0098] In some cases, the lead sequence in the secondary amplified nucleic acid can be at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 60, at least 65, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, or at least 450 nucleotides in length. In some cases, the lead sequence may be no more than 500, no more than 450, no more than 400, no more than 350, no more than 300, no more than 250, no more than 200, no more than 175, no more than 150, no more than 125, no more than 100, no more than 75, no more than 60, no more than 65, no more than 60, no more than 55, no more than 50, no more than 45, no more than 40, no more than 35, no more than 30, no more than 20, or no more than 10 nucleotides in length. Combinations of any of these may also be possible, for example, the lead sequence in the secondary amplified nucleic acid may have a length between 10 and 20 nucleotides, between 10 and 30 nucleotides, between 20 and 40 nucleotides, between 5 and 50 nucleotides, between 10 and 200 nucleotides, or between 25 and 35 nucleotides, between 10 and 300 nucleotides, etc.
[0099] Any number of read sequences may be present in the secondary amplified nucleic acid. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more read sequences may be present in the secondary amplified nucleic acid. When more than one read sequence is present in the secondary amplified nucleic acid, the read sequences may be the same or different. In some cases, for example, the read sequences may all be identical. In addition, the same or different numbers of read sequences may be present independently in the primary amplified nucleic acid and the secondary amplified nucleic acid.
[0100] In certain embodiments, the population of secondary amplified nucleic acids may be made using only two or three of the four naturally occurring nucleotide bases, such as excluding all "G"s or excluding all "C"s in the nucleic acid population. In certain embodiments, sequences lacking "G"s or "C"s may form fewer secondary structures and contribute to more homogeneous and rapid hybridization. Thus, in some cases, the secondary amplified nucleic acids may contain only A, T, and G; only A, T, and C; only A, C, and G; or only T, C, and G.
[0101] In some cases, more than one type of secondary amplification nucleic acid may be applied to a sample, for example, sequentially or simultaneously.For example, at least 2, at least 5, at least 10, at least 25, at least 50, at least 75, at least 100, at least 300, at least 1,000, at least 3,000, at least 10,000, or at least 30,000 distinct secondary amplification nucleic acids may be applied to a sample.In some cases, secondary amplification nucleic acids may be added sequentially.However, in some cases, more than one type of secondary amplification nucleic acid may be added simultaneously.
[0102] Additionally, in certain embodiments, this pattern is instead repeated, e.g., with a tertiary amplification nucleic acid, a quaternary amplification nucleic acid, etc., similar to the discussion above, before the signal-generating entity. Thus, the signal-generating entity can be attached to the final amplified nucleic acid. Thus, by way of non-limiting example, a coded nucleic acid probe can be attached to a target, which can be attached to a primary amplified nucleic acid, which can be attached to a secondary amplified nucleic acid, which can be attached to a tertiary amplified nucleic acid, which can be attached to a signal-generating entity; or a coded nucleic acid probe can be attached to a target, which can be attached to a primary amplified nucleic acid, which can be attached to a secondary amplified nucleic acid, which can be attached to a tertiary amplified nucleic acid, which can be attached to a quaternary amplified nucleic acid, which can be attached to a signal-generating entity, etc. Thus, in all embodiments, the final amplified nucleic acid does not necessarily have to be a secondary amplified nucleic acid.
[0103] In some aspects, cells may be immobilized or fixed to a substrate before determining the genotype, e.g., as discussed below. In some cases, immobilization or fixation of cells may occur after determining the phenotype. According to certain embodiments, this may be useful, for example, to correlate the phenotype of cells in an image with the subsequent genotype of the cells (e.g., determined as discussed below). In some embodiments, cells may also be fixed before measuring the phenotype and before measuring the genotype, rather than after measuring the phenotype.
[0104] Those skilled in the art will be aware of systems and methods for fixing or otherwise immobilizing cells on a substrate. By way of non-limiting example, cells may be fixed using chemicals such as formaldehyde, paraformaldehyde, glutaraldehyde, ethanol, methanol, acetone, acetic acid, and the like. In one embodiment, cells may be fixed using Hepes-glutamate buffer-mediated organic solvent (HOPE). See also U.S. Patent Application No. 62 / 419,033, incorporated herein by reference in its entirety.
[0105] Certain aspects of the present invention are directed to determining samples that may include cell cultures, cell suspensions, biological tissues, biopsies, organisms, etc. The sample may also be acellular, but may nevertheless optionally contain nucleic acids. If the sample contains cells, the cells may be human cells or any other suitable cells, such as mammalian cells, fish cells, insect cells, plant cells, etc. In some cases, more than one cell may be present.
[0106] In a sample, the targets to be determined may include nucleic acids, proteins, etc. The nucleic acids to be determined may include, for example, DNA (e.g., genomic DNA), RNA, or other nucleic acids present in the cell (or in other samples). The nucleic acids may be endogenous to the cell or may be added to the cell. For example, the nucleic acids may be viral or artificially created. In some cases, the nucleic acids to be determined may be expressed by the cell. In some embodiments, the nucleic acid is RNA. The RNA may be coding RNA and / or non-coding RNA. For example, the RNA may encode a protein. Non-limiting examples of RNA that may be studied in a cell include mRNA, siRNA, rRNA, miRNA, tRNA, lncRNA, snoRNA, snRNA, exRNA, piRNA, etc.
[0107] In some cases, a substantial portion of the nucleic acids in a cell can be studied. For example, in some cases, a sufficient amount of RNA present in a cell can be determined to provide a partial or complete cellular transcriptome. In some cases, at least four mRNAs in the cell are determined, and in some cases, at least 3, at least 4, at least 7, at least 8, at least 12, at least 14, at least 15, at least 16, at least 22, at least 30, at least 31, at least 32, at least 50, at least 63, at least 64, at least 72, at least 75, at least 100, at least 127, at least 128, at least 140, at least 255, at least 256, at least 300, at least 315, at least 325, at least 326, at least 327, at least 328, at least 329, at least 400, at least 401, at least 402, at least 403, at least 404, at least 405, at least 406, at least 407, at least 408, at least 409, at least 509, at least 510, at least 511, at least 512, at least 513, at least 514, at least 515, at least 516, at least 518, at least 519 ... At least 500, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 4,000, at least 5,000, at least 7,500, at least 10,000, at least 12,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000, or at least 100,000 mRNAs can be determined.
[0108] In some cases, the transcriptome of cell can be determined.It should be understood that transcriptome generally includes all RNA molecules produced in cell, not just mRNA.Therefore, for example, transcriptome can also include rRNA, tRNA, siRNA, etc. in certain cases.In some embodiments, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% of the transcriptome of cell can be determined.
[0109] In some embodiments, other targets to be determined may include targets linked to nucleic acids, proteins, etc. For example, in one set of embodiments, a binding entity capable of recognizing a target may be conjugated to a nucleic acid probe. The binding entity may be any entity that can recognize a target, for example, specifically or non-specifically. Non-limiting examples include enzymes, antibodies, receptors, complementary nucleic acid strands, aptamers, etc. For example, an antibody linked to an oligonucleotide may be used to determine the target. The target may be capable of binding to an antibody linked to an oligonucleotide, and the oligonucleotide may be determined as discussed herein.
[0110] The determination of a target, such as a nucleic acid, in a cell or other sample may be qualitative and / or quantitative. In addition, the determination may be spatial, for example, the location of a nucleic acid or other target in a cell or other sample may be determined in two or three dimensions. In some embodiments, the location, number, and / or concentration of a nucleic acid or other target in a cell or other sample may be determined.
[0111] In some cases, a substantial portion of the genome of a cell can be determined. The determined genome segments can be continuous or interrupted in the genome. For example, in some cases, at least four genome segments are determined in the cell, and in some cases, at least three, at least four, at least seven, at least eight, at least 12, at least 14, at least 15, at least 16, at least 22, at least 30, at least 31, at least 32, at least 50, at least 63, at least 64, at least 72, at least 75, at least 100, at least 127, at least 128, at least 140, at least 255, at least 25 6, at least 500, at least 1,000, at least 1,500, at least 2,000, at least 2,500, at least 3,000, at least 4,000, at least 5,000, at least 7,500, at least 10,000, at least 12,000, at least 15,000, at least 20,000, at least 25,000, at least 30,000, at least 40,000, at least 50,000, at least 75,000, or at least 100,000 genome segments may be determined.
[0112] In some cases, the entire genome of a cell can be determined. It should be understood that a genome generally encompasses all DNA molecules produced in a cell, not just chromosomal DNA. Thus, for example, a genome can also optionally include, for example, mitochondrial DNA, chloroplast DNA, plasmid DNA, etc., in addition to (or rather than) chromosomal DNA. In some embodiments, at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or 100% of the genome of a cell can be determined.
[0113] As discussed herein, according to certain embodiments, various nucleic acid probes can be used to determine one or more targets in cells or other samples. The probes can include nucleic acids (or entities that can specifically hybridize with nucleic acids, for example), such as DNA, RNA, LNA (locked nucleic acid), PNA (peptide nucleic acid), and / or combinations thereof. In some cases, additional components can also be present in the nucleic acid probe, for example, as discussed herein. In addition, any suitable method can be used to introduce the nucleic acid probe into cells.
[0114] Other components may also be present in the nucleic acid probe or the amplified nucleic acid. For example, in one set of embodiments, one or more primer sequences may be present, for example, to facilitate enzymatic amplification. Those skilled in the art will be familiar with primer sequences suitable for applications such as amplification (e.g., using PCR or other suitable techniques). Many such primer sequences are commercially available. Other examples of sequences that may be present in the primary nucleic acid probe include, but are not limited to, promoter sequences, operons, discriminator sequences, nonsense sequences, etc.
[0115] Typically, a primer is a single-stranded or partially double-stranded nucleic acid (e.g., DNA) that serves as a starting point for nucleic acid synthesis, allowing a polymerase enzyme, such as a nucleic acid polymerase, to extend the primer and replicate the complementary strand. A primer is complementary to and hybridizes with (e.g., is designed to be complementary to and hybridize with) a target nucleic acid. In some embodiments, a primer is a synthetic primer. In some embodiments, a primer is not naturally occurring. A primer typically has a length of 10 to 50 nucleotides. For example, a primer can have a length of 10 to 40, 10 to 30, 10 to 20, 25 to 50, 15 to 40, 15 to 30, 20 to 50, 20 to 40, or 20 to 30 nucleotides. In some embodiments, a primer has a length of 18 to 24 nucleotides.
[0116] In some embodiments, one or more signal-generating entities can be bound to recognition entities on the secondary amplified nucleic acid (or other final amplified nucleic acid). Non-limiting examples of signal-generating entities include, for example, fluorescent entities (fluorophores) or phosphorescent entities, as discussed below. The signal-generating entities can then be determined, for example, to determine the nucleic acid probe or target. In some cases, the determination can be spatial, for example, in two or three dimensions. Additionally, in some cases, the determination can be quantitative, for example, the amount or concentration of the signal-generating entity and / or target can be determined.
[0117] In one set of embodiments, the signal-generating entity may be conjugated to the secondary amplified nucleic acid (or other final amplified nucleic acid). The signal-generating entity may be conjugated to the secondary amplified nucleic acid (or other final amplified nucleic acid) before or after the secondary amplified nucleic acid associates with the target in the sample. For example, the signal-generating entity may be conjugated to the secondary amplified nucleic acid first, or after the secondary amplified nucleic acid has been applied to the sample. In some cases, the signal-generating entity is added and then reacted to conjugate it to the amplified nucleic acid.
[0118] In one set of embodiments, the signal-generating entity may be attached to the nucleotide sequence via a bond that can be cleaved to release the signal-generating entity. For example, after determining the distribution of nucleic acid probes in a sample, the signal-generating entity may be released or inactivated prior to another round of nucleic acid probes and / or amplified nucleic acids. Thus, in some embodiments, the bond may be a cleavable bond, such as a disulfide bond or a photocleavable bond. Examples of photocleavable bonds are discussed in detail herein. In some cases, such bonds may be cleaved upon exposure to, for example, a reducing agent or light (e.g., ultraviolet light). For further details, see below. Other examples of systems and methods for inactivating and / or removing signal-generating entities are discussed in detail herein.
[0119] In certain embodiments, the use of primary and secondary amplification nucleic acids implies that there is a maximum number of signal-generating entities that can be bound to a given nucleic acid probe. For example, there is a maximum number of primary amplification nucleic acids that can be bound to a nucleic acid probe, due to the maximum number of secondary amplification nucleic acids that can be bound to a finite number of primary amplification nucleic acids and / or the maximum number of primary amplification nucleic acids that can be bound to lead sequences on a finite number of nucleic acid probes. While each potential position need not actually be filled with a signal-generating entity, this structure implies the existence of a saturation limit of signal-generating entities beyond which any additional signal-generating entities that may by chance be present cannot associate with the nucleic acid probe or its target.
[0120] Accordingly, certain embodiments of the present invention are generally directed to systems and methods for amplifying a signal indicative of a nucleic acid probe or its target that is saturable, i.e., there is a saturation limit, which is an upper limit on how many signal-generating entities can associate with the nucleic acid probe or its target. Typically, this number is greater than 1. For example, the upper limit of signal-generating entities can be at least 2, at least 3, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, etc. In some cases, the upper limit may be less than 500, less than 400, less than 300, less than 250, less than 200, less than 175, less than 150, less than 125, less than 100, less than 75, less than 50, less than 40, less than 30, less than 25, less than 20, less than 15, less than 10, less than 5, etc. In some cases, the upper limit may be determined as the maximum number of signal-generating entities that can bind to secondary amplified nucleic acids multiplied by the maximum number of secondary amplified nucleic acids that can bind to primary amplified nucleic acids multiplied by the maximum number of primary amplified nucleic acids that can bind to nucleic acid probes that bind to targets. In contrast, techniques such as rolling circle amplification or hairpin unfolding allow for uncontrolled signal amplification, i.e., when sufficient reagents are present, amplification can continue without a predetermined end point or saturation limit. Therefore, such techniques do not have a theoretical upper limit for the number of signal-generating entities that can associate with nucleic acid probes or their targets.
[0121] However, it should be understood that the average number of signal-generating entities actually bound to the nucleic acid probe or its target need not actually be the same as that upper limit, i.e., the signal-generating entities may not actually fully saturate (although fully saturating is possible). For example, the saturating amount (or the number of signal-generating entities bound compared to the maximum number that can be bound) can be less than 97%, less than 95%, less than 90%, less than 85%, less than 80%, less than 75%, etc., and / or at least 50%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, etc. In some cases, allowing time for binding to occur and / or increasing the concentration of reagents can increase the saturating amount.
[0122] Due to a potential upper limit on the number of signal-generating entities actually bound to a nucleic acid probe or its target, binding events distributed, for example, spatially, in a sample may exhibit substantially uniform size and / or brightness, in contrast to uncontrolled amplification, such as the uncontrolled amplification discussed above. For example, due to the specific number of secondary amplified nucleic acids that may bind to a primary amplified nucleic acid, the secondary amplified nucleic acid may not be found beyond a fixed distance from the nucleic acid probe or its target, which may limit the "spot size," or diameter of fluorescence from the signal-generating entity, indicating binding.
[0123] In certain embodiments, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% of the binding events may exhibit substantially the same brightness, size (e.g., apparent diameter), color, etc., which may facilitate distinguishing the binding events from other events, such as non-specific binding, noise, etc.
[0124] Additionally, as previously discussed, certain embodiments of the present invention may use a code space that encodes multiple binding events, but may use error detection and / or correction to determine the binding of nucleic acid probes to their targets. In some cases, a population of nucleic acid probes may contain a specific "lead sequence" that can bind to a specific amplified nucleic acid, as discussed above, but the location of the nucleic acid probe or target may be determined within a specific code space using a signal-generating entity associated with the amplified nucleic acid in the sample, e.g., as discussed herein. See also International Patent Application Publication Nos. WO2016 / 018960 and WO2016 / 018963, each of which is incorporated herein by reference in its entirety. As mentioned, in some cases, a population of lead sequences within a nucleic acid probe may be combined in various combinations, e.g., as discussed herein, such that a relatively small number of lead sequences can be used to determine a relatively large number of different nucleic acid probes.
[0125] Thus, in some cases, each nucleic acid probe population may contain a certain number of lead sequences, some of which are shared among different nucleic acid probes, such that the nucleic acid probe population may contain a certain number of lead sequences. A nucleic acid probe population may have any suitable number of lead sequences. For example, a nucleic acid probe population may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, etc. lead sequences. In some embodiments, more than 20 lead sequences may also be possible. Additionally, in some cases, a nucleic acid probe population may have a total of 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 20 or more, 24 or more, 32 or more, 40 or more, 50 or more, 60 or more, 64 or more, 100 or more, 128 or more, etc., lead sequences out of the possible lead sequences present, although some or all of the probes may each contain more than one lead sequence, as discussed herein. Additionally, in some embodiments, a nucleic acid probe population may have no more than 100, no more than 80, no more than 64, no more than 60, no more than 50, no more than 40, no more than 32, no more than 24, no more than 20, no more than 16, no more than 15, no more than 14, no more than 13, no more than 12, no more than 11, no more than 10, no more than 9, no more than 8, no more than 7, no more than 6, no more than 5, no more than 4, no more than 3, or no more than 2 lead sequences present. Combinations of any of these may also be possible; for example, a nucleic acid probe population may include between 10 and 15 lead sequences in total.
[0126] As a non-limiting example of a combinatorial approach to identifying a relatively large number of nucleic acid probes from a relatively small number of lead sequences contained in the nucleic acid probes, in a population of six different types of nucleic acid probes, each containing one or more lead sequences, the total number of lead sequences in the population may not exceed four. In this example, for ease of explanation, four lead sequences are used, but it should be understood that in other embodiments, a large number of nucleic acid probes can be achieved using, for example, 5, 8, 10, 16, 32, or more lead sequences, or any other suitable number of lead sequences described herein depending on the application. For example, if each nucleic acid probe contains two different lead sequences, a maximum of six probes can be individually identified by using four such lead sequences (A, B, C, and D). In this example, the order of the lead sequences on the nucleic acid probes is not critical, i.e., "AB" and "BA" can be treated as synonymous (although in other embodiments, the order of the lead sequences may be critical, and "AB" and "BA" may not necessarily be synonymous). Similarly, if five lead sequences (A, B, C, D, and E) are used in a population of nucleic acid probes, then up to ten probes (e.g., AB, AC, AD, AE, BC, BD, BE, CD, CE, DE) can be individually identified. For example, assuming that the order of the lead sequences is not critical, one skilled in the art would know that for k lead sequences in a population, with n lead sequences on each probe, up to
[0127] [ka] It will be understood that in certain embodiments, more or less than this number of different probes may also be used, as not all probes need to have the same number of lead sequences, and not all combinations of lead sequences need to be used in every embodiment. In addition, it should also be understood that in some embodiments, the number of lead sequences on each probe need not be the same. For example, some probes may contain two lead sequences, while other probes may contain three lead sequences.
[0128] In some embodiments, the lead sequence and / or binding pattern of a nucleic acid probe in a sample can be used to define an error detection and / or error correction code, e.g., to reduce or prevent misidentification or errors of nucleic acids. Thus, for example, if binding is indicated (e.g., determined using a signal-generating entity), the location can be identified by a "1"; conversely, if binding is not indicated, the location can be identified by a "0" (or vice versa, as the case may be). Multiple rounds of binding determination, e.g., using different nucleic acid probes, can then be used to create a "code word," e.g., for this spatial location. In some embodiments, the code word can be subjected to error detection and / or correction. For example, the code word can be configured such that, for a given set of lead sequences or binding patterns of a nucleic acid probe, if no match is found, the match can be identified as an error, and error correction can be applied to determine the correct target for the nucleic acid probe. In some cases, for example, when each code word encodes a different nucleic acid, a code word may have fewer "letters" or positions than the total number of nucleic acids encoded by the code word.
[0129] Such error detection and / or error correction codes can take a variety of forms. A variety of such codes, such as Golay or Hamming codes, have already been developed in other contexts, such as the telecommunications industry. In one set of embodiments, the read sequences or binding patterns of the nucleic acid probes are assigned such that not all possible combinations are assigned.
[0130] For example, if four lead sequences are possible and a nucleic acid probe contains two lead sequences, a maximum of six nucleic acid probes can be identified, although the number of nucleic acid probes used can be less than six. Similarly, for k lead sequences in a population with n lead sequences on each nucleic acid probe,
[0131] [ka] Although different probes can be generated, the number of nucleic acid probes used is
[0132] [ka] In addition, they may be assigned randomly or in a specific manner that increases the ability to detect and / or correct errors.
[0133] As another example, when multiple rounds of nucleic acid probes are used, the number of rounds can be chosen arbitrarily. If, within each round, each target can have two possible outcomes, such as detection or non-detection, then for n rounds of probes, at most 2 n Although different targets are possible, the number of targets actually used is limited to 2 n For example, if within each round each target can have more than two possible outcomes, such as detection in different color channels, then for n rounds of probes, 2 nMore than (e.g., 3 n , 4 n , , , different targets may be possible. In some cases, the number of targets actually used may be any number less than this number. In addition, these may be assigned randomly or in a specific manner that increases the ability to detect and / or correct errors.
[0134] Code words can be used to define various code spaces. For example, in one set of embodiments, code words or nucleic acid probes can be assigned within the code space such that the assignments are separated by a Hamming distance, which measures the number of incorrect "reads" in a given pattern that will cause a nucleic acid probe to be misinterpreted as a different, legitimate nucleic acid probe. In certain cases, the Hamming distance can be at least 2, at least 3, at least 4, at least 5, at least 6, etc. Additionally, in one set of embodiments, the assignments can be formed as Hamming codes, e.g., a Hamming (7,4) code, a Hamming (15,11) code, a Hamming (31,26) code, a Hamming (63,57) code, a Hamming (127,120) code, etc. In another set of embodiments, the assignments may form a SECDED code, e.g., a SECDED(8,4) code, a SECDED(16,4) code, a SECDED(16,11) code, a SECDED(22,16) code, a SECDED(39,32) code, a SECDED(72,64) code, etc. In yet another set of embodiments, the assignments may form an extended binary Golay code, a full binary Golay code, or a ternary Golay code. In another set of embodiments, the assignments may represent a subset of possible values taken from any of the codes described above.
[0135] For example, an error detection code may be formed by limiting the number of code words used to less than 10%, less than 5%, less than 2%, less than 1%, less than 0.1%, less than 0.01%, or less than 0.001% of the total number of possible code words, so that an incorrect code word is unlikely to exist as another code word used. Thus, any code word detected that does not match a code word used is likely to be incorrect.
[0136] For example, an error-correcting code may be formed by encoding a target using only binary words containing a fixed or constant number of "1" bits (or "0" bits). For example, the code space may include only 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, etc. "1" bits (or "0" bits), e.g., all of the codes have the same number of "1" bits or "0" bits, etc. In another set of embodiments, the assignment may represent a subset of possible values taken from any of the codes described above for purposes of addressing asymmetric readout errors. For example, in some cases, a code in which the number of "1" bits may be fixed for all binary words used may eliminate bias measurements of words with different numbers of "1"s if the proportion of "0" bits measured as "1" or the proportion of "1" bits measured as "0" differs.
[0137] Thus, in some embodiments, once a code word is determined (e.g., as discussed herein), the code word can be compared to known nucleic acid code words. If a match is found, the nucleic acid target can be identified or determined. If no match is found, an error in reading the code word can be identified. Optionally, error correction can also be applied to determine the correct code word, thereby resulting in the correct identification of the nucleic acid target. Optionally, the code word can be selected such that, assuming only one error is present, only one possible correct code word is available, thereby allowing only one correct identification of the nucleic acid target. Optionally, this can also be generalized to larger code word intervals or Hamming distances; for example, the code word can be selected such that if two, three, or four errors (or, in some cases, more errors) are present, only one possible correct code word is available, thereby allowing only one correct identification of the nucleic acid target.
[0138] The error-correcting code may be a binary error-correcting code or may be based on other numbering systems, such as ternary or quaternary error-correcting codes. For example, in one set of embodiments, more than one type of signal-emitting entity may be used and assigned to different numbers within the error-correcting code. Thus, as a non-limiting example, a first signal-emitting entity (or, in some cases, more than one signal-emitting entity) may be assigned as "1," a second signal-emitting entity (or, in some cases, more than one signal-emitting entity) may be assigned as "2" (with "0" indicating the absence of a signal-emitting entity), and the code words may be distributed to define a ternary error-correcting code. Similarly, a third signal-emitting entity may additionally be assigned as "3," to create a quaternary error-correcting code, and so on. Non-limiting examples of such codes include Reed-Solomon erasure codes and generalizations thereof.
[0139] Additionally, in some embodiments, codes may also be selected through random selection of a subset of all possible codewords. For example, a random subset of binary codewords for a length-n code may be selected. In some cases, these codewords may be separated by a Hamming distance, i.e., the number of bits that must be inverted to transform one codeword into another, so that some of the codewords used maintain some degree of error robustness or error correction. In some embodiments, techniques such as next-generation sequencing may be used to determine the random subset of codewords to be used, and error robustness and error correction could be selectively applied to codewords that meet the constraints required for these properties.
[0140] As discussed herein, in certain embodiments, the signal-generating entity is determined, for example, by imaging to determine a nucleic acid probe and / or creating a code word. Examples of signal-generating entities include those discussed herein. In some cases, the signal-generating entity in a sample can be determined, for example, spatially, using various techniques. In some embodiments, the signal-generating entity can be fluorescent, and techniques for determining fluorescence in a sample, such as fluorescence microscopy or confocal microscopy, can be used to spatially identify the location of the signal-generating entity within a cell. In some cases, the location of the entity in a sample can be determined in two or even three dimensions. In addition, in some embodiments, more than one signal-generating entity can be determined simultaneously (e.g., signal-generating entities with different colors or emissions) and / or sequentially.
[0141] Additionally, in some embodiments, a confidence level for a target, e.g., a nucleic acid target, can be determined. For example, the confidence level can be determined using the ratio of the number of exact matches to the number of matches with one or more single-bit errors. In some cases, only matches with a confidence ratio above a certain value can be used. For example, in certain embodiments, a match can be accepted only if the confidence ratio for the match is greater than about 0.01, greater than about 0.03, greater than about 0.05, greater than about 0.1, greater than about 0.3, greater than about 0.5, greater than about 1, greater than about 3, greater than about 5, greater than about 10, greater than about 30, greater than about 50, greater than about 100, greater than about 300, greater than about 500, greater than about 1000, or any other suitable value. Additionally, in some embodiments, a match may be accepted only if the confidence ratio for the target exceeds the internal standard or false positive control by about 0.01, about 0.03, about 0.05, about 0.1, about 0.3, about 0.5, about 1, about 3, about 5, about 10, about 30, about 50, about 100, about 300, about 500, about 1000, or any other suitable value.
[0142] In some embodiments, the spatial location of an entity (and thus a nucleic acid probe that may be associated with the entity) can be determined at a relatively high resolution. For example, the location can be determined at a spatial resolution of greater than about 100 micrometers, greater than about 30 micrometers, greater than about 10 micrometers, greater than about 3 micrometers, greater than about 1 micrometer, greater than about 800 nm, greater than about 600 nm, greater than about 500 nm, greater than about 400 nm, greater than about 300 nm, greater than about 200 nm, greater than about 100 nm, greater than about 90 nm, greater than about 80 nm, greater than about 70 nm, greater than about 60 nm, greater than about 50 nm, greater than about 40 nm, greater than about 30 nm, greater than about 20 nm, or greater than about 10 nm, etc.
[0143] A variety of techniques exist that can optically determine or image the spatial location of an entity, for example, using fluorescence microscopy. In some embodiments, more than one color may be used. In some cases, spatial location may be determined at ultra-high resolution, or at a resolution that exceeds the wavelength or diffraction limit of light. Non-limiting examples include stochastic optical reconstruction microscopy (STORM), stimulated emission depletion microscopy (STED), near-field scanning optical microscopy (NSOM), 4Pi microscopy, structured illumination microscopy (SIM), spatially modulated illumination microscopy (SMI), reversible saturable optically linear fluorescence transition microscopy (RESOLFT), ground state depletion microscopy (GSD), saturated structured-illumination microscopy (SSIM), spectral precision distance microscopy (SPDM), photoactivated localization microscopy (PALM), fluorescence photoactivated localization microscopy (FPALM), 3D light microscopical nanosizing microscopy (LIMON), super-resolution optical fluctuation imaging (SOFI), and the like.See, for example, U.S. Patent No. 7,838,302 to Zhuang et al., entitled "Sub-Diffraction Limit Image Resolution and Other Imaging Techniques," issued on November 23, 2010; U.S. Patent No. 8,564,792 to Zhuang et al., entitled "Sub-Diffraction Limit Image Resolution in Three Dimensions," issued on October 22, 2013; or International Patent Application Publication No. WO2013 / 090360 to Zhuang et al., entitled "High Resolution Dual-Objective Microscopy," published on June 20, 2013, each of which is incorporated herein by reference in its entirety.
[0144] As an illustrative, non-limiting example, in one set of embodiments, a sample may be imaged with a high-numerical aperture, 100x magnification oil-immersion objective and light collected on an electron-multiplying CCD camera. In another example, a sample may be imaged with a high-numerical aperture, 40x magnification oil-immersion objective and light collected by a wide-field academic CMOS camera. In various non-limiting embodiments, different combinations of objectives and cameras may result in a single field of view corresponding to a field of view not exceeding 40x40 microns, 80x80 microns, 120x120 microns, 240x240 microns, 340x340 microns, or 500x500 microns, etc. Similarly, in some embodiments, a single camera pixel may correspond to a sample area not exceeding 80x80 nm, 120x120 nm, 160x160 nm, 240x240 nm, or 300x300 nm, etc. In another example, the sample can be imaged with a low numerical aperture, 10x magnification air lens and light collected by an sCMOS camera. In a further embodiment, the sample can be optically resolved by focusing light through a single or multiple pinholes and illuminating it through a diffraction-limited focused beam generated by a scanning mirror or a rotating disk in one or more scans. In another embodiment, the sample can also be illuminated through a slab of light generated through any one of several methods known to those skilled in the art.
[0145] In one embodiment, the sample can be illuminated with a single Gaussian mode laser beam. In some embodiments, the illumination profile can be flattened by passing these laser beams through a multimode fiber vibrated via piezoelectric or other mechanical means. In some embodiments, the illumination profile can be flattened by passing the single-mode Gaussian beam through various refractive beam shapers, such as a π-shaper, or a series of stacked Powell lenses. In yet another set of embodiments, the Gaussian beam can be passed through a variety of different diffusing elements, such as ground glass or an optical diffuser, which can optionally be spun at high speed to remove residual laser speckle. In yet another embodiment, the laser illumination can be passed through a series of lenslet arrays to produce overlapping illumination images that approximate a planar illumination field.
[0146] In some embodiments, the centroid of the spatial location of the entity can be determined. For example, the centroid of the signal-emitting entity can be determined within an image or within a series of images using image analysis algorithms known to those skilled in the art. In some cases, the algorithm can be selected to determine non-overlapping single emitters and / or partially overlapping single emitters in the sample. Non-limiting examples of suitable techniques include maximum likelihood algorithms, least squares algorithms, Bayesian algorithms, compressed sensing algorithms, etc. In some cases, combinations of these techniques can also be used.
[0147] Additionally, in some cases, the signal-generating entity can be inactivated. For example, in some embodiments, a first secondary nucleic acid probe that can associate with a signal-generating entity (e.g., using an amplified nucleic acid) and that can recognize a first lead sequence (e.g., on a nucleic acid probe) can be applied to a sample, and then the signal-generating entity can be inactivated, for example, before a second secondary nucleic acid probe that can associate with the signal-generating entity (e.g., using an amplified nucleic acid) is applied to the sample. When multiple signal-generating entities are used, the same or different techniques can be used to inactivate the signal-generating entities, and some or all of the multiple signal-generating entities can be inactivated, for example, sequentially or simultaneously.
[0148] Inactivation may be caused by removal of the signal-generating entity (e.g., from the sample or from the nucleic acid probe, etc.) and / or by chemically altering the signal-generating entity in some way (e.g., by photobleaching the signal-generating entity, by photobleaching the signal-generating entity, by chemically altering the structure of the signal-generating entity, e.g., by reduction, etc.). For example, in one set of embodiments, fluorescent signal-generating entities may be inactivated by chemical or optical techniques such as oxidation, photobleaching, chemically bleaching, rigorous washing, or reaction by digestion or exposure to enzymes, dissociating the signal-generating entity from other components (e.g., the probe), chemical reaction of the signal-generating entity (e.g., a reactant capable of altering the structure of the signal-generating entity), etc. For example, bleaching may occur by exposure to oxygen, a reducing agent, or the signal-generating entity may be chemically cleaved from the nucleic acid probe (e.g., using tris(2-carboxyethyl)phosphine) and washed away via fluid flow.
[0149] In some embodiments, for example, using the amplified nucleic acids discussed herein, multiple nucleic acid probes (including primary and / or secondary nucleic acid probes) can be associated with one or more signal-generating entities. When more than one nucleic acid probe is used, the signal-generating entities can be the same or different. In certain embodiments, the signal-generating entity is any entity capable of emitting light. For example, in one embodiment, the signal-generating entity is a fluorescent entity. In other embodiments, the signal-generating entity can be a phosphorescent entity, a radioactive entity, a light-absorbing entity, or the like. In some cases, the signal-generating entity is any entity that can be determined in a sample at a relatively high resolution, for example, at a resolution exceeding the wavelength or diffraction limit of visible light. The signal-generating entity can be, for example, a dye, a small molecule, a peptide, or a protein. In some cases, the signal-generating entity can be a single molecule. When multiple secondary nucleic acid probes are used, the nucleic acid probes can be associated with or contain the same or different signal-generating entities.
[0150] Non-limiting examples of signal-generating entities include fluorescent entities (fluorophores) or phosphorescent entities, such as cyanine dyes (e.g., Cy2, Cy3, Cy3B, Cy5, Cy5.5, Cy7, etc.), Alexa Fluor dyes, Atto dyes, photoswitchable dyes, photoactivatable dyes, fluorescent dyes, metal nanoparticles, semiconductor nanoparticles, or "quantum dots," fluorescent proteins such as GFP (green fluorescent protein), or photoactivatable fluorescent proteins such as PAGFP, PSCFP, PSCFP2, Dendra, Dendra2, EosFP, tdEos, mEos2, mEos3, PAmCherry, PAtagRFP, mMaple, mMaple2, and mMaple3. Other suitable signal-generating entities are known to those skilled in the art. See, for example, U.S. Patent No. 7,838,302 or International Patent Application Publication No. WO 2015 / 160690, each of which is incorporated herein by reference in its entirety.
[0151] In one set of embodiments, the signal-generating entity may be joined to the oligonucleotide sequence via a bond that may be cleaved to release the signal-generating entity. In one set of embodiments, the fluorophore may be conjugated to the oligonucleotide via a cleavable bond, such as a photocleavable bond. Non-limiting examples of photocleavable bonds include 1-(2-nitrophenyl)ethyl, 2-nitrobenzyl, biotin phosphoramidite, acrylphosphoramidite, diethylaminocoumarin, 1-(4,5-dimethoxy-2-nitrophenyl)ethyl, cyclododecyl(dimethoxy-2-nitrophenyl)ethyl, 4-aminomethyl-3-nitrobenzyl, (4-nitro-3-(1-chlorocarbonyloxyethyl)phenyl)methyl-S-acetylthioate, [4-nitro-3-(1-chlorocarbonyloxyethyl)phenyl]methyl-3-(2-pyridyldithiopropionic acid) ester [(4-nitro-3-(1-thlorocarbonyloxyethyl)phenyl)methyl-3-(2-pyridyldithiopropionic acid) ester]. acid)ester], 3-(4,4'-dimethoxytrityl)-1-(2-nitrophenyl)-propane-1,3-diol-[2-cyanoethyl-(N,N-diisopropyl)]-phosphoramidite, 1-[2-nitro-5-(6-trifluoroacetylcaproamidomethyl)phenyl]-ethyl-[2-cyano-ethyl-(N,N-diisopropyl)]-phosphoramidite, 1-[2-nitro-5-(6-(4,4'-dimethoxytrityloxy)butylamidomethyl)phenyl]-ethyl-[2-cyanoethyl-(N,N-diisopropyl)]-phosphoramidite, 1-[2-nitro-5-(6-(N-(4,4'-dimethoxytrityl))-biotinamidocaproamido-methyl)phenyl]-ethyl-[2-cyanoethyl-(N,N-diisopropyl)]-phosphoramidite, or similar linkers. The oligonucleotide sequences can be, for example, primary or secondary (or other) amplified nucleic acids, such as the amplified nucleic acids discussed herein.
[0152] In another set of embodiments, the fluorophore may be conjugated to the oligonucleotide via a disulfide bond. Disulfide bonds can be cleaved by a variety of reducing agents, such as, but not limited to, dithiothreitol, dithioerythritol, beta-mercaptoethanol, sodium borohydride, thioredoxin, glutaredoxin, trypsinogen, hydrazine, diisobutylaluminum hydride, oxalic acid, formic acid, ascorbic acid, phosphoric acid, stannous chloride, glutathione, thioglycolate, 2,3-dimercaptopropanol, 2-mercaptoethylamine, 2-aminoethanol, tris(2-carboxyethyl)phosphine, bis(2-mercaptoethyl)sulfone, N,N'-dimethyl-N,N'-bis(mercaptoacetyl)hydrazine, 3-mercaptopropionate, dimethylformamide, thiopropyl agarose, tri-n-butylphosphine, cysteine, ferrous sulfate, sodium sulfite, phosphite, hypophosphite, phosphorothioates, and the like, and / or any combination thereof. The oligonucleotide sequences can be, for example, primary or secondary (or other) amplified nucleic acids, such as the amplified nucleic acids discussed herein.
[0153] In another embodiment, a fluorophore can be conjugated to an oligonucleotide via one or more phosphorothioate-modified nucleotides, in which sulfur modifications replace bridging and / or non-bridging oxygens. In certain embodiments, a fluorophore can be cleaved from an oligonucleotide via the addition of compounds such as, but not limited to, iodoethanol (iodine mixed in ethanol), silver nitrate, or mercuric chloride. In yet another set of embodiments, a signal-generating entity can be chemically inactivated via reduction or oxidation. For example, in one embodiment, a chromophore such as Cy5 or Cy7 can be reduced to a stable, non-fluorescent state using sodium borohydride. In yet another set of embodiments, a fluorophore can be conjugated to an oligonucleotide via an azo bond, which can be cleaved with 2-[(2-N-arylamino)phenylazo]pyridine. In yet another set of embodiments, the fluorophore can be conjugated to the oligonucleotide via a suitable nucleic acid segment that can be cleaved upon appropriate exposure to a DNase, such as an exodeoxyribonuclease or an endodeoxyribonuclease. Examples include, but are not limited to, DNase I or DNase II. In one set of embodiments, cleavage can occur via a restriction endonuclease. Non-limiting examples of potentially suitable restriction endonucleases include BamHI, BsrI, NotI, XmaI, PspAI, DpnI, MboI, MnlI, Eco57I, Ksp632I, DraIII, AhaII, SmaI, MluI, HpaI, ApaI, BclI, BstEII, TaqI, EcoRI, SacI, HindII, HaeII, DraII, Tsp509I, Sau3AI, PacI, and the like. Over 3000 restriction enzymes have been extensively studied, and of these, over 600 are commercially available. In yet another set of embodiments, the fluorophore may be conjugated to biotin, and the oligonucleotide may be conjugated to avidin or streptavidin.While the interaction of biotin with avidin or streptavidin conjugates the fluorophore to the oligonucleotide, upon sufficient exposure to excess, free biotin may "overcome" the ligation, thereby causing cleavage. Additionally, in another set of embodiments, the probe may be removed using a corresponding "toe-hold probe" that contains the same sequence as the probe as well as an extra number of bases (e.g., 1 to 20 extra bases, e.g., 5 extra bases) that have homology to the coded probe. These probes may remove the labeled readout probe through strand displacement interactions. The oligonucleotide sequence may be, for example, a primary or secondary (or other) amplified nucleic acid, such as the amplified nucleic acids discussed herein.
[0154] As used herein, the term "light" generally refers to electromagnetic radiation having any suitable wavelength (or, equivalently, frequency). For example, in some embodiments, light can include wavelengths in the optical or visible range (e.g., light having a wavelength between about 400 nm and about 700 nm, i.e., "visible light"), infrared wavelengths (e.g., light having a wavelength between about 300 micrometers and 700 nm), ultraviolet wavelengths (e.g., light having a wavelength between about 400 nm and about 10 nm), etc. In certain cases, as discussed in detail below, more than one entity can be used, i.e., chemically distinct or significantly, e.g., structurally distinct entities. However, in other cases, the entities can be chemically identical, or at least substantially chemically identical.
[0155] In one set of embodiments, the signal-generating entity is "switchable," i.e., the entity can be switched between two or more states, at least one of which emits light having a desired wavelength. In the other state(s), the entity may not emit light or may emit light of a different wavelength. For example, the entity can be "activated" to a first state capable of providing light having a desired wavelength, and "inactivated" to a second state incapable of emitting light of the same wavelength. An entity is "photoactivatable" when activated by incident light of the appropriate wavelength. As a non-limiting example, Cy5 can be switched between a fluorescent state and a non-luminescent state in a controlled and reversible manner by light of different wavelengths; i.e., red light at 633 nm (or 642 nm, 647 nm, 656 nm) can switch Cy5 to a stable non-luminescent state or inactivate it, while green light at 405 nm can switch Cy5 to a fluorescent state or activate it back. In some cases, an entity can be reversibly switched between two or more states, for example, upon exposure to an appropriate stimulus. For example, a first stimulus (e.g., light of a first wavelength) can be used to activate the switchable entity, while a second stimulus (e.g., light of a second wavelength) can be used to inactivate the switchable entity, for example, to a non-luminescent state. Any suitable method can be used to activate the entity. For example, in one embodiment, incident light of an appropriate wavelength can be used to activate the entity, causing it to emit light, i.e., the entity is "photoswitchable." Thus, a photoswitchable entity can be switched between different emitting or non-emitting states, for example, by incident light of different wavelengths. The light can be monochromatic (e.g., provided using a laser) or polychromatic. In another embodiment, the entity can be activated when stimulated with an electric and / or magnetic field. In other embodiments, the entity can be activated when exposed to an appropriate chemical environment, for example, by adjusting the pH or inducing a reversible chemical reaction involving the entity.Similarly, any suitable method may be used to inactivate an entity, and the method of activating an entity need not be the same as the method of inactivating an entity, for example, an entity may be inactivated upon exposure to incident light of an appropriate wavelength, or an entity may be inactivated by waiting for a sufficient period of time.
[0156] Typically, a "switchable" entity can be identified by one skilled in the art by determining the conditions under which the entity in a first state will emit light when exposed to an excitation wavelength, switching the entity from the first state to a second state, e.g., by exposing it to light of a switched wavelength, and then demonstrating that when the entity is in the second state, it no longer emits light (or emits light of a reduced intensity) when exposed to the excitation wavelength.
[0157] As discussed, in one set of embodiments, the switchable entity may be switched upon exposure to light. In some cases, the light used to activate the switchable entity may come from an external light source, such as a laser light source, another light-emitting entity in close proximity to the switchable entity, etc. In some cases, the second light-emitting entity may be a fluorescent entity, and in certain embodiments, the second light-emitting entity may also itself be a switchable entity.
[0158] In some embodiments, the switchable entity comprises a first light-emitting moiety (e.g., a fluorophore) and a second moiety that activates or "switches" the first moiety. For example, upon exposure to light, the second moiety of the switchable entity may activate the first moiety, causing the first moiety to emit light. Examples of activator moieties include, but are not limited to, Alexa Fluor 405 (Invitrogen), Alexa Fluor 488 (Invitrogen), Cy2 (GE Healthcare), Cy3 (GE Healthcare), Cy3B (GE Healthcare), Cy3.5 (GE Healthcare), or other suitable dyes. Examples of light-emitting moieties include, but are not limited to, Cy5, Cy5.5 (GE Healthcare), Cy7 (GE Healthcare), Alexa Fluor 647 (Invitrogen), Alexa Fluor 680 (Invitrogen), Alexa Fluor 700 (Invitrogen), Alexa Fluor 750 (Invitrogen), Alexa Fluor 790 (Invitrogen), DiD, DiR, YOYO-3 (Invitrogen), YO-PRO-3 (Invitrogen), TOT-3 (Invitrogen), TO-PRO-3 (Invitrogen), or other suitable dyes.These may be linked together, for example, covalently, e.g., directly, or via a linker, and examples thereof include Cy5-Alexa Fluor 405, Cy5-Alexa Fluor 488, Cy5-Cy2, Cy5-Cy3, Cy5-Cy3.5, Cy5.5-Alexa Fluor 405, Cy5.5-Alexa Fluor 488, Cy5.5-Cy2, Cy5.5-Cy3, Cy5.5-Cy3.5, Cy7-Alexa Fluor 405, Cy7-Alexa Fluor 488, Cy7-Cy2, Cy7-Cy3, Cy7-Cy3.5, Alexa Fluor 647-Alexa Fluor 405, Alexa Fluor 647-Alexa Fluor 488, Alexa Fluor 647-Cy2, Alexa Fluor The activator moiety can be linked to form compounds such as, but not limited to, Alexa Fluor 647-Cy3, Alexa Fluor 647-Cy3.5, Alexa Fluor 750-Alexa Fluor 405, Alexa Fluor 750-Alexa Fluor 488, Alexa Fluor 750-Cy2, Alexa Fluor 750-Cy3, or Alexa Fluor 750-Cy3.5. Those skilled in the art will be familiar with the structures of these and other compounds, many of which are commercially available. The moieties may be linked via a covalent bond or by a linker, such as those described in detail below. Other luminescent or activator moieties may include moieties having two quaternized nitrogen atoms connected by a polymethine chain, where each nitrogen is independently part of a heteroaromatic moiety, such as pyrrole, imidazole, thiazole, pyridine, quinoline, indole, or benzothiazole, or part of a non-aromatic amine. In some cases, there may be 5, 6, 7, 8, 9, or more carbon atoms between the two nitrogen atoms.
[0159] In certain cases, when the light-emitting moiety and the activator moiety are separated from each other, each can be a fluorophore, i.e., an entity that can emit light of a specific emission wavelength when exposed to a stimulus, such as an excitation wavelength. However, when a switchable entity is formed that includes a first fluorophore and a second fluorophore, the first fluorophore forms a first light-emitting moiety, and the second fluorophore forms an activator moiety that activates or "switches" the first moiety in response to a stimulus. For example, the switchable entity can include a first fluorophore directly bonded to a second fluorophore, or the first and second entities can be connected via a linker or a common entity. Whether a pair of a light-emitting moiety and an activator moiety results in a suitable switchable entity can be determined by methods known to those skilled in the art. For example, light of various wavelengths may be used to stimulate the pair and the light emission from the light emitting moiety may be determined to determine whether the pair results in an appropriate switch.
[0160] As a non-limiting example, Cy3 and Cy5 can be linked together to form such an entity. In this example, Cy3 is an activator moiety capable of activating Cy5, the light-emitting moiety. Thus, light at or near the absorption maximum of the activating or second portion of the entity (e.g., light near 532 nm for Cy3) can cause this portion to activate the first light-emitting moiety, thereby causing the first portion to emit light (e.g., near 647 nm for Cy5). See, e.g., U.S. Pat. No. 7,838,302, incorporated herein by reference in its entirety. Optionally, the first light-emitting moiety can then be inactivated by any suitable technique (e.g., by directing 647 nm red light toward the Cy5 portion of the molecule).
[0161] Other non-limiting examples of potentially suitable activator moieties include 1,5 IAEDANS, 1,8-ANS, 4-methylumbelliferone, 5-carboxy-2,7-dichlorofluorescein, 5-carboxyfluorescein (5-FAM), 5-carboxynaphthofluorescein, 5-carboxytetramethylrhodamine (5-TAMRA), 5-FAM (5-carboxyfluorescein), 5-HAT (hydroxytryptamine), 5-hydroxytryptamine (HAT), 5-ROX (carboxy-X-rhodamine), 5-TAMRA (5-carboxytetramethylrhodamine), 6-carboxyrhodamine 6G, 6-CR 6G, 6-JOE, 7-amino-4-methylcoumarin, 7-aminoactinomycin D (7-AAD), 7-hydroxy-4-methylcoumarin, 9-amino-6-chloro-2-methoxyacridine, ABQ, acid fuchsin, ACMA (9-amino-6-chloro-2-methoxyacridine), acridine orange, acridine red, acridine yellow, acriflavine, acriflavine feulgen SITSA, Alexa Fluor 350, Alexa Fluor 405, Alexa Fluor 430, Alexa Fluor 488, Alexa Fluor 500, Alexa Fluor 514, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 610, Alexa Fluor 633, Alexa Fluor 635, Alizarin Complexone, Alizarin Red, AMC, AMCA-S, AMCA (Aminomethylcoumarin), AMCA-X, Aminoactinomycin D, Aminocoumarin, Aminomethylcoumarin (AMCA), Aniline Blue, Anthrosyl Stearate, APTRA-BTC, APTS, Astrazon Brilliant Red 4G, Astrazon Orange R, Astrazon Red 6B, Astrazon Yellow 7 GLL, Atabrine, ATTO 390, ATTO 425, ATTO 465, ATTO 488, ATTO 495, ATTO 520, ATTO 532, ATTO 550, ATTO 565, ATTO590, ATTO 594, ATTO 610, ATTO 611X, ATTO 620, ATTO 633, ATTO 635, ATTO 647, ATTO 647N, ATTO 655, ATTO 680, ATTO 700, ATTO 725, ATTO 740, ATTO-TAG CBQCA, ATTO-TAG FQ, Auramine, Aurophosphine G, Aurophosphine, BAO9 (Bisaminophenyloxadiazole), BCECF (High pH), BCECF (Low pH), Berberine Sulfate, Bimane, Bisbenzamide, Bisbenzimide (Hoechst), Bis-BTC, Blancophor FFG, Blancophor SV, BOBO-1, BOBO-3, Bodipy 492 / 515, Bodipy 493 / 503, Bodipy 500 / 510, Bodipy 505 / 515, Bodipy 530 / 550, Bodipy 542 / 563, Bodipy 558 / 568, Bodipy 564 / 570, Bodipy 576 / 589, Bodipy 581 / 591, Bodipy 630 / 650-X, Bodipy 650 / 665-X, Bodipy 665 / 676, Bodipy Fl, Bodipy FL ATP, Bodipy Fl-ceramide, Bodipy R6G, Bodipy TMR, Bodipy TMR-X conjugate, Bodipy TMR-X, SE, Bodipy TR, Bodipy TR ATP, Bodipy TR-X SE, BO-PRO-1, BO-PRO-3, Brilliant SulphoflavinFF, BTC, BTC-5N, Calcein, Calcein Blue, Calcium Crimson, Calcium Green, Calcium Green-1 Ca 2+ Pigment, Calcium Green-2 Ca 2+ , Calcium Green-5N Ca 2+ , Calcium Green-C18 Ca 2+, Calcium Orange, Calcofluor White, Carboxy-X-rhodamine (5-ROX), Cascade Blue, Cascade Yellow, Catecholamine, CCF2 (GeneBlazer), CFDA, Chromomycin A, Chromomycin A, CL-NERF, CMFDA, Coumarin phalloidin, CPM methylcoumarin, CTC, CTC formazan, Cy2, Cy3.18, Cy3.5, Cy3, Cy5.18, Cyclic AMP fluorosensor (FiCRhR), Dabcyl, Dansyl, Dansylamine, Dansylcadaverine, Dansyl chloride, DansylDHPE, Dansyl fluoride, DAPI, Dapoxyl, Dapoxyl 2, Dapoxyl 3'DCFDA, DCFH (dichlorodihydrofluorescein diacetate), DDAO, DHR (dihydrorhodamine 123), di-4-ANEPPS, di-8-ANEPPS [non-ratio], DiA (4-di-16-ASP), dichlorodihydrofluorescein diacetate (DCFH), DiD (lipophilic tracer), DiD [DiIC18(5)], DIDS, dihydrorhodamine 123 (DHR), DiI [DiIC18(3)], dinitrophenol, DiO [DiOC18(3)], DiR, DiR [DiIC18(7)], DM-NERF (high pH), DNP, dopamine, DTAF, DY-630-NHS, DY-635-NHS, DyLight 405, DyLight 488, DyLight 549, DyLight 633, DyLight 649, DyLight 680, DyLight 800, ELF97, Eosin, Erythrosin, Erythrosin ITC, Ethidium Bromide, Ethidium Homodimer 1 (EthD-1), Euchrisin, EukoLight, Europium(III) Chloride, Fast Blue, FDA, Feulgen (pararosaniline), FIF [Formaldehyde-Induced Fluorescence], FITC, FlazoOrange, Fluo-3, Fluo-4, Fluorescein (FITC), Fluorescein diacetate, Fluoro-Emerald, Fluoro-Gold (hydroxystilbamidine), Fluor-Ruby, FluorX, FM1-43, FM4-46, Fura Red (high pH), Fura Red / Fluo-3, Fura-2, Fura-2 / BCECF, Genacryl Brilliant Red B, Genacryl Brilliant Yellow 10GF, Genacryl Pink 3G, Genacryl Yellow 5GF, GeneBlazer (CCF2), Gloxalic Acid, Granular blue, Hematoporphyrin, Hoechst 33258, Hoechst 33342, Hoechst 34580, HPTS, Hydroxycoumarin, Hydroxystilbamidine (FluoroGold), Hydroxytryptamine, Indo-1 (High Calcium), Indo-1 (Low Calcium), Indodicarbocyanine (DiD), Indotricarbocyanine (DiR), Intrawhite Cf, JC-1, JO-JO-1, JO-PRO-1, LaserPro, Laurodan, LDS 751 (DNA), LDS 751 (RNA), Leucophor PAF, Leucophor SF, Leucophor WS, Lissamine Rhodamine, Lissamine Rhodamine B, Calcein / Ethidium Homodimer, LOLO-1, LO-PRO-1, Lucifer Yellow, Lyso Tracker Blue, Lyso Tracker Blue-White, Lyso Tracker Green, Lyso Tracker Red, Lyso Tracker Yellow, LysoSensor Blue, LysoSensor Green, LysoSensor Yellow / Blue, Mag Green, Magdala Red (Phloxin B), Mag-Fura Red, Mag-Fura-2, Mag-Fura-5, Mag-Indo-1, Magnesium Green, Magnesium Orange, Malachite Green, Marina Blue, Maxilon Brilliant Flavin 10 GFF, MaxilonBrilliant Flavin 8 GFF, merocyanine, methoxycoumarin, Mitotracker Green FM, Mitotracker Orange, Mitotracker Red, mithramycin, monobromobimane, monobromobimane (mBBr-GSH), monochlorobimane, MPS (Methyl Green Pyronine Stilbene), NBD, NBD-amine, Nile Red, nitrobenzoxadiazole, noradrenaline, Nuclear Fast Red, Nuclear Yellow, Nylosan Brilliant Iavin E8G (Nylosan Brilliant Iavin E8G), Oregon Green, Oregon Green 488-X, Oregon Green, Oregon Green 488, Oregon Green 500, Oregon Green 514, Pacific Blue, pararosaniline (Feulgen), PBFI, Phloxin B (Magdala Red), Phorwite AR, Phorwite BKL, Phorwite Rev, Phorwite RPA, Phosphine 3R, PKH26 (Sigma), PKH67, PMIA, Pontochrome Blue Black, POPO-1, POPO-3, PO-PRO-1, PO-PRO-3, primulin, Procion Yellow, propidium iodide (PI), PyMPO, pyrene, pyronine, pyronine B, pyrazole brilliant flavin 7GF, QSY 7, quinacrine mustard, resorufin, RH 414, Rhod-2, rhodamine, rhodamine 110, rhodamine 123, rhodamine 5 GLD, rhodamine 6G, rhodamine B, rhodamine B 200, Rhodamine B Extra, Rhodamine BB, Rhodamine BG, Rhodamine Green, Rhodamine phallicidin, Rhodamine phalloidin, Rhodamine Red, Rhodamine WT, Rose Bengal, S65A, S65C, S65L, S65T, SBFI, Serotonin, Sevron Brilliant Red 2B, Sevron Brilliant Red 4G, Sevron Brilliant RedB, Sevron Orange, Sevron Yellow L, SITS, SITS (primulin), SITS (stilbene isothiosulfonate), SNAFL calcein, SNAFL-1, SNAFL-2, SNARF calcein, SNARF1, Sodium Green, Spectrum Aqua, Spectrum Green, Spectrum Orange, Spectrum Red, SPQ [6-methoxy-N-(3-sulfopropyl)quinolinium], stilbene, sulforhodamine B and C, sulforhodamine Extra, SYTO 11, SYTO 12, SYTO 13, SYTO 14, SYTO 15, SYTO 16, SYTO 17, SYTO 18, SYTO 20, SYTO 21, SYTO 22, SYTO 23, SYTO 24, SYTO 25, SY SYTO 40, SYTO 41, SYTO 42, SYTO 43, SYTO 44, SYTO 45, SYTO 59, SYTO 60, SYTO 61, SYTO 62, SYTO 63, SYTO 64, SYTO 80, SYTO 81, SYTO 82, SYTO 83, SYTO 84, SYTO 85, SYTOX Blue, SYTOX Green, SYTOX Orange, tetracycline, tetramethylrhodamine (TAMRA), Texas Red, Texas Red-X conjugate, thiadicarbocyanine (DiSC3), thiazine red R, thiazole orange, thioflavin 5, thioflavin S, thioflavin TCN, Thiolyte, Thiozole Orange, Tinopol CBS (Calcofluor White), TMR, TO-PRO-1, TO-PRO-3, TO-PRO-5, TOTO-1, TOTO-3, TRITC (tetramethylrhodamine isothiocyanate), True Blue, TruRed, Ultralite, Uranine B, Uvitex SFC, WW 781, X-rhodamine, XRITC, Xylene Orange, Y66F, Y66H, Y66W, YO-PRO-1, YO-PRO-3, YOYO-1, YOYO-3, SYBR Green, thiazole orange (interchelating dye), or combinations thereof.
[0162] Another aspect of the present invention is directed to computer-implemented methods. For example, a computer and / or automated system may be provided that can automatically and / or repeatedly perform any of the methods described herein. As used herein, an "automated" device refers to a device that can operate without human direction; that is, an automated device can perform a function at a time after any human has taken any action to facilitate the function, for example, by inputting instructions into a computer to initiate the process. Typically, an automated apparatus can perform a repetitive function after this point. In some cases, the process steps may also be recorded on a computer-readable medium.
[0163] For example, in some cases, a computer may be used to control the imaging of a sample using, for example, a fluorescence microscope, STORM, or other ultra-high resolution techniques, such as the techniques described herein. In some cases, a computer may also control operations in image analysis, such as drift correction, physical registration, hybridization, and cluster alignment, cluster decoding (e.g., decoding fluorescent clusters), error detection or correction (e.g., as discussed herein), noise reduction, distinguishing foreground features from background features (such as noise or debris in an image), and the like. By way of example, a computer may be used to control the activation and / or excitation of signal-generating entities in a sample and / or the collection of images of the signal-generating entities. In one set of embodiments, a sample may be excited using light having various wavelengths and / or intensities, and using a computer, the sequence of wavelengths of light used to excite the sample may be correlated with images collected of the sample containing the signal-generating entities. For example, the computer may direct light having various wavelengths and / or intensities onto the sample to result in an average number of different signal-generating entities within each region of interest (e.g., one activating entity per location, two activating entities per location, etc.), optionally, this information may be used to construct an image of the signal-generating entities and / or determine the location of the signal-generating entities, optionally at high resolution, as mentioned above.
[0164] In some embodiments, the sample is placed on a microscope. In some cases, the microscope may contain one or more channels, such as microfluidic channels, that direct or control fluid to or from the sample. For example, in one embodiment, nucleic acid probes, such as those discussed herein, may be fluidically introduced and / or removed to or from the sample through one or more channels. In some cases, there may also be one or more chambers or reservoirs for holding fluid, for example, in fluid communication with the channel and / or the sample. Those skilled in the art will be familiar with channels, including microfluidic channels, for moving fluid to or from a sample.
[0165] The following documents are cited: U.S. Patent No. 10,240,146, entitled "Probe Library Construction"; U.S. Patent Application Publication No. 2017 / 0220733, entitled "Systems and Methods for Determining Nucleic Acids"; U.S. Patent Application No. 62 / 779,333, entitled "Amplification Methods and Systems for MERRFISH and Other Applications"; International Patent Application Publication No. WO2016 / 018960, entitled "Systems and Methods for Determining Nucleic Acids"; International Patent Application Publication No. WO2016 / 018963, entitled "Probe Library Construction"; International Patent Application Publication No. WO2018 / 089445, entitled "Matrix Imprinting and Clearing"; and "Systems and Methods for High-Throughput Image-Based International Patent Application Publication No. WO2018 / 218150, entitled "Imaging-Based Pooled CRISPR Screening," and International Patent Application Publication No. WO2018 / 089438, entitled "Multiplexed Imaging Using MERFISH and Expansion Microscopy," are each incorporated herein by reference in their entireties. In addition, the following: U.S. Provisional Patent Application No. 62 / 836,578, filed April 19, 2019, by Zhuang et al., entitled "Imaging-Based Pooled CRISPR Screening," and U.S. Provisional Patent Application No. 62 / 841,715, filed May 1, 2019, by Zhuang et al., entitled "Imaging-Based Pooled CRISPR Screening," are each incorporated herein by reference in their entireties. [Example]
[0166] The following examples are intended to illustrate certain embodiments of the present invention, but do not exemplify the full scope of the invention. [Example 1]
[0167] Pooled library CRISPR screening offers a powerful means for discovering genetic factors involved in cellular processes in a high-throughput manner. However, the phenotypes accessible to pooled library screening are limited. Complex phenotypes, such as cell shape and intracellular molecular organization, as well as their dynamics, require imaging-based readouts and are currently beyond the reach of pooled library CRISPR screening. These examples demonstrate an all-imaging-based pooled library CRISPR screening method that combines high-content phenotypic imaging with high-throughput guide RNA (sgRNA)-based identification in individual cells. In one such approach, sgRNAs are co-delivered into cells with corresponding barcodes placed in the 3' untranslated region (3'UTR) of a reporter gene using a lentiviral delivery system with reduced recombinogenic sgRNA-barcode mispairing. Multiplexed error-robust fluorescence in situ hybridization (MERFISH) can be used to read the barcodes and thereby identify sgRNAs with high accuracy. See, for example, International Patent Application Publication Nos. WO2016 / 018960, WO2016 / 018963, WO2018 / 089445, WO2018 / 218150, WO2018 / 089438, and WO2018 / 089438, each of which is incorporated herein by reference in its entirety. These examples use this approach to screen 162 sgRNAs targeting 54 RNA-binding proteins for their effects on RNA localization to nuclear compartments, and discover previously unknown regulators of RNA nuclear localization. Notably, these screens revealed MALAT1, a positive and negative regulator of long non-coding RNA (lncRNA) localization to nuclear speckles, suggesting dynamic regulation of lncRNA localization within subcellular compartments.
[0168] These examples develop an imaging-based pooled library CRISPR screening method that allows both the phenotype and genotype of individual cells to be read out by high-resolution, high-content imaging. This approach promises to substantially expand the phenotypic space accessible to pooled genetic screens by allowing complex cellular phenotypes, such as cell shape, the subcellular organization of different molecular species, and their dynamics, to be probed. This approach was applied to a screen for genetic factors involved in RNA localization within the nucleus, identifying both positive and negative regulators that control lncRNA localization to nuclear speckles.
[0169] These examples illustrate a method for imaging-based pooled library CRISPR screening in mammalian cells. This method enables both high-content phenotypic imaging of multiple molecular targets within individual cells and highly accurate identification of each cell's genotype. The latter is achieved by associating each sgRNA with a unique barcode and reading the barcodes using MERFISH (multiplexed error-robust fluorescence in situ hybridization). To demonstrate the power of this method, we performed a genetic screen for factors that regulate RNA localization within nuclear compartments. Diverse nuclear RNAs, including small nuclear RNAs (snRNAs), nucleolar RNAs (snoRNAs), and long non-coding RNAs (lncRNAs), are associated nuclear compartments formed by fluid-phase separation, such as nucleoli and nuclear speckles. Insight into the spatial regulation of these RNAs is critical for understanding how they orchestrate diverse nuclear activities and functions, including transcriptional regulation, transcript processing, and genome stability. We screened 162 sgRNAs (targeting 54 genes) for their effects on the localization of six RNA targets, including the lncRNA MALAT1, snRNA U2, and non-coding RNA 7SK, all of which are known to localize to nuclear speckles; the nascent pre-ribosomal RNA and non-coding RNA MRP, all of which are known to localize to the nucleolus; and poly(A)-containing RNAs. These results revealed numerous regulators of RNA localization within the nucleus. In particular, we identified both positive regulators essential for MALAT1 localization to nuclear speckles and negative regulators that reduce MALAT1 localization to nuclear speckles, suggesting dynamic regulation of lncRNA localization. [Example 2]
[0170] This example demonstrates high-throughput, high-precision barcode imaging in mammalian cells. In situ imaging-based pooled library screening, in which the genotype of individual cells is identified through multiplexed FISH imaging of barcodes associated with genetic variants, has recently been performed in bacteria. Due to the small volume of bacterial cells, the diffuse signal from barcode RNA within individual cells is sufficiently strong and easily measured. However, the approximately 1,000-fold larger volume of a mammalian cell makes it difficult to achieve a sufficiently high concentration of barcode RNA to enable reliable measurement. Therefore, new barcode expression and detection schemes are needed for mammalian cells to enhance barcode signals and reduce background.
[0171] To achieve this goal, two independent promoters were used within the same vector to express the sgRNA and reporter gene, and a 12-value ternary barcode was incorporated into the 3' untranslated region (3'UTR) of the reporter gene (Figure 1A). Each value in the ternary barcode (referred to hereafter as a trit) is composed of one of three different readout sequences (30 nucleotides (nt) long) specific to that value, corresponding to the three possible trit values: 0, 1, and 2. For example, the 12 trits represent a total of 3 12The system has the capacity to encode 531,441 barcodes. Because there were a total of 36 different trit sequences (three different sequences for each of the 12 trits), the barcodes were read using sequential hybridization rounds to image 36 pseudocolor channels (18 rounds of hybridization with two-color imaging per round, one pseudocolor channel per trit sequence), resulting in highly multiplexed detection. To increase the signal from the barcodes, a branched DNA amplification scheme was used to amplify the signal for each trit sequence (Figure 1A). To reduce background interference, reporter gene mRNA sequences were co-stained, and both the reporter gene and barcode sequences were detected by single-molecule FISH (smFISH) to examine only barcode signals that colocalized with the reporter gene signal (Figure 1A). For each specific trit, a trit value (0, 1, or 2) was assigned based on the pseudocolor channel that exhibited the greatest ratio of smFISH signal due to reporter mRNA that colocalized with the trit signal. This detection scheme reduced background signal arising from nonspecific binding of barcode FISH probes, which is important for decoding accuracy, as shown in the following section.
[0172] To investigate this barcode identification scheme, a library of vectors, each containing a common reporter gene, luciferase-mCherry, under the control of the same promoter and a unique barcode, was cloned in a pooled fashion (Figure 1B; see below and Figure 7). Although the total number of possible barcodes exceeds 500,000, the library was restricted to approximately 2,000 vectors (for error detection purposes, described below), and barcodes within the library were determined by sequencing. The library was delivered into the genome of U-2 OS cells using lentivirus at a low multiplicity of infection (MOI) so that the majority of transfected cells received only one barcode. Barcode signals for individual cells were measured using the multiplexed detection scheme described above.
[0173] After each round of hybridization, we observed clear barcode signals (Figure 1C) that colocalized with the smFISH signal of the reporter gene (luciferase-mCherry) mRNA. For detection of each trit, three trit values were probed individually (in different pseudocolor channels, as described above), and three distinct cell populations were observed, representing cells expressing barcodes with the three different trit values (Figure 1D; see below and Figure 8). A k-means clustering algorithm was used to separate the three cell populations, and the assignment of trit values to each population was based on which of the three pseudocolor channels assigned to that trit exhibited the greatest proportion of reporter gene mRNA spots that colocalized with the trit signal. Detection of 12 trits using 36 pseudocolor channels allowed us to assign a barcode to each cell. For the majority of cells (approximately 57%), the decoded barcode matched approximately 2,000 barcodes in the library as determined by sequencing (Figure 1E), excluding cells with mismatched barcodes. To assess the improvement in barcode detection accuracy using this reporter gene colocalization method, we also based barcode assignment to each cell on the number of FISH spots detected for the barcode signal alone (without considering colocalization with the reporter gene signal). In this case, we found that the decoded barcode did not match the actual barcode in the library (Figure 1F), likely due to background signal introduced by nonspecific FISH labeling, illustrating the substantial improvement in decoding accuracy by the reporter gene colocalization method.
[0174] Because any numeric readout error would most likely generate an invalid barcode not present in the library, the bottlenecking strategy used (i.e., limiting the total number of vectors in the library to approximately 2000, which represents <0.4% of the total possible number of 12-value ternary barcodes) allowed for error detection. Quantitatively, because only 0.4% of all possible barcodes were present in the library, the probability that any false-positive barcode would match a barcode in the library would be only 0.4%. Therefore, of the 57% of correctly matched barcodes, only 0.3% could result from barcode misidentification (see below for detailed calculations).
[0175] Figure 1 shows imaging-based barcode detection for genotyping in mammalian cells. Figure 1A shows a strategy for high-precision imaging-based barcode detection for genotyping. An sgRNA and a reporter gene are co-delivered with the imaging-based barcode into the genome of a host cell. As sequential hybridization rounds detect each value (trit) of the ternary barcode, the reporter gene portion of the mRNA is detected by single-molecule FISH (smFISH), and the barcode is detected by MERFISH. A 4x4 branched DNA amplification scheme is used to amplify the barcode signal. Figure 1B shows the construction design of a reporter gene-barcode library to probe barcode identification accuracy. Figure 1C, top panel, shows an example image showing the smFISH signal from the reporter mRNA and the signal for each of the three trit values (0, 1, and 2) for a single trit in the barcode. The bottom panel shows an enlargement of the white-boxed area in the top panel, with the reporter gene signal shown on the left and the reporter gene signal and barcode trit signal overlay on the right. For this cell, a trit value of 1 has a high colocalization rate, while trit values of 0 and 2 do not. The scale bar is 10 micrometers. Figure 1D shows the colocalization rates of three trit values measured for one example trit for all cells. Each spot in the plot corresponds to a single cell. The colocalization rate is defined as the number of reporter gene smFISH spots that colocalized with a trit signal spot divided by the total number of reporter gene smFISH spots in the cell. Using a k-means clustering algorithm, cells were divided into three clusters (indicated by different shading) based on their colocalization rate. Each cluster corresponds to cells with a specific trit value. Figure 1E shows a histogram for the number of cells with a different number of mismatched trits in the decoded barcode compared to the valid barcode in the library.The barcodes were decoded as described above using the colocalization of the reporter gene signal and the trit signal. Figure 1F is the same as Figure 1E, but the barcodes were decoded using only the number of measured trit signal spots, without considering the colocalization of the reporter gene signal and the trit signal.
[0176] Figure 7 shows the library cloning strategy for assessing imaging-based barcode detection accuracy. Briefly, barcodes and UMIs (single molecular identifiers) were first assembled from individual fragments of DNA oligos via two-step overlap PCR (see below). The shading for different oligos represents different trit sequences. This barcode-UMI library was then inserted into a digested lentiviral plasmid backbone to form a barcode-UMI lentiviral vector library. A reporter gene cassette was further inserted into the barcode-UMI lentiviral vector library to create the final reporter gene-barcode library.
[0177] Figure 8 shows the colocalization rate analysis for all 12 trits. The colocalization rates of the three values of each trit measured for all cells are presented for all 12 trits. Using a k-means clustering algorithm, cells were divided into three clusters (indicated by different shading) based on their colocalization rates. Each cluster corresponds to cells with one trit value. [Example 3]
[0178] To further verify the low false positive rate, we designed two reporter-barcode libraries (Figure 2A), each expressing the reporter gene luciferase-mCherry with a distinct epitope tag (HA tag or Myc tag) fused to the barcode library as described above. The two libraries were cloned separately. Each library was bottlenecked to contain <0.2% of all possible barcodes, making it extremely unlikely that the same barcode would appear in either library. The identity of the barcode associated with each epitope-tagged reporter gene was determined by sequencing. The two libraries were individually introduced into U-2 OS cells, and then the two libraries were pooled together in approximately equal numbers of cells. Immunofluorescence was used to image the phenotype of each cell, i.e., the expression of the HA or Myc tag (Figure 2B), and sequential hybridization rounds, as described above, were used to associate a barcode with each imaged cell. The rationale was that determining the phenotype of each cell would allow for an estimation of the identity of the barcode for this cell derived from the sequencing results, and then comparison with the barcode determined by imaging would allow for the determination of the proportion of misidentified barcodes.
[0179] Only about 1% of cells were barcode misidentified, as determined by barcode-phenotype mismatch (Figures 2C and 2D). Even this small error, most likely due to errors in cell segmentation, led to phenotyping errors, further supporting the extremely low barcode misidentification rate in these experiments.
[0180] Figure 2 illustrates the assessment of barcode misidentification rates using cells with known phenotype-barcode correspondences. Figure 2A illustrates the construction used to assess barcode detection accuracy. The reporter gene, luciferase-mCherry, was tagged with either an HA tag or a Myc tag, which define two phenotypes, and a nuclear localization signal to concentrate the HA and Myc signals in the nucleus for easy detection. The barcode was placed in the 3'UTR of the reporter gene, and the correspondence between the barcode and the HA or Myc tag was determined by sequencing. The reporter gene was driven by a CMV promoter. Figure 2B illustrates images showing HA and Myc immunostaining signals in two different channels. Nuclei with strong HA signals have weak Myc signals, and vice versa. Cell boundaries are highlighted. The nuclear boundaries of HA- and Myc-expressing cells are also labeled, respectively. The scale bar is 50 micrometers. Figure 2C shows scatter plots of the HA and Myc immunostaining intensities of individual cells. Cells assigned to the HA or Myc library based on imaging-based barcode determination are shown. Cells classified as positive for HA and Myc immunostaining (see below) are indicated by triangles and circles, respectively. Of 1,105 cells with HA-specific barcodes, only 10 were observed to be positive for Myc immunostaining, and of 1,034 cells with Myc-specific barcodes, only 9 were observed to be positive for HA immunostaining, indicating a low barcode misidentification rate. A small proportion of cells (197 of 2,336 cells) could not be clearly identified as HA- or Myc-positive because both their HA and Myc immunostaining signals were below the threshold or both their HA and Myc immunostaining signals were above the threshold (see below). These cells were excluded from the analysis.FIG. 2D shows a histogram of the ratio of HA intensity to Myc intensity for individual cells decoded by barcode imaging to contain either the HA- or Myc-tagged reporter. [Example 4]
[0181] This example illustrates a lentiviral delivery system with reduced recombination effects for accurate sgRNA identification. Another challenge in identifying sgRNAs by pooled barcode imaging arises from the viral system used to deliver sgRNA-reporter gene-barcode vectors to mammalian cells. Lentiviruses are a preferred delivery system for mammalian cells because they allow stable integration of the vector into the genome and transduction at a low MOI, allowing for the introduction of one sgRNA per cell (although other delivery systems can also be used in other cases). However, lentiviruses have two single-stranded RNA genomes, making them prone to recombination during viral transduction, which can result in mismatching between the sgRNA and the barcode. The recombination rate of lentiviruses is approximately one event per kilobase (kb). In these experiments, the sgRNA and reporter gene-barcode combinations were expressed separately under two independent promoters; therefore, in these examples, the barcode and sgRNA sequences were separated by large genomic distances (>1 kb), and thus the probability of recombination-induced barcode-sgRNA mispairing could be substantial.
[0182] This example illustrates a strategy modified from the CROP-seq method to overcome this recombination problem. Specifically, a reporter gene (puromycin-T2A-mCherry) was placed under a strong Pol II promoter (EF1α, EF1-alpha), and the sgRNA was placed under a separate promoter (hU6) downstream of the PPT (polypurine tract) in the lentiviral genome, together with the barcode (Figure 3A). In this way, the protospacer of the sgRNA, a sequence of approximately 20 nt for specific gene targeting, and the barcode sequence could be separated by a minimal distance (approximately 100 bases) in the genome. Although expression of the sgRNA downstream of the reporter gene may be impaired due to interference from the strong EF1α (EF1-alpha) promoter for reporter gene expression, the sgRNA expression cassette is replicated into the 5' LTR of the proviral genome upon integration into the genome, resulting in an additional functional unit for sgRNA expression that is free from interference from the EF1α (EF1-alpha) promoter (Figure 3A). Because transcription of the reporter gene is terminated only at the 3' end of the 3' LTR, the barcode should be expressed within the 3' UTR of the reporter mRNA for imaging-based barcode identification (Figure 3A).
[0183] To assess whether this construction design supports functional lentiviral infection and sgRNA expression, we constructed a library containing both sgRNAs targeting genes essential for cell viability and non-targeting control sgRNAs. Efficient sgRNA expression would deplete cells expressing sgRNAs targeting essential genes. We selected 159 sgRNAs targeting 53 essential ribosomal proteins (three sgRNAs per gene) as well as 51 non-targeting sgRNAs as controls (Dataset S1). A lentiviral library containing these 210 sgRNAs, along with a reporter gene (puromycin-T2A-mCherry) and barcode, was generated by pooled cloning (Figure 3A; see below and S3). U-2 OS cells stably expressing Cas9-BFP were then infected with this lentiviral library. Two days after lentiviral infection, both library-infected cells and cells expressing high levels of Cas9 were sorted based on mCherry and BFP fluorescence, respectively, and these cells were retained for experiments at different time points after infection. The abundance of cells expressing various sgRNAs was determined by sequencing genomic DNA.
[0184] As expected, cells containing sgRNAs targeting essential genes were significantly depleted compared to cells containing non-targeting sgRNAs, and the degree of depletion depended on the length of time after lentiviral infection (Figure 3B), indicating that this viral system can support the expression of functional sgRNAs. In addition, the abundance of cells containing different sgRNAs was measured by imaging-based barcode identification, as described above. The abundance of cells containing individual sgRNAs, as measured by imaging-based barcode identification, correlated well with the abundance of cells measured by direct sgRNA protospacer sequencing (Figure 3C), further supporting accurate barcode detection.
[0185] This experiment was then used to assess the recombination rate of the construct. If recombination occurred, the barcode assigned to the sgRNA of an essential gene could recombine with a non-targeting sgRNA, which would increase the abundance of cells measured by barcode imaging relative to the abundance of cells measured by protospacer sequencing. Similarly, the barcode assigned to a non-targeting sgRNA could recombine with an sgRNA targeting an essential gene, which would decrease the abundance of cells measured by barcode imaging. Therefore, the fold change in relative cell abundance was measured for both cells containing an sgRNA targeting an essential gene and cells containing a non-targeting sgRNA between days 2 and 21 after lentiviral transduction.
[0186] As expected, at day 21, the relative abundance of cells containing sgRNAs targeting essential genes was greatly reduced compared to day 2, whereas the relative abundance of cells containing non-targeting sgRNAs increased substantially at day 21 (Figure 3D). The fold change determined by barcode imaging was slightly smaller than the results obtained by sgRNA sequencing due to recombination (Figure 3D). This difference allowed quantification of the recombination-induced mispairing rate (see below), which was determined to be approximately 8% between the sgRNA protospacer and the sgRNA barcode (Figure 3E).
[0187] In addition, we also measured the recombination-induced mispairing rate between the sgRNA protospacer and the single-molecule identifier (UMI) (Figure 3A), a 20-nt sequence located approximately 500 bases downstream from the protospacer. As expected, the recombination-induced mispairing rate for the region between the protospacer and the UMI was large, approximately 16%, due to the large genomic distance between the UMI and the protospacer (approximately 500 nt) compared to the genomic distance between the barcode and the protospacer (approximately 100 nt) (Figures 3D and 3E). Because there are three possible sequences for any given trit, and the barcodes in the bottlenecked library are a randomly selected subset of all possible barcodes, we noted that for a random pair of barcodes, the probability that these barcodes share the same sequence at any given trit position is approximately 1 / 3. Therefore, we estimated that the recombination rate within the barcode region is approximately 1 / 3 of the recombination rate for perfectly homologous sequences of the same length. Based on the approximately 8% recombination rate measured for the approximately 100 nt genomic region between the barcode and protospacer (the consensus sequence of the sgRNA), the recombination rate within the approximately 400 nt barcode region is estimated to be approximately (400 / 100) × 8% / 3 = 10.7%, which would yield an approximately 8% + 10.7% = 18.7% recombination rate for the genomic region between the protospacer and UMI, consistent with the measured value of approximately 16%. Furthermore, because the barcode library was bottlenecked, recombinations occurring within the barcode region would be less likely to generate new barcodes that matched valid barcodes in the library, thereby resulting in less chance of misidentification of barcodes.
[0188] In summary, the low error rate (<1%) in barcode imaging and the low mismatch rate (approximately 8%) between sgRNAs and barcodes induced by recombination enabled highly accurate identification of sgRNAs by barcode imaging, enabling all-CRISPR screening using imaging-based pooled libraries. The remaining 8% mismatch rate between sgRNAs and barcodes could potentially generate false positives and false negatives in the screening. However, it was noted that the error rate was minimal because we typically probed hundreds of cells carrying the same sgRNA to determine whether the sgRNA had a statistically significant effect. Furthermore, we probed three sgRNAs targeting each gene and considered the gene a hit only if two of the three sgRNAs exhibited a statistically significant effect. Any remaining false positives could be easily identified through validation experiments.
[0189] Figure 3A illustrates the design of a lentiviral delivery method with a small recombination-induced sgRNA-barcode mismatch rate. Figure 3A shows the lentiviral construct used to deliver sgRNA and a barcode for sgRNA identification. The sgRNA cassette (hU6 promoter with sgRNA) and barcode array were placed downstream of a polypurine tract (PPT). A strong Pol II promoter (EF1α, EF1-alpha) drove expression of the reporter gene, puromycin-T2A-mCherry. After integration into the genome, the sgRNA cassette was replicated into the 5' LTR for sgRNA expression, while the barcode was expressed in the 3' UTR along with the reporter gene for barcode imaging. UMI: Unique Molecular Identifier. Figure 3B shows the protospacer counts for each sgRNA at days 8, 21, and 28 after lentiviral transduction plotted against the protospacer counts measured at day 2 after transduction. The protospacer counts at days 8, 21, and 28 were normalized by a factor such that the average counts for non-targeting sgRNAs for these conditions were the same as the average counts for non-targeting sgRNAs at day 2. Protospacer counts were determined by sequencing. As expected, cells expressing sgRNAs targeting essential ribosomal genes were strongly depleted over the time course, resulting in much reduced counts of sgRNAs targeting essential genes compared to non-targeting control sgRNAs. Figure 3C shows the correlation between the number of cells expressing a particular sgRNA, as measured by imaging-based barcode detection, and the sgRNA counts measured by protospacer sequencing at day 21 after lentiviral transduction. sgRNAs targeting essential genes are generally labeled red, and non-targeting control sgRNAs are generally labeled blue.Figure 3D is a violin diagram showing the median fold-change in relative sgRNA abundance measured by protospacer sequencing, imaging-based barcode detection, and UMI sequencing between days 2 and 21 post-lentiviral transduction. The relative abundance of a particular sgRNA is defined as the proportion of all sgRNA reads corresponding to this specific sgRNA (i.e., the protospacer (or UMI) reads for this particular sgRNA determined by sequencing, normalized by the total protospacer (or UMI) reads for all sgRNAs, or the number of cells expressing the barcode corresponding to this sgRNA, determined by imaging, normalized by the total cell number). As expected, the relative abundance of sgRNAs targeting essential genes decreased over the time course, while the relative abundance of non-targeting sgRNAs increased. The fold change due to recombination determined by barcode imaging was slightly smaller than that determined by protospacer sequencing, and the fold change determined by UMI sequencing was slightly smaller than that determined by barcode imaging. Figure 3E shows the median mismatch rates between protospacers and barcodes and between protospacers and UMIs due to recombination determined at 21 and 28 days after lentiviral transduction. Error bars indicate 95% confidence intervals. [Example 5]
[0190] This example illustrates a pooled CRISPR screen for factors that regulate RNA localization within the nucleus. To demonstrate the power of this screening method, potential regulators of RNA localization were screened within the nucleus (Figure 4A). Fifty-four candidate genes involved in nuclear RNA regulation were selected, including hnRNP family proteins, DExD / H box RNA helicases, and genes involved in RNA modification (Dataset S2). A library of 167 sgRNAs was designed, containing three sgRNAs for each of the 54 genes and five non-targeting sgRNAs as controls. A lentiviral library containing these sgRNAs, along with a reporter gene (puromycin-T2A-mCherry) and a barcode, was generated by pooled cloning (see below and Figure 9). To demonstrate the ability of this method to assess complex phenotypes, we used FISH to image the spatial distribution of five specific RNA species: lncRNA MALAT1, U2 snRNA, 7SK, MRP, and nascent pre-ribosomes, as well as poly(A)-containing RNAs. Additionally, we also incorporated the nuclear speckle protein, SON, into phenotypic imaging using immunolabeling with oligonucleotide-conjugated antibodies. These RNA and protein targets were imaged using sequential hybridization rounds with three to four different color channels per round, along with barcode imaging (Figure 4A) (see below for details of the imaging procedure).
[0191] As expected, the protein SON exhibited a clustered distribution, marking nuclear speckles, as well as MRP and pre-ribosomal signals, marking subnuclear compartments (Figure 4B). Based on these images, the boundaries of these structures were identified, and their number, the area covered by them, and their mean signal intensity (i.e., the total signal localized within the identified cluster boundaries divided by the total area covered by these clusters) were determined within individual cells. Next, we quantified the enrichment of MALAT1, U2, 7SK, and poly(A)-containing RNAs within nuclear speckles identified by SON staining (see below).
[0192] For quantification of each of these features, the values determined for cells harboring the targeting sgRNA were compared to the values measured from cells harboring a non-targeting control sgRNA to determine the fold change. Experiments were performed across four biological replicates, decoding a total of approximately 30,000 cells, and hits were determined based on the criterion that at least two of the three sgRNAs targeting a gene exhibited a statistically significant fold change (Dataset S3).
[0193] As a positive control, we detected statistically significant decreases in cluster signal intensity, cluster area, and cluster number associated with SON staining in cells expressing SON-targeting sgRNAs (Figure 4C). Additionally, sgRNAs against several DExD / H-box RNA helicases (DDX10, DDX18, DDX21, DDX24, DDX52, and DDX56) caused statistically significant changes in various features of nascent pre-ribosome staining (Figure 4D), consistent with the known function of these genes in ribosome biogenesis. Potentially, because not all cells expressing sgRNAs have their genomes edited, the magnitude of these phenotypic changes was observed to be modest (Figures 4C and 4D). Therefore, while these quantifications allowed for the identification of genetic perturbations with statistically significant effects, the magnitude of the phenotypic changes was not very informative. We also observed that perturbation of several genes within the hnRNP family caused significant changes in pre-ribosomal and MRP signals within the nucleolus (Dataset S3), potentially through indirect effects.
[0194] Figure 4 shows an imaging-based pooled CRISPR screen for regulators of nuclear RNA localization. Figure 4A shows a scheme of the imaging-based screen. Lentivirus-infected cells expressing sgRNA, barcode, and reporter gene were fixed and imaged. Barcodes were imaged by MERFISH using the 647 nm and 750 nm color channels over 18 rounds of hybridization (rounds 1–18). To increase the accuracy of barcode imaging, reporter gene mRNA was imaged using the 561 nm color channel in all rounds (rounds 1–18) to enable determination of colocalization of the barcode with the reporter gene mRNA signal. Seven protein and RNA targets for phenotypic measurement were imaged in the 488 nm color channel in the first seven rounds (rounds 1–7). The mosaic on the left contains 900 fields of view derived from a single screen. Figure 4B shows phenotypic images of SON, MRP, pre-ribosomes, MALAT1, U2 snRNA, 7SK, and polyA-containing RNA. SON marks nuclear speckles, while pre-ribosomes and MRP mark structures within the nucleolus. The cluster number, cluster area, and cluster intensity of SON, pre-ribosomes, and MRP are quantified. The enrichment of MALAT1, U2 snRNA, 7SK, and polyA-containing RNA within nuclear speckles is quantified. The scale bar is 20 micrometers. Figure 4C shows volcano plots of the effect of each sgRNA on SON cluster intensity, cluster area, and cluster number. Figure 4D shows volcano plots of the effect of each sgRNA on pre-ribosome cluster intensity, cluster area, and cluster number. In Figures 4C and 4D, the fold change induced by each sgRNA is calculated as the average value from all cells containing this sgRNA divided by the average value from cells containing all non-targeting sgRNAs.The horizontal dashed line indicates the p-value (0.05) used to define a hit from the screen. Data points for the indicated hits (i.e., two of the three sgRNAs targeting a gene show a statistically significant (p<0.05) fold change) are shown in a shade that matches the shade of the gene name shown in the caption; data points for other gene-targeting sgRNAs are shown in gray. Data points for non-targeting sgRNAs are shown in black.
[0195] Figure 9 shows the cloning strategy for the lentiviral sgRNA-barcode delivery library. Briefly, barcodes and UMIs were first assembled from individual fragments of DNA oligos via two-step overlap PCR, and then assembled with protospacer sequences and sgRNA non-variable region sequences using overlap PCR to form the sgRNA-barcode-UMI cassette library. The shading for different oligos represents different trit sequences. This library was then inserted into the digested lentiviral backbone containing the reporter gene, along with the hU6 promoter, downstream of a polypurine tract (PTT). [Example 6]
[0196] This example illustrates the involvement of novel factors in regulating MALAT1 localization to nuclear speckles. This screen revealed genes (Dataset S3) involved in regulating the localization of different RNA species to nuclear speckles. More genes were identified that regulated MALAT1 localization compared with 7SK, U2 snRNA, and poly(A)-containing RNA. This discussion focuses on MALAT1. In particular, we identified two groups of genes that inversely regulated MALAT1 localization to nuclear speckles (Fig. 5A; Dataset S3), and this was verified for all but one gene (hnRNPH3) by siRNA-mediated knockdown (Figs. 5B and 5C). Due to the lack of an effective antibody against this protein, we were unable to confirm whether siRNA against hnRNPH3 was effective. Depletion of the first group of genes, DHX15, DDX42, hnRNPK, and hnRNPH1, resulted in a statistically significant reduction in MALAT1 enrichment in nuclear speckles (Figures 5A-5C), suggesting that these genes upregulate MALAT1 localization to nuclear speckles. DHX15 and DDX42 are involved in spliceosome recycling and assembly, respectively, which is consistent with the involvement of mRNA splicing factors in the recruitment of MALAT1 to nuclear speckles. The involvement of hnRNP family proteins, hnRNPH1 and hnRNPK, in the upregulation of MALAT1 localization to nuclear speckles has not yet been reported. These two genes were also found to affect the localization of other RNA species, including U2 snRNA, polyA-containing RNA, pre-ribosomal RNA, and MRP (Dataset S3), which may imply a global perturbation effect of these two genes.
[0197] Unexpectedly, three factors, hnRNPA1, hnRNPL, and PCBP1, were also identified that negatively regulate MALAT1 localization to nuclear speckles. Their depletion by sgRNA or siRNA induced a statistically significant increase in MALAT1 enrichment within nuclear speckles (Figures 5A-5C). The fold change in MALAT1 enrichment induced by siRNA may be an underestimate due to incomplete knockdown. Combined knockdown of all three factors further increased MALAT1 enrichment within nuclear speckles (see Figure 10A below), which interestingly also resulted in an expansion of the nuclear speckle fraction (see Figures 10B and 10C below). In the triple knockdown samples, the composition of each nuclear speckle, as measured by the ratio of MALAT1 and SON levels within the nuclear speckle, also became more heterogeneous: some speckles had a reduced MALAT1 to SON ratio, while others had an increased ratio (see Figure 10D, below). These results indicated that the enhanced localization of MALAT1 to nuclear speckles upon knockdown of the three negative regulators was associated with changes in the shape and composition of nuclear speckles. This suggested a role for MALAT1 in regulating nuclear speckle structure, consistent with the observation that knockdown of MALAT1 can result in a reduction in nuclear speckle size.
[0198] Figure 5 shows genetic factors involved in regulating MALAT1 localization to nuclear speckles. Figure 5A shows a Volcano plot of the effect of each sgRNA on MALAT1 enrichment in nuclear speckles. Fold changes were calculated as described in Figure 4. The horizontal dashed line indicates the p-value (0.05) used to define hits from the screen. Hits confirmed by siRNA knockdown are highlighted with a shade matching the shade of the gene name shown in the caption, and data points for other gene-targeting sgRNAs are shown in gray. Data points for non-targeting sgRNAs are shown in black. Figure 5B shows a box plot showing the effect of siRNA knockdown of seven hit genes on MALAT1 localization alongside data for a control, non-targeting siRNA. The central line indicates the median, the boxes indicate 25-75% quartiles, and the whiskers indicate maximum and minimum values. For each condition, quantify 100-300 cells. Perform a Student's t-test for each condition compared to the control. **** :p<0.0001. Figure 5C shows images of MALAT1 localization upon siRNA knockdown of the seven hit genes. Data from control, non-targeting siRNA is also shown. MALAT1 staining is shown in the upper image, and SON staining is shown in the lower image. The scale bar is 10 micrometers.
[0199] Figure 10 shows that triple knockdown of hnRNPA1, hnRNPL, and PCBP1 affects the shape and composition of nuclear speckles. Figure 10A shows box plots showing the effects of control siRNA and single and triple knockdown (KD) of HNRNPA1, HNRNPL, and PCBP1 on MALAT1 localization. Box plot elements are as described in Figure 5. For each condition, 100 to 300 cells were quantified. Student's t-tests were performed between each single KD and the non-targeting control, and between the triple KD and single KD of hnRNPA1, hnRNPL, or PCBP1. **** p<0.0001. Figure 10B shows an example of an image of a cell showing that a portion of MALAT1-positive nuclear speckles is enlarged after triple KD of hnRNPA1, hnRNPL, and PCBP1 compared to cells transfected with a control, non-targeting siRNA (highlighted by arrows). The scale bar is 10 micrometers. Figure 10C shows the distribution of nuclear speckle size indicating that triple KD of hnRNPA1, hnRNPL, and PCBP1 increases nuclear speckle size. The Kolmogorov-Smirnov test for two samples was used to test for the difference between the two distributions. Figure 10D shows the distribution of log2 (ratio of MALAT1 intensity to SON intensity) within each nuclear speckle for a control siRNA sample and a triple KD sample of hnRNPA1, hnRNPL, and PCBP1. Approximately 300 cells and 7000 speckles are measured for the control and triple KD conditions, respectively. [Example 7]
[0200] It has been previously shown that the localization of MALAT1 to nuclear speckles can be impaired under transcriptional inhibition. However, the genetic factors involved in this process remain largely unknown. Therefore, we investigated whether these three negative regulators play a role in this process. To this end, in this example, we added the drug 5,6-dichloro-1-β(beta)-D-ribofuranosylbenzimidazole (DRB) to inhibit transcription and observed a substantial reduction in the enrichment of MALAT1 within nuclear speckles. Single knockdown of hnRNPA1, hnRNPL, and PCBP1 did not substantially rescue the DRB-induced dissociation of MALAT1 from nuclear speckles (Figures 6A and 11; see also below). On the other hand, double knockdown of two of these three factors or triple knockdown of all three factors largely rescued this DRB-induced dissociation effect (Figures 6A, 6B, and 11; see also below), suggesting that these hnRNP family proteins are required for transcription inhibition-induced dissociation of MALAT1 from nuclear speckles and that these factors likely play redundant roles in this process. These results suggest a potential mechanism for transcription inhibition-induced dissociation of MALAT1 from nuclear speckles. Upon transcription inhibition, RNA-binding proteins such as hnRNPA1 and hnRNPL are released from nascent mRNA transcripts, allowing them to bind to other RNA species. Thus, the released hnRNPA1 and hnRNPL may bind to MALAT1, which may compete with factors that recruit MALAT1 to nuclear speckles, thereby preventing MALAT1 from localizing to nuclear speckles under transcription inhibition.
[0201] Figure 6 shows that hnRNPA1, hnRNPL, and PCBP1 are required for transcription inhibition-induced dissociation of MALAT1 from nuclear speckles. Figure 6A shows quantification of MALAT1 enrichment in nuclear speckles for cells transfected with different siRNA combinations, with or without treatment with the transcription inhibitor DRB (50 micromolar, 1 hour). 100–300 cells were quantified for each condition. Transcription inhibition-induced dissociation of MALAT1 from nuclear speckles was not rescued by single knockdown of hnRNPA1, hnRNPL, or PCBP1, but was rescued by double and triple knockdown of these factors. Figure 6B shows images showing that in cells transfected with control siRNA, MALAT1 dissociates from nuclear speckles upon transcription inhibition; whereas in cells cotransfected with siRNA targeting hnRNPA1, hnRNPL, and PCBP1, transcription inhibition does not dissociate MALAT1 from nuclear speckles. The scale bar is 10 micrometers.
[0202] 11 shows the fold change of enrichment of MALAT1 in nuclear speckles after transcription inhibition under different knockdown conditions. This bar graph shows the fold change of enrichment of MALAT1 in nuclear speckles after transcription inhibition under each knockdown condition. This fold change is defined as the enrichment of MALAT1 in nuclear speckles after transcription inhibition divided by the enrichment before transcription inhibition under the same siRNA treatment. Control siRNA, hnRNPA1, hnRNPL, and PCBP1 single knockdown conditions, as well as hnRNPA1, hnRNPL, and PCBP1 double knockdown conditions and triple knockdown conditions are shown. Error bars: SD. Student's t-test is performed for each condition against the control. ** :p<0.01, ****: p<0.0001, ns: non-significant. n=3, each experiment containing 30 to 100 cells. [Example 8]
[0203] These examples demonstrate the development of an imaging-based pooled library CRISPR screening method that allows genotype-phenotype correspondences to be established for individual cells, enabling high-throughput screening of mammalian cells based on complex phenotypes previously inaccessible to pooled library screening. This imaging-based screening was enabled by the identification of sgRNAs via MERFISH-based barcode detection, demonstrating a low barcode misidentification rate of approximately 1%. A lentiviral delivery scheme was devised that reduced the recombination-induced mismatch rate between sgRNAs and barcodes (mismatch rate <10%). Collectively, these results led to highly accurate identification of sgRNAs via barcode imaging. These approaches substantially expanded the phenotypic space accessible to pooled library screening. Compared with imaging-based screening using an arrayed format in which individual gene perturbations are assayed individually in individual wells, a major advantage of pooled screening is that reagents for gene perturbations, i.e., DNA plasmids and lentiviruses, can be prepared in a pooled manner with standard molecular biology procedures, reducing labor and costs, which is particularly beneficial for large-scale custom-designed libraries. Reagent preparation for arrayed screening typically requires costly multiwell robotic processing systems and more complicated procedures. Another advantage of pooling methods is that, because measurements for all gene perturbations are performed in the same experiment, variations in experimental conditions for different perturbations can be minimized. This is particularly desirable when cells are treated with concentration- or time-sensitive conditions. Furthermore, pooled formats can also simplify multiplexed phenotypic measurements, which require sequential rounds of staining and signal removal via buffer exchange.On the other hand, if the generation of individual gene perturbation reagents is not particularly demanding and the phenotypic measurements are not highly sensitive to variations in sample processing conditions, arrayed screening may be preferable, as the MERFISH barcode readout process substantially increases the complexity of the imaging procedure.
[0204] Current 12-ary ternary barcode libraries contain over 500,000 barcodes. Even with a strict 1% bottlenecking strategy, which allows for robust barcode detection against errors, over 5,000 distinct sgRNAs can still be incorporated into each library; this capacity can be easily expanded by adding more values to the barcodes. The current limitation on the number of sgRNAs that can be screened is the time required to image a large number of cells. This imaging system utilizes a high-magnification (60x) objective to read FISH signals on individual single mRNA molecules for barcode detection, limiting the number of cells that can be imaged within each field of view. However, imaging speed can be substantially improved by: 1) using greater amplification of the barcode signal, using lower magnification objectives, allowing each field of view to be captured at a higher frame rate and / or allowing more cells to be imaged within each field; 2) using multiple cameras for detection, allowing simultaneous detection of fluorescent signals in different color channels. These improvements can achieve a greater than 10-fold improvement in the number of cells and genotypes that can be screened per experiment.
[0205] To demonstrate the power of this method for screening complex phenotypes in mammalian cells, we imaged the subcellular localization of seven different molecular species, including six RNAs and proteins. These screening experiments revealed unknown regulators of RNA localization within the nucleus. Interestingly, we identified both positive and negative regulators of the nuclear speckle localization of lncRNA MALAT1. Positive regulators included the DExD / H-box RNA helicases DHX15 and DDX42 and the hnRNP family genes hnRNPH1 and hnRNPK; whereas negative regulators included hnRNPA1, hnRNPL, and PCBP1. RNAs can localize to subcellular compartments formed by phase separation via at least two mechanisms: 1) RNAs can act as scaffolds that can nucleate phase separation, such as mRNAs in P-bodies and stress granules and pre-ribosomal RNAs in nucleoli; and 2) RNAs can be recruited to phase-separated bodies as clients, which have been shown to contribute to the localization of MALAT1 in nuclear speckles. It is possible that the negative regulators discovered in this screen compete with factors that recruit MALAT1 to nuclear speckles, thereby preventing MALAT1 from localizing to nuclear speckles. The role of these negative regulators in the dissociation of MALAT1 from nuclear speckles induced by transcriptional inhibition was also identified. These results suggest that lncRNA localization can be dynamically regulated by protein factors.
[0206] These results support the ability of this imaging-based screening method to reveal molecular factors involved in cellular processes that can only be assessed by high-resolution imaging. This screening method is broadly applicable to probing genetic factors that control or regulate a broad spectrum of phenotypes, including shape features, molecular organization, and dynamics of cellular structures, as well as cell-cell interactions. This screening method can also be combined with highly multiplexed DNA, RNA, and protein imaging methods, including genome-scale imaging methods, to profile factors involved in gene regulation and other genomic functions in a high-throughput manner. [Example 9]
[0207] This example illustrates the various materials and methods used in these examples.
[0208] Cloning of reporter-barcode and sgRNA-reporter-barcode libraries was performed in a pooled fashion using oligos ordered from IDT (Datasets S1, S2, and S4). These libraries were cloned into pFUGW, a lentiviral vector described below. High-throughput sequencing was used to identify the barcodes present in the libraries and establish barcode-sgRNA correspondence. Lentivirus was produced in LentiX cells (Takara Bio Inc., 632180) using Lenti-X™ Packaging Single Shots (VSV-G) (Takara Bio Inc., 631276). U-2 OS cells were infected with the lentiviral libraries at a low multiplicity of infection (MOI) so that only 10–20% of the cells were infected. Infected cells were sorted based on mCherry expression and Cas9-BFP expression. Sorted cells were fixed, permeabilized, and stained for imaging according to the detailed methods discussed below.
[0209] A custom-built microscope with a Nikon Ti-U microscope body and a Nikon CFI Plan Apo Lambda 60x oil immersion objective with a numerical aperture of 1.4 was used for imaging. For sequential hybridization rounds and imaging, a peristaltic pump (Gilson, MINIPULS 3) aspirated liquid (TCEP buffer for dye cleavage, hybridization buffer with readout probe, or hybridization buffer for sample washing) into a Bioptech FCS2 flow chamber with three valves (Hamilton, MVP and HVXM 8-5) to select the input fluid (see details below). Barcode decoding and phenotype quantification based on collected images are also described in detail below.
[0210] Cell culture: U-2 OS cells were cultured at 37°C in EMEM medium (ATCC, HTB-96) supplemented with 10% FBS (Sigma, F4135-1L) and 1% penicillin / streptomycin (Invitrogen, 15140122) antibiotics. Following lentiviral transduction, U-2 OS cells stably expressing Cas9-BFP were generated via FACS sorting using the BFP signal. To generate the lentiviral vector for Cas9-BFP, the Cas9-BFP sequence was PCR amplified from pLentiCas9-BFP (Addgene product no. 78545) and cloned into the pFUGW backbone along with the SVVF promoter. Two nuclear localization signal sequences were added to enhance nuclear localization of Cas9.
[0211] Cloning of vector libraries: Each 12-valued barcode contained twelve 30-nt sequences, each representing a trit, with the nucleotide "A" separating adjacent trits. Oligos encoding each pair of adjacent 30-nt sequences, oriented alternately in the forward and reverse directions, were ordered from IDT (i.e., Trit1 + Trit2, Trit2 + Trit3 reverse complement, Trit3 + Trit4, Trit4 + Trit5 reverse complement, Trit9 + Trit10, Trit10 + Trit11 reverse complement, Trit11 + Trit12; see Dataset S4). Because each trit had three distinct values, represented by three significantly different 30-nt sequences, each pair of adjacent 30-nt sequences had nine possible distinct sequences, requiring a total of 11 × 9 = 99 oligos to cover all possible pairs of adjacent 30-nt sequences. Two non-variable primer binding sequences were added to both ends of the barcode (represented by oligos 1-9 and 91-99) for PCR amplification purposes. The sequences of these 99 oligos are listed in Dataset S4. The entire barcode library was assembled by two-step overlap PCR. First, the 12 trits were divided into three segments, each generated by the following reactions: Segment 1: oligos 1-36 as template, oligos 1-9 as forward primers, and oligos 28-36 as reverse primers; Segment 2: oligos 37-72 as template, oligos 37-45 as forward primers, and oligos 64-72 as reverse primers; Segment 3: oligos 73-99 as template, oligos 73-81 as forward primers, and oligo 100 as reverse primers. The three PCR products were gel-purified. The three segments were then mixed and subjected to overlap PCR using forward primer oligo 101 and reverse primer oligo 102. The reverse primer in this step contained a 20-base random sequence region used as a single molecule identifier (UMI) for the sequencing step. The sequences of oligos 100-102 are also listed in Dataset S4.All PCR reactions were performed using a real-time qPCR machine that monitored the reactions, stopping them in the logarithmic growth phase to reduce library distortions resulting from PCR bias. The PCR products were assembled into a modified pFUGW backbone via isothermal assembly. The assembled library was electroporated into Endura electrocompetent cells (Lucigen, 60242-2) and then grown overnight under ampicillin selection to amplify the library. The amplified library was purified by Mini Prep (mini prep). This library was named the pFUGW_barcodes_UMI_backbone library.
[0212] The pFUGW_barcodes_UMI_backbone library was then used to generate a library containing a reporter gene (luciferase-mCherry) for barcode imaging. First, a reporter cassette containing a CMV promoter and a reporter open reading frame was generated in an intermediate vector. The open reading frame contained luciferase-mCherry, 2x HA tags, or 2x Myc tags at the N-terminus and a nuclear localization signal at the C-terminus. The reporter cassette was PCR amplified from the intermediate vector using oligos 103 and 104 (sequences presented in Dataset S4). The pFUGW_barcodes_UMI_backbone library was then digested with BstXI, treated with alkaline phosphatase, and assembled with the reporter cassette PCR product using isothermal assembly. The assembled libraries were electroporated into Endura electrocompetent cells and grown overnight under ampicillin selection for amplification. The cells were then diluted to contain the desired number of constructs in each library. These bottlenecked libraries were then purified by Mini Prep. These libraries were named reporter gene_barcodes libraries.
[0213] Cloning of sgRNA-barcode libraries: The sgRNA-barcode libraries were cloned using the following strategy: first, a protospacer-sgRNA nonvariable region-barcode cassette library was generated via multi-step overlap PCR; then, this library was inserted via isothermal assembly into a lentiviral vector containing a U6 promoter downstream of a PPT sequence. To generate the protospacer-sgRNA nonvariable region-barcode cassette library, barcode segments were first generated using a method similar to that described in the "Cloning a Barcode Library for Quantifying Barcode Decoding Accuracy" section. The only difference was that the nonvariable region within the 5' end was altered by replacing oligos 1-9 with oligos 105-113, since the barcode was positioned immediately adjacent to the sgRNA. Oligos 114 and 115 were used to PCR amplify the sgRNA nonvariable region. For PCR amplification, a protospacer library with nonvariable regions on both sides of the protospacer was ordered from IDT. In this work, we created two protospacer libraries: one for essential ribosomal genes and non-targeting sgRNA controls (Dataset S1), which were used to measure recombination rates between sgRNAs and barcodes; and one for targeting genes potentially regulating RNA localization within the nucleus (Dataset S2). The protospacer library was PCR-amplified using oligos 116 and 117, and the PCR products were gel-purified. The protospacer, sgRNA non-variable region, and barcode PCR products were then mixed and subjected to overlap PCR using oligos 116 and 118 as primers. Oligo 118, the reverse primer, contained a 20-base random sequence region used as a single molecule identifier (UMI) for the sequencing step. The sequences of oligos 105–118 are also listed in Dataset S4. All PCR reactions were performed using a real-time qPCR instrument that monitored the reaction so that it was stopped in the logarithmic growth phase.The PCR products were assembled into a modified pFUGW backbone via isothermal assembly, with a U6 promoter located downstream of the PPT sequence. The assembled libraries were electroporated into Endura electrocompetent cells and then grown overnight on ampicillin selection plates for amplification. A certain number of colonies (approximately 3,800 for the essential ribosomal gene library and approximately 2,500 for the RNA localization screening library) were scraped from the plate with LB buffer and cultured overnight in 200 mL of LB buffer. The libraries were purified using Maxi Prep (maxi prep). These libraries were named sgRNA_barcodes libraries.
[0214] Preparation and analysis of sequencing libraries: To determine the identity of the barcodes represented in the library as well as to establish the correspondence between sgRNAs and barcodes, the libraries were analyzed using high-throughput sequencing. It was found that PCR amplification of barcode regions can lead to barcode recombination due to homologous regions between barcodes. Therefore, a ligation-based approach was used to install sequencing adapters into the barcode library. In this approach, the region to be sequenced was digested from the library and then ligated to the adapter using T4 ligase.
[0215] To determine the barcodes within the reporter gene-barcode library, the library was digested with BstXI and BamHI at 37°C for 2-3 hours, and the resulting fragments were purified using a Zymo DNA purification kit (ZD4002). To generate adapters with sticky ends for ligation, oligos 119-124 (sequences presented in Dataset S4) were mixed at 0.5 micromolar each and subjected to five cycles of PCR. The product was purified using a Zymo DNA purification kit to create double-stranded sequences in which the 5' and 3' adapters were separated by BstXI and BamHI digestion sites. The purified product was digested with BstXI and BamHI at 37°C for 2-3 hours and then purified. The resulting mixture, containing adapters with sticky ends for ligation, was mixed with the purified library fragment mixture described above. T4 ligase was added, and the reaction was held at room temperature for 2-4 hours. The reaction mixture was directly subjected to electrophoresis on a 2% agarose gel, and a band corresponding to a size of approximately 400 bp was excised and purified. The purified DNA sample was used for concentration measurement and high-throughput sequencing using the V2-MISeq kit (Illumina, MS-103-1003).
[0216] Because the length from the protospacer to the end of the barcode exceeds 500 bp, the optimal length range for high-quality sequencing using the V2-MISeq kit, two sequencing libraries were generated to determine the sgRNA-barcode correspondence of the sgRNA_barcodes library. In one library, ligation sites were created 5' to the protospacer and 3' to the UMI using BstXI and BamHI, respectively. Because some central barcodes were not accessible by sequencing from either end, sequencing of this library covered the protospacer region, part of the barcode region, and the UMI region. In the second library, ligation sites were created using KpnI and BamHI (the KpnI site was placed immediately after the sgRNA and before the barcode). Sequencing of this library covered the entire barcode region as well as the UMI region. UMI sequences were used to identify protospacers and barcodes within the same construct from these two libraries. Oligos 125-130 and 131-136 (sequences presented in Dataset S4) were used to generate adapters for the first and second libraries, respectively. The procedure was identical to that described for generating the sequencing libraries for the reporter gene-barcode library.
[0217] UMI sequences, protospacer sequences, and barcode sequences were extracted from sequencing reads.Then, the reads were grouped by common UMI and barcode to generate a codebook for the correspondence between sgRNAs and barcodes.Reads with incorrect protospacers or barcodes assigned to multiple sgRNAs were excluded from further analysis.
[0218] To determine the distribution of protospacers and UMIs derived from cell populations at different time points after lentiviral transduction in experiments using sequencing to determine recombination rates, sequencing libraries were prepared from purified genomic DNA by PCR amplification using oligos 137–140 and oligos 141–144 (sequences presented in Dataset S4) as forward and reverse primers, respectively.
[0219] Lentivirus production and transduction: Lentivirus was produced in LentiX cells (Takara Bio, Inc., 632180) using Lenti-X™ Packaging Single Shots (VSV-G) (Takara Bio, Inc., 631276). The produced virus was concentrated using a Lenti-X™ Concentrator (Takara Bio, Inc., 631231) and stored at -80°C. For transfection, the amount of virus was controlled so that 10-30% of the cells were transduced, ensuring that the majority of infected cells were infected with only a single virus particle. Viral transduction was performed using 10 micrograms / mL of polybrene (Sigma, TR-1003-G). The viral titer of the construction in which the U6-sgRNA-barcode array was placed after PPT did not show a clear reduction compared to the viral titer of the construction without an insert after PPT, indicating that the insert did not impair lentiviral transduction.
[0220] siRNA knockdown: All siRNAs were purchased from Dharmacon, and siRNA knockdown was performed according to the Dharmacon protocol. Briefly, U-2 OS cells were seeded on imaging coverslips in 12-well plates at 30,000 cells per well. For siRNA transfection, 1.5 microliters of 20 micromolar siRNA was added to 100 microliters of serum-free, antibiotic-free medium in one tube, and 1 microliter of Dharmacon transfection reagent (Dharmacon, T-2001-01) was added to 100 microliters of serum-free medium in a separate tube. The two tubes were incubated for 5 minutes, then gently mixed and incubated at room temperature for an additional 20 minutes. 800 microliters of serum-containing antibiotic-free medium was mixed with 200 microliters of siRNA and the transfection reagent mix described above to produce 1 mL of transfection medium. The cell culture medium was replaced with 1 mL of transfection medium. The cells were incubated at 37°C for 72 hours before phenotyping.
[0221] Silanization of imaging coverslips: First, imaging coverslips were cleaned with 1M KOH and pure methanol, washed with 70% ethanol, and dried in an oven. For silanization, coverslips were covered with silanization buffer (500 mL distilled water, 1500 microliters of Bind-silane (Sigma, GE17-1330-01), and adjusted to pH 3.5 with ice-cold acetic acid) at room temperature for 1 hour. Then, coverslips were washed with water and dried for storage. Before seeding cells, silanized coverslips were coated with 1% poly-D-lysine (Sigma, P0899) in a 60 mm diameter cell culture dish for 30 minutes, and then washed with water for 1 hour.
[0222] Preparation of imaging samples: U-2 OS cells were seeded on coverslips for 2 days before fixation. For phenotypic imaging in experiments screening for factors involved in regulating nuclear RNA localization, U-2 OS cells were fixed 6 days after lentiviral transduction. Samples were fixed with 4% paraformaldehyde (EMS, 15714) in PBS for 15 minutes and permeabilized in 0.5% Triton-X (Sigma, X100) for 30 minutes. Next, the samples were incubated in blocking buffer (500 microliters of blocking buffer: 50 microliters of 10x PBS, 200 microliters of RNAse-free BSA (ThermoFisher, AM2618), 50 microliters of 25 mg / ml Yeast tRNA (ThermoFisher, 15401029), 5 microliters of mouse RNase inhibitor (NEB, M0314L), 1 microliter of 25% Triton-X, and RNAse-free water to 500 microliters) for 1 hour and stained with anti-SON (ABCAM, ab121759) primary antibody at 1:100 in blocking buffer for 1 hour at room temperature. The samples were washed three times with 1x PBS and incubated with 1:300 oligonucleotide-labeled secondary antibody for 1 hour. The oligonucleotide-labeled secondary antibody was then probed with a readout probe with a sequence complementary to the oligonucleotide sequence on the antibody. The samples were washed three times with 1x PBS and post-fixed with 4% PFA for 30 minutes. Prior to FISH staining, the samples were equilibrated for 5 minutes in 30% formamide in 2x SSC. The FISH hybridization buffer contained 30% formamide (ThermoFisher, AM9342), 60% Stellaris RNA FISH hybridization buffer (Biosearch, SMF-HB1-10), 25 mg / mL 10% yeast tRNA, and 1:100 mouse RNase inhibitor.Samples were stained overnight at 37°C with 300 nM FISH probes for the reporter gene, 300 nM FISH probes for RNA phenotype (i.e., six RNA molecular species) imaging, and 100 nM primary amplification probes for barcode imaging. Each FISH probe for the reporter gene contained a 30-nt targeting sequence capable of binding to the reporter gene mRNA and three 20-nt readout sequences that allow binding of complementary fluorescently labeled readout probes. Each FISH probe for each RNA target in phenotype imaging contained a 30-nt targeting sequence capable of binding to the RNA target and one or two 20-nt readout sequences that allow binding of complementary fluorescently labeled readout probes. Each primary amplification probe for barcode imaging contained a 30-nt targeting sequence capable of binding to one of the 30-nt trit sequences on the barcode, as well as four additional 30-nt identical sequences that allowed binding of secondary amplification probes (Fig. 1A). The samples were then washed twice in 30% formamide in 2x SSC and stained with 100 nM secondary amplification probe for barcode imaging in 10% hybridization buffer (10% formamide, 80% Stellaris RNA FISH hybridization buffer, 10% 25 mg / mL Yeast tRNA, and 1:100 mouse RNase inhibitor) at 37°C for 1 hour. Each secondary amplification probe contained a 30-nt targeting sequence capable of binding to the primary amplification probe, as well as four additional 20-nt identical readout sequences that allowed binding of complementary fluorescently labeled readout probes. This amplification scheme therefore allows for up to 16-fold signal amplification.Samples labeled with FISH probes for phenotype imaging and reporter gene mRNA imaging, and primary and secondary amplification probes for barcode imaging, were washed twice in 30% formamide in 2x SSC and then embedded in a 4% polyacrylamide gel. Protein and lipids were then removed from the samples by overnight incubation at 37°C with protein digestion buffer [for 50 mL of digestion buffer: 5 mL of 8 M guanidine-HCl (ThermoFisher, 24115), 2.5 mL of 1 M Tris pH 8.0 (ThermoFisher, 15569025), 100 microliters of 0.5 M EDTA (ThermoFisher, 15575020), 0.25 mL of Triton-X, and 1:100 proteinase K (ThermoFisher, AM2548)]. This step is referred to below as the sample cleanup step. Because cleavage with protease K resulted in protein digestion (including digestion of the mCherry protein), the fluorescent signal from mCherry was eliminated after digestion and did not interfere with FISH signal detection using the 561 nm channel. FISH probes for oligonucleotides linked to poly(A)-containing RNA, 7SK, MRP, U2 snRNA, and secondary antibodies for SON staining were conjugated with acrydite, which can be crosslinked to polyacrylamide gels and retain these probes, as well as their bound RNAs, in the gel during the sample cleanup step. FISH probes for MALAT1 and preribosomes were not labeled with acrydite because the large size of both MALAT1 and preribosomes would have retained them in the gel during sample cleanup. The reporter gene mRNA was linked to the gel via an acrydite-labeled poly-T probe that could bind to the poly-A tail of the reporter mRNA, allowing the FISH probe for the reporter gene and the FISH probe for barcode imaging to be retained within the gel during cleanup.The sample cleanup step substantially reduces background signals due to cellular autofluorescence and nonspecific binding of FISH probes to proteins and lipids. Samples were then washed with 2x SSC and left in 2x SSC for imaging. The sequences of the FISH probes used are listed in Dataset S5.
[0223] For experiments using two known phenotypes (expression of reporter genes tagged with HA or Myc) to measure barcode identification errors, U-2 OS cells were fixed 6 days after transduction. The tags were stained with primary antibodies [anti-Myc (Abcam, ab9132), anti-HA (Abcam, ab9110)], followed by Alexa 405-labeled anti-mouse secondary antibodies (Abcam, ab175658) and Alexa 488-labeled anti-rabbit secondary antibodies (Invitrogen, A21206). Samples were incubated for 1 hour in 25 mM MA-NHS (Sigma, 730300) in 2x SSC before gel embedding, allowing the MA-NHS-labeled antibody to link to the gel upon gel polymerization. After sample cleanup, the antibody was digested into fragments, and the dye was linked to the gel via the cross-linked antibody fragments. The dyes, Alexa 405 and Alexa 488, persist after polymerization during gel embedding. The remainder of the sample preparation, including immunostaining and barcode staining, was as described above.
[0224] Labeling of antibodies with oligonucleotides: Antibodies were labeled with oligonucleotides using the following strategy. First, the antibody was mixed with DBCO-NHS, which conjugates DBCO to the antibody. Then, the DBCO-labeled antibody was mixed with azide-labeled oligonucleotide to conjugate the oligonucleotide to the antibody. Specifically, 100 micrograms of anti-rabbit antibody (ThermoFisher, 31210) was buffer-exchanged into 100 microliters of PBS using a 50 kD protein concentrator (Millipore, UFC510024). NaHCO3 and DBCO-NHS ester (Kerafast, FCC310) were added to the antibody solution to final concentrations of 50 mM and 100 micromolar, respectively. The reaction was allowed to proceed at room temperature for 1 hour to produce the DBCO-labeled antibody, and excess DBCO was removed via buffer exchange with PBS using a 50 kD protein concentrator. PBS buffer was then added to the DBCO-labeled antibody to bring the solution volume to 100 microliters, and 25 microliters of azide-labeled oligonucleotide (100 micromolar, Dataset S3) was added. The reaction was allowed to proceed overnight at 4°C. After the reaction was completed, excess oligonucleotide was removed via buffer exchange using PBS, and the final oligonucleotide-labeled antibody was aliquoted and stored at -80°C.
[0225] Imaging setup and sequential imaging: The imaging setup was as previously described. See, for example, U.S. Patent Application Publication No. 2017-0220733 or International Patent Application Publication No. WO2018 / 089438, each of which is incorporated herein by reference in its entirety. Briefly, the input fluid was selected using a peristaltic pump (Gilson, MINIPULS 3) aspirating the liquid into a Bioptech FCS2 flow chamber with a sample cover slip and three valves (Hamilton, MVP and HVXM 8-5). A custom-built microscope with a Nikon Ti-U microscope body and a Nikon CFI Plan Apo Lambda 60x oil immersion objective with a numerical aperture of 1.4 was used for imaging. Solid-state single-mode lasers (405 nm laser, Obis 405 nm LX 200 mW, Coherent; 488 nm laser, Genesis MX488-1000, Coherent; 560 nm laser, 2RU-VFL-P-2000-560-B1R, MPB Communications; 647 nm laser, 2RU-VFL-P-1500-647-B1R, MPB Communications; and 750 nm laser, 2RU-VFL-P-500-750-B1R, MPB Communications) were used for irradiation. AOTFs (acousto-optic tunable filters) were used to control the intensity of the 488 nm, 560 nm, and 647 nm lasers; the 405 nm laser was modulated by a direct digital signal; and the 750 nm laser was switched by a mechanical shutter. A custom-made dichroic filter (Chroma, zy405 / 488 / 561 / 647 / 752RP-UF1) and an emission filter (Chroma, ZET405 / 488 / 461 / 647-656 / 752m) were used to separate excitation illumination from fluorescence emission. Emission light was imaged onto a Hamamatsu digital CMOS camera. During acquisition, the sample was translated using a motorized XY stage (Ludl, BioPrecision2) and kept in focus using a homemade autofocus system.
[0226] For experiments requiring high-throughput imaging, a 40x objective (CFI60 Plan Fluor 40x oil immersion objective) was used to collect more cells per field of view (FOV). A four-camera system was used to collect signals from 750 nm, 647 nm, 561 nm, and 488 nm fluorophores individually and simultaneously. Specifically, a four-camera mount (QuadCam LS 1.0x, 89 North) was attached to a Nikon Ti-U microscope body, and four Hamamatsu digital CMOS cameras were attached to the mount. Four dichroic filters (T495lpxr, T562lpxr, T647lpxr, T760lpxr, Chroma Tech) that distribute signals from different fluorophores and four single-band emission filters (ET 450 / 50m, ET 525 / 50m, ET 605 / 75m, and ET 705 / 70m, Chroma Tech) were mounted on the camera mount. To align signals from the four cameras, multicolor beads (FP-0257-2, Spherotech) were imaged, and signals from the 647 nm, 561 nm, and 488 nm color channels were aligned against the signal from the 750 nm color channel using the cp2tform function in MatLab. Cp2tform inferred polynomial spatial transformations for x, y, and z coordinates and applied the transformations to the barcode and phenotype signals. For imaging nuclei in a 4-camera setup, 1:1000 647 nm Nucred dye (R37106, ThermoFisher) was used rather than DAPI.
[0227] Prior to loading into the flow chamber, samples were stained with a 20-nt readout probe (Dataset S5) labeled with Atto565, which has sequence complementarity to the readout sequence on the FISH probe for reporter gene mRNA imaging. Staining was performed in hybridization buffer [10% ethylene carbonate (Sigma, E26258) in 2x SSC] with a readout probe concentration of 3 nM. The readout probe for the reporter gene was introduced only once but was repeatedly imaged in all hybridization rounds. Readout probes for seven molecular targets (SON protein and six RNA targets) for phenotypic imaging and for barcode imaging were introduced into sequential rounds of hybridization. For phenotypic and barcode imaging, 3 nM of a 20-nt readout probe (Bio-Synthesis Inc., Dataset S5) complementary to the oligonucleotide sequence on the SON antibody (Abcam, ab121759), or the readout sequence on the FISH probes for six RNA targets, or the readout sequence on the secondary amplification probe for barcode imaging in hybridization buffer (10% ethylene carbonate in 2x SSC) was flowed into the chamber and left for 15 min, followed by a wash with hybridization buffer. The dyes for these probes, Alexa 488, Cy5, or Alexa 750, were linked to the oligo via a cleavable disulfide bond (Biosynthesis, Dataset S5).Samples were imaged in antibleach buffer (50 mL of antibleach buffer: 50 mg of glucoxidase (Sigma, G2133), 50 mg of (+ / -)-6-hydroxy-2,5,7,8-tetramethylchroman-2-carboxylic acid (Trolox) (Sigma, 238813), 300 microliters of catalase (Sigma, C100-C500MG), 10% w / v glucose (Sigma, G8270), 5 mL of 500 micromolar Trolox quinone, and 50 microliters of mouse RNase inhibitor). For each round of hybridization, fluorescent signals were imaged from four channels (488 nm, 561 nm, 647 nm, and 750 nm if phenotypic imaging was included in the round) or three channels (561 nm, 647 nm, and 750 nm if phenotypic imaging was not included in the round). After each round, the dye on the readout probe was cleaved with 10% tris(2-carboxyethyl)phosphine (TCEP; Sigma, 646547-10X1ML), followed by hybridization of the readout probe for the next round.
[0228] In all rounds, the 561 nm channel was used to detect reporter gene mRNA signals for quantification of colocalization rates between reporter gene and barcode signals and for image registration. In sequential hybridization rounds and rounds 1–18, the 647 nm and 750 nm channels were used for imaging with cleavable Cy5 and Alexa 750 dyes, allowing imaging of all 36 values of the 12-trit barcode. In sequential hybridization rounds and rounds 1–7, the 488 nm channel was used for imaging with cleavable Alexa 488 dye to measure signals for SON and six RNA targets for phenotypic imaging. For phenotypic imaging, images were collected at a slightly higher focal plane (2–3 micrometers), which is optimal for signal from within the nucleus.
[0229] For experiments measuring barcode identification errors using two known phenotypes (expression of reporter genes tagged with HA or Myc), the Myc and HA tags were stained with Alexa 405 and Alexa 488 dye-labeled secondary antibodies and imaged in the 405 nm and 488 nm channels, respectively.
[0230] DAPI staining was imaged and used for cell segmentation and nuclei identification. For experiments measuring two known phenotypes (reporter genes tagged with HA or Myc), DAPI staining was imaged in the last round of imaging. For experiments screening for factors that regulate RNA localization in the nucleus, DAPI staining was imaged in the first round of imaging. The sequences of the dye-labeled readout probes are listed in Dataset S5.
[0231] Inhibition of transcription: For transcription inhibition, 50 micromolar DRB (Sigma, D1916-10MG) was mixed in EMEM and incubated with cells for 1 hour before fixation.
[0232] Barcode decoding analysis: To correct for uneven illumination, all image intensities for a given color channel were divided by the average image intensity for all images for that illumination color. Images across multiple rounds were registered using uncleaved signals for reporter gene mRNA. Cells were segmented using a Watershed algorithm, which used DAPI staining as a seed and cellular autofluorescence (for experiments assessing barcode decoding accuracy and lentiviral recombination) or polyA-containing RNA staining (for experiments screening for factors that regulate RNA localization within the nucleus) to identify cell boundaries. A spot-finding algorithm was used to identify single-molecule signals for reporter gene mRNA and barcodes across all hybridizations. For experiments using four-camera imaging, a segmentation algorithm was used to identify spots. Specifically, pixels whose intensity was greater than a brightness threshold were selected. Clusters of selected pixels were identified using the bwareaopen function in MatLab. Clusters within the bounded region range (2–30 pixels) were retained. Region ranges were determined by visual inspection of raw images. To capture spots of various intensities, this process was repeated using multiple brightness thresholds, e.g., 0.6 × max (pixel intensity within FOV) to 0.3 × max (pixel intensity within FOV), with increments of 0.05 × max (pixel intensity within FOV). The brightness threshold for each trit signal was determined manually. For each iteration, the lower brightness threshold identified two types of clusters: (i) dim clusters that could not be detected at the upper brightness threshold from the previous round, and (ii) large clusters that completely contained one or more clusters identified from the previous round. Any cluster of type (i) was retained only if its area was within the acceptable area range described above.Any cluster of type (ii) was retained if its area was within the allowed region range; otherwise, it was discarded and instead the smaller cluster(s) identified from the previous round that overlapped with this new cluster were retained. The centers of these clusters were identified using the regionprops function in MatLab.
[0233] Single molecule FISH spots were assigned to cells, and the colocalization rate for each of the three trit values within the barcode was calculated as the number of reporter gene smFISH spots colocalizing with the barcode smFISH signal divided by the total number of reporter gene smFISH spots in the cell. To determine each trit value for each cell, cells were clustered based on the three colocalization rates of this trit using k-means clustering, and a trit value was assigned to each cluster based on which of the three average colocalization rates was the highest for this cluster. The same process was repeated for all 12 trits to assign each cell to 12 trit barcodes. For each trit value, the average colocalization rate for the cell population assigned this value was measured to be 0.4; whereas the average colocalization rate due to random colocalization with nonspecifically bound probes, assessed from the two cell populations that were not assigned this trit value, was measured to be 0.1.
[0234] To combine cells based on the barcode signal alone (i.e., without considering the co-localization of the barcode signal and the reporter gene signal), cells were clustered based on the number of barcode signal spots detected for the three trit values in each cell. Using a k-means clustering algorithm, cells were divided into three populations, and a trit value was assigned to each population based on which one of the three trit values had the highest average number of spots. This same process was repeated for all 12 trits.
[0235] To estimate the barcode misidentification rate, because only 0.4% of all possible barcodes were present in the library due to the bottlenecking strategy, there is only a 0.4% chance that any incorrectly decoded barcode will match a barcode in the library. Therefore, among the 57% of barcode matches, only about p = 0.3% could result from barcode misidentification [by solving (p × 57%) / (1-57%) = 0.4% / (1-0.4%)].
[0236] Quantification of Myc and HA signals: To quantify nuclear HA and Myc expression, the nuclear border of each cell was used as a mask for measuring the intensity of the corresponding Myc or HA channel. To enable unambiguous assignment of HA and Myc expression to individual cells, thresholds for HA and Myc expression above which HA or Myc tag expression could be reliably detected were first determined. To determine these thresholds, a k-means clustering algorithm was used to cluster cells into two groups based on the staining intensity of their unthresholded HA and Myc tags. This grouping allowed for approximate separation of cells into HA-expressing and Myc-expressing cells. To estimate the background staining level for the HA tag, the mean and standard deviation of the HA intensity values for cells within the Myc-expressing cluster were calculated, and the threshold for the HA signal was calculated as the mean + 3 standard deviations. The threshold for Myc expression was also determined from the HA-expressing cluster. Cells in which both HA and Myc intensities were below or above their respective thresholds were excluded (197 of 2336 cells). After removing these uncertain cells, the remaining cells were again clustered using the k-means algorithm to obtain the final grouping shown in Figure 2C.
[0237] Calculating recombination rate: Recombination rate a for the i-th sgRNA (with barcode or UMI) iThe calculation is as follows:
[0238] [ka] [where n is the number of days after transduction, which in these experiments was equal to 21 or 28. P iday 2 is the normalized protospacer reads (normalized by the total protospacer reads measured 2 days after transduction) for the i-th sgRNA, as determined by sequencing, at 2 days after transduction. i,day n B is the normalized protospacer reads (normalized by the total protospacer reads measured at day n after transduction) for the i-th sgRNA at day n after transduction, as determined by sequencing. i,day n is the normalized cell number determined by barcode imaging or the normalized UMI reads determined by sequencing (n days after transduction, normalized by total cell number or UMI reads). S i is the survival rate of the i-th sgRNA. C is the average survival rate for all sgRNAs in the library, calculated by considering the abundance weights of different sgRNAs, which is the average survival rate when recombination occurs.] Based on.
[0239] Quantification of phenotypic measurements: Nuclear boundaries were determined by DAPI signal. Cells whose nuclei contacted the edge of the imaging field were removed from further analysis. To identify clusters of MRP, pre-ribosomes, and SON, background intensities of the channels were subtracted, and clusters were identified using functions similar to the spot-finding algorithm for experiments with four-camera imaging: regionprops (MatLab) and bwareaopen (MatLab). Specifically, pixels whose intensity was greater than a brightness threshold were selected. Clusters of selected pixels were identified using the bwareaopen function. Clusters within the bounded region ranges (20–3,000 pixels for SON, 100–5,000 pixels for pre-ribosomes, and 100–6,000 pixels for MRP) were retained. Region ranges were determined by visual inspection of the raw images. To capture clusters with relatively widely varying staining levels, this process was repeated using multiple intensity thresholds (0.9 × maximum (nuclear pixel intensity) to 0.1 × maximum (nuclear pixel intensity) in increments of 0.05 × maximum (nuclear pixel intensity)). The number of finally identified clusters and regions for each cluster was measured using the regionprops function. For MRP, preribosomes, and SON, the number of clusters, average cluster area, and cluster intensity (defined as the total signal within the cluster boundary divided by the total cluster area) were calculated for each cell. To quantify enrichment of MALAT1, 7SK, U2, and poly(A)-containing RNAs within nuclear speckles, the cluster boundaries derived from SON staining were used as a mask for measuring MALAT1, 7SK, U2, and poly(A)-containing RNA signals within the SON cluster boundaries. The nuclear speckle intensity of each of these RNAs was measured as the total signal of that RNA within the SON cluster boundary divided by the total area covered by the SON cluster. The signal intensity outside the speckles was measured as the total signal of RNA within the nucleus but outside the nuclear speckles divided by the total area of the nucleus that was not within the nuclear speckles.Nuclear speckle enrichment was determined as the ratio of the intensity within the nuclear speckle to the signal intensity outside the speckle.
[0240] To identify hits from the screening, four experimental replicates were combined. The values described above (i.e., cluster number, cluster area, cluster intensity, and nuclear speckle enrichment) quantified for each replicate were normalized by the median value of all cells within each replicate before combination. A Student's t-test was used to calculate p-values by testing the values measured for cells carrying one targeting sgRNA against the values measured for all control, non-targeting sgRNA-carrying cells. If at least two sgRNAs targeting a particular gene showed a p-value of <0.05, the gene was listed as a hit. sgRNAs with fewer than 40 cells were removed from the analysis.
[0241] Dataset S1 sgRNA library assessing lentiviral designs to reduce recombination effects. This dataset lists oligo sequences for the protospacers of 159 sgRNAs and 51 non-targeting sgRNAs targeting essential ribosomal genes.
[0242] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8]
[0243] Dataset S2 This dataset lists the protospacer oligo sequences of 162 sgRNAs and five non-targeting sgRNAs targeting candidate genes selected to regulate nuclear RNA localization.
[0244] [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7]
[0245] Dataset S3 Traits quantified and gene hits identified in screens for factors regulating RNA localization within the nucleus. This dataset lists gene hits for each phenotypic trait quantified in screens.
[0246] [Table 3] Dataset S4 DNA oligo sequences used for library cloning and sequencing. This dataset lists the DNA oligo sequences used for library construction and next-generation sequencing (to identify barcodes and determine barcode-sgRNA correspondence within the library, as well as to quantify protospacer and UMI abundance in quantification of recombination).
[0247] [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5] [Table 4-6]
[0248] Dataset S5 FISH probe sequences for barcode imaging, reporter gene imaging, and phenotype imaging. This dataset contains the following separate lists of oligonucleotide probes: Primary amplification probes for barcode imaging Secondary amplification probes for barcode imaging FISH probes for poly(A)-containing RNA FISH probe for MALAT1 FISH probes for 7SK FISH probes for MRP FISH probes for preribosomes FISH probe for U2 snRNA FISH probes for the reporter genes puromycin-T2A-mCherry or mCherry-luciferase Oligonucleotide probes conjugated to antibodies for SON Readout probes for all of the above targets Some probes are modified, and the modification is indicated in the probe sequence.
[0249] [Table 5-1] [Table 5-2]
[0250] While several embodiments of the present invention have been described and illustrated herein, those skilled in the art will readily envision various other means and / or structures for performing the functions and / or obtaining one or more of the results and / or advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the present invention. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and that the actual parameters, dimensions, materials, and / or configurations will depend on the specific application or applications for which the teachings of the present invention are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Accordingly, the foregoing embodiments are presented by way of example only, and it should be understood that, within the scope of the appended claims and their equivalents, the invention may be practiced otherwise than as specifically described and claimed. The present invention is directed to each individual feature, system, item, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, items, materials, kits, and / or methods is also included within the scope of the present invention, to the extent that such features, systems, items, materials, kits, and / or methods are not mutually inconsistent.
[0251] In the event that the present specification and a document incorporated by reference include conflicting and / or inconsistent disclosure, the present specification shall control. In the event that two or more documents incorporated by reference include conflicting and / or inconsistent disclosure with respect to each other, the document having the later publication date shall control.
[0252] All definitions provided and used herein are to be understood to take precedence over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0253] The indefinite articles "a" and "an," as used in this specification and claims, unless indicated to the contrary, shall be understood to mean "at least one."
[0254] As used in the specification and claims, the term "and / or" shall be understood to mean "one or both" of the elements so conjoined, i.e., elements that are present conjunctively in some cases and disjunctively in other cases. Multiple elements listed with "and / or" shall be understood in the same manner, i.e., "one or more" of the elements so conjoined. The "and / or" clause indicates that other elements may be present other than the elements specifically identified, whether related or unrelated to the elements specifically identified. Thus, by way of non-limiting example, a reference to "A and / or B," when used in conjunction with open-ended language such as "comprising," may, in one embodiment, refer only to A (which may include elements other than B); in another embodiment, it may refer only to B (which may include elements other than A); in yet another embodiment, it may refer to both A and B (which may include other elements), etc.
[0255] As used herein and in the claims, "or" shall be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" shall be interpreted as inclusive, i.e., the inclusion of at least one of a number or list of elements, but also including more than one number or list, and may include additional, unlisted items. Only terms clearly indicated to the contrary, such as "only one of" or "exactly one of," or, when used in the claims, "consisting of," will refer to exactly one element of a number or list of elements. Generally, the term "or" as used herein shall be interpreted to indicate exclusive alternatives (i.e., "one or the other, but not both") only when preceded by terms of exclusivity, such as "either," "one of," "only one of," or "exactly one of."
[0256] The phrase "at least one," as used herein and in the claims, in reference to a list having one or more elements, shall be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed in the list of elements, and not excluding any combination of elements in the list of elements. This definition also allows for elements, whether related or unrelated to the specifically identified element, that may be present other than the elements specifically identified in the list of elements to which the phrase "at least one" refers. Thus, by way of non-limiting example, "at least one of A and B" (or, equivalently, "at least one of A or B," or, equivalently, "at least one of A and / or B") can refer, in one embodiment, to at least one A (and may include elements other than B), which may include more than one A, but no B; in another embodiment, to at least one B (and may include elements other than A), which may include more than one B, but no A; in yet another embodiment, to at least one A, which may include at least one A and more than one B, which may include at least one B (and may include other elements), etc.
[0257] When the word "about" is used herein in reference to a number, it should be understood that further embodiments of the present invention also include numbers that are not modified by the word "about."
[0258] It is also to be understood that, unless expressly indicated to the contrary, in any method claimed herein that includes more than one step or act, the order of the method steps or acts is not necessarily limited to the order of the method steps or acts recited.
[0259] In the claims, as well as in the above specification, all transitional phrases such as "comprising," "including," "holding," "having," "containing," "with," "holding," "consisting of," etc., shall be understood to be open-ended, i.e., meaning "including, but not limited to." Only the transitional phrases "consisting of" and "consisting essentially of" shall be closed or semi-closed transitional phrases, respectively, as expressly set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.
Claims
1. (a) introducing into a plurality of cells DNA comprising a guide portion comprising a recognition sequence, a reporter portion, and an identification portion comprising a lead sequence; (b) determining the location of the RNA molecules expressed from the reporter moiety of the introduced DNA in the plurality of cells by determining the reporter moiety; (c) determining a lead sequence on an RNA molecule expressed from DNA comprising a reporter moiety and an identifying moiety introduced into the plurality of cells by exposing the cells to a readout probe capable of binding to the lead sequence; (d) co-localizing the binding of the readout probe with the location of the RNA molecule expressed from the reporter portion of the introduced DNA; (e) repeating (b), (c), and (d) multiple times using different read sequences; and (f) creating a code word corresponding to the binding of the co-localized readout probes, the code word's numerical value being based on the binding of the readout probe to the read sequence. A method comprising:
2. 10. The method of claim 1, further comprising identifying a guide moiety for each individual cell based on the measured code word.
3. The method of claim 1 or 2, wherein the recognition sequence recognizes a DNA sequence.
4. The method of claim 1 , wherein the recognition sequence recognizes an RNA sequence.
5. 5. The method of claim 1, wherein the introduced DNA originates from a library of nucleic acids.
6. The method of claim 5, wherein the library is generated by pooled cloning.
7. 7. The method of claim 1, wherein the identity of pairs of guide moieties and identifying moieties associated on the DNA is determined by sequencing.
8. 8. The method of claim 1, wherein the introduced DNA allows association of the guide moiety with the identification moiety to occur.
9. The method of claim 1 , wherein each lead sequence represents a value for a position within a codeword.
10. 10. The method of claim 1, wherein the lead sequences are determined sequentially.
11. 11. The method of any one of claims 1 to 10, wherein the guide portion further comprises a Cas protein binding sequence.
12. 12. The method of any one of claims 1 to 11, wherein the guide moiety enables targeting of the Cas protein to DNA or RNA to perturb the sequence or expression of the gene.
13. 13. The method of any one of claims 1 to 12, wherein the reporter moiety encodes a protein that is detectable by fluorescence.
14. 14. The method of any one of claims 1 to 13, wherein the reporter moiety encodes a fluorescent protein.
15. 15. The method of any one of claims 1 to 14, wherein the reporter moiety encodes a luciferase.
16. 16. The method of any one of claims 1 to 15, wherein the reporter moiety encodes a protein that is detectable by immunoprecipitation.
17. 17. The method of any one of claims 1 to 16, wherein the reporter moiety encodes a protein that is detectable by immunofluorescence.
18. 18. The method of any one of claims 1 to 17, wherein the reporter moiety encodes a Myc tag.
19. 19. The method of any one of claims 1 to 18, wherein the reporter moiety encodes an HA tag.
20. 20. The method of any one of claims 1 to 19, wherein the reporter moiety comprises a reporter gene.
21. 21. The method of any one of claims 1 to 20, wherein the identifying portion is present within the 3'UTR of the reporter gene.
22. 22. The method of any one of claims 1 to 21, wherein the reporter moiety comprises a first promoter.
23. 23. The method of claim 22, wherein the promoter is a promoter that drives transcription.
24. 24. The method of claim 22 or 23, wherein the promoter comprises a CMV promoter.
25. 25. The method of any one of claims 22 to 24, wherein the recognition sequence comprises a second promoter separated from the first promoter.
26. 26. The method of any one of claims 1 to 25, wherein the recognition sequence and the reporter moiety are separated by less than 1000 bases.
27. 27. The method of any one of claims 1 to 26, wherein the recognition sequence in the DNA and the reporter moiety are separated by less than 100 bases.
28. 28. The method of any one of claims 1 to 27, comprising determining one or more lead sequences using fluorescence.
29. 29. The method of any one of claims 1 to 28, comprising using smFISH targeting the reporter moiety to determine the location of RNA comprising the reporter moiety and the identifying moiety.
30. 30. The method of any one of claims 1 to 29, comprising using a virus to introduce DNA into a plurality of cells.
31. 31. The method of claim 30, wherein the virus is a lentivirus.
32. The method of claim 31 , wherein the guide portion and the identification portion are located adjacent to the 3′ end of a polypurine tract (PPT) sequence within the lentivirus.
33. 33. The method of claim 32, wherein the guide portion is replicated within the 5' region of the lentivirus.
34. 34. The method of any one of claims 1 to 33, wherein introducing the DNA into the plurality of cells comprises electroporating the DNA into the plurality of cells.
35. 35. The method of any one of claims 1 to 34, wherein the cells comprise cells within a tissue.
36. 36. The method of any one of claims 1 to 35, comprising introducing DNA into a plurality of cells such that at least 50% of the cells do not contain more than one type of introduced DNA.
37. 37. The method of any one of claims 1 to 36, comprising introducing DNA into a plurality of cells such that at least 90% of the cells do not contain more than one type of introduced DNA.
38. 38. A method according to any one of claims 1 to 37, comprising introducing the DNA into the genome of the cell.
39. 39. The method of any one of claims 1 to 38, wherein the DNA further comprises a promoter.
40. 40. The method of any one of claims 1 to 39, wherein the lead sequence defines a binary space of values for the codeword.
41. 41. The method of claim 1, wherein the lead sequence defines a ternary space of values for the codeword.
42. Determining one or more lead sequences For each value of the code word, a readout probe corresponding to the value of the code word is applied to a plurality of cells.
42. The method of any one of claims 1 to 41, comprising:
43. 43. The method of claim 42, comprising sequentially applying and removing each of the readout probes to a plurality of cells.
44. 44. The method of any one of claims 1 to 43, wherein for at least some of the created code words, the code words are matched to valid code words, and if no match is found, the code word is rejected or error correction is applied to the code word to form a valid code word.
45. 45. The method of claim 44, wherein the numerical values of the code words define a set of potential code words.
46. The set of potential code words is at least 10 2 46. The method of claim 45, comprising unique sequences.
47. The set of potential code words is at least 10 3 47. The method of claim 46, comprising unique sequences.
48. The set of potential code words is at least 10 4 48. The method of claim 46 or 47, comprising unique sequences.
49. The set of potential code words is at least 10 5 49. The method of any one of claims 46 to 48, comprising unique sequences.
50. The set of potential code words is at least 10 6 50. The method of any one of claims 46 to 49, comprising unique sequences.
51. 51. The method of any one of claims 44 to 50, wherein a portion of the potential code words are valid code words.
52. 52. The method of claim 51, comprising comparing the measured codeword with valid codewords to determine an error in the codeword measurement.
53. 53. A method according to any one of claims 44 to 52, wherein the valid code words are a randomly selected subset of the possible code words.
54. 54. The method of claim 53, wherein less than 10% of the possible code words are valid code words.
55. 55. The method of claim 53 or 54, wherein less than 5% of the possible code words are valid code words.
56. 56. The method of any one of claims 53 to 55, wherein less than 1% of the possible code words are valid code words.
57. 57. The method of any one of claims 53 to 56, wherein less than 0.5% of the possible code words are valid code words.
58. 58. The method of any one of claims 1 to 57, wherein the identifying portion comprises N (N) variable portions, where N is at least 3, and each variable portion has at least two possibilities.
59. 59. The method of claim 58, wherein N is at least 5.
60. 60. The method of claim 58 or 59, wherein N is at least 10.
61. 61. The method of any one of claims 58 to 60, wherein N is at least 15.
62. 62. The method of any one of claims 58 to 61, wherein N is at least 20.
63. 63. The method of any one of claims 1 to 62, wherein the variable portions within the identifying portion are each the same length.
64. 64. The method of any one of claims 1 to 63, wherein the variable portions within the identifying portion each have a length between 5 and 50 nt.
65. 65. The method of any one of claims 1 to 64, wherein the variable portions within the identifying portion each have a length of between 15 and 25 nt.
66. 66. The method of any one of claims 1 to 65, wherein the identification portion comprises an error detecting code.
67. 67. The method of any one of claims 1 to 66, wherein the identification portion comprises an error correcting code.
68. 68. A method according to claim 66 or 67, wherein the error detecting or correcting code comprises a Hamming code.
69. 69. A method according to any one of claims 66 to 68, wherein the error detecting or correcting code comprises an extended Hamming code.
70. 70. A method according to any one of claims 66 to 69, wherein the error detecting or correcting code comprises a Reed-Solomon code.
71. 71. A method according to any one of claims 66 to 70, wherein the error detecting or correcting code is not uniform for all code members.
72. 72. The method of any one of claims 1 to 71, wherein the nucleic acid probe comprises at least eight possible lead sequences.
73. 73. The method of any one of claims 1 to 72, wherein the nucleic acid probe comprises at least 16 possible lead sequences.
74. 74. The method of any one of claims 1 to 73, wherein the nucleic acid probe comprises no more than 32 possible lead sequences.
75. 75. The method of any one of claims 1 to 74, wherein the nucleic acid probe comprises at least 32 possible lead sequences.
76. introducing into the plurality of cells DNA comprising a guide portion comprising a recognition sequence, a reporter portion, and an identification portion comprising a lead sequence; determining the location of the RNA molecule expressed from the reporter portion of the introduced DNA in the plurality of cells by determining the reporter portion; determining the lead sequence in a plurality of cells by exposing the cells to a plurality of readout probes, each capable of binding to a lead sequence; Co-localizing the binding of the readout probe with the location of the RNA molecule expressed from the reporter portion of the introduced DNA; and creating a code word corresponding to the binding of the co-localized readout probes, the code word's numerical value being based on the binding of the readout probe to the read sequence; A method comprising:
77. introducing DNA into a plurality of cells using a lentivirus, wherein the DNA comprises a guide portion comprising a recognition sequence, a reporter portion, and an identification portion comprising a lead sequence; Determining the phenotype of the plurality of cells; and determining the genotype of the plurality of cells; and Determining the correspondence between genotype and phenotype A method comprising:
78. 78. The method of claim 77, wherein determining the phenotype comprises determining a cellular characteristic.
79. 79. The method of claim 77 or 78, wherein determining the phenotype comprises determining the shape of a cellular structure.
80. 80. The method of claim 79, wherein the shape is a whole cell shape.
81. 81. The method of claim 79 or 80, wherein the shape is a subcompartment shape.
82. 82. The method of any one of claims 77 to 81, wherein determining the phenotype comprises determining the shape of more than one cellular structure.
83. 83. The method of any one of claims 77 to 82, wherein determining the phenotype comprises determining the protein using immunofluorescence.
84. 84. The method of any one of claims 77 to 83, wherein determining the phenotype comprises determining the protein using fluorescence.
85. 85. The method of any one of claims 77 to 84, wherein determining the phenotype comprises determining the protein using a fluorescent protein.
86. 86. The method of any one of claims 77 to 85, wherein determining the phenotype comprises determining the protein using an organic dye.
87. 87. The method of any one of claims 77 to 86, wherein determining the phenotype comprises determining the dynamic behavior of the cell.
88. 88. The method of any one of claims 77 to 87, wherein determining the phenotype comprises determining cell-cell interactions.
89. 89. The method of any one of claims 77 to 88, wherein determining the phenotype comprises determining a cell state.
90. 90. The method of any one of claims 77-89, wherein determining the phenotype comprises determining one or more RNAs of the cell using smFISH.
91. 91. The method of any one of claims 77 to 90, wherein determining the phenotype comprises determining a gene expression profile using multiplexed FISH.
92. 92. The method of any one of claims 77 to 91, wherein determining the phenotype comprises determining a gene expression profile using MERFISH.
93. 93. The method of any one of claims 77 to 92, wherein determining the phenotype comprises spatially determining one or more RNAs.
94. 94. The method of any one of claims 77 to 93, wherein determining the phenotype comprises determining at least a portion of the proteome.
95. 95. The method of any one of claims 77 to 94, wherein determining the phenotype comprises determining at least a portion of a chromosome using DNA FISH.
96. 96. The method of any one of claims 77 to 95, wherein determining the phenotype comprises determining at least a portion of a chromosome using multiplexed DNA FISH.
97. 97. The method of any one of claims 77 to 96, wherein determining the phenotype comprises determining at least a portion of a chromosome using CASFISH.
98. 98. The method of any one of claims 77 to 97, wherein determining the phenotype comprises using immunofluorescence to determine protein-modified cells.
99. 99. The method of any one of claims 77 to 98, wherein determining the phenotype comprises determining the interaction of the protein with RNA or DNA.
100. 100. The method of any one of claims 77 to 99, wherein determining the phenotype comprises determining an epigenetic modification.
101. 101. The method of any one of claims 77 to 100, wherein determining the phenotype comprises determining cell proliferation.
102. 102. The method of any one of claims 77 to 101, wherein determining the phenotype comprises determining a change in cell proliferation.
103. 103. The method of any one of claims 77 to 102, wherein determining the phenotype and determining the genotype use a common imaging method.
104. 104. The method of any one of claims 77 to 103, wherein determining the phenotype and determining the genotype use different imaging methods.
105. 105. The method of claim 103 or 104, wherein the imaging method has a resolution of better than 300 nm.
106. 106. The method of any one of claims 103 to 105, wherein determining the phenotype uses multi-color fluorescence imaging.
107. 107. The method of any one of claims 103 to 106, wherein determining the phenotype uses confocal imaging.
108. 108. The method of any one of claims 103 to 107, wherein determining the phenotype uses TIRF imaging.
109. 109. The method of any one of claims 103 to 108, wherein determining the phenotype uses two-photon imaging.
110. 110. The method of any one of claims 103 to 109, wherein determining the phenotype uses STORM.
111. 111. The method of any one of claims 103 to 110, wherein determining the phenotype uses super-resolution methods.
112. 112. The method of claim 111, wherein the super-resolution method is PALM, FPALM, STED, SIM, and / or RESOLFT.
113. 113. The method of any one of claims 77 to 112, wherein determining the genotype comprises determining the sequence of an identifying portion of the introduced DNA.
114. 114. The method of any one of claims 77 to 113, wherein determining the genotype comprises determining the genotype using smFISH.
115. 115. The method of any one of claims 77 to 114, wherein determining the genotype comprises determining the genotype using multiplexed FISH.
116. 116. The method of any one of claims 77 to 115, wherein determining the genotype comprises determining the genotype using MERFISH.
117. 117. The method of any one of claims 77 to 116, wherein determining the genotype comprises determining the genotype using in situ hybridization.
118. 118. The method of any one of claims 77 to 117, wherein determining the genotype comprises determining the genotype using sequential FISH.
119. 119. The method of any one of claims 77 to 118, wherein determining the genotype comprises determining the genotype using CASFISH.
120. 120. The method of any one of claims 77 to 119, wherein determining the genotype comprises determining the genotype using in situ sequencing.
121. introducing into a plurality of cells a nucleic acid, the nucleic acid comprising a guide portion comprising a recognition sequence, a reporter portion, and an identifying portion comprising a lead sequence; imaging a plurality of cells, wherein the cells exhibit an imageable phenotypic difference due to expression of the guide moiety; and collecting a plurality of images of a plurality of cells, the images of the cells exhibiting differences due to differences in identifying portions of nucleic acids within the cells; A method comprising:
122. 122. The method of claim 121, wherein the guide moiety introduced into the plurality of cells is identified based on the correspondence between the guide moiety and the identification moiety.
123. 123. The method of claim 121 or 122, wherein cells with different phenotypes have different appearances when imaged.
124. 124. The method of any one of claims 121 to 123, wherein cells with different phenotypes have different fluorescence when imaged.
125. 125. The method of any one of claims 121 to 124, wherein imaging a plurality of cells comprises collecting images for the plurality of cells using a single imaging modality.
126. 126. The method of any one of claims 121 to 125, wherein imaging a plurality of cells comprises collecting images for the plurality of cells using a plurality of imaging modalities.
127. 127. The method of any one of claims 121 to 126, wherein imaging a plurality of cells comprises acquiring a single image.
128. 128. The method of any one of claims 121 to 127, wherein imaging a plurality of cells comprises acquiring a plurality of images.
129. 129. The method of any one of claims 121 to 128, wherein imaging the plurality of cells comprises imaging the plurality of cells using smFISH.
130. 130. The method of any one of claims 121 to 129, wherein imaging the plurality of cells comprises imaging the plurality of cells using MERFISH.
131. 131. The method of any one of claims 121 to 130, further comprising determining the shape of the cells.
132. 132. The method of claim 131, comprising determining changes in cell shape over time.
133. 133. The method of claim 131 or 132, comprising determining changes in the shape of organelles of a cell.
134. 134. A method according to any one of claims 131 to 133, comprising determining shape changes during cell proliferation.
135. introducing DNA into a plurality of cells using a lentivirus, wherein the DNA comprises a guide portion comprising a recognition sequence and an identification portion comprising a lead sequence; determining the phenotype of multiple cells; determining the genotype of the plurality of cells; and Determining the correspondence between genotype and phenotype A method comprising:
136. 136. The method of claim 135, wherein determining the phenotype comprises determining a cellular characteristic.
137. 137. The method of claim 135 or 136, wherein determining the phenotype comprises determining the shape of a cellular structure.
138. 138. The method of claim 137, wherein the shape is a whole cell shape.
139. 139. The method of claim 137 or 138, wherein the shape is a subcompartment shape.
140. 140. The method of any one of claims 137 to 139, wherein determining the phenotype comprises determining the shape of more than one cellular structure.
141. 141. The method of any one of claims 135 to 140, wherein determining the phenotype comprises determining the protein using immunofluorescence.
142. 142. The method of any one of claims 135 to 141, wherein determining the phenotype comprises determining the protein using fluorescence.
143. 143. The method of any one of claims 135 to 142, wherein determining the phenotype comprises determining the protein using a fluorescent protein.
144. 144. The method of any one of claims 135 to 143, wherein determining the phenotype comprises determining the protein using an organic dye.
145. 145. The method of any one of claims 135 to 144, wherein determining the phenotype comprises determining the dynamic behavior of the cell.
146. 146. The method of any one of claims 135 to 145, wherein determining the phenotype comprises determining cell-cell interactions.
147. 147. The method of any one of claims 135 to 146, wherein determining the phenotype comprises determining a cell state.
148. 148. The method of any one of claims 135 to 147, wherein determining the phenotype comprises determining one or more RNAs using smFISH.
149. 149. The method of any one of claims 135 to 148, wherein determining the phenotype comprises determining a gene expression profile using multiplexed FISH.
150. 150. The method of any one of claims 135 to 149, wherein determining the phenotype comprises determining a gene expression profile using MERFISH.
151. 151. The method of any one of claims 135 to 150, wherein determining the phenotype comprises spatially determining one or more RNAs.
152. 152. The method of any one of claims 135 to 151, wherein determining the phenotype comprises determining at least a portion of the proteome.
153. 153. The method of any one of claims 135 to 152, wherein determining the phenotype comprises determining at least a portion of a chromosome using DNA FISH.
154. 154. The method of any one of claims 135 to 153, wherein determining the phenotype comprises determining at least a portion of a chromosome using multiplexed DNA FISH.
155. 155. The method of any one of claims 135 to 154, wherein determining the phenotype comprises determining at least a portion of a chromosome using CASFISH.
156. 156. The method of any one of claims 135 to 155, wherein determining the phenotype comprises using immunofluorescence to determine protein-modified cells.
157. 157. The method of any one of claims 135 to 156, wherein determining the phenotype comprises determining the interaction of the protein with RNA or DNA.
158. 158. The method of any one of claims 135 to 157, wherein determining the phenotype comprises determining an epigenetic modification.
159. 159. The method of any one of claims 135 to 158, wherein determining the phenotype comprises determining cell proliferation.
160. 160. The method of any one of claims 135 to 159, wherein determining the phenotype comprises determining a change in cell proliferation.
161. 161. The method of any one of claims 135 to 160, wherein determining the phenotype and determining the genotype use a common imaging method.
162. 162. The method of claim 161, wherein the imaging method has a resolution better than 300 nm.
163. 161. The method of any one of claims 135 to 160, wherein determining the phenotype and determining the genotype use different imaging methods.
164. 164. The method of claim 163, wherein at least one of the imaging methods has a resolution better than 300 nm.
165. 165. The method of any one of claims 135 to 164, wherein determining the phenotype uses multi-color fluorescence imaging.
166. 166. The method of any one of claims 135 to 165, wherein determining the phenotype uses confocal imaging.
167. 167. The method of any one of claims 135 to 166, wherein determining the phenotype uses TIRF imaging.
168. 168. The method of any one of claims 135 to 167, wherein determining the phenotype uses two-photon imaging.
169. 169. The method of any one of claims 135 to 168, wherein determining the phenotype uses STORM.
170. 170. The method of any one of claims 135 to 169, wherein determining the phenotype uses super-resolution methods.
171. 171. The method of any one of claims 135 to 170, wherein the super-resolution method is PALM, FPALM, STED, SIM, and / or RESOLFT.
172. 172. The method of any one of claims 135 to 171, wherein determining the genotype comprises determining the sequence of an identifying portion of the introduced DNA.
173. 173. The method of any one of claims 135 to 172, wherein determining the genotype comprises determining the genotype using smFISH.
174. 174. The method of any one of claims 135 to 173, wherein determining the genotype comprises determining the genotype using multiplexed FISH.
175. 175. The method of any one of claims 135 to 174, wherein determining the genotype comprises determining the genotype using MERFISH.
176. 176. The method of any one of claims 135 to 175, wherein determining the genotype comprises determining the genotype using in situ hybridization.
177. 177. The method of any one of claims 135 to 176, wherein determining the genotype comprises determining the genotype using sequential FISH.
178. 178. The method of any one of claims 135 to 177, wherein determining the genotype comprises determining the genotype using CASFISH.
179. 179. The method of any one of claims 135 to 178, wherein determining the genotype comprises determining the genotype using in situ sequencing.