Synthetic barcoding of cell line background genetics
Patent Information
- Application Number
- JP2023550166
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-02-18
- Filing Date
- 2022-02-17
- Publication Date
- 2026-03-02
AI Technical Summary
Current pooled screening approaches are limited by their reliance on sequencing data for genetic identification, which restricts the number of assays that can be performed and are not suitable for large population studies due to scaling challenges and environmental artifacts.
A method involving labeling cells from different genetic backgrounds with unique nucleic acid barcode sequences, combining them into a mixed population, and performing in situ single-cell sequencing to analyze phenotypes, allowing for controlled, multiplexed screening of cells with similar or engineered backgrounds.
Enables large-scale, controlled population genetics-based assays with reduced environmental artifacts, enabling statistical analysis and multiplexing of cell lines without the limitations of sequencing-only readouts.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Application No. 63 / 150,979, entitled "SYNTHETIC BARCODING OF CELL LINE BACKGROUND GENETICS," filed February 18, 2021, the contents of which are incorporated by reference in their entirety for all purposes.
[0002] [Technical field] The present invention relates to a method for pooled screening of cells from different genetic backgrounds.The present invention also relates to a computer-implemented method for registration between a first plurality of images and a second plurality of images. [Background technology]
[0003] In cell-based models, the background genetics of cell lines can have a substantial impact on cell behavior, phenotype, and disease state. Approaches that utilize cell pools have the potential to allow population-based studies to be performed in vitro. However, currently available pooled screening approaches have limitations.
[0004] Some current pooled screening approaches rely on the background genetics of cells as their "barcode", meaning that the assays may ultimately use sequencing as their final readout. This typically comes in the form of genomic DNA (gDNA) sequencing of sorted cells, or as demuxlet analysis of single-cell RNAseq data, allowing genotype identification based on gene variants at the 3' end of single-cell transcripts. This reliance on sequencing data severely limits the number of assays that can be performed in this pooled format.
[0005] An alternative pooled approach utilizes optical barcoding. Current pooled optical barcoding and CRISPR screening strategies generally implement a single cell line containing a variant of Cas9 along with the incorporation of a unique padlock-flanked, otherwise sequenceable barcode. In this way, the perturbing gRNA can be identified via in situ sequencing. Although a substantial technical feature, this method can only utilize a single genetic background at a time.
[0006] Using the methods of pooled screening of cells from different genetic backgrounds provided herein, individual cell lines derived from various healthy and / or patient strains are labeled with unique genetic barcodes designed to be processed via in situ sequencing after any assay. The phenotype can then be classified and linked back to the genotype of the original cell line, allowing statistical genetic analysis to be performed in a highly controlled in vitro assay. This technique also allows pooled screening studies to be applied to cell lines with similar background genetics (i.e., trios and genetically engineered strains) as well as multiplexed with perturbation-based genetic screens. Also provided herein are machine learning-enabled improvements to the genetic barcoding method.
[0007] The methods provided herein solve the problem of being unable to perform large population studies in vitro without facing the challenges of difficult, costly, and artifact-sensitive scaling. The pooling approach significantly reduces the effects of evaporation, temperature, and other plate- or well-like artifacts that can obscure larger screens. These methods also allow population genetics-based assays to be performed on specific cell types in a highly controlled manner, rather than clinical outcomes that are highly variable and influenced by a myriad of convoluting factors. Other methods such as those mentioned above include either sequencing-only readouts or barcoding of perturbagens (pooled optical barcoding), which eliminates the importance of non-coding variants in a given assay.
[0008] All references cited herein, including patent applications, patent publications, and scientific literature, are incorporated by reference in their entirety as if each individual reference was specifically and individually indicated to be incorporated by reference. Summary of the Invention
[0009] Provided herein is a pooled screening method for cells from different genetic backgrounds, comprising: a) labeling two or more populations of cells of different genetic backgrounds with two or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells; b) combining the two or more populations of cells to obtain a single mixed population of cells; c) performing in situ single cell sequencing on the cells; and d) analyzing known phenotypes or identifying new phenotypes of the cells in the mixed population. In some embodiments, the two or more populations of cells are from different cell lines. In some embodiments, the different cell lines are healthy cell lines. In some embodiments, the different cell lines are patient cell lines. In some embodiments, the different cell lines are isogenically engineered cell lines. In some embodiments, the different cell lines include any combination of healthy cell lines, patient cell lines, and isogenically engineered cell lines. In some embodiments, the cells are induced pluripotent stem cells (iPSCs). In some embodiments, the iPSCs are differentiated prior to analyzing the phenotype. In some embodiments, the method further comprises culturing the cells prior to analyzing the phenotype. In some embodiments, the single mixed population of cells is on a substrate or in a three-dimensional culture. In some embodiments, the substrate is a cell culture dish. In some embodiments, the method further comprises performing single cell RNAseq. In some embodiments, the method further comprises, prior to step b), growing the two or more populations of cells for two or more generations. In some embodiments, the method comprises stably integrating the unique nucleic acid barcode sequence into the genome of the two or more populations of cells. In some embodiments, the unique nucleic acid barcode sequence is delivered into the cells using a virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus encodes a selection marker. In some embodiments, the selection marker is an antibiotic resistance gene. In some embodiments, the virus encodes a fluorescent protein. In some embodiments, each unique nucleic acid barcode sequence is at least one base pair in length.In some embodiments, each unique nucleic acid barcode sequence is 1 to about 18 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 8 base pairs in length. In some embodiments, the two or more populations of cells are sequenced prior to labeling with the unique nucleic acid barcode sequences. In some embodiments, the sequencing is whole genome sequencing. In some embodiments, the two or more populations of cells are obtained from related individuals. In some embodiments, the two or more populations of cells are obtained from a human. In some embodiments, the method comprises labeling 10 or more populations of cells of different genetic backgrounds with 10 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells. In some embodiments, analyzing the phenotype of the cells comprises an assay selected from the group consisting of high content imaging, calcium imaging, immunohistochemistry, cell morphology imaging, protein aggregation imaging, cell-cell interaction imaging, live cell imaging, and any other image-based assay modality. In some embodiments, step d) comprises analyzing the phenotype of the cells by capturing a microscopic image or a time series of microscopic images of the cells and evaluating the phenotypic features exhibited in the image or the plurality of images. In some embodiments, the method further includes a computer-implemented method for registration between the first and second multiple images, comprising: a) generating a first coordinate space of reference of a first multiple images of a well on a culture plate; b) extracting a first patch of the first multiple images; c) generating a second coordinate space of reference of a second multiple images of a well on a culture plate; d) extracting a second patch of the second multiple images; e) calculating an affine transformation function between the first and second patches to obtain a plurality of transformation parameters; and f) generating a coordinate transformation function between the first and second coordinate spaces of reference based on the plurality of transformation parameters.
[0010] Further provided herein is a computer-implemented method for registration between a first and a second plurality of images, comprising: generating a first coordinate space of reference of a first plurality of images of a well on a culture plate; extracting a first patch of the first plurality of images; generating a second coordinate space of reference of a second plurality of images of the well on the culture plate; extracting a second patch of the second plurality of images; calculating an affine transformation function between the first and second patches to obtain a plurality of transformation parameters; and generating a coordinate transformation function between the first and second coordinate spaces of reference based on the plurality of transformation parameters. In some embodiments, the first plurality of images are a plurality of bar-coding images. In some embodiments, the second plurality of images are a plurality of marker-based / marker-free readout images. In some embodiments, the first and second plurality of images provide different coverages of the well. In some embodiments, the first and second plurality of images are captured at different times. In some embodiments, the first and second plurality of images have different resolutions. In some embodiments, the first plurality of images are captured by a first microscope and the second plurality of images are captured by a second microscope. In some embodiments, the first microscope is a fluorescent microscope. In some embodiments, the second microscope is a non-fluorescent microscope. In some embodiments, the first plurality of images are captured by a first imager. In some embodiments, the method further includes detecting one or more physical characteristics of the well in the first plurality of images and generating a first coordinate space of reference based on the detected one or more physical characteristics. In some embodiments, the one or more physical characteristics of the well include a shape of the well, an edge of the well, and a position of the well. In some embodiments, two or more images in the first plurality of images are offset from each other by an overlap ratio, and the first coordinate space of reference is generated based on the overlap ratio. In some embodiments, the first coordinate space of reference is generated based on metadata of the first imager.In some embodiments, the method further includes selecting one or more marker images from the first plurality of images based on one or more landmarks captured in the one or more images, where the first patch is obtained from one or more marker images from the first plurality of images. In some embodiments, the one or more landmarks include one or more cells, one or more well borders, one or more beads, one or more nuclei, or any combination thereof. In some embodiments, the second plurality of images is captured by a second imager. In some embodiments, the method further includes detecting one or more physical characteristics of the well in the second plurality of images and generating a second coordinate space of reference based on the detected one or more physical characteristics. In some embodiments, the one or more physical characteristics of the well include a shape of the well, an edge of the well, and a position of the well. In some embodiments, two or more images in the second plurality of images are offset from each other by an overlap ratio, and the second coordinate space of reference is generated based on the overlap ratio. In some embodiments, the second coordinate space of reference is generated based on metadata of the second imager. In some embodiments, the method further includes selecting one or more marker images of a second plurality of images based on one or more landmarks captured in the one or more images, the second patch being obtained from one or more marker images from the second plurality of images. In some embodiments, the one or more landmarks include one or more cells, one or more well boundaries, one or more beads, one or more nuclei, or any combination thereof. In some embodiments, the first patch covers a center of the first image. In some embodiments, the second patch covers a center of the second image. In some embodiments, at least a portion of the first patch and at least one patch of the second patch correspond to the same object. In some embodiments, the transformation parameters include one or more of a translation parameter, a scaling parameter, and a rotation parameter.In some embodiments, the first and second images are obtained from a pooled screening assay of cells from different genetic backgrounds comprising: a) labeling two or more populations of cells of different genetic backgrounds with two or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells; b) combining the two or more populations of cells to obtain a single mixed population of cells; c) performing in situ single cell sequencing on the cells; and d) analyzing known phenotypes or identifying new phenotypes of the cells in the mixed population.
[0011] In some embodiments, the method further comprises utilizing a classifier configured to receive the image and output a classification result. In some embodiments, the classifier comprises multiple layers. In some embodiments, the classifier is a convolutional neural network. In some embodiments, the classifier is a DenseNet classifier. In some embodiments, the classification result is a classification based on the genetic background of each single cell captured in the image. In some embodiments, the method further comprises generating an embedding of the image. In some embodiments, the embedding is generated from an activation output layer before the last layer of the classifier. In some embodiments, the embedding is dimensionally reduced using a linear dimensionality reduction method such as the Uniform Manifold Approximation and Projection for Dimension Reduction (UMAP) algorithm, Principal Component Analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t-SNE), etc. In some embodiments, the embedding is dimensionally reduced using a UMAP algorithm to obtain one or more UMAP plots for visualization of the embedding. In some embodiments, the method further comprises evaluating the treatment based at least in part on the embedding. In some embodiments, the method further comprises evaluating the treatment based at least in part on the one or more UMAP plots.
[0012] It should be understood that one, some, or all of the features of the various embodiments described herein may be combined to form other embodiments of the present invention. These and other aspects of the present invention will become apparent to those skilled in the art. These and other embodiments of the present invention are further described in the following detailed description. [Brief description of the drawings]
[0013] This patent or application contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Patent and Trademark Office upon request and payment of the necessary fee.
[0014] [Figure 1] 1 shows an example of a vector used to introduce a unique nucleic acid barcode sequence into a cell. Features flanking the unique nucleic acid barcode sequence are shown. POSH = Pooled Optical Screening in Human Cells.
[0015] [Diagram 2] FIG. 1 is a schematic diagram showing a first modality using the method of pooled screening of cells from different genetic backgrounds described in this application. In some embodiments, the modality involves performing ultra-throughput drug screening. POSH=Pooled Optical Screening in Human Cells.
[0016] [Diagram 3] 1 is a schematic diagram showing a second modality using the method of pooled screening of cells from different genetic backgrounds described in this application. In some embodiments, the modality involves performing ultra-throughput drug screening. POSH = Pooled optical screening in human cells.
[0017] [Figure 4]FIG. 1 is a schematic diagram showing a third modality that uses the method of pooled screening of cells from different genetic backgrounds described in this application, which in some embodiments involves performing ultra-throughput drug screening.
[0018] [Diagram 5] FIG. 1 is a schematic diagram showing the analytical pipeline used to analyze data obtained from the method of pooled screening of cells from different genetic backgrounds described in this application.
[0019] [Figure 6] 1 illustrates an exemplary computer-implemented process for alignment between a first plurality of images (e.g., from a first image acquisition) and a second plurality of images (e.g., from a second image acquisition), according to some embodiments.
[0020] [Figure 7] 1 depicts an exemplary plate designed for microscopy acquisition, according to some embodiments.
[0021] [Figure 8] 1 illustrates multiple images of a well, according to some embodiments.
[0022] [Figure 9] 1 illustrates multiple images of a well, according to some embodiments.
[0023] [Figure 10] 1 shows two example patches from two different image captures, according to some embodiments.
[0024] [Figure 11] 4 illustrates an example process for calculating transformation parameters, according to some embodiments.
[0025] [Figure 12]1 illustrates an exemplary electronic device according to some embodiments.
[0026] [Figure 13A] , and [Figure 13B] 13A shows exemplary images of neuronal cells in co-culture with astrocytes (e.g., TSC2 knockout (TSC2 ko), wild type (wt), SETD1A heterozygous knockout (SETD1A het)) stained with anti-MAP2 antibody according to some embodiments (FIG. 13A) and corresponding cell nucleus and soma segmentation of a single neuronal cell segmented from the background (FIG. 13B).
[0027] [Figure 14A] FIG. 1 shows exemplary images of genotype-labeled barcoded cells (e.g., TSC2 knockout (TSC2 ko), wild type (wt), SETD1A heterozygous knockout (SETD1A het)) captured during a round of in situ sequencing by synthesis, according to some embodiments.
[0028] [Figure 14B] , and [Figure 14C] Exemplary counts of pooled optical screening (POSH) barcodes for each genotype (e.g., TSC2 knockout (TSC2 ko), wild type (wt), SETD1A heterozygous knockout (SETD1A het)) (FIG. 14B) and corresponding cell counts for each genotype (FIG. 14C) are shown according to some embodiments.
[0029] [Figure 15A] 1 shows an exemplary transformation (e.g., embedding) depicting sequencing and imaging data of untreated cells according to some embodiments. The dashed circle indicates the feature embedding space that is predominantly occupied by TSC2 ko cells.
[0030] [Figure 15B]1 illustrates an exemplary overlay of representative image patches of untreated cells over their embedding coordinates, according to some embodiments.
[0031] [Figure 16A] 1 shows exemplary transformations calculated after treatment of cells (e.g., TSC2 knockout (TSC2 ko), wild type (wt), SETD1A heterozygous knockout (SETD1A het)) with rapamycin or without treatment (No Tr) according to some embodiments. The dashed circle indicates the feature embedding space that is predominantly occupied by untreated TSC2 ko cells.
[0032] [Figure 16B] 1 shows an exemplary transform calculated after treating TSC2 knockout cells with or without rapamycin (Rapamycin), according to some embodiments. The dashed circle indicates the feature embedding space predominantly occupied by untreated TSC2 ko cells.
[0033] [Figure 17] Exemplary scRNAseq embeddings are shown, followed by treatment of cells (e.g., TSC2 knockout (TSC2 ko), TSC2 heterozygous knockout (TSC2 het), wild type (wt), SETD1AG3 heterozygous knockout (SETD1AG3 het) SETD1AG4 heterozygous knockout (SETDAG4 het)) with various compounds (e.g., DMSO, everolimus, iadamstat, lonafarnib, rapamycin), or untreated cells. TSC2 ko neurons appear in cluster 6, as represented by the solid arrow. Clear arrows indicate a shift to a new population of all cells upon treatment with rapamycin. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0034] [I.Definition] In order that this disclosure may be more readily understood, certain terms are first defined. As used in this application, unless otherwise expressly provided herein, each of the following terms shall have the meaning set forth below. Additional definitions are set forth throughout this application.
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure pertains. For example, Concise Dictionary of Biomedical and Molecular Biology, Juo, Pei-Show, 2nd ed, 2002, CRC Press, The dictionary Of Cell And Molecular Biology, 3rd ed, 1999, Academic Press, and Oxford dictionary Of Biochemistry And Molecular Biology, Revised, 2000, Oxford University Press provide one with a general dictionary of many terms used in this disclosure.
[0036] Units, prefixes, and symbols are denoted in the accepted format of the systeme International de Unites (SI). Numerical ranges are inclusive of the numerical values defining the range. The headings provided herein are not intended to limit the various aspects of the disclosure, which can be had by reference to the specification in its entirety. Thus, the terms defined immediately below are more fully defined by reference to the specification in its entirety.
[0037] The term "and / or" as used herein should be interpreted as a specific disclosure of each of the two specified features or components with or without the other. Thus, the term "and / or" used herein in phrases such as "A and / or B" is intended to include "A and B," "A or B," "A" (single), and "B" (single). Similarly, the term "and / or" used in phrases such as "A, B, and / or C" is intended to encompass each of the following aspects: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (single); B (single); and C (single).
[0038] The use of alternatives (e.g., "or") should be understood to mean either one, both, or any combination thereof of the alternatives. As used herein, the indefinite article "a" or "an" should be understood to refer to "one or more" of the list or listed members.
[0039] It is understood that aspects and embodiments of the invention described herein include "comprising," "consisting," and "consisting essentially of" aspects and embodiments.
[0040] The term "about" refers to a value or composition that is within an acceptable error range for a particular value or composition as determined by one of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, "about" can mean a range of up to 20%. Furthermore, particularly with respect to biological systems or processes, the term can mean values up to an order of magnitude or up to 5-fold. When a particular value or composition is provided in this application and claims, unless otherwise specified, the meaning of "about" should be assumed to be within an acceptable error range for that particular value or composition.
[0041] The words "optional" or "optionally" mean that the subsequently described event, circumstance, or substitution may or may not occur, and that the description includes instances where the event or circumstance occurs and instances where it does not occur.
[0042] A "subject" includes any human or non-human animal. The term "non-human animal" includes, but is not limited to, vertebrates such as non-human primates, sheep, dogs, and rodents such as mice, rats, and guinea pigs. In some embodiments, the subject is a human. The terms "subject" and "patient" and "individual" are used interchangeably herein.
[0043] As used herein, a "biological sample" may include whole cells and / or live cells and / or cell debris. The biological sample may contain (or be derived from) a "body fluid". The invention encompasses embodiments in which the body fluid is selected from amniotic fluid, aqueous humor, vitreous humor, bile, blood serum, breast milk, cerebrospinal fluid, earwax (ear wax), chyle, chyme, endolymph, perilymph, exudate, feces, female ejaculate, gastric acid, gastric juice, lymph, mucus (including nasal drainage and phlegm), pericardial fluid, peritoneal fluid, pleural fluid, pus, saliva, sebum (skin oil), semen, phlegm, synovial fluid, sweat, tears, urine, vaginal secretions, vomit, and mixtures of one or more thereof. Biological samples include cell cultures, body fluids, and cell cultures derived from body fluids. Body fluids may be obtained from a mammalian organism, for example, by puncture or other collection or sampling procedures.
[0044] As used herein, "isogenic" refers to an organism or cell characterized by essentially identical genomic DNA, e.g., the genomic DNA is at least about 92%, preferably at least about 98%, and most preferably at least about 99% identical to the genomic DNA of the isogenic organism or cell.
[0045] The term "cell" is used in the broadest sense herein to mean a living organism that is a structural unit of tissue in a multicellular organism, surrounded by a membrane structure that separates it from the outside, that contains genetic information, and that has a mechanism for expressing the genetic information. As used herein, the cell may be a naturally occurring cell or an artificially modified cell (e.g., a fusion cell, a genetically modified cell, etc.).
[0046] The term "differentiated cells" as used herein can refer to cells that have developed from an undifferentiated phenotype to a specialized phenotype. For example, embryonic cells can differentiate into epithelial cells of the intestinal lining. Differentiated cells can be isolated, for example, from fetuses or from born animals.
[0047] As used herein, the term "undifferentiated cell" can refer to a precursor cell that has an undifferentiated phenotype and is capable of differentiation. An example of an undifferentiated cell is a stem cell.
[0048] As used herein, the term "stem cell" refers to a cell that is self-renewing and pluripotent. "Pluripotency" means that the cell can give rise to the three primary germ layers that comprise an adult animal; through its progeny germ cells and all three germ layers, endoderm (internal gastric mucosa, gastrointestinal tract, lungs), mesoderm (muscle, bone marrow, blood, urogenital tract), or ectoderm (epithelial tissue and nervous system). Stem cells herein may be, but are not limited to, embryonic stem (ES) cells, tissue stem cells (also called tissue-specific stem cells or somatic stem cells), or induced pluripotent stem cells. Artificially produced cells (e.g., reprogrammed cells) with the above capabilities may be stem cells.
[0049] As used herein, the term "embryonic stem (ES) cells" can refer to pluripotent cells isolated from an embryo maintained in an in vitro cell culture medium.
[0050] Tissue stem cells are divided into categories based on the site from which they originate, such as the skin system (e.g., epidermal stem cells, hair follicle stem cells), digestive system (e.g., pancreatic stem cells, hepatic stem cells, etc.), bone marrow (e.g., hematopoietic stem cells, mesenchymal stem cells, etc.), and nervous system (e.g., neural stem cells, retinal stem cells, etc.).
[0051] "Induced pluripotent stem cells" (commonly abbreviated as iPS cells or iPSCs) are pluripotent stem cells derived from non-pluripotent cells, typically adult somatic cells, or terminally differentiated cells such as fibroblasts, hematopoietic cells, muscle cells, or neural cells. It refers to a type of pluripotent stem cell artificially prepared by expressing reprogramming factors from transiently differentiated cells such as epithelial cells.
[0052] "Self-renewal" refers to the ability to undergo many cycles of cell division while maintaining an undifferentiated state.
[0053] As used herein, the term "somatic cell" refers to any cell other than a germ cell, such as an egg or sperm. Typically, somatic cells have limited or no pluripotency. As used herein, somatic cells may be natural or genetically modified.
[0054] As used herein, "single-cell RNAseq" or "scRNA-Seq" generally refers to single-cell RNA sequencing methods to obtain expression profiles of individual cells.
[0055] The term "whole genome sequencing (WGS)" as used herein refers to a process by which the entire genome of an organism, e.g., human, dog, mouse, virus, or bacteria, can be sequenced. It is not necessary that the entire genome is actually sequenced.
[0056] The term "sequencing" in this specification refers to a method for determining the nucleotide sequence of polynucleotide, such as genomic DNA.Preferably, the sequencing method includes, as a non-limiting example, next generation sequencing (NGS) method, (NGS), in which clonally amplified DNA templates or single DNA molecules are sequenced in a massively parallel manner (e.g., as described in Volkerding et al Clin Chem 55:641-658 (2009); Metzker M Nature Rev 11:31-46 (2010)).
[0057] As described herein, any concentration range, percentage range, ratio range, or integer range, unless otherwise stated, should be understood to include any integer value within the recited range, and fractions thereof, where appropriate (such as tenths and hundredths of integers).
[0058] Various aspects of the disclosure are described in further detail in the following subsections.
[0059] II. Method of the Invention One aspect of the invention provides a method for pooled screening of cells from different genetic backgrounds, comprising: a) labeling two or more populations of cells of different genetic backgrounds with two or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells; b) combining the two or more populations of cells to obtain a single mixed population of cells; c) performing in situ single cell sequencing on the cells; and d) analyzing known phenotypes or identifying new phenotypes of the cells in the mixed population. In some embodiments, the two or more populations of cells comprise any combination of cultured cells, primary cells, post-mitotic cells (such as neural cells), and tissue sections. In some embodiments, the two or more populations of cells are from different cell lines. In some embodiments, the different cell lines are healthy cell lines. In some embodiments, the different cell lines are patient cell lines. In some embodiments, the different cell lines are isogenically engineered cell lines. In some embodiments, the different cell lines comprise any combination of healthy cell lines, patient cell lines, and isogenically engineered cell lines. In some embodiments, the two or more populations of cells are induced pluripotent stem cells (iPSCs). In some embodiments, the method further comprises culturing the cells prior to analyzing the phenotype. In some embodiments, the two or more populations of cells are obtained from a human. In some embodiments, step d) comprises analyzing the phenotype of the cells by capturing a microscopic image or a time series of microscopic images of the cells and evaluating the phenotypic characteristics exhibited in the image or images.In some embodiments, the method further includes a computer-implemented method for registration between the first and second multiple images, comprising: 1) generating a first coordinate space of reference of a first multiple images of a well on a culture plate; 2) extracting a first patch of the first multiple images; 3) generating a second coordinate space of reference of a second multiple images of a well on a culture plate; 4) extracting a second patch of the second multiple images; 5) calculating an affine transformation function between the first and second patches to obtain a plurality of transformation parameters; and 6) generating a coordinate transformation function between the first and second coordinate spaces of reference based on the plurality of transformation parameters.
[0060] In some embodiments, the method further includes utilizing a classifier configured to receive an image, such as an aligned image generated from the first and second plurality of images, according to the methods provided herein, and output a classification result. The classifier may be a neural network including multiple layers. In some embodiments, the classifier is a convolutional neural network (CNN), such as a DenseNet classifier. In some embodiments, the classifier may be used to obtain a low-dimensional representation of the image, such as an embedding. The embedding may be generated from one of the multiple layers of the classifier. For example, an activation output layer before the last layer of the classifier may be utilized as a low-dimensional representation of the image for visualization. In some embodiments, the embedding is reduced in dimension using a dimensionality reduction method. In some embodiments, the dimensionality reduction method is an unsupervised linear dimensionality reduction method. In some embodiments, the dimensionality reduction method is an unsupervised nonlinear dimensionality reduction method. In some embodiments, the embedding is reduced in dimension using a Uniform Manifold Approximation and Projection for Dimension Reduction (UMAP) algorithm, Principal Component Analysis (PCA), t-Distributed Stochastic Neighbor Embedding (t-SNE), or any other suitable dimensionality reduction method. In some embodiments, this embedding is dimensionally reduced using the UMAP algorithm to obtain a UMAP plot for visualization of the embedding.
[0061] In some embodiments, the method further comprises expanding the two or more populations of cells for two or more generations prior to step b), in some embodiments, the method further comprises expanding the two or more populations of cells for three or more generations, four or more generations, five or more generations, six or more generations, seven or more generations, eight or more generations, nine or more generations, or ten or more generations prior to step b).
[0062] In some embodiments, performing in situ single-cell sequencing involves using fluorescent in situ RNA sequencing (FISSEQ) (Lee et al, Nature Protocols 2015, 10(3):442-58). In this method, mRNA is reverse transcribed in situ using aminoallyl dUTP and adapter sequence tagged random hexamers. The resulting cDNA fragments are immobilized in a cell protein matrix and circularized. The circular templates are amplified by rolling circle amplification (RCA) followed by sequencing and imaging. This method allows for the simultaneous detection of tissue-specific gene expression, RNA splicing, post-transcriptional modifications, and the preservation of their spatial information. It is a relatively unbiased method and can achieve sampling of the entire transcriptome.
[0063] In some embodiments, performing in situ single-cell sequencing involves using the padlock in situ sequencing method (Ke et al. Nature Methods 2013, 10(9)857-60). In this method, after mRNA is reverse transcribed into cDNA, the mRNA is degraded by RNaseH. Padlock probes then bind to the cDNA at the gap between the probe ends on the bases targeted for sequencing. The gaps are filled by DNA polymerization and ligated to form circularized molecules. The circular templates are amplified by RCA, followed by sequencing and imaging. Similar to FISSEQ, padlock in situ sequencing allows for the preservation of spatial information of the analyzed RNA sequences.
[0064] In some embodiments, after RCA, sequencing of the amplified DNA can be achieved using sequencing by ligation or sequencing by synthesis. Sequencing by synthesis relies on DNA polymerase to incorporate four reversible terminator-bound dNTPs. One base is added per cycle, and fluorescently labeled reversible terminators are imaged as each dNTP is added. Sequencing by ligation distinguishes the sequence of interest and uses the mismatch sensitivity of DNA ligase instead of incorporating a pool of fluorescently labeled oligonucleotides of various lengths. Sequencing by ligation has high accuracy, but can face problems with palindromic sequences.
[0065] In some embodiments, rather than in situ sequencing of RNA, the DNA barcode is read directly via methods such as peptide nucleic acid, locked nucleic acid, transposase, zombie, or other in situ transcription methods.
[0066] In some embodiments, rather than in situ sequencing of RNA or DNA, methods such as Procode are used to label individual cell lines using protein tags as barcodes.
[0067] In certain embodiments, multiple barcode sequences within the same cell may be determined by in situ sequencing. Barcode screening methods can also be combined with high-dimensional morphological profiling and in situ multiplexed gene expression analysis. In certain embodiments, phenotypes can be measured in their native spatial context using in situ sequencing of tissue samples.
[0068] In some embodiments, the method further comprises performing single-cell RNAseq (scRNA-seq). In some embodiments, a single unique nucleic acid barcode sequence is used for in situ single-cell sequencing and scRNA-seq. scRNA-seq approaches include 10x Ganomics, Drop-seq, and Seq-well, inDrops, Rhapsody, and Split-Seq. For example, single-cell libraries can be prepared from single-cell suspensions using Chromium with v2 chemistry (10x Ganomics). Such single-cell libraries can be sequenced (e.g., NextSeq 500 (Illumina)). Sequencing reads can be processed by alignment, filtering, de-duplication, and / or conversion to a digital count matrix, for example, using Cell Ranger 1.2 (10x Ganomics).
[0069] In some embodiments, the methods include the application of sc-RNAseq and the use of compressed sensing methodologies for RNA sequencing. Examples include L1000 and hybrid capture.
[0070] In some embodiments, the method further comprises culturing the single mixed population of cells before analyzing the phenotype. In some embodiments, the culture medium used to culture the cells contains serum. In some embodiments, the culture medium used to culture the cells is serum-free. Serum-free medium refers to a medium that does not contain raw or unpurified serum, and thus can include a medium with purified blood-derived components or animal tissue-derived components (e.g., growth factors). In terms of preventing contamination with components from different animals, the serum may be from the same animal as the cells. The medium may or may not contain serum substitutes. Serum substitutes include albumin (albumin substitutes such as lipid-rich albumin, recombinant albumin, plant starch, dextran, and protein hydrolysates), transferrin (or other iron transporters), fatty acids, insulin, collagen precursors, trace elements, 2-mercaptoethanol, or 3'-thioglycerol, or equivalents thereof.
[0071] In some embodiments, a single mixed population of cells is on a substrate or in a three-dimensional culture. In some embodiments, a mixed population of cells is on a substrate. In some embodiments, the substrate is any standard tissue culture vessel, such as a tissue culture plate or flask. In some embodiments, the substrate is a cell culture dish. In some embodiments, the substrate is a tissue culture plate. In some embodiments, the substrate is a Petri dish. In some embodiments, the substrate is a tissue culture flask. In some embodiments, the substrate is a well of a standard microwell plate, such as a 6-well, 12-well, 24-well, 96-well, 384-well, or 1,536-well plate. In some embodiments, the substrate is a 6-well plate. In some embodiments, the substrate is a 12-well plate. In some embodiments, the substrate is a 24-well plate. In some embodiments, the substrate is a 96-well plate. In some embodiments, the substrate is a 384-well plate. In some embodiments, the substrate is a 1,536-well plate. The substrate may be made of any material for imaging using the imaging modalities described herein. In certain embodiments, the plate may be a plastic bottom plate suitable for imaging using the imaging modalities described herein. In certain embodiments, the plate may be a glass bottom plate suitable for imaging using the imaging modalities described herein. In certain embodiments, the substrate may be a culture chamber in an array of culture chambers defined on a microfluidic device, or a droplet generated on a microfluidic device. In certain example embodiments, single cells or cell populations may be cultured on individual microscope slides in culture medium. In some embodiments, a single mixed population of cells is in a three-dimensional culture. In some embodiments, the three-dimensional culture comprises a scaffold. In some embodiments, the scaffold comprises a three-dimensional matrix.In some embodiments, the three-dimensional matrix comprises a material selected from the group consisting of BD Matrigel™ basement membrane matrix (BD Sciences), Cultrex® basement membrane extract (BME; Trevigen), hyaluronic acid, polyethylene glycol (PEG), polyvinyl alcohol (PVA), polylactide-co-glycolide (PLG), and polycaprolactone (PLA). In some embodiments, the three-dimensional matrix comprises BD Matrigel™ basement membrane matrix (BD Sciences). In some embodiments, the three-dimensional matrix comprises Cultrex® basement membrane extract (BME; Trevigen). In some embodiments, the three-dimensional matrix comprises hyaluronic acid. In some embodiments, the three-dimensional matrix comprises polyethylene glycol (PEG). In some embodiments, the three-dimensional matrix comprises polyvinyl alcohol (PVA). In some embodiments, the three-dimensional matrix comprises polylactide-co-glycolide (PLG). In some embodiments, the three-dimensional matrix comprises polycaprolactone (PLA). In some embodiments, the three-dimensional culture does not comprise a scaffold.
[0072] In some embodiments, the method includes stably integrating a unique nucleic acid barcode sequence into the genome of two or more populations of cells. In some embodiments, the unique nucleic acid barcode sequence is delivered into the cells using a virus. In some embodiments, the virus is a retrovirus. In some embodiments, the retrovirus is or is derived from Moloney Murine Leukemia Virus (MMULV), Feline Immunodeficiency Virus (FIV), Harvey Murine Sarcoma Virus (HaMuSV), Mouse Mammary Tumor Virus (MuMTV), Gibbon Ape Leukemia Virus (GaLV), Human Immunodeficiency Virus (HIV), Rous Sarcoma Virus (RSV), or a lentivirus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is derived from a lentivirus. In some embodiments, the U3 sequence from a lentivirus 5' LTR may be replaced with a promoter sequence in the viral construct. This may increase the titer of the virus recovered from the packaging cell line. An enhancer sequence may also be included. In some embodiments, the virus encodes a selectable marker. In some embodiments, the selectable marker is an antibiotic resistance gene. In some embodiments, the antibiotic resistance gene confers resistance to an antibiotic selected from the group consisting of puromycin, hygromycin, bleomycin, neomycin, actinomycin D, and mitomycin C. In some embodiments, the antibiotic resistance gene confers resistance to puromycin. In some embodiments, the antibiotic resistance gene confers resistance to hygromycin. In some embodiments, the antibiotic resistance gene confers resistance to bleomycin. In some embodiments, the antibiotic resistance gene confers resistance to neomycin. In some embodiments, the antibiotic resistance gene confers resistance to actinomycin D. In some embodiments, the antibiotic resistance gene confers resistance to mitomycin C. In some embodiments, the virus encodes one or more fragments of the antibiotic resistance gene. In some embodiments, the virus encodes a fluorescent protein.In some embodiments, the fluorescent protein is selected from the group consisting of green fluorescent protein, red fluorescent protein, blue fluorescent protein, cyan fluorescent protein, yellow fluorescent protein, and orange fluorescent protein.
[0073] In some embodiments, a unique nucleic acid barcode sequence is a short sequence of nucleotides (e.g., DNA or RNA) that is used as an identifier for a target molecule and / or related molecules, such as a target nucleic acid, or as an identifier for the source of the related molecules, such as the cell of origin. Barcode may also refer to any unique, non-naturally occurring nucleic acid sequence that can be used to identify the source of a nucleic acid fragment.
[0074] The unique nucleic acid barcode sequence may be attached, or "tagged," to the target molecule. This attachment can be direct (e.g., covalent or non-covalent attachment of the unique nucleic acid barcode sequence to the target molecule) or indirect (e.g., via an additional molecule).
[0075] A target molecule can be optionally labeled with multiple unique nucleic acid barcode sequences in a combinatorial manner (e.g., using multiple unique nucleic acid barcode sequences bound to one or more specific binding agents that specifically recognize the target molecule), thus greatly increasing the number of unique identifiers that can be in a particular pool of unique nucleic acid barcode sequences. In certain embodiments, the unique nucleic acid barcode sequences are added, e.g., one at a time, to growing barcode concatemers bound to the target molecule. In other embodiments, multiple unique nucleic acid barcode sequences are assembled prior to binding to the target molecule.
[0076] In some embodiments, each unique nucleic acid barcode sequence is at least 1 base pair in length. In some embodiments, each unique nucleic acid barcode sequence is 1 to about 18 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 1 to about 12 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 1 to about 10 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 1 to about 8 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 2 to about 12 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 2 to about 10 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 2 to about 8 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 3 to about 12 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 3 to about 10 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 3 to about 8 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 4 to about 12 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 4 to about 10 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 4 to about 8 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 5 to about 12 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 5 to about 10 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 5 to about 8 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 6 to about 12 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 6 to about 10 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 6 to about 8 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 1 base pair in length.In some embodiments, each unique nucleic acid barcode sequence is 2 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 3 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 4 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 5 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 6 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 7 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 8 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 9 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 10 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 11 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 12 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 13 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 14 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 15 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 16 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 17 base pairs in length. In some embodiments, each unique nucleic acid barcode sequence is 18 base pairs in length. In certain embodiments, the unique nucleic acid barcode sequences may be detected directly using in situ sequencing methods.In certain exemplary embodiments, the unique nucleic acid barcode sequence is detected using fluorescent in situ RNA sequencing (FISSEQ), in situ mRNA-seq, padlock in situ sequencing, sequencing by ligation, SOLiD® sequencing, sequencing by synthesis, peptide nucleic acid, locked nucleic acid, transposase, zombie, other in situ transcription methods, or other protein- or peptide-based barcode technologies such as ProCodes. In certain exemplary embodiments, the mRNA transcripts encoding the unique nucleic acid barcode sequence are sequenced. In certain other exemplary embodiments, a cDNA copy of the mRNA is first generated and then sequenced. In certain other exemplary embodiments, the DNA containing the barcode is directly sequenced.
[0077] In some embodiments, two or more populations of cells are sequenced before labeling with unique nucleic acid barcode sequences. In some embodiments, a portion of the genome of two or more populations of cells is sequenced. In some embodiments, the sequencing is whole genome sequencing. In some embodiments, whole genome sequencing includes the use of next generation sequencing (NGS). NGS techniques for determining the whole human genome sequence have been previously described (Levy et al PLoS Biol 55, e254 (2007); Wheeler et al. Nature 452:872-876 (2008); Bentley et al, Nature 456:53-59 (2008)).
[0078] In some embodiments, the two or more populations of cells are obtained from related individuals. In some embodiments, the related individuals are parents and offspring. In some embodiments, the related individuals are siblings. In some embodiments, the two or more populations of cells are obtained from individuals with low genetic diversity. In some embodiments, the two or more populations of cells are obtained from individuals with high genetic diversity. In some embodiments, the two or more populations of cells are comprised of a combination of groups including low and high genetic diversity.
[0079] In some embodiments, the method comprises labeling three or more populations of cells of different genetic backgrounds with three or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling four or more populations of cells of different genetic backgrounds with four or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling five or more populations of cells of different genetic backgrounds with five or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling six or more populations of cells of different genetic backgrounds with six or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling seven or more populations of cells of different genetic backgrounds with seven or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling eight or more populations of cells of different genetic backgrounds with eight or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling 9 or more populations of cells of different genetic backgrounds with 9 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells. In some embodiments, the method comprises labeling 10 or more populations of cells of different genetic backgrounds with 10 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells. In some embodiments, the method comprises labeling 11 or more populations of cells of different genetic backgrounds with 11 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells.In some embodiments, the method comprises labeling 12 or more populations of cells of different genetic backgrounds with 12 or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling 13 or more populations of cells of different genetic backgrounds with 13 or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling 14 or more populations of cells of different genetic backgrounds with 14 or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling 15 or more populations of cells of different genetic backgrounds with 15 or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling 20 or more populations of cells of different genetic backgrounds with 20 or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling 50 or more populations of cells of different genetic backgrounds with 50 or more unique nucleic acid barcode sequences, each of which corresponds to a different population of cells. In some embodiments, the method comprises labeling 100 or more populations of cells of different genetic backgrounds with 100 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells. In some embodiments, the method comprises labeling 500 or more populations of cells of different genetic backgrounds with 500 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells. In some embodiments, the method comprises labeling 1000 or more populations of cells of different genetic backgrounds with 1000 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells.In some embodiments, the method comprises labeling 2000 or more populations of cells of different genetic backgrounds with 2000 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells. In some embodiments, the method comprises labeling 5000 or more populations of cells of different genetic backgrounds with 5000 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells. In some embodiments, the method comprises labeling 10,000 or more populations of cells of different genetic backgrounds with 10,000 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells. In some embodiments, the method comprises labeling over 100 populations of cells of different genetic backgrounds with over 100 unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells.
[0080] In some embodiments, analyzing the phenotype of the single mixed population of cells comprises an assay selected from the group consisting of high content imaging, calcium imaging, immunohistochemistry, cell morphology imaging, protein aggregation imaging, cell-cell interaction imaging, live cell imaging, and any other image-based assay modality. In some embodiments, analyzing the phenotype of the cells comprises an assay selected from the group consisting of high content imaging, calcium imaging, immunohistochemistry, cell morphology imaging, protein aggregation imaging, cell-cell interaction imaging, and live cell imaging. In some embodiments, analyzing the phenotype of the single mixed population of cells comprises performing high content imaging. In some embodiments, analyzing the phenotype of the single mixed population of cells comprises performing calcium imaging. In some embodiments, analyzing the phenotype of the single mixed population of cells comprises performing immunohistochemistry. In some embodiments, analyzing the phenotype of the single mixed population of cells comprises performing cell morphology imaging. In some embodiments, analyzing the phenotype of the single mixed population of cells comprises performing protein aggregation imaging. In some embodiments, analyzing the phenotype of the single mixed population of cells comprises performing cell-cell interaction imaging. In some embodiments, analyzing the phenotype of the single mixed population of cells comprises performing live cell imaging.
[0081] [Stem cells] In some embodiments, the cells are stem cells. In some embodiments, the stem cells are pluripotent stem cells. In some embodiments, the cells are induced pluripotent stem cells (iPSCs). In some embodiments, iPSCs are generated from somatic cells by introducing one or more known reprogramming factors. In some embodiments, iPSCs are differentiated before analyzing the phenotype. Stem cells are characterized by their ability to renew themselves through mitotic cell division and differentiation into a diverse range of specialized cell types. There are two main types of mammalian stem cells: embryonic stem cells found in blastocysts and adult stem cells found in adult tissues. In the developing embryo, stem cells can differentiate into any specialized embryonic tissue. In adults, stem cells and progenitor cells function as the body's repair system, mobilizing specialized cells as well as maintaining the normal turnover of regenerative organs such as blood, skin, or intestinal tissue. Human embryonic stem cells (hES) can be defined by the presence of several transcription factors and cell surface proteins. The transcription factors Oct4, Nanog, and Sox2 form a core regulatory network that represses genes linked to differentiation and the maintenance of pluripotency. The cell surface antigens most frequently used to identify hES cells include the glycolipids SSEA3 and SSEA4 and the keratan sulfate antigens Tra-1-60 and Tra-1-81.
[0082] The generation of iPSCs depends on the genes used for induction. Factors such as Oct3 / 4, KLF4, Sox2 and / or c-myc or combinations thereof can be used. Nucleic acids encoding these reprogramming factors can be included in monocistronic or multicistronic expression cassettes. Similarly, nucleic acids encoding monocistronic or multicistronic expression cassettes can be included in one reprogramming vector or multiple reprogramming vectors.
[0083] iPSCs are typically generated by transfecting specific stem cell-associated genes into non-pluripotent cells, such as adult fibroblasts or umbilical cord blood cells. Transfection can be accomplished using integrating viral vectors, such as retroviruses (e.g., lentiviruses), or non-integrating viral vectors, such as Sendai virus. Reprogramming may also be performed using virus-free methods, such as episomal reprogramming or mRNA reprogramming. After a critical period, a small number of transfected cells begin to resemble pluripotent stem cells morphologically and biochemically, based on morphological selection, doubling time, reporter gene expression, and / or antibiotic resistance.
[0084] Pluripotent cells can be cultured and maintained in an undifferentiated state using a variety of methods. In some embodiments, matrix components may be included in a given medium to culture and maintain pluripotent cells in a substantially or essentially undifferentiated state. A variety of matrix components can be used to culture and maintain pluripotent cells, such as hESCs or iPSCs. For example, collagen IV, fibronectin, laminin, and vitronectin can be used in combination to provide a solid support for embryonic cell culture and maintenance.
[0085] Matrigel™ may be used to provide a substrate for cell culture and maintenance of pluripotent cells. Matrigel™ is a gelatinous protein mixture secreted by mouse tumor cells and is commercially available from BD Biosciences (New Jersey, USA). This mixture resembles the complex extracellular environment found in many tissues and is used by cell biologists as a substrate for cell culture. It will be appreciated that additional methods of culturing and maintaining iPSCs are known to those skilled in the art and may be used with embodiments of the present invention.
[0086] In some embodiments, the method comprises analyzing differentiation of stem cells. In some embodiments, the method comprises analyzing differentiation of pluripotent stem cells. In some embodiments, the method comprises analyzing differentiation of iPSCs. In some embodiments, the method comprises analyzing cell types derived from stem cell differentiation.
[0087] [Microscope image alignment] In some embodiments, the exemplary platform includes a microscope image registration technique that can be used to obtain conversion between sequencing and marker-based / marker-free readouts for demultiplexing single-cell image-based readouts.
[0088] Image registration is the process of transforming different sets of images into one coordinate system. The different sets of images may be of the same object, but from different image acquisitions; in other words, the different sets of images may be captured by different imagers (e.g., different microscopes) at different times using different settings, and therefore may have different depths, different resolutions, different viewpoints, etc. For example, each of the bases of the barcode may be captured as independent acquisitions spaced apart in time, with potential physical movement between acquisitions. Additionally, the marker / phenotype acquisitions may be captured at different resolutions, times, and / or from different microscopes compared to the barcode image acquisitions.
[0089] Existing methods for alignment between microscopic image acquisitions are performed at either the field level or the well level, and both suffer from drawbacks. For example, field-level alignment cannot be used to acquire in different settings or between different microscopes, and movement of the plate between acquisitions can lead to loss of information. Well-level alignment is computationally expensive and is not scalable for larger well sizes and higher magnifications, as the complexity of the alignment method is O(NlogN), where N is the number of pixels in the image. The drawbacks of current approaches are described in detail herein.
[0090] The image registration technique in this disclosure involves calculating a transformation function (e.g., a fixed affine transformation function) between two acquisitions, where the two acquisitions refer to two sets of images taken using two different acquisition settings, at two different times, and / or using two different imaging devices (e.g., different microscopes). After calculating such a transformation function, an acquisition of overlapping information can be made by independently applying the transformations to the coordinates of field images in the source acquisition to acquire corresponding field images and coordinates in the target acquisition.
[0091] Thus, the image registration techniques in this disclosure can be used to obtain a transformation of coordinates between two different acquisitions of the same object, where the two acquisitions may have occurred in different settings and using different devices, and where potential human interaction may have occurred between the acquisitions. This technique may be useful in a variety of laboratory image-based experimental settings.
[0092] For example, embodiments of the present disclosure can be used to perform physical coordinate alignment between images acquired from two different microscopes (fluorescent / non-fluorescent). For example, one image may be from a microscope inside an incubator in a live imaging setup, while the other may be a fluorescent image acquired after fixation. As another example, embodiments of the present disclosure can be used to perform physical coordinate alignment between images acquired at two different resolutions for multi-scale image analysis and reconstruction. As another example, embodiments of the present disclosure can be used to perform physical coordinate alignment between images obtained from a fixed sample after successive washes, and / or multiple staining procedures involving movement of plates between a microscope and an automated setup, e.g., sequencing by synthesis cycles of barcoding cells in a pooled optical screening analysis.
[0093] Microscopic acquisition for high content imaging, high throughput imaging, generally involves imaging cells cultured in special purpose plastic / glass bottom plates. FIG. 7 depicts an exemplary plate 700 designed for microscopic acquisition, according to some embodiments. Plate 700 includes a plurality of wells (e.g., 6, 24, 86, 384), such as well 702. Each of the plurality of wells includes cells that have been treated with a predetermined condition according to an experimental design. In some embodiments, multiple wells on a plate are treated with the same condition, and in some embodiments, the wells are treated with different conditions.
[0094] Referring to Figure 7, a well is generally imaged in part as a collection of overlapping field-of-view images by a microscope camera. For example, a typical microscope camera has an image output size of about 2,000 x 2,000 pixels. The resolution of this image can vary based on the magnification of the objective lens used for imaging. As shown in Figure 7, three separate field-of-view images 704a, 704b, and 704c of a well 702 can be imaged.
[0095] 8 illustrates multiple images of a well, according to some embodiments. As shown, the multiple images include images capturing the boundaries of the well (e.g., images 802 and 804) and the center of the well (e.g., image 806). The multiple images are the same size, and adjacent images overlap each other. For example, images 802 and 808 overlap each other horizontally by 812. As another example, images 802, 808, and 810 overlap each other by area 814.
[0096] Performing image registration for microscopic image acquisitions can be difficult for many reasons. Image registration is the process of transforming different sets of images into one coordinate system. The two different sets of images (i.e., the two image acquisitions) may be from different sensors (e.g., different microscopes), different times, different depths, different resolutions, and / or different viewpoints. Existing methods for alignment between microscopic image acquisitions are performed either at the field level or at the well level, and both suffer from drawbacks as described below.
[0097] Current field level alignment involves establishing a one-to-one correspondence between field images acquired from two image acquisitions. The one-to-one correspondence is established by the order of image acquisitions within fields and field positions within a well. This aligns a field image captured in a first acquisition to the same field image captured in a second acquisition. However, field level alignment only works for acquisitions with the same settings and the same microscope. It cannot be used to acquire at different settings or between different microscopes, and movement of the plate between acquisitions can lead to loss of information.
[0098] Current well-level alignment involves reconstructing the well image by stitching together smaller field-of-view images and aligning the well images using a Fourier correlation-based or related methodology. This method can work for smaller well sizes and smaller objective magnifications (smaller well image size in terms of number of pixels), but is not scalable for larger well sizes and higher magnifications, since the complexity of the alignment method is O(NlogN), where N is the number of pixels in the image.
[0099] FIG. 6 illustrates an exemplary computer-implemented process 600 for alignment of a first plurality of images (e.g., from a first image acquisition) with a second plurality of images (e.g., from a second image acquisition). The first and second acquisitions may differ in terms of well coverage, image resolution, image size, and imager type, and the relative positions of the wells may be shifted / rotated, but the underlying object being imaged is the same. Process 600 may be performed at least in part using one or more electronic devices. In some embodiments, blocks of 600 may be divided among multiple electronic devices. Some blocks may be optionally combined, the order of some blocks may be optionally changed, and some blocks are optionally omitted. In some examples, additional steps may be performed in combination with the process. Thus, as illustrated, the operations are exemplary in nature and therefore should not be considered limiting.
[0100] In the exemplary process 600, the first and second multiple images are both of a well (e.g., well 702) on a culture plate, but they are from two different image acquisitions. The contents in the well remain the same between the two image acquisitions, but the imaging settings may have changed between the two image acquisitions. For example, the first multiple images may be captured using a first imager (e.g., a fluorescent microscope) and the second multiple images may be captured using a second imager (e.g., a non-fluorescent microscope). As another example, the first multiple images may have a different resolution than the second multiple images. As another example, the first and second multiple images provide different coverages of the well (e.g., the second coverage may be a different size than the first coverage and / or may be shifted compared to the first coverage). As another example, the first and second multiple images are captured at different times.
[0101] In some embodiments, the images are obtained from a pooled screening assay of cells from different genetic backgrounds. In some embodiments, the first plurality of images are a plurality of barcoding images and the second plurality of images are a plurality of marker-based / marker-free readout images. In some embodiments, the first plurality of images are a plurality of marker-based / marker-free readout images and the second plurality of images are a plurality of barcoding images. In some embodiments, the multiple rounds or acquisitions of marker-based / marker-free readout images may be acquired using different imaging / staining paradigms, while the final plurality of images are a plurality of barcoding images.
[0102] In block 602, an exemplary system (e.g., one or more electronic devices) generates a first coordinate space of reference for the first plurality of images. In other words, the system assigns coordinate values in the first coordinate space of reference to pixels of at least one image of the first plurality of images. The first plurality of images may include one or more sets of images corresponding to one or more acquisitions. The coordinate space of reference may be calculated at the well level. A particular location of a well may be assigned a value on the first coordinate space of reference. As an example, the center of the well may be assigned as the origin (i.e., 0) of the first coordinate space of reference. Furthermore, the left and right extreme points of the well may be assigned values of -x and x on the X axis, and the top and bottom extreme points of the well may be assigned values of -y and y on the Y axis, where x and y are predefined numbers. In some embodiments, the X and Y values are based on the physical dimensions of the field of view. In some embodiments, the X and Y values are obtained from image acquisition metadata, which may be provided by the imager.
[0103] In some embodiments, the coordinate values in pixel space can be X=image_size_X (in pixels) and Y=image_size_Y (in pixels). In pixel space, the coordinates are relative to the image dimensions (in pixels). In the physical dimensions (i.e., the first coordinate space), the system establishes a coordinate space relative to the microscope stage. Based on the physical dimensions of the field from the microscope (e.g., field dimensions in micrometers) (e.g., metadata), the system can convert the pixel space to the physical space obtained from the microscope.
[0104] Each image of the first plurality of images can be associated with a first coordinate space of reference. In some embodiments, the system detects one or more physical characteristics of the well in the image (e.g., well edge, well shape, well location) and accordingly associates the image with the first coordinate space of reference. With reference to FIG. 8, the system can detect the left-most and right-most points of the well in the image and assign -X and X values accordingly. Additionally, the system can calculate an overlap ratio (e.g., between images 802 and 808). The overlap ratio may be part of metadata information from the imager (e.g., microscope metadata) or may be calculated using redundant image information between the two images. Based on the overlap ratio and a particular location on the well, the system can assign coordinate values (e.g., X, Y) to pixels of the image.
[0105] In some embodiments, the system can calculate (e.g., using imager metadata) a global coordinate space relative to well positions within the microscope stage in physical dimensions. For example, with reference to FIG. 9, the location of the center of each field image can be measured a priori based on device and acquisition settings, and overlap ratios and well extrema are calculated using the field image positions and known image dimensions. In some embodiments, the first coordinate space can be artificially simulated using the physical dimensions and overlap ratios of pixels. For example, if the microscope device metadata does not include information regarding the positions (i.e., physical coordinate locations) of the field images, the coordinate space can be created by identifying the order of field image acquisition within the well, image size, and overlap ratio.
[0106] The system extracts a first patch of the first plurality of images at 604. In some embodiments, block 604 includes blocks 606 and 608, as described below.
[0107] In block 606, the system selects one or more images from the first plurality of images that capture one or more landmarks (i.e., marker images). The landmarks can be information within the well, such as one or more nuclei, one or more cells, one or more beads (obtained from fluorescent markers or from segmentation from bright field or quantitative phase contrast images). The landmarks can also be well information, such as well boundaries and well centers. In some embodiments, the selection of the marker images can be performed using one or more machine learning models. With reference to FIG. 8, the marker images may include image 802 (capturing well boundaries), image 806 (capturing well centers), and image 816 (capturing objects or landmarks within the wells).
[0108] In block 608, the system extracts a first patch from one or more marker images. The first patch may be one of the marker images or a portion of one of the marker images. The patch may be extracted to capture a particular object or marker, or a particular location of the well (e.g., the center of the well). The size of the patch is determined such that the patch is large enough to capture the maximum allowable tolerance limit of movement / shift between acquisitions, as described below with reference to block 612. In some embodiments, the size and location of the patch are determined empirically based on an allowable tolerance threshold. For example, if the system aims to allow a maximum of 1 mm shift between two acquisitions, the system may take a patch size corresponding to x×1 mm in the corresponding image space where x>1. The location of this patch may be anywhere with a patch size corresponding to the allowable tolerance value. In some embodiments, the system extracts the first patch by sparse sampling of the marker images, either at random locations in the well or at fixed locations between the two acquisitions. For example, a well may be covered by hundreds of images. Sparse sampling involves selecting a subset of these images that correspond to random physical locations within the well to obtain the first patch. A patch can also be constructed by combining multiple images that correspond to a fixed well location (eg, the center of the well).
[0109] In block 610, the system generates a second reference coordinate space for the second plurality of images. In other words, the system assigns coordinate values in the second reference coordinate space to pixels of at least one image of the second plurality of images. The second plurality of images are of the same well on the culture plate. As described above, the second plurality of images can be from a second image acquisition. The first and second acquisitions may differ in terms of well coverage, image resolution, image size, and imager type, and the relative positions of the wells may be shifted / rotated, but the underlying object being imaged is the same. The generation of the second reference coordinate space in block 610 can be performed in a similar manner to block 602, but independently.
[0110] In block 612, the system extracts a second patch of the second plurality of images. Extraction of the second patch may be performed in a manner similar to block 604. For example, one or more marker images may be selected from the second plurality of images. A second patch may then be extracted from the one or more marker images. The second patch may be extracted to capture the same object or marker or the same location of the well (e.g., the center of the well) captured in the first patch.
[0111] The size of the first and second patches is determined such that the patches are large enough to capture the maximum allowable tolerance limits of movement / shift between acquisitions. For example, if the first patch and the second patch both capture the center of the well, then both patches need to be large enough to capture the translation, rotation and scaling about the center of the well in both acquisitions.
[0112] FIG. 10 shows two example patches from two different image acquisitions, according to some embodiments. Circle 1002 represents the position of a well during a first image acquisition, and circle 1004 represents the position of the same well during a second image acquisition. As shown, the well has shifted between the two acquisitions (e.g., due to plate movement or imager position). Patch 1006 is selected from the first image acquisition, where a barcoding image is acquired. Patch 1008 is selected from the second image acquisition, where a marker-based / marker-free readout image is acquired. The two patches also have different resolutions and coverage of the well. In the illustrated example, the two patches are selected to contain the same object. As shown in FIG. 10, the dots at 1008 and 1006 correspond to the same physical point or location. Affine functions T(x) and T-1(x) can be calculated to translate between the coordinate spaces of the two acquisitions, as described below.
[0113] At 614, the system calculates an affine transformation function between the first patch and the second patch to obtain a plurality of transformation parameters. The transformation parameters may include one or more of a translation parameter, a scaling parameter, and a rotation parameter.
[0114] 11 illustrates an example set of transformation functions according to some embodiments. Each of the transformations T1(x) through T5(x) can be expressed as a matrix multiplication operation, where T1(x)=T1×x, where T1 is an affine matrix.
[0115] 11, a first patch 1112 is from a first image acquisition and is associated with a first reference coordinate space 1102, which may be generated in block 602. T1(x) refers to the transformation from field A pixel coordinates (e.g., pixel values) to the reference coordinate space 1102. As mentioned above, the values in the first reference coordinate space 1102 may be based on physical dimensions (e.g., the physical dimensions of the field of view).
[0116] In some embodiments, an additional transformation T2(x) is performed. T2(x) refers to the transformation from the first reference coordinate space 1102 to the well A pixel coordinates. The first reference coordinate space 1102 is a coordinate space with its origin at the microscope stage. The well pixel coordinates refer to a coordinate space with its origin at the top left corner (or bottom left) of the well image (in pixels). The transformation T2(x) between the stage and well pixel coordinates is typically a scale+translation transformation calculated based on the dimensions of the microscope stage and the dimensions of the well image.
[0117] 11, a second patch 1114 is from a second image acquisition and is associated with a second coordinate space of reference 1104, which may be generated in block 610. A pixel in patch 1112 and a pixel in patch 1114 correspond to the same physical point (e.g., the same point on the well, the same object in the well), but are associated with two different coordinate spaces of reference. T5(x) refers to the transformation from the second coordinate space of reference 1104 to field B pixel coordinates. In other words, in block 606, T5 is used to transform the image from the second image acquisition to the second coordinate space of reference 1104 in block 610. -1 As mentioned above, the values in the second reference coordinate space 1104 can be based on physical dimensions (e.g., the physical dimensions of the field of view). Furthermore, T4(x) refers to the transformation from well B pixel coordinates to the reference coordinate space 1104, and T4 -1 (x) refers to the transformation from the second reference coordinate space 1104 to well B pixel coordinates.
[0118] In block 614, T3(x) is calculated by dividing the transformed patch 1112 (transformed with T1 and T2) and the transformed patch 1114 (transformed with T5 -1 and T4 -1T3(x) is calculated based on the pixel coordinates of well A to well B, transformed by (x). T3(x) refers to the transformation from well A pixel coordinates to well B pixel coordinates, calculated from the patch information using a Fast Fourier Transform based registration method. The transformation parameters include a translation matrix, a rotation matrix, a scale matrix, and a shear matrix in the illustrated example. These parameters are calculated using common landmarks in the patch (e.g., specific locations on the well, cells) using standard Fast Fourier Transform based registration methods.
[0119] In other words, T3(x) is generated based on two patches to convert between pixel coordinates in two reference coordinate spaces. Each arrow between blocks in Figure 11 includes a coordinate conversion operation. For example, given a field A pixel coordinate, obtaining the reference space A coordinate includes a conversion operation T1(x).
[0120] In some embodiments, the transform T3 of FIG. 11 may be calculated by sequentially applying four transform matrices, which may be expressed as matrix pre-multiplication operations on the coordinate vector x.
[0121] T3(x)=T3_translate(T3_shear(T3_scale(T3_rotate(x))))
[0122] = T3_translate × T3_shear × T3_scale × T3_rotate × x (where T3_<> is the corresponding matrix shown in Figure 11.)
[0123] =T3×x (where T3=T3_translate×T3_shear×T3_scale×T3_rotate)
[0124] At 616, the system generates a coordinate transformation function between the first and second reference coordinate spaces based on the plurality of transformation parameters. In the illustrated example of FIG. 11, the overall transformation matrix is calculated as T=T5×T4×T3×T2×T1. Thus, the end-to-end transformation function T(x) can be expressed as a series of matrix pre-multiplications as follows:
[0125] T(x)=T5(T4(T3(T2(T1(x)))))=T5×T4×T3×T2×T1×x=T×x (where T=T5×T4×T3×T2×T1)
[0126] The image may be further processed to obtain cell representations, such as assigning cell genetic backgrounds, treatments, and the like to single cells in the image. Any model for feature extraction may be used to obtain cell representations to classify between genetic backgrounds of cells in the image. In some embodiments, the system deploys self-supervised learning (SSL), semi-supervised, or unsupervised learning methods. In some embodiments, the system deploys SSL techniques in which a machine learning model learns from unlabeled sample data, as described in detail herein. For example, an image (e.g., a segmented morphological image) may be input to a trained self-supervised learning model configured to receive an image and output an embedding (i.e., a vector) that represents the image in a latent space. The embedding may be a vector representation of the input image in the latent space. By converting the input image to an embedding, the size and dimensionality of the original data may be significantly reduced. As described herein, the lower dimensional embedding may be used for downstream processing. By obtaining the embedding from the image, the self-supervised model may generate a spatiotemporal topological space in which directionality is available. For example, each image may be converted to an embedding, and the embedding may be mapped to a location in the topological space and time-stamped with the time the image was captured. Thus, directionality across space and / or time can be obtained across multiple embeddings in the topological space.
[0127] In some embodiments, the self-supervised learning model is a DINO visual transformer, a SimCLR model, or any other model that learns from unlabeled sample data. In some embodiments, the unsupervised machine learning model is a trained contrastive learning algorithm. Contrastive learning can refer to a machine learning technique used to learn general features of a dataset without labels by teaching the model which data points are similar or different. A contrastive learning model can extract embeddings from imaging data that linearly predict the labels that might otherwise be assigned to such data. Any contrastive learning model is trained by minimizing the contrastive loss, which maximizes the similarity between embeddings from different extensions of the same sample image and minimizes the similarity between embeddings of different sample images. For example, the model can extract embeddings from images that are invariant to rotation, flipping, cropping, and color jitter.
[0128] In some embodiments, the system deploys a classifier to analyze the image. In some embodiments, the classifier is a convolutional neural network (CNN) with multiple layers, such as a DenseNet classifier. For example, an image (e.g., a segmented morphological image) may be input to a classifier configured to receive an image and output a classification result. In some embodiments, a segmented morphological image of a plurality of cells is input to the classifier. In some embodiments, a genetic background label (e.g., a label identifying a genotype of an individual cell of the plurality of cells) is an output of the classifier. In some embodiments, the classification result may include associating a single cell with a genetic background. An embedding (i.e., a vector) representing the image in a latent space may be obtained from the classifier. In some embodiments, the embedding is obtained from a layer of the classifier's multiple layers prior to the last layer of the classifier. An exemplary DenseNet model architecture is described in "Densely Connected Convolutional Networks," Huang et al. (2016), arXiv:1608.06993, the contents of which are incorporated by reference in their entirety. In some embodiments, the layers are modified to receive a multi-channel fluorescence image as input and output a genetic background label. The embedding can be a vector representation of the input image in latent space. By converting the input image into an embedding using a classifier, the size and dimensionality of the original data can be significantly reduced. The embedding or a part thereof can be used to evaluate treatment responses, such as the treatment response of single cells of different genetic backgrounds. In some embodiments, the obtained embedding is further reduced in dimensionality using the UMAP algorithm to obtain a UMAP plot for visualization of the embedding, which can be used for downstream processing, as described herein. The UMAP plot or a part thereof can be used to evaluate treatment responses, such as the treatment response of single cells of different genetic backgrounds.
[0129] By obtaining embeddings from images, the system can generate a spatio-temporal topological space in which directionality is available. For example, each image can be converted into an embedding, and the embedding can be mapped to a location in the topological space and time-stamped with the time the image was captured. Thus, directionality across space and / or time can be obtained across multiple embeddings in the topological space. The embedding and subsequent UMAP plot can cluster data points representing single cells, such as single cells of different genetic backgrounds that have optionally undergone different treatments (e.g., genetic or chemical treatments), into specific parts of the embedding and / or UMAP plot. By overlaying images of individual cells onto their embedding coordinates, the cells cluster into specific locations in the topological space, making it clearer which features of the cell images (e.g., phenotypes) are causing the separation in the model. In some embodiments, the embedding and / or UMAP plot match the genotypes (e.g., genetic backgrounds) and physiological phenotypes obtained from the classifier as classification results.
[0130] The operations described above with reference to FIG. 6 are optionally implemented by the components shown in FIG. 12. FIG. 12 illustrates an example of a computing device according to an embodiment. The device 1200 may be a host computer connected to a network. The device 1200 may be a client computer or a server. As shown in FIG. 12, the device 1200 may be any suitable type of microprocessor-based device, such as a personal computer, a workstation, a server, or a handheld computing device (portable electronic device) such as a phone or tablet. The device 1600 may include, for example, one or more of a processor 1210, an input device 1220, an output device 1230, a storage device 1240, and a communication device 2160. The input device 1220 and the output device 1230 may generally correspond to those described above and may either be connectable to the computer or integrated with the computer.
[0131] Input device(s) 1220 may be any suitable device for providing input, such as a touch screen, a keyboard or keypad, a mouse, or a voice recognition device. Output device(s) 1230 may be any suitable device for providing output, such as a touch screen, a tactile device, or a speaker.
[0132] The storage device 1240 may be any suitable device that provides storage, such as electrical, magnetic, or optical memory, including RAM, cache memory, a hard drive, or a removable storage disk. The communication device 1260 may include any suitable device capable of sending and receiving signals over a network, such as a network interface chip or device. The components of a computer may be connected in any suitable manner, such as via a physical bus or wirelessly.
[0133] Software 1250 that can be stored in memory 1240 and executed by processor 1210 can include, for example, programming that embodies functions of the present disclosure (e.g., embodied in a device such as those described above).
[0134] The software 1250 may also be stored and / or transported in any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch and execute instructions associated with the software from the instruction execution system, apparatus, or device. In the context of this disclosure, a computer-readable storage medium may be any medium, such as storage device 1240, that includes or can store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0135] The software 1250 can also be propagated in any transport medium for use by or in connection with an instruction execution system, apparatus, or device, such as those described above, that can fetch instructions associated with the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of this disclosure, a transport medium can be any medium that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Transport-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.
[0136] The device 1200 may be connected to a network, which may be any suitable type of interconnected communication system. The network may implement any suitable communication protocol and may be protected by any suitable security protocol. The network may comprise any suitable arrangement of network links capable of implementing the transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.
[0137] Device 1200 may implement any operating system suitable for operating on a network. Software 1250 may be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying functionality of the present disclosure may be deployed in different configurations, for example, in a client / server arrangement, or through a web browser as a web-based application or web service.
[0138] Although the present disclosure and embodiments have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art, and such changes and modifications should be understood as being included within the scope of the disclosure and embodiments as defined by the claims.
[0139] The foregoing description has been described with reference to specific embodiments for purposes of illustration. However, the illustrative discussion above is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described in order to best explain the principles of the technology and their practical applications. In this way, those skilled in the art can best utilize the technology and various embodiments with various modifications suited to the particular use envisioned.
[0140] The present invention will be more fully understood by reference to the following examples. However, they should not be construed as limiting the scope of the present invention. It is understood that the examples and embodiments described herein are for illustrative purposes only, and various modifications or changes thereto may be suggested to those skilled in the art and should be included within the spirit and scope of this application and the scope of the appended claims.
[0141] [Example] Example 1: Modality using the method of pooled screening of cells from different genetic backgrounds The method of pooled screening of cells from different genetic backgrounds described in this application can be used in a variety of different assay modalities. The method of pooled screening of cells from different genetic backgrounds described herein is also called Visual Village in a Dish (ViViD). In each modality, a vector that codes for a unique nucleic acid barcode sequence is introduced into the cell to be assayed. An exemplary vector that codes for a unique nucleic acid barcode sequence is shown in FIG. 1.
[0142] A first modality using the method of pooled screening of cells from different genetic backgrounds described in this application is shown in FIG. 2. In this modality, a large collection of cell lines containing various genetic backgrounds is collected. These individual lineages are then labeled with unique nucleic acid barcode sequences, which are delivered and integrated into the genome. The cells, such as iPSCs, are then pooled and carried through any assay. In some embodiments, the modality includes performing ultra-throughput drug screening. At the end of the assay, the barcodes are read and the measured phenotype is linked to the genomic DNA that matches the barcode. Finally, machine learning or other statistical genetic methods are applied to the phenotype / genotype dataset to reveal the underlying biology, disease phenotype, drug response, etc.
[0143] A second modality using the method of pooled screening of cells from different genetic backgrounds described in this application is shown in Figure 3. In this modality, multiple engineered mutants from one or more healthy / patient strains are individually labeled, rather than pooling several genetic backgrounds as in the first modality. This allows specifically engineered mutants, including purified subclones, to be implemented in the assay. The mutant strains, once barcoded, are pooled and carried through any assay as previously described. In some embodiments, the modality includes performing ultra-throughput drug screening.
[0144] A third modality using the method of pooled screening of cells from different genetic backgrounds described in this application is shown in Figure 4. In this modality, standard ViViD is combined with perturbations. Multiple cell lines from several genetic backgrounds are collected and individually labeled with unique nucleic acid barcode sequences. These cells are also engineered to contain genetic modifiers and then treated with a pooled lentiviral library containing gRNAs targeting a selection or all of the genes. Alternatively, the lentiviral library can be an "all-in-one" vector. The pooled cells are then carried through any assay. In some embodiments, the modality includes performing ultra-throughput drug screening. Finally, a combination of genotypes and perturbations can be detected and linked to any cellular phenotype observed.
[0145] Example 2: Analysis of data obtained from pooled screening of cells from different genetic backgrounds. Described herein is an analysis pipeline that can be used to analyze data obtained from the method of pooled screening of cells from different genetic backgrounds described in this application. The method of pooled screening of cells from different genetic backgrounds described herein is also called Visual Village in a Dish (ViViD). The analysis pipeline includes non-parametric and parametric learnable methods for detecting and demultiplexing cell barcodes, segmenting cells, aligning image coordinates between acquisitions and other downstream tasks for learning meaningful expressions, and discovering unknown biology from data generated from the method of pooled screening of cells from different genetic backgrounds described herein.
[0146] Figure 5 is a schematic diagram showing the analytical pipeline used to analyze the data obtained from the method of pooled screening of cells from different genetic backgrounds described in this application. The analytical pipeline includes different subcomponents, starting from microscopic image acquisition and ending with single cell image-based readout for downstream analysis. The analytical pipeline includes the following steps: (0) Barcoding of images captured using a microscope (Sequencing by Synthesis) and marker-based / marker-free readout images (CellPaint / Antibody Staining / FISH / QPC) (1) Detection and sequencing of barcodes from individual base cycles imaged by sequencing by synthesis (2) Preprocessing of marker-based / marker-free cell images for downstream readout (3) Calculation of conversion between sequencing and marker-based / marker-free readouts to demultiplex single-cell image-based readouts (4) Mapping the sequenced barcodes to a dictionary of barcodes used in the experiment. (5) Segmentation of nuclei and cell boundaries to identify single cells in marker-based / marker-free readout images (6) Assignment of sequenced barcodes to corresponding single cells identified by cell segmentation using the transformation calculated from (3). (7) Marker-based / marker-free readout image tiling (and masking by cell segmentation) for downstream single-cell image analysis (8) A non-exhaustive list of possible single-cell readouts enabled by this pipeline (antibody / fluorescent protein markers, CellPaint, quantitative phase imaging, fluorescent in-situ hybridization).
[0147] Example 3: Pooled screening of cells from different genetic backgrounds using Visual Village in a Dish (ViViD) This example illustrates pooled screening of cells from different genetic backgrounds and its analysis using the methods described herein. Specifically, this example illustrates detection and demultiplexing of cell barcodes, cell segmentation, and alignment of image coordinates between acquisitions from data generated from the methods of pooled screening of cells from different genetic backgrounds described herein.
[0148] [Sample preparation, screening, and cell culture] Wild-type GM25256 induced pluripotent stem cells (iPSCs) (California Institute for Regenerative Medicine (CIRM)) were grown using standard Essential 8 (E8), vitronectin, and ReLeSR culture methods. Cells were engineered to contain the doxycycline-inducible neurogenin-2 (doxNGN2) gene at the AAVS1 safe locus, and the doxNGN2 vector contains a linked geneticin resistance gene, allowing continuous culture under antibiotic selection of these vectors.
[0149] Initial screening was performed using ribonucleoprotein (RNP) transfection of guide RNA (gRNA)-conjugated Cas9 to identify gRNAs effective for knocking out the TSC2 ("TSC2 ko") and SETD1A genes ("SETD1A ko"), respectively. iPSCs were then subcloned in single or double rounds to achieve pure homozygous or heterozygous knockout of the targeted genes. Wt control subclones were harvested through the same process but were derived from cells that did not undergo editing during the RNP transfection step. Homozygous knockouts of SETD1A did not survive in culture due to the essentiality of at least one copy of the gene for cell survival. Thus, a heterozygous knockout of SETD1A ("SETD1AG3 het") was generated using the number-specified SETD1A gRNA 3 and SETD1A gRNA4, resulting in cell lines labeled as heterozygous knockout of SETD1AG3 ("SETD1AG3 het") and heterozygous knockout of SETD1AG4 ("SETD1AG4 het"), respectively, and heterozygous knockout of TSC2 ("TSC2 het").
[0150] A lentiviral vector containing a non-targeting gRNA was synthesized. Each of the engineered strains was labeled with a unique gRNA tag, recorded as a specific "barcode" of the genotype. Cells with lentiviral integration of the gRNA expression transcript were selected with puromycin. The lentiviral vector also contained a feature extraction component, allowing the unique tag to be read during single-cell RNAseq (scRNAseq) analysis.
[0151] After each subclone was uniquely labeled, multiple engineered lines were pooled together for all downstream assays.The cells were then differentiated into neurons using induction of the NGN2 gene.Briefly, cells were dissociated via Accutase and seeded in E8 medium + Y27632 on day -2, then fed only E8 on day -1.On day 0, cells were treated to induce NGN2 expression on days 1 and 2.
[0152] On day 2, the culture plates were treated with 0.1% polyethyleneimine (PEI) in cell culture water and incubated overnight at 37°C. On day 3, the PEI plates were completely aspirated and dried. The PEI plates were then coated with laminin in PBS for 1 h at 37°C. The neural progenitor cells were then dissociated with Accutase for 5 min at 37°C and resuspended in NGN2 medium. The cells were filtered and plated at various densities.
[0153] For neuron-only cultures (e.g., scRNAseq samples), the laminin coating solution was aspirated and neural progenitor cells were plated onto PEI / laminin-coated culture plates in NGN2 medium.
[0154] For neuron-astrocyte co-cultures (e.g., imaging samples), neural progenitor cells were plated on PEI / laminin-coated culture plates simultaneously with rat astrocytes (Lonza). Each well received neurons and rat astrocytes at a 3:2 ratio, all in NGN2 medium. Medium was changed on a Monday / Wednesday / Friday schedule, with day 6 starting on a Monday. On each seeding day, 1 / 3 of the medium was removed from each well and replaced with fresh medium. Starting on day 11, cells were treated with NGN2 medium without doxycycline, continuing the Monday / Wednesday / Friday schedule.
[0155] [Imaging and segmentation] Starting on day 20, cells were treated with either DMSO control or 100 nM rapamycin. Rapamycin treatment was supplemented on days 22 and 24, with final paraformaldehyde (PFA) fixation of plates on day 27 (total treatment period of 7 days). Fixed cells were washed and the respective background genomic labels ("barcodes") were reverse transcribed and amplified by rolling circle amplification with gap filling.
[0156] Cells were stained with corresponding secondary antibodies Hoescht (1:200,000), 1:1,000 rab-phS6 antibody (Cell Signaling Tech), and 1:2,500 ch-MAP2 antibody (Novus) at 4° C. overnight and imaged via Nikon Ti2 with lumencor celesta excitation. Figure 13A shows an exemplary image of stained cells in a co-culture of neurons and astrocytes.
[0157] The cell images were then segmented by nuclei and cell boundaries to identify neuronal single cells in the readout images, as shown in Figure 13B. Briefly, the DAPI channel was thresholded, followed by applying a watershed transformation to obtain labeled nuclei instances, to generate a nuclei segmentation mask. The MAP2 channel was then thresholded, followed by a series of morphological operations to segment neuronal cells by applying a watershed transformation, using the nuclei segmentation mask as a seed.
[0158] [Barcode detection and sequencing] After imaging, the cells were subjected to several cycles of in situ sequencing by synthesis until the entire gRNA barcode label corresponding to the cell genetic background was measured in each cell. An exemplary image of in situ sequencing is shown in Figure 14A. The sequenced barcodes were mapped to the dictionary of barcodes used in the experiment. Barcode sequences were calculated from the sequence by synthesis images using a peak detection algorithm followed by classification by intensity of nucleotide bases based on the peak fluorescence channel. Figure 14B shows the number of POSH (e.g., pooled optical screening) barcodes summed for each cell line genetic background, and the corresponding number of cells based on the number of POSH barcodes is shown in Figure 14C.
[0159] [conversion] We then calculated a transformation between sequencing and marker-based / marker-free readouts for demultiplexing the single-cell image-based readouts. The sequenced barcodes were assigned to the corresponding single cells identified by cell segmentation using the calculated transformation. Briefly, well images from the first cycle of the sequencing-by-synthesis imaging run and the phenotypic (MAP2+DAPI) imaging run were then registered to calculate a transformation function. The barcode locations in the first cycle of the sequencing-by-synthesis were then transformed into phenotypic images using the calculated transformation function and assigned to single cells based on their segmentation labels. Once barcodes were assigned to cells, single neuronal images were generated using the segmentation mask and stored for subsequent modeling. To train a supervised classifier, samples from one well were used to create a test dataset, and the remaining wells were used for training. To classify between genetic backgrounds given single-cell inputs, we trained a supervised DenseNet model using the dataset with a cross-entropy loss function. To obtain a low-dimensional representation of these single-cell images, we utilized the activation output layer before the last layer of the network as a low-dimensional feature embedding for visualization. These embeddings were then further reduced in dimensionality using the UMAP algorithm to obtain two dimensions for visualization (eg, "UMAP1" and "UMAP2").
[0160] The supervised DenseNet model was able to achieve a classification accuracy of 45.5% on the test dataset. Figure 15A shows an example transformation to visualize barcode sequencing and imaging data (e.g., "UMAP1 and UMAP2" embeddings) of untreated ("No Tr") cells. By overlaying representative image tiles of individual cells onto their embedding coordinates (Figure 15B), it becomes clearer which features of the cell images are causing the separation in the model. Notably, the low UMAP1 / high UMAP2 (top left of Figures 15A and 15B) region contained more dense regions of pooled cells, with a mixture of multiple genotypes, resulting in cells of multiple genotypes in a single tile. Alternatively, clear differences in cell body size and neurite shape emerged in low UMAP1 / low UMAP2 (bottom left of Figures 15A and 15B (dashed circle)) compared to high UMAP1 (middle right of Figures 15A and 15B), consistent with the physiological phenotypes exhibited by neurons in patient hamartomas. In summary, UMAP visualization confirmed that TSC2 ko cells have a distinct feature embedding space, indicated by the circle with dashed line in FIG. 15A, compared to the remaining cells.
[0161] Figures 16A-16B show exemplary transformations visualizing barcode sequencing and imaging data (e.g., embeddings) of untreated cells ("No Tr") and cells treated with rapamycin ("Rapamycin"). The bottom left cluster (indicated by the circle with dashed line) of Figures 16A and 16B is primarily populated by untreated TSC2 ko cells, while untreated wt cells, SETD1A het cells, and rapamycin-treated TSC2 ko neurons all fall into other regions of the embedding.
[0162] [scRNAseq analysis] On day 14, neurons were treated with six treatments: rapamycin (100 nM), everolimus (100 nM), lonafarnib (100 nM), iademstat (100 nM), DMSO (1:100,000 dilution to match treatment), and no treatment. Seeding was performed to create target dilutions after 1 / 3 medium change. Cells were treated again with the same treatments on day 16. On day 17, neurons were dissociated with Accutase for 30 min at 37°C. Accutase was inactivated with NGN2, and cells were filtered and counted. Cells were resuspended in medium and subjected to standard 10X scRNAseq processing with PCR amplification feature extraction option. Sequencing libraries were aligned to the Hg38 transcript map using the standard CellRanger pipeline (10X). Genotype-labeled non-targeting gRNAs (e.g., "features") were assigned to cells based on the following: 1) the highest number of barcodes was assigned to a given cell if they constituted more than 55% of the total barcode reads, 2) the cell was labeled "multiple" (in preference to "1") if the second highest number of barcodes constituted more than 40% of the reads, and 3) the cell was labeled "unassigned" if no barcode counts were detected. By labeling the genetic backgrounds and tracking the treatment wells from which each pool originated, a combined dataset of 6 treatments with 14 genetic backgrounds was generated.
[0163] Figure 17 shows an exemplary embedding of pooled ViViD labeled cells with barcodes and transcriptional signatures measured using scRNAseq for untreated cells and cells treated with rapamycin, everolimus, lonafarnib, iadamstat, and DMSO. TSC2 ko neurons are shown in cluster 6 (solid arrow). Rapamycin treatment pushed all cells of all genotypes, including TSC2 ko neurons, into a new population (open arrow).
[0164] Overall, these data demonstrate that the Visual Village in a Dish (ViViD) pooled optical screening platform of cells from different genetic backgrounds can discover novel biology using large-scale population genetic assays based on specific cell types without convoluting factors.
[0165] Exemplary embodiments The following embodiments are illustrative and are not intended to limit the scope of the invention(s) described herein.
[0166] [Embodiment 1] a) labeling two or more populations of cells of different genetic backgrounds with two or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells; b) combining two or more of said populations of cells to obtain a single mixed population of cells; c) performing in situ single-cell sequencing on the cells; d) analyzing known phenotypes or identifying new phenotypes of the cells in the mixed population; and A method for pooled screening of cells from different genetic backgrounds comprising:
[0167] [Embodiment 2] 2. The method of embodiment 1, wherein said two or more populations of cells are from different cell lines.
[0168] [Embodiment 3] 3. The method of embodiment 2, wherein said different cell line is a healthy cell line.
[0169] [Embodiment 4] 3. The method of embodiment 2, wherein said different cell lines are patient cell lines.
[0170] [Embodiment 5] 3. The method of embodiment 2, wherein said different cell lines are isogenically engineered cell lines.
[0171] [Embodiment 6] 3. The method of embodiment 2, wherein said different cell lines comprise any combination of healthy cell lines, patient cell lines, and isogenically engineered cell lines.
[0172] [Embodiment 7] 7. The method of any one of embodiments 1 to 6, wherein the cells are induced pluripotent stem cells (iPSCs).
[0173] [Embodiment 8] 8. The method of embodiment 7, wherein the iPSCs are differentiated prior to analyzing the phenotype.
[0174] [Embodiment 9] 9. The method of embodiment 8, further comprising culturing said cells prior to analyzing said phenotype.
[0175] [Embodiment 10] 10. The method of any one of the preceding claims, wherein the single mixed population of cells is on a substrate or in a three-dimensional culture.
[0176] [Embodiment 11] 11. The method of embodiment 10, wherein the substrate is a cell culture dish.
[0177] [Embodiment 12] 12. The method of any one of embodiments 1 to 11, wherein the method further comprises performing single-cell RNAseq.
[0178] [Embodiment 13] 13. The method of any one of the preceding embodiments, wherein the method further comprises, prior to step b), propagating the two or more populations of cells for two or more generations.
[0179] [Embodiment 14] 14. The method of any one of embodiments 1 to 13, wherein the method comprises stably integrating the unique nucleic acid barcode sequence into the genome of two or more populations of the cells.
[0180] [Embodiment 15] 15. The method of embodiment 14, wherein the unique nucleic acid barcode sequence is delivered into the cell using a virus.
[0181] [Embodiment 16] 16. The method of embodiment 15, wherein the virus is a lentivirus.
[0182] [Embodiment 17] 17. The method of embodiment 15 or 16, wherein the virus encodes a selectable marker.
[0183] [Embodiment 18] 18. The method of embodiment 17, wherein the selectable marker is an antibiotic resistance gene.
[0184] [Embodiment 19] 19. The method of any one of embodiments 15 to 18, wherein the virus encodes a fluorescent protein.
[0185] [Embodiment 20] 20. The method of any one of the preceding claims, wherein each unique nucleic acid barcode sequence is at least one base pair in length.
[0186] [Embodiment 21] 21. The method of embodiment 20, wherein each unique nucleic acid barcode sequence is from 1 to about 18 base pairs in length.
[0187] [Embodiment 22] 21. The method of embodiment 20, wherein each unique nucleic acid barcode sequence is 8 base pairs in length.
[0188] [Embodiment 23] 23. The method of any one of the preceding claims, wherein the two or more populations of cells are sequenced prior to labeling with the unique nucleic acid barcode sequences.
[0189] [Embodiment 24] 24. The method of embodiment 23, wherein said sequencing is whole genome sequencing.
[0190] [Embodiment 25] 25. The method of any one of the preceding embodiments, wherein the two or more populations of cells are obtained from related individuals.
[0191] [Embodiment 26] 26. The method of embodiment 25, wherein the two or more populations of cells are obtained from a human.
[0192] [Embodiment 27] 27. The method of any one of embodiments 1 to 26, comprising labeling 10 or more populations of cells of different genetic backgrounds with 10 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells.
[0193] [Embodiment 28] 28. The method of any one of embodiments 1 to 27, wherein analyzing the phenotype of the cells comprises an assay selected from the group consisting of high content imaging, calcium imaging, immunohistochemistry, cell morphology imaging, protein aggregation imaging, cell-cell interaction imaging, live cell imaging, and any other image-based assay modality.
[0194] [Embodiment 29] 29. The method of any one of the preceding claims, wherein step d) comprises analyzing the phenotype of the cells by capturing a microscopic image or a time series of microscopic images of the cells, and evaluating phenotypic characteristics exhibited in the image or images.
[0195] [Embodiment 30] The method further comprising: generating a first reference coordinate space of a first plurality of images of a well on a culture plate; Extracting a first patch of the first plurality of images; generating a second reference coordinate space of a second plurality of images of the wells on the culture plate; Extracting second patches of the second plurality of images; and Calculating an affine transformation function between the first patch and the second patch to obtain a plurality of transformation parameters; generating a coordinate transformation function between the first reference coordinate space and the second reference coordinate space based on the plurality of transformation parameters; 30. The method of embodiment 29, further comprising computer-implemented techniques for registration between the first and second plurality of images, comprising:
[0196] [Embodiment 31] generating a first reference coordinate space of a first plurality of images of a well on a culture plate; Extracting a first patch of the first plurality of images; generating a second reference coordinate space of a second plurality of images of the wells on the culture plate; Extracting second patches of the second plurality of images; and Calculating an affine transformation function between the first patch and the second patch to obtain a plurality of transformation parameters; generating a coordinate transformation function between the first reference coordinate space and the second reference coordinate space based on the plurality of transformation parameters; 23. A computer-implemented method for registration between the first and second plurality of images, comprising:
[0197] [Embodiment 32] 32. The method of embodiment 31, wherein the first plurality of images is a plurality of barcoding images.
[0198] [Embodiment 33] 33. The method of embodiment 31 or 32, wherein the second plurality of images are a plurality of marker-based / marker-free readout images.
[0199] [Embodiment 34] 34. The method of any one of embodiments 31 to 33, wherein the first plurality of images and the second plurality of images provide different coverage of the well.
[0200] [Embodiment 35] 35. The method of any one of embodiments 33-34, wherein the first plurality of images and the second plurality of images are taken at different times.
[0201] [Embodiment 36] 36. The method of any one of embodiments 31 to 35, wherein the first plurality of images and the second plurality of images have different resolutions.
[0202] [Embodiment 37] 37. The method of any one of embodiments 31 to 36, wherein the first plurality of images are taken by a first microscope and the second plurality of images are taken by a second microscope.
[0203] [Embodiment 38] 38. The method of embodiment 37, wherein the first microscope is a fluorescence microscope.
[0204] [Embodiment 39] 38. The method of embodiment 37, wherein said second microscope is a non-fluorescence microscope.
[0205] [Embodiment 40] 40. The method of any one of embodiments 31 to 39, wherein the first plurality of images are captured by a first imager.
[0206] [Embodiment 41] detecting one or more physical properties of the wells in the first plurality of images; generating the first coordinate space of reference based on the one or more detected physical characteristics; 41. The method of embodiment 40, further comprising:
[0207] [Embodiment 42] 42. The method of embodiment 41, wherein the one or more physical characteristics of the well include a shape of the well, an edge of the well, and a position of the well.
[0208] [Embodiment 43] 43. A method according to any one of embodiments 40 to 42, wherein two or more images in the first plurality of images are offset from each other by an overlap ratio, and the first reference coordinate space is generated based on the overlap ratio.
[0209] [Embodiment 44] 44. A method according to any one of embodiments 40 to 43, wherein the first reference coordinate space is generated based on metadata of the first imager.
[0210] [Embodiment 45] A method as described in any one of embodiments 40 to 44, further comprising selecting one or more marker images from the first plurality of images based on one or more landmarks captured within one or more images, wherein the first patches are obtained from the one or more marker images from the first plurality of images.
[0211] [Embodiment 46] 46. The method of embodiment 45, wherein the one or more landmarks comprise one or more cells, one or more well borders, one or more beads, one or more nuclei, or any combination thereof.
[0212] [Embodiment 47] 47. The method of any one of embodiments 31 to 46, wherein the second plurality of images is captured by a second imager.
[0213] [Embodiment 48] detecting one or more physical properties of the wells in the second plurality of images; and generating the second coordinate space of reference based on the one or more detected physical characteristics; 48. The method of embodiment 47, further comprising:
[0214] [Embodiment 49] 49. The method of embodiment 48, wherein the one or more physical characteristics of the well include the shape of the well, the edge of the well, and the position of the well.
[0215] [Embodiment 50] 50. A method as described in any one of embodiments 47 to 49, wherein two or more images in the second plurality of images are offset from each other by an overlap ratio, and the second reference coordinate space is generated based on the overlap ratio.
[0216] [Embodiment 51] 51. A method according to any one of embodiments 47 to 50, wherein the second reference coordinate space is generated based on metadata of the second imager.
[0217] [Embodiment 52] A method as described in any one of embodiments 47 to 51, further comprising selecting one or more marker images from the second plurality of images based on one or more landmarks captured in the one or more images, wherein the second patches are obtained from the one or more marker images of the second plurality of images.
[0218] [Embodiment 53] 53. The method of embodiment 52, wherein the one or more landmarks comprise one or more cells, one or more well borders, one or more beads, one or more nuclei, or any combination thereof.
[0219] [Embodiment 54] 54. A method according to any one of embodiments 31 to 53, wherein the first patch covers the center of the first image.
[0220] [Embodiment 55] 55. The method of any one of embodiments 31 to 54, wherein the second patch covers the center of the second image.
[0221] [Embodiment 56] A method according to any one of embodiments 31 to 55, wherein at least a portion of the first patch and at least one of the second patches correspond to the same subject.
[0222] [Embodiment 57] 57. The method according to any one of embodiments 31 to 56, wherein the transformation parameters include one or more of a translation parameter, a scaling parameter, and a rotation parameter.
[0223] [Embodiment 58] a) labeling two or more populations of cells of different genetic backgrounds with two or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells; b) combining two or more of said populations of cells to obtain a single mixed population of cells; c) performing in situ single-cell sequencing on the cells; d) analyzing known phenotypes or identifying new phenotypes of the cells in the mixed population; and 58. The method of any one of embodiments 31 to 57, wherein the first image and the second image are obtained from an assay of pooled screening of cells from different genetic backgrounds, comprising:
[0224] [Embodiment 59] 59. The method of any one of embodiments 1 to 58, further comprising utilizing a classifier configured to receive the image and output a classification result.
[0225] [Embodiment 60] 60. The method of embodiment 59, wherein the classifier comprises multiple layers.
[0226] [Embodiment 61] 61. The method of embodiment 59 or 60, wherein the classifier is a convolutional neural network.
[0227] [Embodiment 62] 62. The method of any one of embodiments 59 to 61, wherein the classifier is a DenseNet classifier.
[0228] [Embodiment 63] 63. The method of any one of embodiments 59 to 62, wherein the classification result is a classification based on the genetic background of each single cell captured in the image.
[0229] [Embodiment 64] 64. The method of any one of embodiments 59 to 63, further comprising generating an embedding of the image.
[0230] [Embodiment 65] 65. The method of embodiment 64, wherein the embeddings are generated from an activation output layer before the last layer of the classifier.
[0231] [Embodiment 66] 66. The method of embodiment 64 or 65, wherein the embedding is reduced in dimension using a linear dimensionality reduction method.
[0232] [Embodiment 67] A method according to any one of embodiments 64 to 66, wherein the embedding is dimensionally reduced using a Uniform Manifold Approximation and Projection for Dimension Reduction (UMAP) algorithm to obtain one or more UMAP plots for visualization of the embedding.
[0233] [Embodiment 68] 68. The method of any one of embodiments 64 to 67, further comprising evaluating a process based at least in part on the embedding.
[0234] [Embodiment 69] 69. The method of embodiment 67 or 68, further comprising evaluating the treatment based at least in part on the one or more UMAP plots.
Claims
1. a) labeling two or more populations of cells of different genetic backgrounds with two or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells; b) combining two or more populations of said cells to obtain a single mixed population of cells; c) performing in situ single-cell sequencing on the cells; d) analyzing known phenotypes or identifying new phenotypes of the cells in the mixed population; A method for pooled screening of cells from different genetic backgrounds, comprising:
2. The method of claim 1 , wherein the two or more populations of cells are from different cell lines.
3. 3. The method of claim 2, wherein the different cell lines are selected from the group consisting of healthy cell lines, patient cell lines, isogenically engineered cell lines, and any combination thereof.
4. 4. The method of any one of claims 1 to 3, wherein the cells are induced pluripotent stem cells (iPSCs), and optionally the iPSCs are differentiated before analyzing the phenotype, and optionally further comprising culturing the cells before analyzing the phenotype.
5. 5. The method of any one of claims 1 to 4, wherein the single mixed population of cells is on a substrate or in a three-dimensional culture, optionally wherein the substrate is a cell culture dish.
6. 6. The method of claim 1, further comprising performing single-cell RNA sequencing.
7. 7. The method of any one of claims 1 to 6, wherein the method comprises stably integrating the unique nucleic acid barcode sequence into the genome of two or more populations of the cells.
8. 8. The method of claim 7, wherein the unique nucleic acid barcode sequence is delivered into the cell using a virus, optionally wherein the virus is a lentivirus.
9. The method of claim 8 , wherein the virus encodes a selectable marker.
10. 10. The method of claim 9, wherein the selectable marker is an antibiotic resistance gene.
11. 11. The method of any one of claims 8 to 10, wherein the virus encodes a fluorescent protein.
12. 12. The method of any one of claims 1-11, wherein each unique nucleic acid barcode sequence is at least 1 base pair in length, optionally each unique nucleic acid barcode sequence is 1 to about 18 base pairs in length, and optionally each unique nucleic acid barcode sequence is 8 base pairs in length.
13. 13. The method of any one of claims 1-12, wherein the two or more populations of cells are sequenced prior to labeling with the unique nucleic acid barcode sequences.
14. 14. The method of claim 13, wherein the sequencing is whole genome sequencing.
15. 15. The method of any one of claims 1 to 14, wherein the two or more populations of cells are obtained from related individuals, optionally wherein the two or more populations of cells are obtained from a human.
16. 16. The method of any one of claims 1-15, comprising labeling 10 or more populations of cells of different genetic backgrounds with 10 or more unique nucleic acid barcode sequences, each unique nucleic acid barcode sequence corresponding to a different population of cells.
17. 17. The method of any one of claims 1 to 16, wherein analyzing the phenotype of the cells comprises an assay selected from the group consisting of high-content imaging, calcium imaging, immunohistochemistry, cell morphology imaging, protein aggregation imaging, cell-cell interaction imaging, live-cell imaging, and any other image-based assay modality.
18. 18. The method of any one of claims 1 to 17, wherein step d) comprises analyzing the phenotype of the cells by capturing a microscopic image or a time series of microscopic images of the cells and evaluating phenotypic characteristics exhibited in the image or images.
19. The method comprises: generating a first reference coordinate space of a first plurality of images of wells on a culture plate; Extracting first patches of the first plurality of images; generating a second reference coordinate space of a second plurality of images of the wells on the culture plate; Extracting second patches of the second plurality of images; calculating an affine transformation function between the first patch and the second patch to obtain a plurality of transformation parameters; generating a coordinate transformation function between the first reference coordinate space and the second reference coordinate space based on the plurality of transformation parameters; 20. The method of claim 18, further comprising computer-implemented techniques for registration between the first and second plurality of images, comprising:
20. generating a first reference coordinate space of a first plurality of images of wells on a culture plate; Extracting first patches of the first plurality of images; generating a second reference coordinate space of a second plurality of images of the wells on the culture plate; Extracting second patches of the second plurality of images; calculating an affine transformation function between the first patch and the second patch to obtain a plurality of transformation parameters; generating a coordinate transformation function between the first reference coordinate space and the second reference coordinate space based on the plurality of transformation parameters; a computer-implemented method for registration between the first and second plurality of images, comprising: