Means and methods for linking genetic perturbations or the expression of a gene or RNA of interest with phenotypes of cells
The method addresses inefficiencies in optical pooled perturbation screening by using dual-promoter nucleic acid molecules and in situ sequencing to precisely link genetic perturbations to individual cell phenotypes, enhancing accuracy and applicability to high-complexity cell libraries.
Patent Information
- Application Number
- PCT/EP2025/051042
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-24
AI Technical Summary
Current methods for optical pooled perturbation screening in genetic perturbation screens face limitations such as the need for high transcriptional activity in cells, reliance on mRNA molecules in the cytosol, and challenges in assigning perturbagens to individual cells in densely packed, high-complexity libraries, leading to inefficiencies and inconsistencies.
A method involving the introduction of nucleic acid molecules with dual promoters, where one promoter controls expression in the 5'-3' direction and the other in the 3'-5' direction, followed by fixing and permeabilizing cells, generating barcode sequences, and using in situ sequencing by synthesis to link genetic perturbations to individual cell phenotypes, allowing for precise assignment of perturbagens.
This method enables accurate linking of genetic perturbations to individual cell phenotypes, overcoming limitations of current methods by providing reliable assignment of perturbagens to individual cells, even in densely packed libraries, independent of transcriptional activity and cell size.
Smart Images

Figure IMGF000039_0001 
Figure IMGF000040_0001 
Figure IMGF000040_0002
Abstract
Description
[0001] Means and methods for linking genetic perturbations or the expression of a gene or RNA of interest with phenotypes of cells
[0002] The present invention relates to a method for linking the genetic perturbations of individual cells within a cell population to the phenotype of the individual cell s, comprising (a) introducing a plurality of different nucleic acid molecules into the genome of the cells, wherein each of the nucleic acid molecules comprises a nucleotide sequence encoding a genetic pertubator, a first promoter and a second promoter, wherein the first promoter is 5' of the nucleotide sequence encoding the genetic pertubator and controls the expression of the genetic pertubator in 5'-3' direction in the cells, wherein the second promoter is a reverse phage promoter and / or reverse in vitro transcription promoter 3' of the nucleotide sequence encoding the genetic pertubator and controls the expression of the genetic pertubator in 3'-5' direction in the nucleus of the cells; (b) initiating the expression of the nucleic acid sequences encoding the genetic pertubators in 5' -3' direction from the first promoter in the cells, under conditions wherein genetic perturbations are introduced into the cells of the cell population via the expressed genetic pertubators; (c) fixing and permeabilizing the cells; (d) initiating the expression of the nucleic acid sequences encoding the genetic pertubators in 3' -5' direction from the second promoter in the fixed cells, thereby generating barcode sequences that are presentative for each kind of genetic perturbation as introduced via the different genetic pertubators; (e) optionally fixing the cells, preferably on a surface; (f) reverse transcription of the barcode sequences; (g) optionally amplification of the reversely transcribed barcode sequences; (h) identifying the reversely transcribed barcode sequences in the individual cells within the cell population by in situ sequencing by synthesis thereby assigning a particular genetic perturbation to each individual cell; and (i) linking the genetic perturbation of each individual cell within the phenotype of each individual cell, wherein the phenotype of each individual cell (I) has been imaged or monitored in the living cells by microscopy any time before step (c) and preferably after step (b), wherein the imaged or monitored phenotype of each individual cell is preferably the cell state, the cell motility, the cell shape, the cell-cell interactions, the fluorescence tag of one or more targets in the living cell, the affinity reagent staining of a target in the living cell and / or the fluorescence or color of a cell staining dye, or (II) has been imaged or monitored in the fixed cells by microscopy any time after step (c) and preferably any time before step (f), wherein the imaged or monitored phenotype of each individual cell is preferably the shape of the fixed cell, the fluorescence tag of one or more targets in the fixed cell, the affinity reagent staining of a target in the fixed cell, and / or the sequence-specific detection of one or more RNA and / or DNA sequences in the fixed cell.
[0003] In this specification, a number of documents including patent applications and manufacturer's manuals are cited. The disclosure of these documents, while not considered relevant for the patentability of this invention, is herewith incorporated by reference in its entirety. More specifically, all documents are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference.
[0004] Genetic perturbation screens aim to decipher the relationship between genotype and phenotype by ablating random genes in a pool of cells. Perturbation screens are commonly based on enrichment of cells with a phenotype of interest followed by next generation sequencing of the perturbagens, e.g., CRISPR sgRNAs1 2or gene trap insertions in haploid cells3. Enrichment of specific sgRNA sequences allows to conclude to the biological function of targeted genes. While enrichment is commonly achieved through cell-growth or fluorescence activated cell sorting (FACS)4 5, more sophisticated screening approaches were realized combining perturbation screens with single cell molecular profiling via mass cytometry or RNA sequencing6 7 8. Those methods acquire high-dimensional information on single cell level, but are incapable of monitoring dynamic processes due to their disruptive nature.
[0005] For arrayed CRISPR screening, perturbed cells are physically separated into compartments containing cells with largely homogeneous genotypes. This allows to analyze cells using complex assays, such as fluorescence microscopy9, and thus to monitor cell physiology in a spatially and temporally resolved manner. However, as arrayed screens are laborious and prone to experimental inconsistencies between compartments, they are mostly feasible for small targeted perturbagen libraries.
[0006] Optical pooled perturbation screening addresses these limitations by accessing the bandwidth of phenotypes that can be observed by fluorescence microscopy while working with a complex pool of cells10. Currently, several approaches to optical pooled perturbation screening were developed: For Optical-Based Enrichment, cells with a phenotype of interest are identified by microscopy and are marked by light-based conversion of a photoactivatable fluorescent protein. Subsequently, cells are dissociated, marked cells are enriched by FACS sorting, and perturbagens are analyzed by deep sequencingn. To a similar end, image-based cell sorting12or laser microdissection13can be used to isolate cells displaying a microscopic phenotype of interest for subsequent perturbagen identification, albeit at limited optical resolution. Instead of enriching for a population of cells for lysis and sequencing, Feldman et al.10developed a method for microscopic perturbagen identification based on in-situ sequencing of barcodes contained in cellular mRNAs, using signal amplification by rolling circle amplification (RCA). Using this method, both the phenotype and the genotype are recorded from individual cells using fluorescence microscopy, allowing to categorize cellular phenotypes post-hoc14. This method was recently adapted to screen genome-scale perturbation libraries for pre-defined15and multi-dimensional phenotypes16 17.
[0007] While in-situ sequencing-based screening has the advantage of post-hoc genetic dissection of multidimensional phenotypes, current methods still bear two major limitations: Because sequencing is dependent on mRNA molecules in the cytosol, cells with low transcriptional activity or small volume generate insufficient signal, precluding most cell types from optical pooled screening. Furthermore, since sequencing spots are located in the cytosol, either local clusters of clonal cells bearing the same perturbagen or precise three-dimensional cell boundary detection are required for reliable assignment of perturbagens to individual cells, which is largely incompatible with screening densely packed, high- complexity genome-scale cell libraries.
[0008] The present disclosure addresses and overcomes these shortcomings.
[0009] Accordingly, the present invention relates in a first aspect to a method for linking the genetic perturbations of individual cells within a cell population to the phenotype of the individual cells, comprising (a) introducing a plurality of different nucleic acid molecules into the genome of the cells, wherein each of the nucleic acid molecules comprises a nucleotide sequence encoding a genetic pertubator, a first promoter and a second promoter, wherein the first promoter is 5' of the nucleotide sequence encoding the genetic pertubator and controls the expression of the genetic pertubator in 5'- 3' direction in the cells, wherein the second promoter is a reverse phage promoter and / or reverse in vitro transcription promoter 3' of the nucleotide sequence encoding the genetic pertubator and controls the expression of the genetic pertubator in 3' -5' direction in the nucleus of the cells; (b) initiating the expression of the nucleic acid sequences encoding the genetic pertubators in 5' -3' direction from the first promoter in the cells, under conditions wherein genetic perturbations are introduced into the cells of the cell population via the expressed genetic pertubators; (c) fixing and permeabilizing the cells; (d) initiating the expression of the nucleic acid sequences encoding the genetic pertubators in 3'-5' direction from the second promoter in the fixed cells, thereby generating barcode sequences that are presentative for each kind of genetic perturbation as introduced via the different genetic pertubators; (e) optionally fixing the cells, preferably on a surface; (f) reverse transcription of the barcode sequences; (g) optionally amplification of the reversely transcribed barcode sequences; (h) identifying the reversely transcribed barcode sequences in the individual cells within the cell population by in situ sequencing by synthesis thereby assigning a particular genetic perturbation to each individual cell; and (i) linking the genetic perturbation of each individual cell within the phenotype of each individual cell, wherein the phenotype of each individual cell (I) has been imaged or monitored in the living cells by microscopy any time before step (c) and preferably after step (b), wherein the imaged or monitored phenotype of each individual cell is preferably the cell state, the cell motility, the cell shape, the cell-cell interactions, the fluorescence tag of one or more targets in the living cell, the affinity reagent staining of a target in the living cell and / or the fluorescence or color of a cell staining dye, or (II) has been imaged or monitored in the fixed cells by microscopy any time after step (c) and preferably any time before step (f), wherein the imaged or monitored phenotype of each individual cell is preferably the shape of the fixed cell, the fluorescence tag of one or more targets in the fixed cell, the affinity reagent staining of a target in the fixed cell, and / or the sequence-specific detection of one or more RNA and / or DNA sequences in the fixed cell.
[0010] Genetic perturbation (or gene perturbation) refers to a specific alteration of the genome of a cell, in particular a gene function by a molecule. Genes within cells can be perturbed through a number of means. Gene perturbations can be, for example, through gene mutation (e.g. deletions or substitutions), gene knockout, gene knockdown, gene overexpression, promoter replacement, or 3'UTR disruption. Genetic perturbations are reviewed, for example, in Ishikawa and Saitoh, Biomolecules, 2023; 13(4): 716.
[0011] Accordingly, a genetic pertubator as used herein refers to a molecule being capable to specifically modifying the genome of a cell, in particular gene function. The genetic pertubator is encoded by a nucleic acid sequence and can therefore be a nucleic acid molecule (e.g. RNA or DNA) or a (poly)peptide.
[0012] The term "nucleic acid molecule" in accordance with the present invention includes DNA, such as cDNA or double or single stranded genomic DNA and RNA. In this regard, "DNA" (deoxyribonucleic acid) means any chain or sequence of the chemical building blocks adenine (A), guanine (G), cytosine (C) and thymine (T), called nucleotide bases, that are linked together on a deoxyribose sugar backbone. DNA can have one strand of nucleotide bases, or two complimentary strands which may form a double helix structure. "RNA" (ribonucleic acid) means any chain or sequence of the chemical building blocks adenine (A), guanine (G), cytosine (C) and uracil (U), called nucleotide bases, that are linked together on a ribose sugar backbone. RNA typically has one strand of nucleotide bases, such as mRNA. Included are also single- and double-stranded hybrids molecules, i.e., DNA-DNA, DNA-RNA and RNA-RNA. The nucleic acid molecule may also be modified by many means known in the art. Non-limiting examples of such modifications include methylation, "caps", substitution of one or more of the naturally occurring nucleotides with an analog, and internucleotide modifications such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoroamidates, carbamates, etc.) and with charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.). Nucleic acid molecules, in the following also referred as polynucleotides, may contain one or more additional covalently linked moieties, such as, for example, proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), intercalators (e.g., acridine, psoralen, etc.), chelators (e.g., metals, radioactive metals, iron, oxidative metals, etc.), and alkylators. The polynucleotides may be derivatized by formation of a methyl or ethyl phosphotriester or an alkyl phosphoramidate linkage. Further included are nucleic acid mimicking molecules known in the art such as synthetic or semi-synthetic derivatives of DNA or RNA and mixed polymers. Such nucleic acid mimicking molecules or nucleic acid derivatives according to the invention include phosphorothioate nucleic acid, phosphoramidate nucleic acid, 2'-O-methoxyethyl ribonucleic acid, morpholino nucleic acid, hexitol nucleic acid (HNA), peptide nucleic acid (PNA) and locked nucleic acid (LNA) (see Braasch and Corey, Chem Biol 2001, 8: 1). LNA is an RNA derivative in which the ribose ring is constrained by a methylene linkage between the 2' -oxygen and the 4' -carbon. Also included are nucleic acids containing modified bases, for example thio-uracil, thio-guanine and fluoro-uracil. A nucleic acid molecule typically carries genetic information, including the information used by cellular machinery to make proteins and / or polypeptides. The nucleic acid molecule of the invention may additionally comprise promoters, enhancers, response elements, signal sequences, polyadenylation sequences, introns, 5'- and 3'- noncoding regions, and the like.
[0013] The term "protein" as used herein interchangeably with the term "polypeptide" describes linear molecular chains of amino acids, including single chain proteins or their fragments, containing at least 50 amino acids. The term "peptide" as used herein describes a group of molecules consisting of up to 49 amino acids, whereas the term "polypeptide" (also referred to as "protein") as used herein describes a group of molecules consisting of at least 50 amino acids. The term "peptide" as used herein describes a group of molecules consisting with increased preference of at least 15 amino acids, at least 20 amino acids at least 25 amino acids, and at least 40 amino acids. The group of peptides and polypeptides are referred to together by using the term "(poly)peptide". (Poly)peptides may further form oligomers consisting of at least two identical or different molecules. The corresponding higher order structures of such multimers are, correspondingly, termed homo- or heterodimers, homo- or heterotrimers etc. Furthermore, peptidomimetics of such proteins / (poly)peptides where amino acid(s) and / or peptide bond(s) have been replaced by functional analogues are also encompassed by the invention. Such functional analogues include all known amino acids other than the 20 gene-encoded amino acids, such as selenocysteine. The terms "(poly)peptide" and "protein" also refer to naturally modified (poly)peptides and proteins where the modification is effected e.g. by glycosylation, acetylation, phosphorylation and similar modifications which are well known in the art.
[0014] The cells of the cell population are not particularly limited. They can be prokaryotic cells or eukaryotic cells.
[0015] Suitable prokaryotes (bacteria) useful for the invention are, for example, those generally used for cloning and / or expression like E. coli (e.g., E coli strains BL21, HB101, DH5a, XL1 Blue, Y1090 and J MIDI), Salmonella typhimurium, Serratia marcescens, Burkholderia glumae, Pseudomonas putida, Pseudomonas fluorescens, Pseudomonas stutzeri, Streptomyces lividans, Lactococcus lactis, Mycobacterium smegmatis, Streptomyces coelicolor or Bacillus subtilis. Appropriate culture mediums and conditions for the above-described host cells are well known in the art.
[0016] A suitable eukaryotic cell may be a vertebrate cell, an insect cell, a fungal / yeast cell, a nematode cell or a plant cell. The fungal / yeast cell may a Saccharomyces cerevisiae cell, Pichia pastoris cell or an Aspergillus cell.
[0017] In a different preferred embodiment the cell is a mammalian cell, such as a Chinese Hamster Ovary (CHO) cell, mouse myeloma lymphoblastoid, human embryonic kidney cell (HEK-293), human embryonic retinal cell (Crucell's Per.C6), or human amniocyte cell (Glycotope and CEVEC). The cells are frequently used in the art to produce recombinant proteins. CHO cells are the most commonly used mammalian cells for industrial production of recombinant protein therapeutics for humans.
[0018] A promoter is a sequence of DNA to which proteins bind to initiate transcription of a single RNA transcript from the DNA downstream of the promoter.
[0019] In accordance with the claimed method two promoters are used, first promoter and a second promoter: (1) The first promoter controls the expression of the genetic pertubator in 5' -3' direction in the cells, and (2) the second promoter is a reverse phage promoter and / or reverse in vitro transcription that controls the expression of the genetic pertubator in 3' -5' direction in the nucleus of the cells. With a "normal" promoter transcription occurs in the 5' - 3' direction. This means that it will start at the 3' end of its complementary strand and proceed in the 5' - 3' direction as it copies the gene on DNA into RNA. The first promoter is such a "normal" promoter.
[0020] On the other hand, with a reverse promoter transcription occurs in the 3' - 5' direction. This means that it will start at the 5' end of its complementary strand and proceed in the 3' - 5' direction as it copies the gene on DNA into RNA. The second promoter is a reverse promoter (or reverse complement promoter).
[0021] The second promoter is in addition a phage promoter or in vitro transcription promoter.
[0022] In vitro transcription is the DNA-dependent synthesis of RNA in a test tube. In vitro transcription was established by the laboratory of Douglas A. Melton. He was able to show that large amounts of RNA can be produced using a plasmid vector that has a bacteriophage SP6 promoter followed by an open reading frame (ORF), the SP6 RNA polymerase and four ribonucleotides. The system was subsequently expanded with the promoters of the bacteriophages T3 and T7. The synthesis with the T7 RNA polymerase proved to be particularly efficient.
[0023] Examples of bacteriophage promoters or in vitro transcription promoters are therefore T7, T3, and SP6 that each consist of 23 basepairs numbered -17 to +6, where +1 indicates the first base of the coded transcript. All of T7, T3, and SP6 can be reversed in order to obtain a revere phage promoter or in vitro transcription promoter.
[0024] The second promoter controls the expression of the genetic pertubator in 3' -5' direction in the nucleus of the cells which nuclear expression is achieved by local transcription from genomic DNA using the reverse phage promoter or in vitro transcription promoter (e.g. reverse T7 promoter) and a phage or vitro transcription polymerase (e.g. T7 polymerase).
[0025] The following SEQ. ID NO: 1 is an example of nucleic acid molecule comprising a nucleotide sequence encoding a genetic pertubator (e.g. a guideRNA (gRNA or sgRNA)), a first promoter (e.g. U6 promoter) and a second promoter (e.g. reverse T7 promoter), wherein the first promoter is 5' of the nucleotide sequence encoding the genetic pertubator and controls the expression of the genetic pertubator in 5'- 3' direction in the cells, wherein the second promoter is a reverse phage promoter and / or reverse in vitro transcription promoter 3' of the nucleotide sequence encoding the genetic pertubator and controls the expression of the genetic pertubator in 3'-5' direction in the nucleus of the cells. GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTAGAATTA ATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTTG CAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCT TTATATATCTTGTGGAAAGG ACG AAACACCG NNNNNNNNNNNNNNNNNNN NGTTTTAGAGCTAGAAA TA GCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC\ I I I I I GATCCCTATA GTGAGTCGTATTA.(SEQ ID NO: 1)
[0026] U6 promoter (human) sgRNA target-specific sequence sgRNA constant scaffold
[0027] Polymerase-Ill terminator
[0028] T7 _ _rPrn oter reverse_cpm plern ent_(RN A starts wjt_h_5'-GGG_ATC. „ )
[0029] In this respect it is already noted that a sgRNA comprises two parts; a constant scaffold and targetspecific sequence. The constant scaffold interacts with the CRISPR-Cas nuclease and target-specific sequence interacts with a target sequence, for example within the genome of a cell thereby guiding the CRISPR-Cas nuclease to its target sequence and mediating a target-specific genomic perturbation (or in this case genome editing). As will be further explained herein below, the presence of a polymerase terminator is optional but preferred since it reduces the length of the expressed genetic pertubator to the required length. It is also of note that the sgRNA target-specific sequence of SEQ. ID NO: 1 has 20 nucleotides and that the length of the sgRNA target-specific sequence can be between 18 and 22 nucleotides.
[0030] It is to be understood that steps (a) and (b) of the method of the invention are carried out with living cells and that the fixation and permeabilization in step (c) "kills" the cells.
[0031] In step (a) of the method of the invention a plurality of the above-described nucleic acid molecules are to be introduced into the genome of the cells of the cell population that are different from each other. The plurality of different nucleic acid molecules are different from each other with respect to the nucleic acid sequence encoding a genetic pertubator, so that by the plurality of different nucleic acid molecules a plurality of different a genetic pertubations is introduced into the different cells of the cell population. Means and methods for introducing a plurality of the above-described nucleic acid molecules into the genome of the cells of the cell population are known in the art and preferred examples will be further described herein below.
[0032] In step (b) the genetic pertubations are introduced into the cells of the cell population by initiating the expression of the nucleic acid sequences encoding the genetic pertubators in 5'-3' direction from the first promoter in the cell s, under conditions wherein genetic perturbations are introduced into the cells of the cell population via the expressed genetic pertubator. The conditions wherein genetic perturbations are introduced into the cells of the cell population via the expressed genetic pertubator depend on the nature of the used genetic perturbators to be used and can be set based on the prior art knowledge. For example, in case the genetic perturbations are gRNAs the conditions require the presence of a polymerase that initiates the expression of the gRNAs from the first promoter (e.g. a polymerase III in the case of U promoter) and a CRISPR-Cas nuclease (e.g. Cas9, Cpfl, SRGN3.1 or a small Cas9s; see Schmidt et al., Nature Communications volume 12, Article number: 4219 (2021) and Seo et al., Nature Methods volume 20, pages999-1009 (2023)) that adds the genetic pertubations together with the gRNAs.
[0033] SRGN3.1 and the small Cas9s are smaller in size than other CRISPR-Cas nucleases, such as Cas9 and Cpfl. This has the technical advantage that the same vector (e.g. a lentiviral vector) can encode the genetic pertubator as well as the CRISPR-Cas nuclease without exceeding the optimal or maximum insert size of a vector, thereby obtaining an "all-in-one" vector.
[0034] In step (c) the cells with the introduced genetic pertubations are fixed and permeabilized. Means and method for the fixation and permeabilization are known in the art and preferred fixation and permeabilization steps will be further detailed herein below.
[0035] In step (d) the expression of the nucleic acid sequences encoding the genetic pertubators in 3' -5' direction from the second promoter in the fixed cells is initiated. This reverse transcription does not result in the expression of the genetic pertubators but instead in the expression of the reverse complement of the genetic pertubators that serve barcode sequences that are presentative for each kind of genetic perturbation as introduced via the different genetic pertubators.
[0036] In the optional but preferred step (e) the cells are again fixed, preferably on a surface. Also for step (b) preferred fixation steps will be further detailed herein below. The fixation on a surface is advantageous for cell tracking throughout the method. The nature of the surface is not particularly limited and can be, for example, a glass or (natural or synthetic) polymer (e.g. plastic) surface. Suitable polymers are, for example, poly(tetrafluoroethylene) (PTFE), polystyrene (PS), and polyurethane The surface is preferably a flat or 2D-surface. The flat or 2D-surface can be in the form of a well, well plate or slide, for example.
[0037] In step (f) the barcode sequences are reversely transcribed thereby obtaining reversely transcribed barcode sequences. This is done by a reverse transcriptase. A reverse transcriptase (RT) is an enzyme used to generate complementary DNA (cDNA) from an RNA template, a process termed reverse transcription. Well-studied reverse transcriptases include: HIV-1 reverse transcriptase from human immunodeficiency virus type 1, M-MLV reverse transcriptase from the Moloney murine leukemia virus, AMV reverse transcriptase from the avian myeloblastosis virus and Induro® Reverse Transcriptase which is a group II intron-encoded RT (NEB).
[0038] Step (g) is an optional but preferred step. In step (g) the reversely transcribed barcode sequences are amplified and the amplification is expected to improve the accuracy of the subsequent identification step (h). cDNA amplification protocols are known in the state of the art.
[0039] In step (h) the reversely transcribed barcode sequences are identified in the individual cells within the cell population by in situ sequencing by synthesis, thereby assigning a particular genetic perturbation to each individual cell of the initial cell population.
[0040] Means and method for in situ sequencing by synthesis (SBS) of the reversely transcribed barcode sequences are available in the state of the art and include next generation sequencing protocols. Sequencing by synthesis is a DNA sequencing method in which DNA polymerases and dNTPs are used to replicate (synthesize) the strand to be sequenced. Nucleotides are introduced either on a single time or modified with identifying tags (e.g. fluorophore) so that the base type of the incorporated nucleotide can be recognized as the DNA molecule extends. Besides, these nucleotides are chemically blocked such that each incorporation is a unique event. After the signal of each incorporation step is collected, the blocked group will be removed to prepare the strand for the next incorporation event by DNA polymerase. By continuing this series of steps for a specific number of cycles, the strand of interest is replicated and sequenced. Sequencing by synthesis (SBS) technology generally uses four nucleotides differentially labelled with a unique combination of two to four fluorescent dyes to sequence the tens of millions of clusters on the flow cell surface in parallel and has been developed by Illumina. Finally in step (i) the genetic perturbation of each individual cell is linked within the phenotype of each individual cell. By this linkage it can be determined whether the genetic perturbation has an influence on (e.g. changes) the phenotype of a cell as compared to a cell without the genetic perturbation.
[0041] The phenotype of each individual cell
[0042] (I) has been imaged or monitored in the living cells by microscopy any time before step (c) and preferably after step (b), or
[0043] (II) has been imaged or monitored in the fixed cells by microscopy any time after step (c) and preferably any time before step (f).
[0044] The method of the invention may also additionally comprise the active step of (I) imaging or monitoring in the living cells by microscopy any time before step (c) and preferably after step (b), or (II) imaging or monitoring in the fixed cells by microscopy any time after step (c) and preferably any time before step (f).
[0045] According to option (I) the imaged or monitored phenotype of each individual cell is preferably the cell state, the cell motility, the cell shape, the cell-cell interactions, the fluorescence tag of one or more targets in the living cell, the affinity reagent staining of a target in the living cell and / or the fluorescence or color of a cell staining dye.
[0046] According to option (II) the imaged or monitored phenotype of each individual cell is preferably the shape of the fixed cell, the fluorescence tag of one or more targets in the fixed cell, the affinity reagent staining of a target in the fixed cell, and / or the sequence-specific detection of one or more RNA and / or DNA sequences in the fixed cell.
[0047] With respect to option (I) it is to be understood that "before step (c)" means that the phenotype of living cells is determined and that "before step (c)" and "after step (b)" means that the phenotype of living cells is determined, wherein a genetic perturbation has been introduced.
[0048] With respect to option (II) it is to be understood that "after step (c)" means that the phenotype of fixed cells is determined and that "after step (c)" and "any time before step step (f)" means that the phenotype of fixed cells is determined before the barcode sequences are reversely transcribed.
[0049] The imaged or monitored phenotype of each cell as obtained by microscopy can be stored as images of the cells and / or image information (e.g. as cell or nuclear dimensions, cell movements, coordinates of a surface, fluorophore dye color and / or intensity) on a storage medium. A storage medium is a physical device that receives and retains electronic data for applications and users and makes the data available for retrieval. The storage medium might be inside a computer or other device or attached to a system externally, either directly or over a network. A storage medium may be internal to a computing device, such as a computer's SSD, or a removable device such as an external HDD or universal serial bus (USB) flash drive. There are also other types of storage media, including magnetic tape, compact discs (CDs) and non-volatile memory (NVM) cards.
[0050] It is to be understood that after the phenotype of each cell has been imaged or monitored it is to be ensured that each cell can be tracked throughout the following steps of the methods, for example and as illustrated by the examples, by keeping the cells on the same surface throughout the following steps, so that in the final step (i) the determined genetic perturbation of each individual cell can be linked within the known and predetermined phenotype of each individual cell.
[0051] Depending on the desired phenotype the microscope can be, for example, a bright field microscope and / or a fluorescent microscope.
[0052] As can be taken from the appended examples the method of the invention is also designated nuclear in-situ sequencing (NIS-Seq). By NIS-Seq after phenotyping live cells and fixation, the reversecomplement sequences of perturbation gRNAs can be locally transcribed from genomic DNA using T7 polymerase by an adapted Zombie protocol18. Nuclear clusters of RNA are efficiently sequenced by padlock-based in-situ sequencing. Thereby, NIS-seq circumvents all major drawbacks of prior art pooled optical screening by unambiguously assigning bright signal clusters to nuclei independent of cell size, type, or transcriptional activity. In this context, it is of further note that the use of a U6 promoter and a reserve T7 promoter as illustrated by the appended examples provides several advantages as compared to the conceivable option of using instead a U6 / T7 hybrid promoter (i.e., a T7 promoter embedded within a U6 promoter) as described in Kudo et al., December 26, 2023 on bioRxiv; DOI: 10.1101 / 2023.12.26.573143. (1) By the use of a separate U6 promoter and a separate reverse T7 promoter, sense transcripts from the U6 promoter or from an upstream sense-orientation Pol-Il promoter are excluded from producing sequencing signals, thereby enabling a strictly separate control of the functional expression of the perturbagen and generating sequencing signal. No unwanted U6 or T7 transcripts will be generated. (2) Due to the inverted orientation of the reverse T7 promoter, the seed region of the gRNAs is sequenced in the illustrated method, which is the most important region determining editing specificity using most widely used CRISPR enzymes. This ensures better discrimination between functional and mutated non-functional gRNAs as compared to the use of the discussed hybrid promoter. (3) The use of a separate U6 promoter and a separate reserve T7 promoter allows for higher flexibility in designing experiments, because, for example, the wild-type human U6 promoter can be replaced by a wild-type mouse U6 promoter, or an Hl promoter, or any other, potentially inducible, synthetic, stronger, or otherwise improved promoter. Likewise, the reverse T7 can be replaced with a reverse Sp6 or other reverse phage- or synthetic promoter, not requiring engineering as in the case of a hybrid promoter, where each novel design has to be de novo empirically tested and / or optimized. As such, in Kudo et al., several designs of replacing portions of the human U6 promoter by a T7 promoter were tested and most hybrid promoters did not work satisfactory with respect to perturbagen expression by the U6 promoter and / or sequencing signal generation by the T7 promoter. (4) It can be assumed that the inverted design as shown herein (including the specific RT-primer, padlock and sequencing primer sequences) produces higher signal than the "same-orientation" hybrid promoter design. The reasons for this are assumed to be that (i) the wild-type U6 and T7 promoter sequences were not altered in any way, thus maximizing their efficiency compared to an engineered hybrid promoter, (ii) the hybrid U6 promoter is occupied by proteins in cells, which can be immobilized during fixation steps and hinder efficient T7 RNA polymerase binding and transcription, (iii) the reverse T7 product RNA is predicted to be less self- complementary (structured) than the sense RNA by http: / / www.unafold.org. The latter matters because structures in RNA can interfere with downstream primer annealing, RNA diffusion, and enzymatic reaction steps. In contrast, the sense transcript generated in Kudo et al. is structured because it contains the natural gRNA scaffold binding to the CRSPR-Case nuclease (e.g. Cas9), whereas the antisense RNA has no natural function, so that it has not been evolutionarily optimized to form a stable RNA structure and (iv) finally, for some other perturbagens than gRNAs, like pegRNAs used in PRIME editing, the 3' end is often more informative than the 5' end to determine editing outcomes, especially when many similar editing constructs are used. The closer the T7 promoter is to the „informative" sequence, the more efficient the interesting sequence is transcribed and ultimately sequenced.
[0053] In accordance with a preferred embodiment of the first aspect of the invention the genetic pertubator is a guide RNA (gRNA) and step (b) is carried out in the presence of a CRISPR-Cas nuclease in the cells, thereby introducing genetic perturbations into the cells of the cell population by CRISPR gene editing.
[0054] The method according to this preferred embodiment is illustrated by the appended examples. In the examples the genetic pertubators are gRNAs and the gRNAs along with the CRISPR-Cas nuclease introduce the genetic pertubations into the cells in step (b). The CRISPR-Cas nuclease is enzymatically active; i.e. it displays nuclease activity. In accordance with another preferred embodiment of the first aspect of the invention the genetic pertubator is a guide RNA (gRNA) and step (b) is carried out in the presence of a modified CRISPR-Cas nuclease without endonuclease activity and a nucleobase modifying enzyme being linked to the gRNA or the modified CRISPR-Cas nuclease without endonuclease activity in the cells, thereby introducing genetic perturbations into the cells of the cell population by the nucleobase modifying enzyme, wherein the nucleobase modifying enzyme is preferably nucleobase deaminase enzyme.
[0055] Also according to this preferred embodiment the genetic pertubators are gRNAs. However, this time a modified CRISPR-Cas nuclease without endonuclease activity is used. The CRISPR-Cas nuclease without endonuclease activity is a catalytically inactive (dead) enzyme. However, a nucleobase modifying enzyme is linked to the gRNA or the modified CRISPR-Cas nuclease without endonuclease activity that is capable of introducing the genetic perturbations.
[0056] The nucleobase modifying enzyme is preferably nucleobase deaminase enzyme. Nucleobase deaminases are essential enzymes that modify the DNA or RNA sequences across domains of life, e.g., during antibody- and T-cell-receptor gene diversification in human immune cells, or anti-phage immune defence in bacteria; see, for review, Gaded and Anand, RSC Adv. 2018 Jun 27; 8(42): 23567- 23577.
[0057] The above preferred embodiment is also known as base editing. Base editing combines the powerful DNA-scanning and sequence-identification capabilities of the CRISPR-Cas9 system with a nucleobase modifying enzyme such as a deaminase enzyme, which introduces single nucleotide polymorphisms (SNPs) by chemically altering the target DNA sequence without the intentional generation of a DNA double-strand break (DSB). This chemical modification, known as deamination, consists of the removal of an amino group from a nucleotide, which after DNA repair or replication results in the installation of a new base.
[0058] In accordance with a further preferred embodiment of the first aspect of the invention the genetic pertubator is a prime editing guide RNA (pegRNA) and step (b) is carried out in the presence of a modified CRISPR-Cas nuclease without endonuclease activity and a engineered reverse transcriptase enzyme being linked to the modified CRISPR-Cas nuclease without endonuclease activity in the cells, thereby introducing genetic perturbations into the cells of the cell population by the pegRNA and the engineered reverse transcriptase enzyme. This preferred embodiment is also known as prime editing. Prime editing is a 'search-and-replace' genome editing technology by which the genome of cells may be modified. The technology directly writes new genetic information into a targeted DNA site. It uses a fusion protein, consisting of a catalytically impaired CRISPR-Cas nuclease fused to an engineered reverse transcriptase enzyme, and a prime editing guide RNA (pegRNA), capable of identifying the target site and providing the new genetic information to replace the target DNA nucleotides. It can mediate targeted insertions, deletions, and base-to-base conversions without the need for double strand breaks (DSBs) or donor DNA template.
[0059] The prime editing guide RNA (pegRNA) (i) is capable of identifying the target nucleotide sequence to be edited, and (ii) encodes new genetic information that replaces the targeted sequence. The pegRNA comprises or consists of an extended gRNA containing in addition a primer binding site (PBS) and a reverse transcriptase (RT) template sequence. During genome editing, the primer binding site allows the 3' end of the nicked DNA strand to hybridize to the pegRNA, while the RT template serves as a template for the synthesis of edited genetic information.
[0060] For prime editing the modified CRISPR-Cas nuclease without endonuclease activity is catalytically impaired CRISPR-Cas nuclease that introduces a single strand nick and, hence, may be named a nickase.
[0061] The engineered reverse transcriptase can be, for example, a M-MLV reverse transcriptase.
[0062] In accordance with a yet further preferred embodiment of the first aspect of the invention the genetic pertubator is a non-coding RNA (ncRNA) / gRNA hybrid or the combination of a ncRNA and a gRNA, wherein the ncRNA can be reversely transcribed into a reverse transcribed (RT)-DNA with homology to a genomic locus in the cells, and step (b) is carried out in the presence of a reverse transcriptase and a CRISPR-Cas nuclease in the cells, thereby introducing genetic perturbations into the cells of the cell population by CRISPR gene editing that uses the RT-DNA as a homologous recombination template.
[0063] The above set-up of a ncRNA) / gRNA hybrid and CRISPR-Cas nuclease or the combination of a ncRNA and a gRNA (as separately expressed and encoded molecules) is described in Lopez et al., Nature Chemical Biology volume 18, pagesl99-206 (2022). The ncRNA is both the primer and template for the reverse transcriptase. The ncRNA can be sub-divided into a region that is reverse transcribed (msd) and a region that remains RNA in the final molecule (msr), which may be partially overlapping. The msd comprise one or more nucleotide modifications as compared to the target genome, whereby a genetic pertubation is introduced into the cell. In accordance with a preferred embodiment of the first aspect each individual cell comprises one perturbation only or each cell comprises two or more perturbations.
[0064] In case each individual cell comprises one perturbation the effect of each individual perturbation on the phenotype of each individual cell can be determined. In case each cell comprises two or more perturbations the combined effect of two or more individual perturbations on the phenotype of each individual cell can be determined.
[0065] The present invention relates in a second aspect to a method for linking the expression of a gene or RNA of interest in individual cells within a cell population to the phenotype of the individual cells, comprising (a) introducing a plurality of different nucleic acid molecules into the genome of the cells, wherein each of the nucleic acid molecules comprises a nucleotide sequence encoding a gene or RNA of interest, a first promoter and a second promoter, wherein the first promoter is 5' of the nucleotide sequence encoding the gene or RNA of interest and controls the expression of the gene or RNA of interest in 5' -3' direction in the cells, wherein the second promoter is a reverse phage promoter and / or reverse in vitro transcription promoter 3' of the nucleotide sequence encoding the gene or RNA of interest and controls the expression of the gene or RNA of interest in 3'-5' direction in the nucleus of the cells; (b) initiating the expression of the nucleic acid sequences encoding the gene or RNA of interest in 5' -3' direction from the first promoter in the cells, thereby expressing the gene or RNA of interest in the cell; (c) fixing and permeabilizing the cells; (d) initiating the expression of the nucleic acid sequences encoding the gene or RNA of interest in 3' -5' direction from the second promoter in the fixed cells, thereby generating barcode sequences that are presentative for each kind of gene or RNA of interest; (e) optionally fixing the cells, preferably on a surface; (f) reverse transcription of the barcode sequences; (g) optionally amplification of the reversely transcribed barcode sequences; (h) identifying the reversely transcribed barcode sequences in the individual cells within the cell population by in situ sequencing by synthesis thereby assigning a particular gene or RNA of interest to each individual cell; and (i) linking the gene or RNA of interest of each individual cell within the phenotype of each individual cell, wherein the phenotype of each individual cell (I) has been imaged or monitored in the living cells by microscopy any time before step (c) and preferably after step (b), wherein the imaged or monitored phenotype of each individual cell is preferably the cell state, the cell motility, the cell shape, cell-cell interactions, the fluorescence tag of one or more targets in the living cell, the affinity reagent staining of target in the living cell and / or the fluorescence or color of a cell staining dye, or (II) has been imaged or monitored in the fixed cells by microscopy any time after step (c) and preferably any time before step (f), wherein the imaged or monitored phenotype of each individual cell is preferably the shape of the fixed cell, the fluorescence tag of one or more targets in the fixed cell, the affinity reagent staining of a target in the fixed cell, and / or the sequence-specific detection of one or more RNA and / or DNA sequences in the fixed cell.
[0066] The definitions and preferred embodiments of the first aspect of the invention apply mutatis mutandis to the second aspect of the invention as far as being amendable with the second aspect.
[0067] As discussed above, the purpose of the first aspect of the invention is linking the genetic perturbations of individual cells within a cell population to the phenotype of the individual cells and in connection with this method a plurality of nucleic acid molecules encoding different genetic perturbators are introduced into the cells of a cell population and then the expressed genetic perturbators introduce genetic perturbation into the cells.
[0068] The purpose of the second aspect of the invention is related but different. The purpose is linking the expression of a gene or RNA of interest in individual cells within a cell population to the phenotype of the individual cells and in connection with this method a plurality of nucleic acid molecules encoding different genes or RNAs of interest are introduced into the cells of a cell population and then the different genes or RNAs of interest are expressed without introducing any genetic perturbation into the cells.
[0069] In case a gene of interest is expressed in the cell this expression is in the form of an mRNA that is translated into a (poly)peptide in the cell. Examples of (poly)peptides will be described herein below. In case a RNA of interest is expressed this RNA is a non-coding RNA that is not translated into a(poly)peptide (such as, for example, transfer RNAs (tRNAs), ribosomal RNAs (rRNAs), small RNAs such as microRNAs, siRNAs, piRNAs, snoRNAs, snRNAs, exRNAs, scaRNAs and long ncRNAs such as Xist and HOTAIR).
[0070] Other than that difference the steps of the methods of the first and second aspect of the invention are essentially the same. In particular also the method of the second aspect of the invention relies on the use the same two promoters, barcode sequences, in situ sequencing by synthesis and imaged or monitored cell phenotypes as the method of the first aspect.
[0071] In accordance with a preferred embodiment of the second aspect the RNA of interest is a guide RNA (gRNA) being complementary to the promoter or coding sequences of a target gene and step (b) is carried out in the presence of a modified CRISPR-Cas nuclease without endonuclease activity and a transcriptional activator being linked to the gRNA or the modified CRISPR-Cas nuclease without endonuclease activity in the cells, thereby increasing the expression of the gene of interest.
[0072] The set-up of this preferred embodiment is also called CRISPRa (or CRISPR activation). CRISPRa uses modified versions of CRISPR-Cas nucleases without endonuclease activity, with added transcriptional activators on the CRISPR-Cas nucleases without endonuclease activity or the gRNAs.
[0073] Like for CRISPR interference, the CRISPR effector is guided to the target by a complementary guide RNA. However, CRISPR activation systems are fused to transcriptional activators to increase expression of genes of interest.
[0074] In accordance with another preferred embodiment of the second aspect the RNA of interest is a guide RNA (gRNA) being complementary to the promoter or the exonic sequences of a target gene and step (b) is carried out in the presence of a modified CRISPR-Cas nuclease without endonuclease activity in the cells, thereby sterically repressing the transcription of a target gene by blocking either transcriptional initiation or elongation.
[0075] The set-up of this preferred embodiment is also called CRISPR interference (CRISPRi). CRISPRi is a genetic perturbation technique that allows for sequence-specific repression of gene expression in cells. CRISPRi can sterically repress transcription by blocking either transcriptional initiation or elongation. This is accomplished by designing sgRNA complementary to the promoter or the exonic sequences. The level of transcriptional repression with a target within the coding sequence is strand-specific. The technique provides a complementary approach to RNA interference. The difference between CRISPRi and RNAi, though, is that CRISPRi regulates gene expression primarily on the transcriptional level, while RNAi controls genes on the mRNA level.
[0076] In accordance with a preferred embodiment of the second aspect the gene of interest encodes a protein of interest, an aptamer, a scFv-antibody fragment, a nanobody, a computationally designed functional protein or an antibody mimetic, wherein the antibody mimetic is preferably selected from affibodies, adnectins, anticalins, DARPins, avimers, nanofitins, affilins, Kunitz domain peptides, Fynomers®, trispecific binding molecules and probodies.
[0077] Hence, the gene of interest may encode any poly(peptide) being selected from a protein of interest, an aptamer, a scFv-antibody fragment, nanobody, a computationally designed functional protein or an antibody mimetic, wherein the antibody mimetic is preferably selected from affibodies, adnectins, anticalins, DARPins, avimers, nanofitins, affilins, Kunitz domain peptides, Fynomers®, trispecific binding molecules and probodies.
[0078] In this preferred embodiment aptamers are peptide molecules that bind a specific target molecule. Aptamers are usually created by selecting them from a large random sequence pool, but natural aptamers also exist in riboswitches. Aptamers can be used for both basic research and clinical purposes as macromolecular drugs. Aptamers can be combined with ribozymes to self-cleave in the presence of their target molecule. These compound molecules have additional research, industrial and clinical applications (Osborne et. al. (1997), Current Opinion in Chemical Biology, 1:5-9; Stull & Szoka (1995), Pharmaceutical Research, 12, 4:465-483).
[0079] A single-chain variable fragment (scFv) is a fusion protein of the variable regions of the heavy (VH) and light chains (VL) of immunoglobulins, connected with a short linker peptide of ten to about 25 amino acids. The linker is usually rich in glycine for flexibility, as well as serine or threonine for solubility, and can either connect the N-terminus of the VH with the C-terminus of the VL, or vice versa. This protein retains the specificity of the original immunoglobulin (antibody), despite removal of the constant regions and the introduction of the linker.
[0080] A nanobody (also known as single-domain antibody (sdAb)) is an antibody fragment consisting of a single monomeric variable antibody domain. Like a whole antibody, it is able to bind selectively to a specific antigen. With a molecular weight of only 12-15 kDa, single-domain antibodies are much smaller than common antibodies (150-160 kDa) which are composed of two heavy protein chains and two light chains, and even smaller than Fab fragments (~50 kDa, one light chain and half a heavy chain) and single-chain variable fragments (~25 kDa, two variable domains, one from a light and one from a heavy chain). The first single-domain antibodies were engineered from heavy-chain antibodies found in camelids; these are called VHH fragments. Cartilaginous fishes also have heavy-chain antibodies (IgNAR, 'immunoglobulin new antigen receptor'), from which single-domain antibodies called VNAR fragments can be obtained. [2] An alternative approach is to split the dimeric variable domains from common immunoglobulin G (IgG) from humans or mice into monomers.
[0081] The de novo design of computationally designed functional proteins is described, for example, in Watson et al., Nature, volume 620, pages 1089-1100 (2023). De novo protein design seeks to generate proteins with specified structural and / or functional properties, for example, making a binding interaction with a given target, folding into a particular topology or containing a catalytic site. Denoising diffusion probabilistic models (DDPMs), a powerful class of machine learning models recently demonstrated to generate new photorealistic images in response to text prompts, have several properties well suited to de novo protein design.
[0082] The term "affibody", as used herein, refers to a family of antibody mimetics which is derived from the Z-domain of staphylococcal protein A. Structurally, affibody molecules are based on a three-helix bundle domain which can also be incorporated into fusion proteins. In itself, an affibody has a molecular mass of around 6kDa and is stable at high temperatures and under acidic or alkaline conditions. Target specificity is obtained by randomisation of 13 amino acids located in two alphahelices involved in the binding activity of the parent protein domain (Feldwisch J, Tolmachev V.; (2012) Methods Mol Biol. 899:103-26).
[0083] The term "adnectin" (also referred to as "monobody"), as used herein, relates to a molecule based on the 10th extracellular domain of human fibronectin III (10Fn3), which adopts an Ig-like p-sandwich fold of 94 residues with 2 to 3 exposed loops, but lacks the central disulphide bridge (Gebauer and Skerra (2009) Curr Opinion in Chemical Biology 13:245-255). Adnectins with the desired target specificity can be genetically engineered by introducing modifications in specific loops of the protein.
[0084] The term "anticalin", as used herein, refers to an engineered protein derived from a lipocalin (Beste G, Schmidt FS, Stibora T, Skerra A. (1999) Proc Natl Acad Sci U S A. 96(5):1898-903; Gebauer and Skerra (2009) Curr Opinion in Chemical Biology 13:245-255). Anticalins possess an eight-stranded p-barrel which forms a highly conserved core unit among the lipocalins and naturally forms binding sites for ligands by means of four structurally variable loops at the open end. Anticalins, although not homologous to the IgG superfamily, show features that so far have been considered typical for the binding sites of antibodies: (i) high structural plasticity as a consequence of sequence variation and (ii) elevated conformational flexibility, allowing induced fit to targets with differing shape.
[0085] As used herein, the term "DARPin" refers to a designed ankyrin repeat domain (166 residues), which provides a rigid interface arising from typically three repeated p-turns. DARPins usually carry three repeats corresponding to an artificial consensus sequence, wherein six positions per repeat are randomised. Consequently, DARPins lack structural flexibility (Gebauer and Skerra, 2009).
[0086] The term "avimer", as used herein, refers to a class of antibody mimetics which consist of two or more peptide sequences of 30 to 35 amino acids each, which are derived from A-domains of various membrane receptors and which are connected by linker peptides. Binding of target molecules occurs via the A-domain and domains with the desired binding specificity can be selected, for example, by phage display techniques. The binding specificity of the different A-domains contained in an avimer may, but does not have to be, identical (Weidle UH, et al., (2013), Cancer Genomics Proteomics;
[0087] 10(4):155-68).
[0088] A "nanofitin" (also known as affitin) is an antibody mimetic protein that is derived from the DNA binding protein Sac7d of Sulfolobus acidocaldarius. Nanofitins usually have a molecular weight of around 7kDa and are designed to specifically bind a target molecule by randomising the amino acids on the binding surface (Mouratou B, Behar G, Paillard-Laurance L, Colinet S, Pecorari F., (2012) Methods Mol Biol.; 805:315-31).
[0089] The term "affilin", as used herein, refers to antibody mimetics that are developed by using either gamma-B crystalline or ubiquitin as a scaffold and modifying amino-acids on the surface of these proteins by random mutagenesis. Selection of affilins with the desired target specified is effected, for example, by phage display or ribosome display techniques. Depending on the scaffold, affilins have a molecular weight of approximately 10 or 20kDa. As used herein, the term affilin also refers to di- or multimerised forms of affilins (Weidle UH, et al., (2013), Cancer Genomics Proteomics; 10(4):155-68).
[0090] A "Kunitz domain peptide" is derived from the Kunitz domain of a Kunitz-type protease inhibitor such as bovine pancreatic trypsin inhibitor (BPTI), amyloid precursor protein (APP) or tissue factor pathway inhibitor (TFPI). Kunitz domains have a molecular weight of approximately 6kDA and domains with the required target specificity can be selected by display techniques such as phage display (Weidle et al., (2013), Cancer Genomics Proteomics; 10(4):155-68).
[0091] As used herein, the term "Fynomer®" refers to a non-immunoglobulin-derived binding polypeptide derived from the human Fyn SH3 domain. Fyn SH3-derived polypeptides are well-known in the art and have been described e.g. in Grabulovski et al. (2007) JBC, 282, p. 3196-3204, WO 2008 / 022759, Bertschinger et al (2007) Protein Eng Des Sei 20(2):57-68, Gebauer and Skerra (2009) Curr Opinion in Chemical Biology 13:245-255, or Schlatter et al. (2012), MAbs 4:4, 1-12).
[0092] The term "trispecific binding molecule" as used herein refers to a polypeptide molecule that possesses three binding domains and is thus capable of binding, preferably specifically binding to three different epitopes. At least one of these three epitopes is an epitope of the protein of the fourth aspect of the invention. The two other epitopes may also be epitopes of the protein of the fourth aspect of the invention or may be epitopes of one or two different antigens. The trispecific binding molecule is preferably a TriTac. A TriTac is a T-cell engager for solid tumors which comprised of three binding domains being designed to have an extended serum half-life and be about one-third the size of a monoclonal antibody.
[0093] As used herein, the term "probody" refers to a protease-activatable antibody prodrug. A probody consists of an authentic IgG heavy chain and a modified light chain. A masking peptide is fused to the light chain through a peptide linker that is cleavable by tumor-specific proteases. The masking peptide prevents the probody binding to healthy tissues, thereby minimizing toxic side effects.
[0094] In accordance with a further preferred embodiment of the second aspect the RNA of interest is an aptamer, a siRNA, a shRNA, a miRNA, a ribozyme or an antisense nucleic acid molecule.
[0095] In this preferred embodiment the aptamers are nucleic acid molecules or peptide molecules that bind a specific target molecule.
[0096] In accordance with the present invention, the term "small interfering RNA (siRNA)", also known as short interfering RNA or silencing RNA, refers to a class of 18 to 30, preferably 19 to 25, most preferred 21 to 23 or even more preferably 21 nucleotide-long double-stranded RNA molecules that play a variety of roles in biology. Most notably, siRNA is involved in the RNA interference (RNAi) pathway where the siRNA interferes with the expression of a specific gene. In addition to their role in the RNAi pathway, siRNAs also act in RNAi-related pathways, e.g. as an antiviral mechanism or in shaping the chromatin structure of a genome. siRNAs naturally found in nature have a well-defined structure: a short double-strand of RNA (dsRNA) with 2-nt 3' overhangs on either end. Each strand has a 5' phosphate group and a 3' hydroxyl (-OH) group. This structure is the result of processing by dicer, an enzyme that converts either long dsRNAs or small hairpin RNAs into siRNAs. siRNAs can also be exogenously (artificially) introduced into cells to bring about the specific knockdown of a gene of interest. Essentially any gene for which the sequence is known can thus be targeted based on sequence complementarity with an appropriately tailored siRNA. The double-stranded RNA molecule or a metabolic processing product thereof is capable of mediating target-specific nucleic acid modifications, particularly RNA interference and / or DNA methylation. Exogenously introduced siRNAs may be devoid of overhangs at their 3' and 5' ends, however, it is preferred that at least one RNA strand has a 5'- and / or 3'-overhang. Preferably, one end of the double-strand has a 3'-overhang from 1 to 5 nucleotides, more preferably from 1 to 3 nucleotides and most preferably 2 nucleotides. The other end may be blunt-ended or has up to 6 nucleotides 3'-overhang. In general, any RNA molecule suitable to act as siRNA is envisioned in the present invention. The most efficient silencing was so far obtained with siRNA duplexes composed of 21-nt sense and 21-nt antisense strands, paired in a manner to have a 2-nt 3'- overhang. The sequence of the 2-nt 3' overhang makes a small contribution to the specificity of target recognition restricted to the unpaired nucleotide adjacent to the first base pair. 2'-deoxynucleotides in the 3' overhangs are as efficient as ribonucleotides, but are often cheaper to synthesize and probably more nuclease resistant. Delivery of siRNA may be accomplished using any of the methods known in the art, for example by combining the siRNA with saline and administering the combination intravenously or intranasally or by formulating siRNA in glucose (such as for example 5% glucose) or cationic lipids and polymers can be used for siRNA delivery in vivo through systemic routes either intravenously (IV) or intraperitoneally (IP) (Fougerolles et al. (2008), Current Opinion in Pharmacology, 8:280-285; Lu et al. (2008), Methods in Molecular Biology, vol. 437: Drug Delivery Systems - Chapter 3: Delivering Small Interfering RNA for Novel Therapeutics).
[0097] A short hairpin RNA (shRNA) is a sequence of RNA that makes a tight hairpin turn that can be used to silence gene expression via RNA interference. shRNA uses a vector introduced into cells and utilizes the U6 promoter to ensure that the shRNA is always expressed. This vector is usually passed on to daughter cells, allowing the gene silencing to be inherited. The shRNA hairpin structure is cleaved by the cellular machinery into siRNA, which is then bound to the RNA-induced silencing complex (RISC). This complex binds to and cleaves mRNAs which match the siRNA that is bound to it. si / shRNAs to be used in the present invention are preferably chemically synthesized using appropriately protected ribonucleoside phosphoramidites and a conventional DNA / RNA synthesizer. Suppliers of RNA synthesis reagents are Proligo (Hamburg, Germany), Dharmacon Research (Lafayette, CO, USA), Pierce Chemical (part of Perbio Science, Rockford, IL, USA), Glen Research (Sterling, VA, USA), ChemGenes (Ashland, MA, USA), and Cruachem (Glasgow, UK).
[0098] Further molecules effecting RNAi include, for example, microRNAs (miRNA). Said RNA species are single-stranded RNA molecules. Endogenously present miRNA molecules regulate gene expression by binding to a complementary mRNA transcript and triggering of the degradation of said mRNA transcript through a process similar to RNA interference.
[0099] A ribozyme (from ribonucleic acid enzyme, also called RNA enzyme or catalytic RNA) is an RNA molecule that catalyses a chemical reaction. Many natural ribozymes catalyse either their own cleavage or the cleavage of other RNAs, but they have also been found to catalyse the aminotransferase activity of the ribosome. Non-limiting examples of well-characterised small selfcleaving RNAs are the hammerhead, hairpin, hepatitis delta virus, and in v / tro-selected lead-dependent ribozymes, whereas the group I intron is an example for larger ribozymes. The principle of catalytic self-cleavage has become well established in recent years. The hammerhead ribozymes are characterised best among the RNA molecules with ribozyme activity. Since it was shown that hammerhead structures can be integrated into heterologous RNA sequences and that ribozyme activity can thereby be transferred to these molecules, it appears that catalytic antisense sequences for almost any target sequence can be created, provided the target sequence contains a potential matching cleavage site. The basic principle of constructing hammerhead ribozymes is as follows: A region of interest of the RNA, which contains the GUC (or CUC) triplet, is selected. Two oligonucleotide strands, each usually with 6 to 8 nucleotides, are taken and the catalytic hammerhead sequence is inserted between them. The best results are usually obtained with short ribozymes and target sequences.
[0100] The term "antisense nucleic acid molecule", as used herein, refers to a nucleic acid which is complementary to a target nucleic acid. An antisense molecule in accordance with the invention is capable of interacting with the target nucleic acid, more specifically it is capable of hybridizing with the target nucleic acid. Due to the formation of the hybrid, transcription of the target gene(s) and / or translation of the target mRNA is reduced or blocked. Standard methods relating to antisense technology have been described (see, e.g., Melani et al., Cancer Res. (1991) 51:2897-2901).
[0101] In accordance with a preferred embodiment of the first and second aspect the cells comprise or are immune cells, wherein the immune cells preferably comprise or are macrophages.
[0102] It is shown in the examples that the in-situ sequencing by synthesis protocol as provided herein (NIS- Seq) is compatible with immune cells, in particular macrophages, whereas previously published in- situ sequencing protocols failed in these cell types.
[0103] In accordance with another preferred embodiment of the first and second aspect the population of cells comprises at least 1x10scells, preferably at least lxlO7cells and most preferably at least 1x10scells.
[0104] Hence, the methods of the invention can be applied to populations of millions of cells.
[0105] In accordance with a further preferred embodiment of the first and second aspect the first promoter is a polymerase-lll promoter, preferably a U6 promoter (more preferably a human U6 promoter) and / or the second promoter is a reverse phage promoter, preferably a reverse T7 promoter. In the appended examples a U6 promoter and a reverse T7 promoter were used as first and second promoters. As alternatives for the U6 promoter and the reverse T7 promoter the promoters Hl or mouse U6 and the reverse SP6 promoter are preferred.
[0106] The U6 promoter is a polymerase-lll promoter. In eukaryote cells, RNA polymerase III (also called Pol III) is a protein that transcribes DNA to synthesize 5S ribosomal RNA, tRNA, and other small RNAs. The T7 promoter is a phage promoter.
[0107] The U6 promoter and reverse T7 promoter preferably have a sequence as shown in SEQ. ID NO: 1 herein above or a sequence being with increasing preference at least 80%, at least 85%, at least 90%, at least 95%, and at least 97.5% identical thereto.
[0108] Nucleotide and amino acid sequence analysis and alignment in connection with the present invention are preferably carried out using the NCBI BLAST algorithm (Stephen F. Altschul, Thomas L. Madden, Alejandro A. Schaffer, Jinghui Zhang, Zheng Zhang, Webb Miller, and David J. Lipman (1997), Nucleic Acids Res. 25:3389-3402). BLAST can be used for nucleotide sequences (nucleotide BLAST) and amino acid sequences (protein BLAST). The skilled person is aware of additional suitable programs to align nucleic acid sequences.
[0109] In accordance with a preferred embodiment of the first and second aspect the nucleic acid molecules are (i) genomic integrating vectors, preferably lentiviral or retroviral vectors and most preferably lentiviral CRISPR droplet sequencing (CROP-seq) vectors; or (ii) PiggyBac transposon plasmids or sleeping beauty transposon plasmids.
[0110] As discussed above, in accordance with the the first and second aspect, a plurality of nucleic acid molecules has to be integrated into the genome of cells.
[0111] In order to achieve this genomic integrating vectors can be used. Genomic integrating vectors can insert their genetic material into the host cell's genome. Preferred examples are lentiviral or retroviral vectors and most preferably lentiviral CRISPR droplet sequencing (CROP-seq) vectors.
[0112] Lentiviral vectors are derived from lentiviruses and can act as a vector to insert genes into cells. Unlike other retroviruses, which cannot penetrate the nuclear envelope and can therefore only act on cells while they are undergoing mitosis, lentiviruses can infect cells whether or not they are dividing (shown to be largely due to the capsid protein). Retroviral vectors stably integrate into the dividing target cell genome so that the introduced gene is passed on and expressed in all daughter cells.
[0113] A CROP-seq (CRISPR droplet sequencing) vector has been used in the examples and is particular useful, because it duplicates the gRNA expression cassette during genome integration. The CROP-seq vector is based on a lentiviral vector. In more detail, the CROP-seq vector is a modified version of lentiGuide- Puro, where the U6-promoter-sgRNA expression cassette has been moved into the 3'-LTR region, just downstream of the puromycin gene (Datlinger et al., Nat Methods 14: 297-301 (2017)).
[0114] Alternatively genomic integration can be achieved, for example, by recombinase-based integration (e.g. Flip-IN cells), random integration, (Flip-IN cells or CRISPR- / TALEN- / ZFN- / meganuclease-mediated integration that can use NHEJ, HR, or PRIME. The use of TALEN- / ZFN- / meganuclease-mediated integration is described, for example, in Silva et al., Curr Gene Ther. 2011;ll(l):ll-27; Miller et al., Nature biotechnology. 2011;29(2):143-148, and Klug, Annual review of biochemistry. 2010; 79:213- 231.
[0115] In accordance with a preferred embodiment of the first and second aspect the amplification of the reversely transcribed barcode sequences comprises (i) padlock elongation, ligation and rolling circle amplification; (ii) an isothermal amplification, preferably selected from Loop-Mediated Isothermal Amplification (LAMP), Whole Genome Amplification (WGA), Rolling Circle Amplification (RCA), Strand Displacement Amplification (SDA), Helicase-Dependent Amplification (HDA), Recombinase Polymerase Amplification (RPA), Transcription Mediated Amplification (TMA), Sequence Mediated Amplification of RNA Technology (SMART), Multiple Cross Displacement Amplification (MCDA) and Nucleic Acid Sequences Based Amplification (NASBA); or (iii) linear polymerization, preferably via a Phi29 polymerase or a strand displacing polymerase.
[0116] Within these options (i) to (iii), option (i) is preferred since it is used in the appended examples; see Figure 1A.
[0117] According to option (i) padlock elongation, ligation and rolling circle amplification are used. Padlock elongation uses padlock probes that are single stranded DNA molecules. When the target complementary regions are hybridized to the DNA target, the padlock probes can become circular by polymerase-mediated elongation and ligase-mediated ligation. Rolling circle replication (RCR) is a process of unidirectional nucleic acid replication that can rapidly synthesize multiple copies of circular molecules of DNA or RNA, such as plasmids, the genomes of bacteriophages, and the circular RNA genome of viroids. Some eukaryotic viruses also replicate their DNA or RNA via the rolling circle mechanism.
[0118] According to option (ii) isothermal amplification is used, preferably selected from Loop-Mediated Isothermal Amplification (LAMP), Whole Genome Amplification (WGA), Rolling Circle Amplification (RCA), Strand Displacement Amplification (SDA), Helicase-Dependent Amplification (HDA), Recombinase Polymerase Amplification (RPA), Transcription Mediated Amplification (TMA), Sequence Mediated Amplification of RNA Technology (SMART), Multiple Cross Displacement Amplification (MCDA) and Nucleic Acid Sequences Based Amplification (NASBA); see Zhao et al. (2015), Chem. Rev. 2015, 115, 22, 12491-12545 and Obande and Singh (2020), Infection and Drug Resistance, 13:155-483.
[0119] According to option (iii) linear polymerization is used, preferably via a Phi29 polymerase or a strand displacing polymerase is used. phi29 DNA Polymerase is the replicative polymerase from the Bacillus subtilis phage phi29 (029). This polymerase has exceptional strand displacement and processive synthesis properties. The polymerase has an inherent 3'->5' proofreading exonuclease activity.
[0120] In connection with the amplification step it is also possible to generate amplification products that can be locally immobilized, for example, through the attachment of nuclear proteins or other anchors. In order to immobilize anchors in the nucleus the genome might be tagmented with a Tn5 transposase or nuclear epitopes could be bound by an oligo-coupled affinity reagent, or chemically reactive oligonucleotides could be reacted with the cells, so that the nucleus harbours a plurality of adapter- antenna-oligonucleotides. These oligonucleotides can than capture the amplification products or without amplification the expression products for the second promoter.
[0121] In accordance with a preferred embodiment of the first and second aspect the nucleic acid molecule additionally comprises a terminator for the transcript of the first promoter, which terminator is preferably located 3' of the nucleic acid sequence encoding a genetic pertubator or a gene or RNA of interest and 5' of the second promoter.
[0122] The presence of a polymerase terminator is optional but preferred since it reduces the length of the expressed genetic pertubator (first aspect) or a gene or RNA of interest (second aspect) to the required length. The polymerase terminator is preferably a polymerase III terminator. The polymerase III terminator preferably has a sequence as shown in SEQ. ID NO: 1 herein above or a sequence being with increasing preference at least 80%, at least 85%, at least 90%, at least 95%, and at least 97.5% identical thereto.
[0123] In accordance with a preferred embodiment of the first and second aspect the cells after step (a) and before step (c) are randomly mutated, preferably by UV radiation or mutagenic chemicals.
[0124] Such random mutations may be used to further diversify the cells in addition to the introduction of the genetic pertubator (first aspect) or the expression of a gene or RNA of interest (second aspect) in the cells. Random mutations can be introduced into cells, for example, by UV radiation or mutagenic chemicals.
[0125] In accordance with a preferred embodiment of the first and second aspect the plurality of different nucleic acid molecules comprises at least 1000, more preferably at least 10000, and most preferably at least 50000 different nucleotide sequences encoding a genetic pertubator or a gene or RNA of interest.
[0126] Hence, the methods of the invention can investigate or screen thousands of genetic pertubators or genes or RNAs of interest in parallel.
[0127] In accordance with a preferred embodiment of the first and second aspect the cells are fixed by methanol acetic acid, glutaraldehyde or paraformaldehyde (PFA), wherein in the case of PFA the fixation is preferably followed by proteinase digestion step, preferably with proteinase K.
[0128] The means and methods for the fixation and permeabilization in step (c) and the fixation in step (e) are not particularly limited. In particular, available expansion-microscopy-fixation steps are expected to also work for the methods as described herein.
[0129] In the appended examples cells were fixed and permeabilized in step (c) with methanol acetic acid (preferably 3:1 mixture) and in step (e) were fixed PFA (preferably 4% PFA, e.g. in phosphate buffered saline (PBS)). Accordingly, the options are most preferred for step (c) and / or (e).
[0130] Methanol acetic acid is particularly suitable to fix and permeabilize at the same time. PFA fixation is particularly suitable to crosslink and thereby prevent the loss of small DNA fragments. The PFA fixation may be followed by a post-fixation step, preferably in 3% paraformaldehyde and 0.1% glutaraldehyde.
[0131] The PFA fixation may also be followed by a decrosslinking step. The decrosslinking is preferably a proteinase digestion step, preferably with proteinase K. Proteinase K (EC 3.4.21.64) is a broadspectrum serine protease. Alternatively, the decrosslinking step can be induced by heat (e.g. 65°C for 4h in 0.1M sodium bicarbonate and 0.3M NaCI in water), noting that heat decrosslinking is widely used in the field of Chlp-Seq.
[0132] In accordance with a preferred embodiment of the first and second aspect the phenotype of each individual cell has been imaged or monitored in the living cells by microscopy over time and preferably assisted by a cell assignment algorithm which is preferably a cross-correlation-based search algorithm.
[0133] The use of a cell assignment algorithm, in particular a cross-correlation-based search algorithm using random patterns of cell or nuclear locations as unique fingerprints is illustrated in the appended examples and is based on the Cellpose algorithm (Stringer et al., Nature Methods volume 18, pagesl00-106 (2021)). At sufficient cell density, this algorithm robustly matches cell neighbourhoods across magnifications, microscopes, and time-points
[0134] The cross-correlation-based search algorithm is also suitable for other applications, such as base editing, prime editing, homologous recombination, CRISPRi, CRISPRa, cDNA overexpression, nanobody, or aptamer.
[0135] In accordance with a preferred embodiment of the first and second aspect the in situ sequencing by synthesis uses a two- or three-color sequencing chemistry.
[0136] In the appended examples three-color sequencing chemistry is used (channels 477 nm, 546 nm and 638 nm) which advantageously leaves a fourth channel (405 nm) for a nuclear counterstain by DAPL In principle also a two-color sequencing chemistry is sufficient in order to distinguish four nucleotides by the two colors, the mixed color and no color.
[0137] The present invention relates in a third aspect to a nucleic acid molecule comprising in 5' -3' direction a U6 promoter, a gRNA and a T7 promoter reverse complement. The definitions and preferred embodiments of the first and second aspect of the invention apply mutatis mutandis to the third aspect of the invention as far being amendable with the third aspect.
[0138] The U6 promoter and the T7 promoter reverse complement (also referred to herein as reverse T7 promoter) preferably have a sequence as shown in SEQ ID NO: 1 herein above or a sequence being with increasing preference at least 80%, at least 85%, at least 90%, at least 95%, and at least 97.5% identical thereto.
[0139] Similarly, the gRNA preferably comprises a sgRNA constant scaffold having a sequence as shown in SEQ ID NO: 1 herein above or a sequence being with increasing preference at least 80%, at least 85%, at least 90%, at least 95%, and at least 97.5% identical thereto. The gRNA also preferably comprises a sgRNA target-specific sequence having a length of between 15 and 265 nucleotides.
[0140] It is to be understood that the nucleic acid molecule is a DNA molecule, so that strictly speaking the construct does not comprise a gRNA but a nucleic acid sequence encoding a gRNA. Said gRNA can be expressed from the U6 promoter.
[0141] As alternatives for the U6 promoter and the reverse T7 promoter the promoters Hl or mouse U6 and the reverse SP6 promoter are preferred.
[0142] Also described herein is a nucleic acid molecule comprising in 5' -3' direction a U6 promoter, a nucleic acid sequence encoding a genetic pertubator or a gene or RNA of interest and a T7 promoter reverse complement.
[0143] Suitable genetic pertubators (including gRNAs) or genes or RNAs of interest are described herein above in connection with the first and second aspect.
[0144] Again, as alternatives for the U6 promoter and the reverse T7 promoter the promoters Hl or mouse U6 and the reverse SP6 promoter are preferred.
[0145] In accordance with a preferred embodiment of the third aspect the nucleic acid molecule additionally comprises a polymerase III terminator, which terminator is preferably located 3' of the gRNA and 5' of the second promoter. As discussed, the presence of a polymerase III terminator is optional but preferred since it reduces the length of the expressed gRNA to the required length.
[0146] The polymerase III terminator preferably has a sequence as shown in SEQ. ID NO: 1 herein above or a sequence being with increasing preference at least 80%, at least 85%, at least 90%, at least 95%, and at least 97.5% identical thereto.
[0147] In accordance with another preferred embodiment of the third aspect the nucleic acid molecule is a vector, preferably a lentiviral or retroviral vector, more preferably a duplicate-integrating lentiviral vector and most preferably a lentiviral CRISPR droplet sequencing (CROP-seq) vector.
[0148] The above vectors have been described herein above in connection with the first and second aspect.
[0149] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. In case of conflict, the patent specification including definitions, will prevail.
[0150] Regarding the embodiments characterized in this specification, in particular in the claims, it is intended that each embodiment mentioned in a dependent claim is combined with each embodiment of each claim (independent or dependent) said dependent claim depends from. For example, in case of an independent claim 1 reciting 3 alternatives A, B and C, a dependent claim 2 reciting 3 alternatives D, E and F and a claim 3 depending from claims 1 and 2 and reciting 3 alternatives G, H and I, it is to be understood that the specification unambiguously discloses embodiments corresponding to combinations A, D, G; A, D, H; A, D, I; A, E, G; A, E, H; A, E, I; A, F, G; A, F, H; A, F, I; B, D, G; B, D, H; B, D, I; B, E, G; B, E, H; B, E, I; B, F, G; B, F, H; B, F, I; C, D, G; C, D, H; C, D, I; C, E, G; C, E, H; C, E, I; C, F, G; C, F, H; C, F, I, unless specifically mentioned otherwise.
[0151] Similarly, and also in those cases where independent and / or dependent claims do not recite alternatives, it is understood that if dependent claims refer back to a plurality of preceding claims, any combination of subject-matter covered thereby is considered to be explicitly disclosed. For example, in case of an independent claim 1, a dependent claim 2 referring back to claim 1, and a dependent claim 3 referring back to both claims 2 and 1, it follows that the combination of the subject-matter of claims 3 and 1 is clearly and unambiguously disclosed as is the combination of the subject-matter of claims 3, 2 and 1. In case a further dependent claim 4 is present which refers to any one of claims 1 to 3, it follows that the combination of the subject-matter of claims 4 and 1, of claims 4, 2 and 1, of claims
[0152] 4, 3 and 1, as well as of claims 4, 3, 2 and 1 is clearly and unambiguously disclosed.
[0153] This also holds true for alternatives in different claims that depend from each other. Thus, if claim 1 recites three alternatives of the same category and claim 2 recites three alternatives of a different category as recited in claim 1, and refers back to claim 1, all combinations of the alternatives as recited in claims 1 and 2 are explicitly disclosed herein.
[0154] The above considerations apply mutatis mutandis to all appended claims.
[0155] The figures show.
[0156] Figure 1 | NIS-Seq enables optical barcode identification in any nucleated cell type. (A) Outline of NIS-Seq reaction steps in comparison to previously established in-situ sequencing of barcoded mRNA10. (B) NIS-Seq imaging results in comparison to cytosolic in-situ sequencing results obtained across three cell types. Scale bar, 50 pm. (C) Nuclei assignment between live cell imaging and NIS-Seq data across different objectives and timepoints. Nuclear staining images are mapped by two-dimensional FFT-accelerated high-pass filtered cross-correlation (top right panel). Overlay of nuclear signal reveals slight dislocation of nuclei between imaging timepoints (bottom left). Centers of gravity of CellPose- defined nuclei are assigned to nearest neighbors and ambiguous assignments are removed (bottom center and right). Assigned nuclei are color-coded with the same random color. Scale bar, 50 pm. (D) Raw images of 14 cycles of NIS-Seq barcode sequencing. Nuclear staining was performed at cycles 1, 4, 7, 10, and 13. Scale bar, 10 pm. (E) Quantitative spot intensities obtained from nuclei I and II highlighted in (D). Indicated on top is the base calling result, matching two members of the pooled lentiviral library used. (F) Fraction of nuclei mapping to known library member sequences. Unambiguous mapping is defined as more than two thirds of aggregated library-matching spot intensities within a nucleus mapping to a single library member. (G) Library coverage in transduced THP1 macrophages, measured by PCR-based NGS and NIS-Seq. Each library member covered in PCR- based sequencing is represented by one dot; dropping out sgRNAs, e.g., those targeting essential genes, are not shown. Jitter is added to values to visualize spot density even at low integer values. (H) Genome editing efficiencies of widely-used sgRNA-expressing lentivirus designs compared to constructs with an additional T7 promoter inserted for NIS-Seq in reverse-orientation. Genome editing efficiencies at three independent loci were assessed by NGS19in HeLa-Cas9 cells transduced with indicated lentiviral constructs after four days of Puromycin selection and Cas9 induction. Figure 2 | Genome-scale optical perturbation screening for mediators of NF-kB activation (A) Results of two replicates of genome-scale NIS-Seq perturbation screening in HeLa-Cas9-p65-mNeonGreen cells stimulated with IL-lb. Dots correspond to genes and axis positions correspond to mean pixel-wise Pearson correlation between mNeonGreen and nuclear staining signals. (B) Collages of cellular images from (A) mapped to perturbed genes indicated. Shown is the mNeonGreen signal. (C) Arrayed hit validation in HeLa-Cas9-p65-mNeonGreen cells using alternative sgRNA sequences from the Toronto KO library v3. Top panel, exemplary mNeonGreen images of IL-lb-stimulated cells. Bottom panel, distribution of activation states, quantified by pixel-wise Pearson correlation between mNeonGreen and nuclear staining signals. Scale bar, 50 pm. (D) Results of two replicates of genome-scale NIS-Seq perturbation screening in HeLa-Cas9-p65-mNeonGreen cells stimulated with TNF-a. Dots correspond to genes, and axis positions indicate mean pixel-wise Pearson correlation between mNeonGreen and nuclear staining signals. (E) Collages of cellular images from (D) mapped to perturbed genes indicated. Shown is the mNeonGreen signal. (F) Arrayed hit validation in HeLa-Cas9-p65-mNeonGreen cells using alternative sgRNA sequences from the Toronto KO library v3. Top panel, exemplary mNeonGreen images of TNF-a-stimulated cells. Bottom panel, distribution of activation states, quantified by pixelwise Pearson correlation between mNeonGreen and nuclear staining signals. Scale bar, 50 pm.
[0157] Figure 3 | Genome-scale optical perturbation screening in THPl-derived macrophages for inflammasome activation (A) Results of two replicates of genome-scale NIS-Seq perturbation screening in THPl-Cas9-ASC-GFP-CASPl / 8DKOcells stimulated with Nigericin. Dots correspond to genes and axis positions correspond to mean ratios of high-pass filtered GFP signal relative to the overall GFP signal per cell. (B) Collages of cellular images from (A) mapped to perturbed genes indicated. Shown are membrane stain (red) and ASC-GFP (green) signals. (C) Arrayed hit validation in THPl-Cas9-ASC- GFP cells using alternative sgRNA sequences from the Toronto KO library v3. Shown are fractions of cells with an ASC speck in three replicate wells. (D) Cytokine secretion in response to two inflammasome triggers in wild type or clonal NLRP3-deficient THP1-ASC-GFP cells, measured by IL-lb ELISA. Shown are three replicate wells of stimulated cells.
[0158] Figure 4 | NIS-Seq in further cell lines A panel of six indicated mouse and human cell lines was transduced with a lentiviral gRNA library containing an inverted T7 promoter downstream of the U6- gRNA-terminator cassette. Positively transduced cells were selected using Puromycin. Raw NIS-Seq imaging results from the first sequencing cycle are shown. Sequencing channels are overlayed in red, green, and blue; Hoechst 33342 nuclear staining is shown in grey. Scale bar, 50 pm. The examples illustrate the invention.
[0159] Example 1 - NIS-Seq enables cell type-agnostic optical barcode identification
[0160] The herein developed nuclear in-situ sequencing (NIS-Seq) enables cell-type-agnostic high-density optical CRISPR screening. While published in-situ sequencing mainly detects barcodes from RNA Polymerase Il-expressed mRNAs in the cytosol, NIS-Seq uses T7 in vitro transcription to generate multiple RNA copies in the nucleus (Figure 1A). After subsequent reverse transcription, padlock elongation, ligation, and rolling circle amplification, sgRNAs are identified by 14 cycles of sequencing- by-synthesis using three excitation wavelengths (Figure 1A). Nuclear signals enable unambiguous assignment to cells even at high cell densities as required for genome-scale screening. Not relying on transcriptional activity or cytosolic volume, NIS-Seq was found compatible with THPl-derived and primary human macrophages, whereas the previously published in-situ sequencing protocol failed in these cell types (Figure IB). For cell-accurate mapping of live cell phenotyping data to NIS-Seq data using different objectives and high cell densities, a cross correlation-based search algorithm was optimized that can buffer small movements and distortions of cells and glass plates during live-cell imaging, assigning pairs of nuclei between imaging modalities with high confidence (Figure 1C). Next a genome-scale perturbation library in THPl-derived macrophages was generated. Sequencing this library using NIS-Seq revealed clean nuclear signal intensities over the course of 14 cycles (Figure ID), which can be translated to known sequences of library members (Figure IE). Overall, more than 60% of nuclei were unambiguously assigned to single library members with >66% of aggregated spot intensities (Figure IF, see Methods section for details), while <20% of nuclei had no spots mapping to the library. Overall, library representation was highly correlated between NGS- and NIS-Seq based sequencing (Fig. 1G). Finally, to ensure high genome efficiencies using a NIS-Seq-compatible lentiviral CRISPR vector, different vector designs across three genomic target sites in HeLa-Cas9 cells were compared. Genome editing efficiencies were determined by NGS after Puromycin selection and Cas9 induction19, confirming that inserting an inverted T7 promoter after the Polymerase-Ill terminator of a sgRNA expression cassette mirrors efficiencies of state-of-the-art lentiviral CRISPR vectors (Figure 1H).
[0161] Example 2 - Genome-scale screening in HeLa cells reveals known components of two immune receptor signaling pathways
[0162] To demonstrate genome-scale screening applications of NIS-Seq, HeLa cells were targeted to investigate genes involved in the activation of the Nuclear Factor (NF)-kB, a family of five members of inducible transcription factors functioning as homo- or hetero-dimers20. In an inactive state, NF-kB dimers are retained in the cytosol by inhibitory proteins such as IkBa. Upon activation of the pathway by external stimuli, phosphorylation of the Ikk complex leads to ubiquitination and proteasomal degradation of IkBa, allowing dimeric NF-kB to shuttle into the nucleus20. To study genes involved in this process, a previously established translocation assay based on HeLa cells carrying a fluorescent p65-mNeonGreen reporter was used10. To elucidate pathway members, two replicate screens using 76,441 sgRNAs targeting human protein coding genes as well as 1,000 non-targeting control guides were performed. HeLa cells were stimulated with interleukin lb (IL-lb) or tumor necrosis factor alpha (TNF-a) to target two different pathways of NF-kB activation. After live cell phenotyping, the sgRNA identity present in each cell was determined by NIS-Seq. Mapping of phenotype and NIS-Seq nuclei resulted in >14,000 genes covered by at least 15 cells in each of two replicate screens for each stimulus (Figure 2A, D). Nuclear translocation of p65 was quantified as pixel-wise Pearson correlation coefficient between p65-mNeonGreen and nuclear staining images. Genes with altered mean nuclear translocation across targeted cells corresponded to the receptors IL1R1 and TNFRSF1A as well as multiple known downstream pathway members such as TRAF6, CHUK, and IKBKG, which were confirmed by single-cell collages of targeted cells retrieved from pooled imaging data (Figure 2B, E). Strong outlier genes were further validated using orthogonal sgRNAs from the Toronto KO CRISPR library v321, confirming the screening hits in all cases (Fig. 2C, F).
[0163] Example 3 - Genome-scale optical screening for mediators of inflammasome activation in human monocytes
[0164] NLRP3 inflammasomes are megadalton protein complexes that assemble in response to danger- and pathogen-associated molecular patterns in macrophages, leading to rapid IL-ip and IL-18 cytokine release and a rapid form of programmed cell death termed pyroptosis. Even though this pathway is critically involved in a multitude of age-related diseases, no genetic screening has been performed in human cells to systematically identify genetic components involved in NLRP3 inflammasome assembly. Using NIS-Seq, a genome-scale CRISPR KO library of THP1 monocyte-derived macrophages stimulated with the ionophore Nigericin was screened. The cells used were deficient in Caspase 1 and 8 in order to avoid pyroptotic cell death downstream of early steps of inflammasome assembly. An ASC-GFP reporter enabled monitoring inflammasome assembly in live cells. Inflammasome activation was quantified in individual cells by acquiring Z-stacks, and calculating the ratio of overall GFP intensity to high-pass filtered GFP intensity, the latter originating from smaller objects like ASC specks. Correlating the results of two independent screens in live macrophages, NLRP3 and IKBKB were confirmed as the most critical pathway members22, while ablation of MAP3K7, TRAF6, HSP90B1, or RNF31 resulted in intermediate levels of pathway dysfunction (Fig. 3A). Aggregated images of live cells assigned to hit genes confirmed a reduction in ASC specking as compared to control cells (Fig. 3B). While the relevance of the HSP90 chaperone, which is not a known NF-kB pathway member, for inflammasome activation had been observed before23 24, a specific involvement of the beta isoform has not been described, and was not observed in our previous screening in murine macrophages25. Independent validation of hit genes by lentiviral expression of orthogonal guide RNAs from the Toronto Knock Out (TKO) library confirmed reduced ASC specking in response to Nigericin (Fig. 3C), which was reflected in cytokine secretion levels in a clonal NLRP3 knock-out cell line (Fig. 3D).
[0165] Example 4 - Discussion
[0166] While efficient editing of single genomic loci for functional genomics studies was possible with predecessor technologies of CRISPR, genome-scale perturbation screening has been revolutionized by the programmability of CRISPR nucleases through short lentiviral sgRNA expression cassettes. Perturbation screening enables to systematically discover genetic determinants of biological processes4 5 6 7 8. NIS-Seq significantly extends the applications of perturbation screening by enabling optical phenotyping in living cells of any nucleated type. It is compatible with high cell density, high library complexity and highly dynamic phenotypes. Thus, biological processes can be mapped to involved genes quantitatively and kinetically, not only identifying critical pathway members, but also reading out genetic rheostats or pathway intersections. NIS-Seq is expected to be compatible with any phenotype observable by microscopy, including sub-cellular transport, cell migration, dynamic oscillation, protein complex formation, RNA splicing, or single-molecule RNA localization2S, optionally using newly developed super-resolution27or non-optical microscopy techniques28. Ongoing improvements in image acquisition speed will be critical to observe sufficient numbers of cells and to enable reliable assignment of nuclei between phenotyping and sequencing images. To compensate for long imaging times, an imaging sequence could be developed acquiring low-magnification overview images in regular intervals between high-resolution imaging, keeping track of cellular movements. NIS- Seq will not only enable loss-of-function screening, but also enable CRISPR-activation or cDNA overexpression profiling, which could help to identify optimized iPS-cell differentiation protocols based on optical cell type identification29.
[0167] NIS-Seq requires five days of hands-on time for performing a genome-scale screen. Many manual steps currently rely on specific pipetting techniques and are very repetitive; to enable future scale-up, faster cycling times, and to minimize reagent use, all repetitive steps were successfully automated using an off-the-shelf pipetting robot. To reduce acquisition times of NIS-Seq sequencing cycles, base calling with laser-only switching using a multiband emission filter was successfully tested. In the future, antibody-based base detection might further increase signal intensity, reduce amplification steps, or reduce phasing artifacts30.
[0168] All NIS-Seq image analysis can performed using our open-source analysis website (jsb- lab.bio / opticalscreening), which does not require local software installation, server hardware or coding experience to analyze genome-scale optical perturbation screens.
[0169] Example 5 - Methods
[0170] Library cloning:
[0171] Library cloning was performed as previously described by Joung et al.31. In brief, the Human Brunello CRISPR knockout pooled library, a gift from David Root and John Doench (Addgene #73178)32was PCR- amplified with gRNA_library_fwd and gRNA_library_rev primers. The target vector CROPseq_iT7 was digested with the restriction enzyme Esp31. Subsequently, the gel-purified PCR product was cloned into the digested vector using Gibson assembly. The assembled reactions were pooled and purified via a Zymo DNA Clean & Concetrator-25 column. The purified library was electroporated in eight replicates into Endura electrocompetent cells (Lucigen), each consisting of 50-100 ng / pl DNA and 25 pl cells in 0.1 cm BioRad cuvettes at 1800 V, 10 pF and 600 W. Cells were directly recovered in 975 pL prewarmed recovery medium and incubated 1 hours at 37°C and 300 rpm shaking. Then, the culture was transferred to 1 L LB medium containing 100 pg / ml ampicillin and grown overnight at 37°C 230 rpm shaking. DNA was purified by maxiprep using a Purelink™ HiPure Plasmid Maxiprep Kit. Library coverage was determined by Illumina Next Generation Sequencing using staggered guide-specific primers for a first target amplification PCR and dual indexing barcodes for a second barcoding PCR using NEBNext PCR polymerase.
[0172] Tissue culture:
[0173] HeLa cells were cultivated in DMEM Glutamax media supplemented with 10% FCS and 10 pg / ml Ciprofloxacin in a 37°C incubator with 5% CO2. THP1 cells were cultivated in RPMI Glutamax media supplemented with 10% FCS and 10 pg / ml Ciprofloxacin in a 37°C incubator with 5% CO2. Primary human monocytes were obtained from fresh human blood by Ficoll gradient centrifugation and subsequent CD14-based MACS separation (Miltenyi, 130-050-201) according to the manufacturer's instructions and cultivated for differentiation in RPMI Glutamax media supplemented with 10% FCS, 10 pg / ml Ciprofloxacin (Sigma), and 2.5 pg / ml M-CSF (ImmunoTools). THP1 cells expressing ASC-GFP from an NF-0B-dependent promoter were purchased from Invivogen (thp-ascgfp). Library transduction and quality control:
[0174] 1.5 x 107HEK 293T cells were transfected in a 15 cm tissue culture dish using Lipofectamine 2000 (Invitrogen) with 14.1 pg Brunello_iT7 plasmid library, 7,0 pg lentiviral packaging plasmid pMD2.G and 10.6 pg lentiviral packaging plasmid psPAX2. After 4-6 h of incubation, the medium was changed. 48 hours later, virus-containing supernatant was filtered through a 0.45 pm filter (Merck Millipore), aliquoted, and stored at -80°C. For each of four screening replicates, 8 x 106HeLa-Cas9-p65- mNeonGreen cells10were transduced with 1 ml lentivirus library and 10 pg / ml polybrene (Merck Millipore). After one day, cells were selected and induced with 3 pg / ml Puromycin and 1 pg / ml Doxycycline (Cayman Chemicals). Cells were split 1:3 when reaching confluency. For each of two macrophage screening replicates, 1.6xl07THPl-ASC-GFP-Cas9 CASP1 / 8DKOcells were transduced with
[0175] 2 ml of lentivirus library and 10 pg / mL polybrene. The next day, cells were selected with 3 pg / ml Puromycin. After 4-7 days of selection and induction, one million cells were lysed in 100 pl of direct lysis buffer at 65°C for 10 minutes, and 95°C for 15 minutes19. Genomically integrated guide sequences were amplified using NEBNext 2x PCR master mix (NEB) and staggered guide-specific primers (Table 1) in two replicates per library.
[0176] Table 1
[0177] After secondary barcoding PCR, purification, and Nanodrop-based quantification, libraries were sequenced on an Illumina NextSeq 2000 using a P2 100-cycle cassette. Library members were counted using the web tool www.jsb-lab.bio / LibCounter.htm.
[0178] NIS-Seq:
[0179] For genome-scale NIS-Seq perturbation screens, 24-well glass bottom plates (GreinerBio) were coated with 0.1% poly-L-Lysine (w / v in H2O; Sigma-Aldrich) for 30 minutes at room temperature and washed three times with PBS. For HeLa screens, cells were seeded at 5 or 18 days of doxycycline induction. Per replicate, 4 x 105HeLa-Cas9-p65-mNeonGreen Brunello-iT7 library cells were seeded per well and incubated overnight. The next day, cells were stimulated for 45 minutes with 30 ng / ml hrIL-lb (rcyec- hillb; Invivogen) or 30 ng / ml hrTNF-a (rcyc-htnfa; Invivogen), respectively. Live cell nuclei were stained with 2 pM Hoechst 33342 and membranes were stained with 200 ng / ml CellMask Plasma Membrane Stain Deep Red (Thermo). For THP1 screens, cells were pre-differentiated overnight with 100 ng / ml PMA (Invivogen). The next day, cells were carefully washed, detached, and seeded in coated glass bottom plates with 8 xlO5THPl-ASC-GFP-Cas9 CASP1 / 8KOBrunello-iT7 library cells per well. The next day, cells were primed with 200 ng / ml LPS (Sigma Aldrich) for 3 hours, pre-incubated with 50 pM Z- VAD (MedChemExpress) for 30 minutes and stimulated with 7.5 pg / ml Nigericin (Cayman Chemicals) for 1 hour. Live cell membranes were stained with 200 ng / ml CellMask Plasma Membrane Stain Deep Red (Thermo). All live cell phenotype images were acquired in DMEM Fluorobrite with 10 mM HEPES and 2 pM Hoechst 33342 using a 20x objective for HeLa cells and lOx objective with Z-stacks for THP1 cells. After phenotyping, cells were fixed and permeabilized with a 3:1 mixture of methanol and acetic acid for 20 minutes. The fixation was carefully replaced with IX PBS to avoid dehydration of the cells. Cells were washed with nuclease free water before adding T7 in-vitro transcription mix (T7 MEGAscript, Thermo) for 3 hours at 37°C. After IVT, cells were fixed with 4% paraformaldehyde in PBS for 20 minutes and washed with PBS-T (PBS + 0.1% Tween-20) two times. Reverse transcription and post-fixation steps were performed as described by Feldman et al.10using oRT_CROPseq_iT7 as reverse transcription primer. After post-fixation, cells were washed three times with PBS-T and incubated with a gap-fill Phusion mix (lx Ampligase Buffer, 0.4 U / mL RNase H, 100 nM padlock probe oPD_CROPseq_iT7, 0.0125 U / mL NEB Phusion polymerase, 0.5 U / mL Ampligase, 0.05 mM dNTPs, 0.05 M KCI and 5% formamide) for 30 minutes at 37°C and 45 min at 45°C. Cells were then washed twice with PBS-T and incubated with a rolling circle amplification mix overnight at 30°C10. Cells were washed twice with PBS-T before hybridization of 1 pM of the in-situ sequencing primer oSBS_CROPseq_iT7 in 2x SSC for 5 minutes at 37°C. Cells were washed with MiSeq buffer PR2 (Illumina), and perturbation barcodes were analyzed by sequencing-by-synthesis using incorporation and cleavage mixes from a previously used Illumina NextSeq2000 P2 100-cycle cassette. For each of 14 cycles, cells were incubated with the nucleotide incorporation mix for 3 minutes at 60°C, followed by three rounds of five washes with PR2, each with 5 minutes incubation at 60°C. The incorporated nucleotides were imaged after addition of 200 ng / ml Hoechst 33342 in PR2. After each imaging cycle, fluorescent nucleotides were cleaved and de-blocked by incubation with Illumina cleavage mix for 3 minutes at 60°C, three washes with PR2, incubation for 2 min at 60°C, and three additional washes before addition of incorporation mix for the next cycle. Imaging:
[0180] All images were acquired using a Nikon Ti2 body equipped with a Yokogawa CSU-W1 spinning disc unit connected to Lumencor Celesta multimode lasers with wavelengths of 405 nm (nuclear staining), 477 nm (sequencing channel 1, mNeonGreen, GFP), 546 nm (sequencing channel 2), and 638 nm (sequencing channel 3, CellMask deep red). Emission filters used were Chroma ET450 / 50 (nuclear staining), Chroma ET525 / 50 (sequencing channel 1, mNeonGreen, GFP), 572 / 28 BrightLine HC (sequencing channel 2), and 680 / 42 BrightLine HC (sequencing channel 3, CellMask deep red). Exposure times were 90 ms for all channels except p64-mNeonGreen, which was exposed for 150 ms. Objectives used were a Nikon lOx CFI P-Apo, a Nikon 20x CFI P-Apo, or a Nikon 40x CFI Apo 40x Wl with or without a 1.5x tube lens inserted into the light path. A Hamamatsu Orca Flash4.0 LT+ camera was used in electronic shutter mode at full resolution (2048x2048). All components were controlled using custom scripts.
[0181] Image analysis of NIS-Seq data:
[0182] Raw images of up to 14 NIS-Seq cycles were aligned by FFT-accelerated cross-correlation of nuclear staining images. Spots were detected by summing up all sequencing channels across the first three cycles, high-pass filtering, local maximum detection, and brightness thresholding. Spot sequence information was aggregated across 5x5 pixels for every spot after high-pass filtering and eliminating negative values. Channel unmixing was performed by multiplying the channel vector of each cycle with the inverse matrix of average base-wise channel intensities. Non-G bases were called by the maximum of unmixed channels, whereas Gs were called at cycles with all unmixed intensities below 20% of the spot's maximum unmixed intensity across all cycles. Sequences were assigned to the dictionary of known sequences (Brunello sgRNA sequences reverse complemented), allowing zero or one mismatch and no ambiguities. Dictionary-matched and -corrected spot sequences were assigned to nuclei, whose outlines were defined by CellPose using the "nuclei" model33, requiring the dominant sequence to make up more than two thirds of total intensity of library spots in a given nucleus.
[0183] Image analysis of live phenotyping data:
[0184] Z-stacks were collapsed by averaging where applicable. Cell- and nuclear outlines were defined by CellPose using the "cyto2" and "nuclei" models33. Nuclear translocation of mNeonGreen was quantified by calculating the Pearson correlation between nuclear staining and mNeonGreen fluorescence across pixels pertaining to each cell. GFP specking was quantified by local background subtraction (see below), high-pass filtering of fluorescent images, eliminating negative valued pixels, and calculating the mean fluorescence across pixels pertaining to each cell before and after high-pass filtering. Cells with low mNeonGreen or GFP expression were excluded from downstream analysis. For local background subtraction, images were down-sampled 8x8-fold. For each pixel in the full-resolution image, the local minimum across the 9x9 closest pixels in the down-sampled image was subtracted, after which negative values were eliminated.
[0185] Analysis of HeLa genome-scale NIS-Seq perturbation screening data:
[0186] Pairs of live phenotyping and NIS-Seq nuclear images acquired at corresponding stage positions were fine-mapped using FFT-accelerated cross-correlation. Phenotyping nuclei were assigned to the closest nucleus in shifted NIS-Seq data by centers of gravity, with a maximum movement distance of 11.1 pm and with the nuclear area matching within a twofold margin. Any ambiguously mapping nuclei were removed. Cells mapping to the same perturbed gene were aggregated by calculating the mean phenotype (e.g., the mean of Pearson correlations between nuclear and mNeonGreen signal in individual cells). Genes perturbed in less than 15 cells were excluded. Mean phenotypes from two independent replicate screens were visualized as scatter plots. Genes that displayed an altered phenotype in both independent screens were validated by individual lentiviral transductions.
[0187] Analysis of THP1 genome-scale NIS-Seq perturbation screening data:
[0188] THP1 phenotypes were imaged using a lOx objective with Z-stacking to better cover small ASC specks across the cellular cytosol. Furthermore, live nuclear imaging turned out to be affected by the stimulation of cells. Therefore, the assignment of phenotype and NIS-Seq images described above for HeLa cells was modified: First, instead of using phenotyping nuclear images, cell outlines derived using CellPose from Z-aggregated membrane staining were shrunk by 5 pixels to predict the location of the nuclei. Potential pairs of phenotyping and NIS-Seq imaging fields-of-view were identified based on microscope stage positions. Images were scaled according to the relative magnification used, and coarsely mapped at 8x8-downsampled resolution using FFT-accelerated cross-correlation. Best- correlated pairs of fields-of-view were fine-mapped at 2x2-downsampled resolution using FFT- accelerated cross-correlation. All subsequent analysis steps were performed as described above for HeLa cells.
[0189] Generation of single gene perturbation cell lines for validation assays sgRNAs were selected from the Toronto human knockout pooled library (TKOv3), and ordered as DNA oligonucleotides (IDT, Table 1). Oligos were phosphorylated and annealed at equimolar ratio with T4 PNK (Thermo) in lx T4 DNA Ligase Buffer (Thermo) for 30 minutes at 37 °C, 5 minutes at 95 °C, and ramping to room temperature. Diluted annealed oligos were inserted into the CROPseq_iT7 lentiviral backbone using Golden Gate cloning. Purified and sequence-verified plasmids were used to produce lentivirus in HEK 293T cells. Transduced cells were selected with puromycin for 3-5 days and screened for the phenotype of interest.
[0190] References
[0191] 1. Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013).
[0192] 2. Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013).
[0193] 3. Carette, J. E. et al. Haploid genetic screens in human cells identify host factors used by pathogens. Science 326, 1231-1235 (2009).
[0194] 4. Shalem, O. et al. Genome-scale CRISPR-Cas9 knockout screening in human cells. Science 343, 84- 87 (2014).
[0195] 5. Parnas, O. et al. A Genome-wide CRISPR Screen in Primary Immune Cells to Dissect Regulatory Networks. Cell 162, 675-686 (2015).
[0196] 6. Wroblewska, A. et al. Protein Barcodes Enable High-Dimensional Single-Cell CRISPR Screens. Cell 175, 1141-1155. el6 (2018).
[0197] 7. Dixit, A. et al. Perturb-Seq: Dissecting Molecular Circuits with Scalable Single-Cell RNA Profiling of Pooled Genetic Screens. Cell 167, 1853-1866.el7 (2016).
[0198] 8. Datlinger, P. et al. Pooled CRISPR screening with single-cell transcriptome readout. Nat. Methods 14, 297-301 (2017).
[0199] 9. Guerriero, M. L. et al. Delivering Robust Candidates to the Drug Pipeline through Computational Analysis of Arrayed CRISPR Screens. SLAS Discov. Adv. Life Sci. R D 25, 646-654 (2020).
[0200] 10. Feldman, D. et al. Optical Pooled Screens in Human Cells. Cell 179, 787-799. el7 (2019).
[0201] 11. Yan, X. et al. High-content imaging-based pooled CRISPR screens in mammalian cells. J. Cell Biol. 220, e202008158 (2021).
[0202] 12. Schraivogel, D. et al. High-speed fluorescence image-enabled cell sorting. Science 375, 315-320 (2022).
[0203] 13. Schmacke, N. A. et al. SPARCS, a platform for genome-scale CRISPR screening for spatial cellular phenotypes. http: / / biorxiv.org / lookup / doi / 10.1101 / 2023.06.01.542416 (2023) doi:10.1101 / 2023.06.01.542416.
[0204] 14. Funk, L. et al. The phenotypic landscape of essential human genes. Cell 185, 4634-4653. e22 (2022).
[0205] 15. Carlson, R. J., Leiken, M. D., Guna, A., Hacohen, N. & Blainey, P. C. A genome-wide optical pooled screen reveals regulators of cellular antiviral responses. Proc. Natl. Acad. Sci. U. S. A. 120, e2210623120 (2023).
[0206] 16. Ramezani, M. et al. A genome-wide atlas of human cell morphology. http: / / biorxiv.org / lookup / doi / 10.1101 / 2023.08.06.552164 (2023) doi:10.1101 / 2023.08.06.552164.
[0207] 17. Sivanandan, S. et al. A Pooled Cell Painting CRISPR Screening Platform Enables de novo Inference of Gene Function by Self-supervised Deep Learning. http: / / biorxiv.org / lookup / doi / 10.1101 / 2023.08.13.553051 (2023) doi:10.1101 / 2023.08.13.553051.
[0208] 18. Askary, A. et al. In situ readout of DNA barcodes and single base edits facilitated by in vitro transcription. Nat. Biotechnol. 38, 66-75 (2020).
[0209] 19. Schmid-Burgk, J. L. et al. OutKnocker: a web tool for rapid and simple genotyping of designer nuclease edited cell lines. Genome Res. 24, 1719-1723 (2014).
[0210] 20. Liu, T., Zhang, L., Joo, D. & Sun, S.-C. NF-KB signaling in inflammation. Signal Transduct. Target. Ther. 2, 17023- (2017).
[0211] 21. Hart, T. et al. Evaluation and Design of Genome-Wide CRISPR / SpCas9 Knockout Screens. G3 Bethesda Md 7, 2719-2727 (2017).
[0212] 22. Schmacke, N. A. et al. IKKP primes inflammasome formation by recruiting NLRP3 to the transGolgi network. Immunity 55, 2271-2284. e7 (2022).
[0213] 23. Mayor, A., Martinon, F., De Smedt, T., Petrilli, V. & Tschopp, J. A crucial function of SGT1 and HSP90 in inflammasome activity links mammalian and plant innate immune responses. Nat. Immunol. 8, 497-503 (2007). 24. Nizami, S. etal. Inhibition of the NLRP3 inflammasome by HSP90 inhibitors. Immunology 162, 84- 91 (2021).
[0214] 25. Schmid-Burgk, J. L. et al. A Genome-wide CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) Screen Identifies NEK7 as an Essential Component of NLRP3 Inflammasome Activation. J. Biol. Chem. 291, 103-109 (2016).
[0215] 26. Femino, A. M., Fay, F. S., Fogarty, K. & Singer, R. H. Visualization of single RNA transcripts in situ. Science 280, 585-590 (1998).
[0216] 27. Reinhardt, S. C. M. et al. Angstrom-resolution fluorescence microscopy. Nature 617, 711-716 (2023).
[0217] 28. Weinstein, J. A., Regev, A. & Zhang, F. DNA Microscopy: Optics-free Spatio-genetic Imaging by a Stand-Alone Chemical Reaction. Cell 178, 229-241. el6 (2019).
[0218] 29. Fernandopulle, M. S. et al. Transcription Factor-Mediated Differentiation of Human iPSCs into Neurons. Curr. Protoc. Cell Biol. 79, e51 (2018).
[0219] 30. Drmanac, S. et al. CooIMPS ™ : Advanced massively parallel sequencing using antibodies specific to each natural nucleobase. http: / / biorxiv.org / lookup / doi / 10.1101 / 2020.02.19.953307 (2020) doi:10.1101 / 2020.02.19.953307.
[0220] 31. Joung, J. et al. Genome-scale CRISPR-Cas9 knockout and transcriptional activation screening. Nat. Protoc. 12, 828-863 (2017).
[0221] 32. Doench, J. G. et al. Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9. Nat. Biotechnol. 34, 184-191 (2016).
[0222] 33. Stringer, C., Wang, T., Michaelos, M. & Pachitariu, M. Cellpose: a generalist algorithm for cellular segmentation. Nat. Methods 18, 100-106 (2021).
Claims
CLAIMS1. Method for linking the genetic perturbations of individual cells within a cell population to the phenotype of the individual cells, comprising(a) introducing a plurality of different nucleic acid molecules into the genome of the cells, wherein each of the nucleic acid molecules comprises a nucleotide sequence encoding a genetic pertubator, a first promoter and a second promoter, wherein the first promoter is 5' of the nucleotide sequence encoding the genetic pertubator and controls the expression of the genetic pertubator in 5'-3' direction in the cells, wherein the second promoter is a reverse phage promoter and / or reverse in vitro transcription promoter 3' of the nucleotide sequence encoding the genetic pertubator and controls the expression of the genetic pertubator in 3' -5' direction in the nucleus of the cells;(b) initiating the expression of the nucleic acid sequences encoding the genetic pertubators in 5' -3' direction from the first promoter in the cells, under conditions wherein genetic perturbations are introduced into the cells of the cell population via the expressed genetic pertubators;(c) fixing and permeabilizing the cells;(d) initiating the expression of the nucleic acid sequences encoding the genetic pertubators in 3' -5' direction from the second promoter in the fixed cells, thereby generating barcode sequences that are presentative for each kind of genetic perturbation as introduced via the different genetic pertubators;(e) optionally fixing the cells, preferably on a surface;(f) reverse transcription of the barcode sequences;(g) optionally amplification of the reversely transcribed barcode sequences;(h) identifying the reversely transcribed barcode sequences in the individual cells within the cell population by in situ sequencing by synthesis thereby assigning a particular genetic perturbation to each individual cell; and(i) linking the genetic perturbation of each individual cell within the phenotype of each individual cell, wherein the phenotype of each individual cell(I) has been imaged or monitored in the living cells by microscopy any time before step (c) and preferably after step (b), wherein the imaged or monitored phenotype of eachindividual cell is preferably the cell state, the cell motility, the cell shape, the cell-cell interactions, the fluorescence tag of one or more targets in the living cell, the affinity reagent staining of a target in the living cell and / or the fluorescence or color of a cell staining dye, or(II) has been imaged or monitored in the fixed cells by microscopy any time after step (c) and preferably any time before step (f), wherein the imaged or monitored phenotype of each individual cell is preferably the shape of the fixed cell, the fluorescence tag of one or more targets in the fixed cell, the affinity reagent staining of a target in the fixed cell, and / or the sequence-specific detection of one or more RN A and / or DNA sequences in the fixed cell.
2. The method of claim 1, wherein the genetic pertubator is a guide RNA (gRNA) and step (b) is carried out in the presence of a CRISPR-Cas nuclease in the cells, thereby introducing genetic perturbations into the cells of the cell population by CRISPR gene editing.
3. The method of claim 1, wherein the genetic pertubator is a guide RNA (gRNA) and step (b) is carried out in the presence of a modified CRISPR-Cas nuclease without endonuclease activity and a nucleobase modifying enzyme being linked to the gRNA or the modified CRISPR-Cas nuclease without endonuclease activity in the cells, thereby introducing genetic perturbations into the cells of the cell population by the nucleobase modifying enzyme, wherein the nucleobase modifying enzyme is preferably a nucleobase deaminase enzyme.
4. The method of claim 1, wherein the genetic pertubator is a prime editing guide RNA (pegRNA) and step (b) is carried out in the presence of a modified CRISPR-Cas nuclease without endonuclease activity and an engineered reverse transcriptase enzyme being linked to the modified CRISPR-Cas nuclease without endonuclease activity in the cells, thereby introducing genetic perturbations into the cells of the cell population by the pegRNA and the engineered reverse transcriptase enzyme.
5. The method of claim 1, wherein the genetic pertubator is a non-coding RNA (ncRNA) / gRNA hybrid or the combination of a ncRNA and a gRNA, wherein the ncRNA can be reversely transcribed into a reverse transcribed (RT)-DNA with homology to a genomic locus in the cells, and step (b) is carried out in the presence of a reverse transcriptase and a CRISPR-Cas nuclease in the cells, thereby introducing genetic perturbations into the cells of the cell population by CRISPR gene editing that uses the RT-DNA as a homologous recombination template.
6. The method of any one of claims 1 to 5, wherein each individual cell comprises one perturbation only or each cell comprises two or more perturbations.
7. Method for linking the expression of a gene or RNA of interest in individual cells within a cell population to the phenotype of the individual cells, comprising(a) introducing a plurality of different nucleic acid molecules into the genome of the cells, wherein each of the nucleic acid molecules comprises a nucleotide sequence encoding a gene or RNA of interest, a first promoter and a second promoter, wherein the first promoter is 5' of the nucleotide sequence encoding the gene or RNA of interest and controls the expression of the gene or RNA of interest in 5' -3' direction in the cells, wherein the second promoter is a reverse phage promoter and / or reverse in vitro transcription promoter 3' of the nucleotide sequence encoding the gene or RNA of interest and controls the expression of the gene or RNA of interest in 3' -5' direction in the nucleus of the cells;(b) initiating the expression of the nucleic acid sequences encoding the gene or RNA of interest in 5' -3' direction from the first promoter in the cells, thereby expressing the gene or RNA of interest in the cell;(c) fixing and permeabilizing the cells;(d) initiating the expression of the nucleic acid sequences encoding the gene or RNA of interest in 3' -5' direction from the second promoter in the fixed cells, thereby generating barcode sequences that are presentative for each kind of gene or RNA of interest;(e) optionally fixing the cells, preferably on a surface;(f) reverse transcription of the barcode sequences;(g) optionally amplification of the reversely transcribed barcode sequences;(h) identifying the reversely transcribed barcode sequences in the individual cells within the cell population by in situ sequencing by synthesis thereby assigning a particular gene or RNA of interest to each individual cell; and(i) linking the gene or RNA of interest of each individual cell within the phenotype of each individual cell, wherein the phenotype of each individual cell(I) has been imaged or monitored in the living cells by microscopy any time before step (c) and preferably after step (b), wherein the imaged or monitored phenotype of each individual cell is preferably the cell state, the cell motility, the cell shape, cell-cellinteractions, the fluorescence tag of one or more targets in the living cell, the affinity reagent staining of target in the living cell and / or the fluorescence or color of a cell staining dye, or(II) has been imaged or monitored in the fixed cells by microscopy any time after step (c) and preferably any time before step (f), wherein the imaged or monitored phenotype of each individual cell is preferably the shape of the fixed cell, the fluorescence tag of one or more targets in the fixed cell, the affinity reagent staining of a target in the fixed cell, and / or the sequence-specific detection of one or more RN A and / or DNA sequences in the fixed cell.
8. The method of claim 7, wherein the RNA of interest is a guide RNA (gRNA) being complementary to the promoter or coding sequences of a target gene and step (b) is carried out in the presence of a modified CRISPR-Cas nuclease without endonuclease activity and a transcriptional activator being linked to the gRNA or the modified CRISPR-Cas nuclease without endonuclease activity in the cells, thereby increasing the expression of the gene of interest.
9. The method of claim 7, wherein the RNA of interest is a guide RNA (gRNA) being complementary to the promoter or the exonic sequences of a target gene and step (b) is carried out in the presence of a modified CRISPR-Cas nuclease without endonuclease activity in the cells, thereby sterically repressing the transcription of a target gene by blocking either transcriptional initiation or elongation.
10. The method of claim 7, wherein the gene of interest encodes a protein of interest, an aptamer, a scFv-antibody fragment, a nanobody, a computationally designed functional protein or an antibody mimetic, wherein the antibody mimetic is preferably selected from affibodies, adnectins, anticalins, DARPins, avimers, nanofitins, affilins, Kunitz domain peptides, Fynomers®, trispecific binding molecules and probodies.
11. The method of claim 7, wherein RNA of interest is an aptamer, a siRNA, a shRNA, a miRNA, a ribozyme or an antisense nucleic acid molecule.
12. The method of any one of claims 1 to 11, wherein the first promoter is a polymerase-lll promoter, preferably a U6 promoter and / or the second promoter is a reverse phage promoter, preferably a reverse T7 promoter.
13. The method of any one of claims 1 to 12, wherein the phenotype of each individual cell has been imaged or monitored in the living cells by microscopy over time and preferably assisted by a cell assignment algorithm which is preferably a cross-correlation-based search algorithm.
14. The method of any one of claims 1 to 13, wherein the in situ sequencing by synthesis uses a two- or three-color sequencing chemistry.
15. A nucleic acid molecule comprising in 5'-3' direction a U6 promoter, a gRNA and a T7 promoter reverse complement, wherein the nucleic acid molecule optionally additionally comprises a polymerase III terminator, which terminator is preferably located 3' of the gRNA and 5' of the second promoter.
Citation Information
Patent Citations
Specific and high affinity binding proteins comprising modified SH3 domains of FYN kinase
WO2008022759A2
In SITU cell screening methods and systems
WO2019222284A1