CRISPR guide RNA internal standard

The CRISPR-StAR system addresses the challenge of high cell requirements in CRISPR screening by using recombinase recognition sites for stochastic activation/inactivation of sgRNAs, ensuring robust and reliable screening results even in low cell counts.

JP7834648B2Active Publication Date: 2026-03-24IMBA INSTITUT FUR MOLEKULARE BIOTECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-30
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing CRISPR/Cas systems require a large number of cells for screening, which is difficult to achieve in experiments with primary cell lines, organoids, or in vivo settings due to cell proliferation bottlenecks, heterogeneous growth, and cellular heterogeneity, leading to reduced screen quality and inconsistent results.

Method used

A CRISPR system utilizing recombinase recognition sites to stochastically activate or inactivate sgRNAs, allowing for conditional expression and internal control, reducing the number of cells required by introducing recombinases after bottleneck stages to ensure adequate representation and control.

Benefits of technology

The system enables high-resolution CRISPR screening with fewer cells by providing an internal control, reducing noise and improving screen quality, even in low cell counts, and overcoming bottlenecks related to infection efficiency, cell availability, engraftment, and differentiation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007834648000002
    Figure 0007834648000002
  • Figure 0007834648000003
    Figure 0007834648000003
  • Figure 0007834648000004
    Figure 0007834648000004
Patent Text Reader

Abstract

The present invention provides nucleic acids comprising a sequence encoding a single guide RNA (sgRNA) for a CRISPR / Cas system, as well as methods, transgenic cells, and kits for using such sgRNAs, wherein the sgRNA sequence is interrupted by a guide disruption sequence flanked by a first pair of recombinase recognition sites, the sgRNA sequence further comprises a second pair of recombinase recognition sites having a recombinase recognition sequence different from the first pair of recombinase recognition sites, the guide disruption sequence is not flanked by the second pair of recombinase recognition sites, and the sequences flanked by the first and second recombinase recognition sites overlap.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of CRISPR / Cas systems and DNA editing using such means.

Background Art

[0002] CRISPR screening has become a major method for functionally examining genomes in various assays. In positive selection screens, the enrichment of sgRNAs in a cell population is examined to identify genes that promote cell survival upon knockout. In contrast, in negative selection screens, specific sgRNAs are deleted from the cell population because knockout of the corresponding gene results in cell death (Miles et al. FEBS J. 283, 2016: 3170-3180). These screens are also called essentialome screens. To do so, cell lines expressing the bacterial endonuclease Cas9 are transduced with an sgRNA library to induce loss of functional variation of genes. The sgRNA is a short RNA consisting of a 20bp gene-specific stretch and a 3' scaffold that guides the Cas enzyme to a genomic locus complementary to the sgRNA sequence. Upon binding, Cas causes genetic or regulatory changes.

[0003] For a good evaluation of gene function, it is necessary to transduce specific sgRNAs into multiple independent cells to account for cellular heterogeneity and various editing results. In pooled gene screens, the number of independently targeted cells is typically maintained at 300-1,000 cells / sgRNA or more. Therefore, if each gene is targeted using 5 sgRNAs in a genome-wide (20,000 genes) screening approach, this results in a minimum screen size of 300 × 5 × 20,000 = 30 million cells across the entire experiment. For example, Wang et al. (Science 343, 2014: 80-84) describe a large-scale CRISPR-Cas screen using 73,000 sgRNAs to transduce 90 million target cells, i.e., 1,233 cells / sgRNA. 5-10 sgRNAs / gene is recommended. Such screens essentially work with a large number of highly viable immortalized cells. [Overview of the project] [Problems that the invention aims to solve]

[0004] This requirement of a large number of cells is difficult to meet in some experiments, such as with primary cell lines exhibiting heterogeneous growth, organoids of limited size, and in vivo screening. Therefore, a robust method is needed that can reduce the number of cells required in screening and overcome the cell proliferation bottleneck. [Means for solving the problem]

[0005] Summary of the present invention The present invention provides a nucleic acid comprising a sequence encoding a single guide RNA (sgRNA) for the CRISPR / Cas system, wherein the sgRNA sequence is interrupted by a guide disruption sequence franked by a first pair of recombinase recognition sites, and the sgRNA sequence further comprises a second pair of recombinase recognition sites having a different recombinase recognition sequence from the first pair of recombinase recognition sites, the guide disruption sequence is not franked by the second pair of recombinase recognition sites, and / or the second pair of recombinase recognition sites franks a portion of the sgRNA necessary to form an active sgRNA, and the sequences franked by the first and second recombinase recognition sites overlap.

[0006] In this regard, the present invention provides a nucleic acid comprising a sequence encoding a single guide RNA (sgRNA) for a CRISPR / Cas system, wherein the sgRNA sequence is interrupted by a guide disruption sequence franked by a first pair of recombinase recognition sites, and the sgRNA sequence further comprises a second pair of recombinase recognition sites having a different recombinase recognition sequence from the first pair of recombinase recognition sites, where one recombinase recognition site of the second pair of recombinase recognition sites is located between the first pair of recombinase recognition sites and optionally located downstream of the guide disruption sequence, and another recombinase recognition site of the second pair of recombinase recognition sites is located downstream of the first pair of recombinase recognition sites.

[0007] The present invention further provides a method for expressing sgRNA of the CRISPR / Cas system upon recombinase stimulation, the method comprising: A) providing a plurality of sgRNA-coding nucleic acids of the present invention to a plurality of cells; B) introducing or activating one or more recombinases in cells capable of activating a first and second pair of recombinase recognition sites; and C) wherein the activation of the first pair of recombinase recognition sites and the second pair of recombinase recognition sites are competitive reactions, the activation of the first pair of recombinase recognition sites resulting in the expression of active sgRNA, and the activation of the second pair of recombinase recognition sites inactivating the sgRNA sequence.

[0008] Furthermore, cells comprising nucleic acids containing a sequence encoding a single guide RNA (sgRNA) for the CRISPR / Cas system of the present invention are provided. Furthermore, a kit is provided comprising i) a nucleic acid encoding sgRNA, and ii) a nucleic acid for the expression of a recombinase that activates a pair of recombinase recognition sites of the sgRNA-coding nucleic acid.

[0009] All embodiments of the present invention are described together in the following detailed description, and all preferred embodiments are similarly relevant to all embodiments, aspects, nucleic acids, methods, cells, and kits. For example, the descriptions of nucleic acids, cells, and kits themselves also apply to the nucleic acids and means used in the methods of the present invention. The preferred and detailed descriptions of the methods of the present invention also apply similarly to the compatibility and requirements of the nucleic acids, cells, kits, or products in general, such as expressed sgRNA, of the present invention. Unless otherwise specified, all embodiments can be combined with one another. [Modes for carrying out the invention]

[0010] Detailed description of the invention The main challenge of high-resolution in vivo CRISPR is the display of each sgRNA in multiple independent cells (Miles et al., cited above). Ideally, a gene is targeted by multiple sgRNAs, 5-10 sgRNAs / gene, and each sgRNA is displayed in 300-1,000 cells. This so-called library complexity is easily achieved and maintained in immortalized cell lines in vitro. However, achieving this complexity in primary cells or in vivo is far more difficult, and therefore, so far, impossible in a whole-genome library. Furthermore, growth bottlenecks, such as the selection process that removes a large number of cells from the system, lead to a loss of complexity, and thus reduced screen quality due to decreased or lost sgRNA display. Further display bottlenecks include: i) infection efficiency: how many sgRNAs are successfully transduced into independent cells; some cells are more difficult to infect than others. Inefficient sgRNA infection leads to clonal growth and loss of many sgRNAs in the library before screening can even begin. ii) Cell availability: The amount of cells that can be expanded to increase library complexity. Some cell lines have limited proliferative capacity, making it difficult to display sufficient sgRNA before and during actual screening. iii) Engraftment: The amount of transduced (e.g., tumor) cells that survive after in vivo injection also depends on the site of cell injection. Only a limited amount of sgRNA will be displayed, depending on which cells engraft. iv) Differentiation: The bias in which certain cells differentiate instead of other cells in the population. These factors together contribute to a very broad and probabilistic spread of sgRNA display in in vivo experiments, regardless of biological activity. Therefore, absolute display is not a useful predictor of the phenotype induced by a particular sgRNA, leading to insufficient validation of screening results.

[0011] In addition to these bottlenecks, cellular heterogeneity also plays a crucial role in causing confusion and sometimes contradictory screening results. Cells within a population often possess a viability advantage, which can also occur after gene editing, as can be seen with the addition of reporters or immunofluorescence assays.

[0012] The present invention provides a CRISPR system that can reduce the number of cells required at the time of transfection, or the number of cells surviving at the bottleneck. The method of the present invention is based on the stochastic activation or inactivation of sgRNA, thereby creating both activated and inactivated sgRNA in a cell population. Inactivated sgRNA can serve as a control to activated sgRNA, and vice versa. Importantly, the timing of activation and inactivation can be controlled (usually after such a bottleneck), thereby allowing the creation of a control at the same time as the test sgRNA species, for example, when the cell number has recovered during the proliferation phase, thus avoiding the effects of the bottleneck or cellular heterogeneity.

[0013] Based on stochastic activity, the method of the present invention is also called CRISPR-StAR (recombination-induced stochastic activation). By using a recombination system, it is possible to express sgRNA in an active or inactive state. Alternative recombination of two different pairs or sets of recombinase recognition sites results in either activation or inactivation, generating an internal control (e.g., typically inactive sgRNA) within the cell population (Figure 3).

[0014] For such conditional expression of a single guide RNA (sgRNA) in a CRISPR / Cas system, the present invention provides nucleic acids, such as expression cassettes, comprising a sequence encoding an sgRNA sequence. Within the sgRNA, a recombinase recognition site is located, enabling the activation or inactivation of the present invention. The recombinase recognition site of the sgRNA has already been disclosed in WO2017 / 158153A1 and Chylinski et al, Nature Communications 10, 2019: 5454, and is referred to as the CRISPR-switch. Both references are incorporated herein by reference. The present invention leverages the basic principle of recombinase use in sgRNA modification and further advances this principle several steps to provide a conditional activation / inactivation system using at least two different pairs or sets of recombinase recognition, thereby overcoming the problem of insufficient representation in situations of low cell counts.

[0015] sgRNA is RNA used in combination with Cas enzymes such as Cas1, Cas2, Cas3, Cas9, dCas9, Cas10, or Cas12a in CRISPR / Cas methods such as CRISPRi and CRISPRa. A single guide RNA (sgRNA) contains both crRNA (CRISPR RNA) and tracrRNA (trans-activated crRNA) as a single construct. crRNA is also called guide RNA because it contains a DNA guide sequence. TracrRNA and crRNA can be joined to form a single molecule, i.e., a single guide RNA (sgRNA). TracrRNA and crRNA hybridize in complementary regions. This complementary region can be used for binding, and together they can form a stem-loop, which is referred to herein as a crRNA:tracrRNA stem-loop. This region is sometimes called a Cas-binding element because it most often mediates binding to Cas proteins. Site-specific cleavage occurs at a location determined by both the complementarity of base pairs between the crRNA and the target protospacer DNA, and a short motif [called a protospacer adjacency motif (PAM)] juxtaposed in the complementary region of the target DNA. The target DNA can be any DNA molecule that needs modification; it may be the gene to be modified. Typical uses of the CRISPR / Cas system include introducing mutations or modifications to DNA or altering gene expression. sgRNA design is currently conventional, as reviewed, for example, Ciu et al. (Interdisciplinary Sciences Computational Life Sciences 2018, DOI: 10.1007 / s12539-018-0298-z) or Hwang et al. (BMC Bioinformatics 19, 2018:542). Many tools exist that, when used in accordance with the present invention, can generate sgRNA sequences targeting a target gene that have the desired activity (e.g., activation or inhibition of the gene by CRISPR / Cas action).

[0016] According to the present invention, the sgRNA sequence is interrupted by a guide disruption sequence. This guide disruption sequence interferes with the formation of active sgRNA that can be used by the Cas enzyme. The guide disruption sequence is flanked by a first pair of recombinase recognition sites, allowing the guide disruption sequence to be deleted by recombinase action at these sites. The sgRNA sequence includes a second pair of recombinase recognition sites, which have a different recombinase recognition sequence from the first pair of recombinase recognition sites. The difference from the first pair of recombinase recognition sites means that no recombination mixture or linkage occurs between the two types of recombinase recognition sites. Different recombinases can be used for this effect, but some recombinases, such as Cre, recognize many sites without linking such different sites during recombination, so the same recombinase can also be used for the first and second sites.

[0017] Notably, both the first and second recombinase recognition sites result in deletions during recombination; that is, they are oriented in the same direction (not in opposite directions, which would lead to sequence inversion).

[0018] The main difference between the first and second recombinase recognition sites is that only the first pair of sites flanks the guide disruption sequence, while the second pair does not. This means that recombination in the first pair removes the guide disruption sequence (activates the sgRNA), whereas recombination in the second pair does not remove it (the sgRNA remains inactive). In this specification, reference to "inactivating the sgRNA sequence" means that the sgRNA is inactive and can no longer produce active sgRNA, i.e., no recombination occurs in the second pair of recombinase recognition sites. Furthermore, the sequences flanked by the first pair and the second pair overlap, and recombination in the first pair (which deletes the flanked sequences) removes the necessary recombinase recognition sites in the second pair, and in other cases, recombination in the second pair (which deletes the flanked sequences) removes the necessary recombinase recognition sites in the first pair, so that the recombination of the first pair and the second pair is mutually exclusive. Accordingly, in the sgRNA of the present invention, one recombinase recognition site of the second pair of recombinase recognition sites is located between the first pair of recombinase recognition sites (and preferably downstream of the guide disruption sequence), and the other recombinase recognition site of the second pair of recombinase recognition sites is located downstream of the first pair of recombinase recognition sites. "Downstream" means the 5' to 3' direction on the sgRNA sequence. In other words, the nucleic acid of the present invention can also be defined as a nucleic acid comprising a sequence encoding a single guide RNA (sgRNA) for the CRISPR / Cas system, wherein the sgRNA sequence is interrupted by a guide disruption sequence franked by a first pair of recombinase recognition sites, and wherein the sgRNA sequence further comprises a second pair of recombinase recognition sites having a different recombinase recognition sequence from the first pair of recombinase recognition sites, wherein the guide disruption sequence is not franked by the second pair of recombinase recognition sites, and wherein the sequences franked by the first and second recombinase recognition sites overlap.The mentioned second recombinase recognition site, located between the first pair of recombinase recognition sites, is optionally and preferably also located downstream of the guide disruption sequence. This leaves the guide disruption sequence outside the region flanked by the second pair of recombinase recognition sites, and thus leaves the guide disruption sequence active upon inactivation. This efficiently generates inactivated sgRNA. However, other options exist, such as removing a portion of the tracrRNA necessary to form active sgRNA. Such removal would also inactivate the sgRNA. Of course, this removal of a portion of the tracrRNA can be combined with placing the guide disruption sequence outside the region flanked by the second pair of recombinase recognition sites. Thus, instead of "the guide disruption sequence is not flanked by the second pair of recombinase recognition sites," it is also possible that the second pair of recombinase recognition sites provides nucleic acids that flank a portion of the sgRNA necessary to form active sgRNA, such as the essential tracrRNA portion as described above. Such essential portions of tracrRNA may be Cas binding elements or parts of tracrRNA required for any function of tracrRNA as described herein. This choice, where the second recombinase recognition site is downstream of the guide disruption sequence and one of the first pair of recombinase recognition sites, is particularly applicable to sgRNAs in which the tracr portion follows the 5' to 3' structure of the guide. Some Cas enzymes recognize a different order, such as when the guide is downstream (3' side) of tracr. For these Cas enzymes, the order within the sgRNA is reversed, and the second recombinase recognition site must be upstream (i.e., in the 3' to 5' direction) of the guide disruption sequence.Accordingly, the present invention also provides a nucleic acid comprising a sequence encoding a single guide RNA (sgRNA) for a CRISPR / Cas system, wherein the sgRNA sequence is interrupted by a guide disruption sequence franked by a first pair of recombinase recognition sites, and wherein the sgRNA sequence further comprises a second pair of recombinase recognition sites having a different recombinase recognition sequence from the first pair of recombinase recognition sites, wherein one recombinase recognition site of the second pair of recombinase recognition sites is located between the first pair of recombinase recognition sites (and preferably upstream of the guide disruption sequence), and another recombinase recognition site of the second pair of recombinase recognition sites is located upstream of the first pair of recombinase recognition sites.

[0019] Based on this sequence structure, only one pair of the first or second recombinase recognition sites can trigger a recombination reaction, resulting in a deletion of the sequence between the recombinase recognition sites. Which of the pair of sites, i.e., the first pair or the second pair, will result in recombination ("selection") is inherently probabilistic. While it is possible to select a preferred recombination sequence over others, site selection by the recombinase remains inherently probabilistic. Briefly, as will be described in more detail below along with other options, recombination at shorter flanking sequences is generally preferred for the recombinase enzyme over recombination at longer flanking sequences. When using a population or group of cells having the sgRNA sequence of the present invention, probabilistic recombinase site selection by the recombinase enzyme means that a (first) group of cells has recombination at the first recombinase recognition site (activation of sgRNA), and another (second) group of cells has recombination at the second recombinase recognition site (inactivation of sgRNA). The ratio of the first group of cells to the second group of cells depends on the priority of recombination selection between the first and second pairs of recombinase recognition sites.

[0020] Based on these principles, the present invention provides a method for expressing sgRNA of the CRISPR / Cas system upon recombinase stimulation, the method comprising: A) providing a plurality of nucleic acids encoding the sgRNA of the present invention to a plurality of cells; B) introducing or activating one or more recombinases into a cell capable of activating pairs of first and second recombinase recognition sites, C) wherein activation of the pair of first recombinase recognition sites and the pair of second recombinase recognition sites is a competitive reaction, wherein activation of the pair of first recombinase recognition sites results in expression of an active sgRNA, and activation of the pair of second recombinase recognition sites inactivates the sgRNA sequence.

[0021] One advantage of the present invention is that the practitioner can select the time of introduction or activation of the recombinase. Thus, it is possible to grow the cells to the desired number optimal for the assay or screening method under consideration. This allows, for example, selection of a beneficial time point in the recombination paradigm such as the desired stage of differentiation (e.g., terminal or end differentiation) in a differentiation paradigm when starting from totipotent, pluripotent, or multipotent (stem) cells. On the other hand, the presence of both active and inactive sgRNAs provides an internal control within the cell population, allowing for improvements even when the cell number is low and suboptimal.

[0022] The CRISPR-StAR method of the present invention can avoid stochastic display drift by comparing the abundance of sgRNA to an internal inactive control rather than the abundance of sgRNA prior to the bottleneck. Particularly in the case of screens that pass through a bottleneck in sgRNA expression, this control method is more robust than conventional methods of controlling sgRNA before and after screening (Figs. 8 - 13). CRISPR-StAR reduces noise and allows separation of the population of essential sgRNAs and control sgRNAs even when the number of cells is low.

[0023] To maintain the comparability between the active sgRNA and the inactive sgRNA, it is essential that both, i.e., the inactive sgRNA also maintains at least a part of the guide sequence (corresponding to the crRNA, see above), so that the inactive sgRNA can be assigned to the active sgRNA. Of course, since the said sequence is conserved in both recombination events and is specific to the sgRNA, other sequences can be used to assign the inactive and active sgRNAs derived from the same nucleic acid to each other, provided that they are not confused with other sgRNA sequences having other gene targets.

[0024] In a preferred embodiment of the present invention, cells having an inactive portion of the sgRNA sequence are identified as detecting the presence of the sgRNA sequence. Using the inactive sgRNA as a means to detect the presence of the sgRNA in the experiment, it is possible to identify the sgRNA (especially its guide sequence) that was present and thus tested in the experiment, regardless of disappearance due to bottleneck or other reasons for lack. When there is no active sgRNA in the cell population, such a conclusion is usually not obtained because the lack may be caused by the activity of the sgRNA itself (e.g., harmful to cell survival). This means that the lack of active sgRNA (as in the prior art) may mean that the sgRNA interferes with cell survival (and thus the detection of the sgRNA) or that it was lost during the experiment. The absence of the inactive sgRNA of the present invention probably means that it was lost in the experiment, but since the sgRNA remained inactive, it would not have affected cell survival, so this reason can be excluded. This means that the system of the present invention provides evidence that there is no result (cell survival).

[0025] The system of the present invention enables screening at a lower display than initially required for large-scale screening and will thus overcome the bottleneck. This is particularly beneficial for in vivo gene screens, especially large-scale gene screens.

[0026] To overcome the impact of low cell counts on cell survival or proliferation bottlenecks, it is preferable to introduce or activate a recombinase (actively) after the cells having the nucleic acid of the present invention have been grown to a desired number. For example, in a preferred embodiment of the present invention, after step A) and before step B), the cells are grown (cloned), preferably to a number of at least 250, preferably at least 300, at least 350, or at least 400 cells per different sgRNA sequence used in the experiment. In some cases, a larger number is possible and preferable, such as at least 500, or at least 800, for example, 500 to 5000, or 800 to 2000, or more cells per different sgRNA sequence of the present invention. Due to cellular heterogeneity, many cells are tested in parallel in CRISPR experiments. The method and means of the present invention enable the creation of these cells containing the nucleic acids of the present invention (after any step that may reduce the number of cells, such as transfection or transplantation), and then activation of the recombinase, which causes recombination at the first or second recombinase recognition site, thereby activating or inactivating the sgRNA. After this activation / recombinase action, the genetic or physiological effects of the sgRNA in the cell or organism can be observed. The expression "the pair of first and second recombinase recognition sites can be activated" refers to a recombinase that can induce recombination at the first and second recombinase recognition sites. "Activating the recombinase" means that the recombinase performs recombination. The recombinase can exist in an inactive form and can be activated in the presence of a cofactor or other activator.

[0027] In a preferred embodiment, the recombinase is an inducible recombinase. This facilitates the preparation of transgenic cells containing the recombinase, which can then be activated as described in step B) above. Inducible recombinases can be induced by using an inducible promoter or transcriptional enhancer. Activation of the promoter or enhancer increases the expression and activity of the recombinase. Another example is a recombinase that is inactive (as a protein) and activated by the action of an activator. Such recombinases can be genetically engineered. One example is Cre recombinase, CreER, fused to the estrogen receptor (ER) or to the (mutated) ligand-binding domain of the ER. The Cre enzyme is activated by providing a ligand to the estrogen receptor or domain (e.g., 4OH-tamoxifen or tamoxifen).

[0028] Further methods include conditional gene expression systems such as doxycycline-dependent or photoinducible expression of Cre or Flp recombinases. Similarly, gene expression can be induced at a specific time or place using cell type or stage-specific promoters. Yet another example may be chemical stabilization (shielding) or destabilization (degroning) of the recombinase activity.

[0029] For example, recombinases may be induced or activated (for example, by administering 4OH-tamoxifen) in cells or cell cultures, or in animals containing cells having the nucleic acids of the present invention, after a bottleneck in cell / sgRNA representation, when, for example, the number of cells recovers to, for example, at least 500 or at least 1,000 cells / sgRNA.

[0030] Preferably, the nucleic acids of the present invention are used in cells, i.e., provided to or have been provided to cells in the method of the present invention. Cells should also be able to stably proliferate the nucleic acids along with cell proliferation. This is done, for example, by incorporating the nucleic acids or sgRNA sequences into the cell's genome.

[0031] Preferably, each cell has one copy of the nucleic acid encoding the sgRNA of the present invention. This ensures that only one type of recombinase reaction (either activation or inactivation, but not both) occurs in a given cell. Of course, different cells may exhibit different recombinase reactions according to the stochastic principle described above, providing inactive or active sgRNA populations within the cell. To ensure that a cell has only one copy of sgRNA, one can target a specific genomic locus, such as the AASV1 locus disclosed in Wang et al. (2014, above), but of course, other unique loci are also possible. Only one insertion into the genome is possible per cell.

[0032] The nucleic acid of the present invention preferably comprises an sgRNA sequence and also preferably a promoter operably ligated to the sgRNA sequence for the expression of the sgRNA sequence. The promoter may be a constitutive promoter or an inductive promoter. A constitutive promoter is particularly preferred because the activity of the sgRNA is regulated by the sgRNA sequence construct of the present invention (guide disruption sequence or inactivation recombinase product) itself. Examples of promoters are disclosed in particular in Ma et al. (Molecular Therapy-Nucleic Acids 3, 2014: e161). The promoter may be an RNA polymerase II (Pol II) or RNA polymerase III (Pol III) promoter (see WO2015 / 099850). Preferably, it is a Pol III promoter such as the U6, 7SK, or H1 promoter. The structure of the Pol III promoter is disclosed in Ma et al. 2014. The use of the H1 promoter is shown, for example, in WO2015 / 195621 (incorporated herein by reference), and this method and the design of the construct can be used according to any aspect of the present invention. The preferred promoter is the U6 promoter.

[0033] The Pol II promoter can be selected from the group consisting of the retroviral Roussarcoma virus (RSV) LTR promoter (optionally including an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally including a CMV enhancer), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, the EFla promoter, and any one of the following promoters: CAG, EF1A, CAGGS, PGK, UbiC, CMV, B29, Desmin, Endoglin, FLT-1, GFPA, and SYN1. The Pol II promoter can be used in combination with a Csy4 cleavage site or an autocleavage ribozyme that franks the guide RNA sequence, as disclosed in WO2015 / 099850. The use of pol II in guide expression is further described in WO2015 / 153940.

[0034] The nucleic acid also preferably includes a selection marker. The selection marker can be used to identify, and preferably select or isolate, cells containing the nucleic acid of the present invention. Thus, the success of such transformation of cells having the nucleic acid of the present invention can be confirmed and controlled. The cells having the selection marker, and consequently the sgRNA of the present invention, can then proceed to step A) of the method of the present invention, etc.

[0035] Such selection markers may be any marker known in the art. They may be cell survival markers, such as antibiotic resistance genes, or optical markers, such as genes encoding fluorescent proteins like GFP, BFG, or RFG.

[0036] Preferably, the marker is positioned at a location where it is excised by the first and / or second recombinase activity, i.e., it is flanked by the first and / or second pair of recombinase recognition sites. This removal prevents interference of the selection marker sequence in the formation of active sgRNA, or, in the case of inactive sgRNA, this helps to reduce the size, as the inactive sgRNA is preferably identified by sequencing. Reducing the sequencing size reduces sequencing effort and cost, which is particularly important in large screens where many sgRNA sequences are sequenced. A further advantage of positioning the selection marker between each of the first and second pairs of recombinase recognition sites, i.e., overlapping, is that it is removed by both activation and inactivation, thereby enabling reverse selection against premature recombination before using the nucleic acid of the present invention. Thus, the selection "and" is most preferred, i.e., the marker is preferably flanked by both the first and second pairs of recombinase recognition sites, i.e., within the overlapping portion of the first and second pairs of recombinase recognition sites.

[0037] Preferably, the nucleic acid of the present invention comprises one or more primer or probe binding sites, so that the nucleic acid primer or probe can bind to the nucleic acid for detection of the sgRNA of the present invention in either its non-recombinase-transformed (original) state or in an inactive or active sgRNA state. The primer can also be used to amplify or sequence the sgRNA sequence for detection and preferably identification. The nucleic acid can be bound using the probe, and the sgRNA sequence can be further bound using the probe for sequence identification.

[0038] Preferably, the primer or probe binding site is located outside the first and second pairs of recombinase recognition sites and is therefore maintained during and after the action of the recombinase. Such a probe or primer binding site may, for example, flank the entire guide sequence or sgRNA sequence. Preferably, two probe or primer binding sites are used, one at 5' of the guide sequence or sgRNA sequence and the other at 3' of the guide sequence or sgRNA sequence. One or more probe or primer binding sites are preferably located near the sgRNA sequence, preferably within 20,000 nt (nucleotides) from any end of the sgRNA sequence, preferably within 15,000 nt, 10,000 nt, 5,000 nt, or 1,000 nt from any end of the sgRNA sequence.

[0039] The structure of sgRNA is disclosed, for example, in Jiang and Doudna (Annu. Rev. Biophys. 46, 2017:505-29), Swarts et al. (Molecular Cell 66, 2017: 221-233), WO2015 / 089364, WO2014 / 191521, WO2015 / 065964, and WO2017 / 158153A1. The sgRNA molecule contains a portion corresponding to crRNA, which includes a guide typically 15–30 nt, most often 17–21 nt in length, that mediates target specificity. The crRNA may contain a pseudoknot structure and / or a seed region. This crRNA portion is connected to a portion corresponding to tracrRNA. In sgRNA, the crRNA and tracrRNA portions are fused in a stem-loop region, which typically contains (crRNA) repeats, loops, and (tracrRNA) antirepeats. The stem may also contain mismatched nucleotides in addition to the palindromic sequence. The sgRNA portion corresponding to the tracrRNA further has a loop region and may generally have a 3D folded structure that mediates binding to the Cas enzyme. Inactivation of sgRNA by the action of a recombinase on a second pair of recombinase recognition sites is preferably caused by interfering with the binding of Cas1, Cas2, Cas3, Cas9, dCas9, Cas10, Cas12a, Cas12b, or Cas12c, preferably Cas9, and / or its variants, such as dCas9, to a selected Cas enzyme, for example, by interfering with the folding structure necessary for Cas binding. The deleted region is the region that is flanked by the pair of recombinase recognition sites. Preferably, the deletion results in the deletion of one or more loops or a portion of a loop.

[0040] In the case of active sgRNA, deletion of the first pair of recombinase recognition sites by recombinase action should maintain the active crRNA-tracrRNA structure and establish the Cas binding ability of the sgRNA. Therefore, it is preferable that the first recombinase recognition sites (one of which remains after recombinase action) are located in an inactive region such as a loop. Accordingly, in a preferred embodiment of the present invention, one, preferably two, of the first recombinase recognition sites are located within the loop region of the sgRNA sequence. Preferably, the sgRNA sequence includes a crRNA portion and a tracrRNA portion, and one of the first recombinase recognition sites is located within the crRNA-tracrRNA linker loop, i.e., the loop connecting the crRNA and tracrRNA portions.

[0041] Preferably, the sgRNA includes one or more loops, e.g., 1, 2, 3, 4, or more loops. One loop preferably connects a crRNA-tracrRNA portion. One loop may be contained within a crRNA portion such as a pseudoknot structure. Preferably, the tracrRNA portion includes 1, 2, 3, or more loops that are entirely within the tracrRNA portion (not counting the crRNA-tracrRNA linker loop). Typically, Cas binding requires the first two loops after the crRNA-tracrRNA linker loop, and one, preferably both, of these are deleted or partially deleted to inactivate the sgRNA by recombinase action (e.g., by placing the recombinase recognition site within the loop and deleting one leg of the stem). The loops can be connected to a stem of 3-20 nt in length, etc., where the length of one leg of the stem is counted, i.e., the stem may be twice that number when counting base pairs, but of course, base mismatches are also included. Preferably, the stem of any one of these loops, in particular the complete stem loop of the entire tracrRNA portion, contains 3 to 20 base pairs, e.g., 4 to 15 or 5 to 10 base pairs, e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 base pairs, preferably 4 to 7 base pairs. Preferably, the crRNA-tracrRNA linker loop is a stem loop having a length of 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 base pairs within the stem.

[0042] Guide disruption sequences that interfere with the formation of the active guide in the sgRNA before recombinase action may include transcription termination sequences such as poly(A) sequences, resulting in the remainder of the sgRNA not being transcribed. These may also include sequences that interfere with the folding of active guide RNA or sgRNA, which can interact with the Cas enzyme to form an active CRISPR-Cas complex. This can be achieved, for example, by including or having a sufficiently long folding element that does not bind to the Cas enzyme. Guide disruption sequences may interfere with the formation of any loop in the sgRNA, particularly preferably the crRNA-tracrRNA linker loop. Such sequences may be, for example, selection marker sequences, if they are long enough to interfere with the folding of the active guide RNA. In a particularly preferred embodiment, both a transcription termination sequence and a sequence (of a certain length) that interferes with the folding of the active guide RNA are used. In a preferred embodiment, the transcription termination sequence is located within the loop, particularly preferably within the crRNA-tracrRNA linker loop. Preferably, both of the first flanking recombinase recognition sites are located in the same loop, and as a result, only a portion of that loop is deleted in the activated recombinase reaction. In such a case, as described above, one of the second recombinase recognition sites is also located in the same loop because the first and second flanking sequences overlap. Another of the second recombinase recognition sites is preferably located downstream, and as a result, an essential portion of the sgRNA, particularly its tracrRNA, is deleted upon action of the recombinase on the second pair of recombinase recognition sites. Preferably, the second recombinase recognition site is located entirely within the tracrRNA, or downstream of the tracrRNA portion of the sgRNA, for example, also downstream of the sgRNA, in a loop. This may be in the transcribed region or further downstream, for example, after the transcription end. For example, this may be within 10,000 nt of the 5' end of the sgRNA, preferably within 5,000 nt of the 5' end of the sgRNA.

[0043] In a preferred embodiment, the first and second pairs of recombinase recognition sites are activated by the same recombinase enzyme. Thus, the recombinase may have different recognition sites that do not interact with each other in the recombinase reaction. Such a recombinase is, for example, Cre. The first and second pairs of recombinase recognition sites can be independently selected from lox sites, such as loxP, lox511, lox5171, lox2272, M2, M3, M7, M11, lox71, and lox66.

[0044] Using one recombinase for both pairs of recombinase recognition sites has the advantage that only one recombinase needs to be provided to the cells. In other embodiments, the first and second pairs of recombinase recognition sites are activated by different recombinase enzymes. In this case, the population of cells used in the experiment must have both recombinase enzymes, and individual cells in that population may have one, preferably both, of the two recombinase enzymes. Examples of recombinases with both options (i.e., the same or different recombinases) are site-specific recombinases such as Cre, Hin, Tre, and FLP.

[0045] The recombinase reactivity and the selection priority (and therefore the probabilistic distribution—see above) between the first and second pairs of recombinase recognition sequences can be controlled by structural elements and sequences. Shorter sequences that are flanked by the recombinase recognition site will be preferred over longer sequences that are flanked by the recombinase recognition site, and will therefore result in more deletion events in the flanked region. Thus, the distribution or ratio of active and inactive sgRNAs during recombinase action can be manipulated by selecting the length of the flanked region accordingly. Another option for controlling the distribution or ratio of active and inactive sgRNAs during recombinase action is to add more recombinase recognition sites, which also increases the likelihood of triggering a recombinase reaction. When using three or more recombinase recognition sites, if recombinase activity causes deletion of regions flanked by internal recombinase recognition sequences, at least two recombinase recognition sequences will remain, leading to further deletions, which will continue until only one recombinase recognition site remains, thus resulting in the deletion of the entire sequence portion between the outermost recombinase recognition sites. Therefore, even when a "set" of three or more recombinase recognition sites of the same type (as first and / or second sites) is used in the nucleic acid sequence of the present invention, the flanked and overlapping portions should be selected, for example, by considering the outermost recombinase recognition sites as a "pair" of recombinase recognition sites as described herein.

[0046] Preferably, the nucleic acids are modified to provide an average ratio of active sgRNA to inactive sgRNA of 9:1 to 1:9, preferably 5:1 to 1:5, and particularly preferably 2:1 to 1:2. Such ratios can also be achieved by the methods of the present invention.

[0047] Since recombinase activity is length-dependent, for proper recombinase activity, pairs of recombinase recognition sites are preferably separated by a maximum of 100,000 nt, preferably a maximum of 50,000 nt, particularly preferably a maximum of 10,000 nt, or a maximum of 5,000 nt. This applies to the first and / or second pair, preferably both.

[0048] Preferably, the nucleic acids of the present invention include a unique molecular identifier (UMI) or barcode. The UMI or barcode is a sequence that enables the identification of a specific sgRNA molecule and is different for each molecule, even when targeting the same gene target (same guide sequence). An example of a UMI is a random sequence. Such a UMI must be long enough to distinguish all nucleic acid molecules used. Preferably, the length of the UMI is at least 6 nt, preferably at least 8 nt. For example, it may be 6-40 nt, preferably 8-20 nt in length. Preferably, it is located downstream of the sgRNA sequence. Also preferably, it is located outside both the first and second (or any) pairs of recombinase recognition sites so that it is not deleted during the action of the recombinase and is retained by both active and inactive sgRNA. Other UMIs may also be located on one (but not the other) of the first and second recombinase recognition sites to enable tracking of only active sgRNA or only inactive sgRNA (i.e., it is retained during the action of the recombinase). However, the use of UMIs present in both active and inactive sgRNAs is recommended. Using UMIs allows for the analysis of independent events passing through the bottleneck as independent replications (Michlits et al., Nature Methods 14, 2017: 1191-1197), thus explaining clonal proliferation. Cells with different UMIs can be used as biological replicas, which is a great advantage for highly heterogeneous compositions in assays, such as organoid cultures and in vivo applications. Thus, in the method of this invention, UMIs are used to identify the same sgRNA in different cells. This means that these cells are clones of a single original cell transformed to contain a nucleic acid molecule encoding one specific sgRNA. UMIs detected in the product after recombinase activation may also indicate the degree of the proliferation bottleneck. A lower number of UMIs per guide in the cell population before and after the bottleneck indicates cell loss and to what extent.

[0049] The present invention also provides cells comprising nucleic acids encoding the sgRNA of the present invention. These cells can be used in the methods of the present invention. The cells may be mammalian cells, preferably human or non-human cells. If totipotent cells are used, they are preferably non-human. These may be primate, mouse, bovine, or rodent cells. The cells may be isolated cells or aggregates of cells, e.g., cultures, organoids, or in vivo cells. The in vivo cells of the present invention are preferably not present in humans. The cells may be cell lines and / or pluripotent cells. However, the cells do not need to maintain pluripotency and may differentiate. Recombinase activity (and therefore, activating a portion of the sgRNA according to the stochastic principle) is performed at any point in proliferation or growth. The present invention also relates to cells having such activated or inactivated sgRNA.

[0050] The cells preferably contain one or more nucleic acids, such as expression constructs, for the expression of one or more recombinases, such as Cre. The recombinases must activate the first and / or second recombinase recognition sites, as described above. The expression nucleic acids may include a selection marker. The selection marker can be used to identify and / or isolate cells having an active recombinase. The selection marker may include a specific sequence, such as a length marker, a barcode, or a cell survival marker, such as an antibiotic resistance gene. Length markers can be identified, for example, during sequencing. The markers may optionally or additionally serve as a control in the production of nucleic acids encoding the recombinase protein in viruses. It is possible to use viruses as transfection materials to transform cells that later express recombinases (and are used in the methods of the present invention). A suitable virus having the markers can be selected. As described above, the nucleic acid for recombinase expression, e.g., an expression construct, preferably contains an inducible or alternatively constitutive promoter. However, recombinases can preferably be induced by the selection of a promoter or by using a recombinase that requires controllable activation when expressed (e.g., CreER disclosed above). Photoactivatable Cas9 is also possible (Nihongaki et al., Nature, 2015, 33(7): 755-760). The method of the present invention would then also include the step of photoactivating Cas. In some cases, the recombinase may not be active in all cells ("unreacted"). This is usually not a problem because non-recombinase-activated sgRNAs have different sequences from inactive and active sgRNAs after activation and can therefore be identified and examined. When using a recombinase under a cell type-specific promoter (e.g., CreER), recombination actually also selects cell type specificity, making it possible to measure only the cell type of interest, even if additional cell types are transduced with sgRNA.

[0051] In a preferred embodiment of the present invention, the sequences of the active / inactive / unreacted sgRNA are determined after activation / introduction of the recombinase in step B), and preferably after the effect is observed in the cells after step C). To determine the sequence of the sgRNA, the nucleic acid of the present invention preferably includes a primer binding site as described above. The primer binding site enables sequencing of the sgRNA (including its active / inactive recombinant product) and any UMI, if present, where the primer binding site franks the sgRNA sequence and the UMI sequence.

[0052] The cells preferably contain nucleic acids, such as expression constructs, for expressing Cas, such as Cas9, or any of the Cas enzymes described above. These nucleic acids, for example, expression constructs for Cas expression, may also contain inductive or alternatively constitutive promoters. Inductive promoters are preferred to allow control of Cas enzyme activity. The Cas nucleic acids may also contain similarly but independently selected selection markers, as described above.

[0053] Recombinases and / or Cas enzymes, preferably both, are provided into the cell. For example, commercially available cells in which these are incorporated into the genome are available. Thus, the description of nucleic acids extends to the cell's genome.

[0054] Typically, large-scale screening and other experiments use many cells. Preferably, the cells of the present invention are provided in a population of at least 10,000 cells, more preferably at least 100,000 cells, or at least 1 million cells. Preferably, the cells have different sgRNAs according to the number of cells per sgRNA (i.e., sgRNAs with different guides).

[0055] Cells can be studied for any effect of sgRNA on their proliferation morphology or activity, which may be altered by active sgRNA, compared to cells without active sgRNA, particularly cells with inactive sgRNA. Such studied cells may be wild-type cells or may have mutations. In such cases, the effect of active sgRNA on the effects of the mutation may be observed. Such mutations may be oncogenic mutations, such as activation or upregulation of oncogenes, or repression or inactivation of tumor suppressor genes.

[0056] Accordingly, in a preferred embodiment of the present invention, the cells further express transgenic oncogenes or have repressed tumor suppressor genes. The method of the present invention further includes observing the difference in tumorigenesis after activation in step C) compared with cells not activated in step C), thereby screening the role of genes targeted by sgRNA during tumorigenesis. A portion of the tumor proliferates, i.e., cells containing inactive sgRNAs proliferate after recombinase action. If active sgRNAs corresponding to inactive sgRNAs are not found in the tumor, the presence of these inactive sgRNAs is evidence that the active sgRNAs were initially activated and present but were unable to proliferate in the tumor. Thus, essential genetic targets for tumor growth or its suppression have been identified. As described above, the presence of inactive sgRNAs provides evidence that active sgRNAs are not present.

[0057] In another embodiment, the CRISPR-StAR cassette is incorporated into a germline or cell line of an animal model to enable sparse gene deletions, for example, to generate a tumor model with rare and reproducible loss of tumor suppressors. The present invention also includes testing the effects of candidate compounds in combination with sgRNA activation. Thus, cells can be further treated with the candidate compound, and this method further includes observing differences in cellular activity or morphology after activation in step C) compared to cells not activated in step C), thereby screening for the activity of genes targeted by sgRNA under the influence of the candidate compound. Such a method can be used, for example, in toxicity screening. The candidate compound may be a toxin, and improvement in toxicity can be observed when the sgRNA is active.

[0058] The method of the present invention is particularly suitable for overcoming the bottleneck of low cell numbers, as described above. Such situations occur in in vivo transplantation, organoids, or xenocellular cultures. Therefore, these are preferred applications of the present invention. Preferably, cells are grown in or within tissue aggregates such as organoids. The tissue of the aggregate (e.g., organoid tissue) may be liver, spleen, cerebrum, muscle, heart, kidney, colorectal, bladder, blood vessels, ovaries, testes, or pancreatic tissue. Preferably, cells are also transferred to non-human animals to form allografts or xenografts. Preferably, the animals are rodents, non-human primates, cattle, horses, pigs, mice, hamsters, rats, etc. The introduction or activation of the recombinase occurs in the tissue aggregates, organoids, or non-human animals, i.e., after the transplantation or engraftment bottleneck has passed and the cells of the present invention have preferably grown to the desired number. In other embodiments, cells are grown in cell cultures such as 2D or 3D cell cultures. Recombinase activation occurs when the desired cell number and / or cell differentiation stage is reached. The desired cell number is described above in terms of a specific cell / sgRNA ratio.

[0059] The present invention further provides a kit comprising any means used in the method of the present invention, such as nucleic acids and / or cells. In particular, the kit comprises i) a nucleic acid encoding the sgRNA of the present invention, and ii) a nucleic acid for the expression of one or more recombinases that activate a pair of recombinase recognition sites of the sgRNA. The kit preferably further comprises iii) a nucleic acid encoding a Cas gene. Any such nucleic acid can be further defined as described above and, for example, has a promoter operably linked to the sgRNA, recombinase, and Cas protein. A gene is generally considered to contain a promoter and a coding region.

[0060] The present invention is not limited to these embodiments, but will be further illustrated by the following figures and examples. [Brief explanation of the drawing]

[0061] [Figure 1] Distribution of the log2 multiple change between barcodes before and after pooled CRISPR screening when the number of barcodes per guide in the library is reduced. [Figure 2] A) Schematic diagram of a 2D in vitro gene screen without bottlenecks. In each division of the cell population, the cell / sgRNA representation is maintained at a level higher than 500-1,000 cells / sgRNA, thus maintaining the complexity of the screen. B) Schematic diagram of a complexity bottleneck in the gene screen. After a bottleneck caused by infection efficiency, the limited number of cells, engraftment efficiency, and / or differentiated cells are differentially recovered, and the cell / sgRNA representation decreases. Regardless of clonal size heterogeneity of cells, single-cell derived clones are probabilistically divided into experimental and control populations, indicated by the upper green double arrow (active sgRNA) and the lower red double arrow (inactive sgRNA). [Figure 3] Schematic diagram of a CRISPR-StAR vector encoding (A, B) sgRNA, stop cassette, selection cassette, tracrRNA, and UMI. Recombination yields active (A) sgRNA or inactive (B) sgRNA. [Figure 4]Schematic diagram of the CRISPR-StAR construct series. StAR1 contains two sets of different lox sites. Compared to StAR1, StAR3 contains an extra loxP site, resulting in a longer distance between the Lox5171 site and the stop cassette, and a shorter distance between the tracr and the second Lox5171 site. Removing the extra loxP site yielded construct StAR4. [Figure 5] An experimental outline for determining the frequency of active to inactive recombination in CRISPR StAR constructs. [Figure 6] A schematic diagram of the demonstration experiment. [Figure 7] Benchmark test for CRISPR StAR analysis. Comparison with the conventional day 0 baseline. [Figure 8-1] Correlation of two highly complex biological replicas using conventional methods (active vs. day 0) and CRISPR-StAR analysis (active vs. inactive). Each dot represents one sgRNA. Density plots and stacked histograms show the guide distribution for each replica. Essential replicas are shown in red, and non-essential replicas in blue. [Figure 8-2] Same as the explanation above. [Figure 8-3] Same as the explanation above. [Figure 9-1] Correlation of two low-complexity biological replicas using conventional methods (active vs. day 0) and CRISPR-StAR analysis (active vs. inactive). Each dot represents one sgRNA. Density plots and stacked histograms show the guide distribution of each replica. Essential sgRNAs are shown in red, and non-essential sgRNAs in blue. In addition to a dramatic increase in the spread of neutral (blue) sgRNAs, an additional complete dropout is observed at very low levels. This is due to the fact that sgRNAs are completely lost at the bottleneck. In contrast, CRISPR-StAR scores only sgRNAs that are detected in the inactive conformation and lost in the active conformation. [Figure 9-2] Same as the explanation above. [Figure 9-3] Same as the explanation above. [Figure 10-1]Area under the curve analysis of essential (red) biological replicas compared to non-essential (blue) two biological replicas in the decrease in cell number per guide within the library. [Figure 10-2] Same as the explanation above. [Figure 10-3] Same as the explanation above. [Figure 10-4] Same as the explanation above. [Figure 11-1] Area under the receiver operating characteristic curve (AUROC) analysis for reduced cell number complexity compared to library analysis. CRISPR StAR analysis (active vs. inactive) is green, and conventional analysis (active vs. day 0) is black. [Figure 11-2] Same as the explanation above. [Figure 12] Pearson correlation, delta area (dAUC) analysis of the reduction in cell number complexity compared to the library, and receiver operating characteristic area (AUROC) analysis. Black dots represent values ​​for individual replicas, and bars represent the average of two replicas. [Figure 13-1] Improved robustness of organoid screening. a) Correlation between two biological replicas determined by UMI. Density plots and stacked histograms show the guide distribution for each replica. b) For top-ranked sgRNAs correlated with the same gene, the average number of guides targeting the same gene (y-axis) is shown (x-axis). c) Volcano plot of conventional method (active vs. day 0) and CRISPR-StAR analysis (active vs. inactive) for two biological replicas determined by UMI. Top-ranked genes are shown in blue. Genes scored in other replicas are shown in green. [Figure 13-2] Same as the explanation above. [Figure 13-3] Same as the explanation above. [Figure 14]Correlation plots of in vitro and in vivo CRISPR-StAR screening results. Each dot represents all sgRNAs of a single gene, and the dot size represents the number of UMIs per gene in the in vivo sample. The stacked histogram shows the guide distribution in each sample. The in vivo sample consists of two combined replicas. Essential genes are shown in red, and non-essential genes are shown in black. The majority of essential genes show a decrease in display in both in vitro and in vivo. [Figure 15] The Sleeping Beauty transposon possesses an EGFP-P2A-FAH expression cassette under the control of an EF1a promoter containing a CRISPR-StAR construct. (Left) Liver of FAH- / - mice injected with saline only, maintained in NTBC, and collected 14 days after injection. (Right) Liver of FAH- / - mice injected with the transposon and transposase, and collected 25 days after injection. Nuclei were counterstained with DAPI (blue), and proliferating cells containing the CRISPR-StAR construct were visualized with EGFP (green). [Figure 16] A "Sleeping Beauty" transposon possessing a KrasG12D-P2A-FAH expression cassette under the control of an EF1a promoter containing a CRISPR-StAR construct. (Left) Liver of a wild-type mouse collected 50 days after injection with transposase only. (Right) Liver of a wild-type mouse collected 50 days after injection with both transposase and transposase. The nuclei were counterstained with DAPI (blue), and proliferating cells containing the CRISPR-StAR construct were visualized with EGFP (green). [Examples]

[0062] Example 1: Materials and Method 1.1 Materials

[0063] 1.1.1 Cell lines Tamoxifen-inducible Cre-ERT mouse embryonic stem cells AN3-12 (ESC) Platinum-E cell (Cell Biolabs RV-101) Vil-CreERT2;Rosa-LSL-Cas9-2A-eGFP mouse small intestine organoids

[0064] 1.1.2 Cell culture medium Mouse embryonic stem cell culture medium (ESCM): 450 ml DMEM, 75 ml FCS (Sigma, 025M3347), 5.5 ml penicillin-streptomycin (Sigma), 5.5 ml NEAA (Gibco), 5.5 ml L-glutamine (Gibco), 5.5 ml sodium pyruvate (Sigma), 0.55 ml β-mercaptoethanol (Merck), 7.5 μl LIF (2 mg / ml)

[0065] Complete organoid medium: Advanced DMEM / F12, penicillin / streptomycin, 10 mmol / L hepes, glutamax, 1xN2, 1xB27 (all from Invitrogen), and 1 mmol / L N-acetylcysteine ​​(Sigma), recombinant human Wnt-3A, mouse EGF, mouse noggin, human R-spongin-1, nicotinamide

[0066] 1.1.3 Buffer Lauryl sarcosin lysis buffer: 10 mM Tris-HCl, pH 7.5 (Sigma Aldrich), 10 mM EDTA (Sigma Aldrich), 10 mM NaCl (Sigma Aldrich), 0.5% N-Lauryl Sarcosine (Sigma Aldrich), 1 mg / ml Proteinase K (Thermo Fisher Scientific), 0.1 mg / ml RNase A (Qiagen)

[0067] 2XSDS lysis buffer: 10 mM Tris-HCl, pH 8 (Sigma Aldrich), 1% SDS (in-house), 10 mM EDTA (Sigma Aldrich), 100 mM NaCl (Sigma Aldrich), 0.1 mg / ml RNase A (Qiagen)

[0068] 1.1.4 Primer FW_G_CrSc_5: AATGATACGGCGACCACCGAGATCTACACAGATAACGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 1) FW_G_CrSc_6: AATGATACGGCGACCACCGAGATCTACACAGCTTGCGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 2) FW_G_CrSc_7: AATGATACGGCGACCACCGAGATCTACACAGGACACGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 3) FW_G_CrSc_10: AATGATACGGCGACCACCGAGATCTACACATCACTCGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 4) FW_G_CrSc_12: AATGATACGGCGACCACCGAGATCTACACCAACACCGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 5) FW_G_CrSc_13: AATGATACGGCGACCACCGAGATCTACACCACGCCCGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 6) FW_G_CrSc_15: AATGATACGGCGACCACCGAGATCTACACCATTACCGAGGGCCTATTTCCCATGATTCCTTC (SEQ ID NO: 7) FW_G_CrSc_19: AATGATACGGCGACCACCGAGATCTACACCCCCAACGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 8) FW_G_CrSc_20: AATGATACGGCGACCACCGAGATCTACACCGTCATCGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 9) FW_G_CrSc_21: AATGATACGGCGACCACCGAGATCTACACCTATGCCGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 10) FW_G_CrSc_22: AATGATACGGCGACCACCGAGATCTACACCTCCGCCGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 11) FW_G_CrSc_39: AATGATACGGCGACCACCGAGATCTACACTGCCGACGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 12) FW_G_CrSc_41: AATGATACGGCGACCACCGAGATCTACACTGTAGACGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 13) FW_G_CrSc_42: AATGATACGGCGACCACCGAGATCTACACTTGCCACGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 14) RV_G_CrSc: CAAGCAGAAGACGGCATACGAGATACCGTTGATGAGTAG (Sequence ID 15) NGS_U6: CGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCG (Sequence ID 16)

[0069] 1.2 Method 1.2.1 Mouse Embryo Stem Cell Culture Cells were cultured in ESCM, which was changed daily. Upon reaching confluence, the cells were trypsinized and divided into a 1:10 ratio. For 4-hydroxytamoxifen (4OH) treatment, 0.5 μM 4OH(Sigma) was supplemented to the medium daily.

[0070] 1.2.2 Small Intestine Organoid Culture Small intestinal organoids were established from Vil-CreERT2; Rosa-LSL-Cas9-2A-eGFP (homozygous) mice. For organoid establishment, crypts were isolated from mouse small intestinal epithelium after washing and separation. The isolated crypts were resuspended in Matrigel (Corning) at a crypt density of 150-200 per 20 μl droplet. The droplets were inoculated into 48-well plates (Corning), using 250 μl of medium in each well. For the first two passages, cells were cultured in complete organoid medium supplemented with a Rho kinase inhibitor (Y-27632, R&D Systems). Organoids were divided every 5-7 days in a ratio of 1:5-1:6 by mechanical pipetting.

[0071] 1.2.3 Unicellular clones ESCs were trypsin-treated and counted. 500 cells were inoculated into a 15 cm dish (Sigma Aldrich). The ESCM was changed every two days. After growing colonies for 10 days, they were harvested into a 96U well plate (Thermo Fisher), trypsin-treated, and divided into a 96F well plate (Thermo Fisher). Cells were cultured until confluence and lysed overnight at 37°C in 75 μl of lauryl sarcosine lysis buffer. For amplification, 1 μl of lysate was used in a 25 μl PCR reaction (3 minutes at 95°C, [20 seconds at 95°C, 20 seconds at 65°C (-0.3°C / cycle), 30 seconds at 72°C] × 23, [20 seconds at 95°C, 20 seconds at 58°C, 30 seconds at 72°C] × 30, 3 minutes at 72°C, ∞ at 12°C).

[0072] 1.2.4 Retroviral vectors and ESC infection The CRISPR-StAR library was packaged into platinum-enriched E cells according to the manufacturer's recommendations. 300 million ESCs were infected with a 1:10 dilution of virus-containing supernatant in the presence of 2 μg / ml polyblen. 24 hours after infection, selection of infected cells was initiated using 1 μg / ml blastosidine and puromycin, respectively. To estimate the degree of infection multiplicity, 10,000 cells were plated in a 15 cm dish and selected with G418. For comparison, another 1,000 cells were plated and G418 selection was not performed. Colonies were counted on day 10.

[0073] 1.2.5 Cell Culture Screen ESCs were infected with a retroviral CRISPR-StAR vector and selected for blastosidine and puromycin resistance for 3 days. To mimic a bottleneck, cells were fully counted and inoculated in a library at densities of 1 cell / sgRNA (5870 cells), 4 cells / sgRNA, 16 cells / sgRNA, 64 cells / sgRNA, 256 cells / sgRNA, and 1024 cells / sgRNA. Cells were grown to equal density over 7 days. To induce recombination, ESCs were treated with 5 μM 4OH for 3 days. These were maintained for a further 14 days.

[0074] 1.2.6 Organoid Screen For screen preparation, organoids were expanded in 10 cm dishes (Corning). 50-55 drops were inoculated into each 10 cm dish, with each drop containing approximately 100 organoids, and a total of 10 10 cm dishes were used for the screen. 10 ml of complete medium was added to each dish and replaced with fresh medium every two days. To prepare organoids for viral infection, the organoids were first mechanically broken down into small fragments. After centrifugation (500 g x 5 min) to remove the supernatant (containing old Matrigel), the cells were resuspended in TrypLE (Gibco) and separated into aggregates of 5-8 cells at 37°C. The cells were centrifuged at 300 g for 3 minutes. After removing the supernatant, the cell pellet was resuspended in virus-containing medium and dispensed into 48-well plates. The plates were sealed with Parafilm and spinoculated at 37°C for 1 hour. After spinoculation, the Parafilm was removed and the plates were incubated at 37°C for 6 hours. The cells were then transferred to Eppendorf tubes and centrifuged (300g x 3 minutes). The cell pellet was resuspended in Matrigel and inoculated into 10cm dishes. After a 3-day recovery period, infected organoids were selected for blastosidine resistance at 1 μg / ml for 8 days. Subsequently, the organoids were isolated and the complete medium was replaced with 4OH for 6 hours. The organoids were then cultured in complete medium for 12 days without division. The medium was replaced every 3 days.

[0075] 1.2.7 DNA Collection and NGS Sample Preparation 60 million cells were collected per sample and lysed in SDS lysis buffer containing 1 mg / ml proteinase K and 0.1 mg / ml RNAse A. Genomic DNA was extracted with phenol and chloroform and precipitated with 1 volume of isopropanol. The incorporated sgRNA constructs were flanked by PacI restriction sites. Samples were digested with PacI for 48 hours and then co-digested with BbsI for the last 12 hours. Each sample was PCR amplified in 96 individual 50 μl reactions, each containing 1 μg of DNA (3 min at 95°C, [10 sec at 95°C, 20 sec at 59°C, 30 sec at 72°C] × 36, 3 min at 72°C, and infinity at 4°C). The forward primers were unique to each sample and contained a 6bp experimental index for demultiplexing after NGS (AATGATACGGCGACCACCGAGATCTACAC-NNNNNN-CGAGGGCCTATTTCCCATGATTCCTTC (SEQ ID NO: 17), where the 6bp NNNNNN sequence represents the specific experimental index used to demultiplex the sample after NGS). The reverse primers were the same for each sample. PCR products were purified and size-separated by agarose gel electrophoresis. Two recombinant products were excised separately, purified on a mini-elution column, and mixed in equal volumes. This sample was sequenced using an Illumina HiSeqV4 SR100 dual-index sequencing experiment. sgRNA was sequenced using custom read primers. To distinguish between active and inactive guides, the sequence downstream of the first lox site (TCAGCATAGC for active, TTTTTTT for inactive) was selected.

[0076] Example 2: Conceptual Overview Genetic screening suggests that genome editing can have three major effects: it may result in increased proliferation, decreased proliferation, or no effect on cells targeting specific sgRNAs. Increased proliferation leads to enrichment within the population. Decreased proliferation leads to deletion.

[0077] Pooled CRISPR screens are typically maintained at a complexity of 300–1,000 individually targeted cells per sgRNA. This allows a sufficient number of unique editing events to trigger significant changes in the population. However, this high level of complexity cannot always be maintained. This can happen when the system encounters bottlenecks caused by inefficient infection or limited cell number or differentiation, or when cells recover at different rates, reducing the library display. To illustrate this, we calculated the log2 multiple change (LFC) between the number of barcode reads before and after the CRISPR screen. The number of barcodes represents the number of differently transformed cells, i.e., the number of barcodes per guide represents the number of cells / sgRNA.

[0078] As complexity decreases, the distribution of LFCs widens because there are fewer barcodes present, and population changes have a greater impact. Further decreases in complexity result in a bimodal distribution with a second peak representing strong LFCs (Figure 1). This peak is due to the absence of guides on read 0. In analysis, these guides are mistaken for guides causing strong deletion indications, thus distorting the screening results. This means that, due to insufficient complexity, the number of guide readings before screening does not match the number of readings after screening, and conventional analysis fails.

[0079] The problem caused by insufficient library display at the bottleneck in CRISPR screens can be overcome by the present invention (shown in Figure 2).

[0080] Example 3: sgRNA construct Two sets of interwoven lox sites allow the CRISPR StAR system to produce two distinct recombinant products: inactive sgRNA or active sgRNA. The vector contains sgRNA (library), followed by two pairs of lox sites in the tracr region. Between the lox sites is a blastosidine selection cassette, which prevents premature activation, for example, due to Cre activity during viral packaging or a recombination event. Finally, this contains a set of random nucleotides that function as a unique molecular identifier (UMI). Recombination at the loxP site results in active sgRNA (Figure 3A), while recombination at the lox5171 site results in termination and exclusion of the tracr. As a result, the sgRNA becomes inactive (Figure 3B). The two recombination events are mutually exclusive.

[0081] This system allows for the comparison of active guides with inactive internal controls within the final population of the CRISPR screen. However, if the ratio of active to inactive recombinants is fairly similar, it is beneficial to compare the number of readings for the two recombinant products. In most cases, the ratio of loxP (active) to lox5171 (inactive) recombinants should be between 10:90 and 90:10.

[0082] The recombination probability between two loxP pairs depends on several factors, including the distance at the locus and the DNA structure (primary, secondary, and tertiary). Therefore, it is difficult to predict. Single-cell quantification of recombination probability revealed that the original construct (StAR1) yielded a recombination ratio of 33% active sgRNA and 66% inactive sgRNA. Such a ratio provides an ideal dynamic range and is therefore ideal if the screen is to track the relative enrichment of active sgRNA over inactive sgRNA. However, in the analysis of essential genes, it is also desirable to start with an equal ratio of active and inactive sgRNA, or to start with a bias towards active sgRNA. Therefore, StAR3 and StAR4 were developed by altering the relative distance, the primary sequence, and introducing one additional loxP site (Figure 4). In doing so, we successfully generated a series of constructs yielding various recombination ratios. active inactive StAR1 (SEQ ID NO: 18): 33% 66% StAR3 (SEQ ID NO: 19): 90% 10% StAR4 (SEQ ID NO: 20): 50% 50%

[0083] The ideal setup varies depending on the purpose of the experiment.

[0084] To determine how efficiently any pair of lox sites recombined, sgRNA-infected cells were treated with 4OH for 3 days and then inoculated at clonal density (Figure 5). At this point, recombination occurred, and these clones expressed active or inactive guides. To identify these, PCR was performed using primers that flanked the guide constructs. The recombination products were 580 bp for active and 542 bp for inactive. The frequency of each band size was counted. Most importantly, no non-recombined clones were found, confirming stable Cre expression in the cell line. The above recombination frequencies were determined. In the case of StAR1, out of 288 total clones, recombination produced 97 active sgRNAs and 172 inactive sgRNAs. Twenty-one double bands were found, which were due to either contaminated mixed clones or double infection. These were counted in both events.

[0085] Example 4: Cell Culture

[0086] 4.1 Experimental Design To confirm that CRISPR-StAR overcomes noise in the bottleneck screen, a controlled bottleneck was introduced in cell culture experiments. Thus, mouse embryonic stem cells stably incorporating a Cas9 expression cassette and a CreERT2 expression construct were infected with a retroviral sgRNA StAR1 type library of 5,870 sgRNAs targeting 1,245 genes (Table 1).

[0087] [Table 1]

[0088] 15% of the cells were infected to ensure single infection. After selection for viral integration, cells were counted and diluted to introduce a controlled bottleneck. Complexity decreased to 1 cell / sgRNA (5,870 cells), 4 cells / sgRNA, 16 cells / sgRNA, 64 cells / sgRNA, 256 cells / sgRNA, and 1,024 cells / sgRNA. Cells were grown to an isodensity of over 1,000 cells / sgRNA in 7 days. Subsequently, cells were treated with 4OH to induce Cre recombination, and the cells were maintained for a further 14 days. The experiment was performed with two independent replicas (Figure 6).

[0089] After 14 days, genomic DNA was extracted and digested in PacI using cleavage sites to flank the construct. Next, guide constructs were amplified from the fragmented genome via PCR using primers containing the experimental index and Illumina adapter for each sample, thus allowing direct sequencing of the PCR products. Both recombinant products were gel-extracted separately and mixed in a 1:1 ratio. This pool was then sequenced.

[0090] 4.2 Bioinformatics Pipeline After mapping the NGS reads, the active guide was bioinformatically distinguished from the inactive guide (active: TCAGCATAGC, inactive: TTTTTTT) using a 10 bp stretch immediately downstream of the first loxP site. After mixing the active and inactive recombinant products in a 1:1 ratio and sequencing, it was found that there were twice as many reads from the inactive guide as from the active guide, indicating better sequencing of the inactive construct. Nevertheless, this situation does not pose a disadvantage to the analysis.

[0091] Each cell was infected with a single guide construct. That is, every UMI represents one clone, and the number of UMIs per guide is equal to the number of cells per guide, which is a direct measure of the number of infected cells per guide. To confirm whether cell dilution in the demonstration experiment was sufficient, we calculated the median number of UMIs per inactive guide for the least complex sample (one cell per guide). However, we found a much larger number than the theoretical 1 UMI per guide. We hypothesized there were two reasons for this: firstly, most of these UMIs had only one or two reads, which could be due to base substitution errors in sequencing; secondly, when calculating the distribution of reads per UMI, we found a bimodal distribution. Looking at sgRNA-UMI combinations from the low read rates of this distribution, we were able to find the same sgRNA-UMI combinations with high read counts in different samples. This suggests index hopping, a known problem in Illumina-based sequencing, where the index between adjacent clusters is assigned to the wrong sample. In more complex samples, these issues can be ignored because there are many true UMIs per guide, and therefore the overall impact of these errors is negligible. Thus, this is only relevant to less complex samples (1-16 cells per sgRNA). Here, true reads have a clear distribution with a large number of reads, while errors have a distribution with a small number of reads.

[0092] To isolate true reads from errors, a threshold was defined for the local minimum of this bimodal read distribution in each less complex sample, and all reads below this threshold were discarded. Since the number of UMI reads in the active guide can represent the phenotype, only a cutoff was set for the inactive guide, and the sgRNA-UMI combinations of the active guide were mapped, which further removed datasets with non-existent UMIs.

[0093] Finally, to benchmark the performance of CRISPR-StAR against conventional CRISPR screen analysis, LFCs were calculated for both methods: active guide vs. conventional analysis on day 0, and active guide vs. inactive guide for CRISPR-StAR analysis (Figure 7).

[0094] 4.3 Benchmark Test To benchmark the performance of CRISPR-StAR compared to conventional screening methods, we calculated the Pearson coefficient between replicas, the area under the delta curve (dAUC), and the area under the receiver operating characteristic curve (AUROC).

[0095] 4.3.1 Correlation of Replicates To test the reproducibility of our results, we calculated the correlation coefficient between two biological replicas of essential and non-essential guides. To do this, we defined essential genes (red) using the same library of high complexity, with data from two independent screens performed on the same cell line. We calculated the median deletion for each guide and defined guides with an LFC of less than -3 as essential. Non-essential genes (blue), on the other hand, were defined as having the same number of deletions as essential guides from the same dataset. Next, we correlated the LFCs of the guides in the two independent replicas and determined the Pearson coefficient based on essential and non-essential genes. To better understand the data distribution, we calculated the density and ratio of essential and non-essential data for each replica (Figures 8 and 9, side density plots). Finally, we counted the number of sgRNAs present in each replica and the overlap between both replicas.

[0096] At a high complexity of 64–1,024 cells per sgRNA, good correlations between replicates were found in both conventional and CRISPR-StAR analyses. While the data distribution was slightly broader using conventional analysis than with CRISPR-StAR, essential replicas were clearly distinguishable from non-essential ones. The correlation coefficient ranged from 0.72–0.75 for conventional analysis and 0.80–0.84 for CRISPR-StAR (Figure 8). In this homogeneous system, 64 cells per sgRNA appears to be sufficient complexity for CRISPR screens using conventional analysis.

[0097] Using conventional, less complex analyses of 1–16 cells per sgRNA, we found increased spread of both essential and non-essential guides. With 4 and 1 cells per sgRNA sample, the data distribution is bimodal. This is due to sgRNAs with zero reads in one or both replicas, resulting in a stronger deletion compared to day 0. This deletion may be due to a guide-induced phenotype or the absence of guides in the final population. Guides can be lost, particularly in systems encountering bottlenecks. Conventional analyses cannot distinguish between missing guides and phenotypes. In contrast, using CRISPR StAR analysis, the abundance of active guides is compared to the abundance of inactive control guides in the final population. Thus, guides lost due to bottlenecks are excluded from the analysis. The resulting guide population is smaller, and LFCs are due to guide-induced phenotypes. As a result, in the least complex samples (one cell per sgRNA), the correlation decreased to 0.16 using conventional analysis, but was as high as the most complex samples at 0.83 with CRISPR StAR analysis (Figure 9).

[0098] In conclusion, using conventional analysis, we found that reproducibility decreases as complexity decreases. This is due to the spread of data caused by missing guides. Using CRISPR StAR, missing guides are removed, and only the current guides are considered. Therefore, the results show higher reproducibility even at lower complexity.

[0099] 4.3.2 dAUC Calculating the dAUC of a defined category within a population reveals how well members of each category are separated from one another. This was used to benchmark the performance of CRISPR StAR against conventional analysis for the separation of essential and non-essential guides. To this end, essential and non-essential guides were subsetted into new lists as defined above and ranked by LFC from most deficient to most abundant. Next, the cumulative percentage of the presence of each guide within the categories of the entire ranked list was calculated. In other words, when essential guides are scored, the essential curve rises. The same is true for non-essential guides. For guides to be effective, essential guides must be ranked at the top of the list, resulting in a sharp increase followed by a plateau where non-essential guides are not scored. Non-essential guides, on the other hand, are ranked at the bottom of the list, which is represented by a plateau followed by a sharp increase. Ideally, both categories would be expected to be clearly separated from each other. Thus, a better method shows better separation. To obtain an equivalent measure, the dAUC was calculated by subtracting the AUC of essential guides from the AUC of non-essential guides. The ideal score, when all essential elements are separated from non-essential elements, would be 0.5. Random samples would be diagonalized, with a dAUC score of 0.

[0100] The dAUC of CRISPR-StAR analysis is stable in the range of 0.45–0.47. Even in the least complex samples, the dAUC is 0.46 and 0.45, respectively. In contrast, using conventional analysis reduces complexity and makes it difficult to clearly separate essential and non-essential substances. As mentioned above, this is caused by the broad spread of both essential and non-essential substances (Figure 9). The dAUC drops to 0.14 and 0.09, respectively, in the least complex samples (Figure 10). Therefore, CRISPR-StAR analysis is superior to conventional analysis by clearly identifying essential substances and separating them from non-essential substances.

[0101] 4.3.3 AUROC In receiver operating characteristic (ROC) curves, true positive rates are compared to false positive rates. They quantify how well a method can classify data (in this case, guides) as essential or non-essential. Essential were defined as described above and classified as true positives. Similarly, non-essential items were classified as false positives. For true CRISPR StAR and conventional analysis, AUROC scores were calculated using the pROC package in R with a ranked list of guides by LFC. An ideal score would be 1, and a random score would be 0.5.

[0102] In conventional analysis, as complexity decreases, the AUROC decreases from 0.94 to 0.44, which is the same as the random score (Figure 11). Missing non-essential elements are not included in the analysis. This results in a large LFC, which incorrectly scores non-essential elements as essential. In contrast, in CRISPR StAR analysis, the AU-ROC remains between 0.91 and 0.95. Therefore, even at the lowest complexity, true positives can be clearly distinguished from false positives.

[0103] 4.4 Summary The performance of CRISPR StAR over conventional CRISPR screen analysis was benchmarked by calculating Pearson coefficients, dAUC, and AUROC. Using all three methods, CRISPR StAR was found to be significantly superior to conventional analysis, especially in the least complex samples, as complexity decreased (Figure 12).

[0104] In summary, the presented data confirms that CRISPR-StAR effectively overcomes the gene screen noise introduced by the loss of complexity after the bottleneck in population screening.

[0105] Example 5: Organoid Screen In homogeneous cell populations, conditions supporting high-resolution CRISPR screening can be easily controlled. In more heterogeneous systems such as organoids, this is a significant problem. To specifically test the effects of clonal heterogeneity in a model, we tested CRISPR-StAR in gut organoids. First, delivery of our retroviral library infects only crypt stem cells, which are a small subset of the overall cell population. Thus, infection of organoids is highly inefficient and usually becomes the first bottleneck that needs to be overcome. Second, clonal proliferation is highly heterogeneous.

[0106] Organoids containing CreERT2 and Cas9 transgenes were transduced with sgRNA libraries. These were selected for blastosidine resistance for 8 days, treated with 4OH-tamoxifen to induce Cre recombination, and maintained in culture for a further 12 days.

[0107] To estimate the complexity of infection, the median number of UMIs per guide was calculated. Similar to cell culture screens, a bimodal read distribution caused by index swapping was observed. This was handled in the same way as in cell culture screens: i.e., to separate true reads from errors, a local minimum of this bimodal read distribution was defined as a threshold, and all reads below this threshold were discarded. Since the number of UMI reads for active guides can represent the phenotype, only a cutoff was set for inactive guides, and the sgRNA-UMI combinations of active guides were mapped, which further removed datasets of non-existent UMIs. After the cutoff, it was found that infection occurred with a cell complexity of 30 per sgRNA.

[0108] Because UMIs on the guide construct allow tracking the clonal proliferation of individually marked cells, all UMIs within the same guide represent biological replicas. Therefore, we modified the dataset by splitting it into two groups according to the first letter of the UMI: one group of UMIs beginning with A or T, and the other group of UMIs beginning with C and G. We then used these two groups as biological replicas.

[0109] 5.1 Benchmark Testing To benchmark the performance of CRISPR-StAR in organoids compared to conventional screening methods, we calculated the Pearson coefficient between replicas based on UMI. Next, we analyzed the reproducibility of guides within a ranked list of guides by calculating the number of genes compared to the number of guides, and scored the correlation between two biological replicas as determined by UMI. Finally, we compared the hit lists in both types of analyses within the same two replicas.

[0110] 5.2 Correlation To compare the reproducibility of CRISPR StAR with conventional analysis, Pearson coefficients were calculated between these UMI-based biological replicas. To generate day 0 samples for conventional analysis, replicas of day 0 samples were obtained from demonstration screens, and the average number of reads for each guide was calculated. The same benchmark testing procedure as for cell culture screens could not be applied because the complete set of essential genes for organoids was unknown (Example 2.2). Instead, core essentials defined by Hart (Hart et al., Cell 163(6), 2015: 1515-1526), ​​which should be deleted in all cell types, were used.

[0111] The screening results using conventional analysis were found to have low reproducibility (R=0.27), while CRISPR-StAR analysis on the same dataset generated a more reproducible hit list (R=0.53). Overall, using conventional analysis results in a wider spread of data. In contrast, after identifying 557 missing guides that were excluded from the CRISPR StAR analysis due to a bottleneck, the CRISPR StAR analysis shows a very sharp signal (Figure 13a).

[0112] 5.3 Guide Reproducibility To test the reproducibility of the guides, a ranked list of guides was created using the MAGeCK algorithm. From this list, the average number of sgRNAs present per gene was calculated for all genes hit by each group of ranked guides. For example, if 15 genes were hit among the top 30 sgRNAs, the value was 2. For a random dataset, a value of 1 would be expected. While conventional analyses yield nearly random results, CRISPR-StAR demonstrates high reproducibility of scored genes (Figure 13b).

[0113] 5.4 Gene reproducibility Finally, for comparison at the gene level, we used MAGeCK with both analysis methods to create a combined guide and ranked list of genes. Not only were we able to call up top hits with higher p-values ​​compared to the conventional analysis, but the scored genes were more reproducible across replicas. Furthermore, using CRISPR-StAR analysis, we called up four of the top 10 deletion genes from both replicas, in contrast to the single common deletion gene found using the conventional analysis. These are hits we expect to find because they are essential or specific to organoid proliferation (Egfr, Itgb1, Top2a, Rpl14). Under the top five enriched genes, we found two common to both replicas (Nf2, Cdkn2a), whereas no common genes were found in the conventional analysis (Figure 13c). Additionally, the scored genes in each other's replicas showed high scores in the CRISPR-StAR analysis, but were quite distributed in the conventional analysis.

[0114] CRISPR-StAR is superior to conventional analyses because it can reliably identify screen hits, and we can conclude that it can yield reproducible results even in heterogeneous systems such as gut organoids.

[0115] Example 6: In vitro and in vivo screening 6.1 Materials

[0116] cell line Yumm1.7 450R melanoma cells (received from Obenauf Lab, IMP, Vienna). Lenti-X (Clontech 632180)

[0117] Cell culture medium Yumm1.7 450R melanoma cells: DMEM / F12 supplemented with 10% FCS (Gibco), 1% L-glutamine (Gibco), and 1% penicillin-streptomycin (Sigma). YUMM1.7 450R (Cas9-Cre ERT2The culture medium also contains puromycin (1 μg / ml, Invivogen).

[0118] Lenti-X cells: DMEM supplemented with 10% FCS (Gibco), 1% L-glutamine (Gibco), 1% penicillin-streptomycin (Sigma), 1% non-essential amino acids (NEAA, Gibco), and 1% sodium pyruvate (Sigma).

[0119] buffer solution 2X SDS lysis buffer: 10 mM Tris-HCl, pH 8 (Sigma Aldrich), 1% SDS (in-house), 10 mM EDTA (Sigma Aldrich), 100 mM NaCl (Sigma Aldrich), newly added 1 mg / ml proteinase K (New England Biolabs).

[0120] Primer FW_G_CrSc_2: AATGATACGGCGACCACCGAGATCTACACACCGAACGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 21) FW_G_CrSc_15: AATGATACGGCGACCACCGAGATCTACACCATTACCGAGGGCCTATTTCCCATGATTCCTTC (SEQ ID NO: 22) FW_G_CrSc_20: AATGATACGGCGACCACCGAGATCTACACCGTCATCGAGGGCCTATTTCCCATGATTCCTTC (Sequence ID 23) RV_G_CrSc: CAAGCAGAAGACGGCATACGAGATACCGTTGATGAGTAG (Sequence ID 24) NGS_U6: CGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCG (Sequence ID 25) NGS_customNextSeq_i2_primer: GAAGGAATCATGGGAAATAGGCCCTCG (SEQ ID NO: 26)

[0121] 6.2. Method 6.2.1 Creation of single-cell-derived clones expressing Cas9 and CreERT2

[0122] Yumm1.7 450R cells were generated using Cas9 and CreERT2 via in vivo and in vitro screening. First, cells were sequentially transduced with PX459 pSpCas9(BB)-2A-Puro and pMSCV-GFP-mir30-PGK-CreERT2. Bulk cell populations were selected for puromycin resistance, and single-cell clones were induced by single-cell fluorescence-activated cell sorter analysis (FACS). Subsequently, clones were tested for Cas9 function and leaky creERT2 expression using a CRISPR-Switch containing sgRNA for GFP (Chylinski et al, Nature Communications 10, 2019).

[0123] 6.2.2 Cloning of pooled libraries To create a lentiviral library containing StAR constructs using a drug-enhanced sgRNA library pool, 15,723 sgRNAs were PCR-amplified and cloned into StAR vectors by Golden Gate cloning. Subsequently, plasmids were electroporated into bacteria (Endura electrocompetent cells, Lucigen). After transformation, the bacteria were harvested in LB medium at 37°C for 1 hour, plated on LB agar plates containing ampicillin, and incubated overnight at 37°C. A 3,000-fold coverage of each sgRNA in the library was confirmed. Plasmid DNA was isolated and used to create lentiviral particles.

[0124] 6.2.3 In vitro screening StAR constructs containing drug-enhanced sgRNA library pools (157,23 sgRNAs) were packaged into Lenti-X cells according to the manufacturer's recommendations. Monoclonal YUMM1.7450R(Cas9-Cre ERT2Lentiviral particles were transduced into cells, followed by neomycin selection (geneticin G-418, 500 μg / ml, Gibco) for 4 days. The cells were divided into two groups: in vitro and in vivo screening. The in vitro cells were cultured and creERT2 recombination was induced with 4OH (0.5 μM) for 3 days. The cells were maintained for 21 days after induction.

[0125] 6.2.4 In vivo screening 50 μl (PBS: Matrigel) 1 × 10 6 Cells were subcutaneously injected into the flanks of female mice aged 6-12 weeks. Seven days after cell injection, creERT2 recombination was induced by intraperitoneal injection of 5 mg of tamoxifen per 30 g of mouse tissue. Tumor size was measured weekly, and when the tumor size reached 2 cm, it was recombined. 3 The mice were euthanized when they reached this point (6-13 days after tamoxifen injection).

[0126] 6.2.5 Genomic DNA Extraction and NGS Library Preparation Cells screened in vitro and collected on day 21 were lysed in lysis buffer at 55°C for 24 hours. Tumors collected from mice were lysed in 15-20 ml of lysis buffer at 55°C for 48-72 hours. Both lysed cells and tumors were treated with 0.1 mg / ml RNase A (Qiagen) at 37°C for 1 hour. gDNA was extracted with phenol and chloroform, followed by precipitation with isopropanol and EtOH. To fragment the DNA, samples were digested in BsmBI for 48 hours, and each sample was PCR-amplified in 48 separate 50 μl reactions with 1 μg of DNA per reaction (3 minutes at 95°C, [20 seconds at 95°C, 20 seconds at 59°C, 40 seconds at 72°C] × 33, 3 minutes at 72°C, and infinity at 4°C). Forward primers were unique to each sample and included a 6bp experimental index for remultiplexing after NGS (FW_G_CrSc_2, FW_G_CrSc_15, or FW_G_CrSc_20 primers in the material). Reverse primers were the same for each sample (RV_G_CrSc). PCR products were purified and size-separated by agarose gel electrophoresis. The two recombinant products were excised together and purified on a mini-elution column. This sample was sequenced using P2 SR100 sequencing on an Illumina NextSeq2000. sgRNA was sequenced with a custom read primer (read 1, NGS_U6). Active and inactive sgRNA constructs could be distinguished by analyzing the sequence of the 55bp vector after sgRNA. Another custom primer was used to determine the index (Index2, NGS_customNextSeq_i2_primer).

[0127] 6.3. Results and Discussion When performing in vivo screening, significant challenges must be overcome. Allogeneic graft screening has several technical bottlenecks, including infection and engraftment efficiency. Furthermore, heterogeneity arises both due to endogenous factors dependent on the cell (lineage) and exogenously depending on the cell's location in vivo (e.g., near blood vessels or in the middle of a tumor). These problems lead to unequal sgRNA representation, confusing conventional screening analyses, where comparing sgRNA on the first and last day of screening is inappropriate. One example of this is the loss of some sgRNA because cells containing these sgRNAs failed to engraft in mice. If sgRNA on the first and last day of screening were compared, these sgRNAs would be identified as deleted, thus defining the target gene as essential for tumor growth, which is a false positive result. CRISPR-StAR overcomes such challenges by comparing active and inactive sgRNAs present in cells that engrafted at the end of the screening. This example allows for further elucidation of different genetic dependencies under in vitro and in vivo conditions.

[0128] This example describes a comparison between in vivo and in vitro screens. Cas9 and Cre ERT2 The monoclonal melanoma cell line YUMM1.7 450R was used. Selected cells were screened in vitro or in vivo after viral transduction with a StAR construct containing a drug-enhanced sgRNA library pool (15,723 sgRNAs). 4OH was used to induce Cre recombination in vitro at the start of the screening, while recombination was induced in vivo by intraperitoneal injection of tamoxifen 10 days after cell injection. After a short in vivo screening period of 6–13 days (dependent on tumor growth rate), DNA was extracted from tumor and in vitro screened cells, subjected to next-generation sequencing, and bioinformatics analysis.

[0129] From this in vivo screen, reads could be obtained from inactive and active sgRNA constructs, indicating successful in vivo Cre recombination with the StAR vector. Active sgRNAs targeting essential genes were deleted compared to their corresponding inactive sgRNAs. The effects of sgRNAs in vitro and in vivo are calculated by summing the UMI reads for the same sgRNA, calculating the Log2 multiplier change (LFC) for each UMI, and then calculating the median total LFC for sgRNAs targeting the same gene (Figure 14). Negative control genes (shown in black) show no effect in either in vitro or in vivo. Most essential genes (shown in red) are deleted in either in vitro or in vivo. Dot size represents the number of UMIs per gene in the in vivo sample.

[0130] Example 7: CRISPR screening in mouse liver To perform in vivo CRISPR screening in endogenous tissue, it is necessary to selectively expand library-carrying cells in vivo, similar to in vitro selection of cells using antibiotics. This example demonstrates this expansion in the liver, where hepatocytes can proliferate and regenerate the liver after liver injury. In this case, only a small number of cells carrying the StAR library regrow in the liver, resulting in a sufficient number of cells to obtain the library and perform screening by comparing the ratio of active to inactive sgRNAs. Hepatic regrowth in fumarylacetoacetate hydroxylase (FAH) homozygous knockout (FAH- / -) mice using healthy hepatocytes is an established method for studying liver regeneration (Montini et al. (2002) Molecular Therapy, 6(6), 759-769; Wuestefeld et al. (2013) Cell, 153(2), 389-401; Zhu et al. (2019) Cell, 177(3), 608-621.e12). FAH metabolizes toxic fumarylacetoacetate (FAA) to fumarate and acetoacetate. Mice lacking functional FAH enzymes die from liver failure. However, FAH- / - mice can be maintained by nitisinone (NTBC) treatment. NTBC inhibits 4-hydroxyphenylpyruvate dioxygenase (HPD), an upstream enzyme in this metabolic pathway, thereby preventing the accumulation of FAA. Hepatocytes carrying the functional FAH gene can regrow into FAH- / - liver cells when NTBC is discontinued.

[0131] Figure 15 shows the Sleeping Beauty transposon containing an EGFP-P2A-FAH expression cassette under the control of an EF1a promoter containing a CRISPR-StAR construct. 25 μg of the transposon plasmid and 5 μg of the Sleeping Beauty transposase SB100X plasmid were injected into FAH- / - mice in 0.9% NaCl saline, and maintained with 1.8 mg of NTBC in 250 mL of drinking water. A volume equivalent to 10% of total body weight was injected into the tail vein over 5 seconds. The NTBC concentration decreased to 20% of the original concentration one day after injection. Seven days after injection, NTBC was completely removed from the drinking water. The StAR construct was cloned in the Sleeping Beauty transposon containing the FAH expression cassette. In this way, cells carrying the StAR construct can be regrown in the liver. The Sleeping Beauty transposon and transposase were delivered to the liver by hydrodynamic tail vein injection (Bell et al. (2007) Nature Protocols, 2(12), 3153-3165; Liu et al. (1999) Gene Therapy, 6(7), 1258-1266), and it was confirmed that cells carrying the StAR construct regrow in the liver after NTBC discontinuation. Therefore, CRISPR-StAR screening can be performed by regrowing healthy StAR-containing cells in the liver.

[0132] Another example of proliferating StAR-containing cells in the liver is inducing liver cancer. Here, the StAR construct is cloned in a Sleeping Beauty transposon with a KrasG12D expression cassette, a well-known cancer driver. We confirmed that StAR-containing cells proliferate in healthy livers. Figure 16 shows a Sleeping Beauty transposon with a KrasG12D-P2A-FAH expression cassette under the control of an EF1a promoter containing a CRISPR-StAR construct. WT mice were injected with 15 μg of the transposon plasmid and 3 μg of the Sleeping Beauty transposase SB100X plasmid in 0.9% NaCl saline. A volume equivalent to 10% of total body weight was injected into the tail vein over 5 seconds. To accelerate this expansion, the transposon was used in livers with conditionally deleted p53 (this is Alb-Cre in p53fl / fl mice). ERT2 This is achieved by activating [a specific substance] (which is injected into [a specific area] (Ju et al. (2016) International Journal of Cancer, 138(7), 1601-1608).

[0133] In vivo liver screening is performed in Cas9 and FAH- / - Alb-CreERT2 mice, or p53fl / fl mice. These examples demonstrate two methods for expanding CRISPR-StAR libraries in vivo before inducing recombination and performing the screening.

Claims

1. A nucleic acid comprising a sequence encoding a single guide RNA (sgRNA) of the CRISPR / Cas system, wherein the sgRNA sequence is interrupted by a guide disruption sequence franked by a first pair of recombinase recognition sites, and the sgRNA sequence further comprises a second pair of recombinase recognition sites having a different recombinase recognition sequence from the first pair of recombinase recognition sites, the guide disruption sequence is not franked by the second pair of recombinase recognition sites, and / or the second pair of recombinase recognition sites franks a portion of the sgRNA necessary to form an active sgRNA, and the sequences franked by the first and second recombinase recognition sites are duplicated, thereby resulting in deletion of the sequence between the recombinase recognition sites of the pair during recombination in the pair of recombinase recognition sites.

2. The nucleic acid according to claim 1, wherein one recombinase recognition site of the second pair of recombinase recognition sites is located between the first pair of recombinase recognition sites, and another recombinase recognition site of the second pair of recombinase recognition sites is located downstream of the first pair of recombinase recognition sites.

3. The nucleic acid according to claim 2, wherein one of the pair of second recombinase recognition sites is located downstream of the guide disruption sequence.

4. The nucleic acid according to claim 1, 2, or 3, wherein one of the first recombinase recognition sites is located within the loop region of the sgRNA sequence.

5. The nucleic acid according to claim 4, wherein the sgRNA sequence comprises a crRNA portion and a tracrRNA portion, and one of the first recombinase recognition sites is located within the crRNA-tracrRNA linker loop.

6. The nucleic acid according to any one of claims 1 to 5, wherein the guide disruption sequence includes a transcription disruption sequence or has a length sufficient to prevent folding into active sgRNA folding.

7. The nucleic acid according to any one of claims 1 to 6, wherein the first and second pairs of recombinase recognition sites are activated by the same recombinase enzyme.

8. The nucleic acid according to claim 7, wherein the first and second pairs of recombinase recognition sites are independently selected from lox sites, loxP, lox511, lox5171, lox2272, M2, M3, M7, M11, lox71, and lox66.

9. The nucleic acid according to any one of claims 1 to 8, further comprising a selection marker sequence.

10. The nucleic acid according to claim 9, wherein the selection marker sequence is an antibiotic selection marker sequence.

11. The nucleic acid according to claim 9 or 10, wherein the selection marker is located between pairs of recombinase recognition sites.

12. The nucleic acid according to claim 11, wherein the selection marker is located between both the first and second pairs of recombinase recognition sites.

13. An in vitro method for expressing the sgRNA of the CRISPR / Cas system upon recombinase stimulation, this method is A) To provide a plurality of nucleic acids according to any one of claims 1 to 12 to a plurality of cells, B) In a cell capable of activating a pair of first and second recombinase recognition sites, the method includes introducing or activating one or more recombinases. C) The above method, wherein the activation of the first pair of recombinase recognition sites and the second pair of recombinase recognition sites are competitive reactions, the activation of the first pair of recombinase recognition sites leads to the expression of active sgRNA, and the activation of the second pair of recombinase recognition sites inactivates the sgRNA sequence.

14. The method according to claim 13, wherein each of the plurality of cells has a single copy of the nucleic acid described in claims 1 to 12.

15. The method according to claim 13 or 14, wherein the cells are grown after step A) and before step B).

16. The method according to claim 15, wherein the cells are proliferated to a number of at least 250 cells per number of different sgRNA sequences in the plurality of nucleic acids described in any one of claims 1 to 12.

17. The method according to any one of claims 13 to 16, wherein cells having an inactive portion of the sgRNA sequence are identified in order to detect the presence of the sgRNA sequence.

18. The method according to any one of claims 13 to 17, wherein the cells further express a transgenic oncogene or have a suppressed tumor suppressor gene, The method further provides an excessive amount of the difference in tumorigenesis after activation in step C) compared with cells without activation in step C), thereby screening for the role of genes targeted by sgRNA during tumorigenesis, or further treatment of the cells with a candidate compound. The method further comprises providing an excessive amount of differences in cell activity or morphology after activation in step C) compared to cells without activation in step C), thereby screening for the activity of genes targeted by sgRNA under the influence of a candidate compound.

19. The method according to any one of claims 13 to 18, wherein the nucleic acid according to any one of claims 1 to 12 comprises a unique molecular identifier (UMI) sequence, the UMI being used to identify the same sgRNA in different cells.

20. The method according to any one of claims 13 to 19, wherein the cells contain a nucleic acid sequence for the expression of a recombinase.

21. The method according to claim 20, wherein the recombinase is Cre.

22. The method according to claim 20 or 21, wherein the nucleic acid sequence for the expression of the recombinase also includes a selection marker.

23. A cell comprising nucleic acid according to any one of claims 1 to 12.

24. i) The nucleic acid according to any one of claims 1 to 12, and ii) A kit comprising one or more nucleic acids for the expression of one or more recombinases capable of activating a pair of recombinase recognition sites of both nucleic acids according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Conditional gene knockout method based on CRISPR / Cas9 technology

    CN104404036A

  • Scalable recombinase cascade

    JP2019514377A

  • Conditional crispr sgrna expression

    WO2017158153A1