Methods and Systems for CRISPR Selection
By using viral vector libraries to infect Cas9-positive cells and sequence the read count of gRNAs in the CRISPR experiment, the probability of accidentally observing the number of gRNAs in the target region in the selected cells was calculated, which solved the problem of differences between CRISPR experiments and improved the analysis accuracy of gene screening data.
Patent Information
- Application Number
- CN202080035076.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-18
- Filing Date
- 2020-03-17
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-03-17
AI Technical Summary
The existing technology is difficult to effectively solve the problem of differences between CRISPR experiments, which leads to challenges in the calculation and analysis of gene screening data.
By infecting Cas9-positive cells using viral vector libraries, read counts of gRNA were obtained, and summed based on regions whose read counts exceeded the background threshold, the probability of accidentally observing the number of gRNAs in the target region in the selected cells was calculated to identify genes related to the phenotype.
This method can robustly handle differences in CRISPR experiments, improve the analytical accuracy and consistency of gene screening data, and help identify genes related to phenotypes.
Smart Images

Figure CN113841202B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 820,106, filed on March 18, 2019, which is hereby incorporated by reference in its entirety. Background of the Invention
[0003] Clustered regularly interspaced short palindromic repeats (CRISPR)-Cas9 technology has revolutionized genome engineering. In this system, a guide RNA (gRNA) directs the Cas9 nuclease to induce double-strand breaks in a target genomic region. The 5'-end of the gRNA includes a nucleotide sequence of approximately 20 nucleotides that is complementary to the target region. When the double-strand breaks are repaired by non-homologous end joining (NHEJ), insertions and deletions occur frequently, effectively knocking out the target genomic locus. The development of lentiviral delivery methods has enabled the generation of genome-scale CRISPR / Cas9 knockout libraries. These libraries allow for negative and positive selection screening in mammalian cell lines. In CRISPR / Cas9 knockout screening, each gene is targeted by several gRNAs, and mutant pools carrying different gene knockouts can be determined by high-throughput sequencing. CRISPR activation (CRISPRa) can also be used with a gRNA library, where the activated genes can be determined by high-throughput sequencing.
[0004] Genome-wide CRISPR / Cas9 knockout or gene activation technology is an effective gene perturbation screening technology. The goal is to identify gRNAs associated with a phenotype and thereby identify the corresponding affected genes. However, the data generated from these screens pose several challenges for computational analysis. CRISPR studies are typically conducted in a multi-replicate manner. CRISPR is susceptible to differences because each experiment may not use the same gRNA virus titer in the screening library, the lentiviral infection rate may vary between experiments, and the gRNAs may not target genes with the same efficiency in experiments. Therefore, even for cells with the same phenotype, the observed gRNA abundances are highly variable across experiments. Existing techniques rely on read counts to identify gRNAs associated with a phenotype, particularly using the mean and variance of normalized gRNA read counts to test whether there are significant differences in gRNA abundances between cells with or without a phenotype.
[0005] However, such techniques do not address the inter-experiment differences described above and instead assume a high degree of identity between pre-selection and post-selection experiments. These techniques cannot account for differences within a single CRISPR experiment and / or between CRISPR experiments.
[0006] Therefore, technical improvements to computing technology are needed to address the CRISPR discrimination problem when screening and identifying genes through positive and negative selections. SUMMARY OF THE INVENTION
[0007] It should be understood that the following general description and the following detailed description are both exemplary and explanatory and not restrictive.
[0008] In one embodiment, a method includes (A) infecting a first culture of cas9-positive cells with a library of viral vectors, the library comprising at least 3 guide RNAs (gRNAs) for cleaving target regions of DNA within the genome of the cells; sequencing the cells to obtain read counts for each of the gRNAs; summing (Σ) the number of corresponding gRNAs for each DNA target region for which the read count exceeds a background threshold, where Σ = n; and summing (Σ) the total number of gRNAs in all target regions for which the read count exceeds the background threshold, where Σ = N. The method includes (B) infecting a second culture of cas9-positive cells with the library of viral vectors; sorting the cells of the second culture as having a designated phenotype or not having the designated phenotype; selecting the cells having the designated phenotype and sequencing the selected cells to obtain post-selection read counts for each of the gRNAs; summing (Σ) the number of corresponding gRNAs for each DNA target region in the selected cells for which the post-selection read count exceeds the background threshold, where Σ = n'; and summing (Σ) the total number of gRNAs in all target regions of the selected cells for which the read count exceeds the threshold, where Σ = N'. The method includes (C) for a target region of DNA, calculating the probability of observing n' gRNAs for the target region by chance in the selected cells according to the formula where calculating the number of ways to select x objects from y objects; and for a target region of DNA containing a gene, calculating the probability of observing n' or more gRNAs for the gene by chance in the selected cells according to the formula where
[0009] In one embodiment, a method includes determining, for each of a plurality of target regions of DNA, the corresponding number of guide RNAs (gRNAs) (n) present in each of the plurality of target regions of DNA for which read counts exceed a background threshold, after infecting a first cell population with a vector comprising a library of at least 3 guide RNAs (gRNAs), determining, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA, the total number of gRNAs (N) present in the first cell population in all of the target regions of the plurality of target regions of DNA for which read counts exceed the background threshold, determining, for each of the plurality of target regions of DNA, the corresponding number of guide RNAs (gRNAs) (n') present in each of the plurality of target regions of DNA for which read counts exceed the background threshold, after infecting a second cell population with a vector comprising the library of at least 3 gRNAs, determining, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA, the total number of gRNAs (N') present in the second cell population in all of the target regions of the plurality of target regions of DNA for which read counts exceed the background threshold, for each target region of the plurality of target regions, determining the probability of observing n' gRNAs of the target region by chance in the selected cells, for a target region comprising a sequence of interest, determining the probability of observing n' or more gRNAs of the sequence of interest by chance in the selected cells based on the probability of observing n' gRNAs of the target region by chance in the selected cells, and identifying that the sequence of interest is positively selected based on the probability of observing n' or more gRNAs of the sequence of interest by chance in the selected cells.
[0010] Other advantages will be set forth in part in the description that follows, or may be learned by practice. The advantages will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings incorporated in and forming a part of this specification illustrate embodiments and, together with the description, serve to explain the principles of the method and system:
[0012] Figure 1 is an exemplary method of CRISPR positive selection;
[0013] Figure 2 shows another exemplary method;
[0014] Figure 3 shows an exemplary operating environment;
[0015] Figure 4Shows an exemplary data structure storing artificial read counts of four hypothetical genes, where each gene has three (A - C or D - F) or four (G - J or K - N) gRNAs, from which pre - selection and post - selection read counts for each gRNA are collected, and pre - selection n and post - selection n’ read counts greater than a threshold of 30 are identified and these read counts are summed together (N and N'), and probabilities are determined according to the formula shown;
[0016] Figure 5 Shows an exemplary data structure storing artificial data and artificial results of the disclosed method, according to which initially the number of gRNAs N is used and after 10 days of cell culture, the number of gRNAs N’ is provided. As shown, target region 1 or target region 2 has a number of gRNAs n, and the probability (probability represented as a p - value) that the number of gRNAs in the cells after 10 days of culture is n’ is provided;
[0017] Figure 6 Shows the experimental results involving approximately 21,000 genes (G#), where the read counts for each gRNA (g) from four pre - selection and three post - selection experiments and the summed read counts (Sum) are shown, and read counts greater than a threshold of 30 are identified (sum of presences), and these read counts are summed together (Sum);
[0018] Figure 7 Shows the sample gRNA libraries (Gecko A and Geck B) used in multiple parallel experiments (xp#) at three days (d03), six days (d06), and ten days (d10) of cell culture, after which Tau aggregation measured by FRET fluorescence is determined; the composition of Gecko A and B libraries is shown to include specific gene - targeting gRNAs, specific microRNA - targeting gRNAs, and non - targeting gRNAs and the approximate number of gRNAs for each target;
[0019] Figure 8 Shows the read counts of each gRNA in the Gecko A library before normalization in each experiment (xp#) at days 3, 6, and 10, according to which the gRNAs are used to direct the inactivation of the target;
[0020] Figure 9 Shows Figure 8 the normalization of the read counts of the Gecko A library in the experiment shown, the normalization is based on the median of the read counts, and day 10 is post - selection; and
[0021] Figure 10Shown, since the target may have different numbers of gRNAs, and at day 10 some samples had much more "present" gRNAs than other samples, another way to calculate the frequency of "present" is to calculate the probability of a gene being "present". The probability of five different genes screened at day 10 (after selection) was calculated using the Gecko A library gRNAs. Detailed implementation
[0022] Before disclosing and describing the methods and systems of the present invention, it should be understood that the methods and systems are not limited to specific methods, specific components or specific detailed implementations. It should also be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to be limiting.
[0023] As used in the specification and the appended claims, unless the context clearly dictates otherwise, the singular forms of the words "a", "an" and "the" also include plural referents.
[0024] As used herein, the terms "probe" and "guide RNA (gRNA)" and "guide" are used interchangeably. In one embodiment, the gRNA can also be provided in the form of DNA encoding the gRNA.
[0025] As used herein, "Cas protein" can be a wild-type protein (i.e., those existing in nature), a modified Cas protein (i.e., a Cas protein variant), or a fragment of a wild-type or modified Cas protein. In terms of the catalytic activity of a wild-type or modified Cas protein, the Cas protein can also be an active variant or fragment.
[0026] In a first aspect, the present disclosure features methods for identifying genes or gene products, e.g., methods for regulating the expression of other genes or gene products. For example, these methods can be used to demonstrate positive selection after perturbation using CRISPR guides.
[0027] In one embodiment (as Figure 1As shown), the method includes the following steps: infecting a first culture 110 of cas9-positive cells with a library of viral vectors, the library comprising at least 3 guide RNAs for cleaving target regions of DNA within the genome of the cells, sequencing the cells to obtain read counts 120 for each of the gRNAs, summing (Σ) (Σ = n) the number of corresponding gRNAs for each DNA target region for which the read count exceeds a background threshold and summing (Σ) (Σ = N) the total number of gRNAs in all target regions for which the read count exceeds the background threshold 130, infecting a second culture 140 of cas9-positive cells with the library of viral vectors, classifying the cells of the second culture as having a specified phenotype or not having a specified phenotype 150, selecting the cells having the specified phenotype and sequencing the selected cells to obtain post-selection read counts 160 for each of the gRNAs, summing (Σ) (Σ = n') the number of corresponding gRNAs for each DNA target region in the selected cells for which the post-selection read count exceeds the background threshold, and summing (Σ) (Σ = N') the total number of gRNAs in all target regions of the selected cells for which the read count exceeds the threshold 170, for a target region of DNA, calculating the probability of randomly observing n' gRNAs for the target region in the selected cells according to the following formula
[0028]
[0029] where calculating the number of ways to select x objects from y objects, and for a target region of DNA containing a gene, calculating the probability of randomly observing n' or more gRNAs for the gene in the selected cells according to the formula 190.
[0030] In one embodiment (also as Figure 1As shown, the method includes the following steps: infecting a first culture 110 of cas9-positive cells with a library of viral vectors, the library comprising at least 3 guide RNAs (gRNAs) for enhancing transcription of a target region of DNA within the genome of the cells, sequencing the cells to obtain read counts 120 for each of the gRNAs, summing (Σ) (Σ=n) the number of corresponding gRNAs for each DNA target region for which the read count exceeds a background threshold and summing (Σ) (Σ=N) the total number of gRNAs in all target regions for which the read count exceeds the background threshold 130, infecting a second culture of cas9-positive cells with the library of viral vectors 140, classifying the cells of the second culture as having a specified phenotype or not having the specified phenotype 150, selecting the cells having the specified phenotype and sequencing the selected cells to obtain post-selection read counts 160 for each of the gRNAs, summing (Σ) (Σ=n') the number of corresponding gRNAs for each DNA target region in the selected cells for which the post-selection read count exceeds the background threshold, and summing (Σ) (Σ=N') the total number of gRNAs in all target regions of the selected cells for which the read count exceeds the threshold 170, for a target region of DNA, calculating the probability of randomly observing n' gRNAs for the target region in the selected cells
[0031]
[0032] where calculating the number of ways to select x objects from y objects, and for a target region of DNA containing a gene, calculating the probability of randomly observing n' or more gRNAs for the gene in the selected cells according to the formula 190.
[0033] In one embodiment, the first culture can be infected according to CRISPR technology. In one embodiment, a first culture of cas9-positive cells is infected with a library of viral vectors (e.g., Figure 2 ; 201) and a second culture of cas9-positive cells is infected with a library of viral vectors (e.g., Figure 2; 205) The second culture of cas9-positive cells infected) includes the use of a CRISPR gRNA library. The CRISPR gRNA library can be a knockout library, including, for example, a genome-wide gRNA knockout library that includes one or more gRNAs (e.g., sgRNAs) targeting each gene in the genome, where the genome can be any type of genome. In some embodiments, the gRNA library can include a pooled library. Non-limiting examples of pooled libraries include the Genome-scale CRISPR Knockout (GeCKO) library. See, for example, Shalem O et al. (2014) Science 343:84-7 and Sanjana NE et al. (2014) Nat. Methods 11:783-4. The gRNAs in the library can target any number of target regions (e.g., genes) in the DNA. For example, the gRNAs can target about 50 or more genes, about 100 or more genes, about 200 or more genes, about 300 or more genes, about 400 or more genes, about 500 or more genes, about 1000 or more genes, about 2000 or more genes, about 3000 or more genes, about 4000 or more genes, about 5000 or more genes, about 10000 or more genes, or about 20000 or more genes. In some libraries, the gRNAs can be selected to target genes in a specific signaling pathway. The gRNA library can be administered using a wider range of multiplicities of infection (MOI). In some aspects, a lower MOI can be used to facilitate infection to obtain one gRNA per cell.
[0034] In one embodiment, the Cas positive cells can comprise a Cas protein for cleaving a target region of DNA or a Cas protein for regulating transcription (e.g., enhancing or inhibiting transcription). The Cas protein for cleaving a target region of DNA can comprise an RNA binding domain and a nuclease domain. The Cas protein for regulating transcription is inactivated so that it no longer has nuclease activity. The inactivated Cas (e.g., dCas-9) can be fused with a transcriptional activator or a transcriptional repressor. Thus, Cas-9 positive cells comprising Cas-9 with wild-type activity or inactivated Cas-9 are disclosed in one embodiment. The inactivated Cas-9 can be fused, for example, with at least one transcriptional activation domain. One or more gRNAs can bind to a target sequence upstream of the transcription start site of a gene of interest, rather than cleaving the DNA, and dCas-9 and the transcriptional regulator can play a role in activating the transcription of the gene of interest or inhibiting the transcription of the gene of interest. If dCas-9 is fused with one or more transcriptional activators, the one or more transcriptional activators will recruit transcription factors to the transcription start site of the gene of interest, thereby activating or upregulating transcription. If dCas-9 is fused with a transcriptional repressor, then transcription will be inhibited, downregulated, or suppressed.
[0035] In one embodiment, the target region of DNA can be a gene. In one embodiment, the target region of DNA can be a promoter region or a regulatory element region of a gene. In one embodiment, the target region of DNA regulates a downstream gene or protein. In one embodiment, regulating a downstream gene or protein includes activation or inhibition of the downstream gene or protein.
[0036] In some embodiments, the cells comprise a selectable marker system. For example, as disclosed in the methods herein, the cas-9 positive cells of the first culture can be modified to contain one or more selectable markers. The selectable marker system can involve one or more marker proteins fused or linked to a selectable marker that is only activated when one or more proteins are regulated. The "one or more marker proteins" can be any protein that can be regulated, where being regulated means that the marker protein changes shape, binds to one or more other proteins or nucleic acids, changes in activity, or changes in expression level. The one or more marker proteins can be regulated by a gene targeted by one or more gRNAs, so if the gene is cleaved by Cas9, the one or more marker proteins are not regulated, resulting in the selectable marker not being activated. Thus, cells without an activated selectable marker can be selected as those cells containing the gene that regulates one or more marker proteins, which are fused or linked to the selectable marker in the cell. For example, Tau can be fused or linked to CFP. Another Tau protein can be fused or linked to YFP. When the tau protein aggregates, the light reaching CFP will be emitted as blue light, which excites YFP, and YFP in turn emits yellow light. If the tau protein does not aggregate, the blue light emitted by CFP cannot excite the YFP of the other tau protein, so no yellow light will be emitted. Thus, if the gene regulating tau is targeted and thus cleaved by one or more gRNAs, no yellow light will be emitted. Finally, the gene can be identified as the gene regulating Tau (e.g., causing Tau aggregation). In CRISPRa, if the gRNA binds to the target region, the downstream gene will be activated, so cells with overexpression or excess of the selectable marker can be selected. For example, using the above Tau protein, an increase in the amount of yellow light compared to cells without gRNA can indicate the gene regulating Tau. Any number of proteins present in known pathways (especially disease pathways) can be used as marker proteins to identify the gene regulating the marker protein and can ultimately be found to be related to a specific disease.
[0037] In one embodiment, a cell population can be sorted and selected (e.g., by phenotype) (e.g., Figure 2; 206), which can be carried out for a period of time after initial infection, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 days, etc. after infection. In one embodiment, multiple types of screening / selection mechanisms can be used to exploit the selectable markers contained in the cell set. In one embodiment, the selection mechanism includes one or more of the following: exposing the second cell population to a drug, or exposing the second cell population to a substance that recognizes protein activity or expression level. In one embodiment, selecting cells with a specified phenotype includes sorting the cells according to one or more selectable markers. In one embodiment, the viability of the cells can be used for selection.
[0038] In one embodiment, classifying the cells of the second culture as having a specified phenotype or not having a specified phenotype includes identifying the presence or absence of the specified phenotype in the cells. Classifying the cells of the second culture as having a specified phenotype or not having a specified phenotype can include applying a selection mechanism to the second cell culture. In one embodiment, once the phenotype is identified, cells with the specified phenotype (or without the specified phenotype) can be selected. A phenotype is any observable characteristic or functional effect that can be measured in an assay, such as changes in cell growth, proliferation, morphology, enzyme function, signal transduction, expression pattern, downstream expression pattern, reporter gene activation, hormone release, growth factor release, neurotransmitter release, ligand binding, apoptosis, and product formation. In one embodiment, the specified phenotype can be fluorescence or cell survival.
[0039] Cells can be modified to convey phenotypes that can be directly selected, for example, by marked genomic integration or by the presence of intracellular markers that are not integrated into the genome. As used herein, a "marker" most commonly refers to a biological feature or trait that, when present in a cell (e.g., expressed), results in an attribute or phenotype that renders the cell visible or identifies the cell as containing the marker. Many types of markers are commonly used and can be, for example, visual markers (such as chromogenic, e.g., lacZ complementation (β-galactosidase), or fluorescent, e.g., expression of green fluorescent protein (GFP) or a GFP fusion protein), RFP, BFP, luciferase, β-galactosidase, enhanced green fluorescent protein (eGFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), blue fluorescent protein (BFP), enhanced blue fluorescent protein (eBFP), DsRed, ZsGreen, MmGFP, mPlum, mCherry, tdTomato, mStrawberry, J-Red, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, Cerulean, T-Sapphire, and alkaline phosphatase, phenotypic markers (growth rate, cell morphology, colony color or colony morphology, temperature sensitivity), auxotrophic markers (growth requirements), antibiotic sensitivities and resistances, molecular markers (such as biomolecules distinguishable by antigen sensitivity (e.g., blood group antigens and histocompatibility markers), cell surface markers (e.g., H2KK), enzyme markers, and nucleic acid markers such as restriction fragment length polymorphisms (RFLP), single nucleotide polymorphisms (SNP), and various other amplifiable genetic polymorphisms). Thus, for example, the one or more selectable markers can be a detectable enzyme such as β-galactosidase or luciferase
[0040] "Selective marker" or "screening marker" or "positive selection marker" refers to a marker that, when present in a cell, confers an attribute or phenotype that allows for the selection or isolation of those cells from other cells that do not express the selective marker trait. Many genes can be used as selective markers, for example genes encoding drug resistance or auxotrophic rescue are well known. For example, kanamycin (neomycin) resistance can be used as a trait to select bacteria that have taken up a plasmid carrying a gene encoding bacterial kanamycin resistance (e.g., the enzyme neomycin phosphotransferase II). When the culture is treated with neomycin or a similar antibiotic, untransfected cells will eventually die. This cell population can be used for drug screening to identify genes that confer drug resistance. Cells can be treated with the drug of interest, and the enriched gRNAs are associated with genes that confer drug resistance when mutated. Resistance screening against viral or bacterial pathogens can be used to identify genes that prevent infection or pathogen replication. Similar to drug resistance screening, survival after pathogen exposure provides a strong selection. In cancer, negative selection CRISPR screening can identify "oncogene addiction" in specific cancer subtypes, which can provide a basis for molecular targeted therapies. For developmental studies, screening in human and mouse pluripotent cells can precisely map the genes required for pluripotency or differentiation into different cell types.
[0041] In one embodiment, the cell contains a selective marker system. The selective marker system can involve a fluorescent protein FRET biosensor. For example, Tau can be fused or linked to CFP. Another Tau protein can be fused or linked to YFP. When the tau protein aggregates, the light reaching CFP will be emitted as blue light, which will then excite YFP, and YFP in turn will emit yellow light. If the tau protein does not aggregate, the blue light emitted by CFP cannot excite the YFP of another tau protein, and thus no yellow light will be emitted. Therefore, if the gene regulating tau is targeted by one or more gRNAs and thus cleaved, no yellow light will be emitted. Finally, the gene can be identified as a gene that regulates Tau (e.g., causes Tau aggregation). In CRISPRa, if the gRNA binds to the target region, downstream genes will be activated, so cells with overexpression or excess of the selective marker can be selected. For example, using the above Tau protein FRET biosensor, an increase in the amount of yellow light compared to cells without gRNA can indicate a gene that regulates Tau. Any number of proteins present in a known pathway (especially a disease pathway) can be used as a marker protein to identify the gene that regulates the marker protein and ultimately be found to be related to a specific disease pathway.
[0042] In one embodiment, the cells of the first culture and the selected cells of the second culture can be sequenced (e.g., Figure 2; 202 and 207). The cells of the first culture and the selected cells of the second culture can be sequenced at different times after infection. The cells of the first culture and the selected cells of the second culture can be sequenced at different times after infection to generate read counts (for the cells of the first culture) before selection and read counts (for the selected cells of the second culture) after selection. For example, the cells of the first culture can be sequenced 3 days after infection, and the selected cells of the second culture can be sequenced 10 days after infection. Any available sequencing technology (such as NGS) can be used to sequence the cells. The nucleic acids of the cells can be sequenced to generate sequence data. The sequence data can include read counts. The sequence data can include read counts of one or more gRNAs of the library. Sequencing the cells can generate sequence data that can be stored in a data structure. The data structure can contain one or more nucleic acid sequences and / or sample identifiers.
[0043] The read counts generated by sequencing are affected by traditional biases present in CRISPR screening assays, including frequent dropout of duplicates, differences in gRNA knockout efficiency, and differences in read count distribution. These biases lead to poor results when analyzing read counts according to the negative binomial method, log2 ratio method, and paired t-test method. Such methods require a certain degree of homogeneity / consistency between gRNAs of the same gene and between repetitive sequences. These existing methods cannot handle the large differences between gRNAs and gene repeats of the same gene, which may be due to different infection efficiencies, different gene editing efficiencies, initial viral counts in the screening library, and the presence of other guides with the same phenotype. These biases may exist within a single CRISPR experiment and between multiple CRISPR experiments. The steps currently described for determining the total number of gRNAs for each target region and the total number of gRNAs in all target regions are shown to be robust to large differences in read counts. Since one embodiment of the disclosed method is based on the positive occurrence of guides for each gene in a single experiment rather than the exact read counts of each guide, it offers advantages over existing methods.
[0044] The read counts can be normalized. For example, the read counts from different samples can be median-normalized to adjust for the effects of library size and read count distribution. In one embodiment, assume that N CRISPR / Cas9 gene knockout screening experiments are performed on a set of M gRNAs, and the read count of gRNA i in experiment j is x ij , where 1 ≤ i ≤ M and 1 ≤ j ≤ N. Since the sequencing depth (or library size) may vary between experiments, the read counts can be adjusted by applying the median ratio method to all experiments. In one embodiment, the adjusted read count x' ijIt can be calculated according to the following formula:
[0045]
[0046] where S is the median of for j = 1 to M. In another embodiment, the adjusted read count x' ij can be calculated as the rounded value of x ij / s j where s j is the size factor in experiment j and is calculated as the median of all size factors calculated from single sgRNA read counts:
[0047]
[0048] where x^i is the geometric mean of the read counts of gRNA i:
[0049] Alternatively, the read counts for each gRNA and each gene can be normalized using counts per million, total counts, or size factor normalization. See Anders, S. and Huber, W. (2010) Differential expression analysis for sequence count data. Genome Biol., 11, R106, which is incorporated herein by reference in its entirety.
[0050] In one embodiment, the disclosed method uses sequence data to determine the sum (Σ) of the corresponding number of gRNAs (n') for each DNA target region in selected cells where the read count exceeds a background threshold after selection (e.g., Figure 2 ; 203, 208). The method can also include determining the total number (N') of gRNAs in all target regions of selected cells where the read count exceeds the background threshold (e.g., Figure 2 ; 204, 209).
[0051] Thus, the disclosed method can identify positive occurrences of gRNAs from the library in the sequences of the selected cells. Sequence data including read counts can be analyzed to determine the "sum of presences" or n' after selection of the target region (e.g., gene) of each DNA. The "sum of presences", or the corresponding number of gRNAs present in the target region of each DNA, can be determined by comparing the individual read counts of each gRNA to a background threshold. Determining the corresponding number of gRNAs in each DNA target region whose read counts exceed the background threshold can be performed by a computing device. The background threshold can be any value sufficient to reduce background noise in the sequence data. For example, 30 can be used as the background threshold. Thus, the "sum of presences" represents the number of gRNAs present (read counts exceeding the background threshold), in contrast to the number of gRNAs present represented by the read counts.
[0052] In one embodiment, the steps can be repeated any number of times, the steps including: infecting a second culture of cas9-positive cells with the library of viral vectors; sorting the cells of the second culture as having a designated phenotype or not having a designated phenotype; selecting the cells having the designated phenotype and sequencing the selected cells to obtain a post-selection read count for each of the gRNAs; summing (Σ) the corresponding gRNA numbers for each DNA target region in the selected cells whose post-selection read counts exceed the background threshold, where Σ = n'; and summing (Σ) the total number of gRNAs in all target regions of the selected cells whose read counts exceed the threshold, where Σ = N'.
[0053] In one embodiment, the disclosed method further includes identifying that a target region is under positive selection based on the probability of observing n' or more gRNAs for a gene incidentally in the selected cells (e.g., Figure 2 ; 211). In some embodiments, identifying that a target region is under positive selection includes determining that the probability of observing n' or more gRNAs of the sequence of interest incidentally in the selected cells meets a threshold.
[0054] In some embodiments, the disclosed methods further include identifying a target region as a modifier of a second gene based on the probability of observing n' or more gRNAs of a gene incidentally in a selected cell. In some embodiments, the disclosed methods further include identifying a target region as a therapeutic target based on the probability of observing n' or more gRNAs of a gene incidentally in a selected cell. In some embodiments, the disclosed methods further include identifying a target region associated with a specified phenotype based on the probability of observing n' or more gRNAs of a gene incidentally in a selected cell. In some embodiments, the disclosed methods further include identifying a target region as exhibiting a protective effect based on the probability of observing n' or more gRNAs of a gene incidentally in a selected cell. Accordingly, the disclosed methods can aid in identifying target regions (e.g., genes) involved in disease pathways, involved in the regulation of one or more other genes / proteins, and / or otherwise associated with a phenotype. If the regulation of gene expression of a candidate gene results in a change in a selected phenotype, a region of DNA (e.g., the candidate gene) may be "associated" with the selected phenotype.
[0055] In one embodiment, the disclosed methods further include determining an enrichment score for each gRNA. In some embodiments, determining an enrichment score for each gRNA includes evaluating N / N'. An enrichment score can be determined to exceed a threshold. The threshold can be relative to other enrichment scores of other target regions. Target regions with high enrichment scores and low probability of the gRNA being present incidentally can be used to identify the target region as being associated with a phenotype.
[0056] In one exemplary embodiment, the method and system can be implemented on a computer 301 as shown and described below. Similarly, the method and system can utilize one or more computers to perform one or more functions at one or more locations. Figure 3 is a block diagram showing an exemplary operating environment for performing the method. This exemplary operating environment is only an example of an operating environment and is not intended to impose any limitation on the scope of use or functionality of the operating environment structure. Nor should the operating environment be construed as having any dependency or requirement on any component or combination of components shown in the exemplary operating environment. Figure 3 The present method and system can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of computing systems, environments, and / or configurations that can be adapted to be used with the system and method include, but are not limited to, personal computers, server computers, laptop computer devices, and multiprocessor systems. Additional examples include set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like.
[0057] The present method and system can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of computing systems, environments, and / or configurations that can be adapted to be used with the system and method include, but are not limited to, personal computers, server computers, laptop computer devices, and multiprocessor systems. Additional examples include set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like.
[0058] The processing of the described method and system can be performed by software components. The system and method can be described in the general context of computer-executable instructions, such as program modules, executed by one or more computers or other devices. Generally, program modules include computer code, routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The method can also be practiced in a distributed computing environment according to a grid, where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including memory storage devices.
[0059] In addition, the system and method can be implemented by a computing device in the form of a computer 301. The components of the computer 301 can include, but are not limited to, one or more processors 303, a system memory 312, and a system bus 313 that couples the various system components including the one or more processors 303 to the system memory 312. The system can utilize parallel computing.
[0060] The system bus 313 represents one or more of several possible types of bus structures, including a memory bus or memory controller using any of the various bus structures, a peripheral bus, an accelerated graphics port, or a local bus. The bus 313 and all buses specified in this specification can also be implemented by wired or wireless network connections, and each subsystem, including one or more processors 303, a mass storage device 304, an operating system 305, software 306, data 307, a network adapter 308, a system memory 312, an input / output interface 310, a display adapter 309, a display device 311, and a human-machine interface 302, can be included within one or more remote computing devices 314a, b, c located at physically separate locations, and the remote computing devices are connected by this form of bus, effectively implementing a fully distributed system.
[0061] The computer 301 generally includes various computer-readable media. Exemplary readable media can be any available media accessible by the computer 301, and includes, by way of example and not limitation, volatile and non-volatile media, removable and non-removable media. The system memory 312 includes computer-readable media in the form of volatile memory, such as random access memory (RAM), and / or computer-readable media in the form of non-volatile memory, such as read-only memory (ROM). The system memory 312 typically contains data such as data 307 and / or program modules such as an operating system 305 and software 306, which can be immediately accessed by and / or currently operated on by one or more processors 303.
[0062] In another embodiment, computer 301 may also include other removable / non-removable, volatile / non-volatile computer storage media. By way of example, Figure 3 a mass storage device 304 is shown, which may provide non-volatile storage of computer code, computer-readable instructions, data structures, program modules, and other data for computer 301. By way of example and not limitation, the mass storage device 304 may be a hard disk, a removable magnetic disk, a removable optical disk, a magnetic tape cartridge, or other magnetic storage device, a flash memory card, a CD-ROM, a digital versatile disk (DVD), or other optical memory, a random access memory (RAM), a read-only memory (ROM), and / or an electrically erasable programmable read-only memory (EEPROM).
[0063] Optionally, any number of program modules may be stored on the mass storage device 304, including, for example, an operating system 305 and software 306. Each of the operating system 305 and software 306 (or some combination thereof) may include programming elements and software 306. Data 307 may also be stored on the mass storage device 304. Data 307 may be stored in any one or more databases. Examples of such databases include Access, SQL Server, and / or The databases may be centralized or distributed across multiple systems. Data 307 may include sequencing data. The sequencing data may include sequencing read data (e.g., read counts). Computer 301 may receive, for example, first read count data generated at steps 120 and 202 respectively Figure 1 and Figure 2 Computer 301 may receive, for example, second read count data generated at steps 160 and 207 respectively Figure 1 and Figure 2
[0064] In another embodiment, a user may input commands and information into computer 301 via an input device (not shown). Examples of such input devices include, but are not limited to, a keyboard, a pointing device (e.g., a "mouse"), a microphone, a joystick, a scanner, a haptic input device (e.g., a glove), and / or other body coverings, etc. These and other input devices may be connected to one or more processors 303 via a human-machine interface 302, which is coupled to system bus 313, but may be connected via other interfaces and bus structures such as a parallel port, a game port, an IEEE 1394 port (also known as a Firewire port), a serial port, or a universal serial bus (USB).
[0065] In yet another embodiment, the display device 311 may also be connected to the system bus 313 via an interface such as a display adapter 309. It is contemplated that the computer 301 may have more than one display adapter 309, and the computer 301 may have more than one display device 311. For example, the display device may be a monitor, an LCD (Liquid Crystal Display), or a projector. In addition to the display device 311, other output peripheral devices may include components such as speakers (not shown) and printers (not shown), which may be connected to the computer 301 via the input / output interface 310. Any step and / or result of the method may be output to the output device in any form. Such output may be any form of visual reproduction, including but not limited to text, graphics, animation, audio, and / or tactile, etc. The display 311 and the computer 301 may be part of one device or separate devices.
[0066] The computer 301 may operate in a networked environment using logical connections to one or more remote computing devices 314a, b, c. For example, the remote computing devices may be personal computers, laptop computers, smart phones, servers, routers, network computers, peer devices, or other common network nodes, etc. The logical connection between the computer 301 and the remote computing devices 314a, b, c may be made via a network 315, such as a local area network (LAN) and / or a general wide area network (WAN). Such network connections may be made through a network adapter 308. The network adapter 308 may be implemented in wired and wireless environments. In one embodiment, the system memory 312 may store one or more objects that are accessible by the one or more remote computing devices 314a, b, c via the network 315. Thus, the computer 301 may function as a cloud-based object memory. In another embodiment, one or more of the one or more remote computing devices 314a, b, c may store one or more objects that are accessible by the computer 301 and / or another of the one or more remote computing devices 314a, b, c. Thus, the one or more remote computing devices 314a, b, c may also function as a cloud-based object memory.
[0067] For illustrative purposes, the present disclosure illustrates application programs and other executable program components, such as operating system 305, in discrete blocks, but it should be appreciated that such programs and components reside in different storage components of computing device 301 at different times and are executed by one or more processors 303 of the computer. In one embodiment, at least a portion of software 306 and / or data 307 may be stored on and / or executed on one or more of computing device 301, remote computing devices 314a, b, c, and / or combinations thereof. Thus, software 306 and / or data 307 may operate in a cloud computing environment, whereby access to software 306 and / or data 307 may be performed via network 315 (e.g., the Internet). Additionally, in one embodiment, data 307 may be synchronized on one or more of computing device 301, remote computing devices 314a, b, c, and / or combinations thereof.
[0068] Implementations of software 306 may be stored on or transmitted via some form of computer-readable medium. Any of the methods may be performed by computer-readable instructions included on a computer-readable medium. A computer-readable medium may be any available medium that can be accessed by a computer. By way of example and not limitation, computer-readable media may include "computer storage media" and "communication media". "Computer storage media" includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. Exemplary computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVDs) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer.
[0069] The software 306 can be configured to perform some or all of the steps of the methods disclosed herein. In one embodiment, the software 306 can be configured to determine, for each of a plurality of target regions of DNA, the corresponding number of guide RNAs (gRNAs) (n) present in each of the plurality of target regions of DNA for which read counts exceed a background threshold, based on sequencing of a first cell population after infection with a vector containing a library of at least 3 guide RNAs (gRNAs); determine the total number of gRNAs (N) present in the first cell population in all of the target regions of DNA for which read counts exceed the background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA; determine, for each of the plurality of target regions of DNA, the corresponding number of guide RNAs (gRNAs) (n') present in each of the plurality of target regions of DNA for which read counts exceed the background threshold, based on sequencing of a second cell population after infection with a vector containing the library of at least 3 gRNAs; determine the total number of gRNAs (N') present in the second cell population in all of the target regions of DNA for which read counts exceed the background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA; for each target region of the plurality of target regions, determine the probability of observing n' gRNAs for the target region in a selected cell, based on n, N, n', and N'; for a target region containing a sequence of interest, determine the probability of observing n' or more gRNAs for the sequence of interest in a selected cell, based on the probability of observing n' gRNAs for the target region in a selected cell; and identify that the sequence of interest is positively selected, based on the probability of observing n' or more gRNAs for the sequence of interest in a selected cell.
[0070] Determining the total number of gRNAs (N) present in the first cell population in all of the target regions of DNA for which read counts exceed a background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA, can include counting each gRNA present in the first cell population for which read counts exceed the background threshold.
[0071] Determining the total number of gRNAs (N') present in the second cell population in all of the target regions of DNA for which read counts exceed the background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA, can include counting each gRNA present in selected cells of the second cell population for which read counts exceed the background threshold.
[0072] For each of the plurality of target regions, determining the probability of observing n' guides of the target region in a selected cell based on n, N, n', and N' can include evaluating where Calculating the number of ways to choose x objects from y objects.
[0073] For a target region containing a sequence of interest, determining the probability of observing n' or more gRNAs of the sequence of interest in a selected cell based on the probability of observing n' guides of the target region in the selected cell can include evaluating
[0074] Identifying that the sequence of interest is positively selected based on the probability of observing n' or more gRNAs of the sequence of interest in a selected cell can include determining that the probability of observing n' or more gRNAs of the sequence of interest in the selected cell meets a threshold.
[0075] Software 306 can also be configured to determine an enrichment score for each gRNA. Determining an enrichment score for each gRNA can include evaluating N / N'.
[0076] Software 306 can be configured to store the read counts and quantities of gRNAs that exist as data 307. Figure 4 An exemplary data structure 410 representing an embodiment of data 307 is shown. The data structure 410 can include one or more tables, arrays, and the like. The data structure 410 can include multiple rows and multiple columns. The "Gene #" column can contain identifiers of target regions of unique genes or DNA. The "gRNAs per Gene" column contains identifiers of unique gRNAs that can bind to the target regions of genes or DNA. The "gRNA Read Count (Before Selection)" column contains values of read counts. The "n" column contains the number of gRNAs present for each target region of DNA (e.g., gene) from the "gRNA Read Count (Before Selection)" column. The "gRNA Read Count (After Selection)" column contains values of read counts. The "n'" column contains the number of gRNAs present for each target region of DNA (e.g., gene) from the "gRNA Read Count (After Selection)" column.
[0077] For illustrative purposes, Figure 4The data structure 410 shows artificial values stored as data 307. For gene 1, gRNAs A, B, and C are present in the library and can bind to gene 1. After sequencing, gRNAs A, B, and C are present in the cells with read counts of 100, 200, and 10 respectively. A background threshold 30 can be applied to the read counts, and it can be determined by software 306 that the number of positive occurrences of gRNAs from the library is 2 (n1 = 3), because only gRNAs A and B have read counts exceeding 30. For gene 2, gRNAs D, E, and F are present in the library and can bind to gene 2. After sequencing, gRNAs D, E, and F are present in the cells with read counts of 500, 300, and 200 respectively. A background threshold 30 can be applied to the read counts, and it can be determined by software 306 that the number of positive occurrences of gRNAs from the library is 3 (n2 = 3), because gRNAs D, E, and F have read counts exceeding 30. For gene 3, gRNAs G, H, I, and J are present in the library and can bind to gene 3. After sequencing, gRNAs G, H, I, and J are present in the cells with read counts of 200, 250, 10, and 300 respectively. A background threshold 30 can be applied to the read counts, and it can be determined by software 306 that the number of positive occurrences of gRNAs from the library is 3 (n3 = 3), because only gRNAs G, H, and J have read counts exceeding 30. For gene 4, gRNAs K, L, M, and N are present in the library and can bind to gene 4. After sequencing, gRNAs K, L, M, and N are present in the cells with read counts of 100, 200, 200, and 100 respectively. A background threshold 30 can be applied to the read counts, and it can be determined by software 306 that the number of positive occurrences of gRNAs from the library is 4 (n4 = 4), because gRNAs K, L, M, and N have read counts exceeding 30.
[0078] As Figure 4 shown in the data structure 410, the total number of gRNAs or N is 12. N is derived from the n values 2 + 3 + 3 + 4 (n1 + n2 + n3 + n4) for each gene. Thus, in the original library of 14 different gRNAs (gRNAs A - N) used to infect the cell population, only 12 different gRNAs have a presence count exceeding the background threshold. The values of n and N can be stored in the data structure. In addition to the values of n and N, the data structure can contain one or more nucleic acid sequences (e.g., target regions and / or gRNA sequences) and / or one or more gRNA identifiers.
[0079] The "Gene #" column contains identifiers of target regions of unique genes or DNA. The "gRNA per Gene" column contains identifiers of unique gRNAs that can bind to the target regions of the genes or DNA. The "gRNA Read Count (After Selection)" column contains values of read counts. The "n'" column contains the number of gRNAs present for each DNA target region (e.g., gene) from the "gRNA Read Count (After Selection)" column. As Figure 4 shown in the data structure 410, for Gene 1, gRNAs A, B, and C are present in the library and can bind to Gene 1. After selection, upon sequencing, gRNAs A, B, and C are present (or absent) in the cells with read counts of 0, 50, and 0, respectively. A background threshold of 30 can be applied to the read counts, and it can be determined by the software 306 that the number of positive occurrences of the gRNAs from the library is 1 (n ’1 = 1), since only gRNA B has a read count exceeding 30. For Gene 2, gRNAs D, E, and F are present in the library and can bind to Gene 2. After selection, upon sequencing, gRNAs D, E, and F are present in the cells with read counts of 100, 100, and 100, respectively. A background threshold of 30 can be applied to the read counts, and it can be determined by the software 306 that the number of positive occurrences of the gRNAs from the library is 3 (n'2 = 3), since gRNAs D, E, and F have read counts exceeding 30. For Gene 3, gRNAs G, H, I, and J are present in the library and can bind to Gene 3. After selection, upon sequencing, gRNAs G, H, I, and J are present (or absent) in the cells with read counts of 0, 0, 0, and 20, respectively. A background threshold of 30 can be applied to the read counts, and it can be determined by the software 306 that the number of positive occurrences of the gRNAs from the library is 0 (n'3 = 0), since none of the read counts exceed 30. For Gene 4, gRNAs K, L, M, and N are present in the library and can bind to Gene 4. After selection, upon sequencing, gRNAs K, L, M, and N are present (or absent) in the cells with read counts of 0, 0, 0, and 20, respectively. A background threshold of 30 can be applied to the read counts, and it can be determined by the software 306 that the number of positive occurrences of the gRNAs from the library is 0 (n'4 = 0), since the read counts of the gRNAs do not exceed 30.
[0080] The total number of gRNAs or N' in all target regions present in the selected cells with read counts exceeding the background threshold can be determined by the software 306. As Figure 4As shown in the data structure 410, the total number of gRNAs or N' is 4. N' is derived from the n values 1 + 3 + 0 + 0 (n'1 + n'2 + n'3 + n'4) for each gene. Thus, in the original library of 14 different gRNAs (gRNAs A - N) used to infect the cell population, only 4 different gRNAs are present in the selected cells in numbers exceeding the background threshold. The values of n' and N' can be stored in the data structure. In addition to the values of n' and N', the data structure can include one or more nucleic acid sequences (e.g., target regions and / or gRNA sequences) and / or one or more gRNA identifiers.
[0081] In one embodiment, the data 307 can also be configured to store one or more results of the software 306. Figure 5 An exemplary result data structure 510 produced by the software 306 is shown, e.g., using the data structure 410 as input. The data structure 510 can include one or more tables, arrays, etc. For illustrative purposes, Figure 5 Artificial data and artificial results are shown. Given the library size, the number of guides per gene, and the total number of forward guides in each experiment, formal statistical p - values can be calculated to the number of guides in the experimental replicates for a positive observation. The data structure 510 represents an initial library of 63,950 gRNAs (N). The data structure 510 represents that after selection, 4,946 gRNAs (N') (the number of unique gRNAs, not the number of gRNAs) were retained in the cell population of experiment 1 and 13,606 gRNAs (N') (the number of unique gRNAs, not the number of gRNAs) were retained in the cell population of experiment 2. As shown in the data structure 510, target region 1 has three (3) gRNAs capable of binding to at least a portion of target region 1. For experiment 1, assuming 4,946 gRNAs are retained in the cell population after selection, the probability that all 3 of the 3 gRNAs are present in the cell population after selection by chance is 0.000462378. For experiment 2, assuming 13,606 gRNAs are retained in the cell population after selection, the probability that all 3 of the 3 gRNAs are present in the cell population after selection by chance is 0.009629. The data structure 510 represents the results of the probabilities that 2 out of 3 gRNAs, 1 out of 3 gRNAs, and 0 out of 3 gRNAs are present in the cell population after selection by chance.
[0082] As shown in data structure 510, the target region 2 has four (4) gRNAs that can bind to at least a portion of the target region 2. Data structure 510 represents the results of the probabilities of 4 out of 4, 3 out of 4, 2 out of 4, 1 out of 4, and 0 out of 4 of the 4 gRNAs being present in the selected cell population by chance. Based on the results shown in data structure 510, software 306 can determine that the probability is below a threshold (e.g., small enough). For example, for target region 2, experiment 1, the probability of 4 out of 4 of the 4 gRNAs being present by chance is 3.57411E-05, which indicates that it is very likely that 4 out of 4 of the 4 gRNAs are not just present by chance.
[0083] Example
[0084] A. Example 1. Development of a genome-wide CRISPR / Cas9 screening platform to identify genetic modifiers of Tau aggregation
[0085] To identify genes and pathways that alter the abnormal tau protein aggregation process, a platform for genome-wide screening using a CRISPR nuclease (CRISPRn) sgRNA library was developed. The screening was used to identify genes that regulate the potential of cells to be "seeded" with tau disease-related protein aggregates (i.e., genes that, when disrupted, cause cells to be more likely to form tau aggregates when exposed to a source of tau protofibrillar protein). The screening used a tau biosensor human cell line composed of HEK293T cells stably expressing the tau four-repeat domain tau_4RD, which contains the microtubule-binding domain (MBD) of tau with the P301S pathogenic mutation fused to CFP or YFP. That is, the HEK293T cell line contains two transgenes stably expressing disease-related protein variants fused to the fluorescent protein CFP or the fluorescent protein YFP: tau4RD-CFP / tau4RD-YFP (TCY), where the tau repeat domain (4RD) contains the P301S pathogenic mutation.
[0086] In these biosensor cell lines, tau-CFP / tau-YFP protein aggregation generates a FRET signal, which is the result of fluorescence energy transfer from the donor CFP to the acceptor YFP. FRET-positive cells contain tau aggregates and can be sorted and isolated by flow cytometry. At baseline, unstimulated cells express the reporter gene in a stable, soluble state with minimal FRET signal. Upon stimulation (e.g., lipofection with seed particles), the reporter protein forms aggregates, generating a FRET signal. Cells containing aggregates can be isolated by FACS. A stably propagated cell line containing aggregates, Agg[+], can be isolated by serial dilution cloning of the Agg[-] cell line.
[0087] Several modifications were made to the tau biosensor cell line to make it suitable for genetic screening. First, these tau biosensor cells were modified by introducing a transgene expressing Cas9 (SpCas9) via a lentiviral vector. Clonal transgenic cell lines expressing Cas9 were selected with blasticidin and isolated by serial dilution cloning to obtain single-cell-derived clones. The Cas9 expression levels of the clones were evaluated by qRT-PCR, and the DNA cleavage activity was evaluated by digital PCR.
[0088] Specifically, Cas9 mutation efficiency was evaluated by digital PCR 3 days and 7 days after transduction of the lentivirus encoding gRNAs against two selected target genes. The cleavage efficiency was limited by the Cas9 level in the low-expressing clones. Clones with sufficient Cas9 expression levels were required to achieve maximum activity. Several derived clones with lower Cas9 expression were unable to efficiently cleave the target sequences, while clones with higher expression (including the clone used for screening) were able to generate mutations at the target sequences in the genes PERK and SNCA after three days of culture with an efficiency of approximately 80%. Efficient cleavage was already observed 3 days after gRNA transduction and improved only slightly after 7 days. Clone 7B10-C3 was selected as a high-performance clone for subsequent library screening.
[0089] Then, reagents and methods were developed to make the cells sensitive to tau seeding activity. Tau intercellular propagation may be the result of tau seeding activity secreted by cells containing aggregates. To study the cell proliferation of tau aggregation, subclones of the tau-YFP cell line were obtained, which consisted of HEK293T cells stably expressing the tau repeat domain tau_4RD, which contains the tau microtubule-binding domain (MBD) with the P301S pathogenic mutation fused to YFP.
[0090] Cells in which the tau-YFP protein is stably in the aggregated state (Agg[+]) were obtained by treating these tau-YFP cells with a mixture of recombinant protofibrillar tau and lipofectamine reagent, thereby inoculating the aggregation of the tau-YFP protein stably expressed in these cells. Then the "inoculated" cells were serially diluted to obtain single-cell-derived clones. These clones were then expanded to identify clonal cell lines in which tau-YFP aggregates stably continued to grow and were passaged multiple times over time in all cells. One of these tau-YFP_Agg[+] clones was used to generate conditioned medium by collecting the medium that had been placed on confluent tau-YFP_Agg[+] cells for four days. The conditioned medium (CM) was then applied to the initial biosensor tau-CFP / Tau-YFP cells at a ratio of 3:1 CM:fresh medium to induce tau aggregation in a small fraction of these recipient cells. Lipofectamine was not used. The non-use of lipofectamine was to perform an assay as physiological as possible without using lipofectamine to induce forced / increased tau aggregation in the recipient cells. As measured by flow cytometry, the conditioned medium continuously induced FRET in approximately 0.1% of the cells, and the flow cytometry evaluated the percentage of cells generating a FRET signal as a measure of aggregation.
[0091] B. Example 2. Genome-wide CRISPR / Cas9 screening to identify genetic modifiers of Tau aggregation
[0092] To uncover modifier genes of tau aggregation as sgRNAs enriched in FRET(+) cells, two human genome-wide CRISPR sgRNA libraries (GeCKO A and GeCKO B) were used to transduce aggregate-free Cas9-expressing tau-CFP / tau-YFP biosensor cells (Agg–)( Figure 7) Knockout mutations were introduced at each target gene using a lentiviral delivery method. Each CRISPR sgRNA library targeted the 5' constitutive exon for functional knockout, with each gene covered by approximately 3 sgRNAs on average (a total of 6 gRNAs per gene in the combination of two libraries). The read count distribution of each library (i.e., the representation of each gRNA in the library) was normal and similar. The sgRNAs were designed to avoid off-target effects by avoiding sgRNAs with two or fewer mismatches to off-target genomic sequences. These libraries covered 19,050 human genes and 1,864 miRNAs, as well as 1,000 non-targeting control sgRNAs. The libraries were transduced at a multiplicity of infection (MOI) of <0.3, with a coverage of >300 cells / sgRNA. Tau biosensor cells were grown under puromycin selection to select cells that had integrated and expressed a unique sgRNA per cell. Puromycin selection started 24 hours after transduction at a concentration of 1 μg / mL. Five independent screening replicates were used in the primary screen.
[0093] Samples of the complete transduced cell population were collected at cell passage on days 3 and 6 after transduction. After passage on day 6, the cells were grown in conditioned medium to make them sensitive to seeding activity. On day 10, fluorescence-activated cell sorting (FACS) was used to specifically isolate the subpopulation of FRET[+] cells. The screen consisted of five replicate experiments. DNA isolation and PCR amplification of the integrated sgRNA constructs allowed characterization of the sgRNA library by next-generation sequencing (NGS) at each time point.
[0094] Statistical analysis of the NGS data was able to identify sgRNAs enriched in the FRET[+] subpopulation on day 10 of the five experiments compared to the sgRNA libraries at the earlier time points of days 3 and 6. The first strategy for identifying potential tau modifiers was to use DNA sequencing to generate sgRNA read counts in each sample using the DESeq algorithm to find sgRNAs that were more abundant on day 10 than on day 3 or more abundant on day 10 than on day 6, but not more abundant on day 6 than on day 3 (fold change (fc) ≥1.5 and negative binomial test p < 0.01). Fc ≥1.5 indicates that the ratio (average count on day 10) / (average count on day 3 or day 6) ≥1.5. P < 0.01 indicates that the probability of no statistical difference between the counts on day 10 and day 3 or day 6 is <0.01. The DESeq algorithm is a widely used algorithm for "differential expression analysis of sequence count data". See, for example, Anders et al. (2010) Genome Biology 11:R106, which is incorporated herein by reference.
[0095] Specifically, two comparisons were used in each library to identify significant sgRNAs: day 10 vs. day 3 and day 10 vs. day 6. For each of these four comparisons, the DESeq algorithm was used, and the cut-off thresholds considered significant were fold change ≥ 1.5 and negative binomial test p < 0.01. Once significant guides were identified in each of these comparisons in each library, a gene was considered significant if it met one of the following two criteria: (1) at least two sgRNAs corresponding to the gene were considered significant in one comparison (day 10 vs. day 3 or day 10 vs. day 6); (2) at least one sgRNA was significant in both comparisons (day 10 vs. day 3 and day 10 vs. day 6). Using this algorithm, five genes were identified as significant from the first library and four genes were identified as significant from the second library. See Table 1.
[0096] Table 1. Genes Identified Using Strategy #1.
[0097]
[0098] However, the requirement for a certain level of read count homogeneity within each experimental group in this first strategy may be too strict. For the same sgRNA, many factors can create differences in read counts between samples (day 3, day 6, or day 10 samples) within each experimental group, such as the initial viral count in the screening library, infection or gene editing efficiency, and the relative growth rate after gene editing. Therefore, the second strategy also uses the positive occurrence (read count > 30) of the guide for each gene in each sample at day 10 (after selection) instead of the exact read count.
[0099] The pre-selection CRISPR experiment was repeated four times. As Figure 6As shown, for gene "G1" in the pre-selection experiment "Experiment 1", gRNAs "g1", "g2", and "g3" are present in the cell population, with read counts of 121, 1000, and 302 respectively. For gene "G2", gRNAs "g4", "g5", "g6", and "g7" are present in the cell population, with read counts of 443, 2012, 534, and 150 respectively. This read count data was generated for genes "G1" through "G21,000". For each gene, the "sum of presences" or n was determined. The "sum of presences", or the corresponding number of gRNAs present in the target region of each DNA, was determined by comparing the individual read count of each gRNA to a background threshold. In this case, 30 was used as the background threshold. Thus, the "sum of presences" represents the number of gRNAs present qualitatively, in contrast to the number of gRNAs present represented by the read count. The sum of presences of the gRNAs corresponding to gene G1 is 3, because the read counts of gRNAs g1, g2, and g3 all exceed the background threshold of 30. The sum of presences of the gRNAs corresponding to gene G2 is 4, because the read counts of gRNAs g4, g5, g6, and g7 all exceed the background threshold of 30. The total number of gRNAs or N present in all target regions in the cell population with read counts exceeding the background threshold is represented as 59,010. Thus, in the original library of approximately 64,000 different gRNAs used to infect the cell population, only approximately 59,000 different gRNAs were present in numbers exceeding the background threshold.
[0100] The CRISPR experiment with phenotypic selection was repeated four times. However, prior to sequencing, the cells in the cell population were sorted according to phenotype using fluorescence techniques (e.g., FRET fluorescence). If the Cas9 / CRISPR cuts the target region (e.g., gene) of a cell and the cell does not fluoresce, then the gene has been successfully knocked out. If the cell fluoresces, then the gene has not been knocked out. The fluorescing cells can then be sequenced, while the non-fluorescing cells are not sequenced. The selected cells represent the cells exhibiting a specific phenotype / marker.
[0101] As Figure 6As shown, for gene "G1" in the post-selection experiment "Experiment 1", gRNAs "g1", "g2", and "g3" are present (or absent) in the cell population, with read counts of 0, 8, and 12 respectively. For gene "G2", gRNAs "g4", "g5", "g6", and "g7" are present (or absent) in the cell population, with read counts of 4, 25, 4, and 150 respectively. This read count data was generated for genes "G1" through "G21,000". For each gene, the "sum of presences" or n' is determined. The "sum of presences", or the corresponding number of gRNAs present in the target region of each DNA, is determined by comparing the individual read counts of each gRNA to a background threshold. In this case, 30 is used as the background threshold. Thus, the "sum of presences" represents the number of gRNAs present, as opposed to the number of gRNAs present represented by the read count. The sum of presences of the gRNAs corresponding to gene G1 is 0, since the read counts of gRNAs g1, g2, and g3 do not exceed the background threshold of 30. The sum of presences of the gRNAs corresponding to gene G2 is 1, since only the read count of gRNA g7 exceeds the background threshold of 30. The total number of gRNAs or N' present in all target regions in the cell population with read counts exceeding the background threshold is represented as 4,320. Thus, in the original library of approximately 64,000 different gRNAs used to infect the cell population, only approximately 4,320 different gRNAs are present in numbers exceeding the background threshold.
[0102] Given the library size, the number of guides per gene, and the total number of positive guides in the post-selection sample, a formal statistical p-value is calculated to observe the number of guides in the post-selection sample positively. Once read counts are considered, the probability of randomly observing n' guides in the target region (e.g., gene) of the selected cells (post-selection) is determined according to the formula As an explanation, determines the number of ways to choose x objects from y objects.
[0103] Once the probability of randomly observing n' guides in the target region (e.g., gene) of the selected cells (post-selection) is determined, the probability of randomly observing n' or more gRNAs in the gene in the selected cells (post-selection) is determined according to the formula Once the probability of randomly observing n' or more gRNAs in the target region of the selected cells (post-selection) is determined, the average enriched gRNA is determined at the target region level. The overall enrichment of the read counts of the gene post-selection compared to pre-selection is used as an additional parameter to identify positive genes. The average enrichment is expressed as an enrichment fraction. The enrichment fraction is determined by evaluating N / N'. As
[0104] Once the probability of randomly observing n' or more gRNAs in the target region of the selected cells (post-selection) is determined, the average enriched gRNA is determined at the target region level. The overall enrichment of the read counts of the gene post-selection compared to pre-selection is used as an additional parameter to identify positive genes. The average enrichment is expressed as an enrichment fraction. The enrichment fraction is determined by evaluating N / N'. As Figure 6As shown, the enrichment score is 59010 / 4320 or 13.66.
[0105] The probability of observing n' or more gRNAs in the target region in the selected cells (after selection) by chance can be used to evaluate whether the target region is under positive selection. The enrichment score can additionally be used to evaluate whether the target region is under positive selection. A target region with a probability of observing n' or more gRNAs in the selected cells that is significantly lower than the probability of observing n' or more gRNAs in the selected cells (after selection) by chance can be identified as a positively selected target region. Additionally, a target region with an enrichment score greater than a threshold can indicate a positively selected target region.
[0106] Thus, this second strategy represents a novel and more sensitive CRISPR positive selection assay method. The goal of CRISPR positive selection is to use DNA sequencing to identify genes where sgRNA perturbation is associated with a phenotype. To reduce the noise background, multiple sgRNAs of the same gene and experimental replicates are typically used in these experiments. However, currently commonly used statistical analysis methods require a certain degree of homogeneity / consistency between sgRNAs of the same gene and between technical replicates, and the effect is not good. This is because due to many possible reasons (e.g., different infection or gene editing efficiencies, initial viral counts in the screening library, and the presence of other sgRNAs with the same phenotype), these methods cannot handle the large differences between sgRNAs of the same gene and replicates. In contrast, the method shown in this example is robust to large differences. It is based on the positive occurrence of guides for each gene in a single experiment, rather than the exact read counts for each sgRNA. Given the library size, the number of sgRNAs per gene, and the total number of positive sgRNAs in each experiment, a formal statistical p-value is calculated to positively observe the number of sgRNAs in the experimental replicates. The relative sgRNA sequence read enrichment before and after phenotypic selection is also used as a parameter. The performance of this method is superior to the most recently widely used methods, including DESeq, MAGECK, etc. Specifically, this method includes the following steps:
[0107] (1) For each experiment, identify any present guides in the cells with a positive phenotype.
[0108] (2) At the gene level, calculate the random probability (also known as the p-value) of the presence of a guide in each experiment. The overall probability present in multiple experiments is calculated by Fisher's combined probability test (Reference: Fisher, R.A.; Fisher, R.A (1948). "Questions and answers #14". The American Statistician). That is, first use the p-values from multiple experiments to calculate the test statistic where p k is the p-value calculated for the k-th experiment, and K is the total number of experiments. Then the combined p-value in K experiments is equal to the probability φ of the observed value under the chi-square distribution with 2*K degrees of freedom.
[0109] (3) Calculate the average enrichment of the guide at the gene level: Enrichment score = relative abundance after selection / relative abundance before selection. Relative abundance = read count of the guide for a gene / read count of all guides.
[0110] (4) Select genes that are significantly less than the random probability of presence and greater than a certain enrichment score.
[0111] C. Example 3
[0112] CRISPR / Cas9 activation and inactivation mutagenesis were used to screen for genetic modifiers of Tau and α-synuclein protofibrillation and proliferation using the CRISPRn sgRNA libraries (hGeCKO-A and hGeCKO-B) (targeting coding exons for functional knockout). Figure 7 Sample identifiers, experiment numbers, time before sequencing are shown, and the libraries used for the experiments (Gecko A or Gecko B) are identified.
[0113] The Gecko A library consists of approximately 63,950 gRNAs. Among the 63,950 gRNAs, 56,116 gRNAs target 18,874 genes, and most of the genes are targeted by 3 gRNAs. Among the 63,950 gRNAs, 6,834 gRNAs target 1,795 microRNAs, and most of the microRNAs are targeted by 4 gRNAs.
[0114] The Gecko B library consists of approximately 56,869 gRNAs. Among the 56,869 gRNAs, 55,869 gRNAs target 18,834 genes, and most of the genes are targeted by 3 gRNAs. Among the 56,869 gRNAs, no gRNAs target microRNAs. There are no targets among 1,000 gRNAs.
[0115] On day 3 and day 6, no phenotype was shown. On day 10, the sample was phenotypically positive.
[0116] Figure 8 Shows DNA read counts for samples infected with the Gecko A library. Each bar represents a virus-infected sample. Each sample was sequenced on day 3 (d03), day 6 (d06), or day 10 (d10), and each sample represents the sequencing reads of the gRNA, rather than the whole genome.
[0117] Figure 9 Shows the read counts from samples infected with the Gecko A library normalized to the median. Normalization was performed by dividing the gRNA read count by the sum of the read counts in each sample and multiplying by the median of the sum of the read counts in all samples. The bottom bar graph is qualitative and represents read counts greater than the threshold of 30; day 10 is after selection. Given the similar sum of read counts, samples on day 3 and day 6 have multiple gRNAs, while samples on day 10 have fewer gRNAs.
[0118] Figure 10 Shows the formal statistical p-values calculated given the library size, number of gRNAs per gene, and total number of positive gRNAs in each experiment, for a large number of gRNAs in forward-observed experimental replicates. The p-values for five samples infected with the Gecko A library and sequenced on day 10 are shown.
[0119] Embodiments
[0120] Embodiment 1. A method, comprising:
[0121] (A) Infecting a first culture of cas9-positive cells with a library of viral vectors, the library comprising at least 3 guide RNAs (gRNAs) for cleaving a target region of DNA within the genome of the cells; sequencing the cells to obtain a read count for each of the gRNAs;
[0122] Summing (Σ) the corresponding number of gRNAs for each DNA target region with a read count exceeding a background threshold, where Σ = n; and summing (Σ) the total number of gRNAs in all target regions with a read count exceeding the background threshold, where Σ = N;
[0123] (B) Infecting a second culture of cas9-positive cells with the library of viral vectors;
[0124] Sorting the cells of the second culture as having a designated phenotype or not having the designated phenotype; selecting the cells having the designated phenotype and sequencing the selected cells to obtain post-selection read counts for each of the gRNAs;
[0125] Sum (Σ) the corresponding number of gRNAs for each DNA target region in the selected cells whose read counts exceed the background threshold, where Σ = n'; and
[0126] Sum (Σ) the total number of gRNAs in all target regions of the selected cells whose read counts exceed the threshold, where Σ = N'; (C) For a target region of DNA, calculate the probability of randomly observing n' gRNAs for the target region in the selected cells according to the formula where Calculate the number of ways to choose x objects from y objects; and for a target region of DNA containing a gene, calculate the probability of randomly observing n' or more gRNAs for the gene in the selected cells according to the formula
[0127] Embodiment 2. A method comprising:
[0128] (A) Infect a first culture of cas9-positive cells with a library of viral vectors, the library comprising at least 3 guide RNAs (gRNAs) for enhancing transcription of a target region of DNA within the genome of the cells;
[0129] Sequence the cells to obtain a read count for each of the gRNAs;
[0130] Sum (Σ) the corresponding number of gRNAs for each DNA target region whose read counts exceed a background threshold, where Σ = n; and
[0131] Sum (Σ) the total number of gRNAs in all target regions whose read counts exceed the background threshold, where Σ = N;
[0132] (B) Infect a second culture of cas9-positive cells with the library of viral vectors;
[0133] Classify the cells of the second culture as having a specified phenotype or not having the specified phenotype;
[0134] Select the cells having the specified phenotype and sequence the selected cells to obtain a post-selection read count for each of the gRNAs;
[0135] Sum (Σ) the corresponding number of gRNAs for each DNA target region in the selected cells whose post-selection read counts exceed the background threshold, where Σ = n'; and
[0136] Sum (Σ) the total number of gRNAs in all target regions of the selected cells whose read counts exceed the threshold, where Σ = N';
[0137] (C) For the target region of DNA, according to the formula to calculate the probability of observing n' gRNAs of the target region in the selected cells by chance, where calculate the number of ways to choose x objects from y objects; and
[0138] For the target region of DNA containing a gene, according to the formula to calculate the probability of observing n' or more gRNAs of the gene in the selected cells by chance.
[0139] Embodiment 3. The method according to any one of the preceding embodiments, wherein the first culture is infected according to the CRISPR technique.
[0140] Embodiment 4. The method according to any one of the preceding embodiments, wherein the cas-9 positive cells of the first culture are modified to contain one or more selectable markers.
[0141] Embodiment 5. The method according to any one of the preceding embodiments, wherein the specified phenotype is fluorescence.
[0142] Embodiment 6. The method according to any one of the preceding embodiments, wherein the specified phenotype is cell survival.
[0143] Embodiment 7. The method according to Embodiment 4, wherein the one or more selectable markers include fluorescent markers.
[0144] Embodiment 8. The method according to Embodiment 7, wherein the fluorescent marker is part of a FRET biosensor.
[0145] Embodiment 9. The method according to Embodiment 4, wherein the one or more selectable markers are detectable enzymes.
[0146] Embodiment 10. The method according to Embodiment 9, wherein the detectable enzyme is β-galactosidase.
[0147] Embodiment 11. The method according to Embodiment 9, wherein the detectable enzyme is luciferase.
[0148] Embodiment 12. The method according to any one of the preceding embodiments, wherein the target region contains a gene.
[0149] Embodiment 13. The method according to any one of the preceding embodiments, wherein classifying the cells of the second culture as having the specified phenotype or not having the specified phenotype includes applying a selection mechanism to the second cell culture.
[0150] Embodiment 14. The method according to embodiment 13, wherein the selection mechanism comprises one or more of the following: exposing the second cell population to a drug, or exposing the second cell population to a substance that recognizes protein activity or expression level.
[0151] Embodiment 15. The method according to embodiment 4, wherein selecting the cells having the specified phenotype comprises sorting the cells according to the one or more selection markers.
[0152] Embodiment 16. The method according to any one of the foregoing embodiments, further comprising identifying that the target region is positively selected based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cells.
[0153] Embodiment 17. The method according to embodiment 16, wherein identifying that the target region is positively selected comprises determining that the probability of accidentally observing n' or more gRNAs of the sequence of interest in the selected cells meets a threshold.
[0154] Embodiment 18. The method according to any one of the foregoing embodiments, further comprising determining an enrichment score for each gRNA.
[0155] Embodiment 19. The method according to embodiment 18, wherein determining the enrichment score for each gRNA comprises evaluating N / N'.
[0156] Embodiment 20. The method according to any one of embodiments 2-19, wherein the cas-9 positive cells comprise inactivated cas-9.
[0157] Embodiment 21. The method according to embodiment 20, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0158] Embodiment 22. The method according to any one of the foregoing embodiments, wherein the target region of the DNA regulates a downstream gene or protein.
[0159] Embodiment 23. The method according to embodiment 22, wherein regulating the downstream gene or protein comprises activation or inhibition of the downstream gene or protein.
[0160] Embodiment 24. The method according to any one of the foregoing embodiments, further comprising identifying the target region as a modifier of a second gene based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cells.
[0161] Embodiment 25. The method according to any one of the foregoing embodiments, further comprising identifying the target region as a therapeutic target based on the probability of observing n' or more gRNAs of the gene incidentally in the selected cells.
[0162] Embodiment 26. The method according to any one of the foregoing embodiments, further comprising identifying the target region as being related to the specified phenotype based on the probability of observing n' or more gRNAs of the gene incidentally in the selected cells.
[0163] Embodiment 27. The method according to any one of the foregoing embodiments, further comprising identifying the target region as exhibiting a protective effect based on the probability of observing n' or more gRNAs of the gene incidentally in the selected cells.
[0164] Embodiment 28. A method comprising: determining, for each of a plurality of target regions of DNA, the corresponding number of gRNAs (n) present in each of the plurality of target regions of DNA for which the read count exceeds a background threshold, based on sequencing of a first cell population after infection with a vector comprising a library of at least 3 guide RNAs (gRNAs); determining the total number (N) of gRNAs present in the first cell population in all of the plurality of target regions of DNA for which the read count exceeds the background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA; determining, for each of the plurality of target regions of DNA, the corresponding number of gRNAs (n') present in each of the plurality of target regions of DNA for which the read count exceeds the background threshold, based on sequencing of a second cell population after infection with a vector comprising the library of at least 3 gRNAs;
[0165] determining the total number (N') of gRNAs present in the second cell population in all of the plurality of target regions of DNA for which the read count exceeds the background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA; for each target region of the plurality of target regions, determining the probability of observing n' gRNAs of the target region incidentally in the selected cells based on n, N, n', and N'; for a target region comprising a sequence of interest, determining the probability of observing n' or more gRNAs of the sequence of interest incidentally in the selected cells based on the probability of observing n' gRNAs of the target region incidentally in the selected cells; and identifying the sequence of interest as being positively selected based on the probability of observing n' or more gRNAs of the sequence of interest incidentally in the selected cells.
[0166] Embodiment 29. The method according to embodiment 28, wherein the first cell population and the second cell population comprise cas-9 positive cells.
[0167] Embodiment 30. The method according to embodiment 29, wherein for each of a plurality of target regions of DNA, the number of corresponding gRNAs (n) present in each of the plurality of target regions of DNA for which the read count exceeds a background threshold, as determined by sequencing the first cell population after infection with a vector containing a library of at least 3 gRNAs, comprises: infecting the first cell population with the library of the viral vector; sequencing the cells of the first cell population to obtain the read count for each of the gRNAs; and for each of the plurality of target regions of DNA, counting each gRNA for which the read count exceeds the background threshold.
[0168] Embodiment 31. The method according to embodiment 30, wherein the total number of gRNAs (N) present in the first cell population in all of the plurality of target regions of DNA for which the read count exceeds the background threshold, as determined according to the number of corresponding gRNAs for each of the plurality of target regions of DNA, comprises counting each gRNA present in the first cell population for which the read count exceeds the background threshold.
[0169] Embodiment 32. The method according to embodiment 31, wherein for each of a plurality of target regions of DNA, the number of corresponding gRNAs (n') present in each of the plurality of target regions of DNA for which the read count exceeds the background threshold, as determined by sequencing the second cell population after infection with a vector containing the library of at least 3 gRNAs, comprises: sorting the cells of the second cell population as having a specified phenotype or not having the specified phenotype;
[0170] selecting the cells having the specified phenotype; sequencing the selected cells to obtain the read count for each of the gRNAs; and for each of the plurality of target regions of DNA, counting each gRNA for which the read count exceeds the background threshold.
[0171] Embodiment 33. The method according to any one of embodiments 28-32, wherein the total number of gRNAs (N') present in the second cell population in all of the plurality of target regions of DNA for which the read count exceeds the background threshold, as determined according to the number of corresponding gRNAs for each of the plurality of target regions of DNA, comprises counting each gRNA present in the selected cells of the second cell population for which the read count exceeds the background threshold.
[0172] Embodiment 34. The method according to any one of embodiments 28-33, wherein for each target region among the plurality of target regions, determining the probability of randomly observing n' gRNAs of the target region in the selected cells based on n, N, n', and N' includes evaluating where Calculating the number of ways to select x objects from y objects.
[0173] Embodiment 35. The method according to any one of embodiments 28-34, wherein for a target region containing a sequence of interest, determining the probability of randomly observing n' or more gRNAs of the sequence of interest in the selected cells based on the probability of randomly observing n' gRNAs of the target region in the selected cells includes evaluating
[0174] Embodiment 36. The method according to any one of embodiments 28-35, wherein identifying that the sequence of interest is positively selected based on the probability of randomly observing n' or more gRNAs of the sequence of interest in the selected cells includes determining that the probability of randomly observing n' or more gRNAs of the sequence of interest in the selected cells meets a threshold.
[0175] Embodiment 37. The method according to any one of embodiments 28-36, further comprising determining an enrichment score for each gRNA.
[0176] Embodiment 38. The method according to embodiment 37, wherein determining the enrichment score for each gRNA includes evaluating N / N'.
[0177] Embodiment 39. The method according to any one of embodiments 28-38, wherein the cas-9 positive cells comprise inactivated cas-9.
[0178] Embodiment 40. The method according to embodiment 39, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0179] Embodiment 41. A device, comprising: one or more processors; and a memory storing processor-executable instructions that, when executed by the one or more processors, cause the device to: (A) receive first read count data, wherein the first read count data is generated by: infecting a first culture of cas9-positive cells with a library of viral vectors, the library comprising at least 3 guide RNAs (gRNAs) for cleaving target regions of DNA within the genome of the cells, and sequencing the cells to obtain read counts for each of the gRNAs; summing (Σ) the corresponding number of gRNAs for each DNA target region for which the read count exceeds a background threshold, wherein Σ = n; and summing (Σ) the total number of gRNAs in all target regions for which the read count exceeds the background threshold, wherein Σ = N; (B) receive second read count data, wherein the second read count data is generated by: infecting a second culture of cas9-positive cells with the library of viral vectors, sorting the cells of the second culture as having a specified phenotype or not having the specified phenotype, and selecting the cells having the specified phenotype and sequencing the selected cells to obtain post-selection read counts for each of the gRNAs; summing (Σ) the corresponding number of gRNAs for each DNA target region in the selected cells for which the post-selection read count exceeds the background threshold, wherein Σ = n'; and summing (Σ) the total number of gRNAs in all target regions of the selected cells for which the read count exceeds the threshold, wherein Σ = N'; (C) for a target region of DNA, calculate the probability of randomly observing n' gRNAs for the target region in the selected cells according to the formula wherein calculate the number of ways to select x objects from y objects; and for a target region of DNA containing a gene, calculate the probability of randomly observing n' or more gRNAs for the gene in the selected cells according to the formula
[0180] Embodiment 42. The device according to Embodiment 41, wherein the first culture is infected according to CRISPR technology.
[0181] Embodiment 43. The device according to any one of Embodiments 41-42, wherein the cas-9 positive cells of the first culture are modified to contain one or more selectable markers.
[0182] Embodiment 44. The device according to any one of Embodiments 41-43, wherein the target region contains a gene.
[0183] Embodiment 45. The apparatus according to any one of embodiments 41-44, wherein classifying the cells of the second culture as having or not having the designated phenotype comprises applying a selection mechanism to the second cell culture.
[0184] Embodiment 46. The apparatus according to embodiment 45, wherein the selection mechanism comprises one or more of the following: exposing a second cell population to a drug, or exposing the second cell population to a substance that recognizes protein activity or expression level.
[0185] Embodiment 47. The apparatus according to any one of embodiments 41-46, wherein selecting cells having the designated phenotype comprises sorting the cells according to the one or more selection markers.
[0186] Embodiment 48. The apparatus according to any one of embodiments 41-47, further comprising identifying that the target region is positively selected based on the probability of observing n' or more gRNAs of the gene incidentally in the selected cells.
[0187] Embodiment 49. The apparatus according to embodiment 48, wherein identifying that the target region is positively selected comprises determining that the probability of observing n' or more gRNAs of the sequence of interest incidentally in the selected cells meets a threshold.
[0188] Embodiment 50. The apparatus according to any one of embodiments 41-49, further comprising determining an enrichment score for each gRNA.
[0189] Embodiment 51. The apparatus according to embodiment 50, wherein determining the enrichment score for each gRNA comprises evaluating N / N'.
[0190] Embodiment 52. The apparatus according to any one of embodiments 41-51, wherein the cas-9 positive cells comprise inactivated cas-9.
[0191] Embodiment 53. The apparatus according to embodiment 52, wherein the inactivated cas-9 is fused with at least one transcriptional activation domain.
[0192] Embodiment 54. A non - transitory computer - readable medium for determining the probability of accidentally observing one or more guide RNAs (gRNAs), the non - transitory computer - readable medium storing processor - executable instructions that, when executed by one or more processors, cause the one or more processors to: (A) Receive first read - count data, wherein the first read - count data is generated by infecting a first culture of cas9 - positive cells with a library of viral vectors, the library comprising at least 3 guide RNAs (gRNAs) for cleaving target regions of DNA within the genome of the cells, and sequencing the cells to obtain the read - count for each of the gRNAs; according to the first read - count data, sum (Σ) the number of corresponding gRNAs for each DNA target region for which the read - count exceeds a background threshold, wherein Σ = n; and according to the first read - count data, sum (Σ) the total number of gRNAs in all target regions for which the read - count exceeds the background threshold, wherein Σ = N; (B) Receive second read - count data, wherein the second read - count data is generated by infecting a second culture of cas9 - positive cells with the library of viral vectors, classifying the cells of the second culture as having a designated phenotype or not having the designated phenotype, and selecting the cells having the designated phenotype and sequencing the selected cells to obtain the post - selection read - count for each of the gRNAs; according to the second read - count data, sum (Σ) the number of corresponding gRNAs for each DNA target region in the selected cells for which the post - selection read - count exceeds the background threshold, wherein Σ = n'; and according to the second read - count data, sum (Σ) the total number of gRNAs in all target regions of the selected cells for which the read - count exceeds the threshold, wherein Σ = N'; (C) For a target region of DNA, calculate the probability of accidentally observing n' gRNAs for the target region in the selected cells according to the formula wherein calculates the number of ways to choose x objects from y objects; and for a target region of DNA containing a gene, calculate the probability of accidentally observing n' or more gRNAs for the gene in the selected cells according to the formula
[0193] Embodiment 55. The non - transitory computer - readable medium according to Embodiment 54, wherein the first culture is infected according to CRISPR technology.
[0194] Embodiment 56. The non - transitory computer - readable medium according to any one of Embodiments 54 - 55, wherein the cas - 9 positive cells of the first culture are modified to contain one or more selectable markers.
[0195] Embodiment 57. The non-transitory computer-readable medium according to any one of Embodiments 54-56, wherein the target region contains a gene.
[0196] Embodiment 58. The non-transitory computer-readable medium according to any one of Embodiments 54-57, wherein classifying the cells of the second culture as having or not having the specified phenotype includes applying a selection mechanism to the second cell culture.
[0197] Embodiment 59. The non-transitory computer-readable medium according to Embodiment 58, wherein the selection mechanism includes one or more of the following: exposing a second cell population to a drug, or exposing the second cell population to a substance that recognizes protein activity or expression level.
[0198] Embodiment 60. The non-transitory computer-readable medium according to any one of Embodiments 54-59, wherein selecting cells having the specified phenotype includes sorting the cells according to the one or more selection markers.
[0199] Embodiment 61. The non-transitory computer-readable medium according to any one of Embodiments 54-60, further comprising identifying that the target region is positively selected based on the probability of observing n' or more gRNAs of the gene incidentally in the selected cells.
[0200] Embodiment 62. The non-transitory computer-readable medium according to Embodiment 61, wherein identifying that the target region is positively selected includes determining that the probability of observing n' or more gRNAs of the sequence of interest incidentally in the selected cells meets a threshold.
[0201] Embodiment 63. The non-transitory computer-readable medium according to any one of Embodiments 54-62, further comprising determining an enrichment score for each gRNA.
[0202] Embodiment 64. The non-transitory computer-readable medium according to Embodiment 63, wherein determining the enrichment score for each gRNA includes evaluating N / N'.
[0203] Embodiment 65. The non-transitory computer-readable medium according to any one of Embodiments 54-64, wherein the cas-9 positive cells contain inactivated cas-9.
[0204] Embodiment 66. The non-transitory computer-readable medium according to Embodiment 65, wherein the inactivated cas-9 is fused with at least one transcriptional activation domain.
[0205] Embodiment 67. A device comprising: one or more processors; and a memory storing processor-executable instructions that, when executed by the one or more processors, cause the device to: for each of a plurality of target regions of DNA, determine the corresponding number of guide RNAs (gRNAs) (n) present in each of the plurality of target regions of DNA for which the read count exceeds a background threshold, based on sequencing of a first cell population after infection with a vector containing a library of at least 3 guide RNAs (gRNAs); determine the total number of gRNAs (N) present in the first cell population in all of the target regions of the plurality of target regions of DNA for which the read count exceeds the background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA; for each of the plurality of target regions of DNA, determine the corresponding number of gRNAs (n') present in each of the plurality of target regions of DNA for which the read count exceeds the background threshold, based on sequencing of a second cell population after infection with a vector containing the library of at least 3 gRNAs; determine the total number of gRNAs (N') present in the second cell population in all of the target regions of the plurality of target regions of DNA for which the read count exceeds the background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA; for each of the plurality of target regions, determine the probability of observing n' gRNAs of the target region by chance in the selected cells, based on n, N, n', and N'; for a target region containing a sequence of interest, determine the probability of observing n' or more gRNAs of the sequence of interest by chance in the selected cells, based on the probability of observing n' gRNAs of the target region by chance in the selected cells; and identify that the sequence of interest is positively selected, based on the probability of observing n' or more gRNAs of the sequence of interest by chance in the selected cells.
[0206] Embodiment 68. The device according to Embodiment 67, wherein the first cell population and the second cell population comprise Cas-9 positive cells.
[0207] Embodiment 69. The apparatus according to any one of embodiments 67-68, wherein the processor-executable instructions, when executed by the one or more processors, cause the apparatus to determine, for each of a plurality of target regions of DNA, the corresponding number (n) of gRNAs present in each of the plurality of target regions of DNA for which the read count exceeds a background threshold, based on sequencing performed on a first cell population after infection with a vector containing a library of at least 3 gRNAs, such that the apparatus: receives first read count data generated by infecting the first cell population with the library of viral vectors and sequencing the cells of the first cell population to obtain a read count for each of the gRNAs; and based on the first read count data, for each of the plurality of target regions of DNA, counts each gRNA for which the read count exceeds the background threshold.
[0208] Embodiment 70. The apparatus according to embodiment 69, wherein the processor-executable instructions, when executed by the one or more processors, cause the apparatus to determine, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA, the total number (N) of gRNAs present in the first cell population in all of the plurality of target regions of DNA for which the read count exceeds the background threshold, such that the apparatus counts each gRNA present in the first cell population for which the read count exceeds the background threshold.
[0209] Embodiment 71. The apparatus according to any one of embodiments 67-70, wherein the processor-executable instructions, when executed by the one or more processors, cause the apparatus to determine, for each of a plurality of target regions of DNA, the corresponding number (n') of gRNAs present in each of the plurality of target regions of DNA for which the read count exceeds the background threshold, based on sequencing performed on a second cell population after infection with a vector containing the library of at least 3 gRNAs, such that the apparatus: receives second read count data generated by sorting the cells of the second cell population into those having a specified phenotype or not having the specified phenotype; selecting the cells having the specified phenotype; sequencing the selected cells to obtain a read count for each of the gRNAs; and based on the second read count data, for each of the plurality of target regions of DNA, counts each gRNA for which the read count exceeds the background threshold.
[0210] Embodiment 72. The apparatus according to any one of embodiments 67-71, wherein the processor-executable instructions, when executed by the one or more processors, cause the apparatus to determine the total number (N') of gRNAs present in the second cell population in all target regions of the DNA for which the read count exceeds the background threshold, based on the respective number of gRNAs for each of the plurality of target regions of the DNA, such that the apparatus counts each gRNA present in the selected cells of the second cell population for which the read count exceeds the background threshold.
[0211] Embodiment 73. The apparatus according to any one of embodiments 67-72, wherein the processor-executable instructions, when executed by the one or more processors, cause the apparatus to determine, for each target region of the plurality of target regions, the probability of randomly observing n' gRNAs of the target region in the selected cells, based on n, N, n', and N', such that the apparatus evaluates where Calculate the number of ways to choose x objects from y objects.
[0212] Embodiment 74. The apparatus according to any one of embodiments 67-73, wherein the processor-executable instructions, when executed by the one or more processors, cause the apparatus to determine, for a target region containing a sequence of interest, the probability of randomly observing n' or more gRNAs of the sequence of interest in the selected cells, based on the probability of randomly observing n' gRNAs of the target region in the selected cells, such that the apparatus evaluates
[0213] Embodiment 75. The apparatus according to any one of embodiments 67-74, wherein the processor-executable instructions, when executed by the one or more processors, cause the apparatus to identify that the sequence of interest is positively selected based on the probability of randomly observing n' or more gRNAs of the sequence of interest in the selected cells, such that the apparatus evaluates and determines that the probability of randomly observing n' or more gRNAs of the sequence of interest in the selected cells meets a threshold.
[0214] Embodiment 76. The apparatus according to any one of embodiments 67-76, further comprising processor-executable instructions that, when executed by the one or more processors, cause the apparatus to determine an enrichment score for each gRNA.
[0215] Embodiment 77. The apparatus according to embodiment 76, wherein the processor-executable instructions, when executed by the one or more processors, cause the apparatus to determine an enrichment score for each gRNA, and the processor-executable instructions cause the apparatus to evaluate N / N'.
[0216] Embodiment 78. The device according to any one of embodiments 67-77, wherein the cas-9 positive cells comprise inactivated cas-9.
[0217] Embodiment 79. The device according to embodiment 78, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0218] Embodiment 80. A non-transitory computer-readable medium for determining the probability of accidentally observing one or more guide RNAs (gRNAs), the non-transitory computer-readable medium storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to: determine, for each of a plurality of target regions of DNA, the corresponding number of gRNAs (n) present in each of the plurality of target regions of DNA for which the read count exceeds a background threshold, based on sequencing of a first cell population after infection with a vector comprising a library of at least 3 guide RNAs (gRNAs); determine the total number of gRNAs (N) present in the first cell population in all of the target regions of DNA for which the read count exceeds the background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA; determine, for each of a plurality of target regions of DNA, the corresponding number of gRNAs (n') present in each of the plurality of target regions of DNA for which the read count exceeds the background threshold, based on sequencing of a second cell population after infection with a vector comprising the library of at least 3 gRNAs; determine the total number of gRNAs (N') present in the second cell population in all of the target regions of DNA for which the read count exceeds the background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA; for each target region of the plurality of target regions, determine the probability of accidentally observing n' gRNAs of the target region in the selected cells based on n, N, n', and N'; for a target region comprising a sequence of interest, determine the probability of accidentally observing n' or more gRNAs of the sequence of interest in the selected cells based on the probability of accidentally observing n' gRNAs of the target region in the selected cells; and
[0219] Identify that the sequence of interest is positively selected based on the probability of accidentally observing n' or more gRNAs of the sequence of interest in the selected cells.
[0220] Embodiment 81. The non-transitory computer-readable medium according to embodiment 80, wherein the first cell population and the second cell population comprise cas-9 positive cells.
[0221] Embodiment 82. The non-transitory computer-readable medium according to any one of Embodiments 80-81, wherein the processor-executable instructions, when executed by the one or more processors, cause the device to determine, for each of a plurality of target regions of DNA, the corresponding number of gRNAs (n) present in each of the plurality of target regions of DNA for which the read count exceeds a background threshold, after infecting a first cell population with a vector containing a library of at least 3 gRNAs, such that the one or more processors: receive first read count data generated by infecting the first cell population with the library of viral vectors and sequencing the cells of the first cell population to obtain the read count for each of the gRNAs; and count, for each of the plurality of target regions of DNA, each gRNA for which the read count exceeds the background threshold, based on the first read count data.
[0222] Embodiment 83. The non-transitory computer-readable medium according to any one of Embodiments 80-82, wherein the processor-executable instructions, when executed by the one or more processors, cause the device to determine the total number (N) of gRNAs present in the first cell population in all of the target regions of DNA for which the read count exceeds the background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of DNA, such that the one or more processors count each gRNA present in the first cell population for which the read count exceeds the background threshold.
[0223] Embodiment 84. The non-transitory computer-readable medium according to any one of Embodiments 80-83, wherein the processor-executable instructions, when executed by the one or more processors, cause the device to determine, for each of a plurality of target regions of the DNA, the corresponding number of gRNAs (n') present in each of the plurality of target regions of the DNA for which the read count exceeds the background threshold, based on sequencing performed on a second cell population after infection with a vector comprising the library of at least 3 gRNAs, such that the one or more processors: receive second read count data generated by classifying the cells of the second cell population as having a specified phenotype or not having the specified phenotype; selecting the cells having the specified phenotype; sequencing the selected cells to obtain the read count for each of the gRNAs; and counting, for each of the plurality of target regions of the DNA, each gRNA for which the read count exceeds the background threshold based on the second read count data. Embodiment 85. The non-transitory computer-readable medium according to any one of Embodiments 80-84, wherein the processor-executable instructions, when executed by the one or more processors, cause the device to determine, for each of the plurality of target regions of the DNA, the total number (N) of gRNAs present in the second cell population in all of the target regions of the plurality of target regions of the DNA for which the read count exceeds the background threshold, based on the corresponding number of gRNAs for each of the plurality of target regions of the DNA, such that the one or more processors count each gRNA present in the selected cells of the second cell population for which the read count exceeds the background threshold.
[0224] Embodiment 86. The non-transitory computer-readable medium according to any one of Embodiments 80-85, wherein the processor-executable instructions, when executed by the one or more processors, cause the device to determine, for each target region of the plurality of target regions, the probability of randomly observing n' gRNAs of the target region in the selected cells, based on n, N, n', and N', such that the one or more processors evaluate where Calculate the number of ways to select x objects from y objects.
[0225] Embodiment 87. The non-transitory computer-readable medium according to any one of Embodiments 80-86, wherein the processor-executable instructions, when executed by the one or more processors, cause the device to determine, for a target region comprising a sequence of interest, the probability of randomly observing n' or more gRNAs of the sequence of interest in the selected cells, based on the probability of randomly observing n' gRNAs of the target region in the selected cells, such that the one or more processors evaluate
[0226] Embodiment 88. The non-transitory computer-readable medium according to any one of Embodiments 80-87, wherein the processor-executable instructions, when executed by the one or more processors, cause the device to identify that the sequence of interest is positively selected based on the probability of observing n' or more gRNAs of the sequence of interest in a selected cell, such that the one or more processors evaluate and determine that the probability of observing n' or more gRNAs of the sequence of interest in a selected cell meets a threshold.
[0227] Embodiment 89. The non-transitory computer-readable medium according to any one of Embodiments 80-88, further comprising processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to determine an enrichment score for each gRNA.
[0228] Embodiment 90. The non-transitory computer-readable medium according to Embodiment 89, wherein the processor-executable instructions, when executed by the one or more processors, cause the device to determine an enrichment score for each gRNA including causing the one or more processors to evaluate N / N'.
[0229] Embodiment 91. The non-transitory computer-readable medium according to any one of Embodiments 80-90, wherein the cas-9 positive cells contain inactivated cas-9.
[0230] Embodiment 92. The non-transitory computer-readable medium according to Embodiment 91, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0231] Those skilled in the art will recognize or be able to use, without more than routine experimentation, many equivalents to the specific embodiments of the methods and compositions described herein. Such equivalents are intended to be encompassed by the following claims.
Claims
1. A method, comprising: (A) Infect a first culture of cas9-positive cells with a library of viral vectors, the library comprising at least 3 guide RNAs (gRNAs) for cleaving a target region of DNA within the genome of the cells; Sequence the cells to obtain read counts for each of the gRNAs; Determine the corresponding number of gRNAs (n) for each DNA target region for which the read count exceeds a background threshold; and Determine the total number of gRNAs (N) in all target regions for which the read count exceeds the background threshold; (B) Infect a second culture of cas9-positive cells with the library of viral vectors; Sort the cells of the second culture as having a specified phenotype or not having the specified phenotype; Select the cells having the specified phenotype and sequence the selected cells to obtain post-selection read counts for each of the gRNAs; Determine the corresponding number of gRNAs (n') for each DNA target region in the selected cells for which the post-selection read count exceeds the background threshold; and Determine the total number of gRNAs (N') in all target regions of the selected cells for which the read count exceeds the background threshold; (C) For a target region of DNA, calculate the probability of randomly observing n' gRNAs for the target region in the selected cells according to the following formula wherein calculate the number of ways to select x objects from y objects; and For the target region of DNA containing a gene, according to the formula to calculate the probability of observing n ′ or more gRNAs of the gene incidentally in the selected cells; and Identifying that the gene is under positive selection based on the probability of observing n or more gRNAs of the gene incidentally in the selected cells ′ or more gRNAs of the gene in the selected cells 2. A method, comprising: (A) Infect a first culture of cas9-positive cells with a library of viral vectors, the library comprising at least 3 guide RNAs (gRNAs) for enhancing transcription of a target region of DNA within the genome of the cells; Sequence the cells to obtain read counts for each of the gRNAs; Determine the corresponding number of gRNAs (n) for each DNA target region for which the read count exceeds a background threshold; and Determine the total number of gRNAs (N) in all target regions for which the read count exceeds the background threshold; (B) Infect a second culture of cas9-positive cells with the library of viral vectors; Sort the cells of the second culture as having a specified phenotype or not having the specified phenotype; Select the cells having the specified phenotype and sequence the selected cells to obtain post-selection read counts for each of the gRNAs; Determine the corresponding number of gRNAs (n') for each DNA target region in the selected cells for which the post-selection read count exceeds the background threshold; and Determine the total number of gRNAs (N') in all target regions of the selected cells for which the read count exceeds the background threshold; (C) For a target region of DNA, calculate the probability of randomly observing n' gRNAs for the target region in the selected cells according to the following formula wherein calculate the number of ways to select x objects from y objects; and For the target region of DNA containing a gene, according to the formula to calculate the probability of observing n ′ or more gRNAs of the gene incidentally in the selected cells; and Identifying that the gene is under positive selection based on the probability of observing n or more gRNAs of the gene incidentally in the selected cells ′ or more gRNAs of the gene incidentally in the selected cells 3. The method according to any one of the preceding claims, wherein the first culture is infected according to CRISPR technology.
4. The method according to claim 1 or 2, wherein the cas-9 positive cells of the first culture are modified to contain one or more selectable markers.
5. The method according to claim 1 or 2, wherein the specified phenotype is fluorescence.
6. The method according to claim 1 or 2, wherein the specified phenotype is cell survival.
7. The method according to claim 4, wherein the one or more selective markers include fluorescent markers.
8. The method according to claim 7, wherein the fluorescent marker is part of a FRET biosensor.
9. The method according to claim 4, wherein the one or more selective markers are detectable enzymes.
10. The method according to claim 9, wherein the detectable enzyme is β-galactosidase.
11. The method according to claim 9, wherein the detectable enzyme is luciferase.
12. The method according to claim 1 or 2, wherein the target region contains a gene.
13. The method according to claim 1 or 2, wherein classifying the cells of the second culture as having or not having the designated phenotype comprises applying a selection mechanism to the second cell culture.
14. The method according to claim 13, wherein the selection mechanism comprises one or more of: exposing a second cell population to a drug, or exposing the second cell population to a substance that recognizes protein activity or expression level.
15. The method according to claim 4, wherein selecting cells having the designated phenotype comprises sorting the cells based on the one or more selective markers.
16. The method according to claim 1 or 2, further comprising identifying that the target region is positively selected based on the probability of incidentally observing n ′ or more gRNAs of the gene in the selected cells.
17. The method according to claim 16, wherein identifying that the target region is positively selected comprises determining that the probability of incidentally observing n ′ or more gRNAs of the sequence of interest in the selected cells meets a threshold.
18. The method according to claim 1 or 2, further comprising determining an enrichment score for each gRNA.
19. The method according to claim 18, wherein determining the enrichment score for each gRNA comprises evaluating N / N’.
20. The method according to claim 2, wherein the cas-9 positive cells contain inactivated cas-9.
21. The method according to claim 20, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
22. The method according to claim 1 or 2, wherein the target region of the DNA regulates a downstream gene or protein.
23. The method according to claim 22, wherein modulating the downstream gene or protein comprises activation or inhibition of the downstream gene or protein.
24. The method according to claim 1 or 2, further comprising identifying the target region as a modifier of a second gene based on the probability of observing n or more gRNAs of the gene incidentally in the selected cells. ′ in the selected cells.
25. The method according to claim 1 or 2, further comprising identifying the target region as a therapeutic target based on the probability of observing n or more gRNAs of the gene incidentally in the selected cells. ′ in the selected cells.
26. The method according to claim 1 or 2, further comprising identifying the target region as being associated with the specified phenotype based on the probability of observing n or more gRNAs of the gene incidentally in the selected cells. ′ in the selected cells.
27. The method according to claim 1 or 2, further comprising identifying the target region as exhibiting a protective effect based on the probability of observing n or more gRNAs of the gene incidentally in the selected cells. ′ in the selected cells.
28. A method comprising: Determine, for each of a plurality of target regions of DNA, the corresponding number of gRNAs (n) present in each of the plurality of target regions of DNA for which the read count exceeds a background threshold, based on sequencing of a first population of cells after infection with a vector of a library comprising at least 3 guide RNAs (gRNAs); Determine the total number (N) of gRNAs present in the first cell population in all target regions of the DNA for which the read count exceeds the background threshold, based on the respective number of gRNAs for each of the multiple target regions of the DNA; Determine the respective number of gRNAs (n') present in each of the multiple target regions of the DNA for which the read count exceeds the background threshold, based on sequencing of a second cell population after infection with a vector of the library comprising the at least 3 gRNAs; Determine the total number (N') of gRNAs present in the second cell population in all target regions of the DNA for which the read count exceeds the background threshold, based on the respective number of gRNAs for each of the multiple target regions of the DNA; For each target region of the multiple target regions, determine the probability of observing n' gRNAs of the target region in the selected cells, according to the following formula and based on n, N, n' and N'; wherein calculate the number of ways to select x objects from y objects; For a target region containing the sequence of interest, according to the formula and determine the probability of observing n or more gRNAs of the sequence of interest by chance in the selected cells based on the probability of observing n' gRNAs of the target region by chance in the selected cells ′ ; and Identifying a sequence of interest as being positively selected based on the probability of observing n or more gRNAs of the sequence of interest by chance in a selected cell ′ or more gRNAs of the sequence of interest in a selected cell 29. The method according to claim 28, wherein the first cell population and the second cell population comprise cas-9 positive cells.
30. The method according to claim 29, wherein for each of a plurality of target regions of DNA, determining the number of corresponding gRNAs (n) present in each of the plurality of target regions of DNA for which the read count exceeds a background threshold, based on sequencing of the first cell population after infection with a vector comprising a library of at least 3 gRNAs, comprises: Infect the first cell population with a library of viral vectors; Sequence the cells of the first cell population to obtain a read count for each of the gRNAs; And For each of the multiple target regions of the DNA, count each gRNA for which the read count exceeds the background threshold.
31. The method according to claim 30, wherein determining the total number (N) of gRNAs present in the first cell population among all target regions of the DNA whose read count exceeds the background threshold, based on the respective number of gRNAs for each of the plurality of target regions of the DNA, includes counting each gRNA present in the first cell population whose read count exceeds the background threshold.
32. The method according to claim 31, wherein determining the respective number of gRNAs (n') present in each of the plurality of target regions of the DNA whose read count exceeds the background threshold, based on sequencing of a second cell population after infection with a vector containing the library of at least 3 gRNAs, for each of the plurality of target regions of the DNA, includes: Classify the cells of the second cell population as having a specified phenotype or not having the specified phenotype; Select the cells having the specified phenotype; Sequence the selected cells to obtain a read count for each of the gRNAs; And For each of the multiple target regions of the DNA, count each gRNA for which the read count exceeds the background threshold.
33. The method according to any one of claims 28 - 32, wherein determining the total number (N') of gRNAs present in the second cell population among all target regions of the DNA whose read count exceeds the background threshold, based on the respective number of gRNAs for each of the plurality of target regions of the DNA, includes counting each gRNA present in the selected cells of the second cell population whose read count exceeds the background threshold.
34. The method according to any one of claims 28 - 32, wherein for each target region of the plurality of target regions, determining the probability of randomly observing n' gRNAs of the target region in the selected cells based on n, N, n', and N' includes evaluating where calculating the number of ways to select x objects from y objects.
35. The method according to any one of claims 28 - 32, wherein for a target region containing the sequence of interest, determining the probability of randomly observing n ′ or more gRNAs of the sequence of interest in the selected cells, based on the probability of randomly observing n' gRNAs of the target region in the selected cells, includes evaluating 36. The method according to any one of claims 28 - 32, wherein identifying that the sequence of interest is positively selected based on the probability of randomly observing n ′ or more gRNAs of the sequence of interest in the selected cells includes determining that the probability of randomly observing n ′The probability of one or more gRNAs meets a threshold.
37. The method according to any one of claims 28-32, further comprising determining an enrichment score for each gRNA.
38. The method according to claim 37, wherein determining an enrichment score for each gRNA comprises evaluating N / N'.
39. The method according to any one of claims 29-32, wherein the cas-9 positive cells comprise inactivated cas-9.
40. The method according to claim 39, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
41. An apparatus, comprising: One or more processors; And A memory storing processor-executable instructions that, when executed by the one or more processors, cause the apparatus to: (A) Receive first read count data, wherein the first read count data is generated by: Infecting a first culture of cas9-positive cells with a library of viral vectors, the library comprising at least 3 guide RNAs (gRNAs) for cleaving target regions of DNA within the genome of the cells; And Sequencing the cells to obtain a read count for each of the gRNAs; Determine the respective number of gRNAs (n) for each DNA target region for which the read count exceeds the background threshold, based on the first read count data; and Determine the total number of gRNAs (N) in all target regions for which the read count exceeds the background threshold, based on the first read count data; (B) Receive second read count data, wherein the second read count data is generated by: Infecting a second culture of cas9-positive cells with the library of viral vectors, Classifying the cells of the second culture as having a specified phenotype or not having the specified phenotype, and Select cells having the specified phenotype and sequence the selected cells to obtain post-selection read counts for each of the gRNAs; Based on the second read count data, determine the corresponding number of gRNAs (n') for each DNA target region in the selected cells where the post-selection read count exceeds the background threshold; and Based on the second read count data, determine the total number of gRNAs (N') in all target regions of the selected cells where the read count exceeds the background threshold; (C) For a target region of DNA, calculate the probability of randomly observing n' gRNAs for the target region in the selected cells according to the following formula wherein calculate the number of ways to select x objects from y objects; and For a target region of DNA containing a gene, according to the formula to calculate the probability of observing n ′ or more gRNAs of said gene in the selected cells; and Identifying that the gene is under positive selection based on the probability of observing n or more gRNAs of the gene incidentally in the selected cells ′ or more gRNAs of the gene incidentally in the selected cells 42. The apparatus according to claim 41, wherein the first culture is infected according to CRISPR technology.
43. The apparatus according to any one of claims 41-42, wherein the cas-9 positive cells of the first culture are modified to contain one or more selectable markers.
44. The apparatus according to any one of claims 41-42, wherein the target region contains a gene.
45. The device according to any one of claims 41-42, wherein classifying the cells of the second culture as having the designated phenotype or not having the designated phenotype comprises applying a selection mechanism to the second cell culture.
46. The device according to claim 45, wherein the selection mechanism comprises one or more of the following: exposing a second cell population to a drug, or exposing the second cell population to a substance that recognizes protein activity or expression level.
47. The device according to any one of claims 41-42, wherein selecting cells having the designated phenotype comprises sorting the cells according to the one or more selection markers.
48. The device according to any one of claims 41-42, further comprising identifying that the target region is positively selected based on the probability of accidentally observing n ′ or more gRNAs of the gene in the selected cells.
49. The device according to claim 48, wherein identifying that the target region is positively selected comprises determining that the probability of accidentally observing n ′ or more gRNAs of the sequence of interest in the selected cells meets a threshold.
50. The device according to any one of claims 41-42, further comprising determining an enrichment score for each gRNA.
51. The device according to claim 50, wherein determining the enrichment score for each gRNA comprises evaluating N / N'.
52. The device according to any one of claims 41-42, wherein the cas-9 positive cells comprise inactivated cas-9.
53. The device according to claim 52, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
54. A non-transitory computer-readable medium for determining the probability of accidentally observing one or more guide RNAs (gRNAs), the non-transitory computer-readable medium storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to: (A) Receive first read count data, wherein the first read count data is generated by: Infecting a first culture of cas9 positive cells with a library of viral vectors, the library comprising at least 3 guide RNAs (gRNAs) for cleaving a target region of DNA within the genome of the cells; and Sequence the cells to obtain read counts for each of the gRNAs; Based on the first read count data, determine the corresponding number of gRNAs (n) for each DNA target region where the read count exceeds the background threshold; and Based on the first read count data, determine the total number of gRNAs (N) in all target regions where the read count exceeds the background threshold; (B) Receive second read count data, where the second read count data is generated by infecting a second culture of cas9-positive cells with the library of viral vectors, sorting the cells of the second culture as having the specified phenotype or not having the specified phenotype, and selecting cells having the specified phenotype and sequencing the selected cells to obtain post-selection read counts for each of the gRNAs; Based on the second read count data, determine the corresponding number of gRNAs (n') for each DNA target region in the selected cells where the post-selection read count exceeds the background threshold; and Based on the second read count data, determine the total number of gRNAs (N') in all target regions of the selected cells where the read count exceeds the background threshold; (C) For a target region of DNA, calculate the probability of randomly observing n' gRNAs for the target region in the selected cells according to the following formula wherein calculate the number of ways to select x objects from y objects; and For a target region of DNA containing a gene, according to the formula to calculate the probability of observing n ′ or more gRNAs of the gene in the selected cells; Identifying that the gene is under positive selection based on the probability of observing n or more gRNAs of the gene incidentally in the selected cells ′ or more gRNAs of the gene incidentally in the selected cells 55. The non-transitory computer-readable medium according to claim 54, wherein the first culture is infected according to CRISPR technology.
56. The non-transitory computer-readable medium according to any one of claims 54-55, wherein the cas-9 positive cells of the first culture are modified to contain one or more selectable markers.
57. The non-transitory computer-readable medium according to any one of claims 54-55, wherein the target region contains a gene.
58. The non-transitory computer-readable medium according to any one of claims 54-55, wherein classifying the cells of the second culture as having or not having the specified phenotype comprises applying a selection mechanism to the second cell culture.
59. The non-transitory computer-readable medium according to claim 58, wherein the selection mechanism comprises one or more of the following: exposing a second cell population to a drug, or exposing the second cell population to a substance that recognizes protein activity or expression level.
60. The non-transitory computer-readable medium according to any one of claims 54-55, wherein selecting cells having the specified phenotype comprises sorting the cells according to the one or more selectable markers.
61. The non-transitory computer-readable medium according to any one of claims 54-55, further comprising identifying that the target region is positively selected based on the probability of observing n or more gRNAs of the gene incidentally in the selected cells. ′ or more gRNAs of the gene incidentally in the selected cells.
62. The non-transitory computer-readable medium according to claim 61, wherein identifying that the target region is positively selected comprises determining that the probability of observing n or more gRNAs of the sequence of interest incidentally in the selected cells meets a threshold. ′ or more gRNAs of the sequence of interest incidentally in the selected cells meets a threshold.
63. The non-transitory computer-readable medium according to any one of claims 54-55, further comprising determining an enrichment score for each gRNA.
64. The non-transitory computer-readable medium according to claim 63, wherein determining the enrichment score for each gRNA comprises evaluating N / N'.
65. The non-transitory computer-readable medium according to any one of claims 54-55, wherein the cas-9 positive cells comprise inactivated cas-9.
66. The non-transitory computer-readable medium according to claim 65, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
67. An apparatus, comprising: one or more processors; and a memory storing processor-executable instructions that, when executed by the one or more processors, cause the apparatus to perform the method according to any one of claims 28-40.
68. A computer-readable medium having computer-executable instructions adapted to cause a computer system to perform the method according to any one of claims 28-40.