Method for selecting crispr
The method enhances CRISPR selection by accounting for variability in gRNA infection and efficiency, enabling accurate identification of genes associated with phenotypes through probabilistic analysis of read counts and threshold-based sequencing.
Patent Information
- Application Number
- JP2025044318
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-03-18
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing CRISPR/Cas9 technology faces challenges in addressing variability between experiments due to inconsistent gRNA viral titer, infection rates, and target efficiency, leading to unreliable identification of genes associated with phenotypes.
A method involving infecting cell cultures with a gRNA library, sequencing to determine read counts exceeding a threshold, categorizing cells by phenotype, and calculating the probability of gRNA occurrence to identify positively selected genes, using formulas to account for random observations.
This approach provides a robust method to identify genes associated with phenotypes by accounting for variability, improving the accuracy of CRISPR selection and reducing noise in computational analysis.
Smart Images

Figure 2025098101000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to methods and systems for CRISPR selection. Cross - reference to related applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 820,106, filed on March 18, 2019, the entire content of which is incorporated herein by reference.
Background Art
[0002] Clustered regularly interspaced short palindromic repeat (CRISPR) - Cas9 technology has revolutionized genome engineering. In this system, a guide RNA (gRNA) induces double - strand breaks in a target genomic region in Cas9 nuclease. The 5' end of the gRNA contains a nucleotide sequence of about 20 nucleotides complementary to the target region. When double - strand breaks are repaired by non - homologous end joining (NHEJ), insertions and deletions occur frequently, thus efficiently knocking out the target genomic locus. The development of lentiviral delivery methods has enabled the creation of libraries of genome - scale CRISPR / Cas9 knockouts. These libraries allow both negative and positive selection screening to be performed on mammalian cell lines. In a CRISPR / Cas9 knockout screen, each gene is targeted by several gRNAs, and mutant pools carrying different gene knockouts can be determined by high - throughput sequencing. CRISPR activation (CRISPRa) can also be used with a gRNA library where the activated genes can be determined by high - throughput sequencing.
[0003] Genome-wide CRISPR / Cas9 knockout or gene activation technology is an effective gene perturbation screening technology. The goal is to identify gRNAs related to phenotypes and thus the corresponding disease genes. However, the data generated by these screens pose several challenges for computational analysis. CRISPR studies are often carried out using multiple replicates. CRISPR is susceptible to variability where each experiment may not use the same gRNA viral titer within the screening library, the lentiviral infection rate may vary between experiments, and the gRNAs may not target genes with the same efficiency between experiments. Therefore, the observed gRNA abundances are highly variable between experiments even when using cells of the same phenotype. Existing techniques rely on read counts to identify gRNAs related to phenotypes, specifically using the mean and variance of normalized gRNA read counts to test whether the abundances of gRNAs are significantly different between cells with or without the phenotype.
[0004] However, these techniques do not address the problem of variability between experiments described above and rather assume a high degree of homogeneity between pre-selection and post-selection experiments. These techniques do not address variability within a single CRISPR experiment and / or between CRISPR experiments.
Summary of the Invention
Problems to be Solved by the Invention
[0005] Therefore, there is a need for technical improvements in computing technologies to address the problem of CRISPR variability when identifying genes through positive and negative selection screens.
Means for Solving the Problems
[0006] It should be understood that both the following overview and the following embodiments for carrying out the invention are merely illustrative and explanatory and not limiting. In one embodiment, the method includes: (A) infecting a first culture of cas9-positive cells with a library of viral vectors, the library including at least three guide RNAs (gRNAs) for cleaving target regions of DNA within the genome of the cells, sequencing the cells to obtain the read count for each of the gRNAs, summing (Σ) for each target region of DNA the number of each gRNA for which the read count exceeds a background threshold, where Σ=n, and summing (Σ) the total number of gRNAs, where Σ=N. The method includes: (B) infecting a second culture of cas9-positive cells with a library of viral vectors, categorizing the cells of the second culture as either having a specified phenotype or not having the specified phenotype, selecting the cells having the specified phenotype, sequencing the selected cells to obtain the post-selection read count for each of the gRNAs, summing (Σ) for each target region of DNA the number of each gRNA in the selected cells for which the post-selection read count exceeds the background threshold, where Σ=n’, and summing (Σ) for the entire target region the total number of gRNAs in the selected cells for which the read count exceeds the threshold, where Σ=N’. The method includes: (C) calculating, for a target region of DNA, the probability of randomly observing n’ gRNAs in the target region in the selected cells according to the formula
[0007]
Number
[0008] wherein
[0009]
Number
[0010] calculates the number of ways to select x objects from y objects, and for a target region of DNA containing a gene, according to the formula
[0011]
Number
[0012] including calculating the probability of randomly observing n' or more gRNAs of a gene in a selected cell according to [the relevant content]. In one embodiment, the method comprises determining, based on sequencing of a first cell population after infection with a vector comprising a library of at least three guide RNAs (gRNAs) for each of a plurality of target regions of DNA, the respective number (n) of gRNAs present for each of the plurality of target regions of DNA, wherein the read count exceeds a background threshold; determining, based on the respective number of gRNAs for each of the plurality of target regions of DNA, the total number (N) of gRNAs present in the first cell population across all target regions of the plurality of target regions of DNA, wherein the read count exceeds a background threshold; determining, based on sequencing of a second cell population after infection with a vector comprising a library of at least three guide RNAs (gRNAs) for each of a plurality of target regions of DNA, the respective number (n') of gRNAs present for each of the plurality of target regions of DNA, wherein the read count exceeds a background threshold; determining, based on the respective number of gRNAs for each of the plurality of target regions of DNA, the total number (N') of gRNAs present in the second cell population across all target regions of the plurality of target regions of DNA, wherein the read count exceeds a background threshold; determining, based on n, N, n', and N', for each target region of the plurality of target regions, the probability of randomly observing n' gRNAs for the target region in the selected cell; determining, for a target region containing the sequence of interest, the probability of randomly observing n' or more gRNAs for the target region in the selected cell based on the probability of randomly observing n' gRNAs for the target region in the selected cell; and identifying that the sequence of interest is positively selected based on the probability of randomly observing n' or more gRNAs for the sequence of interest in the selected cell.
[0013] Additional advantages will be described in part below or will be apparent from practice. These advantages will be realized and achieved by the elements and combinations particularly pointed out in the appended claims.
Advantages of the Invention
[0014] As described above, according to the present invention, a method and a system for CRISPR selection can be provided.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
[0016] The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate embodiments and, together with the description, serve to explain the principles of the method and system. Prior to the disclosure and description of the present method and system, it should be understood that the present method and system are not limited to a particular method, a particular component, or a particular implementation form. It should also be understood that the terms used in this specification are for the purpose of describing particular embodiments only and are not intended to be limiting.
[0017] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise.
[0018] As used herein, the terms "probe" and "guide RNA (gRNA)" and "guide" are used interchangeably. In one embodiment, the gRNA may also be provided in the form of DNA encoding the gRNA.
[0019] As used herein, a "Cas protein" can be a wild-type protein (i.e., one that occurs in nature), a modified Cas protein (i.e., a Cas protein variant), or a fragment of a wild-type or modified Cas protein. The Cas protein can also be an active variant or fragment with respect to the catalytic activity of a wild-type or modified Cas protein.
[0020] In a first aspect, the present disclosure features a method of identifying, for example, a gene or gene product that regulates the expression of other genes or gene products. The method is useful, for example, for demonstrating positive selection after perturbation using a CRISPR guide.
[0021] In one embodiment (shown in FIG. 1), the method comprises step 110 of infecting a first culture of cas9-positive cells with a library of viral vectors, the library comprising at least three guide RNAs for cleaving a target region of DNA within the genome of the cells; step 120 of sequencing the cells to obtain the read count for each of the gRNAs; step 130 of summing (Σ) (Σ=n) the number of each of the gRNAs for which the read count exceeds a background threshold per target region of DNA, and summing (Σ) (Σ=N) the total number of gRNAs for which the read count exceeds the background threshold across all target regions; step 140 of infecting a second culture of cas9-positive cells with a library of viral vectors; step 150 of categorizing the cells of the second culture as either having a designated phenotype or not having a designated phenotype; step 160 of selecting the cells having the designated phenotype, sequencing the selected cells to obtain the post-selection read count for each of the gRNAs; step 170 of summing (Σ) (Σ=n’) the number of each of the gRNAs for which the post-selection read count exceeds the background threshold in the selected cells per target region of DNA, and summing (Σ) (Σ=N’) the total number of gRNAs for which the read count exceeds the threshold across all target regions in the selected cells; step 180 of calculating, for a target region of DNA, the probability of randomly observing n’ gRNAs in the selected cells, according to the formula
[0022] [Number]
[0023] wherein step 180 is a step of calculating the probability of randomly observing n’ gRNAs in the target region in the selected cells, wherein in the formula
[0024] [Number]
[0025] is a step of calculating, calculating the number of ways to select x objects from y objects, and for a target region of DNA containing a gene, according to the formula
[0026] [Number]
[0027] step 190 of calculating the probability of accidentally observing n' or more gRNAs of genes in the cells selected according to [Number], In one embodiment (also shown in FIG. 1), the method is step 110 of infecting a first culture of cas9-positive cells with a library of viral vectors, the library comprising at least three guide RNAs (gRNAs) for enhancing the transcription of target regions of DNA in the genome of the cells; the step of infecting, step 120 of sequencing the cells to obtain the read count of each of the gRNAs; step 130 of summing (Σ) (Σ=n) the number of each of the gRNAs for which the read count exceeds the background threshold per target region of DNA, and summing (Σ) (Σ=N) the total number of gRNAs for which the read count exceeds the background threshold over the entire target region; step 140 of infecting a second culture of cas9-positive cells with a library of viral vectors; step 150 of categorizing the cells of the second culture as either having a designated phenotype or not having a designated phenotype; step 160 of selecting the cells having the designated phenotype, sequencing the selected cells to obtain the post-selection read count of each of the gRNAs; step 170 of summing (Σ) (Σ=n') the number of each of the gRNAs in the selected cells for which the post-selection read count exceeds the background threshold per target region of DNA, and summing (Σ) (Σ=N') the total number of gRNAs in the selected cells for which the read count exceeds the threshold over the entire target region; for the target region of DNA, the formula
[0028] [Number]
[0029] according to [Number], step 180 of calculating the probability of accidentally observing n' gRNAs for the target region in the selected cells, where in the formula
[0030] [Number]
[0031] is a step of calculating and computing the number of ways to select an x object from a y object, and for a target region of DNA containing a gene, the formula
[0032]
Number
[0033] including step 190 of calculating the probability of accidentally observing n' or more gRNAs of a gene in a cell selected according to In one embodiment, the first culture can be infected according to CRISPR technology. In one embodiment, infecting a first culture of cas9-positive cells with a library of viral vectors (e.g., FIG. 2; 201), and infecting a second culture of cas9-positive cells with a library of viral vectors (e.g., FIG. 2; 205) includes using a CRISPR gRNA library. The CRISPR gRNA library may be a knockout library, for example, a genome-wide gRNA knockout library that includes one or more gRNAs (e.g., sgRNAs) targeting each gene in the genome, where the genome can be any type of genome. In some embodiments, the gRNA library may include a pooled library. Non-limiting examples of pooled libraries include the genome-scale CRISPR knockout (GeCKO) library. See, for example, Shalem O et al. (2014) Science 343:84-7 and Sanjana NE et al. (2014) Nat. Methods 11:783-4. The gRNAs in the library can target any number of target regions (e.g., genes) in the DNA. For example, the gRNA can target about 50 or more genes, about 100 or more genes, about 200 or more genes, about 300 or more genes, about 400 or more genes, about 500 or more genes, about 1000 or more genes, about 2000 or more genes, about 3000 or more genes, about 4000 or more genes, about 5000 or more genes, about 10000 or more genes, or about 20000 or more genes. In some libraries, the gRNAs can be selected to target genes in a specific signaling pathway. The gRNA library can be administered using a wide range of multiplicity of infection (MOI). In some aspects, a lower MOI can be used to promote infection resulting in one gRNA per cell.
[0034] In one embodiment, the Cas-positive cells may contain a Cas protein for cleaving a target region of DNA, or a Cas protein for regulating transcription (e.g., enhancing or suppressing transcription). The Cas protein for cleaving a target region of DNA may contain an RNA-binding domain and a nuclease domain. The Cas protein for regulating transcription is inactivated so as to no longer have nuclease activity. The inactivated Cas (e.g., dCas-9) may be fused to a transcriptional activator or a transcriptional repressor. Thus, in one embodiment, Cas-9 positive cells containing wild-type active Cas-9 or inactivated Cas-9 are disclosed. The inactivated Cas-9 may be fused to, for example, at least one transcriptional activation domain. One or more gRNAs can bind to a target sequence upstream of the transcription start site of a target gene. Instead of cleaving the DNA, the dCas-9 and the transcriptional regulator can play a role in either activating or suppressing the transcription of the target gene. When the dCas-9 is fused to one or more transcriptional activators, the one or more transcriptional activators recruit transcription factors to the transcription start site of the target gene, thereby activating or upregulating the transcription. When the dCas-9 is fused to a transcriptional repressor, the transcription is inhibited, downregulated, or suppressed.
[0035] In one embodiment, the target region of DNA may be a gene. In one embodiment, the target region of DNA may be a promoter region or a regulatory factor region of a gene. In one embodiment, the target region of DNA regulates a downstream gene or protein. In one embodiment, regulating a downstream gene or protein includes activating or inhibiting the downstream gene or protein.
[0036] In some embodiments, the cells comprise a selection marker system. For example, as disclosed in the methods herein, the cas-9 positive cells of the first culture can be modified to contain one or more selection markers. The selection marker system can involve a marker protein fused or linked to a selection marker that is only activated by a regulated protein. The "marker protein" can be any protein that can be regulated, and being regulated means that the marker protein changes shape, binds to one or more other proteins or nucleic acids, has a change in activity, or has a change in expression level. The marker protein can be regulated by a gene targeted by one or more gRNAs, such that when the gene is cleaved by Cas9, the marker protein is not regulated and the selection marker is not activated. Thus, cells without an activated selection marker can be selected as cells containing a gene that regulates a marker protein fused or linked to a selection marker within the cell. For example, tau can be fused or linked to CFP. Another tau protein can be fused or linked to YFP. Upon aggregation of the tau protein, the light hitting the CFP is emitted as blue light, exciting the YFP and emitting yellow light. In the absence of tau protein aggregation, the blue light emitted from the CFP cannot excite the YFP of other tau proteins, so no yellow light is emitted. Thus, when the gene regulating tau is targeted and thus cleaved by one or more gRNAs, no yellow light is emitted. Ultimately, this gene can be identified as the gene that regulates tau (e.g., causes tau aggregation). In CRISPRa, when the gRNA binds to the target region, the downstream gene is activated, and thus, cells with overexpression or an excessive amount of the selection marker can be selected. For example, using the tau protein described above, an increased amount of yellow light can indicate the gene that regulates tau compared to cells without the gRNA.Any number of proteins present in a known pathway, specifically a disease pathway, can be used as a marker protein to regulate the marker protein and ultimately identify genes that can be found to be involved in a specific disease.
[0037] In one embodiment, the set of cells can be categorized (e.g., by phenotype), selected (e.g., FIG. 2; 206), and performed at some time after initial infection, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 days after infection, etc. In embodiments, many types of screening / selection mechanisms that utilize selectable markers contained within the set of cells can be used. In one embodiment, the selection mechanism includes exposing a second cell population to a drug or exposing the second cell population to one or more substances that identify protein activity or expression levels. In one embodiment, selecting cells having a designated phenotype includes sorting the cells based on one or more selectable markers. In one embodiment, the cell viability can be used for selection.
[0038] In one embodiment, categorizing the cells of the second culture as either having or not having a designated phenotype includes identifying the presence or absence of the designated phenotype in the cells. Categorizing the cells of the second culture as either having or not having a designated phenotype can include applying a selection mechanism to the second cell culture. In one embodiment, once the phenotype is identified, selecting cells having (or not having) the designated phenotype can be performed. The phenotype can be any observable feature or functional effect measurable in an assay such as changes in cell growth, proliferation, morphology, enzyme function, signaling, expression pattern, downstream expression pattern, reporter gene activation, hormone release, growth factor release, neurotransmitter release, ligand binding, apoptosis, and product formation. In one embodiment, the designated phenotype can be fluorescence or cell survival.
[0039] Cells can be modified to convey a phenotype that can be directly selected, for example, by genomic integration of a marker or by the presence of an intracellular marker not integrated into the genome. As used herein, a "marker" most commonly refers to a biological feature or trait that, when present (e.g., expressed) in a cell, confers an attribute or phenotype that allows that cell to be visualized or identified as containing that marker. A variety of marker types are commonly used and include, for example, visible markers such as chromogenic expression, such as lacZ complementation (β-galactosidase), or fluorescent, such as green fluorescent protein (GFP) or GFP fusion proteins, RFP, BFP, luciferase, β-galactosidase, enhanced green fluorescent protein (eGFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), blue fluorescent protein (BFP), enhanced blue fluorescent protein (eBFP), DsRed, ZsGreen, MmGFP, mPlum, mCherry, tdTomato, mStrawberry, J-Red, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, Cerulean, T-Sapphire, and expression of alkaline phosphatase, phenotypic markers (growth rate, cell morphology, colony color or colony morphology, temperature sensitivity), auxotrophic markers (growth requirements), antibiotic sensitivity and resistance, molecular markers such as biomolecules distinguishable by antigen sensitivity (e.g., blood group antigens and histocompatibility markers), cell surface markers (e.g., H2KK), enzyme markers, and nucleic acid markers such as restriction fragment length polymorphisms (RFLPs), single nucleotide polymorphisms (SNPs), and other various amplifiable genetic polymorphisms. Thus, for example, one or more selectable markers can be a detectable enzyme such as β-galactosidase or luciferase.
[0040] A "selection marker" or "screening marker" or "positive selection marker" refers to a marker that, when present in a cell, confers an attribute or phenotype that enables the selection or isolation of those cells from other cells that do not express the selection marker trait. Various genes can be used as selection markers; for example, genes encoding drug resistance or auxotrophic rescue are well known. For example, kanamycin (neomycin) resistance can be used as a trait to select bacteria that have taken up a plasmid carrying the gene encoding bacterial kanamycin resistance (e.g., the enzyme neomycin phosphotransferase II). Non-transfected cells will ultimately die when the culture is treated with neomycin or a similar antibiotic. A set of cells can be used in a drug screen to identify genes that confer drug resistance. The cells may be treated with the drug of interest, and the enriched gRNAs can be associated with genes that confer drug resistance upon mutation. Screens for resistance to viral or bacterial pathogens can be used to identify genes that prevent infection or pathogen replication. As in a drug resistance screen, survival after pathogen exposure provides a strong selection. In cancer, negative selection CRISPR screens can identify "cancer gene dependencies" in specific cancer subtypes that can provide a basis for molecularly targeted chemotherapy. For developmental studies, screening in human and mouse pluripotent cells can accurately indicate genes required for pluripotency or differentiation into distinct cell types.
[0041] In one embodiment, the cell comprises a selection marker system. The selection marker system may comprise a fluorescent protein FRET biosensor. For example, tau can be fused or linked to CFP. Another tau protein can be fused or linked to YFP. With the aggregation of tau proteins, the light hitting CFP is emitted as blue light, YFP is excited, and yellow light is emitted. In the absence of tau protein aggregation, the blue light emitted from CFP cannot excite the YFP of other tau proteins, so no yellow light is emitted. Thus, when the gene regulating tau is targeted and thus cleaved by one or more gRNAs, no yellow light is emitted. Ultimately, this gene can be identified as the gene regulating tau (e.g., causing tau aggregation). In CRISPRa, when the gRNA binds to the target region, the downstream gene is activated, and thus cells with overexpression or an excessive amount of the selection marker can be selected. For example, using the tau protein FRET biosensor described above, an increased amount of yellow light can be shown for the gRNA bound to the gene regulating tau compared to cells without the gRNA. Any number of proteins present in known pathways, specifically disease pathways, can be used as marker proteins to regulate that marker protein and ultimately identify genes that can be found to be involved in a specific disease pathway.
[0042] In one embodiment, the cells of the first culture and the selected cells of the second culture may be sequenced (e.g., FIGS. 2; 202 and 207). The cells of the first culture and the selected cells of the second culture may be sequenced at different times after infection. The cells of the first culture and the selected cells of the second culture may be sequenced at different times after infection to generate a pre-selection read count (cells of the first culture) and a post-selection read count (selected cells of the second culture). For example, the cells of the first culture may be sequenced 3 days after infection, and the selected cells of the second culture may be sequenced 10 days after infection. The cells can be sequenced using any available sequencing technology such as NGS. The nucleic acids of the cells can be sequenced to generate sequence data. The sequence data can include read counts. The sequence data can include read counts for one or more gRNAs of the library. Sequencing the cells can generate sequence data that can be stored in a data structure. The data structure can include one or more nucleic acid sequences and / or sample identifiers.
[0043] The number of reads resulting from array determination suffers from the conventional biases present in CRISPR screen analysis, including the frequent absence of replication, the variability of gRNA knockout efficiency, and the variability of read number distribution. These biases lead to poor results when analyzing read numbers according to the negative binomial approach, the log2-ratio approach, and the paired t-test approach. These approaches require a certain degree of homogeneity / agreement between gRNAs for the same gene, as well as between replicates. These existing approaches cannot handle the large variations between gRNAs for the same gene, as well as between replicates, that can arise, for example, from different infection efficiencies, different gene editing efficiencies, the initial virus number in the screening library, and the presence of other guides with the same phenotype. These biases can exist within a CRISPR experiment and between multiple CRISPR experiments. The step of determining the total number of gRNAs per target region and the total number of gRNAs across all target regions, described herein, represents a read number process that is robust to large variations in read numbers. One embodiment of the method of the present disclosure provides an advantage over existing approaches because it is based on the positive occurrence of guides per gene in an individual experiment, instead of the exact read number of each guide.
[0044] The read numbers can be normalized. For example, read numbers from different samples can be median-normalized to adjust for the effects of library size and read number distribution. In an embodiment, a given N CRISPR / Cas9 knockout screening experiment is performed with a set of MgRNAs, and the read number i of gRNA during experiment j is x ij , where 1 ≤ i ≤ M and 1 ≤ j ≤ N. Since the sequencing depth (or library size) may vary between experiments, the read numbers can be adjusted by applying the ratio median method to all experiments. In an embodiment, the adjusted read number x' ij is given by the equation:
[0045]
Number
[0046] It may be calculated according to, where S is for j = 1 to M
[0047]
Number
[0048] is the median value of. In another embodiment, the adjusted read count x' ij is x ij / S j may be calculated as the rounded value of, where s j is the dimensional coefficient in experiment j and is calculated as the median value of all dimensional coefficients calculated from the individual sgRNA read counts:
[0049]
Number
[0050] where
[0051]
Number
[0052] is the geometric mean of the read count i of the gRNA.
[0053]
Number
[0054] Alternatively, the read count per gRNA and the read count per gene may be normalized using any of counts per million, total counts, or dimensional coefficient normalization. See Anders, S. and Huber, W. (2010) Differential expression analysis for sequence count data. Genome Biol., 11, R106 (which is incorporated herein by reference in its entirety).
[0055] In one embodiment, the method of the present disclosure uses the array data to determine, per target region of DNA, the sum (Σ) of the respective number (n') of each gRNA in selected cells whose read count exceeds a background threshold after its selection (e.g., FIG. 2; 203, 208). The method may also include determining the total number (N') of gRNAs in selected cells whose read count exceeds the background threshold across all target regions (e.g., FIG. 2; 204, 209).
[0056] Thus, the disclosed method can identify the positive occurrence of gRNAs from the library in the sequences of the selected cells. By analyzing the sequence data including the read counts, for each target region of DNA (e.g., gene), the "sum of presence" or n' after selection can be determined. By comparing the individual read count of each gRNA with the background threshold, the "sum of presence", or the respective number of gRNAs present per target region of DNA, can be determined. Determining the respective number of gRNAs per target region of DNA whose read count exceeds the background threshold can be performed by a computing device. The background threshold can be any value sufficient to reduce background noise in the sequence data. For example, 30 may be used as the background threshold. Thus, as contrasted with the amount of gRNAs present as indicated by the read counts, the "sum of presence" indicates the number of gRNAs present (whose read count exceeds the background threshold).
[0057] In one embodiment, a step of infecting a second culture of cas9-positive cells with a library of viral vectors; a step of categorizing the cells of the second culture as either having a specified phenotype or not having the specified phenotype; a step of selecting the cells having the specified phenotype, sequencing the selected cells, and obtaining the number of reads after each selection of gRNA; a step of summing (Σ) the respective numbers of gRNA in the selected cells in which the number of reads after selection exceeds a background threshold per target region of DNA, where Σ = n' in the formula; a step of summing (Σ) the total number of gRNA of the selected cells in which the number of reads exceeds the threshold over the entire target region, where Σ = N' can be repeated any number of times.
[0058] In one embodiment, the method of the present disclosure further includes identifying that a target region is positively selected based on the probability of accidentally observing n' or more gRNAs of a gene in the selected cells (e.g., FIG. 2; 211). In some embodiments, identifying that a target region is positively selected includes determining that the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cells meets a threshold.
[0059] In some embodiments, the method of the present disclosure further includes identifying a target region as a modifier of a second gene based on the probability of randomly observing n' or more gRNAs of a gene in a selected cell. In some embodiments, the method of the present disclosure further includes identifying a target region as a therapeutic target based on the probability of randomly observing n' or more gRNAs of a gene in a selected cell. In some embodiments, the method of the present disclosure further includes identifying that a target region correlates with a designated phenotype based on the probability of randomly observing n' or more gRNAs of a gene in a selected cell. In some embodiments, the method of the present disclosure further includes identifying that a target region exhibits a protective effect based on the probability of randomly observing n' or more gRNAs of a gene in a selected cell. Thus, the disclosed method can assist in the identification of target regions (e.g., genes) involved in disease pathways, target regions (e.g., genes) involved in the regulation of one or more other genes / proteins, and / or target regions (e.g., genes) associated with phenotypes. If the modulation of the gene expression of a candidate gene results in a change in a selected phenotype, the DNA region (e.g., the candidate gene) may be "associated" with the selected phenotype.
[0060] In one embodiment, the disclosed method further includes determining a enrichment score for each gRNA. In some embodiments, determining an enrichment score for each gRNA includes evaluating N / N'. The enrichment score may be determined to exceed a threshold value. The threshold value may be related to the other enrichment scores of other target regions. A target region having a high enrichment score and a low probability of randomly having a gRNA may be used to identify the target region as being associated with a phenotype.
[0061] In an exemplary embodiment, the method and system can be implemented on computer 301 as shown in FIG. 3 and described below. Similarly, the method and system can utilize one or more computers to perform one or more functions in one or more locations. FIG. 3 is a block diagram showing an exemplary operating environment for implementing the method. This exemplary operating environment is merely an example of an operating environment and is not intended to suggest any limitations regarding the use or functionality scope of operating environment architectures. Also, no operating environment should be construed as having any dependencies or requirements related to any one or combination of the components illustrated in the exemplary operating environment.
[0062] The present method and system can be operable in a number of other general-purpose or special-purpose computing system environments or configurations. Examples of computing systems, environments, and / or configurations that may be suitable for use with the system and method include, but are not limited to, personal computers, server computers, laptop devices, and multiprocessor systems. Additional examples include set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices.
[0063] The processing of the method and system can be implemented via software components. The system and method can be described in the general context of computer-executable instructions, such as program modules executed via one or more computers or other devices. Generally, program modules include computer code, routines, programs, objects, components, data structures, etc., by which specific tasks are executed or specific abstract data types are implemented. Also, the method can be practiced in grid-based and distributed computing environments where tasks are performed via remote processing devices linked via a communication network. In a distributed computing environment, program modules can be located on both local and remote computer storage media including memory storage devices.
[0064] Furthermore, the system and method can be implemented via a computing device in the form of a computer 301. The components of computer 301 can include, but are not limited to, one or more processors 303, a system memory 312, and a system bus 313 that couples various system components including the one or more processors 303 to the system memory 312. The system can utilize parallel computing.
[0065] System bus 313 represents one or more of several possible types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, or a local bus, using any of a variety of bus architectures. Bus 313, and all buses specified in this description, can also be implemented via a wired or wireless network connection, and each subsystem including one or more processors 303, a mass storage device 304, an operating system 305, software 306, data 307, a network adapter 308, system memory 312, an input / output interface 310, a display adapter 309, a display device 311, and a human machine interface 302 can be housed within one or more remotely located computing devices 314a, b, c physically separate and connected via this form of bus, implementing a virtually fully distributed system.
[0066] Computer 301 typically includes various computer-readable media. Exemplary readable media can be any available media accessible by computer 301, including, by way of example and not limitation, both volatile and nonvolatile media, removable and non-removable media. System memory 312 includes computer-readable media in the form of volatile memory such as random access memory (RAM) and / or nonvolatile memory such as read-only memory (ROM). System memory 312 typically includes data such as data 307 and / or program modules such as an operating system 305 and software 306 that are immediately accessible to and / or currently being operated on by one or more processors 303.
[0067] In another embodiment, computer 301 may also include other removable / non-removable, volatile / non-volatile computer storage media. By way of example, FIG. 3 illustrates a mass storage device 304 that can provide non-volatile storage of computer code, computer readable instructions, data structures, program modules, and other data for computer 301. For example, but not limited to, mass storage device 304 may be a hard disk, a removable magnetic disk, a removable optical disk, a magnetic cassette or other magnetic storage device, a flash memory card, a CD-ROM, a digital versatile disk (DVD) or other optical storage, a random access memory (RAM), a read only memory (ROM), and / or an electrically erasable programmable read only memory (EEPROM).
[0068] Optionally, any number of program modules can be stored in the mass storage device 304, including, by way of example, the operating system 305 and software 306. Each of the operating system 305 and software 306 (or some combination thereof) can include programming and elements of the software 306. Data 307 can also be stored in the mass storage device 304. The data 307 can be stored in any one or more databases. Examples of such databases include DB2®, MICROSOFT® Access, MICROSOFT® SQL Server, ORACLE®, and / or MYSQL®, POSTGRESQL®. The database can be centralized or distributed across multiple systems. The data 307 may include array determination data. The array determination data may include determining an array of read data (e.g., the number of reads). The computer 301 may receive, for example, first read count data generated in respective steps 120 and 202 of FIGS. 1 and 2. The computer 301 may receive, for example, second read count data generated in respective steps 160 and 207 of FIGS. 1 and 2.
[0069] In another embodiment, the user can input commands and information into computer 301 via an input device (not shown). Examples of such input devices include, but are not limited to, a keyboard, a pointing device (e.g., a "mouse"), a microphone, a joystick, a scanner, a tactile input device such as a glove, and / or other body coverings. These and other input devices can be connected to one or more processors 303 via a human-machine interface 302 connected to the system bus 313, but can also be connected by other interfaces and bus structures such as a parallel port, a game port, an IEEE 1394 port (also referred to as a Firewire (registered trademark) port), a serial port, or a universal serial bus (USB).
[0070] In yet another embodiment, the display device 311 can also be connected to the system bus 313 via an interface such as a display adapter 309. It is expected that two or more display adapters 309 can be provided for computer 301, and two or more display devices 311 can also be provided for computer 301. For example, the display device can be a monitor, a liquid crystal display (LCD), or a projector. In addition to the display device 311, other output peripheral devices can include components such as a speaker (not shown) and a printer (not shown) that can be connected to computer 301 via an input / output interface 310. Any step and / or result of the method can be output to an output device in any form. Such output can be in any form of visual representation including, but not limited to, text, graphical, animation, audio, and / or tactile. The display 311 and the computer 301 can be part of one device or separate devices.
[0071] Computer 301 can operate in a network environment using logical connections to one or more remote computing devices 314a, b, c. As an example, the remote computing device can be a personal computer, a portable computer, a smartphone, a server, a router, a network computer, a peer device, or other common network nodes, etc. The logical connection between computer 301 and remote computing devices 314a, b, c can be made via a network 315 such as a local area network (LAN) and / or a general wide area network (WAN). Such network connections can be via network adapter 308. Network adapter 308 can be implemented in both wired and wireless environments. In one embodiment, system memory 312 can store one or more objects that are accessible to one or more remote computing devices 314a, b, c via network 315. Thus, computer 301 can function as a cloud-based object storage. In another embodiment, one or more of the one or more remote computing devices 314a, b, c can store one or more objects that are permitted access to computer 301, and / or one or more objects that are permitted access to the other of the one or more remote computing devices 314a, b, c. Thus, the one or more remote computing devices 314a, b, c can also function as cloud-based object storage.
[0072] Such programs and components exist at various times within different storage components of computing device 301 and are recognized as being executed via one or more processors 303 of the computer, but for purposes of illustration, other executable program components such as application programs and operating system 305 are shown as discrete blocks herein. In one embodiment, at least a portion of software 306 and / or data 307 can be stored and / or executed on one or more of computing device 301, remote computing devices 314a, b, c, and / or combinations thereof. Accordingly, software 306 and / or data 307 can operate within a cloud computing environment that can implement access to software 306 and / or data 307 via network 315 (e.g., the Internet). Further, in one embodiment, data 307 can be synchronized across one or more of computing device 301, remote computing devices 314a, b, c, and / or combinations thereof.
[0073] The implementation form of software 306 may be stored on a computer-readable medium in some form, or may be transmitted via the computer-readable medium. Any of the methods can be implemented by computer-readable instructions embodied on a computer-readable medium. The computer-readable medium can be any available medium accessible by a computer. By way of example, and not limitation, the computer-readable medium may include "computer storage media" and "communication media". "Computer storage media" includes volatile and non-volatile removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Exemplary computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD), or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, or any other medium that can be used for storing the desired information and is accessible by a computer.
[0074] Software 306 may be configured to perform some or all of the steps of the methods disclosed herein. In one embodiment, software 306 is configured to determine, based on sequencing of a first cell population after infection with a vector comprising a library of at least three guide RNAs (gRNAs) for each of a plurality of target regions of DNA, the respective number (n) of gRNAs present for each of the plurality of target regions of DNA whose read count exceeds a background threshold; determine, based on the respective number of gRNAs for each of the plurality of target regions of DNA, the total number (N) of gRNAs present in the first cell population across all target regions of the plurality of target regions of DNA whose read count exceeds the background threshold; determine, based on sequencing of a second cell population after infection with a vector comprising a library of at least three guide RNAs (gRNAs) for each of a plurality of target regions of DNA, the respective number (n') of gRNAs present for each of the plurality of target regions of DNA whose read count exceeds a background threshold; determine, based on the respective number of gRNAs for each of the plurality of target regions of DNA, the total number (N') of gRNAs present in the second cell population across all target regions of the plurality of target regions of DNA whose read count exceeds the background threshold; determine, based on n, N, n', and N', for each target region of the plurality of target regions, the probability of randomly observing n' gRNAs for the target region in a selected cell; determine, for a target region comprising a sequence of interest, the probability of randomly observing n' or more gRNAs for the target region in a selected cell based on the probability of randomly observing n' gRNAs for the target region in a selected cell; and identify that the sequence of interest is positively selected based on the probability of randomly observing n' or more gRNAs for the sequence of interest in a selected cell.
[0075] Determining the total number (N) of gRNAs present in a first cell population across all target regions of a plurality of target regions of DNA, based on the respective number of each gRNA for each of the plurality of target regions of DNA, where the read count exceeds a background threshold, can include counting each gRNA present in the first cell population where the read count exceeds the background threshold.
[0076] Determining the total number (N’) of gRNAs present in a second cell population across all target regions of a plurality of target regions of DNA, based on the respective number of each gRNA for each of the plurality of target regions of DNA, where the read count exceeds a background threshold, can include counting each gRNA present in selected cells of the second cell population where the read count exceeds the background threshold.
[0077] For each target region of the plurality of target regions, determining the probability of randomly observing n’ guides for the target region in the selected cells, based on n, N, n’, and N’, can
[0078]
Number
[0079] include evaluating, where in the formula
[0080]
Number
[0081] calculates the number of ways to select x objects from y objects. For a target region containing the target sequence, determining the probability of randomly observing n’ or more gRNAs for the target sequence in the selected cells, based on the probability of randomly observing n’ guides for the target region in the selected cells, can
[0082]
Number
[0083] may include evaluating. Identifying that a target sequence is positively selected based on the probability of randomly observing n' or more gRNAs of the target sequence in the selected cells may include determining that the probability of randomly observing n' or more gRNAs of the target sequence in the selected cells meets a threshold.
[0084] Software 306 may further be configured to determine an enrichment score for each gRNA. Determining an enrichment score for each gRNA may include evaluating N / N'.
[0085] Software 306 may be configured to store the read count and number of gRNAs that exist as data 307. FIG. 4 shows an exemplary data structure 410 representing an embodiment of data 307. The data structure 410 may comprise one or more tables, arrays, etc. The data structure 410 may include a plurality of rows and a plurality of columns. The column "Gene #" may include an identifier for a unique gene or target region of DNA. The column "gRNAs per Gene" includes identifiers for unique gRNAs that can bind to a gene or target region of DNA. The column "gRNA Read Count (before selection)" contains a value for the read count. The column "n" includes the number of gRNAs present per target region (e.g., gene) of DNA derived from the column "gRNA Read Count (before selection)". The column "gRNA Read Count (after selection)" contains a value for the read count. The column "n'" includes the number of gRNAs present per target region (e.g., gene) of DNA derived from the column "gRNA Read Count (after selection)".
[0086] The data structure 410 of FIG. 4 is shown including artificial values stored as data 307 for illustrative purposes. For gene 1, gRNAs A, B, and C are present in the library and may bind to gene 1. After sequencing, gRNAs A, B, and C were present in the cell with read counts of 100, 200, and 10 respectively. A background threshold 30 may be applied to the read counts, and the number of positive occurrences of gRNAs from the library may be determined by software 306 to be 2 (n1 = 3) since only gRNAs A and B have read counts exceeding 30. For gene 2, gRNAs D, E, and F are present in the library and may bind to gene 2. After sequencing, gRNAs D, E, and F were present in the cell with read counts of 500, 300, and 200 respectively. A background threshold 30 may be applied to the read counts, and the number of positive occurrences of gRNAs from the library may be determined by software 306 to be 3 (n2 = 3) since gRNAs D, E, and F have read counts exceeding 30. For gene 3, gRNAs G, H, I, and J are present in the library and may bind to gene 3. After sequencing, gRNAs G, H, I, and J were present in the cell with read counts of 200, 250, 10, and 300 respectively. A background threshold 30 may be applied to the read counts, and the number of positive occurrences of gRNAs from the library may be determined by software 306 to be 3 (n3 = 3) since only gRNAs G, H, and J have read counts exceeding 30. For gene 4, gRNAs K, L, M, and N are present in the library and may bind to gene 4. After sequencing, gRNAs K, L, M, and N were present in the cell with read counts of 100, 200, 200, and 100 respectively. A background threshold 30 may be applied to the read counts, and the number of positive occurrences of gRNAs from the library may be determined by software 306 to be 4 (n4 = 4) since gRNAs K, L, M, and N have read counts exceeding 30.
[0087] As shown in the data structure 410 of FIG. 4, the total number of gRNAs or N is 12. N is derived from 2 + 3 + 3 + 4 (n1 + n2 + n3 + n4) from the values of n for each gene. Thus, of the original library of 14 different gRNAs (gRNAs A - N) used to infect the cell population, only 12 different gRNAs were present in amounts exceeding the background threshold. The values of n and N may be stored in the data structure. In addition to the values of n and N, the data structure may include one or more nucleic acid sequences (e.g., target regions and / or gRNA sequences) and / or one or more gRNA identifiers.
[0088] The column "Gene #" contains identifiers of unique genes or target regions of DNA. The column "gRNA per Gene" contains identifiers of unique gRNAs that can bind to genes or target regions of DNA. The column "gRNA Read Count (after selection)" contains values of read counts. The column "n'" contains the number of gRNAs present for each target region (e.g., gene) of DNA derived from the column "gRNA Read Count (after selection)". As shown in the data structure 410 of FIG. 4, for Gene 1, gRNAs A, B, and C are present in the library and may bind to Gene 1. After selection and sequencing, gRNAs A, B, and C were present in the cell with read counts of 0, 50, and 0, respectively. A background threshold of 30 may be applied to the read counts, and the number of positive occurrences of gRNAs from the library may be determined by software 306 to be 1 (n'1 = 1) because only gRNA B has a read count exceeding 30. For Gene 2, gRNAs D, E, and F are present in the library and may bind to Gene 2. After selection and sequencing, gRNAs D, E, and F were present in the cell with read counts of 100, 100, and 100, respectively. A background threshold of 30 may be applied to the read counts, and the number of positive occurrences of gRNAs from the library may be determined by software 306 to be 3 (n'2 = 3) because gRNAs D, E, and F have read counts exceeding 30. For Gene 3, gRNAs G, H, I, and J are present in the library and may bind to Gene 3. After selection and sequencing, gRNAs G, H, I, and J were present in the cell with read counts of 0, 0, 0, and 20, respectively. A background threshold of 30 may be applied to the read counts, and the number of positive occurrences of gRNAs from the library may be determined by software 306 to be 0 (n'3 = 0) because none of them have a read count exceeding 30. For Gene 4, gRNAs K, L, M, and N are present in the library and may bind to Gene 4. After selection and sequencing, gRNAs K, L, M, and N were present in the cell with read counts of 0, 0, 0, and 20, respectively.Background threshold 30 may be applied to the read count, and the number of positive occurrences of gRNA from the library may be determined by software 306 to be 0 (n’4 = 0) since none of the gRNAs have a read count exceeding 30.
[0089] The total number of gRNAs or N’ across all target regions present in the selected cells whose read count exceeds the background threshold may be determined by software 306. As shown in the data structure 410 of FIG. 4, the total number of gRNAs or N’ is 4. N’ is derived from 1 + 3 + 0 + 0 from the values of n for each gene (n’1 + n’2 + n’3 + n’4). Thus, out of the original library of 14 different gRNAs (gRNAs A - N) used to infect the cell population, only 4 different gRNAs were present in the selected cells at an amount exceeding the background threshold. The values of n’ and N’ may be stored in a data structure. In addition to the values of n’ and N’, the data structure may include one or more nucleic acid sequences (e.g., target regions and / or gRNA sequences) and / or one or more gRNA identifiers.
[0090] In an embodiment, the data 307 may also be configured to store one or more results of the software 306. FIG. 5 shows, for example, an example of a result data structure 510 generated by the software 306 using a data structure 410 as an input. The data structure 510 may include one or more tables, arrays, etc. FIG. 5 shows artificial data and artificial results for illustrative purposes. Formal statistical p-values may be calculated to positively observe the number of guides over experimental repeats considering library size, number of guides per gene, and total number of positive guides in each experiment. The data structure 510 shows an initial library (N) of 63,950 gRNAs. The data structure 510 shows that after selection, 4,946 gRNAs (N') (number of unique gRNAs, not amount of gRNAs) remained in the cell population of experiment 1 and 13,606 gRNAs (N') (number of unique gRNAs, not amount of gRNAs) remained in the cell population of experiment 2. As shown in the data structure 510, target region 1 has three (3) gRNAs that can bind to at least a portion of target region 1. In experiment 1, the probability that all 3 of the 3 gRNAs are present by chance in the cell population after selection is 0.000462378 considering that 4,946 gRNAs remain in the cell population after selection. In experiment 2, the probability that all 3 of the 3 gRNAs are present by chance in the cell population after selection is 0.009629 considering that 13,606 gRNAs remain in the cell population after selection. The data structure 510 shows the results of the probabilities that 2 out of 3, 1 out of 3, and 0 out of 3 gRNAs are present by chance in the cell population after selection.
[0091] As shown in data structure 510, target region 2 has four (4) gRNAs that can bind to at least a portion of target region 2. Data structure 510 shows the results of the probabilities that 4 out of 4, 3 out of 4, 2 out of 4, 1 out of 4, and 0 out of 4 gRNAs will randomly be present in the selected cell population. From the results shown in data structure 510, the software 306 can determine that the probability is below a threshold (e.g., sufficiently small). For example, for target region 2, experiment 1, the probability that all 4 out of 4 gRNAs are randomly present is 3.57411E-05, indicating that it is unlikely that all 4 out of 4 gRNAs are simply randomly present.
Example
[0092] A. Example 1: Development of a Genome-Wide CRISPR / Cas9 Screening Platform to Identify Genetic Modifiers of Tau Aggregation To identify genes and pathways that alter the process of abnormal tau protein aggregation, a platform was developed to perform a genome-wide screen using a CRISPR nuclease (CRISPRn) sgRNA library, and genes were identified that regulate the potential for cells to be "seeded" by tau disease-related protein aggregates (i.e., genes that, when disrupted, increase the sensitivity of cells to tau aggregate formation when exposed to a source of tau fibrillized protein). The screen employed a tau biosensor human cell line consisting of HEK293T cells that stably express a tau 4 repeat domain (tau_4RD) containing the P301S pathogenic mutation fused to either CFP or YFP. That is, the HEK293T cell line contains two transgenes that stably express a disease-related protein variant fused to the fluorescent protein CFP or the fluorescent protein YFP:tau4RD-CFP / tau4RD-YFP (TCY), and the tau repeat domain (4RD) contains the P301S pathogenic mutation.
[0093] In these biosensor lines, tau-CFP / tau-YFP protein aggregation produces a FRET signal, which is the result of fluorescence energy transfer from the donor CFP to the acceptor YFP. FRET-positive cells containing tau aggregates can be sorted and isolated by flow cytometry. At baseline, unstimulated cells express a reporter in a stable soluble state with minimal FRET signal. Upon stimulation (e.g., liposomal transfection with seed particles), the reporter protein forms aggregates and produces a FRET signal. Aggregate-containing cells can be isolated by FACS. A stably growing aggregate-containing cell line, Agg[+], can be isolated by clonal serial dilution of the Agg[-] cell line.
[0094] To make it useful for genetic screening, several modifications were made to this tau biosensor cell line. First, these tau biosensor cells were modified by introducing the Cas9-expressing transgene (SpCas9) via a lentiviral vector. Clonal transgenic cell lines expressing Cas9 were selected with blasticidin and isolated by clonal serial dilution to obtain single-cell-derived clones. The clones were evaluated for the level of Cas9 expression by qRT-PCR and for DNA cleavage activity by digital PCR.
[0095] Specifically, Cas9 mutation efficiency was evaluated by digital PCR 3 days and 7 days after transduction of lentivirus encoding gRNA against two selected target genes. Cleavage efficiency was limited by Cas9 levels in the lower-expressing clones. Clones with sufficient levels of Cas9 expression were necessary to achieve maximal activity. Some induced clones with lower Cas9 expression were unable to efficiently cleave the target sequence, while clones with higher expression (including those used in the screening) were able to generate mutations at the target sequences of genes PERK and SNCA with an efficiency of approximately 80% three days after culture. Efficient cleavage was already observed 3 days after gRNA transduction, with only minor improvement after 7 days. Clone 7B10-C3 was selected as a high-performance clone for use in subsequent library screens.
[0096] Second, reagents and methods were developed to render cells sensitive to tau seeding activity. Tau intercellular propagation may result from tau aggregation activity secreted by aggregate-containing cells. To test the cell proliferation of tau aggregation, subclones of a tau YFP cell line consisting of HEK293T cells stably expressing tau repeat domain, tau_4RD, containing the P301S pathogenic mutation in the tau microtubule-binding domain (MBD) fused to YFP were obtained.
[0097] Cells in which tau-YFP protein stably exists in an aggregated state (Agg[+]) were obtained by treating these tau-YFP cells with recombinant fibrillar tau mixed with Lipofectamine reagent to seed the aggregation of tau-YFP protein stably expressed by these cells. Then, to obtain single-cell-derived clones, the "seeded" cells were serially diluted. These clones were then grown to identify clonal cell lines in which tau YFP aggregates persisted stably in all cells in growth passages over time and multiple subcultures. Using one of these tau-YFP_Agg[+] clones, the medium in which tau-YFP_Agg[+] cells were confluent for four days was collected to produce conditioned medium. The conditioned medium (CM) was then applied to naive biosensor tau-CFP / tau-YFP cells at a ratio of 3:1 CM:fresh medium to induce tau aggregation in a small percentage of these recipient cells. Lipofectamine was not used. Lipofectamine was not used in order to have an assay as physiological as possible without tricking the recipient cells with Lipofectamine to force / increase tau aggregation. The conditioned medium consistently induced FRET in approximately 0.1% of the cells, as measured by using flow cytometry to assess the percentage of cells producing an FRET signal as a measure of aggregation.
[0098] B. Example 2: Genome-wide CRISPR / Cas9 screening to identify genetic modifiers of tau aggregation To identify genes whose alteration causes tau aggregation as enriched sgRNAs in FRET(+) cells, Cas9-expressing tau-CFP / tau-YFP biosensor cells without aggregates (Agg[-]) were transduced using a lentiviral delivery approach with two human genome-wide CRISPR sgRNA libraries (GeCKO A and GeCKO B) (Figure 7) to introduce knockout mutations at each target gene. Each CRISPR sgRNA library targets 5’ constitutive exons for functional knockout with an average coverage of approximately 3 sgRNAs per gene (a total of 6 gRNAs per gene when the two libraries are combined). The read number distribution (i.e., the representation of each gRNA in the library) was normal and similar for each library. sgRNAs were designed to avoid off-target effects by avoiding sgRNAs with less than two mismatches to off-target genomic sequences. This library covers 19,050 human genes and 1,864 miRNAs with 1,000 non-target control sgRNAs. The library was transduced at a multiplicity of infection (MOI) <0.3 with coverage of >300 cells per sgRNA. Tau biosensor cells were grown under puromycin selection to select cells with integration and expression of unique sgRNAs per cell. Puromycin selection was started at 1 μg / mL 24 hours after transduction. Five independent screening replicates were used in the primary screen.
[0099] Samples of the fully transduced cell population were taken at the time of cell passage on days 3 and 6 after transduction. After passage on day 6, the cells were grown in conditioned medium and made sensitive to seeding activity. On day 10, fluorescence-activated cell sorting (FACS) was used to specifically isolate a subpopulation of FRET[+] cells. The screen consisted of five replicated experiments. DNA isolation and PCR amplification of the integrated sgRNA constructs enabled characterization of the sgRNA repertoire at each time point by next-generation sequencing (NGS).
[0100] By statistical analysis of the NGS data, it became possible to identify sgRNAs enriched in the FRET[+] subpopulation on day 10 of five experiments, compared to the sgRNA repertoire at early time points on days 3 and 6. The first strategy to identify potential tau modifiers was to use DNA sequencing to generate sgRNA read counts in each sample using the DESeq algorithm and find sgRNAs that were more abundant on day 10 vs. day 3, or day 10 vs. day 6, but not on day 6 vs. day 3 (fold change (fc) ≥ 1.5 and negative binomial test p < 0.01). Fc ≥ 1.5 means a ratio of (average of day 10 counts) / (average of day 3 or day 6 counts) ≥ 1.5. P < 0.01 means there is no statistical difference between day 10 counts and day 3 counts, or the probability that day 6 counts < 0.01. The DESeq algorithm is an algorithm widely used for "analysis of differential expression of sequence count data". See, for example, Anders et al. (2010), Genome Biology, 11:R106, which is incorporated herein by reference.
[0101] Specifically, two comparisons were used in each library and significant sgRNAs were identified: day 10 vs. day 3, and day 10 vs. day 6. For each of these four comparisons, using the DESeq algorithm, the cut-off threshold considered significant was a fold change ≥ 1.5 in addition to a negative binomial test p < 0.01. When significant guides were identified in each of these comparisons for each library, a gene was considered significant if it met one of the following two criteria: (1) at least two sgRNAs corresponding to that gene were considered significant in one comparison (either day 10 vs. day 3, or day 10 vs. day 6). And (2) at least one sgRNA was significant in both comparisons (day 10 vs. day 3, and day 10 vs. day 6). Using this algorithm, five genes were identified as significant from the first library and four genes were identified from the second library. See Table 1.
[0102]
Table 1
[0103] However, this first strategy requires a very strict level of homogeneity in the number of leads within each experimental group. For the same sgRNA, many factors such as the initial virus number in the screening library, the infection or gene editing efficiency, and the relative growth rate after gene editing can cause variability in the number of leads among samples within each experimental group (day 3, day 6, or day 10 samples). Therefore, a second strategy was also used based on the positive occurrence of guides per gene (>30 leads) in each sample at day 10 (after selection) instead of the exact number of leads.
[0104] The pre-selection CRISPR experiments were repeated four times. As shown in Figure 6, for gene "G1" in the pre-selection experiment "Exp1", gRNAs "g1", "g2", and "g3" were present in the cell population with read counts of 121, 1000, and 302 respectively. For gene "G2", gRNAs "g4", "g5", "g6", and "g7" were present in the cell population with read counts of 443, 2012, 534, and 150 respectively. Such read count data were generated for genes from "G1" to "G21,000". For each gene, the "total presence" or n was determined. By comparing the individual read counts of each gRNA with the background threshold, the "total presence", which is the number of each gRNA present per target region of DNA, was determined. In this case, 30 was used as the background threshold. Thus, as contrasted with the amount of gRNA present as indicated by the read counts, the "total presence" indicates the number of gRNAs that are qualitatively present. Since the read counts of gRNAs g1, g2, and g3 each exceed the background threshold of 30, the total presence of gRNAs corresponding to gene G1 is 3. Since the read counts of gRNAs g4, g5, g6, and g7 each exceed the background threshold of 30, the total presence of gRNAs corresponding to gene G2 is 4. The total number or N of gRNAs across all target regions present in the cell population with read counts exceeding the background threshold is shown as 59,010. Thus, out of the original library of approximately 64,000 different gRNAs used to infect the cell population, only approximately 59,000 different gRNAs were present in an amount exceeding the background threshold.
[0105] The CRISPR experiments involving phenotypic selection were repeated four times. However, prior to sequencing, the cells in the cell population were sorted according to phenotype using fluorescence techniques (e.g., FRET fluorescence). When Cas9 / CRISPR cleaved the target region (e.g., gene) of the cell and the cell did not emit fluorescence, the gene was successfully knocked out. When the cell emitted fluorescence, the gene was not knocked out. The fluorescent cells could then be sequenced, and the non-fluorescent cells were not sequenced. The selected cells represent cells showing a specific phenotype / marker.
[0106] As shown in FIG. 6, for gene "G1" in the post-selection experiment "Exp1", gRNAs "g1", "g2", and "g3" were present (or not present) in the cell population with read counts of 0, 8, and 12, respectively. For gene "G2", gRNAs "g4", "g5", "g6", and "g7" were present (or not present) in the cell population with read counts of 4, 25, 4, and 150, respectively. Such read count data was generated for genes from "G1" to "G21,000". For each gene, the "total presence" or n' was determined. By comparing the individual read counts of each gRNA with the background threshold, the "total presence", which is the number of each gRNA present per target region of DNA, was determined. In this case, 30 was used as the background threshold. Thus, as contrasted with the amount of gRNA present as indicated by the read count, the "total presence" indicates the number of gRNAs present. Since the read counts of gRNAs g1, g2, and g3 do not exceed the background threshold of 30 each, the total presence of gRNAs corresponding to gene G1 is 0. Since only the read count of gRNA g7 exceeds the background threshold of 30, the total presence of gRNAs corresponding to gene G2 is 1. The total number of gRNAs or N' across all target regions present in the cell population where the read count exceeded the background threshold is shown as 4,320. Thus, out of the original library of approximately 64,000 different gRNAs used to infect the cell population, only approximately 4,320 different gRNAs were present in an amount exceeding the background threshold.
[0107] Formal statistical p-values were calculated to positively observe the number of guides in the post-selection sample, taking into account the library size, the number of guides per gene, and the total number of positive guides in the post-selection sample. When read counts are considered, the probability of accidentally observing n' guides for a target region (e.g., a gene) within the selected cells (post-selection) is given by the formula
[0108]
Number
[0109] It was determined according to. As an explanation,
[0110]
Number
[0111] This determines the number of ways to select the x object from the y object. After determining the probability of accidentally observing n' guides for a target region (e.g., a gene) in the selected cell (after selection), the probability of accidentally observing n' or more gRNAs of the gene in the selected cell (after selection) is given by the formula
[0112]
Number
[0113] It was determined according to. Once the probability of accidentally observing n' or more gRNAs of the target region in the selected cell (after selection) was determined, the average enriched gRNA was determined at the target region level. The overall enrichment of the read count of the gene after selection compared to before selection was used as an additional parameter to identify positive genes. The average enrichment was represented by an enrichment score. The enrichment score was determined by evaluating N / N'. Referring to Figure 6, the enrichment score is 59010 / 4320, or 13.66.
[0114] The probability of accidentally observing n' or more gRNAs of the target region in the selected cell (after selection) can be used to evaluate whether the target region is positively selected. The enrichment score can further be used to evaluate whether the target region is positively selected. A target region having a probability of observing n' or more gRNAs of the target region in the selected cell that is significantly lower than the probability of accidentally observing n' or more gRNAs of the target region in the selected cell may be identified as a positively selected target region. Further, a target region having an enrichment score exceeding a threshold may indicate a positively selected target region.
[0115] Thus, this second strategy represents a new and more sensitive assay for CRISPR positive selection. The goal of CRISPR positive selection is to use DNA sequencing to identify genes where perturbation by an sgRNA correlates with a phenotype. To reduce the noise background, multiple sgRNAs for the same gene are typically used in these experiments, along with replicates of the experiment. However, currently available commonly used statistical analysis methods that require a certain degree of homogeneity / agreement between sgRNAs for the same gene, as well as between technical replicates, do not work well. This is because these methods cannot handle the large variation between sgRNAs and repeats for the same gene due to many possible reasons (e.g., different infection or gene editing efficiencies, initial virus numbers in the screening library, and the presence of other sgRNAs with the same phenotype). In contrast, the method shown in this example is robust to large variations. This is based on the positive occurrence of guides per gene in individual experiments, rather than the exact read counts of each sgRNA. Formal statistical p-values are calculated to positively observe the number of sgRNAs across experimental replicates, taking into account the library size, the number of sgRNAs per gene, and the total number of positive sgRNAs in each experiment. The relative sgRNA sequence read enrichment before and after phenotype selection is also used as a parameter. This method performs better than currently widely used state-of-the-art methods, including DESeq, MAGECK, and others. Specifically, this method includes the following steps.
[0116] (1) For each experiment, identifying any present guides in cells with a positive phenotype. (2) At the gene level, calculate the random possibility (also known as p-value) where a guide exists in each experiment. The overall possibility existing across multiple experiments is calculated by Fisher's combined probability test (see: Fisher, R.A. Fisher, R.A. (1948). "Question and Answer #14". The American Statistician). That is, first, calculate the test statistic φ using the p-values from multiple experiments:
[0117] [Number]
[0118] where p k is the p-value calculated for the k-th experiment, and K is the total number of experiments. Then, the combination of p-values across K experiments is equal to the probability of observing the value of φ under the chi-square distribution with 2*K degrees of freedom.
[0119] (3) Calculate the average enrichment of the guide at the gene level: Enrichment score = relative abundance after selection / relative abundance before selection. Relative abundance = number of reads of the guide for the gene / total number of reads of all guides.
[0120] (4) Select genes that are significantly lower than the existing random possibility and exceed a specific enrichment score. C. Example 3 Using CRISPR / Cas9 activation and inactivating mutagenesis, a CRISPRn sgRNA library (hGeCKO-A and hGeCKO-B) targeting the coding exons for functional knockout was used to screen for genetic modifiers of tau and α-synuclein fibrillation and proliferation. Figure 7 shows the sample identifier, experiment number, time to sequencing, and identifies which library (Gecko A or Gecko B) was used in the experiment.
[0121] The Gecko A library was composed of approximately 63,950 gRNAs. Of the 63,950 gRNAs, 56,116 gRNAs targeted 18,874 genes, and the majority of genes were targeted by three gRNAs. Of the 63,950 gRNAs, 6,834 gRNAs targeted 1,795 microRNAs (miRNAs), and the majority of microRNAs were targeted by four gRNAs.
[0122] The Gecko B library was composed of approximately 56,869 gRNAs. Of the 56,869 gRNAs, 55,869 gRNAs targeted 18,834 genes, and the majority of genes were targeted by three gRNAs. None of the gRNAs targeted miRNAs. There were no targets for 1,000 gRNAs.
[0123] On days 3 and 6, no phenotypes were shown. On day 10, the samples were phenotype positive. Figure 8 shows the DNA read counts from samples infected with the Gecko A library. Each bar represents a sample infected with the virus. Each sample was sequenced on day 3 (d03), day 6 (d06), or day 10 (d10), and each sample represents the sequencing readout of gRNAs rather than the whole genome.
[0124] Figure 9 shows the normalization of read counts to the median from samples infected with the Gecko A library. Normalization was performed by dividing the gRNA read counts by the sum of the read counts in each sample and multiplying by the median of the sum of the read counts across all samples. The lower bar graph is qualitative and shows read counts above a threshold of 30, and day 10 is after selection. Considering similar sums of read counts, the samples on days 3 and 6 had a wide variety of gRNAs, and the samples on day 10 had far fewer types of gRNAs.
[0125] Figure 10 shows the formal statistical p-values calculated to positively observe the number of gRNAs across experimental repeats, taking into account library size, the number of gRNAs per gene, and the total number of positive gRNAs in each experiment. The p-values are shown for five samples infected with the Gecko A library and sequenced on day 10.
[0126] Embodiment Embodiment 1. (A) Infecting a first culture of cas9-positive cells with a library of viral vectors, wherein the library comprises at least three guide RNAs (gRNAs) for cleaving a target region of DNA within the genome of the cells, infecting, sequencing the cells to obtain the read count for each of the gRNAs, Summing (Σ) the number of each of the gRNAs for which the read count exceeds a background threshold per target region of DNA, where Σ = n in the formula, summing, and summing (Σ) the total number of gRNAs for which the read count exceeds the background threshold across all target regions, where Σ = N in the formula; (B) Infecting a second culture of cas9-positive cells with a library of viral vectors, Categorizing the cells of the second culture as either having a specified phenotype or not having the specified phenotype, selecting the cells having the specified phenotype, sequencing the selected cells to obtain the post-selection read count for each of the gRNAs, Summing (Σ) the number of each of the gRNAs for which the post-selection read count exceeds the background threshold in the selected cells, where Σ = n' in the formula, summing, and Summing (Σ) the total number of gRNAs in the selected cells for which the read count exceeds the threshold across all target regions, where Σ = N' in the formula; (C) For a target region of DNA, the formula
[0127]
Number
[0128] According to this, it is to calculate the probability of accidentally observing n' gRNAs for a target region in a selected cell, where in the formula
[0129]
Number
[0130] is to calculate, compute, and for a target region of DNA containing a gene, according to the formula
[0131]
Number
[0132] to calculate the probability of accidentally observing n' or more gRNAs of a gene in a selected cell, according to the formula, a method comprising. Embodiment 2 (A) Infecting a first culture of cas9-positive cells with a library of viral vectors, wherein the library contains at least three guide RNAs (gRNAs) for enhancing the transcription of a target region of DNA in the genome of the cells, infecting; Sequencing the cells to obtain the read count of each of the gRNAs; For each target region of DNA, summing (Σ) the number of each gRNA whose read count exceeds a background threshold, where Σ = n in the formula, and summing; and Summing (Σ) the total number of gRNAs whose read count exceeds the background threshold over all target regions, where Σ = N in the formula; summing; (B) Infecting a second culture of cas9-positive cells with a library of viral vectors; Categorizing the cells of the second culture as either having a designated phenotype or not having a designated phenotype; Selecting cells having a specified phenotype and sequencing the selected cells to obtain the number of reads after each selection of the gRNAs; For each target region of DNA, summing (Σ) the number of each of the gRNAs in the selected cells in which the number of reads after selection exceeds a background threshold, where Σ = n' in the formula, and summing; and Summing (Σ) the total number of gRNAs of the selected cells in which the number of reads exceeds the threshold over all target regions, where Σ = N' in the formula; (C) For a target region of DNA, calculating the probability of randomly observing n' gRNAs in the target region in the selected cells according to the formula
[0133] [Number]
[0134] wherein the formula calculates the number of ways to select x objects from y objects, calculating, and for a target region of DNA containing a gene, calculating the probability of randomly observing n' or more gRNAs of the gene in the selected cells according to the formula
[0135] [Number]
[0136] which calculates the number of ways to select x objects from y objects, calculating, and a method comprising, for a target region of DNA containing a gene, calculating the probability of randomly observing n' or more gRNAs of the gene in the selected cells according to the formula
[0137] [Number]
[0138] wherein the formula calculates the probability of randomly observing n' or more gRNAs of the gene in the selected cells according to the formula. Embodiment 3 A method according to any one of Embodiments 1 to 2, wherein the first culture is infected according to CRISPR technology.
[0139] Method according to any one of Embodiments 1 to 3, wherein the cas-9 positive cells of the first culture are modified to contain one or more selection markers. Method according to any one of Embodiments 1 to 4, wherein the designated phenotype is fluorescence.
[0140] Method according to any one of Embodiments 1 to 5, wherein the designated phenotype is cell survival. Method according to Embodiment 4, wherein one or more selection markers include a fluorescence marker.
[0141] Method according to Embodiment 7, wherein the fluorescence marker is part of a FRET biosensor. Method according to Embodiment 4, wherein one or more selection markers are detectable enzymes.
[0142] Method according to Embodiment 9, wherein the detectable enzyme is β-galactosidase. Method according to Embodiment 9, wherein the detectable enzyme is luciferase.
[0143] Method according to any one of Embodiments 1 to 11, wherein the target region contains a gene. Method according to any one of Embodiments 1 to 12, wherein categorizing the cells of the second culture as either having the designated phenotype or not having the designated phenotype includes applying a selection mechanism to the second cell culture.
[0144] Method according to Embodiment 13, wherein the selection mechanism includes one or more of exposing the second cell population to a drug or exposing the second cell population to a substance that identifies protein activity or expression level.
[0145] Method according to Embodiment 4, wherein selecting cells having the designated phenotype includes sorting the cells based on one or more selection markers. The method according to any one of Embodiments 1 to 15, further comprising identifying that the target region is positively selected based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cell.
[0146] Embodiment 17 The method according to Embodiment 16, wherein identifying that the target region is positively selected includes determining that the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cell meets a threshold.
[0147] Embodiment 18 The method according to any one of Embodiments 1 to 17, further comprising determining an enrichment score for each gRNA. Embodiment 19 The method according to Embodiment 18, wherein determining an enrichment score for each gRNA includes evaluating N / N'.
[0148] Embodiment 20 The method according to any one of Embodiments 2 to 19, wherein the cas-9 positive cells contain inactivated cas-9. Embodiment 21 The method according to Embodiment 20, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0149] Embodiment 22 The method according to any one of Embodiments 1 to 21, wherein the target region of the DNA regulates a downstream gene or protein. Embodiment 23 The method according to Embodiment 22, wherein regulating a downstream gene or protein includes activating or inhibiting the downstream gene or protein.
[0150] Embodiment 24 The method according to any one of Embodiments 1 to 23, further comprising identifying the target region as a modifying factor for a second gene based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cell.
[0151] Embodiment 25 The method according to any one of Embodiments 1 to 24, further comprising identifying the target region as a therapeutic target based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cell.
[0152] Embodiment 26. The method according to any one of Embodiments 1 to 25, further comprising identifying that the target region correlates with the specified phenotype based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cell.
[0153] Embodiment 27. The method according to any one of Embodiments 1 to 26, further comprising identifying that the target region exhibits a protective effect based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cell.
[0154] Embodiment 28 A method comprising: determining, for each of a plurality of target regions of DNA, the respective number (n) of gRNAs present for each of the plurality of target regions of DNA, based on sequencing of a first cell population after infection using a vector comprising a library of at least three guide RNAs (gRNAs) for each of the plurality of target regions of DNA, wherein the read count exceeds a background threshold; determining the total number (N) of gRNAs present in the first cell population, across all target regions of the plurality of target regions of DNA, based on the respective number of gRNAs for each of the plurality of target regions of DNA, wherein the read count exceeds a background threshold; determining, for each of a plurality of target regions of DNA, the respective number (n') of gRNAs present for each of the plurality of target regions of DNA, based on sequencing of a second cell population after infection using a vector comprising a library of at least three guide RNAs (gRNAs) for each of the plurality of target regions of DNA, wherein the read count exceeds a background threshold; determining the total number (N') of gRNAs present in the second cell population, across all target regions of the plurality of target regions of DNA, based on the respective number of gRNAs for each of the plurality of target regions of DNA, wherein the read count exceeds a background threshold; determining, for each target region of the plurality of target regions, the probability of randomly observing n' gRNAs for the target region in a selected cell, based on n, N, n', and N'; determining, for a target region comprising a sequence of interest, the probability of randomly observing n' or more gRNAs for the sequence of interest in a selected cell, based on the probability of randomly observing n' gRNAs for the target region in a selected cell; and identifying that the sequence of interest is positively selected, based on the probability of observing n' or more gRNAs for the sequence of interest in a selected cell.
[0155] Embodiment 29 The method according to embodiment 28, wherein the first cell population and the second cell population comprise cas-9 positive cells. Based on the sequencing of a first cell population after infection using a vector comprising a library of at least three gRNAs for each of a plurality of target regions of DNA, determining, for each of the plurality of target regions of DNA, the respective number (n) of gRNAs present for which the read count exceeds a background threshold, comprising infecting the first cell population with a library of viral vectors, sequencing the cells of the first cell population to obtain the read count for each of the gRNAs, and counting, for each of the plurality of target regions of DNA, each gRNA for which the read count exceeds the background threshold, the method according to embodiment 29.
[0156] Based on the respective number of gRNAs for each of a plurality of target regions of DNA, determining the total number (N) of gRNAs present in the first cell population over all target regions of the plurality of target regions of DNA for which the read count exceeds a background threshold, comprising counting each gRNA present in the first cell population for which the read count exceeds the background threshold, the method according to embodiment 30.
[0157] Based on the sequencing of a second cell population after infection using a vector comprising a library of at least three gRNAs for each of a plurality of target regions of DNA, determining, for each of the plurality of target regions of DNA, the respective number (n') of gRNAs present for which the read count exceeds a background threshold, comprising categorizing the cells of the second cell population as either having a specified phenotype or not having a specified phenotype, selecting the cells having the specified phenotype, sequencing the selected cells to obtain the read count for each of the gRNAs, and counting, for each of the plurality of target regions of DNA, each gRNA for which the read count exceeds the background threshold, the method according to embodiment 31.
[0158] Method according to any one of Embodiments 28 to 32, which comprises determining the total number (N') of gRNAs present in a second cell population across all target regions of a plurality of DNA target regions, the read count of which exceeds a background threshold, based on the respective number of each gRNA for each of the plurality of DNA target regions, which comprises counting each gRNA present in a selected cell of the second cell population, the read count of which exceeds the background threshold.
[0159] Embodiment 34 For each target region of a plurality of target regions, determining the probability of accidentally observing n' gRNAs for the target region in a selected cell, based on n, N, n', and N',
[0160]
Number
[0161] which comprises evaluating, where
[0162]
Number
[0163] is the number of ways to select x objects from y objects, is a method according to any one of Embodiments 28 to 33. Embodiment 35 For a target region containing a target sequence, determining the probability of accidentally observing n' or more gRNAs for the target sequence in a selected cell, based on the probability of accidentally observing n' gRNAs for the target region in the selected cell,
[0164]
Number
[0165] which comprises evaluating, is a method according to any one of Embodiments 28 to 34. Embodiment 36. Based on the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cells, identifying that the target sequence is positively selected includes determining that the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cells meets a threshold value. The method according to any one of Embodiments 28 to 35.
[0166] Embodiment 37. The method according to any one of Embodiments 28 to 36, further comprising determining a enrichment score for each gRNA. Embodiment 38. The method according to Embodiment 37, wherein determining a enrichment score for each gRNA includes evaluating N / N'.
[0167] Embodiment 39. The method according to any one of Embodiments 28 to 38, wherein the cas-9 positive cells contain inactivated cas-9. Embodiment 40. The method according to Embodiment 39, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0168] Embodiment 41 An apparatus comprising one or more processors and processor-executable instructions that, when executed by the one or more processors, cause the apparatus to: (A) receive first read count data, where the first read count data is generated by infecting a first culture of cas9-positive cells with a library of viral vectors, the library containing at least three guide RNAs (gRNAs) for cleaving a target region of DNA in the genome of the cells, and sequencing the cells to obtain the read count for each of the gRNAs; based on the first read count data, sum (Σ) for each gRNA the number of times its read count exceeds a background threshold per target region of DNA, where Σ = n, and based on the first read count data, sum (Σ) for the entire target region the total number of gRNAs whose read count exceeds the background threshold, where Σ = N; (B) receive second read count data, where the second read count data is generated by infecting a second culture of cas9-positive cells with a library of viral vectors, categorizing the cells of the second culture as having either a specified phenotype or not having the specified phenotype, and selecting the cells having the specified phenotype and sequencing the selected cells to obtain the post-selection read count for each of the gRNAs; based on the second read count data, sum (Σ) for each gRNA the number of times its post-selection read count exceeds a background threshold in the second cells per target region of DNA, where Σ = n', and based on the second read count data, sum (Σ) for the entire target region the total number of gRNAs in the selected cells whose read count exceeds the threshold, where Σ = N'; (C) for a target region of DNA, calculate the probability of randomly observing n' gRNAs in the target region in the selected cells according to the formula
[0169] [Number]
[0170] such that the probability of randomly observing n' gRNAs in the target region in the selected cells is calculated according to the formula, where
[0171]
Mathematics
[0172] calculates the number of ways to select an x object from a y object, and for a target region of DNA containing a gene, according to the formula
[0173]
Mathematics
[0174] a memory storing processor-executable instructions that cause, for a target region of DNA containing a gene, the probability of accidentally observing n' or more gRNAs of the gene in a selected cell to be calculated Embodiment 42 The apparatus according to Embodiment 41, wherein a first culture is infected according to CRISPR technology.
[0175] Embodiment 43 The apparatus according to any one of Embodiments 41 to 42, wherein the cas-9 positive cells of the first culture are modified to contain one or more selection markers. Embodiment 44 The apparatus according to any one of Embodiments 41 to 43, wherein the target region contains a gene.
[0176] Embodiment 45 The apparatus according to any one of Embodiments 41 to 44, wherein applying a selection mechanism to a second cell culture includes categorizing the cells of the second culture as either having a specified phenotype or not having a specified phenotype.
[0177] Embodiment 46 The apparatus according to Embodiment 45, wherein the selection mechanism includes one or more of exposing a second cell population to a drug or exposing a second cell population to a substance that identifies protein activity or expression level.
[0178] Embodiment 47 The apparatus according to any one of Embodiments 41 to 46, wherein selecting cells having a specified phenotype includes sorting the cells based on one or more selection markers.
[0179] Embodiment 48. The apparatus according to any one of Embodiments 41 to 47, further comprising identifying that the target region is positively selected based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cell.
[0180] Embodiment 49. The apparatus according to Embodiment 48, wherein identifying that the target region is positively selected includes determining that the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cell satisfies a threshold value.
[0181] Embodiment 50. The apparatus according to any one of Embodiments 41 to 49, further comprising determining a enrichment score for each gRNA. Embodiment 51. The apparatus according to Embodiment 50, wherein determining a enrichment score for each gRNA includes evaluating N / N'.
[0182] Embodiment 52. The apparatus according to any one of Embodiments 41 to 51, wherein the cas-9 positive cells contain inactivated cas-9. Embodiment 53. The apparatus according to Embodiment 52, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0183] Embodiment 54 A non-transitory computer-readable medium for determining the probability of accidentally observing one or more guide RNAs (gRNAs), which is processor-executable instructions that, when executed by one or more processors, cause the one or more processors to: (A) receive first read count data, where the first read count data is generated by infecting a first culture of cas9-positive cells with a library of viral vectors, the library containing at least three guide RNAs (gRNAs) for cleaving a target region of DNA in the genome of the cells, and sequencing the cells to obtain the read count for each of the gRNAs; based on the first read count data, sum (Σ) the number of each gRNA whose read count exceeds a background threshold per target region of DNA, where Σ=n, and based on the first read count data, sum (Σ) the total number of gRNAs whose read count exceeds the background threshold over the entire target region, where Σ=N; (B) receive second read count data, where the second read count data is generated by infecting a second culture of cas9-positive cells with a library of viral vectors, categorizing the cells of the second culture as having either a specified phenotype or not having the specified phenotype, and selecting the cells having the specified phenotype and sequencing the selected cells to obtain the post-selection read count for each of the gRNAs; based on the second read count data, sum (Σ) the number of each gRNA in the second cells whose post-selection read count exceeds the background threshold per target region of DNA, where Σ=n', and based on the second read count data, sum (Σ) the total number of gRNAs in the selected cells whose read count exceeds the threshold over the entire target region, where Σ=N'; (C) for the target region of DNA, calculate the probability of accidentally observing n' gRNAs in the target region in the selected cells according to the formula
[0184]
Number
[0185] such that the probability of accidentally observing n' gRNAs in the target region in the selected cells is calculated according to the formula, where
[0186]
Number
[0187] calculates the number of ways to select an x object from a y object, and for a target region of DNA containing a gene, according to formula n’, the gene in the selected cell
[0188]
Number
[0189] A non - transitory computer - readable medium storing processor - executable instructions that cause the probability of accidentally observing the above gRNA to be calculated. Embodiment 55 The non - transitory computer - readable medium according to embodiment 54, wherein the first culture is infected according to CRISPR technology.
[0190] Embodiment 56 The non - transitory computer - readable medium according to any one of embodiments 54 - 55, wherein the cas - 9 positive cells of the first culture are modified to contain one or more selectable markers.
[0191] Embodiment 57 The non - transitory computer - readable medium according to any one of embodiments 54 - 56, wherein the target region contains a gene. Embodiment 58 Applying a selection mechanism to a second cell culture, including categorizing the cells of the second culture as having either a specified phenotype or not having a specified phenotype. The non - transitory computer - readable medium according to any one of embodiments 54 - 57.
[0192] Embodiment 59 The non - transitory computer - readable medium according to embodiment 58, wherein the selection mechanism includes one or more of exposing a second cell population to a drug or exposing a second cell population to a substance that identifies protein activity or expression level.
[0193] Embodiment 60 The non - transient computer - readable medium according to any one of Embodiments 54 - 59, wherein selecting a cell having a specified phenotype includes sorting the cells based on one or more selection markers.
[0194] Embodiment 61 The non - transient computer - readable medium according to any one of Embodiments 54 - 60, further comprising identifying that a target region is positively selected based on the probability of accidentally observing n' or more gRNAs of a gene in a selected cell.
[0195] Embodiment 62 The non - transient computer - readable medium according to Embodiment 61, wherein identifying that a target region is positively selected includes determining that the probability of accidentally observing n' or more gRNAs of a target sequence in a selected cell meets a threshold.
[0196] Embodiment 63 The non - transient computer - readable medium according to any one of Embodiments 54 - 62, further comprising determining an enrichment score for each gRNA. Embodiment 64 The non - transient computer - readable medium according to Embodiment 63, wherein determining an enrichment score for each gRNA includes evaluating N / N'.
[0197] Embodiment 65 The non - transient computer - readable medium according to any one of Embodiments 54 - 64, wherein the cas - 9 positive cells include inactivated cas - 9. Embodiment 66 The non - transient computer - readable medium according to Embodiment 65, wherein the inactivated cas - 9 is fused to at least one transcriptional activation domain.
[0198] Embodiment 67 An apparatus, comprising: one or more processors; processor-executable instructions which, when executed by the one or more processors, cause the apparatus to: based on the sequencing of a first cell population after infection using a vector comprising a library of at least three guide RNAs (gRNAs) for each of a plurality of target regions of DNA, determine, for each of the plurality of target regions of DNA, the respective number (n) of gRNAs present for which the read count exceeds a background threshold; based on the respective number of gRNAs for each of the plurality of target regions of DNA, determine the total number (N) of gRNAs present in the first cell population across all target regions of the plurality of target regions of DNA for which the read count exceeds the background threshold; based on the sequencing of a second cell population after infection using a vector comprising a library of at least three guide RNAs (gRNAs) for each of a plurality of target regions of DNA, determine, for each of the plurality of target regions of DNA, the respective number (n') of gRNAs present for which the read count exceeds a background threshold; based on the respective number of gRNAs for each of the plurality of target regions of DNA, determine the total number (N') of gRNAs present in the second cell population across all target regions of the plurality of target regions of DNA for which the read count exceeds the background threshold; for each target region of the plurality of target regions, determine the probability of randomly observing n' gRNAs in the target region in the selected cell based on n, N, n', and N'; for the target region containing the sequence of interest, determine the probability of randomly observing n' or more gRNAs of the sequence of interest in the selected cell based on the probability of randomly observing n' gRNAs in the target region in the selected cell; and based on the probability of randomly observing n' or more gRNAs of the sequence of interest in the selected cell, identify that the sequence of interest is positively selected; a memory storing the processor-executable instructions.
[0199] Embodiment 68 The apparatus according to embodiment 67, wherein the first cell population and the second cell population comprise cas-9 positive cells. Embodiment 69 Processor-executable instructions that, when executed by one or more processors, cause the apparatus to determine, for each of a plurality of target regions of DNA, the respective number (n) of gRNAs present based on sequencing of a first cell population after infection using a vector comprising a library of at least 3 gRNAs for each of the plurality of target regions of DNA, where the read count exceeds a background threshold. The processor-executable instructions cause the apparatus to receive first read count data generated by infecting the first cell population with a library of viral vectors and sequencing the cells of the first cell population to obtain the read count for each gRNA, and to count each gRNA for which the read count exceeds the background threshold for each of the plurality of target regions of DNA based on the first read count data. The apparatus according to any one of Embodiments 67 to 68.
[0200] Embodiment 70 Processor-executable instructions that, when executed by one or more processors, cause the apparatus to determine, for all target regions of a plurality of target regions of DNA, the total number (N) of gRNAs present in a first cell population, where the read count exceeds a background threshold, based on the respective number of gRNAs for each of the plurality of target regions of DNA. The processor-executable instructions cause the apparatus to count each gRNA present in the first cell population for which the read count exceeds the background threshold. The apparatus according to Embodiment 69.
[0201] Embodiment 71 Processor-executable instructions that, when executed by one or more processors, cause the device to determine, based on the sequencing of a first cell population after infection using a vector containing a library of at least three gRNAs for each of a plurality of target regions of DNA, the respective number (n') of gRNAs present for each of the plurality of target regions of DNA whose read count exceeds a background threshold. The device is caused to categorize the cells of a second cell population as either having a specified phenotype or not having the specified phenotype, select the cells having the specified phenotype, sequence the selected cells to obtain the read count for each gRNA, receive the second read count data generated thereby, and count, for each of the plurality of target regions of DNA, each gRNA whose read count exceeds the background threshold based on the second read count data. The device according to any one of Embodiments 67 to 70.
[0202] Embodiment 72 Processor-executable instructions that, when executed by one or more processors, cause the device to determine, based on the respective number of gRNAs for each of a plurality of target regions of DNA, the total number (N') of gRNAs present in a second cell population across all target regions of the plurality of target regions of DNA whose read count exceeds a background threshold. The device is caused to count each gRNA present in the selected cells of the second cell population whose read count exceeds the background threshold. The device according to any one of Embodiments 67 to 71.
[0203] Embodiment 73 Processor-executable instructions that, when executed by one or more processors, cause the device to determine, for each target region of a plurality of target regions, the probability of accidentally observing n' gRNAs for the target region in the selected cells based on n, N, n', and N'. The device is caused to
[0204]
Number
[0205] to evaluate, in the formula
[0206] [Number]
[0207] is the device according to any one of Embodiments 67 to 72 that calculates the number of ways to select an x object from a y object. Embodiment 74 Processor-executable instructions that, when executed by one or more processors, cause the device to determine, for a target region containing a target sequence in a selected cell, the probability of randomly observing n' or more gRNAs of the target sequence in the selected cell based on the probability of randomly observing n'gRNAs for the target region within the selected cell. The device has
[0208] [Number]
[0209] to evaluate, the device according to any one of Embodiments 67 to 73. Embodiment 75 Processor-executable instructions that, when executed by one or more processors, cause the device to identify that the target sequence is positively selected based on the probability of randomly observing n' or more gRNAs of the target sequence in a selected cell. The device is caused to determine that the probability of randomly observing n' or more gRNAs of the target sequence in the selected cell satisfies a threshold value, the device according to any one of Embodiments 67 to 74.
[0210] Embodiment 76 The device according to any one of Embodiments 67 to 76, further comprising processor-executable instructions that, when executed by one or more processors, cause the device to determine a enrichment score for each gRNA.
[0211] Embodiment 77 Processor-executable instructions that, when executed by one or more processors, cause the apparatus to determine a enrichment score for each gRNA, the apparatus according to Embodiment 76 that causes the apparatus to evaluate N / N'.
[0212] Embodiment 78 The apparatus according to any one of Embodiments 67 to 77, wherein the cas-9 positive cells contain inactivated cas-9. Embodiment 79 The apparatus according to Embodiment 78, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0213] Embodiment 80 A non-transitory computer-readable medium for determining the probability of accidentally observing one or more guide RNAs (gRNAs), the processor-executable instructions which, when executed by one or more processors, cause the one or more processors to, based on the sequencing of a first cell population after infection using a vector containing a library of at least three guide RNAs (gRNAs) for each of a plurality of target regions of DNA, determine the respective number (n) of gRNAs present for each of the plurality of target regions of DNA whose read count exceeds a background threshold; based on the respective number of gRNAs for each of the plurality of target regions of DNA, determine the total number (N) of gRNAs present in the first cell population across all target regions of the plurality of target regions of DNA whose read count exceeds a background threshold; based on the sequencing of a second cell population after infection using a vector containing a library of at least three gRNAs for each of the plurality of target regions of DNA, determine the respective number (n') of gRNAs present for each of the plurality of target regions of DNA whose read count exceeds a background threshold; based on the respective number of gRNAs for each of the plurality of target regions of DNA, determine the total number (N') of gRNAs present in the second cell population across all target regions of the plurality of target regions of DNA whose read count exceeds a background threshold; for each target region of the plurality of target regions, determine the probability of accidentally observing n' gRNAs in the target region in the selected cell based on n, N, n', and N'; based on the probability of accidentally observing n' or more gRNAs for the target sequence in the selected cell for the target region containing the sequence of the subject, determine the probability of accidentally observing n' or more gRNAs for the target sequence of the subject in the selected cell, and, processor-executable instructions for identifying that the target sequence is positively selected based on the probability of accidentally observing n' or more gRNAs for the target sequence in the selected cell. A non-transitory computer-readable medium storing the instructions.
[0214] Embodiment 81 The non-transitory computer-readable medium according to Embodiment 80, wherein the first cell population and the second population comprise cas-9 positive cells. Embodiment 82 Processor-executable instructions that, when executed by one or more processors, cause the apparatus to determine, for each of a plurality of target regions of DNA, the respective number (n) of gRNAs present for which the read count exceeds a background threshold, based on sequencing of a first cell population after infection using a vector containing a library of at least three gRNAs for each of the plurality of target regions of DNA, by causing the one or more processors to infect the first cell population with a library of viral vectors, sequence the cells of the first cell population to obtain the read count for each gRNA, receive the first read count data generated thereby, and, based on the first read count data, count each gRNA for which the read count exceeds the background threshold for each of the plurality of target regions of DNA. A non-transitory computer-readable medium according to any one of Embodiments 80-81.
[0215] Embodiment 83 Processor-executable instructions that, when executed by one or more processors, cause the apparatus to determine the total number (N) of gRNAs present in a first cell population over all target regions of a plurality of target regions of DNA for which the read count exceeds a background threshold, based on the respective number of gRNAs for each of the plurality of target regions of DNA, by causing the one or more processors to count each gRNA present in the first cell population for which the read count exceeds the background threshold. A non-transitory computer-readable medium according to any one of Embodiments 80-82.
[0216] Embodiment 84 Processor-executable instructions which, when executed by one or more processors, cause the apparatus to determine, for each of a plurality of target regions of DNA, the respective number (n') of gRNAs present for which the read count exceeds a background threshold, based on sequencing of a second cell population after infection using a vector comprising a library of at least three gRNAs for each of the plurality of target regions of DNA, cause the one or more processors to categorize the cells of the second cell population as either having a specified phenotype or not having the specified phenotype, select the cells having the specified phenotype, sequence the selected cells to obtain the read count for each gRNA, receive second read count data generated thereby, and, based on the second read count data, cause each gRNA for which the read count exceeds the background threshold to be counted for each of the plurality of target regions of DNA, a non-transitory computer-readable medium according to any one of Embodiments 80 to 83.
[0217] Embodiment 85 Processor-executable instructions which, when executed by one or more processors, cause the apparatus to determine, for all target regions of a plurality of target regions of DNA for which the read count exceeds a background threshold, the total number (N') of gRNAs present in a second cell population based on the respective number of gRNAs for each of the plurality of target regions of DNA, cause the one or more processors to count each gRNA present in the selected cells of the second cell population for which the read count exceeds the background threshold, a non-transitory computer-readable medium according to any one of Embodiments 80 to 84.
[0218] Embodiment 86 Processor-executable instructions which, when executed by one or more processors, cause the apparatus to determine, for each target region of a plurality of target regions, the probability of randomly observing n' gRNAs for the target region in the selected cells based on n, N, n', and N', cause the one or more processors to
[0219]
Number
[0220] to evaluate, in the formula
[0221] [Number]
[0222] is a non - transitory computer - readable medium according to any one of Embodiments 80 to 85 that calculates the number of ways to select an x object from a y object. Embodiment 87 Processor - executable instructions which, when executed by one or more processors, cause the apparatus to determine, for a target region containing a target sequence in a selected cell, the probability of randomly observing n or more gRNAs of the target sequence in the selected cell, based on the probability of randomly observing n'gRNAs in the target region within the selected cell. The processor - executable instructions are provided to one or more processors
[0223] [Number]
[0224] to evaluate, a non - transitory computer - readable medium according to any one of Embodiments 80 to 86. Embodiment 88 Processor - executable instructions which, when executed by one or more processors, cause the apparatus to determine, based on the probability of randomly observing n or more gRNAs of a target sequence in a selected cell, that the target sequence is positively selected. The processor - executable instructions cause one or more processors to evaluate that the probability of randomly observing n' or more gRNAs of the target sequence in the selected cell meets a threshold. A non - transitory computer - readable medium according to any one of Embodiments 80 to 87.
[0225] A non-transitory computer-readable medium according to any one of Embodiments 80 to 88, further comprising processor-executable instructions that, when executed by one or more processors, cause the one or more processors to determine an enrichment score for each gRNA.
[0226] A non-transitory computer-readable medium according to Embodiment 89, wherein when executed by one or more processors, the processor-executable instructions cause the apparatus to determine an enrichment score for each gRNA and cause the one or more processors to evaluate N / N', where N / N' is included.
[0227] A non-transitory computer-readable medium according to any one of Embodiments 80 to 90, wherein the cas-9 positive cells contain inactivated cas-9. A non-transitory computer-readable medium according to Embodiment 91, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0228] Those skilled in the art can recognize or confirm a number of equivalents of specific embodiments of the methods and compositions described herein by using only ordinary experiments. Such equivalents are intended to be encompassed by the following claims.
[0229] The technical ideas that can be grasped from the above embodiments are described below as supplementary notes. [Supplementary Note 1] (A) Infecting a first culture of cas9-positive cells with a library of viral vectors, wherein the library comprises at least three guide RNAs (gRNAs) for cleaving a target region of DNA within the genome of the cells, Sequencing the cells to obtain the read count for each of the gRNAs, Summing (Σ) the number of each gRNA whose read count exceeds a background threshold per target region of DNA, where Σ = n in the formula, Summing (Σ) the total number of gRNAs whose read count exceeds the background threshold across all target regions, where Σ = N in the formula; (B) Infecting a second culture of cas9-positive cells with the library of viral vectors; Categorizing the cells of the second culture as either having the designated phenotype or not having the designated phenotype; Selecting the cells having the designated phenotype, sequencing the selected cells to obtain the post-selection read count for each of the gRNAs; Summing (Σ) the respective number of gRNAs per target region of DNA in the selected cells whose post-selection read count exceeds the background threshold, where Σ = n’ in the formula, and; Summing (Σ) the total number of gRNAs of the selected cells whose read count exceeds the threshold across all target regions, where Σ = N’ in the formula; (C) For a target region of DNA, Formula
[0230]
Number
[0231] Calculating the probability of randomly observing n’ gRNAs for the target region in the selected cells according to the formula, where
[0232]
Number
[0233] Calculating, where is the number of ways to select x objects from y objects, and; For a target region of DNA containing a gene, formula
[0234]
Number
[0235] comprising calculating the probability of accidentally observing n' or more gRNAs of a gene in the selected cell according to the following; a method. [Appendix 2] (A) infecting a first culture of cas9-positive cells with a library of viral vectors, wherein the library comprises at least three guide RNAs (gRNAs) for enhancing the transcription of a target region of DNA in the genome of the cell; sequencing the cells to obtain the read count of each of the gRNAs; summing (Σ) the number of each gRNA whose read count exceeds a background threshold per target region of DNA, where Σ = n in the formula; summing, and summing (Σ) the total number of gRNAs whose read count exceeds the background threshold across all target regions, where Σ = N in the formula; summing; and (B) infecting a second culture of cas9-positive cells with the library of viral vectors; categorizing the cells of the second culture as either having a designated phenotype or not having the designated phenotype; selecting the cells having the designated phenotype, sequencing the selected cells to obtain the post-selection read count of each of the gRNAs; in the selected cells whose post-selection read count exceeds the background threshold, summing (Σ) the number of each gRNA per target region of DNA, where Σ = n' in the formula; summing, and summing (Σ) the total number of gRNAs of the selected cells whose read count exceeds the threshold across all target regions, where Σ = N' in the formula; summing; and (C) for a target region of DNA, the formula
[0236] [Formula]
[0237] According to this, it is to calculate the probability of accidentally observing n' gRNAs for the target region in the selected cells, where in the formula
[0238] [Number]
[0239] is to calculate, compute the number of ways to select x objects from y objects, and For the target region of DNA containing the gene, the said formula
[0240] [Number]
[0241] According to this, calculating the probability of accidentally observing n' or more gRNAs of the gene in the selected cells, and a method comprising the same. [Appendix 3] The method according to any one of Appendices 1 to 2, wherein the first culture is infected according to the CRISPR technique.
[0242] [Appendix 4] The method according to any one of Appendices 1 to 3, wherein the cas-9 positive cells of the first culture are modified to contain one or more selection markers.
[0243] [Appendix 5] The method according to any one of Appendices 1 to 4, wherein the specified phenotype is fluorescence. [Appendix 6] The method according to any one of Appendices 1 to 5, wherein the specified phenotype is cell survival.
[0244] [Appendix 7] The method according to Appendix 4, wherein the one or more selection markers include a fluorescent marker. [Appendix 8] The method according to Supplementary Note 7, wherein the fluorescent marker is part of a FRET biosensor.
[0245] [Supplementary Note 9] The method according to Supplementary Note 4, wherein the one or more selection markers are detectable enzymes. [Supplementary Note 10] The method according to Supplementary Note 9, wherein the detectable enzyme is β-galactosidase.
[0246] [Supplementary Note 11] The method according to Supplementary Note 9, wherein the detectable enzyme is luciferase. [Supplementary Note 12] The method according to any one of Supplementary Notes 1 to 11, wherein the target region contains a gene.
[0247] [Supplementary Note 13] The method according to any one of Supplementary Notes 1 to 12, including applying a selection mechanism to the second cell culture by categorizing the cells of the second culture as either having the specified phenotype or not having the specified phenotype.
[0248] [Supplementary Note 14] The method according to Supplementary Note 13, wherein the selection mechanism includes one or more of exposing the second cell population to a drug or exposing the second cell population to a substance that identifies protein activity or expression level.
[0249] [Supplementary Note 15] The method according to Supplementary Note 4, including sorting the cells based on the one or more selection markers to select the cells having the specified phenotype.
[0250] [Supplementary Note 16] The method according to any one of Supplementary Notes 1 to 15, further including identifying that the target region is positively selected based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cells.
[0251] [Appendix 17] The method according to Appendix 16, comprising determining that the target region is positively selected when the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cells meets a threshold value.
[0252] [Appendix 18] The method according to any one of Appendices 1 to 17, further comprising determining an enrichment score for each gRNA.
[0253] [Appendix 19] The method according to Appendix 18, wherein determining the enrichment score for each gRNA includes evaluating N / N'.
[0254] [Appendix 20] The method according to any one of Appendices 2 to 19, wherein the cas-9 positive cells contain inactivated cas-9.
[0255] [Appendix 21] The method according to Appendix 20, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0256] [Appendix 22] The method according to any one of Appendices 1 to 21, wherein the target region of the DNA regulates a downstream gene or protein.
[0257] [Appendix 23] The method according to Appendix 22, wherein regulating a downstream gene or protein includes activating or inhibiting the downstream gene or protein.
[0258] [Appendix 24] The method according to any one of Appendices 1 to 23, further comprising identifying the target region as a modifying factor for a second gene based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cells.
[0259] [Appendix 25] The method according to any one of Appendices 1 to 24, further comprising identifying the target region as a therapeutic target based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cell.
[0260] [Appendix 26] The method according to any one of Appendices 1 to 25, further comprising identifying that the target region correlates with the designated phenotype based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cell.
[0261] [Appendix 27] The method according to any one of Appendices 1 to 26, further comprising identifying that the target region exhibits a protective effect based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cell.
[0262] [Appendix 28] Determining, for each of a plurality of target regions of DNA, the respective number (n) of gRNAs present for each of the plurality of target regions of DNA, the read count of which exceeds a background threshold, based on sequencing a first cell population after infection with a vector comprising a library of at least three guide RNAs (gRNAs) for each of the plurality of target regions of DNA; Determining the total number (N) of gRNAs present in the first cell population over all target regions of the plurality of target regions of DNA, the read count of which exceeds the background threshold, based on the respective number of the gRNAs for each of the plurality of target regions of DNA; Determining, for each of a plurality of target regions of DNA, the respective number (n') of gRNAs present for each of the plurality of target regions of DNA, the read count of which exceeds the background threshold, based on sequencing a second cell population after infection with a vector comprising the library of at least three gRNAs for each of the plurality of target regions of DNA; Determining the total number (N’) of gRNAs present in the second cell population across all target regions of the plurality of target regions of the DNA, where the read count exceeds the background threshold, based on the respective number of each of the gRNAs for each of the plurality of target regions of the DNA; For each target region of the plurality of target regions, determining the probability of accidentally observing n’ gRNAs for the target region in the selected cell based on n, N, n’, and N’; For a target region containing a target sequence, determining the probability of accidentally observing n’ or more gRNAs of the target sequence in the selected cell based on the probability of accidentally observing n’ gRNAs for the target region in the selected cell; Identifying that the target sequence is positively selected based on the probability of accidentally observing n’ or more gRNAs of the target sequence in the selected cell, the method comprising.
[0263] [Appendix 29] The method according to Appendix 28, wherein the first cell population and the second population comprise cas-9 positive cells.
[0264] [Appendix 30] Based on sequencing a first cell population after infection with a vector containing a library of at least three gRNAs for each of a plurality of target regions of DNA, determining the respective number (n) of each of the gRNAs present for each of the plurality of target regions of the DNA, where the read count exceeds the background threshold; Infecting the first cell population with the library of viral vectors; Sequencing the cells of the first cell population to obtain the read count of each of the gRNAs; Counting each gRNA whose read count exceeds the background threshold for each of the plurality of target regions of the DNA, the method according to Appendix 29 comprising.
[0265] [Appendix 31] Determining the total number (N) of gRNAs present in the first cell population over all target regions of the plurality of target regions of the DNA, where the read count of each gRNA exceeds the background threshold, based on the respective number of each of the gRNAs for each of the plurality of target regions of the DNA, including counting each gRNA present in the first cell population whose read count exceeds the background threshold, according to the method described in Supplementary Note 30.
[0266] [Supplementary Note 32] Based on the sequencing of a second cell population after infection using a vector containing a library of at least three gRNAs for each of the plurality of target regions of the DNA, determining the respective number of gRNAs (n') present for each of the plurality of target regions of the DNA, where the read count of each gRNA exceeds the background threshold, categorizing the cells of the second cell population as either having a specified phenotype or not having the specified phenotype, selecting the cells having the specified phenotype, sequencing the selected cells to obtain the read count of each of the gRNAs, including counting each gRNA whose read count exceeds the background threshold for each of the plurality of target regions of the DNA, according to the method described in Supplementary Note 31.
[0267] [Supplementary Note 33] Determining the total number (N') of gRNAs present in the second cell population over all target regions of the plurality of target regions of the DNA, where the read count of each gRNA exceeds the background threshold, based on the respective number of each of the gRNAs for each of the plurality of target regions of the DNA, including counting each gRNA present in the selected cells of the second cell population whose read count exceeds the background threshold, according to any one of Supplementary Notes 28 to 32.
[0268] [Supplementary Note 34] For each of the plurality of target regions, determining the probability of accidentally observing n' gRNAs in the target region in the selected cell based on n, N, n', and N'
[0269]
Number
[0270] includes evaluating, where in the formula
[0271]
Number
[0272] is the method according to any one of Appendices 28 to 33 for calculating the number of ways to select an x object from a y object. [Appendix 35] For a target region containing a target sequence, determining the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cell based on the probability of accidentally observing n' gRNAs in the target region in the selected cell
[0273]
Number
[0274] includes evaluating, and is the method according to any one of Appendices 28 to 34. [Appendix 36] Identifying that the target sequence is positively selected based on the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cell includes determining that the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cell meets a threshold value, and is the method according to any one of Appendices 28 to 35.
[0275] [Appendix 37] The method according to any one of Appendices 28 to 36, further comprising determining an enrichment score for each gRNA.
[0276] [Appendix 38] The method according to Appendix 37, wherein determining an enrichment score for each gRNA comprises evaluating N / N'.
[0277] [Appendix 39] The method according to any one of Appendices 28 to 38, wherein the cas-9 positive cells comprise inactivated cas-9.
[0278] [Appendix 40] The method according to Appendix 39, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0279] [Appendix 41] An apparatus comprising: One or more processors; Processor-executable instructions that, when executed by the one or more processors, cause the apparatus to: (A) Receive first read count data generated by: Infecting a first culture of cas9 positive cells with a library of viral vectors, the library comprising at least three guide RNAs (gRNAs) for cleaving target regions of DNA within the genome of the cells; and Sequencing the cells to obtain the read count for each of the gRNAs; Based on the first read count data, sum the number of each gRNA whose read count per target region of DNA exceeds a background threshold, where Σ = n in the formula; and Based on the first read count data, sum the total number of gRNAs whose read count exceeds the background threshold across all target regions, where Σ = N in the formula; (B) Second read count data generated by: Infecting a second culture of cas9-positive cells with the library of viral vectors, and categorizing the cells of the second culture as either having or not having a designated phenotype; receiving second read count data generated by selecting the cells having the designated phenotype, sequencing the selected cells, and obtaining the number of reads per selection for each of the gRNAs; summing, for each of the gRNAs, the number of each in the selected cells, per target region of DNA, where the number of reads after selection exceeds the background threshold, based on the second read count data, where Σ = n’ in the formula; and summing the total number of gRNAs in the selected cells, where the number of reads exceeds the threshold, across all target regions, based on the second read count data, where Σ = N’ in the formula; (C) For a target region of DNA, the formula
[0280]
Number
[0281] calculating the probability of randomly observing n’ gRNAs in the target region in the selected cells according to the formula, where
[0282]
Number
[0283] calculating the number of ways to select x objects from y objects, and for a target region of DNA containing a gene, the formula
[0284]
Number
[0285] A memory storing processor-executable instructions for causing a calculation of the probability of accidentally observing n' or more gRNAs of a gene in the selected cell according to the above. An apparatus comprising the memory. [Appendix 42] The apparatus according to Appendix 41, wherein the first culture is infected according to CRISPR technology.
[0286] [Appendix 43] The apparatus according to any one of Appendices 41 to 42, wherein the cas-9 positive cells of the first culture are modified to contain one or more selection markers.
[0287] [Appendix 44] The apparatus according to any one of Appendices 41 to 43, wherein the target region contains a gene. [Appendix 45] The apparatus according to any one of Appendices 41 to 44, wherein applying a selection mechanism to the second cell culture includes categorizing the cells of the second culture as either having a specified phenotype or not having the specified phenotype.
[0288] [Appendix 46] The apparatus according to Appendix 45, wherein the selection mechanism includes one or more of exposing the second cell population to a drug or exposing the second cell population to a substance that identifies protein activity or expression level.
[0289] [Appendix 47] The apparatus according to any one of Appendices 41 to 46, wherein selecting the cells having the specified phenotype includes sorting the cells based on the one or more selection markers.
[0290] [Appendix 48] The apparatus according to any one of Appendices 41 to 47, further comprising identifying that the target region is positively selected based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cell.
[0291] [Appendix 49] The apparatus according to Appendix 48, wherein identifying that the target region is positively selected includes determining that the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cells meets a threshold value.
[0292] [Appendix 50] The apparatus according to any one of Appendices 41 to 49, further comprising determining a enrichment score for each gRNA.
[0293] [Appendix 51] The apparatus according to Appendix 50, wherein determining a enrichment score for each gRNA includes evaluating N / N'.
[0294] [Appendix 52] The apparatus according to any one of Appendices 41 to 51, wherein the cas-9 positive cells contain inactivated cas-9.
[0295] [Appendix 53] The apparatus according to Appendix 52, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0296] [Appendix 54] A non-transitory computer-readable medium for determining the probability of accidentally observing one or more guide RNAs (gRNAs), the processor-executable instructions, when executed by one or more processors, cause the one or more processors to, (A) First read count data, infecting a first culture of cas9 positive cells with a library of viral vectors, the library containing at least three guide RNAs (gRNAs) for cleaving a target region of DNA within the genome of the cells, and sequencing the cells to obtain the read count of each of the gRNAs, and receiving the first read count data generated thereby. The sum (Σ), where Σ = n, of the respective numbers of gRNAs per target region of DNA, for which the read count exceeds the background threshold, based on the first read count data, and The sum (Σ), where Σ = N, of the total number of gRNAs, for which the read count exceeds the background threshold, across all target regions, based on the first read count data; (B) Second read count data, Infecting a second culture of cas9-positive cells with the library of viral vectors, Categorizing the cells of the second culture as either having the designated phenotype or not having the designated phenotype, Selecting the cells having the designated phenotype, sequencing the selected cells to obtain the post-selection read count for each of the gRNAs, and receiving the second read count data generated thereby, The sum (Σ), where Σ = n’, of the respective numbers of gRNAs per target region of DNA in the selected cells, for which the post-selection read count exceeds the background threshold, based on the second read count data, and The sum (Σ), where Σ = N’, of the total number of gRNAs in the selected cells, for which the read count exceeds the threshold, across all target regions, based on the second read count data; (C) For a target region of DNA, The formula
[0297]
Number
[0298] Calculating the probability of randomly observing n’ gRNAs in the target region in the selected cells according to the formula, where
[0299]
Number
[0300] calculates the number of ways to select an x object from a y object, causes a calculation, and For a target region of DNA containing a gene, the formula
[0301] [Number]
[0302] A non-transitory computer-readable medium storing processor-executable instructions for causing a calculation of the probability of randomly observing n' or more gRNAs of a gene in the selected cell according to the formula. [Appendix 55] The non-transitory computer-readable medium according to Appendix 54, wherein the first culture is infected according to the CRISPR technique.
[0303] [Appendix 56] The non-transitory computer-readable medium according to any one of Appendices 54 to 55, wherein the cas-9 positive cells of the first culture are modified to contain one or more selection markers.
[0304] [Appendix 57] The non-transitory computer-readable medium according to any one of Appendices 54 to 56, wherein the target region contains a gene.
[0305] [Appendix 58] The non-transitory computer-readable medium according to any one of Appendices 54 to 57, wherein applying a selection mechanism to the second cell culture includes categorizing the cells of the second culture as having either the designated phenotype or not having the designated phenotype.
[0306] [Appendix 59] The non-transitory computer-readable medium according to Appendix 58, wherein the selection mechanism includes one or more of exposing the second cell population to a drug or exposing the second cell population to a substance that identifies protein activity or expression level.
[0307] [Appendix 60] The non-transitory computer-readable medium according to any one of Appendices 54 to 59, wherein selecting the cells having the specified phenotype includes sorting the cells based on the one or more selection markers.
[0308] [Appendix 61] The non-transitory computer-readable medium according to any one of Appendices 54 to 60, further comprising identifying that the target region is positively selected based on the probability of accidentally observing n' or more gRNAs of the gene in the selected cells.
[0309] [Appendix 62] The non-transitory computer-readable medium according to Appendix 61, wherein identifying that the target region is positively selected includes determining that the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cells meets a threshold.
[0310] [Appendix 63] The non-transitory computer-readable medium according to any one of Appendices 54 to 62, further comprising determining an enrichment score for each gRNA.
[0311] [Appendix 64] The non-transitory computer-readable medium according to Appendix 63, wherein determining the enrichment score for each gRNA includes evaluating N / N'.
[0312] [Appendix 65] The non-transitory computer-readable medium according to any one of Appendices 54 to 64, wherein the cas-9 positive cells contain inactivated cas-9.
[0313] [Appendix 66] The non-transitory computer-readable medium according to Appendix 65, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0314] [Appendix 67] An apparatus comprising one or more processors, and processor-executable instructions that, when executed by the one or more processors, cause the apparatus to determine, for each of a plurality of target regions of DNA, the respective number (n) of guide RNAs (gRNAs) present for that target region of DNA, based on sequencing a first cell population after infection with a vector comprising a library of at least three gRNAs for each of the plurality of target regions of DNA, wherein the read count for that target region exceeds a background threshold; determine, for all target regions of the plurality of target regions of DNA, the total number (N) of gRNAs present in the first cell population, across all target regions of the plurality of target regions of DNA, based on the respective number of gRNAs for each of the plurality of target regions of DNA, wherein the read count for that target region exceeds the background threshold; determine, for each of a plurality of target regions of DNA, the respective number (n’) of gRNAs present for that target region of DNA, based on sequencing a second cell population after infection with a vector comprising a library of at least three gRNAs for each of the plurality of target regions of DNA, wherein the read count for that target region exceeds the background threshold; determine, for all target regions of the plurality of target regions of DNA, the total number (N’) of gRNAs present in the second cell population, across all target regions of the plurality of target regions of DNA, based on the respective number of gRNAs for each of the plurality of target regions of DNA, wherein the read count for that target region exceeds the background threshold; for each target region of the plurality of target regions, determine the probability of accidentally observing n’ gRNAs for that target region in the selected cell, based on n, N, n’, and N’; for a target region comprising a sequence of interest, determine the probability of accidentally observing n’ or more gRNAs for the sequence of interest in the selected cell, based on the probability of accidentally observing n’ gRNAs for that target region in the selected cell; and A memory storing processor-executable instructions for causing the selection of the target sequence to be positively identified based on the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cell. An apparatus comprising the memory.
[0315] [Appendix 68] The apparatus according to Appendix 67, wherein the first cell population and the second population include cas-9 positive cells.
[0316] [Appendix 69] The processor-executable instructions which, when executed by the one or more processors, cause the apparatus to determine, for each of the plurality of target regions of DNA, the number (n) of each gRNA present based on sequencing a first cell population after infection with a vector containing a library of at least 3 gRNAs for each of the plurality of target regions of DNA, wherein the read count exceeds a background threshold. The processor-executable instructions cause the apparatus to Infect the first cell population with the library of viral vectors; Sequence the cells of the first cell population to obtain the read count of each of the gRNAs, and receive first read count data generated thereby, and Based on the first read count data, cause the apparatus to count each gRNA for which the read count exceeds the background threshold for each of the plurality of target regions of DNA. The apparatus according to any one of Appendices 67 to 68.
[0317] [Appendix 70] The processor-executable instructions which, when executed by the one or more processors, cause the apparatus to determine, for all target regions of the plurality of target regions of DNA, the total number (N) of gRNAs present in the first cell population, based on the number of each gRNA for each of the plurality of target regions of DNA, wherein the read count exceeds the background threshold. The processor-executable instructions cause the apparatus to count each gRNA present in the first cell population for which the read count exceeds the background threshold. The apparatus according to Appendix 69.
[0318] [Appendix 71] The processor-executable instructions which, when executed by the one or more processors, cause the apparatus to determine, for each of the plurality of target regions of the DNA, the respective number (n') of gRNAs present therein based on sequencing a second cell population after infection using a vector containing a library of at least three gRNAs for each of the plurality of target regions of the DNA, where the read count thereof exceeds the background threshold, cause the apparatus to categorize the cells of the second cell population as either having a designated phenotype or not having the designated phenotype, select the cells having the designated phenotype, sequence the selected cells to obtain the read count for each of the gRNAs, and receive second read count data generated thereby, and cause the apparatus to count, for each of the plurality of target regions of the DNA, each gRNA whose read count exceeds the background threshold, based on the second read count data, the apparatus according to any one of Appendices 67 to 70.
[0319] [Appendix 72] The processor-executable instructions which, when executed by the one or more processors, cause the apparatus to determine, based on the respective number of the gRNAs for each of the plurality of target regions of the DNA, the total number (N') of gRNAs present in the second cell population over all target regions of the plurality of target regions of the DNA whose read count exceeds the background threshold, cause the apparatus to count each gRNA present in the selected cells of the second cell population whose read count exceeds the background threshold, the apparatus according to any one of Appendices 67 to 71.
[0320] [Appendix 73] Processor-executable instructions that, when executed by the one or more processors, cause the device to determine, for each target region of the plurality of target regions, the probability of randomly observing n' gRNAs in the selected target region within the cell based on n, N, n', and N'. The processor-executable instructions cause the device
[0321] [Number]
[0322] to evaluate, where in the formula
[0323] [Number]
[0324] is a device according to any one of Appendices 67 to 72 for calculating the number of ways to select an x object from a y object. [Appendix 74] Processor-executable instructions that, when executed by the one or more processors, cause the device to determine, for a target region containing a target sequence, the probability of randomly observing n' or more gRNAs of the target sequence in the selected cell based on the probability of randomly observing n' gRNAs in the selected target region within the cell. The processor-executable instructions cause the device
[0325] [Number]
[0326] to evaluate, which is a device according to any one of Appendices 67 to 73. [Appendix 75] The processor-executable instructions, when executed by the one or more processors, cause the apparatus to identify that the target sequence is positively selected based on the probability of randomly observing n' or more gRNAs of the target sequence in the selected cell. The processor-executable instructions cause the apparatus to determine that the probability of randomly observing n' or more gRNAs of the target sequence in the selected cell meets a threshold. The apparatus according to any one of Appendices 67 to 74.
[0327] [Appendix 76] The apparatus according to any one of Appendices 67 to 76, further comprising processor-executable instructions that, when executed by the one or more processors, cause the apparatus to determine a enrichment score for each gRNA.
[0328] [Appendix 77] The apparatus according to Appendix 76, wherein the processor-executable instructions, when executed by the one or more processors, cause the apparatus to determine a enrichment score for each gRNA, and cause the apparatus to evaluate N / N'.
[0329] [Appendix 78] The apparatus according to any one of Appendices 67 to 77, wherein the cas-9 positive cell contains inactivated cas-9.
[0330] [Appendix 79] The apparatus according to Appendix 78, wherein the inactivated cas-9 is fused to at least one transcriptional activation domain.
[0331] [Appendix 80] A non-transitory computer-readable medium for determining the probability of randomly observing one or more guide RNAs (gRNAs), comprising processor-executable instructions that, when executed by one or more processors, cause the one or more processors to Sequencing a first cell population after infection using a vector containing a library of at least three guide RNAs (gRNAs) for each of a plurality of target regions of DNA, and determining, for each of the plurality of target regions of DNA, the respective number (n) of gRNAs present for which the read count exceeds a background threshold value. Based on the respective number of the gRNAs for each of the plurality of target regions of the DNA, determining the total number (N) of gRNAs present in the first cell population across all target regions of the plurality of target regions of the DNA for which the read count exceeds the background threshold value. Sequencing a second cell population after infection using the vector containing the library of at least three gRNAs for each of the plurality of target regions of DNA, and determining, for each of the plurality of target regions of DNA, the respective number (n') of gRNAs present for which the read count exceeds the background threshold value. Based on the respective number of the gRNAs for each of the plurality of target regions of the DNA, determining the total number (N') of gRNAs present in the second cell population across all target regions of the plurality of target regions of the DNA for which the read count exceeds the background threshold value. For each target region of the plurality of target regions, determining the probability of randomly observing n' gRNAs for the target region in the selected cell based on n, N, n', and N'. For a target region containing a sequence of interest, determining the probability of randomly observing n' or more gRNAs for the sequence of interest in the selected cell based on the probability of randomly observing n' gRNAs for the target region in the selected cell, and Based on the probability of randomly observing n' or more gRNAs for the sequence of interest in the selected cell, identifying that the sequence of interest is positively selected, a non-transitory computer-readable medium storing processor-executable instructions.
[0332] [Appendix 81] The non-transitory computer-readable medium according to appendix 80, wherein the first cell population and the second population comprise cas-9 positive cells.
[0333] [Appendix 82] The processor-executable instructions that, when executed by the one or more processors, cause the apparatus to determine, for each of a plurality of target regions of DNA, the respective number (n) of gRNAs present for which the read count exceeds a background threshold, based on sequencing a first cell population after infection with a vector comprising a library of at least three gRNAs for each of the plurality of target regions of DNA, cause the one or more processors to infect the first cell population with the library of viral vectors; sequence the cells of the first cell population to obtain the read count for each of the gRNAs, and receive first read count data generated thereby; and The non-transitory computer-readable medium according to any one of appendices 80 to 81, which causes, for each of the plurality of target regions of DNA, the apparatus to count each gRNA for which the read count exceeds the background threshold, based on the first read count data.
[0334] [Appendix 83] The processor-executable instructions that, when executed by the one or more processors, cause the apparatus to determine, for all target regions of the plurality of target regions of DNA for which the read count exceeds the background threshold, the total number (N) of gRNAs present in the first cell population, based on the respective number of each of the gRNAs for each of the plurality of target regions of DNA, cause the one or more processors to count each gRNA present in the first cell population for which the read count exceeds the background threshold. The non-transitory computer-readable medium according to any one of appendices 80 to 82.
[0335] [Appendix 84] Processor-executable instructions that, when executed by the one or more processors, cause the apparatus to determine, based on sequencing a second cell population after infection using a vector containing a library of at least three gRNAs for each of a plurality of target regions of the DNA, the respective number (n') of gRNAs present for each of the plurality of target regions of the DNA whose read count exceeds the background threshold, cause the one or more processors to categorize the cells of the second cell population as either having a specified phenotype or not having the specified phenotype; select the cells having the specified phenotype; receive second read count data generated by sequencing the selected cells to obtain the read count of each of the gRNAs; and cause, based on the second read count data, each gRNA whose read count exceeds the background threshold to be counted for each of the plurality of target regions of the DNA, the non-transitory computer-readable medium according to any one of Appendices 80 to 83.
[0336] [Appendix 85] Processor-executable instructions that, when executed by the one or more processors, cause the apparatus to determine, based on the respective number of gRNAs for each of the plurality of target regions of the DNA, the total number (N') of gRNAs present in the second cell population over all target regions of the DNA whose read count exceeds the background threshold, cause the one or more processors to count each gRNA present in the selected cells of the second cell population whose read count exceeds the background threshold, the non-transitory computer-readable medium according to any one of Appendices 80 to 84.
[0337] [Appendix 86] Processor-executable instructions that, when executed by the one or more processors, cause the apparatus to determine, for each target region of the plurality of target regions, the probability of randomly observing n' gRNAs in the target region within the selected cell based on n, N, n', and N' are the one or more processors
[0338] [Number]
[0339] to evaluate, where in the formula
[0340] [Number]
[0341] is a non-transitory computer-readable medium according to any one of Appendices 80 to 85 for calculating the number of ways to select an x object from a y object [Appendix 87] Processor-executable instructions that, when executed by the one or more processors, cause the apparatus to determine, for a target region containing a target sequence, the probability of randomly observing n' or more gRNAs of the target sequence in the selected cell based on the probability of randomly observing n' gRNAs in the target region within the selected cell are the one or more processors
[0342] [Number]
[0343] to evaluate, which is a non-transitory computer-readable medium according to any one of Appendices 80 to 86 [Appendix 88] Processor-executable instructions which, when executed by the one or more processors, cause the device to identify that the target sequence is positively selected based on the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cell, the processor-executable instructions cause the one or more processors to determine that the probability of accidentally observing n' or more gRNAs of the target sequence in the selected cell meets a threshold, a non-transitory computer-readable medium according to any one of appendices 80 to 87.
[0344] [Appendix 89] Processor-executable instructions which, when executed by the one or more processors, cause the one or more processors to determine an enrichment score for each gRNA, further included in a non-transitory computer-readable medium according to any one of appendices 80 to 88.
[0345] [Appendix 90] Processor-executable instructions which, when executed by the one or more processors, cause the device to determine an enrichment score for each gRNA, the processor-executable instructions cause the one or more processors to evaluate N / N', a non-transitory computer-readable medium according to appendix 89.
[0346] [Appendix 91] The cas-9 positive cell contains inactivated cas-9, a non-transitory computer-readable medium according to any one of appendices 80 to 90.
[0347] [Appendix 92] The inactivated cas-9 is fused to at least one transcriptional activation domain, a non-transitory computer-readable medium according to appendix 91.
Claims
1. infecting a first culture of cas9 positive cells with a library of viral vectors, the library comprising at least three guide RNAs (gRNAs) for cleaving a target region of DNA within the genome of the cells; sequencing the cells to obtain a number of reads for each of the gRNAs; Summarizing the number of qualitatively present gRNAs (n) per target region of DNA whose read count exceeds the background threshold (Σ); and summing (Σ) the total number (N) of qualitatively present gRNAs whose read counts exceed the background threshold across all target regions; infecting a second culture of cas9 positive cells with the library of viral vectors; categorizing cells of the second culture as either having a designated phenotype or not having the designated phenotype based on applying a selection mechanism to cells of the second culture, wherein cas9 positive cells of the second culture having the designated phenotype express a selection marker that comprises a signal transduction pathway protein; selecting cells having the specified phenotype and sequencing the selected cells to obtain a post-selection read count for each of the gRNAs; Summarizing the number of gRNAs (n') that are qualitatively present per target region of DNA in the selected cells whose post-selection read counts exceed the background threshold (Σ); and summing (Σ) the total number (N′) of gRNAs qualitatively present in the selected cells whose read counts exceed the background threshold across all target regions; For a target region of DNA: formula [0010] calculating the probability of observing a n'gRNA for the target region in the selected cell by chance according to [0025] calculate the number of ways to select an x object from a y object; For a target region of DNA containing a gene, the formula [0030] Calculating the probability of observing n' or more gRNAs of a gene in the selected cell by chance according to: Identifying that the target region affects the signaling pathway based on the probability of observing n' or more gRNAs of the gene in the selected cells by chance; A method comprising:
2. 2. The method of claim 1, wherein the target region of DNA regulates a downstream gene or protein, and regulating the downstream gene or protein comprises activating or inhibiting the downstream gene or protein.
3. The method of claim 1 or claim 2, wherein the designated phenotype is at least one of fluorescence or cell survival, and the selection mechanism includes one or more of exposing the cas9 positive cells of the second culture to a drug or a substance that identifies protein activity or expression levels.
4. The method of any one of claims 1 to 3, wherein the signaling pathway comprises a disease pathway.
5. Identifying that the target region affects the signaling pathway based on the probability of observing n' or more gRNAs of the gene in the selected cell by chance includes: identifying the target regions as positively selected based on the probability of observing n' or more gRNAs of the gene in the selected cells by chance; and determining, based on the target region being positively selected, that the target region modulates a protein of the signal transduction pathway; The method according to any one of claims 1 to 4, comprising:
6. identifying the target region as positively selected based on the probability of observing n' or more gRNAs of the gene in the selected cells by chance; identifying the target region as a modifier of a second gene based on the probability of observing n' or more gRNAs of the gene in the selected cells by chance; identifying the target region as a therapeutic target based on the probability of observing n' or more gRNAs of the gene in the selected cells by chance; identifying the target region as correlated with the specified phenotype based on the probability of observing n' or more gRNAs of the gene in the selected cells by chance; or identifying the target region as exhibiting a protective effect based on the probability of observing n' or more gRNAs of the gene in the selected cells by chance; The method of any one of claims 1 to 5, further comprising one or more of:
7. calculating an enrichment score for the target region by evaluating N / N'; and identifying said target regions as positively selected target regions based on said target regions having an enrichment score above a threshold; The method of any one of claims 1 to 6, further comprising:
8. 8. The method of claim 7, wherein identifying the target region as affecting the signaling pathway is further based on the target region having an enrichment score above the threshold.
9. summing (Σ) the number of guide RNAs (gRNAs) qualitatively present in the cas9 positive cells of the first culture after infection with the library of viral vectors per target region of DNA, the number of gRNAs having a read count above a background threshold; and wherein the library comprises at least three gRNAs for cleaving target regions of DNA in the genome of the cells; Summing up the total number (N) of qualitatively present gRNAs whose read counts exceed the background threshold across all target regions (Σ); categorizing cells of the second culture as either having a designated phenotype or not having said designated phenotype based on applying a selection mechanism to cas9 positive cells of the second culture following infection with a library of viral vectors, wherein cas9 positive cells of the second culture having said designated phenotype express a selection marker comprising a protein of a signal transduction pathway; summing (Σ) the number of each gRNA (n') that is qualitatively present per targeted region of DNA in cas9 positive cells of the second culture having the specified phenotype whose post-selection read count exceeds the background threshold; and summing the total number (N') of gRNAs qualitatively present in cas9 positive cells of the second culture whose read counts exceed the background threshold across all target regions (Σ); For a target region of DNA: formula [0045] calculating the probability of observing a n'gRNA for the target region by chance in a cas9 positive cell of the second culture having the specified phenotype according to the formula: [0050] calculate the number of ways to select an x object from a y object; For a target region of DNA containing a gene, the formula [006] calculating the probability of observing, by chance, n' or more gRNAs of a gene in cas9 positive cells of the second culture having the specified phenotype according to identifying the target region as affecting the signaling pathway based on the probability of observing, by chance, n' or more gRNAs for the gene in cas9 positive cells of the second culture having the specified phenotype; A method comprising:
10. infecting a first culture of cas9 positive cells with a library of viral vectors, the library comprising at least three guide RNAs (gRNAs) for cleaving a target region of DNA within the genome of the cells; and sequencing cas9 positive cells of said first culture to obtain a number of reads for each of said gRNAs; The method of claim 9 further comprising:
11. infecting a second culture of cas9 positive cells with the library of viral vectors; and selecting cas9 positive cells of the second culture having a designated phenotype and sequencing the selected cas9 positive cells to obtain a post-selection number of reads for each of the gRNAs; 11. The method of claim 9 or claim 10, further comprising:
12. 12. The method of any one of claims 9 to 11, wherein the target region of DNA regulates a downstream gene or protein, and regulating the downstream gene or protein comprises activating or inhibiting the downstream gene or protein.
13. 13. The method of any one of claims 9 to 12, wherein the designated phenotype is at least one of fluorescence or cell survival, and the selection mechanism comprises one or more of exposing the cas9 positive cells of the second culture to a drug or a substance that identifies protein activity or expression levels.
14. The method of any one of claims 9 to 13, wherein the signaling pathway comprises a disease pathway.
15. Identifying the target region as affecting the signaling pathway based on the probability of observing n' or more gRNAs for the gene in cas9 positive cells of the second culture by chance having the specified phenotype comprises: identifying the target region as positively selected based on the probability of observing by chance n' or more gRNAs for the gene in cas9 positive cells of the second culture having the specified phenotype; and determining, based on the target region being positively selected, that the target region modulates a protein of the signal transduction pathway; The method according to any one of claims 9 to 14, comprising:
16. calculating an enrichment score for the target region by evaluating N / N'; and identifying said target regions as positively selected target regions based on said target regions having an enrichment score above a threshold; The method of any one of claims 9 to 15, further comprising:
17. 17. The method of claim 16, wherein identifying the target region as affecting the signaling pathway is further based on the target region having an enrichment score above the threshold.
18. summing the number of each guide RNA (gRNA) qualitatively present in the cas9 positive cells of the first culture following infection with a library of viral vectors per target region of DNA, the number of each gRNA whose read count exceeds a background threshold, wherein the library comprises at least three gRNAs for cleaving a target region of DNA in the genome of the cas9 positive cells; and Summarizing the total number of gRNAs qualitatively present in cas9 positive cells of said first culture after infection with a library of viral vectors whose read counts exceed said background threshold across all target regions; categorizing the cas9 positive cells of the second culture as either having a designated phenotype or not having said designated phenotype based on applying a selection mechanism to the cas9 positive cells of the second culture following infection with a library of viral vectors, wherein the cas9 positive cells of the second culture having said designated phenotype express a selection marker comprising a signal transduction pathway protein; summing the number of each gRNA qualitatively present per targeted region of DNA in cas9 positive cells of the second culture having the specified phenotype whose post-selection read count exceeds the background threshold; and summing the total number of gRNAs qualitatively present in cas9 positive cells of said second culture having said specified phenotype whose read counts exceed said background threshold across all target regions; calculating the probability, for a target region of DNA, of observing by chance a number of gRNAs qualitatively present for said target region in cas9 positive cells of said second culture having said specified phenotype; For a target region of DNA that contains a gene, calculating the probability of observing by chance the number of gRNAs for that gene in cas9 positive cells of the second culture having the specified phenotype; and identifying the target region as affecting the signaling pathway based on the probability of observing by chance the number of gRNAs for the gene in cas9 positive cells of the second culture having the specified phenotype; A method comprising:
19. infecting a first culture of cas9 positive cells with a library of viral vectors, the library comprising at least three guide RNAs (gRNAs) for cleaving a target region of DNA within the genome of the cells; sequencing cas9 positive cells of said first culture to obtain a number of reads for each of said gRNAs; infecting a second culture of cas9 positive cells with the library of viral vectors; and selecting cas9 positive cells of the second culture having a designated phenotype and sequencing the selected cas9 positive cells to obtain a post-selection number of reads for each of the gRNAs; 20. The method of claim 18, further comprising:
20. Calculating the probability, for a target region of DNA, of observing by chance the number of gRNAs for the target region in a cas9 positive cell of the second culture having the specified phenotype, formula [0070] (In the ceremony [0080] calculates the number of ways to select an x object from a y object) and Calculating the probability of observing, by chance, the number of gRNAs for a gene in cas9 positive cells of the second culture having the specified phenotype for a target region of DNA that contains a gene, formula [0090] 20. The method of claim 18 or 19, comprising evaluating:
21. 21. The method of any one of claims 18 to 20, wherein the target region of DNA regulates a downstream gene or protein, and regulating the downstream gene or protein comprises activating or inhibiting the downstream gene or protein.
22. 22. The method of any one of claims 18-21, wherein the designated phenotype is at least one of fluorescence or cell survival, and the selection mechanism comprises one or more of exposing the cas9 positive cells of the second culture to a drug or substance that identifies protein activity or expression levels.
23. The method of any one of claims 18 to 22, wherein the signaling pathway comprises a disease pathway.
24. Identifying the target region as affecting the signaling pathway based on the probability of observing a certain number of gRNAs for the gene in cas9 positive cells of the second culture by chance includes: identifying the target region as positively selected based on the probability of observing a certain number of gRNAs for the gene in cas9 positive cells of the second culture by chance; and determining, based on the target region being positively selected, that the target region modulates a protein of the signal transduction pathway; The method according to any one of claims 18 to 23, comprising:
25. calculating an enrichment score for the target region by evaluating N / N'; and identifying a target region as a positively selected target region based on the target region having an enrichment score above a threshold, wherein identifying the target region as affecting the signaling pathway is further based on the target region having an enrichment score above the threshold; The method of any one of claims 18 to 24, further comprising:
Citation Information
Patent Citations
Assays for massively combinatorial perturbation profiling and cellular circuit reconstruction
WO2017075294A1