Crispr / CAS screening platform to identify genetic modifiers of tau seeding or aggregation
The CRISPR/Cas screening platform effectively identifies genetic modifiers of tau aggregation in neurodegenerative diseases by using biosensor cells and conditioned medium to detect guide RNAs that enhance or disrupt tau aggregation, addressing the limitations of existing screening methods.
Patent Information
- Application Number
- JP2025185469
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-03-18
- Filing Date
- 2025-11-04
- Publication Date
- 2026-01-23
AI Technical Summary
Current methods are inadequate for identifying genetic modifiers of tau aggregation, which is a key factor in neurodegenerative diseases such as Alzheimer's and Parkinson's, as they do not effectively screen for genes that alter protein aggregation or intercellular propagation.
A CRISPR/Cas-based screening platform is developed to identify genetic modifiers of tau aggregation by using Cas-tau biosensor cells and conditioned medium to induce or sensitize tau aggregation, allowing for the detection of genetic modifiers through enrichment of specific guide RNAs in aggregation-positive cell populations.
The platform efficiently identifies genetic modifiers of tau aggregation, providing insights into disease pathogenesis and potential therapeutic targets by enriching guide RNAs that enhance or disrupt tau aggregation.
Smart Images

Figure 2026012398000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Application No. 62 / 820,086, filed March 18, 2019, which is incorporated herein by reference in its entirety for all purposes.
[0002] Reference to a sequence listing submitted as a text file via EFS WEB The sequence listing set forth in file 544635SEQLIST.txt is 75.7 kilobytes, was created on March 16, 2020, and is incorporated herein by reference. [Background technology]
[0003] Abnormal protein aggregation or fibrillation is a defining feature of many diseases, including several neurodegenerative disorders, particularly Alzheimer's disease (AD), Parkinson's disease (PD), frontotemporal dementia (FTD), amyotrophic lateral sclerosis (ALS), chronic traumatic encephalopathy (CTE), and Creutzfeldt-Jakob disease (CJD). In many of these diseases, fibrillation of certain proteins into insoluble aggregates is not only a hallmark of the disease but also implicated as a causative factor in neurotoxicity. Furthermore, these diseases are characterized by the propagation of aggregation pathology through the central nervous system following a stereotypical pattern, a process that correlates with disease progression. Therefore, identifying genes and genetic pathways that alter the process of abnormal protein aggregation or the intercellular propagation of aggregates would be of great value in better understanding the pathogenesis of neurodegenerative diseases as well as in devising strategies for therapeutic intervention. Summary of the Invention [Means for solving the problem]
[0004] Provided herein are methods for screening genetic modifiers of tau aggregation, methods for producing conditioned medium for inducing or sensitizing tau aggregation, and methods for producing tau aggregation-positive cell populations.Also provided herein are Cas-tau biosensor cells or populations of such cells, and in vitro cultures of Cas-tau biosensor cells and conditioned medium.Also provided herein are CRISPR / Cas synergistic activation mediator (SAM)-tau biosensor cells or populations of such cells, and in vitro cultures of SAM-tau biosensor cells and conditioned medium.
[0005] In one embodiment, a method for screening genetic modifiers of tau aggregation is provided. Some such methods (CRISPRn) include: (a) providing a cell population comprising a Cas protein, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter; (b) introducing a library comprising multiple unique guide RNAs targeting multiple genes into the cell population; (c) culturing the cell population to allow genome editing and expansion, wherein the multiple unique guide RNAs form complexes with the Cas protein, and the Cas protein cleaves the multiple genes to knock out gene function, generating a genetically modified cell population; (d) contacting the genetically modified cell population with a tau seeding agent to generate a seeded cell population; and (e) culturing the seeded cell population to allow tau aggregate formation, wherein the first tau repeat domain and the second tau repeat domain are linked to a second reporter. (f) determining the abundance of each of a plurality of unique guide RNAs in the aggregation-positive cell population identified in step (e) relative to the genetically modified cell population in step (c), wherein enrichment of the guide RNA in the aggregation-positive cell population identified in step (e) relative to the cultured cell population in step (c) indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, and disruption of the gene targeted by the guide RNA enhances tau aggregation or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and disruption of the gene targeted by the guide RNA is expected to enhance tau aggregation.
[0006] In some such methods, the first and / or second tau repeat domains are human tau repeat domains. In some such methods, the first and / or second tau repeat domains comprise a pro-aggregation mutation. Optionally, the first and / or second tau repeat domains comprise a tau P301S mutation.
[0007] In some such methods, the first tau repeat domain and / or the second tau repeat domain comprises a tau 4 repeat domain. In some such methods, the first tau repeat domain and / or the second tau repeat domain comprises SEQ ID NO: 11. In some such methods, the first tau repeat domain and the second tau repeat domain are the same. In some such methods, the first tau repeat domain and the second tau repeat domain are the same, each comprising a tau 4 repeat domain comprising the tau P301S mutation.
[0008] In some such methods, the first reporter and the second reporter are fluorescent proteins. Optionally, the first reporter and the second reporter are a fluorescence resonance energy transfer (FRET) pair. Optionally, the first reporter is cyan fluorescent protein (CFP), and the second reporter is yellow fluorescent protein (YFP).
[0009] In some such methods, the Cas protein is a Cas9 protein. Optionally, the Cas protein is Streptococcus pyogenes Cas9. Optionally, the Cas protein comprises SEQ ID NO: 21. Optionally, the Cas protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO: 22.
[0010] In some such methods, the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter are stably expressed in the cell population. In some such methods, nucleic acids encoding the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter are genomically integrated in the cell population.
[0011] In some such methods, the cell is a eukaryotic cell. Optionally, the cell is a mammalian cell. Optionally, the cell is a human cell. Optionally, the cell is a HEK293T cell.
[0012] In some such methods, the plurality of unique guide RNAs are introduced at a concentration selected so that the majority of cells receive only one of the unique guide RNAs. In some such methods, the plurality of unique guide RNAs target 100 or more genes, 1000 or more genes, or 10,000 or more genes. In some such methods, the library is a genome-wide library.
[0013] In some such methods, a plurality of target sequences are targeted, on average, in each of the plurality of targeted genes. Optionally, at least three target sequences are targeted, on average, in each of the plurality of targeted genes. Optionally, about three to about six target sequences (e.g., about three, about four, or about six) are targeted, on average, in each of the plurality of targeted genes.
[0014] In some such methods, each guide RNA targets a constitutive exon.Optionally, each guide RNA targets a 5' constitutive exon.In some such methods, each guide RNA targets the first exon, the second exon, or the third exon.
[0015] In some such methods, multiple unique guide RNAs are introduced into cell population by viral transduction.Optionally, each of multiple unique guide RNAs is in separate viral vector.Optionally, multiple unique guide RNAs are introduced into cell population by lentiviral transduction.In some such methods, cell population is infected with a multiplicity of infection of less than about 0.3.
[0016] In some such methods, the multiple unique guide RNAs are introduced into the cell population along with a selection marker, and step (b) further comprises selecting cells containing the selection marker. Optionally, the selection marker confers resistance to a drug. Optionally, the selection marker confers resistance to puromycin or zeocin. Optionally, the selection marker is selected from neomycin phosphotransferase, hygromycin B phosphotransferase, puromycin-N-acetyltransferase, and blasticidin S deaminase. Optionally, the selection marker is selected from neomycin phosphotransferase, hygromycin B phosphotransferase, puromycin-N-acetyltransferase, blasticidin S deaminase, and bleomycin resistance protein.
[0017] In some such methods, the cell population into which multiple unique guide RNAs are introduced in step (b) comprises more than about 300 cells per unique guide RNA.
[0018] In some such methods, step (c) is from about 3 days to about 9 days. Optionally, step (c) is about 6 days.
[0019] In some such methods, step (d) comprises culturing the genetically modified cell population in the presence of conditioned medium collected from cultured tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state. Optionally, the conditioned medium is collected after being placed on the confluent tau aggregation-positive cells for about 1 to about 7 days. Optionally, the conditioned medium is collected after being placed on the confluent tau aggregation-positive cells for about 4 days. Optionally, step (d) comprises culturing the genetically modified cell population in about 75% conditioned medium and about 25% fresh medium. In some such methods, the genetically modified cell population is not co-cultured with tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state.
[0020] In some such methods, step (e) is performed for about 2 days to about 6 days. Optionally, step (e) is performed for about 4 days. In some such methods, the first reporter and the second reporter are a fluorescence resonance energy transfer (FRET) pair, and the agglutination-positive cell population in step (e) is identified by flow cytometry.
[0021] In some such methods, the abundance is determined by next-generation sequencing. In some such methods, a guide RNA is considered enriched if the abundance of the guide RNA relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population in step (e) relative to the cultured cell population in step (c).
[0022] In some such methods, step (f) comprises determining the abundance of each of the plurality of unique guide RNAs in the aggregation-positive cell population in step (e) relative to the cultured cell population in step (c) at the first time point in step (c) and / or the second time point in step (c). Optionally, the first time point in step (c) is the first passage of culturing the cell population, and the second time point is midway through culturing the cell population to allow genome editing and expansion. Optionally, the first time point in step (c) is after about 3 days of culturing, and the second time point in step (c) is after about 6 days of culturing.
[0023] In some such methods, a gene is considered a genetic modifier of tau aggregation, and disruption of the gene enhances tau aggregation (or is a candidate genetic modifier of tau aggregation, and disruption of the gene is expected to enhance tau aggregation) if: (1) the abundance of guide RNAs targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population at step (e) relative to the cultured cell population at step (c) at both the first time point in step (c) and the second time point in step (c); and / or (2) the abundance of at least two unique guide RNAs targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population at step (e) relative to the cultured cell population at step (c) at either the first time point in step (c) or the second time point in step (c).
[0024] Some such methods include the steps of: (1) identifying which of a plurality of unique guide RNAs are present in the aggregation-positive cell population generated in step (e); (2) calculating the random probability that a guide RNA identified in step (f)(1) is present using the formula nCn'*(x-n')C(mn) / xCm, where x is the different unique guide RNAs introduced into the cell population in step (b), m is the different unique guide RNAs identified in step (f)(1), n is the different unique guide RNAs introduced into the cell population in step (b) that target a gene, and n' is the different unique guide RNAs identified in step (f)(1) that target a gene; and (3) calculating the average of the guide RNAs identified in step (f)(1). (3) calculating an enrichment score, where the enrichment score for the guide RNA is the relative abundance of the guide RNA in the aggregation-positive cell population generated in step (e) divided by the relative abundance of the guide RNA in the cultured cell population in step (c), where the relative abundance is the read count of the guide RNA divided by the read count of the total population of multiple unique guide RNAs; and (4) selecting genes where guide RNAs targeting the gene are significantly below random probability of being present and above an enrichment score threshold, are performed in step (f) to identify genes as genetic modifiers of tau aggregation, where disruption of the gene enhances tau aggregation (or is a candidate genetic modifier of tau aggregation, where disruption of the gene is expected to enhance tau aggregation).
[0025] Some such methods (CRISPRa) include: (a) providing a cell population comprising a chimeric Cas protein comprising a nuclease-inactive Cas protein fused to one or more transcription activation domains, a chimeric adaptor protein comprising an adaptor protein fused to one or more transcription activation domains, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter; (b) introducing into the cell population a library comprising a plurality of unique guide RNAs targeting a plurality of genes; (c) culturing the cell population to allow for transcription activation and expansion, wherein the plurality of unique guide RNAs form a complex with the chimeric Cas protein and the chimeric adaptor protein, and the complex activates transcription of the plurality of genes resulting in increased gene expression, generating a genetically modified cell population; and (d) contacting the genetically modified cell population with a tau seeding agent to generate a seeded cell population. (e) culturing the seeded cell population to allow for the formation of tau aggregates, wherein aggregates of the first tau repeat domain and the second tau repeat domain form in a subset of the seeded cell population to generate an aggregation-positive cell population; and (f) determining the abundance of each of a plurality of unique guide RNAs in the aggregation-positive cell population identified in step (e) relative to the genetically modified cell population in step (c), wherein enrichment of the guide RNA in the aggregation-positive cell population identified in step (e) relative to the cultured cell population in step (c) indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, and transcriptional activation of the gene targeted by the guide RNA enhances tau aggregation or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and transcriptional activation of the gene targeted by the guide RNA is expected to enhance tau aggregation.
[0026] In some such methods, the first and / or second tau repeat domains are human tau repeat domains. In some such methods, the first and / or second tau repeat domains comprise a pro-aggregation mutation. Optionally, the first and / or second tau repeat domains comprise a tau P301S mutation.
[0027] In some such methods, the first tau repeat domain and / or the second tau repeat domain comprises a tau 4 repeat domain. In some such methods, the first tau repeat domain and / or the second tau repeat domain comprises SEQ ID NO: 11. In some such methods, the first tau repeat domain and the second tau repeat domain are the same. In some such methods, the first tau repeat domain and the second tau repeat domain are the same, each comprising a tau 4 repeat domain comprising the tau P301S mutation.
[0028] In some such methods, the first reporter and the second reporter are fluorescent proteins. Optionally, the first reporter and the second reporter are a fluorescence resonance energy transfer (FRET) pair. Optionally, the first reporter is cyan fluorescent protein (CFP), and the second reporter is yellow fluorescent protein (YFP).
[0029] In some such methods, the Cas protein is a Cas9 protein. Optionally, the Cas protein is Streptococcus pyogenes Cas9. In some such methods, the chimeric Cas protein comprises a nuclease-inactive Cas protein fused to a VP64 transcription activation domain, and optionally, the chimeric Cas protein comprises, from N-terminus to C-terminus, the nuclease-inactive Cas protein, a nuclear localization signal, and a VP64 transcription activator domain. In some such methods, the adaptor protein is an MS2 coat protein, and the one or more transcription activation domains in the chimeric adaptor protein comprise a p65 transcription activation domain and an HSF1 transcription activation domain, and optionally, the chimeric adaptor protein comprises, from N-terminus to C-terminus, the MS2 coat protein, a nuclear localization signal, a p65 transcription activation domain, and an HSF1 transcription activation domain. In some such methods, the chimeric Cas protein comprises SEQ ID NO:36, and optionally, the chimeric Cas protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO:38. In some such methods, the chimeric adapter protein comprises SEQ ID NO:37, and optionally, the chimeric adapter protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO:39.
[0030] In some such methods, the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter are stably expressed in the cell population. In some such methods, the nucleic acid encoding the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter is genomically integrated in the cell population.
[0031] In some such methods, the cell is a eukaryotic cell. Optionally, the cell is a mammalian cell. Optionally, the cell is a human cell. Optionally, the cell is a HEK293T cell.
[0032] In some such methods, the plurality of unique guide RNAs are introduced at a concentration selected so that the majority of cells receive only one of the unique guide RNAs. In some such methods, the plurality of unique guide RNAs target 100 or more genes, 1000 or more genes, or 10,000 or more genes. In some such methods, the library is a genome-wide library.
[0033] In some such methods, a plurality of target sequences are targeted, on average, in each of the plurality of targeted genes. Optionally, at least three target sequences are targeted, on average, in each of the plurality of targeted genes. Optionally, about three to about six target sequences (e.g., about three, about four, or about six) are targeted, on average, in each of the plurality of targeted genes. Optionally, about three target sequences are targeted, on average, in each of the plurality of targeted genes.
[0034] In some such methods, each guide RNA targets a guide RNA target sequence within 200 bp upstream of the transcription start site. In some such methods, each guide RNA comprises one or more adapter binding elements to which a chimeric adapter protein can specifically bind. Optionally, each guide RNA comprises two adapter binding elements to which a chimeric adapter protein can specifically bind. Optionally, a first adapter binding element is within a first loop of each of the one or more guide RNAs, and a second adapter binding element is within a second loop of each of the one or more guide RNAs. Optionally, the adapter binding elements comprise the sequence set forth in SEQ ID NO: 33. In some such methods, each of the one or more guide RNAs is a single guide RNA comprising a CRISPR RNA (crRNA) portion fused to a trans-activating CRISPR RNA (tracrRNA) portion, wherein the first loop is a tetraloop corresponding to residues 13-16 of SEQ ID NO: 17, and the second loop is stem-loop 2 corresponding to residues 53-56 of SEQ ID NO: 17.
[0035] In some such methods, multiple unique guide RNAs are introduced into cell population by viral transduction.Optionally, each of multiple unique guide RNAs is in separate viral vector.Optionally, multiple unique guide RNAs are introduced into cell population by lentiviral transduction.In some such methods, cell population is infected with a multiplicity of infection of less than about 0.3.
[0036] In some such methods, the multiple unique guide RNAs are introduced into the cell population along with a selection marker, and step (b) further comprises selecting cells containing the selection marker. Optionally, the selection marker confers resistance to a drug. Optionally, the selection marker confers resistance to puromycin or zeocin. Optionally, the selection marker is selected from neomycin phosphotransferase, hygromycin B phosphotransferase, puromycin-N-acetyltransferase, and blasticidin S deaminase. Optionally, the selection marker is selected from neomycin phosphotransferase, hygromycin B phosphotransferase, puromycin-N-acetyltransferase, blasticidin S deaminase, and bleomycin resistance protein.
[0037] In some such methods, the cell population into which multiple unique guide RNAs are introduced in step (b) comprises more than about 300 cells per unique guide RNA.
[0038] In some such methods, step (c) is from about 3 days to about 9 days. Optionally, step (c) is about 6 days.
[0039] In some such methods, step (d) comprises culturing the genetically modified cell population in the presence of conditioned medium collected from cultured tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state. Optionally, the conditioned medium is collected after being placed on the confluent tau aggregation-positive cells for about 1 to about 7 days. Optionally, the conditioned medium is collected after being placed on the confluent tau aggregation-positive cells for about 4 days. Optionally, step (d) comprises culturing the genetically modified cell population in about 75% conditioned medium and about 25% fresh medium. In some such methods, the genetically modified cell population is not co-cultured with tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state.
[0040] In some such methods, step (e) is performed for about 2 days to about 6 days. Optionally, step (e) is performed for about 4 days. In some such methods, the first reporter and the second reporter are a fluorescence resonance energy transfer (FRET) pair, and the agglutination-positive cell population in step (e) is identified by flow cytometry.
[0041] In some such methods, the abundance is determined by next-generation sequencing. In some such methods, a guide RNA is considered enriched if the abundance of the guide RNA relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population in step (e) relative to the cultured cell population in step (c).
[0042] In some such methods, step (f) comprises determining the abundance of each of the plurality of unique guide RNAs in the aggregation-positive cell population in step (e) relative to the cultured cell population in step (c) at the first time point in step (c) and / or the second time point in step (c). Optionally, the first time point in step (c) is the first passage of culturing the cell population, and the second time point is midway through culturing the cell population to allow genome editing and expansion. Optionally, the first time point in step (c) is after about 3 days of culturing, and the second time point in step (c) is after about 6 days of culturing.
[0043] In some such methods, a gene is considered a genetic modifier of tau aggregation, and transcriptional activation of the gene enhances tau aggregation (or is a candidate genetic modifier of tau aggregation, and transcriptional activation of the gene is expected to enhance tau aggregation) if: (1) the abundance of a guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population at step (e) relative to the cultured cell population at step (c) at both the first time point in step (c) and the second time point in step (c); and / or (2) the abundance of at least two unique guide RNAs targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population at step (e) relative to the cultured cell population at step (c) at either the first time point in step (c) or the second time point in step (c).
[0044] Some such methods include the following steps: (1) identifying which of a plurality of unique guide RNAs are present in the aggregation-positive cell population generated in step (e); (2) calculating the random probability of a guide RNA identified in step (f)(1) being present using the formula nCn'*(x-n')C(mn) / xCm, where x are different unique guide RNAs introduced into the cell population in step (b), m are different unique guide RNAs identified in step (f)(1), n are different unique guide RNAs introduced into the cell population in step (b) that target a gene, and n' are different unique guide RNAs identified in step (f)(1) that target a gene; and (3) calculating the average enrichment score of the guide RNAs identified in step (f)(1). (3) calculating a score, where the enrichment score for the guide RNA is the relative abundance of the guide RNA in the aggregation-positive cell population generated in step (e) divided by the relative abundance of the guide RNA in the cultured cell population in step (c), where the relative abundance is the read count for the guide RNA divided by the read count for the total population of multiple unique guide RNAs; and (4) selecting genes where the guide RNAs targeting the gene are significantly below random probability of being present and above an enrichment score threshold, are performed in step (f) to identify genes as genetic modifiers of tau aggregation, where transcriptional activation of the gene enhances tau aggregation (or as candidate genetic modifiers of tau aggregation, where transcriptional activation of the gene is expected to enhance tau aggregation).
[0045] In another aspect, additional methods of screening for genetic modifiers of tau aggregation are provided. Some such methods (CRISPRn) include (a) providing a cell population comprising a Cas protein, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter; (b) introducing a library comprising a plurality of unique guide RNAs targeting a plurality of genes into the cell population; (c) culturing the cell population to allow genome editing and expansion, wherein the plurality of unique guide RNAs form a complex with the Cas protein, and the Cas protein cleaves the plurality of genes resulting in knockout of gene function, to generate a genetically modified cell population; (d) contacting the genetically modified cell population with a tau seeding agent to generate a seeded cell population; and (e) culturing the seeded cell population to allow formation of tau aggregates, wherein aggregates of the first tau repeat domain and the second tau repeat domain form in a first subset of the seeded cell population to generate an aggregate-positive cell population and a second subset of the seeded cell population to generate an aggregate-negative cell population. (f) determining the abundance of each of a plurality of unique guide RNAs in the aggregation-positive cell population identified in step (e) relative to the agglutination-negative cell population identified in step (e) and / or the plated cell population in step (d), and / or determining the abundance of each of a plurality of unique guide RNAs in the aggregation-positive cell population identified in step (e) relative to the agglutination-negative cell population identified in step (e) and / or the plated cell population in step (d), wherein enrichment of the guide RNA in the aggregation-negative cell population identified in step (e) relative to the agglutination-positive cell population identified in step (e) and / or the plated cell population in step (d), or depletion of the guide RNA in the aggregation-positive cell population identified in step (e) relative to the agglutination-negative cell population identified in step (e) and / or the plated cell population in step (d), indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation;Disruption of the gene targeted by the guide RNA prevents tau aggregation or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), where disruption of the gene targeted by the guide RNA is expected to prevent tau aggregation, and / or enrichment of guide RNA in the aggregation-positive cell population identified in step (e) relative to the aggregation-negative cell population identified in step (e) and / or the cell population seeded in step (d), or depletion of guide RNA in the aggregation-negative cell population identified in step (e) relative to the aggregation-positive cell population identified in step (e) and / or the cell population seeded in step (d), indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, where disruption of the gene targeted by the guide RNA promotes or enhances tau aggregation or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), where disruption of the gene targeted by the guide RNA is expected to promote or enhance tau aggregation.
[0046] In some such methods, the Cas protein is a Cas9 protein. Optionally, the Cas protein is Streptococcus pyogenes Cas9. In some such methods, the Cas protein comprises SEQ ID NO:21, and optionally, the Cas protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO:22.
[0047] In some such methods, the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter are stably expressed in the cell population. In some such methods, nucleic acids encoding the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter are genomically integrated in the cell population.
[0048] In some such methods, each guide RNA targets a constitutive exon.Optionally, each guide RNA targets a 5' constitutive exon.In some such methods, each guide RNA targets the first exon, the second exon, or the third exon.
[0049] Some such methods (CRISPRa) include: (a) providing a cell population comprising a chimeric Cas protein comprising a nuclease-inactive Cas protein fused to one or more transcription activation domains, a chimeric adaptor protein comprising an adaptor protein fused to one or more transcription activation domains, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter; (b) introducing into the cell population a library comprising a plurality of unique guide RNAs targeting a plurality of genes; (c) culturing the cell population to allow for transcription activation and expansion, wherein the plurality of unique guide RNAs form a complex with the chimeric Cas protein and the chimeric adaptor protein, and the Cas protein complex activates transcription of the plurality of genes resulting in increased gene expression, generating a genetically modified cell population; (d) contacting the genetically modified cell population with a tau seeding agent to generate a seeded cell population; and (e) culturing the seeded cell population to allow for the formation of tau aggregates. (f) culturing a first subset of the plated cell population such that aggregates of the first tau repeat domain and the second tau repeat domain form in a first subset of the plated cell population to generate an aggregation-positive cell population and do not form in a second subset of the plated cell population to generate an aggregation-negative cell population; and (f) determining the abundance of each of the plurality of unique guide RNAs in the aggregation-positive cell population identified in step (e) relative to the aggregation-negative cell population identified in step (e) and / or the plated cell population in step (d). and determining the abundance of each of a plurality of unique guide RNAs in the agglutination-negative cell population identified in step (e) relative to the agglutination-positive cell population identified in step (e) and / or the seeded cell population in step (d), wherein the enrichment of the guide RNAs in the agglutination-negative cell population identified in step (e) relative to the agglutination-positive cell population identified in step (e) and / or the seeded cell population in step (d) or the agglutination-negative cell population identified in step (e) and / or the seeded cell population in step (d)Depletion of the guide RNA in the aggregation-positive cell population identified in step (e) indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, transcriptional activation of the gene targeted by the guide RNA prevents tau aggregation, or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), transcriptional activation of the gene targeted by the guide RNA is expected to prevent tau aggregation, and / or is a candidate genetic modifier of tau aggregation, relative to the aggregation-negative cell population identified in step (e) and / or the seeded cell population in step (d), relative to the aggregation-positive cell population identified in step (e). Enrichment of guide RNA in, or depletion of guide RNA in, the aggregation-negative cell population identified in step (e) relative to the aggregation-positive cell population identified in step (e) and / or the seeded cell population in step (d) indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, transcriptional activation of the gene targeted by the guide RNA promotes or enhances tau aggregation, or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), where transcriptional activation of the gene targeted by the guide RNA is expected to promote or enhance tau aggregation.
[0050] In some such methods, the Cas protein is a Cas9 protein. Optionally, the Cas protein is Streptococcus pyogenes Cas9. In some such methods, the chimeric Cas protein comprises a nuclease-inactive Cas protein fused to a VP64 transcription activation domain, and optionally, the chimeric Cas protein comprises, from N-terminus to C-terminus, the nuclease-inactive Cas protein, a nuclear localization signal, and a VP64 transcription activator domain. In some such methods, the adaptor protein is an MS2 coat protein, and the one or more transcription activation domains in the chimeric adaptor protein comprise a p65 transcription activation domain and an HSF1 transcription activation domain, and optionally, the chimeric adaptor protein comprises, from N-terminus to C-terminus, the MS2 coat protein, a nuclear localization signal, a p65 transcription activation domain, and an HSF1 transcription activation domain. In some such methods, the chimeric Cas protein comprises SEQ ID NO:36, and optionally, the chimeric Cas protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO:38. In some such methods, the chimeric adapter protein comprises SEQ ID NO:37, and optionally, the chimeric adapter protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO:39.
[0051] In some such methods, the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter are stably expressed in the cell population. In some such methods, the nucleic acid encoding the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter is genomically integrated in the cell population.
[0052] In some such methods, each guide RNA targets a guide RNA target sequence within 200 bp upstream of the transcription start site. In some such methods, each guide RNA comprises one or more adapter binding elements to which a chimeric adapter protein can specifically bind. Optionally, each guide RNA comprises two adapter binding elements to which a chimeric adapter protein can specifically bind. Optionally, a first adapter binding element is within a first loop of each of the one or more guide RNAs, and a second adapter binding element is within a second loop of each of the one or more guide RNAs. Optionally, the adapter binding elements comprise the sequence set forth in SEQ ID NO: 33. Optionally, each of the one or more guide RNAs is a single guide RNA comprising a CRISPR RNA (crRNA) portion fused to a trans-activating CRISPR RNA (tracrRNA) portion, wherein the first loop is a tetraloop corresponding to residues 13-16 of SEQ ID NO: 17, and the second loop is stem-loop 2 corresponding to residues 53-56 of SEQ ID NO: 17.
[0053] In some such methods, step (c) is from about 3 days to about 13 days. In some such methods, step (c) is from about 7 days to about 10 days, about 7 days, or about 10 days.
[0054] In some such methods, step (d) comprises culturing the genetically modified cell population in the presence of a conditioned medium containing a cell lysate from cultured tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state. Optionally, the cell lysate in the medium is at a concentration of about 1 to about 5 μg / mL. In some such methods, the medium containing the cell lysate further comprises lipofectamine or another transfection reagent. Optionally, the medium containing the cell lysate contains lipofectamine at a concentration of about 1.5 to about 4 μL / mL. In some such methods, the genetically modified cell population is not co-cultured with tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state.
[0055] In some such methods, step (e) takes about 1 day to about 3 days. Optionally, step (e) takes about 2 days. In some such methods, the first reporter and the second reporter are a fluorescence resonance energy transfer (FRET) pair, and the agglutination-positive and agglutination-negative cell populations in step (e) are identified by flow cytometry. In some such methods, abundance is determined by next-generation sequencing.
[0056] In some such methods, a guide RNA is considered enriched in the aggregate-negative cell population at step (e) if the abundance of the guide RNA relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregate-negative cell population at step (e) relative to the aggregate-positive cell population at step (e) and / or the plated cell population at step (d), and a guide RNA is considered depleted in the aggregate-positive cell population at step (e) if the abundance of the guide RNA relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregate-positive cell population at step (e) relative to the aggregate-negative cell population at step (e) and / or the plated cell population at step (d). A guide RNA is considered enriched in the aggregation-positive cell population at step (e) if the abundance of the guide RNA relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population at step (e) relative to the aggregation-negative cell population at step (e) and / or the seeded cell population at step (d), and a guide RNA is considered depleted in the aggregation-negative cell population at step (e) if the abundance of the guide RNA relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregation-negative cell population at step (e) relative to the aggregation-positive cell population at step (e) and / or the seeded cell population at step (d).
[0057] In some such methods, step (f) comprises determining the abundance of each of the plurality of unique guide RNAs in the aggregation-negative cell population at step (e) relative to the aggregation-positive cell population at step (e), the cultured cell population at the first time point in step (c), and the seeded cell population at the second time point in step (d); and / or step (f) comprises determining the abundance of each of the plurality of unique guide RNAs in the aggregation-positive cell population at step (e) relative to the aggregation-negative cell population at step (e), the cultured cell population at the first time point in step (c), and the seeded cell population at the second time point in step (d). Optionally, the first time point in step (c) is the first passage of culturing the cell population. Optionally, the first time point in step (c) is after about 3 days of culturing, and the second time point in step (c) is after about 7 days of culturing or about 10 days of culturing.
[0058] In some such methods, a gene is considered to be a genetic modifier of tau aggregation, and disruption (CRISPRn) or transcriptional activation (CRISPRa) of the gene is determined if: (1) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-negative cell population at step (e) relative to the aggregation-positive cell population at step (e), the cultured cell population at the first time point of step (c), and the plated cell population at the second time point of step (d); and / or (2) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-negative cell population at step (e) relative to the aggregation-positive cell population at step (e) and the plated cell population at the second time point of step (d). and / or (3) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregation-positive cell population at step (e) relative to the aggregation-negative cell population at step (e), the cultured cell population at the first time point of step (c), and the plated cell population at the second time point of step (d); and / or (4) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregation-positive cell population at step (e) relative to the aggregation-negative cell population at step (e) and the plated cell population at the second time point of step (d).In some such methods, a gene is considered to be a genetic modifier of tau aggregation, and disruption (CRISPRn) or transcriptional activation (CRISPRa) of the gene is determined when: (1) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population at step (e) relative to the aggregation-negative cell population at step (e), the cultured cell population at the first time point of step (c), and the plated cell population at the second time point of step (d); and / or (2) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population at step (e) relative to the aggregation-negative cell population at step (e) and the plated cell population at the second time point of step (d); and / or or (3) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregation-negative cell population at step (e) relative to the aggregation-positive cell population at step (e), the cultured cell population at the first time point of step (c), and the plated cell population at the second time point of step (d); and / or (4) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregation-negative cell population at step (e) relative to the aggregation-positive cell population at step (e) and the plated cell population at the second time point of step (d).
[0059] Some such methods include the steps of: (1) identifying which of a plurality of unique guide RNAs are present in the agglutination-negative cell population generated in step (e); (2) calculating the random probability that a guide RNA identified in step (f)(1) is present using the formula nCn'*(x-n')C(mn) / xCm, where x are different unique guide RNAs introduced into the cell population in step (b), m are different unique guide RNAs identified in step (f)(1), n are different unique guide RNAs introduced into the cell population in step (b) that target a gene, and n' are different unique guide RNAs identified in step (f)(1) that target a gene; and (3) calculating the average enrichment score of the guide RNAs identified in step (f)(1), where the enrichment score of the guide RNA is higher than the enrichment score of the guide RNA identified in step (e). (4) calculating the relative abundance of the guide RNA in the generated aggregation-negative cell population divided by the relative abundance of the guide RNA in the aggregation-positive cell population generated in step (e) or the seeded cell population in step (d), where the relative abundance is the read count of the guide RNA divided by the read count of the total population of multiple unique guide RNAs; and (5) selecting genes where the guide RNAs targeting the gene are significantly below random probability of being present and above an enrichment score threshold, in step (f) to identify genes as genetic modifiers of tau aggregation, where disruption of the gene (CRISPRn) or transcriptional activation (CRISPRa) prevents tau aggregation (or as candidate genetic modifiers of tau aggregation, where disruption of the gene (CRISPRn) or transcriptional activation (CRISPRa) is expected to prevent tau aggregation).Some such methods include the steps of: (1) identifying which of a plurality of unique guide RNAs are present in the aggregation-positive cell population generated in step (e); (2) calculating the random probability that a guide RNA identified in step (f)(1) is present using the formula nCn'*(x-n')C(mn) / xCm, where x are different unique guide RNAs introduced into the cell population in step (b), m are different unique guide RNAs identified in step (f)(1), n are different unique guide RNAs introduced into the cell population in step (b) that target a gene, and n' are different unique guide RNAs identified in step (f)(1) that target a gene; and (3) calculating the average enrichment score of the guide RNAs identified in step (f)(1), where the enrichment score of the guide RNA is calculated based on the average enrichment score of the guide RNAs in step (f)(1). calculating the relative abundance of the guide RNA in the aggregation-positive cell population generated in step (e) divided by the relative abundance of the guide RNA in the aggregation-negative cell population generated in step (e) or the seeded cell population in step (d), where the relative abundance is the read count of the guide RNA divided by the read count of the total population of multiple unique guide RNAs; and (4) selecting genes where the guide RNAs targeting the gene are significantly below random probability of being present and above an enrichment score threshold, performed in step (f) to identify genes as genetic modifiers of tau aggregation, where gene disruption (CRISPRn) or transcriptional activation (CRISPRa) promotes or enhances tau aggregation (or as candidate genetic modifiers of tau aggregation, where gene disruption or transcriptional activation is expected to promote or enhance tau aggregation).
[0060] In some such methods, the first and / or second tau repeat domains are human tau repeat domains. In some such methods, the first and / or second tau repeat domains comprise a pro-aggregation mutation. Optionally, the first and / or second tau repeat domains comprise a tau P301S mutation.
[0061] In some such methods, the first tau repeat domain and / or the second tau repeat domain comprises a tau 4 repeat domain. In some such methods, the first tau repeat domain and / or the second tau repeat domain comprises SEQ ID NO: 11. In some such methods, the first tau repeat domain and the second tau repeat domain are the same. In some such methods, the first tau repeat domain and the second tau repeat domain are the same, each comprising a tau 4 repeat domain comprising the tau P301S mutation.
[0062] In some such methods, the first reporter and the second reporter are fluorescent proteins. Optionally, the first reporter and the second reporter are a fluorescence resonance energy transfer (FRET) pair. Optionally, the first reporter is cyan fluorescent protein (CFP), and the second reporter is yellow fluorescent protein (YFP).
[0063] In some such methods, the cell is a eukaryotic cell. Optionally, the cell is a mammalian cell. Optionally, the cell is a human cell. Optionally, the cell is a HEK293T cell.
[0064] In some such methods, the plurality of unique guide RNAs are introduced at a concentration selected so that the majority of cells receive only one of the unique guide RNAs. In some such methods, the plurality of unique guide RNAs target 100 or more genes, 1000 or more genes, or 10,000 or more genes. In some such methods, the library is a genome-wide library.
[0065] In some such methods, a plurality of target sequences are targeted, on average, in each of the plurality of targeted genes. Optionally, at least three target sequences are targeted, on average, in each of the plurality of targeted genes. Optionally, about three to about six target sequences (e.g., about three, about four, or about six) are targeted, on average, in each of the plurality of targeted genes. Optionally, about three target sequences are targeted, on average, in each of the plurality of targeted genes.
[0066] In some such methods, multiple unique guide RNAs are introduced into cell population by viral transduction.Optionally, each of multiple unique guide RNAs is in separate viral vector.Optionally, multiple unique guide RNAs are introduced into cell population by lentiviral transduction.In some such methods, cell population is infected with a multiplicity of infection of less than about 0.3.
[0067] In some such methods, the multiple unique guide RNAs are introduced into the cell population along with a selection marker, and step (b) further comprises selecting cells containing the selection marker. Optionally, the selection marker confers resistance to a drug. Optionally, the selection marker confers resistance to puromycin or zeocin. Optionally, the selection marker is selected from neomycin phosphotransferase, hygromycin B phosphotransferase, puromycin-N-acetyltransferase, and blasticidin S deaminase. Optionally, the selection marker is selected from neomycin phosphotransferase, hygromycin B phosphotransferase, puromycin-N-acetyltransferase, blasticidin S deaminase, and bleomycin resistance protein.
[0068] In some such methods, the cell population into which multiple unique guide RNAs are introduced in step (b) comprises more than about 300 cells per unique guide RNA.
[0069] In another embodiment, a method of screening for genetic modifiers of tau aggregation and / or disaggregation is provided. Some such methods (CRISPRn) include: (a) providing a cell population comprising a Cas protein, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter, wherein the cells are tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state; (b) introducing into the cell population a library comprising a plurality of unique guide RNAs targeting a plurality of genes; (c) culturing the cell population to allow genome editing and expansion, wherein the plurality of unique guide RNAs form a complex with the Cas protein, and the Cas protein cleaves the plurality of genes resulting in knockout of gene function, generating a genetically modified cell population; and culturing, wherein the culturing results in an aggregation-positive cell population and an aggregation-negative cell population; (d) identifying the aggregation-positive cell population and the aggregation-negative cell population; and (e) performing a step on the aggregation-negative cell population identified in step (d) and / or the cultured cell population at one or more time points of step (c). determining the abundance of each of a plurality of unique guide RNAs in the aggregation-positive cell population identified in step (d), and / or determining the abundance of each of a plurality of unique guide RNAs in the aggregation-negative cell population identified in step (d) relative to the aggregation-positive cell population identified in step (d) and / or the cultured cell population at one or more time points of step (c), wherein enrichment of guide RNA in the aggregation-negative cell population identified in step (d) relative to the aggregation-positive cell population identified in step (d) and / or the cultured cell population at one or more time points of step (c), or depletion of guide RNA in the aggregation-positive cell population identified in step (d) relative to the aggregation-negative cell population identified in step (d) and / or the cultured cell population at one or more time points of step (c), indicates that the gene targeted by the guide RNA is a genetic modifier of tau disaggregation, and disruption of the gene targeted by the guide RNA promotes tau disaggregation or is a candidate genetic modifier of tau disaggregation (e.g.,(e.g., for further testing via secondary screening), disruption of the gene targeted by the guide RNA is expected to promote tau disaggregation, and / or enrichment of the guide RNA in the aggregation-positive cell population identified in step (d) relative to the aggregation-negative cell population identified in step (d) and / or the cultured cell population at one or more time points of step (c), or depletion of the guide RNA in the aggregation-negative cell population identified in step (d) relative to the aggregation-positive cell population identified in step (d) and / or the cultured cell population at one or more time points of step (c), indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, and disruption of the gene targeted by the guide RNA promotes or enhances tau aggregation or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), disruption of the gene targeted by the guide RNA is expected to promote or enhance tau aggregation.
[0070] In some such methods, the Cas protein is a Cas9 protein. Optionally, the Cas protein is Streptococcus pyogenes Cas9. In some such methods, the Cas protein comprises SEQ ID NO:21, and optionally, the Cas protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO:22.
[0071] In some such methods, the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter are stably expressed in the cell population. In some such methods, nucleic acids encoding the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter are genomically integrated in the cell population.
[0072] In some such methods, each guide RNA targets a constitutive exon.Optionally, each guide RNA targets a 5' constitutive exon.In some such methods, each guide RNA targets the first exon, the second exon, or the third exon.
[0073] Some such methods (CRISPRa) include: (a) providing a cell population comprising a chimeric Cas protein comprising a nuclease-inactive Cas protein fused to one or more transcription activation domains, a chimeric adaptor protein comprising an adaptor protein fused to one or more transcription activation domains, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter, wherein the cells are tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state; (b) introducing into the cell population a library comprising a plurality of unique guide RNAs targeting a plurality of genes; (c) culturing the cell population to allow transcription activation and expansion, wherein the plurality of unique guide RNAs form a complex with the chimeric Cas protein and the chimeric adaptor protein, and the Cas protein complex activates transcription of the plurality of genes resulting in increased gene expression, generating a genetically modified cell population; and (d) culturing resulting in an aggregation-positive cell population and an aggregation-negative cell population. (e) determining the abundance of each of the plurality of unique guide RNAs in the aggregation-positive cell population identified in step (d) relative to the aggregation-negative cell population identified in step (d) and / or the cultured cell population at one or more time points of step (c); and / or determining the abundance of each of the plurality of unique guide RNAs in the aggregation-negative cell population identified in step (d) relative to the aggregation-positive cell population identified in step (d) and / or the cultured cell population at one or more time points of step (c); wherein enrichment of guide RNA in the aggregation-negative cell population identified in step (d) relative to the aggregation-positive cell population identified in step (d) and / or the cultured cell population at one or more time points of step (c), or depletion of guide RNA in the aggregation-positive cell population identified in step (d) relative to the aggregation-negative cell population identified in step (d) and / or the cultured cell population at one or more time points of step (c), indicates that the gene targeted by the guide RNA is a genetic modifier of tau disaggregation;Transcriptional activation of the gene targeted by the guide RNA is expected to promote tau disaggregation or is a candidate genetic modifier of tau disaggregation (e.g., for further testing via secondary screening), and / or enrichment of the guide RNA in the aggregation-positive cell population identified in step (d), or the aggregation-negative cell population identified in step (d), relative to the aggregation-negative cell population identified in step (d) and / or the cultured cell population at one or more time points of step (c). and / or for the cultured cell population at one or more time points of step (c), depletion of the guide RNA in the aggregation-negative cell population identified in step (d) indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, transcriptional activation of the gene targeted by the guide RNA promotes or enhances tau aggregation, or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and transcriptional activation of the gene targeted by the guide RNA is expected to promote or enhance tau aggregation.
[0074] In some such methods, the Cas protein is a Cas9 protein. Optionally, the Cas protein is Streptococcus pyogenes Cas9. In some such methods, the chimeric Cas protein comprises a nuclease-inactive Cas protein fused to a VP64 transcription activation domain, and optionally, the chimeric Cas protein comprises, from N-terminus to C-terminus, the nuclease-inactive Cas protein, a nuclear localization signal, and a VP64 transcription activator domain. In some such methods, the adaptor protein is an MS2 coat protein, and the one or more transcription activation domains in the chimeric adaptor protein comprise a p65 transcription activation domain and an HSF1 transcription activation domain, and optionally, the chimeric adaptor protein comprises, from N-terminus to C-terminus, the MS2 coat protein, a nuclear localization signal, a p65 transcription activation domain, and an HSF1 transcription activation domain. In some such methods, the chimeric Cas protein comprises SEQ ID NO:36, and optionally, the chimeric Cas protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO:38. In some such methods, the chimeric adapter protein comprises SEQ ID NO:37, and optionally, the chimeric adapter protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO:39.
[0075] In some such methods, the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter are stably expressed in the cell population. In some such methods, the nucleic acid encoding the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter is genomically integrated in the cell population.
[0076] In some such methods, each guide RNA targets a guide RNA target sequence within 200 bp upstream of the transcription start site. In some such methods, each guide RNA comprises one or more adapter binding elements to which a chimeric adapter protein can specifically bind. Optionally, each guide RNA comprises two adapter binding elements to which a chimeric adapter protein can specifically bind. Optionally, a first adapter binding element is within a first loop of each of the one or more guide RNAs, and a second adapter binding element is within a second loop of each of the one or more guide RNAs. Optionally, the adapter binding elements comprise the sequence set forth in SEQ ID NO: 33. Optionally, each of the one or more guide RNAs is a single guide RNA comprising a CRISPR RNA (crRNA) portion fused to a trans-activating CRISPR RNA (tracrRNA) portion, wherein the first loop is a tetraloop corresponding to residues 13-16 of SEQ ID NO: 17, and the second loop is stem-loop 2 corresponding to residues 53-56 of SEQ ID NO: 17.
[0077] In some such methods, step (c) is from about 3 days to about 14 days. Optionally, step (c) is from about 10 days to about 14 days or from about 12 days to about 14 days.
[0078] In some such methods, step (d) comprises synchronizing cell cycle progression to obtain a cell population primarily enriched in S phase. Optionally, synchronization is achieved by double thymidine block.
[0079] In some such methods, the first reporter and the second reporter are a fluorescence resonance energy transfer (FRET) pair, and the agglutination-positive and agglutination-negative cell populations in step (d) are identified by flow cytometry. In some such methods, abundance is determined by next-generation sequencing.
[0080] In some such methods, a guide RNA is considered enriched in the aggregation-negative cell population at step (d) if the abundance of the guide RNA relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-negative cell population at step (d) relative to the aggregation-positive cell population at step (d) and / or the cultured cell population at one or more time points of step (c), and a guide RNA is considered depleted in the aggregation-positive cell population at step (d) if the abundance of the guide RNA relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregation-positive cell population at step (d) relative to the aggregation-negative cell population at step (d) and / or the cultured cell population at one or more time points of step (c). A guide RNA is considered enriched in the aggregation-positive cell population at step (d) if the abundance of the guide RNA relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population at step (d) relative to the aggregation-negative cell population at step (d) and / or the cultured cell population at one or more time points of step (c), and a guide RNA is considered depleted in the aggregation-negative cell population at step (d) if the abundance of the guide RNA relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregation-negative cell population at step (d) relative to the aggregation-positive cell population at step (d) and / or the cultured cell population at one or more time points of step (c).
[0081] In some such methods, step (e) comprises determining the abundance of each of the plurality of unique guide RNAs in the aggregation-negative cell population at step (d) relative to the aggregation-positive cell population at step (d), the cultured cell population at the first time point in step (c), and the cultured cell population at the second time point in step (c); and / or step (e) comprises determining the abundance of each of the plurality of unique guide RNAs in the aggregation-positive cell population at step (d) relative to the aggregation-negative cell population at step (e), the cultured cell population at the first time point in step (c), and the cultured cell population at the second time point in step (c). Optionally, the first time point in step (c) is the first passage of culturing the cell population, and the second time point is midway through culturing the cell population to allow genome editing and expansion or transcriptional activation and expansion. Optionally, the first time point in step (c) is after about 7 days of culturing, and the second time point in step (c) is after about 10 days of culturing.
[0082] In some such methods, a gene is considered to be a genetic modifier of tau disaggregation, and disruption (CRISPRn) or transcriptional activation (CRISPRa) of the gene is determined if (1) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-negative cell population at step (d) relative to the aggregation-positive cell population at step (d), the cultured cell population at the first time point of step (c), and the cultured cell population at the second time point of step (c); and / or (2) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-negative cell population at step (d) relative to the aggregation-positive cell population at step (d) and the cultured cell population at the second time point of step (c). and / or (3) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregation-positive cell population at step (d) relative to the aggregation-negative cell population at step (d), the cultured cell population at the first time point of step (c), and the cultured cell population at the second time point of step (c); and / or (4) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregation-positive cell population at step (d) relative to the aggregation-negative cell population at step (d) and the cultured cell population at the second time point of step (c).In some such methods, a gene is considered to be a genetic modifier of tau aggregation, and disruption (CRISPRn) or transcriptional activation (CRISPRa) of the gene is determined when: (1) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population at step (d) relative to the aggregation-negative cell population at step (d), the cultured cell population at the first time point of step (c), and the cultured cell population at the second time point of step (c); and / or (2) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher in the aggregation-positive cell population at step (d) relative to the aggregation-negative cell population at step (d) and the cultured cell population at the second time point of step (c); and / or or (3) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregation-negative cell population at step (d) relative to the aggregation-positive cell population at step (d), the cultured cell population at the first time point of step (c), and the cultured cell population at the second time point of step (c); and / or (4) the abundance of the guide RNA targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold lower in the aggregation-negative cell population at step (d) relative to the aggregation-positive cell population at step (d) and the cultured cell population at the second time point of step (c).
[0083] Some such methods include the steps of: (1) identifying which of a plurality of unique guide RNAs are present in the agglutination-negative cell population identified in step (d); (2) calculating the random probability that a guide RNA identified in step (e)(1) is present using the formula nCn'*(x-n')C(mn) / xCm, where x is the different unique guide RNAs introduced into the cell population in step (b), m is the different unique guide RNAs identified in step (e)(1), n is the different unique guide RNAs introduced into the cell population in step (b) that target a gene, and n' is the different unique guide RNAs identified in step (e)(1) that target a gene; and (3) calculating the average enrichment score of the guide RNAs identified in step (e)(1), where the enrichment score of the guide RNA is greater than or equal to the average enrichment score of the guide RNAs identified in step (d). (4) calculating the relative abundance of the guide RNA in the aggregation-negative cell population identified in step (c) divided by the relative abundance of the guide RNA in the aggregation-positive cell population identified in step (d) or the cultured cell population at the first or second time point in step (c), where the relative abundance is the read count of the guide RNA divided by the read count of the total population of multiple unique guide RNAs; and (5) selecting genes where the guide RNAs targeting the gene are significantly below random probability of being present and above an enrichment score threshold, performed in step (e) to identify genes as genetic modifiers of tau disaggregation, where disruption (CRISPRn) or transcriptional activation (CRISPRa) of the gene promotes tau disaggregation (or are candidate genetic modifiers of tau disaggregation, where disruption or transcriptional activation of the gene is expected to promote tau disaggregation).Some such methods include the following steps: (1) identifying which of a plurality of unique guide RNAs are present in the aggregation-positive cell population identified in step (d); (2) calculating the random probability that a guide RNA identified in step (e)(1) is present using the formula nCn'*(x-n')C(mn) / xCm, where x is the different unique guide RNAs introduced into the cell population in step (b), m is the different unique guide RNAs identified in step (e)(1), n is the different unique guide RNAs introduced into the cell population in step (b) that target a gene, and n' is the different unique guide RNAs identified in step (e)(1) that target a gene; and (3) calculating the average enrichment score of the guide RNAs identified in step (e)(1), where the enrichment scores of the guide RNAs are the same as those of the guide RNAs identified in step (d). (3) calculating the relative abundance of the guide RNA in the aggregation-positive cell population identified in step (d) divided by the relative abundance of the guide RNA in the aggregation-negative cell population identified in step (c) or the cultured cell population at the first or second time point, where the relative abundance is the read count of the guide RNA divided by the read count of the total population of multiple unique guide RNAs; and (4) selecting genes where guide RNAs targeting the gene are significantly below random probability of being present and above an enrichment score threshold, performed in step (e) to identify genes as genetic modifiers of tau aggregation, where disruption (CRISPRn) or transcriptional activation (CRISPRa) of the gene promotes or enhances tau aggregation (or are candidate genetic modifiers of tau aggregation, where disruption or transcriptional activation of the gene is expected to promote or enhance tau aggregation).
[0084] In some such methods, the first and / or second tau repeat domains are human tau repeat domains. In some such methods, the first and / or second tau repeat domains comprise a pro-aggregation mutation. Optionally, the first and / or second tau repeat domains comprise a tau P301S mutation.
[0085] In some such methods, the first tau repeat domain and / or the second tau repeat domain comprises a tau 4 repeat domain. In some such methods, the first tau repeat domain and / or the second tau repeat domain comprises SEQ ID NO: 11. In some such methods, the first tau repeat domain and the second tau repeat domain are the same. In some such methods, the first tau repeat domain and the second tau repeat domain are the same, each comprising a tau 4 repeat domain comprising the tau P301S mutation.
[0086] In some such methods, the first reporter and the second reporter are fluorescent proteins. Optionally, the first reporter and the second reporter are a fluorescence resonance energy transfer (FRET) pair. Optionally, the first reporter is cyan fluorescent protein (CFP), and the second reporter is yellow fluorescent protein (YFP).
[0087] In some such methods, the cell is a eukaryotic cell. Optionally, the cell is a mammalian cell. Optionally, the cell is a human cell. Optionally, the cell is a HEK293T cell.
[0088] In some such methods, the plurality of unique guide RNAs are introduced at a concentration selected so that the majority of cells receive only one of the unique guide RNAs. In some such methods, the plurality of unique guide RNAs target 100 or more genes, 1000 or more genes, or 10,000 or more genes. In some such methods, the library is a genome-wide library.
[0089] In some such methods, a plurality of target sequences are targeted, on average, in each of the plurality of targeted genes. Optionally, at least three target sequences are targeted, on average, in each of the plurality of targeted genes. Optionally, about three to about six target sequences (e.g., about three, about four, or about six) are targeted, on average, in each of the plurality of targeted genes. Optionally, about three target sequences are targeted, on average, in each of the plurality of targeted genes.
[0090] In some such methods, multiple unique guide RNAs are introduced into cell population by viral transduction.Optionally, each of multiple unique guide RNAs is in separate viral vector.Optionally, multiple unique guide RNAs are introduced into cell population by lentiviral transduction.In some such methods, cell population is infected with a multiplicity of infection of less than about 0.3.
[0091] In some such methods, the multiple unique guide RNAs are introduced into the cell population along with a selection marker, and step (b) further comprises selecting cells containing the selection marker. Optionally, the selection marker confers resistance to a drug. Optionally, the selection marker confers resistance to puromycin or zeocin. Optionally, the selection marker is selected from neomycin phosphotransferase, hygromycin B phosphotransferase, puromycin-N-acetyltransferase, and blasticidin S deaminase. Optionally, the selection marker is selected from neomycin phosphotransferase, hygromycin B phosphotransferase, puromycin-N-acetyltransferase, blasticidin S deaminase, and bleomycin resistance protein.
[0092] In some such methods, the cell population into which multiple unique guide RNAs are introduced in step (b) comprises more than about 300 cells per unique guide RNA.
[0093] In another embodiment, a Cas-tau biosensor cell or population of such cells is provided, wherein some such cells comprise a population of one or more cells comprising a Cas protein, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter.
[0094] In some such cells, the first tau repeat domain and / or the second tau repeat domain is a human tau repeat domain. In some such cells, the first tau repeat domain and / or the second tau repeat domain comprises an aggregation-promoting mutation. Optionally, the first tau repeat domain and / or the second tau repeat domain comprises a tau P301S mutation.
[0095] In some such cells, the first tau repeat domain and / or the second tau repeat domain comprises a tau 4 repeat domain. In some such cells, the first tau repeat domain and / or the second tau repeat domain comprises SEQ ID NO: 11. In some such cells, the first tau repeat domain and the second tau repeat domain are the same. In some such cells, the first tau repeat domain and the second tau repeat domain are the same, each comprising a tau 4 repeat domain comprising the tau P301S mutation.
[0096] In some such cells, the first reporter and the second reporter are fluorescent proteins.Optionally, the first reporter and the second reporter are a fluorescence resonance energy transfer (FRET) pair.Optionally, the first reporter is cyan fluorescent protein (CFP), and the second reporter is yellow fluorescent protein (YFP).
[0097] In some such cells, the Cas protein is a Cas9 protein. Optionally, the Cas protein is Streptococcus pyogenes Cas9. Optionally, the Cas protein comprises SEQ ID NO: 21. Optionally, the Cas protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO: 22.
[0098] In some such cells, the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter are stably expressed in the cell. In some such cells, the nucleic acid encoding the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter is genomically integrated in the cell.
[0099] In some such cells, the cell is a eukaryotic cell. Optionally, the cell is a mammalian cell. Optionally, the cell is a human cell. Optionally, the cell is a HEK293T cell. In such cells, the cell is in vitro.
[0100] In some such cells, the first tau repeat domain linked to the first reporter and the second tau repeat domain linked to the second reporter are not stably present in an aggregated state, and in some such cells, the first tau repeat domain linked to the first reporter and the second tau repeat domain linked to the second reporter are stably present in an aggregated state.
[0101] In another aspect, an in vitro culture of Cas-tau biosensor cells and conditioned medium is provided. Some such in vitro cultures comprise any of the cell populations described above or elsewhere herein and a culture medium comprising conditioned medium harvested from cultured tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state.
[0102] In some such in vitro cultures, the conditioned medium was harvested after about 1 to about 7 days of incubation on confluent tau aggregate-positive cells. Optionally, the conditioned medium was harvested after about 4 days of incubation on confluent tau aggregate-positive cells.
[0103] In some such in vitro cultures, the culture medium comprises about 75% conditioned medium and about 25% fresh medium. In some such in vitro cultures, the cell population is not co-cultured with cultured tau aggregate-positive cells in which the tau repeat domains are stably present in an aggregated state.
[0104] In another embodiment, a SAM-tau biosensor cell or population of such cells is provided, wherein some such cells comprise a population of one or more cells comprising a chimeric Cas protein comprising a nuclease-inactive Cas protein fused to one or more transcription activation domains, a chimeric adaptor protein comprising an adaptor protein fused to one or more transcription activation domains, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter.
[0105] In some such cells, the first tau repeat domain and / or the second tau repeat domain is a human tau repeat domain. In some such cells, the first tau repeat domain and / or the second tau repeat domain comprises an aggregation-promoting mutation. Optionally, the first tau repeat domain and / or the second tau repeat domain comprises a tau P301S mutation.
[0106] In some such cells, the first tau repeat domain and / or the second tau repeat domain comprises a tau 4 repeat domain. In some such cells, the first tau repeat domain and / or the second tau repeat domain comprises SEQ ID NO: 11. In some such cells, the first tau repeat domain and the second tau repeat domain are the same. In some such cells, the first tau repeat domain and the second tau repeat domain are the same, each comprising a tau 4 repeat domain comprising the tau P301S mutation.
[0107] In some such cells, the first reporter and the second reporter are fluorescent proteins.Optionally, the first reporter and the second reporter are a fluorescence resonance energy transfer (FRET) pair.Optionally, the first reporter is cyan fluorescent protein (CFP), and the second reporter is yellow fluorescent protein (YFP).
[0108] In some such cells, the Cas protein is a Cas9 protein. Optionally, the Cas protein is Streptococcus pyogenes Cas9. In some such cells, the chimeric Cas protein comprises a nuclease-inactive Cas protein fused to a VP64 transcription activation domain, and optionally, the chimeric Cas protein comprises, from N-terminus to C-terminus, the nuclease-inactive Cas protein, a nuclear localization signal, and a VP64 transcription activator domain. In some such cells, the adaptor protein is an MS2 coat protein, and the one or more transcription activation domains in the chimeric adaptor protein comprise a p65 transcription activation domain and an HSF1 transcription activation domain, and optionally, the chimeric adaptor protein comprises, from N-terminus to C-terminus, the MS2 coat protein, a nuclear localization signal, a p65 transcription activation domain, and an HSF1 transcription activation domain. In some such cells, the chimeric Cas protein comprises SEQ ID NO:36, and optionally, the chimeric Cas protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO:38. In some such cells, the chimeric adapter protein comprises SEQ ID NO:37, and optionally, the chimeric adapter protein is encoded by a coding sequence comprising the sequence set forth in SEQ ID NO:39.
[0109] In some such cells, the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter are stably expressed in the cell. In some such cells, the nucleic acid encoding the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter is genomically integrated in the cell.
[0110] In some such cells, the cell is a eukaryotic cell. Optionally, the cell is a mammalian cell. Optionally, the cell is a human cell. Optionally, the cell is a HEK293T cell. In such cells, the cell is in vitro.
[0111] In some such cells, the first tau repeat domain linked to the first reporter and the second tau repeat domain linked to the second reporter are not stably present in an aggregated state, and in some such cells, the first tau repeat domain linked to the first reporter and the second tau repeat domain linked to the second reporter are stably present in an aggregated state.
[0112] In another aspect, an in vitro culture of SAM-tau biosensor cells and conditioned medium is provided. Some such in vitro cultures comprise any of the cell populations described above or elsewhere herein and a culture medium comprising conditioned medium harvested from cultured tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state.
[0113] In some such in vitro cultures, the conditioned medium was harvested after about 1 to about 7 days of incubation on confluent tau aggregate-positive cells. Optionally, the conditioned medium was harvested after about 4 days of incubation on confluent tau aggregate-positive cells.
[0114] In some such in vitro cultures, the culture medium comprises about 75% conditioned medium and about 25% fresh medium. In some such in vitro cultures, the cell population is not co-cultured with cultured tau aggregate-positive cells in which the tau repeat domains are stably present in an aggregated state.
[0115] In another aspect, an in vitro culture is provided of Cas-tau biosensor cells or a population of such cells or SAM-tau biosensor cells, and a culture medium comprising cell lysate from a cultured tau aggregation-positive cell population in which the tau repeat domains are stably present in an aggregated state. Some such in vitro cultures comprise any of the cell populations described above or elsewhere herein.
[0116] In some such in vitro cultures, the cell lysate in the medium is at a concentration of about 1 to about 5 μg / mL. In some such in vitro cultures, the medium containing the cell lysate further contains lipofectamine or another transfection reagent. Optionally, the medium containing the cell lysate contains lipofectamine at a concentration of about 1.5 to about 4 μL / mL. In some such in vitro cultures, the cell lysate is generated by sonicating tau aggregate-positive cells for about 2 to about 4 minutes after harvesting the cells in a buffer containing protease inhibitors.
[0117] In another aspect, methods for producing a conditioned medium for inducing or sensitizing tau aggregation are provided, some of which include (a) providing a tau aggregation-positive cell population in which tau repeat domains are stably present in an aggregated state, (b) culturing the tau aggregation-positive cell population in a medium to produce a conditioned medium, and (c) harvesting the conditioned medium.
[0118] In some such methods, the tau aggregate-positive cells are cultured until confluent in step (b). Optionally, the conditioned medium is harvested in step (c) after being placed on the confluent tau aggregate-positive cells for about 1 to about 7 days. Optionally, the conditioned medium is harvested in step (c) after being placed on the confluent tau aggregate-positive cells for about 4 days.
[0119] In another aspect, methods for generating a tau aggregation-positive cell population are provided. Some such methods include (a) generating a conditioned medium to induce tau aggregation according to any of the methods described above or elsewhere herein, and (b) culturing a cell population comprising a protein comprising a tau repeat domain in a culture medium containing the conditioned medium to generate a tau aggregation-positive cell population.
[0120] In some such methods, the culture medium comprises about 75% conditioned medium and about 25% fresh medium. In some such methods, the cell population is not co-cultured with the tau aggregate-positive cells used in the method to generate the conditioned medium.
[0121] In some such methods, the tau repeat domain comprises an aggregation-promoting mutation. In some such methods, the tau repeat domain comprises a tau P301S mutation. In some such methods, the tau repeat domain comprises a tau 4 repeat domain. In some such methods, the tau repeat domain comprises SEQ ID NO: 11.
[0122] In another aspect, a method for producing a medium containing a cell lysate from cultured tau aggregation-positive cells for inducing tau aggregation is provided. Some such methods include (a) providing a population of tau aggregation-positive cells in which tau repeat domains are stably present in an aggregated state, (b) collecting the tau aggregation-positive cells in a buffer containing a protease inhibitor, (c) sonicating the tau aggregation-positive cells for about 2 to about 4 minutes to produce a cell lysate, and (d) adding the cell lysate to a growth medium.
[0123] In some such methods, the cell lysate in the growth medium is at a concentration of about 1 to about 5 μg / mL. Some such methods further comprise, in step (d), adding lipofectamine or another transfection reagent to the growth medium. Optionally, step (d) comprises adding lipofectamine at a concentration of about 1.5 to about 4 μL / mL.
[0124] In another aspect, methods for generating a tau aggregate-positive cell population are provided. Some such methods include (a) generating a medium containing a cell lysate from cultured tau aggregate-positive cells according to any of the methods described above, and (b) culturing a cell population containing a protein comprising a tau repeat domain in the medium containing the cell lysate from the cultured tau aggregate-positive cells.
[0125] In some such methods, the cell population is not co-cultured with the tau aggregate-positive cells used in the method to generate the conditioned medium.
[0126] In some such methods, the tau repeat domain comprises an aggregation-promoting mutation. In some such methods, the tau repeat domain comprises a tau P301S mutation. In some such methods, the tau repeat domain comprises a tau 4 repeat domain. In some such methods, the tau repeat domain comprises SEQ ID NO: 11. [Brief explanation of the drawings]
[0127] [Figure 1] A schematic diagram of the tau isoform 2N4R (not to scale) is shown. The tau biosensor line contains only tau4RD-YFP and tau4RD-CFP as transgenes, rather than the full 2N4R. [Figure 2] A schematic diagram of how aggregate formation is monitored by fluorescence resonance energy transfer (FRET) in a tau biosensor cell line is shown. Tau4RD-CFP protein is excited with violet light and emits blue light. Tau4RD-YFP fusion protein is excited with blue light and emits yellow light. In the absence of aggregation, excitation with violet light does not result in FRET. In the presence of tau aggregation, excitation with violet light results in FRET and yellow emission. [Figure 3A]Figure 1 shows relative Cas9 mRNA expression in Tau4RD-CFP / Tau4RD-YFP (TCY) biosensor cell clones transduced with a lentiviral Cas9 expression construct relative to clone Cas9H1, a previously isolated suboptimal control Cas9-expressing TCY clone. [Figure 3B] Figure 1 shows the cleavage efficiency at the PERK and SNCA loci in Cas9 TCY clones 3 and 7 days after transduction with sgRNAs targeting PERK and SNCA, respectively. [Figure 4] Figure 1 shows a schematic of the strategy for disrupting target genes in Cas9 TCY biosensor cells using a genome-wide CRISPR / Cas9 sgRNA library. [Figure 5] Figure 1 shows a schematic diagram illustrating the induction of stably growing Tau4RD-YFP Agg[+] subclones containing tau aggregates when Tau4RD-YFP cells are seeded with Tau4RD fibrils. Fluorescence microscopy images showing the subclones containing tau aggregates are also shown. [Figure 6] This is a schematic diagram showing that conditioned medium from Tau4RD-YFP Agg[+] subclones, collected after 3 days in confluent cells, can provide a source of tau aggregation activity, whereas medium from Tau4RD-YFP Agg[-] subclones does not. Conditioned medium was applied to recipient cells as 75% conditioned medium and 25% fresh medium. Respective fluorescence-activated cell sorting (FACS) analysis images are shown. The x-axis represents CFP (405 nm laser excitation) and the y-axis represents FRET (excitation from CFP emission). The upper right quadrant is FRET[+], the lower right quadrant is CFP[+], and the lower left quadrant is double negative. [Figure 7] FIG. 1 is a schematic diagram showing the strategy of genome-wide CRISPR nuclease (CRISPRn) screening to identify modifier genes that promote tau aggregation. [Figure 8] FIG. 1 is a schematic illustrating the concept of abundance and enrichment for next-generation sequencing (NGS) analysis using genome-wide CRISPRn screening. [Figure 9] Figure 1 shows a schematic diagram of the secondary screening of target genes 1 to 14 identified in the genome-wide screening of modifier genes that promote tau aggregation. [Figure 10] Figure 1 shows FRET induction by tau aggregate-conditioned medium in Cas9 TCY biosensor cells transduced with lentiviral expression constructs of sgRNAs targeting target genes 1 to 14. A secondary screen confirmed that target genes 2 and 8 modulate cellular susceptibility to tau seeding / aggregation. [Figure 11] FACS analysis images of Cas9 TCY biosensor cells transduced with lentiviral expression constructs for target gene 2 gRNA1, target gene 8 gRNA5, a non-targeting gRNA, and no gRNA are shown. Cells were cultured in conditioned medium or fresh medium. The x-axis represents CFP (405 nm laser excitation), and the y-axis represents FRET (excitation from CFP emission). The upper right quadrant represents FRET [+], the lower right quadrant represents CFP [+], and the lower left quadrant represents double negative. Disruption of target gene 2 or 8 increases tau aggregate formation in response to tau aggregate-conditioned medium but not to fresh medium. [Figure 12] Figure 1 shows a schematic diagram of secondary screening in Cas9 TCY biosensor cells transduced with lentiviral expression constructs of sgRNAs targeting target genes 2 and 8, including mRNA expression analysis, protein expression analysis, and FRET analysis. Two sgRNAs were used against target gene 2 (g1 and g3), one sgRNA was used against target gene 8 (g5), and a non-targeting sgRNA (g3) was used as a non-targeting control. [Figure 13] Figure 1 shows the relative expression of target gene 2 and target gene 8 in Cas9 TCY biosensor cells as assessed by qRT-PCR 6 days after transduction with lentiviral sgRNA expression constructs. [Figure 14]FIG. 1 shows the expression of protein 2 (encoded by target gene 2) and protein 8 (encoded by target gene 8) in Cas9 TCY biosensor cells as assessed by Western blot 13 days after transduction with lentiviral sgRNA expression constructs. [Figure 15] Figure 1 shows tau aggregation as measured by percent FRET[+] cells in Cas9 TCY biosensor cells 10 days after transduction with lentiviral sgRNA expression constructs. Lipofectamine was not used. [Figure 16] 1 shows the expression of target gene 2 and target gene 8 in knockdown Cas9 TCY cell clones as assessed by Western blot. [Figure 17] 1 shows the expression of tau aggregation in target gene 2 and target gene 8 knockdown Cas9 TCY cell clones as assessed by FRET. [Figure 18] 1 shows the expression of target gene 2 and target gene 8 in knockdown Cas9 TCY cell clones as assessed by Western blot, and the phosphorylation of tau at positions S262 and S356 in those clones as assessed by Western blot. [Figure 19] We demonstrate that whole-cell lysates from tau-YFP Agg[+] clone 18 can induce tau aggregation and FRET signals in tau biosensor cells. Different amounts of whole-cell lysate were tested, as were different sonication conditions for generating the lysates. [Figure 20] We show that whole cell lysate from tau-YFP Agg[+] clone 18 can induce tau aggregation and FRET signals in tau biosensor cells. Different amounts of whole cell lysate and different amounts of Lipofectamine were tested. [Figure 21]We demonstrate that whole-cell lysates from tau-YFP Agg[+] clone 18 can induce tau aggregation and FRET signals in tau biosensor cells, whereas whole-cell lysates from Agg[-] clones cannot. Different amounts of whole-cell lysates and different amounts of Lipofectamine were tested. [Figure 22] A schematic diagram illustrating the strategy of genome-wide CRISPR nuclease (CRISPRn) screening to identify modifier genes that prevent tau aggregation. [Figure 23] Graph showing identification of genes specifically enriched for sgRNAs in FRET[-] samples. [Figure 24] Graph showing identification of genes specifically depleted by sgRNA in FRET[-] samples. [Figure 25] FIG. 1 shows a schematic diagram illustrating a secondary screening strategy to validate identified modifier genes that prevent tau aggregation. [Figure 26] A schematic diagram illustrating the strategy of genome-wide CRISPR activation (CRISPRa) screening to identify modifier genes that prevent tau aggregation is shown. [Figure 27] A schematic diagram illustrating the strategy of genome-wide CRISPR nuclease (CRISPRn) screening to identify modifier genes that promote tau disaggregation is shown. [Figure 28] The gating used to sort Agg[+], speckle[+], and Agg[-] cell populations is shown. [Figure 29] Figure 1 shows a schematic diagram of the thymidine block strategy used in genome-wide CRISPR nuclease (CRISPRn) screening to identify modifier genes that promote tau disaggregation. DETAILED DESCRIPTION OF THE INVENTION
[0128] definition The terms "protein," "polypeptide," and "peptide," used interchangeably herein, include polymeric forms of amino acids of any length, including coded and non-coded amino acids and amino acids that are chemically or biochemically modified or derivatized. These terms also include modified polymers, such as polypeptides with modified peptide backbones. The term "domain" refers to any portion of a protein or polypeptide having a specific function or structure.
[0129] Proteins are said to have an "N-terminus" and a "C-terminus". The term "N-terminus" refers to the beginning of a protein or polypeptide terminated by an amino acid with a free amine group (-NH2). The term "C-terminus" refers to the end of an amino acid chain (protein or polypeptide) terminated by a free carboxyl group (-COOH).
[0130] The terms "nucleic acid" and "polynucleotide," used interchangeably herein, include polymeric forms of nucleotides of any length, containing ribonucleotides, deoxyribonucleotides, or analogs or modified versions thereof. These include single-, double-, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers that contain purine bases, pyrimidine bases, or other natural, chemically modified, biochemically modified, non-natural, or derivatized nucleotide bases.
[0131] Nucleic acids are said to have a "5' end" and a "3' end" because mononucleotides react to form oligonucleotides in a manner such that the 5' phosphate of one mononucleotide pentose ring is unidirectionally linked to the 3' oxygen of its neighbor via a phosphodiester bond. The end of an oligonucleotide is called the "5' end" if the 5' phosphate of the mononucleotide pentose ring is not linked to the 3' oxygen of the mononucleotide pentose ring. The end of an oligonucleotide is called the "3' end" if its 3' oxygen is not linked to the 5' phosphate of another mononucleotide pentose ring. A nucleic acid sequence can also be said to have 5' and 3' ends, even if it is internal to a larger oligonucleotide. In either a linear or circular DNA molecule, separate elements are referred to as "upstream" or "downstream" 5' or 3' elements.
[0132] The term "genomically integrated" refers to a nucleic acid that has been introduced into a cell such that the nucleotide sequence is integrated into the genome of the cell. Any protocol may be used for stable integration of a nucleic acid into the genome of a cell.
[0133] The term "targeting vector" refers to a recombinant nucleic acid that can be introduced by homologous recombination, non-homologous end joining-mediated ligation, or any other means of recombination into a target location in the genome of a cell.
[0134] The term "viral vector" refers to a recombinant nucleic acid that contains at least one element of viral origin and contains elements sufficient for or that allow packaging into a viral vector particle. The vector and / or particle can be used to transfer DNA, RNA, or other nucleic acids into cells either ex vivo or in vivo. Many forms of viral vectors are known.
[0135] The term "wild-type" includes entities having a structure and / or activity as found in a normal state or context (as opposed to mutant, diseased, altered, etc.). Wild-type genes and polypeptides often exist in multiple alternative forms (e.g., alleles).
[0136] The term "endogenous sequence" refers to a nucleic acid sequence that occurs naturally within a cell or organism. For example, an endogenous MAPT sequence of a cell or organism refers to the native MAPT sequence that occurs naturally at the MAPT locus of the cell or organism.
[0137] An "exogenous" molecule or sequence includes a molecule or sequence that is not normally present in a cell in that form. Normal presence includes presence in relation to a particular developmental stage and environmental conditions of the cell. An exogenous molecule or sequence may include, for example, a mutated version of a corresponding endogenous sequence in a cell, such as a humanized version of an endogenous sequence, or may include a sequence that corresponds to an endogenous sequence within a cell but in a different form (i.e., not within a chromosome). In contrast, an endogenous molecule or sequence includes a molecule or sequence that is normally present in that form, in a particular cell, at a particular developmental stage, and under particular environmental conditions.
[0138] The term "heterologous" when used in the context of a nucleic acid or protein indicates that the nucleic acid or protein contains at least two segments that do not naturally occur together in the same molecule. For example, the term "heterologous" when used with reference to a segment of a nucleic acid or a segment of a protein indicates that the nucleic acid or protein contains two or more subsequences that are not found in the same relationship to each other (e.g., linked together) in nature. As an example, a "heterologous" region of a nucleic acid vector is a segment of nucleic acid within or attached to another nucleic acid molecule that is not found in association with that other molecule in nature. For example, a heterologous region of a nucleic acid vector can include a coding sequence that is adjacent to a sequence not found in association with the coding sequence in nature. Similarly, a "heterologous" region of a protein is a segment of amino acids within or attached to another peptide molecule that is not found in association with the other peptide molecule in nature (e.g., a fusion protein or a tagged protein). Similarly, a nucleic acid or protein can include a heterologous tag or a heterologous secretion or localization sequence.
[0139] The term "locus" refers to a specific location of a gene (or key sequence), DNA sequence, polypeptide coding sequence, or location on a chromosome of an organism's genome. For example, "MAPT locus" can refer to a specific location of the MAPT gene, a MAPT DNA sequence, a microtubule-associated protein tau coding sequence, or a MAPT location on a chromosome of an organism's genome identified with respect to where such a sequence resides. "MAPT locus" can include regulatory elements of the MAPT gene, including, for example, an enhancer, a promoter, a 5' and / or 3' untranslated region (UTR), or a combination thereof.
[0140] The term "gene" refers to a DNA sequence in a chromosome that encodes a product (e.g., an RNA product and / or a polypeptide product), including coding regions interrupted by non-coding introns and sequences located adjacent to the coding region at both the 5' and 3' ends, such that the gene corresponds to the full-length mRNA (including 5' and 3' untranslated sequences). The term "gene" also includes other non-coding sequences, including regulatory sequences (e.g., promoters, enhancers, and transcription factor binding sites), polyadenylation signals, internal ribosome entry sites, silencers, insulating sequences, and matrix attachment regions. These sequences may be close to (e.g., within 10 kb) or distant from the coding region of a gene, and they affect the level or rate of transcription and translation of the gene.
[0141] The term "allele" refers to variant forms of a gene. Some genes have different forms that are located at the same position, or locus, on a chromosome. Diploid organisms have two alleles at each locus. Each pair of alleles represents a genotype at a particular locus. A genotype is described as homozygous if there are two identical alleles at a particular locus, and as heterozygous if the two alleles are different.
[0142] A "promoter" is a regulatory region of DNA that typically contains a TATA box that can direct RNA polymerase II to begin RNA synthesis at the appropriate transcription start site for a particular polynucleotide sequence. A promoter may further contain other regions that affect the rate of transcription initiation. The promoter sequences disclosed herein regulate the transcription of an operably linked polynucleotide. The promoter may be active in one or more cell types disclosed herein (e.g., human cells, pluripotent cells, one-cell stage embryos, differentiated cells, or a combination thereof). The promoter may be, for example, a constitutively active promoter, a conditional promoter, an inducible promoter, a temporally restricted promoter (e.g., a developmentally regulated promoter), or a spatially restricted promoter (e.g., a cell-specific or tissue-specific promoter). Examples of promoters can be found, for example, in WO2013 / 176772, incorporated herein by reference in its entirety for all purposes.
[0143] "Operable linkage" or "operably linked" includes the juxtaposition of two or more components (e.g., a promoter and another sequence element) that allows for both components to function normally and for at least one of the components to mediate the function of at least one of the other components. For example, a promoter can be operably linked to a coding sequence if it controls the level of transcription of the coding sequence depending on the presence or absence of one or more transcriptional regulatory factors. Operable linkage can include proximity of such sequences to each other or acting in trans (e.g., regulatory sequences can act at a distance to control transcription of the coding sequence).
[0144] The term "variant" refers to a nucleotide sequence (eg, one nucleotide) that differs from the most common sequence in a population, or a protein sequence (eg, one amino acid) that differs from the most common sequence in a population.
[0145] The term "fragment," when referring to a protein, refers to a protein that is shorter or has fewer amino acids than the full-length protein. The term "fragment," when referring to a nucleic acid, refers to a nucleic acid that is shorter or has fewer nucleotides than the full-length nucleic acid. A fragment can be, for example, an N-terminal fragment (i.e., removal of a portion of the C-terminus of the protein), a C-terminal fragment (i.e., removal of a portion of the N-terminus of the protein), or an internal fragment.
[0146] "Sequence identity" or "identity" in the context of two polynucleotide or polypeptide sequences refers to the residues of the two sequences that are the same when aligned for maximum correspondence over a specified comparison window. When using percentage sequence identity in proteins, non-identical residue positions often differ by conservative amino acid substitutions, in which an amino acid residue is replaced with another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity) and thus does not alter the functional properties of the molecule. When sequences differ by conservative substitutions, the percent sequence identity may be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." Means for making this adjustment are well known. Typically, this involves scoring conservative substitutions as partial rather than complete mismatches, thereby increasing the percentage sequence identity. Thus, for example, where identical amino acids are given a score of 1 and non-conservative substitutions are given a score of zero, conservative substitutions are given a score between zero and 1. Scoring of conservative substitutions is calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, Calif.).
[0147] "Percentage of sequence identity" includes a value determined by comparing two optimally aligned sequences (maximum number of perfectly matched residues) over a comparison window, where the portion of the polynucleotide sequence in the comparison window may contain additions or deletions (i.e., gaps) when compared to a reference sequence (without additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions where the same nucleic acid base or amino acid residue occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Unless otherwise specified (e.g., the shorter sequence includes a concatenated non-homologous sequence), the comparison window is the entire length of the shorter of the two sequences being compared.
[0148] Unless otherwise specified, sequence identity / similarity values include values obtained using GAP version 10 with the following parameters: % identity and % similarity for nucleotide sequences using a GAP weight of 50 and a length weight of 3, and the nwsgapdna.cmp scoring matrix; % identity and % similarity for amino acid sequences using a GAP weight of 8 and a length weight of 2, and the BLOSUM62 scoring matrix; or any equivalent program. "Equivalent program" includes any sequence comparison program that produces alignments with identical nucleotide or amino acid residue matches and identical percent sequence identity for any two sequences in question when compared to corresponding alignments produced by GAP version 10.
[0149] The term "conservative amino acid substitution" refers to the substitution of an amino acid normally present in a sequence with a different amino acid of similar size, charge, or polarity. Examples of conservative substitutions include the substitution of a non-polar (hydrophobic) residue such as isoleucine, valine, or leucine for another non-polar residue. Similarly, examples of conservative substitutions include the substitution of one polar (hydrophilic) residue for another, such as between arginine and lysine, between glutamine and asparagine, or between glycine and serine. In addition, the substitution of a basic residue such as lysine, arginine, or histidine for another, or the substitution of one acidic residue such as aspartic acid or glutamic acid for another, are additional examples of conservative substitutions. Examples of non-conservative substitutions include the substitution of a non-polar (hydrophobic) amino acid residue, such as isoleucine, valine, leucine, alanine, or methionine, for a polar (hydrophilic) residue, such as cysteine, glutamine, glutamic acid, or lysine, and / or a polar residue for a non-polar residue. Typical amino acid classifications are summarized below. [Table 1]
[0150] A "homologous" sequence (e.g., a nucleic acid sequence) includes a sequence that is identical to or substantially similar to a known reference sequence, e.g., at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the known reference sequence. Homologous sequences can include, for example, orthologous and paralogous sequences. For example, homologous genes typically originate from a common ancestral DNA sequence through either speciation events (orthologous genes) or gene duplication events (paralogous genes). "Orthologous" genes include genes from different species that have evolved from a common ancestral gene through speciation. Orthologs typically retain the same function during evolution. "Paralogous" genes include genes that are related by duplication within a genome. Paralogs can evolve new functions during evolution.
[0151] The term "in vitro" includes an artificial environment and processes or reactions that occur within an artificial environment (e.g., a test tube or an isolated cell or cell line). The term "in vivo" includes a natural environment (e.g., a cell or organism or body) and processes or reactions that occur within a natural environment. The term "ex vivo" includes cells removed from an individual's body and processes or reactions that occur within such cells.
[0152] The term "reporter gene" refers to a nucleic acid having a sequence encoding a gene product (typically an enzyme) that is easily and quantitatively assayed when a construct containing the reporter gene sequence operably linked to a heterologous promoter and / or enhancer element is introduced into a cell that contains (or can be made to contain) the factors necessary for activation of the promoter and / or enhancer element. Examples of reporter genes include, but are not limited to, the gene encoding beta-galactosidase (lacZ), the bacterial chloramphenicol acetyltransferase (cat) gene, the firefly luciferase gene, the gene encoding beta-glucuronidase (GUS), and genes encoding fluorescent proteins. "Reporter protein" refers to the protein encoded by the reporter gene.
[0153] As used herein, the term "fluorescent reporter protein" refers to a reporter protein that is detectable based on fluorescence, which can be from either the reporter protein directly, the activity of a fluorogenic substrate of the reporter protein, or a protein with affinity for binding to a fluorescently tagged compound. Examples of fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, and ZsGreenl), yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, and ZsYellowl), blue fluorescent proteins (e.g., BFP, eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, and T-sapphire), cyan fluorescent proteins (e.g., CFP, eCFP, Cerulean, CyPet, AmCyanl, and Midori). mishi-Cyan), red fluorescent proteins (e.g., RFP, mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRaspberry, mStrawberry, and Jred), orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, and tdTomato), and any other suitable fluorescent proteins whose presence within a cell can be detected by flow cytometry methods.
[0154] Repair in response to double-strand breaks (DSBs) mainly occurs through two conserved DNA repair pathways: homologous recombination (HR) and non-homologous end joining (NHEJ). See Kasparek & Humphrey (2011) Seminars in Cell & Dev. Biol. 22:886-897, which is incorporated herein by reference in its entirety for all purposes. Similarly, repair of target nucleic acid mediated by exogenous donor nucleic acid can include any process of exchanging genetic information between two polynucleotides.
[0155] The term "recombination" includes any process of exchanging genetic information between two polynucleotides and can occur by any mechanism. Recombination can occur via homology-directed repair (HDR) or homologous recombination (HR). HDR or HR involves forms of nucleic acid repair that may require nucleotide sequence homology and use a "donor" molecule as a template for repair of a "target" molecule (i.e., a molecule that has experienced a double-strand break), resulting in the transfer of genetic information from the donor to the target. Without wishing to be bound by any particular theory, such transfer may involve mismatch correction of heteroduplex DNA formed between the broken target and donor, and / or synthesis-dependent strand annealing, and / or related processes, where the donor resynthesizes the genetic information that will become part of the target. In some cases, the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide is integrated into the target DNA. See Wang et al. (2013) Cell 153:910-918, Mandalos et al. (2012) PLOS ONE 7:e45768:1-9, and Wang et al. (2013) Nat Biotechnol. 31:530-532, each of which is incorporated by reference in its entirety for all purposes.
[0156] NHEJ involves the repair of double-strand breaks in nucleic acids by directly ligating the cut ends to each other or to an exogenous sequence without the need for a homologous template. Ligation of non-adjacent sequences by NHEJ can often result in deletions, insertions, or translocations near the site of the double-strand break. For example, NHEJ can also result in targeted integration of an exogenous donor nucleic acid by directly ligating the cut ends to the ends of the exogenous donor nucleic acid (i.e., NHEJ-based capture). Such NHEJ-mediated targeted integration may be preferable for the insertion of an exogenous donor nucleic acid when the homology-directed repair (HDR) pathway is not readily available (e.g., non-dividing cells, primary cells, and cells that perform poorly in homology-based DNA repair). Furthermore, in contrast to homology-directed repair, knowledge of large regions of sequence identity adjacent to the break site is not required, which can be beneficial when attempting targeted insertion into an organism with a genome with limited knowledge of the genome sequence. Integration can be carried out by blunt-end ligation between exogenous donor nucleic acid and cleaved genome sequence, or by sticky-end ligation (i.e., having 5' or 3' overhang) using exogenous donor nucleic acid adjacent to the overhang that matches that generated by nuclease agent in cleaved genome sequence.See, for example, US2011 / 020722, WO2014 / 033644, WO2014 / 089290, and Maresca et al.(2013)Genome Res.23(3):539-546, each of which is incorporated herein by reference in its entirety for all purposes.When blunt-end ligation is carried out, the target and / or donor may need to be excised due to the generation of microhomology regions required for fragment binding, which may cause undesirable changes in the target sequence.
[0157] A composition or method "comprising" or "including" one or more recited elements may include other elements not specifically recited. For example, a composition "comprises" or "includes" a protein may contain the protein alone or in combination with other ingredients. The transitional phrase "consisting essentially of" means that the claim scope shall be construed to include the specific elements recited in the claim as well as elements that do not materially affect the basic and novel characteristics of the claimed invention. Thus, the term "consisting essentially of," when used in the claims of the present invention, is not intended to be construed as equivalent to "comprising."
[0158] "Optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes examples when the event or circumstance occurs and examples when it does not occur.
[0159] The specification of a range of values includes all integers within or defining the range, and all subranges defined by integers within the range.
[0160] Unless otherwise clear from the context, the term "about" encompasses values within the standard error of measurement (eg, SEM) of the stated value.
[0161] The term "and / or" refers to and includes all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or").
[0162] The term "or" refers to any one member of a particular list and also includes any combination of members of that list.
[0163] The singular articles "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "protein" or "at least one protein" can include a plurality of proteins, including mixtures thereof.
[0164] Statistically significant means p≦0.05. [Mode for Carrying Out the Invention]
[0165] I. Overview Cas protein-enabled tau biosensor cells and methods of making and using such cells to screen for genetic modifiers of tau seeding or aggregation are provided. CRISPR / Cas synergistic activation mediator (SAM)-enabled tau biosensor cells and methods of making and using such cells to screen for genetic modifiers of tau seeding or aggregation are provided. Cas protein-enabled tau biosensor cells and methods of making and using such cells to screen for genetic modifiers of tau disaggregation or aggregation are provided. CRISPR / Cas synergistic activation mediator (SAM)-enabled tau biosensor cells and methods of making and using such cells to screen for genetic modifiers of tau disaggregation are provided. Reagents and methods for sensitizing such cells to tau seeding activity or tau aggregation are also provided. Reagents and methods for inducing tau aggregation are also provided.
[0166] To identify genes and pathways that alter the process of abnormal tau protein aggregation, a platform has been developed to screen using CRISPR (e.g., CRISPR / Cas9) nuclease (CRISPRn) sgRNA libraries to identify genes that regulate the likelihood of a cell being "seeded" by tau disease-associated protein aggregates (e.g., genes that, when disrupted, render cells more susceptible to tau aggregate formation when exposed to a source of tau fibrillary protein). To further identify genes and pathways that alter the process of abnormal tau protein aggregation, a platform has been developed to screen using CRISPR-activated (CRISPRa) sgRNA libraries to identify genes that regulate the likelihood of a cell being "seeded" by tau disease-associated protein aggregates (e.g., genes that, when transcriptionally activated, render cells more susceptible to tau aggregate formation when exposed to a source of tau fibrillary protein). Similarly, a platform has been developed to screen using CRISPR (CRISPR / Cas9) nuclease (CRISPRn) sgRNA libraries to identify genes that, when disrupted, prevent tau aggregation or promote tau disaggregation. Similarly, a platform has been developed to screen using CRISPR-activated (CRISPRa) sgRNA libraries to identify genes that, when transcriptionally activated, prevent tau aggregation or promote tau disaggregation. "Seed" refers to one or more proteins that nucleate the aggregation of other proteins with similar aggregation domains. The seeding activity of a sample refers to the sample's ability to nucleate (i.e., induce) the aggregation of proteins with similar aggregation domains. Identification of such genes may elucidate the mechanisms and genetic pathways of tau intercellular aggregate growth that govern neuronal susceptibility to tau aggregate formation in the context of neurodegenerative disease.
[0167] The screen uses a tau biosensor cell line (e.g., a human cell line, or HEK293T) consisting of cells stably expressing a tau repeat domain (e.g., tau 4 repeat domain, tau_4RD) with a pathogenic mutation (e.g., the P301S pathogenic mutation) linked to a unique reporter that can act together as an intracellular biosensor to generate a detectable signal upon aggregation. In one non-limiting example, the cell line contains two transgenes that stably express disease-associated protein variants fused to the fluorescent protein CFP (e.g., eCFP) or the fluorescent protein YFP (eYFP): tau4RD-CFP / tau4RD-YFP (TCY), where the tau repeat domain (4RD) contains the P301S pathogenic mutation. In these biosensor lines, aggregation of the tau-CFP / tau-YFP proteins generates a fluorescence resonance energy transfer (FRET) signal, which is the result of the transfer of fluorescence energy from the donor CFP to the acceptor YFP. As used herein, the term CFP (cyan fluorescent protein) includes eCFP (enhanced cyan fluorescent protein), and the term YFP (yellow fluorescent protein) includes eYFP (enhanced yellow fluorescent protein). FRET-positive cells containing tau aggregates can be sorted and isolated by flow cytometry. At baseline, unstimulated cells express the reporter in a stable, soluble state with minimal FRET signal. Upon stimulation (e.g., liposomal transfection of seed particles), the reporter protein forms aggregates and generates a FRET signal. Aggregate-containing cells can be isolated by FACS. Stably growing aggregate-containing cell lines, Agg[+], can be isolated by clonal serial dilution of Agg[-] cell lines.
[0168] This tau biosensor cell line was modified in several ways to make it useful for genetic screening using CRISPRn libraries. First, these tau biosensor cells were modified by introducing a Cas-expressing transgene (e.g., Cas9 or SpCas9) for use in CRISPRn screening. Next, reagents and methods were developed to sensitize cells to tau seeding activity and tau aggregation. A cell line was developed in which tau aggregates stably persisted in all cells, proliferating over time and passaged multiple times. These cells were used to generate conditioned medium by collecting media present on confluent cells for a period of time. This conditioned medium could then be applied to naive tau biosensor tau cells at a ratio that would induce tau aggregation in a small percentage of these recipient cells, thereby sensitizing them to tau seeding activity and tau aggregation. Conditioned medium without co-culture has not previously been used in this context as a seeding agent. However, because in vitro-generated tau fibrils are a limited source, conditioned medium is particularly useful for large-scale genome-wide screening. Furthermore, conditioned medium is more physiologically relevant because it is produced and secreted by cells rather than in vitro.
[0169] We used these cell lines to develop a screening method in which a CRISPR guide RNA library was introduced into aggregate-free Cas-expressing tau biosensor cells (Agg[-]) and knockout mutations were introduced into each target gene. After culturing the cells to allow for genome editing and expansion, we grew them in conditioned medium to sensitize them to seeding activity and identify cells that developed tau aggregation. To identify genes that can modulate the susceptibility of cells to tau seeding when exposed to an external source of tau seeding activity, we identified guide RNAs enriched in the aggregation-positive subpopulation at early time points during genome editing and expansion.
[0170] Similarly, several modifications were made to this tau biosensor cell line to make it useful for genetic screening using the CRISPRa library (e.g., for use in the CRISPR / Cas synergistic activation mediator (SAM) system). In an exemplary SAM system, several activation domains interact to cause greater transcriptional activation than can be induced by any one factor alone. For example, an exemplary SAM system includes a chimeric Cas protein comprising a nuclease-inactive Cas protein fused to one or more transcriptional activation domains (e.g., VP64), and a chimeric adaptor protein comprising an adaptor protein (e.g., MS2 coat protein (MCP)) fused to one or more transcriptional activation domains (e.g., fused to p65 and HSF1). MCP naturally binds to the MS2 stem loop. In the exemplary SAM system, MCP interacts with the MS2 stem loop engineered into the CRISPR-associated sgRNA, thereby shuttles the bound transcription factor to the appropriate genomic location.
[0171] First, these tau biosensor cells were modified by introducing one or more transgenes expressing a chimeric Cas protein comprising a nuclease-inactive Cas protein fused to one or more transcription activation domains (e.g., VP64) and a chimeric adaptor protein comprising an adaptor protein (e.g., MS2 coat protein (MCP)) fused to one or more transcription activation domains (e.g., fused to p65 and HSF1). While the SAM system is described herein, other CRISPRa systems, such as a nuclease-inactive Cas protein fused to one or more transcription activation domains, can also be used, and such systems would also not include a chimeric adaptor protein. In such cases, the tau biosensor cells would be modified by introducing a transgene expressing the chimeric Cas protein.
[0172] Next, reagents and methods were developed to sensitize cells to tau seeding activity and tau aggregation. Cell lines were developed in which tau aggregates stably persisted in all cells, proliferating over time and passaged multiple times. These cells were used to generate conditioned medium by collecting media present on confluent cells for a period of time. This conditioned medium could then be applied to naive tau biosensor tau cells at a ratio sufficient to induce tau aggregation in a small percentage of these recipient cells, thereby sensitizing them to tau seeding activity and tau aggregation. Conditioned medium without co-culture has not previously been used in this context as a seeding agent. However, because in vitro-generated tau fibrils are a limited source, conditioned medium is particularly useful for large-scale genome-wide screening. Furthermore, conditioned medium is more physiologically relevant because it is produced and secreted by cells rather than in vitro.
[0173] We used these cell lines to develop a screening method for transfecting aggregate-free SAM-expressing tau biosensor cells (Agg[-]) with a CRISPRa guide RNA library and transactivating each target gene. After culturing the cells to allow for genome editing and expansion, we grew them in conditioned medium to sensitize them to seeding activity and identify cells that developed tau aggregation. To identify genes that can modulate the susceptibility of cells to tau seeding when exposed to an external source of tau seeding activity, we identified guide RNAs enriched in the aggregation-positive subpopulation at early time points during genome editing and expansion.
[0174] II. Cas / Tau Biosensor and SAM / Tau Biosensor Cell Lines and Methods of Generation A. Cas / Tau biosensor cells and SAM / Tau biosensor cells Disclosed herein are cells that express not only a first tau repeat domain (e.g., comprising a tau microtubule-binding domain (MBD)) linked to a first reporter and a second tau repeat domain linked to a second reporter, but also express a Cas protein, such as Cas9. Also disclosed herein are cells that express not only a first tau repeat domain (e.g., comprising a tau microtubule-binding domain (MBD)) linked to a first reporter and a second tau repeat domain linked to a second reporter, but also express a chimeric Cas protein comprising a nuclease-inactive Cas protein fused to one or more transcription activation domains, and a chimeric adaptor protein comprising an adaptor protein fused to one or more transcription activation domains. The first tau repeat domain linked to the first reporter can be stably expressed, and the second tau repeat domain linked to the second reporter can be stably expressed. For example, DNA encoding a first tau repeat domain linked to a first reporter can be genomically integrated, and DNA encoding a second tau repeat domain linked to a second reporter can be genomically integrated. Similarly, a Cas protein can be stably expressed in a Cas / tau biosensor cell. For example, DNA encoding a Cas protein can be genomically integrated. Similarly, a chimeric Cas protein and / or a chimeric adaptor protein can be stably expressed in a SAM / tau biosensor cell. For example, DNA encoding a chimeric Cas protein can be genomically integrated, and / or DNA encoding a chimeric adaptor protein can be genomically integrated. The cell can be tau aggregation negative or tau aggregation positive.
[0175] 1. Tau and tau repeat domains linked to reporters The microtubule-associated protein tau, which promotes microtubule assembly and stability, is primarily expressed in neurons. Tau stabilizes neuronal microtubules and thus promotes axonal elongation. In Alzheimer's disease (AD) and a family of related neurodegenerative disorders called tauopathies, tau protein becomes abnormally hyperphosphorylated and aggregates into bundles of filaments (paired helical filaments) that appear as neurofibrillary tangles. Tauopathies are a heterogeneous group of neurodegenerative conditions characterized by abnormal tau deposition in the brain.
[0176] The tau repeat domain can be derived from tau protein from any animal or mammal, such as human, mouse, or rat. In one specific example, the tau repeat domain is derived from human tau protein. An exemplary human tau protein has been assigned the UniProt accession number P10636. Tau protein is the product of alternative splicing from a single gene called MAPT (microtubule-associated protein tau) in humans. The tau repeat domain carries sequence motifs involved in aggregation (i.e., it is the aggregation-prone domain from tau). Depending on splicing, the repeat domain of tau protein has either three or four repeat regions that constitute the aggregation-prone core of the protein, often referred to as the repeat domain (RD). Specifically, the tau repeat domain represents the core of the microtubule-binding region and harbors the R2 and R3 hexapeptide motifs involved in tau aggregation. There are six tau isoforms in the human brain, ranging in length from 352 to 441 amino acids. These isoforms vary in the presence or absence of one or two insert domains at the amino terminus, as well as the presence of either three or four repeat domains (R1-R4) at the carboxyl terminus. The repeat domains located in the carboxyl-terminal half of tau are thought to be important for microtubule binding and the pathological aggregation of tau into paired helical filaments (PHFs), the core components of the neurofibrillary tangles seen in tauopathies. Exemplary sequences of the four repeat domains (R1-R4) are provided in SEQ ID NOS: 1-4, respectively. Exemplary coding sequences for the four repeat domains (R1-R4) are provided in SEQ ID NOS: 5-8. An exemplary sequence of the tau4 repeat domain is provided in SEQ ID NOS: 9. An exemplary coding sequence of the tau4 repeat domain is provided in SEQ ID NOS: 10. An exemplary sequence of the tau4 repeat domain with a P301S mutation is provided in SEQ ID NOS: 11. An exemplary coding sequence for a tau 4 repeat domain with a P301S mutation is provided in SEQ ID NO:12.
[0177] The tau repeat domain used in the Cas / tau or SAM / tau biosensor cells may include a tau microtubule-binding domain (MBD). The tau repeat domain used in the Cas / tau or SAM / tau biosensor cells may include one or more or all of the four repeat domains (R1-R4). For example, the tau repeat domain may comprise, consist essentially of, or consist of one or more or all of SEQ ID NOS: 1, 2, 3, and 4, or a sequence that is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NOS: 1, 2, 3, and 4. In one specific example, the tau repeat domain is the tau 4 repeat domain (R1-R4) found in some tau isoforms. The tau 4 repeat domain may be used in place of full-length tau to reliably form fibrils in cultured cells. For example, a tau repeat domain can comprise, consist essentially of, or consist of SEQ ID NO:9 or SEQ ID NO:11, or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:9 or SEQ ID NO:11. In one particular example, a nucleic acid encoding a tau repeat domain can comprise, consist essentially of, or consist of SEQ ID NO:12, or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:12, optionally, the nucleic acid encodes a protein comprising, consisting essentially of, or consisting of SEQ ID NO:11. In another specific example, the nucleic acid encoding the second tau repeat domain linked to the second reporter can comprise, consist essentially of, or consist of SEQ ID NO: 10 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 10, and optionally, the nucleic acid encodes a protein comprising, consisting essentially of, or consisting of SEQ ID NO: 9. The first and second tau repeat domains in the cells disclosed herein can be the same, similar, or different.
[0178] One or both of the first tau repeat domain linked to the first reporter and the second tau repeat domain linked to the second reporter can be stably expressed in the cells. For example, a nucleic acid encoding one or both of the first tau repeat domain linked to the first reporter and the second tau repeat domain linked to the second reporter can be genomically integrated in a cell population and operably linked to a promoter active in the cells.
[0179] The tau repeat domain used in the cells disclosed herein may also contain a tau pathogenic mutation, such as an aggregation-promoting mutation. Such a mutation may be, for example, a mutation associated with (e.g., segregating with) or causing a tauopathy. As an example, the mutation may be an aggregation-sensitizing mutation that sensitizes tau to seeding but does not cause tau to readily aggregate by itself. For example, the mutation may be a disease-associated P301S mutation. The P301S mutation refers to the human tau P301S mutation or the corresponding mutation of another tau protein when optimally aligned with the human tau protein. The P301S mutation of tau exhibits high susceptibility to seeding but does not readily aggregate by itself. Thus, at baseline, a tau reporter protein containing the P301S mutation exists in a stable, soluble form within the cell, but exposure to exogenous tau species causes the tau reporter protein to aggregate. Other tau mutations include, for example, K280del, P301L, V337M, P301L / V337M, and K280del / I227P / I308P.
[0180] The first tau repeat domain can be linked to a first reporter and the second tau repeat domain can be linked to a second reporter by any means, for example, the reporter can be fused to the tau repeat domain (e.g., as part of a fusion protein).
[0181] The first and second reporters can be a unique reporter pair that can act together as an intracellular biosensor that generates a detectable signal when the first and second proteins aggregate. As an example, the reporters can be fluorescent proteins, and protein aggregation can be measured using fluorescence resonance energy transfer (FRET). Specifically, the first and second reporters can be a FRET pair. Examples of FRET pairs (donor and acceptor fluorophores) are well known. See, for example, Bajar et al. (2016) Sensors (Basel) 16(9):1488, incorporated herein by reference in its entirety for all purposes. Typical fluorescence microscopy techniques rely on the absorption of a certain wavelength of light (excitation) by a fluorophore, followed by the subsequent emission of a secondary fluorescence at a longer wavelength. The mechanism of fluorescence resonance energy transfer involves a donor fluorophore in an excited electronic state, which can transfer its excitation energy to a nearby acceptor chromophore in a non-radiative manner via a long-range dipole-dipole interaction. For example, the FRET energy donor can be the first reporter, and the FRET energy acceptor can be the second reporter. Alternatively, the FRET energy donor can be the second reporter, and the FRET energy acceptor can be the first reporter. In a particular example, the first and second reporters are CFP and YFP. Exemplary protein and coding sequences for CFP are provided, for example, in SEQ ID NOs: 13 and 14, respectively. Exemplary protein and coding sequences for YFP are provided, for example, in SEQ ID NOs: 15 and 16, respectively. As a particular example, the CFP can comprise, consist essentially of, or consist of SEQ ID NO: 13 or a sequence that is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 13. As another specific example, the YFP can comprise, consist essentially of, or consist of SEQ ID NO:15 or a sequence that is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:15.
[0182] As another example, protein fragment complementation strategies can be used to detect aggregation. For example, split luciferase can be used to generate bioluminescence from a substrate, and the first and second reporters can be amino-(NLuc) and carboxy-(CLuc) terminal fragments of luciferase. Examples of luciferases include Renilla, firefly, click beetle, and Metridia luciferase.
[0183] In one non-limiting example, the biosensor cells disclosed herein contain two transgenes (Tau4RD-CFP / Tau4RD-YFP(TCY)) that stably express a disease-associated tau protein variant fused to the fluorescent protein CFP or the fluorescent protein YFP, respectively, where the tau repeat domain (4RD) contains the P301S pathogenic mutation. In these biosensor lines, aggregation of the tau-CFP / tau-YFP proteins generates a FRET signal, which is the result of the transfer of fluorescence energy from the donor CFP to the acceptor YFP. FRET-positive cells containing tau aggregates can be sorted and isolated by flow cytometry. At baseline, unstimulated cells express the reporter in a stable, soluble state with minimal FRET signal. Upon stimulation (e.g., liposomal transfection of seed particles), the reporter protein forms aggregates and generates a FRET signal.
[0184] The Cas / tau biosensor cells disclosed herein can be aggregation-positive (Agg[+]) cells in which the tau repeat domains are stably present in an aggregated state, meaning that the tau repeat domains remain stable in all cells over time and through expansion and multiple passages. Alternatively, the Cas / tau biosensor cells disclosed herein can be aggregation-negative (Agg[-]).
[0185] 2. Cas proteins and chimeric Cas proteins The Cas / tau biosensor cells disclosed herein also contain a nucleic acid (DNA or RNA) encoding a Cas protein. Optionally, the Cas protein is stably expressed. Optionally, the cell contains a genomically integrated Cas coding sequence. Similarly, the SAM / tau biosensor cells disclosed herein also contain a nucleic acid (DNA or RNA) encoding a chimeric Cas protein comprising a nuclease-inactive Cas protein fused to one or more transcription activation domains (e.g., VP64). Optionally, the chimeric Cas protein is stably expressed. Optionally, the cell contains a genomically integrated chimeric Cas coding sequence.
[0186] Cas proteins are part of the clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated (Cas) system. CRISPR / Cas systems include transcripts and other elements involved in the expression of or directing the activity of Cas genes. CRISPR / Cas systems can be, for example, type I, type II, type III, or type V systems (e.g., subtype VA or subtype VB). The methods and compositions disclosed herein can use CRISPR / Cas systems by utilizing CRISPR complexes (including guide RNAs (gRNAs) complexed with Cas proteins) for site-specific binding or cleavage of nucleic acids.
[0187] The CRISPR / Cas system used in the compositions and methods disclosed herein can be non-naturally occurring. "Non-naturally occurring" systems include those that show the involvement of human assistance, such as one or more components of the system that are modified or mutated from their naturally occurring state, or that are at least substantially free from at least one other component that they are essentially associated with in nature, or that are associated with at least one other component that they are not associated with in nature.For example, some CRISPR / Cas systems use non-naturally occurring CRISPR complexes that include non-naturally occurring gRNA and Cas protein together, or use non-naturally occurring Cas protein, or use non-naturally occurring gRNA.
[0188] Cas proteins generally contain at least one RNA recognition or binding domain capable of interacting with a guide RNA. Cas proteins may also contain a nuclease domain (e.g., a DNase or RNase domain), a DNA-binding domain, a helicase domain, a protein-protein interaction domain, a dimerization domain, and other domains. Some such domains (e.g., a DNase domain) may be derived from native Cas proteins. Other such domains can be added to create modified Cas proteins. The nuclease domain has catalytic activity for nucleic acid cleavage, including covalent cleavage of nucleic acid molecules. Cleavage can generate blunt or staggered ends, which can be single-stranded or double-stranded. For example, wild-type Cas9 proteins typically generate blunt cleavage products. Alternatively, wild-type Cpf1 proteins (e.g., FnCpf1) can produce cleavage products containing a 5-nucleotide 5' overhang, with cleavage occurring 18 base pairs from the PAM sequence on the non-targeted strand and 23 bases on the targeted strand. A Cas protein can have full cleavage activity and create a double-stranded break at the target genomic locus (e.g., a double-stranded break with a blunt end), or it can be a nickase that creates a single-stranded break at the target genomic locus.
[0189] Examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csn1 or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (CasA), Cse2 (CasB), Cse3 (CasE), These include Cse4 (CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cu1966, and homologs or modified versions thereof.
[0190] Exemplary Cas protein is Cas9 protein or a protein derived from Cas9 protein.Cas9 protein is derived from type II CRISPR / Cas system, and typically shares four important motifs with conserved architecture.Modifications 1, 2, and 4 are RuvC-like motifs, and motif 3 is HNH motif. Exemplary Cas9 proteins are Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Neisseria meningitidis, or Campylobacter jejuni. Additional examples of Cas9 family members are described in WO2014 / 131833, which is incorporated herein by reference in its entirety for all purposes. Cas9 from S. pyogenes (SpCas9), assigned SwissProt accession number Q99ZW2, is an exemplary Cas9 protein. An exemplary SpCas9 protein and coding sequence are set forth in SEQ ID NOs: 21 and 22, respectively. S.Cas9 from Staphylococcus aureus (SaCas9), assigned the UniProt accession number J7RUA5, is another exemplary Cas9 protein. Cas9 from Campylobacter jejuni (CjCas9), assigned the UniProt accession number Q0P897, is another exemplary Cas9 protein. See, e.g., Kim et al. (2017) Nat. Comm. 8:14500, incorporated herein by reference in its entirety for all purposes. SaCas9 is smaller than SpCas9, and CjCas9 is smaller than both SaCas9 and SpCas9. Cas9 from Neisseria meningitidis (Nme2Cas9) is another exemplary Cas9 protein. See, e.g., Edraki et al. (2019) Mol. Cell 73(4):714-726, incorporated herein by reference in its entirety for all purposes. Cas9 proteins from Streptococcus thermophilus (e.g., Streptococcus thermophilus LMD-9 Cas9 encoded by the CRISPR1 locus (St1Cas9) or Streptococcus thermophilus Cas9 from the CRISPR3 locus (St3Cas9)) are other exemplary Cas9 proteins. Cas9 from Francisella novicida (FnCas9) or the RHA Francisella novicida Cas9 variant (E1369R / E1449H / R1556A substitutions) that recognizes an alternative PAM are other exemplary Cas9 proteins. These and other exemplary Cas9 proteins are reviewed, for example, in Cebrian-Serrano and Davies (2017) Mamm. Genome 28(7):247-261, which is incorporated herein by reference in its entirety for all purposes.
[0191] As one example, the Cas protein can be a Cas9 protein. For example, the Cas9 protein can be a Streptococcus pyogenes Cas9 protein. As one particular example, the Cas protein can comprise, consist essentially of, or consist of SEQ ID NO:21 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:21. As another particular example, a chimeric Cas protein comprising a nuclease-inactive Cas protein and one or more transcription activation domains can comprise, consist essentially of, or consist of SEQ ID NO:36 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:36.
[0192] Another example of a Cas protein is the Cpf1 (CRISPR of Prevotella and Francisella 1) protein. Cpf1 is a large protein (approximately 1300 amino acids) that contains a RuvC-like nuclease domain homologous to the corresponding domain in Cas9, along with a counterpart of Cas9's characteristic arginine-rich cluster. However, Cpf1 lacks the HNH nuclease domain present in the Cas9 protein, and in contrast to Cas9, which contains a long insert containing the HNH domain, the RuvC-like domain is adjacent in the Cpf1 sequence. See, e.g., Zetsche et al. (2015) Cell 163(3):759-771, incorporated herein by reference in its entirety for all purposes. Exemplary Cpf1 proteins are Francisella tularensis 1, Francisella tularensis subsp.novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacteria GW2011_GWA2_33_10, Parcubacteria bacteria GW2011_GWC2_44_17, Smithella sp.SCADC, Acidaminococcus sp.BV3L6, Lachnospiraceae bacteria MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacteria ND2006, Porphyromonas crevioricanis 3, Prevotella Cpf1 from Francisella novicida U112 (FnCpf1; assigned UniProt accession number A0Q7Q2) is an exemplary Cpf1 protein.
[0193] The Cas protein can be a wild-type protein (i.e., naturally occurring), a modified Cas protein (i.e., a Cas protein variant), or a fragment of a wild-type or modified Cas protein. The Cas protein can also be a variant or fragment that is active with respect to catalytic activity of the wild-type or modified Cas protein. A catalytically active variant or fragment can contain at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with the wild-type or modified Cas protein or a portion thereof, and the active variant retains the ability to cleave at the desired cleavage site and thus retains double-strand break-inducing activity. Assays for double-strand break-inducing activity are known and generally measure the overall activity and specificity of the Cas protein on a DNA substrate containing the cleavage site.
[0194] One example of a modified Cas protein is the modified SpCas9-HF1 protein, which is a high-fidelity variant of Streptococcus pyogenes Cas9 (N497A / R661A / Q695A / Q926A) containing modifications designed to reduce nonspecific DNA contact. See, e.g., Kleinstiver et al. (2016) Nature 529(7587):490-495, incorporated herein by reference in its entirety for all purposes. Another example of a modified Cas protein is the modified eSpCas9 variant (K848A / K1003A / R1060A) designed to reduce off-target effects. See, e.g., Slaymaker et al. (2016) Science 351(6268):84-88, incorporated herein by reference in its entirety for all purposes. Other SpCas9 variants include K855A and K810A / K1003A / R1060A. These and other modified Cas proteins are reviewed, for example, in Cebrian-Serrano and Davies (2017) Mamm. Genome 28(7):247-261, which is incorporated herein by reference in its entirety for all purposes. Another example of a modified Cas9 protein is xCas9, which is an SpCas9 variant that can recognize an extended range of PAM sequences. See, for example, Hu et al. (2018) Nature 556:57-63, which is incorporated herein by reference in its entirety for all purposes.
[0195] Cas proteins can be modified to increase or decrease one or more of nucleic acid binding affinity, nucleic acid binding specificity, and enzymatic activity. Cas proteins can also be modified to alter any other activity or property of the protein, such as stability. For example, Cas proteins can be truncated to remove domains that are not essential for protein function or to optimize (e.g., enhance or reduce) the activity or property of the Cas protein. As another example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated (e.g., for use in SAM / tau biosensor cells containing nuclease-inactive Cas proteins).
[0196] Cas proteins can contain at least one nuclease domain, such as a DNase domain. For example, wild-type Cpf1 proteins generally contain a RuvC-like domain, likely in a dimeric conformation, that cleaves both strands of target DNA. Cas proteins can also contain at least two nuclease domains, such as a DNase domain. For example, wild-type Cas9 proteins generally contain a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC domain and the HNH domain can each cleave different strands of double-stranded DNA, creating a double-strand break in DNA. See, e.g., Jinek et al. (2012) Science 337:816-821, the entire contents of which are incorporated herein by reference for all purposes.
[0197] One or more or all of the nuclease domains can be deleted or mutated, rendering them nonfunctional or with reduced nuclease activity. For example, if one of the nuclease domains is deleted or mutated in a Cas9 protein, the resulting Cas9 protein is called a nickase and can generate single-strand breaks in double-stranded target DNA, but not double-strand breaks (i.e., it can cleave either the complementary or non-complementary strand, but not both). If both nuclease domains are deleted or mutated, the resulting Cas protein (e.g., Cas9) has reduced ability to cleave both strands of double-stranded DNA (e.g., a nuclease-null or nuclease-inactive Cas protein, or a Cas protein without catalytic activity (dCas)). An example of a mutation that converts Cas9 into a nickase is the D10A (aspartate to alanine at position 10 of Cas9) mutation in the RuvC domain of Cas9 from S. pyogenes. Similarly, H939A (histidine to alanine at amino acid position 839), H840A (histidine to alanine at amino acid position 840), or N863A (asparagine to alanine at amino acid position N863) in the HNH domain of Cas9 from S. pyogenes can convert Cas9 into a nickase. Other examples of mutations that convert Cas9 into a nickase include corresponding mutations in Cas9 from S. thermophilus. See, e.g., Sapranauskas et al. (2011) Nucleic Acids Res. 39(21):9275-9282 and WO2013 / 141680, each of which is incorporated by reference in its entirety for all purposes. Such mutations can be generated using methods such as site-directed mutagenesis, PCR-mediated mutagenesis, or total gene synthesis. Other examples of nickase-making mutations can be found, for example, in WO2013 / 176772 and WO2013 / 142578, each of which is incorporated by reference herein in its entirety for all purposes.When all of the nuclease domains are deleted or mutated in a Cas protein (e.g., when both nuclease domains are deleted or mutated in a Cas9 protein), the resulting Cas protein (e.g., Cas9) has a reduced ability to cleave both strands of double-stranded DNA (e.g., a nuclease-null or nuclease-inactive Cas protein). One particular example is the D10A / H840A S. pyogenes Cas9 double mutant, or the corresponding double mutant of a Cas9 from another species when optimally aligned with S. pyogenes Cas9. Another particular example is the D10A / N863A S. pyogenes Cas9 double mutant, or the corresponding double mutant of a Cas9 from another species when optimally aligned with S. pyogenes Cas9.
[0198] Examples of inactivating mutations in the catalytic domain of xCas9 are the same as those described above for SpCas9. Examples of inactivating mutations in the catalytic domain of the Staphylococcus aureus Cas9 protein are also known. For example, the Staphylococcus aureus Cas9 enzyme (SaCas9) can include a substitution at position N580 (e.g., an N580A substitution) and a substitution at position D10 (e.g., a D10A substitution) to generate a nuclease-inactive Cas protein. See, for example, WO2016 / 106236, the entire contents of which are incorporated herein by reference for all purposes. Examples of inactivating mutations in the catalytic domain of Nme2Cas9 are also known (e.g., a combination of D16A and H588A). Examples of inactivating mutations in the catalytic domain of St1Cas9 are also known (e.g., a combination of D9A, D598A, H599A, and N622A). Examples of inactivating mutations in the catalytic domain of St3Cas9 are also known (e.g., a combination of D10A and N870A). Examples of inactivating mutations in the catalytic domain of CjCas9 are also known (e.g., a combination of D8A and H559A). Examples of inactivating mutations in the catalytic domain of FnCas9 and RHA FnCas9 are also known (e.g., N995A).
[0199] Examples of inactivating mutations in the catalytic domain of the Cpf1 protein are also known. For the Cpf1 proteins from Francisella novicida U112 (FnCpf1), Acidaminococcus sp. BV3L6 (AsCpf1), Lachnospiraceae bacterium ND2006 (LbCpf1), and Moraxella bovoculi 237 (MbCpf1 Cpf1), such mutations can include mutations at positions 908, 993, or 1263 of AsCpf1 or corresponding positions in Cpf1 orthologs, or at positions 832, 925, 947, or 1180 of LbCpf1 or corresponding positions in Cpf1 orthologs. Such mutations can include, for example, one or more of the mutations D908A, E993A, and D1263A in AsCpf1 or corresponding mutations in Cpf1 orthologs, or D832A, E925A, D947A, and D1180A in LbCpf1 or corresponding mutations in Cpf1 orthologs. See, e.g., US2016 / 0208243, incorporated herein by reference in its entirety for all purposes.
[0200] Cas proteins can also be operably linked to heterologous polypeptides as fusion proteins. For example, Cas proteins can be fused to a cleavage domain, an epigenetic modification domain, a transcriptional activation domain, or a transcriptional repressor domain. See WO2014 / 089290, the entire contents of which are incorporated herein by reference for all purposes. For example, Cas proteins can be operably linked or fused to a transcriptional activation domain for use in SAM / tau biosensor cells. Examples of transcriptional activation domains include the herpes simplex virus VP16 activation domain, VP64 (a tetrameric derivative of VP16), the NFκB p65 activation domain, p53 activation domains 1 and 2, the CREB (cAMP response element binding protein) activation domain, the E2A activation domain, and the NFAT (nuclear factor of activated T cells) activation domain. Other examples include activation domains from Oct1, Oct-2A, SP1, AP-2, CTF1, P300, CBP, PCAF, SRC1, PvALF, ERF-2, OsGAI, HALF-1, C1, AP1, ARF-5, ARF-6, ARF-7, ARF-8, CPRF1, CPRF4, MYC-RP / GP, TRAB1PC4, and HSF1. See, for example, US2016 / 0237456, EP3045537, and WO2011 / 146121, each of which is incorporated by reference in its entirety for all purposes. In some cases, a transcription activation system comprising a dCas9-VP64 fusion protein paired with MS2-p65-HSF1 can be used. Guide RNAs for such systems can be designed using an aptamer sequence appended to an sgRNA tetraloop and stem loop 2 designed to bind to the dimerized MS2 bacteriophage coat protein. See, e.g., Konermann et al. (2015) Nature 517(7536):583-588, incorporated herein by reference in its entirety for all purposes.Examples of transcriptional repressor domains include the inducible cAMP early repressor (ICER) domain, the Kruppel-associated box A (KRAB-A) repressor domain, the YY1 glycine-rich repressor domain, the Sp1-like repressor, the E(spl) repressor, the IκB repressor, and MeCP2. Other examples include transcriptional repressor domains from A / B, KOX, TGF-beta inducible early gene (TIEG), v-erbA, SID, SID4X, MBD2, MBD3, DNMT1, DNMG3A, DNMT3B, Rb, and ROM2. See, for example, EP3045537 and WO2011 / 146121, each of which is incorporated by reference in its entirety for all purposes. Cas proteins can also be fused to heterologous polypeptides that provide increased or decreased stability. The fusion domain or heterologous polypeptide can be located at the N-terminus, C-terminus, or internally within the Cas protein.
[0201] Cas proteins can also be operably linked to heterologous polypeptides as fusion proteins. For example, Cas proteins can be fused to one or more heterologous polypeptides that provide subcellular localization. Such heterologous polypeptides include one or more nuclear localization signals (NLSs), such as the monopartite SV40 NLS and / or the bipartite alpha-importin NLS for nuclear targeting, a mitochondrial localization signal for mitochondrial targeting, an ER retention signal, etc. See, e.g., Lange et al. (2007) J. Biol. Chem. 282:5101-5105, incorporated herein by reference in its entirety for all purposes. Such subcellular localization signals can be located at the N-terminus, C-terminus, or anywhere within the Cas protein. The NLS can include a stretch of basic amino acids and can be a monopartite or bipartite sequence. Optionally, the Cas protein may contain two or more NLSs, including an N-terminal NLS (e.g., an alpha-importin NLS or a monokaryotic NLS) and a C-terminal NLS (e.g., an SV40 NLS or a bikaryotic NLS). The Cas protein may also contain two or more NLSs at the N-terminus and / or two or more NLSs at the C-terminus.
[0202] Cas proteins can also be operably linked to a cell penetration domain or protein transduction domain. For example, the cell penetration domain can be derived from the HIV-1 TAT protein, the TLM cell penetration motif from human hepatitis B virus, MPG, Pep-1, VP22, the cell penetration peptide from herpes simplex virus, or a polyarginine peptide sequence. See, for example, WO2014 / 089290 and WO2013 / 176772, each of which is incorporated herein by reference in its entirety for all purposes. The cell penetration domain can be located at the N-terminus, C-terminus, or anywhere within the Cas protein.
[0203] The Cas protein can also be operably linked to a heterologous polypeptide, such as a fluorescent protein, a purification tag, or an epitope tag, to facilitate tracking or purification. Examples of fluorescent proteins include green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, Emerald, Azami Green, Monomeric Azami). Green, CopGFP, AceGFP, ZsGreenl), yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, ZsYellowl), blue fluorescent proteins (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., eCFP, Cerulean, CyPet, AmCyanl, Midoriishi-Cyan), red fluorescent proteins (e.g., mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred), orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, Monomeric Examples of tags include glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, hemagglutinin (HA), nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, histidine (His), biotin carboxyl carrier protein (BCCP), and calmodulin.
[0204] The Cas protein can be provided in any form. For example, the Cas protein can be provided in the form of a protein. For example, the Cas protein can be provided as a Cas protein complexed with a gRNA. Alternatively, the Cas protein can be provided in the form of a nucleic acid encoding the Cas protein, such as RNA (e.g., messenger RNA (mRNA)) or DNA. Optionally, the nucleic acid encoding the Cas protein can be codon-optimized for efficient translation into protein in a particular cell or organism. For example, the nucleic acid encoding the Cas protein can be modified to replace codons with higher usage frequencies in bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest, compared to the naturally occurring polynucleotide sequence. For example, the nucleic acid encoding the Cas protein can be codon-optimized for expression in human cells. When the nucleic acid encoding the Cas protein is introduced into a cell, the Cas protein can be expressed in the cell transiently, conditionally, or constitutively.
[0205] Cas proteins provided as mRNA can be modified to improve stability and / or immunogenicity. Modifications can be made to one or more nucleosides within the mRNA. Examples of chemical modifications to mRNA nucleobases include pseudouridine, 1-methyl-pseudouridine, and 5-methyl-cytidine. For example, capped and polyadenylated Cas mRNA containing N1-methylpseudouridine can be used. Similarly, Cas mRNA can be modified by depleting uridines using synonymous codons.
[0206] The nucleic acid encoding the Cas protein can be stably integrated into the genome of the cell and operably linked to a promoter active within the cell. In one particular example, the nucleic acid encoding the Cas protein can comprise, consist essentially of, or consist of SEQ ID NO:22 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:22, and optionally, the nucleic acid encodes a protein comprising, consisting essentially of, or consisting of SEQ ID NO:21. In another particular example, the nucleic acid encoding a chimeric Cas protein comprising a nuclease-inactive Cas protein and one or more transcription activation domains can comprise, consist essentially of, or consist of SEQ ID NO:38 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:38, and optionally, the nucleic acid encodes a protein comprising, consisting essentially of, or consisting of SEQ ID NO:36. Alternatively, the nucleic acid encoding the Cas protein can be operably linked to a promoter in an expression construct. Expression constructs include any nucleic acid construct that can direct the expression of a gene or other nucleic acid sequence of interest (e.g., a Cas gene) and transfer such a nucleic acid sequence of interest into target cells. Promoters that can be used in expression constructs include, for example, promoters active in one or more of eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, pluripotent cells, embryonic stem (ES) cells, adult stem cells, developmentally restricted progenitor cells, induced pluripotent stem (iPS) cells, or one-cell stage embryos. Such promoters can be, for example, conditional promoters, inducible promoters, constitutive promoters, or tissue-specific promoters.
[0207] 3. Chimeric Adapter Proteins The SAM / tau biosensor cells disclosed herein can contain nucleic acids (DNA or RNA) encoding a chimeric Cas protein comprising a nuclease-inactive Cas protein fused to one or more transcription activation domains (e.g., VP64), as well as, optionally, a chimeric adaptor protein comprising an adaptor protein (e.g., MS2 coat protein (MCP)) fused to one or more transcription activation domains (e.g., fused to p65 and HSF1). Optionally, the chimeric Cas protein and / or chimeric adaptor protein are stably expressed. Optionally, the cells contain a genomically integrated chimeric Cas protein coding sequence and / or a genomically integrated chimeric adaptor protein coding sequence.
[0208] Such a chimeric adapter protein comprises (a) an adapter (i.e., an adapter domain or adapter protein) that specifically binds to an adapter-binding element in the guide RNA, and (b) one or more heterologous transcriptional activation domains. For example, such a fusion protein can comprise one, two, three, four, five, or more transcriptional activation domains (e.g., two or more heterologous transcriptional activation domains). In one example, such a chimeric adapter protein can comprise (a) an adapter (i.e., an adapter domain or adapter protein) that specifically binds to an adapter-binding element in the guide RNA, and (b) two or more transcriptional activation domains. For example, the chimeric adapter protein can comprise (a) an MS2 coat protein adapter that specifically binds to one or more MS2 aptamers in the guide RNA (e.g., two MS2 aptamers at separate locations in the guide RNA), and (b) one or more (e.g., two or more transcriptional activation domains). For example, the two transcriptional activation domains can be p65 and HSF1 transcriptional activation domains or functional fragments or variants thereof. However, chimeric adapter proteins in which the transcription activation domain comprises other transcription activation domains or functional fragments or variants thereof are also provided.
[0209] One or more transcription activation domains can be fused directly to the adapter. Alternatively, one or more transcription activation domains can be linked to the adapter via a linker or a combination of linkers, or via one or more additional domains. Similarly, when two or more transcription activation domains are present, they can be fused directly to each other, or linked to each other via a linker or a combination of linkers, or via one or more additional domains. Linkers that can be used in these fusion proteins can include any sequence that does not interfere with the function of the fusion protein. Exemplary linkers are short (e.g., 2-20 amino acids) and typically flexible (e.g., containing flexible amino acids such as glycine, alanine, and serine).
[0210] The one or more transcription activation domains and adapters can be in any order within the chimeric adapter protein. As an alternative, the one or more transcription activation domains can be at the C-terminus of the adapter, and the adapter can be at the N-terminus of the one or more transcription activation domains. For example, the one or more transcription activation domains can be at the C-terminus of the chimeric adapter protein, and the adapter can be at the N-terminus of the chimeric adapter protein. However, the one or more transcription activation domains can be at the C-terminus of the adapter without being at the C-terminus of the chimeric adapter protein (e.g., when a nuclear localization signal is at the C-terminus of the chimeric adapter protein). Similarly, the adapter can be at the N-terminus of one or more transcription activation domains without being at the N-terminus of the chimeric adapter protein (e.g., when a nuclear localization signal is at the N-terminus of the chimeric adapter protein). As another alternative, the one or more transcription activation domains can be at the N-terminus of the adapter, and the adapter can be at the C-terminus of one or more transcription activation domains. For example, the one or more transcription activation domains can be at the N-terminus of the chimeric adapter protein, and the adapter can be at the C-terminus of the chimeric adapter protein. As a further alternative, when a chimeric adapter protein comprises two or more transcription activation domains, the two or more transcription activation domains can flank the adapter.
[0211] The chimeric adapter protein can also be operably linked or fused to an additional heterologous polypeptide. The fused or linked heterologous polypeptide can be located at the N-terminus, C-terminus, or anywhere within the chimeric adapter protein. For example, the chimeric adapter protein can further include a nuclear localization signal. A specific example of such a protein includes an MS2 coat protein (adapter) linked (either directly or via an NLS) to the C-terminus of the p65 transcription activation domain of the MS2 coat protein (MCP) and to the C-terminus of the HSF1 transcription activation domain of the p65 transcription activation domain. Such a protein can include, from N-terminus to C-terminus, an MCP, a nuclear localization signal, a p65 transcription activation domain, and an HSF1 transcription activation domain. For example, a chimeric adapter protein can comprise, consist essentially of, or consist of an amino acid sequence that is at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the MCP-p65-HSF1 chimeric adapter protein sequence set forth in SEQ ID NO: 37. Similarly, a nucleic acid encoding a chimeric adapter protein can comprise, consist essentially of, or consist of a sequence that is at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the MCP-p65-HSF1 chimeric adapter protein coding sequence set forth in SEQ ID NO: 39.
[0212] An adaptor (i.e., an adaptor domain or adaptor protein) is a nucleic acid-binding domain (e.g., a DNA-binding domain and / or an RNA-binding domain) that specifically recognizes and binds to a distinct sequence (e.g., binds to a distinct DNA and / or RNA sequence, such as an aptamer, in a sequence-specific manner). Aptamers include nucleic acids that can bind to target molecules with high affinity and specificity through their ability to adopt a specific three-dimensional conformation. Such adaptors can, for example, bind to specific RNA sequences and secondary structures. These sequences (i.e., adaptor-binding elements) can be engineered into guide RNAs. For example, an MS2 aptamer can be engineered into a guide RNA to specifically bind to MS2 coat protein (MCP). For example, an adaptor can comprise, consist essentially of, or consist of an amino acid sequence that is at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the MCP sequence set forth in SEQ ID NO: 40. Similarly, the nucleic acid encoding the adapter can comprise, consist essentially of, or consist of an amino acid sequence that is at least 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the MCP-encoding sequence set forth in SEQ ID NO: 41. Specific examples of adapters and targets include, for example, RNA-binding protein / aptamer combinations present among the diversity of bacteriophage coat proteins. See, e.g., US2019-0284572 and WO2019 / 183123, each of which is incorporated herein by reference in its entirety for all purposes.
[0213] The chimeric adaptor protein disclosed herein comprises one or more transcription activation domains. Such transcription activation domains can be naturally occurring transcription activation domains, functional fragments or functional variants of naturally occurring transcription activation domains, or engineered or synthetic transcription activation domains. Transcription activation domains that can be used include, for example, those described in US2019-0284572 and WO2019 / 183123, each of which is incorporated herein by reference in its entirety for all purposes.
[0214] 4.Cell type The Cas / tau biosensor cells disclosed herein can be of any cell type and can be in vitro, ex vivo, or in vivo. The Cas / tau biosensor cell line or cell population can be a monoclonal cell line or cell population. Similarly, the SAM / tau biosensor cells disclosed herein can be of any cell type and can be in vitro, ex vivo, or in vivo. The SAM / tau biosensor cell line or cell population can be a monoclonal cell line or cell population. The cells can be from any source. For example, the cells can be eukaryotic cells, animal cells, plant cells, or fungal (e.g., yeast) cells. Such cells can be fish or bird cells, or such cells can be mammalian cells, such as human cells, non-human mammalian cells, rodent cells, mouse cells, or rat cells. Mammals include, for example, humans, non-human primates, monkeys, apes, cats, dogs, horses, bulls, deer, bison, sheep, rodents (e.g., mice, rats, hamsters, guinea pigs), livestock (e.g., bovine species such as cows and steers, ovine species such as sheep and goats, and porcine species such as pigs and wild boars). Birds include, for example, chickens, turkeys, ostriches, geese, and ducks. Livestock and agricultural animals are also included. The term "non-human animals" excludes humans. In certain examples, the Cas / tau biosensor cells are human cells (e.g., HEK293T cells). Similarly, in certain examples, the SAM / tau biosensor cells are human cells (e.g., HEK293T cells).
[0215] The cells may be, for example, totipotent or pluripotent cells (e.g., embryonic stem (ES) cells such as rodent ES cells, mouse ES cells, or rat ES cells). Totipotent cells include undifferentiated cells that can give rise to any cell type, while pluripotent cells include undifferentiated cells that have the ability to develop into multiple differentiated cell types. Such pluripotent and / or totipotent cells may be, for example, ES cells or ES-like cells, such as induced pluripotent stem (iPS) cells. ES cells include embryo-derived totipotent or pluripotent cells that can contribute to any tissue of the developing embryo when introduced into the embryo. ES cells may be derived from the inner cell mass of a blastocyst and can differentiate into cells of any of the three vertebrate germ layers (endoderm, ectoderm, and mesoderm).
[0216] The cells may be primary somatic cells or cells that are not primary somatic cells. Somatic cells may include any cells that are not gametes, germ cells, gamete cells, or undifferentiated stem cells. The cells may also be primary cells. Primary cells include cells or cell cultures isolated directly from an organism, organ, or tissue. Primary cells include cells that have not been transformed or immortalized. They include any cells obtained from an organism, organ, or tissue that have not previously been passaged in tissue culture, or that have previously been passaged in tissue culture but cannot be passaged indefinitely in tissue culture. Such cells can be isolated by conventional techniques and include, for example, somatic cells, hematopoietic cells, endothelial cells, epithelial cells, fibroblasts, mesenchymal cells, keratinocytes, melanocytes, monocytes, mononuclear cells, adipocytes, preadipocytes, neurons, glial cells, hepatocytes, skeletal myoblasts, and smooth muscle cells. For example, primary cells may be derived from connective tissue, muscle tissue, nervous system tissue, or epithelial tissue.
[0217] Such cells include those that do not normally proliferate indefinitely, but due to mutations or modifications, can avoid normal cell aging and continue to divide instead. Such mutations or changes can occur naturally or can be intentionally induced. Examples of immortalized cells include Chinese hamster ovary (CHO) cells, human embryonic kidney cells (e.g., HEK293T cells), and mouse embryonic fibroblasts (e.g., 3T3 cells). Many types of immortalized cells are well known. Immortalized or primary cells typically include cells used for culturing or expressing recombinant genes or proteins.
[0218] The cell can also be a differentiated cell, such as a neuronal cell (eg, a human neuronal cell).
[0219] B. Methods for generating Cas / tau and SAM / tau biosensor cells The Cas / tau biosensor cells disclosed herein can be generated by any known means. The first tau repeat domain linked to a first reporter, the second tau repeat domain linked to a second reporter, and the Cas protein can be introduced into cells in any form (e.g., DNA, RNA, or protein) by any known means. Similarly, the SAM / tau biosensor cells disclosed herein can be generated by any known means. The first tau repeat domain linked to a first reporter, the second tau repeat domain linked to a second reporter, the chimeric Cas protein, and the chimeric adaptor protein can be introduced into cells in any form (e.g., DNA, RNA, or protein) by any known means. "Introducing" includes presenting a nucleic acid or protein to a cell so that the sequence is accessible to the interior of the cell. The methods provided herein do not depend on a particular method for introducing a nucleic acid or protein into a cell, but only on the ability of the nucleic acid or protein to access the interior of at least one cell. Methods for introducing nucleic acids and proteins into various cell types are known and include, for example, stable transfection, transient transfection, and viral-mediated methods. Optionally, a targeting vector can be used.
[0220] Transfection protocols and protocols for introducing nucleic acids or proteins into cells can vary.Non-limiting transfection methods include chemical-based transfection methods using liposomes, nanoparticles, calcium phosphate (Graham et al. (1973) Virology 52(2):456-67, Bacchetti et al. (1977) Proc. Natl. Acad. Sci. USA 74(4):1590-4, and Kriegler, M (1991). Transfer and Expression: A Laboratory Manual. New York: W.H. Freeman and Company. pp.96-97), dendrimers, or cationic polymers such as DEAE-dextran or polyethyleneimine.Non-chemical methods include electroporation, sonoporation, and phototransfection.Particle-based transfection includes the use of a gene gun or magnet-assisted transfection (Bertram (2006) Current Pharmaceutical Biotechnology 7,277-28). Viral methods can also be used for transfection.
[0221] Introduction of nucleic acids or proteins into cells can also be mediated by electroporation, intracytoplasmic injection, viral infection, adenovirus, adeno-associated virus, lentivirus, retrovirus, transfection, lipid-mediated transfection, or nucleofection. Nucleofection is an improved electroporation technique that allows nucleic acid substrates to be delivered not only to the cytoplasm but also through the nuclear membrane to the nucleus. Furthermore, the use of nucleofection in the methods disclosed herein typically requires far fewer cells than conventional electroporation (e.g., only about 2 million cells compared to 7 million cells required for conventional electroporation). In one example, nucleofection is performed using the LONZA® NUCLEOFECTOR™ system.
[0222] The introduction of nucleic acid or protein into cells can also be achieved by microinjection.The microinjection of mRNA is preferably into the cytoplasm (for example, to directly deliver mRNA to the translational machinery), while the microinjection of protein or DNA encoding protein is preferably into the nucleus.Alternatively, microinjection can be performed by injection into both the nucleus and the cytoplasm, and the needle can be first introduced into the nucleus and the first amount can be injected, and the second amount can be injected into the cytoplasm while the needle is removed from the cell.Methods for performing microinjection are well known. See, for example, Nagy et al. (Nagy A, Gertsenstein M, Wintersten K, Behringer R., 2003, Manipulating the Mouse Embryo. Cold Spring Harbor, New York: Cold Spring Harbor Laboratory Press), Meyer et al. (2010) Proc. Natl. Acad. Sci. USA 107:15022-15026, and Meyer et al. (2012) Proc. Natl. Acad. Sci. USA 109:9354-9359.
[0223] Other methods for introducing nucleic acids or proteins into cells can include, for example, vector delivery, particle-mediated delivery, exosome-mediated delivery, lipid nanoparticle-mediated delivery, cell-penetrating peptide-mediated delivery, or implantable device-mediated delivery. Methods for administering nucleic acids or proteins to a subject to modify cells in vivo are disclosed elsewhere herein.
[0224] In one example, a first tau repeat domain linked to a first reporter, a second tau repeat domain linked to a second reporter, and a Cas protein can be introduced via viral transduction, such as lentiviral transduction.
[0225] Screening of cells containing a first tau repeat domain linked to a first reporter, a second tau repeat domain linked to a second reporter, and a Cas protein can be done by any known means.
[0226] As an example, reporter genes can be used to screen cells with a Cas protein, a first tau repeat domain linked to a first reporter, or a second tau repeat domain linked to a second reporter. Exemplary reporter genes include those encoding luciferase, β-galactosidase, green fluorescent protein (GFP), enhanced green fluorescent protein (eGFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), blue fluorescent protein (BFP), enhanced blue fluorescent protein (eBFP), DsRed, ZsGreen, MmGFP, mPlum, mCherry, tdTomato, mStrawberry, J-Red, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, Cerulean, T-Sapphire, and alkaline phosphatase. For example, if the first and second reporters are fluorescent proteins (e.g., CFP and YFP), cells containing these reporters can be selected by flow cytometry to select double-positive cells. The double-positive cells can then be combined to generate polyclonal lines, or monoclonal lines can be generated from a single double-positive cell.
[0227] As another example, a selectable marker can be used to screen for cells with a Cas protein, a first tau repeat domain linked to a first reporter, or a second tau repeat domain linked to a second reporter. Exemplary selectable markers include neomycin phosphotransferase (neor), hygromycin B phosphotransferase (hygr), puromycin-N-acetyltransferase (puror), blasticidin S deaminase (bsrr), xanthine / guanine phosphoribosyltransferase (gpt), or herpes simplex virus thymidine kinase (HSV-k). Another exemplary selectable marker is the bleomycin resistance protein encoded by the Sh ble gene (Streptoalloteichus hindustanus bleomycin gene), which confers resistance to zeocin (phleomycin D1).
[0228] Aggregation-positive (Agg[+]) cells, in which the tau repeat domains stably exist in an aggregated state, meaning that the tau repeat domains stably persist in all cells and undergo proliferation and multiple passages over time, can be generated, for example, by seeding with tau aggregates. For example, naive aggregation-negative (Agg[-]) Cas / tau biosensor cells disclosed herein can be treated with recombinant fibrillar tau (e.g., recombinant fibrillar tau repeat domains) to seed aggregates of the tau repeat domain protein stably expressed by these cells. Similarly, naive aggregation-negative (Agg[-]) SAM / tau biosensor cells disclosed herein can be treated with recombinant fibrillar tau (e.g., recombinant fibrillar tau repeat domains) to seed aggregates of the tau repeat domain protein stably expressed by these cells. The fibrillar tau repeat domain can be the same as, similar to, or different from the tau repeat domain stably expressed by the cells. Optionally, recombinant fibrillar tau can be mixed with Lipofectamine reagent. The plated cells can then be serially diluted to obtain single-cell-derived clones, allowing for the identification of clonal cell lines in which tau repeat domain aggregates stably persist in all cells and are expanded over time and passaged multiple times.
[0229] As another example, aggregation-positive (Agg[+]) cells, in which the tau repeat domains are stably present in an aggregated state, meaning that the tau repeat domains stably persist in all cells and undergo proliferation and multiple passages over time, can be generated, for example, by seeding cells (e.g., tau aggregation-negative cells) with cell lysate from tau aggregation-positive cells. This is the "maximum seeding" described in the Examples herein. For example, cells can be seeded using medium containing the cell lysate (e.g., fresh medium containing the cell lysate). "Maximum seeding" can refer to seeding that, by itself, induces tau aggregation in the majority of aggregation-negative tau biosensor cells. "Minimum seeding" can refer to seeding that, by itself, is insufficient to induce tau aggregation in aggregation-negative tau biosensor cells (or only minimally induces tau aggregation), but sensitizes such cells to the induction of aggregation.
[0230] The amount or concentration of the cell lysate in the medium can be any suitable amount or concentration. For example, the concentration of the cell lysate in the medium (e.g., fresh culture medium) can be about 0.1 μg / mL to about 50 μg / mL, about 0.1 μg / mL to about 25 μg / mL, about 0.1 μg / mL to about 10 μg / mL, about 0.1 μg / mL to about 5 μg / mL, about 0.1 μg / mL to about 4.5 μg / mL, about 0.1 μg / mL to about 4 μg / mL, about 0.1 μg / mL to about 3.5 μg / mL, about 0.1 μg / mL to about 3 μg / mL, about 0.1 μg / mL to about 2.5 μg / mL, about 0.1 μg / mL to about 2 μg / mL, about 0.1 μg / mL to about 1 μg / mL. .5μg / mL, about 0.1μg / mL to about 1μg / mL, about 0.5μg / mL to about 50μg / mL, about 0.5μg / mL to about 25μg / mL, about 0.5μg / mL to about 10μg / mL, about 0.5μg / mL to about 5μg / mL, about 0.5μg / mL to about 4.5μg / mL mL, approximately 0.5 μg / mL to approximately 4 μg / mL, approximately 0.5 μg / mL to approximately 3.5 μg / mL, approximately 0.5 μg / mL to approximately 3 μg / mL, approximately 0.5 μg / mL to approximately 2.5 μg / mL, approximately 0.5 μg / mL to approximately 2 μg / mL, approximately 0.5 μg / mL to approximately 1.5 μg / mL, approximately 0.5μg / mL to approximately 1μg / mL, approximately 1μg / mL to approximately 50μg / mL, approximately 1μg / mL to approximately 25μg / mL, approximately 1μg / mL to approximately 10μg / mL, approximately 1μg / mL to approximately 5μg / mL, approximately 1μg / mL to approximately 4.5μg / mL, approximately 1μg / mL to approximately 4μg / mL , about 1μg / mL to about 3.5μg / mL, about 1μg / mL to about 3μg / mL, about 1μg / mL to about 2.5μg / mL, about 1μg / mL to about 2μg / mL, about 1μg / mL to about 1.5μg / mL, about 1.5μg / mL to about 50μg / mL, about 1.5μg / mL to about 2 5μg / mL, approximately 1.5μg / mL to approximately 10μg / mL, approximately 1.5μg / mL to approximately 5μg / mL, approximately 1.5μg / mL to approximately 4.5μg / mL, approximately 1.5μg / mL to approximately 4μg / mL, approximately 1.5μg / mL to approximately 3.5μg / mL, approximately 1.5μg / mL to approximately 3μg / m L, about 1.5μg / mL to about 2.5μg / mL, about 1.5μg / mL to about 2μg / mL, about 2μg / mL to about 50μg / mL, about 2μg / mL to about 25μg / mL, about 2μg / mL to about 10μg / mL, about 2μg / mL to about 5μg / mL, about 2μg / mL to about 4.The concentration may be 5 μg / mL, about 2 μg / mL to about 4 μg / mL, about 2 μg / mL to about 3.5 μg / mL, about 2 μg / mL to about 3 μg / mL, about 2 μg / mL to about 2.5 μg / mL, about 2.5 μg / mL to about 50 μg / mL, about 2.5 μg / mL to about 25 μg / mL, about 2.5 μg / mL to about 10 μg / mL, about 2.5 μg / mL to about 5 μg / mL, about 2.5 μg / mL to about 4.5 μg / mL, about 2.5 μg / mL to about 4 μg / mL, about 2.5 μg / mL to about 3.5 μg / mL, or about 2.5 μg / mL to about 3 μg / mL. For example, the cell lysate in the culture medium may be at a concentration of about 1 μg / mL to about 5 μg / mL, or at a concentration of about 1.5 μg / mL, about 2 μg / mL, about 2.5 μg / mL, about 3 μg / mL, about 3.5 μg / mL, about 4 μg / mL, about 4.5 μg / mL, or about 5 μg / mL. Optionally, the cell lysate may be in a buffer solution such as phosphate-buffered saline. Optionally, the buffer solution may contain protease inhibitors. Examples of protease inhibitors include, but are not limited to, AEBSF, aprotinin, bestatin, E-64, leupeptin, pepstatin A, and ethylenediaminetetraacetic acid (EDTA). The buffer solution may contain any one or any combination of these inhibitors (e.g., the buffer solution may contain all of these protease inhibitors).
[0231] The cells to generate a lysate can be collected in a buffer solution such as phosphate-buffered saline. Optionally, the buffer solution can contain a protease inhibitor. Examples of protease inhibitors include, but are not limited to, AEBSF, aprotinin, bestatin, E-64, leupeptin, pepstatin A, and ethylenediaminetetraacetic acid (EDTA). The buffer solution can contain any of these inhibitors or any combination thereof (e.g., the buffer solution can contain all of these protease inhibitors).
[0232] Cell lysates can be collected, for example, by sonicating tau aggregate-positive cells (e.g., cells collected in buffer and protease inhibitors as described above) for any suitable time. For example, cells can be sonicated for about 1 to about 6 minutes, about 1 to about 5 minutes, about 1 to about 4 minutes, about 1 to about 3 minutes, about 2 to about 6 minutes, about 2 to about 5 minutes, about 2 to about 4 minutes, about 2 to about 3 minutes, about 2 to about 6 minutes, about 3 to about 5 minutes, or about 3 to about 4 minutes. For example, cells can be sonicated for about 2 to about 4 minutes, or about 3 minutes.
[0233] Optionally, the medium contains lipofectamine or liposomes (e.g., cationic liposomes) or phospholipids or another transfection agent. Optionally, the medium contains lipofectamine. Optionally, the medium does not contain lipofectamine or liposomes (e.g., cationic liposomes) or phospholipids or another transfection agent. Optionally, the medium does not contain lipofectamine. The amount or concentration of lipofectamine or liposomes (e.g., cationic liposomes) or phospholipids or other transfection agent in the medium can be any suitable amount or concentration. For example, the concentration of lipofectamine or liposomes (e.g., cationic liposomes) or phospholipids or other transfection agents in the medium may be about 0.5 μL / mL to about 10 μL / mL, about 0.5 μL / mL to about 5 μL / mL, about 0.5 μL / mL to about 4.5 μL / mL, about 0.5 μL / mL to about 4 μL / mL, about 0.5 μL / mL to about 3.5 μL / mL of medium (e.g., fresh medium). L, approximately 0.5 μL / mL to approximately 3 μL / mL, approximately 0.5 μL / mL to approximately 2.5 μL / mL, approximately 0.5 μL / mL to approximately 2 μL / mL, approximately 0.5 μL / mL to approximately 1.5 μL / mL, approximately 0.5 μL / mL Approximately 1μL / mL, approximately 1μL / mL to approximately 10μL / mL, approximately 1μL / mL to approximately 5μL / mL, approximately 1μL / mL to approximately 4.5μL / mL, approximately 1μL / mL to approximately 4μL / mL, approximately 1μL / mL to approximately 3.5μ L / mL, approximately 1μL / mL to approximately 3μL / mL, approximately 1μL / mL to approximately 2.5μL / mL, approximately 1μL / mL to approximately 2μL / mL, approximately 1μL / mL to approximately 1.5μL / mL, approximately 1.5μL / mL to approximately 10μ L / mL, about 1.5 μL / mL to about 5 μL / mL, about 1.5 μL / mL to about 4.5 μL / mL, about 1.5 μL / mL to about 4 μL / mL, about 1.5 μL / mL to about 3.5 μL / mL, about 1.5 μL / mL mL ~ approx. 3 μL / mL, approx. 1.5 μL / mL ~ approx. 2.5 μL / mL, approx. 1.5 μL / mL ~ approx. 2 μL / mL, approx. 2 μL / mL ~ approx. 10 μL / mL, approx. 2 μL / mL ~ approx. 5 μL / mL, approx. 2 μL / m L can be about 4.5 μL / mL, about 2 μL / mL to about 4 μL / mL, about 2 μL / mL to about 3.5 μL / mL, about 2 μL / mL to about 3 μL / mL, or about 2 μL / mL to about 2.5 μL / mL.For example, the concentration of lipofectamine or liposomes (e.g., cationic liposomes) or phospholipids or other transfection agents in the medium can be about 1.5 μL / mL to about 4 μL / mL, or it can be about 1.5 μL / mL, about 2 μL / mL, about 2.5 μL / mL, about 3 μL / mL, about 3.5 μL / mL, or about 4 μL / mL.
[0234] Intercellular proliferation of tau can also result from tau aggregation activity secreted by aggregate-containing cells. For example, Agg[+] cells, or cells sensitized to become Agg[+] cells (e.g., sensitized to tau seeding or tau aggregation activity), can be generated by co-culturing Agg[-]Cas / tau biosensor cells with Agg[+] cells. Similarly, Agg[+] cells, or cells sensitized to become Agg[+] cells (e.g., sensitized to tau seeding or tau aggregation activity), can be generated by co-culturing Agg[-]SAM / tau biosensor cells with Agg[+] cells.
[0235] Agg[+] cells, or cells sensitized to becoming Agg[+] cells (e.g., sensitized to tau seeding or tau aggregation activity), can also be generated using conditioned medium collected from cultured tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state, as described herein. This is the "minimal seeding" method disclosed in the Examples herein. Conditioned medium refers to spent medium collected from cultured cells. It contains metabolites, growth factors, and extracellular matrix proteins secreted into the medium by the cultured cells. The use of conditioned medium does not involve co-culture with Agg[+] cells (i.e., naive Agg[-] cells are not co-cultured with Agg[+] cells). For example, conditioned medium can be generated by collecting medium from confluent Agg[+] cells. The medium can be left on the confluent Agg[+] cells for about 12 hours, about 24 hours, about 2 days, about 3 days, about 4 days, about 5 days, about 6 days, about 7 days, about 8 days, about 9 days, or about 10 days. For example, the medium can be left on the confluent Agg[+] cells for about 1 to about 7 days, about 2 to about 6 days, about 3 to about 5 days, or about 4 days. The conditioned medium can then be combined with fresh medium and applied to naive (Agg[-]) Cas / tau biosensor cells. Similarly, the conditioned medium can then be combined with fresh medium and applied to naive (Agg[-]) SAM / tau biosensor cells. The ratio of conditioned medium to fresh medium can be, for example, about 10:1, about 9:1, about 8:1, about 7:1, about 6:1, about 5:1, about 4:1, about 3:1, about 2:1, about 1:1, about 1:2, about 1:3, about 1:4, about 1:5, about 1:6, about 1:7, about 1:8, about 1:9, or about 1: 10. For example, the ratio of conditioned medium to fresh medium can be about 5:1 to about 1:1, about 4:1 to about 2:1, or about 3:1.For example, it may be about 90% conditioned medium and about 10% fresh medium, about 85% conditioned medium and about 15% fresh medium, about 80% conditioned medium and about 20% fresh medium, about 75% conditioned medium and about 25% fresh medium, about 70% conditioned medium and about 30% fresh medium, about 65% conditioned medium and about 35% fresh medium, about 60% conditioned medium and about 40% fresh medium, about 55% conditioned medium and about 45% fresh medium, about 50% conditioned medium and about 50% fresh medium The method may include culturing the genetically modified cell population in about 45% conditioned medium and about 55% fresh medium, about 40% conditioned medium and about 60% fresh medium, about 35% conditioned medium and about 65% fresh medium, about 30% conditioned medium and about 70% fresh medium, about 25% conditioned medium and about 75% fresh medium, about 20% conditioned medium and about 80% fresh medium, about 15% conditioned medium and about 85% fresh medium, or about 10% conditioned medium and about 90% fresh medium. In one example, it may include culturing the genetically modified cell population in medium containing at least about 50% conditioned medium and about 50% or less fresh medium. In a particular example, it may include culturing the genetically modified cell population in about 75% conditioned medium and about 25% fresh medium. Optionally, the conditioned medium is applied to naive Agg[-] cells without lipofectamine, without liposomes (e.g., cationic liposomes), or without phospholipids. Optionally, the genetically modified cell population is not co-cultured with tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state.
[0236] Conditioned medium without co-culture has not previously been used as a seeding agent in this context. However, because in vitro-generated tau fibrils are a limited source, conditioned medium is particularly useful for large-scale genome-wide screening. Furthermore, conditioned medium is more physiologically relevant because it is produced and secreted by cells rather than in vitro. Use of conditioned medium as described herein provides a boost in tau seeding activity (e.g., approximately 0.1%, as measured by FRET induction as disclosed elsewhere herein) to sensitize cells to tau aggregation.
[0237] C. In vitro cultures and conditioned media Also disclosed herein are in vitro cultures or compositions comprising the Cas / tau biosensor cells disclosed herein and media for culturing those cells. Also disclosed herein are in vitro cultures or compositions comprising the SAM / tau biosensor cells disclosed herein and media for culturing those cells. The cells can be Agg[-] cells or Agg[+] cells. For example, the culture or composition can include Agg[-] cells. In one example, the media includes conditioned media from Agg[+] cells, as disclosed elsewhere herein. Optionally, the cells in the culture or composition are Agg[-] cells and have not been co-cultured with Agg[+] cells. The media can include a mixture of conditioned media and fresh media. For example, the ratio of conditioned medium to fresh medium can be, for example, about 10:1, about 9:1, about 8:1, about 7:1, about 6:1, about 5:1, about 4:1, about 3:1, about 2:1, about 1:1, about 1:2, about 1:3, about 1:4, about 1:5, about 1:6, about 1:7, about 1:8, about 1:9, or about 1:10. For example, the ratio of conditioned medium to fresh medium can be about 5:1 to about 1:1, about 4:1 to about 2:1, or about 3:1. For example, it may be about 90% conditioned medium and about 10% fresh medium, about 85% conditioned medium and about 15% fresh medium, about 80% conditioned medium and about 20% fresh medium, about 75% conditioned medium and about 25% fresh medium, about 70% conditioned medium and about 30% fresh medium, about 65% conditioned medium and about 35% fresh medium, about 60% conditioned medium and about 40% fresh medium, about 55% conditioned medium and about 45% fresh medium, about 50% conditioned medium and about 50% fresh medium The method may include culturing the genetically modified cell population in about 45% conditioned medium and about 55% fresh medium, about 40% conditioned medium and about 60% fresh medium, about 35% conditioned medium and about 65% fresh medium, about 30% conditioned medium and about 70% fresh medium, about 25% conditioned medium and about 75% fresh medium, about 20% conditioned medium and about 80% fresh medium, about 15% conditioned medium and about 85% fresh medium, or about 10% conditioned medium and about 90% fresh medium.In one example, it may involve culturing the genetically modified cell population in a medium containing at least about 50% conditioned medium and no more than about 50% fresh medium. In a particular example, it may involve culturing the genetically modified cell population in about 75% conditioned medium and about 25% fresh medium. Optionally, the medium contains lipofectamine, liposomes (e.g., cationic liposomes), or phospholipids. Optionally, the medium does not contain lipofectamine, liposomes (e.g., cationic liposomes), or phospholipids. Optionally, the medium does not contain lipofectamine.
[0238] D. In vitro cultures and media containing lysates from tau aggregate-positive cells Also disclosed herein are in vitro cultures or compositions comprising the Cas / tau biosensor cells disclosed herein and media for culturing those cells. Also disclosed herein are in vitro cultures or compositions comprising the SAM / tau biosensor cells disclosed herein and media for culturing those cells. In one example, the media contains cell lysate from cultured tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state. The cells can be Agg[-] cells or Agg[+] cells. For example, the culture or composition can contain Agg[-] cells. Optionally, the cells in the culture or composition are Agg[-] cells and are not co-cultured with Agg[+] cells. The media can contain a mixture of fresh media and cell lysate.
[0239] The amount or concentration of the cell lysate in the medium can be any suitable amount or concentration. For example, the concentration of the cell lysate in the medium (e.g., fresh culture medium) can be about 0.1 μg / mL to about 50 μg / mL, about 0.1 μg / mL to about 25 μg / mL, about 0.1 μg / mL to about 10 μg / mL, about 0.1 μg / mL to about 5 μg / mL, about 0.1 μg / mL to about 4.5 μg / mL, about 0.1 μg / mL to about 4 μg / mL, about 0.1 μg / mL to about 3.5 μg / mL, about 0.1 μg / mL to about 3 μg / mL, about 0.1 μg / mL to about 2.5 μg / mL, about 0.1 μg / mL to about 2 μg / mL, about 0.1 μg / mL to about 1 μg / mL. .5μg / mL, about 0.1μg / mL to about 1μg / mL, about 0.5μg / mL to about 50μg / mL, about 0.5μg / mL to about 25μg / mL, about 0.5μg / mL to about 10μg / mL, about 0.5μg / mL to about 5μg / mL, about 0.5μg / mL to about 4.5μg / mL mL, approximately 0.5 μg / mL to approximately 4 μg / mL, approximately 0.5 μg / mL to approximately 3.5 μg / mL, approximately 0.5 μg / mL to approximately 3 μg / mL, approximately 0.5 μg / mL to approximately 2.5 μg / mL, approximately 0.5 μg / mL to approximately 2 μg / mL, approximately 0.5 μg / mL to approximately 1.5 μg / mL, approximately 0.5μg / mL to approximately 1μg / mL, approximately 1μg / mL to approximately 50μg / mL, approximately 1μg / mL to approximately 25μg / mL, approximately 1μg / mL to approximately 10μg / mL, approximately 1μg / mL to approximately 5μg / mL, approximately 1μg / mL to approximately 4.5μg / mL, approximately 1μg / mL to approximately 4μg / mL , about 1μg / mL to about 3.5μg / mL, about 1μg / mL to about 3μg / mL, about 1μg / mL to about 2.5μg / mL, about 1μg / mL to about 2μg / mL, about 1μg / mL to about 1.5μg / mL, about 1.5μg / mL to about 50μg / mL, about 1.5μg / mL to about 2 5μg / mL, approximately 1.5μg / mL to approximately 10μg / mL, approximately 1.5μg / mL to approximately 5μg / mL, approximately 1.5μg / mL to approximately 4.5μg / mL, approximately 1.5μg / mL to approximately 4μg / mL, approximately 1.5μg / mL to approximately 3.5μg / mL, approximately 1.5μg / mL to approximately 3μg / m L, about 1.5μg / mL to about 2.5μg / mL, about 1.5μg / mL to about 2μg / mL, about 2μg / mL to about 50μg / mL, about 2μg / mL to about 25μg / mL, about 2μg / mL to about 10μg / mL, about 2μg / mL to about 5μg / mL, about 2μg / mL to about 4.The concentration may be 5 μg / mL, about 2 μg / mL to about 4 μg / mL, about 2 μg / mL to about 3.5 μg / mL, about 2 μg / mL to about 3 μg / mL, about 2 μg / mL to about 2.5 μg / mL, about 2.5 μg / mL to about 50 μg / mL, about 2.5 μg / mL to about 25 μg / mL, about 2.5 μg / mL to about 10 μg / mL, about 2.5 μg / mL to about 5 μg / mL, about 2.5 μg / mL to about 4.5 μg / mL, about 2.5 μg / mL to about 4 μg / mL, about 2.5 μg / mL to about 3.5 μg / mL, or about 2.5 μg / mL to about 3 μg / mL. For example, the cell lysate in the culture medium may be at a concentration of about 1 μg / mL to about 5 μg / mL, or at a concentration of about 1.5 μg / mL, about 2 μg / mL, about 2.5 μg / mL, about 3 μg / mL, about 3.5 μg / mL, about 4 μg / mL, about 4.5 μg / mL, or about 5 μg / mL. Optionally, the cell lysate may be in a buffer solution such as phosphate-buffered saline. Optionally, the buffer solution may contain protease inhibitors. Examples of protease inhibitors include, but are not limited to, AEBSF, aprotinin, bestatin, E-64, leupeptin, pepstatin A, and ethylenediaminetetraacetic acid (EDTA). The buffer solution may contain any one or any combination of these inhibitors (e.g., the buffer solution may contain all of these protease inhibitors).
[0240] The cells to generate a lysate can be collected in a buffer solution such as phosphate-buffered saline. Optionally, the buffer solution can contain a protease inhibitor. Examples of protease inhibitors include, but are not limited to, AEBSF, aprotinin, bestatin, E-64, leupeptin, pepstatin A, and ethylenediaminetetraacetic acid (EDTA). The buffer solution can contain any of these inhibitors or any combination thereof (e.g., the buffer solution can contain all of these protease inhibitors).
[0241] Cell lysates can be collected, for example, by sonicating tau aggregate-positive cells (e.g., cells collected in buffer and protease inhibitors as described above) for any suitable time. For example, cells can be sonicated for about 1 to about 6 minutes, about 1 to about 5 minutes, about 1 to about 4 minutes, about 1 to about 3 minutes, about 2 to about 6 minutes, about 2 to about 5 minutes, about 2 to about 4 minutes, about 2 to about 3 minutes, about 2 to about 6 minutes, about 3 to about 5 minutes, or about 3 to about 4 minutes. For example, cells can be sonicated for about 2 to about 4 minutes, or about 3 minutes.
[0242] Optionally, the medium contains lipofectamine or liposomes (e.g., cationic liposomes) or phospholipids or another transfection agent. Optionally, the medium contains lipofectamine. Optionally, the medium does not contain lipofectamine or liposomes (e.g., cationic liposomes) or phospholipids or another transfection agent. Optionally, the medium does not contain lipofectamine. The amount or concentration of lipofectamine or liposomes (e.g., cationic liposomes) or phospholipids or other transfection agent in the medium can be any suitable amount or concentration. For example, the concentration of lipofectamine or liposomes (e.g., cationic liposomes) or phospholipids or other transfection agents in the medium may be about 0.5 μL / mL to about 10 μL / mL, about 0.5 μL / mL to about 5 μL / mL, about 0.5 μL / mL to about 4.5 μL / mL, about 0.5 μL / mL to about 4 μL / mL, about 0.5 μL / mL to about 3.5 μL / mL of medium (e.g., fresh medium). L, approximately 0.5 μL / mL to approximately 3 μL / mL, approximately 0.5 μL / mL to approximately 2.5 μL / mL, approximately 0.5 μL / mL to approximately 2 μL / mL, approximately 0.5 μL / mL to approximately 1.5 μL / mL, approximately 0.5 μL / mL Approximately 1μL / mL, approximately 1μL / mL to approximately 10μL / mL, approximately 1μL / mL to approximately 5μL / mL, approximately 1μL / mL to approximately 4.5μL / mL, approximately 1μL / mL to approximately 4μL / mL, approximately 1μL / mL to approximately 3.5μ L / mL, approximately 1μL / mL to approximately 3μL / mL, approximately 1μL / mL to approximately 2.5μL / mL, approximately 1μL / mL to approximately 2μL / mL, approximately 1μL / mL to approximately 1.5μL / mL, approximately 1.5μL / mL to approximately 10μ L / mL, about 1.5 μL / mL to about 5 μL / mL, about 1.5 μL / mL to about 4.5 μL / mL, about 1.5 μL / mL to about 4 μL / mL, about 1.5 μL / mL to about 3.5 μL / mL, about 1.5 μL / mL mL ~ approx. 3 μL / mL, approx. 1.5 μL / mL ~ approx. 2.5 μL / mL, approx. 1.5 μL / mL ~ approx. 2 μL / mL, approx. 2 μL / mL ~ approx. 10 μL / mL, approx. 2 μL / mL ~ approx. 5 μL / mL, approx. 2 μL / m L can be about 4.5 μL / mL, about 2 μL / mL to about 4 μL / mL, about 2 μL / mL to about 3.5 μL / mL, about 2 μL / mL to about 3 μL / mL, or about 2 μL / mL to about 2.5 μL / mL.For example, the concentration of lipofectamine or liposomes (e.g., cationic liposomes) or phospholipids or other transfection agents in the medium can be about 1.5 μL / mL to about 4 μL / mL, or it can be about 1.5 μL / mL, about 2 μL / mL, about 2.5 μL / mL, about 3 μL / mL, about 3.5 μL / mL, or about 4 μL / mL.
[0243] III. Guide RNA Knockout Library The CRISPRn screening method disclosed herein utilizes a CRISPR guide RNA (gRNA) knockout library, such as a genome-wide gRNA knockout library. Cas nucleases, such as Cas9, can be programmed to induce double-strand breaks at specific genomic loci via gRNAs designed to target specific target sequences. The targeting specificity of the Cas protein is conferred by short gRNAs, enabling pooled genome-scale functional screening. Such libraries offer several advantages over libraries such as shRNA libraries, which reduce protein expression by targeting mRNA. In contrast, gRNA libraries achieve knockout via frameshift mutations introduced into the genomic coding region of genes.
[0244] The CRISPRa screening method disclosed herein utilizes a CRISPR guide RNA (gRNA) transcription activation library, such as a genome-wide gRNA transcription activation library. The SAM system can be programmed to activate transcription of genes at specific genomic loci via gRNAs designed to target specific target sequences. Because the targeting specificity of the Cas protein is conferred by short gRNAs, pooled genome-scale functional screening is possible.
[0245] The gRNA of the library can target any number of genes.For example, the gRNA can target about 50 or more genes, about 100 or more genes, about 200 or more genes, about 300 or more genes, about 400 or more genes, about 500 or more genes, about 1000 or more genes, about 2000 or more genes, about 3000 or more genes, about 4000 or more genes, about 5000 or more genes, about 10000 or more genes, or about 20000 or more genes.In some libraries, the gRNA can be selected to target genes of a specific signal transduction pathway.Some libraries are genome-wide libraries.
[0246] Genome-wide library comprises one or more gRNAs (e.g., sgRNAs) that target each gene in genome.The genome that is targeted can be any type of genome.For example, genome can be eukaryotic genome, mammalian genome, non-human mammalian genome, rodent genome, mouse genome, rat genome, or human genome.In one example, the genome that is targeted is human genome.
[0247] The gRNA can target any number of sequences within each individual targeted gene. In some libraries, multiple target sequences are targeted, on average, in each of the multiple targeted genes. For example, an average of about 2 to about 10, about 2 to about 9, about 2 to about 8, about 2 to about 7, about 2 to about 6, about 2 to about 5, about 2 to about 4, or about 2 to about 3 unique target sequences may be targeted in each of the multiple targeted genes. For example, an average of at least about 2, at least about 3, at least about 4, at least about 5, or at least about 6 unique target sequences may be targeted in each of the multiple targeted genes. As a specific example, an average of about 6 target sequences may be targeted in each of the multiple targeted genes. As another specific example, an average of about 3 to about 6 or about 4 to about 6 target sequences may be targeted in each of the multiple targeted genes.
[0248] For example, libraries can target genes with an average coverage of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 gRNAs per gene. In particular examples, libraries can target genes with an average coverage of about 3-4 gRNAs per gene, or about 6 gRNAs per gene.
[0249] The gRNA can be targeted to any desired location within a target gene. CRISPRn gRNAs can be designed to target the coding region of a gene so that cleavage by the corresponding Cas protein results in a frameshift insertion / deletion (indel) mutation, resulting in a loss-of-function allele. More specifically, frameshift mutations can be achieved by targeted DNA double-strand breaks followed by mutagenic repair via the non-homologous end joining (NHEJ) pathway, which generates indels at the break site. The indels introduced into the DSB are random, and some indels result in frameshift mutations that cause premature termination of the gene.
[0250] In some CRISPRn libraries, each gRNA targets a constitutive exon, if possible. In some CRISPRn libraries, each gRNA targets a 5' constitutive exon, if possible. In some methods, each gRNA targets the first exon, second exon, or third exon (from the 5' end of the gene), if possible.
[0251] As an example, gRNAs in a CRISPRn library can target constitutive exons. Constitutive exons are exons that are consistently conserved after splicing. Exons expressed across all tissues can be considered constitutive exons for gRNA targeting. gRNAs in the library can target constitutive exons near the 5' end of each gene. Optionally, the first and last exons of each gene can be excluded as potential targets. Optionally, any exons containing alternative splice sites can be excluded as potential targets. Optionally, the earliest two exons that meet the above criteria can be selected as potential targets. Optionally, exons 2 and 3 can be selected as potential targets (e.g., if no constitutive exons are identified). Furthermore, gRNAs in the library can be selected and designed to minimize off-target effects.
[0252] In a specific example, a genome-wide CRISPRn gRNA library contains sgRNAs targeting the 5' constitutive exons of >18,000 genes in the human genome, with an average coverage of approximately 6 sgRNAs per gene, and each target site selected to minimize off-target modifications.
[0253] CRISPRa gRNAs can be designed to target sequences adjacent to the transcription start site of a gene. For example, target sequences can be 1000, 900, 800, 700, 600, 500, 400, 300, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, or 1 base pair within the transcription start site. For example, each gRNA in a CRISPRa library can target a sequence within 200 bp upstream of the transcription start site. Optionally, the target sequence is within the region 200 base pairs upstream of the transcription start site and 1 base pair downstream (-200 to +1) of the transcription start site.
[0254] The gRNAs of the genome-wide library can be in any form. For example, the gRNA library can be packaged in a viral vector, such as a retroviral vector, a lentiviral vector, or an adenoviral vector. In a specific example, the gRNA library is packaged in a lentiviral vector. The vector may further include a reporter gene or selection marker to facilitate selection of cells that receive the vector. Examples of such reporter genes and selection markers are disclosed elsewhere herein. As an example, the selection marker may confer resistance to drugs such as neomycin phosphotransferase, hygromycin B phosphotransferase, puromycin-N-acetyltransferase, and blasticidin S deaminase. Another exemplary selection marker is the bleomycin resistance protein encoded by the Sh ble gene (Streptoalloteichus hindustanus bleomycin gene), which confers resistance to zeocin (phleomycin D1). For example, cells can be selected with a drug (e.g., puromycin) so that only cells transduced with the guide RNA construct are preserved for use in performing screening.
[0255] A. Guide RNA A "guide RNA" or "gRNA" is an RNA molecule that binds to a Cas protein (e.g., a Cas9 protein) and targets the Cas protein to a specific location within a target DNA. A guide RNA may contain two segments: a "DNA-targeting segment" and a "protein-binding segment." A "segment" includes a section or region of a molecule, such as a consecutive stretch of nucleotides in an RNA. Some gRNAs, such as those of Cas9, may contain two separate RNA molecules: an "activator RNA" (e.g., tracrRNA) and a "targeter RNA" (e.g., CRISPR RNA or crRNA). Other gRNAs are single RNA molecules (single RNA polynucleotides), which may also be referred to as "single-molecule gRNA," "single guide RNA," or "sgRNA." See, for example, WO2013 / 176772, WO2014 / 065596, WO2014 / 089290, WO2014 / 093622, WO2014 / 099750, WO2013 / 142578, and WO2014 / 131833, each of which is incorporated herein by reference in its entirety for all purposes. For example, in the case of Cas9, a single guide RNA may comprise a crRNA fused to a tracrRNA (e.g., via a linker). For example, in the case of Cpf1, only the crRNA is required to achieve binding to the target sequence. The terms "guide RNA" and "gRNA" include both double-molecule (i.e., modular) gRNAs and single-molecule gRNAs.
[0256] Exemplary bimolecular gRNAs include a crRNA-like ("CRISPR RNA" or "targeter RNA" or "crRNA" or "crRNA repeat") molecule and a corresponding tracrRNA-like ("trans-acting CRISPR RNA" or "activator RNA" or "tracrRNA") molecule. The crRNA comprises both the DNA-targeting segment (single strand) of the gRNA and a stretch of nucleotides that forms one half of the dsRNA duplex of the protein-binding segment of the gRNA. An example of a crRNA tail located downstream (3') of the DNA-targeting segment comprises, consists essentially of, or consists of GUUUUAGAGCUAUGCU (SEQ ID NO: 23). Any of the DNA-targeting segments disclosed herein can be attached to the 5' end of SEQ ID NO: 23 to form a crRNA.
[0257] The corresponding tracrRNA (activator RNA) contains a stretch of nucleotides that forms the other half of the dsRNA duplex of the protein-binding segment of the gRNA. The stretch of nucleotides in the crRNA is complementary to and hybridizes with the stretch of nucleotides in the tracrRNA to form the dsRNA duplex of the protein-binding domain of the gRNA. Thus, each crRNA can be said to have a corresponding tracrRNA. An example of a tracrRNA sequence comprises, consists essentially of, or consists of AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUU (SEQ ID NO: 24). Other examples of other tracrRNA sequences comprise, consist essentially of, or consist of AAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 28), or GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 29).
[0258] In systems where both crRNA and tracrRNA are required, the crRNA and the corresponding tracrRNA hybridize to form a gRNA. In systems where only crRNA is required, the crRNA can be the gRNA. The crRNA also provides a single-stranded DNA targeting segment that hybridizes to the complementary strand of the target DNA. When used for intracellular modification, the precise sequence of a given crRNA or tracrRNA molecule can be designed to be specific for the species in which the RNA molecule is used. See, e.g., Mali et al. (2013) Science 339:823-826, Jinek et al. (2012) Science 337:816-821, Hwang et al. (2013) Nat. Biotechnol. 31:227-229, Jiang et al. (2013) Nat. Biotechnol. 31:233-239, and Cong et al. (2013) Science 339:819-823, each of which is incorporated by reference in its entirety for all purposes.
[0259] The DNA-targeting segment (crRNA) of a given gRNA contains a nucleotide sequence that is complementary to a sequence on the complementary strand of the target DNA, as described in more detail below. The DNA-targeting segment of a gRNA interacts with the target DNA in a sequence-specific manner through hybridization (i.e., base pairing). Therefore, the nucleotide sequence of the DNA-targeting segment may vary and determines the location within the target DNA where the gRNA and target DNA interact. The DNA-targeting segment of a given gRNA can be modified to hybridize to any desired sequence within the target DNA. Naturally occurring crRNAs vary depending on the CRISPR / Cas system and organism, but often contain a targeting segment 21-72 nucleotides long flanked by two direct repeats (DRs) that are 21-46 nucleotides long (see, e.g., WO2014 / 131833, incorporated herein by reference in its entirety for all purposes). In S. pyogenes, the DRs are 36 nucleotides long and the targeting segment is 30 nucleotides long. The 3'-located DR is complementary to and hybridizes with the corresponding tracrRNA, which in turn binds to the Cas protein.
[0260] A DNA-targeting segment can have a length of, for example, at least about 12, 15, 17, 18, 19, 20, 25, 30, 35, or 40 nucleotides. Such a DNA-targeting segment can have a length of, for example, about 12 to about 100, about 12 to about 80, about 12 to about 50, about 12 to about 40, about 12 to about 30, about 12 to about 25, or about 12 to about 20 nucleotides. For example, a DNA-targeting segment can be about 15 to about 25 nucleotides (e.g., about 17 to about 20 nucleotides, or about 17, 18, 19, or 20 nucleotides). See, e.g., US2016 / 0024523, incorporated herein by reference in its entirety for all purposes. For Cas9 derived from S. pyogenes, a typical DNA-targeting segment is 16 to 20 nucleotides in length, or 17 to 20 nucleotides in length. For Cas9 from S. aureus, a typical DNA-targeting segment is 21-23 nucleotides in length. For Cpf1, a typical DNA-targeting segment is at least 16 nucleotides in length or at least 18 nucleotides in length.
[0261] TracrRNA can be in any form (e.g., full-length tracrRNA or active partial tracrRNA) and of various lengths, including primary transcripts or processed forms. For example, tracrRNA (as part of a single-guide RNA or as a separate molecule as part of a bimolecular gRNA) can comprise, consist essentially of, or consist of all or a portion of the wild-type tracrRNA sequence (e.g., about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides of the wild-type tracrRNA sequence). Examples of wild-type tracrRNA sequences from S. pyogenes include 171-, 89-, 75-, and 65-nucleotide versions. See, e.g., Deltcheva et al. (2011) Nature 471:602-607, WO2014 / 093661, each of which is incorporated by reference in its entirety for all purposes. Examples of tracrRNA within a single guide RNA (sgRNA) include the tracrRNA segments found within the +48, +54, +67, and +85 versions of the sgRNA, where "+n" indicates that up to +n nucleotides of the wild-type tracrRNA are included in the sgRNA. See US8,697,359, which is incorporated by reference in its entirety for all purposes.
[0262] The percent complementarity between the DNA-targeting segment of the guide RNA and the complementary strand of the target DNA can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%). The percent complementarity between the DNA-targeting segment and the complementary strand of the target DNA can be at least 60% over approximately 20 contiguous nucleotides. As an example, the percent complementarity between the DNA-targeting segment and the complementary strand of the target DNA can be 100% over 14 contiguous nucleotides at the 5'-end of the complementary strand of the target DNA and can be as low as 0% over the remainder. In such a case, the DNA-targeting segment can be considered to be 14 nucleotides in length. As another example, the percent complementarity between the DNA-targeting segment and the complementary strand of the target DNA can be 100% over 7 contiguous nucleotides at the 5'-end of the complementary strand of the target DNA and can be as low as 0% over the remainder. In such cases, the DNA targeting segment can be considered to be 7 nucleotides in length. In some guide RNAs, at least 17 nucleotides in the DNA targeting segment are complementary to the complementary strand of target DNA. For example, the DNA targeting segment can be 20 nucleotides in length and can contain 1, 2, or 3 mismatches with the complementary strand of target DNA. In one example, the mismatch is not adjacent to the region of the complementary strand corresponding to the protospacer adjacent motif (PAM) sequence (i.e., the reverse complement of the PAM sequence) (e.g., the mismatch is at the 5' end of the DNA targeting segment of the guide RNA, or the mismatch is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 base pairs away from the region of the complementary strand corresponding to the PAM sequence).
[0263] The protein-binding segment of the gRNA may contain two stretches of nucleotides that are complementary to each other. The complementary nucleotides of the protein-binding segment hybridize to form a double-stranded RNA duplex (dsRNA). The protein-binding segment of the target gRNA interacts with the Cas protein, and the gRNA directs the bound Cas protein to a specific nucleotide sequence within the target DNA via the DNA-targeting segment.
[0264] A single guide RNA can include a DNA-targeting segment and a scaffold sequence (i.e., the protein-binding or Cas-binding sequence of the guide RNA). For example, such a guide RNA can have a 5' DNA-targeting segment linked to a 3' scaffold sequence. Exemplary scaffold sequences comprise, consist essentially of, or consist of GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCU (version 1; SEQ ID NO: 17), GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 2; SEQ ID NO: 18), GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 3; SEQ ID NO: 19), and GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 4; SEQ ID NO: 20). Other exemplary scaffold sequences comprise, consist essentially of, or consist of GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU (Version 5; SEQ ID NO: 30), GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (Version 6; SEQ ID NO: 31), or GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (Version 7; SEQ ID NO: 32).A guide RNA targeting any of the guide RNA target sequences disclosed herein may, for example, comprise a DNA-targeting segment at the 5' end of the guide RNA fused to any of the exemplary guide RNA scaffold sequences at the 3' end of the guide RNA. That is, any of the DNA-targeting segments disclosed herein can be attached to the 5' end of any one of the above scaffold sequences to form a single guide RNA (chimeric guide RNA).
[0265] Guide RNAs may contain modifications or sequences that provide additional desirable characteristics (e.g., modified or regulated stability; intracellular targeting; tracking by fluorescent labeling; binding sites for proteins or protein complexes, etc.). Examples of such modifications include, for example, a 5' cap (e.g., a 7-methylguanylate cap (m7G)), a 3' polyadenylation tail (i.e., a 3' poly(A) tail), a riboswitch sequence (e.g., to allow for regulated stability and / or regulated accessibility by proteins and / or protein complexes), a stability control sequence, a sequence that forms a dsRNA duplex (i.e., a hairpin), a modification or sequence that targets the RNA to a subcellular location (e.g., the nucleus, mitochondria, chloroplasts, etc.), a modification or sequence that provides tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows fluorescent detection, etc.), a modification or sequence that provides a binding site for a protein (e.g., a protein that acts on DNA, including a transcriptional activator, a transcriptional repressor, a DNA methyltransferase, a DNA demethylase, a histone acetyltransferase, a histone deacetylase, etc.), and combinations thereof. Other examples of modifications include an engineered stem-loop duplex, an engineered bulge region, an engineered hairpin 3' of a stem-loop duplex, or any combination thereof. See, e.g., US2015 / 0376586, incorporated herein by reference in its entirety for all purposes. A bulge can be an unpaired region of nucleotides within a duplex comprised of a crRNA-like region and a minimal tracrRNA-like region. A bulge can comprise an unpaired 5'-XXXY-3' on one side of the duplex, where X is any purine and Y can be a nucleotide that can form a wobble pair with a nucleotide on the opposite strand and with a region of unpaired nucleotides on the other side of the duplex.
[0266] In some cases, a transcription activation system comprising a dCas9-VP64 fusion protein paired with MS2-p65-HSF1 can be used. The guide RNA of such a system can be designed using an aptamer sequence added to stem loop 2 and an sgRNA tetraloop designed to bind to the dimerized MS2 bacteriophage coat protein. See, e.g., Konermann et al. (2015) Nature 517(7536):583-588, the entire contents of which are incorporated herein by reference for all purposes.
[0267] Unmodified nucleic acids may be prone to degradation. Exogenous nucleic acids may also elicit an innate immune response. Modifications can help introduce stability and reduce immunogenicity. Guide RNAs may contain modified nucleosides and nucleotides, including, for example, one or more of the following: (1) modification or substitution of one or both of the non-linked phosphate oxygens and / or one or more of the linked phosphate oxygens in the phosphodiester backbone linkage; (2) modification or substitution of components of the ribose sugar, such as modification or substitution of the 2' hydroxyl of the ribose sugar; (3) replacement of the phosphate moiety with a dephosphoryl linker; (4) modification or substitution of naturally occurring nucleobases; (5) substitution or modification of the ribose-phosphate backbone; (6) modification of the 3' or 5' end of the oligonucleotide (e.g., removal, modification, or substitution of the terminal phosphate group, or conjugation of a moiety); and (7) sugar modifications. Other possible guide RNA modifications include modification or substitution of uracil or poly-uracil tracts. See, for example, WO2015 / 048577 and US2016 / 0237455, each of which is incorporated herein by reference in its entirety for all purposes. Similar modifications can be made to Cas-encoding nucleic acids, such as Cas mRNA. For example, Cas mRNA can be modified by using synonymous codons to deplete uridines.
[0268] As one example, the nucleotides at the 5' or 3' end of the guide RNA may contain phosphorothioate linkages (e.g., the base may have a modified phosphate group that is a phosphorothioate group). For example, the guide RNA may contain phosphorothioate linkages between the two, three, or four terminal nucleotides at the 5' or 3' end of the guide RNA. As another example, the nucleotides at the 5' and / or 3' end of the guide RNA may have 2'-O-methyl modifications. For example, the guide RNA may contain 2'-O-methyl modifications in the two, three, or four terminal nucleotides at the 5' and / or 3' end (e.g., the 5' end) of the guide RNA. See, for example, WO2017 / 173054A1 and Finn et al. (2018) Cell Reports 22:1-9, each of which is incorporated herein by reference in its entirety for all purposes. Other possible modifications are described in more detail elsewhere herein. In a specific example, the guide RNA contains 2'-O-methyl analogs and 3' phosphorothioate internucleotide linkages at the first three 5' and 3' terminal RNA residues. Such chemical modifications provide the guide RNA with greater stability and protection from exonucleases, allowing them to persist in cells longer than unmodified guide RNAs. Such chemical modifications can also protect against, for example, natural intracellular immune responses that can actively degrade RNA or trigger immune cascades that lead to cell death.
[0269] In some guide RNAs (e.g., single-guide RNAs), at least one loop (e.g., two loops) of the guide RNA is modified by the insertion of a distinct RNA sequence that binds to one or more adapters (i.e., adapter proteins or domains). Such adapter proteins can be used to further recruit one or more heterologous functional domains, such as transcription activation domains (e.g., for use in CRISPRa screening in SAM / tau biosensor cells). Examples of fusion proteins containing such adapter proteins (i.e., chimeric adapter proteins) are disclosed elsewhere herein. For example, the MS2 binding loop ggccAACAUGAGGAUCACCCAUGUCUGCAGggcc (SEQ ID NO: 33) can replace nucleotides +13 to +16 and nucleotides +53 to +56 of the sgRNA scaffold (backbone) shown in SEQ ID NO: 17 or SEQ ID NO: 19 (or SEQ ID NO: 30 or 31) or the sgRNA backbone described in WO2016 / 049258 and Konermann et al. (2015) Nature 517(7536):583-588, each of which is incorporated by reference in its entirety for all purposes. See also US2019-0284572 and WO2019 / 183123, each of which is incorporated by reference in its entirety for all purposes. As used herein, guide RNA numbering refers to the numbering of nucleotides in the guide RNA scaffold sequence (i.e., the sequence downstream of the DNA-targeting segment of the guide RNA). For example, the first nucleotide of the guide RNA scaffold is +1, the second nucleotide of the scaffold is +2, etc. Residues corresponding to nucleotides +13 to +16 of SEQ ID NO: 17 or SEQ ID NO: 19 (or SEQ ID NO: 30 or 31) are the loop sequence of the region spanning nucleotides +9 to +21 of SEQ ID NO: 17 or SEQ ID NO: 19 (or SEQ ID NO: 30 or 31), a region referred to herein as the tetraloop.Residues corresponding to nucleotides +53 to +56 of SEQ ID NO:17 or SEQ ID NO:19 (or SEQ ID NO:30 or 31) are the loop sequence of the region spanning nucleotides +48 to +61 of SEQ ID NO:17 or SEQ ID NO:19 (or SEQ ID NO:30 or 31), a region referred to herein as stem-loop 2. Other stem-loop sequences in SEQ ID NO:17 or SEQ ID NO:19 (or SEQ ID NO:30 or 31) include stem-loop 1 (nucleotides +33 to +41) and stem-loop 3 (nucleotides +63 to +75). The resulting structure is an sgRNA scaffold in which each of the tetraloop and stem-loop 2 sequences has been replaced by an MS2-binding loop. The tetraloop and stem-loop 2 protrude from the Cas9 protein such that adding the MS2-binding loop should not interfere with any Cas9 residues. Furthermore, the proximity of the tetraloop and stem-loop 2 sites to DNA indicates that localization to these locations may result in a high degree of interaction between the DNA and any recruited proteins, such as transcriptional activators. Thus, in some sgRNAs, the nucleotides corresponding to +13 to +16 and / or +53 to +56 of the guide RNA scaffold set forth in SEQ ID NO:17 or SEQ ID NO:19 (or SEQ ID NO:30 or 31), or the corresponding residues when optimally aligned with either of these scaffolds / backbones, are replaced by a separate RNA sequence capable of binding to one or more adaptor proteins or domains. Alternatively or additionally, adaptor binding sequences may be added to the 5' or 3' end of the guide RNA. An exemplary guide RNA scaffold containing a tetraloop and an MS2-binding loop in the stem-loop 2 region may comprise, consist essentially of, or consist of the sequence set forth in SEQ ID NO:34. An exemplary generic single guide RNA containing a tetraloop and an MS2-binding loop in the stem-loop 2 region may comprise, consist essentially of, or consist of the sequence set forth in SEQ ID NO:35.
[0270] The guide RNA can be provided in any form. For example, the gRNA can be provided in the form of RNA, either as two molecules (separate crRNA and tracrRNA) or as one molecule (sgRNA), and optionally in the form of a complex with a Cas protein. The gRNA can also be provided in the form of DNA encoding the gRNA. The DNA encoding the gRNA can encode a single RNA molecule (sgRNA) or separate RNA molecules (e.g., separate crRNA and tracrRNA). In the latter case, the DNA encoding the gRNA can be provided as one DNA molecule or as separate DNA molecules encoding the crRNA and tracrRNA, respectively.
[0271] When the gRNA is provided in the form of DNA, the gRNA can be expressed transiently, conditionally, or constitutively in the cell. The DNA encoding the gRNA can be stably integrated into the genome of the cell and operably linked to a promoter active in the cell. Alternatively, the DNA encoding the gRNA can be operably linked to a promoter in an expression construct. For example, the DNA encoding the gRNA can be within a vector containing a heterologous nucleic acid, such as a nucleic acid encoding a Cas protein. Alternatively, it can be a vector or plasmid separate from the vector containing the nucleic acid encoding the Cas protein. Promoters that can be used in such expression constructs include, for example, promoters active in one or more of eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, pluripotent cells, embryonic stem (ES) cells, adult stem cells, developmentally restricted progenitor cells, induced pluripotent stem (iPS) cells, or one-cell embryos. Such promoters can be, for example, conditional promoters, inducible promoters, constitutive promoters, or tissue-specific promoters. Such promoters can also be, for example, bidirectional promoters. Specific examples of suitable promoters include RNA polymerase III promoters such as the human U6 promoter, the rat U6 polymerase III promoter, or the mouse U6 polymerase III promoter.
[0272] Alternatively, gRNA can be prepared by various other methods. For example, gRNA can be prepared by in vitro transcription using, for example, T7 RNA polymerase (see, for example, WO2014 / 089290 and WO2014 / 065596, each of which is incorporated herein by reference in its entirety for all purposes). Guide RNA can also be a synthetically produced molecule prepared by chemical synthesis. For example, guide RNA can be chemically synthesized to include 2'-O-methyl analogs and 3' phosphorothioate internucleotide linkages in the first three 5' and 3' terminal RNA residues.
[0273] The guide RNA (or nucleic acid encoding the guide RNA) can be in a composition comprising one or more guide RNAs (e.g., 1, 2, 3, 4, or more guide RNAs) and a carrier that increases the stability of the guide RNA (e.g., extends the period below a threshold such that degradation products remain below 0.5% by weight of the starting nucleic acid or protein under given storage conditions (e.g., -20°C, 4°C, or ambient temperature) or increases stability in vivo). Non-limiting examples of such carriers include poly(lactic acid) (PLA) microspheres, poly(D,L-lactic-coglycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipid cochleates, and lipid microtubules. Such compositions can further comprise a Cas protein, such as a Cas9 protein, or a nucleic acid encoding a Cas protein.
[0274] B. Guide RNA target sequence The target DNA of a guide RNA includes a nucleic acid sequence present in the DNA to which the DNA target segment of the gRNA binds when conditions sufficient for binding exist. Suitable DNA / RNA binding conditions include physiological conditions normally present in cells. Other suitable DNA / RNA binding conditions (e.g., conditions in cell-free systems) are known in the art (see, e.g., Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001), incorporated herein by reference in its entirety for all purposes). The strand of the target DNA that is complementary to and hybridizes with the gRNA can be referred to as the "complementary strand," and the strand of the target DNA that is complementary to the "complementary strand" (and therefore not complementary to the Cas protein or gRNA) can be referred to as the "non-complementary strand" or "template strand."
[0275] The target DNA includes both the sequence on the complementary strand to which the guide RNA hybridizes and the corresponding sequence on the non-complementary strand (e.g., adjacent to the protospacer adjacent motif (PAM)). As used herein, the term "guide RNA target sequence" specifically refers to the sequence on the non-complementary strand that corresponds to (i.e., its reverse complement) the sequence to which the guide RNA hybridizes on the complementary strand. That is, the guide RNA target sequence refers to the sequence on the non-complementary strand adjacent to the PAM (e.g., upstream or 5' of the PAM in the case of Cas9). The guide RNA target sequence is equivalent to the DNA-targeting segment of the guide RNA, but contains thymine instead of uracil. As an example, the guide RNA target sequence for the SpCas9 enzyme may refer to the sequence upstream of the 5'-NGG-3' PAM on the non-complementary strand. The guide RNA is designed to have complementarity to the complementary strand of the target DNA, and hybridization between the DNA-targeting segment of the guide RNA and the complementary strand of the target DNA promotes the formation of a CRISPR complex. Full complementarity is not necessarily required, as long as there is sufficient complementarity to cause hybridization and promote the formation of CRISPR complexes.When guide RNA is referred to herein as targeting guide RNA target sequence, it means that guide RNA hybridizes with the complementary strand sequence of target DNA, which is the reverse complement of guide RNA target sequence on non-complementary strand.
[0276] The target DNA or guide RNA target sequence can comprise any polynucleotide, and can be located in the nucleus or cytoplasm of a cell, or in a cell organelle such as mitochondria or chloroplast.The target DNA or guide RNA target sequence can be any nucleic acid sequence that is endogenous or exogenous to a cell.The guide RNA target sequence can be a sequence that encodes a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory sequence), or can comprise both.
[0277] For CRISPRa and SAM-based systems, it may be preferable for the target sequence to be adjacent to the transcription start site of the gene. For example, the target sequence can be 1000, 900, 800, 700, 600, 500, 400, 300, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, or 1 base pair within the transcription start site. Optionally, the target sequence is within the region 200 base pairs upstream of the transcription start site and 1 base pair downstream (-200 to +1) of the transcription start site.
[0278] Site-specific binding and cleavage of target DNA by a Cas protein can occur at a location determined by both (i) base-pairing complementarity between the guide RNA and the complementary strand of the target DNA and (ii) a short motif, called a protospacer adjacent motif (PAM), on the non-complementary strand of the target DNA. The PAM can be adjacent to the guide RNA target sequence. Optionally, the guide RNA target sequence can be adjacent to the PAM at its 3' end (e.g., in the case of Cas9). Alternatively, the guide RNA target sequence can be adjacent to the PAM at its 5' end (e.g., in the case of Cpf1). For example, the cleavage site of the Cas protein can be about 1 to about 10 or about 2 to about 5 base pairs (e.g., 3 base pairs) upstream or downstream (e.g., within the guide RNA target sequence) of the PAM sequence. In the case of SpCas9, the PAM sequence (i.e., on the non-complementary strand) can be 5'-N1GG-3', where N1 is any DNA nucleotide, and PAM is immediately 3' to the guide RNA target sequence on the non-complementary strand of the target DNA. Thus, the sequence corresponding to the PAM on the complementary strand (i.e., the reverse complement) is 5'-CCN2-3', where N2 is any DNA nucleotide, and is immediately 5' to the sequence where the DNA-targeting segment of the guide RNA hybridizes to the complementary strand of the target DNA. In some such cases, N1 and N2 can be complementary, and the N1-N2 base pair can be any base pair (e.g., N1=C and N2=G, N1=G and N2=C, N1=A and N2=T, or N1=T and N2=A). For Cas9 from S. aureus, the PAM can be NNGRRT or NNGRR, where N can be A, G, C, or T, and R can be G or A. For Cas9 from C. jejuni, the PAM can be, for example, NNNNACAC or NNNNRYAC, where N can be A, G, C, or T, and R can be G or A. In some cases (e.g., for FnCpfl), the PAM sequence is 5' upstream and can have the sequence 5'-TTN-3'.
[0279] An example of a guide RNA target sequence is a 20-nucleotide DNA sequence immediately preceding the NGG motif recognized by the SpCas9 protein. For example, two examples of guide RNA target sequences and PAMs are GN19NGG (SEQ ID NO: 25) or N20NGG (SEQ ID NO: 26). See, e.g., WO2014 / 165825, incorporated herein by reference in its entirety for all purposes. A guanine at the 5' end can promote transcription by RNA polymerase in cells. Another example of a guide RNA target sequence and PAM may include two guanine nucleotides at the 5' end (e.g., GGN20NGG; SEQ ID NO: 27) to promote efficient transcription by T7 polymerase in vitro. See, e.g., WO2014 / 065596, incorporated herein by reference in its entirety for all purposes. Other guide RNA target sequences and PAMs may have a length of 4 to 22 nucleotides, as shown in SEQ ID NOs: 25 to 27, including a 5' G or GG and a 3' GG or NGG. Still other guide RNA target sequences and PAMs may have a length of 14 to 20 nucleotides of SEQ ID NOs: 25 to 27.
[0280] Formation of a CRISPR complex hybridized to the target DNA can result in cleavage of one or both strands of the target DNA within or near the region corresponding to the guide RNA target sequence (i.e., the guide RNA target sequence on the non-complementary strand of the target DNA and its reverse complement on the complementary strand to which the guide RNA hybridizes). For example, the cleavage site can be within the guide RNA target sequence (e.g., at a defined position relative to the PAM sequence). A "cleavage site" includes the location in the target DNA where the Cas protein generates a single-stranded or double-stranded break. The cleavage site can be single-stranded only (e.g., when a nickase is used) or on both strands of double-stranded DNA. The cleavage sites can be at the same position on both strands (generating blunt ends; e.g., Cas9) or at different sites on each strand (generating staggered ends (i.e., overhangs); e.g., Cpf1). Staggered ends can be generated, for example, by using two Cas proteins, each generating a single-stranded break at a different cleavage site on a different strand, thereby generating a double-stranded break. For example, a first nickase can create a single-stranded break on a first strand of double-stranded DNA (dsDNA), and a second nickase can create a single-stranded break on a second strand of the dsDNA such that an overhang sequence is created. In some cases, the guide RNA target sequence or cleavage site of the first strand nickase is separated from the guide RNA target sequence or cleavage site of the second strand nickase by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, or 1,000 base pairs.
[0281] IV. Methods for Screening for Genetic Modifiers of Tau Seeding or Aggregation The Cas / tau biosensor cell lines disclosed herein can be used in methods of screening for genetic modifiers of tau seeding or aggregation. Such methods can include providing a population of Cas / tau biosensor cells as disclosed elsewhere herein, introducing a library containing a plurality of unique guide RNAs, and assessing tau seeding or aggregation in target cells.
[0282] As an example, a method can include providing a population of Cas / tau biosensor cells (e.g., a cell population comprising a Cas protein, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter), introducing a library containing multiple unique guide RNAs targeting multiple genes into the cell population, and culturing the cell population to allow genome editing and expansion. The multiple unique guide RNAs form a complex with the Cas protein, which cleaves the multiple genes, resulting in knockout of gene function, to generate a genetically modified cell population. The genetically modified cell population can then be contacted with a tau seeding agent to generate a seeded cell population. The seeded cell population can be cultured to allow the formation of tau aggregates, and aggregates of the first tau repeat domain and the second tau repeat domain form in a subset of the seeded cell population to generate an aggregate-positive cell population. Finally, the abundance of each of the multiple unique guide RNAs can be determined in the aggregation-positive cell population relative to the cell population cultured after introduction of the guide RNA library. Enrichment of the guide RNA in the aggregation-positive cell population relative to the cell population cultured after introduction of the guide RNA library indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, and disruption of the gene targeted by the guide RNA enhances tau aggregation or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and disruption of the gene targeted by the guide RNA is expected to enhance tau aggregation.
[0283] Similarly, the SAM / tau biosensor cell lines disclosed herein can be used in methods of screening for genetic modifiers of tau seeding or aggregation. Such methods can include providing a population of SAM / tau biosensor cells as disclosed elsewhere herein, introducing a library containing a plurality of unique guide RNAs, and assessing tau seeding or aggregation in target cells.
[0284] As an example, a method can include providing a population of SAM / tau biosensor cells (e.g., a cell population comprising a chimeric Cas protein comprising a nuclease-inactive Cas protein fused to one or more transcription activation domains, a chimeric adapter protein comprising an adapter protein fused to one or more transcription activation domains, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter), introducing a library comprising multiple unique guide RNAs targeting multiple genes into the cell population, and culturing the cell population to allow transcription activation and expansion. The multiple unique guide RNAs form a complex with the chimeric Cas protein and the chimeric adapter protein, and the complex activates transcription of the multiple genes, resulting in gene expression and expansion of the modified cell population. The modified cell population can then be contacted with a tau seeding agent to generate a seeded cell population. The seeded cell population can be cultured to allow tau aggregate formation, and aggregates of the first tau repeat domain and the second tau repeat domain form in a subset of the seeded cell population to generate an aggregation-positive cell population. Finally, the abundance of each of the multiple unique guide RNAs can be determined in the aggregation-positive cell population relative to the cell population cultured after introduction of the guide RNA library. Enrichment of the guide RNA in the aggregation-positive cell population relative to the cell population cultured after introduction of the guide RNA library indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, and transcriptional activation of the gene targeted by the guide RNA enhances tau aggregation or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and transcriptional activation of the gene targeted by the guide RNA is expected to enhance tau aggregation.
[0285] The Cas / tau biosensor cell used in this method can be any of the Cas / tau biosensor cells disclosed elsewhere herein. Similarly, the SAM / tau biosensor cell used in this method can be any of the SAM / tau biosensor cells disclosed elsewhere herein. The first tau repeat domain and the second tau repeat domain can be different, similar, or the same. The tau repeat domain can be any of the tau repeat domains disclosed elsewhere herein. For example, the first tau repeat domain and / or the second tau repeat domain can be a wild-type tau repeat domain or can include an aggregation-promoting mutation (e.g., a pathogenic, aggregation-promoting mutation) such as the tau P301S mutation. The first tau repeat domain and / or the second tau repeat domain can include a tau 4 repeat domain. As one particular example, the first tau repeat domain and / or the second tau repeat domain may comprise, consist essentially of, or consist of SEQ ID NO: 11 or a sequence that is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 11. In one particular example, the nucleic acid encoding the tau repeat domain may comprise, consist essentially of, or consist of SEQ ID NO: 12 or a sequence that is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 12, optionally, the nucleic acid encodes a protein that comprises, consists essentially of, or consists of SEQ ID NO: 11.
[0286] The first tau repeat domain can be linked to a first reporter, and the second tau repeat domain can be linked to a second reporter by any means. For example, the reporter can be fused to the tau repeat domain (e.g., as part of a fusion protein). The reporter proteins can be any pair of reporter proteins that generate a detectable signal when the first tau repeat domain linked to the first reporter aggregates with the second tau repeat domain linked to the second reporter. As an example, the first and second reporters can be split luciferase proteins. As another example, the first and second reporter proteins can be a fluorescence resonance energy transfer (FRET) pair. FRET is a physical phenomenon in which an excited donor fluorophore nonradiatively transfers its excitation energy to an adjacent acceptor fluorophore, thereby causing the acceptor to emit its characteristic fluorescence. Examples of FRET pairs (donor and acceptor fluorophores) are well known. See, e.g., Bajar et al. (2016) Sensors (Basel) 16(9):1488, incorporated herein by reference in its entirety for all purposes. As one specific example of a FRET pair, the first reporter can be a cyan fluorescent protein (CFP) and the second reporter can be a yellow fluorescent protein (YFP). As a specific example, the CFP can comprise, consist essentially of, or consist of SEQ ID NO: 13 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 13. As another specific example, the YFP can comprise, consist essentially of, or consist of SEQ ID NO: 15 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 15.
[0287] In the case of a Cas / tau biosensor cell, the Cas protein can be any Cas protein disclosed elsewhere herein. As an example, the Cas protein can be a Cas9 protein. For example, the Cas9 protein can be a Streptococcus pyogenes Cas9 protein. As one particular example, the Cas protein can comprise, consist essentially of, or consist of SEQ ID NO:21 or a sequence that is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:21.
[0288] One or more or all of the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter can be stably expressed in the cell population. For example, a nucleic acid encoding one or more or all of the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter can be genomically integrated in the cell population. In one particular example, the nucleic acid encoding the Cas protein can comprise, consist essentially of, or consist of SEQ ID NO:22 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:22, and optionally, the nucleic acid encodes a protein comprising, consisting essentially of, or consisting of SEQ ID NO:21.
[0289] In the case of SAM / tau biosensor cells, the Cas protein can be any Cas protein disclosed elsewhere herein. As an example, the Cas protein can be a Cas9 protein. For example, the Cas9 protein can be a Streptococcus pyogenes Cas9 protein. In one specific example, the chimeric Cas protein can comprise a nuclease-inactive Cas protein fused to a VP64 transcriptional activation domain. For example, the chimeric Cas protein can comprise, from N- to C-terminus, a nuclease-inactive Cas protein, a nuclear localization signal, and a VP64 transcriptional activator domain. In one specific example, the adaptor protein can be an MS2 coat protein, and one or more transcriptional activation domains in the chimeric adaptor protein can comprise a p65 transcriptional activation domain and an HSF1 transcriptional activation domain. For example, the chimeric adaptor protein can comprise, from N- to C-terminus, an MS2 coat protein, a nuclear localization signal, a p65 transcriptional activation domain, and an HSF1 transcriptional activation domain. In one particular example, the nucleic acid encoding the chimeric Cas protein can comprise, consist essentially of, or consist of SEQ ID NO: 38 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 38, and optionally, the nucleic acid encodes a protein comprising, consisting essentially of, or consisting of SEQ ID NO: 36. In one particular example, the nucleic acid encoding the chimeric adapter protein can comprise, consist essentially of, or consist of SEQ ID NO: 39 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 39, and optionally, the nucleic acid encodes a protein comprising, consisting essentially of, or consisting of SEQ ID NO: 37.
[0290] One or more or all of the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to the first reporter, and the second tau repeat domain linked to the second reporter can be stably expressed in the cell population. For example, nucleic acids encoding one or more or all of the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to the first reporter, and the second tau repeat domain linked to the second reporter can be genomically integrated in the cell population.
[0291] As disclosed elsewhere herein, the cell can be of any cell type, for example, the cell can be a eukaryotic cell, a mammalian cell, or a human cell (e.g., a HEK293T cell or a neuronal cell).
[0292] The multiple unique guide RNAs can be introduced into a cell population by any known means. In some methods, the guide RNAs are introduced into a cell population by viral transduction, such as retroviral, adenoviral, or lentiviral transduction. In certain examples, the guide RNAs can be introduced by lentiviral transduction. Each of the multiple unique guide RNAs can be in a separate viral vector. The cell population can be infected at any multiplicity of infection. For example, the multiplicity of infection can be about 0.1 to about 1.0, about 0.1 to about 0.9, about 0.1 to about 0.8, about 0.1 to about 0.7, about 0.1 to about 0.6, about 0.1 to about 0.5, about 0.1 to about 0.4, or about 0.1 to about 0.3. Alternatively, the multiplicity of infection may be less than about 1.0, less than about 0.9, less than about 0.8, less than about 0.7, less than about 0.6, less than about 0.5, less than about 0.4, less than about 0.3, or less than about 0.2. In particular examples, the multiplicity of infection may be less than about 0.3.
[0293] The guide RNA can be introduced into a cell population along with a selection marker or reporter gene to select cells harboring the guide RNA, and the method can further include selecting cells containing the selection marker or reporter gene. Examples of selection markers and reporter genes are provided elsewhere herein. As an example, the selection marker can be one that confers resistance to drugs such as neomycin phosphotransferase, hygromycin B phosphotransferase, puromycin-N-acetyltransferase, and blasticidin S deaminase. Another exemplary selection marker is the bleomycin resistance protein encoded by the Sh ble gene (Streptoalloteichus hindustanus bleomycin gene), which confers resistance to zeocin (phleomycin D1). For example, cells can be selected with a drug (e.g., puromycin) so that only cells transduced with the guide RNA construct are preserved for use in performing screening. For example, the drug can be puromycin or zeocin (phleomycin D1).
[0294] In some methods, multiple unique guide RNAs are introduced at a concentration selected so that most cells receive only one of the unique guide RNAs.For example, when guide RNAs are introduced by viral transduction, cells can be infected at a low multiplicity of infection, ensuring that most cells receive only one viral construct with a high probability.In one specific example, the multiplicity of infection can be less than about 0.3.
[0295] The cell population into which multiple unique guide RNAs are introduced can be any suitable number of cells.For example, the cell population can comprise more than about 50, more than about 100, more than about 200, more than about 300, more than about 400, more than about 500, more than about 600, more than about 700, more than about 800, more than about 900, or more than about 1000 cells per unique guide RNA.In a specific example, the cell population comprises more than about 300 cells or more than about 500 cells per unique guide RNA.
[0296] Multiple unique guide RNAs can target any number of genes.For example, multiple unique guide RNAs can target about 50 or more genes, about 100 or more genes, about 200 or more genes, about 300 or more genes, about 400 or more genes, about 500 or more genes, about 1000 or more genes, about 2000 or more genes, about 3000 or more genes, about 4000 or more genes, about 5000 or more genes, about 10000 or more genes, or about 20000 or more genes.In some methods, guide RNA can be selected to target the gene of a specific signal transduction pathway.In some methods, the library of unique guide RNA is a genome-wide library.
[0297] Multiple unique guide RNAs can target any number of sequences within each individual targeted gene. In some methods, multiple target sequences are targeted, on average, in each of the targeted multiple genes. For example, an average of about 2 to about 10, about 2 to about 9, about 2 to about 8, about 2 to about 7, about 2 to about 6, about 2 to about 5, about 2 to about 4, or about 2 to about 3 unique target sequences may be targeted in each of the targeted multiple genes. For example, an average of at least about 2, at least about 3, at least about 4, at least about 5, or at least about 6 unique target sequences may be targeted in each of the targeted multiple genes. As a specific example, an average of about 6 target sequences may be targeted in each of the targeted multiple genes. As another specific example, an average of about 3 to about 6 or about 4 to about 6 target sequences may be targeted in each of the targeted multiple genes.
[0298] The guide RNA can target any desired location within the target gene. In some CRISPRn methods using Cas / tau biosensor cells, each guide RNA targets a constitutive exon, if possible. In some methods, each guide RNA targets a 5' constitutive exon, if possible. In some methods, each guide RNA targets the first exon, second exon, or third exon (from the 5' end of the gene), if possible. In some CRISPRa methods using SAM / tau biosensor cells, each guide RNA can target a guide RNA target sequence within 200 bp upstream of the transcription start site, if possible. In some CRISPRa methods using SAM / tau biosensor cells, each guide RNA can include one or more adapter binding elements to which a chimeric adapter protein can specifically bind. In one example, each guide RNA comprises two adaptor-binding elements to which a chimeric adaptor protein can specifically bind, optionally with a first adaptor-binding element in the first loop of each of the one or more guide RNAs and a second adaptor-binding element in the second loop of each of the one or more guide RNAs. For example, the adaptor-binding elements may comprise the sequence set forth in SEQ ID NO: 33. In a particular example, each of the one or more guide RNAs is a single guide RNA comprising a CRISPR RNA (crRNA) portion fused to a trans-activating CRISPR RNA (tracrRNA) portion, wherein the first loop is a tetraloop corresponding to residues 13-16 of SEQ ID NO: 17, 19, 30, or 31, and the second loop is stem-loop 2 corresponding to residues 53-56 of SEQ ID NO: 17, 19, 30, or 31.
[0299] The step of culturing the cell population to allow for genome editing and expansion can be for any suitable period of time. For example, the culture can be for about 2 to about 10 days, about 3 to about 9 days, about 4 to about 8 days, about 5 to about 7 days, or about 6 days. Similarly, the step of culturing the cell population to allow for transcription activation and expansion can be for any suitable period of time. For example, the culture can be for about 2 to about 10 days, about 3 to about 9 days, about 4 to about 8 days, about 5 to about 7 days, or about 6 days.
[0300] Any suitable tau seeding agent can be used to generate the seeded cell population. Suitable tau seeding agents are disclosed elsewhere herein. Some suitable seeding agents include, for example, a tau repeat domain that can be different from, similar to, or the same as the first and / or second tau repeat domains. In one example, the seeding step includes culturing the genetically modified cell population in the presence of conditioned medium harvested from cultured tau aggregation-positive cells in which the tau repeat domains are stably present in an aggregated state. For example, the conditioned medium can be harvested from confluent tau aggregation-positive cells after being placed on the confluent cells for about 1 to about 7 days, about 2 to about 6 days, about 3 to about 5 days, or about 4 days. The seeding step can include culturing the genetically modified cell population in any suitable ratio of conditioned medium and fresh medium. For example, it may be about 90% conditioned medium and about 10% fresh medium, about 85% conditioned medium and about 15% fresh medium, about 80% conditioned medium and about 20% fresh medium, about 75% conditioned medium and about 25% fresh medium, about 70% conditioned medium and about 30% fresh medium, about 65% conditioned medium and about 35% fresh medium, about 60% conditioned medium and about 40% fresh medium, about 55% conditioned medium and about 45% fresh medium, about 50% conditioned medium and about 50% fresh medium The method may include culturing the genetically modified cell population in about 45% conditioned medium and about 55% fresh medium, about 40% conditioned medium and about 60% fresh medium, about 35% conditioned medium and about 65% fresh medium, about 30% conditioned medium and about 70% fresh medium, about 25% conditioned medium and about 75% fresh medium, about 20% conditioned medium and about 80% fresh medium, about 15% conditioned medium and about 85% fresh medium, or about 10% conditioned medium and about 90% fresh medium. In one example, it may include culturing the genetically modified cell population in medium containing at least about 50% conditioned medium and about 50% or less fresh medium. In a particular example, it may include culturing the genetically modified cell population in about 75% conditioned medium and about 25% fresh medium.Optionally, the genetically modified cell population is not co-cultured with tau aggregate-positive cells in which the tau repeat domains are stably present in an aggregated state.
[0301] The step of culturing the seeded cell population to allow tau aggregate formation can be for any suitable period of time, with aggregates of the first tau repeat domain and the second tau repeat domain forming in a subset of the seeded cell population to generate an aggregation-positive cell population. For example, the culture can be for about 1 to about 7 days, about 2 to about 6 days, about 3 to about 5 days, or about 4 days. Aggregation can be determined by any suitable means, depending on the reporter used. For example, in methods in which the first and second reporters are a fluorescence resonance energy transfer (FRET) pair, the aggregation-positive cell population can be identified by flow cytometry.
[0302] The abundance of guide RNA can be determined by any suitable means.In certain examples, the abundance is determined by next-generation sequencing.Next-generation sequencing refers to the high-throughput DNA sequencing technology that is not based on Sanger sequencing.For example, determining the abundance of guide RNA can include measuring the read count of guide RNA.
[0303] In some methods, guide RNA is considered to be enriched when the abundance of guide RNA relative to the total population of multiple unique guide RNAs is at least about 1.5 times higher in the aggregation-positive cell population than in the cell population that is cultured after introducing guide RNA library.Various enrichment thresholds can also be used.For example, enrichment thresholds can be set higher to be more stringent (e.g., at least about 1.6 times, at least about 1.7 times, at least about 1.8 times, at least about 1.9 times, at least about 2 times, at least about 2.5 times, or at least about 2.5 times).Alternatively, enrichment thresholds can be set lower to be less stringent (e.g., at least about 1.4 times, at least about 1.3 times, or at least about 1.2 times).
[0304] In one example, the step of determining abundance may include determining the abundance of multiple unique guide RNAs in an aggregation-positive population for a cell population cultured after introduction of the guide RNA library at a first time point during culture and / or a second time point during culture. For example, the first time point may be the first passage of culturing the cell population, and the second time point may be midway through culturing the cell population to allow genome editing and expansion. For example, the first time point may be after sufficient time for the guide RNAs to form a complex with the Cas protein and after sufficient time for the Cas protein to cleave multiple genes, resulting in knockout of gene function (CRISPRn) or transcriptionally activate multiple genes (CRISPRa). However, the first time point should ideally be the first cell passage immediately after infection (i.e., before further expansion and genome editing) to determine gRNA library representation, as well as to determine whether each gRNA representation evolves from the first time point to the second time point and further time points up to the final time point. This allows for the elimination of gRNAs / targets that have been enriched due to the advantages of cell growth during the screening process by confirming that the abundance of gRNAs does not change between the first and second time points. As a specific example, the first time point can be about 1, 2, 3, or 4 days after culturing and expansion, and the second time point can be about 3, 4, 5, or 6 days after culturing and expansion. For example, the first time point can be about 3 days after culturing and expansion, and the second time point can be about 6 days after culturing and expansion. In some methods, if the abundance of a guide RNA targeting a gene relative to the total population of multiple unique guide RNAs is at least 1.5-fold higher (or higher than a different selected enrichment threshold) in the aggregation-positive cell population relative to the cell population cultured after introduction of the guide RNA library at both the first and second time points, the gene can then be considered a genetic modifier of tau aggregation, and disruption (CRISPRn) or transcriptional activation (CRISPRa) of the gene enhances (or is expected to enhance) tau aggregation.Alternatively or additionally, if the abundance of at least two unique guide RNAs targeting the gene relative to the total population of the plurality of unique guide RNAs is at least 1.5-fold higher (or higher than a different selected enrichment threshold) in the aggregation-positive cell population relative to the cell population cultured after introduction of the guide RNA library at either the first time point or the second time point, the gene may be considered a genetic modifier of tau aggregation, and disruption (CRISPRn) or transcriptional activation (CRISPRa) of the gene enhances (or is expected to enhance) tau aggregation.
[0305] Some CRISPRn methods involve the following steps to identify genes as genetic modifiers of tau aggregation: gene disruption (CRISPRn) or transcriptional activation (CRISPRa) enhances (or is expected to enhance) tau aggregation. Similarly, some CRISPRa methods involve the following steps to identify genes as genetic modifiers of tau aggregation: gene transcriptional activation enhances tau aggregation. The first step involves identifying which of multiple unique guide RNAs is present in the aggregation-positive cell population. The second step involves calculating the random probability that the identified guide RNA is present using the formula nCn'*(x-n')C(mn) / xCm, where x is the different unique guide RNAs introduced into the cell population, m is the different unique guide RNAs identified in step (1), n is the different unique guide RNAs introduced into the cell population that target the gene, and n' is the different unique guide RNAs identified in step (1) that target the gene. The third step involves calculating the average enrichment score of the guide RNAs identified in step (1). The guide RNA enrichment score is the relative abundance of the guide RNA in the aggregation-positive cell population divided by the relative abundance of the guide RNA in the cell population cultured after introduction of the guide RNA library. The relative abundance is the read count of the guide RNA divided by the read count of the total population of multiple unique guide RNAs. The fourth step involves selecting genes where the guide RNAs targeting the gene are significantly below random probability of being present and above an enrichment score threshold. Possible threshold enrichment scores are discussed above. As a specific example, the threshold enrichment score can be set to approximately 1.5x.
[0306] When used in the phrase "variety of unique guide RNAs," diversity refers to the number of unique guide RNA sequences. It is not abundance, but a qualitative "presence" or "absence." Variety of unique guide RNAs refers to the number of unique guide RNA sequences. Variety of unique guide RNAs is determined by next-generation sequencing (NGS) to identify all unique guide RNAs present in a cell population. This is done by using two primers that recognize the constant regions of the viral vector and amplify the gRNA between the constant regions, and one primer that recognizes the constant region for sequencing. Each unique guide RNA present in a sample is counted using a sequencing primer. The NGS results also include the sequence and the number of reads corresponding to the sequence. The read counts are used to calculate an enrichment score for each guide RNA, and the presence of each unique sequence indicates which guide RNAs are present. For example, if a gene has three unique guide RNAs before selection and all three are retained after selection, both n and n' are 3. These numbers are used to calculate statistics but are not used to calculate the actual read count. However, the read counts for each guide RNA (in one example, 100, 200, 50 corresponding to each of the three unique guide RNAs) are used to calculate the enrichment score.
[0307] V. Methods for screening genetic modifiers of tau aggregation that prevent tau aggregation The Cas / tau biosensor cell lines disclosed herein can be used in methods of screening for genetic modifiers of tau aggregation (e.g., that prevent or are expected to prevent tau aggregation). Such methods can include providing a population of Cas / tau biosensor cells as disclosed elsewhere herein, introducing a library containing a plurality of unique guide RNAs, and assessing tau aggregation in target cells.
[0308] As an example, a method can include providing a population of Cas / tau biosensor cells (e.g., a cell population comprising a Cas protein, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter), introducing a library containing multiple unique guide RNAs targeting multiple genes into the cell population, and culturing the cell population to allow genome editing and expansion. The multiple unique guide RNAs form a complex with the Cas protein, which cleaves the multiple genes, resulting in knockout of gene function, generating a genetically modified cell population. The genetically modified cell population can then be contacted with a tau seeding agent to generate a seeded cell population. For example, the tau seeding agent can be a "maximal seeding" agent, such as a cell lysate from tau aggregate-positive cells, as described elsewhere herein. The seeded cell population can be cultured to allow for the formation of tau aggregates, wherein aggregates of the first tau repeat domain and the second tau repeat domain form in a subset of the seeded cell population to generate an aggregation-positive cell population, and wherein aggregates do not form in a second subset of the seeded cell population to generate an aggregation-negative cell population. Finally, the abundance of each of the plurality of unique guide RNAs can be determined in the aggregation-positive cell population relative to the aggregation-negative cell population, and / or relative to the cell population that has been seeded, and / or relative to the cell population that has been cultured after introduction of the guide RNA library. Enrichment of the guide RNA in the aggregation-negative cell population relative to the aggregation-positive cell population, and / or relative to the cell population that has been seeded, and / or relative to the cell population that has been cultured after introduction of the guide RNA library indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, and disruption of the gene targeted by the guide RNA prevents tau aggregation or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and disruption of the gene targeted by the guide RNA is expected to prevent tau aggregation.Similarly, depletion of guide RNA in an aggregation-positive cell population relative to an aggregation-negative cell population, and / or relative to a cell population that has been seeded, and / or relative to a cell population that has been cultured after introduction of the guide RNA library, indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, and disruption of the gene targeted by the guide RNA prevents tau aggregation or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and disruption of the gene targeted by the guide RNA is expected to prevent tau aggregation. Enrichment of the guide RNA in the aggregation-positive cell population relative to the aggregation-negative cell population, and / or relative to the cell population that has been seeded, and / or relative to the cell population that has been cultured after introduction of the guide RNA library, indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, disruption of the gene targeted by the guide RNA promotes or enhances tau aggregation, or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and disruption of the gene targeted by the guide RNA is expected to promote or enhance tau aggregation. Similarly, depletion of guide RNA in an aggregation-negative cell population relative to an aggregation-positive cell population, and / or relative to a cell population that has been seeded, and / or relative to a cell population that has been cultured after introduction of the guide RNA library, indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, disruption of the gene targeted by the guide RNA promotes or enhances tau aggregation, or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and disruption of the gene targeted by the guide RNA is expected to promote or enhance tau aggregation.
[0309] Similarly, the SAM / tau biosensor cell lines disclosed herein can be used in methods of screening for genetic modifiers of tau aggregation (e.g., that prevent or are expected to prevent tau aggregation). Such methods can include providing a population of SAM / tau biosensor cells as disclosed elsewhere herein, introducing a library containing a plurality of unique guide RNAs, and assessing tau aggregation in the target cells.
[0310] As an example, a method can include providing a population of SAM / tau biosensor cells (e.g., a cell population comprising a chimeric Cas protein comprising a nuclease-inactive Cas protein fused to one or more transcription activation domains, a chimeric adapter protein comprising an adapter protein fused to one or more transcription activation domains, a first tau repeat domain linked to a first reporter, and a second tau repeat domain linked to a second reporter), introducing a library comprising multiple unique guide RNAs targeting multiple genes into the cell population, and culturing the cell population to allow transcription activation and expansion. The multiple unique guide RNAs form a complex with the chimeric Cas protein and the chimeric adapter protein, and the complex activates transcription of the multiple genes, resulting in gene expression and expansion of the modified cell population. The modified cell population can then be contacted with a tau seeding agent to generate a seeded cell population. For example, the tau seeding agent can be a "maximum seeding" agent as described elsewhere herein, such as cell lysate from tau aggregate-positive cells. The seeded cell population can be cultured to allow for the formation of tau aggregates, wherein aggregates of the first tau repeat domain and the second tau repeat domain form in a subset of the seeded cell population to generate an aggregation-positive cell population, and wherein aggregates do not form in a second subset of the seeded cell population to generate an aggregation-negative cell population. Finally, the abundance of each of the plurality of unique guide RNAs can be determined in the aggregation-positive cell population relative to the aggregation-negative cell population, and / or relative to the cell population that has been seeded, and / or relative to the cell population that has been cultured after introduction of the guide RNA library.Enrichment of guide RNAs in the aggregation-negative cell population relative to the aggregation-positive cell population, and / or relative to the cell population that has been seeded, and / or relative to the cell population that has been cultured after introduction of the guide RNA library, indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, transcriptional activation of the gene targeted by the guide RNA prevents tau aggregation or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and transcriptional activation of the gene targeted by the guide RNA is expected to prevent tau aggregation. Similarly, depletion of guide RNAs in an aggregation-positive cell population relative to an aggregation-negative cell population, and / or relative to a cell population that has been seeded, and / or relative to a cell population that has been cultured after introduction of the guide RNA library, indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, transcriptional activation of the gene targeted by the guide RNA prevents tau aggregation, or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and transcriptional activation of the gene targeted by the guide RNA is expected to prevent tau aggregation. Enrichment of guide RNAs in the aggregation-positive cell population relative to the aggregation-negative cell population, and / or relative to the cell population that has been seeded, and / or relative to the cell population that has been cultured after introduction of the guide RNA library, indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, transcriptional activation of the gene targeted by the guide RNA promotes or enhances tau aggregation, or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and transcriptional activation of the gene targeted by the guide RNA is expected to promote or enhance tau aggregation.Similarly, depletion of guide RNAs in an aggregation-negative cell population relative to an aggregation-positive cell population, and / or relative to a cell population that has been seeded, and / or relative to a cell population that has been cultured after introduction of the guide RNA library, indicates that the gene targeted by the guide RNA is a genetic modifier of tau aggregation, transcriptional activation of the gene targeted by the guide RNA promotes or enhances tau aggregation, or is a candidate genetic modifier of tau aggregation (e.g., for further testing via secondary screening), and transcriptional activation of the gene targeted by the guide RNA is expected to promote or enhance tau aggregation.
[0311] The Cas / tau biosensor cell used in this method can be any of the Cas / tau biosensor cells disclosed elsewhere herein. Similarly, the SAM / tau biosensor cell used in this method can be any of the SAM / tau biosensor cells disclosed elsewhere herein. The first tau repeat domain and the second tau repeat domain can be different, similar, or the same. The tau repeat domain can be any of the tau repeat domains disclosed elsewhere herein. For example, the first tau repeat domain and / or the second tau repeat domain can be a wild-type tau repeat domain or can include an aggregation-promoting mutation (e.g., a pathogenic, aggregation-promoting mutation) such as the tau P301S mutation. The first tau repeat domain and / or the second tau repeat domain can include a tau 4 repeat domain. As one particular example, the first tau repeat domain and / or the second tau repeat domain may comprise, consist essentially of, or consist of SEQ ID NO: 11 or a sequence that is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 11. In one particular example, the nucleic acid encoding the tau repeat domain may comprise, consist essentially of, or consist of SEQ ID NO: 12 or a sequence that is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 12, optionally, the nucleic acid encodes a protein that comprises, consists essentially of, or consists of SEQ ID NO: 11.
[0312] The first tau repeat domain can be linked to a first reporter, and the second tau repeat domain can be linked to a second reporter by any means. For example, the reporter can be fused to the tau repeat domain (e.g., as part of a fusion protein). The reporter proteins can be any pair of reporter proteins that generate a detectable signal when the first tau repeat domain linked to the first reporter aggregates with the second tau repeat domain linked to the second reporter. As an example, the first and second reporters can be split luciferase proteins. As another example, the first and second reporter proteins can be a fluorescence resonance energy transfer (FRET) pair. FRET is a physical phenomenon in which an excited donor fluorophore nonradiatively transfers its excitation energy to an adjacent acceptor fluorophore, thereby causing the acceptor to emit its characteristic fluorescence. Examples of FRET pairs (donor and acceptor fluorophores) are well known. See, e.g., Bajar et al. (2016) Sensors (Basel) 16(9):1488, incorporated herein by reference in its entirety for all purposes. As one specific example of a FRET pair, the first reporter can be a cyan fluorescent protein (CFP) and the second reporter can be a yellow fluorescent protein (YFP). As a specific example, the CFP can comprise, consist essentially of, or consist of SEQ ID NO: 13 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 13. As another specific example, the YFP can comprise, consist essentially of, or consist of SEQ ID NO: 15 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 15.
[0313] In the case of a Cas / tau biosensor cell, the Cas protein can be any Cas protein disclosed elsewhere herein. As an example, the Cas protein can be a Cas9 protein. For example, the Cas9 protein can be a Streptococcus pyogenes Cas9 protein. As one particular example, the Cas protein can comprise, consist essentially of, or consist of SEQ ID NO:21 or a sequence that is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:21.
[0314] One or more or all of the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter can be stably expressed in the cell population. For example, a nucleic acid encoding one or more or all of the Cas protein, the first tau repeat domain linked to a first reporter, and the second tau repeat domain linked to a second reporter can be genomically integrated in the cell population. In one particular example, the nucleic acid encoding the Cas protein can comprise, consist essentially of, or consist of SEQ ID NO:22 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:22, and optionally, the nucleic acid encodes a protein comprising, consisting essentially of, or consisting of SEQ ID NO:21.
[0315] In the case of SAM / tau biosensor cells, the Cas protein can be any Cas protein disclosed elsewhere herein. As an example, the Cas protein can be a Cas9 protein. For example, the Cas9 protein can be a Streptococcus pyogenes Cas9 protein. In one specific example, the chimeric Cas protein can comprise a nuclease-inactive Cas protein fused to a VP64 transcriptional activation domain. For example, the chimeric Cas protein can comprise, from N- to C-terminus, a nuclease-inactive Cas protein, a nuclear localization signal, and a VP64 transcriptional activator domain. In one specific example, the adaptor protein can be an MS2 coat protein, and one or more transcriptional activation domains in the chimeric adaptor protein can comprise a p65 transcriptional activation domain and an HSF1 transcriptional activation domain. For example, the chimeric adaptor protein can comprise, from N- to C-terminus, an MS2 coat protein, a nuclear localization signal, a p65 transcriptional activation domain, and an HSF1 transcriptional activation domain. In one particular example, the nucleic acid encoding the chimeric Cas protein can comprise, consist essentially of, or consist of SEQ ID NO: 38 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 38, and optionally, the nucleic acid encodes a protein comprising, consisting essentially of, or consisting of SEQ ID NO: 36. In one particular example, the nucleic acid encoding the chimeric adapter protein can comprise, consist essentially of, or consist of SEQ ID NO: 39 or a sequence at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 39, and optionally, the nucleic acid encodes a protein comprising, consisting essentially of, or consisting of SEQ ID NO: 37.
[0316] One or more or all of the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to the first reporter, and the second tau repeat domain linked to the second reporter can be stably expressed in the cell population. For example, nucleic acids encoding one or more or all of the chimeric Cas protein, the chimeric adaptor protein, the first tau repeat domain linked to the first reporter, and the second tau repeat domain linked to the second reporter can be genomically integrated in the cell population.
[0317] As disclosed elsewhere herein, the cell can be of any cell type, for example, the cell can be a eukaryotic cell, a mammalian cell, or a human cell (e.g., a HEK293T cell or a neuronal cell).
[0318] The multiple unique guide RNAs can be introduced into a cell population by any known means. In some methods, the guide RNAs are introduced into a cell population by viral transduction, such as retroviral, adenoviral, or lentiviral transduction. In certain examples, the guide RNAs can be introduced by lentiviral transduction. Each of the multiple unique guide RNAs can be in a separate viral vector. The cell population can be infected at any multiplicity of infection. For example, the multiplicity of infection can be about 0.1 to about 1.0, about 0.1 to about 0.9, about 0.1 to about 0.8, about 0.1 to about 0.7, about 0.1 to about 0.6, about 0.1 to about 0.5, about 0.1 to about 0.4, or about 0.1 to about 0.3. Alter...
Claims
[Claim 1] The invention described in the specification.