Regulatory elements for targeted expression of cargo genes in neurons
Nucleic acid constructs with cell type-specific regulatory elements drive targeted gene expression in neural cell types, addressing the lack of effective tools in nonhuman primates, facilitating research and treatment of neurological disorders.
Patent Information
- Application Number
- PCT/US2025/017688
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-27
- Filing Date
- 2025-02-27
- Publication Date
- 2025-09-04
AI Technical Summary
Current methods lack effective tools for manipulating specific neural cell types in nonhuman primates, limiting the understanding of complex cognitive functions and circuit basis, and there is a need for targeted gene expression in neural cell types to study neurological and psychiatric disorders.
Development of nucleic acid constructs containing cell type-specific regulatory elements (REs) linked to transgenes, using viruses like AAV and lentivirus, to drive targeted expression of cargo genes in specific neural cell types, such as those in the cortex and basal ganglia, with high homology to human cell types.
The method achieves selective gene expression in specific neural cell types with minimal off-target effects, enabling research and potential treatments for neurological and psychiatric disorders, and provides tools for studying causal relationships between cell types and behavior.
Smart Images

Figure IMGF000044_0001 
Figure IMGF000046_0001 
Figure IMGF000054_0001
Abstract
Description
[0001] REGULATORY ELEMENTS FOR TARGETED EXPRESSION OF CARGO GENES IN NEURONS
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims benefit of priority from U.S. Provisional Application Serial No. 63 / 558,436, filed February 27, 2024. The disclosure of the prior application is considered part of (and is incorporated by reference in) the disclosure of this application.
[0004] STATEMENT AS TO FEDERALLY SPONSORED RESEARCH
[0005] This invention was made with government support under MH120094 and MH130881 awarded by the National Institutes of Health. The government has certain rights in the invention.
[0006] SEQUENCE LISTING
[0007] This application contains a Sequence Listing that has been submitted electronically as an XML file named 21600-0146W01_SL_ST26.xml. The XML file, created on February 26, 2025, is 162,104 bytes in size. The material in the XML file is hereby incorporated by reference in its entirety.
[0008] TECHNICAL FIELD
[0009] This document relates to methods and materials for expressing one or more desired polypeptides in a specific type or population of cells. For example, this document relates to nucleic acids that contain a cell type-specific RE coupled to a transgene, and to methods for using such nucleic acids to drive or enhance expression of transgenes in specific neural cell types.
[0010] BACKGROUND
[0011] Cell type specific neural circuits are crucial for behavior and cognition. The current gold standard for targeting specific cortical and striatal cell types are transgenic mouse lines. The lack of effective tools for manipulating cell types in nonhuman primates (NHPs) is a significant limitation to understanding the circuit basis of sophisticated cognitive functions.
[0012] SUMMARY
[0013] This document provides methods and materials that can be used to target vertebrate neuron subtypes. In general, the tools provided herein include viruses that express a gene in specific cell types in cognitive and reward systems structures, including the cortex and basal ganglia. The viruses use cell type specific regulatory elements (REs) to drive expression of cargo in cell types that are important for reward, movement, and cognition. These elements are small enough to fit well within the virus packaging limits. The cell type specificity of the REs allows for targeted expression of the genes delivered by the virus. Specificity helps to avoid off target side effects. The distal REs, or enhancers, described herein provide an avenue for manipulating specific neural cell types in mammals such as nonhuman primates (NHPs), and can be used, for example, for research purposes or for treatment of neurological and psychiatric disorders. For example, the cell type specific tools provided herein can enable the study of causal relationships between particular cell types and behavior. Moreover, because the molecular definitions of the cell populations that the REs are designed to target are derived from NHPs, they have a high degree of homology with human cell types and therefore have a high degree of translational potential. In some cases, as described herein, the cell type specific enhancers can be used in concert with Designer Receptors Exclusively Activated by Designer Drugs (DREADDs) (Roth, Neuron, 89:683-694, 2016) to selectively activate or inhibit neuron subtypes underlying a variety of neurological and psychiatric disorders, including Parkinson’s disease and major depressive disorder.
[0014] As demonstrated herein, single nucleus RNA and ATAC seq were used to define NHP neuron types and cell type specific open chromatin regions (OCRs), respectively, in the striatum and prefrontal cortex. Machine learning (ML) algorithms identified potent OCRs as candidate cell type specific enhancers. Enhancers were shown to drive selective transgene expression in cortical layer 3 of mouse and Rhesus monkey, and DI -matrix enhancers were identified in mice. In addition, a potential striosome enhancer infected striatal neurons projecting to potential striosome-dendron bouquets. Very few off target effects were observed, suggesting that the ML-assisted enhancer identification, operating on NHP single cell data, was efficient at identifying cell type specific enhancers.
[0015] In a first aspect, this document features a nucleic acid construct containing (a) a transgene that includes a nucleotide sequence encoding a polypeptide of interest, and (b) at least one RE specific for a selected cell type, wherein the RE includes the nucleotide sequence set forth in any of SEQ ID NOS: 1 to 88, or a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOS: 1 to 88, and wherein the RE is operably linked to the nucleotide sequence encoding the polypeptide of interest, and is effective to drive expression of the polypeptide of interest in the selected cell type. The polypeptide of interest can be a clustered regularly interspaced short palindromic repeats- (CRISPR-) associated (Cas) nuclease or SunlGFP The polypeptide of interest can be a Designer Receptor Exclusively Activated by Designer Drug (DREADD) polypeptide. The DREADD polypeptide can be a hM4Di polypeptide, a hM3Dq polypeptide, a PSAM4-GlyR polypeptide, or a PSAM4-5HT3 polypeptide. The polypeptide of interest can be a channelrhodopsin polypeptide. The nucleic acid construct can further include a nucleotide sequence encoding a tag polypeptide, such that when the nucleotide sequences encoding the polypeptide of interest and the tag polypeptide are expressed, the polypeptide of interest is coupled to the tag polypeptide. The tag polypeptide can be a fluorescent polypeptide. The fluorescent polypeptide can be selected from the group consisting of green fluorescent protein (GFP), nuclear localization sequence-GFP, superfolder GFP, enhanced GFP, mCherry, mCitrine, mRuby, tdTomato, dTomato, SunlGFP, yellow fluorescent protein (YFP), and cyan fluorescent protein (CFP). The fluorescent polypeptide can include an amino acid sequence having at least 95% sequence identity with the superfolder GFP sequence set forth in SEQ ID NO: 100. The RE can be specific for a specific subtype of neuron, such that the transgene is expressed in a majority of neurons of the specific subtype transduced with the nucleic acid construct, but is not expressed in at least 90% of neurons of other subtypes of neurons transduced with the nucleic acid construct. The specific subtype of neuron can include core cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 1, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 1. The specific subtype of neuron can include DI core cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO:2, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:2. The specific subtype of neuron can include DI hybrid cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO:3 or SEQ ID NO:4, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:3 or SEQ ID NO:4. The specific subtype of neuron can include Dl.NUDAP cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:5 or SEQ ID NO:6. The specific subtype of neuron can include DI. Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9. The specific subtype of neuron can include Dl.Striosome cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NOTO or SEQ ID NO: 11. The specific subtype of neuron can include D2 cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 12 or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 12. The specific subtype of neuron can include D2. Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14. The specific subtype of neuron can include dSTR cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 15, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 15. The specific subtype of neuron can include Matrix cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18. The specific subtype of neuron can include Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20. The specific subtype of neuron can include Striosome cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64. The specific subtype of neuron can include DI. Matrix cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30. The specific subtype of neuron can include D2. Matrix cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35. The specific subtype of neuron can include DI cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41. The specific subtype of neuron can include D2 cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51. The specific subtype of neuron can include L3.CUX2.RORB cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:65 to 76, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs 65 to 76. The specific subtype of neuron can include L5.POU3F1 cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88. The nucleic acid construct can further include virus sequences. The virus sequences can be adeno-associated virus (AAV) sequences or lentivirus sequences. In another aspect, this document features a virus particle containing a nucleic acid construct as described herein. The virus particle can be AAV. The nucleic acid construct can include (a) a transgene that includes a nucleotide sequence encoding a polypeptide of interest, and (b) at least one RE specific for a selected cell type, wherein the RE includes the nucleotide sequence set forth in any of SEQ ID NOS:1 to 88, or a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOS: 1 to 88, and wherein the RE is operably linked to the nucleotide sequence encoding the polypeptide of interest, and is effective to drive expression of the polypeptide of interest in the selected cell type. The polypeptide of interest can be a Cas nuclease or SunlGFP The polypeptide of interest can be a DREADD polypeptide. The DREADD polypeptide can be a hM4Di polypeptide, a hM3Dq polypeptide, a PSAM4-GlyR polypeptide, or a PSAM4-5HT3 polypeptide. The polypeptide of interest can be a channelrhodopsin polypeptide. The nucleic acid construct can further include a nucleotide sequence encoding a tag polypeptide, such that when the nucleotide sequences encoding the polypeptide of interest and the tag polypeptide are expressed, the polypeptide of interest is coupled to the tag polypeptide. The tag polypeptide can be a fluorescent polypeptide. The fluorescent polypeptide can be selected from the group consisting of GFP, nuclear localization sequence-GFP, superfolder GFP, enhanced GFP, mCherry, mCitrine, mRuby, tdTomato, dTomato, SunlGFP, YFP, and CFP The fluorescent polypeptide can include an amino acid sequence having at least 95% sequence identity with the superfolder GFP sequence set forth in SEQ ID NO: 100. The RE can be specific for a specific subtype of neuron, such that the transgene is expressed in a majority of neurons of the specific subtype transduced with the nucleic acid construct, but is not expressed in at least 90% of neurons of other subtypes of neurons transduced with the nucleic acid construct. The specific subtype of neuron can include core cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 1, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 1. The specific subtype of neuron can include DI core cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO:2, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:2. The specific subtype of neuron can include DI hybrid cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO:3 or SEQ ID NO:4, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 3 or SEQ ID NON. The specific subtype of neuron can include Dl.NUDAP cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:5 or SEQ ID NO:6. The specific subtype of neuron can include DI. Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9. The specific subtype of neuron can include Dl.Striosome cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11. The specific subtype of neuron can include D2 cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 12 or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 12. The specific subtype of neuron can include D2. Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14. The specific subtype of neuron can include dSTR cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 15, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 15. The specific subtype of neuron can include Matrix cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18. The specific subtype of neuron can include Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20. The specific subtype of neuron can include Striosome cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID N0s:21 to 26 and 52 to 64. The specific subtype of neuron can include DI. Matrix cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30. The specific subtype of neuron can include D2.Matrix cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35. The specific subtype of neuron can include DI cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41. The specific subtype of neuron can include D2 cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51. The specific subtype of neuron can include L3.CUX2.RORB cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:65 to 76, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs 65 to 76. The specific subtype of neuron can include L5.POU3F1 cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88. The nucleic acid construct can further include virus sequences. The virus sequences can be adeno- associated virus (AAV) sequences or lentivirus sequences.
[0016] In another aspect, this document features a method for labeling a selected cell type within a population of different cell types. The method can include, or consist essentially of, introducing into the population of different cell types a nucleic acid construct that contains (a) a transgene that includes a nucleotide sequence encoding a polypeptide of interest and (ii) a sequence encoding a tag polypeptide, wherein expression of the transgene yields the polypeptide of interest coupled to the tag polypeptide, and (b) at least one RE specific for a selected cell type, wherein the RE includes the nucleotide sequence set forth in any of SEQ ID NOS: 1 to 88, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOS: 1 to 88, wherein the RE is operably linked to the nucleotide sequence encoding the polypeptide of interest and the tag polypeptide, and is effective to drive expression of the polypeptide of interest coupled to the tag polypeptide in the selected cell type, and wherein the tagged polypeptide of interest is expressed in and thereby labels the selected cell type. The polypeptide of interest can be a Cas nuclease or SunlGFP The polypeptide of interest can be a DREADD polypeptide. The DREADD polypeptide can be a hM4Di polypeptide, a hM3Dq polypeptide, a PSAM4-GlyR polypeptide, or a PSAM4-5HT3 polypeptide. The polypeptide of interest can be a channelrhodopsin. The tag polypeptide can be a fluorescent polypeptide. The fluorescent polypeptide can be selected from the group consisting of GFP, nuclear localization sequence-GFP, superfolder GFP, enhanced GFP, mCherry, mCitrine, mRuby, tdTomato, dTomato, SunlGFP, YFP, and CFP. The fluorescent polypeptide can include an amino acid sequence having at least 95% sequence identity with the superfolder GFP sequence set forth in SEQ ID NO: 100. The RE can be specific for a specific subtype of neuron, such that the transgene is expressed in a majority of neurons of the specific subtype transduced with the nucleic acid construct, but is not expressed in at least 90% of neurons of other subtypes of neurons transduced with the nucleic acid construct. The specific subtype of neuron can include core cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 1, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 1. The specific subtype of neuron can include DI core cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO:2, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:2. The specific subtype of neuron can include DI hybrid cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO:3 or SEQ ID NO:4, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 3 or SEQ ID NO:4. The specific subtype of neuron can include Dl.NUDAP cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:5 or SEQ ID NO:6. The specific subtype of neuron can include DI. Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9. The specific subtype of neuron can include Dl.Striosome cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11. The specific subtype of neuron can include D2 cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 12 or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 12. The specific subtype of neuron can include D2. Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14. The specific subtype of neuron can include dSTR cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 15, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 15. The specific subtype of neuron can include Matrix cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18. The specific subtype of neuron can include Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20. The specific subtype of neuron can include Striosome cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64. The specific subtype of neuron can include DI. Matrix cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30. The specific subtype of neuron can include D2.Matrix cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35. The specific subtype of neuron can include DI cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41. The specific subtype of neuron can include D2 cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51. The specific subtype of neuron can include L3.CUX2.RORB cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:65 to 76, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs 65 to 76. The specific subtype of neuron can include L5.POU3F1 cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88. The construct can further include virus sequences. The virus sequences can be AAV sequences or lentivirus sequences.
[0017] In another aspect, this document features a method for expressing a transgene in a selected neuron cell type within a population of different neuron cell types. The method can include, or consist essentially of, introducing into the population of different cell types a nucleic acid construct containing (a) a nucleic acid that includes an exogenous transgene that includes a nucleotide sequence encoding a heterologous polypeptide, and (b) at least one RE specific for the selected neuron cell type, wherein the RE includes the nucleotide sequence set forth in any of SEQ ID NOS: 1 to 88, or a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOS: 1 to 88, and wherein the RE is operably linked to the nucleotide sequence encoding the heterologous polypeptide and is effective to drive expression of the heterologous polypeptide in a majority of neurons of the selected neuron cell type into which the nucleic acid construct was introduced, but is not expressed in at least 90% of neurons of other neuron cell types of neurons into which the nucleic acid construct was introduced. The exogenous transgene can further include a nucleotide sequence encoding a tag polypeptide, such that when the nucleotide sequences encoding the heterologous polypeptide and the tag polypeptide are expressed, the heterologous polypeptide is coupled to the tag polypeptide. The tag polypeptide can be a fluorescent polypeptide. The fluorescent polypeptide can be selected from the group consisting of GFP, nuclear localization sequence-GFP, superfolder GFP, enhanced GFP, mCherry, mCitrine, mRuby, tdTomato, dTomato, Sun 1 GFP, YFP, and CFP. The fluorescent polypeptide can include an amino acid sequence having at least 95% sequence identity with the superfolder GFP sequence set forth in SEQ ID NO: 100. The heterologous polypeptide can be a Cas nuclease or Sun 1 GFP. The heterologous polypeptide can be a DREADD polypeptide. The DREADD polypeptide can be a hM4Di polypeptide, a hM3Dq polypeptide, a PSAM4-GlyR polypeptide, or a PSAM4-5HT3 polypeptide. The polypeptide of interest can be a channelrhodopsin. The construct can further include virus sequences. The virus sequences can be AAV sequences or lentivirus sequences. The specific subtype of neuron can include core cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 1, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 1. The specific subtype of neuron can include DI core cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO:2, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:2. The specific subtype of neuron can include DI hybrid cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO:3 or SEQ ID NO:4, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 3 or SEQ ID NO:4. The specific subtype of neuron can include Dl.NUDAP cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 5 or SEQ ID NO: 6, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:5 or SEQ ID NO:6. The specific subtype of neuron can include DI. Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9. The specific subtype of neuron can include Dl.Striosome cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11. The specific subtype of neuron can include D2 cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 12 or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 12. The specific subtype of neuron can include D2. Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14. The specific subtype of neuron can include dSTR cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 15, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 15. The specific subtype of neuron can include Matrix cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18. The specific subtype of neuron can include Shell cells, and the at least one RE can include the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20. The specific subtype of neuron can include Striosome cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64. The specific subtype of neuron can include DI. Matrix cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30. The specific subtype of neuron can include D2.Matrix cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35. The specific subtype of neuron can include DI cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41. The specific subtype of neuron can include D2 cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51. The specific subtype of neuron can include L3.CUX2.RORB cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:65 to 76, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs 65 to 76. The specific subtype of neuron can include L5.POU3F1 cells, and the at least one RE can include the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Although methods and materials similar or equivalent to those described herein can be used to practice the invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting.
[0019] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the invention will be apparent from the description and drawings, and from the claims. DESCRIPTION OF DRAWINGS
[0020] FIGS. 1A-1G indicate how the genomic data was analyzed to identify and prioritize enhancer candidates for cell type-specific labeling. Uniform Manifold Approximation and Projection (UMAP) dimension reduction plots show the molecularly identified cell populations for the striatum (FIG. 1A) and the cortex (FIG. IB) The UMAPs were generated separately for gene expression (left panels) and open chromatin (right panels). Candidate cell type-specific enhancers were identified based on open cell type-specific open chromatin for striatum (FIG. 1C, left panel) and cortex (FIG. ID, left panel). The numbers of enhancers specific to each cell type are visualized in the right panels. FIG. IE shows labels of Rhesus macaque cell types compared to human single cell cortical data, demonstrating a strong conservation of cellular identities. FIGS. IF and 1G include elbow plots visualizing striatal (FIG. IF) and cortical (FIG. 1G) enhancer candidates based on their ranking score, which is a summary of the machine learning model output.
[0021] FIG. 2A is a schematic depicting an in vivo screening strategy used to test for cell type specific enhancers. Enhancer sequences were placed upstream of the HSP68 minimal promoter, GFP transgene, and 500 nt, enhancer specific, DNA barcodes. The DNA barcodes enabled simultaneous testing of expression from multiple candidate enhancers. FIG. 2B includes schematic and representative images illustrating a method for and results of AAV injections into macaque monkey brain. The cartoon at the upper left shows an injection strategy for the macaque prefrontal cortex. An MRI contrast agent (gadolinium) was included in the AAV to enable post-injection location confirmation. The top three images show MRI sections from the same subject, postinjection. The gadolinium-enhanced injections were viewed as bright spots (circled) in the horizontal, coronal, sand sagittal planes. The bottom images show MRI sections as in the top portion, but for striatal injections.
[0022] FIG. 3A includes micrographs of mouse somatosensory cortex following injection of layer 3 enhancer AAVs driving expression of GFP (top) and a control AAV-hSyn driving expression of dTomato (middle). The overlaid images (bottom) showed that GFP expression was restricted to deep layer 2 / 3. FIG. 3B includes micrographs showing that the fluorescent in situ hybridization- (FISH-) labeled GFP transcript was restricted to deep layer 2 / 3. FIG. 3C includes a micrograph showing multichannel FISH labeling of a monkey dorsolateral prefrontal cortex following injection of layer 3 enhancer AAVs. Nulcei (gray) were labeled with DAPI. CUX2 (which fluoresced green) was used as a layer 2 / 3 marker. L3PN Enh.1 (which fluoresced magenta) was used as the DNA barcode for the top-ranked layer 3 enhancer. RORB (blue) was used as a layer 4 marker. The merged image shows the overlap of CUX2, L3PN Enh.l, and RORB, and demonstrates that L3PN Enh.l drove expression in layer 3. FIG. 3D shows validation of the top ranked L3PN enhancer in a monkey. The image at the left is a micrograph of a coronal section of a NHP brain. AAV92YF-L3PN.Enh001-GFP was injected along the ventral bank of the principal sulcus (PS). The injection track is indicated with a white line. The white boxes in the center and right images indicate the region of interest (ROI). The center image shows nuclei labelling in ROI. Dense bands of cells indicate the borders of layers 2 and 4. At the right is a merged image of nuclei and GFP expression, which clearly shows that expression was restricted to L2 / 3. FIG. 3E shows validation of the top ranked L3PN enhancer in a second monkey. The images are analogous to those of FIG. 3D, except that the virus was AAVX1. l-L3PN.Enh001-GFP, and injection was made in the dorsal bank (46d). This demonstrated that the enhancer worked consistently and was not dependent on capsid identity.
[0023] FIG. 4A shows a micrograph of a coronal section through a mouse brain. A control AAV expressing red fluorescent protein was injected into the left striatum (circled). FIG. 4B includes micrographs showing results for a DI enhancer. The image at the upper left is a micrograph of a coronal section of a mouse brain showing GFP positive terminals in GPi (which receives projections from DI neurons) and no such terminals in GPe (which receives projections from D2 neurons). The images at the right are micrographs of FISH labeling for DRD1, DI enhancer Barcode, DRD2, and a merge of those images, as indicated. The DI Enhancer barcode colocalized with DI neurons (arrowheads), but not D2 neurons. FIG. 4C is a graph plotting the percentages of D1+ neurons, D2+ neurons, and non-MSNs that were labeled with barcodes.
[0024] FIG. 5A is an Allen Brain Atlas image showing a coronal section through a mouse midbrain. FIG. 5B is a micrograph of mouse midbrain regions. The dopamine neurons were labeled magenta, and the pars reticulata was labeled blue. The GFP signal between them (boxed) indicated expression in striosome MSNs, which project to striosome-dendron bouquets. FIG. 5C includes micrographs analogous to those shown in FIG. 5B, but from another mouse and labeled with different colors (white for dopamine neurons and red for pars reticulata, with the GFP signal between them indicating expression in striosome MSNs. FIG. 5D includes a micrograph from GFP mouse 2, showing expression in striosome dendron bouquets (left panel). The striosome connection to the bouquets is described elsewhere (see, for example, Crittenden et al., Proc. Natl. Acad. Sci. USA, 113(40): 11318-11323, 2016).
[0025] FIG. 6 shows the amino acid sequences of 88 REs (SEQ ID NOS: 1-88) that were identified as described herein.
[0026] FIG. 7 shows representative nucleotide and amino acid sequences for the hM4Di DREADD (SEQ ID NOS:89 and 90, respectively).
[0027] FIG. 8 shows representative nucleotide and amino acid sequences for the hM3Dq DREADD (SEQ ID NOS:91 and 92, respectively).
[0028] FIG. 9 shows representative nucleotide and amino acid sequences for the PSAM4-GlyR DREADD (SEQ ID NOS:93 and 94, respectively).
[0029] FIG. 10A shows representative nucleotide and amino acid sequences for the PSAM4-5HT3 high conductance DREADD (SEQ ID NOS:95 and 96, respectively). FIG. 10B shows representative nucleotide and amino acid sequences for the PSAM4- 5HT3 low conductance DREADD (SEQ ID NOS:97 and 98, respectively).
[0030] FIG. 11 shows representative nucleotide and amino acid sequences for channelrhodopsin (SEQ ID NOS:99 and 100, respectively).
[0031] FIG. 12 shows representative amino acid sequences for the indicated fluorescent polypeptides (SEQ ID NOS: 101-112, respectively).
[0032] FIGS. 13A and 13B show representative nucleotide (FIG. 13A) and amino acid (FIG. 13B) sequences for the Streptococcus pyogenes Cas polypeptide (SEQ ID NOS: 113 and 114, respectively).
[0033] FIGS. 14A-14G A\ow Rhesus macaque dorsolateral prefrontal cortex (DLPFC) principal neuron types. FIG. 14A is a diagram showing dissected DLPFC regions. The shaded area between aspd and pspd indicates sample location. Axes indicate dorsal (D), ventral (V), rostral (R), caudal (C), medial (M), and lateral (L). FIG. 14B is a UMAP plot showing clustering of neuronal and non-neuronal DLPFC snRNA-Seq data from 5 rhesus monkeys. FIG. 14C is a heatmap showing marker gene expression in each cell class. FIG. 14D is a violin plot showing marker gene expression in each cell class. FIG. 14E is a UMAP plot showing 11 distinct excitatory subtypes in DLPFC. FIG. 14F is a violin plot showing relative normalized marker gene expression in each excitatory cell type. FIG. 14G shows the results of fluorescence in situ hybridization (FISH) profiling of laminar organization of each excitatory cell type. The left enlarged image shows the original FISH signals in the boxed area. The right Nissl image indicates layer boundaries, aspd, anterior supraprincipal dimple; pspd posterior supra-principal dimple; PS: principal sulcus.
[0034] FIGS. 15A-15G show excitatory neurons in primate cortex. FIG. 15A is a UMAP plot showing clustering of neuronal and non-neuronal DLPFC snRNA-Seq data in 5 rhesus monkeys. FIG. 15B is a feature plot of the neuron-specific marker SLC17A7. FIG. 15C includes feature plots of excitatory neuron subtype markers. FIG. 15D is a heat map showing the cosine similarity within and between the eleven types of excitatory neurons. FIG. 15E is a comparison of monkey excitatory neuron subtypes with a human prefrontal cortex (PFC) dataset. FIG. 15F is a comparison of monkey excitatory neuron subtypes with a human medial temporal gyrus (MTG) dataset. FIG. 15G is a comparison of monkey excitatory neuron subtypes with a human primary motor cortex (Ml) dataset.
[0035] FIGS. 16A-16F show inhibitory neurons in primate DLPFC. FIG. 16A shows UMAP visualization of NHP DLPFC inhibitory neuron subtypes, distinguished according to subtypes (left) or individual subjects (right). FIG. 16B is a heat map of differentially expressed genes. FIG. 16C is a heat map showing the cosine similarity within and between the ten subtypes of inhibitory neurons. FIG. 16D includes images showing cell profiler processed marker gene expression in the PFC coronal sections. FIG. 16E includes images showing FISH analysis of 6 major inhibitory cell types. White squares indicate cortical layer borders. FIG. 16F is a pair of graphs plotting cell type specific OCRs before and after ML ranking and selection for inhibitory neurons. The top plot shows the number of OCRs before ML ranking, and the bottom plot shows the number after ML ranking.
[0036] FIGS. 17A-17D show NHP DLPFC enhancer identification using single nuclear open chromatin (snATAC-seq). FIG. 17A is a UMAP plot showing clustering of excitatory neurons in DLPFC snATAC-Seq data from three rhesus monkeys. FIG. 17B includes graphs plotting ML ranking scores for OCRs in each cell type. FIG. 17C is a graph plotting numbers of cell type specific OCRs before ML ranking and selection (left), and after ML ranking and selection (right). FIG. 17D is a track plot showing an example peak for L5 P0U3F1+ neurons.
[0037] FIGS. 18A-18F show the results of screening top enhancer candidates in NHP brain. FIG. 18A is a schematic summary of the screening procedure, with L3PN screening library 1 shown as an example. Each enhancer plasmid and the control plasmid was packaged individually into AAV9-2YF. The top 6 enhancer candidate AAVs and the control AAV were then mixed in equal fractions to create the L3PN screening library 1. FIG. 18B is an image showing post-injection MRI scanning of Gadolinium-doped AAV in the coronal plane, which clearly shows dorsal and ventral injection sites. FIG. 18C includes a series of images from FISH analysis of ten enhancers and positive (hSynapsin promotor) and negative controls (minimal promotor HSP68 alone) driving transgene expression in monkey SM33. FIG. 18D is a graph plotting the distribution of transgene expression in cortical layers. FIG. 18E shows representative FISH of RMacL3-01 driven expression with layer 2 and 3 excitatory neuron marker CUX2 and layer 3 and layer 4 excitatory neuron marker RORB. FIG. 18F shows FISH analysis of L5ET enhancer candidates in monkey SM35.
[0038] FIG. 19 shows results of enhancer screening with FISH. The original FISH images show enhancer expression in individual sections. The numbers below the images indicate brain slice section numbers.
[0039] FIGS. 20A-20E show enhancer expression in rodents. FIG. 20A includes representative images showing L3PN enhancer-driven GFP expression (library 1) and hSynapsin driven dTomato expression in mouse cortex. FIG. 20B includes representative images showing L3PN enhancer-driven GFP expression (library 1) in cortical layers of two additional mice. FIG. 20C shows FISH analysis of enhancer expression of RMacL3-01 and -02, demonstrating that their expression was enriched in layer 2 / 3. FIG. 20D shows L5ET enhancer-driven GFP expression (library 1) in mouse cortex. FIG. 20E shows FISH analysis of RMacL5ET-01 and L5ET marker Fam84b, demonstrating colocalization. FIGS. 21A-21O show results from one-at-a-time validation of top enhancer candidates in NHP brain. FIG. 21A is a schematic showing injection sites in subject SM38 of RMacL3-01 packaged in AAV9-2YF and PHP.eB. FIG. 21B is an image of post-injection MRI scanning of Gadolinium, showing the injection sites of RMacL3- 01 packaged in AAV9-2YF and PHP.eB. FIG. 21C is an image showing enhancer- driven GFP expression for RMacL3-01 packaged in AAV9-2YF and PHP.eB. The inset shows layer-specific PHP.eB expression 1 mm caudal to the injection site. FIG. 21D includes high resolution images of enhancer-driven GFP expression, used for quantification, for RMacL3-01 packaged in AAV9-2YF. FIG. 21E shows layer distribution of enhancer-driven GFP expression for RMacL3-01 packaged in AAV9- 2YF. Dashed lines represent counts of GFP+ cells from individual sections, and the solid line represents mean ± SEM. FIG. 21F is a schematic showing injection sites in subject SM41 of RMacL3-01 packaged in AAV9 XI.1 and 2YF. FIG. 21G illustrates the rostral-caudal distribution of RMacL3 -01 -driven GFP expression. The inset is an enlarged image of the white box area. FIG. 21H includes FISH images aligned with anatomical MRI and with the Nissl sections. FIG. 211 shows a 3-D reconstruction of expression in subject SM41 in MRI space. FIG. 21 J is a graph plotting the layer distribution of enhancer-driven GFP expression for RMacL3-01 in SM41. Dashed lines represent counts of individual sections, and the solid line represents mean ± SEM. FIG. 21K is a schematic showing SM53 injection of RMacL3-01 packaged in AAV9 XI.1 at different titers. FIG. 21L is a graph plotting layer distribution of enhancer-driven GFP expression for RMacL3-01 at medium titer. Other conventions follow FIG. 21 J. FIG. 21M is a graph plotting layer distribution of enhancer-driven GFP expression for RMacL3-01 at high titer. Other conventions follow FIG. 21J. FIG. 21N is an image showing dorsal and ventral banks of the cingulate sulcus with RMacL5ET-driven expression and L5ET marker POU3F1. FIG. 210 includes enlarged views of FIG. 21N showing the colocalization of RMacL5ET-driven expression and L5ET marker POU3F1.
[0040] FIGS. 22A-22E show optogenetic control of enhancer-driven ChR2 in NHP brain. FIG. 22A is a cartoon modification of an MRI scan depicting a 2x3 cm cranial window opened to expose the principal sulcus (PS) and arcuate sulcus (sAS), two months after injection. Injection sites are shown for RMacL3-01 driving ChR2, packaged into AAV9-2YF (rostral, black dots) and AAV9-X1.1 (caudal, white dot). FIG. 22B is an image showing injection sites for RMacL3-01 driving ChR2, packaged into AAV9-2YF (rostral, solid circles) and AAV9-X1.1 (caudal, dashed circle), visualized in the cortex with a handheld UV light and USB camera during the optical stimulation experiment. FIG. 22C shows simultaneous voltage traces recorded on channels 3, 5, 7, 11, and 16 of custom-made 16-channel linear electrode array with four windows. Optical pulse trains (shaded vertical areas) evoked neural activity along the length of the array. FIG. 22D is a graph plotting voltage traces of optically evoked single unit waveforms. The curve is the average waveform. FIG. 22E includes example traces showing that optically evoked activity followed the frequency of the laser stimulus, which was applied at 4, 5, and 10 Hz.
[0041] DETAILED DESCRIPTION
[0042] This document provides methods and materials that can be used, for example, to target vertebrate neuron subtypes. In general, the tools provided herein include viruses that express a gene in specific cell types in cognitive and reward systems structures, including the cortex and basal ganglia. The viruses use cell type specific REs to drive targeted expression of cargo in specific cell types (e.g., in cells that are important for reward, movement, and cognition), with few or no off target side effects. The REs described herein provide an avenue for manipulating specific neural cell types in mammals such as NHPs, and can be used, for example, for research purposes or for treatment of neurological and psychiatric disorders. For example, the cell type specific tools provided herein can enable the study of causal relationships between particular cell types and behavior. Moreover, because the REs are derived from NHPs, they have a high degree of homology with human cell types and therefore have a high degree of translational potential.
[0043] In some cases, this document provides recombinant nucleic acid constructs containing (a) a cell type-specific RE, and (b) a nucleotide sequence encoding a desired polypeptide, where the cell type-specific RE can drive expression of the desired polypeptide in the specific cell type. For example, this document provides nucleic acid constructs containing (a) a RE specific for a particular type of neuron (e.g., core, matrix, shell, striosome, DI, DI core, DI hybrid, DI. matrix, Dl.NUDAP, DI. Shell, Dl.Striosome, D2, D2.matrix, D2.shell, dSTR, L3CUX2.RORB, or L5POU3F1 neurons), and (b) a transgene encoding a desired polypeptide, where the RE is operably linked to the transgene and therefore can drive expression of the desired polypeptide in the selected neuronal cells. In some cases, a recombinant nucleic acid construct provided herein can contain a minimal promoter in addition to the RE and the nucleotide sequence encoding a desired polypeptide.
[0044] The terms “nucleic acid” and “polynucleotide” can be used interchangeably, and refer to both RNA and DNA, including cDNA, genomic DNA, synthetic (e.g., chemically synthesized) DNA, and DNA (or RNA) containing nucleic acid analogs. Polynucleotides can have any three-dimensional structure. A nucleic acid can be double-stranded or single-stranded (i.e., a sense strand or an antisense single strand). Non-limiting examples of polynucleotides include genes, gene fragments, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers, as well as nucleic acid analogs.
[0045] As used herein, “isolated,” when in reference to a nucleic acid, refers to a nucleic acid that is separated from other nucleic acids that are present in a genome, including nucleic acids that normally flank one or both sides of the nucleic acid in the genome. The term “isolated” as used herein with respect to nucleic acids also includes any non-naturally occurring sequence, since such non-naturally occurring sequences are not found in nature and do not have immediately contiguous sequences in a naturally occurring genome.
[0046] An isolated nucleic acid can be, for example, a DNA molecule, provided one of the nucleic acid sequences normally found immediately flanking that DNA molecule in a naturally occurring genome is removed or absent. Thus, an isolated nucleic acid includes, without limitation, a DNA molecule that exists as a separate molecule (e.g., a chemically synthesized nucleic acid, or a cDNA or genomic DNA fragment produced by PCR or restriction endonuclease treatment) independent of other sequences, as well as DNA that is incorporated into a vector, an autonomously replicating plasmid, a virus (e.g., a pararetrovirus, a retrovirus, lentivirus, adenovirus, or herpes virus), or the genomic DNA of a prokaryote or eukaryote. In addition, an isolated nucleic acid can include a recombinant nucleic acid such as a DNA molecule that is part of a hybrid or fusion nucleic acid. A nucleic acid existing among hundreds to millions of other nucleic acids within, for example, cDNA libraries or genomic libraries, or gel slices containing a genomic DNA restriction digest, is not to be considered an isolated nucleic acid.
[0047] A nucleic acid can be made by, for example, chemical synthesis or polymerase chain reaction (PCR). PCR refers to a procedure or technique in which target nucleic acids are amplified. PCR can be used to amplify specific sequences from DNA as well as RNA, including sequences from total genomic DNA or total cellular RNA. Various PCR methods are described, for example, in PCR Primer: A Laboratory Manual, Dieffenbach and Dveksler, eds., Cold Spring Harbor Laboratory Press, 1995. Generally, sequence information from the ends of the region of interest or beyond is employed to design oligonucleotide primers that are identical or similar in sequence to opposite strands of the template to be amplified. Various PCR strategies also are available by which site-specific nucleotide sequence modifications can be introduced into a template nucleic acid.
[0048] The term “RE” (which may be used interchangeably with the terms “regulatory region,” “control element,” and “expression control sequence”) refers to a nucleotide sequence that influences transcription or translation initiation and rate. REs can include, without limitation, promoter sequences, enhancer sequences, response elements, protein recognition sites, inducible elements, promoter control elements, protein binding sequences, 5' and 3' untranslated regions (UTRs), and / or transcriptional start sites.
[0049] As used herein, “operably linked” means incorporated into a genetic construct so that expression control sequences effectively control expression of a coding sequence of interest. A coding sequence is “operably linked” and “under the control” of expression control sequences in a cell when RNA polymerase is able to transcribe the coding sequence into RNA, which if an mRNA, then can be translated into the protein encoded by the coding sequence. Thus, a RE can modulate (e.g., enhance, regulate, facilitate, or drive) transcription in the cell in which it is desired to express a selected nucleic acid. For example, a cell-specific RE that confers transcription only or predominantly in a particular cell type (e.g., a particular type of neuron) can be used. Such REs can be identified using methods such as those described herein and can be used in the nucleic acid constructs provided herein.
[0050] In some cases, a RE described herein can serve as a promoter. In some cases, a regulator element described herein can be used in combination with a promoter (e.g., a minimal promoter that allows for formation of a transcription initiation complex). A promoter is an expression control sequence composed of a region of a DNA molecule, typically (but not always) within 100 nucleotides upstream of the point at which transcription starts (generally near the initiation site for RNA polymerase II). Promoters are involved in recognition and binding of RNA polymerase and other to initiate and modulate transcription. To bring a coding sequence under the control of a promoter, it typically is necessary to position the translation initiation site of the translational reading frame of the polypeptide between one and about fifty nucleotides downstream of the promoter. A promoter can, however, be positioned as much as about 5,000 nucleotides upstream of the translation start site, or about 2,000 nucleotides upstream of the transcription start site. A promoter typically includes at least a core (basal) promoter. A promoter also may include at least one control element such as an upstream element. Such elements include upstream activation regions (UARs) and, optionally, other DNA sequences that affect transcription of a polynucleotide such as a synthetic upstream element.
[0051] Any appropriate cell type-specific RE can be included in the nucleic acid constructs provided herein. An RE can have any appropriate length. For example, an RE can have a length from about 100 nucleotides to about 1000 nucleotides (e.g., from about 100 to about 150, from about 125 to about 175, from about 150 to about 200, from about 200 to about 250, from about 250 to about 300, from about 300 to about 350, from about 350 to about 400, from about 400 to about 450, from about 450 to about 500, from about 450 to about 550, from about 550 to about 600, from about 600 to about 650, from about 650 to about 700, from about 700 to about 800, from about 800 to about 900, or from about 900 to about 1000 nucleotides).
[0052] Non-limiting examples of suitable REs are set forth in SEQ ID NOS: 1-88 herein (see, FIG. 6). For example, a RE can have the sequence set forth in any of SEQ ID NOS: 1-88. In some cases, a RE can have the sequence set forth in any of SEQ ID NOS:26, 65, 66, 67, 68, and 77. In some cases, a RE can have a sequence that is at least 75% identical (e.g., at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical) to the sequence set forth in any of SEQ ID NOS: 1-88. For example, a RE can have a nucleotide sequence that is at least 90% identical to the nucleotide sequence set forth in any of SEQ ID NOS: 1-88, at least 95% identical to the nucleotide sequence set forth in any of SEQ ID NOS: 1-88, or at least 98% identical to the nucleotide sequence set forth in any of SEQ ID NOS: 1-88. In some cases, a RE can have a nucleotide sequence that is at least 90% identical to the nucleotide sequence set forth in any of SEQ ID NOS: 26, 65, 66, 67, 68, and 77, at least 95% identical to the nucleotide sequence set forth in any of SEQ ID NOS: 26, 65, 66, 67, 68, and 77, or at least 98% identical to the nucleotide sequence set forth in any of SEQ ID NOS: 26, 65, 66, 67, 68, and 77.
[0053] The percent sequence identity between a particular nucleic acid or amino acid sequence and a nucleic acid or amino acid sequence referenced by a particular sequence identification number is determined as follows. First, a nucleic acid or amino acid sequence is compared to the sequence set forth in a particular sequence identification number using the BLAST 2 Sequences (B12seq) program from the stand-alone version of BLASTZ containing BLASTN version 2.0.14 and BLASTP version 2.0.14. This stand-alone version of BLASTZ can be obtained from Fish & Richardson’s web site (e.g., www.fr.com / blast / ) or the U.S. government’s National Center for Biotechnology Information web site (www.ncbi.nlm.nih.gov). Instructions explaining how to use the B12seq program can be found in the readme file accompanying BLASTZ. B12seq performs a comparison between two sequences using either the BLASTN or BLASTP algorithm. BLASTN is used to compare nucleic acid sequences, while BLASTP is used to compare amino acid sequences. To compare two nucleic acid sequences, the options are set as follows: -i is set to a file containing the first nucleic acid sequence to be compared (e.g., C:\seql.txt); -j is set to a file containing the second nucleic acid sequence to be compared (e.g., C:\seq2.txt); - p is set to blastn; -o is set to any desired file name (e.g., C:\output.txt); -q is set to -1; - r is set to 2; and all other options are left at their default setting. For example, the following command can be used to generate an output file containing a comparison between two sequences: C:\B12seq -i c:\seql.txt -j c:\seq2.txt -p blastn -o c:\output.txt -q -1 -r 2. To compare two amino acid sequences, the options of B12seq are set as follows: -i is set to a file containing the first amino acid sequence to be compared (e.g., C:\seql.txt); -j is set to a file containing the second amino acid sequence to be compared (e.g., C:\seq2.txt); -p is set to blastp; -o is set to any desired file name (e.g., C:\output.txt); and all other options are left at their default setting. For example, the following command can be used to generate an output file containing a comparison between two amino acid sequences: C:\B12seq -i c:\seql.txt -j c:\seq2.txt -p blastp -o c:\output.txt. If the two compared sequences share homology, then the designated output file will present those regions of homology as aligned sequences. If the two compared sequences do not share homology, then the designated output file will not present aligned sequences.
[0054] Once aligned, the number of matches is determined by counting the number of positions where an identical nucleotide or amino acid residue is presented in both sequences. A matched position refers to a position in which an identical nucleotide or amino acid residue occurs at the same position in aligned sequences. The percent sequence identity is determined by dividing the number of matches by the length of the sequence set forth in the identified sequence (e.g., SEQ ID NO: 1), followed by multiplying the resulting value by 100. For example, an amino acid sequence that has 480 matches when aligned with the sequence set forth in SEQ ID NO: 1 is 95.8 percent identical to the sequence set forth in SEQ ID NO: 1 (i.e., 480 501 x 100 = 95.8). It is noted that the percent sequence identity value is rounded to the nearest tenth. For example, 75.11, 75.12, 75.13, and 75.14 are rounded down to 75.1, while 75.15, 75.16, 75.17, 75.18, and 75.19 are rounded up to 75.2. It also is noted that the length value will always be an integer.
[0055] Any appropriate method can be used to identify cell type-specific REs for inclusion in the nucleic acid constructs provided herein. As described in Example 1 herein, for example, SNAIL methodology was then used to identify a number of REs as being likely to selectively target particular types of neurons. SNAIL methodology is described in detail in, for example, WO 2020 / 257520 (see, e.g., pages 4-6 and 21- 26). In addition to a cell type-specific RE, the nucleic acid constructs described herein can contain a nucleotide sequence encoding a desired polypeptide. The term “polypeptide” as used herein refers to a compound of two or more subunit amino acids, regardless of post-translational modification (e.g., phosphorylation or glycosylation). The subunits may be linked by peptide bonds or other bonds such as, for example, ester or ether bonds. The term “amino acid” refers to either natural and / or unnatural or synthetic amino acids, including D / L optical isomers.
[0056] A nucleic acid construct provided herein can contain any appropriate coding sequence for a polypeptide that is to be expressed in a particular type of cells (e.g., core, matrix, shell, striosome, DI, DI core, DI hybrid, DI. matrix, Dl.NUDAP, DI. Shell, DI. Striosome, D2, D2.matrix, D2.shell, dSTR, L3CUX2.RORB, or L5POU3F1 neurons). In some cases, a polypeptide encoded by a nucleic acid construct provided herein can be a DREADD polypeptide. DREADD polypeptides are a class of artificially engineered protein receptors that can be selectively activated by certain ligands. Examples of DREADD polypeptides include, without limitation, hM4Di, hM3Dq, PSAM4-GlyR, and PSAM4-5HT3. For example, the PSAM4 DREADDs are fusion proteins that include a mutated ligand binding domain from the human nicotinic alpha 7 receptor and an ion channel domain from the human glycine receptor (GlyR) or the human 5HT3 receptor (which has high conductance and low conductance forms). The mutations (e.g., Y115F, Q79R, Q139G, Q139V, Q139W, Q139Y, L141A, L141Q, and / or L 14 IS) in the ligand binding domain can decrease the binding affinity of the endogenous ligand acetylcholine. See, e.g., Magnus et al., Science, 364(6436):eaav5282, 2019. The PSAM4 DREADDs bind with relatively high affinity and selectivity to varenicline (an FDA approved smoking cessation aid that can block the effects of nicotine in the brain), as well as to novel ligands described elsewhere (see, e.g., Magnus et al., supra). hM3Dq and hM4Di are forms of human muscarinic receptors that have been mutated to have a much lower affinity for acetylcholine and a higher affinity for clozapine and related compounds. hM3Dl and hM4Di are second messenger coupled receptors that increase or decrease the membrane potential, respectively, by causing the modification of ion channels, which results in the activation or inhibition, respectively, of neurotransmitter release. See, e.g., Armbruster et al., Proc Natl Acad Sci USA, 104(12):5163-5168, 2007; Roth, Neuron, 89:694-693, 2016; and Gomez et al., Science, 357(6350):503-507, 2017.
[0057] An exemplary nucleotide sequence encoding a hM4Di polypeptide is set forth in SEQ ID NO:89, and an exemplary hM4Di amino acid sequence is set forth in SEQ ID NO: 90 (FIG. 7). In some cases, a polypeptide coding sequence in a nucleic acid construct provided herein can have a nucleotide sequence that is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the nucleotide sequence set forth in SEQ ID NO: 89. In some cases, a polypeptide encoded by a nucleic acid construct provided herein can have an amino acid sequence that is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the amino acid sequence set forth in SEQ ID NO:90.
[0058] An exemplary nucleotide sequence encoding a hM3Dq polypeptide is set forth in SEQ ID NO:91, and an exemplary hM3Dq amino acid sequence is set forth in SEQ ID NO: 92 (FIG. 8). In some cases, a polypeptide coding sequence in a nucleic acid construct provided herein can have a nucleotide sequence that is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the nucleotide sequence set forth in SEQ ID NO:91. In some cases, a polypeptide encoded by a nucleic acid construct provided herein can have an amino acid sequence that is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the amino acid sequence set forth in SEQ ID NO:92.
[0059] An exemplary nucleotide sequence encoding a PSAM4-GlyR polypeptide is set forth in SEQ ID NO:93, and an exemplary PSAM4-GlyR amino acid sequence is set forth in SEQ ID NO: 94 (FIG. 9). In some cases, a polypeptide coding sequence in a nucleic acid construct provided herein can have a nucleotide sequence that is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the nucleotide sequence set forth in SEQ ID NO: 93. In some cases, a polypeptide encoded by a nucleic acid construct provided herein can have an amino acid sequence that is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the amino acid sequence set forth in SEQ ID NO:94. An exemplary nucleotide sequence encoding a PSAM4-5HT3 high conductance (HC) polypeptide is set forth in SEQ ID NO:95, and an exemplary PSAM4-5HT3 HC amino acid sequence is set forth in SEQ ID NO:96 (FIG. 10A). An exemplary nucleotide sequence encoding a PSAM4-5HT3 low conductance (LC) polypeptide is set forth in SEQ ID NO:97, and an exemplary PSAM4-5HT3 amino acid sequence is set forth in SEQ ID NO: 98 (FIG. 10B). In some cases, a polypeptide coding sequence in a nucleic acid construct provided herein can have a nucleotide sequence that is at least 90% (e.g., at least at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, 96%, at least 97%, at least 98%, or at least 99%) identical to the nucleotide sequence set forth in SEQ ID NO: 95. In some cases, a polypeptide encoded by a nucleic acid construct provided herein can have an amino acid sequence that is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the amino acid sequence set forth in SEQ ID NO:96. In some cases, a polypeptide coding sequence in a nucleic acid construct provided herein can have a nucleotide sequence that is at least 90% (e.g., at least at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, 96%, at least 97%, at least 98%, or at least 99%) identical to the nucleotide sequence set forth in SEQ ID NO:97. In some cases, a polypeptide encoded by a nucleic acid construct provided herein can have an amino acid sequence that is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the amino acid sequence set forth in SEQ ID NO:98.
[0060] In some cases, a polypeptide encoded by a nucleic acid construct provided herein can be a channelrhodopsin polypeptide. The channelrhodopsins are a subfamily of retinylidene proteins (rhodopsins) that function as light-gated ion channels and serve as sensory photoreceptors in unicellular green algae. Exemplary nucleotide coding sequences and amino acid sequences for a representative channelrhodopsin are set forth in SEQ ID NOS:99 and 100, respectively (FIG. 11). Other examples of channelrhodopsins include Jaws, Chrimson, Chronos, ChETA, and step-function opsin. In some cases, a polypeptide coding sequence in a nucleic acid construct provided herein can have a nucleotide sequence that is at least 90% (e.g., at least at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, 96%, at least 97%, at least 98%, or at least 99%) identical to the nucleotide sequence set forth in SEQ ID NO:99. In some cases, a polypeptide encoded by a nucleic acid construct provided herein can have an amino acid sequence that is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the amino acid sequence set forth in SEQ ID NO: 100.
[0061] In some cases, a polypeptide encoded by a nucleic acid construct provided herein can be a tag (also referred to as a detectable marker). In some cases, the nucleic acid constructs provided herein can contain a nucleic acid sequence encoding a polypeptide as described above and a nucleic acid sequence encoding a polypeptide tag. In such cases, the nucleotide sequence encoding the polypeptide tag can be linked to the polypeptide coding sequence, such that expression of the polypeptide and tag coding sequences results in a fusion polypeptide containing the desired polypeptide coupled to the tag. The inclusion of a tag, either on its own or coupled to another polypeptide, can enable cells and / or nuclei containing the expressed tag (either on its own or as part of a fusion polypeptide) to be identified and isolated away from cells and / or nuclei that do not express the tag.
[0062] Any appropriate polypeptide tag can be encoded by the nucleic acid constructs provided herein. In some cases, a polypeptide tag can be a fluorescent polypeptide. For example, a nucleic acid construct can include a nucleotide sequence encoding green fluorescent protein (GFP; SEQ ID NO: 101), or a modified GFP such as a superfolder GFP (e.g., the superfolder GFP having the amino acid sequence set forth in SEQ ID NO: 102 or having an amino acid sequence at least 95% identical to the amino acid sequence set forth in SEQ ID NO: 102). Other examples of fluorescent polypeptide tags include, without limitation, mCherry (SEQ ID NO: 103), mCitrine (SEQ ID NO: 104), m-Ruby (SEQ ID NO: 105), nuclear localization sequence-GFP (SEQ ID NO: 106), enhanced GFP (eGFP; SEQ ID NO: 107), tdTomato (SEQ ID NO: 108), dTomato (SEQ ID NO: 109), yellow fluorescent protein (YFP; SEQ ID NO: 110), cyan fluorescent protein (CFP; SEQ ID NO: 111), and Sun 1 GFP (SEQ ID NO: 112). Representative amino acid sequences for these fluorescent polypeptides are set forth in FIG. 12.
[0063] In some cases, a polypeptide encoded by a nucleic acid construct provided herein can be a clustered regularly interspaced short palindromic repeats- (CRISPR-) associated (Cas) nuclease. The CRISPR / Cas system includes components of a prokaryotic adaptive immune system that is functionally analogous to eukaryotic RNA interference, using RNA base pairing to direct DNA or RNA cleavage. The Cas protein functions as an endonuclease, and CRISPR RNA (crRNA) and tracer RNA (tracrRNA) sequences complex with the Cas enzyme and direct it to a target DNA sequence. See, e.g., Makarova et al., Nat Rev Microbiol 9(6):467-477, 2011. In some cases, crRNA and tracrRNA can be engineered as a single cr / tracrRNA hybrid (also referred to as a “guide RNA” or “gRNA”) to direct Cas9 cleavage activity (see, e.g., Jinek et al., Science, 337(6096):816-821, 2012). The combination of Cas, crRNA, and tracrRNA (or Cas and gRNA) can then cleave linear or circular dsDNA targets that are complementary to a spacer within the CRISPR cluster. By pairing an RE for a particular type of neuron with CRISPR / Cas components targeted to a particular gene, genes within the selected type of neurons can be edited or inactivated when the RE drives or enhances expression of the CRISPR RNAs and the Cas nuclease in the selected neurons. For example, the vesicular glutamate transporter 2 gene (VLGUT2) in excitatory neurons can be targeted using (a) an RE that promotes or enhances expression in excitatory neurons in combination with (b) CRISPR RNA targeted to one or more VGLUT2 sequences.
[0064] The homology region within the crRNA sequence (the sequence that targets the crRNA to a desired DNA sequence) can be from about 10 to about 40 (e.g., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40) nucleotides in length. The tracrRNA hybridizing region within each crRNA sequence can be from about 8 to about 20 (e.g., 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) nucleotides in length. The overall length of a crRNA sequence can be, for example, from about 20 to about 80 (e.g., 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, or 80) nucleotides, while the overall length of a tracrRNA can be, for example, from about 10 to about 30 (e.g., 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, or 30) nucleotides. The overall length of a gRNA sequence, which includes a homology region and a stem loop region that contains a crRNA / tracrRNA hybridizing region and a linker-loop sequence, can be from about 30 to about 110 (e.g., 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, or 130) nucleotides. A CRISPR / Cas system can be targeted to any appropriate gene within any appropriate type of neuron. In some cases, for example, the HTT gene, the presenilin 1 gene, the presenilin 2 gene, or the APP gene can be targeted in neurons by a CRISPR / Cas system in mammals (e.g., humans) having Huntington’s disease.
[0065] A representative nucleic acid sequence encoding a Cas polypeptide (in particular, a Cas9 polypeptide from Streptococcus pyogenes is set forth in SEQ ID NO: 113, and a representative Cas9 amino acid sequence is set forth in SEQ ID NO: 114 (FIGS. 13A and 13B). See, also, NCBI Ref. NC_017053.1 and GENBANK® accession no. AKP81606.1 for the nucleotide and amino acid sequences, respectively. It is to be noted, however, that there are multiple types of Cas nucleases (e.g., Casl2a, Cas3, and Cas 10) that can be used in the methods provided herein.
[0066] The polypeptides encoded by the nucleic acid constructs provided herein can have any appropriate length. For example, a polypeptide encoded by a nucleic acid construct provided herein can be between about 20 amino acids and about 750 amino acids (e.g., about 20 to 50 amino acids, about 50 to about 100 amino acids, about 100 to about 200 amino acids, about 200 to about 400 amino acids, about 400 to about 600 amino acids, or about 600 to about 750 amino acids) in length.
[0067] In some cases, a recombinant nucleic acid provided herein can integrate into the genome of a cell. For example, a recombinant nucleic acid provided herein can integrate into the genome of a cell via illegitimate (random, non-homologous, nonsite-specific) recombination. In some cases, a recombinant nucleic acid provided herein can be designed to integrate into the genome of a cell via homologous recombination. Nucleic acid sequences designed for integration via homologous recombination can be flanked on both sides with sequences that are similar or identical to endogenous target nucleotide sequences, which can facilitate integration of the recombinant nucleic acid at the particular site(s) in the genome containing the endogenous target nucleotide sequences. In some cases, nucleic acid sequences adapted for integration via homologous recombination also can include a recognition site for a sequence-specific nuclease. Alternatively, the recognition site for a sequence-specific nuclease can be located in the genome of the cell to be transformed.
[0068] In some cases, a nucleic acid construct can be included in a vector suitable for transformation of cells. Recombinant vectors can be made using, for example, standard recombinant DNA techniques (see, e.g., Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor, NY). This document also provides recombinant nucleic acid constructs (e.g., vectors) containing the REs and polypeptide coding sequences described herein. A “vector” is a replicon, such as a plasmid, phage, or cosmid, into which another DNA segment may be inserted so as to bring about the replication of the inserted segment. Vector backbones include, for example, plasmids, viruses, artificial chromosomes, bacterial artificial chromosomes (BACs), yeast artificial chromosomes (YACs), and phage artificial chromosomes (PACs), as well as RNA vectors, and linear or circular DNA or RNA molecules that include chromosomal, non-chromosomal, semi-synthetic, or synthetic nucleic acids. Vectors include those capable of autonomous replication (episomal vectors) and / or expression of nucleic acids to which they are linked (expression vectors). Generally, a vector is capable of replication when associated with the proper control elements. The term “vector” includes cloning and expression vectors, as well as viral vectors and integrating vectors. An “expression vector” is a vector that includes one or more expression control sequences to control and regulate the transcription and / or translation of another DNA sequence. Suitable expression vectors include, without limitation, plasmids and viral vectors derived from, for example, bacteriophage, baculoviruses, tobacco mosaic virus, herpes viruses, cytomegalovirus, retroviruses, vaccinia viruses, adenoviruses, and adeno-associated viruses. Numerous vectors and expression systems are commercially available.
[0069] Viral vectors include, without limitation, retrovirus (e.g., lentivirus), adenovirus, parvovirus (e.g., adeno associated viruses), coronavirus, negative strand RNA viruses such as orthomyxovirus (e.g., influenza virus), rhabdovirus (e.g., rabies and vesicular stomatitis virus), paramyxovirus (e.g., measles and Sendai), positive strand RNA viruses such as picornavirus and alphavirus, and double-stranded DNA viruses including adenovirus, herpesvirus (e.g., Herpes Simplex virus types 1 and 2, Epstein-Barr virus, cytomegalovirus), and poxvirus (e.g., vaccinia, fowlpox and canarypox). Other viruses include Norwalk virus, togavirus, flavivirus, reoviruses, papovavirus, hepadnavirus, and hepatitis virus, for example. Examples of retroviruses include avian leukosis-sarcoma, mammalian C-type, B-type viruses, D type viruses, HTLV-BLV group, lentivirus, spumavirus (Coffin, “Retroviridae: The viruses and their replication,” in Fundamental Virology, Third Edition, B. N. Fields, et al., eds., Lippincott-Raven Publishers, Philadelphia, 1996).
[0070] Without being bound by any particular mechanism of action, viral delivery of the nucleic acid constructs provided herein can provide flexibility across cell types and species. When using a virus (e.g., adeno-associated virus or lentivirus) delivery method, the nucleic acid constructs can be introduced into any appropriate mammals (e.g., humans, NHPs, mice, rats, sheep, pigs, or dogs) through intravenous injection, direct injection into the brain (e.g., direct injection into the brain parenchyma), or any other appropriate method. This delivery procedure can provide a time and resourceefficient way to introduce nucleic acids into particular cell populations, particularly when compared with the intricacies of transgenic breeding. The method also can minimize the number of collateral animals that are bred for an experiment but cannot be used due to undesirable genotypes.
[0071] This document also provides methods for using the nucleic acid constructs described herein to label and isolate nuclei and / or cells of a selected type. The methods can include, for example providing a nucleic acid construct described herein, introducing the construct into a population of cells, and culturing or incubating the cells under conditions in which a polypeptide encoded by the construct is expressed in the selected cell type (and is not expressed in cells that are not of the selected type). In some cases (e.g., when the encoded polypeptide is fused to a tag), expression of the polypeptide can result in labeling of the selected cell type. It is to be noted that in some cases, the population of cells is within a mammal.
[0072] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural core cells, where the construct includes a RE having the nucleotide sequence set forth in SEQ ID NO: 1, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in SEQ ID NO: 1.
[0073] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural DI core cells, where the construct includes a RE having the nucleotide sequence set forth in SEQ ID NO:2, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in SEQ ID NO:2.
[0074] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural DI hybrid cells, where the construct includes a RE having the nucleotide sequence set forth in SEQ ID NO: 3 or SEQ ID NON, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in SEQ ID NON or SEQ ID NON.
[0075] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural Dl.NUDAP cells, where the construct includes a RE having the nucleotide sequence set forth in SEQ ID NON or SEQ ID NON, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in SEQ ID NON or SEQ ID NON.
[0076] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural DI. Shell cells, where the construct includes a RE having the nucleotide sequence set forth in SEQ ID NON, SEQ ID NON, or SEQ ID NON, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in SEQ ID NON, SEQ ID NON, or SEQ ID NON.
[0077] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural Dl.Striosome cells, where the construct includes a RE having the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11.
[0078] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural D2 cells, where the construct includes a RE having the nucleotide sequence set forth in SEQ ID NO: 12 or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in SEQ ID NO: 12.
[0079] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural D2. Shell cells, where the construct includes a RE having the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14.
[0080] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural dSTR cells, where the construct includes a RE having the nucleotide sequence set forth in SEQ ID NO: 15, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in SEQ ID NO: 15.
[0081] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural matrix cells, and where the construct includes a RE having the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18.
[0082] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural shell cells, where the construct includes a RE having the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20.
[0083] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural striosome cells, where the construct includes a RE having the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in any of SEQ ID N0s:21 to 26 and 52 to 64.
[0084] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural DI. Matrix cells, where the construct includes a RE having the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30.
[0085] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural D2. Matrix cells, where the construct includes a RE having the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35.
[0086] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural DI cells, where the construct includes a RE having the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41.
[0087] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural D2 cells, where the construct includes a RE having the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51.
[0088] In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural L3.CUX2.RORB cells, where the construct includes a RE having the nucleotide sequence set forth in any of SEQ ID NOs:65 to 76, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in any of SEQ ID NOs 65 to 76. In some cases, a method provided herein can include introducing a nucleic acid construct into a population of cells that includes neural L5.POU3F1 cells, where the construct includes a RE having the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88, or a nucleotide sequence with at least 75% sequence identity (e.g., at least 80%, at least 85%, at least 90%, or at least 95% sequence identity) to the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88.
[0089] Any appropriate method can be used to introduce a nucleic acid construct provided herein into a population of cells. As used herein, “transformed” and “transfected” encompass the introduction of one or more nucleic acid molecules (e.g., one or more expression vectors) into a cell by any of a number of techniques. Suitable methods for transforming and transfecting host cells can be found, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual (2nd edition), Cold Spring Harbor Laboratory, New York (1989). For example, calcium phosphate precipitation, electroporation, heat shock, lipofection, microinjection, and virus- mediated nucleic acid transfer can be used introduce nucleic acid molecules into cells. In addition, naked DNA can be delivered directly to cells in vivo (see, e.g., U.S. Patent Nos. 5,580,859 and 5,589,466). The isolated nucleic acid molecule transformed into a host cell can be integrated into the genome of the cell or maintained in an episomal state. Thus, host cells can be stably or transiently transfected with a construct containing an isolated nucleic acid molecule provided herein. In some cases, one or more nucleic acid constructs can be incorporated into virus particles (e.g., AAV or lentivirus particles), and the virus particles can be delivered to cells in order to transfer their nucleic acid contents to the cells. When a nucleic acid construct is introduced into a population of cells within a mammal, the construct can be delivered into the cells by intrathecal injection, systemic transduction, or intraspinal injection.
[0090] This document also provides methods for modifying a selected cell type (e.g., core, matrix, shell, striosome, DI, DI core, DI hybrid, DI. matrix, D1.NUDAP, DI. Shell, DI. Striosome, D2, D2.matrix, D2.shell, dSTR, L3CUX2.RORB, or L5POU3F1 neurons) within a population of different cell types (e.g., neurons within the brain). A method can include, for example, introducing a nucleic acid construct provided herein into a population of cells that includes different cell types. The nucleic acid construct can include a sequence encoding a desired polypeptide, and at least one RE (e.g., one, two, three, or more than three REs) that is operably linked to the sequence encoding the polypeptide and is specific for a particular type of cell within the population. In general, the RE can be effective to drive expression of the sequence encoding the polypeptide in the selected cell type, such that the polypeptide is expressed in the selected cell type. For example, the selected cell type can be DI. matrix neurons, and the RE can be specific for DI. matrix neurons. The use of a RE specific for a particular type of neurons (e.g., DI. matrix neurons) means that the sequence encoding the polypeptide is expressed in a majority of the targeted neurons into which the nucleic acid construct was introduced but is not expressed in at least 90% (e.g., at least 95%, at least 97%, at least 98%, or at least 99%) of neurons of other subtypes into which the nucleic acid construct was introduced. In some cases, the RE can include the nucleotide sequence set forth in any of SEQ ID NOS: 1-88. In some cases, the RE can include a nucleotide sequence that is at least 90% (e.g., at least 95%) identical to the nucleotide sequence set forth in any of SEQ ID NOS: 1-88.
[0091] In addition, this document provides methods for selectively activating or inhibiting particular subtypes of neurons that underlie neurological and / or psychiatric disorders (e.g., Parkinson’s disease, major depressive disorder, Huntington’s disease, epilepsy, Alzheimer’s disease, and abuse / sub stance use disorders) in a mammal (e.g., a human or a NHP). In some cases, a method provided herein can be used to treat a mammal having Parkinson’s disease by, for example, stimulating or repressing specific implicated cell types of the motor circuit (e.g., direct pathway MSN or indirect pathway medium spiny neurons and / or Betz cells). In some cases, a method provided herein can be used to treat a mammal having Huntington’s disease by, for example, editing or suppressing expression of genes (e.g., HTT) that cause or modulate the disease in specific susceptible populations of cells (e.g., indirect pathway MSN and / or striosome MSNs). In some cases, a method provided herein can be used to treat a mammal having epilepsy by, for example, activating or suppressing specific subtypes of neurons (e.g., pyramidal neurons) to modulate overall neural activity. In some cases, a method provided herein can be used to treat a mammal having Alzheimer’s disease by, for example, increasing expression of neuroprotective genes in vulnerable cell types (e.g., layer 3 pyramidal neurons) and / or by selectively stimulating vulnerable cell types to relieve symptoms. In some cases, a method provided herein can be used to treat a mammal having an addiction or substance use disorders by, for example, modulating (e.g., stimulating or suppressing) specific cell populations that underlie addiction behavior (e.g., shell medium spiny neurons and / or core / matrix medium spiny neurons).
[0092] A method provided herein can include, for example, administering a nucleic acid construct provided herein to a mammal (e.g., a human or a NHP), where the nucleic acid construct contains a sequence encoding a polypeptide, and at least one RE (e.g., one, two, three, or more than three REs) that is operably linked to the sequence encoding the polypeptide and is specific for a selected cell type. In some cases, the administering can include introducing a virus containing the nucleic acid construct into the mammal. Again, the RE can be effective to drive expression of the sequence encoding the polypeptide in the selected cell type, such that the polypeptide is expressed in the selected cell type. For example, the selected cell type can be core neurons and the RE can be specific for core neurons, or the selected cell type can be DI neurons and the RE can be specific for DI neurons, or the selected cell type can be DI. shell neurons and the RE can be specific for DI. shell neurons, etc. In some cases, the RE can include the nucleotide sequence set forth in any of SEQ ID NOS: 1- 88. In some cases, the RE can include a nucleotide sequence that is at least 90% (e.g., at least 95%) identical to the nucleotide sequence set forth in any of SEQ ID NOS: 1- 88.
[0093] In some cases, for example, excitatory neuron subtypes that are important for transmission of mechanical allodynia (painful sensation caused by innocuous stimuli such as light touch) can be targeted using methods provided herein in order to prevent the neurons from releasing transmitter. For example, nucleic acid encoding a DREADD that causes inhibition of transmitter release (e.g., hM4Di or PSAM4-GlyR) can be introduced into neurons in a construct that also includes an RE specific for excitatory neurons, such that the RE can drive expression of the DREADD in the excitatory neurons. In some cases, nucleic acid encoding CRISPR / Cas components targeted to VGLUT2 can be introduced into neurons in a construct that also includes an RE specific for excitatory neurons, such that the RE can drive expression of the CRISPR / Cas components in the excitatory neurons, leading to modification or deletion of the VGLUT2 gene so that glutamate cannot be packaged into synaptic vesicles and released from those neurons. In some cases, nucleic acid encoding TeTx can be introduced into neurons in a construct that also includes an RE for excitatory neurons, such that the RE can drive expression of TeTx in the excitatory neurons so neurotransmitter cannot be released. In some cases, inhibitory neuron cell types that are important for preventing mechanical allodynia under normal conditions can be targeted. For example, nucleic acid encoding a DREADD that can increase transmitter release (PSAM4-5HT3 or hM3Dq) can be introduced into neurons in a construct that also includes an RE specific for inhibitory neurons, such that the RE can drive expression of the DREADD in the inhibitor neurons.
[0094] A nucleic acid construct provided herein can be administered to a mammal by any appropriate route and in any appropriate amount. For example, a nucleic acid construct packaged into a virus (e.g., an AAV) and administered directly into the brain via intraparenchymal administration, or can be administered into cerebrospinal fluid or systemically (e.g., intravenously). The amount administered can vary depending on the route of administration. When a nucleic acid is administered via intraparenchymal injection, for example, the amount administered also can depend on the region to be targeted. In generally, an area of about 4 mm2square can be covered with an amount of about 20 microliters of virus. Injections into the CSF and systemic injections typically utilize higher volumes (e.g., about 0.5 mL to about 10 mL).
[0095] The invention will be further described in the following example, which does not limit the scope of the invention described in the claims.
[0096] EXAMPLES
[0097] Example 1 - Cell Type Specific AAVs for Targeting Cognitive and Reward Systems Specific Nuclear-Anchored Independent Labeling (SNAIL) is an ML-based approach that leverages the cell type specificity of REs to control gene expression in specific neuronal cell types in the brain (Lawler et. al., eLife, ll:e69571, 2022). SNAIL was used in deep neural networks to identify candidate RE sequences with highly specific activity in neuronal cell types of interest. Candidate REs were then synthesized and packaged into an engineered adeno-associated virus (AAV) along with a transgene (e.g., a reporter gene such as GFP, or a gene that controls neuronal firing) to be expressed in the neuronal cell types of interest. In particular, the studies described herein were conducted to develop enhancer-driven AAVs optimized for achieving cell type specific transgene expression in non-human primates (NHPs). The genome was screened with SNAIL to identify candidates, and candidates were tested in the NHP cognitive and reward areas.
[0098] Cell type and OCR identification and machine learning (ML) assisted enhancer selection: Single nucleus RNA and DNA data from rhesus monkeys was used to identify cell types (He et al., Curr Biol, 31(24):5473-5486, 2021), including include layer 3 pyramidal neurons, striatal DI -matrix medium spiny neurons (MSNs), and striatal Dl-striosome MSNs, and to identify cell type specific open chromatin regions (OCRs). The SNAIL model was used to scan the OCRs and rank them as potential cell type specific enhancers (FIGS. 1A-1G).
[0099] In vivo screening strategy: The enhancers (SEQ ID NOS: 1-88; FIG. 6) were cloned into AAV libraries, which were injected into mice and monkey brain parenchyma (FIGS. 2A-2B). In particular, the enhancers were packaged in an AAV9 backbone together with an enhancer-specific DNA barcode. Pools of barcoded enhancer virus were injected in NHP dorsolateral prefrontal cortex or striatum. Animals were euthanized 4-5 weeks later. Cell type specificity was assessed using direct visualization of GFP expression, as well as fluorescent in situ hybridization against barcodes and cell type specific markers. Cell type specific expression of reporter genes was observed (FIGS. 3A-3E and 4A-4C). The candidate layer 3 enhancers had a strong preference to be strong active enhancers in that layer of the cortex in both mouse and macaque (FIGS. 3A-3C). DI striosome was distinguishable relative to DI matrix in mouse based on projection targets. DI -matrix labeled neurons tend to project to the substantia nigra, while Dl-striosome neurons project to the GPi (FIGS. 4A-4C and 5A-5D) (Crittenden et al., Proc Natl Acad Sci USA, 113: 11318- 11323, 2016). Dl-matrix neuron subtypes were further validated using fluorescence in situ hybridization (FISH) for the DRD1 transcript (FIG. 4B).
[0100] Example 2 - Cell Type Specific Enhancers for Dorsolateral Prefrontal Cortex The results in this Example re-present and expand on at least some of the results provided in Example 1.
[0101] MA TERIALS AND METHODS NHPs: Rhesus and Cynomolgus Monkeys were single- or pair-housed with a 12h-12h light-dark cycle. Animals were as listed in TABLE 1.
[0102] TABLE 1: Study animals Brain surgery for single cell experiments'. Brain surgery was performed as described elsewhere (He et al., Curr Biol, 31 :5473-5486, 2021). Briefly, animals (monkeys SMI 1, SM16, SM19, SM20, SM45) were anesthetized with ketamine (15 mg / kg IM) within a cage and transported to a surgery suite. The animals were further kept anesthetized with isoflurane and head-fixed in a stereotaxic instrument (Kopf Instruments). Vitals, including body temperature and respirations, were under constant monitoring. All surgical tools were sterilized and wiped with RNase decontaminant (Apex, Cat# 10-228). To maximize the viability of the cells, the skull was removed first, before perfusing the monkeys with approximately 4 liters of oxygenated ice-cold artificial cerebrospinal fluid. The dura was then opened and the brain was removed. The prefrontal cortex was cut under a dissection microscope for nuclei isolation. Monkeys SMI 8 and SM22, for FISH, were perfused first with phosphate-buffered saline (PBS, Fisher Scientific, Cat# BP243820) to flush out blood and then with 4% paraformaldehyde (PF A, Sigma-Aldrich, Cat# P6148) supplemented with 10% sucrose (Sigma-Aldrich, Cat# S8501) to fix the tissue. The brain was post-fixed with 4% PFA and cryopreserved with a gradient of sucrose (10%, 20%, 30%) in PBS.
[0103] Brain surgery for viral injections'. Monkeys were scanned with a T1 weighted scan using a Siemens MAGNETOM 7T Plus at least one week before surgery. These scans (between 3 and 5 individual scans per animal) were later re-oriented into the coronal plane and averaged before loading into BRAINSIGHT®. This approach allowed for accurate surgical targeting in the absence of a stereotaxic MRI. When the averaged and oriented scan was loaded into BRAINSIGHT®, the program allowed for 3D image reconstruction based on the MRI. Using landmarks on the skull and a 3D brain model, the intended targets were located and programmed prior to surgery.
[0104] On the day of surgery, monkeys were placed into a stereotaxic frame to mimic the coronal positioning of the MRI scan. A sagittal opening was made in the skin and fascia using a cauterizing tool. The temporalis muscle’s internal capsule was opened to enable muscle retraction for adequate access to the target sites. A laser pointer was attached to the surgical robot, and two camera capture points were projected on the exposed skull. This registration was compared to the 3D rendering created from the MRI file prior to surgery. The program then overlaid the 3D rendering and planned targets to the monkey’s exposed skull for targeting. Once registration was complete, the laser was removed, burr holes were drilled, and viruses were infused at a slow controlled rate with a Harvard Apparatus Pump 11 elite Nanomite. Each track generally consisted of multiple injection sites at different dorsal / ventral locations until the most dorsal injection site was reached. At the end of the track, the needle stayed in place for 5-10 minutes to allow the virus to diffuse into the target tissue before being removed from the tissue. After the end of all injections, the soft tissue was closed in anatomical layers from muscle capsule to skin. All procedures were conducted using aseptic techniques in a dedicated operating suite. At least 12 hours before surgery, subjects were treated with antibiotics (ceftriaxone; 50 mg / kg at 350 mg / mL, administered intramuscularly (IM)) and steroidal anti-inflammatory (dexamethasone; 0.5 mg / kg at 4 mg, administered orally (PO)) to minimize the risk of post-operative infection and inflammation. To minimize nausea, animals were dosed with Cerenia (1 mg / kg at 10 mg / mL, administered subcutaneously (SQ)) the night before and the night of the surgical procedure. Antibiotic and steroid treatment were continued for 3- 7 days post-surgery along with analgesics such as Mel oxicam (0.2 mg / kg at 1.5 mg, PO) and Buprenorphine (0.02 mg / kg at 0.4 mg / mL, IM). For the one-at-a-time validation experiments, the injection details are shown in TABLE 2.
[0105] TABLE 2: Injection details
[0106] Optogenetics-. On the day of surgery, monkeys were placed into a stereotaxic frame, and the opening was the same as that for injection surgeries. However, once the skull was cleaned, a rectangular window was opened over the cortex, which had been injected with the virus in the previous procedure. Using surgical blade #11, an incision was made in the dura, exposing the cortex of interest. To stabilize cortical pulsations, 3% agarose (Invitrogen) solution was coated over the exposed tissue during data acquisition. Optogenetic stimulation and recording were collected using a Plexon custom S-Probe, 16 recording channels with four light channels. The probe was connected to a Saphire 488-500 LPX LDRH Laser System with an LC adaptor. The laser was paired with an adjustable and focusable laser beam coupler and a single-mode fiber cable. Stimulations were fed into the laser coupler from a A-M Systems Isolated Pulse Stimulator (Model 2100). A record of the stimulations and recordings was saved using the Spike2 Program and later analyzed with MATLAB. Once all the data were collected, the soft tissue was closed and the animal was prepared for perfusion. All procedures were conducted using aseptic techniques in a dedicated operating suite. Isoflurane sedation was kept at a minimum during the procedure, with additional doses of ketamine (15 mg / kg at 100 mg / mL) provided every 6 hours.
[0107] Single nucleus RN A sequencing'. lOx Chromium Single Cell 3’ Reagent kits, v3.1 Chemistry (lOx Genomics, Cat# PN- 1000121) were used for monkeys SM11 and SM16. Nuclei were isolated as described elsewhere (He et al., supra). The standard lOx protocol for v3.1 chemistry was followed for library generation. Briefly, mRNAs were reverse transcribed within Gel beads-in-emulsion (GEMs) after running through a lOx Genomics Chromium controller. The emulsion was broken with a recovery agent (lOx Genomics, Cat# 220016) and cDNAs were purified with Dynabeads (lOx Genomics, Cat# 2000048). cDNAs were then amplified and subsequently purified using the SPRIselect reagent (Beckman Coulter, Cat# B23318). After analyzing the cDNA quality using an Agilent Bioanalyzer 2100, libraries were prepared following fragmentation, end repair, A-tailing, adaptor ligation, and sample index PCR. The libraries were quantified by qPCR using a KAPA Library Quantification Kit (KAPA Biosystems, Cat# KK4824). Libraries from individual monkeys were pooled together and loaded onto a NovaSeq S4 Flow Cell Chip. Samples were sequenced to a depth of 200,000 reads per nucleus.
[0108] Single nucleus RN A andATAC multiomics sequencing'. The lOx Chromium Single Cell Multiome Library & Gel Bead Kit (10X Genomics, Cat# PN-1000283) was used for monkeys SMI 9, SM20, and SM45. Briefly, the tissue was broken into small pieces with a wide-bore pipette tip followed by a regular-bore pipette tip, then filtered through a 30 pm MACS SmartStrainer. After spin down, the cells were lysed with 0. IX lysis buffer (10 mM Tri s-HCl, lO mM NaCl, 3 mM MgCh, 0.1% TWEEN®-20, 0.1% NP-40, 0.01% Digitonin, 1% BSA, 1 mM DTT, 1 U / pL RNase inhibitor). After washing, the nuclei were resuspended with diluted nuclei buffer (lOx Genomics, PN-2000153). The nuclei concentration was adjusted to 2,900-7,260 nuclei / pL, based on the targeted nuclei recovery of 9,000. The nuclei were transposed and run through the lOx chromium controller, followed by reverse transcription. The snATAC-Seq and snRNA-Seq libraries were generated according to standard protocols. Libraries were pooled together and loaded onto a NovaSeq 6000 S4 Flow Cell. Each sample from monkeys SM19, SM20, and SM45 was sequenced to a depth of 150,000 x 150 bp paired-end reads per nuclei. snRNA-Seq analysis'. After sequencing, the BCL reads were converted to fastq files and reads were aligned to a custom transcriptome reference. The monkey dataset was integrated using a standard Seurat v3 pipeline. Briefly, ambient RNA, ribosomal genes, and doublets were removed. Standard log-normalization and a variance stabilizing transformation were then performed to identify variable features individually for each monkey’s dataset using Seurat’s FindVariableFeatures function. Next, anchors were identified using the FindlntegrationAnchors function with default parameters, and the anchors were passed to the IntegrateData function. This returned a Seurat object with an integrated expression matrix for all nuclei. The integrated data were scaled with the ScaleData function, PCA was run using the RunPCA function, and the results were visualized with UMAP. Louvain clustering was used, and a resolution was chosen that reflected the major cell classes of the PFC, including excitatory neurons, inhibitory neurons, and astrocytes. Differentially expressed genes were calculated for each cell class with the FindMarkers function and, based on the marker genes, the major cell classes were annotated.
[0109] Because the variation with excitatory neurons was masked when co-clustering with other cell types, excitatory neuron clusters that express excitatory neuron marker genes such as SLC17A7 and TBR1 were isolated. The principal components (PCs) were re-calculated, UMAP dimension reduction was performed, and a resolution that separated clusters that were distinct in UMAP space was chosen. These clusters were annotated based on their maker genes and mapping to laminar locations. To analyze the inhibitory neuron populations, clusters were isolated based on inhibitory neuron markers GAD1 and GAD2. Similarly, PCs were re-calculated and UMAP dimensionality reduction was performed on the first 20 PCs. Louvain clustering was used and the resulting clusters were annotated based on known interneuron markers.
[0110] Cosine similarity was calculated for excitatory and inhibitory neurons within and between clusters based on PCA space. A permutation test was used on the cosine similarity between pairs of clusters. The nuclei were randomly shuffled to mask the nuclei identity and recalculated cosine similarity. These were repeated 10,001 times and within group and between group variance ratio was used to determine a P value. To compare the cell types between monkey and human excitatory neurons, the labeling of human PFC, MTG, and Ml subclass and cell type was transferred to the reference monkey dataset with Seurat TransferData function. The cell number was normalized for each human subclass / cell type and a heatmap was plotted to show their correspondence relationship. snATAC-Seq alignment and cell quality control (QC) filtering'. snATAC-Seq reads from all samples were aligned to the rheMaclO genome using the fast snATAC- Seq mapper chromap (Warren et al., Science 370(6523), 2020, doi: 10.1126 / science.abc6617; and Zhang et al., Nat Commun, 12:6566, 2021, doi: 10.1038 / s41467-021-26865-w). Arrow files were created from the fragment files for downstream analyses with the ArchR package (Granja et al., Nat Genet, 53:403-411, 2021). The previous approach was followed to map the higher quality GRCh38.pl3 human RefSeq gene annotations onto the rheMaclO genome with the liftOff tool, and this was used to create a custom rheMaclO ArchR gene and genome annotation (He et al., supra'. Granja et al., supra, and Shumate and Salzberg, Bioinformatics, 37: 1639- 1643, 2021). Using this liftOff annotation, ArchR gene activity scores were computed around the orthologous rhesus macaque regions of human genes, enabling scores to be obtained for substantially more genes and complete gene bodies than otherwise would have been obtained using base rheMaclO gene annotations. Genes, exons, and transcription start sites within 1Mb from chromosome ends were excluded, as they could have created errors during per-cell gene score computation. Within ArchR, two rounds of low-quality nuclei removal were performed. This 2-step filtering process excluded 6-12% of unfiltered nuclei and removed nuclei that tend to form clusters entirely made of low-QC cells or single nuclei that do not cluster with other high- quality nuclei. The first round used cut-offs to exclude nuclei with a high probability of being doublets (cutEnrich = 0.5, cutScore = -logl0(.05), filterRatio = 1) and low number of unique per cell fragments (nFrags < 10A3.5). The second round used the MASS R-package negative-binomial generalized linear model to learn the relationship between per-cell metrics TSSenrichment and PromoterRatio and DoubletEnrichment with the number of unique fragments for each sample, glm.nb(nFrags ~ TSSEnrichment + PromoterRatio + DoubletEnrichment) (Bon et al., Philos Trans A Math Phys Eng Sci, 381 : 20220156, 2023 , doi: 10.1098 / rsta.2022.0156). The standardized residual was calculated from the curve fit and nuclei with residuals more than two standard deviations away from the fit were excluded.
[0111] Single nuclei AT AC clustering, cell annotation, and differential analyses'. Iterative latent-semantic index clustering was performed with 4 iterations and 30,000- 230,000 variable features, followed by Harmony batch correction (Korsunsky et al., Nat Methods, 16: 1289-1296, 2019), UMAP visualization, and Louvain clustering, where default settings were used for all methods. snATAC-Seq clusters were labeled by first identifying glial and neuronal clusters by marker gene activity scores (RBFOX3, GAD1, GAD2, AQP4, CX3CR1, PDGFRA, MOG, etc ). Glial and neuronal clusters were subsetted and iterative LSI embedding and clustering were reperformed as described above. To annotate the snATAC-Seq nuclei, Seurat’s canonical correlation analysis (CCA) was applied on ArchR gene activity scores with the corresponding RNA profiles from the same animal and brain region, which was confirmed with the barcode link associated with the multiomics kit. Pseudo-bulk profiles were created across the annotated cell types and replicates with ArchR function addGroupCoverage(minReplicates = 5, maxReplicates = 24, minCells = 40, maxCells = 1000), peaks across replicates were called with macs2 (Zhang et al., Genome Biol, 9:R137, 2008, doi: 10.1186 / gb-2008-9-9-rl37), reproducible peaks between replicates were identified for each cell type, a consensus peak set of 501bp fixed-width, summit-centered peaks across cell types was created, and a peak by cell, “PeakMatrix,” was created through ArchR.
[0112] Machine models of cell type specificity. Peaks were further split into a test set, validation set, and training set. Support vector machine (SVM) models were trained with Isgkm gkmpredict function with the parameters -t 4 -1 7 -k 6 -d 1 -M 50 -H 50 -m 40000 -s -T 4, which were found to be robust to train accurate models as described elsewhere (Lawler et al., supra, doi: 10.7554 / eLife.69571; and Lee et al., Bioinformatics, 32:2196-2198, 2016). A grid search was performed for the regularization -c and weight -w parameters to train SVM models, and the models with the best F half score evaluated on the validation set were selected. Models that did not have area under the receiver operator characteristic (auROC) curve > 0.65 and accuracy > 0.65 on the validation set were excluded. Convolutional neural network (CNN) models were trained with a 5-convolutional layer architecture and one-cycle policy training regime described elsewhere (Srinivasan et al., JNeurosci, 41 :9008- 9030, 2021). A batch size of 64 sequences and 30 epochs was used, and models that did not have area under the auROC curve > 0.65 and accuracy > 0.65 on the validation set were excluded.
[0113] Prioritization of cell type specific enhancers'. Differential peaks were scored for each cell type with the collection of best SVMs and CNNs successfully trained for each comparison, and scores from both SVM and CNN were averaged for the same comparison and across comparisons to rank the degree of cell type specificity. Peaks that were predicted to have off-target effects in any other cell type were screened out by removing any peak with negative predictions across comparisons. In addition, for each candidate enhancer, it was determined whether the accessibility of these enhancers correlated with increased gene expression of nearby genes. Using the subset of nuclei with same-cell ATAC and RNA profiles, the ArchR Peak2GeneLinkage function was used to identify peaks statistically co-accessible with measured gene expression. Further priority was added to peaks where the associated gene was also separately identified as a differentially expressed marker gene. The geometric mean was applied to the Peak2Gene correlation with gene expression, MarkerGene log2-fold change, cell type specific motif Z-score, and the average machine learning (ML) score to create the composite score of whether the candidate peak contains cell type specific DNA sequences based on ML and motif analyses likely drives expression of cell type specific genes by co-accessibility and differentially expressed gene analyses. Finally, genome track plots of the open chromatin were inspected around the candidate enhancer and gene expression of nearby genes to select the top 12 candidates to test for cell type specificity in the rhesus macaque brain.
[0114] Plasmid Cloning and AAV purification'. To test the effects of these putative cell type specific enhancers, plasmids were generated with each enhancer upstream of the minimal promoter HSP68 to drive expression of GFP. A control plasmid was also generated in which the human synapsin promoter drove the expression of dTomato. For all plasmids, a 500bp barcode was included downstream of the transgene to allow for FISH analysis of activity. These inserts were synthesized and cloned into backbone pEMS2115 (Addgene Plasmid #49140) with restriction enzymes EcoRI and Notl. Plasmids were packaged into AAV9 2YF, PHP.eB, and XI .1 serotypes for viral purification (Vectorbuilder).
[0115] FISH stain and imaging'. After removing excess solution from the brain surface, the brain was air-dried for 15 minutes, embedded in optimal cutting temperature compound (OCT), and stored at -80°C until sectioning. 15 pm free- floating sections were collected from monkeys SMI 8 and SM22 and mounted on 2- inch><3-inch pre-coated slides. The mounted sections were preserved in a -80°C freezer. FISH was performed with Multiplex Fluorescent Detection Reagents v2 (ACD, Cat# 323110) according to the manufacturer’s protocol, with slight modifications for monkey brain tissue. Slides were retrieved from the -80°C freezer and equilibrated the to room temperature for 30 minutes. The slides were washed briefly with water and the tissue was fixed in 4% PFA buffer for 30 minutes. The sections were then dehydrated in 50%, 70%, and 100% ethanol and baked at 60°C for 20 minutes. The brain sections were incubated with hydrogen peroxide for 10 minutes to quench endogenous horseradish peroxidase (ACD, Cat# 322335) and then target retrieval was performed with RNAscope Target Retrieval Reagents (ACD, Cat# 322000) for 8 minutes at 99°C. The slides were dehydrated in 100% alcohol for 3 minutes and then baked at 60°C for 10 minutes. The slides were then incubated with protease III (ACD, Cat# 322337) for 30 minutes at 40°C to increase probe penetration, and then hybridized with probes for 2 hours. After signal amplification with AMP 1, 2, and 3, probes were conjugated with different HRP channels and fluorophores, including Opal 520 (PerkinElmer, Cat# FP1487A), Opal 570 (PerkinElmer, Cat# FP1488A), and Opal 650 (PerkinElmer, Cat# FP1496A). Sections were counterstained with DAPI and mounted with Prolong Gold Antifade Mountant (Life technologies, Cat# P36930). Sections were scanned using a Nikon Eclipse Ti2 under 20x objective. Imaged and Adobe Photoshop were used to adjust brightness and overlay images.
[0116] QUANTIFICATION AND STATISTICAL ANALYSIS
[0117] Enhancer distribution across layers'. To assess the enhancer expression across layers, Cellpose was used to detect the cells and quantify their distribution (Stringer et al., Nat Methods, 18: 100-106, 2021). Briefly, the orientation of the images was first adjusted to be vertical and of equal width. Then Cellpose was run to detect the cells using the default parameters. Cellpose output the positions, sizes, and signal intensities of cells. A consistent threshold was set for size and signal intensity to remove noise signals. The layers were divided into 50 parts and the number of cells was counted in each segment. The cell number was normalized to the maximum number for each image and their layer distribution was plotted. The layer boundaries were judged based on the original DAPI images.
[0118] Enhancer cell type specificity / efficiency by colocalization with marker genes'. To determine the specificity and efficiency of enhancer-driven expression, FISH was performed against enhancer-specific barcodes and marker genes. Using CellProfiler cell image analysis software (cellprofiler.org), intensities and locations for each RNA signal were defined for each cell. Cell counts were performed for colocalization of FISH signal for the RMacL3-01 barcode and CUX2. Specificity was calculated as the percentage of enhancer barcode-expressing cells that also co-expressed the CUX2. Efficiency was calculated as the percentage of CUX2-expressing cells that also coexpressed the enhancer barcode. Results are reported as mean ± SD (n=4).
[0119] Brain-wide FISH mapping'. The image reconstruction approach used a generative probabilistic model incorporating diffeomorphic spatial warping and contrast adjustments. Mapping to a common coordinate system was achieved through maximum a posteriori estimation, enabling precise reconstruction of both 3D and 2D datasets. To minimize tissue distortion, the Tape-Transfer cryosectioning method, which preserved the tissue’s original spatial orientation with 99% distortion-free accuracy, was employed (Lin et al., Elife, 8, 2019, doi: 10.7554 / eLife.40042. CrossRefGoogle Scholar; and Pinskiy et al., PLoS One, 10:e0102363, 2015). This framework supports data from MRI, various staining methods, and partial tissue regions. By integrating different types of histology and FISH data, the workflow provided a unified coordinate framework for quantification of ground-truth distribution across datasets for brain mapping. Histology sections (downsample) were first preprocessed to match MRI resolution using scattering operators. This approach preserved critical texture information necessary for accurate multimodal alignment. A Rigid Alignment was then adopted from each image acquisition to minimize a multimodal cost function, which accounted for contrast transformations and the likelihood of pixel artifacts. To address signal inhomogeneity and large inter-modality differences, parameters were optimized locally in small regional blocks. The optimization was carried out iteratively with a multi-resolution approach to avoid local minima. For initial 3D reconstruction, the 3D structure of histology sections and regional tissues were reconstructed by rigidly aligning them with neighboring sections. This alignment was adjusted to account for any anomalous pixels, ensuring accurate reconstruction. For MRI-guided registration, diffeomorphic mappings between MRI and histology sections were created using a sequence of spatial warps based on the Large Deformation Diffeomorphic Metric Mapping (LDDMM) model (Lee et al., J Comp Neurol, 529:281-295, 2021). For multimodal registration, FISH data were aligned to the transformed histology sections using rigid alignment. Partial samples could also be aligned manually with a rigid initialization framework if needed. All transformations between spatial domains were saved in a standardized format. Throughout the process, uncertainties were measured by positioning corresponding points, curves, and surfaces and by evaluating the true distances between them for further analysis.
[0120] Sources of key resources are listed in TABLE 3.
[0121] TABLE 3: Key resources
[0122] *Schneider et al., Nature Methods, 9:671-675, 2012
[0123] **Carpenter et al., Genome Biol, 7(10):R100, 2006
[0124] RESULTS DLPFC neuron phenotypes'. To define DLPFC neuronal phenotypes that could be targeted by enhancers, fresh tissue from the DLPFC in five rhesus monkeys was dissected. The rostral and caudal extents of the anterior and posterior supra-principal dimples guided the dissections (FIG. 14A, top). Coronal sections were cut around these landmarks and the cortical gray matter was separated from the underlying white matter (FIG. 14A, bottom). Well-validated single cell pipelines were used to perform single nucleus RNA with sequencing (snRNA-Seq) and single nucleus assay for transposase-accessible chromatin with sequencing (snATAC-Seq). The transcriptomic data from snRNA-Seq were used to define cell types and subtypes, and the open chromatin data from snATAC-Seq were used to identify enhancers to target those cell types and subtypes.
[0125] UMAP dimensionality reduction revealed well-separated clusters (FIG. 14B), and each of the five biological replicates contributed nuclei to each cluster (FIG. 15A). Based on cluster-specific differential marker genes (FIG. 14C), seven major cell classes were identified, including principal excitatory neurons (PN), inhibitory interneurons (In), astrocytes (A), microglia (pG), oligodendrocytes (O), oligodendrocyte precursors (OP), and endothelial cells (E). Violin plots revealed class specific marker gene expression levels (FIG. 14D). Together, these results indicated that the dataset contained five high-quality biological replicates from the rhesus macaque DLPFC.
[0126] The primary objective of these studies was to find enhancers that target the principal neurons of the DLPFC. Principal neuron clusters were defined as those enriched for marker genes of excitatory neurons, including SLC17A7 (FIG. 15B) (Ma et al., Science, 377:eabo7257, 2022, doi: 10.1126 / science.abo7257; Lei et al., Nature Commun, 13 :6747, 2022; and Chen et al., Cell, 186:3726-3743, 2023). This subset was re-analyzed and 11 distinct excitatory subtype clusters were identified (FIG. 14E). Differential gene analysis was used to identify excitatory subtype specific marker genes for each cluster (FIGS. 14F and 15C), and FISH against those marker genes was used to identify the anatomical position of each of those subtypes (FIG. 14G). The FISH analysis revealed that the 11 clusters mapped onto subtypes that were layer specific and subtypes that spanned adjacent layers. Layer specific subtypes included L3PNs that expressed marker genes CUX2 and RORB and L5PNs that expressed marker gene POU3F1. Subtypes that spanned adjacent layers included a layer 4 / 5 subtype that expressed TBX15 and a layer 5 / 6 subtype that expressed NR4A2 (FIG. 14G). Hierarchical clustering and cosine similarity measures both showed that shallow- and deep-layer subtypes were molecularly separable, and that superficial layer subtypes were more similar to one another whereas layer 5 subtypes were highly distinct (FIG. 15D) (p < 0.0001, Permutation tests). These excitatory neuron subtypes closely corresponded to cell type taxonomies previously defined in humans, including near projecting (NP), corti co-thal ami c (CT), inter-tel encephalic (IT), and extra-tel encephalic (ET) or distinct cell types described elsewhere (Velmeshev et al., Science, 364:685-689, 2019; and Hodge et al., Nature, 573:61-68, 2019), such as L6b and Car3+ L6IT PNs (FIGS. 15E-15G).
[0127] A similar approach was followed to identify inhibitory interneuron subtypes in the rhesus macaque DLPFC. Analysis of inhibitory neuron clusters revealed 10 distinct subtypes (FIG. 16A), which segregated based on their developmental origins from either the medial ganglionic eminence, marked by LHX6 expression, or the caudal ganglionic eminence, marked by ADARB2 expression (FIGS. 16B and 16C). Differential gene analysis identified subtype-specific markers for molecularly distinct populations of LAMP5+ and SST+ interneurons, as well as a discrete population of PVALB+ basket cells that co-expressed TH (FIG. 16C). Multiplexed FISH on coronal DLPFC sections was used to determine the laminar distribution of these interneuron subtypes (FIG. 16D). The FISH analysis revealed that some interneurons subtypes showed layer preferences. For example, NDNF+ interneurons were restricted to LI, VIP+ interneurons were enriched in superficial layers, and the PVALB+ / TH+ interneuron subtype populated L5-6 (FIG. 16E). The remaining subtypes were distributed across multiple cortical layers. Together with the principal neuron taxonomies, these results establish a comprehensive taxonomy of DLPFC neuron types that can be used to identify subtype-specific enhancers.
[0128] Enhancer Identification'. Cell type specific OCRs may contain cell type specific enhancers. Single nucleus OCR data were collected using a multi-omic assay that revealed both the transcriptome and OCRs in each cell, and thereby enabled direct annotation of OCRs with the cell type labels. Cell type labels from above (FIGS. 14A-14G) were used to identify reproducible, cell type specific, open chromatin profiles for all excitatory and inhibitory neuron subtypes (FIGS. 17A and 16F). This process revealed thousands of cell type specific OCRs for each neuron type (3,738 ± 1,048, mean ± standard error).
[0129] Most OCRs are unlikely to elicit cell type specific gene expression. This is because most OCRs contain enhancers that lack the desired specificity, or that contain insufficient regulatory grammar to achieve desired levels of transgene expression (Vormstein-Schneider et al., Nat Neurosci, 23: 1629-1636, 2020; Lawler et al., supra, and Hrvatin et al., Elife, 8, 2019, doi: 10.7554 / eLife.48089). To efficiently screen for cell type specific enhancers, an ML-based approach (SNAIL) was developed to select OCRs for the capacity to elicit cell type specific gene expression in PV neurons relative to other neuron subtypes and glial cell types, as described elsewhere (Lawler et al., supra). Here, SNAIL was generalized to all neuron types. Support vector machines (SVMs) and convolutional neural networks (CNNs) were used to train an ensemble of ML models for each neuron subtype. Consistent with models for PV neurons (Lawler et al., supra), the areas under the receiver operator curves (auROC) ranged from 0.867 to 0.951, and the areas under the precision recall curves (auPRC) ranged from 0.876 to 0.958. The subtype ensembles were used to summarize the overall likelihoods for cell type specificity and off-target expression as ‘ML Enhancer Ranking Scores’ for all enhancer candidates. The ML Enhancer Ranking Scores were used to prioritize candidates for in vivo testing (FIG. 17B). The ML models substantially restricted the number of candidate enhancers for in vivo testing (3,738 ± 1,048 OCRs vs 233 ± 106 candidates; FIG. 17C). However, using this stringent ML approach, no candidate enhancers were identified for the L5ET POU3F1+ neuron subtype. Therefore, the criterion was relaxed and paired with domain heuristic approaches to select a panel of candidate enhancers for L5ET POU3F1+ neurons (FIG. 17D). The results of these ML-assisted rankings were testable sets of high- priority enhancer candidates.
[0130] Enhancer screening'. The top twelve enhancer candidates for L3PNs and for L5ET POU3F1+ neurons were selected for in vivo testing. The enhancers were cloned upstream of the HSP68 minimal promoter and GFP was placed in the open reading frame (ORF). No upstream enhancer was included with GFP, as a control for the activity of the minimal promoter. The human synapsin promoter (hSyn) was used with dTomato as a positive control. For each enhancer and the controls, a unique 500 nt barcode was inserted downstream of the transgene. To ensure that all enhancers were at the same final titer, each plasmid was packaged individually into a backbone for AAV9-2YF. The top 6 enhancer candidate AAVs, the positive control AAV, and the minimal promoter-only AAV were then mixed in equal fractions to create L3PN screening library 1 (FIG. 18A). The same procedure was carried out for the L3PN enhancers ranked 7-12, and for the L5ET enhancer candidates. This process resulted in two L3PN and two L5ET AAV screening libraries (TABLE 1). The L3PN libraries were injected into the dorsal and ventral banks of the principal sulcus in one rhesus monkey (FIG. 18B). The L5ET libraries were injected in the same locations, but in a different monkey. After one month, both animals were euthanized and DLPFCs were sectioned for FISH and immunohistochemistry.
[0131] FISH probes against the 500 nt, candidate-specific barcodes were used to measure enhancer performance. Two or three barcode probes were used on each section, and this was performed on multiple adjacent sections to assess the layerspecificity of all candidate cell type enhancers and controls (FIG. 19). Representative images for each enhancer candidate demonstrated that many were detected in superficial layers, but for many, especially in library 2, the detection level was low and centered on deep layers (FIG. 18C). This deep-layer expression resembled the basal expression of the HSP68 minimal promoter when no enhancer was present (FIG. 18C, right, and FIG. 19 S3). For each AAV, the proportion of positive cells across DLPFC cortical layers was quantified by counting labeled nuclei in each layer (FIG. 18D) This analysis revealed that the hSyn positive control drove expression across layers 2-6, with a peak in layer 5. The analysis also highlighted two promising enhancer candidates, RMacL3-01 and RMacL3-06. The greatest activity for both enhancers was centered on cells in layer 3. RMacL3-01 (which was the sequence the ML models ranked first and had the strongest expression in L3 neurons) was selected for deeper cell type profiling. In a subsequent round of FISH, marker genes for superficial (CUX2) and deeper (RORB) layers were included to identify layer 3 PN. This tissue processing showed that the majority of the RMacL3-01 signal was in L3PN, with some signal detected in both layers 2 and 4 (FIG. 18E). Crosstalk between enhancer-AAVs has been documented as described elsewhere (Coughlin et al., bioRxiv, 2023, doi: 10.1101 / 2023.12.23.573214), and given the subsequent results from one-at-time screening, this likely caused some of the non-specific labeling seen here. The FISH analysis for the L5ET enhancer candidates was repeated (FIGS. 18F, 20E, and 20E).
[0132] To further validate this screening approach, studies were conducted to investigate whether these primate-derived enhancers could maintain their layer specificity in rodents. Despite known differences in cortical organization between primates and rodents, particularly in upper layers, the pool of RMacL3 AAVs successfully drove layer 2 / 3 -specific GFP expression in mouse cortex (FIGS. 20A- 20B), while the control hSyn promoter showed broad expression across cortical layers (FIG. 20A). FISH analysis against enhancer-specific barcodes revealed that RMacL3- 01 and RMacL3-02 were sufficient to drive layer 2 / 3 -specific transcription in mouse cortex (FIG. 20C). Similarly, the RMacL5ET AAV pool restricted GFP expression to mouse L5 (FIG. 20D), with RMacL5ET-01 transcripts specifically co-localizing with the ET-specific marker Fam84b (FIG. 20E). These promising results from both the multiplexed NHP screening and validation in mouse warranted rigorous testing of the top enhancer candidates, RMacL3-01 and RMacL5ET-01.
[0133] One-at-a-time enhancer validation'. One-at-a-time injections were used to validate the cell type specific expression of the best enhancers for L3PNs (RMacL3- 01) and L5ETs (RMacL5ET-01). RMacL3-01 was packaged into three AAV9 variants: PHP.eB (Deverman et al., Nat Biotechnol, 34:204-209, 2016), 2YF (Byrne et al., Mol Ther, 23:290-296, 2015), and XI.1 (Chen et al., Nature Commun, 14:3345, 2023, doi: 10.1038 / s41467-023-38582-7). All injections were made directly into the cortex via MRI-guided syringe pumps (TABLE 2). In subject SM38, the PHP.eB and 2YF vectors were injected into dorsal and ventral banks of the principal sulcus (FIGS. 21A and 21B). At high titer and close to the dorsal bank injection site, the PHP.eB vector resulted in non-specific GFP expression (FIG. 21C). Further from the injection site, however, where the effective titer was lower, GFP expression was more specific to layers 2 and 3 (FIG. 21C, inset). The 2YF vector generated highly restricted expression visible between the top of layer 2 and the granular layer 4 (FIGS. 21C- 21E). Thus, enhancers in both AAV capsids were layer specific, but 2YF was more specific at the tested titer. In subject SM41, the XI.1 and 2YF capsid variants were injected into rostral and caudal aspects of the dorsal bank, respectively (FIG. 21F). Inspection of the XI.1 injection site and surrounding tissue showed that the virus spread ~3.5 mm in the rostro-caudal direction (FIG. 21G). GFP expression appeared highly restricted to layers 2 and 3 (FIG. 21G (inset) and FIG. 21 J). FISH was performed against the barcode for RMacL3-01 and layer 4 marker RORB. These images were registered with the anatomical MRI and with the Nissl sections (FIG. 21H). This enabled creation of a 3-D reconstruction of the expression derived from the XI.1 vector (FIG. 211). All of these measures indicated that RMacL3-01 drove expression that was restricted to layers 2 and 3. In subject SM53, three titers of the XI .1 vector were injected (FIG. 21K). It was observed that the highest titer resulted in non-specific expression, a tenfold dilution of that titer resulted in specific expression in layers 2 and 3, and a further 10-fold dilution was undetectable (FIGS. 21K-21M). In most cases, there was non-specific expression near visible injection sites, as observed elsewhere (Lawler et al., supra). Together, these results indicated that enhancer RMacL3-01, when injected at the appropriate titer (FIG. 21L), reproducibly restricts expression to layers 2 and 3.
[0134] To quantify the cell type specificity of RMacL3-01, automated cell counting was used to examine the colocalization of the RMacL3-01 barcode with CUX2, a marker gene for layer 2 / 3 pyramidal neurons. This was done on tissue sections from SM41, where the XI.1 vector was injected. The specificity for CUX2 labeled cells was 77.2 ± 19.9% (mean ± SD) and the efficiency was 73.0 ± 12.4% (mean ± SD). This result indicated that not only was the enhancer layer specific, but it also was specific for layer 2 / 3 pyramidal neurons.
[0135] The most promising L5ET cell enhancer (RMacL5ET-01) was packaged into AAV9-2YF and injected into the ventral bank of the principal sulcus in subject SM39. As with RMacL3-01, non-specific labeling was observed near the injection site. However, further from the injection site where the effective titer was diminished by diffusion, highly restricted expression was observed in L5 cells. In fact, the highest specificity to layer 5 was detected in medial wall structures around the cingulate sulcus, which is about 3 mm from the injection site (FIG. 21N). Multi-channel FISH against the RMacL5ET-01 barcode and a L5ET marker gene, POUF31, showed high colocalization (FIG. 210). This result indicated that RMacL5ET-01 effectively targeted layer 5 projection neurons.
[0136] Functional testing of RMacL 3 -01 with optogenetics'. The enhancer-driven AAVs were developed to enable circuit-breaking studies in NHPs, especially with light-gated ion channels such as ChR2. Therefore, in vivo validation was performed using RMacL3-01 to drive expression of ChR2. The GFP was exchanged with ChR2(H134R)-p2A-GFP. The opsin-containing plasmids were packaged into AAV9- 2YF and AAV9-X1.1 capsids and injected into the dorsal and ventral banks of the principal sulcus in SM63 (FIG. 22A). Two months after injection, a 2x3 cm cranial window was opened to expose the principal and arcuate sulci (FIG. 22A). A handheld royal blue light, 500 nm long pass filter, and USB camera were used to visualize the injection sites (FIG. 22B). A custom-made 16-channel linear electrode array with one lightguide and four windows also was used; one window was located between electrodes 3 and 4, another between 6 and 7, 9 and 10, and 12 and 13. The lightguide was attached to a 473 nm, 500 mW laser. The array was approximately aligned to the principal sulcus, aimed near injection sites, and advanced until spikes were detected on the most superficial channel (Ch 1). Optical pulse trains readily evoked activity along the length of the array (FIG. 22C). The activity included multi-unit and single unit action potentials (FIG. 22D). Optically evoked activity followed the frequency of the laser stimulus (FIG. 22E). These results demonstrated that the RMacL3-01 enhancer can drive functional levels of ChR2 expression. Thus, RMacL3-01 vectors are suitable for immediate application to circuit-breaking studies in the NHP DLPFC.
[0137] OTHER EMBODIMENTS It is to be understood that while the invention has been described in conjunction with the detailed description thereof, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.
Claims
WHAT IS CLAIMED IS:
1. A nucleic acid construct comprising:(a) a transgene that comprises a nucleotide sequence encoding a polypeptide of interest, and(b) at least one RE specific for a selected cell type, wherein the RE comprises the nucleotide sequence set forth in any of SEQ ID NOS: 1 to 88, or a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOS: 1 to 88, and wherein the RE is operably linked to the nucleotide sequence encoding the polypeptide of interest, and is effective to drive expression of the polypeptide of interest in the selected cell type.
2. The nucleic acid construct of claim 1, wherein the polypeptide of interest is a clustered regularly interspaced short palindromic repeats- (CRISPR-) associated (Cas) nuclease or SunlGFP3. The nucleic acid construct of claim 1, wherein the polypeptide of interest is a Designer Receptor Exclusively Activated by Designer Drug (DREADD) polypeptide.
4. The nucleic acid construct of claim 3, wherein the DREADD polypeptide is a hM4Di polypeptide, a hM3Dq polypeptide, a PSAM4-GlyR polypeptide, or a PSAM4-5HT3 polypeptide.
5. The nucleic acid construct of claim 1, wherein the polypeptide of interest is a channelrhodopsin polypeptide.
6. The nucleic acid construct of any one of claims 1 to 5, wherein the nucleic acid construct further comprises a nucleotide sequence encoding a tag polypeptide, such that when the nucleotide sequences encoding the polypeptide of interest and the tag polypeptide are expressed, the polypeptide of interest is coupled to the tag polypeptide.
7. The nucleic acid construct of claim 6, wherein the tag polypeptide is a fluorescent polypeptide.
8. The nucleic acid construct of claim 7, wherein the fluorescent polypeptide is selected from the group consisting of green fluorescent protein (GFP), nuclear localization sequence-GFP, superfolder GFP, enhanced GFP, mCherry, mCitrine, mRuby, tdTomato, dTomato, Sun 1 GFP, yellow fluorescent protein (YFP), and cyan fluorescent protein (CFP).
9. The nucleic acid construct of claim 7, wherein fluorescent polypeptide comprises an amino acid sequence having at least 95% sequence identity with the superfolder GFP sequence set forth in SEQ ID NO: 100.
10. The nucleic acid construct of any one of claims 1 to 9, wherein the RE is specific for a specific subtype of neuron, such that the transgene is expressed in a majority of neurons of the specific subtype transduced with the nucleic acid construct, but is not expressed in at least 90% of neurons of other subtypes of neurons transduced with the nucleic acid construct.
11. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises core cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 1, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 1.
12. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises DI core cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO:2, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:2.
13. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises DI hybrid cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO:3 or SEQ ID NON, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NON or SEQ ID NON.
14. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises Dl.NUDAP cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NON or SEQ ID NON, or a nucleotidesequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:5 or SEQ ID NO:6.
15. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises DI. Shell cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9.
16. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises Dl.Striosome cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11.
17. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises D2 cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 12 or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 12.
18. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises D2. Shell cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14.
19. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises dSTR cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 15, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 15.
20. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises Matrix cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18, or anucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18.
21. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises Shell cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20.
22. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises Striosome cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64.
23. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises DI. Matrix cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30.
24. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises D2. Matrix cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35.
25. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises DI cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41.
26. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises D2 cells, and wherein the at least one RE comprises the nucleotidesequence set forth in any of SEQ ID NOs:42 to 51, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51.
27. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises L3.CUX2.RORB cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:65 to 76, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs 65 to 76.
28. The nucleic acid construct of claim 10, wherein the specific subtype of neuron comprises L5.POU3F1 cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88.
29. The nucleic acid construct of any one of claims 1 to 28, wherein the construct further comprises virus sequences.
30. The nucleic acid construct of claim 29, wherein the virus sequences are adeno- associated virus (AAV) sequences or lentivirus sequences.
31. A virus particle comprising the nucleic acid construct of any one of claims 1 to 30.
32. The virus particle of claim 31, wherein the virus is AAV.
33. A method for labeling a selected cell type within a population of different cell types, comprising introducing into the population of different cell types a nucleic acid construct comprising:(a) a transgene that comprises a nucleotide sequence encoding a polypeptide of interest and (ii) a sequence encoding a tag polypeptide, wherein expression of the transgene yields the polypeptide of interest coupled to the tag polypeptide, and(b) at least one RE specific for a selected cell type, wherein the RE comprises the nucleotide sequence set forth in any of SEQ ID NOS: 1 to 88, or a nucleotidesequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOS: 1 to 88, wherein the RE is operably linked to the nucleotide sequence encoding the polypeptide of interest and the tag polypeptide, and is effective to drive expression of the polypeptide of interest coupled to the tag polypeptide in the selected cell type, and wherein the tagged polypeptide of interest is expressed in and thereby labels the selected cell type.
34. The method of claim 33, wherein the polypeptide of interest is a Cas nuclease or SunlGFP.
35. The method of claim 33, wherein the polypeptide of interest is a Designer Receptor Exclusively Activated by Designer Drug (DREADD) polypeptide.
36. The method of claim 35, wherein the DREADD polypeptide is a hM4Di polypeptide, a hM3Dq polypeptide, a PSAM4-GlyR polypeptide, or a PSAM4-5HT3 polypeptide.
37. The method of claim 33, wherein the polypeptide of interest is a channelrhodopsin.
38. The method of any one of claims 33 to 37, wherein the tag polypeptide is a fluorescent polypeptide.
39. The method of claim 38, wherein the fluorescent polypeptide is selected from the group consisting of green fluorescent protein (GFP), nuclear localization sequence-GFP, superfolder GFP, enhanced GFP, mCherry, mCitrine, mRuby, tdTomato, dTomato, SunlGFP, yellow fluorescent protein (YFP), and cyan fluorescent protein (CFP).
40. The method of claim 38, wherein fluorescent polypeptide comprises an amino acid sequence having at least 95% sequence identity with the superfolder GFP sequence set forth in SEQ ID NO: 100.
41. The method of any one of claims 33 to 40, wherein the RE is specific for a specific subtype of neuron, such that the transgene is expressed in a majority of neurons of the specific subtype transduced with the nucleic acid construct, but is not expressed in at least 90% of neurons of other subtypes of neurons transduced with the nucleic acid construct.
42. The method of claim 41, wherein the specific subtype of neuron comprises core cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 1, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 1.
43. The method of claim 41, wherein the specific subtype of neuron comprises DI core cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO:2, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:2.
44. The method of claim 41, wherein the specific subtype of neuron comprises DI hybrid cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO:3 or SEQ ID NON, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NON or SEQ ID NON.
45. The method of claim 41, wherein the specific subtype of neuron comprises Dl.NUDAP cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NON or SEQ ID NON, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NON or SEQ ID NON.
46. The method of claim 41, wherein the specific subtype of neuron comprises DI. Shell cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO:7, SEQ ID NON, or SEQ ID NON, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NON, SEQ ID NON, or SEQ ID NON.
47. The method of claim 41, wherein the specific subtype of neuron comprises Dl.Striosome cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11.
48. The method of claim 41, wherein the specific subtype of neuron comprises D2 cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 12 or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 12.
49. The method of claim 41, wherein the specific subtype of neuron comprises D2. Shell cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14.
50. The method of claim 41, wherein the specific subtype of neuron comprises dSTR cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 15, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 15.
51. The method of claim 41, wherein the specific subtype of neuron comprises Matrix cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO:18.
52. The method of claim 41, wherein the specific subtype of neuron comprises Shell cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20.
53. The method of claim 41, wherein the specific subtype of neuron comprises Striosome cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID N0s:21 to 26 and 52 to 64.
54. The method of claim 41, wherein the specific subtype of neuron comprises DI. Matrix cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30.
55. The method of claim 41, wherein the specific subtype of neuron comprises D2. Matrix cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35.
56. The method of claim 41, wherein the specific subtype of neuron comprises DI cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41.
57. The method of claim 41, wherein the specific subtype of neuron comprises D2 cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51.
58. The method of claim 41, wherein the specific subtype of neuron comprises L3.CUX2.RORB cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:65 to 76, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs 65 to 76.
59. The method of claim 41, wherein the specific subtype of neuron comprises L5.POU3F1 cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88.
60. The method of any one of claims 33 to 59, wherein the construct further comprises virus sequences.
61. The method of claim 60, wherein the virus sequences are AAV sequences or lentivirus sequences.
62. A method for expressing a transgene in a selected neuron cell type within a population of different neuron cell types, comprising introducing into the population of different cell types a nucleic acid construct comprising:(a) a nucleic acid comprising an exogenous transgene that comprises a nucleotide sequence encoding a heterologous polypeptide, and(b) at least one RE specific for the selected neuron cell type, wherein the RE comprises the nucleotide sequence set forth in any of SEQ ID NOS: 1 to 88, or a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOS: 1 to 88, and wherein the RE is operably linked to the nucleotide sequence encoding the heterologous polypeptide and is effective to drive expression of the heterologous polypeptide in a majority of neurons of the selected neuron cell type into which the nucleic acid construct was introduced, but is not expressed in at least 90% of neurons of other neuron cell types of neurons into which the nucleic acid construct was introduced.
63. The method of claim 62, wherein the exogenous transgene further comprises a nucleotide sequence encoding a tag polypeptide, such that when the nucleotide sequences encoding the heterologous polypeptide and the tag polypeptide are expressed, the heterologous polypeptide is coupled to the tag polypeptide.
64. The method of claim 63, wherein the tag polypeptide is a fluorescent polypeptide.
65. The method of claim 64, wherein the fluorescent polypeptide is selected from the group consisting of green fluorescent protein (GFP), nuclear localization sequence-GFP, superfolder GFP, enhanced GFP, mCherry, mCitrine, mRuby, tdTomato, dTomato, SunlGFP, yellow fluorescent protein (YFP), and cyan fluorescent protein (CFP).
66. The method of claim 64, wherein fluorescent polypeptide comprises an amino acid sequence having at least 95% sequence identity with the superfolder GFP sequence set forth in SEQ ID NO: 100.
67. The method of any one of claims 62 to 66, wherein the heterologous polypeptide is a Cas nuclease or SunlGFP.
68. The method of any one of claims 62 to 66, wherein the heterologous polypeptide is a Designer Receptor Exclusively Activated by Designer Drug (DREADD) polypeptide.
69. The method of claim 68, wherein the DREADD polypeptide is a hM4Di polypeptide, a hM3Dq polypeptide, a PSAM4-GlyR polypeptide, or a PSAM4-5HT3 polypeptide.
70. The method of any one of claims 62 to 69, wherein the polypeptide of interest is a channelrhodopsin.
71. The method of any one of claims 62 to 70, wherein the construct further comprises virus sequences.
72. The method of claim 71, wherein the virus sequences are AAV sequences or lentivirus sequences.
73. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises core cells, and wherein the at least one RE comprises the nucleotidesequence set forth in SEQ ID NO: 1, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 1.
74. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises DI core cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO:2, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:2.
75. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises DI hybrid cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO:3 or SEQ ID NO:4, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:3 or SEQ ID NO:4.
76. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises Dl.NUDAP cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO:5 or SEQ ID NO:6, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:5 or SEQ ID NO:6.
77. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises DI. Shell cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 7, SEQ ID NO: 8, or SEQ ID NO: 9, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO:7, SEQ ID NO:8, or SEQ ID NO:9.
78. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises Dl.Striosome cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 10 or SEQ ID NO: 11.
79. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises D2 cells, and wherein the at least one RE comprises the nucleotidesequence set forth in SEQ ID NO: 12 or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 12.
80. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises D2. Shell cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 13 or SEQ ID NO: 14.
81. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises dSTR cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 15, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 15.
82. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises Matrix cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 16, SEQ ID NO: 17, or SEQ ID NO: 18.
83. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises Shell cells, and wherein the at least one RE comprises the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 19 or SEQ ID NO:20.
84. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises Striosome cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:21 to 26 and 52 to 64.
85. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises DI. Matrix cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30, or a nucleotidesequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:27 to 30.
86. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises D2. Matrix cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:31 to 35, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID N0s:31 to 35.
87. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises DI cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:36 to 41.
88. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises D2 cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:42 to 51.
89. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises L3.CUX2.RORB cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:65 to 76, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs 65 to 76.
90. The method of any one of claims 62 to 72, wherein the selected neuron cell type comprises L5.POU3F1 cells, and wherein the at least one RE comprises the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88, or a nucleotide sequence having at least 95% sequence identity to the nucleotide sequence set forth in any of SEQ ID NOs:77 to 88.
Citation Information
Patent Citations
Specific nuclear-anchored independent labeling system
US20220235102A1