Artificial expression constructs for regulating gene expression in intrathalamic neurons
Patent Information
- Application Number
- JP2023572122
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-21
- Filing Date
- 2022-05-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-20
AI Technical Summary
Existing methods for labeling and perturbing specific cell types in the central nervous system, particularly in humans, are costly, require breeding of transgenic animals, and are limited by the need for germline transgenic animals, making them infrequently available and not applicable to humans.
Development of artificial expression constructs using enhancer elements to induce gene expression in targeted central nervous system cell populations, including thalamic neurons and other specific cell types, utilizing concatemerization of enhancer cores to enhance expression levels and specificity.
The artificial expression constructs enable rapid and high-level gene expression in targeted CNS cell types, providing a cost-effective and applicable method for labeling and perturbing these cells, including in human models.
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 191,832, filed May 21, 2021, the contents of which are incorporated by reference in their entirety as if set forth herein.
[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT This invention was made with government support under Grant No. MH114126 awarded by the National Institutes of Health (NIH). The United States Government has certain rights in this invention.
[0003] The present disclosure provides artificial expression constructs for regulating gene expression in specific targeted central nervous system cells. The artificial expression constructs of the present invention can be used to express synthetic genes or regulate gene expression in the thalamus.
[0004] Sequence Listing Reference The sequence listing accompanying this application is provided in text format rather than hard copy, and is incorporated herein by reference. The text file containing the sequence listing is named A166-0028PCT_ST25.txt. The text file is 178kb in size, was created on May 20, 2022, and was submitted electronically via EFS-Web. [Background technology]
[0005] To fully understand the biology of the brain, it is necessary to distinguish between different cell types, define them, and investigate them in detail, as well as to identify artificial expression constructs that can label and perturb those cells. In mice, driver lines expressing recombinases have been used successfully to label cell populations that share marker gene expression. However, the creation, maintenance, and use of such lines that can label specific cell types with high specificity is costly and often requires crossing transgenic animals between three species, which only results in low frequency of obtaining the desired experimental animals. Moreover, these tools cannot be applied to humans because of the need for germline transgenic animals. Summary of the Invention [Means for solving the problem]
[0006] The present disclosure provides an artificial expression construct for inducing gene expression in targeted central nervous system cell populations. Targeted central nervous system cell populations include thalamic neurons, including GABAergic neurons in the thalamus (Gata / Dlx5-6), GABAergic neurons in the thalamic reticular nucleus (TRN) of the thalamus, thalamic reticular nucleus cells, glutamatergic neurons in the thalamus, glutamatergic neurons in the thalamus (Prkcd-Grin2c (core, LGN)), glutamatergic neurons in the thalamus (Rxfp1-Epb4 (matrix)), and glutamatergic neurons in the parafascicular nucleus (Pf) of the thalamus. In certain embodiments, the artificial expression constructs described herein induce gene expression in a second type of cell in addition to inducing gene expression in the thalamus. The second type of cells includes striatal medium spiny neurons (including all types), Purkinje cells in the cerebellum, deep cerebellar nucleus (DCN) cells in the cerebellum, molecular layer interneuron (MLI) cells in the cerebellum, parvalbumin (Pvalb)-positive neurons, chandelier cells, glutamatergic layer 5 neurons that project outside the telencephalon (L5 ET) cells in the neocortex, and Vip-positive neurons in the neocortex.
[0007] In certain embodiments of the artificial expression constructs of the invention, the following enhancers are utilized to induce gene expression in targeted central nervous system cell populations. The enhancer and target cell population combinations used in the invention are listed below in the order of enhancer / target cell population. eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_ 830h, eHGT_831h, eHGT_832h, eHGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHG T_577h, 3xcore3_eHGT_577h, 3xcore2_eHGT_606h, 3xcore3_eHGT_606h, core4_eHGT_577h, core6_eHGT_606h, 3xcore_eHGT_121h, eHGT_590m or eHGT_976h / glutamatergic neurons in the thalamus; MGT_E117 or MGT_E118 / glutamatergic neurons in the thalamus (Prkcd-Grin2c (core, LGN)); MGT_E119 or MGT_E120 / GABAergic neurons in the thalamus (Gata / Dlx5-6); MGT_E121 / glutamatergic neurons in the thalamus (Rxfp1-Epb4(matrix)); 3xCore2-eHGT_367h / glutamatergic neurons in the parafascicular nucleus (Pf) of the thalamus and striatal medium spiny neurons (including all types); eHGT_359h / glutamatergic neurons in the thalamus, Purkinje cells in the cerebellum, and Pvalb-positive neurons in the neocortex; eHGT_479m / glutamatergic neurons in the thalamus, Purkinje cells in the cerebellum, and chandelier cells in the neocortex; eHGT_453m / GABAergic neurons in the thalamic reticular nucleus (TRN) of the thalamus, deep cerebellar nucleus (DCN) cells of the cerebellum, and glutamatergic layer 5 neurons (L5 ET) cells projecting outside the telencephalon in the neocortex; eHGT_140h / GABAergic and Pvalb-positive neurons in the thalamic reticular nucleus (TRN) of the thalamus; eHGT_356h / GABAergic neurons in the thalamic reticular nucleus (TRN) of the thalamus, DCN cells of the cerebellum, Vip-positive neurons and thalamic reticular nucleus cells of the neocortex; eHGT_128h / glutamatergic and Pvalb-positive neurons in the thalamus; eHGT_369h / glutamatergic neurons in the thalamus, molecular layer interneuron (MLI) cells in the cerebellum, and Pvalb-positive neurons in the neocortex; and eHGT_710m / glutamatergic neurons in the thalamus, MLI cells, chandelier cells and GABAergic interneurons in the molecular layer of the cerebellum.
[0008] In certain embodiments, multiple copies of the enhancer are concatenated (linearly linked) or multiple copies of the enhancer core are concatenated. Examples include the core regions of eHGT_367h, eHGT_121h, eHGT_577h and / or eHGT_606h, or concatenated versions of these core regions. By using such artificial enhancer elements, the transgene can be expressed more rapidly and higher than when the original (natural) full-length enhancer is used alone.
[0009] In certain embodiments, the enhancer core comprises a sequence as set forth in any of SEQ ID NO: 18, SEQ ID NO: 26, SEQ ID NO: 31, SEQ ID NO: 33, SEQ ID NO: 35, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, and SEQ ID NO: 41. In certain embodiments, these enhancer cores are concatemerized and have 2, 3, 4, 5, 6, 7, 8, 9, or 10 copies of the core sequence. In certain embodiments, a concatemer comprising 3 copies of a selected enhancer core comprises a sequence as set forth in any of SEQ ID NO: 17, SEQ ID NO: 25, SEQ ID NO: 30, SEQ ID NO: 32, SEQ ID NO: 34, SEQ ID NO: 36, and SEQ ID NO: 40.
[0010] Particular embodiments of the enhancer core utilize Core2-eHGT_367h, coreB_eHGT121h, core2_eHGT_577h, core3_eHGT_577h, core2_eHGT_606h, core3_eHGT_606h, core4_eHGT_577h, core6_eHGT_606h or core_eHGT_121h. Particular embodiments of the concatemerized enhancer core utilize 3xCore2-eHGT_367h, eHGT_369h (3xcoreB_eHGT121h), 3xcore2_eHGT_577h, 3xcore3_eHGT_577h, 3xcore2_eHGT_606h, 3xcore3_eHGT_606h or 3xcore_eHGT_121h. In this disclosure, eHGT_369h can be exchanged for (3xCoreB)eHGT_121h.
[0011] Certain embodiments provide artificial expression constructs that include the features of vectors described herein, including CN2415, CN2416, CN2417, CN2418, CN2436, CN3000, CN3001, CN3003, CN3004, CN3005, CN3007, CN3009, AiP1335, AiP1336, AiP1337, AiP1338, AiP1339, AiP2000, AiP2001, AiP2002, AiP2003, AiP2004, AiP2005, AiP2006, AiP2007, AiP2009, AiP2008, AiP2009, AiP2009, AiP2006, AiP2008, AiP2009 ... These include iP1338, AiP1339, CN2555, CN2045, CN2258, CN2251, CN1633, CN2043, CN1621, CN2216, CN2717, CN3639, CN3050, CN3051, CN3056, CN3057, CN4001, CN4003, CN2786, CN2840, CN3460, and CN2650. [Brief description of the drawings]
[0012] Some of the drawings submitted in this application may be more easily understood in color, and applicants hereby contemplate color versions of these drawings as part of the original application and reserve the right to submit color images of such drawings in subsequent proceedings.
[0013] [Figure 1A-1C]Overview of the selection of enhancer candidates for the thalamus. (Figure 1A) Examples of marker genes for thalamic neurons using in situ hybridization (ISH) data from the Allen Brain Atlas. Both Rgs16 and Plekhg1 are shown to be thalamus-specifically labeled in the brain. Additionally, Gad1, a marker gene for inhibitory neurons, and Slc17a7, a marker gene for excitatory neurons, are shown for comparison. (Figure 1B) Dot plot summary of marker gene expression levels comparing excitatory neuron clusters (exc) and inhibitory neuron clusters (inh) in the thalamus from DropVis browser (dropviz.org) single-cell RNA-seq data from mice. Darker circles indicate higher gene expression levels. Note that Rgs16 and Plekhg1 are specifically expressed in excitatory neurons, but not inhibitory neurons. (Figure 1C) An example of candidate peak selection using the Brain Open Chromatin Atlas dataset (bendlj01.u.hpc.mssm.edu / multireg / ) displayed on the UCSC genome browser. The open chromatin peaks shown (highlighted in grey) indicate that the medial dorsal thalamic nucleus (MDT) is accessible, unlike various other regions such as the neocortex (CTX), amygdala (AMY), hippocampus (HIPP), nucleus accumbens (NAC) and putamen (PUT). The MDT peak is located near the Rgs16 locus and is predicted to be an enhancer for a subclass of thalamic excitatory neurons.
[0014] [Figure 2A-2B] (FIG. 2A) Inverted epifluorescence microscope image showing expression of native SYFP2 in the visual cortex 2 months after retro-orbital delivery of AAV vector #CN2415 (eHGT_576h) at 1.0×1012 viral genome copies. Scale bar: 1 mm. (FIG. 2B) High magnification image of the thalamic region. Scale bar: 500 microns.
[0015] [Figure 3A-3B](FIG. 3A) An inverted epifluorescence microscope image showing native SYFP2 expression in the visual cortex 2 months after retro-orbital delivery of AAV vector #CN2416 (eHGT_577h) at 1.0×1012 viral genome copies. Scale bar: 1 mm. (FIG. 3B) A high magnification image of the thalamic region. Scale bar: 500 microns.
[0016] [Figure 4A-4B] (FIG. 4A) An inverted epifluorescence microscope image showing native SYFP2 expression in the visual cortex 2 months after retro-orbital delivery of AAV vector #CN2417 (eHGT_578h) at 1.0×1012 viral genome copies. Scale bar: 1 mm. (FIG. 4B) A high magnification image of the thalamic region. Scale bar: 500 microns.
[0017] [Figure 5A-5B] (FIG. 5A) An inverted epifluorescence microscope image shows native SYFP2 expression in the brain 2 months after retro-orbital delivery of AAV vector #CN2216 (3x(coreB)eHGT121h) at approximately 1.0×1012 viral genome copies. Mice were transduced by intravenous delivery of AAV packaged with PHP.eB capsids. Scale bar: 1 mm. (FIG. 5B) A high magnification image of the thalamic region is shown. Scale bar: 500 microns.
[0018] [Figure 6A-6B] (FIG. 6A) Expression of the fluorescent reporter in mouse brain tissue after retro-orbital injection of CN2555, serotype PHPeB. Scale bar: 1 mm. (FIG. 6B) High magnification image of SYFP2 expression in the dorsal striatum region boxed in FIG. 6A. Scale bar: 500 microns.
[0019] [Figure 7A-7G]CN2555 virus of serotype PHP.eB was injected into C57Bl / 6J (wild type) mice via the retro-orbital sinus at 1.0×1012 viral genome copies. Representative two-photon computed tomography (TissueCyte) images showing native SYFP2 fluorescence in coronal sections of the brain 4 weeks after injection are shown. Figures 7A, 7C, and 7E show overviews of coronal sections in various planes of the brain (moving from anterior to posterior), and Figures 7B, 7D, and 7F show high magnification images of the boxed areas in the corresponding Figures 7A, 7C, and 7E, respectively. (Figure 7G) TissueCyte imaging datasets were processed and aligned to CCFv3.0 (Wang et al., Cell 2020, 181(4):936-953.e20) using a previously established informatics pipeline (Oh et al., Nature 2014, 508: 207-214). Analysis was performed using segmented pixel counts or voxels for each brain region, and analyzed data was shown in density dot plots for whole cortical (left = left hemisphere, right = right hemisphere) and subcortical structures. Large dark black circles = brain regions with highest SYFP2 signal; small light circles = brain regions with little or no SYFP2 signal. See Figure 16 for abbreviations indicating brain structures.
[0020] [Figure 8A-8I]Vector: CN2045, Enhancer: eHGT_359h. (Figures 8A-8E) Animal: Mouse 200910-08. (Figure 8A) Fluorescence composite image of native SYFP2 in sagittal sections of whole mouse brain and (Figure 8B) mouse cerebellum, and (Figure 8C) high magnification image of the mouse cerebellum, showing that SYFP2 is selectively expressed in cells with Purkinje cell morphology. (Figure 8D) Fluorescence signal of native SYFP2 in sagittal section of cerebellum showing that Purkinje cells are labeled, and (Figure 8E) image of native SYFP2 fluorescence signal superimposed with expression of Pvalb mRNA (arrow). Virus was administered to adult mice by intravascular (IV) injection (retro-orbital) of CN2045 virus packaged with PHP.eB capsid. (Figures 8F, 8G) Animal: Rat 585761. (Figure 8F) Fluorescence composite image of native SYFP2 in a sagittal section of a whole rat brain and (Figure 8G) a magnified image of the rat cerebellum show that SYFP2 is selectively expressed in cells with Purkinje cell morphology. Virus was administered to neonatal rats one day after birth by intracerebroventricular (ICV) injection of CN2045 virus packaged with PHP.eB capsid. (Figure 8H, Figure 8I) Animal: Monkey Q21.26.022. (Figure 8H) Fluorescence composite image of native SYFP2 in a sagittal section of a rhesus monkey cerebellum and (Figure 8I) a magnified image of the rhesus monkey cerebellum show that SYFP2 is selectively expressed in cells with Purkinje cell morphology. Virus was administered by intraparenchymal injection of CN2045 virus packaged with PHP.eB capsid.
[0021] [Figure 9A-9D](Figure 9A) Mouse genomic coordinates for the eHGT_453m enhancer (shaded region) and corresponding single-cell ATAC-seq peaks are shown for each cell classification and subclassification. Arrows indicate open chromatin peaks in L5 ET neurons of the subclassification. (Figure 9B) Epifluorescence microscopy images showing expression of native SYFP2 in the visual cortex 2 months after retro-orbital delivery of 7.00 × 1011 genome copies of AAV vector #CN2251. Scale bar: 100 microns. (Figure 9C) Mapping of single-cell transcriptome profiles of SYPF2+ cells sorted from the visual cortex region of mouse brain after retro-orbital injection of CN2251 virus packaged with PHP.eB capsids. The number of cells mapped to the final branch point is shown in the bar graph at the bottom of the phylogenetic tree. The cell types whose transcriptomes were investigated are shown at the bottom. The data show that reporter expression is selectively induced by the eHGT_453m enhancer in all four L5 ET neurons in the visual cortex. From left to right, the lower strings are: 169 L2 / 3 IT VISp Rrad, 168 L2 / 3 IT VISp Adamts2, 167 L2 / 3 IT VISp Agmay, 164 L4 IT VISp Rspo1, 163 L5 IT VISp Hsd11b1 Endou, 162 L5 IT VISp Whrn Tox2, 160 L5 IT VISp Batf3, 158 L5 IT VISp Col6a1 Fezf1, 157 L5 IT VISp Col27a1, 154 L6 IT VISp Penk Col27a1, 153 L6 IT VISp Penk Fst, L6 IT VISp Col23a1 Adamts2, 149 L6 IT VISp Col18a1, 146 L6 IT VISp Car3, 144 L5 PT VISp Chrna6, 143 L5 PT VISp Lgr5, 142 L5 PT VISp C1ql2 Ptgfr, 141 L5 PT VISp C1qI2 Cdh13, 140 L5 PT VISp Krt80, 134 L5 NP VISp Trhr Cpne7, 133 L5 NP VISp Trhr Met, L5 CT Nxph2Sla, 130 L5 CT VISp Krt80 Sla, L5 CT VISp Nxph2.Wls, 127 L5 CT VISp Ctxn3 Brinp3, 126 L5 CT VISp Ctxn3 Sla, 122 L5 CT VISp Gpr139, 120 L6b Col8a1 Rprm, 119 L6b VISp Mup5, 118 L6b VISp Col8a1 Rxfp1, 115 L6b P2ry12, L6b VISp Crh, 110 Lamp5 Krt73, Lamp5 Fam19a1 Pax6, 108 Lamp5 Fam19a1 Tmem182, 106 Lamp5 Ntn1 Npy2r, 105 Lamp5 Plch2 Dock5, 101 Lamp5 Lsp1, 100 Lamp5 Lhx6, Sncg Slc17a8, 96 Sncg Vip Nptx2, 95 Sncg Gpr50、93 Sncg Vip Itih5、90 Serpinf1 Clrn1、89 Serpinf1 Aqp5 Vip、85 Vip Igfbp6 Car10、84 Vip Igfbp6 Pltp、Vip Lmo1 Fam159b、Vip Lmo1 Myl1、79 Vip Igfbp4 Mab21l1、78 Vip Arhgap36 Hmcn1、77 Vip Gpc3 Slc18a3、74 Vip Ptprt Pkp2、73 Vip Rspo4 Rxfp1 Chat、71 Vip Lect1 Oxtr、70 Vip Rspo1 Itga4、67 Vip Chat Htr1f、66 Vip Pygm C1qI1、61 Vip CrispId2 Htr2c、60 Vip CrispId2 Kcne4、58 Vip Col15a1 Pde1a、54 Sst Chodl、53 Sst Mme Fam114a1、52 Sst Tac1 Htr1d、50 Sst Tac1 Tacr3、49 Sst Calb2 Necab1、48 Sst Calb2 Pdlim5、46 Sst Nr2f2 Necab1、45 Sst Myh8 Etv1、44 Sst Chrna1 Glra3、42 Sst Myh8 Fibin、40 Sst Chrna2 Ptgdr、39 Sst Tac2 Myn4、37 Sst HpseSema3c, 36 Sst Hpse Cbln4, 34 Sst Crh2 Efemp1, 33 Sst Crh2 4930553C11Rik, 31 Sst Esm1, 29 Sst Tac2 Tacstd2, 28 Sst Rxfp1 Eya1, 27 Sst Rxfp1 Prdm8, 23 Sst Nts, Pvalb Gabrg1, 20 Pvalb Th Sst, 18 Pvalb Calb1 Sst, 17 Pvalb Akr1c18 Ntf3, 16 Pvalb Sema3e Kank4, 14 Pvalb Gpr149 IsIr, 11 Pvalb ReIn Itm2a, 10 Pvalb ReIn Tac1, 9 Pvalb Tpbg, 4 Pvalb Vlpr2, Meis2 Adamts19, 170 Astro Aqp4, 171 OPC Pdgfra Grm5, Oligo Serpinb1a, 174 Oligo Synpr, VLMC Osr1 Cd74, VLMC Osr1 Mc5r, VLMC Spp1 Col15a1, Peri Kcnj8, SMC Acta2, Endo Ctla2a and 181 Microglia Siglech. (Figure 9D) Epifluorescence microscopy image of native SYFP2 fluorescence in fixed brain sections from rhesus monkey primary motor cortex 64 days after stereotaxic injection of enhancer AAV vector CN2251 of serotype PHP.eB. Genomic coordinates are mapped to build Hg38 and figures are taken from the UCSC Genome Browser. Scale bar: 500 μm.
[0022] [Figure 10A-10B] Vector: CN2251, Enhancer: eHGT_453m, Animal: Mouse. (Figure 10A) Fluorescent composite image of native SYFP2 in a coronal section of a mouse cerebellum and (Figure 10B) a magnified image of the mouse cerebellum are shown, demonstrating that SYFP2 is selectively expressed in cells of the deep cerebellar nuclei. Adult mice were administered virus by intravascular (IV) injection (retroorbital) of CN2251 virus packaged with PHP.eB capsids.
[0023] [Figures 11A-11C] (Figure 11A) Fluorescence expression of CN1633 virus (eHGT_140h) shown in black in sagittal sections of whole mouse brain. (Figure 11B) High-resolution image of mRNA expression of Gad1 and Pvalb, markers of GABAergic neurons, overlaid with SYFP2 fluorescence of CN1633 virus. Arrows indicate cells labeled with SYFP2. (Figure 11C) Single-cell transcriptome characterization of SYFP2 fluorescent cells isolated from mouse V1. After performing single-cell gene expression analysis, each cell was mapped to the existing taxonomic classification of mouse V1 cells. Note that nearly all cells are Pvalb positive neurons.
[0024] [Figures 12A-12C] (Figure 12A) Fluorescence expression of CN1621 virus (eHGT_128h) shown in black in sagittal sections of whole mouse brain. (Figure 12B) High-resolution image of Pvalb mRNA expression, a marker of GABAergic neurons, overlaid with SYFP2 fluorescence of CN1621 virus. Arrows indicate SYFP2-labeled cells. (Figure 12C) Single-cell transcriptome characterization of SYFP2-fluorescent cells isolated from mouse V1. After performing single-cell gene expression analysis, each cell was mapped to the existing taxonomic classification of mouse V1 cells. Note that nearly all cells are Pvalb-positive neurons.
[0025] [Figure 13A-13B] CN2717(eHGT_710m) in mouse neocortex. (FIG. 13A) Inverted epifluorescence microscope image showing native SYFP2 expression in neocortex 40 days after retro-orbital delivery of AAV vector #CN2717(eHGT_710m) at 6.0×1011 viral genome copies. Scale bar: 200 microns. (FIG. 13B) High magnification image of sparse cell bodies and characteristic chandelier cell axon cartridge. Scale bar: 50 microns.
[0026] [Figure 14A-14B]CN2717 (eHGT_710m) in the frontal cortex of rhesus monkeys. (FIG. 14A) Epifluorescence microscopy (inverted) image showing native SYFP2 expression in the superior frontal cortex region 43 days after stereotaxic injection of 3.46×1011 viral genome copies of AAV vector #CN2717 (eHGT_710m) into adult rhesus monkeys in vivo. Scale bar: 200 microns. (FIG. 14B) High magnification image of the cell body and axon cartridge of a characteristic chandelier cell. Scale bar: 50 microns.
[0027] [Figure 15A-15B] Vector: CN2717, Enhancer: eHGT_710m, Animal: Mouse C57BL6J-560070. (FIG. 15A) Fluorescent composite image of native SYFP2 in a coronal section of a mouse cerebellum and (FIG. 15B) a magnified image of the mouse cerebellum showing that SYFP2 is selectively expressed in cells with the morphology of small interneurons in the molecular layer. Adult mice were administered virus by intravascular (IV) injection (retroorbital) of CN2717 virus packaged with PHP.eB capsids.
[0028] [Figure 16A-16B] (FIG. 16A) An inverted epifluorescence microscope image shows native SYFP2 expression in the brain 2 months after retro-orbital delivery of AAV vector #CN2437 (eHGT_607h) at approximately 1.0×10 viral genome copies. Mice were transduced by intravenous delivery of AAV packaged with PHP.eB capsids. Scale bar: 1 mm. (FIG. 16B) A high magnification image of the thalamic region is shown. Scale bar: 500 microns.
[0029] [Figure 17A-17B](FIG. 17A) An inverted epifluorescence microscope image shows native SYFP2 expression in the brain 2 months after retro-orbital delivery of AAV vector #CN3000 (eHGT_827h) at approximately 1.0×10 viral genome copies. Mice were transduced by intravenous delivery of AAV packaged with PHP.eB capsids. Scale bar: 1 mm. (FIG. 17B) A high magnification image of the thalamic region is shown. Scale bar: 500 microns.
[0030] [Figure 18A-18B] (FIG. 18A) An inverted epifluorescence microscope image shows native SYFP2 expression in the brain 2 months after retro-orbital delivery of AAV vector #CN3003 (eHGT_830h) at approximately 1.0×10 viral genome copies. Mice were transduced by intravenous delivery of AAV packaged with PHP.eB capsids. Scale bar: 1 mm. (FIG. 18B) A high magnification image of the thalamic region is shown. Scale bar: 500 microns.
[0031] [Figure 19A-19B] (FIG. 19A) An inverted epifluorescence microscope image shows native SYFP2 expression in the brain 2 months after retro-orbital delivery of AAV vector #CN3005 (eHGT_832h) at approximately 1.0×10 viral genome copies. Mice were transduced by intravenous delivery of AAV packaged with PHP.eB capsids. Scale bar: 1 mm. (FIG. 19B) A high magnification image of the thalamic region is shown. Scale bar: 500 microns.
[0032] [Figure 20A-20B] (FIG. 20A) An inverted epifluorescence microscope image shows native SYFP2 expression in the brain 2 months after retro-orbital delivery of AAV vector #CN3007 (eHGT_834h) at approximately 1.0×10 viral genome copies. Mice were transduced by intravenous delivery of AAV packaged with PHP.eB capsids. Scale bar: 1 mm. (FIG. 20B) A high magnification image of the thalamic region is shown. Scale bar: 500 microns.
[0033] [Figure 21A-21B] (FIG. 21A) An inverted epifluorescence microscope image shows native SYFP2 expression in the brain 2 months after retro-orbital delivery of AAV vector #CN3009 (eHGT_836h) at approximately 1.0×10 viral genome copies. Mice were transduced by intravenous delivery of AAV packaged with PHP.eB capsids. Scale bar: 1 mm. (FIG. 21B) A high magnification image of the thalamic region is shown. Scale bar: 500 microns.
[0034] [Fig. 22A-22B] (FIG. 22A) An inverted epifluorescence microscope image shows native SYFP2 expression in the brain 2 months after retro-orbital delivery of AAV vector #CN2786 (3xcore_eHGT_121h) at approximately 1.0×1012 viral genome copies. Mice were transduced by intravenous delivery of AAV packaged with PHP.eB capsids. Scale bar: 1 mm. (FIG. 22B) A high magnification image of the thalamic region is shown. Scale bar: 500 microns.
[0035] [Figure 23] A table of abbreviations for brain structures is shown.
[0036] [Figure 24]The following sequences are provided in support of this disclosure: eHGT_576h (SEQ ID NO: 1), eHGT_577h (SEQ ID NO: 2), eHGT_578h (SEQ ID NO: 3), eHGT_579h (SEQ ID NO: 4), eHGT_606h (SEQ ID NO: 5), eHGT_827h (SEQ ID NO: 137), eHGT_828h (SEQ ID NO: 138), eHGT_830h (SEQ ID NO: 6), eHGT_831h (SEQ ID NO: 7), eHGT_832h (SEQ ID NO: 8), eHGT_834h (SEQ ID NO: 9), eHGT_836h (SEQ ID NO: 10), MGT_E117 (SEQ ID NO: 11), MGT_E118 (SEQ ID NO: 12), MGT_E119 (SEQ ID NO: 13), MGT_E120 (SEQ ID NO: 14), MGT_E121 (SEQ ID NO: 15), eHGT_717h (SEQ ID NO: 16), 3xCore2-eHGT_367h (SEQ ID NO: 17), the core region of eHGT_367h (SEQ ID NO: 18), eHGT_359h (SEQ ID NO: 19), eHGT_479m (SEQ ID NO: 20), eHGT_453m (SEQ ID NO: 21), eHGT_140h (SEQ ID NO: 22), eHGT_356h (SEQ ID NO: 23), eHGT_128h (SEQ ID NO: 24), eHGT_369h(also called 3xcoreB_eHGT121h) (SEQ ID NO: 25), coreB_eHGT121h (SEQ ID NO: 26), the 300 bp core region of eHGT_369h (2xcoreB_eHGT121h) (SEQ ID NO: 27), eHGT_710m (SEQ ID NO: 28), eHGT_895h (SEQ ID NO: 29), 3xcore2_eHGT_577h (SEQ ID NO: 30), core2_eHGT_577h (SEQ ID NO: 31), 3xcore3_eHGT_577h (SEQ ID NO: 32), core3_eHGT_577h (SEQ ID NO: 33), 3xcore2_eHGT_606h (SEQ ID NO: 34), core2_eHGT_606h (SEQ ID NO: row number 35), 3xcore3_eHGT_606h (SEQ ID NO: 36), core3_eHGT_606h (SEQ ID NO: 37), core4_eHGT_577h (SEQ ID NO: 38), core6_eHGT_606h (SEQ ID NO: 39), 3xcore_eHGT_121h (SEQ ID NO: 40), core_eHGT_121h (SEQ ID NO: 41), eHGT_590m (SEQ ID NO: 42), eHGT_976h (SEQ ID NO: 43), β-globin minimal promoter (pBGmin / minBGlobin / minBGprom) (SEQ ID NO: 45), minCMV promoter (SEQ ID NO: 46), mutated minCMV promoter (SacIRE site removed) (SEQ ID NO: 47), minRho promoter (SEQ ID NO: 48), minRho* promoter (SEQ ID NO: 49), Hsp68 minimal promoter (proHsp68) (SEQ ID NO: 50), SYFP2 (SEQ ID NO: 51), EGFP (SEQ ID NO: 52), optimized Flp recombinase (FlpO) (SEQ ID NO: 53), improved Cre recombinase (iCre) (SEQ ID NO: 54), SP10 insulator (SP10ins) (SEQ ID NO: 55), 3xSP10ins (SEQ ID NO: 56), WPRE3 (SEQ ID NO: 57), WPRE (SEQ ID NO: 58), BGHpA (SEQ ID NO: 59), No. 59), HGHpA (SEQ ID NO: 60), 3XFLAG (SEQ ID NO: 61), hsA2 (SEQ ID NO: 62), 10a.a. (SEQ ID NO: 63), H2B (SEQ ID NO: 64), P2A (SEQ ID NO: 65 or SEQ ID NO: 66), T2A (SEQ ID NO: 67), E2A (SEQ ID NO: 68), F2A (SEQ ID NO: 69), representative plasmid backbone 1 - left ITR (SEQ ID NO: 70), representative plasmid backbone 1 - right ITR (SEQ ID NO: 71), representative plasmid backbone 2 - left ITR (SEQ ID NO: 72), representative plasmid backbone 2 - right ITR (SEQ ID NO: 73), PHP.eB capsid (SEQ ID NO: 74), AAV9 VP1 capsid protein (SEQ ID NO: 75), tet-transactivator version 2 (tTA2) (SEQ ID NO: 76), GTPase HRas [Homo sapiens] (SEQ ID NO: 77), substance P consisting of residues 58 to 68 of protachykinin-1 [Homo sapiens] (SEQ ID NO: 78), oxytocin-neurophysin complex 1 [Homosapiens] consisting of amino acids 20 to 28, GCaMP6m (SEQ ID NO: 80), GCaMP6s (SEQ ID NO: 81), GCaMP6f (SEQ ID NO: 82), CN2415 (SEQ ID NO: 83), CN2416 (SEQ ID NO: 84), CN2417 (SEQ ID NO: 85), CN2418 (SEQ ID NO: 86), CN2436 (SEQ ID NO: 87), CN3000 (SEQ ID NO: 139), CN3001 (SEQ ID NO: 140), CN3003 (SEQ ID NO: 88), CN3004 (SEQ ID NO: 89), CN3005 (SEQ ID NO: 90), CN3007 (SEQ ID NO: 91), CN3009 (SEQ ID NO: 92), AiP1335 (SEQ ID NO: 93), AiP1336 (SEQ ID NO: 94), AiP1337 (SEQ ID NO: 95), AiP1338 (SEQ ID NO: 96), AiP1 339 (SEQ ID NO: 97), CN2555 (SEQ ID NO: 99), CN2045 (SEQ ID NO: 100), CN2258 (SEQ ID NO: 101), CN2251 (SEQ ID NO: 102), CN1633 (SEQ ID NO: 103), CN2043 (SEQ ID NO: 104), CN1621 (SEQ ID NO: 105), CN2216 (SEQ ID NO: 106), CN2717 (SEQ ID NO: 107), CN3639 (SEQ ID NO: 108), CN3050 (SEQ ID NO:109), CN3051 (SEQ ID NO:110), CN3056 (SEQ ID NO:111), CN3057 (SEQ ID NO:112), CN4001 (SEQ ID NO:113), CN4003 (SEQ ID NO:114), CN2786 (SEQ ID NO:115), CN2840 (SEQ ID NO:116), CN3460 (SEQ ID NO:117) and CN2650 (SEQ ID NO:118). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0037] To fully understand the biology of the brain, we need to distinguish, define, and dissect different cell types, as well as identify artificial expression constructs that can label and perturb them (Tasic, Curr. Opin. Neurobiol. 50, 242-249 (2018); Zeng & Sanes, Nat. Rev. Neurosci. 18, 530-546 (2017)). In mice, driver lines expressing recombinases have been used successfully to label cell populations that share marker gene expression (Daigle et al., Cell 174, 465-480.e22 (2018); Taniguchi, et al., Neuron 71, 995-1013 (2011); Gong et al., J. Neurosci. 27, 9817-9823 (2007)). However, the creation, maintenance and use of such lines capable of labeling specific cell types with high specificity is costly and often requires crossing of transgenic animals between three species, resulting in low frequency of obtaining the desired experimental animals. Moreover, these tools cannot be applied to humans because of the need for germline transgenic animals.
[0038] The present disclosure provides an artificial expression construct for inducing gene expression in targeted central nervous system cell populations. The targeted central nervous system cell populations include GABAergic neurons in the thalamus (Gata / Dlx5-6), GABAergic neurons in the thalamic reticular nucleus (TRN) of the thalamus, thalamic reticular nucleus cells, glutamatergic neurons in the thalamus, glutamatergic neurons in the thalamus (Prkcd-Grin2c (core, LGN)), glutamatergic neurons in the thalamus (Rxfp1-Epb4 (matrix)), and glutamatergic neurons in the parafascicular nucleus (Pf) of the thalamus. In certain embodiments, the artificial expression constructs described herein induce gene expression in a second type of target cell. Representative second type cells include striatal medium spiny neurons (including all types), Purkinje cells in the cerebellum, deep cerebellar nucleus (DCN) cells in the cerebellum, molecular layer interneuron (MLI) cells in the cerebellum, parvalbumin (Pvalb)-positive neurons, chandelier cells, glutamatergic layer 5 neurons that project outside the telencephalon (L5 ET) cells in the neocortex, and Vip-positive neurons in the neocortex.
[0039] In certain embodiments of the artificial expression constructs of the invention, the following enhancers are utilized to induce gene expression in targeted central nervous system cell populations. The enhancer and target cell population combinations used in the invention are listed below in the order of enhancer / target cell population. eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_ 830h, eHGT_831h, eHGT_832h, eHGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHG T_577h, 3xcore3_eHGT_577h, 3xcore2_eHGT_606h, 3xcore3_eHGT_606h, core4_eHGT_577h, core6_eHGT_606h, 3xcore_eHGT_121h, eHGT_590m or eHGT_976h / glutamatergic neurons in the thalamus; MGT_E117 or MGT_E118 / glutamatergic neurons in the thalamus (Prkcd-Grin2c (core, LGN)); MGT_E119 or MGT_E120 / GABAergic neurons in the thalamus (Gata / Dlx5-6); MGT_E121 / glutamatergic neurons in the thalamus (Rxfp1-Epb4(matrix)); 3xCore2-eHGT_367h / glutamatergic neurons in the parafascicular nucleus (Pf) of the thalamus and striatal medium spiny neurons (including all types); eHGT_359h / glutamatergic neurons in the thalamus, Purkinje cells in the cerebellum, and Pvalb-positive neurons in the neocortex; eHGT_479m / glutamatergic neurons in the thalamus, Purkinje cells in the cerebellum, and chandelier cells in the neocortex; eHGT_453m / GABAergic neurons in the thalamic reticular nucleus (TRN) of the thalamus, deep cerebellar nucleus (DCN) cells of the cerebellum, and glutamatergic layer 5 neurons (L5 ET) cells projecting outside the telencephalon in the neocortex; eHGT_140h / GABAergic and Pvalb-positive neurons in the thalamic reticular nucleus (TRN) of the thalamus; eHGT_356h / GABAergic neurons in the thalamic reticular nucleus (TRN) of the thalamus, DCN cells of the cerebellum, Vip-positive neurons and thalamic reticular nucleus cells of the neocortex; eHGT_128h / glutamatergic and Pvalb-positive neurons in the thalamus; eHGT_369h / glutamatergic neurons in the thalamus, molecular layer interneuron (MLI) cells in the cerebellum, and Pvalb-positive neurons in the neocortex; and eHGT_710m / glutamatergic neurons in the thalamus, MLI cells, chandelier cells and GABAergic interneurons in the molecular layer of the cerebellum.
[0040] In certain embodiments, multiple copies of the enhancer are concatenated (linearly linked) or multiple copies of the enhancer core are concatenated. Examples include the core regions of eHGT_367h, eHGT_121h, eHGT_577h and / or eHGT_606h, or concatenated versions of these core regions. By using such artificial enhancer elements, the transgene can be expressed more rapidly and higher than when the original (natural) full-length enhancer is used alone.
[0041] In certain embodiments, the enhancer core comprises a sequence as set forth in any of SEQ ID NO: 18, SEQ ID NO: 26, SEQ ID NO: 31, SEQ ID NO: 33, SEQ ID NO: 35, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39, and SEQ ID NO: 41. In certain embodiments, these enhancer cores are concatemerized and have 2, 3, 4, 5, 6, 7, 8, 9, or 10 copies of the core sequence. In certain embodiments, a concatemer comprising 3 copies of a selected enhancer core comprises a sequence as set forth in any of SEQ ID NO: 17, SEQ ID NO: 25, SEQ ID NO: 30, SEQ ID NO: 32, SEQ ID NO: 34, SEQ ID NO: 36, and SEQ ID NO: 40.
[0042] Particular embodiments of the enhancer core utilize Core2-eHGT_367h, coreB_eHGT121h, core2_eHGT_577h, core3_eHGT_577h, core2_eHGT_606h, core3_eHGT_606h, core4_eHGT_577h, core6_eHGT_606h or core_eHGT_121h. In certain embodiments of the concatemerized enhancer cores, 3xCore2-eHGT_367h, eHGT_369h (3xcoreB_eHGT121h), 3xcore2_eHGT_577h, 3xcore3_eHGT_577h, 3xcore2_eHGT_606h, 3xcore3_eHGT_606h, core6_eHGT_606h or 3xcore_eHGT_121h are utilized. In the present disclosure, eHGT_369h can be replaced with 3xCoreB-eHGT_121h.
[0043] Certain embodiments provide artificial expression constructs that include the features of vectors described herein, including CN2415, CN2416, CN2417, CN2418, CN2436, CN3000, CN3001, CN3003, CN3004, CN3005, CN3007, CN3009, AiP1335, AiP1336, AiP1337, AiP1338, AiP1339, AiP2000, AiP2001, AiP2002, AiP2003, AiP2004, AiP2005, AiP2006, AiP2007, AiP2009, AiP2008, AiP2009, AiP2009, AiP2006, AiP2008, AiP2009 ... These include iP1338, AiP1339, CN2555, CN2045, CN2258, CN2251, CN1633, CN2043, CN1621, CN2216, CN2717, CN3639, CN3050, CN3051, CN3056, CN3057, CN4001, CN4003, CN2786, CN2840, CN3460, and CN2650.
[0044] Various aspects of the disclosure are described in more detail below with further options. Various aspects of the disclosure are described under the following headings: (i) artificial expression constructs and vectors for targeted expression of genes in target cells; (ii) compositions for administration; (iii) cell lines containing artificial expression constructs; (iv) transgenic animals; (v) methods of use; (vi) kits and commercial packages; (vii) representative embodiments; and (viii) conclusion. These headings are provided for organizational purposes only and are not intended to limit the scope or interpretation of the disclosure.
[0045] (i) Artificial expression constructs and vectors for targeted expression of genes in target cells The artificial expression constructs disclosed herein comprise (i) an enhancer sequence that induces targeted expression of a coding sequence in a targeted central nervous system cell, (ii) a coding sequence to be expressed, and (iii) a promoter. The artificial expression constructs of the invention may further comprise other regulatory elements as needed or beneficial.
[0046] In certain embodiments, an "enhancer" or "enhancer element" is a cis-acting sequence that increases the amount of transcription associated with a promoter, can function in either the forward or reverse orientation relative to the promoter and the coding sequence to be transcribed, and can be located upstream or downstream relative to the promoter or the coding sequence to be transcribed. There are various methods or techniques known in the art for measuring the function of enhancer element sequences. Specific examples of enhancer sequences used in the artificial expression constructs disclosed herein include eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_830h, eHGT_831h, eHGT_832h, eHGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHGT_577h, 3xcore3_eHGT_577h, 3xcore2 ... eHGT_606h, 3xcore3_eHGT_606h, core4_eHGT_577h, core6_eHGT_606h, 3xcore_eHGT_121h, eHGT_590m, eHGT_976h, MGT_E117, MGT_E118, MGT_E119, MGT_E120, MGT_E121, 3xCore2-eHGT_367h, eHGT_359h, eHGT_479m, eHGT_453m, eHGT_140h, eHGT_356h, eHGT_128h, eHGT_369h and eHGT_710m.
[0047] In certain embodiments, the enhancer used in targeting central nervous system cells is the enhancer that is only utilized in targeting central nervous system cells, or is the enhancer that is mainly utilized in targeting central nervous system cells.The enhancer used in targeting central nervous system cells is the enhancer that enhances the expression of genes in targeting central nervous system cells.In certain embodiments, the enhancer used in targeting central nervous system cells enhances the expression of genes in targeting central nervous system cells, but does not substantially induce the expression of genes in other non-targeting cells, and therefore is also the targeting central nervous system enhancer that has cell type-specific transcription activity.
[0048] When a heterologous coding sequence operably linked to an enhancer disclosed herein is expressed in a targeted cell, the administered heterologous coding sequence is expressed in cells of the intended type.
[0049] When a heterologous coding sequence is preferentially expressed in selected cells, the administered heterologous coding sequence is expressed in the intended type of cells, but is not substantially expressed in other types of cells. This is described in more detail below. In certain embodiments, not substantially expressed in other types of cells means less than 50% expression in the reference cells compared to the target cells; less than 40% expression in the reference cells compared to the target cells; less than 30% expression in the reference cells compared to the target cells; less than 20% expression in the reference cells compared to the target cells; or less than 10% expression in the reference cells compared to the target cells. In certain embodiments, "reference cells" refers to non-target cells. Non-target cells may be in the same anatomical structure as the target cells and / or may project to a common anatomical region. In certain embodiments, the reference cells are in an anatomical structure adjacent to the anatomical structure in which the target cells are included. In certain embodiments, the reference cells are non-target cells that have a different gene expression profile than the target cells.
[0050] In certain embodiments, the transcription product of the coding sequence may be expressed at a low level in unselected cells, for example, at less than 1% or 1%, 2%, 3%, 5%, 10%, 15% or 20% of the transcription product expression in selected cells. In certain embodiments, the targeted central nervous system cells are the only type of cells that can express the correct combination of various transcription factors that can bind to the enhancers disclosed herein and induce gene expression. Thus, in certain embodiments, expression occurs only in the targeted cell type.
[0051] In certain embodiments, target cells (e.g., neuronal and / or non-neuronal cells) can be identified based on transcriptional profiles, such as those described in Tasic et al., Nature 563, 72-78 (2018) and Hodge et al., Nature 573, 61-68 (2019). For reference, various types of cells and their salient features are described below.
[0052] Classification and subclassification of thalamic GABAergic neurons: · Overall: expresses the GABA synthesis genes Gad1 / GAD1 and / or Gad2 / GAD2. Thalamic reticular nucleus (TRN) neurons express the GABA synthesis genes Gad1 / GAD1 and Pvalb / PVALB.
[0053] Classification and subclassification of thalamic glutamatergic neurons: Whole glutamatergic neurons: express the glutamate transporters Slc17a6 / SLC17A6 and / or Slc17a7 / SLC17A7. Glutamatergic neurons lack expression of Gad1 / Gad2 and express one or more of the following marker genes: Synpo2 / SYNPO2, Rgs16 / RGS16, Plekhg1 / PLEKHG1 and Prkcd / PRKCD. Glutamatergic neurons in the parafascicular nucleus (Pf): The parafascicular nucleus (Pf) is the posterior component of the intralaminar thalamic nucleus. It plays a role in the feedback system of the basal ganglia-thalamo-cortical circuit, which is crucial for cognitive processes.
[0054] Subclassification of GABAergic neurons in the neocortex: · Overall: expresses the GABA synthesis genes Gad1 / GAD1 and / or Gad2 / GAD2. · Lamp5-, Sncg-, Serpinf1- and Vip-positive GABAergic neurons: neurons that arise from neural precursor cells derived from the caudal ganglia primordium (CGE) or preoptic area (POA) during development. · Sst and Pvalb positive GABAergic neurons: neurons that arise from neural precursor cells derived from the medial ganglia primordium (MGE) during development. Lamp5-positive GABAergic neurons are found in numerous neocortical layers, especially in the upper layers (L1-L2 / 3), and mainly have neurogliaform and single bouquet cell morphologies. · Lamp5_Lhx6 positive GABAergic neurons: a subset of Lamp5 positive GABAergic neurons that co-express Lamp5 and Lhx6. Sncg-positive GABAergic neurons are found in many neocortical layers and contain molecules common to Lamp5-positive and Vip-positive cells, but the expression of Lamp5 and vasoactive intestinal peptide (Vip) is inconsistent, whereas the expression of Sncg is consistent. Serpinf1-positive GABAergic neurons are found in many neocortical layers and share molecules common to Sncg-positive and Vip-positive cells, but the expression of Sncg and Vip is inconsistent, whereas the expression of Serpinf1 is consistent. Vip-positive GABAergic neurons: Found in many neocortical layers, but are particularly common in the upper layers (L1-L4) and highly express the neurotransmitter Vip. · Sst positive GABAergic neurons: found in many neocortical layers, but especially frequent in the lower layers (L5-L6). They highly express the neurotransmitter somatostatin (Sst) and frequently block dendritic inputs to postsynaptic neurons. This subclass includes sleep-active Sst Chodl neurons (which additionally express Nos1 and Tacr1), which are significantly different from other Sst neurons, but express some shared marker genes, including Sst. Expression of the SST gene in humans is frequently detected in a subtype of LAMP5+ GABAergic neurons in layer 1. Pvalb-positive GABAergic neurons are found in many neocortical layers, but are especially prevalent in the lower layers (L5-L6). These neurons highly express the calcium-binding protein parvalbumin (Pvalb) and the neuropeptide Tac1, which often dampen the output of postsynaptic neurons. Most GABAergic neurons with fast firing properties highly express Pvalb. This subclass includes chandelier cells, which have a characteristic chandelier-like morphology and express the markers Cpne5 and Vipr2 in mice and NOG and UNC5B in humans. Meis2: A distinct subclass defined by one type of cell, neocortical GABAergic neurons, that express the Meis2 gene but do not express several other genes expressed by other neocortical GABAergic neurons (e.g. Thy1 and Scn2b). These cells are found in L6b and subcortical white matter.
[0055] Subclassification of neocortical glutamatergic neurons: Overall: These neurons express the glutamatergic transmitters Slc17a6 and / or Slc17a7. Both neurons express Snap25 and lack Gad1 / Gad2 expression. · L2 / 3 IT glutamatergic neurons: located mainly in layers 2 and 3, with predominantly intratelencephalic (intracortical) projections. · L4 IT glutamatergic neurons: located mainly in layer 4, with mainly local or intratelencephalic (intracortical) projections. L5 IT glutamatergic neurons: mainly found in layer 5, with predominantly intratelencephalic (intracortical) projections. Also called L5a. · L5 PT glutamatergic neurons: mainly found in layer 5, mainly with cortico-subcortical (pyramidal or corticofugal) projections. Also called L5b or L5 CF (corticofugal) or L5 ET (extratelencephalic). This subclassification includes cells present in the primary motor cortex and adjacent areas, and is a corticospinal projection type neuron associated with motor neuron / movement disorders (e.g. ALS). This subclassification includes thick tufted pyramidal neurons, including specialized subtypes found only in certain areas, e.g. Betz, Meynert and von Economo cells. · L5 NP glutamatergic neurons: mainly located in layer 5 and project mainly to nearby areas. · L6 CT glutamatergic neurons: mainly found in layer 6, mainly project to the corticothalamus. · L6 IT glutamatergic neurons: located mainly in layer 6 and project mainly intratelencephalic (intracerebrocortical) L6 IT Car3 glutamatergic neurons: Most densely present in the claustrum and endopyriform nucleus, but sparsely present throughout L6 in many cortical areas, including the primary visual cortex. These neurons project primarily intratelencephalic (intracerebrospinal). Additional marker genes for claustrum-enriched neurons include Gnb4 and Ntng2. · L6b glutamatergic neurons: mainly present in the neocortical subplate (L6b), project locally (near the cell body), and some also project cortically from the VISp to the anterior cingulate bundle and cortico-subcortically to the thalamus. CR neurons: a distinctive subclass defined by a single type found in L1. Cajal-Retzius cells express the characteristic molecular markers Lhx5 and Trp73. ·Cerebellar Purkinje cells: Large GABAergic neurons that are the only projection neurons that originate exclusively from the cerebellum. The cell bodies of cerebellar Purkinje cells form a single layer known as the "Purkinje cell layer" and express parvalbumin. Deep cerebellar nucleus (DCN) neurons: neurons present in the deep cerebellar nucleus structures. These neurons include glutamatergic and GABAergic cells that express the Pvalb gene. · Molecular layer interneuron cells (MLIs): intracerebellar neurons that participate in spatially structured networks via chemical and electrical synapses. · Chandelier cells: specialized GABAergic interneurons that selectively innervate pyramidal neurons. Striatal medium spiny neurons: These are the main striatal neurons. They receive synaptic input from glutamatergic and dopaminergic afferents.
[0056] Subclassification of non-neuronal cells: Astrocytes: Glial cells derived from the neuroectoderm that express the marker Aqp4 and often also GFAP, but not the neuronal marker SNAP25. Astrocytes can have a characteristic star-shaped morphology and are involved in the metabolic support of other cells in the brain. Many types of astrocyte morphology are observed in mice and humans. Oligodendrocytes: Neuroectoderm-derived glial cells that express the Sox10 marker. This category includes oligodendrocyte precursor cells (OPCs). Oligodendrocytes are a subcategory primarily responsible for myelination of neurons. · VLMC: Vascular leptomeningeal cells (VLMCs), which are part of the meninges surrounding the outer layer of the cortex, express the marker genes Lum and Col1a1. Pericytes: Blood vessel-associated cells that express the marker genes Kcnj8 and Abcc9. Pericytes surround endothelial cells and are important in regulating capillary blood flow and are involved in blood-brain barrier permeability. Smooth muscle cells (SMCs): Specialized smooth muscle cells, vascular associated cells that express the Acta2 marker gene. Smooth muscle cells line arterioles in the brain and are involved in blood-brain barrier crossing. Endothelial cells: These are the cells that line the blood vessels in the brain. Endothelial cells express the Tek and PDGF-B markers. Microglia: immune cells derived from hematopoietic cells, macrophages localized in brain tissue, perivascular macrophages (PVM) that can migrate from the blood and associate with brain tissue, and can be seen as a by-product of brain dissection procedures. Microglia are known to express Cx3cr1, Tmem119, and PTPRC (CD45).
[0057] In certain embodiments, the coding sequence is a heterologous coding sequence that codes for an effector element.Effector element is the sequence that is expressed to obtain an intended effect, and the intended effect is actually achieved by this effector element.Examples of effector elements include reporter genes / proteins and functional genes / proteins.
[0058] Representative reporter genes / proteins include those expressed by Addgene ID No. 83894 (pAAV-hDlx-Flex-dTomato-Fishell_7), ID No. 83895 (pAAV-hDlx-Flex-GFP-Fishell_6), ID No. 83896 (pAAV-hDlx-GiDREADD-dTomato-Fishell-5), ID No. 83898 (pAAV-mDlx-ChR2-mCherry-Fishell-3), ID No. 83899 (pAAV-mDlx-GCaMP6f-Fishell-2), ID No. 83900 (pAAV-mDlx-GFP-Fishell-1), and ID No. 89897 (pcDNA3-FLAG-mTET2(N500)). Representative reporter genes include, inter alia, expressible fluorescent proteins or expressible biotin; blue fluorescent proteins (e.g., eBFP, eBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-sapphire); cyan fluorescent proteins (e.g., eCFP, Cerulean, CyPet, AmCyanl, Midoriishi-Cyan, mTurquoise); green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, Emerald, Azami Green, Monomeric Azami Green (mAzamigreen), CopGFP, AceGFP, avGFP, ZsGreen1, Oregon Green, etc.) TM (Thermo Fisher Scientific); luciferase; orange fluorescent proteins (mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, tdTomato, dTomato); red fluorescent proteins (mKate, mKate2, mPlum, DsRed monomer, mCherry, mRuby, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred, Texas Red) TM(Thermo Fisher Scientific); far-red fluorescent proteins (e.g., mPlum and mNeptune); yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, SYFP2, Venus, YPet, PhiYFP, ZsYellow1); or reporter genes encoding tandemly linked complexes.
[0059] GFP is composed of 238 amino acids (26.9 kDa) and was first isolated from the jellyfish Aequorea victoria / Aequorea aequorea / Aequorea forskalea, which fluoresces green when exposed to blue light. GFP isolated from A. victoria has a major excitation peak at a wavelength of 395 nm and a minor excitation peak at 475 nm. Its emission peak is at 509 nm, which is in the low wavelength range of green light in the visible spectrum. GFP from the sea pansy (Renilla reniformis) has one major excitation peak at 498 nm. Due to its wide range of applications and the demand from researchers for further improvements, various GFP variants have been created. The first major improvement was a single point mutation (S65T) reported in Nature by Roger Tsien in 1995. This mutation dramatically improved the spectral properties of GFP, increasing its fluorescence and photostability, and shifting the major excitation peak to 488 nm while maintaining the emission peak at 509 nm. Enhanced GFP (EGFP) was obtained by adding a point mutation (F64L) to GFP that improves folding efficiency at 37 °C. EGFP has an extinction coefficient (denoted ε) of 55,000 L / mol cm, which is 9.13 × 10 per molecule. -21 m 2 Also known as the optical cross section of GFP, Superfolder GFP was reported in 2006 as a series of GFP mutants that can rapidly fold and mature even when fused to poorly folded peptides.
[0060] "Yellow fluorescent protein" (YFP) is a genetic variant of the green fluorescent protein derived from Aequorea victoria. Its excitation peak is at 514 nm and its emission peak is at 527 nm.
[0061] Representative functional molecules include ion transporters, cell transport proteins, enzymes, transcription factors, neurotransmitters, calcium reporters, channelrhodopsins, guide RNAs, nucleases, microRNAs, or designer receptors activated only by designer drugs (DREADDs), each of which has a function.
[0062] Ion transporters are transmembrane proteins responsible for the transport of ions across cell membranes. Ion transporters are found in the majority of cells and are important in regulating cellular excitability and homeostasis. Ion transporters are involved in numerous cellular processes, including action potentials, synaptic transmission, hormone secretion, and muscle contraction. Many biological processes important to living cells involve the transport of calcium ions (Ca) through ion channels. 2+ ), potassium ion (K + ), sodium ion (Na + In certain embodiments, ion transporters include voltage-gated sodium channels (e.g., SCN1A), potassium channels (e.g., KCNQ2), and calcium channels (e.g., CACNA1C).
[0063] Representative enzymes, transcription factors, receptors, membrane proteins, cell transport proteins, signaling molecules and neurotransmitters include enzymes such as lactase, lipase, helicase, α-glucosidase, aromatic L-amino acid decarboxylase (AADC), and amylase; transcription factors such as SP1, AP-1, heat shock factor protein 1, C / EBP (CCAAT / enhancer binding protein), and Oct-1; transforming growth factor receptor β1, platelet-derived growth factor receptor, epidermal growth factor receptor, vascular endothelial growth factor receptor, interleukin-1 (IL-1), and IL-2; These include receptors such as leukin-8 receptor α; membrane proteins and cell trafficking proteins such as clathrin, dynamin, caveolin, Rab4A, and Rab-11A; signaling molecules such as nerve growth factor (NGF), glial cell line-derived neurotrophic factor (GDNF), platelet-derived growth factor (PDGF), transforming growth factor beta (TGFβ), epidermal growth factor (EGF), GTPases, and HRas; and neurotransmitters such as cocaine- and amphetamine-regulated transcript, substance P, oxytocin, and somatostatin.
[0064] In certain embodiments, the functional molecules include reporters that indicate cell function and state, such as calcium reporters. Intracellular calcium concentration is an important predictor of many cellular activities, such as neuronal activation, muscle cell contraction, and second messenger signaling. A sensitive and simple technique for monitoring intracellular calcium concentration is the use of genetically encoded calcium indicators (GECIs). Among GECIs, a green fluorescent protein (GFP)-based calcium sensor, named GCaMP, is highly efficient and widely used. GCaMP is formed by fusing M13 and calmodulin proteins to the N- and C-termini of circularly permuted GFP. Some types of GCaMP exhibit characteristic fluorescence emission spectra (Zhao et al., Science, 2011, 333(6051): 1888-1891). Representative GECIs that exhibit green fluorescence include GCaMP3, GCaMP5G, GCaMP6s, GCaMP6m, GCaMP6f, jGCaMP7s, jGCaMP7c, jGCaMP7b, jGCaMP7f, jGCaMP8s, jGCaMP8m, and jGCaMP8f. In addition, GECIs that exhibit red fluorescence include jRGECO1a and jRGECO1b. AAV products containing GECIs are commercially available.For example, AAV8-CAG-GCaMP3 (catalog no. BS4-CX3AAV8), AAV8-Syn-FLEX-GCaMP6s-WPRE (catalog no. BS1-NXSAAV8), AAV8-Syn-FLEX-GCaMP6s-WPRE (catalog no. BS1-NXSAAV8), AAV9-CAG-FLEX-GCaMP6m-WPRE (catalog no. BS2-CXMAAV9), AAV9-Syn-FLEX-jGCaMP7s-WPRE ( AAV products available include AAV9-CAG-FLEX-jGCaMP7f-WPRE (Catalog No. BS12-CXFAAV9), AAV9-Syn-FLEX-jGCaMP7b-WPRE (Catalog No. BS12-NXBAAV9), AAV9-Syn-FLEX-jGCaMP7c-WPRE (Catalog No. BS12-NXCAAV9), AAV9-Syn-FLEX-NES-jRGECO1a-WPRE (Catalog No. BS8-NXAAAV9), and AAV8-Syn-FLEX-NES-jRCaMP1b-WPRE (Catalog No. BS7-NXBAAV8).
[0065] In certain embodiments, the calcium reporter includes the genetically encoded calcium indicator (GECI) NTnC; a myosin light chain kinase-GFP-calmodulin chimera; the calcium indicator TN-XXL; a BRET-based auto-luminescent calcium indicator; and / or the calcium indicator protein OeNL(Ca2+)-18μ.
[0066] In certain embodiments, the functional molecules include modulators of neuron-active channelrhodopsins (e.g., channelrhodopsin 1, channelrhodopsin 2, and variants thereof). Channelrhodopsins are a subfamily of retinylidene proteins (rhodopsins) that function as light-gated ion channels. In addition to channelrhodopsin 1 (ChR1) and channelrhodopsin 2 (ChR2), several channelrhodopsin variants have been developed. For example, Lin et al. (Biophys J, 2009, 96(5): 1803-14) describe the creation of transmembrane domain chimeras of ChR1 and ChR2 using site-directed mutagenesis. Zhang et al. (Nat Neurosci, 2008, 11(6): 631-3) describe a red-light shifted channelrhodopsin variant, VChR1. VChR1 has reduced photosensitivity and reduced membrane trafficking and expression. Other known channelrhodopsin variants include the ChR2 variant described in Nagel, et al., Proc Natl Acad Sci USA, 2003, 100(24): 13940-5, ChR2 / H134R (Nagel, G., et al., Curr Biol, 2005, 15(24): 2279-84) and ChD / ChEF / ChIEF (Lin, JY, et al., Biophys J, 2009, 96(5): 1803-14), all of which are activated by blue light (470 nm) but are insensitive to orange / red light. Other variants are described in Lin, Experimental Physiology, 2010, 96.1: 19-25; Knopfel et al., The Journal of Neuroscience, 2010, 30(45): 14998-15004; and Mardinly et al., Nat Neurosci. 2018, 21(6):881-893.
[0067] In certain embodiments, the functional molecule includes DNA and RNA editing tools, such as CRISPR / Cas (e.g., guide RNA and nuclease such as Cas, Cas9, cpf1).Furthermore, the functional molecule includes recombinant Cpf1 as described in US Patent Publication No. 2018 / 0030425, US Patent Publication No. 2016 / 0208243, WO / 2017 / 184768 and Zetsche et al. (2015) Cell 163: 759-771; single-stranded gRNA (see, for example, Jinek et al. (2012) Science 337:816-821; Jinek et al. (2013) eLife 2:e00471; Segal (2013) eLife 2:e00563), editase, guide RNA molecule, microRNA, or homologous recombination donor cassette.
[0068] In certain embodiments, the functional molecule includes a localization cassette. In certain embodiments, the localization cassette is used to localize a molecule (e.g., a vector, a protein, a sensor) to a specific subcellular compartment, such as the cell body, axon, or dendrite of a neuron. In certain embodiments, the localization cassette includes a somatic tag (e.g., soma (EE-RR)) for localization to the cell body; an axon tag (e.g., derived from GAP43) or synaptophysin (sy) for localization to the axon; a hydrophobic tail for localization to the cell membrane; and a hydrophobic or alkyl chain for localization to the endoplasmic reticulum. In certain embodiments, the localization cassette is fused to a sensor molecule, such as a GECI. In certain embodiments, the fusion protein of the localization cassette and the GECI includes soma-jGCaMP8s, axon-jRGECO1a, syGCaMP5G, and soma-jGCaMP7s.
[0069] In certain embodiments, the functional molecule includes a tag cassette. Examples of tag cassettes include His tag (HHHHHH; SEQ ID NO: 125), Flag tag (DYKDDDDK; SEQ ID NO: 126), Xpress tag (DLYDDDDK; SEQ ID NO: 127), Avi tag (GLNDIFEAQKIEWHE; SEQ ID NO: 128), calmodulin tag (KRRWKKNFIAVSAANRFKKISSSGAL; SEQ ID NO: 129), polyglutamic acid tag, HA tag (YPYDVPDYA; SEQ ID NO: 130), Myc tag (EQKLISEEDL; SEQ ID NO: 131), Strep tag (meaning the original STREP (registered trademark) tag) (WRHPQFGG; SEQ ID NO: 132), STREP tag II (WSHPQFEK; SEQ ID NO: 133; (Institut fur Bioanalytik (IBA) GmbH, Germany; see, for example, U.S. Patent Publication No. 7,981,632), Softag 1 (SLAELLNAGLGGS; SEQ ID NO: 134), Softag 3 (TQDPSRVG; SEQ ID NO: 135), and V5 tag (GKPIPNPLLGLDST; SEQ ID NO: 136). In certain embodiments, the tag cassette includes a fusion tag cassette such as 3XFLAG. In certain embodiments, the 3XFLAG includes the sequence shown in SEQ ID NO: 61.
[0070] The sequences of the aforementioned functional molecules have been published, for example, lactase (e.g., GenBank: EAX11622.1), lipase (e.g., GenBank: AAA60129.1), helicase (e.g., GenBank: AMD82207.1), amylase (e.g., GenBank: AAA51724.1), α-glucosidase (e.g., GenBank: ABI53718.1), transcription factor SP1 (e.g., UniProtKB / Swiss-Prot: P08047.3), transcription factor AP-1 (e.g., NP_002219.1), heat shock factor protein 1 (e.g., UniProtK B / Swiss-Prot:Q00613.1), CCAAT / enhancer-binding protein (C / EBP) beta isoform a (e.g. NP_005185.2), Oct-1 (e.g. UniProtKB / Swiss-Prot:P14859.2), TGF-β (e.g. GenBank:CAF02096.2), glial cell line-derived neurotrophic factor (GDNF) (e.g. NP_001177397.1), platelet-derived growth factor receptor (e.g. GenBank:AAA60049.1), epidermal growth factor receptor (e.g. GenBank:CAA25 240.1), vascular endothelial growth factor receptor (e.g., GenBank: AAC16449.2), interleukin-8 receptor α (e.g., GenBank: AAB59436.1), caveolin (e.g., GenBank: CAA79476.1), dynamin (e.g., GenBank: AAA88025.1), clathrin heavy chain 1 isoform 1 (e.g., NP_004850.1), clathrin heavy chain 2 isoform 1 (e.g., NP_009029.3), clathrin light chain A isoform a (e.g., NP_001824.1), clathrin light chain B isoform a (e.g. NP_001825.1), ras-related protein Rab-4A isoform 1 (e.g. NP_004569.2), ras-related protein Rab-11A (e.g. UniProtKB / Swiss-Prot:P62491.3), platelet-derived growth factor (e.g. GenBank:AAA60552.1), transforming growth factor beta 3 (e.g. GenBank:AAA61161.1), nerve growth factor (e.g. GenBank:CAA37703.1), EGF (e.g. GenBank:CAA34902.2), cocaine-amphetamine regulated transcript (A chain) (e.g. PDB:1HY9_A), protachykinin-1 (e.g. UniProtKB-P20366), oxytocin neurophysin 1 (e.g. UniProtKB-P01178), somatostatin (e.g. GenBank:AAH32625.1), genetically encoded green calcium indicator NTnC (A chain) [synthetic construct] (e.g. PDB:5MWC_A), calcium indicator TN-XXL [synthetic construct] (e.g. GenBank:ACF93133.1), BRET-based auto-luminescent calcium indicator [synthetic construct] (e.g., GenBank: ADF42668.1), calcium indicator protein OeNL(Ca2+)-18μ [synthetic construct] (e.g., GenBank: BBB18812.1), myosin light chain kinase, green fluorescent protein, calmodulin chimera (A chain) [synthetic construct] (e.g., PDB: 3EKJ_A), channelopsin 1 (e.g., UniProtKB-F8UVI5), channelopsin 1 (e.g., GenBank: AER58217.1), channelrhodopsin 2 (e.g., UniProtKB-B4Y105), channelrhodopsin 2 [synthetic construct] (e.g., GenBank: ABO64386.1), CRISPR-associated proteins (Cas) (e.g., GenBank: AKG27598.1), Cas9 [synthetic construct] (e.g., GenBank: AST09977.1), CRISPR-associated endonuclease Cpf1 (e.g., UniProtKB / Swiss-Prot: U2UMQ6.1), ribonuclease Examples of such proteins include ribonuclease 4 or ribonuclease L (e.g. UniProtKB / Swiss-Prot:Q05823.2), deoxyribonuclease IIβ (e.g. GenBank:AAF76893.1), sodium channel protein type 1 subunit α (e.g. UniProtKB-P35498), member 2 of the voltage-gated potassium channel subfamily KQT (e.g. UniProtKB-O43526) and voltage-gated L-type calcium channel subunit α-1C (e.g. UniProtKB-Q13936).
[0071] Further effector elements include Cre, iCre, dgCre, FlpO and tTA2. iCre refers to codon-improved Cre. dgCre is a GFP / Cre recombinase fusion gene enhanced by the N-terminal fusion of the first 159 amino acids of the dihydrofolate reductase gene (DHFR or folA) of the E. coli K12 chromosome, which has a G67S mutation and a destabilization domain mutation R12Y / Y100I upon recombination. FlpO is a codon-optimized form of FLPe, which significantly improves protein expression and FRT recombination efficiency in mouse cells. The FLP / FRT system is widely used for gene expression, as is the Cre / LoxP system (the generation of conditional knockout mice using the FLP / FRT system is also widely practiced). tTA2 refers to tetracycline transactivator.
[0072] Representative expressible elements include expression products that do not include effector elements, such as non-functional or defective proteins. In certain embodiments, such expressible elements can be used to perform methods for testing the effect of their corresponding functional molecules. In certain embodiments, the expressible elements are non-functional or defective due to recombinant mutations that abolish their function. In these aspects, the non-expressible elements are as similar in structure as possible to their corresponding functional molecules.
[0073] A representative self-cleaving peptide is the 2A peptide, which allows two proteins to be produced from one mRNA. The 2A sequence is a short sequence (e.g., 20 amino acids long) and is often used in size-restricted constructs. Specific examples include P2A, T2A, E2A, and F2A. In certain embodiments, the artificial expression construct comprises an internal ribosome entry site (IRES) sequence. The IRES can initiate ribosome translation from a second internal site on the mRNA molecule, allowing two proteins to be produced from one mRNA.
[0074] The artificial expression construct may encode nuclear transport proteins such as histone H1, histone H2A, histone H2B, histone H3, histone H4, histone-like proteins HPhA, H2B*.
[0075] Coding sequences encoding the molecules (e.g., RNA and proteins) described herein can be obtained from publicly available databases and publications. The coding sequences may further contain various sequence polymorphisms, mutations and / or sequence variants, and such changes do not affect the function of the encoded molecule. "Encode" refers to the property of a nucleic acid sequence, such as a vector, plasmid, gene, cDNA, mRNA, etc., to function as a template for the synthesis of other molecules, such as proteins.
[0076] The term "gene" may include not only coding sequences, but also regulatory regions such as promoters, enhancers, insulators and / or post-transcriptional regulatory elements (e.g., termination regions). In addition, the term may include any introns and other DNA sequences spliced from the mRNA transcript, as well as variants resulting from alternative splice sites. These sequences may further include sequences or degenerate codons of a reference sequence that may be introduced to confer codon preference in a particular type of organism or cell.
[0077] The promoter may be a general promoter, a tissue-specific promoter, a cell-specific promoter, and / or a cytoplasm-specific promoter. The promoter may be a strong promoter, a weak promoter, a constitutive expression promoter, and / or an inducible promoter. An inducible promoter induces expression in response to a specific condition, signal, or cellular event. For example, the promoter may be an inducible promoter that requires a specific ligand, small molecule, transcription factor, or hormone protein to induce transcription from the promoter. Specific examples of promoters include minBglobin (also called minBGprom), CMV, minCMV, minCMV* (minCMV* is minCMV with the SacI restriction site removed), minRho, minRho* (minRho* is minRho with the SacI restriction site removed), SV40 immediate early promoter, Hsp68 minimal promoter (proHSP68), and Rous sarcoma virus (RSV) long terminal repeat (LTR) promoter. A minimal promoter does not have the activity of inducing gene expression by itself, but when linked to an enhancer element nearby, it is activated and can induce gene expression.
[0078] In certain embodiments, the expression construct is provided in a vector. A "vector" refers to a nucleic acid molecule capable of transferring or transporting another nucleic acid molecule, such as an expression construct. The transferred nucleic acid is usually linked to, e.g., inserted into, the nucleic acid molecule of the vector. The vector may contain a sequence that induces autonomous replication in the cell or may contain a sequence that allows integration into the DNA of the host cell. Useful vectors include, for example, plasmids (e.g., DNA and RNA plasmids), transposons, cosmids, bacterial artificial chromosomes, and viral vectors.
[0079] The term "viral vector" is used broadly to refer to a nucleic acid molecule that contains components derived from a virus that facilitate the transfer and expression of non-natural nucleic acid molecules in cells. An "adeno-associated viral vector" refers to a viral vector or plasmid that contains structural and functional genetic elements or portions thereof that are primarily derived from AAV. A "retroviral vector" refers to a viral vector or plasmid that contains structural and functional genetic elements or portions thereof that are primarily derived from a retrovirus. A "lentiviral vector" refers to a viral vector or plasmid that contains structural and functional genetic elements or portions thereof that are primarily derived from a lentivirus or the like. A "hybrid vector" refers to a vector that contains structural and functional genetic elements from two or more viruses.
[0080] "Adenoviral vector" refers to a construct that contains sufficient adenoviral sequences (a) to facilitate packaging of an artificial expression construct, and (b) to express a coding sequence cloned in the sense or antisense orientation. Recombinant adenoviral vectors include genetically engineered forms of adenovirus. The genetic makeup of adenovirus is a 36 kb linear double-stranded DNA virus, allowing replacement of large portions of adenoviral DNA with up to 7 kb of foreign sequence. Adenoviral DNA can replicate episomally without causing genotoxicity, and thus, unlike retroviruses, is not integrated into chromosomes upon infection of a host cell with adenovirus. Furthermore, adenoviruses are structurally stable, and no genome rearrangements have been detected after extensive amplification.
[0081] Adenoviruses are particularly suitable for use as gene transfer vectors due to their moderate genome size, ease of manipulation, high titer, wide target cell range and high infectivity. Both ends of the adenovirus genome contain inverted repeats (ITRs) of 100-200 base pairs in length, which are cis elements required for viral DNA replication and packaging. The early (E) and late (L) regions of the adenovirus genome contain various transcription units that are divided by the initiation of viral DNA replication. The E1 region (E1A and E1B) encodes proteins responsible for regulating the transcription of the adenoviral genome and several cellular genes. Expression of the E2 region (E2A and E2B) results in the synthesis of proteins for viral DNA replication. These proteins are involved in DNA replication, expression of late genes and shut-off of host cell protein biosynthesis. Late gene products, including the majority of adenovirus capsid proteins, are expressed only after significant processing of a single primary transcript driven by the major late promoter (MLP). The MLP is particularly efficient at the late stages of infection, when all mRNAs driven by this promoter have a tripartite 5'-leader (TPL) sequence and are preferentially selected by the mRNA for translation.
[0082] Other than the requirement that the adenoviral vector be replication-deficient or at least conditionally-deficient, the characteristics of the adenoviral vector are not believed to be critical to the successful implementation of certain embodiments disclosed herein. The adenovirus may be of any of the 42 known serotypes or subgenuses A-F. In certain embodiments, adenovirus of serotype 5 of the C subgenus is preferred as the starting material for obtaining a conditionally replication-deficient adenoviral vector for use in certain embodiments, since type 5 adenovirus is a human adenovirus with a large amount of known biochemical and genetic information and has been used historically in the majority of constructions using adenoviruses as vectors.
[0083] As described herein, the vector is generally replication-defective and lacks the E1 region of adenovirus. Therefore, it is most convenient to introduce the polynucleotide encoding the gene of interest into the position where the coding sequence of the E1 region has been deleted. However, the insertion position of the construct within the adenovirus sequence is not critical. The polynucleotide encoding the gene of interest may be inserted into the deleted E3 region of the E3 replacement vector, or into the E4 region, and the defect in the E4 region is complemented by a helper cell line or a helper virus.
[0084] Adeno-associated virus (AAV) is a parvovirus found as a contaminant of adenovirus stocks. AAV is a ubiquitous virus (85% of the US population has anti-AAV antibodies) and does not cause disease. Furthermore, AAV is classified as a dependovirus because its replication depends on the presence of a helper virus (e.g., adenovirus). Various serotypes have been isolated, of which AAV-2 is the most extensively characterized. AAV has a single-stranded linear DNA that is packaged with the capsid proteins VP1, VP2, and VP3 to form icosahedral virions with a diameter of 20-24 nm.
[0085] The length of AAV DNA is 4.7 kilobases. AAV DNA contains two open reading frames, flanked by two ITRs. There are two main genes in the AAV genome: rep and cap. The rep gene codes for the proteins responsible for AAV viral replication, and the cap gene codes for the capsid proteins VP1-3. Each ITR forms a T-shaped hairpin structure. These terminal repeats are the only cis components of AAV required for chromosomal integration. Thus, AAV can be used as a vector in which all viral coding sequences can be removed and replaced with gene cassettes for delivery. Three AAV viral promoters have been identified and named p5, p19, and p40, respectively, based on their map location. Transcription from p5 and p19 results in the production of the rep protein, and transcription from p40 produces the capsid protein.
[0086] AAV is outstanding for use in the present disclosure because it has a good safety profile and can be expressed in target cell populations by modifying capsid and genome.scAAV refers to self-complementary AAV.pAAV refers to plasmid adeno-associated virus.rAAV refers to recombinant adeno-associated virus.
[0087] Other viral vectors may also be used, for example vectors derived from viruses such as vaccinia virus, poliovirus or herpes virus, which offer beneficial characteristics for a variety of mammalian cells.
[0088] Retroviruses are commonly used tools for gene delivery. A retrovirus is an RNA virus whose genomic RNA is reverse transcribed to produce a double-stranded linear DNA copy, which is then covalently integrated into the host genome. Once integrated into the host genome, the retrovirus is called a provirus. The provirus functions as a template for RNA polymerase II to induce the expression of RNA molecules that code for structural proteins and enzymes required for the production of new viral particles.
[0089] Examples of retroviruses suitable for use in certain embodiments include Moloney murine leukemia virus (M-MuLV), Moloney murine sarcoma virus (MoMSV), Harvey murine sarcoma virus (HaMuSV), mouse mammary tumor virus (MuMTV), gibbon ape leukemia virus (GaLV), feline leukemia virus (FLV), spumavirus, Friend murine leukemia virus, murine stem cell virus (MSCV), Rous sarcoma virus (RSV), and lentiviruses.
[0090] "Lentivirus" refers to the complex retrovirus group (or complex retrovirus genus). Examples of lentiviruses include HIV (human immunodeficiency virus; including HIV type 1 and HIV type 2); Visna-Maedi virus (VMV); Caprine arthritis-encephalomyelitis virus (CAEV); Equine infectious anemia virus (EIAV); Feline immunodeficiency virus (FIV); Bovine immunodeficiency virus (BIV); and Simian immunodeficiency virus (SIV). In certain embodiments, a vector backbone based on HIV (i.e., HIV cis-acting sequence elements) can be used.
[0091] In some types of vectors, safety can be improved by replacing the U3 region of the 5'LTR, which induces the transcription of the viral genome in the production of viral particles, with a heterologous promoter. Examples of heterologous promoters that can be used for this purpose include, for example, Simian Virus 40 (SV40) (e.g., early or late) promoter, Cytomegalovirus (CMV) (e.g., immediate early) promoter, Moloney Murine Leukemia Virus (MoMLV) promoter, Rous Sarcoma Virus (RSV) promoter, and Herpes Simplex Virus (HSV) (thymidine kinase) promoter. Conventional promoters can induce high levels of transcription independent of Tat. Replacement of the U3 region with a heterologous promoter removes the complete U3 sequence from the virus production system, reducing the possibility of recombination that produces a replicable virus. In certain embodiments, a heterologous promoter has the additional advantage of being able to control the way the viral genome is transcribed. For example, the heterologous promoter can be an inducible promoter, such that the entire viral genome or a portion thereof is transcribed only when an inducer is present. Inducers include one or more compounds or physiological conditions, such as the culture temperature or pH of the host cell.
[0092] In certain embodiments, the viral vector comprises a TAR element. "TAR" refers to the "transactivation response" gene element present in the R region of the LTR of lentivirus. This element interacts with the transactivator (tat) gene element of lentivirus to enhance viral replication. However, this element is not required in the embodiment that replaces the U3 region of 5'LTR with a heterologous promoter.
[0093] The "R region" refers to the region in the retroviral LTR from the start of the cap site (i.e., the transcription start site) to just before the start of the poly(A) tail. The R region is also defined as the region between the U3 and U5 regions. The R region plays a role in moving nascent DNA from one end of the genome to the other during reverse transcription.
[0094] In certain embodiments, the expression of heterologous sequences in viral vectors can be increased by incorporating post-transcriptional regulatory elements and efficient polyadenylation sites into the viral vector, and a transcription termination signal may also be incorporated into the viral vector. Various post-transcriptional regulatory elements can increase the expression of heterologous nucleic acids. Examples of post-transcriptional regulatory elements include the Woodchuck Hepatitis Virus post-transcriptional regulatory element (WPRE; Zufferey et al., 1999, J. Virol., 73:2886); the Hepatitis B virus post-transcriptional regulatory element (HPRE) (Smith et al., Nucleic Acids Res. 26(21):4818-4827, 1998); and other post-transcriptional regulatory elements (Liu et al., 1995, Genes Dev., 9:1766). In certain embodiments, the vector includes a post-transcriptional regulatory element such as a WPRE or HPRE. In certain embodiments, the vector lacks or does not include a post-transcriptional regulatory element such as a WPRE or HPRE.
[0095] Expression of heterologous genes can be increased by elements capable of inducing efficient transcription termination and polyadenylation of heterologous nucleic acid transcripts. Transcription termination signals are usually found downstream of polyadenylation signals. In certain embodiments, vectors contain a polyadenylation signal at the 3' end of a polynucleotide encoding an expressed molecule (e.g., a protein). "Poly(A) site" or "poly(A) sequence" refers to a DNA sequence that induces both transcription termination and polyadenylation of a nascent RNA transcript transcribed by RNA polymerase II. Polyadenylation sequences can improve mRNA stability by adding a poly(A) tail to the 3' end of a coding sequence, thereby contributing to improved translation efficiency. In certain embodiments, BGHpA, hGHpA, or SV40pA may be utilized. In certain embodiments, a preferred embodiment of an expression construct includes a terminator element. Terminator elements can increase the amount of transcription and minimize transcription from the construct to another plasmid sequence by read-through.
[0096] In certain embodiments, the viral vector further comprises one or more insulator elements. The insulator elements may protect sequences expressed from the viral vector, such as effector elements and expressible elements, from integration site effects. Integration site effects occur through cis-acting elements in genomic DNA, meaning that the imported sequence is either expressed or not expressed (i.e., position effects; see, for example, Burgess-Beusse et al., PNAS., USA, 99:16433, 2002; and Zhan et al., Hum. Genet., 109:471, 2001). In certain embodiments, the viral import vector comprises one or more insulator elements in the 3'LTR, and upon provirus integration into the host genome, this insulator is integrated into both the 5'LTR and the 3'LTR during the replication of the 3'LTR. Insulators suitable for use in certain embodiments include the chicken β-globin insulator (see Chung et al., Cell 74:505, 1993; Chung et al., PNAS USA 94:575, 1997; and Bell et al., Cell 98:387, 1999), the SP10 insulator (Abhyankar et al., JBC 282:36143, 2007), or other small CTCF recognition sequences that function as enhancer-blocking insulators (Liu et al., Nature Biotechnology, 33:198, 2015).
[0097] In addition to the above, various types of suitable expression vectors are also known to those skilled in the art. These known expression vectors include commercially available expression vectors designed for general recombinant manipulation, such as plasmids that contain one or more reporter genes and the regulatory elements required for the reporter gene to be expressed in cells. Many vectors are commercially available, such as from Invitrogen, Stratagene, Clontech, etc., and are described in various accompanying guidebooks. In certain embodiments, suitable expression vectors include any plasmid, cosmid, or phage construct that can express the encoded gene in mammalian cells, such as the pUC plasmid system and the Bluescript plasmid system.
[0098] Particular embodiments of the vectors disclosed herein include those set forth in the table below. [Table 1]
[0099] Subcomponent sequences within a larger vector sequence can be readily identified by one of skill in the art based on the disclosure herein (see FIG. 24). The nucleotides between the identifiable sequences in the table above are restriction enzyme recognition sites used in construct assembly (cloning) and, in some cases, additional nucleotides with no identifiable function. These segments of the complete vector sequence can be adjusted using different cloning techniques and / or different vectors. Short palindromic sequences of six bases usually represent vector construction artifacts that are not critical to the function of the vector.
[0100] In certain embodiments, a vector (e.g., AAV) is selected that has a capsid that can cross the blood-brain barrier (BBB). In certain embodiments, the vector is engineered to contain a capsid that crosses the blood-brain barrier. Examples of AAVs with viral capsids that can cross the blood-brain barrier include AAV9 (Gombash et al., Front Mol Neurosci. 2014; 7:81), AAVrh.10 (Yang, et al., Mol Ther. 2014; 22(7): 1299-1309), AAV1R6, AAV1R7 (Albright et al., Mol Ther. 2018; 26(2): 510), rAAVrh.8 (Yang et al., supra), AAV-BR1 (Marchio et al., EMBO Mol Med. 2016; 8(6): 592), AAV-PHP.S (Chan et al., Nat Neurosci. 2017; 20(8): 1172), and AAV-PHP.B (Deverman et al., Nat Biotechnol. 2016; 34(2): 204), AAV-PPS (Chen et al., Nat Med. 2009; 15: 1215) and PHP.eB. In certain embodiments, the capsid of PHP.eB differs from that of AAV9 in that the amino acid residues from position 586 onwards, S-AQ-A (SEQ ID NO: 119), are changed to S-DGTLAVPFK-A (SEQ ID NO: 120) when comparing AAV9 as a reference. In certain embodiments, PHP.eb refers to the sequence of SEQ ID NO: 74.
[0101] AAV9 is a naturally occurring AAV serotype that, unlike many other naturally occurring serotypes, is able to cross the blood-brain barrier (BBB) upon intravenous injection. AAV9 transduces a wide area of the central nervous system (CNS), allowing for minimally invasive therapeutic approaches (Naso et al., BioDrugs. 2017; 31(4): 317). Such cases have been reported, for example, in connection with AveXis' clinical trial for the treatment of spinal muscular atrophy (SMA) syndrome (AVXS-101, NCT03505099) and the clinical trial for the treatment of CLN3-associated neuronal ceroid lipofuscinosis (NCT03770572).
[0102] AAVrh.10 is an AAV originally isolated from rhesus macaques that has weak human seroreactivity compared to other common serotypes used in gene delivery applications (Selot et al., Front Pharmacol. 2017; 8: 441) and has been evaluated in several clinical trials (LYS-SAF302, LYSOGENE and NCT03612869).
[0103] AAV1R6 and AAV1R7 are two variants isolated from a library of chimeric AAV vectors in which the capsid domain of AAVrh.10 is replaced by that of AAV1, which retain the ability to cross the BBB and transduce the central nervous system but show significantly reduced transduction of the liver and vascular endothelium.
[0104] rAAVrh.8, also an AAV isolated from rhesus macaques, showed widespread transduction of glial and neuronal cells in clinically relevant areas following peripheral administration, with reduced peripheral tissue tropism compared to other vectors.
[0105] AAV-BR1 is an AAV2 variant that displays the NRGTEWD epitope (SEQ ID NO: 121) and was isolated during in vivo screening of a random AAV-display peptide library. AAV-BR1 exhibits high specificity with high transgene expression in the brain and minimal off-target affinity, including to the liver (Korbelin et al., EMBO Mol Med. 2016; 8(6): 609).
[0106] AAV-PHP.S (Addgene, Watertown, MA) is a variant of AAV9 generated by the CREATE method that encodes the 7-mer sequence QAVRTSL (sequence number 122) and transduces neurons of the enteric nervous system and potently transduces peripheral sensory afferents that project to the spinal cord and brainstem.
[0107] AAV-PHP.B (Addgene, Watertown, MA) is a variant of AAV9 generated by the CREATE method that encodes the 7-mer sequence TLAVPFK (SEQ ID NO: 123). AAV-PHP.B transfers genes throughout the central nervous system more efficiently than AAV9, transducing a large proportion of astrocytes and neurons in multiple central nervous system regions.
[0108] AAV-PPS is an AAV2 variant created by inserting the DSPAHPS epitope (SEQ ID NO: 124) into the capsid of AAV2, and exhibits dramatically improved brain tropism compared to AAV2.
[0109] For more information regarding capsids crossing the blood-brain barrier, see Chan et al., Nat. Neurosci. 2017 Aug: 20(8): 1172-1179.
[0110] (ii) Composition for Administration The artificial expression constructs and vectors (herein referred to as bioactive components) of the present disclosure can be formulated with carriers suitable for administration to cells, tissue slices, animals (e.g., mice and non-human primates), or humans. The bioactive components contained in the compositions described herein can be prepared in a neutral form, can be prepared as a free base, or can be prepared as a pharmacologically acceptable salt.
[0111] Pharmaceutically acceptable salts include the acid addition salts (formed from the free amino groups of the protein) which are formed with inorganic acids such as, for example, hydrochloric or phosphoric acid, or organic acids such as acetic, oxalic, tartaric, mandelic, etc. Also, salts formed with the free carboxyl groups can be derived from inorganic bases such as, for example, sodium, potassium, ammonium, calcium, or ferric hydroxides, or organic bases such as isopropylamine, trimethylamine, histidine, procaine, and the like.
[0112] Carriers for biologically active ingredients include solvents, dispersion media, vehicles, coating agents, diluents, isotonicity agents, absorption delaying agents, buffers, solutions, suspensions, colloids, etc. The use of such carriers for biologically active ingredients is well known in the art. Except insofar as a conventional media or agent is incompatible with the biologically active ingredients of the present invention, any conventional media or agent can be used in combination with the compositions described herein.
[0113] A "pharmacologically acceptable carrier" refers to a carrier that does not produce an allergic or similar untoward reaction when administered to a human, and in certain embodiments, when administered intravenously (e.g., into the retro-orbital plexus).
[0114] In certain embodiments, compositions of the invention can be formulated for intravenous, intraparenchymal, intraocular, intravitreal, parenteral, subcutaneous, intraventricular, intramuscular, intrathecal, intraspinal, intraperitoneal, oral or nasal inhalation, or for direct injection or administration into one or more cells, tissues or organs.
[0115] The compositions of the present invention may comprise liposomes, lipids, lipid complexes, microspheres, microparticles, nanospheres and / or nanoparticles.
[0116] The formation of liposomes and their use are widely known to those skilled in the art. Liposomes have been developed to improve serum stability and blood half-life (see, for example, U.S. Patent No. 5,741,516). In addition, various methods have been reported for using liposomes and liposome-like preparations as potential drug carriers (see, for example, U.S. Patent No. 5,567,434; U.S. Patent No. 5,552,157; U.S. Patent No. 5,565,213; U.S. Patent No. 5,738,868; and U.S. Patent No. 5,795,587).
[0117] The present disclosure also provides pharma- ceutically acceptable nanocapsule formulations of the bioactive ingredients of the present invention. In general, nanocapsule formulations can encapsulate compounds in a stable and reproducible manner (Quintanar-Guerrero et al., Drug Dev Ind Pharm 24(12):1113-1128, 1998; Quintanar-Guerrero et al., Pharm Res. 15(7):1056-1062, 1998; Quintanar-Guerrero et al., J. Microencapsul. 15(1):107-119, 1998; Douglas et al., Crit Rev Ther Drug Carrier Syst 3(3):233-261, 1987). To avoid side effects caused by large amounts of macromolecules being taken up into cells, such ultrafine particles can be designed using polymers that can be degraded in vivo. Biodegradable polyalkyl cyanoacrylate nanoparticles that meet such requirements are also envisioned for use in the present disclosure.Such microparticles can be easily prepared, and are described, for example, in Couvreur et al., J Pharm Sci 69(2):199-202, 1980; Couvreur et al., Crit Rev Ther Drug Carrier Syst. 5(1)1-20, 1988; zur Muhlen et al., Eur J Pharm Biopharm, 45(2):149-155, 1998; Zambaux et al., J Control Release 50(1-3):31-40, 1998; and U.S. Patent No. 5,145,684.
[0118] Injectable compositions include sterile aqueous solutions or dispersions and sterile powders for extemporaneous preparation of sterile injectable solutions or dispersions (US Pat. No. 5,466,468). Injectable compositions delivered by injection are in the form of a sterile fluid to the extent that they can be delivered using a syringe. In certain embodiments, injectable compositions are usually stable during manufacturing and storage, and may contain one or more preservative compounds to prevent the contaminating action of microorganisms such as bacteria and fungi. The carrier may be a solvent or dispersion medium, which may include, for example, water, ethanol, polyol (e.g., glycerol, propylene glycol, liquid polyethylene glycol, and the like), and suitable mixtures thereof, and / or vegetable oils. To maintain proper fluidity, a coating agent such as lecithin may be used, or in the case of dispersions, the particle size may be maintained to the required size, and / or a surfactant may be used. To prevent the action of microorganisms, various antibacterial and / or antifungal agents may be used, such as, for example, parabens, chlorobutanol, phenol, sorbic acid, thimerosal, and the like. In various embodiments, the injectable composition contains an isotonic agent, such as, for example, sugars or sodium chloride. Prolonged absorption of the injectable composition can be achieved by incorporating an agent that delays absorption, such as, for example, aluminum monostearate or gelatin, into the injectable composition. If necessary, an appropriate buffer may be added to the injectable composition, and the diluted liquid is first made isotonic with sufficient saline or glucose.
[0119] Dispersions may also be prepared in glycerol, liquid polyethylene glycols or mixtures thereof, or oils. As described herein, under ordinary conditions of storage and use, such preparations may contain a preservative to prevent the growth of microorganisms.
[0120] Sterile compositions can be prepared by mixing the physiologically active ingredient with any other ingredients (e.g., those mentioned above) in an appropriate amount of solvent and sterilizing by filtration. Dispersions are usually prepared by dispersing various sterilized physiologically active ingredients in a sterile solvent containing a basic dispersion medium and other necessary ingredients (e.g., those mentioned above). In the case of sterile powders for preparing sterile injectable solutions, a preferred method is to sterilize a solution containing the physiologically active ingredient and other desired ingredients in advance by filtration, and then vacuum drying or freeze-drying the solution to prepare a powder containing the physiologically active ingredient and other desired ingredients.
[0121] Oral compositions may be in liquid form, such as solutions, syrups or suspensions, and may be provided as pharmaceutical products to be reconstituted with water or other suitable solvents before use. Such liquid preparations may be prepared by conventional methods using pharmaceutically acceptable additives, such as suspending agents (e.g., sorbitol syrup, cellulose derivatives or hydrogenated edible fats); emulsifying agents (e.g., lecithin or gum arabic); non-aqueous solvents (e.g., almond oil, ester oils or fractionated vegetable oils); and preservatives (e.g., methyl p-hydroxybenzoate, propyl p-hydroxybenzoate or sorbic acid). The composition of the present invention may be prepared, for example, in the form of tablets or capsules, by conventional methods using pharma- ceutically acceptable excipients, such as, for example, binders (e.g., pregelatinized corn starch, polyvinylpyrrolidone or hydroxypropylmethylcellulose); fillers (e.g., lactose, microcrystalline cellulose or calcium hydrogen phosphate); lubricants (e.g., magnesium stearate, talc or silica); disintegrants (e.g., potato starch or sodium starch glycolate); and wetting agents (e.g., sodium lauryl sulfate). Tablets may be coated by methods known in the art.
[0122] Compositions for inhalation can be delivered in the form of an aerosol spray from a pressurized pack or nebulizer using a suitable propellant, such as, for example, dichlorodifluoromethane, trichlorofluoromethane, dichlorotetrafluoroethane, carbon dioxide or other suitable gas. In the case of a pressurized aerosol, the dosage unit may be determined by providing a valve to deliver a metered amount. The compositions may be formulated into capsules or cartridges (e.g., gelatin capsules or cartridges) for use in an inhaler or nebulizer, which contain a powder mix of the compositions described herein and a suitable powder base, such as lactose or starch.
[0123] Additionally, compositions of the present invention include microchip devices (U.S. Pat. No. 5,797,898), ophthalmic formulations (Bourlais et al., Prog Retin Eye Res, 17(1):33-58, 1998), transdermal matrices (U.S. Pat. Nos. 5,770,219 and 5,783,208) and feedback controlled delivery (U.S. Pat. No. 5,697,899).
[0124] Supplementary active ingredients can also be included in the compositions of the present invention.
[0125] Typically, the compositions of the present invention may contain at least 0.1% or more of the physiologically active ingredient, but it goes without saying that the percentage of the physiologically active ingredient may vary and may conveniently range from 1% or 2% to 70% or 80% or more, or may range from 0.5 to 99%, based on the total weight or volume of the composition of the present invention. Of course, the amount of physiologically beneficial physiologically active ingredient in each composition may be adjusted so that a given unit dose of said compound provides an appropriate dosage. Factors such as solubility, bioavailability, biological half-life, route of administration, shelf life of the product, and other pharmacological considerations will be considered by those skilled in the art responsible for preparing pharmaceutical formulations, and therefore various compositions and dosages may be desirable.
[0126] In certain embodiments, for human administration, compositions of the invention should meet sterility, pyrogenicity, general safety and purity standards as required by the U.S. Food and Drug Administration (FDA) or other relevant regulatory authorities in other countries.
[0127] (iii) Cell lines containing the artificial expression constructs The present disclosure includes cells comprising the artificial expression constructs described herein. Cells transformed with the artificial expression constructs can be used for a variety of purposes, such as neuronal cell studies, evaluation of functional and / or non-functional proteins, drug screening to evaluate the regulatory properties of enhancers, etc.
[0128] While a variety of host cell lines can be used, in certain embodiments the host cell is a mammalian cell. In certain embodiments, the artificial expression construct is selected from the group consisting of eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_830h, eHGT_831h, eHGT_832h, eHGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHGT_577h, 3xcore3_eHGT_ 577h, 3xcore2_eHGT_606h, 3xcore3_eHGT_606h, core4_eHGT_577h, core6_eHGT_606h, 3xcore_eHGT_121h, eHGT_590 m, eHGT_976h, MGT_E117, MGT_E118, MGT_E119, MGT_E120, MGT_E121, 3xCore2-eHGT_367h, eHGT_359h, eHGT_479m, eHG T_453m, eHGT_140h, eHGT_356h, eHGT_128h, eHGT_369h and eHGT_710m, and / or CN2415, CN2416, CN2417, CN2418, CN2436, CN3000, CN3001, CN3003, CN3004, CN3005, CN3007, CN3009, AiP1335, AiP1336, AiP1337, AiP1338, AiP1339, CN255 The enhancer and / or vector sequence may be selected from the group consisting of CN1633, CN2043, CN1621, CN2216, CN2717, CN3639, CN3050, CN3051, CN3056, CN3057, CN4001, CN4003, CN2786, CN2840, CN3460 and CN2650, and the host cell line is a human cell, a primate cell or a mouse cell. Additionally, cell lines that can be used for gene transfer in the present disclosure include primary cell lines derived from living tissues such as rat or mouse brain, and organotypic cell cultures such as brain slices from animals, such as rat, mouse, non-human primate or human neurosurgical tissues.The PC12 cell line (available from the American Type Culture Collection (ATCC), Manassas, VA) has been shown to express many neuronal marker proteins in response to nerve growth factor (NGF). The PC12 cell line is believed to be a neuronal cell line and is applicable for use in the present disclosure. JAR cells (available from the ATCC) are a platelet-derived cell line that express several neuronal genes, such as the serotonin transporter gene, and may be used in the embodiments described herein.
[0129] WO91 / 13150 describes various cell lines, including neuronal cell lines, and methods for their production. Similarly, WO97 / 39117 describes neuronal cell lines and methods for the production of such cell lines. The neuronal cell lines disclosed in these patent applications are applicable for use in the present disclosure.
[0130] In certain embodiments, the term "neuronal cell" is used to describe any neuronal cell, related to neuronal cells, or including neuronal cells. Neuronal cells are defined by the characteristic of having an axon and a dendrite. The term "neuron-specific" refers to something that is found in a neuronal cell or cells derived therefrom, but is not found or is substantially absent in non-neuronal cells (e.g., glial cells such as astrocytes and oligodendrocytes); or activity that occurs in a neuronal cell or cells derived therefrom, but is not found or is substantially absent in non-neuronal cells (e.g., glial cells such as astrocytes and oligodendrocytes).
[0131] In certain embodiments, non-neuronal cell lines such as mouse embryonic stem cells may be used. Cultured mouse embryonic stem cells can be transiently transfected with plasmid constructs to analyze the expression of gene constructs. Mouse embryonic stem cells are pluripotent undifferentiated cells. Mouse embryonic stem cells can be maintained in an undifferentiated state by leukemia inhibitory factor (LIF). Mouse embryonic stem cells can be induced to differentiate by removing LIF. Mouse embryonic stem cells form various types of differentiated cells in culture. Differentiation of mouse embryonic stem cells occurs through the expression of tissue-specific transcription factors, which allows the evaluation of the function of enhancer sequences (see, for example, Fiskerstrand et al., FEBS Lett 458: 171-174, 1999).
[0132] The method of differentiating stem cells into neural cells includes replacing the stem cell culture medium with a medium containing basic fibroblast growth factor (bFGF), heparin, N2 supplement (e.g., transferrin, insulin, progesterone, putrescine and selenite), laminin and polyornithine. The method of producing myelinating oligodendrocytes from stem cells is described in Hu, et al., 2009, Nat. Protoc. 4:1614-22. Bibel, et al., 2007, Nat. Protoc. 2:1034-43 describes a protocol for producing glutamatergic neurons from stem cells, and Chatzi, et al., 2009, Exp. Neurol. 217:407-16 describes a procedure for producing GABAergic neurons. This procedure includes exposing stem cells to all-trans retinoic acid for 3 days. GABAergic neurons, which account for 95% of all cells, are then obtained by culturing in a serum-free neuronal induction medium such as neurobasal medium supplemented with B27, bFGF and EGF.
[0133] US Patent Publication No. 2012 / 0329714 describes the use of prolactin to increase neural stem cell numbers, and US Patent Publication No. 2012 / 0308530 describes a culture surface with amino groups that promotes differentiation of neural cells into neurons, astrocytes, and oligodendrocytes. Thus, the fate of neural stem cells can be controlled by various extracellular factors. Commonly used extracellular factors include brain-derived growth factor (BDNF; Shetty and Turner, 1998, J. Neurobiol. 35:395-425); fibroblast growth factor (bFGF; U.S. Patent No. 5,766,948; FGF-1, FGF-2); neurotrophin 3 (nt-3) and neurotrophin 4 (nt-4) (Caldwell, et al., 2001, Nat. Biotechnol. 1;19:475-9); ciliary neurotrophic factor (CNTF); BMP-2 (U.S. Pat. Nos. 5,948,428 and 6,001,654); isobutyl-3-methylxanthine; leukemia growth inhibitory factor (LIF; U.S. Pat. No. 6,103,530); somatostatin; amphiregulin; neurotrophins (e.g., cyclic adenosine monophosphate); epidermal growth factor (EGF); dexamethasone (a glucocorticoid hormone); forskolin; ligands for the GDNF family of receptors; potassium; retinoic acid (U.S. Pat. No. 6,395,546); tetanus toxoid; and transforming growth factors alpha and TGF-beta (U.S. Pat. Nos. 5,851,832 and 5,753,506).
[0134] In certain embodiments, the yeast one-hybrid system may be used to identify compounds that inhibit specific protein-DNA interactions, such as eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_830h, eHGT_831h, eHGT_832h, eHGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHGT_577h, 3xcore3_eHGT_577h, 3xcore4_eHGT_577h, 3xcore5_eHGT_577h, 3xcore6_eHGT_577h, 3xcore7_eHGT_577h, 3xcore8_eHGT_577h, 3xcore9_eHGT_577h, 3xcore10_eHGT_577h, 3xcore11_eHGT_577h, 3xcore12_eHGT_577h, 3xcore13_eHGT_577h, 3xcore14_eHGT_577h, 3xcore15_eHGT_577h, 3xcore16_eHGT_577h, 3xcore17_eHGT_577h, 3xcore18_eHGT_577h, 3xcore19_eHGT_577h, 3xcore19_eHGT_577h, 3xcore19_eHGT_577h, 3xcore10_eHGT_577h, 3xcore11_eHGT_577h, 3xcore12_eHGT_57 Examples of such transcription factors include xcore2_eHGT_606h, 3xcore3_eHGT_606h, core4_eHGT_577h, core6_eHGT_606h, 3xcore_eHGT_121h, eHGT_590m, eHGT_976h, MGT_E117, MGT_E118, MGT_E119, MGT_E120, MGT_E121, 3xCore2-eHGT_367h, eHGT_359h, eHGT_479m, eHGT_453m, eHGT_140h, eHGT_356h, eHGT_128h, eHGT_369h or eHGT_710m.
[0135] Transgenic animals are described below. Cell lines may be derived from such transgenic animals. For example, cell lines having artificial expression constructs integrated into their genomes can be obtained from primary tissue cultures derived from transgenic mice (e.g., as described below) (see, e.g., MacKenzie & Quinn, Proc Natl Acad Sci USA 96: 15251-15255, 1999).
[0136] (iv) Transgenic animals Another aspect of the disclosure is to provide a method for the preparation of a nucleotide sequence comprising administering to a subject a nucleotide sequence selected from the group consisting of eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_830h, eHGT_831h, eHGT_832h, eHGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHGT_577h, 3xcore3_eHGT_577h, 3xcore2_eHGT_606h, 3xcore3_eHGT_577h, 3xcore3 ...577h, 3xcore3_eHGT_606h, 3xcore3_ xcore3_eHGT_606h, core4_eHGT_577h, core6_eHGT_606h, 3xcore_eHGT_121h, eHGT_590m, eHGT_976h, MGT_E117, MGT_E118, MGT_E119, MGT_E120, MGT_E121, 3xCore2-eHGT_367h, eHGT_359h, eHGT_479m, eHGT_453m, eHGT_140h, eHGT_356h, eHGT_128h, eHGT_369h and / or eHGT_710m. In certain embodiments, the transgenic animal is selected from the group consisting of CN2415, CN2416, CN2417, CN2418, CN2436, CN3000, CN3001, CN3003, CN3004, CN3005, CN3007, CN3009, AiP1335, AiP1336, AiP1337, AiP1338, AiP1339, CN 2555, CN2045, CN2258, CN2251, CN1633, CN2043, CN1621, CN2216, CN2717, CN3639, CN3050, CN3051, CN3056, CN3057, CN4001, CN4003, CN2786, CN2840, CN3460 and / or CN2650 in the genome.In certain embodiments, when a non-integrating vector is utilized, the transgenic animal may be any of eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_830h, eHGT_831h, eHGT_832h, eHGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHGT_577h, 3xcore3_eHGT_577h, 3xcore2_eHGT_ 606h, 3xcore3_eHGT_606h, core4_eHGT_577h, core6_eHGT_606h, 3xcore_eHGT_121h, eHGT_590m, eHGT_976h, MGT_E117, MGT_E118, MGT_E119, MGT_E120, MGT_E121, 3xCore2-eHGT_367h, eHGT_359h, eHGT_479m, eHGT_453m, eHGT_140h, eHGT_356h, eHGT_128h, eHGT_369h, and / or eHGT_710m, and / or CN2415, CN2416, CN2417, CN2418, CN2436, CN3000, CN3001, CN3003, CN3004, CN3005, CN3007, CN3009, AiP1335, AiP1336, AiP1337, AiP1338, AiP1339, CN2555, CN2045, C The one or more cells contain an artificial expression construct comprising N2258, CN2251, CN1633, CN2043, CN1621, CN2216, CN2717, CN3639, CN3050, CN3051, CN3056, CN3057, CN4001, CN4003, CN2786, CN2840, CN3460, and / or CN2650.
[0137] A detailed description of the methods for producing transgenic animals is provided in U.S. Patent No. 4,736,866. The transgenic animals may be of any non-human species, but are preferably non-human primates (NHPs), sheep, horses, cows, pigs, goats, dogs, cats, rabbits, chickens, or rodents, such as guinea pigs, hamsters, gerbils, rats, mice, and ferrets.
[0138] In certain embodiments, by producing transgenic animals, organisms are obtained in which recombinant constructs are introduced into the same genome integration site of every cell.Therefore, the cell lines derived from such transgenic animals have consistent characteristics in that they have recombinant constructs in the same genome integration site of every cell, and therefore all of these cells undergo the same variegated position effect.In contrast, when gene is introduced into cell lines or primary cell cultures, heterologous expression of constructs is obtained.This method has the disadvantage that the expression of introduced DNA is affected by the specific genetic background of host animal.
[0139] As previously described in connection with cell lines, the artificial expression constructs of the present disclosure can be used to genetically modify mouse embryonic stem cells using techniques known in the art. Typically, the artificial expression constructs are introduced into cultured mouse embryonic stem cells. The transformed ES cells are then injected into blastocysts from a host mother, and the host embryo is reimplanted into the host mother. This procedure results in chimeric mice with tissues composed of cells derived from both embryonic stem cells present in the cultured cell line and embryonic stem cells present in the host embryo. Typically, mice are selected for isolating cultured ES cells used for gene transfer that have a different coat color from the host mouse whose embryo was injected with the transformed cells. Thus, the chimeric mice have a mixed coat color. If at least a portion of the germline tissue is derived from the genetically modified cells, the chimeric mice can then be crossed with an appropriate line to obtain offspring carrying the transgene.
[0140] In addition to the delivery methods described above, other methods of delivering artificial expression constructs to target cells or tissues or organs of animals, particularly cells, organs or tissues of mammalian vertebrates, are contemplated, including sonophoresis (e.g., ultrasound as described in U.S. Pat. No. 5,656,016); intraosseous injection (U.S. Pat. No. 5,779,708); microchip devices (U.S. Pat. No. 5,797,898); ophthalmic formulations (Bourlais et al., Prog Retin Eye Res, 17(1):33-58, 1998); transdermal matrices (U.S. Pat. Nos. 5,770,219 and 5,783,208); feedback controlled delivery (U.S. Pat. No. 5,697,899), and other delivery methods available and / or described elsewhere in this disclosure.
[0141] (v) How to use In certain embodiments, a composition comprising a bioactive ingredient described herein is administered to a subject to produce a physiological effect.
[0142] In certain embodiments, the present disclosure includes the use of the artificial expression constructs described herein to regulate the expression of a heterologous gene encoded in part or in its entirety downstream of an enhancer in a recombinant sequence. Accordingly, provided herein are methods of using the artificial expression constructs of the present disclosure in the research, study and future development of pharmaceuticals for the prevention, treatment or alleviation of symptoms of a disease, dysfunction or disorder.
[0143] Certain embodiments include a method of selectively expressing genes in selected cells described herein and inducing expression of genes in target cells by administering to a subject an artificial expression construct, the artificial expression construct being selected from the group consisting of eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_830h, eHGT_831h, eHGT_832h, eHGT_831h, eHGT_833h, eHGT_834h, eHGT_835h, eHGT_836h, eHGT_837h, eHGT_838h, eHGT_839 ... GT_832h, eHGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHGT_577h, 3xcore3_eHGT_577h, 3xcore2_eHGT_606h, 3xc ore3_eHGT_606h, core4_eHGT_577h, core6_eHGT_606h, 3xcore_eHGT_121h, eHGT_590m, eHGT_976h, MGT_E117, MGT_E118, MGT _E119, MGT_E120, MGT_E121, 3xCore2-eHGT_367h, eHGT_359h, eHGT_479m, eHGT_453m, eHGT_140h, eHGT_356h, eHGT_128h, eHGT_369h or eHGT_710m, and / or CN2415, CN2416, CN2417, CN2418, CN2436, CN3000, CN3001, CN3003, CN3004, CN3005, CN3007, CN3009, AiP1335, AiP1336, AiP1337, AiP1338, AiP1339, CN2555, CN2045, CN2258, CN2251, CN1633, CN2043, CN1621, CN2216, CN2717, CN3639, CN3050, CN3051, CN3056, CN3057, CN4001, CN4003, CN2786, CN2840, CN3460 and / or CN2650. The subject may be an isolated cell, a network of cells, a tissue section, an experimental animal, a veterinary animal or a human.
[0144] As is well known in the medical arts, the dose administered to a subject will depend on a variety of factors, such as the subject's body size, surface area and age, the particular compound administered, sex, duration and route of administration, general health, other drugs being administered concomitantly, etc. Although doses of the compounds of the present disclosure may vary, in certain embodiments, doses of the artificial expression constructs of the present disclosure may be administered in doses of 10 to 20 mg / kg. 5 ~10 100 In certain embodiments, patients receiving intravenous, intraparenchymal, intraspinal, retroorbital or intrathecal administration may receive 10 copies of the 6 ~10 22 A copy of the artificial expression construct can be injected.
[0145] "Effective amount" is the amount of a composition required to produce a desired physiological change in a subject. Effective amounts are often administered for research purposes. The effective amount disclosed herein is an amount that can produce a statistically significant effect in animal models, human studies, in vivo assays, or in vitro assays.
[0146] The dose of the expression construct and the duration of administration of such compositions are determined by those skilled in the art who have the benefit of the teachings of the present invention. However, it is contemplated that administration of an effective amount of the composition of the present disclosure may be performed by a single administration, for example, by a single injection of a sufficient number of infectious particles to confer an effect on the subject. Alternatively, in some circumstances, it may be desirable to administer multiple or sequential administrations of the artificial expression construct composition or other genetic constructs over a relatively short or long period of time, and the decision to administer such may be determined by the person overseeing the administration of such compositions. For example, the number of infectious particles administered to a mammal may be as little as 10 or more times as necessary to achieve the intended effect. 7 pieces / ml, 10 8 pieces / ml, 10 9 pieces / ml, 10 10 pieces / ml, 10 11 pieces / ml, 10 12 pieces / ml, 10 13This may be in the form of a single dose or two or more divided doses of cells / ml or more, and in certain embodiments, it may actually be desirable to administer two or more expression constructs in combination to achieve the desired effect.
[0147] In certain circumstances, it may be desirable to deliver the artificial expression construct in the form of an appropriately formulated composition as disclosed herein using a pipette or by retro-orbital injection, subcutaneous administration, intraocular administration, intravitreal administration, parenteral administration, subcutaneous administration, intravenous administration, intraparenchymal administration, intraventricular administration, intramuscular administration, intrathecal administration, intraspinal administration, intraperitoneal administration, oral administration, nasal inhalation, or direct administration or injection into one or more cells, tissues, or organs. Methods of administration may include those described in U.S. Patent No. 5,543,158; U.S. Patent No. 5,641,515, and U.S. Patent No. 5,399,363.
[0148] (vi) Kits and commercial packages The kits and commercial packages include an artificial expression construct as described herein. The artificial expression construct can be isolated. In certain embodiments, the components of the expression product can be separated from each other. In certain embodiments, the expression product is found within a vector, a viral vector, a cell, a tissue section or tissue sample, and / or a transgenic animal. Such kits may further include one or more reagents, restriction enzymes, peptides, therapeutic agents, pharmaceutical compounds, or a means for delivery of the compositions of the invention (e.g., a syringe, injection, etc.).
[0149] Embodiments of the kit or commercial package further include instructions for use of the components included in the kit or commercial package, e.g., in basic research, electrophysiological studies, neuroanatomical studies, and / or in the study and / or treatment of a disorder, disease or condition.
[0150] (vii) Representative Embodiments The following representative embodiments are set forth to illustrate specific embodiments of the present disclosure. Those skilled in the art having reference to this disclosure will appreciate that various modifications may be made to the specific embodiments disclosed herein while still achieving the same or similar results without departing from the spirit and scope of the present disclosure. 1. An artificial enhancer comprising the core region of the eHGT_367h enhancer, the core region of the eHGT_121h enhancer, the core region of the eHGT_577h enhancer and / or the core region of the eHGT_606h enhancer. 2. The artificial enhancer according to embodiment 1, wherein the core region has a sequence as set forth in SEQ ID NO:18, SEQ ID NO:26, SEQ ID NO:31, SEQ ID NO:33, SEQ ID NO:35, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39 or SEQ ID NO:41, or has a sequence that has at least 90% sequence identity with the sequence as set forth in SEQ ID NO:18, SEQ ID NO:26, SEQ ID NO:31, SEQ ID NO:33, SEQ ID NO:35, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39 or SEQ ID NO:41. 3. The artificial enhancer according to embodiment 1 or 2, wherein the core region comprises 2, 3, 4, 5, 6, 7, 8, 9 or 10 copies of the sequence shown in SEQ ID NO: 18, SEQ ID NO: 26, SEQ ID NO: 31, SEQ ID NO: 33, SEQ ID NO: 35, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39 or SEQ ID NO: 41, or a sequence having at least 90% sequence identity with the sequence shown in SEQ ID NO: 18, SEQ ID NO: 26, SEQ ID NO: 31, SEQ ID NO: 33, SEQ ID NO: 35, SEQ ID NO: 37, SEQ ID NO: 38, SEQ ID NO: 39 or SEQ ID NO: 41. 4. An artificial enhancer according to embodiment 3, comprising three copies of the sequence shown in SEQ ID NO: 18. 5. An artificial enhancer according to embodiment 3, comprising three copies of the sequence shown in SEQ ID NO: 26. 6. An artificial enhancer according to embodiment 3, comprising three copies of the sequence shown in SEQ ID NO: 31. 7. An artificial enhancer according to embodiment 3, comprising three copies of the sequence shown in SEQ ID NO: 33. 8. An artificial enhancer according to embodiment 3, comprising three copies of the sequence shown in SEQ ID NO: 35. 9. An artificial enhancer according to embodiment 3, comprising three copies of the sequence shown in SEQ ID NO: 37. 10. An artificial enhancer according to embodiment 2, having one copy of the sequence shown in SEQ ID NO: 38. 11. An artificial enhancer according to embodiment 2, having one copy of the sequence shown in SEQ ID NO: 39. 12. An artificial enhancer according to embodiment 3, comprising three copies of the sequence shown in SEQ ID NO: 41. 13. An artificial enhancer according to embodiment 4, having a sequence as set forth in SEQ ID NO: 17 or having at least 95% sequence identity with the sequence as set forth in SEQ ID NO: 17. 14. An artificial enhancer according to embodiment 5, having a sequence as set forth in SEQ ID NO: 25 or having at least 95% sequence identity with the sequence as set forth in SEQ ID NO: 25. 15. An artificial enhancer according to embodiment 6, having a sequence as set forth in SEQ ID NO: 30 or having at least 95% sequence identity with the sequence as set forth in SEQ ID NO: 30. 16. An artificial enhancer according to embodiment 7, having a sequence as set forth in SEQ ID NO: 32 or having at least 95% sequence identity with the sequence as set forth in SEQ ID NO: 32. 17. An artificial enhancer according to embodiment 8, having a sequence as set forth in SEQ ID NO: 34 or having at least 95% sequence identity with the sequence as set forth in SEQ ID NO: 34. 18. An artificial enhancer according to embodiment 9, having a sequence as set forth in SEQ ID NO: 36 or having at least 95% sequence identity with the sequence as set forth in SEQ ID NO: 36. 19. An artificial enhancer according to embodiment 12, having a sequence as set forth in SEQ ID NO: 40 or having at least 95% sequence identity with the sequence as set forth in SEQ ID NO: 40. 20. An artificial expression construct comprising: (i)eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_830h, eHGT_831h, eHGT_832h, e HGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHGT_577h, 3xcore3_eHGT_577h, 3xcore2_eHGT_606h, 3xcore3_eHGT _606h, core4_eHGT_577h, core6_eHGT_606h, 3xcore_eHGT_121h, eHGT_590m, eHGT_976h, MGT_E117, MGT_E118, MGT_E119, MGT_E120, MGT_E121, 3xCore2-eHGT_367h, eHGT_359h, eHGT_479m, eHGT_453m, eHGT_140h, eHGT_356h, eHGT_128h, eHGT_369h and eHGT_710m; (ii) a promoter; and (iii) heterologous coding sequence An artificial expression construct comprising: 21. The artificial expression construct of embodiment 20, wherein the heterologous coding sequence encodes an effector element or an expressible element. 22. The artificial expression construct of embodiment 20 or 21, wherein the effector element comprises a reporter protein or a functional molecule. 23. The artificial expression construct of embodiment 22, wherein the reporter protein is a fluorescent protein. 24. The artificial expression construct of embodiment 22 or 23, wherein the functional molecule is a functional ion transporter, a functional enzyme, a functional transcription factor, a functional receptor, a functional membrane protein, a functional cell trafficking protein, a functional signaling molecule, a functional neurotransmitter, a functional calcium reporter, a functional channelrhodopsin, a functional CRISPR / Cas molecule, a functional editase, a functional guide RNA molecule, a functional microRNA, a functional homologous recombination donor cassette, or a functional designer receptor activated only by designer drugs (DREADD). 25. The artificial expression construct of embodiment 21, wherein the expressible element is a non-functional molecule. 26. The artificial expression construct of embodiment 25, wherein the non-functional molecule is a non-functional ion transporter, a non-functional enzyme, a non-functional transcription factor, a non-functional receptor, a non-functional membrane protein, a non-functional cell transport protein, a non-functional signaling molecule, a non-functional neurotransmitter, a non-functional calcium reporter, a non-functional channelrhodopsin, a non-functional CRISPR / Cas molecule, a non-functional editase, a non-functional guide RNA molecule, a non-functional microRNA, a non-functional homologous recombination donor cassette, or a non-functional designer receptor activated only by designer drugs (DREADD). 27. An artificial expression construct according to any one of embodiments 20 to 26, which is associated with a capsid that crosses the blood-brain barrier. 28. The artificial expression construct of embodiment 27, wherein the capsid is PHP.eB, AAV-BR1, AAV-PHP.S, AAV-PHP.B or AAV-PPS. 29. An artificial expression construct according to any one of embodiments 20 to 28, comprising or encoding a skipping element. 30. The artificial expression construct of embodiment 29, wherein said skipping element is a 2A peptide and / or an internal ribosome entry site (IRES). 31. The artificial expression construct of embodiment 30, wherein the 2A peptide is T2A, P2A, E2A or F2A. 32.eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_830h, eHGT_ 831h, eHGT_832h, eHGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHGT_577h, 3xcore3_eHGT_57 7h, 3xcore2_eHGT_606h, 3xcore3_eHGT_606h, core4_eHGT_577h, core6_eHGT_606h, 3xcore_eHGT_121h, eHGT_590m, eHGT_976h, MGT_E117, MGT_E118, MGT_E119, MGT_E120, MGT_E121, 3xCore2-eHGT_367h, eHGT_3 32. The artificial expression construct according to any one of embodiments 20 to 31, comprising or encoding a combination of features selected from: 59h, eHGT_479m, eHGT_453m, eHGT_140h, eHGT_356h, eHGT_128h, eHGT_369h, eHGT_710m, AAV, scAAV, rAAV, pAAV, minBglobin, CMV, minCMV, minCMV*, minRho, minRho*, fluorescent proteins (e.g., EGFP, SYFP, GFP), hsA2, Cre, iCre, dgCre, FlpO, tTA2, SP10ins (e.g., 3xSP10ins), tag cassette, 10aa, nuclear transport protein, self-cleaving peptide, WPRE, WPRE3, hGHpA and / or BGHpA. 33.eHGT_576h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_577h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_578h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; 3xSP10ins-eHGT_579h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_606h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_827h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_828h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_830h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_831h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_832h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_834h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_836h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; MGT_E117-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; MGT_E118-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; MGT_E119-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; MGT_E120-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; MGT_E121-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; 3xSP10ins-3xcore2_eHGT_367h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; 3xSP10ins-eHGT_359h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_479m-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_453m-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; hsA2-eHGT_140h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; 3xSP10ins-eHGT_356h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; hsA2-eHGT_128h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; 3xcoreB_eHGT121h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_710m-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_895h-[minimal promoter]-[heterologous coding sequence]-P2A-3XFLAG-10aa-H2B-WPRE3-BGHpA; 3xcore2_eHGT_577h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; 3xcore3_eHGT_577h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; 3xcore2_eHGT_606h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; 3xcore3_eHGT_606h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; core4_eHGT_577h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; core6_eHGT_606h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; 3xcore_eHGT_121h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_590m-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_976h-[minimal promoter]-[heterologous coding sequence]-P2A-3XFLAG-10aa-H2B-WPRE3-BGHpA; eHGT_717h-[minimal promoter]-[heterologous coding sequence]-WPRE3-BGHpA; eHGT_576h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_577h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_578h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; 3xSP10ins-eHGT_579h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_606h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_827h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_828h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_830h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_831h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_832h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_834h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_836h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; MGT_E117-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; MGT_E118-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; MGT_E119-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; MGT_E120-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; MGT_E121-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; 3xSP10ins-3xcore2_eHGT_367h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; 3xSP10ins-eHGT_359h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_479m-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_453m-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; hsA2-eHGT_140h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; 3xSP10ins-eHGT_356h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; hsA2-eHGT_128h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; 3xcoreB_eHGT121h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_710m-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_895h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; 3xcore2_eHGT_577h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; 3xcore3_eHGT_577h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; 3xcore2_eHGT_606h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; 3xcore3_eHGT_606h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; core4_eHGT_577h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; core6_eHGT_606h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; 3xcore_eHGT_121h-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_590m-[minimal promoter]-[heterologous coding sequence]-[post-transcriptional regulatory element]; eHGT_976h - [minimal promoter] - [heterologous coding sequence] - [post-transcriptional regulatory element]; and eHGT_717h - [minimal promoter] - [heterologous coding sequence] - [post-transcriptional regulatory element] 33. The artificial expression construct according to any one of embodiments 20 to 32, comprising or encoding a combination of features selected from: 34. A vector comprising an artificial expression construct according to any one of embodiments 20 to 33. 35. The vector according to embodiment 34, which is a viral vector. 36. The vector of embodiment 34 or 35, wherein the viral vector is a recombinant adeno-associated viral (AAV) vector. 37. An adeno-associated virus (AAV) vector comprising at least one heterologous coding sequence, the heterologous coding sequence being selected from the group consisting of eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_830h, eHGT_831h, eHGT_832h, eHGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHGT_577h, 3xcore3_eHGT_577h, 3xcore2_eHGT_606h, 3xcore3_eHGT_717 ... eHGT_590m, eHGT_976h, MGT_E117, MGT_E118, MGT_E119, MGT_E120, MGT_E121, 3xCore2-eHGT_367h, eHGT_359h, eHGT_479m, eHGT_453m, eHGT_140h, eHGT_356h, eHGT_128h, eHGT_369h and eHGT_710m. 38. The AAV vector of embodiment 37, wherein the heterologous coding sequence encodes an effector element or an expressible element. 39. The AAV vector of embodiment 38, wherein the effector element comprises a reporter protein or functional molecule. 40. The AAV vector of embodiment 39, wherein the reporter protein is a fluorescent protein. 41. The AAV vector of embodiment 39, wherein the functional molecule is a functional ion transporter, a functional enzyme, a functional transcription factor, a functional receptor, a functional membrane protein, a functional cell transport protein, a functional signaling molecule, a functional neurotransmitter, a functional calcium reporter, a functional channelrhodopsin, a functional CRISPR / Cas molecule, a functional editase, a functional guide RNA molecule, a functional microRNA, a functional homologous recombination donor cassette, or a functional designer receptor activated only by designer drugs (DREADD). 42. The AAV vector of embodiment 38, wherein the expressible element is a non-functional molecule. 43. The AAV vector of embodiment 42, wherein the non-functional molecule is a non-functional ion transporter, a non-functional enzyme, a non-functional transcription factor, a non-functional receptor, a non-functional membrane protein, a non-functional cell trafficking protein, a non-functional signaling molecule, a non-functional neurotransmitter, a non-functional calcium reporter, a non-functional channelrhodopsin, a non-functional CRISPR / Cas molecule, a non-functional editase, a non-functional guide RNA molecule, a non-functional microRNA, a non-functional homologous recombination donor cassette, or a non-functional designer receptor activated only by designer drugs (DREADD). 44. A transgenic cell comprising an artificial expression construct or vector according to any one of the preceding embodiments. 45. The transgenic cell of embodiment 44, which is a neuron in the thalamus. 46. The transgenic cell of embodiment 44, which is a GABAergic or glutamatergic neuron in the thalamus. 47. The transgenic cell of embodiment 46, wherein the GABAergic neuron is a Gata / Dlx5-6 cell. 48. The transgenic cell of embodiment 46, wherein the glutamatergic neuron is a Prkcd-Grin2c cell. 49. The transgenic cell of embodiment 46, wherein the glutamatergic neuron is an Rxfp1-Epb4 cell. 50. The transgenic cell of embodiment 46, wherein the GABAergic neurons include thalamic reticular nucleus (TRN) cells of the thalamus. 51. The transgenic cell of embodiment 46, wherein the glutamatergic neurons comprise the parafascicular nucleus (Pf) of the thalamus. 52. The transgenic cell of embodiment 44, which is a second type of cell. 53. The transgenic cell of embodiment 52, wherein the second type of cell is a striatal medium spiny neuron. 54. The transgenic cell of embodiment 52, wherein the second type of cell is a Purkinje cell of the cerebellum. 55. The transgenic cell of embodiment 52, wherein the second type of cell is a deep cerebellar nucleus (DCN) cell of the cerebellum. 56. The transgenic cell of embodiment 52, wherein the second type of cell is a cerebellar molecular layer interneuron cell (MLI). 57. The transgenic cell of embodiment 52, wherein the second type of cell is a Pvalb positive neuron. 58. The transgenic cell of embodiment 52, wherein the second type of cell is a chandelier cell. 59. The transgenic cell of embodiment 52, wherein the second type of cell is a neocortical glutamatergic L5 ET cell. 60. The transgenic cell of embodiment 52, wherein the second type of cell is a Vip-positive neuron of the neocortex. 61. The transgenic cell of embodiment 44, which is a mouse cell, a human cell or a non-human primate cell. 62. A non-human transgenic animal comprising an artificial expression construct, vector and / or transgenic cell according to any one of the preceding embodiments. 63. The non-human transgenic animal of embodiment 62, which is a mouse or a non-human primate. 64. An administrable composition comprising an artificial expression construct, a vector and / or a transgenic cell according to any one of the preceding embodiments. 65. A kit comprising an artificial expression construct, a vector, a transgenic cell, a non-human transgenic animal and / or an administrable composition according to any one of the preceding embodiments. 66. A method for expressing a gene in a cell population in or derived from the thalamus, in vivo or in vitro, comprising providing to a sample or subject comprising a cell population in or derived from the thalamus an administrable composition according to embodiment 64 in a sufficient dose and for a sufficient period of time, thereby expressing the gene in the cell population. 67. The method of embodiment 66, wherein the gene encodes an effector element or an expressible element. 68. The method of embodiment 67, wherein the effector element comprises a reporter protein or functional molecule. 69. The method of embodiment 68, wherein the reporter protein is a fluorescent protein. 70. The method of embodiment 68, wherein the functional molecule is a functional ion transporter, a functional enzyme, a functional transcription factor, a functional receptor, a functional membrane protein, a functional cell transport protein, a functional signaling molecule, a functional neurotransmitter, a functional calcium reporter, a functional channelrhodopsin, a functional CRISPR / Cas molecule, a functional editase, a functional guide RNA molecule, a functional microRNA, a functional homologous recombination donor cassette, or a functional designer receptor activated only by designer drugs (DREADD). 71. The method of embodiment 67, wherein the expressible element is a non-functional molecule. 72. The method of embodiment 71, wherein the non-functional molecule is a non-functional ion transporter, a non-functional enzyme, a non-functional transcription factor, a non-functional receptor, a non-functional membrane protein, a non-functional cell transport protein, a non-functional signal transduction molecule, a non-functional neurotransmitter, a non-functional calcium reporter, a non-functional channelrhodopsin, a non-functional CRISPR / Cas molecule, a non-functional editase, a non-functional guide RNA molecule, a non-functional microRNA, a non-functional homologous recombination donor cassette, or a non-functional designer receptor activated only by designer drugs (DREADD). 73. The method of any one of embodiments 66-72, wherein the providing step comprises pipetting. 74. The method of embodiment 73, wherein the pipetting is performed on a brain slice. 75. The method of embodiment 74, wherein the brain slice comprises neurons in the thalamus. 76. The method of embodiment 74, wherein the brain slice comprises GABAergic or glutamatergic neurons in the thalamus. 77. The method of embodiment 74, wherein the brain slice comprises GABAergic Gata / Dlx5-6 neurons. 78. The method of embodiment 74, wherein the brain slice comprises glutamatergic Prkcd-Grin2c neurons. 79. The method of embodiment 74, wherein the brain slice comprises glutamatergic Rxfp1-Epb4 neurons. 80. The method of embodiment 74, wherein the brain slice comprises GABAergic neurons in the thalamic reticular nucleus (TRN) cells. 81. The method of embodiment 74, wherein the brain slice comprises glutamatergic neurons of the parafascicular thalamic nucleus (Pf). 82. The method of embodiment 74, wherein the brain slice comprises a second type of cell. 83. The method of embodiment 82, wherein the second type of cells are striatal medium spiny neurons. 84. The method of embodiment 82, wherein the second type of cells are Purkinje cells of the cerebellum. 85. The method of embodiment 82, wherein the second type of cells are deep cerebellar nuclei (DCN) cells of the cerebellum. 86. The method of embodiment 82, wherein the second type of cells are cerebellar molecular layer interneuron cells (MLI). 87. The method of embodiment 82, wherein the second type of cells are Pvalb positive neurons. 88. The method of embodiment 82, wherein the second type of cells are chandelier cells. 89. The method of embodiment 82, wherein the second type of cells are neocortical glutamatergic L5 ET cells. 90. The method of embodiment 82, wherein the second type of cells are Vip-positive neurons of the neocortex. 91. The method of embodiment 74, wherein the brain slice is a mouse, human or non-human primate brain slice. 92. The method of any one of embodiments 66-91, wherein the providing step comprises administration to a living subject. 93. The method of embodiment 92, wherein the living subject is a human, a non-human primate, or a mouse. 94. The method of embodiment 92 or 93, wherein administration to the living subject is by injection. 95. The method of embodiment 94, wherein the injection is an intravenous injection, an intraparenchymal injection into brain tissue, an intracerebroventricular (ICV) injection, an intracisternal (ICM) injection or an intrathecal injection. 96. An artificial expression construct comprising: or comprising the sequence set forth in SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:88, SEQ ID NO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, SEQ ID NO:93, SEQ ID NO:94, SEQ ID NO:95, SEQ ID NO:96, SEQ ID NO:97, SEQ ID NO:98, SEQ ID NO:99, SEQ ID NO:100, SEQ ID NO:101, SEQ ID NO:102, SEQ ID NO:103, SEQ ID NO:104, SEQ ID NO:105, SEQ ID NO:106, SEQ ID NO:107, SEQ ID NO:108, SEQ ID NO:109, SEQ ID NO:110, SEQ ID NO:111, SEQ ID NO:112, SEQ ID NO:113, SEQ ID NO:114, SEQ ID NO:115, SEQ ID NO:116, SEQ ID NO:117 or SEQ ID NO:118; comprising a sequence having at least 90% sequence identity to the sequence set forth in SEQ ID NO:83, SEQ ID NO:84, SEQ ID NO:85, SEQ ID NO:86, SEQ ID NO:87, SEQ ID NO:139, SEQ ID NO:140, SEQ ID NO:88, SEQ ID NO:89, SEQ ID NO:90, SEQ ID NO:91, SEQ ID NO:92, SEQ ID NO:93, SEQ ID NO:94, SEQ ID NO:95, SEQ ID NO:96, SEQ ID NO:97, SEQ ID NO:98, SEQ ID NO:99, SEQ ID NO:100, SEQ ID NO:101, SEQ ID NO:102, SEQ ID NO:103, SEQ ID NO:104, SEQ ID NO:105, SEQ ID NO:106, SEQ ID NO:107, SEQ ID NO:108, SEQ ID NO:109, SEQ ID NO:110, SEQ ID NO:111, SEQ ID NO:112, SEQ ID NO:113, SEQ ID NO:114, SEQ ID NO:115, SEQ ID NO:116, SEQ ID NO:117 or SEQ ID NO:118, Artificial expression constructs.
[0151] (viii) Conclusion Variants of the sequences disclosed and referenced herein are also included herein. Guidelines for determining which amino acid residues can be substituted, inserted or deleted without losing biological activity can be determined using computer programs well known in the art, such as DNASTAR. TM Software (Madison, WI, USA) can be used to find the amino acid changes in the protein variants disclosed herein. The amino acid changes are preferably conservative amino acid changes, i.e., substitutions of similarly charged amino acids with each other or of uncharged amino acids with each other. Conservative amino acid changes include substitutions with members of a family of amino acids whose side chains are related.
[0152] Suitable conservative substitutions of amino acids in peptides or proteins are known to those skilled in the art, and such conservative substitutions can be made without generally altering the biological activity of the resulting molecule.Those skilled in the art will be familiar with the fact that generally, a single amino acid substitution in a non-essential region of a polypeptide will not substantially alter the biological activity (see, for example, Watson et al. Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. Co., p. 224). Naturally occurring amino acids are generally classified into conservative substitution families, specifically: Group 1: alanine (Ala), glycine (Gly), serine (Ser), and threonine (Thr); Group 2: (acidic): aspartic acid (Asp) and glutamic acid (Glu); Group 3: (acidic; also classified as polar, negatively charged residues and their amides): asparagine (Asn), glutamine (Gln), Asp, and Glu; Group 4: Gln and Asn; Group 5: (basic; also classified as polar, positively charged residues): arginine (Arg), lysine (Lys), and histidine (His); Group 6 (large aliphatic nonpolar residues): isoleucine (Ile), leucine (Leu), and ketone (K). Group 7 (polar uncharged): tyrosine (Tyr), Gly, Asn, Gln, Cys, Ser and Thr; Group 8 (large aromatic residues): phenylalanine (Phe), tryptophan (Trp) and Tyr; Group 9 (non-polar): proline (Pro), Ala, Val, Leu, Ile, Phe, Met and Trp; Group 11 (aliphatic): Gly, Ala, Val, Leu and Ile; Group 10 (small aliphatic residues that are non-polar or slightly polar): Ala, Ser, Thr, Pro and Gly; and Group 12 (sulfur-containing residues): Met and Cys. Further information can be found in Creighton (1984) Proteins, WH Freeman and Company.
[0153] In making such changes, the hydropathic index of amino acids may be taken into consideration. The importance of the hydropathic index of amino acids in conferring interactive biological function on a protein is widely understood in the art (Kyte and Doolittle, 1982, J. Mol. Biol. 157(1), 105-32). Each amino acid has been assigned a hydropathic index on the basis of its hydrophobicity and charge characteristics (Kyte and Doolittle, 1982). The hydrophobicity index of each amino acid is Ile (+4.5); Val (+4.2); Leu (+3.8); Phe (+2.8); Cys (+2.5); Met (+1.9); Ala (+1.8); Gly (-0.4); Thr (-0.7); Ser (-0.8); Trp (-0.9); Tyr (-1.3); Pro (-1.6); His (-3.2); glutamic acid (-3.5); Gln (-3.5); aspartic acid (-3.5); Asn (-3.5); Lys (-3.9); and Arg (-4.5).
[0154] It is well known in the art that substitution of a particular amino acid with another amino acid having a similar hydrophobicity index or hydrophobicity degree can also result in a protein with similar biological activity, i.e., a protein with biologically equivalent functionality. When making such changes, substitution of amino acids with hydrophobicity indices within ±2 is preferred, substitution of amino acids with hydrophobicity indices within ±1 is particularly preferred, and substitution of amino acids with hydrophobicity indices within ±0.5 is even more particularly preferred. Furthermore, it is well known in the art that substitution of similar amino acids can be effectively carried out based on hydrophilicity.
[0155] As detailed in U.S. Patent No. 4,554,101, each amino acid residue is assigned a hydrophilicity value, which is as follows: Arg (+3.0); Lys (+3.0); Aspartic acid (+3.0±1); Glutamic acid (+3.0±1); Ser (+0.3); Asn (+0.2); Gln (+0.2); Gly (0); Thr (-0.4); Pro (-0.5±1); Ala (-0.5); His (-0.5); Cys (-1.0); Met (-1.3); Val (-1.5); Leu (-1.8); Ile (-1.8); Tyr (-2.3); Phe (-2.5); Trp (-3.4). It is well known that certain amino acids can be substituted with other amino acids having a similar hydrophilicity value, and that such substitutions will result in biologically equivalent proteins, and in particular immunologically equivalent proteins. When making such changes, substitutions between amino acids whose hydrophilicity values are within the range of ±2 are preferred, substitutions between amino acids whose hydrophilicity values are within the range of ±1 are particularly preferred, and substitutions between amino acids whose hydrophilicity values are within the range of ±0.5 are even more particularly preferred.
[0156] As outlined above, amino acid substitutions may be made on the basis of the relative similarity of the amino acid side-chain substituents, for example, their hydrophobicity, hydrophilicity, charge, size, and the like.
[0157] As described elsewhere herein, variants of a gene sequence include codon-optimized variants, sequence polymorphisms, splice variants, and / or mutations that have no statistically significant effect on the function of the encoded product.
[0158] Variants of the protein, nucleic acid and gene sequences disclosed herein also include sequences having at least 70% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, at least 96% sequence identity, at least 97% sequence identity, at least 98% sequence identity or at least 99% sequence identity to the protein, nucleic acid or gene sequences disclosed herein.
[0159] "Percent sequence identity" refers to the relatedness of two or more sequences, as determined by comparing the sequences. In the art, "identity" also means the degree of relatedness between protein, nucleic acid or gene sequences, as determined by the matching between strings of protein, nucleic acid or gene sequences. "Identity" (often referred to as "similarity") can be readily calculated by known methods, including those described in Computational Molecular Biology (Lesk, AM, ed.) Oxford University Press, NY (1988); Biocomputing: Informatics and Genome Projects (Smith, DW, ed.) Academic Press, NY (1994); Computer Analysis of Sequence Data, Part I (Griffin, AM, and Griffin, HG, eds.) Humana Press, NJ (1994); Sequence Analysis in Molecular Biology (Von Heijne, G., ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J., eds.) Oxford University Press, NY (1992). Methods for determining identity are preferably designed to give the best match between the sequences tested. Methods for determining identity and similarity are codified in publicly available computer programs. Sequence alignment and identity calculations may be performed using the Megalign program (DNASTAR, Inc., Madison, Wis.) in the LASERGENE suite of bioinformatics computing software.Multiple alignment of sequences can also be performed using the Clustal format alignment method (Higgins and Sharp CABIOS, 5, 151-153 (1989) using default parameters (gap penalty=10, gap length penalty=10)). Related programs further include the GCG suite of programs (Wisconsin package version 9.0, Genetics Computer Group (GCG), Madison, Wisconsin); BLASTP, BLASTN, BLASTX (Altschul, et al., J. Mol. Biol. 215:403-410 (1990)); DNASTAR (DNASTAR, Inc., Madison, Wisconsin); and the FASTA program incorporating the Smith-Waterman algorithm (Pearson, Comput. Methods Genome Res., [Proc. Int. Symp.] (1994), Meeting Date 1992, 111-20. Editor(s): Suhai, Sandor. Publisher: Plenum, New York, NY). In this disclosure, when sequence analysis software is used for analysis, the analysis results are interpreted as being based on the "default values" that are the basis of the program. In this specification, "default values" refers to a set of numerical values or parameters that are preregistered in the software at the time of initialization of the software.
[0160] Variants also include nucleic acid molecules that hybridize to the sequences disclosed herein under stringent hybridization conditions and have the same function as the reference sequences. Exemplary stringent hybridization conditions include overnight incubation at 42°C in a solution containing 50% formamide, 5xSSC (750mM NaCl, 75mM trisodium citrate), 50mM sodium phosphate (pH 7.6), 5xDenhardt's solution, 10% dextran sulfate, and 20μg / ml denatured salmon sperm DNA that has been fragmented, followed by washing the filter at 50°C with 0.1xSSC. Hybridization stringency and signal detection are altered primarily by adjusting the concentration of formamide (lower percentage of formamide results in lower stringency), salt conditions, or temperature. For example, moderately stringent conditions include overnight incubation at 37° C. in 6×SSPE (20×SSPE=3M NaCl; 0.2M NaH2PO4; 0.02M EDTA, pH 7.4), 0.5% SDS, 30% formamide, 100 μg / ml blocking salmon sperm DNA, followed by washing with 1×SSPE and 0.1% SDS at 50° C. Even lower stringency is achieved by performing stringent post-hybridization washes at high salt concentrations (e.g., 5×SSC). The above conditions can be varied by adding and / or substituting other blocking reagents used to reduce the background of hybridization experiments. Common blocking reagents include Denhardt's reagent, BLOTTO, heparin, denatured salmon sperm DNA, and commercially available proprietary preparations. The addition of certain blocking reagents may require some modification of the hybridization conditions described above due to compatibility issues.
[0161] The term "concatemerize" is used in a broad sense and means to link in a chain or to link in a series. The term is used to describe the linking of multiple nucleotide sequences to obtain a single nucleotide sequence or the linking of multiple amino acid sequences to obtain a single amino acid sequence. Also, "concatemerize" is understood to refer to "concatenation." As will be understood by one of ordinary skill in the art, each embodiment disclosed herein comprises, consists essentially of, or consists of the particular components, steps, materials, or ingredients described. Thus, the terms "comprise" or "comprising" should be interpreted to mean "comprise, consist essentially of, or consist of." The transitional phrase "comprising" means that unrecited components, steps, materials, or ingredients are included, even if in greater amounts. The transitional phrase "consisting of" excludes all unrecited components, steps, materials, or ingredients. The transitional phrase "consisting essentially of" limits the scope of the embodiment to the recited components, steps, materials, or ingredients and those components, steps, materials, or ingredients that do not materially affect the embodiment. Significant effects were defined as a statistically significant decrease in targeted expression as measured by scRNA-Seq in the target cell population, and included the following enhancer and target cell population combinations: eHGT_576h, eHGT_577h, eHGT_578h, eHGT_579h, eHGT_606h, eHGT_827h, eHGT_828h, eHGT_830h, eHGT_831h, eHGT_832h, eHGT_834h, eHGT_836h, eHGT_717h, eHGT_895h, 3xcore2_eHGT_577h, 3xcore3_eHGT_577h, 3xcore2_eHGT_606h, 3xcore3_eHGT_606h, core4_eHGT_ 577h, core6_eHGT_606h, 3xcore_eHGT_121h, eHGT_590m or eHGT_976h / glutamatergic neurons in the thalamus; MGT_E117 or MGT_E118 / glutamatergic neurons in the thalamus (Prkcd-Grin2c (Core, LGN)); MGT_E119 or MGT_E120 / GABAergic neurons in the thalamus (Gata / Dlx5-6); MGT_E121 / glutamatergic neurons in the thalamus (Rxfp1-Epb4 (Matrix)); 3xCore2-eHGT_367h / glutamatergic neurons in the parafascicular nucleus (Pf) of the thalamus and striatal medium spiny neurons (all types included);eHGT_359h / glutamatergic neurons in the thalamus, Purkinje cells in the cerebellum, and Pvalb-positive neurons in the neocortex; eHGT_479m / glutamatergic neurons in the thalamus, Purkinje cells in the cerebellum, and chandelier cells in the neocortex; eHGT_453m / GABAergic neurons in the thalamic reticular nucleus (TRN) of the thalamus, deep cerebellar nucleus (DCN) cells in the cerebellum, and glutamatergic layer 5 neurons (L5 eHGT_140h / GABAergic neurons and Pvalb-positive neurons in the thalamic reticular nucleus (TRN) of the thalamus; eHGT_356h / GABAergic neurons in the thalamic reticular nucleus (TRN) of the thalamus, DCN cells of the cerebellum, Vip-positive neurons and thalamic reticular nucleus cells of the neocortex; eHGT_128h / glutamatergic neurons and Pvalb-positive neurons in the thalamus; eHGT_369h / glutamatergic neurons in the thalamus, molecular layer interneuron (MLI) cells of the cerebellum and Pvalb-positive neurons of the neocortex; and eHGT_710m / glutamatergic neurons in the thalamus, MLI cells of the cerebellum, chandelier cells and molecular layer GABAergic interneurons of the cerebellum.
[0162] In certain embodiments, "artificial" means not naturally occurring.
[0163] Unless otherwise indicated, all numerical values expressing quantities or properties of materials, such as molecular weight and reaction conditions, in the specification and claims are to be construed in all instances as modified by the term "about." Accordingly, unless otherwise indicated, the numerical parameters set forth in the specification and appended claims are approximations that may vary depending upon the desired properties sought to be obtained by the present invention. Without intending to limit the scope of the doctrine of equivalents to the scope of the claims, each numerical parameter should, at the very least, be construed in light of the number of reported significant digits and by applying ordinary rounding procedures. For clarity, the term "about," when used in conjunction with a stated value or range, has a meaning that would be reasonably interpreted by one of ordinary skill in the art, i.e., within ±20% of the stated value; within ±19% of the stated value; within ±18% of the stated value; within ±17% of the stated value; within ±16% of the stated value; within ±15% of the stated value; within ±14% of the stated value; within ±13% of the stated value; within ±12% of the stated value; within ±11% of the stated value; within ±10% of the stated value; within ±9% of the stated value; within ±8% of the stated value; within ±7% of the stated value; within ±6% of the stated value; within ±5% of the stated value; within ±4% of the stated value; within ±3% of the stated value; within ±2% of the stated value; or within ±1% of the stated value.
[0164] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the invention are approximations and approximate ranges, the numerical values set forth in the specific examples are reported as precisely as possible, however, all numerical values inherently contain certain errors necessarily resulting from the standard deviation associated with their respective testing measurements.
[0165] In the description of the present invention (particularly in the description of the claims below), the terms "a", "an", "the" and similar modifiers are intended to include both the singular and the plural unless otherwise indicated or the context clearly indicates otherwise. Numerical ranges described herein are intended to be a shorthand way of referring to each numerical value falling within the range individually. Unless otherwise indicated, each numerical value is described herein as if it were individually described herein. Any method described herein can be performed in any suitable order unless otherwise indicated or the context clearly indicates otherwise. The use of any examples or language of examples (e.g., "etc.") provided herein is intended to be for the purpose of illustrating the invention only and does not limit the scope of the invention as described in the claims. No term described herein should be construed as indicating any non-claimed element essential to the practice of the invention.
[0166] Groupings of other elements of the invention disclosed herein or of various embodiments of the invention should not be construed as limiting the invention. Members of each group may be described herein or in the claims individually or in combination with other members of the group or other elements described herein. It is anticipated that for reasons of convenience and / or patentability, one or more members of a group may be added to another group, or one or more members may be deleted from a group. When such additions or deletions are made, the specification includes groups that are constructed to satisfy the recitation of all Markush groups set forth in the appended claims.
[0167] Specific embodiments of the present invention are described herein, including those embodiments known to the inventors to be the best mode for carrying out the invention. Of course, those skilled in the art will readily appreciate that the embodiments described herein may be modified in various ways upon review of the above detailed description. The inventors anticipate that such modifications may be adopted by those skilled in the art, and intend that the present invention may be practiced in other ways than as specifically described herein. Accordingly, the present invention includes all modifications of the subject matter recited in the appended claims and all equivalents of the subject matter of the present invention to the extent permitted within the scope of applicable law. Moreover, the present invention includes all combinations of the above-described elements in any and all variations thereof, unless otherwise indicated or the context clearly dictates otherwise.
[0168] Additionally, throughout this specification, various patents, publications, journal articles and other documents are cited (references herein). Each reference cited herein is individually incorporated herein by reference for the teachings thereof as if it were a part of this specification.
[0169] Finally, the embodiments of the invention disclosed herein are to be considered as illustrative of the principles of the invention. Other modifications may be adopted within the scope of the invention. Thus, by way of example, but not by way of limitation, alternative configurations of the invention may be utilized in accordance with the teachings herein. Thus, the invention is not to be limited to what has been precisely shown and described herein.
[0170] The details described herein are by way of example and are presented solely for the purpose of illustrating preferred embodiments of the present invention, to provide what is believed to be the most useful, and to facilitate an understanding of the principles and conceptual aspects of various embodiments of the present invention. In this regard, no structural details of the present invention are described in more detail than is necessary for a basic understanding of the present invention, and those skilled in the art will be able to easily understand how to actually embody some forms of the present invention by reading the description of the present invention in conjunction with the drawings and / or examples.
[0171] The definitions and explanations used in this disclosure are intended to control future interpretations, unless clearly and definitively changed in the following examples, or unless the interpretation is rendered meaningless or substantially meaningless by the meaning of the term. If the definition of a term is rendered meaningless or substantially meaningless by the interpretation of the term, please refer to the definition of the term from a dictionary known to those skilled in the art, such as Webster's Dictionary (3rd Edition) or Oxford Dictionary of Biochemistry and Molecular Biology (Ed. Anthony Smith, Oxford University Press, Oxford, 2004).
Claims
1. An artificially expressed construct, (i) having a sequence selected from the sequences shown in SEQ ID NO:2, SEQ ID NO:1, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, SEQ ID NO:34, SEQ ID NO:36, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:16, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14 and SEQ ID NO:15, or having a sequence having at least 95% sequence identity with the sequence shown in SEQ ID NO:2, SEQ ID NO:1, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:137, SEQ ID NO:138, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:32, SEQ ID NO:34, SEQ ID NO:36, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:16, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14 or SEQ ID NO:15; (ii) a promoter; and (iii) a coding sequence An artificially expressed construct comprising the above.
2. The artificially expressed construct according to claim 1, wherein the coding sequence encodes a fluorescent protein or a neurotransmitter.
3. The artificially expressed construct according to claim 1, which is associated with a capsid that passes through the blood-brain barrier.
4. The artificially expressed construct according to claim 3, wherein the capsid comprises PHP.eB, AAV-BR1, AAV-PHP.S, AAV-PHP.B, AAV9, AAVrh.10 or AAV-PPS.
5. The artificially expressed construct according to claim 1, which comprises a skipping element or encodes a skipping element.
6. The artificially expressed construct according to claim 5, wherein the skipping element comprises a T2A peptide, a P2A peptide, an E2A peptide, an F2A peptide and / or an internal ribosome entry site (IRES).
7. The artificially expressed construct according to claim 1, which is within a viral vector.
8. A composition comprising an artificially expressed construct for expressing a coding sequence in a cell population in the thalamus or a cell population derived from the thalamus in vivo or in vitro, the artificial expression construct has (i) an enhancer having a sequence selected from the sequences shown in SEQ ID NO: 2, SEQ ID NO: 1, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 32, SEQ ID NO: 34, SEQ ID NO: 36, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 16, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14 and SEQ ID NO: 15, or having a sequence having at least 95% sequence identity with the sequence shown in SEQ ID NO: 2, SEQ ID NO: 1, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 137, SEQ ID NO: 138, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 29, SEQ ID NO: 30, SEQ ID NO: 32, SEQ ID NO: 34, SEQ ID NO: 36, SEQ ID NO: 38, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 42, SEQ ID NO: 43, SEQ ID NO: 16, SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14 or SEQ ID NO: 15; (ii) a promoter; and (iii) a coding sequence and is used to express a coding sequence in a cell population in a sample or subject comprising a cell population within the thalamus or a cell population derived from the thalamus, by being provided to the sample or subject in a sufficient dose and for a sufficient period of time. A composition characterized by that.
9. The composition according to claim 8, wherein the coding sequence encodes a fluorescent protein or a neurotransmitter.
10. The composition according to claim 8, wherein the providing includes pipetting performed on a brain slice.
11. The composition according to claim 10, wherein the brain slice includes GABAergic neurons within the thalamus or glutamatergic neurons within the thalamus.
12. The composition according to claim 10, wherein the brain slice includes GABAergic Gata / Dlx5-6 neurons.
13. The composition according to claim 10, wherein the brain slice includes glutamatergic Prkcd-Grin2c neurons or glutamatergic Rxfp1-Epb4 neurons.
14. The composition according to claim 8, wherein the providing includes administration to a living subject.
15. The composition according to claim 14, wherein the living subject is a human, a non-human primate or a mouse.
16. The composition according to claim 14, wherein the administration to the biological subject is carried out by injection.
17. The composition according to claim 16, wherein the injection is intravenous injection, intracerebral injection into brain tissue, intracerebroventricular (ICV) injection, intracisteral (ICM) injection or intrathecal injection.
18. An artificial enhancer comprising 2, 3, 4, 5, 6, 7, 8, 9 or 10 copies of the sequence shown in SEQ ID NO: 31, SEQ ID NO: 33, SEQ ID NO: 38, SEQ ID NO: 35, SEQ ID NO: 37, SEQ ID NO: 39 or SEQ ID NO: 41, or comprising 2, 3, 4, 5, 6, 7, 8, 9 or 10 copies of a sequence having at least 95% sequence identity with the sequence shown in SEQ ID NO: 31, SEQ ID NO: 33, SEQ ID NO: 38, SEQ ID NO: 35, SEQ ID NO: 37, SEQ ID NO: 39 or SEQ ID NO:
41.
19. The artificial enhancer according to claim 18, having the sequence shown in SEQ ID NO: 30, SEQ ID NO: 32, SEQ ID NO: 34, SEQ ID NO: 36 or SEQ ID NO: 40, or having a sequence having at least 95% sequence identity with the sequence shown in SEQ ID NO: 30, SEQ ID NO: 32, SEQ ID NO: 34, SEQ ID NO: 36 or SEQ ID NO: 40.