Constructs, Vectors, and Systems and Their Uses
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-02-24
- Publication Date
- 2026-03-16
AI Technical Summary
Current therapies for neurodegenerative diseases, such as ALS and frontotemporal dementia, face challenges due to the complex molecular mechanisms involved and the risk of adverse effects on healthy cells, limiting their efficacy and regulatory approval.
The development of a construct configured with specific regulatory domains and transgene sequences that are selectively expressed in cells with nuclear depletion of splicing factors of the hnRNP family, such as TDP-43, ensuring targeted protein production only in affected cells.
This approach allows for the selective production of functional proteins in affected cells, minimizing side effects in healthy cells and enhancing the therapeutic efficacy and safety of treatments for neurodegenerative diseases.
Smart Images

Figure 00000146_0000 
Figure 00000146_0001 
Figure 00000146_0002
Abstract
Description
[Technical field]
[0001] Constructs, Vectors, and Systems and Their Uses [Background technology]
[0002] background Neurodegenerative diseases are often fatal and, with very few exceptions, lack long-term effective treatments.Therefore, there is an urgent need for new therapies and treatments for neurodegenerative diseases, but progress has been hindered by a lack of understanding of the complex molecular mechanisms that underpin these diseases.
[0003] Although these disease mechanisms remain largely unknown, many neurodegenerative diseases have been identified as involving dysregulation of RNA-binding proteins (RBPs), including heterogeneous nuclear ribonucleoproteins (hnRNPs). hnRNPs are normally present in the nucleus and are involved in many steps of RNA metabolism, but also play a role in regulating alternative splicing that leads to exon skipping or intron retention.
[0004] One such protein of the hnRNP family is the TAR DNA-binding protein (TDP-43). Although initially identified as a DNA-binding protein, TDP-43 has been well characterized as a member of the hnRNP family of proteins and has important roles in neurodegenerative diseases, with TDP-43 showing abnormal localization in approximately 97% of cases of amyotrophic lateral sclerosis (ALS, a motor neuron disease), approximately half of cases of frontotemporal dementia, and the majority of cases of inclusion body myopathy (IBM). Furthermore, TDP-43 pathology has also been observed in Alzheimer's disease (AD) and other neurodegenerative diseases, including Parkinson's disease (PD) and Perry syndrome, suggesting that its role in neurodegeneration extends beyond ALS / FTD. Furthermore, a small proportion of ALS cases are caused by mutations in the TARDBP gene, which encodes TDP-43. TDP-43 in particular plays many roles in regulating RNA, from RNA transcription to RNA decay. Perhaps its best-characterized function is its role as a regulator of splicing, typically a splicing repressor. When localized near splice sites, TDP-43 has been shown to inhibit and silence splicing. It was first shown to regulate the splicing of the CFTR transcript in 2001, but numerous subsequent studies have shown that TDP-43 regulates many transcripts, including its own transcript. Neurodegenerative diseases that exhibit TDP-43 pathology typically show both cytoplasmic aggregation and nuclear depletion of TDP-43.
[0005] For example, local injection of a virus in combination with a cell type-specific transcription promoter (such as the synapsin promoter) can be used to target expression of a protein to a specific cell type, but this has the drawback that expression occurs in both diseased and non-diseased cells. Transgenic expression of these proteins can therefore significantly damage otherwise healthy cells, increasing the risk of adverse events (e.g., during clinical trials), increasing the side effects of any treatment, and reducing the likelihood of regulatory approval. These risks can be lowered by reducing the expression of the transgenic protein in the patient, but this also has the effect of reducing efficacy in diseased cells. Summary of the Invention
[0006] Thus, there is a need to overcome some of the drawbacks associated with conventional techniques and to develop new tools to further understand, target and correct the molecular mechanisms of dysregulation associated with neurodegenerative diseases. [Means for solving the problem]
[0007] Summary of the Invention In a first aspect, Start codon, A regulatory domain that includes: a first splice acceptor site and a first splice donor site; a binding domain for a splicing factor of the heterogeneous nuclear ribonucleoprotein (hnRNP) family located within 150 nucleotides of the first splice donor site and / or the first splice acceptor site; and / or located between the first splice donor site and the first splice acceptor site; and Transgene sequence Including, (i) when nuclear depletion of a splicing factor is localized in a cell, splicing of the first splice acceptor site and the first donor site is not inhibited and a functional protein is produced from the transgene sequence; (ii) when localized in cells without nuclear depletion of splicing factors, splicing of the first splice acceptor site and / or the first splice donor site is suppressed and no functional protein is produced from the transgene sequence. A construct configured to:
[0008] In an embodiment of the second aspect or the first aspect, Start codon, A regulatory domain that includes: a first splice acceptor site and a first splice donor site defining a cryptic exon sequence; an intron region defined by a second splice acceptor site and a second splice donor site, wherein the cryptic exon sequence is located within the intron region; a binding domain for a splicing factor of the heterogeneous nuclear ribonucleoprotein (hnRNP) family located within 150 nucleotides of the first splice donor site and / or the first splice acceptor site; and Transgene sequence Including, (i) when localized in a cell depleted of splicing factors, splicing of the first splice acceptor site and the first donor site is not repressed, the cryptic exon sequence is present in the mRNA product of the construct, and a functional protein is produced from the transgene sequence; (ii) when localized in a cell that is not depleted of splicing factors, splicing of the first splice acceptor site and / or the first donor site is repressed, the cryptic exon sequence is absent in the mRNA product of the construct, and no functional protein is produced from the transgene sequence. A construct configured to:
[0009] In embodiments of the second aspect, the transgene sequence is entirely downstream of the regulatory domain. These are described as "Design 1" embodiments described herein.
[0010] In another embodiment of the second aspect, at least a portion of the transgene sequence is encoded by a cryptic exon sequence. These are referred to as "Design 2" embodiments described herein.
[0011] In an embodiment of the third aspect or the first aspect, Start codon, A regulatory domain that includes: a first splice donor site and a first acceptor donor site defining a single regulatory intron; a binding domain for a splicing factor of the heterogeneous nuclear ribonucleoprotein (hnRNP) family located within 150 nucleotides of the first splice donor site and / or the first splice acceptor site and / or located between the first splice donor site and the first splice acceptor site; and Transgene sequence Including, (i) when localized in cells depleted of splicing factors, splicing of the first splice acceptor site and the first splice donor site is not repressed, and the single regulatory intron is spliced out to produce a functional protein from the transgene sequence; (ii) when localized in a cell that is not depleted of splicing factors, splicing of the first splice acceptor site and / or the first donor site is inhibited and the single regulatory intron is not spliced or is not properly spliced, resulting in the failure of a functional protein to be produced from the transgene sequence. A construct configured to:
[0012] In a fourth aspect of the invention there is provided a vector comprising a construct according to the above aspect.
[0013] In a fifth aspect of the present invention, there is provided a pharmaceutical composition comprising the construct of the above aspect, or the vector of the above aspect.
[0014] In a sixth aspect of the invention there is provided a system comprising any of the constructs and cells described herein, or any of the vectors and cells described herein, (i) when the cell nucleus is depleted of RNP family splicing factors (i.e., in diseased cells), the system produces functional proteins from the transgene sequence; (ii) in the absence of depletion of hnRNP family splicing factors from the cell nucleus (i.e., in healthy cells), the system does not produce functional protein from the transgene sequence; A system is provided.
[0015] In a seventh aspect of the invention there is provided any construct, vector or pharmaceutical composition described herein for use in therapy.
[0016] In an eighth aspect of the invention, a method for the treatment of a disease associated with a depletion of a splicing factor of the hnRNP family, the treatment comprising contacting a cell with a construct, vector or pharmaceutical composition, (i) In cells with nuclear depletion of RNP family splicing factors, the cells produce functional proteins, (ii) in cells without nuclear depletion of RNP family splicing factors, the cells do not produce functional proteins; Any construct, vector or pharmaceutical composition described herein is provided.
[0017] In some embodiments, the disease is a neurodegenerative disease or a muscular disease. In some embodiments, the neurodegenerative disease is amyotrophic lateral sclerosis (ALS) or frontotemporal dementia (FTD). In a preferred embodiment, the splicing factor of the hnRNP family is TDP-43.
[0018] In a ninth aspect of the invention there is provided the use of any construct as described herein, any vector as described herein or any pharmaceutical composition as described herein in a method for selectively producing a functional protein in a diseased cell where there is nuclear depletion of a splicing factor of the hnRNP family.
[0019] Also herein, Start codon, A regulatory domain that includes: a first splice acceptor site and a first splice donor site; a binding domain for a splicing factor of the heterogeneous nuclear ribonucleoprotein (hnRNP) family located within 150 nucleotides of the first splice donor site or the first splice acceptor site and / or located between the first splice donor site and the first splice acceptor site; and Transgene sequence, Including, (i) when placed in an in vitro system in which the splicing factor is depleted (i.e., absent), splicing of the first splice acceptor site and the first splice donor site is not inhibited and a functional protein is produced from the transgene sequence; (ii) when placed in an in vitro system in which splicing factors are not depleted (i.e., are present), splicing of the first splice acceptor site and / or the first splice donor site is inhibited and no functional protein is produced from the transgene sequence; Disclosed is a construct configured as follows:
[0020] The in vitro system must contain components that allow for transcription, splicing and translation. In some embodiments, these components are provided by the cells.
[0021] Also herein, Start codon, A regulatory domain that includes: a first splice acceptor site and a first splice donor site; a binding domain for a splicing factor of the heterogeneous nuclear ribonucleoprotein (hnRNP) family located within 150 nucleotides of the first splice donor site and / or the first splice acceptor site and / or located between the first splice donor site and the first splice acceptor site; and Transgene sequences encoding functional proteins Including, (i) when localized in cells with nuclear depletion of splicing factors, splicing of the first splice acceptor site and the first donor site is not inhibited and a functional protein is produced from the mRNA product of the construct; (ii) when localized in cells without nuclear depletion of splicing factors, splicing of the first splice acceptor site and / or the first donor site is suppressed and no functional protein is produced from the mRNA product of the construct; Also disclosed are constructs configured so as to
[0022] Also, Start codon, A regulatory domain that includes: a first splice acceptor site and a first splice donor site defining a cryptic exon sequence; an intron region defined by a second splice acceptor site and a second splice donor site, wherein the cryptic exon sequence is located within the intron region; a binding domain for a splicing factor of the heterogeneous nuclear ribonucleoprotein (hnRNP) family located within 150 nucleotides of the first splice donor site and / or the first splice acceptor site; and Transgene sequences encoding functional proteins Including, (iii) when localized in a cell depleted of a splicing factor, splicing of the first splice acceptor site and the first donor site is not inhibited, the cryptic exon sequence is present in the mRNA product of the construct, and a functional protein is produced from the mRNA product of the construct; (iv) when localized in a cell that is not depleted of splicing factors, splicing of the first splice acceptor site and / or the first donor site is repressed, the cryptic exon sequence is absent in the mRNA product of the construct, and no functional protein is produced from the mRNA product of the construct; Also provided is a construct configured to:
[0023] Also, Start codon, A regulatory domain that includes: a first splice donor site and a first acceptor donor site defining a single regulatory intron; a binding domain for a splicing factor of the heterogeneous nuclear ribonucleoprotein (hnRNP) family located within 150 nucleotides of the first splice donor site and / or the first splice acceptor site and / or located between the first splice donor site and the first splice acceptor site; and Transgene sequences encoding functional proteins Including, (iii) when localized in cells depleted of splicing factors, splicing of the first splice acceptor site and the first splice donor site is not inhibited, the single regulatory intron is spliced, and a functional protein is produced from the mRNA product of the construct; (iv) when localized in a cell that is not depleted of splicing factors, splicing of the first splice acceptor site and / or the first donor site is repressed and the single regulatory intron is not spliced or is not properly spliced, such that no functional protein is produced from the mRNA product of the construct. Also provided is a construct configured to:
[0024] In addition, in this specification, (i) when the cell nucleus is depleted of RNP family splicing factors (i.e., in diseased cells), the system produces a functional protein from the mRNA product of the construct; (ii) in the absence of depletion of hnRNP family splicing factors from the cell nucleus (i.e., in healthy cells), the system does not produce functional protein from the mRNA product of the construct; Also provided is a system comprising any of the constructs and cells described herein, or any of the vectors and cells described herein.
[0025] Also provided herein as a further aspect or embodiment of the first and second aspects is a method comprising a transgene sequence and a regulatory domain, the regulatory domain being (from upstream to downstream): The exon immediately upstream of the splice donor site, a splice donor site (i.e., a second splice donor site), A first portion of the intron region, a splice acceptor site (i.e., the first splice acceptor site), a cryptic exon sequence (i.e., embedded within an intron region between the first splice acceptor site and the first splice donor site); a splice donor site (i.e., the first splice donor site), a second portion of the intron region, and a splice acceptor site (i.e., a second splice acceptor site), and The exon immediately downstream of the splice acceptor site Including, Also disclosed are constructs, wherein the regulatory domain comprises a binding site for a splicing factor of the hnRNP family within a first portion of the intron region, a cryptic exon sequence, and / or a second portion of the intron region.
[0026] The splicing factor is preferably TDP-43. In some embodiments, the transgene sequence is completely downstream of the regulatory domain. In some embodiments, the transgene sequence is at least partially encoded by a cryptic exon sequence, optionally encoded by an exon immediately upstream of the splice donor site and / or an exon immediately downstream of the splice acceptor site.
[0027] Also provided herein as a further aspect or embodiment of the first and second aspects are (from upstream to downstream): exon sequence (i.e., immediately upstream of the splice donor site); a splice donor site (i.e., a second splice donor site), a first portion of the intron region (i.e., or the first intron); a splice acceptor site (i.e., the first splice acceptor site), a cryptic exon sequence (i.e., embedded within an intron region between the first splice acceptor site and the first splice donor site); a splice donor site (i.e., the first splice donor site), A second portion of the intron region, a splice acceptor site (i.e., a second splice acceptor site), and the exon sequence immediately downstream of the splice acceptor site, an optional proteolytic or autocleavage site; and the transgene sequence (i.e., the complete transgene sequence). Also disclosed is a construct comprising: The construct comprises a binding domain for a splicing factor of the hnRNP family within a first portion of an intron region, a cryptic exon sequence, and / or a second portion of an intron region.
[0028] The first splice acceptor site and the first splice donor site are repressed by splicing factors of the hnRNP family.
[0029] This construct may be described herein as the "Design 1" construct. The splicing factor is preferably TDP-43. In one embodiment, the start codon may be present in the exon sequence upstream of the cryptic exon, the cryptic exon being of a length not divisible by 3 to introduce a frameshift, and the construct is configured such that the start codon is in-frame with the downstream transgene sequence only if the cryptic exon is included. Alternatively, in another embodiment, the start codon (i.e., required for expression of the transgene) may be present within the cryptic exon itself.
[0030] Also provided herein as a further aspect or embodiment of the first and second aspects is a transgene comprising a transgene sequence (i.e. a transgene sequence encoding a functional protein) and a regulatory domain, said transgene sequence being arranged (from upstream to downstream) as follows: exon sequences (i.e., immediately upstream of the splice donor site and optionally encoding a portion of the transgene sequence); a splice donor site (i.e., a second splice donor site), a first portion of the intron region (i.e., the first intron); a splice acceptor site (i.e., the first splice acceptor site), a cryptic exon sequence (i.e., embedded within the intron region between the first splice acceptor site and the first splice donor site) encoding at least a portion of the transgene; a splice donor site (i.e., the first splice donor site), a second portion of the intron region (i.e., a second intron); and a splice acceptor site (i.e., a second splice acceptor site), and Also disclosed are constructs comprising an exon sequence (i.e., immediately downstream of the splice acceptor site and optionally encoding a portion of a transgene), the construct comprising a binding domain for a splicing factor of the hnRNP family within a first portion of an intron region, a cryptic exon sequence, and / or a second portion of an intron region.
[0031] The first splice acceptor site and the first splice donor site are repressed by splicing factors of the hnRNP family.
[0032] This construct may be described herein as the construct of Design 2. The splicing factor is preferably TDP-43.
[0033] Also provided herein as a further aspect or embodiment of the first and third aspects is a transgene comprising a transgene sequence (i.e. a transgene sequence encoding a functional protein) and a regulatory domain, the regulatory domain being (from upstream to downstream): exon sequence (i.e., immediately upstream of the splice donor site); a splice donor site (i.e., the first splice donor site), A single regulatory intron, a splice acceptor site (i.e., the first splice acceptor site), and Exon sequence (i.e., immediately downstream of the splice acceptor site) Also disclosed is a construct comprising: This regulatory domain contains binding domains for splicing factors of the hnRNP family located within exonic sequences upstream of the splice donor site, a single regulatory intron and / or exonic sequences downstream of the splice acceptor site.
[0034] The splicing factor is preferably TDP-43. In some embodiments, the transgene sequence is entirely downstream of the regulatory domain. In some embodiments, the transgene sequence is encoded by the exon sequence immediately upstream of the splice donor site and the exon immediately downstream of the splice acceptor site.
[0035] The first splice acceptor site and the first splice donor site are repressed by splicing factors of the hnRNP family. In some embodiments, the construct further comprises an alternative splice donor site and / or an alternative splice acceptor site that is not repressed by hnRNP splicing factors. The alternative splice acceptor site may be within a single regulatory intron or downstream of the first splice acceptor site. The alternative splice donor site may be within a single regulatory intron or upstream of the first splice donor site.
[0036] Aspects or embodiments of the invention have one or more of the following advantages.
[0037] Aspects of the present invention provide a mechanism for selectively expressing transgenic proteins in diseased cells associated with depletion of hnRNP splicing factors. This has the immense therapeutic advantage that therapeutic proteins such as chaperones, nuclear transport receptors, or gene editing enzymes such as Cas9 nuclease can be specifically expressed with improved safety and efficacy in cells depleted of hnRNP splicing factors (e.g., diseased cells depleted of TDP-43), while being minimally, lowly, or not expressed at all in healthy cells. Furthermore, the constructs and systems can also be used to express diagnostic proteins such as secreted luciferase that can be used to aid in the detection of patients with cells depleted of hnRNP splicing factors, e.g., cells with TDP-43 pathology.
[0038] The constructs and systems of the present invention also have utility in enabling preventative treatments, whereby at-risk patients are treated before pathology is detectable. Importantly, the constructs and systems are activated only when pathology (e.g., significant TDP-43 pathology in neurons) occurs, and are automatically inactivated when the pathology resolves intracellularly.
[0039] Thus, the present invention provides an improved tool for specifically targeting diseased cells with hnRNP depletion, which can be used as a therapy for neurodegenerative diseases. The constructs, vectors and pharmaceutical compositions described herein are designed to express proteins only in diseased cells, so there is no need to selectively administer the construct to a specific cell type. This means that more general and less invasive administration methods can be used.
[0040] In all of the above aspects and embodiments described herein, the binding domain can be to TDP-43, where the splicing factor of the hnRNP family is TDP-43, which is useful for studying, detecting, and treating cells with TDP-43 pathology, which is involved in many neurodegenerative and muscular diseases.
[0041] In all of the above aspects and embodiments described herein, a transgene sequence and a coding sequence (CDS) can be distinguished. By convention, a CDS refers to the entire sequence between a start codon and a stop codon. A CDS may code for an amino acid sequence that has no obvious protein function (in addition to encoding a functional protein) as a result of the regulatory nucleotide sequences used, and may be separated from the functional protein by a cleavage site. In contrast, a transgene sequence as defined herein is defined as a region of a CDS that codes for a functional protein (i.e., a protein that one wishes to express in diseased cells). For example, in a construct of "Design 1", the start codon may be within or upstream of a cryptic exon, and in such a case, the CDS includes at least a portion of the cryptic exon, but the transgene region of the CDS (i.e., the region of the CDS that codes for a functional protein) may be entirely downstream of the cryptic exon.
[0042] In all of the above aspects and embodiments described herein, the transgene sequence may code for a therapeutic protein. Thus, the construct can be used to code for a protein that is missing or abnormal in diseased cells. In some embodiments, the transgene sequence may code for a regulatory protein. A regulatory protein is a protein that alters the expression of an additional transgene or endogenous gene. Thus, the construct can be used to regulate the expression of an additional gene.
[0043] In all of the above aspects and embodiments described herein, the transgene sequence may encode a diagnostic protein. The constructs can be used to further understand, probe and diagnose cells that are depleted of hnRNP splicing factors.
[0044] In all of the above aspects and embodiments described herein, depletion of a member of the hnRNP family of splicing factors activates the expression of the transgene (i.e., results in the expression of the protein product encoded by the transgene). Here, the term "depletion" can refer in some embodiments to the general depletion of the splicing factor from the cell (i.e., "knockdown" via, for example, expression of an shRNA targeting the splicing factor's mRNA) and / or in some embodiments to the depletion of the splicing factor in the nucleus (i.e., occurs when TDP-43 aggregates in the cytoplasm). Since splicing occurs primarily in the nucleus, in both cases the lack of splicing regulation conferred by the splicing factor occurs, and thus in both cases activates the expression of a functional protein from the constructs described herein.
[0045] In all of the above aspects and embodiments described herein, the sequence defined by the first splice acceptor site and the first splice donor site can be a frameshift-inducing sequence. Whether splicing occurs (i.e., diseased cells) or is suppressed (i.e., healthy cells) determines whether the frameshift-inducing sequence is incorporated into the mRNA product of the construct to introduce a frameshift relative to the start codon. In such an embodiment, the construct may further comprise a premature termination codon (PTC) downstream of the regulatory domain, and the construct is configured such that (i) in cells with nuclear depletion of hnRNP splicing factors, the PTC is out-of-frame with the start codon in the mRNA product of the construct, and (ii) in cells without nuclear depletion of hnRNP splicing factors, the PTC is in-frame with the start codon in the mRNA product of the construct. This results in the formation of a truncated protein in cells not depleted of hnRNP splicing factors (i.e., healthy cells), whereas functional proteins are produced in cells depleted of hnRNP splicing factors, thereby providing a method for selectively expressing proteins in cells depleted of hnRNPs. In such an embodiment, the construct may further comprise an additional intron sequence (i.e., in an exon sequence context), with the PTC being at least 40 nucleotides upstream of the additional intron sequence. The splicing of the additional intron sequence promotes the attachment of the exon junction complex (EJC) to the mRNA product, which induces nonsense-mediated decay of the mRNA when it encounters the PTC. In contrast, if the PTC is not encountered, nonsense-mediated decay does not occur. Thus, the presence of the additional intron sequence further improves the safety of the construct (since peptides (e.g., truncated peptides) produced in healthy cells may otherwise accumulate, aggregate, or become toxic).
[0046] In some embodiments of the above aspect, the sequence between the first acceptor splice site and the first donor splice site is a cryptic exon sequence, and the regulatory domain further comprises an intron region (i.e., defined by a second splice donor site and a second splice acceptor site), and the cryptic exon sequence is located within said intron region. Thus, in such a construct, the regulatory domain is regulated by cryptic splicing, and the construct is configured such that the cryptic exon sequence is incorporated into the construct's mRNA product in diseased cells (i.e., where there is nuclear depletion of hnRNP splicing factors), but is absent from the construct's mRNA product in healthy cells (i.e., where there is no nuclear depletion of hnRNP splicing factors). In some embodiments, the cryptic exon sequence may be a frameshift-inducing cryptic exon sequence, thereby regulating expression of the transgene as above. Additionally or alternatively, the cryptic exon sequence may encode a portion of the transgene. This means that the complete transgene sequence is only fully present in the mature mRNA and capable of producing a functional protein if the cryptic exon is integrated into the mRNA product of the construct in diseased cells, but not if the cryptic exon is not integrated in healthy cells. In some embodiments and examples herein, the intron region is derived from the human AARS1 intron region between exon 4 and exon 5. In some embodiments and examples herein, the intron region is a synthetic sequence that is not derived from a natural intron sequence.
[0047] Constructs containing the cryptic exon sequences described herein may have the design "Design 1" or "Design 2" described herein, as shown in Figure 1 or Figure 2, respectively. In the Design 1 construct, the transgene sequence is completely downstream of the regulatory domain. One advantage of this design is that it can be easily modified on the basis to control the expression of a variety of different proteins by including different complete transgenes or protein coding sequences downstream of the regulatory sequence. Such embodiments may further include a proteolytic or autocleavage site between the regulatory domain and the transgene sequence. The presence of this site has the advantage of ensuring that the transgene is expressed without extra N-terminal sequence, which in some cases can improve the functionality of the protein product of the transgene.
[0048] In the Design 2 constructs, the cryptic exon sequence encodes at least a portion of the transgene sequence. This may be an N-terminal, internal, or C-terminal portion of the transgene sequence. The Design 2 constructs also have many advantages. Compared to the Design 1 constructs, the construct sequence may be smaller. Furthermore, unlike Design 1, where unwanted peptides are produced in diseased cells from upstream regulatory regions, which may be N-terminal sequences or short free peptides attached to the transgene protein product, the Design 2 constructs do not produce unwanted peptides. Finally, unlike the Design 1 constructs, the full-length transgene sequence is present in the mature mRNA only if the cryptic exon is included, reducing the chance of "leaky expression" of the full-length protein in healthy cells (i.e., cells where the cryptic exon is not expressed), for example via leaky scanning.
[0049] In some embodiments of the above aspects, the first splice donor site is upstream of the first acceptor site, and the first splice donor site and the first splice acceptor site define a single regulatory intron. Constructs comprising a single regulatory intron sequence as described herein may be according to "Design 3" as illustrated by FIG. 3. The construct is configured such that in cells depleted of splicing factors, the single regulatory intron is spliced, and in cells not depleted of splicing factors, the single regulatory intron is (i) not spliced or (ii) not properly spliced. This has the effect that only in cells depleted of splicing factors (e.g., not containing TDP-43) will the intron be properly spliced and the start codon be in frame with the uninterrupted coding transgene sequence of the protein to be expressed. In some embodiments, the transgene sequence is entirely downstream of the regulatory domain and / or the single regulatory intron. In another embodiment, the transgene sequence may be encoded by exon sequences upstream and downstream of a single regulatory intron.
[0050] In some embodiments, multiple regulatory domains may be included in the same construct. For example, the transgene sequence may be split into multiple cryptoexons (i.e., may contain multiple regulatory domains in a "design 2" manner) or may feature multiple regulatory introns (i.e., may contain multiple regulatory domains in a "design 3" manner). In some embodiments, regulatory domains in designs 1, 2, and / or 3 may be present in the same vector. The use of multiple control domains may reduce the risk of leaky expression and improve safety.
[0051] BRIEF DESCRIPTION OF THE DRAWINGS The following disclosure is illustrated with reference to the following non-limiting examples and drawings. [Brief description of the drawings]
[0052] [Figure 1]FIG. 1 shows an example of a construct of the present invention of Design 1. The construct is designed such that inhibition of the first splice acceptor site (2) and the first splice donor site (3) by binding to the binding domain (4) of a splicing factor of the hnRNP family results in inhibition of splicing of the first splice acceptor site and / or the first splice donor site. Thus, in healthy cells, the cryptic exon is not included in the mRNA product. In diseased cells, splicing is not inhibited and the mRNA product of the construct in diseased cells contains the cryptic exon. Inclusion or exclusion of the cryptic exon sequence can control the expression of the transgene sequence (5). The construct shown in FIG. 1 includes an initiation codon (1) and a cryptic exon sequence (CE) defined by the first splice acceptor site (2) and the first splice donor site (3). The construct comprises a binding domain (4) for an hnRNP splicing factor, which regulates splicing of the first acceptor site (2) and / or the first splice donor site (3). The cryptic exon sequence (CE) is embedded within an intron region (6) defined by a splice donor site (7) and a splice acceptor site (8). A first portion of the intron region is upstream of the cryptic exon sequence, and a second portion of the intron region is downstream of the cryptic exon sequence. The exon sequence (12) is further flanked by an intron region. The transgene sequence (5) is completely downstream of the regulatory domain and the cryptic exon sequence (CE). The transgene sequence comprises a stop codon (10) at the end of the sequence. An optional cleavage site (9) may be between the cryptic exon sequence and the transgene sequence (5). Optionally, the transgene sequence (5) further comprises a premature termination codon (PTC) at least in the middle of the sequence. Optionally, downstream of the transgene sequence, there is an additional intron sequence (11) in an exon context. Optionally, the cryptic exon sequence is a frameshift cryptic exon sequence.In healthy cells where hnRNP splicing factors are not depleted, splicing of the cryptic exon is suppressed by binding of the splicing factor and binding domains. The complete intron region (6) containing the cryptic exon sequence (CE) is spliced out (i.e., between 7 and 8), and therefore the mRNA product of the construct does not contain a cryptic exon. In this example without a cryptic exon, a premature stop codon (PTC) is in frame with the start codon (1), forming a truncated protein. Furthermore, as a result of further intron sequences downstream of the transgene, the exon junction complex (EJC) attaches to the mRNA product of the construct, which induces nonsense-mediated decay of the mRNA. On the other hand, in diseased cells where hnRNP splicing factors are depleted, splicing of the cryptic exon is not suppressed. The first part of the intron region is spliced out (i.e., between 7 and 2) and the second part of the intron region is also spliced out (i.e., between 3 and 8), so that the cryptic exon sequence (i.e., between 2 and 3) is included in the mRNA (i.e., the mature mRNA). In this example, this introduces a frameshift, so that the PTC is out of frame with the start codon and the transgene can be fully translated and a functional protein produced. The cleavage site (9) releases the transgene's protein separately from the peptide produced from the exon sequence (12) adjacent to the intron region (6). Since the PTC is absent in diseased cells, the ribosome removes the exon junction complex (EJC), which means that NMD does not occur. In another embodiment (not shown), the cryptic exon itself may contain the start codon. [Diagram 2]FIG. 2 shows an example of a construct of the invention, Design 2. As in Design 1, the mRNA product of diseased cells contains a cryptic exon, but the repression of the first splice acceptor site (2) and the first splice donor site (3) caused by binding of the binding domain (4) to a splicing factor of the hnRNP family means that the mRNA product of healthy cells does not contain a cryptic exon. However, unlike Design 1, the (CE) sequence itself encodes a portion of the transgene sequence (5). This construct contains a cryptic exon sequence (CE) defined by a start codon (1) and a first splice acceptor site (2) and a first splice donor site (3). This construct also contains a binding domain (4) for an hnRNP splicing factor, which regulates splicing of the first splice acceptor site (2) and / or the first splice donor site (3). The cryptic exon sequence (CE) is also embedded within an intron region (6) defined by a splice donor site (7) and a splice acceptor site (8). In this example, a first portion of the transgene sequence is encoded by exon sequence upstream of the intron region, a second portion of the transgene is a cryptic exon sequence, and a third portion of the transgene is encoded by exon sequence downstream of the cryptic exon sequence. The portion of the transgene downstream of the cryptic exon sequence optionally also contains a premature termination codon (PTC) at least in the middle of the sequence, and the CE is a frameshifted CE sequence. Optionally, downstream of the transgene sequence is an additional intron sequence (11) in an exonic context. As in design 1, in healthy cells that are not depleted of hnRNP splicing factors, splicing of the cryptic exon is suppressed and the mRNA product of the construct does not contain the cryptic exon. This means that the mature mRNA product of the construct in healthy cells does not contain the full-length sequence encoding the protein to be expressed. In contrast, diseased cells express the mature mRNA that encodes the entire transgene protein product.Furthermore, in this example, similar to design 1, due to the frameshift CE sequence, the premature stop codon (PTC) is in frame with the start codon (1) in the mRNA product of the construct in healthy cells, but not in diseased cells. Due to the presence of additional intronic sequences, the exon junction complex (EJC) triggers nonsense-mediated decay of the mRNA product in healthy cells, but in diseased cells, the ribosome removes the EJC and nonsense-mediated decay does not occur. In another embodiment (not shown), the cryptic exon may instead code for the N- or C-terminal region of the protein product. Additionally or alternatively, the PTC does not need to be present in the transgene downstream of the regulatory domain (not shown), since the absence of the cryptic exon in the mRNA product of the construct may result in the production of a nonfunctional protein product. [Diagram 3]FIG. 3 shows an example of a construct of the invention of design 3. In healthy cells, the construct is designed such that inhibition of the first splice donor site (3) and the first splice acceptor site (2) resulting from binding of a splicing factor of the hnRNP family with the binding domain (4) results in inhibition of splicing of the single regulatory intron, which is not spliced or is not properly spliced. In diseased cells, splicing is not inhibited and the mRNA product of the construct does not contain any part of the single regulatory intron. In this example, the construct contains a start codon (1) and a single regulatory intron sequence (intron) defined by the first splice donor site (3) and the first splice acceptor site (2). The construct contains a binding domain for an hnRNP splicing factor (4), which regulates splicing of the first splice donor site (3) and / or the first splice acceptor site (2). In this example, the transgene sequence is encoded in two parts (5) by exon sequences both upstream and downstream of a single regulatory intron, although in another embodiment (not shown) the transgene (5) is entirely downstream of a single regulatory intron. The construct may further comprise an alternative splice acceptor site and / or an alternative splice donor site (not shown). The alternative splice acceptor site and / or alternative splice donor site may also be referred to herein as a "decoy" splice site. In some embodiments, the alternative splice site is configured to be preferentially spliced in cells that are not depleted of hnRNP family splicing factors (e.g., TDP-43) and the first donor and / or acceptor splice site is spliced (i.e., the decoy splice site is not used) in cells that are depleted of hnRNP family splicing factors (e.g., TDP-43). In some embodiments, the alternative splice site is a donor splice site located upstream of the first donor splice site.In some embodiments, the alternative splice site is a donor splice site or an acceptor splice site located between the first donor splice site and the first acceptor splice site. In some embodiments, the alternative splice site is an acceptor splice site located downstream of the first acceptor splice site. As described for the constructs of Design 1 and Design 2, the construct may further include one or more premature stop codons (PTCs), and the construct may optionally further include an additional intron sequence (11) downstream of the transgene (5). This promotes attachment of EJC and NMD to the mRNA product in healthy cells. In healthy cells, the intron is either fully retained (see, e.g., E) or partially retained (see, e.g., B or D) due to suppression of both splice sites (see, e.g., E), or is not properly spliced due to suppression of one splice site (see, e.g., A and C). This means that a non-functional protein is produced in healthy cells and a functional protein is produced in diseased cells. Optionally, a premature stop codon (PTC) is present in the portion of the transgene sequence (5) downstream of the single regulatory intron sequence. In certain embodiments, for example, where either intron retention or improper splicing introduces a frameshift (see, e.g., A, B, and C), the construct is configured such that the PTC is in frame with the start codon when at least a portion of the intron is included in the mRNA product of the construct, but is out of frame with the start codon when the intron is absent from the mRNA product of the construct. This further results in the formation of a truncated or non-functional protein in healthy cells and a functional protein in diseased cells. Optionally, in addition or instead, the PTC may be present in an intron that is in frame with the start codon, and where full or partial intron retention causes the PTC to be in frame with the start codon in the mRNA product of the construct (see, e.g., D and E). The combination of the PTC and attached EJC results in NMD, preventing the expression of truncated, non-functional proteins that would otherwise be toxic to the cell.The presence of a PTC in a construct, and thereby in the mRNA product of the construct in healthy cells (i.e., the mature mRNA product), is not an essential part of the invention, since retention or improper splicing of an intron may produce a non-functional protein product (e.g., due to internal truncation by improper splicing, or due to inclusion of a disruptive amino acid sequence that impairs folding). [Figure 4] FIG. 4A shows mCherry fluorescent signals from four cryptic exon-containing vectors. "AARS1-based reporter" corresponds to the Design 1 construct, Example 1A, featuring a frameshifted upstream AARS1-derived cryptic exon / intron regulatory sequence and downstream mCherry sequence. "Synthetic-1 / 2 / 3" corresponds to the Design 2 constructs, Examples 2A-2C, featuring computer-generated cryptic exon sequences and encoding internal portions of the mCherry sequence flanked by computer-generated intron sequences. Numbers indicate the ratio of signal in TDP-43 knockdown cells to control cells. FIG. 4B shows mScarlet fluorescent signals from cells co-transfected with an mScarlet-encoding plasmid containing a "poison exon" flanked by LoxP sites with a plasmid encoding Cre recombinase (wherein a portion of the Cre recombinase sequence is encoded by a synthetic cryptic exon) flanked by AARS1-derived intron sequences (i.e., the construct described in Example 3, another Design 2 construct). Numbers indicate the ratio of signals in TDP-43 knockdown cells to control cells. Y-axis values refer to "Scale Values" from Flow-Jo. [Diagram 5] Figure 5 shows the signal from secreted luciferase when using the example construct of Design 1, Example Construct 1B. "-ve control" refers to cells transfected with a vector encoding mCherry. [Figure 6]Figure 6 shows TDP-43-dependent genome editing. Figure 6A shows Western blots showing expression of FLAG-tagged Cas9, TDP-43 and α-tubulin with or without TDP-43 knockdown in cells transfected with Cas9 expression vectors containing a cryptic exon corresponding to Example 4, an example construct of Design 2 (left) or with a constitutive Cas9 expression vector (right). Figure 6B shows the percentage of Illumina reads with indels in the targeted CDK4 locus. "-ve control" = cells transfected with vector encoding mCherry. [Figure 7] Figure 7 shows suppression of cryptic exons and autoregulation. A: RT-PCR analysis of cells transfected with the INSR cryptic exon minigene and optionally co-transfected with a plasmid expressing a cryptic TDP-43-RAVER1 fusion protein (i.e. according to the example constructs of Design 1, Example 1C or Mutant 1C). The "mutant" protein is defective in RNA binding. Doxycycline induces TDP-43 knockdown. B: As described for A, except that the RT-PCR target is the AARS1-derived frameshift cryptic exon, thus demonstrating autoregulation of this construct. [Figure 8] Figure 8 shows results using a Cas9 / AARS1 mCherry reporter corresponding to Example 1D, an example construct of Design 1: mCherry fluorescence was assessed by fluorescence microscopy from cells transfected with a construct containing a downstream mCherry transgene regulated by an upstream frameshift cryptic exon; the cryptic exon is the novel sequence-encoding portion of S. pyogenes Cas9 flanked by AARS1-derived intronic regions. Left: cells not depleted of TDP-43; Right: cells depleted of TDP-43. [Figure 9]FIG. 9 shows the results of mCherry fluorescence assessed by fluorescence microscopy for SK-N-DZ cells transfected with the AARS1-mCherry-FLAG intron-retaining construct, a Design 3 construct corresponding to Example 5, with doxycycline-induced TDP-43 knockdown. [Figure 10] FIG. 10 shows the levels of STMN2 cryptic exons relative to TDP-43 protein levels. The percent inclusion (PSI) of STMN2 cryptic exons as assessed by RNA sequencing is shown relative to the remaining TDP-43 levels as assessed by Western blot (% TDP-43 protein remaining is shown on the x-axis). Because these cells show properly localized TDP-43, the total levels of TDP-43 protein are equivalent to the total levels of nuclear TDP-43. This indicates that the presence of STMN2 cryptic inclusions is indicative of TDP-43 depletion in the nucleus. This further indicates that relatively mild depletion of TDP-43 (e.g., 77% remaining with 23% depletion) can significantly increase the levels of cryptic splicing. [Figure 11] Figure 11 shows the distribution of Splice AI scores (log scale) as determined by the Splice AI algorithm in human transcripts of 500 genes, none of which were included in the original training set of the Splice AI algorithm. The dashed line corresponds to a cutoff of 0.01, which corresponds to approximately the 99.8 percentile rank of the splice sites. [Figure 12] 12 shows fluorescence microscopy images of SK-N-DZ cells transfected with a design type 1 mCherry construct reporter (Example 1A) or various synthetic design type 2 mScarlet construct reporters (Examples 2D-2J). Doxycycline induces TDP-43 knockdown. These images are inverted for clarity. [Figure 13]13 shows A) fluorescence microscopy images, B) fluorescence microscopy quantification, and C) nanopore sequencing of SK-N-DZ cells transfected with constitutively expressing mCherry vector or Example 1A, both with and without TDP-43 knockdown. Numbers above the bars indicate the log2 fold increase upon TDP-43 depletion. [Figure 14] Figure 14 shows A) fluorescence microscopy images and B) nanopore sequencing of SK-N-DZ cells transfected with various Design 2 constructs (i.e., Examples 2D-2J) encoding mScarlet, both with and without TDP-43 knockdown. Numbers above the bars indicate the log2 fold increase upon TDP-43 depletion. [Figure 15] 15 shows A) fluorescence microscopy images and B) nanopore sequencing of SK-N-DZ cells transfected with various Design 3 constructs (i.e., Examples 6A-6D) with or without TDP-43 knockdown. Numbers above the bars indicate the log2 fold increase upon TDP-43 depletion. [Figure 16] Figure 16 shows example nanopore traces obtained from SK-N-DZ cells transfected with a Design 2 construct (i.e., Example 2E) and various Design 3 constructs (i.e., Examples 6A, 6B, and 6D), each encoding mScarlet. Asterisks highlight the usage of cryptic splice sites. The expected splicing patterns are shown above, and for the Design 3 constructs, the "decoy" splice sites are also shown. [Figure 17] FIG. 17 shows RT-PCR analysis of SK-N-DZ cells transfected with the F2L mutants of the Design 2 constructs (i.e., Examples 7A and 7B) with or without TDP-43 knockdown. [Figure 18]FIG. 18 shows RT-PCR analysis of SK-N-DZ cells transfected with Design 2 constructs (i.e., Examples 7A and 7B) containing a functional TDP-43 sequence (i.e., no F2L mutation) with or without TDP-43 knockdown. [Figure 19] 19 shows A) RT-PCR of the cryptic exon region of the endogenously expressed UNC13A transcript for SK-N-DZ cells expressing the Design 2 constructs (i.e., Examples 7A and 7B), and B) quantification of the above RT-PCR for UNC13A and the equivalent RT-PCR for the ELAVL3 cryptic exon (not shown) with or without knockdown of endogenous TDP-43. For each sample, the left bar shows quantification for untreated cells and the right bar shows quantification for dox-treated (i.e., shTDP-43) cells. [Figure 20] FIG. 20 shows A) a diagram of the vectors (bottom) and control (top) of Example 8; B) RT-PCR analysis of splicing of the vectors in A with and without TDP-43 knockdown in SK-N-DZ cells with and without TDP-43 knockdown; and C) analysis of genome editing at expressed loci by nanopore amplicon sequencing in these cells. [Figure 21] FIG. 21 shows A) luciferase activity from culture medium of SK-N-DZ cells transfected with Example 9 constructs with or without TDP-43 knockdown, and B) nanopore tracing of these cells. [Figure 22] 22 shows A) a schematic diagram of the triple cryptic exon Cre-recombinase vector of Example 10 (exons 2, 4, and 6 are "cryptic") and B) quantification of nanopore reads for the number of cryptic exons contained in each transcript in SK-N-DZ cells without (NT) or with doxycycline-induced knockdown of TDP-43. Error bars indicate standard error of triplicates. [Diagram 23]Figure 23 shows air nanopore traces from i3 iPSCs expressing the triple cryptic exon Cre-recombinase vector of Example 10, with or without treatment with a Halotag-based "protac" sequence that depletes endogenous Halotag TDP-43. The predicted locations of the three cryptic exons are indicated by striped boxes. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0053] Detailed Description For the SEQ ID NOs disclosed herein, the complementary sequences of each SEQ ID NO are also disclosed. Also disclosed herein are constructs having sequences complementary to those described herein, which can be used to encode the constructs described herein.
[0054] As used herein, the term "treatment" refers to an approach for obtaining beneficial or desired results in a subject, and includes both prophylactic and therapeutic benefits.
[0055] "Therapeutic benefit" refers to the eradication, amelioration, or slowing of the progression of the underlying disease being treated. Therapeutic benefit is also achieved by eradication or amelioration of one or more physiological symptoms associated with the underlying disease such that an improvement is observed in the subject, even though the patient is still afflicted by the underlying disease.
[0056] "Prophylactic benefit" refers to delaying or eliminating the appearance of a disease or condition, delaying or eliminating the onset of symptoms of a disease or condition, slowing, halting or reversing the progression of a disease or condition, or a combination thereof. In the context of the present invention, a prophylactic benefit or effect can include prevention of a condition or disease. The construct, vector or pharmaceutical composition can be administered to a subject at risk of developing a particular disease, or to a subject who complains of one or more physiological symptoms of a disease, even if the subject has not been diagnosed with the disease.
[0057] The term "subject" refers to any suitable subject, including animals, such as mammals. In preferred embodiments described herein, the subject is a human.
[0058] The term "comprising" (including related terms such as "comprise" or "comprises" or "having" or "including") includes embodiments, e.g., materials, compositions, methods, or steps of any composition, "consisting of" or "consisting essentially of" the described features, unless the context clearly indicates otherwise. The terms "comprises" or "comprising" can be used interchangeably with "includes".
[0059] The term "RNA-seq" as referred to herein, also known as "RNA sequencing," refers to a next-generation sequencing technology that reveals the presence and amount of RNA in a sample that can be used to analyze the cellular transcriptome.
[0060] "Construct" as described herein has its usual meaning in the art and refers to a synthetic nucleic acid sequence that contains genetic material that codes for a gene of interest. It is intended that the construct is not a complete nucleic acid sequence in nature, i.e., a nucleic acid sequence as found in the genome of an organism (the construct itself may contain components derived from a naturally occurring sequence). The construct may have a maximum length, i.e., the construct may contain less than 50,000 nucleotides, or less than 40,000 nucleotides, or less than 30,000 nucleotides, or less than 20,000 nucleotides, or in some examples, less than 10,000 nucleotides, or less than 5000 nucleotides, or less than 2500 nucleotides.
[0061] "Vector" has its usual meaning in the art and refers to a synthetic piece of nucleic acid that contains a construct (i.e., as defined above) and has the function of delivering the construct into a cell.
[0062] As used herein, a "nucleotide" refers to a component part of a nucleic acid sequence. A nucleotide comprises a nucleobase (e.g., A, G, T, and C in DNA, or A, G, U, and C in RNA, although other nucleobases can be used) linked to a sugar (e.g., deoxyribose in DNA, ribose in RNA, although other sugars can be used). In DNA and RNA, the sugars are linked by a phosphodiester backbone to form the nucleic acid sequence, although other backbones can be used.
[0063] As described herein, "nuclear depletion of splicing factors" can be defined as cells that have at least 20% loss, or at least 25% loss, or preferably at least 50% loss of splicing factors in the nucleus of the cells (or as an average of a population of cells) compared to the same type of healthy cells (or as an average of a population of healthy cells). Splicing factor depletion can be determined by standard methods such as Western blotting. In some instances, the term "nuclear depletion of splicing factors" can be replaced or interchangeable with the term "absence of binding of splicing factors to splicing factor binding domains," and the term "no nuclear depletion of splicing factors" can be replaced or interchangeable with the term "presence of binding of splicing factors to splicing factor binding domains." When the splicing factor is TDP-43, nuclear depletion can be determined by determining the presence of STMN2 cryptic splicing events (i.e., the presence of STMN2 cryptic exons) in the cell transcripts, which can be determined by RNA sequencing. This is because the presence of STMN2 cryptic exons in the mRNA transcript indicates nuclear depletion of TDP-43 (see FIG. 10). Depletion of TDP-43 refers to depletion of "normal" or wild-type TDP-43 and may not include pathological or mutant TDP-43. Pathological TDP-43 may be hyperphosphorylated, ubiquitinated, or truncated TDP-43, TDP-43 with reduced solubility, or misfolded TDP-43, mutant TDP-43, or TDP-43 with altered subcellular location.
[0064] Cells that have nuclear depletion of hnRNP family splicing factors may be referred to herein as "diseased cells." Cells that do not have nuclear depletion of hnRNP family splicing factors may be referred to herein as "healthy cells."
[0065] The splicing factor described herein is intended to refer to hnRNP family splicing factor or splicing repressor protein. As defined herein, hnRNP refers to heterogeneous nuclear ribonucleoprotein, and includes TDP-43 as a family member. The term hnRNP splicing factor can be used interchangeably with the term hnRNP splicing repressor protein. The term hnRNP family splicing factor can also be used interchangeably with the term hnRNP splicing factor.
[0066] As defined herein, TDP-43 refers to TAR DNA binding protein 43 (Transactive response DNA binding protein 43 kDa), a protein that in humans is encoded by the TARDBP gene. TDP-43 binds to both DNA and RNA and has been shown to have multiple functions in transcriptional repression, pre-mRNA splicing, and translational regulation, among other functions.
[0067] Splicing, as defined herein, refers to the process by which pre-mRNA is converted into mature mRNA, introns are removed and exons are joined.
[0068] As used herein, synonymous codons refer to different codons that code for the same amino acid.
[0069] As defined herein, "in frame" refers to a situation where the number of nucleotides between codons is divisible by 3. "Out of frame" refers to a situation where the number of nucleotides between codons is not divisible by 3.
[0070] A cryptic exon, as defined herein, refers to a splice variant that is incorporated into a mature mRNA (i.e., upon depletion of the associated splicing factor) and introduces, among other changes, a frameshift or a stop codon into the resulting mRNA. In other words, a cryptic exon is a nucleotide sequence that is preferentially spliced out upon depletion of the associated splicing factor. Cryptic exons are also sometimes referred to herein or elsewhere in the art as "CEs," "cryptics," "cryptic exon sequences," or "cryptic events."
[0071] A single regulatory intron, as defined herein, refers to a splice variant that is at least partially incorporated into a mature mRNA (i.e., by alternative splicing of the intron) and introduces, among other changes, a frameshift or a stop codon into the resulting mRNA. In other words, a single regulatory intron is a nucleotide sequence that is differentially spliced in cells depleted of the relevant splicing factor, and alternative splicing of this intron can introduce, among other changes, a frameshift or a stop codon into the resulting mRNA.
[0072] Sequence complementarity, as disclosed herein, refers to Watson-Crick base pairing in nucleic acids, e.g., A binds to T (or U or modified variants thereof) and C binds to G (or modified variants thereof).
[0073] All genomic or chromosomal locations mentioned herein refer to locations on the human genome and associated transcriptome (hg38).
[0074] When a range is used herein, it is intended to include all combinations and subcombinations of the range and specific embodiments. The term "about" or "~" when referring to a numerical value or numerical range means that the numerical value or numerical range referred to is an approximation within experimental variation (or within statistical experimental error), and thus the numerical value or numerical range may vary. Typical experimental variations may be due to, for example, changes and adjustments required during laboratory experiments and scale-up from manufacturing settings to large scale.
[0075] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include the plural forms unless the context clearly dictates otherwise.
[0076] The binding domain for a splicing factor described herein refers to the sequence encoding the binding domain in the mRNA. For example, when referring to a TG or UG-rich motif for a TDP-43 binding domain, the TG-rich motif is present in the DNA construct and the UG-rich motif is present in the RNA.
[0077] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Abbreviations used herein have their conventional meaning in the chemical and biological arts unless otherwise noted.
[0078] The splice score described herein refers to the splice score determined by the Splice AI algorithm. The splice score determined by the Splice AI algorithm is determined by calculating the probability of splicing at a given position given a particular sequence context. The sequence adjacent to the splice site may include the entire construct (i.e., from start to end, or in a vector context, from the end of the promoter to the start of the polyadenylation signal), since the sequence of the flanking region (e.g., up to a 10,000 nucleotide interval) may affect the splicing prediction at a given position. The Splice AI algorithm can be found at the following link https: / / github.com / Illumina / SpliceAI and can be used according to the description in Jaganathan et al., 2019, Cell, 176, 535-548, “Predicting Splicing from Primary Sequence with Deep Learning”, the contents of which are incorporated herein by reference. The version of the Splice AI algorithm and pre-trained network weights used may be version 1.3.1. A score of 0.01 is the 99.8th percentile of scores generated by the Splice AI algorithm (see Figure 11), which corresponds to a very high probability of splicing (i.e., compared to a random location in the genome), whereas the majority of true native splice sites obtain scores well below 1, as described in the Jaganathan et al. reference and shown in Figure 11. In particular, splice sites that are alternatively spliced in different tissues (e.g., constitutively spliced in neurons but not in hepatocytes) usually receive low Splice AI scores, even though they act as strong splice sites in certain cell types.
[0079] A splice site, as understood in the art, is the boundary between an intron sequence and an exon sequence. During splicing, a nucleotide sequence is cleaved at the splice site, i.e., the nucleotide sequence is cleaved at the boundary between the intron sequence and the exon sequence.
[0080] A splice acceptor site is a splice site that occurs between an intron and an exon, i.e., immediately upstream of the exon sequence, where the intron is upstream of the exon sequence. A splice acceptor site is characterized by a splice site that contains the dinucleotide "AG" upstream of the splice site (i.e., at the end of the intron sequence upstream of the exon).
[0081] A splice donor site is a splicing site that occurs between an exon and an intron, i.e., an exonic sequence where the exon is upstream of the intron. A splice donor site is characterized by a splice site that contains the dinucleotide "GT" downstream of the splice site (i.e., at the start of the intron sequence downstream of the exon).
[0082] Splicing factors are proteins involved in splicing, the removal of introns from mRNA and the joining of their exons.
[0083] It is envisaged that any embodiment described herein can be combined with any other embodiment described herein, unless the context explicitly indicates otherwise. For example, the embodiments described for the hnRNP binding domain, or more specifically the TDP-43 binding domain, can be readily combined with other embodiments described herein, including but not limited to construct designs (e.g., designs 1, 2, or 3), cryptic exon sequences (if any), single regulatory introns (if any), first splice acceptor sites, first splice donor sites, PTCs, further intron sequences, intron regions (if any), etc. Similarly, features of any dependent claim can be readily combined with features of any of the independent claims or other dependent claims, unless the context clearly indicates otherwise.
[0084] As described herein, whether or not a functional protein is produced from a transgene sequence refers to whether or not a functional protein is produced from the mRNA product of the construct.
[0085] Constructs The constructs described herein are synthetic nucleotide sequences. In some embodiments, the constructs preferably comprise DNA nucleotide sequences. The constructs may comprise double-stranded or single-stranded DNA. In some embodiments, the constructs comprise linear or circular DNA. The nucleotides may comprise or be formed from unmodified nucleobases (e.g., C, T, A, or G in DNA), but may also comprise modified nucleobases (e.g., but not limited to, 5-methylcytosine, 6-methyladenosine, deoxyuridine), so long as they do not interfere with Watson-Crick base pairing, transcription, and splicing. Although DNA nucleotide sequences are preferred, any other suitable nucleotide sequence may be used, i.e., may comprise nucleotides with different sugars or different backbones, so long as they do not interfere with Watson-Crick base pairing, transcription, and splicing.
[0086] Regulatory domains A first splice acceptor site and a first splice donor site The regulatory domain comprises a first splice acceptor site and a first splice donor site.
[0087] In some embodiments, the sequence around the first splice acceptor site is HAG / N, where / represents the splice site, H=C, T or A, and N is C, T, A or G. In some embodiments or examples, the construct comprises a polypyrimidine tract upstream of the first splice acceptor site (i.e., upstream of the splice acceptor site, e.g., within the intron region upstream of HAG / N). In some embodiments, the polypyrimidine tract is upstream of the first splice acceptor site, more preferably up to 40 nucleotides upstream of the first splice acceptor site, or up to 20 nucleotides upstream of the first splice acceptor site. A polypyrimidine tract as defined herein can be described as a pyrimidine-rich region, defined as a 20 nucleotide region containing at least 70% pyrimidines, or defined as a 30 nucleotide region containing at least 80% pyrimidines.
[0088] In some embodiments, the regulatory domain further comprises a branch site comprising an adenosine upstream of the first splice acceptor site and the polypyrimidine tract (i.e., within the intron region upstream of the splice acceptor site). A P, where N is any nucleotide, P is a pyrimidine (i.e., C or T), and the underlined A is, for example, a branch point (e.g., CTG). A C). The branch site may be located up to 45 nucleotides upstream of the first splice acceptor, preferably up to 35 nucleotides upstream of the first splice acceptor, preferably 20-35 nucleotides upstream of the first splice acceptor.
[0089] In some embodiments, the sequence before and after the first splice donor site is N / GT, where / represents the splice site and N is C, T, A, or G. In some examples described herein, the sequence before and after the first donor splice site is CAG / GT, where / represents the splice site.
[0090] In some embodiments, the first splice acceptor site and / or the first splice donor site has a splice score of 0.01 or greater as determined by the Splice AI algorithm. In some embodiments, the first splice acceptor site and / or the first splice donor site has a splice score of 0.05 or greater, or at least 0.1 or greater, or at least 0.2 or greater, or at least 0.3 or greater, or at least 0.4 or greater, or at least 0.5 or greater, or at least 0.6 or greater, or at least 0.7 or greater, or at least 0.8 or greater, or at least 0.9 or greater as determined by the Splice AI algorithm.
[0091] The first splice acceptor site and the first splice donor site define a sequence. In some embodiments, the sequence is a frameshift-inducing sequence, i.e., a sequence that includes a number of nucleotides that is not divisible by 3. Thus, splicing results in the introduction of a frameshift-inducing sequence in the mRNA product of the construct compared to when no splicing occurs. In some embodiments, the construct further includes a premature termination codon (PTC) downstream of the regulatory region, configured such that (i) in cells with nuclear depletion of splicing factors, the PTC is out of frame with the start codon in the mRNA product of the construct, and (ii) in cells without nuclear depletion of splicing factors, the PTC is in frame with the start codon in the mRNA product of the construct. This may result in the formation of a truncated protein in cells without nuclear depletion of splicing factors, but selectively produces a functional protein in cells with nuclear depletion of splicing factors. In some embodiments, the construct includes an additional intron sequence at least 40 nucleotides downstream of the PTC. This additional intron sequence is in an exon context. If the presence of the additional intron sequence downstream of the PTC inhibits splicing of the first splice acceptor and / or the first splice donor (i.e., the PTC is in frame with the start codon), the exon junction complex (EJC) attaches to the resulting mRNA and promotes nonsense-mediated decay. If splicing is not inhibited, the PTC codon will not be in frame with the start codon in the mRNA product of the construct, and therefore the ribosome will remove the EJC and nonsense-mediated decay will not occur. The presence of the additional intron sequence increases the safety and selectivity of the construct.
[0092] In some embodiments or aspects, the first splice acceptor site is upstream of the first splice donor site, and the first splice acceptor site and the first splice donor site define a cryptic exon sequence. In some embodiments, the cryptic exon sequence is a frameshift-inducing cryptic exon sequence, which thus alters expression of the transgene as described above.
[0093] In further or alternative embodiments, the cryptic exon sequence encodes at least a portion of the transgene. Thus, suppression of splicing can produce a non-functional protein in cells without nuclear depletion of splicing factors. In further or alternative embodiments, the start codon is present within the cryptic exon sequence.
[0094] In some embodiments or aspects, the first splice donor site is upstream of the first splice acceptor site, and the first splice donor site and the first acceptor donor site define a single regulatory intron. Thus, when splicing is suppressed, the mRNA construct in cells without nuclear depletion of splicing factors contains at least a portion of the intron, resulting in frameshifting and blocking expression of the transgene as described above. Alternatively or additionally, full or partial intron retention may introduce a PTC into the sequence if a PTC is present in the intron itself. Alternatively or additionally, inappropriate splicing or retention of the intron (full or partial) may disrupt the function of the protein product through the introduction of a disruptive amino acid sequence or truncation of the amino acid sequence without the need for a PTC or frameshift. In contrast, splicing in the absence of depletion of hnRNP splicing factors results in the removal of the intron sequence in the mRNA product of the construct. This results in a completely encoded and / or uninterrupted transgene sequence leading to protein production in healthy cells.The above aspects and embodiments are described in further detail below.
[0095] In some embodiments, the construct comprises one single regulatory domain, however, the construct may comprise two or more, or three or more, or four or more regulatory domains as described herein. The presence of multiple regulatory domains may increase the selectivity of expression in diseased cells and / or minimize leaky expression in healthy cells. In some embodiments, the construct may comprise one, or at least two, or at least three, or at least four, or at least five, or at least six, or at least seven, or at least eight, or at least nine, or at least ten cryptic exons and / or regulatory introns.
[0096] Binding domain The regulatory domain comprises a binding domain for a splicing factor of the hnRNP family. The splicing factor of the hnRNP family may be referred to or limited to a splicing repressor protein of the hnRNP family. Such proteins generally have a structure that includes at least one (e.g., two) RNA recognition motifs flanked by N- and C-terminal regions. These proteins generally contain a nuclear localization sequence (NLS) that allows localization within the nucleus. In some embodiments, the splicing factor of the hnRNP family may have a molecular weight of 30 kDa to 120 kDa, more preferably 30 kDa to 50 kDa. In a preferred embodiment, the splicing factor is an endogenous splicing factor, i.e., a splicing factor derived from within the cell.
[0097] In some embodiments, the splicing factor may be any member of the hnRNP family that is associated with depletion in disease, for example, a neurodegenerative or muscular disease.
[0098] In some embodiments, the binding domain is within 150 nucleotides of the first splice acceptor site and / or the first splice donor site. In some embodiments, the binding domain is within 100 nucleotides of the first splice acceptor site and / or the first splice donor site, or within 50 nucleotides of the first splice acceptor site or the first splice donor site, or within 25 nucleotides of the first splice acceptor site or the first splice donor site, or within 10 nucleotides of the first splice acceptor site or the first splice donor site. Binding of the binding domain to a splicing factor of the hnRNP family inhibits the first splice acceptor site and / or the first splice donor site, thus modulating splicing (e.g., of the sequence between the first splice acceptor site and the first splice donor site). Additionally or alternatively, the binding domain may be between the first splice donor site and the first splice acceptor site (e.g., within a single regulatory intron sequence in a Design 3 construct or within a cryptic exon sequence in a Design 1 or 2 construct).
[0099] In some embodiments, the binding domain comprises at least 6 nucleotides, more preferably at least 10 nucleotides. In some embodiments, the binding domain is between 6 and 700 nucleotides, or between 6 and 150 nucleotides, or between 10 nucleotides and 150 nucleotides, or between 15 and 50 nucleotides, or between 6 and 45 nucleotides, or between 10 and 45 nucleotides, or between 10 and 20 nucleotides, and in some instances between 20 and 45 nucleotides.
[0100] In some embodiments, the binding domain is upstream of the first splice acceptor site and / or the first splice donor site. In some embodiments, the binding domain is downstream of the first splice donor site and / or the first splice acceptor site. In some embodiments, the binding domain is between the first splice acceptor site and the first splice donor site (i.e., within the sequence defined by the first splice acceptor site and the first splice donor site, in some embodiments within a cryptic exon sequence, or in other embodiments within a single regulatory intron). In embodiments where the construct comprises a cryptic exon defined by a first splice acceptor site and a first splice donor site (e.g., a Design 1 or Design 2 construct), the binding domain can be upstream of the cryptic exon (i.e., within a first portion of the intron region), downstream of the cryptic exon (i.e., within a second portion of the intron region), or within the cryptic exon sequence. In embodiments where the construct comprises a single regulatory intron defined by a first splice donor site and a first splice acceptor site (e.g., a Design 3 construct), the binding domain can be upstream or downstream of the single regulatory intron (i.e., within an exonic region adjacent to the single regulatory intron), or the binding domain can be within the single regulatory intron. In some embodiments, the construct comprises two binding domains for a splicing factor of the hnRNP family (e.g., one upstream of the first splice acceptor site and one downstream of the first splice donor site).
[0101] The binding domain of the construct may code for any known binding site for splicing factors in RNA. For example, the sequence characteristics that promote the binding of TDP-43 are described in Lukavsky et al., 2013 (NSMB, 20, pp. 1443-1449), which is incorporated herein by reference. Known binding sites for splicing factors are identified by transcriptome mapping of splicing factors, for example, as determined by immunoprecipitation, and this transcriptome mapping may be performed on the human genome.
[0102] In a preferred embodiment, the binding domain is a TDP-43 binding domain and the hnRNP family splicing factor is TDP-43.
[0103] In some embodiments, the TDP-43 binding domain comprises a region of at least 6 nucleotides, or preferably at least 10 nucleotides, or at least 20 nucleotides, that is statistically significantly enriched for TG dinucleotides and / or TGNNTG hexanucleotides, where N is A, T, C, or G. In some embodiments, the TDP-43 binding domain comprises a region of 6 to 150 nucleotides that is statistically significantly enriched for TG dinucleotides and / or TGNNTG hexanucleotides, where N is A, T, C, or G, where statistically significant enrichment is defined as less than a 0.2% probability that a random sequence of nucleotides of the same length will feature the same number of TG dinucleotides and / or TGNNTG hexanucleotides. In some embodiments, statistically significant enrichment is defined as the probability that random sequences of nucleotides of the same length feature the same number of TG dinucleotides and / or TGNNTG hexanucleotides of 0.15% or less, or 0.1% or less, or 0.05% or less, or 0.01% or less, or 0.003% or less, or 0.001% or less, or 0.0003% or less, or 0.0001% or less. These definitions include both short sequences highly enriched for UG and long sequences broadly enriched for UG, both of which have been shown to preferentially bind TDP-43. In some embodiments and examples, statistically significant enrichment is greater than 1×10 -5 or less, or 1×10 -6 or less, or 1×10 -7 or less, or 1×10 -8 or less, or 1×10 -9 or less, or 1×10 -10 It is defined as the probability:
[0104] Exemplary TDP-43 binding domains include the TDP-43 binding region within UNC13A, which suppresses the inclusion of the UNC13A cryptic exon (SEQ ID NO:1).
[0105] SEQ ID NO:1 TAGATAAAAGGATGGATGGAGAGATGGGTGAGTACATGGATGGATAGATGGATGAGTTGGTGGGTAGATTCGTGGCTAGATGGATGATGGATGGATGGACA, This is approximately a 0.01% probability score that random sequences of nucleotides of the same length will feature the same number of TG dinucleotides.
[0106] Other exemplary TDP-43 binding domains include TGTGTG, which has a probability score of 0.02%, and TGNNTGTG, which has a probability score of 0.15%. Exemplary TDP-43 binding domains described herein include SEQ ID NO: 2: TGTGTGTGTGTGTGTGAATGTGTGTGTGTGTGTGTG , which means that the probability that random sequences of nucleotides of the same length will feature the same number of TG dinucleotides is 5 × 10 -20 This is a modified version (with over 90% sequence identity) of the binding domain found in the human AARS1 gene.
[0107] In some embodiments, the TDP-43 binding domain comprises a sequence enriched in TG dinucleotides. In some embodiments, enrichment in TG dinucleotides is defined as a sequence that contains at least 6 nucleotides of 100% TG dinucleotides (i.e., TGTGTG), or one or more regions having at least 6 nucleotides of 100% TG dinucleotides. In some embodiments, enrichment in TG dinucleotides is defined as a sequence (or one or more regions having at least 8 nucleotides) that contains at least 80% TG dinucleotides (e.g., TGAATGTG), or at least 85%, or at least 90%, or at least 95%, or at least 8 nucleotides of 100% TG dinucleotides (i.e., TGTGTGTG). In some embodiments, enrichment in TG dinucleotides is defined as a sequence (or one or more regions having at least 10 nucleotides) that contains at least 60% TG dinucleotides (e.g., TGAATGAATG (SEQ ID NO: 3)), or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 100% TG dinucleotides. In some embodiments, enrichment in TG dinucleotides is defined as a sequence (or one or more regions having at least 15 nucleotides) that contains at least 53% TG dinucleotides (e.g., TGAATGAATG (SEQ ID NO: 4)), or at least 60%, or at least 65%, or at least 70%, or at least 75%, or at least 80%, or at least 85%, or at least 90%, or at least 95%, or 100% TG dinucleotides.
[0108] In some embodiments, the TDP-43 binding domain comprises a sequence comprising at least one of TGTGTG, or TGTGTGTGTG, or TGTGTGTGTG (SEQ ID NO:5), or TGTGTGTGTGTG (SEQ ID NO:6), or TGTGTGTGTGTGTG (SEQ ID NO:7), or TGTGTGTGTGTGTGTG (SEQ ID NO:8), or TGTGTGTGTGTGTGTGTG (SEQ ID NO:9), or any combination thereof. In some examples, the TDP-43 binding domain comprises a sequence having at least 80% sequence identity to SEQ ID NO:2, or at least 85%, or at least 90% sequence identity, or at least 95% sequence identity, or 100% sequence identity to SEQ ID NO:2-TGTGTGTGTGTGTGAATGTGTGTGTGTGTGTGTG.
[0109] In some examples, the TDP-43 binding domain has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:1-9, or SEQ ID NO:115.
[0110] Although TDP-43 can bind to a wide variety of UG / TG-rich sequences, the binding domain does not need to bind to pure UG / TG-repeats. This is in part because the protein does not contact some RNA residues in its binding footprint, and in part because of multivalent protein-protein interactions that enhance binding to large regions of UG-rich RNA. This means that in some embodiments, the TDP-43 binding domain does not need to be pure UG-repeats. Example sequences include: SEQ ID NO:159 TIFF2025513046000001.tif13162 sequence number 160 TIFF2025513046000002.tif16161
[0111] The construct is configured such that when localized in a cell with nuclear depletion of the splicing factor (e.g., in the absence of binding of the splicing factor to the binding domain), splicing of the first splice acceptor site or the first donor site is not inhibited, and when localized in a cell without nuclear depletion of the splicing factor (e.g., in the presence of binding of the splicing factor to the binding domain), splicing of the first splice acceptor site or the first donor site is inhibited, thereby altering the sequence that is incorporated into the mRNA product of the construct and modulating whether the mRNA product of the construct produces a functional protein.
[0112] In some embodiments, the first splice acceptor site is upstream of the first splice donor site, and the first splice acceptor site and the first splice donor site define a cryptic exon sequence (e.g., in the Design 1 or 2 constructs described herein). In some embodiments, the cryptic exon sequence is a frameshift-inducing cryptic exon sequence, i.e., the exon comprises a length of nucleotides not divisible by 3 (e.g., in the Design 1 or 2 constructs described herein). Additionally or alternatively, the cryptic exon sequence may comprise a start codon. Additionally or alternatively, the cryptic exon sequence may encode at least a portion of the transgene sequence (e.g., in the Design 2 constructs described herein).
[0113] In another embodiment, the first splice donor site is upstream of the first splice acceptor site. In some embodiments, the sequence between the first splice donor site and the first splice acceptor site is a single regulatory intron (e.g., the Design 3 constructs described herein). In some embodiments, production of a functional protein from the transgene can be regulated (i.e., switched off or on) by inclusion or exclusion of at least a portion of the intron in the mRNA product of the construct.
[0114] Start codon The construct comprises a start codon, or multiple start codons or start codon arrays (i.e., in frame with each other). In some embodiments, the start codon may be upstream of a regulatory domain. In some embodiments, the start codon may be within a regulatory domain (e.g., in embodiments comprising a cryptic exon, the start codon may be within a cryptic exon). In some embodiments, the start codon is provided in the form of a Kozak sequence or a Kozak-like sequence. In preferred embodiments, the start codon comprises an ATG. In some examples, the construct comprises a sequence encoding a start codon having at least 80% sequence identity, or at least 85% sequence identity, or at least 90% sequence identity, or at least 95% sequence identity, or at least 100% sequence identity to SEQ ID NO:28.
[0115] Approximately half of human mRNAs have an upstream start codon in the 5' untranslated region that does not initiate translation of the canonical coding sequence of the mRNA. Many of these start codons initiate translation of an upstream open reading frame. Despite the presence of an upstream start codon, these mRNAs result in the expression of a canonical protein from a downstream canonical start codon by a variety of proposed mechanisms including leaky scanning and reinitiation. Thus, the start codon mentioned in the above embodiment does not necessarily have to be the 5'-most start codon in the mRNA product.
[0116] Transgene The construct comprises a transgene sequence (e.g., a protein-encoding sequence), which may be formed from one or more exon sequences (or portions) that together form a complete transgene sequence. In some embodiments, at least a portion of the transgene sequence is downstream of a regulatory domain. In some embodiments, the complete transgene sequence may be uninterrupted. In some embodiments, the complete transgene sequence is downstream of a regulatory domain. In some embodiments, the transgene sequence may be interrupted (i.e., broken into portions). In some embodiments, the transgene sequence may be broken into two or more portions, or three or more portions, or four or more portions, or five or more portions, or six or more portions, or seven or more portions, or eight or more portions, or nine or more portions, or ten or more portions. In some embodiments, at least a portion of the transgene sequence is upstream of the regulatory domain and downstream of the regulatory domain. In some embodiments, i.e., embodiments comprising a cryptic exon defined by a first splice acceptor site and a first splice donor site, the cryptic exon may form part of the transgene sequence. In such embodiments, at least a portion of the transgene sequence may be upstream of the regulatory domain, at least a portion of the transgene sequence may be encoded by a partial cryptic exon sequence, and at least a portion of the transgene sequence may be downstream of the regulatory domain.
[0117] In some embodiments, the complete transgene is of (i.e., encodes) a diagnostic protein. The diagnostic protein may be any suitable diagnostic protein known in the art. The construct may then be used as a biomarker (e.g., to monitor depletion of hnRNP splicing factors). In some embodiments, the diagnostic protein is a fluorescent protein, a luminescent protein, or a protein with a detectable antibody binding tag (e.g., a protein with a peptide or polypeptide tag).
[0118] The fluorescent protein may be any suitable fluorescent protein known in the art. In some embodiments, the fluorescent protein is a monomeric red fluorescent protein (mRFP), such as mCherry or mScarlet. In some embodiments, the fluorescent protein is a green fluorescent protein (GFP) or an enhanced derivative (eGFP). In some embodiments, the green fluorescent protein is mNeonGreen or mGreenLantern. In some embodiments, the fluorescent protein is a blue fluorescent protein. In some embodiments, the fluorescent protein is an orange fluorescent protein. In some embodiments, the fluorescent protein is a yellow fluorescent protein.
[0119] The photoprotein may be any suitable photoprotein known in the art. In some embodiments, the photoprotein is a luciferase protein (e.g., firefly luciferase or Renilla luciferase). In some examples, the luciferase protein is Gaussia luciferase (gLuc), i.e., Gaussia princeps luciferase.
[0120] A protein having a detectable antibody binding tag may have any suitable tag. In some embodiments, the tag is a peptide tag. In some embodiments, the peptide tag is a FLAG tag (including, for example, DYKDDDDK (SEQ ID NO: 10) or DDDDK (SEQ ID NO: 11)), a His tag (HHHHHH, SEQ ID NO: 12)), a HA tag (YPYDVPDYA (SEQ ID NO: 13)), a Myc tag (EQKLISEEDL (SEQ ID NO: 14)), a V5 tag (GKPIPNPLLGLDST (SEQ ID NO: 15)), a S tag (KETAAAKFERQHMDS (SEQ ID NO: 16)), an E tag (GAPVPYPDPLEPR (SEQ ID NO: 17)), a T7 tag (MAS MTGQQMG (SEQ ID NO: 18), VSV-G tag (YTDIEMNRLGK (SEQ ID NO: 19)), Glu-Glu tag (EEEEYMPME (SEQ ID NO: 20)), Strep tag II (WSHPQFEK (SEQ ID NO: 21)), HSV tag (QPELAPEDPED (SEQ ID NO: 22)), chitin binding domain (TTNPGVSAWQVNTAYTAGQLVIYNGKTYK (SEQ ID NO: 23)), calmodulin binding domain (KRRWKKNFIAVSAANRFKKISSSGAL (SEQ ID NO: 24)). In some embodiments, the tag is a polypeptide tag. In some embodiments, the polypeptide tag is a glutathione-S-transferase (GST) tag, a maltose binding protein (MBP) tag, or a thioredoxin (Trx) tag.
[0121] In some embodiments, the transgene is of (i.e., encodes) a therapeutic protein (i.e., a protein that has a therapeutic effect on a cell). The therapeutic protein can be a protein that is deficient or abnormal in the diseased cell. The therapeutic protein can be any suitable therapeutic protein known in the art. In some embodiments, the therapeutic protein is a neuroprotective protein. In some embodiments, the therapeutic protein can be a nuclease, a chaperone, a proteasome protein, a recombinase protein, a splicing regulator, or a transcription factor, or any combination thereof. In some embodiments, the therapeutic protein is a regulatory protein. The regulatory protein can be selected from a recombinase protein, a splicing regulator, a transcription factor, or any combination thereof.
[0122] The nuclease may be any suitable nuclease known in the art. In some embodiments, the nuclease is a Cas nuclease, such as a Cas9 or Cas13 nuclease, or a catalytically inactive derivative of a Cas nuclease, or a modified mutant of a Cas family nuclease with improved specificity or activity, or a nicking Cas9 nuclease. In some embodiments, the Cas family nuclease, or a mutant thereof, is fused to a second protein (e.g., a nicking Cas9 nuclease fused to a reverse transcriptase to allow "primed editing").
[0123] The chaperone protein may be any suitable chaperone protein known in the art. In some embodiments, the chaperone protein is a foldase protein. In some embodiments, the chaperone protein is a heat shock protein. In some embodiments, the heat shock protein is selected from, but not limited to, HSPB1, HSP104, HSP40, or HSP70. In some embodiments, the chaperone is a cyclophilin, e.g., cyclophilin A. In some embodiments, the chaperone is any protein of the DnaJ family.
[0124] The recombinase protein may be any suitable recombinase protein used in the art. In some examples, the recombinase protein is Cre recombinase. In some examples, the recombinase protein is Flp recombinase. In some examples, the recombinase protein is Vika recombinase.
[0125] In some examples, the recombinase protein is a Dre recombinase.
[0126] The proteasome protein may be any suitable proteasome protein known in the art.
[0127] The transcription factor may be any suitable transcription factor known in the art. In some embodiments, the transcription factor may or may not be derived from a human or mammalian transcription factor (e.g., as a truncated or fusion protein). In some embodiments, the transcription factor may be a synthetic engineered transcription factor having, for example, a transcription activator-like effector (TALE), or a zinc finger domain, or a DNA binding domain based on a modified Cas family enzyme (e.g., the CRISPRa system). In some embodiments, the transcription factor may be an activator or repressor of transcription. In some embodiments, the transcription factor may feature a characterized transcriptional regulatory domain, for example, a VP16 domain, or a KRAB domain.
[0128] The splicing regulator may be any suitable splicing regulator known in the art. In some embodiments, the splicing regulator is or comprises a splicing inhibitor. In some embodiments, the splicing regulator is hnRNPA1 or RAVER1. In some embodiments, the splicing regulator further comprises (i.e., is fused to) a splicing regulator, such as a TDP-43 binding domain fused to a splicing regulator, such as a TDP-43 binding domain fused to RAVER1 (e.g., a TDP-43 RNA binding domain fused to RAVER1). In some embodiments, the transgene is configured to autoregulate and / or suppress cryptic splicing (i.e., upon depletion of endogenous hnRNP splicing factors such as TDP-43).
[0129] In some embodiments, the construct may include a single transgene. In other embodiments, the construct may include at least two transgenes. The at least two transgenes may include a first transgene encoding a first therapeutic protein and a second transgene encoding a diagnostic protein, or a first transgene encoding a first therapeutic protein and a second transgene encoding a second therapeutic protein. The two transgenes may be separated by a proteolytic or autocleavage site, for example including any of the sequences of the proteolytic or autocleavage sites described elsewhere herein. In some examples described herein, the two transgenes are separated by a T2A cleavage site.
[0130] The transgene sequence may include a stop codon at the end of the transgene sequence (i.e., unless linked to an additional downstream transgene). In embodiments where the construct includes an additional intron sequence (e.g., a constitutively spliced intron), the stop codon is no more than 55 nucleotides, preferably no more than 50 nucleotides, or no more than 40 nucleotides upstream of the additional intron sequence, or the stop codon is downstream of the additional intron sequence.
[0131] In some embodiments, the transgene is a known sequence that encodes a protein, i.e., a naturally occurring sequence, hi some embodiments, the known sequence is modified by replacing naturally occurring codons with synonymous codons.
[0132] Optional features of the construct In some embodiments, the sequence defined by the first acceptor splice site and the first donor splice site is a frameshift-inducing sequence. In such embodiments (e.g., when the sequence between the first splice acceptor site and the first splice donor site is a frameshift-inducing sequence), the construct may further comprise a premature stop codon (PTC). The premature stop codon may be selected from TAG, TAA, or TGA. The PTC may be downstream of the regulatory domain and upstream of at least a portion of the transgene sequence. In some embodiments, the PTC may be located within at least a portion of the transgene that is downstream of the regulatory domain. In other embodiments, the PTC may not be present in at least a portion of the transgene, e.g., the PTC may be present within another sequence that includes the PTC. In some embodiments, i.e., embodiments that include a single regulatory intron, the PTC may be present within the single regulatory intron.
[0133] The PTCs are positioned and configured to be in-frame with the start codon in the construct's mRNA product when splicing is inhibited (i.e., in healthy cells) and out-of-frame in the construct's mRNA product when splicing is not inhibited (i.e., in diseased cells). The PTC that is in-frame with the start codon results in the production of a truncated protein. This results in the production of a functional protein upon nuclear depletion of splicing factors and no functional protein in the absence of nuclear depletion of splicing factors. This selectively results in the formation of a truncated protein in cells without nuclear depletion.
[0134] Additional intron sequences (e.g., constitutively spliced intron sequences) In some embodiments, the construct may further comprise an additional intron sequence downstream of the regulatory domain. The additional intron sequence is in an exonic context or preceded or followed by an exonic context (e.g., flanked by exonic sequences). In a preferred embodiment, the additional intron sequence comprises a constitutively spliced intron sequence. The additional intron sequence is at least 40 nucleotides downstream of the PTC, but in a preferred embodiment, the PTC is at least 50 nucleotides upstream of the additional intron sequence, or at least 55 nucleotides upstream of the additional intron sequence. In some embodiments, the PTC is 40-55 nucleotides upstream of the additional intron sequence, or 50-55 nucleotides upstream of the additional intron sequence. In some embodiments, the additional intron sequence is downstream of the complete transgene sequence. In another embodiment, the additional intron sequence is downstream of the regulatory domain and upstream of at least a portion of the transgene sequence.
[0135] The presence of additional intronic sequences downstream of the PTC promotes attachment of the exon junction complex (EJC) to the resulting mRNA and promotes nonsense-mediated decay when splicing of the first splice acceptor and / or first splice donor is suppressed (i.e., the PTC is in-frame with the start codon). When splicing is not suppressed, the PTC codon will not be in-frame with the start codon in the mRNA product of the construct, and therefore the ribosome will remove the EJC and nonsense-mediated decay will not occur.
[0136] In the examples described herein, the additional intron sequence and surrounding exon context are derived from human RPS24, although any suitable intron and exon sequence can be used. In some embodiments, the additional intron sequence comprises any natural intron and exon sequence (e.g., any intron and exon from the human genome). In another embodiment, the additional intron sequence and exon are formed from or derived from synthetic sequences. These sequences can be designed using the Splice AI algorithm, i.e., the splice sites that define the additional intron sequence have a splice score of at least 0.01, or at least 0.05, preferably at least 0.1, or at least 0.5, or more preferably at least 0.9. Furthermore, synthetic sequences can be designed using "algorithm 1" described herein.
[0137] Protease cleavage site or autocleavage site In some embodiments (e.g., certain Design 1 and Design 3 constructs described herein), the construct further comprises a protease cleavage site or autocleavage site. In some embodiments, the protease cleavage site or autocleavage site can be downstream of the regulatory domain and upstream of at least a portion of the transgene sequence. In another embodiment, the protease cleavage site or autocleavage site can be between the transgene sequences. The protease cleavage site or autocleavage site can be selected from P2A, T2A, F2A, E2A, furin, PCSK1, PCSK6, PCSK7, cathepsin B, granzyme B, factor XA, enterokinase, genenase, sortase, prescission protease, thrombin, TEV protease, or elastase 1. In some examples described herein, the cleavage site is P2A or T2A. The protease cleavage site allows cleavage of the protein encoded by the transgene from any peptide encoded by the regulatory domain, or cleavage of the protein encoded by a first transgene and, optionally, a protein encoded by a second transgene.
[0138] Regulation of constructs The construct and regulatory domain are configured such that (i) when localized in a cell with nuclear depletion of splicing factors of the hnRNP family (e.g., in the absence of binding of the splicing factor and the binding domain), splicing of the first splice acceptor site and the first donor site is not inhibited and a functional protein is produced from the transgene sequence. In other words, a functional protein (i.e., a functional protein encoded by the complete and uninterrupted transgene sequence) is produced from the mRNA product of the construct. A functional protein may be defined herein as a protein produced when the complete and uninterrupted transgene sequence is present in the mRNA product, in frame with the start codon, and there is no in-frame stop codon between the start codon and the transgene sequence. A functional protein may additionally or alternatively be defined as a polypeptide chain of at least 30, preferably 50, more preferably 100 amino acids that can act alone or in tandem with one or more additional proteins (e.g., as a heterodimer) to fulfill a therapeutic, diagnostic, or regulatory role in a cell. For example, a functional protein can be a full-length GFP protein that has intrinsic fluorescence, or one component of a split GFP system that has fluorescence when bound to a second component of the split GFP system, or a mutant or truncated GFP fragment that does not emit detectable fluorescence in an assay such as Western blotting.
[0139] The construct and regulatory domain are also configured such that (ii) when localized in a cell that does not have nuclear depletion of hnRNP family splicing factors (e.g., when there is binding of the splicing factor and the binding domain), splicing of the first splice acceptor site and / or the first donor site is suppressed and no functional protein is produced from the complete transgene sequence. In some embodiments, this can occur because at least a portion of the transgene sequence is not in frame with the start codon (e.g., the sequence defined by the first splice acceptor site and the first splice donor site is a frameshift-inducing sequence). In some embodiments, this can occur because at least a portion of the transgene sequence is not present in the mRNA product of the construct (i.e., in an embodiment where a cryptic exon sequence codes for a portion of the transgene and the cryptic exon sequence is not present in the mRNA product of the construct in a healthy cell, the transgene sequence is not fully transcribed). In some embodiments, this may occur because a sequence is introduced into the mRNA product of the construct that interrupts the transgene sequence (e.g., in embodiments where the first splice donor site and the first splice acceptor site define a single regulatory intron and there is no depletion of splicing factors, the mRNA product of the construct in healthy cells incorporates at least a portion of the intron, or alternatively, the mRNA product of the construct in healthy cells does not include a portion of the transgene sequence). In this last embodiment, this interruption may include the introduction of a PTC and / or a disruptive amino acid sequence that inhibits protein function.
[0140] The cell may be any suitable cell. In some embodiments, the cell is a mammalian cell, more preferably a human cell. In preferred embodiments, the cell is nuclear depleted of hnRNP splicing factors (e.g., TDP-43 depleted). In some embodiments, the cell is a brain cell. In some embodiments, the cell is a neuron or neural cell. In some embodiments, the cell is a microglial cell or an astrocyte. In some embodiments, the cell is a muscle cell.
[0141] In a first embodiment of the first aspect or according to the second aspect of the invention, the regulatory sequence is regulated by cryptic splicing. In such an embodiment, the regulatory sequence comprises a cryptic exon sequence between a first splice acceptor site and a first splice donor site, the cryptic exon being embedded within an intron region. This embodiment is described in more detail below and illustrated by the embodiment shown in Figures 1 and 2. The construct comprises: (i) When localized to cells with nuclear depletion of hnRNP family splicing factors, the cryptic exon sequence was present in the mRNA product of the construct; (ii) The cryptic exon is absent in the mRNA product of the construct when localized in cells without nuclear depletion of hnRNP family splicing factors. It is configured as follows.
[0142] In an embodiment of the first aspect or according to the third aspect of the invention, the regulatory sequence is regulated by splicing of a single regulatory intron.
[0143] In such embodiments, the intron sequence is between the first splice donor site and the first splice acceptor site. (i) When localized in cells with nuclear depletion of splicing factors, the single regulatory intron is spliced out to produce a functional protein; (ii) when localized in cells without nuclear depletion of splicing factors, the single regulatory intron is not properly spliced or is not spliced at all, and no functional protein is produced; It is configured as follows.
[0144] Each of the above embodiments or aspects is described in further detail below. All such embodiments importantly comprise a binding domain for a splicing factor of the hnRNP family, a first splice acceptor site, a first splice donor site, and a transgene sequence (i.e., a transgene sequence encoding a functional protein). The construct is configured such that binding of the splicing factor to the binding domain regulates splicing of the first splice acceptor site or the first splice donor site. Splicing is not inhibited in cells depleted of the splicing factor, but is inhibited in cells not depleted of the splicing factor. This then regulates whether the transgene is sufficiently expressed or encoded to produce a functional protein.
[0145] Constructs in which the regulatory domain is regulated by cryptic splicing In an embodiment of the second aspect or the first aspect, Start codon, A regulatory domain that includes: a first splice acceptor site and a first splice donor site defining a cryptic exon sequence; an intron region defined by a second splice donor site and a second splice acceptor site, wherein the cryptic exon sequence is located within the intron region; and a binding domain for a splicing factor of the heterogeneous nuclear ribonucleoprotein (hnRNP) family located within 150 nucleotides of the first splice donor site or the first splice acceptor site, and Transgene sequence Including, (i) when localized in a cell that is depleted of splicing factors, splicing of the first splice acceptor site and the first donor site is not repressed, the cryptic exon sequence is present in the mRNA product of the construct, and a functional protein is produced from the transgene sequence (i.e., a functional protein is produced from an mRNA product in which the functional protein is encoded by an intact, uninterrupted transgene sequence); (ii) when localized in a cell in which the splicing factor is not depleted, splicing of the first splice acceptor site and / or the first donor site is repressed, the cryptic exon sequence is absent in the mRNA product of the construct, and no functional protein is produced from the transgene sequence (i.e., a functional protein is produced from an mRNA product encoded by an intact, uninterrupted transgene sequence). A construct configured to:
[0146] The binding domain for a splicing factor of the heterologous nuclear ribonucleoprotein (hnRNP) family, the premature stop codon, the first splice acceptor site, the first splice donor site, and the transgene sequence are as described elsewhere herein. In embodiments in which the regulatory domain comprises a cryptic exon, the first splice acceptor site and the first splice donor site can be referred to as "cryptic splice sites."
[0147] Intron Regions The intron region is defined by a second splice donor site and a second splice acceptor site. The intron region comprises (from upstream to downstream) a first portion of the intron region, a cryptic exon sequence, and a second portion of the intron region. The intron region comprises a binding domain for a splicing factor of the hnRNP family located at most 150 nucleotides upstream or downstream from the first splice acceptor and / or the first splice donor site (as described above). The binding domain may be within the first portion of the intron region, within the cryptic exon sequence, or within the second portion of the intron region.
[0148] The first portion of the intron region may be described as a "first intron" and the second portion of the intron region may be described as a "second intron." In some embodiments, the first portion of the intron region and / or the second portion of the intron region each comprise at least 50 nucleotides, preferably at least 70 nucleotides, or at least 100 nucleotides, or at least 150 nucleotides. In some embodiments, the first portion of the intron region and / or the second portion of the intron region each comprise between 70 nucleotides and 5000 nucleotides, or between 70 and 1000 nucleotides, or between 70 and 500 nucleotides, and in some examples, between 125 nucleotides and 250 nucleotides.
[0149] In some embodiments, the second splice donor site and / or the second splice acceptor site have a splice score of 0.01 (Splice AI score 99.8 percentile, see FIG. 11) or greater as determined by the Splice AI algorithm. In preferred embodiments, the second splice donor site and / or the second splice acceptor site have a splice score of 0.05 or greater, or at least 0.1 or greater, or at least 0.2 or greater, or at least 0.3 or greater, or at least 0.4 or greater, or at least 0.5 or greater, or at least 0.6 or greater, or at least 0.7 or greater, or at least 0.8 or greater, or at least 0.9 or greater, more preferably at least 0.95, or at least 0.96, or at least 0.97, or at least 0.98, or at least 0.99 or greater, as determined by the Splice AI algorithm.
[0150] In some embodiments, the intron region may be derived from a naturally occurring intron region (e.g., from the human genome) that contains a cryptic exon, and the cryptic exon is regulated by a splicing factor of the hnRNP family (e.g., TDP-43). In some embodiments, the intron region may be at least 80% identical to at least a portion of a naturally occurring intron region (e.g., from the human genome) that contains a cryptic exon, or at least 85% identical to at least a portion of a naturally occurring intron region (e.g., from the human genome) that contains a cryptic exon, or at least 90% identical to at least a portion of a naturally occurring intron region (e.g., from the human genome) that contains a cryptic exon, or at least 95% identical to at least 100% identical to at least a portion of a naturally occurring intron region (e.g., from the human genome) that contains a cryptic exon. In some embodiments, the intron region may be modified by truncation (i.e., a portion of the intron region upstream and downstream of the cryptic exon may contain fewer nucleotides than found in the human genome). The intron region may be modified by the insertion, deletion, or substitution of one or more nucleotides, e.g., two, three, four, five, or more than five nucleotides. In some embodiments, the intron region may be modified by (i) mutating nucleotides in the intron region to remove one or more premature stop codons, and / or (ii) inserting or deleting one or two nucleotides in a cryptic exon sequence to introduce a frameshift. In some embodiments, the intron region is selected from the group consisting of AACSP1, AARS1, ABCB1, ABCD1, AC002310.11, AC002310.7, AC002456.2, AC008543.1, AC008676.3, AC009133.12, AC010531.1, AC015712.1, AC015712.6, AC022387.2, AC022966.1, AC025165.6, AC064807.1, AC092073.1, AC138932.1, AC245041.2, ACSF2, ACTL6B, ACTR1A, ADARB1, ADARB2, ADCY1, ADCY7, ADCY8, ADGRB1, ADGRL1, ADSSL1,AGK、AGRN、AHNAK、AKT3、AL023775.2、AL031282.2、AL035461.3、AL121845.3、AL157392.3、AL157392.5、AL354696.2、AL360181.3、AL645568.1、AL669831.3、AL672142.1、ALDH3B1、AMPD2、ANKRD19P、ANKRD44、ANOS2P、AP000662.4、AP006621.8、AP4M1、ARAP3、ARF1、ARHGAP22、ARHGAP23、ARHGEF16、ARHGEF19、ASGR1、ATAD5、ATG4B、ATP5MG、ATP8A2、ATXN1、ATXN10、BCL2L11、BCL2L13、BLCAP、BMP8B、BNIP3P11、BRD1、BTN3A3、C16orf95、C20orf194、C2orf81、C4orf36、C5orf66、CACNB2、CACNG5、CAMK2B、CAMTA1、CASP8、CASTOR1、CBY1、CCDC102B、CCDC150、CCDC183-AS1、CCDC33、CCT2、CDHR2、CDK11A、CDKAL1、CDON、CELF5、CENPBD1P1、CENPK、CENPS-CORT、CEP152、CEP290、CEP72、CEP83、CH17-189H20.1、CH507-154B10.1、CHD8、CHFR、CHGB、CHRNA5、CHRNB3、CLCN6、CLSPN、CLTCL1、CNGA3、CNPY1、CORO6、CPVL、CREB3L4、CRLS1、CRTC1、CSMD2、CTC-490E21.12、CTD-2014B16.3、CTD-2054N24.2、CTD-2162K18.4、CTD-2554C21.2、CTD-2561J22.3、CU634019.6、CUL9、CYFIP2、CYP2C8、DACH2、DACT3-AS1、DAGLA、DAPK1、DELE1、DENND2B、DGKA、DLG5、DLGAP1、DNAJC12、DNAJC25-GNG10、DNMT3A、DNMT3B、DOCK1、DPF1、DUXAP9、EBF1、ECEL1、EHD2、EIF2A、EIF2AK1、EIF4ENIF1、ELAVL3、EML6、ENAH、ENTPD6、EP300、EP400、EPB41L1、<h2 style=";text-align:left;direction:ltr">EPB41L4A、EPS8L2、ETV5、F12、FADS2、FAM114A2、FAM156A、FAM182B、FAM66D 、FAM66E、FBL、FBXL19、FGFR4、FIRRE、FKBP14-AS1、FOXK1、FRYL、G2E3、G3BP 1、GALNT12、GAS6、GATA2、GLIPR2、GMPPA、GOLGA7B、GOLGA8A、GPHN、GPSM2、G PX7、GRAMD1A、GREB1、GRIN2D、GSTCD、GTF2H2、GTF2IP13、HAUS2、HDAC6、HDGF L2、HDLBP、HECTD4、HERC2P2、HIPK1、HROB、HULC、ICA1、IFT122、IGSF21、IGS F9、IK、IL15、INPP4A、INSR、INTS11、IQCE、IQCK、ISL2、ISYNA1、ITGA3、ITGA7 、ITPR3、KALRN、KATNA1、KCNIP1、KCNIP2、KCNK15-AS1、KCNQ2、KCNT1、KDM1B 、KDM4D、KIAA1211、KIAA1217、KIF14、KIF21A、KLC1、KMT5A、KNDC1、KRT8、L3M BTL1、LCOR、LIAS、LINC00265、LINC00342、LINC00475、LINC01002、LINC012 24、LINC01322、LINC01503、LINC01572、LINC01684、LINC02082、LINC02202、 LINC02506、LINGO1、LMNA、LRP1B、LRP8、LSM12、LSS、LTBP2、MACROD1、MADD、 MANBAL、MAP2K6、MAPKAPK5、MATK、MBP、MC1R、MCM9、MDC1、MED12、MED13L、MEI S2, METTL8, MGAT5B, MIER3, MMAA, MRPL34, MTRR, MTX1P1, NAA38, NADSYN1, NAT1, NBEA, NBPF9, NDUFB9, NFKBIZ, NFYC, NIPSNAP3B, NPIPB11, NPLOC4, NSFL 1C、NTRK2、NTRK3、NUP188、NUP210、OBSCN、OPCML、PAOX、PATJ、PCBP3、PCBP4 、PCDH11X、PCSK1N、PDCD2L、PDCD6、PDE2A、PDE9A、PER3、PHF2、PHF5A、PI4KA、PIGG、PIGU、PKD1P3、PKN1、PLCE1、PLAY1、PLAY6、PLKHG2、PLKHG4、PL EKHM2、POLD1、POLR2F、POU2F2、PPCDC、PPIP5K1、PPM1N、PPP1R14B-AS1、PRDM 8、PRELID3A、PREX1、PRKG2、PROX1-AS1、PRPF40B、PRRT4、PRUNE2、PSPC1、PT K2、PTPN13、PTPN21、PTPRN2、PTPRT、PUDP、PUS7L、PWWP3A、PXDN、RAB20、RAB2 7A, RALGAPA2, RANBP17, RASGRP2, RBMXL1, RC3H1, RCAN3, RET, RFLNA, RGMA RHOQ、RP1-120G22.12、RP1-138B7.8、RP1-283E3.8、RP1-59M18.2、RP11-101 E3.5、RP11-108K14.8、RP11-108L7.4、RP11-124N2.1、RP11-155D18.12、RP 11-155G14.5、RP11-155G14.6、RP11-206L10.2、RP11-30K9.6、RP11-345P4. 10、RP11-411B6.6、RP11-436D23.1、RP11-465B22.3、RP11-479O9.4、RP11- 505D17.1、RP11-511P7.6、RP11-566K11.4、RP11-613M10.9、RP11-61L23.2、 RP11-718O11.1、RP11-739N20.2、RP11-73M18.2、RP11-761B3.1、RP11-795 F19.5、RP11-977G19.10、RP4-583P15.15、RP5-967N21.13、RPGRIP1L、RSF1、 RTL1、SCN9A、SCUBE3、SDAD1、SEC14L1、SEC31B、SEMA4D、SEMA6C、SEMA6D、SE PT11、SEPT7P2、SEPTIN11、SEPTIN3、SEPTIN6、SEPTIN7P2、SERGEF、SERP1、SE TD5、SFXN2、SGMS1、SH2B1、SH3BP5-AS1、SH3PXD2B、SHANK1、SHLD2、SIPA1L3 、SIX1、SLC12A5、SLC1A6、SLC24A3、SLC25A14、SLC25A22、SLC2A11、SLC35G1、SLC38A7, SLC41A2, SLC4A3, SMAD4, SMG1P7, SPATA17, SPATS2, SPEG, SPIN1, SRRM4, ST5, STMN2, STOX2, STRA6, STX BP5L, SUPT3H, SVEP1, SYDE1, SYNE1, SYNGR3, SYNJ2, SYT7, TAF6, TAFA2, TBCD, TBL1XR1, TENM3, TEX9, TGFB3, THUM PD3-AS1, TM6SF2, TMEM117, TMEM175, TMEM189, TMEM191A, TMEM198B, TMEM214, TMEM230, TMEM88, TPRA1, TRAF3, T RAPPC12, TRIM16, TRIM6, TRIO, TRRAP, TSHZ3, TSPAN3, TTC39C-AS1, TTLL4, TTTY14, TUBB3, TUBB6, TUBGCP6, TXLN GY, UNC13A, UNK, USP10, USP28, USP36, VAX2, VPS29, VPS50, VPS53, WARS2, WASL, WDFY2, WDR19, WDR37, WDR4, WWOX , ZBTB18, ZC2HC1C, ZCCHC4, ZDHHC1, ZFAT, ZFP91, ZFP91-CNTF, ZGPAT, ZNF195, ZNF202, ZNF236, ZNF320, ZNF382, ZNF394, ZNF420, ZNF423, ZNF429, ZNF43, ZNF48, ZNF527, ZNF571-AS1, ZNF583, ZNF594-DT, ZNF598, ZNF692, ZNF696, ZNF700, ZNF737, ZNF785, ZNF789, ZNF81, ZNF814, ZNF826P, ZNF875, ZNHIT1, ZRANB3, ZSCAN12, at least in part.
[0151] In some embodiments and examples described herein, at least a portion of the intron region is derived from AARS1, i.e., the intron region between exons 4 and 5 of AARS1. In some embodiments, the first and second portions of the intron region are derived from AARS1, i.e., the intron region between exons 4 and 5 of the human genome. The first portion of the intron region derived from AARS1 may correspond to at least a portion of the intron region between exons 4 and 5 of AARS1 of the human genome, upstream of the AARS1 cryptic exon. The second portion of the intron region derived from AARS1 may correspond to at least a portion of the intron region between exons 4 and 5 of AARS1 of the human genome, downstream of the AARS1 cryptic exon.
[0152] In some embodiments, the first portion of the intron region may comprise a sequence that is at least 80% identical to one of SEQ ID NO:30, SEQ ID NO:70, SEQ ID NO:76, SEQ ID NO:82, SEQ ID NO:119, SEQ ID NO:125, SEQ ID NO:131, SEQ ID NO:137, SEQ ID NO:143, SEQ ID NO:149, or SEQ ID NO:155, or at least 85%, or at least 90%, or at least 95%, or at least 100% identical to one of SEQ ID NO:30, SEQ ID NO:70, SEQ ID NO:76, SEQ ID NO:82, SEQ ID NO:119, SEQ ID NO:125, SEQ ID NO:131, SEQ ID NO:137, SEQ ID NO:143, SEQ ID NO:149, or SEQ ID NO:155, or SEQ ID NO:179, or SEQ ID NO:185, or SEQ ID NO:191, or SEQ ID NO:197. In some embodiments, the second portion of the intron region may comprise a sequence that is at least 80% identical to one of SEQ ID NO:32, or SEQ ID NO:72, or SEQ ID NO:78, or SEQ ID NO:84, or SEQ ID NO:121, or SEQ ID NO:127, or SEQ ID NO:133, or SEQ ID NO:139, or SEQ ID NO:145, or SEQ ID NO:151, or SEQ ID NO:157, or at least 85%, or at least 90%, or at least 100% identical to SEQ ID NO:32, SEQ ID NO:72, or SEQ ID NO:78, or SEQ ID NO:84, or SEQ ID NO:121, or SEQ ID NO:127, or SEQ ID NO:133, or SEQ ID NO:139, or SEQ ID NO:145, or SEQ ID NO:151, or SEQ ID NO:157, or SEQ ID NO:181, or SEQ ID NO:187, or SEQ ID NO:193, or SEQ ID NO:199. In some examples, a first portion of the intron sequence is at least 80%, or at least 85%, or at least 90%, or at least 95% identical to SEQ ID NO: 30, and a second portion of the intron sequence is at least 80%, or at least 85%, or at least 90%, or at least 95% identical to SEQ ID NO: 32, which is derived from the AARS1 intron region between exon 4 and exon 5 of the human genome.
[0153] In other examples, the first and second portions of the intron region are synthetic. In some embodiments, the intron region is designed such that the intron region begins with GT(AAG) and ends with (C)AG. In some embodiments and examples, the first and second portions of the intron region can be selected such that the first acceptor splice site and / or the first donor splice site have a splice score (as determined by the Splice AI algorithm) of at least 0.01, or at least 0.05, or at least 0.1, or at least 0.3, or between 0.01 and 0.8, and / or the second acceptor splice site and / or the second splice donor site have a splice score of at least 0.01, preferably at least 0.5, or at least 0.9, or at least 0.95, as determined by the Splice AI algorithm. In some embodiments, the intron region (i.e., the first portion of the intron region, the cryptic exon sequence, or the second portion of the intron region) is designed to include a binding domain for a splicing factor of the hnRNP family (e.g., TDP-43). In some embodiments, the binding domain is for TDP-43, and the intron sequence includes a sequence that is at least 80% identical, or at least 85% identical, or at least 90% identical, or at least 95% identical, or at least 100% identical to SEQ ID NO:2 or SEQ ID NO:115, or includes a TDP-43 binding domain as described elsewhere herein. In a preferred embodiment or example, the intron region is designed such that the intron region (e.g., the first portion of the intron region) includes a polypyrimidine tract. A polypyrimidine tract as defined herein can be described as a pyrimidine-rich 20-nucleotide region, defined as a 20-nucleotide region that is at least 70% pyrimidines, or a 30-nucleotide region that is at least 80% pyrimidines.
[0154] As indicated above, the intron region is defined by a second splice donor site and a second splice acceptor site. The second splice donor site and the second splice donor site are generally at least 150 nucleotides in length, more preferably at least 200 nucleotides in length. In some embodiments, the sequence before and after the second splice acceptor site is HAG / N, where / represents the splice site, H=C, T or A, and N is C, T, A or G. In some embodiments or examples, the construct comprises a polypyrimidine tract upstream of the second splice acceptor site (i.e., within the cryptic exon sequence upstream of the second splice acceptor site, e.g., upstream of HAG / N). In some embodiments, the polypyrimidine tract is upstream of the second splice acceptor site, more preferably up to 40 nucleotides upstream of the second splice acceptor site or up to 20 nucleotides upstream of the first splice acceptor site. A polypyrimidine tract as defined herein can be described as a pyrimidine-rich region, defined as a region of 20 nucleotides that is at least 70% pyrimidines and a region of 30 nucleotides that is at least 80% pyrimidines.
[0155] In some instances, the sequences before and after the second donor splice are CAG / GT, where / represents the splice site.
[0156] In some embodiments, the intron region includes one or more branch sites containing an adenosine upstream of the first and / or second splice acceptor site and a polypyrimidine tract (i.e., within the intron region upstream of the second splice acceptor site). A P, where N is any nucleotide, P is a pyrimidine (i.e., C or T), and the underlined A is, for example, a branch point (e.g., CTG). AC). The branch site may be located up to 45 nucleotides upstream of the first and / or second splice acceptor, preferably up to 35 nucleotides upstream of the first and / or second splice acceptor, preferably 20-35 nucleotides upstream of the first and / or second splice acceptor.
[0157] Cryptic exon A cryptic exon sequence is defined by (i.e., between) a first splice acceptor site and a first splice donor site. In some embodiments, the first splice donor site and / or the first splice acceptor site have a splice score of 0.01 (SpliceAI score 99.8 percentile) or greater as determined by the Splice AI algorithm, or in some embodiments, 0.05 or greater, or in some embodiments, 0.1 or greater. In some embodiments, the first splice donor site and / or the first splice acceptor site that define a cryptic exon have a splice score of 0.01-0.7, or 0.05-0.7, or 0.1-0.7. In a preferred embodiment, the splice score of the first splice acceptor site and the first splice donor site may be lower than the splice score of the second splice acceptor site and the second splice donor site. In a preferred embodiment, the intron region (i.e., defined by the second splice donor site and the second splice acceptor site) does not include other splice sites identified as having a splice score of 0.2 or higher. In a preferred embodiment, the first splice acceptor site and the first splice donor site have the highest Splice AI score in the intron region (i.e., defined by the second splice donor site and the second splice acceptor site, but not including the second splice donor site and the second splice acceptor site). In a preferred embodiment, the first splice acceptor site and the first splice donor site have the highest Splice AI score in the cryptic exon sequence. In some embodiments, the first splice acceptor site and the first splice donor site have a highest Splice AI score within 100 nucleotides, or 50 nucleotides, or 25 nucleotides of the first acceptor and first splice donor site.
[0158] In some embodiments, the cryptic exon sequence comprises from about 10 nucleotides to about 2000 nucleotides, preferably from 30 to 500 nucleotides, or in some instances, from 44 nucleotides to about 200 nucleotides.
[0159] In some embodiments, the cryptic exon sequence is a frameshift-inducing cryptic exon sequence, i.e., the exon sequence comprises a number of nucleotides not divisible by 3. The construct may be (i.e., with respect to the mRNA product of the construct) (i) When localized in cells depleted of hnRNP family splicing factors, the entire transgene sequence is in frame with the start codon, (ii) at least a portion of the transgene sequence is out-of-frame with the initiation codon when localized in cells that are not depleted of hnRNP family splicing factors; It is configured as follows.
[0160] In such an embodiment, the construct may further comprise a premature stop codon downstream of the regulatory domain and the cryptic exon sequence. When localized in a cell with nuclear depletion of splicing factors, the cryptic exon sequence is included in the construct's mRNA and the start codon will be out of frame with the premature stop codon. When localized in a cell without nuclear depletion of splicing factors, the cryptic exon sequence is not included in the construct's mRNA and the start codon will be in frame with the premature stop codon. In such an embodiment, the construct may further comprise additional intron sequences downstream of the regulatory domain and the transgene sequence, as described elsewhere herein.
[0161] In another embodiment, the cryptic exon sequence is not a frameshift-inducing cryptic exon sequence, i.e., the nucleotide sequence comprises a number of nucleotides divisible by 3. Such an embodiment can be used, for example, when the cryptic exon comprises a start codon. Such an embodiment can be used when the cryptic exon encodes at least a portion of a transgene. In such a construct, the construct or transgene sequence may not comprise a PTC (i.e., it is associated with regulating protein expression).
[0162] In some embodiments, the cryptic exon sequence is a known cryptic exon regulated by a splicing factor of the hnRNP family, such as TDP-43. In some embodiments, the cryptic exon sequence is a cryptic exon sequence of a human gene, such as AACSP1, AARS1, ABCB1, ABCD1, AC002310.11, AC002310.7, AC002456.2, AC008543.1, AC008676.3, AC009133.12, AC010531.1, AC015712.1, AC015712.6, AC022387.2, AC022966.1, AC025165.6, AC064807.1, AC092073.1 , AC138932.1, AC245041.2, ACSF2, ACTL6B, ACTR1A, ADARB1, ADARB2, ADCY1, ADCY7, ADCY8, ADGRB1, ADGRL1, ADSSL1, AGK, AGRN, AHNAK, AKT 3, AL023775.2, AL031282.2, AL035461.3, AL121845.3, AL157392.3, AL157392.5, AL354696.2, AL360181.3, AL645568.1, AL669831.3, AL6 72142.1, ALDH3B1, AMPD2, ANKRD19P, ANKRD44, ANOS2P, AP000662.4, AP006621.8, AP4M1, ARAP3, ARF1, ARHGAP22, ARHGAP23, ARHGEF16, AR HGEF19, ASGR1, ATAD5, ATG4B, ATP5MG, ATP8A2, ATXN1, ATXN10, BCL2L11, BCL2L13, BLCAP, BMP8B, BNIP3P11, BRD1, BTN3A3, C16orf95, C20or f194, C2orf81, C4orf36, C5orf66, CACNB2, CACNG5, CAMK2B, CAMTA1, CASP8, CASTOR1, CBY1, CCDC102B, CCDC150, CCDC183-AS1, CCDC33, CCT 2, CDHR2, CDK11A, CDKAL1, CDON, CELF5, CENPBD1P1, CENPK, CENPS-CORT, CEP152, CEP290, CEP72, CEP83, CH17-189H20.1, CH507-154B10.1,CHD8, CHFR, CHGB, CHRNA5, CHRNB3, CLCN6, CLSPN, CLTCL1, CNGA3, CNPY1, CO RO6, CPVL, CREB3L4, CRLS1, CRTC1, CSMD2, CTC-490E21.12, CTD-2014B16.3 CTD-2054N24.2, CTD-2162K18.4, CTD-2554C21.2, CTD-2561J22.3, CU634 019.6 CUL9, CYFIP2, CYP2C8, DACH2, DACT3-AS1, DAGLA, DAPK1, DELE1, DEN ND2B, DGKA, DLG5, DLGAP1, DNAJC12, DNAJC25-GNG10, DNMT3A, DNMT3B, DOCK 1, DPF1, DUXAP9, EBF1, ECEL1, EHD2, EIF2A, EIF2AK1, EIF4ENIF1, ELAVL3, E ML6, ENAH, ENTPD6, EP300, EP400, EPB41L1, EPB41L4A, EPS8L2, ETV5, F12, F ADS2, FAM114A2, FAM156A, FAM182B, FAM66D, FAM66E, FBL, FBXL19, FGFR4, F IRRE, FKBP14-AS1, FOXK1, FRYL, G2E3, G3BP1, GALNT12, GAS6, GATA2, GLIPR 2, GMPPA, GOLGA7B, GOLGA8A, GPHN, GPSM2, GPX7, GRAMD1A, GREB1, GRIN2D, G STCD, GTF2H2, GTF2IP13, HAUS2, HDAC6, HDGFL2, HDLBP, HECTD4, HERC2P2, H.S IPK1, HROB, HULC, ICA1, IFT122, IGSF21, IGSF9, IK, IL15, INPP4A, INSR, IN TS11, IQCE, IQCK, ISL2, ISYNA1, ITGA3, ITGA7, ITPR3, KALRN, KATNA1, KCNI P1, KCNIP2, KCNK15-AS1, KCNQ2, KCNT1, KDM1B, KDM4D, KIAA1211, KIAA1217 KIF14, KIF21A, KLC1, KMT5A, KNDC1, KRT8, L3MBTL1, LCOR, LIAS, LINC0026 5. LINC00342, LINC00475, LINC01002, LINC01224, LINC01322, LINC01503.LINC01572、LINC01684、LINC02082、LINC02202、LINC02506、LINGO1、LMNA、 LRP1B、LRP8、LSM12、LSS、LTBP2、MACROD1、MADD、MANBAL、MAP2K6、MAPKAPK5 、MATK、MBP、MC1R、MCM9、MDC1、MED12、MED13L、MEIS2、METTL8、MGAT5B、WEDNESDAY 3、MMAA、MRPL34、MTRR、MTX1P1、NAA38、NADSYN1、NAT1、NBEA、NBPF9、NDUFB9、 NFKBIZ、NFYC、NIPSNAP3B、NPIPB11、NPLOC4、NSFL1C、NTRK2、NTRK3、NUP188 、NUP210、OBSCN、OPCML、PAOX、PATJ、PCBP3、PCBP4、PCDH11X、PCSK1N、PDCD2L 、PDCD6、PDE2A、PDE9A、PER3、PHF2、PHF5A、PI4KA、PIGG、PIGU、PKD1P3、PKN1 、PLCE1、PLAY1、PLAY6、PLAYG2、PLAYG4、PLAYHM2、POLD1、POLR2F、POU 2F2、PPCDC、PPIP5K1、PPM1N、PPP1R14B-AS1、PRDM8、PRELID3A、PREX1、PRKG 2、PROX1-AS1、PRPF40B、PRRT4、PRUNE2、PSPC1、PTK2、PTPN13、PTPN21、PTPRN 2、PTPRT、PUDP、PUS7L、PWWP3A、PXDN、RAB20、RAB27A、RALGAPA2、RANBP17、R ASGRP2、RBMXL1、RC3H1、RCAN3、RET、RFLNA、RGMA、RHOQ、RP1-120G22.12、RP1 -138B7.8、RP1-283E3.8、RP1-59M18.2、RP11-101E3.5、RP11-108K14.8、RP 11-108L7.4、RP11-124N2.1、RP11-155D18.12、RP11-155G14.5、RP11-155G1 4.6、RP11-206L10.2、RP11-30K9.6、RP11-345P4.10、RP11-411B6.6、RP11- 436D23.1、RP11-465B22.3、RP11-479O9.4、RP11-505D17.1、RP11-511P7.6、<h2 style=";text-align:left;direction:ltr">RP11-566K11.4、RP11-613M10.9、RP11-61L23.2、RP11-718O11.1、RP11-73 9N20.2、RP11-73M18.2、RP11-761B3.1、RP11-795F19.5、RP11-977G19.10、 RP4-583P15.15、RP5-967N21.13、RPGRIP1L、RSF1、RTL1、SCN9A、SCUBE3、SD AD1、SEC14L1、SEC31B、SEMA4D、SEMA6C、SEMA6D、SEPT11、SEPT7P2、SEPTIN1 1、SEPTIN3、SEPTIN6、SEPTIN7P2、SERGEF、SERP1、SETD5、SFXN2、SGMS1、SH2 B1、SH3BP5-AS1、SH3PXD2B、SHANK1、SHLD2、SIPA1L3、SIX1、SLC12A5、SLC1A 6, SLC24A3, SLC25A14, SLC25A22, SLC2A11, SLC35G1, SLC38A7, SLC41A2, SLC4A3, SMAD4, SMG1P7, SPATA17, SPATS2, SPEG, SPIN1, SRRM4, ST5, STMN2, STO X2, STRA6, STXBP5L, SUPTH3H, SVEP1, SYDE1, SYNE1, SYNGR3, SYNJ2, SYT7, TAF6, TAFA2, TBCD, TBL1XR1, TENM3, TEX9, TGFB3, THUMPD3-AS1, TM6SF2, TMEM 117、TMEM175、TMEM189、TMEM191A、TMEM198B、TMEM214、TMEM230、TMEM88、T PRA1、TRAF3、TRAPPC12、TRIM16、TRIM6、TRIO、TRRAP、TSHZ3、TSPAN3、TTC39C -AS1, TTLL4, TTTY14, TUBB3, TUBB6, TUBGCP6, TXLNGY, UNC13A, UNK, USP10, USP28, USP36, VAX2, VPS29, VPS50, VPS53, WARS2, WASL, WDFY2, WDR19, WDR3 7、WDR4、WWOX、ZBTB18、ZC2HC1C、ZCCHC4、ZDHHC1、ZFAT、ZFP91、ZFP91-CNTF 、ZGPAT、ZNF195、ZNF202、ZNF236、ZNF320、ZNF382、ZNF394、ZNF420、ZNF423、The cryptic exon may be derived from at least a portion of ZNF429, ZNF43, ZNF48, ZNF527, ZNF571-AS1, ZNF583, ZNF594-DT, ZNF598, ZNF692, ZNF696, ZNF700, ZNF737, ZNF785, ZNF789, ZNF81, ZNF814, ZNF826P, ZNF875, ZNHIT1, ZRANB3, or ZSCAN12. In some embodiments, the known cryptic exon may be mutated by the insertion or deletion of nucleotides (e.g., the addition or deletion of any number of nucleotides not divisible by 3, e.g., preferably the addition or deletion of one or two nucleotides), and thus the cryptic exon is a frameshift-induced cryptic exon. In one example described herein, the cryptic exon is derived from the human AARS1 cryptic exon sequence but contains additional nucleotides, e.g., additional adenosine nucleotides, increasing its length from 87 to 88 nucleotides.
[0163] In some embodiments, the cryptic exon sequence has a sequence with at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO: 31. This sequence is derived from the cryptic exon sequence between exons 4 and 5 of the human AARS1 gene, but with an additional nucleotide inserted. In the example described herein, the additional nucleotide is adenosine. In another embodiment, the cryptic exon sequence is a synthetic exon sequence. The cryptic exon sequence can be designed using the Splice AI algorithm as described above (i.e., comprising a sequence such that the splice sites adjacent to the cryptic exon sequence have a probability score of at least 0.01, or at least 0.05, or at least 0.1 as determined by the Splice AI algorithm) and / or using "Algorithm 1" as described herein. Note that cryptic exon splice sites are expected to be weaker than constitutively spliced splice sites and therefore can be selected to have lower Splice AI scores. In some embodiments, the synthetic cryptic exon sequence encodes a portion of a transgene, where the portion of the transgene is modified to contain synonymous codons.
[0164] In some examples, the cryptic exon sequence has at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:31, SEQ ID NO:49, SEQ ID NO:51-64, SEQ ID NO:71, SEQ ID NO:77, SEQ ID NO:83, SEQ ID NO:88, SEQ ID NO:92, SEQ ID NO:120, SEQ ID NO:126, SEQ ID NO:132, SEQ ID NO:138, SEQ ID NO:14.
[0165] In some embodiments, a regulatory domain may include the following features from upstream to downstream: a splice donor site (i.e., a second splice donor site), A first portion of the intron region, a splice acceptor site (i.e., the first splice acceptor site), Cryptic exon sequence, a splice donor site (i.e., the first splice donor site), a second portion of the intron region, and a splice acceptor site (i.e., a second splice acceptor site), And, the binding domain for a splicing factor (ie, of the hnRNP family) can be within the first part of the intron region, the cryptic exon sequence, or the second part of the intron region.
[0166] In some embodiments, the construct may further comprise an exon sequence or region immediately upstream of the second splice donor site and / or an exon sequence or region immediately downstream of the second splice acceptor site. In some embodiments, the exon immediately upstream of the first splice acceptor site and / or the exon immediately downstream of the first splice donor site may encode at least a portion of the transgene sequence. In other embodiments, the exon immediately upstream of the first splice acceptor site and / or the exon immediately downstream of the first splice donor site may encode a peptide sequence that does not encode a portion of the transgene sequence.
[0167] In some embodiments, a regulatory domain may include the following features from upstream to downstream: the exon sequence immediately upstream of the splice donor site, a splice donor site (i.e., a second splice donor site), A first portion of the intron region, a splice acceptor site (i.e., the first splice acceptor site), Cryptic exon sequences embedded within intron regions, a splice donor site (i.e., the first splice donor site), A second portion of the intron region; and a splice acceptor site (i.e., a second splice acceptor site), and Exon sequence immediately downstream of the splice acceptor site.
[0168] The binding domain for a splicing factor (i.e., of the hnRNP family) may be within the first portion of the intron region, the cryptic exon sequence, or the second portion of the intron region. In some embodiments, the exon sequence immediately upstream of the splice donor site and the exon sequence immediately downstream of the splice acceptor site may code for a portion of the transgene sequence. In another embodiment, the exon sequence immediately upstream of the splice donor site and the exon sequence immediately downstream of the splice acceptor site may code for a peptide distinct from the protein produced by the transgene.
[0169] Constructs containing cryptic exon sequences according to "Design 1" In some embodiments of the construct, one or more exons encoding a transgene are all downstream of a cryptic exon sequence and / or a regulatory domain. Such constructs are referred to herein as "Design 1" constructs and are illustrated in FIG.
[0170] Exemplary constructs may include a regulatory domain and a transgene sequence (i.e., a transgene sequence that encodes a functional protein), where the regulatory domain is, from upstream to downstream: the exon sequence immediately upstream of the splice donor site, a splice donor site (i.e., a second splice donor site), A first portion of the intron region, a splice acceptor site (i.e., the first splice acceptor site), Cryptic exon sequences embedded within intron regions, a splice donor site (i.e., the first splice donor site), a second portion of the intron region, and a splice acceptor site (i.e., a second splice acceptor site), and Exon sequence immediately downstream of the splice acceptor site All of these features may be as described elsewhere herein. The binding domain for the hnRNP family splicing factor may be within the intron region, the first portion of the cryptic exon sequence, or the second portion of the intron region. The transgene may be encoded by at least a portion of the exon sequence downstream of the second splice acceptor site (i.e., downstream of the regulatory domain). In some embodiments, the transgene may be downstream of the regulatory domain. In some embodiments, the transgene may be encoded by the cryptic exon sequence. In such embodiments, the transgene may be encoded by the cryptic exon sequence and the exon sequence immediately upstream of the splice donor site and / or the exon sequence immediately downstream of the splice acceptor site.
[0171] The constructs of Design 1 may further include one or more optional features. A sequence containing a start codon upstream of the regulatory domain A premature termination codon (PTC) downstream of the cryptic exon sequence, which may be present in the transgene sequence (but out of frame) Additional intronic sequences downstream of the PTC A protease cleavage site or autocleavage site sequence (e.g., upstream of the transgene sequence and downstream of the regulatory domain)
[0172] In such an embodiment, the construct comprises the following features from upstream to downstream: an optional sequence including a start codon, the exon sequence immediately upstream of the splice donor site, a splice donor site (i.e., a second splice donor site), A first portion of the intron region, a splice acceptor site (i.e., the first splice acceptor site), a cryptic exon sequence (i.e., embedded within an intron region between the first splice acceptor site and the first splice donor site); a splice donor site (i.e., the first splice donor site), A second portion of the intron region, a splice acceptor site (i.e., a second splice acceptor site), and the exon sequence immediately downstream of the splice acceptor site, an optional proteolytic or autocleavage site; A transgene sequence, optionally including a PTC (i.e., the complete transgene sequence), Optional additional intron sequences (i.e. downstream of the transgene sequence and in an exon context).
[0173] The binding domain for a splicing factor of the hnRNP family can be within the first part of the intron region, the cryptic exon sequence, or the second part of the intron region.
[0174] In some embodiments, the start codon is upstream of the regulatory domain, in other embodiments, the start codon is within the regulatory domain (e.g., within the exon sequence immediately upstream of the second splice donor site), and in some embodiments, the start codon is within a cryptic exon sequence.
[0175] The above features may have any of the same features as described elsewhere herein. In some examples described herein, the exon immediately upstream of the splice donor site, the first part of the intron region, the cryptic exon sequence, the second part of the intron region, and the exon immediately downstream of the splice donor site are all derived from the human AARS1 gene or a modified variant thereof. In other examples, the exon immediately upstream of the splice donor site, the first part of the intron region, the cryptic exon sequence, the second part of the intron region, and the exon immediately downstream of the splice donor site are alternatively synthetic sequences. In some examples, the additional intron sequence and surrounding exon context are derived from RPS24. In some examples, the self-cleavage site is P2A. In some examples, the transgene encodes a diagnostic protein (e.g., mCherry, or Gaussia luciferase). In other examples, the transgene encodes a therapeutic protein (e.g., a splicing regulator such as a TDP-43 binding domain fused to RAVER 1, more particularly, a TDP-43 RNA binding domain fused to RAVER 1). In some examples described herein, the binding domain for the hnRNP family is TDP-43 and the splicing factor is TDP-43. In some embodiments, the binding domain is a functional binding domain or a mutant binding domain.
[0176] In some examples, the construct has a sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:25 or SEQ ID NO:47.
[0177] In some examples, the first portion of the intron region has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:30, or SEQ ID NO:70, or SEQ ID NO:76, or SEQ ID NO:82, or SEQ ID NO:119, or SEQ ID NO:125, or SEQ ID NO:131, or SEQ ID NO:137, or SEQ ID NO:143, or SEQ ID NO:149, or SEQ ID NO:155, or SEQ ID NO:179, or SEQ ID NO:185, or SEQ ID NO:191, or SEQ ID NO:197.
[0178] In some examples, the second portion of the intron region has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:32, or SEQ ID NO:72, or SEQ ID NO:78, or SEQ ID NO:84, or SEQ ID NO:121, or SEQ ID NO:127, or SEQ ID NO:133, or SEQ ID NO:139, or SEQ ID NO:145, or SEQ ID NO:151, or SEQ ID NO:157, or SEQ ID NO:181, or SEQ ID NO:187, or SEQ ID NO:193, or SEQ ID NO:199.
[0179] In some examples, the TDP-43 binding domain has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NOs: 1-9, or SEQ ID NO: 115, SEQ ID NO: 159, or SEQ ID NO: 160.
[0180] In some examples, the additional intron sequence has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:36.
[0181] In some examples, the cryptic exon sequence has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:31, SEQ ID NO:49, or SEQ ID NO:51-64, or SEQ ID NO:71, or SEQ ID NO:77, or SEQ ID NO:83, or SEQ ID NO:120, or SEQ ID NO:126, or SEQ ID NO:132, or SEQ ID NO:138, or SEQ ID NO:144, or SEQ ID NO:150, or SEQ ID NO:156.
[0182] In some examples, the self-cleavage site has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:34.
[0183] In some examples, the exon sequence immediately upstream of the first splice acceptor site has at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:29 or SEQ ID NO:48.
[0184] In some examples, the exon sequence immediately downstream of the first splice donor site has at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:33, or SEQ ID NO:50.
[0185] Constructs following "Design 2" In another embodiment, the cryptic exon sequence may code for at least a portion of the transgene. The cryptic exon sequence may code for an internal portion of a protein, an N-terminal portion of a protein, or a C-terminal portion of a protein. Such constructs are described herein as "Design 2" constructs and are illustrated in FIG. 2. The construct may include additional exon sequences that code for other portions of the transgene protein. In some embodiments, the construct may include other portions of the transgene sequence downstream of the cryptic exon and / or upstream of the cryptic exon. In some examples described herein, the transgene sequence is formed from at least three portions that together form the complete transgene sequence. In some embodiments, the transgene sequence may be split into two or more portions, or three or more portions, or four or more portions, or five or more portions, or six or more portions, or seven or more portions, or eight or more portions, or nine or more portions, or ten or more portions. The transgene may be split into portions such that the first donor acceptor site, the first splice acceptor site, the second splice acceptor site, and the second splice donor site have a splicing score of at least 0.01 as determined by the Splice AI algorithm or according to other splicing scores determined by the Splice AI algorithm described herein. In some embodiments, the transgene sequence may be modified to include synonymous codon sequences.
[0186] In some embodiments, a regulatory domain may include the following features from upstream to downstream: a splice donor site (i.e., a second splice donor site), A first portion of the intron region, a splice acceptor site (i.e., the first splice acceptor site), A cryptic exon sequence encoding at least a portion of a transgene; a splice donor site (i.e., the first splice donor site), A second portion of the intron region; and The splice acceptor site (i.e., the second splice acceptor site).
[0187] The binding domain for a splicing factor (ie, one of the hnRNP family) may be within the first part of the intron region, the cryptic exon sequence, or the second part of the intron region.
[0188] All of these features may be as described elsewhere herein.
[0189] Exemplary constructs may include a transgene and a regulatory domain, where the regulatory domain includes the following features from upstream to downstream: an exon immediately upstream of the splice donor site (i.e., optionally encoding a portion of the transgene); a splice donor site (i.e., a second splice donor site), A first portion of the intron region, a splice acceptor site (i.e., the first splice acceptor site), at least a portion of the transgene, embedded within an intron region, and optionally a cryptic exon sequence encoding a first portion or a second portion of the transgene; a splice donor site (i.e., the first splice donor site), A second portion of the intron region, a splice acceptor site (i.e., a second splice acceptor site), and An exon immediately downstream of the splice acceptor site, optionally encoding a portion of the transgene.
[0190] The binding domain for a splicing factor of the hnRNP family can be within the first part of the intron region, the cryptic exon sequence, or the second part of the intron region.
[0191] The Design 2 construct may also include one or more further optional features. A sequence containing a start codon upstream of the regulatory domain A premature termination codon (PTC) downstream of a cryptic exon sequence, which may be present in the transgene sequence Additional intronic sequences downstream of the PTC A protease cleavage site or self-cleavage site sequence (e.g., between two different transgene sequences)
[0192] Thus, an example construct may have the following characteristics from upstream to downstream: an optional start codon sequence, the exon immediately upstream of the splice donor site (i.e., optionally the location of the transgene, e.g., encoding the first part of the transgene); a splice donor site (i.e., a second splice donor site), a first portion of the intron region (i.e., or the first intron); a splice acceptor site (i.e., the first splice acceptor site), a cryptic exon sequence encoding at least a portion of the transgene (e.g., a second portion of the transgene), embedded within an intron region; a splice donor site (i.e., the first splice donor site), a second portion of the intron region (i.e., a second intron); and a splice acceptor site (i.e., a second splice acceptor site), and an exon immediately downstream of the splice acceptor site, optionally encoding a portion of the transgene (e.g., a third portion of the transgene); Optional additional intron sequences downstream of the transgene.
[0193] The binding domain for a splicing factor of the hnRNP family can be within the first part of the intron region, the cryptic exon sequence, or the second part of the intron region.
[0194] In some embodiments, the start codon is upstream of the regulatory domain, in other embodiments, the start codon is within the regulatory domain, and in some embodiments, the start codon is within a cryptic exon sequence.
[0195] These features may be as described elsewhere herein. In some examples described herein, the exon immediately upstream of the splice donor site, the first portion of the intron region, and the second portion of the intron region are derived from the human AARS1 gene or a modified variant thereof. In some examples, the exons encoding the transgene together encode a diagnostic protein (e.g., mCherry), or a therapeutic protein (e.g., a nuclease such as Cas 9), or a recombinase protein (e.g., Cre recombinase). In some examples, the optional intron sequence and the optional exon sequence downstream of one or more exons that together encode the transgene are derived from RPS24. In the examples described herein, the binding domain is for TDP-43 and the splicing factor (i.e., of the hnRNP family) is TDP-43.
[0196] In some examples, the construct has a sequence having at least 80%, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:68, SEQ ID NO:74, SEQ ID NO:80, SEQ ID NO:86, SEQ ID NO:90, SEQ ID NO:117, SEQ ID NO:123, SEQ ID NO:129, SEQ ID NO:135, SEQ ID NO:141, SEQ ID NO:147, SEQ ID NO:153.
[0197] In some examples, the first portion of the intron region has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:30, or SEQ ID NO:70, or SEQ ID NO:76, or SEQ ID NO:82, or SEQ ID NO:119, or SEQ ID NO:125, or SEQ ID NO:131, or SEQ ID NO:137, or SEQ ID NO:143, or SEQ ID NO:149, or SEQ ID NO:155, or SEQ ID NO:179, or SEQ ID NO:185, or SEQ ID NO:191, or SEQ ID NO:197.
[0198] In some examples, the second portion of the intron region has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:32, or SEQ ID NO:72, or SEQ ID NO:78, or SEQ ID NO:84, or SEQ ID NO:121, or SEQ ID NO:127, or SEQ ID NO:133, or SEQ ID NO:139, or SEQ ID NO:145, or SEQ ID NO:151, or SEQ ID NO:157, or SEQ ID NO:181, or SEQ ID NO:187, or SEQ ID NO:193, or SEQ ID NO:199.
[0199] In some examples, the TDP-43 binding domain has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:1-9, or SEQ ID NO:115, or SEQ ID NO:159 or SEQ ID NO:160.
[0200] In some examples, the additional intron sequence has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:36.
[0201] In some examples, the cryptic exon is placed in a prime editing vector. In some embodiments, the prime editing vector uses a H840A mutant S. pyogenes Cas9. In some embodiments, a first portion of the intron has a sequence with at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO: 191, and has a sequence with at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO: 193.
[0202] Constructs in which the regulatory domain is regulated by splicing of a single regulatory intron In an embodiment of the third aspect or the first aspect, Start codon, A regulatory domain that includes: a first splice donor site and a first acceptor donor site defining a single regulatory intron; a binding domain for a splicing factor of the heterologous nuclear ribonucleoprotein (hnRNP) family located within 150 nucleotides of the first splice donor site or the first splice acceptor site and / or located between the first splice donor site and the first splice acceptor site; and A transgene sequence (i.e., a transgene sequence that encodes a functional protein) Including, (i) when localized in cells depleted of splicing factors, first splice acceptor site and / or first donor site splicing is not repressed and the single regulatory intron is spliced out to produce a functional protein from the transgene sequence (i.e., in the mRNA product of the construct where a functional protein is encoded by the complete, uninterrupted transgene sequence, a functional product is produced); (ii) when localized in cells that are not depleted of splicing factors, the single regulatory intron is not spliced, or is not properly spliced, such that a functional protein is not produced from the transgene sequence (i.e., a functional product is produced in the mRNA product of the construct when a functional protein is encoded by the complete, uninterrupted transgene sequence); A construct configured to:
[0203] Such constructs are referred to herein as "Design 3" constructs and are illustrated in Figure 3. Design 3 constructs are configured such that the intron is properly spliced only in cells that have nuclear depletion of hnRNP splicing factors. This has the effect that in cells that have depletion of hnRNP splicing factors, no part of the intron sequence is present in the mRNA product of the construct. In contrast, in cells that do not have nuclear depletion of hnRNP splicing factors, the intron is not spliced or is not properly spliced. This has the effect that at least a portion of the intron is present in the mRNA product of the construct, thereby interrupting the transgene sequence, resulting in a non-functional protein, and / or essential parts of the transgene sequence are not included in the mature mRNA (see, e.g., Figures 3A-D). Additionally or alternatively, inclusion of all or a portion of the intron in the mature mRNA and / or absence of a portion of the transgene sequence in the mature mRNA induces a frameshift, and the transgene contains a premature stop codon that is in frame with the start codon in the mRNA product of the construct only if at least a portion of the single regulatory intron is incorporated into the mRNA product of the construct and / or if a portion of the transgene sequence is not included in the mature mRNA. Additionally or alternatively, the portion of the single regulatory intron incorporated into the mRNA product contains a premature stop codon that is in frame with the start codon in the mRNA product of the construct (see, e.g., Figures 3D and E). Additionally or alternatively, the portion of the single regulatory intron incorporated into the mRNA product contains a disruptive amino acid sequence.
[0204] In some embodiments, at least a portion of the transgene sequence is downstream of a single regulatory intron. In some embodiments, the complete transgene sequence is downstream of a regulatory domain. In some embodiments, a portion of the transgene sequence is upstream of a single regulatory intron and a portion of the transgene sequence is downstream of a single regulatory intron. Other embodiments of the transgene sequence are as described herein. The transgene may be split into portions such that the first donor acceptor site portion and the first splice acceptor site have a splicing score of at least 0.01 as determined by the Splice AI algorithm or according to other splicing scores determined by the Splice AI algorithm as described herein. In some embodiments, the transgene sequence may be modified to include synonymous codon sequences.
[0205] In some embodiments, the binding domain for the splicing factor of the hnRNP family is within a single regulatory intron. In some embodiments, the binding domain for the splicing factor of the hnRNP family is upstream of the single regulatory intron (i.e., within the exonic sequence upstream of the first splice donor site). In some embodiments, the binding domain for the splicing factor of the hnRNP family is downstream of the single regulatory intron (i.e., within the exonic sequence downstream of the first splice acceptor site). In some examples, the binding domain is a TDP-43 binding domain and the hnRNP splicing factor is TDP-43. Other aspects of the hnRNP binding domain and / or the TDP-43 binding domain are described elsewhere herein. The first splice donor site, the first splice acceptor site, and other aspects of the transgene are as described herein.
[0206] In some embodiments, the first splice acceptor site and / or the first splice donor site has a splice score of 0.01 or greater as determined by the Splice AI algorithm. In some embodiments, the first splice acceptor site and / or the first splice donor site has a splice score of 0.05 or greater, or at least 0.1 or greater, or at least 0.2 or greater, or at least 0.3 or greater, or at least 0.4 or greater, or at least 0.5 or greater, or at least 0.6 or greater, or at least 0.7 or greater, or at least 0.8 or greater, or at least 0.9 or greater as determined by the Splice AI algorithm.
[0207] In some examples, the construct has a sequence having at least 80% sequence identity to SEQ ID NO:95.
[0208] In some examples, the single regulatory intron sequence has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:97, or SEQ ID NO:163, or SEQ ID NO:168, or SEQ ID NO:171, or SEQ ID NO:175.
[0209] In some examples, the exon sequence upstream of the first splice donor site has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:29.
[0210] In some examples, the exon sequence downstream of the first splice acceptor site has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:33.
[0211] In some examples, the TDP-43 binding domain has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:1-9, or SEQ ID NO:115, or SEQ ID NO:159 or SEQ ID NO:160.
[0212] In some examples, the additional intron sequence has a sequence having at least 80% sequence identity, or at least 85%, or at least 90%, or at least 95%, or at least 100% sequence identity to SEQ ID NO:36.
[0213] In some embodiments, in cells without depletion of hnRNP splicing factors, no splicing occurs and the intron is retained in the mRNA product of the construct. The construct is configured such that in cells without depletion of hnRNP splicing factors, the complete single regulatory intron is incorporated in the mRNA product of the construct (i.e., splicing of the first splice donor site and / or the first splice acceptor site is suppressed), but in cells with depletion of hnRNP splicing factors, the complete single regulatory intron is not incorporated in the mRNA product of the construct (i.e., splicing of the first splice donor site and / or the first splice acceptor site is not suppressed).
[0214] In some instances, the regulatory domain comprises: a splice donor site (i.e., the first splice donor site), a single regulatory intron, and Splice acceptor site (i.e., the first splice acceptor site) Includes.
[0215] In some examples, the construct includes a transgene sequence (i.e., a transgene sequence that encodes a functional protein) and a regulatory domain, the regulatory domain comprising (from upstream to downstream): an optional coding sequence, including a start codon; exon sequence (i.e., immediately upstream of the splice donor site); a splice donor site (i.e., the first splice donor site), A single regulatory intron, a splice acceptor site (i.e., the first splice donor site), and Exon sequence (i.e., immediately downstream of the splice acceptor site) Includes.
[0216] The transgene sequence may be completely downstream of the regulatory domain. In other embodiments, the transgene sequence may be encoded by an exon sequence. In some examples, the construct further comprises an additional intron sequence downstream of the exon sequence. The binding domain for the hnRNP splicing factor may be within the single regulatory intron, upstream of the single regulatory intron in the exon sequence immediately upstream of the splice donor site, or downstream of the single regulatory intron immediately downstream of the splice acceptor site.
[0217] In some instances, the constructs are (from upstream to downstream): an optional coding sequence, including a start codon; Exon sequences (i.e., optionally encoding at least a portion of the transgene); a splice donor site (i.e., the first splice donor site), A single regulatory intron, a splice acceptor site (i.e., the first splice donor site), and Exon sequence, a proteolytic or autocleavage site, and Complete transgene sequence Includes.
[0218] In some examples, the construct further comprises additional intron sequences downstream of the exon sequence. The binding domain for the hnRNP splicing factor can be within the single regulatory intron, upstream of the single regulatory intron in the exon sequence immediately upstream of the splice donor site, or downstream of the single regulatory intron immediately downstream of the splice acceptor site.
[0219] In some instances, the constructs are (from upstream to downstream): an optional coding sequence, including a start codon; an exon sequence (i.e., encoding the first part of the transgene); a splice donor site (i.e., the first splice donor site), A single regulatory intron, a splice acceptor site (i.e., the first splice donor site), and Exon sequence (i.e., encoding the second part of the transgene) Includes.
[0220] In some examples, the construct further comprises additional intron sequences downstream of the exon sequence. The binding domain for the hnRNP splicing factor can be within the single regulatory intron, upstream of the single regulatory intron in the exon sequence immediately upstream of the splice donor site, or downstream of the single regulatory intron immediately downstream of the splice acceptor site.
[0221] In some instances, the constructs are (from upstream to downstream): an optional coding sequence, including a start codon; Exon sequences (i.e., encoding at least a portion of the transgene); a splice donor site (i.e., the first splice donor site), A single regulatory intron, a splice acceptor site (i.e., the first splice acceptor site), and Exon sequences (i.e., encoding at least a portion of the transgene) Includes.
[0222] In another embodiment, inappropriate or alternative splicing occurs in cells without nuclear depletion of hnRNP splicing factors. In such an embodiment, the construct and regulatory domain may include an alternative splice donor site and / or an alternative splice acceptor site. In some embodiments, the alternative splice donor site may be upstream of the first splice donor site or within a single regulatory intron sequence (i.e., between the first splice donor site and the first splice acceptor site). In some embodiments, the alternative splice acceptor site may be downstream of the first acceptor site or within a single regulatory intron sequence (i.e., between the first splice donor site and the first splice acceptor site). The alternative splice acceptor site and / or alternative splice donor site may be any splice donor site having a median Splice AI score of at least 0.01 (Splice AI score 99.8 percentile), or at least 0.05, or at least 0.1, or at least 0.5, or at least 0.9, as determined by the Splice AI algorithm as described elsewhere herein. The alternative splice acceptor site and / or alternative splice donor site is not inhibited by hnRNP splicing factors (e.g., TDP-43). In some embodiments, the alternative splice acceptor site and / or alternative splice donor site is further away from the binding domain than the first splice acceptor site and the first splice donor site. In some embodiments, the alternative splice acceptor site and / or the alternative splice donor site may be at least 20 nucleotides away from the binding domain, or at least 50 nucleotides away from the binding domain, or at least 100 nucleotides away from the binding domain, or at least 150 nucleotides away from the binding domain, or at least 200 nucleotides away from the binding domain.
[0223] In some embodiments, the construct is configured such that in cells without nuclear depletion of hnRNP splicing factors (i.e., splicing of the first splice donor site or the first splice acceptor site is inhibited), at least a portion of the single regulatory intron is incorporated into the mRNA product, but in cells with nuclear depletion of hnRNP splicing factors (i.e., splicing of the first splice donor site or the first splice acceptor site is not inhibited), no portion of the single regulatory intron is incorporated into the mRNA product of the construct.
[0224] Additionally or alternatively, the construct is configured such that in cells where there is no nuclear depletion of hnRNP splicing factors (i.e., first splice donor site or first splice acceptor site splicing is inhibited), at least a portion of the transgene sequence is not included in the RNA product, but in cells where there is nuclear depletion of hnRNP splicing factors (i.e., first splice donor site or first splice acceptor site splicing is not inhibited), all of the transgene sequence is present in the mRNA product of the construct.
[0225] In cells with nuclear depletion of hnRNP splicing factors, the intron is fully spliced and removed to provide a complete, uninterrupted transgene sequence that is in frame with the start codon and does not contain a premature stop codon in frame with the start codon in the mRNA product of the construct, and a functional protein is produced.
[0226] In some instances, the regulatory domain comprises: a splice donor site (i.e., the first splice donor site), a single regulatory intron defined by a first splice donor site and a first splice acceptor site; a splice acceptor site (i.e., the first splice acceptor site), and an alternative splice donor and / or alternative splice acceptor site, which may be located within the single regulatory intron upstream of the splice donor site or downstream of the splice acceptor site; Includes.
[0227] In some instances, the constructs are (from upstream to downstream): an optional coding sequence, including a start codon; exon sequence (i.e., immediately upstream of the splice donor site); a splice donor site (i.e., the first splice donor site), a single regulatory intron (i.e., defined by a first splice donor site and a first splice acceptor site); a splice acceptor site (i.e., the first splice acceptor site), and Exon sequence (immediately downstream of the splice acceptor site) Includes.
[0228] The binding domain for the hnRNP splicing factor may be within the single regulatory intron, or upstream or downstream of the single regulatory intron (i.e., within exonic sequences adjacent to the single regulatory intron). The transgene may be entirely downstream of the regulatory domain, or may be encoded by exonic sequences upstream and downstream of the single regulatory intron. The alternative splice acceptor site may be within the single regulatory intron or downstream of the first splice acceptor site. The alternative splice donor site may be within the single regulatory intron or upstream of the first splice donor site. In some examples, the construct further comprises an additional intron sequence downstream of the exonic sequence.
[0229] In some instances, the constructs are (from upstream to downstream): an optional coding sequence, including a start codon; Exon sequence, a splice donor site (i.e., the first splice donor site), a single regulatory intron (i.e., defined by a first splice donor site and a first splice acceptor site); a splice acceptor site (i.e., the first splice acceptor site), Exon sequence, an optional proteolytic or autocleavage site; Complete transgene sequence Binding domains for hnRNP splicing factors, which may be within the single regulatory intron, or upstream or downstream of the single regulatory intron (i.e., within exonic sequences adjacent to the single regulatory intron). The alternative splice acceptor site may be within the single regulatory intron or downstream of the first splice acceptor site. The alternative splice donor site may be within the single regulatory intron or upstream of the first splice donor site. In some examples, the construct further comprises an additional intron sequence downstream of the exon sequence.
[0230] In some instances, the constructs are (from upstream to downstream): an optional coding sequence, including a start codon; an exon sequence (i.e., encoding the first part of the transgene); a splice donor site (i.e., the first splice donor site), a single regulatory intron (i.e., defined by a first splice donor site and a first splice acceptor site); a splice acceptor site (i.e., the first splice acceptor site), and Exon sequence (encoding the second part of the transgene) Includes.
[0231] The binding domain for the hnRNP splicing factor may be within the single regulatory intron, or upstream or downstream of the single regulatory intron (i.e., within exon sequences adjacent to the single regulatory intron). The alternative splice acceptor site may be within the single regulatory intron or downstream of the first splice acceptor site. The alternative splice donor site may be within the single regulatory intron or upstream of the first splice donor site. In some examples, the construct further comprises additional intron sequences downstream of the exon sequences.
[0232] Optional Features In all of the above embodiments, the single regulatory intron, or at least a portion of the single regulatory intron (i.e., the portion of the single regulatory intron that is not properly spliced in cells without nuclear depletion of hnRNP splicing factors and therefore not included in the mRNA product), may contain a premature termination codon (PTC) that is in frame with the start codon. This has the effect that in cells without nuclear depletion of hnRNP splicing factors, at least a portion of the intron is present in the mRNA product of the construct and encounters the PTC, but in cells with nuclear depletion of hnRNP splicing factors, this intron is not present in the mRNA product of the construct and does not encounter the PTC.
[0233] In some embodiments, at least a portion of the transgene sequence downstream of the single regulatory intron contains a PTC that is out of frame with the start codon if the intron is properly spliced, but is in frame with the start codon if the intron is not spliced or is not properly spliced.
[0234] In some embodiments, a single regulatory intron of length not divisible by 3, i.e., incorporation of the single regulatory intron into the mRNA product of the construct introduces a frameshift. In such embodiments, the construct may include a PTC downstream of the regulatory domain, configured such that when a portion of the single regulatory intron is not incorporated into the construct mRNA product (i.e., when the intron is "properly" spliced), the PTC will be out of frame with the start codon, and when the single regulatory intron is incorporated into the construct mRNA product (i.e., when the intron is not spliced), the PTC will be in frame with the start codon.
[0235] In some embodiments, the single regulatory intron comprises a disruptive amino acid sequence.
[0236] In some embodiments, the construct further comprises an additional intron sequence at least 40 nucleotides downstream of the PTC, which provides for the attachment of the EJC complex and promotes NMD of the mRNA when the PTC is in-frame with the start codon.
[0237] In some embodiments, i.e., those in which the transgene is entirely downstream of the regulatory domain, the construct may further comprise a protease cleavage site or an autocleavage site.
[0238] vector Disclosed herein is a vector comprising a construct according to any aspect or embodiment disclosed herein. In some embodiments, the vector is a DNA vector. In some embodiments, the vector is a circular vector, for example in the form of a plasmid. In some embodiments, the vector is a single-stranded or double-stranded vector, for example a double-stranded vector.
[0239] In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is a retrovirus, lentivirus, adenovirus (AV), or adeno-associated virus (AAV), a chimeric AAV vector, or a herpes simplex virus vector. The viral vector may be derived from any suitable serotype or subgroup. The viral vector may be a human viral vector or a non-human viral vector. In some embodiments, the AAV vector is a recombinant AAV vector.
[0240] In some embodiments, the viral vector comprises one or more regions comprising the constructs described herein and inverted terminal repeat (ITR) sequences flanking the construct. In some embodiments, the sequences are operably linked to a promoter. Any suitable promoter can be used. In some examples, the promoter is a cytomegalovirus (CMV) promoter, a CMV enhancer, a CAG promoter, an SV40 promoter, a JeT promoter, a PGK promoter, and a chicken β-actin (CBA) promoter, an eEF1A promoter, a synapsin promoter, a ChAT promoter, a TRE promoter, a calcium / calmodulin-dependent protein kinase II promoter, a tubulin αI promoter, a neuron-specific enolase promoter, or a platelet-derived growth factor β chain promoter, or a fusion of the above.
[0241] In some embodiments, the promoter is a tissue-specific (e.g., CNS-specific) promoter. In some embodiments, the neuron-specific promoter is derived from neuron-specific enolase (NSE) (e.g., see EMBL HSEN02, X51956); aromatic amino acid decarboxylase (MDC) promoter; neurofilament promoter (e.g., see GenBank HUMNFL, L04147); synapsin promoter (e.g., see GenBank HUMSYNIB, M55301); athy-1 promoter; serotonin receptor promoter (e.g., see GenBank S62283); tyrosine hydroxylase promoter (TH); L7 promoter; DNMT promoter; enkephalin promoter; myelin basic protein (MBP) promoter; Ca2+ calmodulin-dependent protein kinase 11-alpha (CamKIM) promoter; CMV enhancer / platelet-derived growth factor-p promoter.
[0242] In some embodiments, the vector comprises a polyadenylation site downstream of the construct. In some embodiments, the vector may comprise a post-transcriptional regulatory element (PRE) downstream of the construct.
[0243] Pharmaceutical Compositions In one aspect of the invention, there is provided a pharmaceutical composition comprising a construct or vector as disclosed herein and a pharma- ceutically acceptable excipient.
[0244] system In one aspect of the invention there is provided a system comprising the cells and constructs, vectors or pharmaceutical compositions described herein, the system comprising: (i) When the cell nucleus is depleted of hnRNP family splicing factors, the system produces functional proteins; (ii) If the hnRNP family of splicing factors is not depleted from the cell nucleus, the system does not produce functional proteins. It is configured as follows.
[0245] In this system, cells selectively express functional proteins only when the splicing factors are depleted from the nucleus (e.g., in diseased cells), but when the splicing factors are not depleted from the nucleus (e.g., in healthy cells), no functional protein is produced.
[0246] The cell may be any suitable cell. In some embodiments, the cell is a mammalian cell, more preferably a human cell. In preferred embodiments, the cell has a nuclear depletion of hnRNP splicing factors (e.g., depletion of TDP-43). In some embodiments, the cell is a brain cell. In some embodiments, the cell is a neuron or a neural cell. In some embodiments, the cell is a microglial cell or an astrocyte. In some embodiments, the cell is a muscle cell.
[0247] Constructs, vectors and pharmaceutical compositions for use in therapy and related methods In a further aspect, there is provided a construct as described herein, a vector as described herein, or a pharmaceutical composition as described herein for use in therapy.
[0248] Also provided herein is a construct as described herein, a vector as described herein, or a pharmaceutical composition as described herein for use in treating a disease associated with depletion of hnRNP family splicing factors. In some embodiments, the disease is a neurodegenerative disease. In some embodiments, the disease is a muscle disease or muscle disorder, such as a neuromuscular disease.
[0249] In a further aspect, there is provided a construct as described herein, a vector as described herein, or a pharmaceutical composition as described herein for use in treating a disease associated with depletion of TDP-43. In some embodiments, the disease is a neurodegenerative disease. In some embodiments, the disease is a muscle disease, e.g., a neuromuscular disease.
[0250] In some embodiments, the disease (e.g., a neurodegenerative disease) is selected from amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), Parkinson's disease, Alzheimer's disease, inclusion body myopathy, or Perry syndrome.
[0251] In a further aspect, there is provided a construct as described herein, a vector as described herein, or a pharmaceutical composition as described herein for use in treating a neuromuscular disease associated with depletion of a splicing factor of the hnRNP family. In some embodiments, the splicing factor of the hnRNP family is TDP-43.
[0252] The constructs, vectors or pharmaceutical compositions described herein may be administered using any suitable method.
[0253] In some embodiments, treating the disease comprises contacting a cell with a construct, vector, or pharmaceutical composition disclosed herein. (i) in cells that are nuclear-depleted of splicing factors (i.e., when the cell nucleus is depleted of splicing factors), the cells produce functional proteins; (ii) In cells where the nucleus is not depleted of splicing factors (i.e., the cell nucleus is depleted of splicing factors), the cells do not produce functional proteins.
[0254] Also disclosed herein is a method for treating a disease associated with depletion of hnRNP splicing factors (e.g., a neurodegenerative disease or a muscular disease, e.g., associated with depletion of TDP-43), comprising contacting a cell with a construct, vector, or pharmaceutical composition disclosed herein. In a preferred embodiment, the disease is associated with depletion of TDP-43. In the method of treatment, (i) In cells in which the nucleus is depleted of splicing factors, the cells produce functional proteins, (ii) In cells where the nucleus is not depleted of splicing factors, the cells do not produce functional proteins.
[0255] Also disclosed herein is a construct described herein, a vector described herein, or a pharmaceutical composition described herein for use in the manufacture of a medicament.The medicament can be used for the treatment of a disease associated with the depletion of hnRNP splicing factors (e.g., a neurodegenerative disease or a neuromuscular disease, e.g., associated with the depletion of TDP-43), the treatment comprising contacting a cell with a construct, vector, or pharmaceutical composition disclosed herein.In a preferred embodiment, the disease is associated with the depletion of TDP-43.
[0256] In this treatment method, (i) In cells in which the nucleus is depleted of splicing factors, the cells produce functional proteins, (ii) In cells where the nucleus is not depleted of splicing factors, the cells do not produce functional proteins.
[0257] In a further aspect, there is provided the use of the construct, the vector, or the pharmaceutical composition disclosed herein in a method for selectively producing functional protein in diseased cells with nuclear depletion of hnRNP family splicing factors.In a preferred embodiment, the hnRNP family splicing factor is TDP-43.The cell may be in vivo or in vitro.
[0258] In vitro systems Also herein, Start codon, A regulatory domain that includes: a first splice acceptor site and a first splice donor site; a binding domain for a splicing factor of the heterogeneous nuclear ribonucleoprotein (hnRNP) family located within 150 nucleotides of the first splice donor site or the first splice acceptor site and / or located between the first splice donor site and the first splice acceptor site; and A transgene sequence (i.e., a transgene sequence that encodes a functional protein) Also disclosed is a construct comprising: (i) when placed in an in vitro system in which splicing factors are depleted, splicing of the first splice acceptor site and the first donor site is not inhibited and a functional protein is produced from the transgene sequence (i.e., a functional protein is produced from the transgene sequence in the mRNA product of a construct in which the functional protein is encoded by a complete, uninterrupted transgene sequence); (ii) when placed in an in vitro system in which splicing factors are not depleted, splicing of the first splice acceptor site and / or the first donor site is inhibited and no functional protein is produced from the transgene sequence (i.e., a functional protein is produced from the transgene sequence in the mRNA product of a construct in which a functional protein is encoded by an intact, uninterrupted transgene sequence); It is structured as follows.
[0259] The in vitro system must contain components that allow for transcription, splicing and translation. In some embodiments, these components are provided by the cells.
[0260] In some embodiments, the use of the construct in an in vitro system is provided for selectively producing functional protein in the absence of hnRNP family splicing factor.In a preferred embodiment, the hnRNP family splicing factor is TDP-43. EXAMPLES
[0261] Design 1 Example 1 An exemplary construct of the invention has a structure according to "Design 1" as shown in Figure 1. The Design 1 construct comprises a regulatory domain comprising an intron sequence comprising a TDP-43 binding domain and a cryptic exon sequence embedded within the intron region. The cryptic exon sequence is defined by a first splice acceptor site and a first splice donor site (i.e., the "cryptic splice site"), and the intron region is defined by a second splice donor site and a second splice acceptor. The construct further comprises a transgene sequence downstream of the regulatory domain encoding a protein (e.g., a functional or diagnostic protein).
[0262] In this construct, binding of TDP-43 to the binding domain suppresses splicing of the cryptic splice acceptor and / or cryptic splice donor sites. Due to the role of exon definition in determining splicing, suppression of one cryptic splice site may also suppress other cryptic splice sites. This results in the absence of cryptic exon sequences in the mRNA product of the construct in healthy cells (i.e., splicing factors are not depleted). In contrast, in diseased cells (i.e., splicing factors are depleted), the cryptic exon sequences are present in the mRNA product of the construct. This can be used to control the expression of downstream transgenes.
[0263] Example 1A In this example, the regulatory domain is based on a modified portion of the AARS1 sequence between exons 4 and 5, and the transgene is a sequence encoding mCherry (a red fluorescent protein).
[0264] The first example construct (SEQ ID NO:25) includes the following features, listed 5' to 3': A sequence that encodes the start codon A regulatory domain (SEQ ID NO:26) comprising: 3' exon sequence (here based on exon 4 of AARS1) A cryptic exon sequence embedded in an intron region, the cryptic exon sequence being defined by a splice acceptor site and a splice donor site, at least one of which is repressed by TDP-43 binding. The intron region itself is defined by a second splice donor site and a second splice acceptor site. The intron region includes a first intron portion upstream of the cryptic exon sequence and a second intron portion downstream of the cryptic exon sequence, and includes a TDP-43 binding domain. A complete intron sequence includes, from 5' to 3', the first intron portion, the cryptic exon, and the second intron portion, if the cryptic exon is not included. 5' exon sequence (here based on exon 5 of AARS1 with a single point mutation) - Sequence of the protease cleavage site or autocleavage site (here, the P2A autocleavage site) The complete transgene sequence (here, encoding mCherry) Further intronic sequences including downstream introns within an exonic context (here, based on human RPS24)
[0265] In this example, the regulatory domain is based on a modified AARS1 gene. Compared to the native sequence, a portion of the intron region was removed (reduced from 6.5 kb to 0.6 kb), and the intron region only contains the cryptic exon, the cryptic exon sequence, and the region adjacent to the cryptic splice site (i.e., forming the first splice acceptor site and the first splice donor site in the construct), and the constitutive splice site (i.e., forming the second splice acceptor site and the second splice donor site). In addition, the TG repeat region (i.e., the TDP-43 binding sequence) was slightly modified to achieve more efficient gene synthesis, and an "AA" was inserted in the middle of the TG repeat to make it less repetitive. Next, the 5' exon sequence based on exon 5 of AARS1 was mutated to avoid a premature stop codon. The cryptic exon sequence was also modified to include an additional adenosine in the sequence compared to that found in nature. This resulted in a total length of the cryptic exon (CE) of 88 nucleotides (instead of 87 nucleotides), which is not divisible by 3. This had the effect that the cryptic exon, when included in the mRNA product of the construct, could perform a frameshifting function. In diseased cells, inclusion of the cryptic exon sequence means that the premature stop codon downstream of the cryptic exon is no longer in-frame with the start codon, resulting in the production of a functional protein. In healthy cells, the cryptic exon sequence is not included and the premature stop codon is in-frame with the start codon, so it will encounter a premature stop codon. This results in the formation of a truncated, non-functional protein with no amino acid similarity to mCherry due to the frameshift.
[0266] In this example, the cryptic splice acceptor site (i.e., the first acceptor splice site) has a splice score of 0.05 as determined by the Splice AI algorithm, and the cryptic splice donor site (i.e., the first splice donor site) has a splice score of 0.19 as determined by the Splice AI algorithm.
[0267] The sequences used in the example constructs are shown in the table below.
[0268] TIFF2025513046000003.tif235162TIFF2025513046000004.tif246163TIFF2025513046000005.tif246163TIFF2025513046000006.tif179163
[0269] The constructs of the above examples were incorporated into a plasmid, which, in addition to the above features, further incorporated an enhancer sequence and a promoter sequence (CMV enhancer and CMV promoter, respectively) upstream of the construct, and a polyadenylation site (here, the SV40 late polyA site) downstream of the construct.
[0270] The plasmid of this example also incorporates sequence elements for propagation in bacteria, namely an origin of replication (in this case the ColE1 origin) and an antibiotic selection gene (in this case AmpR for ampicillin resistance), features which are not relevant for use in mammalian cells and can therefore be omitted.
[0271] The sequence of the plasmid was the sequence shown in SEQ ID NO:41 below. TIFF2025513046000007.tif233161TIFF2025513046000008.tif160161
[0272] Example 1B The construct was prepared exactly as described in Example 1A, except that the transgene was a gene encoding Gaussia princeps luciferase (Gluc), codon-optimized for mammalian cells, and two methionines were changed to leucines. The sequence is shown below.
[0273] TIFF2025513046000009.tif39163TIFF2025513046000010.tif152164
[0274] Example 1C A construct was then prepared as described in Example 1A, except that the transgene was a gene encoding a TDP-43-based fusion protein, i.e., TDP-43 / Raver1. An RNA-binding-deficient mutant of the same construct was also generated, in which two phenylalanines in the RNA recognition domain 1 of TDP-43 were mutated to leucines. Both sequences are shown below.
[0275] TIFF2025513046000011.tif51163TIFF2025513046000012.tif250163TIFF2025513046000013.tif37162
[0276] Example 1D It was found that the construct of Example 1A could be edited using sequences different from both the cryptic exon and the flanking exons, where the cryptic exon and flanking exon sequences were replaced with sequences encoding fragments of the Streptococcus pyogenes Cas9 enzyme. The construct was otherwise prepared as described in Example 1A and contained the mCherry transgene sequence.
[0277] To aid in the design of this construct, a splicing prediction computer program (Splice AI, see https: / / github.com / Illumina / SpliceAI) was used to identify sequences with high splicing potential. Cryptic exon sequences were identified that had a variety of synonymous codons and moderate Splice AI scores for the cryptic donor and acceptor (i.e., >0.01 and <0.5). There were no other predicted splice sites within the cryptic exon. For example, in the following synthetic sequence, the cryptic acceptor had a score of 0.31 and the cryptic donor had a score of 0.42:
[0278] TIFF2025513046000014.tif110163TIFF2025513046000015.tif171164
[0279] Example 1E To further explore whether different regulatory sequences could be used for the reporter of Design 1, a high-throughput assay was designed to test the splicing behavior of a number of different synthetic cryptic exons in the regulatory upstream sequence of Design 1 format. To enable this, a library of plasmids featuring different cryptic sequences was prepared. Each cryptic exon encoded the same amino acid sequence (fragment of Cas9) but featured a different combination of synonymous codons. The surrounding sequences were similar to the upstream regulatory sequence of Example 1D.
[0280] We then performed high-throughput RNA sequencing to determine the splicing behavior of each cryptic exon sequence. The results showed that a variety of sequences showed increased cryptic exon expression upon TDP-43 knockdown, with the majority of them having no detectable leakage in normal cells (i.e., cells without TDP-43 knockdown). For some of these sequences, the Splice AI scores assigned to each cryptic splice site and the inclusion rates of cryptic exons upon TDP-43 knockdown (KD) are detailed below. Although there are some examples with low inclusion rates, they are still sufficient to selectively obtain good protein expression in diseased cells.
[0281] TIFF2025513046000016.tif190168TIFF2025513046000017.tif152167
[0282] As demonstrated in Example 1E, a variety of synthetic cryptic exon sequences with a broad range of Splice AI predicted splice scores were found to enable cryptic exon splicing.
[0283] Comments on the Design 1 construct All of Examples 1A-1D were constructs according to "Design 1" as shown in Figure 1. Example 1E had an AARS1-based intron sequence like Examples 1A-1D, but had no downstream transgene and instead had a 12-nt barcode sequence.
[0284] The constructs of Design 1 have many advantages. The main advantage of this design is that it can be very easily modified to control the expression of different proteins by simply including another complete transgene or protein coding sequence downstream of the regulatory sequences. This is shown in Examples 1A-1C above. A variety of cryptic exon and intron sequences can be used, as shown in Examples 1D-1E.
[0285] The above "Design 1" example includes many preferred or optional features.
[0286] For example, the example "Design 1" construct above includes a P2A cleavage site downstream of the cryptic exon, but this feature is not essential as some transgenes may function correctly with additional N-terminal sequence encoded by an upstream regulatory domain. The presence of a cleavage site (e.g., P2A, etc.) has the advantage of ensuring that the transgene can be expressed without additional N-terminal sequence, which in some cases may improve the function of the protein produced from the transgene. It is envisioned that the P2A cleavage site can be replaced with a variety of alternative proteolytic or autocleavage sites that provide the same advantages, as described above.
[0287] Furthermore, while each of the Design 1 constructs above includes an intron region based on AARS1, below (e.g., Examples 2A-2C) we show that different intron sequences, including cryptic exons, can be designed that are not based on existing sequences. Indeed, the synthetic intron / cryptic exon sequences of Examples 2A-2C can be directly used as regulatory domains in the Design 1 constructs, since the cryptic exons cause frameshifts. Thus, the intron sequences of the Design 1 constructs are not limited to sequences derived from AARS1, but can be any suitable intron sequence that may or may not be based on a naturally occurring cryptic exon / intron context.
[0288] In the above example, the construct of "Design 1" contains a premature stop codon (PTC) in frame with the start codon if the mRNA product does not contain a cryptic exon sequence, but out of frame with the start codon if the mRNA product contains a cryptic exon sequence. The PTC sequence is any sequence selected from TGA, TAA, and TAG, downstream and in frame with the start codon. However, if the cryptic exon itself contains a start codon, the construct does not need to contain a premature stop codon. This means that only in diseased cells (i.e., cells depleted of hnRNP splicing factors) will the entire downstream transgene be translated. In cells that are not depleted of hnRNP splicing factors, the translated protein may be an out-of-frame peptide or an N-terminally truncated version of the transgene-encoded protein, depending on the position of the start codon in the mRNA product without the cryptic exon.
[0289] The above example construct includes an additional intron sequence downstream of the cryptic exon (in this example, from RPS24). Although not required, placing the intron downstream is preferred as it promotes the deposition of an exon junction complex (EJC) into the resulting mRNA. The absence of the cryptic exon sequence in the mRNA transcript, and thus the presence of a premature stop codon, triggers nonsense-mediated decay (NMD) of the transcript, further improving the safety of the construct in healthy cells (where otherwise the peptides produced in healthy cells may accumulate and aggregate or become potentially toxic). In contrast, in cells in which hnRNP splicing factors are absent (i.e., diseased cells lacking TDP-43), splicing is not suppressed and the cryptic exon is included in the mRNA product. In these cases, the PTC codon (e.g., the PTC or stop codon in the transgene sequence) is not in the same frame as the start codon, so the ribosome removes the EJC and does not result in nonsense-mediated decay. In this example, additional intron sequence within this exon is present downstream of the transgene, although additional intron sequence could be present within the transgene itself. In this example, the intron flanking sequences are derived from the human RPS24 gene (selected because it is highly expressed, constitutively spliced, and short in length), however, since there are hundreds of short mammalian introns that are constitutively spliced that could be readily selected by one of skill in the art and used in a similar manner, it is envisioned that any number of suitable introns and their flanking sequences could be used instead.
[0290] Additionally, although the constructs in the above examples contain a cryptic exon (e.g., a sequence having a number of nucleotides not divisible by 3) that induces a frameshift, regulation can be achieved without the need for a frameshift if the cryptic exon itself contains the start codon required for transgene expression.
[0291] In the above construct, the TDP-43 binding domain contains TG / UG repeats (including small "AA" interruptions). However, it is known in the art that TDP-43 can bind to other TG / UG rich sequences that are not pure repeats. Structural biology studies have demonstrated that many bases within the TDP-43 binding footprint are degenerate, and TDP-43 can bind to "UG rich" sequences, such as SEQ ID NO: 65 GUGUGAAUGAAU, with similar affinity to pure UG repeats. Furthermore, there are well-characterized examples of TDP-43-regulated cryptic exons that feature TDP-43-binding domains that are UG-rich but do not contain extended UG repeats. A clear example is the TDP-43-regulated cryptic exon of UNC13A (see SEQ ID NO: 66). Although significant UG richness has been observed in the region near the cryptic exon where TDP-43 binds (as shown via iCLIP studies), there are no UG repeats greater than 3 (UGUGUG) within 400 nt of the cryptic exon, and no TG repeats greater than 4 (UGUGUGUG) anywhere within the annotated intron that contains this cryptic exon. Thus, the TDP-43 binding domain may contain a TG / UG-rich region.
[0292] UNC13A intron containing cryptic exon (cryptic in bold, TG-rich region (SEQ ID NO: 67) in italics) SEQ ID NO:66 TIFF2025513046000018.tif100160
[0293] The constructs described herein contain a TDP-43 binding domain and are regulated by TDP-43, but the binding domain can be switched to other hnRNP splicing factors whose binding domains are known in the art.
[0294] Design 2 Next, a construct with a different design from the construct shown in Example 1 was designed. The construct of Design 2 is illustrated in Figure 2. The construct of Design 2 is composed of a regulatory domain including an intron sequence containing a TDP-43 binding domain and a cryptic exon sequence (defined by a splice acceptor site and a splice donor site) embedded within the intron region, but the cryptic exon sequence itself encodes a portion of a transgene that encodes a protein (e.g., a functional or diagnostic protein).
[0295] Example 2 The construct contains the following features, listed 5' to 3': - Sequence containing the start codon The first exon, which encodes the first part of the transgene (here, mCherry) A regulatory domain that includes: A cryptic exon sequence embedded within an intron region, the cryptic exon sequence encodes the second part of the transgene (here mCherry). The cryptic exon sequence is defined by a splice acceptor site and a splice donor site, one of these splice sites is regulated by TDP-43 binding. The intron region itself is defined by a second splice donor site and an acceptor site, and is split into two parts, the first part is upstream of the cryptic exon sequence and the second part is downstream of the cryptic exon sequence. The intron region is composed of a TDP-43 binding domain A third exon that encodes the third part of the transgene (here, mCherry) Further intronic sequences, including an intron within an exon (here from RPS24)
[0296] In the constructs of Examples 2A-C, all of the exon sequences encode mCherry, with the cryptic exon sequences encoding the internal portion of mCherry and the N- and C-terminal sequences of mCherry being encoded by the upstream exon (i.e., the first exon) and downstream exon (i.e., the third exon), respectively.
[0297] Unlike the Design 1 constructs, the cryptic exon sequences encode a portion of the transgene. Unlike the Design 2 constructs, the cryptic exons and, in some cases, the surrounding intronic regions that form regulatory domains are also fully synthetic. These were designed using a splicing prediction computer program (Splice AI, see https: / / github.com / Illumina / SpliceAI).
[0298] To design these fully synthetic cryptic exons and surrounding introns, an algorithm was used and developed (see Materials and Methods). To generate the introns, random sequences were generated in which each base was equally likely to be A, C, G, or T, and GT (AAG) and (C)AG were added to the 5' and 3' ends, respectively. In addition, TG-rich regions (e.g., sequences with at least 80% identity to SEQ ID NO:2 and / or SEQ ID NO:115) and / or randomized pyrimidine-rich regions (defined as a 30-nucleotide region with 80% pyrimidine probability) were added to form TDP-43 binding sites or polypyrimidine regions, respectively. As a result, the resulting intron sequences were entirely synthetic and not derived from pre-existing intron sequences. To generate the cryptic exon sequences, a portion of the mCherry transgene sequence was selected and reverse-translated. The intron and cryptic exon were then joined and combined with the upstream and downstream mCherry coding sequences to form the initial sequence.
[0299] Splice AI was then used to predict and edit the splicing properties of the initial sequence. Sequences were mutated randomly, except for coding regions, with only synonymous mutations (i.e., mutations that do not change the encoded amino acid sequence). After each round of mutation, Splice AI was used to predict the splicing behavior. Splicing predictions were compared to the putative ideal scenario (splice sites upstream and downstream of introns (i.e., second splice donor and second splice acceptor sites) had high scores of around 1.00 (e.g., >0.95), splice sites defining cryptic exons had slightly lower splicing scores (e.g., 0.8), and no other predicted splice sites had scores >0.01). If the predicted splicing of the mutated sequence was closer to the ideal scenario than the previous optimal sequence, the new mutated sequence was used as a template for the subsequent round of mutation. The mutated sequence was discarded if it was neither better nor worse than the previous optimal sequence. As such, the algorithm can be viewed as a Darwinian directed evolutionary approach to generating optimized sequences.
[0300] Three different constructs encoding mCherry were prepared. The first two examples (Examples 2A and 2B) had a TDP-43 binding domain (TG-rich region) upstream of the cryptic exon. In Example 2C, the TDP-43 binding domain (TG-rich region) was downstream of the cryptic exon. The Splice AI scores of the cryptic splice sites were as follows:
[0301] TIFF2025513046000019.tif41162
[0302] The sequences and components of the constructs in the examples are as follows:
[0303] Example 2A TIFF2025513046000020.tif26162TIFF2025513046000021.tif239163TIFF2025513046000022.tif75163
[0304] The sequence, including the start codon and additional intron sequences (eg, based on RPS2), was the same as that described in Example 1A.
[0305] Example 2B TIFF2025513046000023.tif142163TIFF2025513046000024.tif236163
[0306] The sequence, including the start codon and additional intron sequences (eg, based on RPS24), was the same as that described in Example 1A.
[0307] Example 2C TIFF2025513046000025.tif223164TIFF2025513046000026.tif125163
[0308] The sequence, including the start codon and additional intron sequences (eg, based on RPS24), was the same as that described in Example 1A.
[0309] Examples 2D-2J The following examples are all Design 2 format constructs expressing mScarlet (i.e., part of the mScarlet coding sequence is within a cryptic exon). Importantly, these constructs have different TDP-43 binding domains with shorter TG repeats than those shown in other examples (e.g., Example 1A) that contain an intronic region based on AARS1.
[0310] In all of Examples 2D-2J, the construct further comprises: SEQ ID NO: 116 - GACTACAAGGACGATGATGACAAG The FLAG tag contained a C-terminal
[0311] Each of the constructs of Examples 2D-2J further contains a constitutive downstream intron having the same sequence as described for the Example 1A construct.
[0312] Example 2D The construct in this example contains short TG repeats on either side of the cryptic exon.
[0313] TIFF2025513046000027.tif232161TIFF2025513046000028.tif12161
[0314] Example 2E Similar to Example 2D, the construct contains short TG repeats on either side of the cryptic exon.
[0315] TIFF2025513046000029.tif216163TIFF2025513046000030.tif36163
[0316] Example 2F The construct in this example has a downstream TDP-43 binding domain.
[0317] TIFF2025513046000031.tif194160TIFF2025513046000032.tif54160
[0318] Example 2G This example also has a downstream TDP-43 binding domain.
[0319] TIFF2025513046000033.tif185161TIFF2025513046000034.tif82160
[0320] Example 2H This example has short TG repeats on either side of the cryptic exon.
[0321] TIFF2025513046000035.tif145162TIFF2025513046000036.tif126161
[0322] Example 2I The construct in this example did not have any expanded TG repeats, but instead was TG-rich, with TGs dispersed throughout the intron.
[0323] TIFF2025513046000037.tif100163TIFF2025513046000038.tif174163
[0324] Example 2J Similar to Example 2I, the construct in this example did not have expanded TG repeats, but instead had TGs dispersed throughout the intron, making it TG-rich, but with relatively weak cryptic splice sites.
[0325] TIFF2025513046000039.tif36160TIFF2025513046000040.tif239160
[0326] Example 3 The constructs in the following examples are also of "Design 2", except that the transgene encodes the Cre recombinase with the SV40 nuclear localization signal fused to mNeonGreen (a fluorescent protein) and separated by a T2A self-cleaving sequence. Unlike in Example 2, the intron regions (both the first and second parts), the TDP-43 binding domain, and the further intron sequences were of the same sequence as described in Example 1A.
[0327] The construct contains the following features, listed 5' to 3': The first exon that encodes the first part of the transgene (here, the Cre recombinase with the nuclear localization signal from the SV40 virus) including the initiation codon A regulatory domain that includes: A cryptic exon sequence, encoding the second part of the transgene (here, Cre recombinase), embedded within an intron region. The cryptic exon sequence is defined by a splice acceptor site and a splice donor site, one of these splice sites being regulated by TDP-43 binding. The intron region itself is defined by a second splice donor site and an acceptor site and is split into two parts. The first part is upstream of the cryptic exon sequence and the second part is downstream of the cryptic exon sequence. Here, the intron region is composed of the TDP-43 binding domain and is based on AARS1. A third exon that encodes the third part of the transgene (here, the Cre recombinase) -Sequence containing the T2A cleavage site A sequence encoding the second transgene (mNeonGreen) Downstream intron and exon sequences (here, from RPS24)
[0328] The Cre recombinase transgene was split into three parts: the first exon upstream of the regulatory domain, the second exon containing a cryptic exon sequence, and the third exon downstream of the regulatory domain. The transgene was split into three exons that could be efficiently spliced as predicted using the Splice AI algorithm. First, suitable splice sites were identified in the sequence coding for Cre recombinase by searching for tandem consensus exon splice site motifs ([C / A / G]AG-G). Next, the sequences between the tandem splice motifs that would result in cryptic exons were randomly mutated (only synonymous mutations were used) and sequences with a Splice AI score of approximately 0.3 were selected.
[0329] TIFF2025513046000041.tif33161TIFF2025513046000042.tif250161TIFF2025513046000043.tif245160TIFF2025513046000044.tif46162
[0330] Example 4 Example 4 was similar to Example 3, except for the exon encoding the Cas9 protein, a nuclear localization signal, a tri-flag tag, and an N-terminal T2A-mCherry with a C-terminal flag. The transgene was split into three exons that could be efficiently spliced as predicted using the Splice AI algorithm. Again, suitable splice sites were identified in the Cas9 coding sequence by searching for tandem consensus exon splice site motifs ([C / A / G]AG-G). The sequence between the tandem splice motifs (resulting in a cryptic exon) was then randomly mutated (using only synonymous mutations).
[0331] In selected examples, the splice score of the cryptic splice acceptor site (i.e., the first acceptor splice site) as determined by the Splice AI algorithm was 0.06, and the splice score of the cryptic splice donor site (i.e., the first splice donor site) as determined by the Splice AI algorithm was 0.17.
[0332] TIFF2025513046000045.tif93162TIFF2025513046000046.tif250163TIFF20255130460 00047.tif250165TIFF2025513046000048.tif248166TIFF2025513046000049.tif151166
[0333] The above construct was incorporated into a plasmid, which, in addition to the above features, further contains an enhancer sequence and a promoter sequence (here, the CMV enhancer and the CMV promoter, respectively) upstream of the construct and a polyadenylation site downstream of the construct.
[0334] The sequence of the complete plasmid containing the above Cas9 construct is shown below (SEQ ID NO:94). TIFF2025513046000050.tif49160TIFF2025513046000051.tif246160TIFF2025513046000052.tif246158TIFF2025513046000053.tif132158
[0335] Comments on the Design 2 construct Similar to Design 1, expression of the Design 2 construct can be switched "on" or "off" depending on the presence of a splicing repressor (e.g., TDP-43) that is or is not depleted in neurodegenerative disease. When TDP-43 is present, as in healthy cells, splicing of the cryptic exon is repressed and the resulting transcribed mRNA is absent. During subsequent translation, the ribosome encounters a premature stop codon that leads to a non-functional truncated protein. When TDP-43 is depleted, as in diseased cells, the cryptic exon is retained in the resulting transcribed mRNA. Because the cryptic exon causes a frameshift (i.e., the length of its sequence is not divisible by 3), the premature stop codon is no longer in frame with the start codon, allowing translation of the full-length translated protein. However, frameshifting may not be necessary if the cryptic exon encodes an essential portion of the transgene without which the protein product is not functional (e.g., the catalytic domain) or if the cryptic exon contains the start codon of the transgene.
[0336] The constructs of Design 2 have many advantages. Compared to Design 1, the constructs have smaller sequences. In addition, in Design 1, unwanted peptides are generated from the upstream regulatory region in diseased cells, which could be either N-terminal sequences added to the transgene protein product or short release peptides, whereas in Design 2, unwanted peptides are not generated. Furthermore, if the cryptic exon is expressed, there is a low possibility of full-length protein leakage. The constructs of Design 2 are guaranteed to have zero leakage in the absence of cryptic exons, since the complete and uninterrupted transgene sequence is not present. In contrast, in Design 1, the complete and uninterrupted transgene sequence is present in both healthy and diseased cells, so leakage may occur in healthy cells, for example, due to leaky ribosome scanning or alternative transcription initiation.
[0337] The constructs of Design 2 can contain an intron (here including downstream exon sequences from the RSP24 gene) downstream of the regulatory domain, which is not a required feature of the construct but is preferred because it allows for induction of nonsense-mediated decay (NMD) of transcripts that do not contain cryptic exon sequences (i.e., those produced in healthy cells) similar to the constructs of Design 1. This therefore further improves the safety of the construct.
[0338] Although the above example constructs show exons encoding portions of the protein upstream and downstream of the cryptic exon, this need not be present if a start codon is included in the cryptic exon sequence itself. The cryptic exon may encode the N-terminal, internal, or C-terminal portion of the protein.
[0339] The constructs in the above examples use a cryptic exon sequence that induces a frameshift. If the cryptic exon itself contains an initiation codon, it can be regulated without the need for a frameshift. Alternatively, if the cryptic exon sequence is selected to code for an essential portion of the protein (such as the catalytic domain), then a cryptic exon that induces a frameshift is not required. In healthy cells where the cryptic exon is not included in the mRNA product of the construct, a truncated, non-functional transgene is produced.
[0340] It is envisioned that other TDP-43 binding domains can also be used, as described in Design 1.
[0341] Example 5 An exemplary construct was designed according to "Design 3." This exemplary construct contains (from upstream to downstream): - Sequence containing the start codon A regulatory domain containing 3' exon sequences (here based on exon 4 of AARS1) Splice donor site A single regulatory intronic region (here based on the intronic region between exons 4 and 5 of AARS1, containing the TDP-43 binding domain) Splice acceptor site 5' exon sequence (here based on exon 5 of AARS1) ·P2A cleavage site FLAG-mCherry transgene Further intronic sequences, including introns within exons (here based on RPS24)
[0342] TIFF2025513046000054.tif144162TIFF2025513046000055.tif251162TIFF2025513046 000056.tif250162TIFF2025513046000057.tif245163TIFF2025513046000058.tif81163
[0343] The additional intron sequences, the P2A cleavage sequence, the 3' exon sequences and the 5' exon sequences are otherwise as described in Example 1A.
[0344] Comments on Design 3 Although the transgene is shown here to be completely downstream of the regulatory domain, in other embodiments, the transgene sequence may be upstream and downstream of a single regulatory intron (i.e., as shown in FIG. 3 and Examples 6A-D). Similarly, although the binding domain is shown here to be within a single regulatory intron, the binding domain may instead be upstream or downstream of a single regulatory intron. As described in Design 1, the use of other TDP-43 binding domains is also envisioned. As with the constructs of Design 1 and Design 2, the P2A cleavage site, premature stop codon, and additional intron sequences are merely optional features and may be omitted.
[0345] Results and Discussion Direct and indirect TDP-43-dependent expression of fluorescent proteins As described above, we generated a series of TDP-43-dependent expression vectors that express fluorescent proteins in response to TDP-43 knockdown based on existing and novel cryptic exons. First, we generated a vector featuring an upstream frameshift cryptic exon based on AARS1 (but with a shorter intron and an extra adenosine within the cryptic exon sequence) fused to mCherry with an N-terminal P2A site. We transfected this vector into SK-N-DZ cells with doxycycline-dependent TDP-43 knockdown and analyzed the fluorescence by flow cytometry (see Methods). In untreated cells, only a small amount of expression leakage was detected, whereas a large increase in mCherry signal was detected in doxycycline-treated cells (Figure 4A, "AARS1-based reporter", mean mCherry signal fold change = 8.2-fold).
[0346] Next, three fully synthetic cryptic exons and surrounding introns were generated using a splicing prediction computer program. The generated exon and intron sequences were not derived from or based on existing sequences (see Examples 2A-2C). In each case, the cryptic exon sequence encodes an internal portion of mCherry, and the N- and C-terminal mCherry sequences are encoded by the upstream and downstream exons, respectively, so that the inclusion of the cryptic exon results in expression of a complete mCherry transcript. Designs 1 and 2 feature a TDP-43 binding domain, which consists of a TG-rich region upstream of the cryptic exon, whereas Design 3 features a TG-rich region of the TDP-43 binding domain. All three vectors showed increased mCherry expression upon TDP-43 knockdown, with a 2.2-fold increase in Design 3 and a 16.1-fold increase in Design 2 (Figure 4A).
[0347] Further synthetic cryptic exons and surrounding introns were then purified (see Examples 2D-2J). In all cases, the cryptic exon sequence encodes an internal portion of mScarlet, while the N- and C-terminal mScarlet sequences are encoded by the upstream and downstream exons, respectively, so that the inclusion of the cryptic exon results in the expression of the complete mScarlet transcript. Notably, these constructs contained short TG repeats in the intronic regions adjacent to the cryptic exon or long TG-rich sequences in the intronic regions adjacent to the cryptic exon. These are outlined below. All example constructs showed increased expression in Dox-treated cells compared to untreated cells. These results are shown in FIG. 12.
[0348] TIFF2025513046000059.tif122161
[0349] Next, we designed a vector encoding Cre recombinase. In this vector, an internal portion of the Cre recombinase sequence was encoded by a novel cryptic exon sequence (see Example 4). It was flanked by the same AARS1-derived intronic regions used in Example 1A. To optimize this vector, we used splicing prediction computer software. To evaluate the expression and activity of Cre recombinase in cells, we co-transfected a plasmid encoding mScarlet, which features a constitutive "poison exon" (an exon containing a premature stop codon) flanked by two LoxP sites. Thus, efficient mScarlet expression requires Cre recombinase-mediated excision of the poison exon. In cells without TDP-43 knockdown or cells not transfected with Cre recombinase, mScarlet expression was minimal, whereas cells transfected with both plasmids and with TDP-43 knockdown showed a 15.7-fold increase in the average mScarlet signal (Figure 4, part B). Furthermore, the results indicate that novel and different cryptic exon sequences can be inserted into an intron derived from AARS1 and still function as cryptic exons.
[0350] Finally, a construct containing a single regulatory intron (i.e., according to Example 5) was developed to prove the concept of the "Design 3" construct. In such a design, the regulatory domain was composed of a single regulatory intron, and transgenic expression was determined by whether intron splicing was suppressed. Cells without TDP-43 knockdown showed minimal mCherry expression indicating that the intron was retained in the mRNA product, whereas cells with Dox-induced TDP-43 knockdown showed a marked increase in signal, indicating that the intron was effectively spliced out (see Figure 9).
[0351] TDP-43-dependent Gaussia princeps luciferase expression Gaussia princeps luciferase (GLuc) is a secreted luciferase, and therefore is suitable for use in biomarker research, including minimally invasive biomarker research in vivo. A vector encoding GLuc was designed (see Example 1B above). The other configurations were similar to those described in Example 1A.
[0352] This vector was transfected into SK-N-DZ cells with or without TDP-43 knockdown as described above. The levels of secreted GLuc enzyme were then assessed by removing 20 μL of medium from the cell cultures and assessing chemiluminescence (see Methods). Supernatants from cells transfected with a vector not encoding GLuc or from cells without TDP-43 knockdown did not show a strong signal, whereas a significantly higher signal was detected from cells transfected with the cryptic GLuc vector and with TDP-43 knockdown (Figure 5).
[0353] TDP-43-Dependent Gene Editing We next evaluated whether a cryptic exon could be used to restrict gene editing by the Cas9 enzyme to only cells depleted of TDP-43. Using splicing prediction computer software, we designed a mammalian Streptococcus pyogenes (S. pyogenes) Cas9 expression vector in which an internal portion of the Cas9 coding sequence was encoded by a novel "cryptic exon" flanked on both sides by intronic sequences from AARS1 (see Example 4). We then co-transfected this vector with a vector encoding a single guide RNA (sgRNA) targeting the human CDK4 gene into SK-N-DZ cells with or without doxycycline-dependent TDP-43 knockdown. We then analyzed the expression of FLAG-tagged Cas9 enzyme by Western blotting and gene editing by amplicon Illumina sequencing.
[0354] Full-length FLAG-tagged Cas9 enzyme was detected only in cells transfected with a vector containing a cryptic exon with TDP-43 knockdown, but not in cells transfected with a constitutive FLAG-tagged Cas9 expression plasmid in both conditions (Figure 6A). Consistent with these results, we observed significantly increased indel numbers only in cells transfected with a constitutive Cas9 expression vector or a cryptic exon Cas9 vector with TDP-43 knockdown (Figure 6B).
[0355] Expression and autoregulation of TDP-43 fusion proteins One way to correct the loss of nuclear function of TDP-43 is to express a splicing repressor that binds to the same target sequence as TDP-43. This can be achieved by transgenic expression of TDP-43, but it may exacerbate cytoplasmic aggregation and toxicity. Therefore, an alternative approach is to express the RNA-binding domain of TDP-43 fused to a different splicing repressor. This avoids the risks associated with the expression of the C-terminal domain of TDP-43, which is deeply involved in cytoplasmic aggregation and toxicity. However, because overexpression of TDP-43 can be toxic in vivo, it is expected that similar toxicity may result from the expression of TDP-43-based fusion proteins, even if the toxic C-terminal domain is replaced with a safer alternative.
[0356] Instead, the construct of the present invention presents a possible solution to this problem, since expression of the transgenic protein depends on the loss of nuclear function of TDP-43. As a result, if the therapeutic transgene was a TDP-43-based splicing repressor fusion protein, the expression system could be autoregulated, since expression of the transgene inhibits further expression of the transgene by suppressing the incorporation of cryptic exons required for protein expression.
[0357] To test this idea, we fused the AARS1-based frameshift system used in Example 1A to the mCherry reporter, replacing mCherry with a TDP-43 / Raver1 fusion (see Example 1C). This protein was previously shown to partially rescue TDP-43 loss of function. We also generated an RNA-binding-deficient mutant of the same construct, in which the two phenylalanines in the RNA recognition domain 1 of TDP-43 were mutated to leucines (see mutants in Example 1C).
[0358] These constructs were co-transfected into SK-N-DZ cells containing inducible TDP-43 knockdown in combination with a minigene plasmid of the cryptic exon present in the human INSR gene [Ling et al. Science, 2015, 349 (6248); 650-5 (hereby incorporated by reference)]. We found that TDP-43 knockdown increased the inclusion of the cryptic exon (Figure 7A, lane 3 vs. lane 6). In cells co-transfected with the cryptic TDP-43-RAVER1 fusion, the inclusion rate of the cryptic exon was reduced, indicating rescue of the loss of TDP-43-derived splicing repression (Figure 7A, lane 4). As expected, no rescue was detected in the RNA-binding-deficient mutant (Figure 7A, lane 5).
[0359] TIFF2025513046000060.tif23164TIFF2025513046000061.tif250164TIFF2025513046000062.tif249163TIFF2025513046000063.tif50163
[0360] We next investigated whether this construct could autoregulate. The inclusion of the AARS1-derived cryptic exon was reduced in cells transfected with the cryptic TDP-43-RAVER1 fusion, but not in cells transfected with an RNA-binding-deficient mutant or an AARS1-based TDP-43-dependent mCherry expression vector (Figure 7B). Considering that expression of the fusion protein depends on the inclusion of the AARS1 cryptic exon, this indicates that our system can autoregulate expression if the expressed transgene is a TDP-43-based splicing repressor.
[0361] 17 and 18 provide further evidence of autoregulation of similar constructs.
[0362] Materials and Methods cell culture SK-N-DZ cells containing doxycycline-inducible shRNA targeting TDP-43 were cultured in DMEM / F12 medium supplemented with Glutamax and 10% FBS in 24-well dishes. TDP-43 knockdown was achieved by 1 μg / ml doxycycline treatment for 5 days. Transfections were performed using Lipofectamine3000 (ThermoScientific) on the third day of treatment, using a total of 500 ng of DNA per well. To reduce transfection variability between conditions, equivalent transfections were performed for untreated and doxycycline-treated cells using the same transfection master mix. For smaller samples (e.g., cells cultured in 96-well), the amount of DNA and Lipofectamine were adjusted according to the surface area of the vessel.
[0363] Flow cytometry analysis of cells expressing fluorescent proteins Mammalian expression vectors for fluorescent proteins were co-transfected into SK-N-DZ cells with 100 ng of mammalian Halo-tag expression vector (Promega). After overnight incubation with the Halo-tag-compatible far-red JaneliaFluor646 dye (Promega) 48 hours after transfection, cells were washed with PBS and analyzed on a BD LSRFortessa™ X-20 cell analyzer. Transfected cells were selected for analysis by gating on cells with high JaneliaFluor646 signal. Untransfected cells incubated in parallel with JaneliaFluor646 dye were used as a negative control for gating. Dead cells were filtered using 4',6-diamidino-2-phenylindole (DAPI) staining. mCherry signal was quantified for transfected cells, and background subtraction was performed by analyzing the level of mCherry signal from comparable untransfected cells of similar size (assessed by forward and side scatter height, width, and area values).
[0364] TIFF2025513046000064.tif40161
[0365] Luciferase analysis The cells cultured as described above were transfected. 48 hours after transfection, 20 μL of medium was collected and luminescence was assessed using the Pierce™ Gaussia Luciferase Glow Assay Kit (Thermo Scientific) as described in the manual.
[0366] Cas9 transfection, Western blotting, and indel analysis The cells were cultured as described above and transfected. Each well was transfected with 300 ng of Cas9 expression vector and 200 ng of sgRNA expression vector. Western blotting was performed using NuPage 4-12% gel (Thermo Scientific). Antibodies used were 10782-2-AP (Proteintech) for TDP-43, FLAG M2 antibody (Sigma Aldrich) for FLAG, and A11126 (Thermo Scientific) for tubulin. The Cas9 guide sequence was 5'-CACTCTTGAGGGCCACAAAG-3' (SEQ ID NO: 104). Genomic DNA was purified using primers 10782-2-AP (Proteintech), ... SEQ ID NO: 105 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGACGAACTGTGCTGATGGGA-3', and SEQ ID NO: 106 5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGGGTGCCTATGGGACAGTGTA-3' and sequenced on an Illumina MiSeq machine using PE250.
[0367] RT-PCR analysis of splicing Cultured cells were transfected as described above. 100 ng of minigene plasmid was mixed with 400 ng of TDP-43 / RAVER1 plasmid. 48 hours after transfection, RNA was extracted using the RNeasy Plus kit (Qiagen) according to the manufacturer's protocol. Random hexamer reverse transcription was performed with SuperscriptIV (Thermo Scientific), followed by the primers for AARS1: SEQ ID NO: 107 5'-CGATCCTACCATCCACTCG-3' and SEQ ID NO:108 5'-TTAATGATGGCCATGTTGTC-3', Primer for INSR SEQ ID NO: 109 5'-CTTCTTGGTGCCAGCTTATCAGAACTACTCCTTCTATGCCCTTGG-3' and SEQ ID NO:110 5'-GGCCTCGGATCCAGTTTACGCCTCTTTGTAGAACAGCATG-3' PCR was performed using
[0368] Determination of STMN2 cryptic splicing events SH-SY5Y cells were cultured in DMEM / F12 with Glutamax supplemented with 10% FBS. For induction of shRNA against TDP-43, cells were treated with doxycycline hyclate (Sigma D9891) at concentrations of 12.5ng / mL, 18.75ng / mL, 21ng / mL, 25ng / mL, and 75ng / mL. After 10 days, cells were harvested and RNA sequencing was performed. RNA was isolated using a QIAGEN RNeasy mini kit under conditions including an optional DNAse step according to the manufacturer's instructions. A sequence library was generated with polyA enrichment using the TruSeq Stranded mRNA Prep Kit (Illumina) and sequenced (2x150bp) using an Illumina HiSeq 2500 machine.
[0369] Samples were qualitatively trimmed with Fastp using the parameters "qualified_quality_phred:10" and aligned to the GRCh38 genome build using STAR (v2.7.0f) with GENCODEv31 gene models. STAR aligned BAMs were input into MAJIQ (v2.1) for splicing analysis using the GRCh38 reference genome. Results from the PSI module were then analyzed using a custom R script to calculate the probability of PSI and alteration for each junction. Cryptic splicing was defined as junctions with PSI < 5% in control samples and PSI > 10% in the 25ng / mL condition; junctions were not annotated in GENCODEv31.
[0370] TDP-43 protein levels were assessed as described in Brown, AL., Wilkins, OG, Keuss, MJ et al. TDP-43 deletion and ALS risk SNPs cause UNC13A missplicing and depletion. Nature 603, 131-137(2022), the contents of which are incorporated herein by reference.
[0371] Sequence of the plasmid encoding mScarlet (used in Example 4A) containing the "poison exon" flanked by LoxP sites SEQ ID NO:111 TIFF2025513046000065.tif35165TIFF2025513046000066.tif51167
[0372] Generation and analysis of barcoded Cas9 cryptic mutants (i.e., those used in Example 1E) To generate cryptic exons encoding parts of the Cas9 enzyme with various synonymous mutations, we designed oligos containing degenerate bases at the perturbed positions (i.e., third positions) of the relevant codons, which were then introduced via Gibson assembly into a plasmid with a 12-nt barcode (generated via whole-plasmid PCR using partially degenerate primers).
[0373] The resulting plasmids were then transfected into SK-N-DZ cells as described above. After RNA extraction, reverse transcription was performed with Superscript IV (Thermo Scientific) using specific reverse transcription primers against the construct's RNA, followed by amplification of the relevant cDNA by PCR and the addition of Illumina compatible overhangs. After sequencing using an Illumina MiSeq machine (Paired End 250), the reads were analyzed using custom R scripts.
[0374] Fluorescence microscopy Cells were imaged using an Olympus CKX53 microscope at 20x magnification with green illumination and filtered excitation in the red channel. Relevant settings (exposure, illumination levels, objective lens) were consistent between images.
[0375] "Algorithm 1" for designing synthetic cryptic exons (used in Example 2C) TIFF2025513046000067.tif53160TIFF2025513046000068.tif246159TIFF20255130460 00069.tif246160TIFF2025513046000070.tif244160TIFF2025513046000071.tif235160
[0376] Supplementary Example Fluorescence microscopy and nanopore sequencing of Example 1A Example 1A is a construct of Design 1, which fuses an upstream cryptic exon based on AARS1 to a downstream mCherry sequence (see above for design and sequence details).
[0377] Fluorescence microscopy and nanopore sequencing were performed on SK-N-DZ cells transfected with a constitutively expressing mCherry vector or Example 1A (both with and without TDP-43 knockdown). Figure 13A shows the fluorescence microscopy images. Figure 13B shows the quantification of the images shown in Figure 13A, where the numbers correspond to the log2 fold change in the fluorescence signal in cells with TDP-43 knockdown. Figure 13C shows the nanopore analysis of splicing of the cells in Figure 13A. A "productive transcript" is defined as a transcript that allows expression of the transgene (in this case mCherry). In the case of Example 1A ("cryptic mCherry"), this means a transcript that includes the upstream cryptic exon.
[0378] Fluorescence microscopy and nanopore sequencing of Examples 2D-2J Fluorescence microscopy and nanopore sequencing were performed on SK-N-DZ cells transfected with various Design 2 constructs (Examples 2D-2J) encoding mScarlet (both with and without TDP-43 knockdown). Figure 14A shows quantification of fluorescence microscopy images, where the numbers indicate the log2 fold change in fluorescence upon TDP-43 knockdown (e.g., 7.2 corresponds to a 2^7.2=147-fold increase in fluorescence). Figure 14B shows nanopore analysis of splicing of the cells in Figure 14A. "Productive transcripts" are defined as those that allow expression of the transgene mScarlet (i.e., those that contain a cryptic exon and are spliced as predicted).
[0379] Example 6: Further design 3 vectors An additional Design 3 vector was designed that contains a single TG-rich intron.
[0380] Single intron designs have been found to be highly effective when they contain a "decoy" or alternative splice site in addition to the first or cryptic splice site. The alternative splicing site or decoy splice site is the splice site that is preferentially used in normal cells (i.e., cells that contain TDP-43), but upon TDP-43 depletion, the first or cryptic splice site (i.e., flanked by TG-rich sequences) competes with the decoy or alternative splice site (i.e., not flanked by TG-rich sequences). Such designs can be computationally generated using an approach very similar to that described above.
[0381] Example 6A This Design 3 construct features a single intron with a cryptic acceptor splice site. The intron has a TG-rich region at its 3' end (i.e., on either side of the acceptor splice site) and a decoy splice site further downstream. When productively spliced, it encodes mScarlet. Splice AI predicted the splice strength of the cryptic acceptor splice site to be 25%.
[0382] TIFF2025513046000072.tif187161TIFF2025513046000073.tif82161
[0383] Example 6B The Design 3 construct contains a single intron with a cryptic acceptor splice site. The intron contains a TG-rich region at its 3' end (i.e., on either side of the acceptor splice site). Further downstream is a decoy splice site that encodes mScarlet when productively spliced. The cryptic acceptor splice site has a Splice AI predicted splice strength of 75%.
[0384] TIFF2025513046000074.tif123163TIFF2025513046000075.tif128163
[0385] Example 6C This further construct of Design 3 features a single intron with a cryptic acceptor splice site. There is a TG-rich region spanning the entire intron, with no long pure TG repeats. Further downstream is a decoy splice site that encodes mScarlet when productively spliced. The cryptic acceptor splice site has a Splice AI predicted splice strength of 75%.
[0386] TIFF2025513046000076.tif71162TIFF2025513046000077.tif203162
[0387] Example 6D This further construct of Design 3 contains a single intron with a cryptic donor splice site. A TG-rich region spans the entire 5' end of the intron. Upstream of the cryptic donor splice site is a decoy splice site that encodes mScarlet when productively spliced. The splice strength of the cryptic acceptor splice site predicted by Splice AI is 20%.
[0388] TIFF2025513046000078.tif245165TIFF2025513046000079.tif38165
[0389] Fluorescence microscopy and nanopore sequencing of Examples 6A-6D Fluorescence microscopy and nanopore sequencing were performed on SK-N-DZ cells transfected with the constructs of Examples 6A-6D above, both with and without TDP-43 knockdown. Figure 15A shows the fluorescence intensity of normal SK-N-DZ cells or SK-N-DZ cells with TDP-43 knockdown, where the numbers indicate the log2 fold change in fluorescence upon TDP-43 knockdown (triplicate determinations). Figure 15B shows nanopore analysis of splicing from cells. Error bars indicate standard deviation of triplicate determinations.
[0390] Nanopore sequencing trace Nanopore sequencing provides an excellent demonstration that splicing occurred as expected, since it can provide the sequence of the entire mRNA transcript in a single read from a single molecule. Quantification by nanopore sequencing is shown in Figures 13C, 14B, and 15B for constructs of designs 1, 2, and 3, respectively, and Figure 16 illustrates four nanopore traces of mScarlet from Examples 2E, 6A, 6B, and 6D. A. mScarlet vector "A11" [Example 2E], "Design 2" construct. Cryptic exon is indicated. Knockdown of TDP-43 results in inclusion of the cryptic exon and increased full or partial retention of the intron. B. mScarlet vector "B11" [Example 6A], "Design 3" construct with cryptic exon. In normal cells, it uses a decoy splice site. The cryptic acceptor is used at detectable levels only upon TDP-43 knockdown. C. mScarlet vector "B12" [Example 6B], "Design 3" construct with cryptic acceptor. In normal cells, the decoy splice site is predominantly used. The cryptic acceptor is used at very high levels upon TDP-43 knockdown. D. mScarlet vector "C3" [Example 6D], a "Design 3" construct with a cryptic donor. In normal cells, a decoy splice site is used. The cryptic donor is only used at detectable levels upon TDP-43 knockdown.
[0391] In both cases, the top displays the predicted splicing pattern as designed. Figures 16B-D show the alternative exons created by the splice sites in the second row. Cryptic exons or cryptic splice sites are highlighted (i.e., those predicted to be used upon TDP-43 knockdown).
[0392] In both cases, expression of mScarlet requires the use of a cryptic exon or a cryptic splice site. In both cases, the use of a cryptic splice site is highlighted with a STAR symbol. In both cases, the constitutively spliced intron from RPS24 is confirmed to be correctly spliced at the 3' end of the transcript.
[0393] Example 7: Design 2 TDP-43 / Raver1 fusion vector We designed additional "Design 2" format vectors (i.e., featuring an internal cryptic exon encoding part of the transgenic gene) that encode the TDP-43 / Raver1 fusion protein. These vectors express high levels of the TDP-43 / Raver1 fusion protein only when TDP-43 is depleted.
[0394] Importantly, these vectors express a TDP-43 / Raver1 fusion and therefore can rescue cryptic splicing events (i.e., compensate for the loss of endogenous TDP-43), including in the vector itself; in other words, the vector can autoregulate (i.e., block its own cryptic splicing).
[0395] Therefore, to measure cryptic splicing properties dependent on endogenous TDP-43, a vector expressing a mutant protein is used so that the transgenic protein does not interfere with splicing regulation. The mutant used has two phenylalanine to leucine (F2L) mutations in the RNA-binding domain of TDP-43, which prevents the vector-expressed protein from suppressing splicing and underestimating the inclusion of cryptic exons. This was achieved by two T-to-C point mutations, highlighted in the sequence below:
[0396] For therapeutic applications, non-mutated proteins are used so that splicing can be rescued, therefore, in the following experiments, both F2L mutant and wild-type versions of each example are analyzed to analyze cryptic splicing of the respective vectors or to analyze autoregulation and splicing rescue.
[0397] The sequences of the constructs for Example 7 are shown below. CDS refers to coding sequence. Example 7A has slightly weaker expression of the cryptic exons, resulting in a lower risk of leaky expression. Example 7B shows stronger expression, resulting in a higher maximum protein expression level, and is designed to provide better splicing rescue.
[0398] TIFF2025513046000080.tif234162TIFF2025513046000081.tif251163TIFF2025513046000082.tif249165TIFF2025513046000083.tif251165
[0399] Results and Discussion of Design 2 TDP-43 / Ravel1 Fusion Vector (i.e., Examples 7A and 7B) RT-PCR analysis of F2L mutants First, we analyzed splicing by the F2L mutants of each example (to avoid underestimating the inclusion of cryptic exons due to autoregulation).SK-N-DZ cells were transfected with either vector, with or without TDP-43 depletion, and splicing of the cryptic exons was analyzed using reverse transcription PCR (RT-PCR).
[0400] Figure 17 shows that in cells without TDP-43 depletion, the cryptic exon is barely detectable in Example 7A and only weakly expressed in Example 7B. However, under TDP-43 depletion conditions, the cryptic exon is strongly expressed in both examples (with stronger overall expression in Example 7B). It should be noted that shRNA against TDP-43 (shTDP) targets only endogenous TDP-43.
[0401] RT-PCR analysis of wild-type sequences We next performed a similar experiment using a functional TDP-43 sequence (i.e., without the F2L mutation). The results are shown in Figure 18. In this case, the TDP-43 / Raver1 construct suppresses its own cryptic splicing, and depletion of endogenous TDP-43 results in minimal inclusion of cryptic exons. These vectors are therefore termed "autoregulatory."
[0402] Splicing rescue of endogenous transcripts We next determined whether the constructs, in addition to autoregulating (i.e., “rescuing” their own cryptic splicing), could also rescue cryptic splicing of endogenous transcripts.
[0403] For this purpose, RT-PCR was performed on endogenous UNC13A and ELAVL3 in cells expressing example constructs 7A or 7B or constitutive expression vectors of the TDP-43 / Raver1 fusion protein (+ve control) or mScarlet (-ve control).
[0404] As expected, cryptic exon inclusion is essentially undetectable in normal cells. In cells expressing mScarlet, knockdown of endogenous TDP-43 leads to high levels of cryptic exon inclusion. However, in cells expressing a TDP-43 / Raver1 construct, the levels of cryptic exon inclusion are significantly reduced, indicating that cryptic splicing is rescued.
[0405] These results are shown in Figure 19. Figure 19A shows RT-PCR of endogenously expressed UNC13A transcripts in cells expressing the above constructs. Figure 19B shows quantification of the above RT-PCR for UNC13A and equivalent RT-PCR (not shown) performed on the ELAVL3 cryptic exon with or without knockdown of endogenous TDP-43.
[0406] Example 8: Design 2 Prime Editing Vector Another construct of Design 2 expressing Cas9 was designed with some improvements. This particular construct contained a different cryptic exon that showed higher expression upon TDP-43 depletion, allowing for more efficient genome editing. Furthermore, this sequence was placed within a "prime editing" vector, which allowed for a clearer demonstration of its activity (note that prime editing uses the H840A mutant Cas9).
[0407] The intron used in this case is a modified version of the intron containing a truncated AARS1 cryptic exon, which results in a higher downstream TG frequency and improves TDP-43 binding. The sequence details of this construct are as follows:
[0408] TIFF2025513046000084.tif127162TIFF2025513046000085.tif248163TIFF2025513046000086.tif247165TIFF2025513046 000087.tif247165TIFF2025513046000088.tif248165TIFF2025513046000089.tif246165TIFF2025513046000090.tif12166
[0409] As an example of a genome editing event, we targeted the cryptic exon donor splice site of UNC13A and used nanopore sequencing to measure the editing efficiency of this locus.
[0410] Results and Discussion SK-N-DZ cells were transfected with constitutive or cryptic PE-Max ("PE" = prime edit) vectors with or without TDP-43 knockdown. The cryptic exon was found to be expressed at detectable levels only upon TDP-43 knockdown. High levels of genome editing were seen at the expected locus in both conditions when the constitutive PE-Max vector was used, but only upon TDP-43 knockdown when the cryptic vector was used. Figure 20 shows A) a diagram of the vectors used, B) RT-PCR analysis of splicing with and without TDP-43 knockdown, and C) analysis of genome editing at the locus predicted by nanopore amplicon sequencing.
[0411] Example 9 - Luciferase Vectors An additional construct of Design 2 was designed that encodes luciferase. Figure 21 shows A) luciferase activity from culture medium of SK-N-DZ cells with and without TDP-43 knockdown. Error bars indicate standard deviation of technical triplicates, and B) nanopore traces from these cells.
[0412] The sequence of this construct is detailed below.
[0413] TIFF2025513046000091.tif96166TIFF2025513046000092.tif152166
[0414] Example 10: Combining multiple cryptic exons in a single vector This example demonstrates that multiple cryptic exons can be included in a single vector. In the construction of Example 10 below, the coding sequence for Cre recombinase is split into seven exons, three of which are flanked by TG-rich sequences, resulting in four constitutive exons and three cryptic exons.
[0415] TIFF2025513046000093.tif51161TIFF2025513046000094.tif227161
[0416] Nanopore sequencing revealed that while normal cells occasionally contained one or two cryptic exons, all three cryptic exons (required for Cre recombinase expression) were included at detectable levels only in TDP-43-depleted cells. Figure 22 shows A) a schematic of the three cryptic exon Cre recombinase vectors. Exons 2, 4, and 6 are "cryptic" and B) the number of nanopore reads of cryptic exons included in each transcript in SK-N-DZ cells without (NT) or with doxycycline-induced TDP-43 knockdown. Error bars indicate standard error of triplicate determinations.
[0417] We tested this vector in another cell line with inducible TDP-43 knockdown, in this case i3 iPSCs with Halo-tagged endogenous TDP-43, where we were able to deplete TDP-43 by adding a chimeric (Protac) molecule that targets a protein degradation product that matches the Halo tag. We used nanopore to analyze splicing of this construct in these cells. Duplicate measurements were performed for untreated and Protac-treated cells. As shown in Figure 23, the three cryptic exons were not detected in untreated cells, but were detected at high levels in Protac-treated cells.
[0418] Materials and Methods for Capture Examples Cloning Methods Gene fragments were ordered as eBlock or gBlock gene fragments from IDT. All cloning was performed using NEB HiFi assembly master mix and Stbl3 competent E. coli. Sequences were verified by Sanger or nanopore sequencing.
[0419] Automated Fluorescence Microscopy Transfection of the relevant plasmids and preparation of SK-N-DZ cells were performed as described above in 96-well dishes. Red fluorescence was imaged using an Incucyte S3 machine. Fluorescence intensity was integrated using CellProfiler.
[0420] Nanopore Sequencing RNA was extracted and reverse transcribed with an oligo(dT) primer using Superscript IV. Amplicons were generated using PCR with Q5 polymerase using primers unique to the 5' and 3' ends of the constructs. Primers had 20 nt 5' barcodes to allow for individual sample identification. PCR products were purified and nanopore libraries were prepared using the LSK-109 kit according to the manufacturer's instructions and sequenced on a Minion or Flongle machine. Base calling was performed using Guppy using high accuracy mode. Reads were demultiplexed using primer barcodes using a custom script and then reads were aligned to a reference sequence using Minimap2. Reads were confirmed to align primarily to the predicted sequence specified by the barcode. Splice junctions were then extracted using pysam and a custom Python script, and the level of productive splicing in each sample was quantified using R.
[0421] Example 7 - TDP-43 / Raver1 Materials and Methods Transfection of the relevant plasmids and preparation of SK-N-DZ cells were performed as described above. RNA was extracted from cells and reverse transcribed with SuperscriptIV. RT-PCR was performed using One-Taq polymerase (NEB) and visualized on a Qiaxcel automated capillary electrophoresis instrument (Qiagen). Primers targeting regions containing cryptic exons in each plasmid were used.
[0422] For rescue of endogenous transcripts, PiggyBac expression vectors of the relevant sequences were generated using Gibson assembly, and these plasmids were then co-transfected with PiggyBac transposon expression vectors to generate stable SK-N-DZ polyclonal lines, which were then selected with blasticidin (10 μg / ml) for 4 weeks, followed by doxycycline treatment for 5 days. RT-PCR was performed as described above, using PCR primers against the UNC13A and ELAVL3 sequences.
[0423] Example 8 - Prime editing materials and methods Prime editing vectors were cloned as described above. SK-N-DZ cells were co-transfected with the relevant prime editing vector and a prime editing guide RNA expression plasmid (spacer sequence: SEQ ID NO: 202 "TAAAAGCATGGATGGAGAGA", extension sequence: SEQ ID NO: 203 "ATGgACTCACgCATCTCTCCATCCATGC"), and a plasmid expressing mScarlet and a blasticidin resistance gene. Transfected cells were selected with 10 μg / mL blasticidin for 5 days with or without 1 μg / mL doxycycline to induce TDP-43 depletion.
[0424] RT-PCR was performed as described above using primers targeting the primed editing vector. Genome editing was assessed by PCR of genomic DNA using primers targeting the UNC13A cryptic exon locus, followed by nanopore sequencing (as described above).
[0425] Example 9 - Materials and methods for luciferase experiments Transfection of the relevant plasmids and preparation of SK-N-DZ cells were performed as described above. Luciferase activity in the culture medium of SK-N-DZ cells was assessed in technical triplicates each using the Pierce Gaussia Luciferase Glow Assay kit. Nanopore sequencing was performed as described above.
[0426] Example 10 - Cre Recombinase Materials and Methods SK-N-DZ cells were transfected with the relevant plasmids with or without TDP-43 knockdown. Nanopore sequencing was performed as described above and technical triplicates were assessed.
[0427] i3 iPSCs with Halo-tagged TDP-43 were electroporated using a P3 Primary Cell 4D-Nucleofector. 24 hours after electroporation, Halo-tagged Protac molecules (HaloPROTAC3, Promega) were added at a final concentration of 300 nM. After 3 days of treatment, RNA was extracted and targeted nanopore sequencing was performed as described above.
Claims
1. Start codon, Regulatory domains including the following: The first splice acceptor site and the first splice donor site, A binding domain to a splicing factor of the hnRNP family, located within 150 nucleotides of the first splice donor site and / or the first splice acceptor site; and / or located between the first splice donor site and the first splice acceptor site, and Transgene sequence Includes, (i) When nuclear depletion of splicing factors occurs in a cell, splicing of the first splice acceptor site and the first donor site is not suppressed, and a functional protein is produced from the transgene sequence. (ii) When localized to cells without nuclear depletion of splicing factors, splicing of the first splice acceptor site and / or the first donor site is suppressed, and no functional protein is produced from the transgene sequence. A structure that is configured in such a way.
2. The construct according to claim 1, wherein the binding domain to the splicing factor is a TDP-43 binding domain, and the splicing factor of the hnRNP family is TDP-43.
3. The construct according to claim 1, wherein the TDP-43 binding domain comprises a region of at least six nucleotides in which TG dinucleotides and / or TGNNTG hexanucleotides are statistically significantly enriched, where N is A, T, C, or G, and statistically significant enrichment is defined as the probability that random sequences of nucleotides of the same length are characterized by the same number of TG dinucleotides and / or TGNNTG hexanucleotides being less than 0.2%.
4. The construct according to claim 2 or 3, wherein the TDP-43 binding domain comprises the sequence TGTGTG, more preferably TGTGTGTG, and even more preferably TGTGTGTG.
5. The construct according to claim 1, wherein the binding domain of the hnRNP family to the splicing factor is located within 150 nucleotides of the first splice donor site and / or the first splice acceptor site, optionally within 100 nucleotides of the first splice donor site and / or the first splice acceptor site, and optionally within 50 nucleotides of the first splice donor site and / or the first splice acceptor site.
6. The binding domain of the hnRNP family to the splicing factor is (i) Upstream of the first splice acceptor site and the first splice donor site, (ii) Between the first splice acceptor site and the first splice donor site, (iii) Downstream of the first splice acceptor site and the first splice donor site The structure described in claim 1, located in [location].
7. The construct according to claim 1, wherein the introduced gene is a gene for a diagnostic protein, and optionally the diagnostic protein is a fluorescent protein, a luminescent protein, or a protein having a detectable antibody-binding tag.
8. The construct according to claim 1, wherein the introduced gene is a gene for a therapeutic protein, and optionally the therapeutic protein is a nuclease, a chaperone, a proteasome protein, a recombinase protein, a splicing regulator, or a transcription factor, and optionally the chaperone is a heat shock protein or a foldase.
9. The structure according to claim 1, wherein the first acceptor splice portion and the first donor splice portion have a splice score of 0.01 or higher when determined by the Splice AI algorithm, more preferably a splice score of 0.05 or higher when determined by the Splice AI algorithm.
10. The regulatory domain further includes an immature stop codon (PTC), (i) When nuclear depletion of splicing factors is localized in a cell, the PTC becomes out of frame with the start codon in the mRNA product of the construct. (ii) When localized in cells without nuclear depletion of splicing factors, the PTC becomes in-frame with the start codon in the mRNA product of the construct. The structure according to claim 1, configured in such a manner.
11. The construct according to claim 10, further comprising a further intron sequence downstream of the regulatory domain, wherein the PTC is at least 40 nucleotides upstream of the further intron sequence.
12. The construct according to claim 1, wherein the start codon is located upstream of the regulatory domain.
13. The first splice acceptor region and the first splice donor region define a cryptic exon sequence, the regulatory domain further includes an intron region, and the cryptic exon sequence is located within the intron region. (i) When nuclear depletion of splicing repressor proteins is localized in a cell, the cryptic exon sequence is present in the mRNA product of the construct. (ii) When localized to cells without nuclear depletion of splicing repressor proteins, the cryptic exon sequence is not present in the mRNA product of the construct. The structure according to claim 1, configured in such a manner.
14. The cryptic exon is a frameshift cryptic exon sequence whose nucleotide length is not divisible by 3. (i) When nuclear depletion of splicing factors is localized in a cell, the complete transgene sequence becomes in-frame with the start codon, (ii) When localized to cells without nuclear depletion of splicing factors, at least a portion of the transgene sequence is out of frame with the start codon. The structure according to claim 13, configured in such a manner.
15. The construct according to claim 13, wherein the intron region is formed from a first portion located upstream of a first splice acceptor site and a second portion located downstream of a first splice donor site, the first portion and the second portion originate from AARS1, and optionally the first portion has a sequence that is at least 80% identical to sequence number 30 and the second portion has a sequence that is at least 80% identical to sequence number 32.
16. The construct according to claim 13, wherein the introduced gene sequence is entirely downstream of the regulatory domain.
17. The construct according to claim 16, further comprising an autocleavage site or a protease cleavage site between the regulatory domain and the transgene sequence, wherein the cleavage site is optionally selected from P2A, T2A, F2A, E2A, furin, PCSK1, PCSK6, PCSK7, cathepsin B, granzyme B, factor XA, enterokinase, genenase, saltase, precision protease, thrombin, TEV protease, or elastase 1.
18. The construct according to claim 13, wherein at least a portion of the introduced gene sequence is encoded by a cryptographic exon sequence.
19. The construct according to claim 18, wherein the cryptic exon sequence codes for the N-terminal portion, the internal portion, the C-terminal portion, or any combination thereof of the introduced gene sequence.
20. The regulatory domain includes a single regulatory intron between the first splice donor site and the first splice acceptor site. (i) When the splicing factor is localized to a cell that is depleted, the single regulatory intron is spliced, (ii) If localized in cells where splicing factors are not depleted, the single regulatory intron is (i) not spliced, or (ii) not properly spliced. The structure according to claim 1, configured in such a manner.
21. A vector comprising the structure described in claim 1.
22. The invention comprises cells and the construct according to claim 1, or the vector according to claim 21. (i) When the hnRNP family splicing factors are depleted from the cell nucleus, the system produces functional proteins, (ii) If the splicing factors of the hnRNP family are not depleted, the system will not produce functional proteins. A system that is configured in such a way.
23. A construct according to claim 1, or a vector according to claim 21, for use in therapy.
24. A construct or vector for use in the treatment of diseases associated with the depletion of splicing factors of the hnRNP family, The treatment includes contacting cells with the construct or vector, (i) In cells with nuclear depletion of hnRNP family splicing factors, the cells produce functional proteins, (ii) In cells without nuclear depletion of hnRNP family splicing factors, the cells do not produce functional proteins. The construct according to claim 1, or the vector according to claim 21, wherein the disease is, in some cases, a neurodegenerative disease or a muscular disease, and further in some cases the neurodegenerative disease is amyotrophic lateral sclerosis (ALS) or frontotemporal dementia (FTD).
25. A method for selectively producing a functional protein in diseased cells with nuclear depletion of hnRNP family splicing factors, using the construct according to claim 1 or using the vector according to claim 21.