Methods for gene size reduction
By deleting non-functional sequences and replacing them with flexible linkers, the method addresses AAV's size limit, enabling effective genetic therapy for conditions like autism and epilepsy using AAV vectors.
Patent Information
- Application Number
- PCT/IL2025/050203
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-18
- Filing Date
- 2025-03-03
- Publication Date
- 2025-09-25
AI Technical Summary
The limitation of Adeno-Associated Virus (AAV) for gene therapy is its DNA size limit of 4.8kb, preventing the packaging of larger genes like IQSEC2, which is associated with conditions such as autism and epilepsy, rendering current methods ineffective for treating these disorders.
A method to reduce gene size by deleting non-functional sequences and replacing them with flexible linkers or sequences from different species, maintaining protein function while allowing packaging in AAV vectors.
Enables the use of AAV vectors for larger genes, facilitating effective genetic therapy for conditions like autism and epilepsy by maintaining protein function and compatibility with AAV packaging limits.
Smart Images

Figure IL2025050203_25092025_PF_FP_ABST
Abstract
Description
[0001] METHODS FOR GENE SIZE REDUCTION
[0002] FIELD OF THE INVENTION
[0003] The present disclosure is generally directed to genetic engineering and genetic therapy. More specifically, the method relates to reducing of gene size while maintaining protein function.
[0004] BACKGROUND OF THE INVENTION
[0005] Gene therapy using Adeno-Associated Virus (AAV) is emerging as a means for providing clinically meaningful benefit for an increasing number of diseases caused by mutations in a single gene. Most of the AAV DNA can be replaced with the therapeutic gene. However, a major limitation in the use of AAV is a DNA size limit for efficient packaging of the virus. The maximum size of a DNA cargo for generation of a high AAV titer (1014viral genomes / ml) needed for human gene therapy is 4.8kb. Viral genomes of greater than 5.0 kb are associated with AAV titers that can be orders of magnitude lower, not sufficient for gene therapy. Accordingly, using AAV for gene therapy of larger genes is not possible.
[0006] The X-linkcd IQSEC2 gene has an open reading frame of about 4.46 kb, encoding a protein of 1488 amino acids found at the post-synaptic density (PSD) of both excitatory (E) and inhibitory (I) neurons throughout the brain. The function of IQSEC2 is to promote GDP:GTP exchange on ARF6 which in turn mediates AMPA receptor trafficking and dendritic spine maturation. As such, the IQSEC2 protein plays a major role in controlling synaptic transmission, E / I balance and memory consolidation. Pathological mutations in IQSEC2 account for ~2-5% of all children presenting with frequent seizures, severe intellectual disability and autism spectrum disorder. The vast majority of affected males are unresponsive to antiseizure medication; the few responders require as many as 5-6 different medications to attain only partial control. Children with IQSEC2 mutations are minimally verbal or nonverbal, have an IQ below 50, and will likely require 24-hour access to an adult who can care for them for the rest of their lives.
[0007] However, the IQSEC2 gene size, together with necessary regulatory elements exceeds the size limitation of the AAV virus.
[0008] Accordingly, there is a need for methods for reducing gene size without affecting its function, in order to facilitate a wider use of the AAV platform.
[0009] SUMMARY OF INVENTION
[0010] The following embodiments and aspects thereof are described and illustrated in conjunction with compositions and methods which are meant to be exemplary and illustrative, not limiting in scope. In various embodiments, one or more of the above-described problems have been reduced or eliminated, while other embodiments are directed to other advantages or improvements.
[0011] In some embodiments, there is provided a modified nucleic acid molecule including a modified gene sequence encoding a modified protein having a reduced length compared to a parental protein of the modified protein and essentially the same function as the parental protein, the modified protein differing from the parental protein by lacking a non-functional sequence of the parental protein, wherein the non-functional sequence has at least three features selected from: (1) having no defined function, (2) having a disordered structure, (3) not including positions associated with disease-causing mutations, and (4) not being conserved in at least one different species.
[0012] In some embodiments, the at least three features include: (1) having no defined function, (2) having a disordered structure, and (3) not including positions associated with disease-causing mutations.
[0013] In some embodiments, the parental protein is at least about 1000 amino acids long. In some embodiments, the parental protein has at least two regions having a known function. In some embodiments, the parental protein has at least one disordered region. In some embodiments, the parental protein includes as least one position associated with a disease-causing mutation. In some embodiments, the parental protein has at least two regions which are conserved in at least one different species.
[0014] In some embodiments, the non-functional sequence is internal to the parental gene sequence encoding the parental protein, being flanked by two boundary amino acids.
[0015] In some embodiments, the intramolecular distance between the two boundary amino acids is not more than about 20 Angstroms.
[0016] In some embodiments, the non-functional sequence has a length of at least about 20-2000 amino acids. In some embodiments, the non-functional sequence is located between two regions having defined structures. In some embodiments, the non-functional sequence is located between two regions having defined functions. In some embodiments, the non-functional sequence is located between two regions which are conserved in at least one different species.
[0017] In some embodiments, the modified nucleic acid molecule further includes at least one sequence encoding a replacement sequence shorter than the non-functional sequence, wherein the replacement sequence replaces the non-functional sequence in the modified protein and is not part of the parental protein.
[0018] In some embodiments, the size difference between the non-functional sequence and the replacement sequence is about 20-2000 amino acids. In some embodiments, the replacement sequence has a flexible and / or a disordered structure. In some embodiments, the replacement sequence is a flexible linker.
[0019] In some embodiments, the flexible linker includes at least 50%, 60%, 70%, 80%, or 90% glycine residues.
[0020] In some embodiments, the replacement sequence has at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% sequence identity with a sequence from a different species located at a region corresponding to the location of the non-functional sequence in the parent protein.
[0021] In some embodiments, the parental protein is associated with a disease or condition.
[0022] In some embodiments, the disease or condition is autism, epilepsy, and / or IQSEC2-related disorder.
[0023] In some embodiments, the parental protein is IQSEC2.
[0024] In some embodiments, the non-functional sequence includes or is included in a sequence between amino acids 67 and 193, 73 and 193, 114 and 193, 391 and 627, and / or 381 and 481, of the human IQSEC2 protein.
[0025] In some embodiments, the nucleic acid further includes a sequence encoding a replacement sequence shorter than the non-functional sequence and replacing the non-functional sequence, wherein the replacement sequence has at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% sequence identity to a sequence of a different species located between amino acid positions corresponding to amino acids 67-193 of the human IQSEC2 protein.
[0026] In some embodiments, the different species is Xenopus laevis.
[0027] In some embodiments, the nucleic acid further includes a sequence encoding a replacement sequence shorter than the non-functional sequence and replacing the non-functional sequence, wherein the replacement sequence includes a stretch of at least 4, 5, 6, 7, 8, or 9 glycine residues.
[0028] In some embodiments, the nucleic acid includes a sequence at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 2 or 5, or encodes a sequence at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 1 or 4.
[0029] In some embodiments, there is provided a method for preparing a compact nucleic acid molecule including a modified gene sequence encoding a modified protein which has essentially the same function as a parental protein, the method including:
[0030] (a) providing a nucleic acid molecule including a parental gene sequence encoding the parental protein; and
[0031] (b) deleting from the parental gene sequence at least one sequence encoding a non-functional sequence, thereby generating the modified gene sequence encoding the modified protein, wherein the non-functional sequence has at least three features selected from: (1) having no defined function, (2) having a disordered structure, (3) not including positions associated with disease-causing mutations, and (4) not being conserved in at least one different species.
[0032] In some embodiments, the method further includes, prior to step (b), a step of mapping the parental protein for at least three features selected from: (1) regions having a defined function, (2) regions having a disordered structure, (3) disease-causing mutations, and (4) conservations across species.
[0033] In some embodiments, the at least three features include: (1) having no defined function, (2) having a disordered structure, and (3) not including positions associated with disease-causing mutations.
[0034] In some embodiments, the parental protein is at least about 1000 amino acids long. In some embodiments, the parental protein has at least two regions having a known function. In some embodiments, the parental protein has at least one disordered region. In some embodiments, the parental protein includes as least one position associated with disease-causing mutation. In some embodiments, the parental protein has at least two regions which are conserved in at least one different species.
[0035] In some embodiments, the non-functional sequence is internal to the parental gene sequence encoding the parental protein, being flanked by two boundary amino acids. In some embodiments, the intramolecular distance between the two boundary amino acids is not more than about 20 Angstroms. In some embodiments, the linear distance between the two boundary amino acids is at least about 50 amino acids. In some embodiments, the non-functional sequence has a length of at least about 20-2000 amino acids. In some embodiments, the non-functional sequence is located between two regions having defined structures. In some embodiments, the non-functional sequence is located between two regions having defined functions. In some embodiments, the nonfunctional sequence is located between two regions which are conserved in at least one different species.
[0036] In some embodiments, the method further includes, following step (b), a step of adding to the modified gene sequence at least one sequence encoding a replacement sequence shorter than the non-functional sequence, which replaces in the modified protein the sequence encoding the non-functional sequence.
[0037] In some embodiments, the size difference between the non-functional sequence and the replacement sequence is about 20-2000 amino acids.
[0038] In some embodiments, the replacement sequence has a flexible and / or a disordered structure. In some embodiments, the replacement sequence is a flexible linker. In some embodiments, the flexible linker includes at least 50%, 60%, 70%, 80%, or 90% glycine residues. In some embodiments, the replacement sequence has at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% sequence identity with a sequence of a different species located at a region corresponding to the location of the non-functional sequence in the parent protein.
[0039] In some embodiments, the method further includes, following step (b), additional steps for evaluating the modified protein, including: c) obtaining for the parental protein at least a first value of at least one parameter selected from: i) a binding energy with at least one interacting component, and ii) an intermolecular distance between an amino acid of the protein and an interacting component; d) obtaining for the modified protein at least a second value of the at least one parameter calculated in (c) for the parental protein; and e) comparing the first value and the second value, wherein the second value is not more than about 30% higher or 30% lower than the first value.
[0040] In some embodiments, the non-functional sequence is internal to the parental protein and is flanked by two boundary amino acids.
[0041] In some embodiments, the method further includes, following step (b), a step of adding to the modified gene sequence a sequence encoding a replacement sequence shorter than the nonfunctional sequence, which replaces in the modified protein the sequence encoding the nonfunctional sequence; and the at least one parameter further includes an intramolecular distance between the two boundary amino acids.
[0042] In some embodiments, the parental protein is associated with a disease or condition.
[0043] In some embodiments, the disease or condition is autism, epilepsy, and / or IQSEC2-related disorder.
[0044] In some embodiments, there is provided a modified nucleic acid molecule prepared by the method disclosed herein, including a modified gene sequence encoding a modified protein.
[0045] In some embodiments, there is provided a modified protein encoded by the modified gene sequence included in the modified nucleic acid molecule disclosed herein.
[0046] In some embodiments, the modified protein includes a sequence at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 1 or 4.
[0047] In some embodiments, there is provided an adeno-associated virus (AAV)-based construct including the modified nucleic acid molecule disclosed herein or encoding the modified protein disclosed herein.
[0048] In some embodiments, there is provided a viral particle or viral-like particle (VLP) including the AAV-based construct disclosed herein.
[0049] In some embodiments, there is provided a host cell including the modified gene sequence disclosed herein or the AAV-based construct disclosed herein, or expressing the modified protein disclosed herein.
[0050] In some embodiments, there is provided a system including the modified gene sequence disclosed herein, the modified nucleic acid molecule disclosed herein, or the modified protein disclosed herein, and further including an inhibitory RNA or a sequence encoding an inhibitory RNA, wherein the inhibitory RNA is directed to a nucleotide sequence included in the sequence encoding the non-functional sequence and not in the modified protein sequence.
[0051] In some embodiments, the inhibitory RNA is a short hairpin RNA (shRNA), small interfering RNA (siRNA) or a microRNA (miRNA).
[0052] In some embodiments, the inhibitory RNA is an miRNA at least 85%, 90%, 95%, or 99% identical to a sequence selected from SEQ ID NO: 7, 8, 9, 10, and 11, or an shRNA at least 85%, 90%, 95%, or 99% identical to SEQ ID NO: 12.
[0053] In some embodiments, there is provided a composition including the modified nucleic acid molecule disclosed herein, the AAV-based construct disclosed herein, the viral particle or VLP disclosed herein, the host cell disclosed herein, or the system disclosed herein.
[0054] In some embodiments, the viral particle or VLP has a titer of at least about 1012or 1013viral genomes / ml composition.
[0055] In some embodiments, the composition is a pharmaceutical composition, further including a pharmaceutically acceptable carrier.
[0056] In some embodiments, there is provided a method of treating a disease or condition by genetic therapy in a subject in need thereof, including administering to the subject a therapeutically-effective dose of the pharmaceutical composition disclosed herein.
[0057] In some embodiments, the disease or condition is selected from autism, epilepsy, and / or IQSEC2-related disorder.
[0058] In some embodiments, there is provided the pharmaceutical composition disclosed herein for use in a method of genetic therapy of a disease or a condition in a subject in need thereof.
[0059] In some embodiments, there is provided the system disclosed herein for use in the treatment of a disease or condition associated with a dominant-negative mutation.
[0060] In addition to the exemplary aspects and embodiments described above, further aspects and embodiments will become apparent by reference to the figures and by study of the following detailed descriptions. BRIEF DESCRIPTION OF DRAWINGS
[0061] The invention will now be described in relation to certain examples and embodiments with reference to the following illustrative figures.
[0062] Figs. 1A-1C show that binding of ApoCM to the IQ domain of IQSEC2 regulates accessibility of the IQSEC2 GEF catalytic Sec7 domain to ARF6. Fig. 1A. When ApoCM is bound to IQSEC2, access to the Sec7 domain is blocked. This figure shows the molecular dynamic equilibrated 3D structures in three projections of ApoCM-IQSEC2 complexes for a WT human IQSEC2 protein. As shown, the N terminal coiled region (dark blue-residues 1-67) of IQSEC2 blocks accessibility of the Sec7 domain (red 746-939) to ARF6 (not shown) when ApoCM (yellow) is bound to the IQ domain (green-residues 347-376) of IQSEC2. Fig. IB. shows that when ApoCM dissociates from IQSEC2, there is free access to the Sec7 domain. As shown, when ApoCM is not bound to IQSEC2 (as a result of Calcium binding to ApoCM) the N terminal domain (blue) pivots away from the Sec7 domain and allows the Sec7 domain to interact with ARF6, thereby allowing formation of ARF6-GTP (Fig. 1C). Fig. 1C shows 3D structures of ARF6-Sec 7 complexes. Complexes are presented in two projections. Free energies of binding AG are calculated by FoldX server. The identified three amino acid pairs that dominate the stability of the ARF6-IQSEC2 wild type complex are presented in the right side of figure. Interactions between corresponding functional groups of aa pairs are connected by arrows and labeled by distance values in A. Color scheme: ARF6 - blue, Sec 7 (749-1094 aa residues) - red, the rest of the fragments of IQSEC2 are grey. Residues dominating stability of the ARF6-Sec 7 complexes are element colored. A IQSEC2 3D structure (1-760 aa residues) was previously predicted by RaptorX server.
[0063] Fig. 2 shows a BLAST alignment of human and xenopus (frog) IQSEC2 proteins, and the location of conserved domains and human mutations. Gray vertical arrows above the human sequence demarcate the deleted region, and black vertical arrows below the frog sequence demarcate region inserted into the human gene instead the deleted region, to make the human-frog hybrid minigene. Regions corresponding to coiled coil, IQ, SEC7, PH and PDZ domains are enclosed in rectangles. Identities are indicated by a vertical line; and asterisks denote sites of disease causing mutations.
[0064] Figs. 3A-3B show a comparison of molecular structure and binding energies between WT IQSEC2 (Fig. 3A) and IQSEC2-hf (Fig. 3B) bound to ApoCM. For both the WT and miniprotein the N-terminal domain of IQSEC2 prevents ARF6 (GTPase) binding to Sec7 (guanine nucleotide exchange factor - GEF) when ApoCM is bound to IQSEC2. SEC7, coiled coil and IQ domains and apoCM are presented as their Vander Waals surfaces. The equilibrium state of the complex was simulated by accelerated molecular dynamics (aMD). The free energy of binding AG for the apoCM IQSEC2 complex (WT or miniprotein) (shown in figure) was calculated by the FoldX algorithm implemented in the YASARA Structure software. Color scheme: N-end (l-70)-dark blue; IQ domain (347-376) - orange; SEC7 (746-939) - red; PH (951-1085) - cyan; the remaining IQSEC2 fragments - grey; apoCM - magenta.
[0065] Fig. 4 shows a quantitative assessment of binding of WT IQSEC2, IQSEC2-hf (mini), and two IQSEC2 mutants (A350V and S1474Qfs) with the known interactors apoCM (CALM), p53, and PSD95, and as a control for non-specific interaction with GFP. The Y axis scale on the left applies to the interaction of the different IQSEC2 proteins with Calm, P53 and GFP and the Y axis scale on the right applies to the interaction of IQSEC2 with PSD95.
[0066] Fig. 5 shows a quantitative assessment of binding of WT IQSEC2 and IQSEC2-hf (mini) to Arf6, and their dependence on EDTA and calcium.
[0067] Fig. 6 shows luciferase activity for wild type (wt) and modified (hf) IQSEC2 cDNA with a N-terminal luciferase tag transfected into HEK 293T cells with miRNA50 at different ratios between the constructs. There was a 90% reduction in the expression of luciferase using the full length IQSEC2 construct, and an approximately 20-30% reduction of luciferase using the modified construct.
[0068] Fig. 7 shows luciferase activity for human wild type (wt) or shRNA resistant (wt shRes) IQSEC2 cDNA with an N terminal luciferase tag transfected into HEK 293T cells with shRNA 126 (directed in both human and mouse to the sequence encoding amino acids 126-132) at different ratios between the constructs. As shown, there was a 65% reduction in the expression of luciferase using the wild type IQSEC2 construct, but the shRNA did not reduce expression of luciferase tagging a IQSEC2 construct in the codon-modified, shRNA resistant, IQSEC2. m = mouse shRNA, h - human shRNA.
[0069] Figs. 8A-8B show the conditional Knocked-in (conKI) mouse locus before (Fig. 8A) and after (Fig. 8B) removal of the mouse exons 13-15 by the Cre recombinase. Fig. 8A From left to right: mouse endogenous IQSEC2 exons 11 and 12, first loxP site, mouse exons 13, 14, and 15 (including a natural stop codon), second loxP site, mouse exons 13, 14, and a partial exon 15 up to murine codon P1464 followed by human sequence from Q1474 including the frame shift and the 133 amino acid C terminal extension caused by the frameshift ending with a new stop codon. In Fig. 8B, the mouse exons 13, 14, and 15 (including the stop codon are removed by CRE recombinase, such that a mutant sequence is expressed. DETAILED DESCRIPTION OF THE INVENTION
[0070] In the following description, various aspects of the disclosure will be described. For the purpose of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the different aspects of the disclosure. However, it will also be apparent to one skilled in the art that the disclosure may be practiced without specific details being presented herein. Furthermore, well-known features may be omitted or simplified in order not to obscure the disclosure.
[0071] The present invention, in embodiments thereof, provides a method for reducing the size of a protein without affecting its function. The method is based on assessments of various parameters of the protein, such as mapping function to sequences and finding regions having no known function; simulating the 2D and 3D structure of the protein and finding regions having a disordered structure; mapping known disease-causing mutations to the protein; and determining conservation across species by sequence alignments. Following these assessments, candidates sequences for deletion may be selected based on features including having no known function, having a disordered structure, not including positions associated with disease-causing mutations, and not being conserved at least in other species.
[0072] In the next step, the candidate sequence is deleted (in silico or in vitro) and optionally replaced by a shorter sequence, which may be a flexible linker or a corresponding sequence from a different species (which is shorter). The new modified protein sequence obtained by the deletion and optional replacement is assessed for several features including: binding energy to known binding partners of the protein; intermolecular distance between residues which contact the binding partner; intramolecular distance between the two residues flanking the deleted sequence. Such parameters indicate the functionality of the modified protein and should generally remain similar to the same parameters in the original (parental) protein.
[0073] The present invention, in embodiments thereof, also provides specific examples for such a modified protein for the IQSEC2 protein, which is associated with IQS EC2 -related disorder, involving autism and epilepsy.
[0074] The advantages of the modified protein of the invention, which is shorter than the parental protein, include the ability to clone the gene encoding the protein into a vector having cloning size limitations, such as the adeno-associated virus (AAV).
[0075] Methods for preparing nucleic acids encoding modified proteins
[0076] In some embodiments, there is provided a method for preparing a modified nucleic acid molecule including a modified gene sequence encoding a modified protein having a reduced length compared to a parental protein of the modified protein and essentially the same function as a parental protein, the method including: a) providing a nucleic acid molecule including a parental gene sequence encoding the parental protein; and b) deleting from the parental gene sequence at least one sequence encoding a non-functional sequence, thereby generating the modified gene sequence encoding the modified protein, wherein the non-functional sequence has at least three features selected from: (1) having no defined function, (2) having a disordered structure, (3) not including positions associated with disease-causing mutations, and (4) not being conserved in at least one different species.
[0077] The term “parental protein”, as used herein, relates to a protein for which it is desired to obtain a shorter version (the modified protein) having the same function as the parental protein. The parental protein and the modified protein relate to two versions of the same protein, with the differences defined herein. An example for a reason to obtain a shorter version may be for cloning purposes - when the gene encoding the parental protein is too long for cloning in a desired vector. The parental protein is generally a full-length, or a wild type protein, but it does not have to be, and may be a protein that was previously modified or synthesized. The term “parental gene” relates to a nucleic acid sequence which encodes the parental protein, regardless of whether this is the native gene or has the same nucleic acid sequence as the native gene, as long as it encodes the parental protein. In some embodiments, the parental gene includes at least one intron. In some embodiments, the parental gene is intronless.
[0078] In some embodiments, the modified gene has a shorter length compared to the parental gene.
[0079] In some embodiments, the parental gene has a length which exceeds the length that is compatible with packaging in a desired cloning vector. In some embodiments, the cloning vector is an AAV-based vector.
[0080] In some embodiments, the parental protein has a length of at least about 1000, 1500, 2000, or 2500 amino acids.
[0081] In some embodiments, the parental gene has a length of at least about 4, 4.5, 5, 5,5, or 6 kb.
[0082] In some embodiments, the parental protein includes a sequence that is conserved in a different species, namely a species different from the species from which the parental protein is derived. In some embodiments, the parental protein includes a sequence that is conserved in a species from a different kingdom.
[0083] In some embodiments, the parental protein is associated with a disease or condition.
[0084] In some embodiments, the disease or condition is treatable by gene transfer.
[0085] In some embodiments, the disease or condition is autism, epilepsy, and / or IQS EC2 -related disorder.
[0086] The term “modified protein”, as used herein, relates to a protein that is derived from the parental protein, and is shorter than the parental protein, while maintaining the parental protein function. In other words, the modified protein has essentially the same function as the parental protein. The term “modified gene” relates to a nucleic acid sequence encoding the modified protein.
[0087] The term “modified nucleic acid”, as used herein, relates to a nucleic acid which includes at least the modified gene as defined herein, and possibly additional elements. In some embodiments, the modified nucleic acid has a size compatible with packaging in an AAV-based vector.
[0088] The expression “essentially the same function”, as used herein, is meant to indicate that the modified protein has not significantly lost functional capabilities compared to the parental protein. The meaning of “same” is qualitative rather than quantitative. According to the invention, the definition of having essentially the same function is manifested by maintaining functional and structural features as well as binding capabilities, as further explain herein and as defined in the claims. For quantitative purposes, having the same function is defined for measurable functions as a deviation of up to about 30% between the parental protein function and the modified protein function.
[0089] In some embodiments, the modified protein has an activity that is not more than about 30% higher or 30% lower than activity of the parental protein.
[0090] In some embodiments, the modified protein is obtained from the parental protein by the method disclosed herein.
[0091] In some embodiments, the at least one sequence encoding the non-functional sequence does not include an intron. In some embodiments, the at least one sequence encoding the non-functional sequence includes at least one intron.
[0092] In some embodiments, the at least one sequence encoding the non-functional sequence is more than one different sequence. In some embodiments, the at least one sequence encoding the non-functional sequence is two different sequences each encoding a non-functional sequence.
[0093] In some embodiments, the at least one non-functional sequence is at the N-terminal end of the parental protein. In some embodiments, the at least one non-functional sequence is at the C- terminal end of the parental protein. In some embodiments, the at least one non-functional sequence is internal to the parental protein, i.e., is neither at the N-terminal end of the parental protein, nor at the C-terminal end of the parental protein, and is therefore flanked by parental protein sequences. In some embodiments, the modified protein has a length not higher than about 2000, 1500, 1400, 1300, or 1200 amino acids.
[0094] In some embodiments, the modified gene has a length not higher than about 3, 3.5, 4, 4.5, 5,
[0095] 5.5, or 6 kb.
[0096] In some embodiments, the modified nucleic acid has a length not higher than about 3, 3.5,
[0097] 4. 4.5, 5, 5,5, or 6 kb.
[0098] In some embodiments, the modified protein is about 20-2000, 20-1000, 20-500, 50-500, 50- 400, 50-300, or 50-200 amino acids shorter than the parental protein. In some embodiments, the modified protein is at least about 20, 30, 50, 60, 70, 80, 90, 100, 150, or 200 amino acids shorter than the parental protein.
[0099] The term “region”, as used herein, relates to a stretch of sequence that is included in the protein or the gene. The sequence may be an amino acid sequence or a nucleic acid sequence, depending on whether the region is part of a protein or a gene, respectively.
[0100] The nucleic acid molecule including the parental gene sequence provided in step (a) may be in any form suitable for genetic manipulation. For example, the nucleic acid molecule may be in a vector suitable for genetic engineering.
[0101] In step (b), at least one nucleotide sequence encoding a non-functional protein sequence is deleted from the parental gene.
[0102] The term “non-functional sequence” is used herein to refer to a parental protein sequence selected for deletion based on having the features described below. Briefly, the non-functional sequence is generally selected for not having features associated with a known function or structure, and therefore most-likely serving as a “linker” or “space filling” sequence. This is an indication that the non-functional sequence may be deleted and optionally replaced with a different sequence without affecting the function of the parental protein. In some embodiments, more than one non-functional sequence is deleted from the parental protein. In some embodiments, two different non-functional sequences are deleted from the parental protein.
[0103] In some embodiments, the three features include: (1) having no defined function, (2) having a disordered structure, and (3) not including positions associated with disease-causing mutations.
[0104] Sequences having the above features may be identified by annotating the parental protein sequence for the relevant features by any suitable tool or method. Such tools or methods are generally in silico tools or methods which may be based on sequence analysis or on experimental data.
[0105] The expression “regions of known function”, as used herein, include regions known to have a defined function (e.g., catalytic activity, binding, dimerization, localization, etc.). Nonlimiting examples for methods for mapping of functional regions include existing database annotations (such as gene ontology (GO) annotations, and alignment with known proteins (such as by using basic local alignment search tool (BLAST)).
[0106] In some embodiments, the parental protein includes as least two regions having a known function.
[0107] The expression “disordered structure”, as used herein, relates to regions which do not have a stable 3D structure and are naturally in an unfolded state, and therefore can assume many conformations-providing great flexibility. Nonlimiting examples for methods for mapping of structures include analysis of crystal structures-based 3D molecular modelling methods, such as the deep convolutional residual neural networks (ResNet) method and folding simulations using molecular dynamics. Specifically, disorder regions may be predicted by using the D2P2database of disordered protein predictions (at d2p2.pro).
[0108] In some embodiments, the parental protein includes as least two regions having a defined structure. In some embodiments, the parental protein includes as least one region having a disordered structure.
[0109] The expression “positions associated with disease-causing mutations”, as used herein, relates to amino acid positions for which there are known substitutions, insertions, or deletions, which have been found to be associated with, or to cause, disease.
[0110] In some embodiments, the at least three features further include an additional feature (5): not including positions of known amino acid substitutions which are associated with, or affect, protein function. This additional feature relates to substitutions which have not been found to cause disease, but have been found, or are predicted, to affect function of the protein. Most likely, such substitutions are found in conserved regions, or in regions having a known function.
[0111] In some embodiments, the parental protein includes as least one position associated with a disease-causing mutation.
[0112] The expression “conservations across species”, as used herein, relate to sequences which have a high % of identity with other species, which usually indicates that the conserved sequence has a function. Sequence conservation may be found by sequence alignment (such as by using BLAST), more specifically by the conserved domain database (CCD) at the national center for biotechnology information (NCBI).
[0113] In some embodiments, a “conserved region“ (or “conserved sequence“) is defined as a sequence of at least 30, 40, or 50 aa, which is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95% identical to a sequence in a different species. In some embodiments, a conserved region is defined as a sequence of at least 50 aa, which is at least about 50% identical to a sequence in a different species. In some embodiments, a conserved region is defined as a sequence of at least 50 aa, which is at least about 80% identical to a sequence in a different species.
[0114] In some embodiments, the parental protein includes at least two regions which are conserved in at least one different species.
[0115] The term “different species”, as used herein, relates to species which are not the same as the species from which the parental protein is derived. In some embodiments, the different species belongs to a kingdom, phylum, class, order, family, or genus, different from that of the parental protein. In some embodiments, the parental protein is a human protein. In some embodiments, the different species is a non-human species. In some embodiments, the different species is a nonmammal species. In some embodiments, the different species is selected from Xenopus laevis and Caenorhabditis elegans.
[0116] In some embodiments, the method includes a step, prior to step (b), including annotating the parental protein sequence with, or mapping to the parental protein sequence, the above at least three features, such as known function, 2D or 3D structure, known mutations or functional substitutions, and conservation across species. These methods are carried out by any suitable method or tool, as explained above.
[0117] In some embodiments, the non-functional sequence has a length of about 20-2000, 20-1000, 20-500, 50-500, 50-400, 50-300, or 50-200 amino acids. In some embodiments, the non-functional sequence has a length of at least about 20, 30, 50, 60, 70, 80, 90, 100, 120, 130, 150, or 200 amino acids. In some embodiments, the non-functional sequence has a length of at least about 50 amino acids.
[0118] In some embodiments, the non-functional sequence is internal to the parental protein.
[0119] The term “boundary amino adds” is used herein to define the two amino acids which flank the non-functional sequence in the parental protein, when the non-functional sequence is internal to the parental protein. In some embodiments, the boundary amino acids remain in the modified protein.
[0120] In some embodiments, the intramolecular distance between the two boundary amino acids is not more than about 20 Angstroms. In some embodiments, the intramolecular distance between the two boundary amino acids is not more than about 10 Angstroms. In some embodiments, the intramolecular distance is less than about 20%, 10%, or 5% different between folded and unfolded conformations of the parent protein.
[0121] The term “intramolecular distance”, as used herein, relates to the shortest physical distance between two amino acids in a 3D structure of a protein. In contrast, the term “linear distance”, as used here, relates to a distance in terms of number of amino acids in the amino acid sequence (or chain) between two amino acids.
[0122] In some embodiments, the linear distance between the two boundary amino acids is at least about 40, 50, 60, 70, 80, 90, or 100 amino acids. In some embodiments, the linear distance between the two boundary amino acids is at least about 50 amino acids.
[0123] In some embodiments, the non-functional sequence is located between two regions having defined structures. In some embodiments, the non-functional sequence is a disordered region located between two ordered regions.
[0124] In some embodiments, the non-functional sequence is located between two regions having known or defined functions.
[0125] In some embodiments, the non-functional sequence is located between two regions which are conserved in at least one different species.
[0126] In some embodiments, the method further includes, following step (b), a step (bl) of adding to the modified gene sequence at least one sequence encoding a replacement sequence, and the sequence encoding the replacement sequence replaces in the modified protein the deleted sequence encoding the non-functional sequence.
[0127] The term “replacement sequence” is used herein to refer to the protein sequence which is used according to the invention to replace the non-functional sequence in the modified protein. Specific features of the replacement sequence are defined herein.
[0128] In some embodiments, the replacement sequence is shorter than the non-functional sequence.
[0129] In some embodiments, the sequence encoding a non-functional sequence is more than one sequence, and each deleted non-functional sequence is replaced by a replacement sequence. In some embodiments, the sequence encoding a non-functional sequence is more than one sequence, and at lease one deleted non-functional sequence is not replaced by a replacement sequence.
[0130] In some embodiments, the term “replaces” means that the replacement sequence is inserted at the same position from which the non-functional sequence was deleted. For example, if the nonfunctional sequence was internal to the protein, then the replacement sequence is positioned (or inserted) between the boundary amino acids which flank the non-functional sequence in the parental protein.
[0131] It is appreciated that if a replacement sequence is added, then the difference in size between the parental protein and the modified protein is a result of the differences in sizes between the nonfunctional sequence plus any additional non-functional sequences that are deleted, and the replacement sequence plus any additional replacement sequences that are inserted or added.
[0132] Deleting from the parental gene the sequence encoding the non-functional sequence, and inserting, or adding, the sequence encoding the replacement sequence may be done by any suitable genetic engineering method, including digestion by restriction enzymes and ligation, synthesis of a desired sequence, recombination, CRISPR technology, etc.
[0133] In some embodiments, the at least one sequence encoding the replacement sequence does not include an intron. In some embodiments, the at least one sequence encoding the replacement sequence includes at least one intron.
[0134] In some embodiments, the replacement sequence is not part of the parental protein sequence. In some embodiments, the at least one sequence encoding the replacement sequence is not part of the parental gene sequence.
[0135] In some embodiments, the replacement sequence is about 20-2000, 20-1000, 20-500, SO- SOO, 50-400, 50-300, or 50-200 amino acids shorter than the non-functional sequence. In some embodiments, the replacement sequence is at least about 20, 30, 50, 60, 70, 80, 90, 100, 120, 130, 150, or 200 amino acids shorter than the non-functional sequence.
[0136] The role of the replacement sequence, besides being shorter than the non-functional sequence it replaces, is mainly to provide flexibility and to allow the functional regions of the parental protein to function in the modified protein similarly to their function in the parental protein.
[0137] In some embodiments, the replacement sequence has a flexible and / or a disordered structure.
[0138] In some embodiments, the replacement sequence is a flexible linker. In some embodiments, the flexible linker has a size of about 4-20, 5-10, or about 10 amino acids. In some embodiments, the flexible linker has at least 50%, 60%, 70%, 80%, or 90% glycine residues. In some embodiments, the flexible linker has at least about 4, 5, 6, 7, 8, 9, or 10 glycine residues. In some embodiments, the flexible linker has about 5-15, or 6-13 glycine residues.
[0139] Alternatively, the replacement sequence may be derived from a different species, in which the sequences surrounding the non-functional region are conserved, and are joined by a sequence that is shorter than the non-functional region. The sequence joining the conserved regions in the different species, or a sequence similar to it, may be used as a replacement sequence.
[0140] Accordingly, in some embodiments, the replacement sequence has at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% sequence identity with a sequence from a different species located at a region corresponding to the location of the non-functional sequence in the parent protein.
[0141] The expression “a sequence from a different species located at a region corresponding to the location of the non-functional sequence in the parent protein”, as used herein, relates to a sequence of a protein homolog of the parental protein, which is located in a similar position in the parental protein and the homolog. One example is that the homolog sequence can be directly aligned with the parent sequence in a sequence alignment, and another example is when the homolog sequence and the parent sequence appear in the same position in a sequence alignment of the two proteins, either flanked by aligning sequences, or next to aligning sequences, such that they are located in similar (corresponding) positions in the folded protein.
[0142] In some embodiments, the method further includes a step of transferring the sequence encoding the modified protein sequence to a different nucleic acid molecule.
[0143] In order to predict whether the modified protein obtained by replacing the non-functional sequence with the replacement sequence is functional as the parental protein, several parameters, as detailed below, are evaluated and compared between the modified protein and the parental protein. It is predicted that if these parameters are similar between the parental protein and the modified protein, then the modified protein is likely to function similarly to the parental protein.
[0144] Accordingly, in some embodiments, the method further includes, following step (b), additional steps for evaluating the modified protein, including: c) obtaining for the parental protein at least a first value of at least one parameter selected from: i) a binding energy with at least one interacting component, and ii) an intermolecular distance between an amino acid of the protein and an interacting component; d) obtaining for the modified protein at least a second value of the at least one parameter calculated in (c) for the parental protein; and e) comparing the first value and the second value, wherein the second value is not more than about 30% higher or 30% lower than the first value.
[0145] In some embodiments, the method further includes, after step (b), a step (bl) of adding to the modified gene a sequence encoding a replacement sequence, as described above, and the additional steps (c-e) of evaluating the modified protein are carried out after adding the replacement sequence at step (bl).
[0146] In some embodiments, the non-functional sequence is internal to the parental protein and the method further includes a step (bl) of adding to the modified gene a sequence encoding a replacement sequence (as described above), and the at least one parameter further includes an intramolecular distance between the two boundary amino acids.
[0147] The calculation or evaluation of the parameters may be conducted in silico or in vitro.
[0148] The expression “interacting component”, as used herein, relates to a molecule which naturally interacts with the parental protein. Such molecules may be selected from proteins, peptides, nucleic acids, small molecules, and ions. More specific nonlimiting examples for interacting molecules include substrates and enzymes, cofactors, receptors and ligands.
[0149] In some embodiments, the interacting component is a protein.
[0150] In some embodiments, the at least one interacting components is one interacting components. In some embodiments, the at least one interacting components is more than one interacting components.
[0151] The binding energy may be evaluated by any suitable method. Nonlimiting examples for methods for evaluating binding energy include the modelling and simulation software such as the YAS ARA Structure software, and specifically the FoldX algorithm.
[0152] In some embodiments, the binding energy of the modified protein to an interacting component is not more than about 30%, 25%, 20%, 15%, or 10% higher or lower than the binding energy of the parental protein to the same interacting component.
[0153] In some embodiments, the interacting components are selected from Arf6 and apoCM.
[0154] Calculation of intermolecular distances may be applied to distances between at least one amino acid of the protein and a specific residue of the interacting component. The corresponding residue may be an amino acid, if the interacting component is a protein, or it may be a different molecule or atom, if the interacting component is not a protein. In some embodiments, the interacting component is an interacting protein, and intermolecular distance is calculated between an amino acid of the parental or modified protein and an amino acid of the interacting protein.
[0155] The intermolecular distances may be evaluated by any suitable method. Nonlimiting examples for methods for determining interacting residues (or interacting amino acids) and for calculating intermolecular distance between two interacting residues include 3D modelling methods as disclosed above.
[0156] In some embodiments, the intermolecular distance between an amino acid of the modified protein and an interacting component is not more than about 30%, 25%, 20%, 15%, or 10% higher or lower than the intermolecular distance between the same amino acid of the parental protein and the same interacting component. It is appreciated that since the non-functional sequence does not have a known function, it does not include amino acids interacting with an interacting components. Accordingly, the interacting amino acids in the parental protein remain the same amino acids in the modified protein, although their position numbering may be different.
[0157] In some embodiments, the intramolecular distance between the boundary amino acids of the modified protein is not more than about 30%, 25%, 20%, 15%, or 10% higher or lower than the intramolecular distance between the boundary amino acids of the parental protein.
[0158] In some embodiments, the modified protein is IQSEC2.
[0159] In some embodiments, the non-functional sequence includes or is included in a sequence between amino acids 67 and 193, 73 and 193, 114 and 193, 391 and 627, and / or 381 and 481, of the human IQSEC2 protein. In some embodiments, the modified protein includes two nonfunctional sequences, one included in a sequence between amino acids 67 and 193, and the other included in a sequence between amino acids 391 and 627 or 381 and 481.
[0160] In some embodiments, the replacement sequence has at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% sequence identity to a sequence from a different species located at a region corresponding to positions 67-193, 73-193, 114-193, 391-627, and / or 381-481of the human IQSEC2 protein.
[0161] Additional more specific embodiments related to the modified gene / protein, the nonfunctional sequence and the replacement sequence are provided below.
[0162] Nucleic acids encoding reduced size modified proteins and proteins encoded by them
[0163] In some embodiments, there is provided a modified nucleic acid molecule including the modified gene sequence disclosed herein, encoding the modified protein disclosed herein, and prepared by the method disclosed herein.
[0164] In some embodiments, there is provided a modified nucleic acid molecule including a modified gene sequence encoding a modified protein having a reduced length compared to a parental protein of the modified protein and having essentially the same function as the parental protein, the modified protein differing from the parental protein by lacking a non-functional sequence of the parental protein, wherein the non-functional sequence has at least three features selected from: (1) having no defined function, (2) having a disordered structure, (3) not including positions associated with disease-causing mutations, and (4) not being conserved in at least one different species.
[0165] Definitions and embodiments mentioned above and which may be relevant to the nucleic acid-related embodiments also apply here, and vice versa. Some particularly relevant embodiments may be pointed out or explicitly repeated. For terms used herein, unless stated otherwise, their definition and embodiments are intended to be the same as above (mutatis mutandis).
[0166] The expression “differing from the parental protein by lacking...” is intended to mean that the deletion is not necessarily the only difference between the parental protein and the modified protein, and there may be additional differences between the parental protein and the modified protein.
[0167] In some embodiments, the only difference between the parental protein and the modified protein is in that the modified protein lacks the non-functional sequence.
[0168] In some embodiments, the modified nucleic acid molecule includes at least one sequence encoding a replacement sequence as disclosed herein, wherein the replacement sequence replaces the non-functional sequence in the modified protein and is not part of the parental protein.
[0169] In some embodiments, the replacement sequence is shorter than the non-functional sequence. In some embodiments, the only difference between the parental protein and the modified protein is in that the modified protein includes the replacement sequence instead of the nonfunctional sequence.
[0170] To demonstrate the method disclosed above, a modified human IQSEC2 gene compatible with the packaging limits of AAV (about 4.8 kb), was generated according to the invention, and as described in more detail in the experimental section. Briefly, a modified IQSEC2 gene encoding a modified protein 90 aa shorter than the parental IQSEC2 protein was generated by deleting a sequence encoding a -125 aa sequence between positions 67 and 193 (not inclusive) that was identified based on the criteria detailed and explained above, and replacing it with the corresponding sequence from frog.
[0171] An AAV expressing the modified IQSEC2 gene is shown to be highly efficient in reducing lethal seizures in A350V IQSEC2 mice from 22% (15 / 69) to 0% (0 / 31), see Example 5. This is in comparison to a non- significant reduction from 34% in untreated mice to 21% in treated mice, obtained with a full-length IQSEC2 construct ((the parental gene, Mehta et al., 2021, infra). The titer obtained for an AAV including the modified IQSEC2 gene appears to be significantly superior (approximately 100 times higher) to that observed with an oversized AAV carrying the Mehta full length IQSEC2 gene. One possible explanation for this may be the improved quality and quantity of virus that injected with the modified, shorter, gene. This hypothesis is supported by the demonstration that the amount of viral mRNA made from the oversized virus was highly correlated with efficacy suggesting that the amount and possibly the quality of virus may have been rate limiting. It is well established that oversized AAV have problems with incomplete genomes, empty heads and cannot be made at high titers without clumping. Accordingly, the shorter gene facilitated providing a higher titer of the AAV virus, thereby obtaining a better effect.
[0172] An additional modified IQSEC2 gene was prepared, also encoding a protein 90 aa shorter than the parental protein, by replacing a sequence encoding 99 aa between positions 381 and 481 (not including) and replacing it with a peptide of 9 glycines.
[0173] Additional deletions and replacements were also prepared, including deleting amino acids 392-626 and replacing them with a pentamer (FSKQV, SEQ ID NO: 6) having a length of about 6 Angstroms corresponding to the intramolecular distance between the flanking (boundary) amino acids (391 and 627); and a deletion of amino acids 114 and 193 (partial to the above aa 67-193 deletion), showing dimerization of the deleted construct.
[0174] In some embodiments, the modified gene encodes a human modified IQSEC2 protein.
[0175] In some embodiments, the modified gene includes at least one intron. In some embodiments, the modified gene does not include an intron. In some embodiments, the non-functional sequence includes a sequence between amino acids 67 and 193, 73 and 193, 114 and 193, 391 and 627, and / or 381 and 481, of the human IQSEC2 protein.
[0176] In some embodiments, the non-functional sequence is included in the sequence between amino acids 67 and 193, 73 and 193, 114 and 193, 391 and 627, and / or 381 and 481, of the human IQSEC2 protein.
[0177] In some embodiments, the non-functional sequence is an amino acid sequence immediately flanked by amino acids 67 and 193, 73 and 193, 114 and 193, 391 and 627, and / or 381 and 481, of the human IQSEC2 protein, or a combination of such sequences.
[0178] In some embodiments, the replacement sequence has at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% sequence identity to a sequence from a different species located at a region corresponding to amino acids 67-193 of the human IQSEC2 protein.
[0179] In some embodiments, the nucleic acid further includes a sequence encoding a replacement sequence shorter than the non-functional sequence and replacing the non-functional sequence, wherein the replacement sequence has at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% sequence identity to a sequence of a different species located between amino acid positions corresponding to amino acids 114-193 of the human IQSEC2 protein.
[0180] In some embodiments, the nucleic acid further includes a sequence encoding a replacement sequence shorter than the non-functional sequence and replacing the non-functional sequence, wherein the replacement sequence has at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% sequence identity to a sequence of a different species located between amino acid positions corresponding to amino acids 391-627 of the human IQSEC2 protein.
[0181] In some embodiments, the nucleic acid further includes a sequence encoding a replacement sequence shorter than the non-functional sequence and replacing the non-functional sequence, wherein the replacement sequence has at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% sequence identity to a sequence of a different species located between amino acid positions corresponding to amino acids 381-481 of the human IQSEC2 protein.
[0182] In some embodiments, the replacement sequence has at least about 85% sequence identity to a sequence from a different species located at a region corresponding to amino acids 67-193, 73-193, 114-193, 391-627, and / or 381-481 of the human IQSEC2 protein.
[0183] In some embodiments, the replacement sequence has at least about 95% sequence identity to a sequence from a different species located at a region corresponding to amino acids 67-193, 73-193, 114-193, 391-627, and / or 381-481 of the human IQSEC2 protein.
[0184] In some embodiments, the different species is a non-human species. In some embodiments, the different species is a non-mammal species. In some embodiments, the different species is Xenopus laevis.
[0185] In some embodiments, the replacement sequence includes a stretch of at least 5, 6, 7, 8, or 9 glycine amino acids. In some embodiments, the replacement sequence is GGGGGGGGG (SEQ ID NO: 3).
[0186] In some embodiments, the replacement sequence is FSKQV (SEQ ID NO: 6).
[0187] In some embodiments, the modified nucleic acid molecule includes a nucleotide sequence at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 2 or 5.
[0188] It is appreciated that the nucleic acid, or the modified gene, even when encoding the same amino acids as in the IQSEC2 protein, may not be identical to the parental gene sequence, by using different codons for encoding the same amino acids.
[0189] In some embodiments, there is provided a modified protein encoded by the modified gene sequence disclosed herein.
[0190] In some embodiments, there is provided a modified protein having essentially the same function as a parental protein, the modified protein differing from the parental protein by lacking a non-functional sequence of the parental protein and by having a reduced length compared to the parental protein, wherein the non-functional sequence has at least three features selected from: (1) having no defined function, (2) having a disordered structure, (3) not including positions associated with disease-causing mutations, and (4) not being conserved in at least one different species.
[0191] Definitions and embodiments mentioned above and which may be relevant to the protein- related embodiments also apply here, and vice versa. Some particularly relevant embodiments may be pointed out or explicitly repeated. For terms used herein, unless stated otherwise, their definition and embodiments are intended to be the same as above (mutatis mutandis).
[0192] In some embodiments, the protein includes a replacement sequence shorter than the nonfunctional sequence, wherein the replacement sequence replaces the non-functional sequence in the modified protein and is not part of the parental protein
[0193] In some embodiments, the modified protein is a human modified IQSEC2 protein.
[0194] In some embodiments, the modified protein includes an amino acid sequence at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 1 or 4. Vectors, viral particles, and host cells encoding or including the reduced size modified proteins
[0195] In some embodiments, there is provided a vector including the modified nucleic acid molecule disclosed herein or encoding the modified protein disclosed herein.
[0196] Definitions and embodiments mentioned above and which may be relevant to the vector- related embodiments also apply here, and vice versa. Some particularly relevant embodiments may be pointed out or explicitly repeated. For terms used herein, unless stated otherwise, their definition and embodiments are intended to be the same as above (mutatis mutandis).
[0197] As explained above, an important advantage of the present invention is facilitating packaging of genes which exceed size limitations of certain vectors without affecting the function of the proteins encoded by these genes.
[0198] Accordingly, in some embodiments, the vector has a size limitation. In some embodiments, the vector has a size of not more than about 6, 5.5, 5, 4.8, 4.7, or 4.5 kb.
[0199] In some embodiments, the vector is a viral vector.
[0200] In some embodiments, the vector is an AAV vector.
[0201] In some embodiments, there is provided an AAV-based construct including the modified nucleic acid molecule disclosed herein or encoding the modified protein disclosed herein.
[0202] In some embodiments, the vector includes a promoter for driving expression of the modified gene. In some embodiments, the promoter is a tissue specific promoter, for directing expression to a desired tissue. In some embodiments, the promoter is a promoter specific for expression in neurons. In some embodiments, the promoter is selected from an EFla promoter, a human minicalmodulin (CAM1) promoter (e.g., Wang et al, 2023, Cell Reports 42, 113348), a pseudorabies (PRV) miniLAP (latency-associated) promoter (e.g., Gene Therapy (2024) 31:335-344), and a human synapsin (Hsyn) promoter.
[0203] In some embodiments, the vector further includes a 3 ’-untranslated region (UTR) region sequence. In some embodiments, the 3’-UTR region includes base pairs 5521-5811 of NCBI Reference Sequence: NM_001111125.3.
[0204] In some embodiments, there is provided a viral particle, or a viral-like particle (VLP), including the modified nucleic acid molecule, the vector, and / or the modified protein disclosed herein.
[0205] In some embodiments, there is provided a host cell including the modified nucleic acid molecule disclosed herein, the vector disclosed herein, or the viral particle or VLP disclosed herein, or expressing the modified protein disclosed herein. Systems including nucleic acids of the invention
[0206] When treating a disease caused by a gain of function (dominant-negative) mutation, the mutant protein must be both silenced and replaced with a functional protein. The present invention provides a perfect solution for this issue, since the deletion of a protein sequence from the parental sequence provides a sequence that appears only in the mutant protein (corresponding to the parental protein) and not in the modified protein. This sequence may be used for silencing the mutant protein without affecting the modified protein.
[0207] This provides yet an additional advantage for using the method of the invention, regardless of cloning issues. In this case, the modified protein is not necessarily shorter than the parental protein.
[0208] In some embodiments, there is provided a system including (i) the modified gene sequence disclosed herein, the modified nucleic acid disclosed herein, or the modified protein disclosed herein; and (ii) an inhibitory RNA or a sequence encoding an inhibitory RNA, wherein the inhibitory RNA is directed to a nucleotide sequence included in the sequence encoding the nonfunctional sequence and not in the modified protein sequence.
[0209] In some embodiments, there is provided a system including (i) a modified protein or a modified nucleic acid molecule including a modified gene sequence encoding a modified protein, wherein the modified protein has essentially the same function as a parental protein, the modified protein differing from the parental protein by lacking a non-functional sequence of the parental protein, wherein the non-functional sequence has at least three features selected from: (1) having no defined function, (2) having a disordered structure, (3) not including positions associated with disease-causing mutations, and (4) not being conserved in at least one different species; and (ii) an inhibitory RNA or a sequence encoding an inhibitory RNA, wherein the inhibitory RNA is directed to a nucleotide sequence included in the sequence encoding the non-functional sequence and not in the modified protein sequence.
[0210] Definitions and embodiments mentioned above and which may be relevant to the system- related embodiments also apply here, and vice versa. Some particularly relevant embodiments may be pointed out or explicitly repeated. For terms used herein, unless stated otherwise, their definition and embodiments are intended to be the same as above (mutatis mutandis).
[0211] The elements disclosed in the system embodiments are the same as those disclosed above, with the only difference being that the length of the modified protein is not necessarily reduced compared to the length of the parental protein, and the length of the replacement protein, when present, is not necessarily shorter than the length of the non-functional sequence it replaces.
[0212] In some embodiments, the inhibitory RNA and the modified gene are included in a single nucleic acid molecule or in a single vector. In some embodiments, the inhibitory RNA and the modified gene are included together in a single nucleic acid molecule or in a single AAV vector. In some embodiments, the inhibitory RNA and the modified gene are included in separate vectors.
[0213] In some embodiments, the inhibitory RNA is a short hairpin RNA (shRNA), small interfering RNA (siRNA) or microRNA (miRNA).
[0214] In some embodiments, the inhibitory RNA is an miRNA;
[0215] In some embodiments, the inhibitory RNA is an shRNA;
[0216] In some embodiments, the inhibitory RNA is included in the composition disclosed herein.
[0217] In some embodiments, the inhibitory RNA is provided in a separate compositions from the composition disclosed herein.
[0218] In some embodiments, the inhibitory RNA has a sequence of between about 18-28, 18-24, or 20-24 nucleotides. In some embodiments, the inhibitory RNA sequence is identical or complementary to a sequence at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to a sequence included in the sequence encoding amino acid positions 67-193 or 126-132 of human IQSEC2 (UniProt accession no. Q5JU85). In some embodiments, the inhibitory RNA sequence is identical or complementary to a sequence at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to a sequence encoding positions 114-193, 391-627, or 381- 481 of human IQSEC2 (UniProt accession no. Q5JU85). In some embodiments, the inhibitory RNA sequence is identical or complementary to a sequence at least 85% identical to a sequence encoding positions 114-193, 391-627, or 381- 481 of human IQSEC2 (UniProt accession no. Q5JU85). In some embodiments, the inhibitory RNA sequence is identical or complementary to a sequence at least 95% identical to a sequence encoding positions 114-193, 391-627, or 381- 481 of human IQSEC2 (UniProt accession no. Q5JU85).
[0219] In some embodiments, the inhibitory RNA sequence is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to a sequence selected from (Y stands for C or T, R stands for A or G):
[0220] GGTTCTGGTACTGGCTYTCGCGGC (SEQ ID NO: 7);
[0221] TAACCGGTAGTGTCCTGRAG (SEQ ID NO: 8);
[0222] GTAACCGGTAGTGTTCCTGRA (SEQ ID NO: 9);
[0223] TGTAACCGGTAGTGTTCCTGR (SEQ ID NO: 10);
[0224] TCGTGGTGCAGGTGGCAYTG (SEQ ID NO: 11); or
[0225] CCCGGAGGCTGTGTATCGGGACAATTCAAGAGATTGTCCCGATACACAGCCTC
[0226] CTTTTTTGGAAAT (SEQID NQ: 12)
[0227] In some embodiments, the inhibitory RNA sequence is at least 85% identical to a sequence selected from SEQ ID NO: 7, 8, 9, 10, 11, and 12. In some embodiments, the inhibitory RNA sequence is at least 95% identical to a sequence selected from SEQ ID NO: 7, 8, 9, 10, 11, and 12.
[0228] In some embodiments, the inhibitory RNA sequence is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 7. In some embodiments, the inhibitory RNA sequence is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 8. In some embodiments, the inhibitory RNA sequence is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 9. In some embodiments, the inhibitory RNA sequence is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 10. In some embodiments, the inhibitory RNA sequence is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 11. In some embodiments, the inhibitory RNA sequence is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 12.
[0229] It is sometimes useful to include a replacement sequence which is not recognized by the miRNA or shRNA (resistant to the miRNA or shRNA) by modifying codon usage, i.e., using different amino acid codons for the same amino acids, such that the replacement nucleotide sequence is not identical to the replaced (non-functional) nucleotide sequence, but encodes the same protein sequence as the replaced sequence.
[0230] In some embodiments, the nucleotide sequence encoding the replacement sequence and the nucleotide sequence encoding the non-functional sequence are not identical, but encode an identical amino acid sequence.
[0231] Compositions including the modified gene or protein, and use thereof
[0232] In some embodiments, there is provided a composition including the modified nucleic acid molecule disclosed herein, the vector disclosed herein, the viral particle or viral-like particle disclosed herein, the host cell disclosed herein, or the system disclosed herein.
[0233] In some embodiments, the composition is a pharmaceutical composition, further including a pharmaceutically acceptable carrier.
[0234] Definitions and embodiments mentioned above and which may be relevant to the composition-related embodiments also apply here, and vice versa. Some particularly relevant embodiments may be pointed out or explicitly repeated. For terms used herein, unless stated otherwise, their definition and embodiments are intended to be the same as above (mutatis mutandis).
[0235] The pharmaceutical composition of the present invention may be formulated in any conventional manner using one or more physiologically or pharmaceutically acceptable carriers or excipients. The carrier(s) must be "acceptable" in the sense of being compatible with the other ingredients of the composition, not being deleterious to the recipient thereof, and not significantly interfering with the activity of the polypeptide of the invention, or of any other active ingredient in the composition. The term “carrier” refers to a diluent, adjuvant, excipient, or vehicle with which the active agent is administered. The carriers in the composition may include a binder, such as microcrystalline cellulose, polyvinylpyrrolidone (polyvidone or povidone), gum tragacanth, gelatin, starch, lactose or lactose monohydrate; a disintegrating agent, such as alginic acid, maize starch and the like; a lubricant or surfactant, such as magnesium stearate, or sodium lauryl sulphate; and a glidant, such as colloidal silicon dioxide.
[0236] In some embodiments, the composition further includes an inhibitory RNA as disclosed herein.
[0237] In some embodiments, the composition includes the viral particle or viral-like particle disclosed herein, having a titer of at least about 1012or 1013viral genomes / ml.
[0238] Methods of treating gain-of-function mutations by nucleic acids of the invention
[0239] In some embodiments, there is provided a method of treating a disease or condition by genetic therapy in a subject in need thereof, including administering to the subject a therapeutically-effective dose of the pharmaceutical composition disclosed herein or of the system disclosed herein.
[0240] Definitions and embodiments mentioned above and which may be relevant to the treatment- related embodiments also apply here, and vice versa. Some particularly relevant embodiments may be pointed out or explicitly repeated. For terms used herein, unless stated otherwise, their definition and embodiments are intended to be the same as above (mutatis mutandis).
[0241] In some embodiments, the disease or condition is selected from autism, epilepsy, and / or IQSEC2-related disorder.
[0242] In some embodiments, the disease or condition is an IQSEC2-related disorder caused by a missense mutation selected from Arg43Pro, Ser257Asn, Ala350Thr, Ala350Val, Ala350Asp, Arg359Cys, Arg359His, Arg563Gln, Asn706Ser, Arg758Gln, Gly760Ser, Asn765Asp, Gly771Asp, Pro785Leu, Ala789Val, Gln801Pro, Ala836Val, Ser861Thr, Arg863Trp, Val918Phe, Ala953Thr, Arg970His, Arg995Trp, Leu999Phe, Aspl002Gly, Leul004Prp, PhelOlOLeu, LyslO12Glu, ArglO22His, Argl069Pro, ArglO69Gln, Serl079Pro, Argl l22Cys, Glyl l38Arg, Alal l39Thr, Argl l55Trp, Argl402Thr, and Serl474Gln fs*133. The amino acid positions relate to the 1488 amino acids isoform of the human IQSEC2 protein, UniProt accession No. Q5JU85.
[0243] In some embodiments, the disease or condition is an IQSEC2-related disorder caused by a mutation selected from c.55_151delinsAT p.Alal9Ile fs*32, c.83_85del p.Asp28del, c.88_90delATC p.Ile30del, c.97C>T p.Gln33*, C.184OT p.Arg62*, c.267C>Gp. Thr89*, c.273_282del p.Asn91Lys fs* 112, c.316C > T p.GlnlO6*, c.325delC insGC p. GlnlO84Ala fs*22, c.443,4 44dup p.Alal49Gln fs*58, C.556C > A p. S189*, c.588_610del p.Argl97Ala fs*34, c.675dup p.Ser226 fs, c.737+10385del p, c.738-lG>A Splicing, c.804delC p.Try269Thr fs*3, c.847_848del insT, p.Gly283Ser fs*23, c.854del p.Pro285Leu fs*21, c.895C>T p.Gln299*, c.928G>T p.Glu310*, c.999+8A>G, c,1000_1034del splicing, c,1170dupG p.Gln391Ala fs*5, c.1405,1406 del p.Lys469Val fs*4, c,1417G> T p.Glu473*, chrX: 53251127,53251128 dup p.Ser484*, c.1459_1460delAT p.Met487Val fs*2, C.1510OT p.Gln504*, c.1556_1599delACCT, p.Tyr519Trp fs*87, c.1567,2199 del ins GGC p.Thr523_Thr 733del ins Gly, C.1591O T, p. Arg531*, C.1618OT p.Gln540*, c.1744,1763 del p.Arg582Cys fs*9, c,1813_1814del, p.Asp605Pro fs*3, c.l861dup, c.l881delC p.His629Met fs*4, c.1983,1999 del p.Leu662Gln fs*25, c.2026del p.Ala676Leu fs*46, c.2052_2053 delCG p.Cys684*, c.2078delG р.Gly693Val*29, C.2184C > G p.Tyr728*, c.2203C>T p.Gln735*, C.2272OT p.Arg758*, с.2295_2297del p.Asn765del, c.2317C>T p.Gln773*, c.2317_2332del p.Gln773Glyfs*25, c.2329G>T p.Glu777*, c.2459+21C>T splicing, C.2521C > T p.Glu841*, c.2563C>T p.Arg855*, c.2662dup p.Ile888Asn fs*16, c.2679,2680 ins A p.Asp894 fs*10, c.2776C>T p.Arg926*, c.2799C>G p.Try933*, c.2846_2852 del CCCAGGT p.Ser949Cys fs*7, c.2854C>T p.Gln952*, C.2911C > T, p. Arg971*, c.2962C>T p.Gln988*, c.3079delC p.LeulO27Ser fs*75, c.3097C>T р.GlnlO33*, g.88032_88033del splicing, c.3163C>T p.Argl055*, c.3277+2T>G Splicing, с.3277+5G>A Splicing, c.3278C>A p.Serl093*, c.33OOdup p.Metl lOlTyr fs*5, c.3322C>T р.GlnllO8*, c.3387C>A p.Tyrl l29*, c.3433C>T p.Argl l45*, c.3457del p.Argl l53Gly fs*244, с.3613_3613delC p. Leul205Trp fs*192, c.3669_3733del p.Thrl225Ser fs*4 , c.3780delG р.Glnl261Ser fs*136, c.38Ol_38O8dup p.Glnl270Arg fs*130 „ c.3859C>T p.Glnl287*, с.4039dupG p.Alal347Gly fs*40, c.4110-4111 del p. Tyrl371Gln fs*15, c.4164dupC p.Ilel389
[0244] His fs*218, c.4246_4247insG p.Serl416 fs, c.4401del p.Glyl468 Ala fs*27, c.4419 Serl474Val fs*21, c.4419_4420 insC, Serl474Gln fs*133, c.4419_4431del p.Serl474Arg fs*17, chrX:52789239_53368927dup (579 kb), chrX:52911287,53315010dup (403 kb), chrX:52920728_53321125del (0.4 Mb), chrX:52954520_53315542dup (361 kb), chrX:53276030-53298472 dup / truncation, chrX:53283513- 53325282 dup / truncation, 46 X,t(X;20)(pl l.2;ql l.2) translocation, Dup on chr4: g.l83693432_18375617dup gain 62 kb Insertion point on chrX g. 53318362,53318363, when “p.” indicates protein positions in UniProt accession No. Q5JU85, and “c.” indicates cDNA positions while +1 corresponds to the A of the ATG translation initiation codon in the reference sequence for the IQSEC2 gene (GenBank: NM_001111125.2). * indicates introduction of a termination codon after the indicated number of amino acids where relevant, dup = duplication; ins=insertion; del=deletion; fs=frameshift.
[0245] In some embodiments, disease or a condition are associated with a dominant-negative mutation and the method includes administering the system of the invention to the subject.
[0246] Examples for dominant-negative mutations for the IQSEC2 gene include the ISQSEC2 mutations selected from A350V, A350D, A350T, and S1474Q fs*133.
[0247] In some embodiments, the disease or condition is an IQSEC2-related disorder caused by an A350V substitution.
[0248] The term “treating” or “treatment”, as used herein, refers to means of obtaining a desired physiological effect. The effect may be therapeutic in terms of partially or completely curing a disease or condition and / or symptoms attributed to the disease or condition. The term includes inhibiting the disease or condition, i.e. arresting its development; or ameliorating the disease or condition, i.e. causing regression of the disease or condition, e.g., by eliminating or ameliorating its symptoms.
[0249] The administering may be by any method or route suitable for treating the disease or condition. The dosage and frequency of administration may be determined by a physician based on the type and severity of the disease or condition.
[0250] In some embodiments, the administration is an intraventricular administration to the brain of the subject.
[0251] In some embodiments, there is provided the pharmaceutical composition disclosed herein, or the system disclosed herein, for use in a method of genetic therapy of a disease or a condition in a subject in need thereof.
[0252] In some embodiments, there is provided the use of the pharmaceutical composition disclosed herein, or the system disclosed herein, in a method of genetic therapy of a disease or a condition in a subject in need thereof.
[0253] In some embodiments, there is provided the use of the pharmaceutical composition disclosed herein, or the system disclosed herein, in the preparation of a medicament for treating a disease or a condition in a subject in need thereof.
[0254] Definitions and embodiments mentioned above and which may be relevant to the use-related embodiments also apply here, and vice versa. Some particularly relevant embodiments may be pointed out or explicitly repeated. For terms used herein, unless stated otherwise, their definition and embodiments are intended to be the same as above (mill al is mutandis). An S1474Q IQSEC2 model mouse
[0255] Example 12 and Figs. 8A-8B describe a conditional IQSEQ2 model mouse for the human IQSEQ2 S1474Q fsl33 mutation. The mouse includes flanking the endogenous murine IQSEC2 exons 13-15 (exon 15 including a stop codon) by loxP sites, followed by murine exons 13, 14, and partially 15, when exon 15 includes murine sequences up to codon Pl 464 followed by the human sequences from the mutant codon Q1474 and including the human 3’-UTR sequences expressed in the mutant due to a frame shift. Until a Cre recombinase is added, the endogenous wild type mouse gene is expressed, and adding a Cre recombinase, either by crossing to a mouse carrying the gene or by supplying the gene or enzyme, causes removal of the endogenous murine sequences and expression of the mutant IQSEC2 protein.
[0256] In some embodiments, there is provided an IQSEQ2 model mouse including a conditional human IQSEQ2 S1474Q fs mutation. In some embodiments, there is provided an IQSEQ2 model mouse including a conditional human IQSEQ2 S1474Q fs 133 mutation.
[0257] In some embodiments, the model mouse includes loxP or flippase recognition target (FRT) sites flanking at least mouse IQSEC2 exon 15. In some embodiments, the model mouse includes loxP or flippase recognition target (FRT) sites flanking mouse IQSEC2 exons 13-15.
[0258] In some embodiments, the model mouse includes human IQSEC2 3’-UTR sequences.
[0259] In some embodiments, there is provided a nucleic acid molecule for preparing an IQSEQ2 model mouse including a conditional human IQSEQ2 S1474Q fs 133 mutation, the nucleic acid molecule including: (i) mouse IQSEC2 exon 15 flanked by loxP sites; (ii) a chimeric mouse / human IQSEC2 exon 15 sequence including the human S 1474Q fs 133 mutation downstream from (i); and (iii) a 5’ and a 3’ homology arm having at least 85%, 90%, or 95% identity to mouse sequences 5’ and 3’ of mouse exon 15, respectively, for integration by homologous recombination replacing endogenous mouse exon 15.
[0260] In some embodiments, there is provided a nucleic acid molecule for preparing an IQSEQ2 model mouse including a conditional human IQSEQ2 S1474Q fs 133 mutation, the nucleic acid molecule including: (i) mouse IQSEC2 exons 13-15 flanked by loxP sites; (ii) mouse IQSEC2 exons 13-14 downstream from (i); (iii) a chimeric mouse / human IQSEC2 exon 15 sequence including the human S1474Q fsl33 mutation downstream from (ii); and (iv) a 5’ and a 3’ homology arm having at least 85%, 90%, or 95% identity to mouse sequences 5’ of mouse exon 13 and 3’ of mouse exon 15, respectively, for integration by homologous recombination replacing endogenous mouse exons 13-15.
[0261] In some embodiments, the chimeric mouse / human IQSEC2 exon 15 sequence includes mouse exon 15 in which sequence encoding the C-terminus of the IQSEC2 protein up to the end of exon 15 is replaced by human IQSEC2 mutant sequence encoding the S1474Q fsl33 mutation followed by human 3’-UTR sequences at least until an in-frame stop codon.
[0262] In some embodiments, the mouse exon 15 sequence in the chimeric mouse / human exon 15 includes murine sequences up to codon Pl 464.
[0263] In some embodiments, the model mouse expresses a cDNA at least 85%, 90%, or 95% identical to SEQ ID NO: 27.
[0264] In some embodiments, the model mouse expresses a protein at least 85%, 90%, or 95% identical to SEQ ID NO: 28.
[0265] In some embodiments, the model mouse further includes a gene encoding a recombinase. In some embodiments, the recombinase is a Cre recombinase. In some embodiments, the Cre recombinase is under an inducible promoter. In some embodiments, the Cre recombinase is inducible by Tamoxifen.
[0266] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention pertains.
[0267] All amino acid positions in the application relate to UniProt accession :Q5JU85, unless indicated otherwise.
[0268] The term “bp”, as used herein, means base pairs.
[0269] The term “aa”, as used herein, means amino acids.
[0270] The term "a" and "an" refers to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.
[0271] The term "about", when referring to a measurable value such as an amount, a ratio, and the like, is meant to encompass variations of ±10% of the indicated value, as such variations are also suitable to perform the disclosed invention. Any numerical values appearing in the application are intended to be construed as if preceded by “about”, unless indicated otherwise.
[0272] While certain embodiments of the invention have been illustrated and described, it will be clear that the invention is not limited to the embodiments described herein. Numerous modifications, changes, variations, substitutions, and equivalents will be apparent to those skilled in the art without departing from the spirit and scope of the present invention as described by the claims, which follow.
[0273] The following examples are presented in order to more fully illustrate some embodiments of the invention. They should in no way be construed, however, as limiting the broad scope of the invention. One skilled in the art can readily devise many variations and modifications of the principles disclosed herein without departing from the scope of the invention.
[0274] EXAMPLES
[0275] Materials and Methods
[0276] Molecular modeling and molecular dynamics simulations of wild type and mini IQSEC2 protein structures and their interactions with apoCM and ARF6
[0277] The IQSEC2 is an Intrinsically Disordered Protein (IDP). The 3D crystal structure of ARF1 complex with a small fragment (729-1099) of human IQSEC2 wild type containing SEC7 and PH domains is available as 6FAE.pdb. In Shokhen et al., (2023, infra) the IQSEC2 fragment (749- 1094) was separated from the crystal and then relaxed by conventional molecular dynamics of AMBR20 package. The deep convolutional residual neural networks (ResNet) method was used for predicting protein 3D structure implemented in the RaptorX server for the generation of a model of wild type IQSEC2 residues 1-760 (the server allowed for no more than 1000 amino acids). Finally, the combined 3D structure containing residues 1-1094 of the wild type IQSEC2 free and its complex with apoCM were assembled by molecular modeling, as described in Shokhen et al., (2023), infra. In both cases the generated IQSEC2 1-1094 fragment has an extended conformation, so this work used a standard protocol of accelerated molecular dynamics (aMD) implemented in AMBER20 package for the simulation of the folding process to get compact 3D structures of both IQSEC2 free and its complex with apoCM. An initial 3D structure of the frog mini-protein (IQSEC2-hf) - apoCM complex was constructed by in silico mutation of the 68-192 residues in the folded structure of IQSEC2-apoCM on the 68-102 original frog residues by YAS ARA Structure molecular modeling package. In the following step the IQSEC2-hf - apoCM complex was relaxed to equilibrium state in aMD simulations. ARF6 complexes with the above mentioned structures were generated by protein-protein docking applying the HADDOCK web server. Finally, geometry of the ARF6 complexes were optimized by YAS ARA Structure software in periodic simulation cell filled with explicit water molecules and Na+ and Cl- ions in physiological concentration.
[0278] DNA constructs used in this study
[0279] The WT human IQSEC2 gene was cloned 3' to renilla luciferase gene and three copies of the HA tag in pcDNA3.1 Zeo. The pcDNA3.1 Zeo construct expresses full length (1488 aa) human IQSEC2 with an N-terminal renilla luciferase and HAx3 tag under the control of a CMV promoter. DNA constructs containing N-terminal FLAG (for IQSEC2) or C-terminal FLAG (all interactors except IQSEC2) tags to monitor interaction of Calmodulin (CAM), PSD95, IRP53, IQSEC2 or GFP with the Luc-IQSEC2 construct were generated in a pcDNA 3.1 Hygro backbone. The interaction of IQSEC2 with GFP-FLAG was used to determine background non-specific interaction. Specific human mutations were introduced into IQSEC2 ORF in the renilla luciferase wild type IQSEC2 vector (S1474Qfs or A350V) to serve as negative controls for the interaction of Euc IQSEC2 with PSD95 and CAM respectively. All IQSEC2 constructs were verified by DNA sequencing. The genes for calmodulin (human Caimi (NM_006888), Calm2 (NM_001743) and Calm3 (NM_ 005184), PSD95 (NM_001128827.4), and iRP53 (NM_001144888.2) were obtained from a human ORFeome library.
[0280] Lumier assay to assess interactions between IQSEC2 and Calmodulin, PSD95 and iRP53 in vitro.
[0281] Preparation of assay plates. Ninety six well plates (Greiner white Lumitrac) were coated withlOO ul of 0.1 mg / ml anti flag antibody in PBS (Sigma F1804) overnight at 4°C. Plates were washed five times with PBS containing 0.5% Tween-20 and blocked with PBS containing 5% sucrose, 3% BSA, and 0.5% Tween-20 for 90 minutes at room temperature. Plates were washed 5 times with wash buffer and kept at 4°C until further use.
[0282] Cell preparation and transfection. Cells (HeK-293T) were plated at a density of 20,000 cells / well in a 96-well plate in DMEM containing 10% FCS, 100 I.U. / mE penicillin, 100 pg / mE streptomycin, and 2 mM glutamine, and incubated at 37°C, 5% CO2 overnight. For each well two mixtures were prepared: Mixture A contained 200 ng IQSEC2-luc expressing plasmid, 200 ng interacting protein-flag expressing plasmid Calm / PSD or iRP53), 22.5 ul OptiMEM (Gibco) medium heated to 37°C. Mixture B contained 22.5 ul OptiMEM medium heated to 37°C and 0.9 ul polyethylenimine (PEI). Mixtures A and B were mixed and incubated for 25 min at room temperature and then added to the cultured cells (45 ul / well). Plates were incubated for 48 hours.
[0283] Lumier assay. Transfected plates were put on ice and the medium removed. Extraction buffer (100 ul of ice cold buffer containing 50 mM Hepes pH 7.9, 150 mM NaCl, 2 mM EDTA, 0.5% Triton X-100, and 5% glycerol with proteinase inhibitors (APExBIO cat. no. K101) and phosphatase inhibitors (Sigma P5726). Cells were incubated for 10 min and then 85 ul was transferred to the anti-flag antibody coated plate and incubated for 1 hour rotating at 750 rpm at 4°C. Plates were washed 5 times with extraction buffer and the amount of luciferase activity adhered to the plate was measured by adding 20 ul lx luciferase lysis buffer and 50 ul luciferase substrate mix (Promega Renilla Luciferase Assay kit E2820) to each well and obtaining the results using a BIO Tek synergyx HTX luminescence reader. Afterwards the substrate mixture was removed (no washing) and 100 ul of HRP conjugated anti flag antibody 1:10000 (Abeam cat. no. AB49763) in PBS containing 1% goat serum, and 5% Tween-20 was added to each well and incubated at room temperature with mild mixing. The plate was washed 5 times and the luminescence read using Sartorius reagents EZ-ECL solutions A and B. Results are expressed as the luciferase signal divided by the flag signal.
[0284] Lumier assay to assess the calcium dependent interaction ofIQSEC2 with ARF6
[0285] Preparation of assay plates. Ninety six well plates (Greiner white Lumitrac) were coated withlOO ul of 0.1 mg / ml anti flag antibody in PBS (Sigma F1804) overnight at 4°C. Plates were washed five times with PBS containing 0.5% Tween-20 and blocked with PBS containing 5% sucrose, 3% BSA, and 0.5% Tween-20 for 90 minutes at room temperature. Plates were washed 5 times with wash buffer and kept at 4°C until further use.
[0286] Cell preparation and transfection. Cells (HeK-293T) were plated at a density of 20,000 cells / well in a 96-well plate in DMEM containing 10% FCS, 100 lU / mL penicillin, 100 pg / mL streptomycin, and 2 mM glutamine, and incubated at 37°C, 5% CO2 overnight. For each well two mixtures were prepared: Mixture A contained 200 ng IQSEC2-luc expressing plasmid, 200 ng delNARF6-flag expressing plasmid, 22.5 ul OptiMEM (Gibco) medium heated to 37°C. Mixture B contained 22.5 ul OptiMEM medium heated to 37 °C and 0.9 ul polyethyleneimine (PEI). Mixtures A and B were mixed and incubated for 25 min at room temperature and then added to the cultured cells (45 ul / well). Plates were incubated for 48 hours.
[0287] For detection of interaction in presence of high Ca+2. lonomycin was added to the cell medium at a final concentration of 10 uM and incubated for 5 min at 37°C, 5% CO2. Plates were removed from the incubator and put on ice and the medium removed. Calcium plus extraction buffer [100 ul of ice cold buffer containing 50 mM Hepes pH 7.9, 150 mM NaCl, 1 mM calcium, 0.5% Triton X-100, 5% glycerol, calmodulin (0.15 mg / ml) (Ocean Biologies), proteinase inhibitors (APExBIO cat. no. K101) phosphatase inhibitors (Sigma P5726), PMSF (1 mM), and benzamidine (1 mM)] was added to the cells and incubated on ice for 10 minutes. Aliquots (85 ul) were transferred to the anti-flag antibody coated plate and incubated for 1 hour rotating at 750 rpm at 4°C. Plates were washed 5 times with calcium plus extraction buffer.
[0288] For detection of interaction without Ca+2. Plates were removed from the incubator and put on ice and the medium removed. Calcium negative extraction buffer (100 ul of ice cold buffer containing 50 mM Hepes pH 7.9, 150 mM NaCl, 2mM EDTA, 0.5% Triton X-100, 5% glycerol, calmodulin (0.15 mg / ml), proteinase inhibitors (APExBIO cat. no. K101) phosphatase inhibitors (Sigma P5726), PMSF (1 mM), and benzamidine (1 mM)] was added to the cells and incubated on ice for 10 minutes. Aliquots (85 ul) were transferred to the anti-flag antibody coated plate and incubated for 1 hour rotating at 750 rpm at 4°C. Plates were washed 5 times with calcium negative extraction buffer.
[0289] Measurement of luciferase and flag-tagged protein. The amount of luciferase activity adhered to the plate was measured by adding 20 ul luciferase lysis buffer and 50 ul luciferase substrate mix (Promega Renilla Luciferase Assay kit E2820) to each well and the luminescence measured using a BIO Tek synergyx HTX luminescence reader. Afterwards the substrate mixture was removed by flicking the plate upside down (no washing) and 100 ul of HRP conjugated anti flag antibody 1:10000 (Abeam cat. no. AB49763) in PBS containing 1% goat serum, and 5% Tween-20 was added to each well and incubated at room temperature with mild mixing. The plate was washed 5 times and the luminescence read using Sartorius reagents EZ-ECL solutions A and B. Results are expressed as the luciferase signal divided by the flag signal.
[0290] Generation and maintenance of A350V and S1474Qfsll3 IQSEC2 mutant mice.
[0291] The generation of A350V IQSEC2 mice using CRISPR is described in Rogers et al. (2019), Front. Mol. Neurosci. 12: 43. All mice were in a C57BL6J background and housed in a specific pathogen free (SPF) environment. Mutant and wild type littermates were generated by crossing heterozygous female IQSEC2 mice with IQSEC2 wild type mice. Only male mice were used for studies to avoid the confounding effect of X inactivation of the IQSEC2 gene in female mice. All protocols used for this study were approved by the Institutional Animal Care and Use Committee (IACUC) of the Technion.
[0292] For generating S1474Q fsl33 mice it was found that the knock-in (KI) females do not care for their offspring so KI males and WT littermates were generated by crossing conditional Knock- in (conKI) females with Cre mice. As for A350V only male mice are used for analysis.
[0293] Preparation of IQSEC2 AAV
[0294] For the generation of a AAV expressing full length IQSEC2 the plasmid AM004 encoding rat IQSEC2 from the Tabuchi lab was used (Mehta et al. (2021), Cells 10: 2724). This vector has ITRs, an EFl alpha promoter and a synthetic polyadenylation site. For the generation of a AAV with the frog hybrid minigene, the AM004 backbone (keeping the same ITRs, EFla promoter and synthetic polyA site) was used with the 1488 amino acid human IQSEC2 ORF sequence (refseq mRNA accession: NM_001111125; UniProt accession :Q5JU85-2) instead of the rat ortholog. The region between L67 and G193 of the human gene was deleted and replaced with R69-Q103 of Xenopus IQSEC2 (accession number XP_002941016). AAV (rAAV) virus was generated at the ELSC vector core facility at the Hebrew University of Jerusalem. In brief, the rAAV viral genomes were packaged into AAV9 capsids via PEI-max triple transfection of 293T cells, using pAAV Rep-Cap9 and pHelper-Delta6 plasmids together with the pAAV ITR-containing transfer construct. rAAVs were harvested at 72 hours post-transfection from both cell and culture media fractions and were purified on an iodixanol ultracentrifugation gradient, followed by PBS-MK buffer exchange and concentration using Amicon Ultra-4 centrifugal filter units (PES, 100, 000 MWCO). The titer of the AAV was determined by digital PCR (dPCR) on the DNAse resistant viral genomic DNA by using ITR primers and was reported as genomic copies (GC) per ml (GC / ml).
[0295] Method of injection of AAV
[0296] Mice were injected on P1-P2 (post-natal days 1-2) via the intracerebroventricular route (icv) with a hand held Hamilton syringe with a 10mm 32g needle with a 45% bevel (Ophir Analytical Rosh Ayin Israel) as previously described. Mice were anesthetized on a chilled platform for 3 minutes prior to injection. Mice were injected with 2 ul in each ventricle with in PBS-MK buffer [IxPBS, ImM MgC12, 2.5 mM KC1, 120 mM NaCl and 0.01% pluronic] with 0.05% trypan blue and after three minutes at room temperature were returned to their cages with their mothers.
[0297] Assessment of Male Female Ultrasonic Vocalizations
[0298] Ultrasonic vocalizations were recorded using a condenser ultrasound microphone (Polaroid / CMPA, Avisoft Bioacoustics, Glienicke / Nordbahn, Germany). The microphone was connected to an ultrasound recording interface (Ultrasound Gate 116Hme, Avisoft Bioacoustics, Glienicke / Nordbahn, Germany), which was plugged into a computer equipped with the recording software Avisoft Recorder USG (sampling frequency: 250 kHz; FFT-length 1024 points; 16-bit format). Vocalizations were recorded in 8-10-week-old male wild type or A350V mice during a 5 min interaction with a female C57B1 / 6 stimulus, following a 15 min of habituation to the arena.13Ultrasonic vocalizations were analyzed using our TrackUSF custom-made software version 1.0 (Wagner lab, Haifa, Israel).35Data for vocalizations are reported as ultrasonic fragments (USFs), the number of which is directly proportional to the total time of ultrasonic vocalizations.
[0299] Quantitation of endogenous mouse IQSEC2 and viral IQSEC2 RNA
[0300] RNA extraction and real-time PCR. Brains were collected into RNA / ater solution (Invitrogen, AM7020). Total RNA was extracted using Hybrid-R™ total RNA Isolation kit (GeneAll Biotechnology), and DNA removed using a Clean-Up RNA concentrator (A&A Biotechnology). cDNAs were obtained using the qPCR cDNA Synthesis kit (Tamar, PB30.11-10), and real-time PCRs were performed with a Fast SYBR Green master mix (Applied Biosystems AB-4385612). The primers used for the real-time PCR are listed in Table 1 shown below. Mice phosphoglycerate kinase 1 (PGK1) was used as housekeeping gene for normalization. The real- time PCR program consisted of an initial 20 s at 95 °C, and then 30 cycles as follows: 95 °C for 1 s and 60 °C for 20 s Quantitative real-time-polymerase chain reaction (qRT-PCR) was performed on a StepOnePlus Real-Time PCR system (ThermoFisher 4376600).
[0301] Primer design. All primers were designed by the NCBI primer design tool (primer-BLAST) to have similar annealing temperature (optimal melting temperature was set to 60 °C). Efficiency of all of the primers was assessed using the standard curve method using serial dilutions of the cDNA and plasmid target.
[0302] Table 1: primers
[0303] Example 1: The human IQSEC2 gene structure and functional protein domains
[0304] The human IQSEC2 protein (Uniprot accession no. Q5JU85) has 1488 amino acids with a greater than 80% sequence identity between proteins from species as diverse as man, fish and the frog. There are 6 regions in the IQSEC2 gene of all organisms which encode functional protein domains that may be assayed quantitatively as outlined in Table 2. IQSEC2 is a guanine nucleotide exchange factor (GEF) for Arf6 with the GEF activity mediated by the Sec7 domain exchanging GDP for GTP on Arf6. The Sec7 activity is inhibited by the binding of apocalmodulin (ApoCM) to the IQ domain of IQSEC2 with calcium resulting in the disengagement of ApoCM from the IQ region and release of inhibition of Sec7 activity. The PDZ domain is responsible for the tethering of IQSEC2 to PSD95 and the residence of the IQSEC2 at the PSD. The pleckstrin homology domain (PH-domain) is believed to be responsible for the membrane localization of IQSEC2 juxtaposed to Arf6. The proline rich domain interacts with the IRp53 protein, and has been implicated in dendritic spine maturation. The N-terminal coiled coil region has been implicated in the dimerization of IQSEC2.
[0305] Table 2. Structural domains of the human IQSEC2 protein Example 2: Molecular modeling of IQSEC2 protein structure
[0306] The only region of IQSEC2 for which detailed structural information is available is the Sec7 domain, based on the crystal structure of the IQSEC2 Sec7 domain with Arf6. The IQ and coiled coil regions are both predicted to have an alpha-helical secondary structure. The remaining 78% of the IQSEC2 protein is disordered, as verified by using the D2P2 platform (at d2p2.pro), thus allowing for a high degree of fluidity and flexibility as required for the regulated intramolecular interaction between the alpha-helical IQ and coiled coil regions with the Sec7 domain. The molecular modeling of the wild type (WT) IQSEC2 protein and how its Sec7 catalytic activity is allosterically regulated by the binding of apocalmodulin (ApoCM) to the IQ domain of the IQSEC2 gene has been described in detail by the inventors (Shokhen et al., J Biomol. Structure 2023; 20: 1-12). The molecular dynamics equilibrated structures of IQSEC2 with and without bound ApoCM have now also been generated (shown in Figs. 1A and IB). The model suggests that when apoCM is bound to the IQ domain of IQSEC2, the N-terminal coiled region of IQSEC2 pivots around a glycine hinge (residues 342 / 343) to interact with the Sec7 domain blocking the accessibility of Sec7 to ARF6 thereby inhibiting the GEF activity of IQSEC2 (Fig. 1A). However, upon Ca+2 influx into the neuron mediated by neurotransmitter action, apoCM binds Ca+2 forming CaCM which dissociates from the IQ region of IQSEC2. With the ApoCM no longer bound to IQSEC2 the N terminal fragment pivots away from the Sec7 region allowing free access of the Sec7 region to ARF6 and thereby permitting the GEF activity of IQSEC2 to exchange GDP for GTP on Arf6 which activates ARF6 (Figs. IB).
[0307] Example 3: Molecular modeling of an IQSEC2 miniprotein (IQSEC2-hf)
[0308] In order to fit the size limitations of the AAV vector, the size of the human IQSEC2 ORF had to be reduced by at least 270 bp (encoding at least 90 aa).
[0309] In order to select the regions for deletion, the following criteria were used:
[0310] (1) deleted region must have no known function (as per Table 2);
[0311] (2) deleted region must be structurally disordered;
[0312] (3) deleted region must not include any missense mutations known to cause IQSEC2 disease;
[0313] (4) deleted region should not be conserved between diverse species;
[0314] (5) the binding energy of the interaction of the IQSEC2 full length protein with its interactors (Arf6 and apoCM) as determined using molecular dynamics equilibrated structures must be preserved in the IQSEC2 miniprotein generated by the deletion.
[0315] The IQSEC2 gene from organisms ranging from C. Elegans to man has all six functional IQSEC2 domains which are highly conserved (close to 100% identity for IQ, Sec7 and PH domains). Accordingly, it was assumed that the non-conserved sequences were unnecessary for IQSEC2 function and were therefore lost or became divergent, marking them as potential candidates for deletion while preserving function of the IQSEC2 protein.
[0316] Comparing the human IQSEC2 to the xenopus homolog, the latter appeared considerably shorter than the human protein. A basic local alignment search tool (BLAST) alignment conducted at the NCBI between the human and xenopus (GenBank accession No. XP_002941016) IQSEC2 protein sequences identified a large region of human IQSEC2 that was deleted in the xenopus protein, and that was outside of any known functional domain and outside of where any human mutation causing IQSEC2 disease had been observed (Fig. 2). This region was also highly disordered, as evaluated by the D2P2database (percent of predicted disordered residues, (PPDR) > 30%). The N-terminal boundary of sequence divergence between the xenopus and human sequences (leucine at position 67 (L67) of human IQSEC2) corresponded to the C terminal end of the coiled coil fragment (Fig. 2). Therefore the DNA sequence encoding the human amino acid sequence between L67 and G193 (not including these positions) was replaced with the corresponding xenopus DNA sequence encoding the amino acid sequence between R69 and Q103 (including these positions) of the xenopus protein, to generate a human-frog hybrid IQSEC2 (IQSEC2-hf) minigene encoding a miniprotein which is 90 amino acids shorter than the natural human IQSEC2 protein. The deleted region is colored magenta in Fig. 1.
[0317] IQSEC2-hf protein sequence (SEQ ID NO: 1, the replacement sequence appears in bold and underlined): meagsgppggpgsespnravevllelnniiesqqqlletqrrrieelegqldqltqenrdlreesqlrvereqhraeyrgesggggcgtv ghhhhhrgdnpqgvgprpprergqlsrgasrssspgaggghstststspattlqrksdgensrtvsvegdapgsdlstavdspgsqpp yrlsqlppssshmggppagvglpwaqrarlqpasvalrkqeeeeikrskalsdsyelstdlqdkkvemlerkyggsflsrraartiqtafr qyrmnknferlrssasesrmsrriilsnmrmqfsfeeyekaqnpayfegkpasldegamagarshrlerglpyggscgggidgggss vttsgefsnditeledsfskqvkslaesidealnchpsgpmseepgsaqlekreskeqqedssatsfsdlplylddtvpqqsperlpstep ppqgrpefwapaplppvpppvpsgtredgsreegtrrgpgclecrdfrlraahlplltieppsdssvdlsdrsdrgsvhrqlvyeadgcs phgtlkhkgppgrapiphrhypapegpapappgplppapnsgtgpsgvaggrrlgkceaagensdggdneslesssnsnetincssg sssrdslreppatglckqtyqretrhswdspafnndvvqrrhyriglnlfnkkpekgiqyliergflsdtpvgvahfilerkglsrqmigef Ignrqkqfnrdvldcvvdemdfssmdlddalrkfqshirvqgeaqkverlieafsqrycvcnpalvrqfmpdtifilafaiillntdmys psvkaerkmklddfiknlrgvdngediprdllvgiyqriqgrelrtnddhvsqvqavermivgkkpvlslphrrlvcccqlyevpdpn rpqrlglhqrevflfndllvvtkifqkkkilvtysfrqsfplvemhmqlfqnsyyqfgikllsavpggerkvliifnapslqdrlrftsdlres iaevqemekyrveselekqkgmmrpnasqpggakdsvngtmarssledtygagdglkrgalssslrdlsdagkrgrrnsvgsldstie gsvissprphqrmpppppppppeeyksqrpvsnsssflgslfgskrgkgpfqmpppptgqasassssassthhhhhhhhhghshg glgvlpdgqsklqalhaqycqgpgpapppylppqqpslppppqqppplpqlgsippppasappvgphrhfhahgpvpgpqhytlg rpgraprrgagghpqfaphgrhplhqptsplplyspapqhppahkqgpkhfifshhpqmmpaagaaggpgsrppggsyshphhp qsplsphspipphpsypplpppsphtphsplpptsphgplhasgppgtanppsanpkakpsristvv
[0318] IQSEC2-hf DNA sequence (SEQ ID NO: 2): atggaggcggggtcggggcccccgggcggcccgggatccgagagcccaaatcgggccgtggagtacctgctggagctgaacaacatc atcgagagccagcagcagctgctggaaacccagcggcggcgcatcgaggagctggagggccagctggaccagctcacccaggagaa ccgcgacctgcgagaggagagccagctgagagtggaacgggaacagcacagagccgagtaccggggcgagagcggcggaggagg ctgcggcacctacggccaccaccaccaccatagaggcgacaaccctcagggcgtggggccgcggccaccgcgggagcggggccag ctgagccgtggcgcatccaggagctccagtcccggcgccggcggaggccacagcaccagtaccagcaccagcccggccacgaccctc cagagaaaatctgatggtgagaattccagaacagtcagtgtggagggtgatgccccaggcagtgacctgagcacagcggttgatagtcct gggagccaacccccctaccggctgagccagctgcccccctccagcagccacatggggggcccccctgctggagtgggccttccctggg ctcagcgggcacgcctccagccagccagtgtcgccctgaggaagcaggaggaggaggagataaagcgctccaaggccctatcggaca gctatgaactctccacagacctgcaggacaagaaggtggaaatgctggagaggaagtatgggggctccttcctgagccgcagggctgcc aggaccatccagacagccttccgccagtaccgcatgaacaagaactttgagcggctacgcagctcagcctcagagagccgcatgtcccg ccgcatcatcctttccaacatgcggatgcagttctcctttgaggagtatgagaaggcacagaaccccgcgtacttcgagggcaagcctgcct cgctggacgagggtgccatggctggtgcccggagccaccggcttgaacgggggctcccatatggaggctcctgtggtgggggcatcgat ggtggtggaagctccgtcaccacatctggagagttttctaatgacatcacagaacttgaggactccttctccaaacaggtaaagtctctggct gaatccatcgacgaagccctgaactgccacccgtcagggcccatgtctgaggagccagggtcagcccagctggagaagcgggagtcaa aggaacagcaagaggacagctcagccacatccttcagtgatcttcccctctacctggatgacacagttccccaacaatcccctgagcgact gcccagcacagaacccccaccccagggccggcccgagttctgggcgccagcccctctcccgccagttcctccaccagtgccatcagga acccgggaagacggtagccgtgaggaaggcactcgcaggggtcccgggtgcttggagtgccgggatttccggctgcgggctgcccacc ttcccctgcttaccattgagcctcctagtgacagctccgtggacctgagtgaccgctcagatcgcggctctgtccaccgccagctggtgtatg aggctgatggctgcagcccccatgggaccctgaagcacaaggggccaccaggcagggccccgatcccacaccgccactacccagccc ctgaaggcccagccccagccccaccagggcccctgccaccagcccccaacagtggcactgggcccagtggtgtggctgggggtcgga ggttggggaagtgcgaggcagcaggcgagaactctgatggtggagataacgagagccttgagagctccagcaattccaatgagaccatc aactgcagctccggctcctcttctcgggacagtctaagggagcctccggctaccggcctgtgcaagcagacttaccagcgggagacaagg catagctgggactcgccagctttcaacaatgatgtggtccagaggcggcactaccgaatcggcctcaacctcttcaacaagaagccagaga agggtatccagtatctgatcgagcggggcttcctgtcagacacaccggtgggagtggctcacttcatcctggagcggaaaggcctcagcc ggcagatgataggggaattcctagggaaccggcagaagcagttcaacagagacgtgttggactgtgtggtggatgagatggacttctcctc catggatctggatgatgcgctccggaagttccagtcccatatccgggttcagggtgaggcccagaaagtggagcgactcatcgaagccttc agccagcggtactgtgtctgtaacccagccctcgtgcgccagttccggaacccagacaccatcttcatccttgcttttgccatcatcctcctca ataccgacatgtacagtcccagcgtcaaagctgaacgaaagatgaaactagatgacttcatcaagaacctgagaggggttgacaatggtga agacatcccccgagacctcctggtgggcatctaccagcgcatccaggggcgtgaactgcggaccaacgatgaccatgtgtcccaggtgc aggctgtggagcgcatgattgttggaaagaaaccagtcctgtctctccctcaccgtcgactggtttgctgctgccagctctacgaggtgcca gatccaaaccgcccccagaggctagggttgcatcagcgggaggtcttcctcttcaatgatctccttgtggtcaccaaaattttccagaagaag aagatcttggtgacgtacagtttccgtcagtctttccccctcgtggaaatgcacatgcagctcttccagaattcatattaccagtttgggatcaag ttgctgtctgcagtacctggtggggagcgaaaagtcctcatcatcttcaatgcccccagcctccaggaccggctgcgctttacatccgacctg cgcgagtccattgcggaggtgcaggagatggagaaataccgtgtggagtcggagctggagaagcagaaaggtatgatgcggcctaacg cctcacagcctggaggggccaaggactcagtgaatgggacgatggcccgcagtagcctggaggacacttacggggcaggcgatgggc tcaaacggggcgcactcagcagttccctgcgagacctctctgatgcagggaagcgggggcggcgtaacagcgtgggatcgctggacag caccatcgaagggtctgttattagcagtccacgccctcaccagaggatgccacctccgcccccacccccgccgccagaggagtacaaga gccagaggcccgtctccaactcctcatccttcctgggctccctatttggaagcaagcggggcaaggggcccttccagatgccaccaccgc caacaggccaggcctctgcctcctcttcatctgcttcttccacgcaccaccaccaccaccaccaccatcatggccatagccacggtggcctg ggggtgctgcctgatgggcagtccaagctccaggccctgcatgcccagtattgccaaggaccgggccctgccccgccaccctacctccc accccagcagccctctcttcccccacctccccagcagcccccacccttgccccagctgggctccattccaccgcctcccgcctcagcccca cctgtggggccacatcgccacttccacgcccatggcccagtcccagggccccaacactataccttgggccggccaggcagggcaccca gacggggggctggaggacaccctcagtttgctccacatggccgccaccccctgcaccagcccacatccccactgcccctgtacagtcctg ccccccagcaccctccagcccacaaacagggccctaagcacttcatcttcagccaccacccacagatgatgccagcagcaggcgcggct gggggccctggatcccggccaccagggggctcctactcccacccccaccacccccagtcaccattgtcaccacactcacccatcccacc ccacccctcctatccacccctccccccaccctcccctcacaccccgcactcaccccttccacccacctccccccatggcccgctgcacgcc tctgggccccctggcacagccaacccccccagtgcaaaccccaaggccaagccaagccggatcagcaccgtggtctga
[0319] Equilibrated molecular dynamic simulations of the structure of IQSEC2-hf demonstrated that it behaves the same as the WT protein with respect to the movement of the N terminal region to inhibit accessibility of the Sec7 region to Arf6 when binding to ApoCM (Fig. 3). The binding energies and estimated Kd of the IQSEC2-hf interaction with apoCM were found to be similar to those found for the WT protein for apoCM (Fig. 3).
[0320] The intermolecular distances between ARF6 and IQSEC2 amino acids critical to the GEF function in ARF6-IQSEC2 complexes with WT IQSEC2 (free or bound to apoCM) or IQSEC2-hf (bound to apoCM) were measured by as described in the methods section. As can be seen from Table 3, the intramolecular distances between K69 of ARF6 and D879 of IQSEC2 (WT), which stabilize the complex, are greatly increased in both the WT IQSEC2-hf and upon binding of apoCM, which results in a reduced stability of the complex.
[0321] Table 3. Intermolecular distances (A) between amino acid critical for the GEF function in ARF6-IQSEC2 complexes
[0322] L: ARF6; R: IQSEC2; free: unbound to apoCM; positions in IQSEC2 are based on WT sequence
[0323] Example 4: Experimentally assessing function of the protein product of IQSEC2-hf in vitro.
[0324] The above in silico assessments regarding the similar molecular dynamics of the IQSEC2- hf miniprotein and the WT protein were tested in vitro.
[0325] This was done using the assays described in Table 2 for several functional domains of IQSEC2. WT IQSEC2, IQSEC2-hf (mini), and two IQSEC2 mutants: A350V (A350V) and Serl474Glnfs (fs stands for frameshift, resulting in change of reading frame) (S1474Qfs) were N- terminally tagged with luciferase. PSD95, apoCM (CALM), green fluorescent protein (GFP) and IRp53 were C-terminally tagged with FLAG. The pairwise interaction between the different Luc tagged IQSEC2s and the different FLAG tagged proteins was assessed following a transient transfection of plasmid constructs expressing these potential interactors into 293T cells in a robotic Lumier assay. In this assay, the degree of interaction is quantified by assessing the amount of Luc activity captured on ELISA plates coated with anti-FLAG antibody. A representative experiment is shown in Fig. 4.
[0326] As seen in Fig. 4, both the IQSEC2-hf and the WT IQSEC2 demonstrated significant specific interactions with PSD95, CALM and iRP53 (compared to non-specific interactions with GFP) and in replicate experiments there was no consistent difference between IQSEC2-hf and WT proteins in their interaction with ApoCM, PSD95 or CALM. Binding of apoCM to the A350V mutant and PSD95 to the S1474Qfs mutant was not significantly different from non-specific binding of GFP to these IQSEC2 mutants.
[0327] To compare the binding to Arf6 and its dependence on calcium, WT IQSEC2 and the IQSEC2-hf were N-terminally tagged with luciferase, and Arf6 was C-terminally tagged with FLAG and transfected into 293T cells as described above. 48 hours after transfection the interaction strength (luc / flag) between Arf6 and the IQSEC2 proteins was measured by Lumier (as described above) in the presence of 5 mM EDTA or 1 mM CaC12. As shown in Fig. 5, There was no significant difference between IQSEC2-hf and the WT protein in their interactions with Arf6 in the absence (EDTA) or presence of free Ca+2. Furthermore, both the WT and IQSEC2-hf demonstrated a significant increase in their interaction with Arf6 with increased Ca+2.
[0328] Example 5: Assessment of the ability of the IQSEC2-hf to rescue seizures and behavioral abnormalities in A350V mice The IQSEC2-hf coding sequence was cloned into an AAV virus (a total of 4.8 kb ITR to ITR), and transfected into HEK 293T cells as described in the methods section. Viral particles were produced, and the obtained titer was 2X1013vg (viral genomes) / ml with no clumping observed. This is a much higher titer compared to a published construct (Mehta et al. (2021), Cells 10: 2724) including a full length rat 1QSEC2 gene (encoding an 1488 aa IQSEC2 protein) in an AAV vector (total length of 5113 kb, ITR to ITR). The Mehta construct provided a titer of only IX 1011vg / ml, and clumping occurred upon attempts to increase concentration.
[0329] IQSEC2-hf-containing AAV particles were administered by intracerebroventricular injections at day 1-2 into 31 mice carrying the A350V IQSEC2 mutation (described in Rogers et al. (2019), Front. Mol. Neurosci. 12: 43). 100% survival (0% lethality) was observed (p<0.01 compared to untreated mice) during the Pl 5-21 period, compared to 22% (15 / 69) lethal seizures observed in untreated A350V IQSEC2 mice.
[0330] In comparison, the Mehta FL-IQSEC2-AAV (Mehta et al. (2021), supra) was tested for ability to rescue the A350V IQSEC2 mutation in the Rogers’ mouse model by an intraventricular injection into the lateral ventricles of a 2-3 day old A350V mice. This construct was found to not significantly reduce lethal seizures from 34% (96 / 208) in A350V mice who did not receive the construct to 21% (6 / 29) in A350V mice who received the construct (p=0.174).
[0331] Example 6: Molecular modeling of an IQSEC2 miniprotein including a Gly-rich linker (IQSEC2-gly)
[0332] The DNA sequence encoding the human amino acid sequence between amino acid at positions 381-481 (not including the indicated positions) was replaced with a stretch of 9 glycines (GGGGGGGGG, SEQ ID NO: 3), to generate a human IQSEC2 (IQSEC2-gly) minigene encoding a miniprotein which is 90 amino acids shorter than the natural human IQSEC2 protein.
[0333] IQSEC2-gLy protein (SEQ ID NO: 4, the replacement sequence appears in bold and underlined): meagsgppggpgsespnraveyllelnniiesqqqlletqrrrieelegqldqltqenrdlreesqlhrgelhrdphgardspgresqyqn Iretqfhhrelresqfhqaardvgypnregayqnreavyrdkerdasyplqdttgytarerdvaqchlhhenpalgrerggreagpahp grekeagysaavgvgprpprergqlsrgasrssspgaggghstststspattlqrksdgensrtvsvegdapgsdlstavdspgsqppyrl sqlppssshmggppagvglpwaqrarlqpasvalrkqeeeeikrskalsdsyelstdlqdkkvemlerkyggsflsrraartiqtafrqy rmnknferlrssasesrmsrgggggggggchpsgpmseepgsaqlekreskeqqedssatsfsdlplylddtvpqqsperlpsteppp qgrpefwapaplppvpppvpsgtredgsreegtrrgpgclecrdfrlraahlplltieppsdssvdlsdrsdrgsvhrqlvyeadgcsph gtlkhkgppgrapiphrhypapegpapappgplppapnsgtgpsgvaggrrlgkceaagensdggdneslesssnsnetincssgss srdslreppatglckqtyqretrhswdspafnndvvqrrhyriglnlfnkkpekgiqyliergflsdtpvgvahfilerkglsrqmigeflg nrqkqfnrdvldcvvdemdfssmdlddalrkfqshirvqgeaqkverlieafsqrycvcnpalvrqfrnpdtifilafaiillntdmysps vkaerkmklddfiknlrgvdngediprdllvgiyqriqgrelrtnddhvsqvqavermivgkkpvlslphrrlvcccqlyevpdpnrp qrlglhqrevflfndllvvtkifqkkkilvtysfrqsfplvemhmqlfqnsyyqfgikllsavpggerkvliifnapslqdrlrftsdlresia evqemekyrveselekqkgmmrpnasqpggakdsvngtmarssledtygagdglkrgalssslrdlsdagkrgrmsvgsldstieg s vis sprphqrmpppppppppeeyksqrp v snsssflgslfg skrgkgpfqmpppptgqas as s s s as sthhhhhhhhhghshggl gvlpdgqsklqalhaqycqgpgpapppylppqqpslppppqqppplpqlgsippppasappvgphrhfhahgpvpgpqhytlgrp graprrgagghpqfaphgrhplhqptsplplyspapqhppahkqgpkhfifshhpqmmpaagaaggpgsrppggsyshphhpqs plsphspipphpsypplpppsphtphsplpptsphgplhasgppgtanppsanpkakpsristvv
[0334] IQSEC2-GLy DNA (SEQ ID NO: 5) atggaggcggggtcggggcccccgggcggcccgggatccgagagcccaaatcgggccgtggagtacctgctggagctgaacaacatc atcgagagccagcagcagctgctggaaacccagcggcggcgcatcgaggagctggagggccagctggaccagctcacccaggagaa ccgcgacctgcgagaggagagccagctgcaccgcggggagctgcaccgggacccccacggcgcgcgggatagcccgggccgcga gagccagtaccagaacctgcgcgagacccagttccaccaccgcgagctgcgggagagccagttccaccaggcggcccgggacgtggg ctacccgaaccgggaaggcgcctaccagaatcgggaggctgtgtatcgggacaaggagcgggacgcctcctacccgctccaggacact accggttacacagcccgcgagcgtgacgtggcccagtgccacctgcaccacgagaacccagccctgggtcgcgagcgtggcgggcgg gaggccgggccggcgcacccgggccgcgagaaggaagcgggctattcggcggcggtgggcgtggggccgcggccaccgcgggag cggggccagctgagccgtggcgcatccaggagctccagtcccggcgccggcggaggccacagcaccagtaccagcaccagcccggc cacgaccctccagagaaaatctgatggtgagaattccagaacagtcagtgtggagggtgatgccccaggcagtgacctgagcacagcgg ttgatagtcctgggagccaacccccctaccggctgagccagctgcccccctccagcagccacatggggggcccccctgctggagtgggc cttccctgggctcagcgggcacgcctccagccagccagtgtcgccctgaggaagcaggaggaggaggagataaagcgctccaaggcc ctatcggacagctatgaactctccacagacctgcaggacaagaaggtggaaatgctggagaggaagtatgggggctccttcctgagccgc agggctgccaggaccatccagacagccttccgccagtaccgcatgaacaagaactttgagcggctacgcagctcagcctcagagagccg catgtcccgcggcggcggaggcggtggcggcggcggctgccacccgtcagggcccatgtctgaggagccagggtcagcccagctgg agaagcgggagtcaaaggaacagcaagaggacagctcagccacatccttcagtgatcttcccctctacctggatgacacagttccccaac aatcccctgagcgactgcccagcacagaacccccaccccagggccggcccgagttctgggcgccagcccctctcccgccagttcctcca ccagtgccatcaggaacccgggaagacggtagccgtgaggaaggcactcgcaggggtcccgggtgcttggagtgccgggatttccggc tgcgggctgcccaccttcccctgcttaccattgagcctcctagtgacagctccgtggacctgagtgaccgctcagatcgcggctctgtccac cgccagctggtgtatgaggctgatggctgcagcccccatgggaccctgaagcacaaggggccaccaggcagggccccgatcccacacc gccactacccagcccctgaaggcccagccccagccccaccagggcccctgccaccagcccccaacagtggcactgggcccagtggtgt ggctgggggtcggaggttggggaagtgcgaggcagcaggcgagaactctgatggtggagataacgagagccttgagagctccagcaat tccaatgagaccatcaactgcagctccggctcctcttctcgggacagtctaagggagcctccggctaccggcctgtgcaagcagacttacc agcgggagacaaggcatagctgggactcgccagctttcaacaatgatgtggtccagaggcggcactaccgaatcggcctcaacctcttca acaagaagccagagaagggtatccagtatctgatcgagcggggcttcctgtcagacacaccggtgggagtggctcacttcatcctggagc ggaaaggcctcagccggcagatgataggggaattcctagggaaccggcagaagcagttcaacagagacgtgttggactgtgtggtggat gagatggacttctcctccatggatctggatgatgcgctccggaagttccagtcccatatccgggttcagggtgaggcccagaaagtggagc gactcatcgaagccttcagccagcggtactgtgtctgtaacccagccctcgtgcgccagttccggaacccagacaccatcttcatccttgcttt tgccatcatcctcctcaataccgacatgtacagtcccagcgtcaaagctgaacgaaagatgaaactagatgacttcatcaagaacctgagag gggttgacaatggtgaagacatcccccgagacctcctggtgggcatctaccagcgcatccaggggcgtgaactgcggaccaacgatgac catgtgtcccaggtgcaggctgtggagcgcatgattgttggaaagaaaccagtcctgtctctccctcaccgtcgactggtttgctgctgccag ctctacgaggtgccagatccaaaccgcccccagaggctagggttgcatcagcgggaggtcttcctcttcaatgatctccttgtggtcaccaa aattttccagaagaagaagatcttggtgacgtacagtttccgtcagtctttccccctcgtggaaatgcacatgcagctcttccagaattcatatta ccagtttgggatcaagttgctgtctgcagtacctggtggggagcgaaaagtcctcatcatcttcaatgcccccagcctccaggaccggctgc gctttacatccgacctgcgcgagtccattgcggaggtgcaggagatggagaaataccgtgtggagtcggagctggagaagcagaaaggt atgatgcggcctaacgcctcacagcctggaggggccaaggactcagtgaatgggacgatggcccgcagtagcctggaggacacttacg gggcaggcgatgggctcaaacggggcgcactcagcagttccctgcgagacctctctgatgcagggaagcgggggcggcgtaacagcg tgggatcgctggacagcaccatcgaagggtctgttattagcagtccacgccctcaccagaggatgccacctccgcccccacccccgccgc cagaggagtacaagagccagaggcccgtctccaactcctcatccttcctgggctccctatttggaagcaagcggggcaaggggcccttcc agatgccaccaccgccaacaggccaggcctctgcctcctcttcatctgcttcttccacgcaccaccaccaccaccaccaccatcatggccat agccacggtggcctgggggtgctgcctgatgggcagtccaagctccaggccctgcatgcccagtattgccaaggaccgggccctgcccc gccaccctacctcccaccccagcagccctctcttcccccacctccccagcagcccccacccttgccccagctgggctccattccaccgcct cccgcctcagccccacctgtggggccacatcgccacttccacgcccatggcccagtcccagggccccaacactataccttgggccggcc aggcagggcacccagacggggggctggaggacaccctcagtttgctccacatggccgccaccccctgcaccagcccacatccccactg cccctgtacagtcctgccccccagcaccctccagcccacaaacagggccctaagcacttcatcttcagccaccacccacagatgatgcca gcagcaggcgcggctgggggccctggatcccggccaccagggggctcctactcccacccccaccacccccagtcaccattgtcaccac actcacccatcccaccccacccctcctatccacccctccccccaccctcccctcacaccccgcactcaccccttccacccacctccccccat ggcccgctgcacgcctctgggccccctggcacagccaacccccccagtgcaaaccccaaggccaagccaagccggatcagcaccgtg gtctga
[0335] Example 7: Assessment of the ability of the IQSEC2-gly to rescue seizures and behavioral abnormalities in A350V mice
[0336] The IQSEC2-gly coding sequence is cloned into an AAV virus (a total of about 4.8 kb ITR to ITR), and transfected into cells as described in the methods section. Viral particles are produced.
[0337] IQSEC2-gly-containing AAV particles are administered by intracerebroventricular injections at day 1-2 into 31 mice carrying the A350V IQSEC2 mutation (described in Rogers et al. (2019), Front. Mol. Neurosci. 12: 43).
[0338] Example 8: Further molecular modeling and miniprotein design The inventors have further identified multiple stable structures of the IQSEC2 protein, having different conformations, grouped them into folded and unfolded conformations, and ranked their stability. Next, potential pairs of boundary amino acids, which encompass a putative deletion, were selected and the intramolecular distances between the pair in the most stable folded and unfolded conformations were calculated and compared by applying YAS ARA Structure software to aMD equilibrated structure of IQSEC2. Pairs of boundary amino acids having a short intramolecular distance (i.e., up to 20 Angstroms) but a large linear distance (so as to have a large as possible deletion, such as greater than 50 amino acids) and for which the intramolecular distance was similar between folded and unfolded forms - were selected as best candidates to be the boundary amino acids of a deletion.
[0339] The analyses yielded a pair of amino acids - at positions 391 (Gin) and 627(Ser), that were at an intramolecular distance of only 6 Angstroms in both the folded and unfolded forms.
[0340] As a replacement sequence to replace the deleted sequence (aa392-626), a short artificial linker sequence predicted to have an alpha helix structure of 5 amino acids (Phe-Ser-Lys-Gln-Val, or FSKQV, SEQ ID NO: 6) is selected, which spans the 6 Angstroms needed to maintain the distance between the boundary amino acids. Additional flexible linkers, including Gly-Gly were also used, and the conservation of the structure and molecular interactions of this new minigene has been confirmed by molecular dynamics and in vitro protein-protein interaction analysis (using the Lumier assay described above).
[0341] An additional candidate for a deletion region was also identified, between positions 114 and 193 (not inclusive), which is a partial deletion of the previously identified deletion between positions 67 and 193. As dimerization potential may be important to the function of the protein and therefore rescue of the mutant phenotype, the dimerization ability of the partial deletion was tested by fluorescent resonance energy transfer (FRET) assay, using two chromophores (GFP and CHERRY for example) and testing for dimerization by FRET using flow cytometry or Confocal microscopy. It was found that the partial deletion maintained a dimerization ability similar to that of wild-type.
[0342] Example 9: Adaptation of AAV genome
[0343] In order to further minimize the AAV size, adaptation was made to additional elements of the construct, namely the promoter and the 3’ untranslated region of the gene (3’-UTR).
[0344] As for a promoter, four promoters were considered, trying to minimize size, as well as direct expression to neurons. While an AAV EFla promoter (212 bp) was used in the above examples, the three additional promoters tested were neuron specific: a calmodulin (CAM) promoter (CAMI) (90 bp), a pseudorabies (PRV) miniLAP (latency-associated) promoter (278 bp), and a human synapsin (Hsyn) promoter (440 bp).
[0345] Additionally, in order to achieve physiological levels of the transgene mRNA, the inventors have identified, using TARGETSCAN, a small region of 290bp within the 1343 bp 3’-UTR of IQSEC2 mRNA (NM_001111125.3), which contains all of the miRNA target sequences which have been predicted or experimentally shown to be involved in regulation of IQSEC2 by miRNAs 135-5p, 17-5p, 33-5p, and 31-5p, as well as a TTATTTATT element that has been shown to regulate transcript stability for many other transcripts. This 290 bp region NM_001111125.3 positions 5521-5811 is inserted 3’ of the modified gene sequence.
[0346] Example 10: Effect of silencing miRNA for treating gain-of-function mutations
[0347] For several IQSEC2 mutations (A350V, A350D, A350T) the mutation results in a gain of function (dominant-negative mutation) and thus it is necessary to knock down the mutant gene and replace it with a functional gene. To this end, a silencing miRNA (miRNA50, GGTTCTGGTACTGGCTCTCGCGGC, SEQ ID NO: 7 (y=c)) is used, which is complementary to a region in the IQSEC2 open reading frame that has been deleted in the IQSEC2-hf construct, i.e., the region encoding amino acids flanked by positions 67 and 193 of IQSEC2. As a result, miRNA50 binds to and signal for degradation mRNA from the endogenous (mutant) IQSEC2 gene but not to mRNA derived from the AAV-IQSEC2-hf construct, which lacks sequence complementary to the miRNA. It is noted that there is a single nucleotide difference (bold and underlined) between the human and the mouse in the region complementary to miRNA50 (the human corresponding sequence to miRNA50 sequence being GGTTCTGGTACTGGCTTTCGCGGC, SEQ ID NO: 7 (y=t). As can be seen below, miRNA50, having a mismatch with the human sequence, was nevertheless capable of silencing the human mRNA.
[0348] To test the silencing effect of miRNA50 on wild type and modified IQSEC2 transcripts, a plasmid with a CMV promoter driving expression of an N-terminal Luciferase tag linked either to a wild type human IQSEC2 gene or the IQSEC2-hf minigene was co-transfected into HEK 293T cells together with a plasmid with a U6 RNA pol III promoter driving expression of the miRNA50 having a T6 termination sequence, at different ratios (1:1 -1:10 as shown in Fig. 6) between the plasmids. The transfection was carried out with a PEI transfection reagent in a 96 well format. After 48 hours the amount of luciferase activity was assessed as a readout of the amount of IQSEC2 mRNA that remained at that time.
[0349] As seen from Fig. 6, the miRNA50 efficiently knocked down the wild type gene, while only marginally affecting expression of the IQSEC2-hf modified gene, proving that such a strategy could be used for treating dominant-negative mutations.
[0350] It is further noted that the modified IQSEC2 gene may be further reduced in size to allow adding the miRNA cassette to the same AAV vector. Alternatively, the treatment could be carried out by using two vectors, one for the knockdown (carrying the miRNA) and the other for the replacement, carrying the modified gene.
[0351] Similar experiments are being carried out with additional miRNA sequences which correspond to the same region deleted in IQSEC2-hf: miRNA218 TAACCGGTAGTGTCCTGRAG (SEQ ID NO: 8); miRNA219 GTAACCGGTAGTGTTCCTGRA (SEQ ID NO: 9); miRNA220 TGTAACCGGTAGTGTTCCTGR (SEQ ID NO: 10); and miRNA263 TCGTGGTGCAGGTGGCAYTG (SEQ ID NO: 11), where Y stands for C or T, and R stands for A or G.
[0352] Example 11: shRNA for treating gain-of- function mutations
[0353] Two constructs of AAV were prepared (human / mouse shRNA 126), expressing shRNA recognizing the sequence encoding amino acids 126-132 of the IQSEC2 gene (REAVYRD (SEQ ID NO: 16) in human), as shown:
[0354] Human shRNA 126 (SEQ ID NO: 12):
[0355] CCCGGAGGCTGTGTATCGGGACAATTCAAGAGATTGTCCCGATACACAGCCTCCTTT TTTGGAAAT; and
[0356] Mouse shRNA 126 (SEQ ID NO: 13):
[0357] CCCGGAAGCTATCTATCGGGATAATTCAAGAGATTATCCCGATAGATAGCTTCCTTT TTTGGAAAT.
[0358] An shRNA resistant construct was prepared by changing the human DNA sequence encoding amino acids 126-132 from GGAGGCTGTGTATCGGGAC (SEQ ID NO: 14) to aGAaGCcGTaTAcaGaGAt (SEQ ID NO: 15), encoding the same protein sequence between amino acids 126-132 but resistant to shRNA126 (lower case nucleotides are where changes were made to make the sequence shRNA resistant).
[0359] As can be seen from Fig. 7, while expression of the wild type construct is reduced by human shRNA 126, but not mouse shRNA 126 (which is not identically complementary to the human mRNA), expression of the shRNA resistant construct is not affected by the human or mouse shRNA. Example 12: An S1474Q IQSEC2 model mouse
[0360] To assess the efficacy of the AAV therapy, a model mouse of the most common human missense mutation of IQSEC2, S1474Q, was generated. The mutation is conditional, facilitating it to be generated at any time during development or in any cell in the body by providing a Cre recombinase.
[0361] A model mouse for a humanized mutated Iqsec2 was generated. Briefly, a targeting construct was prepared in a plasmid vector such as pcDNA3.1 with hygromycin neomycin cassettes, including mouse genomic Iqsec2 exons 13 to 15 inserted (“floxed”) between loxP sites, with a knock-in mouse exons 13 and 14 and a chimeric exon 15 following the floxed region (see Fig. 8A). The chimeric exon 15 consisted of murine sequence up to murine codon Pl 464 and human sequence from human codon S1474 onwards. A frame-shift was introduced by the insertion of a C nucleotide immediately 5’ of S1474. An additional 107bp of human genomic sequence and an exogenous Sv40 pA sequence was also introduced after chimeric exon 15. The targeting construct was electroporated into male C57BL / 6J ES cells. Homologous recombinant ES cell clones were selected using hygromycin and neomycin, confirmed by qPCR and injected into goGermline™ male blastocysts and the blastocysts were implanted in uterus of female mice. Male chimeric mice were obtained and crossed to C57BL / 6J flp females to establish heterozygous germline offspring with the conditional IQSEC2 allele (conKI) with hygromycin and neomycin selection cassettes excised. In order to obtain mice which expressed the mutant IQSEC2 allele (converting the conditional allele to the mutant allele) conKI heterozygote females were crossed with male mice ubiquitously expressing CRE recombinase (homozygous at the ROSA locus). Cre- mediated recombination excised murine exons 13-15 of the conKI allele to allow for expression of humanized Iqsec2 with the S1474Q fs mutation (the KI allele). It is noted that the Cre recombinase may also be provided locally or at a certain time during development, e.g., by providing the Cre recombinase gene on a vector having a specific expression pattern, or under the control of an inducible promoter, e.g. with expression of the Cre recombinase controlled by a specific drug such as tamoxifen.
[0362] The targeting construct was fully sequenced by Sanger sequencing prior to electroporation into ES cells. The complete cDNA and protein product sequence of the mutant IQSEC2 is provided in SEQ ID NO: 27 (CDNA) and SEQ ID NO: 28 (protein), respectively.
[0363] The mutant mouse line was tested and found to exhibit delayed growth, early mortality from seizures (25% mortality rate by 2 months of age), and behavioral abnormalities (lack of ultrasonic vocalizations), data not shown.
Claims
CLAIMS1. A modified nucleic acid molecule comprising a modified gene sequence encoding a modified protein having a reduced length compared to a parental protein of the modified protein and essentially the same function as the parental protein, the modified protein differing from the parental protein by lacking at least one non-functional sequence of the parental protein, wherein the non-functional sequence has at least three features selected from: (1) having no defined function, (2) having a disordered structure, (3) not comprising positions associated with disease-causing mutations, and (4) not being conserved in at least one different species.
2. The modified nucleic acid molecule of claim 1, wherein the at least three features comprise: (1) having no defined function, (2) having a disordered structure, and (3) not comprising positions associated with disease-causing mutations.
3. The modified nucleic acid molecule of claim 1 or 2, wherein the parental protein is at least about 1000 amino acids long.
4. The modified nucleic acid molecule of any one of claims 1-3, wherein the parental protein has at least two regions having a known function.
5. The modified nucleic acid molecule of any one of claims 1-4, wherein the parental protein has at least one disordered region.
6. The modified nucleic acid molecule of any one of claims 1-5, wherein the parental protein comprises as least one position associated with a disease-causing mutation.
7. The modified nucleic acid molecule of any one of claims 1-6, wherein the parental protein has at least two regions which are conserved in at least one different species.
8. The modified nucleic acid molecule of any one of claims 1-7, wherein the non-functional sequence is internal to the parental gene sequence encoding the parental protein, and is flanked by two boundary amino acids.
9. The modified nucleic acid molecule of claim 8, wherein the intramolecular distance between the two boundary amino acids is not more than about 20 Angstroms.
10. The modified nucleic acid molecule of any one of claims 1-9, wherein the non-functional sequence has a length of at least about 20-2000 amino acids.
11. The modified nucleic acid molecule of any one of claims 1-10, wherein the non-functional sequence is located between two regions having defined structures.
12. The modified nucleic acid molecule of any one of claims 1-11, wherein the non-functional sequence is located between two regions having defined functions.
13. The modified nucleic acid molecule of any one of claims 1-12, wherein the non-functional sequence is located between two regions which are conserved in at least one different species.
14. The modified nucleic acid molecule of any one of claims 1-13, further comprising a sequence encoding a replacement sequence shorter than the non-functional sequence, wherein the replacement sequence replaces the non-functional sequence in the modified protein and is not part of the parental protein.
15. The modified nucleic acid molecule of claim 14, wherein the size difference between the nonfunctional sequence and the replacement sequence is about 20-2000 amino acids.
16. The modified nucleic acid molecule of claim 14 or 15, wherein the replacement sequence has a flexible and / or a disordered structure.
17. The modified nucleic acid molecule of claim 16, wherein the replacement sequence is a flexible linker.
18. The modified nucleic acid molecule of claim 17, wherein the flexible linker comprises at least 50%, 60%, 70%, 80%, or 90% glycine residues.
19. The modified nucleic acid molecule of any one of claims 14-18, wherein the replacement sequence has at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% sequence identity with a sequence from a different species located at a region corresponding to the location of the non-functional sequence in the parent protein.
20. The modified nucleic acid molecule of any one of claims 1-19, wherein the parental protein is associated with a disease or condition.
21. The modified nucleic acid molecule of claim 20, wherein the disease or condition is autism, epilepsy, and / or IQSEC2-related disorder22. The modified nucleic acid molecule of claim 21, wherein the parental protein is IQSEC2.
23. The modified nucleic acid molecule of claim 22, wherein the non-functional sequence comprises or is comprised in a sequence between amino acids 67 and 193, 73 and 193, 114 and 193, 391 and 627, and / or 381 and 481, of the human IQSEC2 protein.
24. The modified nucleic acid molecule of claim 22 or 23, further comprising a sequence encoding a replacement sequence shorter than the non-functional sequence and replacing the nonfunctional sequence, wherein the replacement sequence has at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% sequence identity to a sequence of a different species located between amino acid positions corresponding to amino acids 67-193, 73-193, 114-193, 391-627, and / or 381-481 of the human IQSEC2 protein.
25. The modified nucleic acid molecule of claim 24, wherein the different species is Xenopus laevis.
26. The modified nucleic acid molecule of any one of claims 14-25, wherein the replacement sequence comprises a stretch of at least 4, 5, 6, 7, 8, or 9 glycine residues.
27. The modified nucleic acid molecule of any one of claims 1-26, comprising a sequence at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 2 or 5, or encoding a sequence at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 1 or 4.
28. A method for preparing a modified nucleic acid molecule comprising a modified gene sequence encoding a modified protein having a reduced length compared to a parental protein of the modified protein and essentially the same function as the parental protein, the method comprising: a) providing a nucleic acid molecule comprising a parental gene sequence encoding the parental protein; and b) deleting from the parental gene sequence at least one sequence encoding a non-functional protein sequence, thereby generating the modified gene sequence encoding the modified protein, wherein the non-functional sequence has at least three features selected from: (1) having no defined function, (2) having a disordered structure, (3) not comprising positions associated with disease-causing mutations, and (4) not being conserved in at least one different species.
29. The method of claim 28, further comprising prior to step (b) a step of mapping the parental protein for at least three features selected from: (1) regions having a defined function, (2)regions having a disordered structure, (3) positions associated with disease-causing mutations, and (4) conservations across species.
30. The method of claim 28 or 29, further comprising, following step (b), a step of adding to the modified gene sequence at least one sequence encoding a replacement sequence shorter than the non-functional sequence, which replaces in the modified protein the sequence encoding the non-functional sequence.
31. The method of any one of claims 28-30, further comprising, following step (b), additional steps for evaluating the modified protein, comprising: c) obtaining for the parental protein at least a first value of at least one parameter selected from: i) a binding energy with at least one interacting component, and ii) an intermolecular distance between an amino acid of the protein and an interacting component; d) obtaining for the modified protein at least a second value of the at least one parameter calculated in (c) for the parental protein; and e) comparing the first value and the second value, wherein the second value is not more than about 30% higher or 30% lower than the first value.
32. The method of claim 31, wherein the non-functional sequence is internal to the parental protein and is flanked by two boundary amino acids, and the method further comprises, following step (b), a step of adding to the modified gene sequence at least one sequence encoding a replacement sequence shorter than the non-functional sequence, which replaces in the modified protein the sequence encoding the non-functional sequence; and the at least one parameter further comprise an intramolecular distance between the two boundary amino acids.
33. A modified nucleic acid molecule prepared by the method of any one of claims 28-32, comprising a modified gene sequence encoding a modified protein.
34. A modified protein encoded by the modified gene sequence comprised in the modified nucleic acid molecule of any one of claims 1-27 and 33.
35. The modified protein of claim 34, comprising a sequence at least about 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to SEQ ID NO: 1 or 4.
36. An adeno-associated virus (AAV)-based construct comprising the modified nucleic acid molecule of any one of claims 1-27 and 33, or encoding the modified protein of claim 34 or37. A viral particle or viral-like particle (VLP) comprising the AAV-based construct of claim 36.
38. A host cell comprising the modified nucleic acid of any one of claims 1-27 or the AAV-based construct of claim 36, or expressing the modified protein of claim 34 or 35.
39. A system comprising the modified nucleic acid molecule of any one of claims 1-27, or the modified protein of claim 34 or 35, and further comprising an inhibitory RNA or a sequence encoding an inhibitory RNA, wherein the inhibitory RNA is directed to a nucleotide sequence comprised in the sequence encoding the non-functional sequence and not in the modified protein sequence.
40. The system of claim 39, wherein the inhibitory RNA is a short hairpin RNA (shRNA), small interfering RNA (siRNA) or a microRNA (miRNA).
41. The system of claim 39 or 40, wherein the inhibitory RNA is an miRNA at least 85%, 90%, 95%, or 99% identical to a sequence selected from SEQ ID NO: 7, 8, 9, 10, and 11, or an shRNA at least 85%, 90%, 95%, or 99% identical to SEQ ID NO: 12.
42. The system of any one of claims 39-41, for use in the treatment of a disease or condition associated with a dominant-negative mutation.
43. A composition comprising the modified nucleic acid molecule of any one of claims 1-27, the modified protein of claim 34 or 35, the AAV-based construct of claim 36, the viral particle or VLP of claim 37, the host cell of claim 38, or the system of any one of claims 39-42.
44. The composition of claim 43, wherein the viral particle or VLP has a titer of at least about 1012or 1013viral genomes / ml composition.
45. The composition of claim 43 or 44, further comprising a pharmaceutically acceptable carrier.
46. The composition of claim 45, for use in a method of genetic therapy of a disease or a condition in a subject in need thereof.
47. A method of treating a disease or condition by genetic therapy in a subject in need thereof, comprising administering to the subject a therapeutically-effective dose of the composition of claim 45.
8. The system for use of claim 42, the composition for use of claim 46, or the method of claim47, wherein the disease or condition is selected from autism, epilepsy, and / or IQSEC2-related disorder.