Methods for modifying gene expression for hereditary disorders

A novel gene editing method using rare-cutting endonucleases and bimodule transgenes addresses the challenges of delivering large transgenes and correcting gain-of-function mutations, achieving precise gene modification and therapeutic benefits for monogenic disorders.

JP2026077682APending Publication Date: 2026-05-13BLUEALLELE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
BLUEALLELE CORP
Filing Date
2026-02-06
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Existing gene therapy methods face challenges in delivering large transgenes and addressing alternative splicing patterns and gain-of-function mutations in monogenic disorders, particularly for diseases like spinocerebellar ataxia and Parkinson's disease, due to size limitations and the need for precise gene editing.

Method used

A novel approach using rare-cutting endonucleases and bimodule bidirectional transgenes, integrated via homologous recombination or non-homologous end joining, to incorporate silencing-resistant coding sequences into endogenous genes, allowing precise modification of protein products, even when standard vectors are insufficient.

Benefits of technology

Enables precise correction of mutations and reduction of pathogenic gene expression, effectively treating disorders with gain-of-function mutations and multiple isoforms, compatible with current delivery vehicles like adeno-associated virus vectors and lipid nanoparticles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026077682000003
    Figure 2026077682000003
  • Figure 2026077682000004
    Figure 2026077682000004
  • Figure 2026077682000005
    Figure 2026077682000005
Patent Text Reader

Abstract

Providing a method for modifying gene expression for hereditary disorders. [Solution] Methods and compositions for modifying the expression of endogenous genes or modifying the coding sequences of endogenous genes using rare-cutting endonucleases and transposases. In one embodiment, the methods described herein provide a novel approach for treating gain-of-function disorders, wherein pathogenic alleles and non-pathogenic alleles are silenced and protein expression is replaced with a silencing-resistant coding sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Reference to Related Applications This application claims priority to prior applications and co-pending applications USSN62 / 754,548, filed November 1, 2018, USSN62 / 755,755, filed November 5, 2018, USSN62 / 756,175, filed November 6, 2018, and USSN62 / 799,615, filed January 31, 2019, the contents of each of which are hereby incorporated by reference in their entirety.

[0002] Sequence Listing This application includes a sequence listing submitted in ASCII format via EFS-Web, which is hereby incorporated by reference in its entirety. The ASCII copy is named SEQ_LISTING_BA2018-5_P12988 and was created on October 29, 2019 and is 507,904 bytes in size.

[0003] This document is in the field of genome editing and gene therapy. More specifically, this document relates to the targeted modification of endogenous genes or the reduction of endogenous gene expression along with gene expression from transgenes.

Background Art

[0004] Monogenic disorders, such as sickle cell disease (hemoglobin-β gene), cystic fibrosis (cystic fibrosis transmembrane conduction regulator gene), and Tay-Sachs disease (β-hexosaminidase A gene), are caused by one or more mutations in a single gene. The treatment of monogenic disorders has garnered interest in gene therapy because replacing the defective gene with a functional copy may offer therapeutic benefits. However, one obstacle to developing effective therapies is the size of the functional copy of the gene. Many delivery methods, including those using viruses, have size limitations that hinder the delivery of large transgenes. Furthermore, many genes exhibit alternative splicing patterns, resulting in a single gene encoding multiple proteins. Methods to modify the region of the defective gene may offer additional means for treating monogenic disorders. [Overview of the project] [Means for solving the problem]

[0005] Gene editing holds promise for correcting mutations found in hereditary disorders, but many challenges remain in creating effective therapies for individual disorders, including those caused by gain-of-function mutations or those requiring precise repair. These challenges are seen in disorders such as spinocerebellar ataxia² and Parkinson's disease, which are associated with gain-of-function mutations.

[0006] In one embodiment, the methods described herein provide a novel approach to treat gain-of-function disorders in which pathogenic and non-pathogenic alleles are silenced and protein expression is replaced using a silencing-resistant coding sequence. These methods can be used for genes producing one or more isoforms. In one embodiment, a rare-cutting endonuclease or transposon can be used to incorporate a transgene containing a silencing sequence and a complete or partial coding sequence for silencing resistance into an endogenous gene (Figures 12-17). If the transgene contains a silencing-resistant partial coding sequence, the transgene may further include a splice acceptor or splice donor operably ligated to the partial coding sequence. The transgene may further include a promoter operably ligated to the silencing-resistant coding sequence (when targeting the 5' region of the gene), or a terminator operably ligated to the silencing-resistant coding sequence (when targeting the 3' region of the gene). Gain-of-function mutations include HD (Huntington's disease), SBMA (spinal and bulbar muscular atrophy), SCA1 (spinocerebellar ataxia type 1), SCA2 (spinocerebellar ataxia type 2), SCA3 (spinocerebellar ataxia type 3 or Machado-Joseph disease), SCA6 (spinocerebellar ataxia type 6), SCA7 (spinocerebellar ataxia type 7), Fragile X syndrome, Fragile XE intellectual disability, Friedreich's ataxia, myotonic dystrophy type 1, myotonic dystrophy type 2, spinal This mutation may result in a disease selected from the group consisting of ataxia 8, ataxia 12, spinal and bulbar muscular atrophy, JPH3, amyotrophic lateral sclerosis (ALS), hereditary motor-sensory neuropathy type IIC, postsynaptic slow-channel congenital myasthenic syndrome, PRPS1 hyperactivity, Parkinson's disease, tubular aggregate myopathy, achondroplasia, Rubus X-linked intellectual disability syndrome, and autosomal dominant retinitis pigmentosa.

[0007] In another aspect, the methods described herein provide a novel approach to correct mutations found at the 5' end of a gene. These methods are partly based on the design of bimodule bidirectional transgenes compatible with integration through multiple repair pathways. The transgenes described herein can be integrated into a gene via homologous recombination (HR), non-homologous end joining (NHEJ), or both homologous and non-homologous end joining pathways, or via transposition. Furthermore, the results of any integration (HR, NHEJ forward, NHEJ reverse, forward transposition, or reverse transposition) may result in precise correction / modification of the protein product of the target gene. The transgenes described herein can be used to fix or introduce mutations in the 5' region of a target gene. These methods are particularly useful when precise gene editing is required, or when the targeted mutated endogenous gene cannot be "replaced" by synthetic copying because it exceeds the size capacity of a standard vector or viral vector. The methods described herein can be used for applied research (e.g., gene therapy) or basic research (e.g., creation of animal models or understanding of gene function).

[0008] The methods described herein are compatible with current in vivo delivery vehicles (e.g., adeno-associated virus vectors and lipid nanoparticles) and address several challenges by achieving precise modification of gene products (particularly those with gain-of-function mutations and those producing multiple isoforms).

[0009] In one embodiment, this document features a method for incorporating a transgene into an endogenous gene. The method may include delivery of the transgene, which harbors first and second splice donor sequences, first and second coding sequences, and one bidirectional promoter, or first and second promoters (Figure 1). In another embodiment, the transgene may also include first and second terminators. In some embodiments, the first and second terminators may be replaced with a single bidirectional terminator. The method further includes administering a rare-cutting endonuclease that targets a site within the endogenous gene. As a result of the method, the transgene is incorporated into the endogenous gene, and regardless of orientation (e.g., forward or reverse), the incorporation will result in a precise modification of the amino acid sequence of the protein produced from the endogenous gene (Figures 3 and 4). The method may include the use of any suitable rare-cutting endonuclease, including CRISPR, TAL effector nuclease, zinc finger nuclease, or meganuclease. The rare-cutting endonuclease may target sequences within introns or exons of an endogenous gene. The endogenous gene may include the ATXN2 gene, and the rare-cutting endonuclease may target intron 1 or exon 1 of the ATXN2 gene. In some embodiments, the CRISPR nuclease may be a CRISPR / Cas12a nuclease or a CRISPR / Cas9 nuclease. In other embodiments, the first and second coding sequences may encode amino acids homologous to the amino acids encoded by the reporter gene, purified tag, or endogenous gene. The first and second coding sequences encode the same amino acid by possessing the same nucleic acid sequence or by possessing different nucleic acid sequences (e.g., using codon degeneracy). The transgene can be synthesized on a viral vector (e.g., an adenovirus vector, an adeno-associated virus vector, or a lentiviral vector). Alternatively, the transgene can be synthesized on a non-viral vector.The embodiments described above may result in targeted integration of the transgene in either a forward or reverse direction, while both products still yield the desired outcome.

[0010] In one embodiment, this document features a method for incorporating a transgene into an endogenous gene. This method may include the delivery of a transgene, the transgene having first and / or second homologous arms, first and second rare-cutting endonuclease target sites, first and second promoters or one bidirectional promoter, first and second splice donor sequences, first and second coding sequences, and optionally, first and second terminators. In some embodiments, the first and second terminators may be replaced with a single bidirectional terminator. The method further includes administering rare-cutting endonucleases targeting sites in the endogenous gene and two sites in the transgene. As a result of this method, the transgene will be incorporated into the endogenous gene, and regardless of orientation (e.g., forward or reverse), the incorporation will result in a precise modification of the amino acid sequence of the protein produced from the endogenous gene. The method may include the use of any suitable rare-cutting endonuclease, including CRISPR, TAL effector nuclease, zinc finger nuclease, or meganuclease. The rare-cutting endonuclease may target sequences within introns or exons of an endogenous gene. The endogenous gene may include the ATXN2 gene, and the rare-cutting endonuclease may target intron 1 or exon 1 of the ATXN2 gene. In some embodiments, the CRISPR nuclease may be a CRISPR / Cas12a nuclease or a CRISPR / Cas9 nuclease. In other embodiments, the first and second coding sequences may encode amino acids homologous to the amino acids encoded by the reporter gene, purified tag, or endogenous gene. The first and second coding sequences encode the same amino acid by possessing the same nucleic acid sequence or by possessing different nucleic acid sequences (e.g., using codon degeneracy). The transgene can be synthesized on a viral vector (e.g., an adenovirus vector, an adeno-associated virus vector, or a lentiviral vector). Alternatively, the transgene can be synthesized on a non-viral vector.The embodiments described above may result in targeted integration of the transgene in either a forward or reverse direction, while both products still yield the desired outcome.

[0011] In further embodiments, this document features a double-stranded polynucleotide. A double-stranded polynucleotide may include first and second splice donor sequences, first and second coding sequences, a bidirectional promoter, or first and second promoters. The double-stranded polynucleotide may further include first and / or second homologous arms, first and second rare-cutting endonuclease target sites, and first and second terminators. In some embodiments, the first and second terminators may be replaced by a single bidirectional terminator. The coding sequences on the double-stranded polynucleotide may be reverse complementary in orientation. The coding sequences may encode the same amino acid sequence. The coding sequences may consist of the same nucleotide sequence or different nucleic acid sequences (e.g., due to codon degeneracy). The first and second promoters may be reverse complementary in orientation.

[0012] In further embodiments, this document features a method for incorporating a transgene into ATXN2. This method may involve administering a polynucleotide encoding a rare-cutting endonuclease targeting a site within the ATXN2 gene, and a transgene to be incorporated into the ATXN2 gene after cleavage by the rare-cutting endonuclease. In another embodiment, the rare-cutting endonuclease may be delivered in the form of a protein (e.g., Cas9 or Cas12a protein, or TALEN protein) or a ribonucleoprotein complex (e.g., Cas9 or Cas12a together with the corresponding gRNA). The transgene can be incorporated into cells including induced pluripotent stem cells, Purkinje cells, granule cells, neuronal cells, or glial cells. The transgene incorporated into the ATXN2 gene may possess the coding sequence of exon 1 of the ATXN2 gene. The transgene can be incorporated into intron 1 or exon 1 of the ATXN2 gene. The transgene may further include a promoter upstream of the coding sequence. The integration of the transgene can be facilitated using any suitable rare-cutting endonuclease, including CRISPR, TAL effector nuclease, zinc finger nuclease, or meganuclease. The transgene can be synthesized on a viral vector (e.g., adenovirus vector, adeno-associated virus vector, or lentiviral vector). Alternatively, the transgene can be synthesized on a non-viral vector.

[0013] In another embodiment, this document features a method for modifying the expression of an endogenous gene, the method comprising administering a transgene, the transgene comprising a first and second promoter, or a bidirectional promoter, a first nucleic acid sequence that reduces the expression of the endogenous gene, and a second nucleic acid sequence encoding a protein homologous to the protein produced by the endogenous gene. The second nucleic acid sequence may comprise a different nucleic acid sequence compared to the first nucleic acid sequence (e.g., due to codon degeneracy or absence of a sequence). The transgene described herein may further comprise first and second terminators operably ligated to the first and second nucleic acid sequences. The transgene may be used when at least one allele contains a gain-of-function mutation. Gain-of-function mutations include HD (Huntington's disease), SBMA (spinal and bulbar muscular atrophy), SCA1 (spinocerebellar ataxia type 1), SCA2 (spinocerebellar ataxia type 2), SCA3 (spinocerebellar ataxia type 3 or Machado-Joseph disease), SCA6 (spinocerebellar ataxia type 6), SCA7 (spinocerebellar ataxia type 7), Fragile X syndrome, Fragile XE intellectual disability, Friedreich's ataxia, myotonic dystrophy type 1, myotonic dystrophy type 2, spinal The mutation may result in a disease selected from the group consisting of ataxia 8, ataxia 12, spinal and bulbar muscular atrophy, JPH3, amyotrophic lateral sclerosis (ALS), hereditary motor-sensory neuropathy type IIC, postsynaptic slow-channel congenital myasthenic syndrome, PRPS1 hyperactivity, Parkinson's disease, tubular aggregate myopathy, achondroplasia, Rubus X-linked intellectual disability syndrome, and autosomal dominant retinitis pigmentosa. The transgene may be carried in a viral vector, including an adenovirus vector, an adeno-associated virus vector, or a lentiviral vector. The transgene may be 4.7 kb or less in size. The transgene may be on a non-viral vector. The transgene may be integrated into the cell genome.

[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art in which the invention relates. The invention can be carried out using methods and materials similar or equivalent to those described herein, but preferred methods and materials are described below. All publications, patent applications, patents, and other references referenced herein are incorporated by reference in their entirety. In case of any conflict, including definitions, this specification shall prevail. Furthermore, materials, methods, and examples are illustrative and not intended to limit the scope.

[0015] Details of one or more embodiments of the present invention are shown in the following description. Other features, purposes, and advantages of the present invention will become apparent from the description and claims. In embodiments of the present invention, for example, the following items are provided. (Item 1) A method for integrating an introduced gene into an endogenous gene, a. Administering a transgene, wherein the transgene is i. First and second splice donor sequences, ii. The first and second subcode sequences, and iii. Administering a transgene containing one bidirectional promoter, or a first and second promoter, The procedure includes administering at least one rare-cutting endonuclease that targets a site within the endogenous gene, A method for incorporating the introduced gene into the endogenous gene. (Item 2) The method according to item 1, wherein the first splice donor is operably connected to the first subcode sequence, and the second splice donor is operably connected to the second subcode sequence. (Item 3) The method according to item 2, wherein the first subcode sequence is operably coupled to the first promoter, and the second subcode sequence is operably coupled to the second promoter. (Item 4) The method according to item 2, wherein the first and second subcode sequences are operably coupled to a bidirectional promoter. (Item 5) The method according to item 3, wherein the first and second splice donors, the first and second subcode sequences, and the first and second promoters are oriented in a head-to-head orientation. (Item 6) The method according to item 5, wherein the transgene further comprises one or more rare-cutting endonuclease first and second target sites, the target sites being adjacent to the first and second splice donors. (Item 7) The method according to item 5, wherein the introduced gene further comprises first and second homologous arms adjacent to the first and second splice donors. (Item 8) The method according to item 5, wherein the transgene is contained within an adeno-associated virus vector. (Item 9) The method according to item 7, wherein the transgene further comprises first and second target sites of one or more rare-cutting endonucleases, the target sites being adjacent to the first and second splice donors. (Item 10) The method according to item 9, wherein the first and second target sites are adjacent to the first and second homologous arms. (Item 11) The method according to item 1, wherein the introduced gene is incorporated into an intron or exon-intron junction of the endogenous gene. (Item 12) The method according to item 1, wherein the introduced gene is incorporated into an intron of the ATXN2 gene or the SNCA gene, or into an exon-intron junction. (Item 13) The method according to item 12, wherein the transgene comprises first and second partial coding sequences that encode a peptide produced by exon 1 of the non-pathogenic ATXN2 gene. (Item 14) The method according to item 12, wherein the introduced gene comprises first and second partial coding sequences encoding peptides produced by exon 2 of the non-pathogenic SNCA gene. (Item 15) The method according to item 1, wherein the nuclease is a CRISPR / Cas12a nuclease or a CRISPR / Cas9 nuclease. (Item 16) The method according to item 1, wherein the first and second partial coding sequences encode the same amino acids. (Item 17) The method according to item 1, wherein the first and second coding sequences are different in the nucleic acid sequence but encode the same amino acids. (Item 18) The method according to item 1, wherein the introduced gene is carried by a vector, and the form of the vector is selected from double-stranded linear DNA, double-stranded circular DNA, or a viral vector. (Item 19) The method according to item 18, wherein the viral vector is selected from an adenovirus vector, an adeno-associated virus vector, or a lentivirus vector. (Item 20) The method according to item 19, wherein the introduced gene is 4.7 kb or less. (Item 21) The method according to item 1, wherein the endogenous gene is the wild-type gene of the partial coding sequence. (Item 22) The method according to item 21, wherein the endogenous gene is abnormal or pathogenic, and the partial coding sequence encodes a partial protein produced from a functional version of the endogenous gene. (Item 23) The method according to item 22, wherein the first and second partial coding sequences are different in the nucleic acid sequence as compared to the corresponding endogenous gene. (Item 24) The method according to item 1, wherein the endogenous gene is selected from SOD1, TRPV4, CHRNA1, CHRND, CHRNE, CHRNB1, PRPS1, LRRK2, STIM1, FGFR3, MECP2, SNCA, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, TBP, HTT, AR, FXN, DMPK, PABPN1, ATXN8, RHO, or C9orf72. (Item 25) The method according to item 1, wherein the introduced gene further comprises a first and a second terminator. (Item 26) A method for integrating an introduced gene into an endogenous gene, a. Administering a transgene, wherein the transgene is i. Splice donor sequence, ii. Partial code array, iii. Promoter, iv. One RNA interference cassette, and v. Optionally administering a transgene including the first and second homologous arms or the left and right transposon terminals, b. Administering at least one rare-cutting endonuclease or transposase that targets a site within the endogenous gene, A method for incorporating the introduced gene into the endogenous gene. (Item 27) A method for integrating an introduced gene into an endogenous gene, a. Administering a transgene, wherein the transgene is i. The left and right transposon terminals, ii. First and second splice donor sequences, iii. The first and second subcode sequences, iv. One bidirectional promoter, or a first and second promoter, and v. Optionally administering a transgene containing the first and second terminators, b. Administering a transposase, A method for incorporating the introduced gene into the endogenous gene. (Item 28) A method for integrating an introduced gene into an endogenous gene, a. Administering a transgene, wherein the transgene is i. Splice acceptor array, ii. Partial code array, iii. Terminator, and iv. One RNA interference cassette, and v. Administering a transgene that includes, optionally, the first and second homologous arms, or the left and right transposon terminals. b. Administering at least one rare-cutting endonuclease or transposase that targets a site within the endogenous gene, A method for incorporating the introduced gene into the endogenous gene. [Brief explanation of the drawing]

[0016] [Figure 1] This diagram shows an exemplary transgene for targeted insertion into an endogenous gene and repair of its 5' end. TS1: Target site 1, SD1: Splice donor site 1, CDS1: Code sequence 1, P1: Promoter 1, TS2: Target site 2, SD2: Splice donor site 2, CDS2: Code sequence 2, P2: Promoter 2, HA1: Homologous arm 1, HA2: Homologous arm 2, T1: Terminator 1, T2: Terminator 2, AS1: Add-on sequence 1, AS2: Add-on sequence 2. [Figure 2] This figure illustrates the integration of a transgene into an intron of an exemplary gene. The transgene contains two target sites for one or more rare-cutting endonucleases, two splice donor sequences, two coding sequences (1.1 and 1.2), and two promoters. Integration proceeds via non-homologous end joining (NHEJ). ATG: start codon, TAA: stop codon. [Figure 3]This figure illustrates the integration of a transgene into an exemplary gene. The transgene contains two homologous arms, two target sites for one or more rare-cutting endonucleases, two splice donor sequences, two coding sequences (1.1 and 1.2), and two promoters. Integration proceeds via either homologous recombination (HR) or non-homologous end joining (NHEJ). [Figure 4] This figure illustrates the integration of a transgene into an exemplary gene. The transgene contains two homologous arms, two target sites for one or more rare-cutting endonucleases, two splice donor sequences, two coding sequences (1.1 and 1.2), and two promoters. Integration proceeds via either homologous recombination (HR) or non-homologous end joining (NHEJ). [Figure 5] This figure shows the gene products produced after integration of the transgene described herein. RNA hairpins and dsRNAs may be formed when the first and second partial coding sequences in the transgene are homologous to the coding sequence of the endogenous gene (top). RNA pairing may be reduced when the first and second partial coding sequences are codon-modified to reduce homology to the coding sequence of the endogenous gene (bottom). T1: Transcript 1, T2: Transcript 2, T3: Transcript 3, +1: RNA synthesis initiation site, S: sense, AntiS: antisense. [Figure 6] This diagram shows exons 1-3 of the ATXN2 gene. The transgenes pB1012-D1 and pBA1141, which are incorporated into the ATXN2 gene, are also shown. [Figure 7] This figure shows the results of incorporating the pB1012-D1 or pBA1141 transgene into the ATXN2 gene. [Figure 8] This diagram illustrates the integration of a transgene into the exon of an exemplary gene. The transgene contains two homologous arms, two target sites for one or more rare-cutting endonucleases, two splice donor sequences, two coding sequences (1.1 and 1.2), and two promoters. Integration proceeds via either homologous recombination (HR) or non-homologous end joining (NHEJ). [Figure 9] This is a diagram of a transgene containing a silencing sequence and a silencing resistance coding sequence. Two scenarios are shown. Scenario 1 is a diagram illustrating an approach to silencing both alleles of an endogenous gene while producing a WT protein substitution. Scenario 2 is a diagram illustrating an approach to silencing two alleles while producing a protein substitution, one allele having a gain-of-function mutation and the other allele having a WT sequence. The silencing sequence may be an RNAi cassette. A silencing-resistant CDS may have mutations within the silencing target sequence to block binding. Alternatively, the CDS may remove the sequence. [Figure 10] This figure shows the structure of an transgene for silencing the SOD1 allele in cells with a gain-of-function mutation in one allele. The transgene also includes a codon-modulating sequence for expressing the substituted SOD1 protein. [Figure 11] This figure shows examples of transgene structures for exemplary endogenous gene silencing and endogenous gene protein product replacement. [Figure 12] This figure shows a common approach to silencing a gain-of-function allele while replacing protein production. A partial coding sequence with a mutation to block silencing by the RNAi cassette is incorporated into the gene. If incorporated at the 5' or 3' end of the gene, the result may be: Result 1: Silencing of the endogenous gene; Result 2: Modification of one of the alleles within the endogenous gene; Result 3: Production of a new protein derived from the incorporation event, where the mRNA is silencing-resistant and the protein product contains the same or different sequence as the original gene. [Figure 13]This diagram shows transgenes used to silence the expression of endogenous genes and replace protein production. CDS1 and CDS2 may be partial coding sequences of endogenous genes. CDS may contain mutations in the corresponding target of the RNAi cassette, or the sequence may be excluded. The target for integration may be located within an intron, but after the endogenous splice donor sequence of the intron. Alternatively, the target for integration may be located at the intron-exon junction. [Figure 14] This diagram shows transgenes used to silence the expression of endogenous genes and replace protein production. CDS1 and CDS2 may be the complete coding sequences of endogenous genes. CDS may contain mutations in the corresponding target of the RNAi cassette, or the sequence may be excluded. The target for integration may be located within an intron, but after the endogenous splice donor sequence of the intron. Alternatively, the target for integration may be located at the intron-exon junction. [Figure 15] This diagram shows transgenes used to silence the expression of endogenous genes and replace protein production. CDS1 and CDS2 may be the complete coding sequences of endogenous genes. CDS may contain mutations in the corresponding target of the RNAi cassette or have sequences excluded. The target for integration may be located within an exon. [Figure 16] This diagram shows a transgene for silencing the expression of an endogenous gene and replacing protein production. CDS1 and CDS2 may be the complete coding sequences of the endogenous gene. CDS may contain mutations or have sequences excluded from the corresponding target of the RNAi cassette. The target for integration may be located within the 5'UTR. The target for integration may be an intron within the 5'UTR region, but requires a splice acceptor operably ligated to CDS. [Figure 17]This diagram shows a transgene used to silence the expression of an endogenous gene and replace protein production. CDS1 and CDS2 may be partial coding sequences of an endogenous gene. CDS may contain mutations in the corresponding target of the RNAi cassette or exclude sequences. The embedded target may be anywhere between the start and stop codons, but not within an endogenous splice acceptor or downstream of the last endogenous splice acceptor. [Figure 18] Images of the gel used to detect the integration of the transgenes described herein. 1: 1kb ladder, 2: 3'HR junction of pBA1141 with an expected size of 1594bp, 3: 3'HR junction of pBA1141 with an expected size of 1775bp, 4: 3'HR junction of pBA1141 with an expected size of 1775bp, 5: 3'NHEJ reverse of pBA1141 with an expected size of 2067bp, 6: 3'NHEJ of pBA1142 with an expected size of 813bp Forward junction, 7: 3'HR junction of pBA1143 with an expected size of 1225bp, 8: 3'HR junction of pBA1143 with an expected size of 1407bp, 9: 3'HR junction of pBA1143 with an expected size of 1225bp, 10: 3'HR junction of pBA1143 with an expected size of 1407bp, 11: 1kb ladder, 12: primer oNJB201+oNJB190 WT DNA control by: 13: WT DNA control with primers oNJB202+oNJB191: 14: WT DNA control with primers oNJB197+oNJB191: 15: WT DNA control with primers oNJB202+oNJB211: 16: 1kb ladder: 17: genomic DNA control of pBA1141+Cas9 transfection: 18: genomic DNA control of pBA1142 transfection: 19: genomic DNA control of pBA1143+Cas9 transfection: 20: genomic DNA control of pBA1141+Cas12a transfection: 21: genomic DNA control of pBA1142+Cas12a transfection: 22: genomic DNA control of pBA1143+Cas12a transfection: 23: WT control: 24: control without DNA. [Modes for carrying out the invention]

[0017] Methods and compositions for modifying the coding sequences of endogenous genes are disclosed herein. In some embodiments, the method comprises inserting a transgene into an endogenous gene, the transgene providing a partial coding sequence that replaces the coding sequence of the endogenous gene. Methods and compositions for expressing a substitution protein while reducing the expression of the endogenous gene are also disclosed herein.

[0018] In one embodiment, this document describes a method for incorporating a transgene into an endogenous gene and modifying the mRNA or protein product. The method comprises administering a transgene comprising first and second splice donor sequences, first and second partial coding sequences, one bidirectional promoter, or first and second promoters, and optionally, first and second terminators, wherein the transgene is administered together with at least one rare-cutting endonuclease targeting a site within the endogenous gene, and the transgene is incorporated into the endogenous gene. The endogenous gene may be located in a eukaryotic cell, including a human cell. The transgene may have a first splice donor operably ligated to a first partial coding sequence, and a second splice donor may be operably ligated to a second partial coding sequence. The first partial coding sequence may also be operably ligated to a first promoter, and the second partial coding sequence may also be operably ligated to a second promoter. Alternatively, the first and second partial coding sequences may be operably linked to a bidirectional promoter. The transgene having the first and second splice donors, the first and second partial coding sequences, and the first and second promoters may be oriented in a head-to-head orientation. These transgenes may be encapsulated within an adeno-associated virus vector and incorporated into the endogenous gene via NHEJ-mediated incorporation to target double-strand breaks. The transgene may further include first and second target sites of one or more rare-cutting endonucleases, the target sites being adjacent to the first and second splice donors. Alternatively, the transgene may further include left and right homologous arms adjacent to the first and second splice donors. The transgene may have both first and second target sites of one or more rare-cutting endonucleases, the target sites being adjacent to the first and second splice donors. The first and second target sites can be adjacent to the first and second homologous arms. The transgene described in this method can be incorporated into an intron or exon-intron junction of an endogenous gene.The endogenous gene may be ATXN2 or SNCA, and the site for integration may be within an intron of the ATXN2 or SNCA gene, or at an exon-intron junction. When integrated into ATXN2, the transgene may contain first and second partial coding sequences encoding a peptide produced by exon 1 of the non-pathogenic ATXN2 gene. When integrated into SNCA, the transgene may contain first and second partial coding sequences encoding a peptide produced by exon 2 of the non-pathogenic SNCA gene. Integration may occur via the use of a CRISPR / Cas12a nuclease or a CRISPR / Cas9 nuclease. The first and second partial coding sequences may encode the same amino acid. The first and second coding sequences may differ in nucleic acid sequence (e.g., through codon degeneracy), but still encode the same amino acid. The transgene described in this method may be contained in a vector, the form of which can be selected from double-stranded linear DNA, double-stranded circular DNA, or a viral vector. The transgene may be contained in a viral vector selected from adenovirus vectors, adeno-associated virus vectors, or lentiviral vectors. The transgene may have a full length of 4.7 kb or less. The method may include using a transgene having a partial coding sequence that encodes a peptide produced by a target endogenous gene. The partial coding sequence may be a WT version of the target endogenous gene, and the target endogenous gene may be a mutated gene or a gene containing a pathogenic mutation. In one embodiment, the host gene is a gene in which protein expression is abnormal; in other words, the protein is expressed at a level lower or higher than a functional protein that is not expressed, or the protein or a part thereof is expressed in a non-functional manner that causes harm in the host. The transgene used in this method may have first and second partial coding sequences whose nucleic acid sequences differ from those of the corresponding endogenous gene. In other words, the partial coding sequence may be modified (via codon degeneracy) to have minimal homology to the endogenous gene.This method can be used to modify genes involved in gain-of-function disorders, including SOD1, TRPV4, CHRNA1, CHRND, CHRNE, CHRNB1, PRPS1, LRRK2, STIM1, FGFR3, MECP2, SNCA, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, TBP, HTT, AR, FXN, DMPK, PABPN1, ATXN8, RHO, or C9orf72.

[0019] In another embodiment, this document describes a method for incorporating a transgene into an endogenous gene and modifying its mRNA or protein product. The method comprises administering a transgene, which comprises left and right transposon ends, first and second splice donor sequences, first and second partial coding sequences, one bidirectional promoter or first and second promoters, and optionally, first and second terminators, wherein the transgene is administered together with at least one transposase targeting a site within the endogenous gene, and the transgene is incorporated into the endogenous gene. The endogenous gene may be located in a eukaryotic cell, including a human cell. The transgene may have a first splice donor operably ligated to a first partial coding sequence, and a second splice donor operably ligated to a second partial coding sequence. The first partial coding sequence may also be operably ligated to a first promoter, and the second partial coding sequence may also be operably ligated to a second promoter. Alternatively, the first and second partial coding sequences may be operably linked to a bidirectional promoter. The transgene having the first and second splice donors, the first and second partial coding sequences, and the first and second promoters may be oriented in a head-to-head orientation. The transgene may further include the left and right transposon ends adjacent to the first and second splice donors. The transposase may be a CRISPR transposase, which contains Cas12k or Cas6 protein. These transgenes may be contained within an adeno-associated virus vector. The transgenes described in this method may be incorporated into an intron or exon-intron junction of an endogenous gene. The endogenous gene may be ATXN2 or SNCA, and the site for integration may be an intron or exon-intron junction of the ATXN2 or SNCA gene. When incorporated into ATXN2, the transgene may contain first and second partial coding sequences that encode peptides produced by exon 1 of the non-pathogenic ATXN2 gene.When incorporated into an SNCA, the transgene may contain first and second partial coding sequences encoding a peptide produced by exon 2 of the non-pathogenic SNCA gene. The first and second partial coding sequences may encode the same amino acid. The first and second coding sequences may differ in their nucleic acid sequences (e.g., through codon degeneracy) but still encode the same amino acid. The transgene described in this method may be contained in a vector, the form of which can be selected from double-stranded linear DNA, double-stranded circular DNA, or viral vectors. The transgene may be contained in a viral vector selected from adenovirus vectors, adeno-associated virus vectors, or lentiviral vectors. The transgene may have a total length of 4.7 kb or less. This method may include using a transgene having a partial coding sequence encoding a peptide produced by a target endogenous gene. The partial coding sequence may be a WT version of the target endogenous gene, which may be a mutated gene or a gene containing a pathogenic mutation. The transgene used in this method may have first and second subcoding sequences that differ in nucleic acid sequence from the corresponding endogenous gene. In other words, the subcoding sequence can be modified (via codon degeneracy) to have minimal homology to the endogenous gene. Using this method, genes involved in gain-of-function disorders, including SOD1, TRPV4, CHRNA1, CHRND, CHRNE, CHRNB1, PRPS1, LRRK2, STIM1, FGFR3, MECP2, SNCA, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, TBP, HTT, AR, FXN, DMPK, PABPN1, ATXN8, RHO, or C9orf72, can be modified.

[0020] This document also describes a method for incorporating a transgene into an endogenous gene and modifying the mRNA or protein product. The method comprises administering a transgene comprising a splice acceptor sequence, a partial coding sequence, a terminator, and a single RNA interference cassette, wherein the transgene is administered together with at least one rare-cutting endonuclease or transposase targeting a site within the endogenous gene, and the transgene is incorporated into the endogenous gene. The partial coding sequence may contain mutations that prevent silencing by the RNAi cassette. The endogenous gene may be located in eukaryotic cells, including human cells. The transgene may have a splice acceptor operably linked to the partial coding sequence. The partial coding sequence may also be operably linked to a terminator. These transgenes are encapsulated within an adeno-associated virus vector and can be incorporated into the endogenous gene via NHEJ-mediated target double-strand breaks or homologous recombination. The transgene may further include left and right homologous arms. The transgenes described in this method can be incorporated within an intron or intron-exon junction of the endogenous gene. The RNAi cassette may be a promoter operably ligated to a sequence homologous to the endogenous gene. The RNAi cassette can generate shRNA or siRNA. The RNAi cassette may contain sequences homologous to the endogenous gene, and the partial coding sequences within the transgene may contain the same sequences as the endogenous gene, but the target site of the RNAi cassette may be mutated to prevent silencing of expression by the incorporated transgene (e.g., having synonymous single nucleotide polymorphisms, insertions, or deletions). Incorporation may occur via the use of a CRISPR / Cas12a nuclease or a CRISPR / Cas9 nuclease, or via the use of a CRISPR-related transposase.When a CRISPR-related transposase is used, the transgene may contain the left and right transposon ends instead of homologous arms. The CRISPR-related transpose may contain the Cas6 protein or the Cas12k protein. The transgene described in this method may be contained in a vector, the form of which can be selected from double-stranded linear DNA, double-stranded circular DNA, or viral vectors. The transgene may be contained in a viral vector selected from adenovirus vectors, adeno-associated virus vectors, or lentiviral vectors. The transgene may have a full length of 4.7 kb or less. This method may include using a transgene having a partial coding sequence that encodes a peptide produced by a target endogenous gene. The partial coding sequence may be the WT version of the target endogenous gene, which may be a mutated gene or a gene containing a pathogenic mutation. This method can be used to modify genes involved in gain-of-function disorders, including CACNA1A, ATXN3, SOD1, TRPV4, CHRNA1, CHRND, CHRNE, CHRNB1, PRPS1, LRRK2, STIM1, FGFR3, MECP2, SNCA, ATXN1, ATXN2, CACNA1A, ATXN7, TBP, HTT, AR, FXN, DMPK, PABPN1, ATXN8, RHO, or C9orf72.

[0021] This document also describes a method for incorporating a transgene into an endogenous gene and modifying its mRNA or protein product. The method comprises administering a transgene comprising a splice acceptor sequence, first and second partial coding sequences, a terminator, and one RNA interference cassette, wherein the transgene is administered together with at least one rare-cutting endonuclease or transposase targeting a site within the endogenous gene, and the transgene is incorporated into the endogenous gene. The first and second partial coding sequences may contain mutations that prevent silencing by the RNAi cassette. The endogenous gene may be located in a eukaryotic cell, including a human cell. The transgene may have a first splice acceptor operably ligated to the first partial coding sequence and a second splice acceptor operably ligated to the second partial coding sequence. Furthermore, the first partial coding sequence may be operably ligated to the first terminator, and the second partial coding sequence may be operably ligated to the second terminator. The partial coding sequences may be tail-to-tail orientation using an RNAi cassette between the two terminators. These transgenes are encapsulated within an adeno-associated virus vector and may be incorporated into the endogenous gene via NHEJ-mediated target double-strand breaks or homologous recombination. The transgene may further include left and right homologous arms. The transgenes described in this method may be incorporated within an intron or intron-exon junction of the endogenous gene. The RNAi cassette may be a promoter operably ligated to a sequence homologous to the endogenous gene. The RNAi cassette may generate shRNA or siRNA. RNAi cassettes can contain sequences homologous to endogenous genes, and partial coding sequences within the transgene can contain the same sequences as the endogenous gene, but the target sites of the RNAi cassette can be mutated to block silencing. Incorporation can occur via the use of CRISPR / Cas12a nuclease or CRISPR / Cas9 nuclease, or via use with CRISPR-related transposases.When a CRISPR-related transposase is used, the transgene may contain the left and right transposon ends instead of homologous arms. The CRISPR-related transpose may contain the Cas6 protein or the Cas12k protein. The transgene described in this method may be contained in a vector, the form of which can be selected from double-stranded linear DNA, double-stranded circular DNA, or viral vectors. The transgene may be contained in a viral vector selected from adenovirus vectors, adeno-associated virus vectors, or lentiviral vectors. The transgene may have a full length of 4.7 kb or less. This method may include using a transgene having a partial coding sequence that encodes a peptide produced by a target endogenous gene. The partial coding sequence may be the WT version of the target endogenous gene, which may be a mutated gene or a gene containing a pathogenic mutation. This method can be used to modify genes involved in gain-of-function disorders, including CACNA1A, ATXN3, SOD1, TRPV4, CHRNA1, CHRND, CHRNE, CHRNB1, PRPS1, LRRK2, STIM1, FGFR3, MECP2, SNCA, ATXN1, ATXN2, CACNA1A, ATXN7, TBP, HTT, AR, FXN, DMPK, PABPN1, ATXN8, RHO, or C9orf72.

[0022] This document also describes a method for incorporating a transgene into an endogenous gene and modifying its mRNA or protein product. The method involves administering a transgene comprising a splice donor sequence, a partial coding sequence, a promoter, and an RNA interference cassette, wherein the transgene is administered with at least one rare-cutting endonuclease or transposase targeting a site within the endogenous gene, and the transgene is incorporated into the endogenous gene. The partial coding sequence may contain mutations that prevent silencing by the RNAi cassette. For example, if the RNAi cassette is designed to target a sequence in a transcript produced by the endogenous gene, the partial coding sequence (found in the transgene) may contain the same coding sequence as the endogenous gene and the corresponding RNAi target, thereby subjecting the modified endogenous gene to the same interference by the RNAi cassette. The partial coding sequence within the transgene can be mutated to minimize or prevent silencing of the modified endogenous gene. The endogenous gene may be located in eukaryotic cells, including human cells. The transgene may have a splice donor operably ligated to a subcoding sequence. The subcoding sequence may also be operably ligated to a promoter. These transgenes are encapsulated within an adeno-associated virus vector and can be incorporated into the endogenous gene via NHEJ-mediated target double-strand breaks or homologous recombination. The transgene may further include left and right homologous arms. The transgenes described in this method can be incorporated within an intron or exon-intron junction of the endogenous gene. The RNAi cassette may be a promoter operably ligated to a sequence homologous to the endogenous gene. The RNAi cassette can generate shRNA or siRNA. The RNAi cassette may contain sequences homologous to the endogenous gene, and the subcoding sequence within the transgene may contain the same sequence as the endogenous gene, but the target site of the RNAi cassette may be mutated to prevent silencing.The endogenous gene may be ATXN2 or SNCA, and the site for integration may be within an intron of the ATXN2 or SNCA gene, or at an exon-intron junction. When integrated into ATXN2, the transgene may include a subcoding sequence encoding a peptide produced by exon 1 of the non-pathogenic ATXN2 gene. The RNAi cassette can be designed to target the transcription sequence from exon 1 of the ATXN2 gene, and the corresponding sequence within the subcoding sequence can be mutated to block silencing. When integrated into SNCA, the transgene may include a subcoding sequence encoding a peptide produced by exon 2 of the non-pathogenic SNCA gene. The RNAi cassette can be designed to target the transcription sequence from exon 2 of the SNCA gene, and the corresponding sequence within the subcoding sequence can be mutated to block silencing. Integration may occur via the use of a CRISPR / Cas12a nuclease or a CRISPR / Cas9 nuclease, or via the use of a CRISPR-related transposase. When a CRISPR-related transposase is used, the transgene may contain the left and right transposon ends instead of homologous arms. The CRISPR-related transpose may contain the Cas6 protein or the Cas12k protein. The transgene described in this method may be contained in a vector, the form of which can be selected from double-stranded linear DNA, double-stranded circular DNA, or viral vectors. The transgene may be contained in a viral vector selected from adenovirus vectors, adeno-associated virus vectors, or lentiviral vectors. The transgene may have a full length of 4.7 kb or less. This method may include using a transgene having a partial coding sequence that encodes a peptide produced by a target endogenous gene. The partial coding sequence may be the WT version of the target endogenous gene, which may be a mutated gene or a gene containing a pathogenic mutation.This method can be used to modify genes involved in gain-of-function disorders, including CACNA1A, ATXN3, SOD1, TRPV4, CHRNA1, CHRND, CHRNE, CHRNB1, PRPS1, LRRK2, STIM1, FGFR3, MECP2, SNCA, ATXN1, ATXN2, CACNA1A, ATXN7, TBP, HTT, AR, FXN, DMPK, PABPN1, ATXN8, RHO, or C9orf72.

[0023] This document also describes a method for incorporating a transgene into an endogenous gene and modifying its mRNA or protein product. The method comprises administering a transgene comprising first and second splice donor sequences, first and second partial coding sequences, first and second promoters (or bidirectional promoters), and an RNA interference cassette, wherein the transgene is administered together with at least one rare-cutting endonuclease or transposase targeting a site within the endogenous gene, and the transgene is incorporated into the endogenous gene. The partial coding sequences may contain mutations that prevent silencing by the RNAi cassette. The endogenous gene may be located in a eukaryotic cell, including a human cell. The transgene may have a first splice donor operably ligated to the first partial coding sequence and a second splice donor operably ligated to the second partial coding sequence. Furthermore, a first partial coding sequence may be operably ligated to a first promoter, and a second partial coding sequence may be operably ligated to a second promoter. The partial coding sequences may be head-to-head oriented, and the RNAi cassette may be positioned between the first and second promoters. These transgenes are encapsulated within an adeno-associated virus vector and may be incorporated into the endogenous gene via NHEJ-mediated target double-strand breaks or homologous recombination. The transgene may further include left and right homologous arms. The transgenes described in this method may be incorporated within an intron or exon-intron junction of the endogenous gene. The RNAi cassette may be a promoter operably ligated to a sequence homologous to the endogenous gene. The RNAi cassette may generate shRNA or siRNA. The RNAi cassette may contain sequences homologous to the endogenous gene, and the partial coding sequences within the transgene may contain the same sequences as the endogenous gene, but the target site of the RNAi cassette may be mutated to prevent silencing. The endogenous gene may be ATXN2 or SNCA, and the site for integration may be within an intron of the ATXN2 or SNCA gene, or at an exon-intron junction.When incorporated into ATXN2, the transgene may include a subcoding sequence encoding a peptide produced by exon 1 of the non-pathogenic ATXN2 gene. The RNAi cassette can be designed to target the transcription sequence from exon 1 of the ATXN2 gene, and the corresponding sequence within the subcoding sequence can be mutated to block silencing. When incorporated into SNCA, the transgene may include a subcoding sequence encoding a peptide produced by exon 2 of the non-pathogenic SNCA gene. The RNAi cassette can be designed to target the transcription sequence from exon 2 of the SNCA gene, and the corresponding sequence within the subcoding sequence can be mutated to block silencing. Incorporation may occur via the use of a CRISPR / Cas12a nuclease or a CRISPR / Cas9 nuclease, or via the use of a CRISPR-related transposase. When a CRISPR-related transposase is used, the transgene may include the left and right transposon ends instead of homologous arms. CRISPR-related transposes may include the Cas6 protein or the Cas12k protein. The transgenes described in this method may be contained in a vector, the form of which can be selected from double-stranded linear DNA, double-stranded circular DNA, or viral vectors. The transgene may be contained in a viral vector selected from adenovirus vectors, adeno-associated virus vectors, or lentiviral vectors. The transgene may have a full length of 4.7 kb or less. This method may involve using a transgene having a partial coding sequence that encodes a peptide produced by a target endogenous gene. The partial coding sequence may be a WT version of the target endogenous gene, which may be a mutated gene or a gene containing a pathogenic mutation. The transgenes used in this method may have first and second partial coding sequences whose nucleic acid sequences differ from those of the corresponding endogenous gene. In other words, the partial coding sequence may be modified (via codon degeneracy) to have minimal homology to the endogenous gene.This method can be used to modify genes involved in gain-of-function disorders, including CACNA1A, ATXN3, SOD1, TRPV4, CHRNA1, CHRND, CHRNE, CHRNB1, PRPS1, LRRK2, STIM1, FGFR3, MECP2, SNCA, ATXN1, ATXN2, CACNA1A, ATXN7, TBP, HTT, AR, FXN, DMPK, PABPN1, ATXN8, RHO, or C9orf72.

[0024] The implementation of this method, in addition to the preparation and use of the compositions disclosed herein, will utilize, unless otherwise indicated, conventional techniques in molecular biology, biochemistry, chromatin structure and analysis, computational chemistry, cell culture, recombinant DNA, and related fields within the art. These techniques are fully described in the literature, for example, Sambrook et al., MOLECULAR CLONING: A LABORATORY MANUAL, Second edition, Cold Spring Harbor Laboratory Press, 1989 and 3rd edition 2001; Ausubel et al., CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley & Sons, New York, 1987 and regularly updated series, METHODS IN ENZYMOLOGY, Academic Press, San See Diego; Wolffe, CHROMATIN STRUCTURE AND FUNCTION, Third edition, Academic Press, San Diego, 1998; METHODS IN ENZYMOLOGY, Vol. 304, “Chromatin” (PM Wassarman and AP Wolffe, eds.), Academic Press, San Diego, 1999; and METHODS IN MOLECULAR BIOLOGY, Vol. 119, “Chromatin Protocols” (PB Becker, ed.), Humana Press, Totowa, 1999.

[0025] As used herein, the terms “nucleic acid” and “polynucleotide” may be used interchangeably. Nucleic acid and polynucleotide may refer to deoxyribonucleotides or ribonucleotide polymers, either in linear or cyclic three-dimensional structures, and in either single-stranded or double-stranded forms. These terms should not be interpreted as restrictive with respect to polymer length. The terms may encompass known analogues of natural nucleotides, as well as nucleotides modified with bases, sugars, and / or phosphate moieties.

[0026] The terms “polypeptide,” “peptide,” and “protein” can be used interchangeably to refer to covalently linked amino acid residues. These terms also apply to proteins in which one or more amino acids are chemical analogues or modified derivatives of corresponding naturally occurring amino acids.

[0027] The terms “operatively linked” and “operably linked” are used interchangeably and refer to the proximal proximity of two or more components (such as sequence elements) to the coding sequence, where both components function correctly and are positioned to allow for the possibility that at least one component may mediate a function that acts on at least one of the other components. For example, a transcriptional regulatory sequence (e.g., a promoter) is operably linked to a coding sequence if the transcriptional regulatory sequence controls the level of transcription of the coding sequence in response to the presence or absence of one or more transcriptional regulators. Transcriptional regulatory sequences are generally cis-operably linked to coding sequences, but do not need to be directly adjacent to the coding sequence. For example, enhancers are transcriptional regulatory sequences that are operably linked to a coding sequence even though they are not contiguous.

[0028] As used herein, the term “cleavage” refers to the cleavage of the covalent backbone of a nucleic acid molecule. Cleavage can be initiated by a variety of methods, including, but not limited to, enzymatic or chemical hydrolysis of phosphodiester bonds. Cleavage can refer to both single-stranded nicks and double-stranded breaks. Double-stranded breaks can result from two different single-stranded nicks. Nucleic acid cleavage can result in the production of either blunt or adherent ends. In certain embodiments, rare-cutting endonucleases are used for cleaving target double-stranded or single-stranded DNA.

[0029] "Exogenous" molecules can refer to small molecules (e.g., sugars, lipids, amino acids, fatty acids, phenolic compounds, alkaloids), or macromolecules (e.g., proteins, nucleic acids, carbohydrates, lipids, glycoproteins, lipoproteins, polysaccharides), or any modified derivatives of the above molecules, or any complex containing one or more of the above molecules that are produced or present outside the cell or are not normally present inside the cell. Exogenous molecules can be introduced into cells. Methods for introducing exogenous molecules into cells include lipid-mediated transport, electroporation, direct injection, cell fusion, particle impaction, calcium phosphate coprecipitation, DEAE-dextran-mediated transport, and viral vector-mediated transport.

[0030] "Endogenous" molecules are small or large molecules present in specific cells at specific developmental stages under specific environmental conditions. Endogenous molecules can include nucleic acids, chromosomes, mitochondria, chloroplasts, or other organelle genomes, or naturally occurring episomal nucleic acids. Additional endogenous molecules may include proteins, such as transcription factors and enzymes.

[0031] As used herein, “gene” refers to the DNA region that encodes a gene product, including all DNA regions that regulate the production of the gene product. Thus, a gene includes, but is not limited to, the promoter sequence, terminators, translation regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix binding sites, and locus regulatory regions.

[0032] "Endogenous genes" refer to DNA regions that are normally present within a specific cell and code for a gene product, as well as all DNA regions that regulate the production of gene products.

[0033] "Gene expression" refers to the process of converting the information contained in a gene into a gene product. A gene product can be the direct transcript of a gene. For example, a gene product can be, but is not limited to, mRNA, tRNA, rRNA, antisense RNA, ribozymes, structural RNA, or proteins produced by the translation of mRNA. Gene products also include RNA modified by processes such as cap formation, polyadenylation, methylation, and editing, as well as proteins modified by processes such as methylation, acetylation, phosphorylation, ubiquitination, ADP-ribosylation, myristylation, and glycosylation.

[0034] "Encoding" refers to the process of converting information contained in nucleic acids into products, which can arise from direct transcripts of nucleic acid sequences. For example, products may include, but are not limited to, mRNA, tRNA, rRNA, antisense RNA, ribozymes, structural RNA, or proteins produced by the translation of mRNA. Gene products also include RNA modified by processes such as cap formation, polyadenylation, methylation, and editing, as well as proteins modified by processes such as methylation, acetylation, phosphorylation, ubiquitination, ADP-ribosylation, myristylation, and glycosylation.

[0035] The "target site" or "target sequence" is a nucleic acid sequence to which a binding molecule binds, provided that sufficient conditions for binding are present, for example, if an endonuclease or transposase, including a rare-cutting endonuclease or CRISPR-related transposase, is present. The target site may be an endogenous gene and may be native or heterologous to the cell.

[0036] As used herein, the term “recombination” refers to the process of exchanging genetic information between two polynucleotides. The term “homologous recombination (HR)” refers to a specific form of recombination that may occur, for example, during the repair of a double-strand break. Homologous recombination requires nucleotide sequence homology present on a “donor” molecule. The donor molecule can be used by a cell as a template for the repair of a double-strand break. Information within the donor molecule that differs from the genomic sequence at or near the double-strand break site can be stably incorporated into the cell’s genomic DNA.

[0037] As used herein, the term “homologous” means a sequence of nucleic acid or amino acid that has similarity to a second sequence of nucleic acid or amino acid. In some embodiments, homologous sequences may have at least 80% sequence identity with one another (e.g., 81%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity).

[0038] The "target site" or "target sequence" defines the portion of nucleic acid to which a rare-cutting endonuclease or CRISPR-related transposase binds, provided sufficient conditions for binding are present.

[0039] As used herein, the term “transgene” refers to a sequence of nucleic acid that can be transported into an organism or cell. A transgene may include a gene or nucleic acid sequence that is not normally present in the target organism or cell. Furthermore, a transgene may include a copy of a gene or nucleic acid sequence that is normally present in the target organism or cell. A transgene may be an exogenous DNA sequence introduced into the cytoplasm or nucleus of a target cell. In one embodiment, the transgene described herein includes a partial coding sequence that codes for a portion of a protein produced by a gene in the host cell.

[0040] As used herein, the term “pathogenic” refers to all things that can cause disease. A pathogenic mutation may refer to a modification in a gene that causes disease. A pathogenic gene refers to a gene that contains a modification that causes disease. For example, the pathogenic ATXN2 gene in a patient with spinocerebellar ataxia 2 refers to the ATXN2 gene with an extended CAG trinucleotide repeat, and the extended CAG trinucleotide repeat causes the disease.

[0041] As used herein, the term “tail-to-tail” refers to the orientation of two units in opposite and reverse directions. The two units may be two sequences on a single nucleic acid molecule, with the 3' ends of each sequence adjacent to each other. For example, a first nucleic acid having element [splice acceptor 1]-[partial coding sequence 1]-[terminator 1] and a second nucleic acid having element [splice acceptor 2]-[partial coding sequence 2]-[terminator 2] can be arranged in a tail-to-tail orientation from 5' to 3', resulting in [splice acceptor 1]-[partial coding sequence 1]-[terminator 1]-[terminator 2 RC]-[partial coding sequence 2 RC]-[splice acceptor 2 RC] (where RC stands for reverse complement).

[0042] As used herein, the term “head-to-head” refers to the orientation of two units in opposite and reverse directions. The two units may be two sequences on a single nucleic acid molecule, with the 5' ends of each sequence adjacent to each other. For example, a first nucleic acid having an element: [promoter 1]-[partial coding sequence 1]-[splice donor 1] and a second nucleic acid having an element: [promoter 2]-[partial coding sequence 2]-[splice donor 2] can be arranged in a head-to-head orientation from 5' to 3', resulting in [splice donor 1 RC]-[partial coding sequence 1 RC]-[promoter 1 RC][promoter 2]-[partial coding sequence 2]-[splice donor 2] (where RC stands for reverse complement).

[0043] As used herein, the term “integrating” refers to the process of adding DNA to a target region of DNA. As described herein, integration can be facilitated by several different means, including non-homologous end joining, homologous recombination, or targeted rearrangement. For example, integration of a user-provided DNA molecule into a target gene can be facilitated by non-homologous end joining, where a targeted double-strand break occurs within the target gene and the user-provided DNA molecule is administered. The user-provided DNA molecule may contain exposed DNA ends to facilitate capture during the repair of the target gene by non-homologous end joining. The exposed ends may be present on the DNA molecule at the time of administration (i.e., when a linear DNA molecule is administered) or may be created at the time of administration to a cell (i.e., a rare-cutting endonuclease cleaves the user-provided DNA molecule in the cell to expose the ends). Furthermore, the user-provided DNA molecule may be contained in a viral vector, including an adeno-associated virus vector. In another example, integration occurs via homologous recombination, where the user-provided DNA may contain left and right homologous arms. In another example, integration occurs via transposition, where the user-provided DNA holds both ends of a transposon.

[0044] The term “intron-exon junction” refers to a specific location within a gene. This specific location lies between the last nucleotide of an intron and the first nucleotide of the subsequent exon. When incorporating a transgene as described herein, the transgene may be incorporated within an “intron-exon junction.” If the transgene contains cargo, the cargo is incorporated immediately after the last nucleotide of the intron. In some cases, incorporation of a transgene within an intron-exon junction may result in the removal of a sequence within an exon (e.g., incorporation via HR, and replacement of the sequence within the exon with cargo in the transgene).

[0045] The term “exon-intron junction” refers to a specific location within a gene. This specific location lies between the last nucleotide of an exon and the first nucleotide of the subsequent intron. When incorporating a transgene as described herein, the transgene may be incorporated within an “exon-intron junction.” If the transgene contains a cargo, the cargo is incorporated immediately before the first nucleotide of the intron. In some cases, incorporating a transgene within an exon-intron junction may result in the removal of a sequence within the exon (e.g., incorporation via HR, and replacement of the sequence within the exon with a cargo in the transgene).

[0046] As used herein, the term “partial coding sequence” refers to a sequence of nucleic acids that encodes a partial protein. A partial coding sequence may encode a protein containing one or fewer amino acids compared to a wild-type protein or a functional protein. A partial coding sequence may encode a partial protein homologous to a wild-type protein or a functional protein. When referring to a “partial coding sequence” operably ligated to a promoter, the term “partial coding sequence” refers to a sequence of nucleotides encoding the N-terminus of the protein of interest. For example, a partial coding sequence of the ATXN2 gene, which contains 25 exons, may include nucleotides encoding peptides produced by exons 1, 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12, 1-13, 1-14, 1-15, 1-16, 1-17, 1-18, 1-19, 1-20, 1-21, 1-22, 1-23, or 1-24. When referring to a “partial coding sequence” that is operably ligated to a terminator, the term “partial coding sequence” refers to a sequence of nucleotides that code for the C-terminus of the protein of interest. For example, a partial coding sequence of the ATXN2 gene may include nucleotides that code for peptides produced by exons 2-25, 3-25, 4-25, 5-25, 6-25, 7-25, 8-25, 9-25, 10-25, 11-25, 12-25, 13-25, 14-25, 15-25, 16-25, 17-25, 18-25, 19-25, 20-25, 21-25, 22-25, 23-25, 24-25, or 25.

[0047] The terms "silencing-resistant coding sequence" or "silencing-resistant subcoding sequence" refer to nucleic acid sequences that, when RNA is produced using that sequence as a template, cannot be silenced, or are unlikely to be silenced, by the corresponding RNAi molecule. This may be due to mutations within the RNAi target site or the absence of such a site.

[0048] The methods and compositions described herein may use transgenes having cargo sequences. The term "cargo" may refer to elements such as the complete or partial coding sequence of a gene, a partial sequence of a gene possessing single-nucleotide polymorphisms compared to the WT or modified target, a splice acceptor, a splice donor, a promoter, a terminator, a transcriptional regulatory element, an RNAi cassette, a purification tag (e.g., glutathione-S-transferase, poly(His), maltose-binding protein, Strep tag, Myc tag, AviTag, HA tag, or chitin-binding protein), or a reporter gene (e.g., GFP, RFP, lacZ, cat, luciferase, puro, neomycin). Where defined herein, "cargo" may refer to a sequence within a transgene that is incorporated into a target site. For example, "cargo" may refer to a sequence on a transgene between two homologous arms, between the target sites of two rare-cutting endonucleases, or between the left and right transposon ends.

[0049] The term “homology sequence” refers to a sequence of nucleic acid that contains homology to a second nucleic acid. A homology sequence may exist on a donor molecule, for example, as an “arm of homology” or “homology arm.” A homology arm may be a sequence of nucleic acid within the donor molecule that promotes homologous recombination with the second nucleic acid. In one embodiment, the homology sequence or homology arm has homology to an endogenous gene. Whereof the foregoing is defined, a homology arm may also be referred to as an “arm.” In a donor molecule having two homology arms, the homology arms may be referred to as “arm 1” and “arm 2.” In one embodiment, a cargo sequence may be adjacent to the first and second homology arms.

[0050] The term "bidirectional terminator" refers to a terminator that can terminate RNA polymerase transcription in either the sense or antisense direction. In contrast to two unidirectional terminators that are tail-to-tail oriented, bidirectional terminators may contain non-chimeric sequences of DNA. Examples of bidirectional terminators include ARO4, TRP1, TRP4, ADH1, CYC1, GAL1, GAL7, and GAL10 terminators.

[0051] The term “bidirectional promoter” refers to a promoter that can initiate RNA polymerase transcription in either the sense direction or the antisense direction. In contrast to two unidirectional promoters with head-to-head orientation, bidirectional promoters may contain non-chimeric sequences of DNA. An example of a bidirectional promoter is the promoter described in Trinklein et al., Genome Res. 14:62-66, 2004 (the entire disclosure is incorporated herein by reference, except for any definitions, disclaimers, denials, and inconsistencies).

[0052] The 5' or 3' end of a nucleic acid molecule refers to the directional and chemical orientation of the nucleic acid. Where defined herein, the “5' end of a gene” may include an exon with a start codon but not an exon with a stop codon. Where defined herein, the “3' end of a gene” may include an exon with a stop codon but not an exon with a start codon.

[0053] The term "RNAi" refers to RNA interference, a process that uses RNA molecules to inhibit or reduce gene expression or translation. RNAi can be induced using small interfering RNA (siRNA) or small hairpin RNA (shRNA).

[0054] The term "ATXN2" gene refers to the gene that codes for the enzyme attaxin-2. A representative sequence of the ATXN2 gene can be found in the NCBI reference sequence: NG_011572.3 and the corresponding sequence number 56. The exon and intron boundaries can be defined by the sequences provided in sequence number 56. Specifically, exon 1 contains sequences 282-532. Exon 2 contains sequences 43397-43433. Exon 3 contains sequences 45099-45158. Exon 4 contains sequences 46339-46410. Exon 5 contains sequences 46886-47036. Exon 6 contains sequences 74000-74124. Exon 7 contains sequences 78343-78434. Exon 8 contains sequences 79240-79437. Exon 9 contains the sequence 80889~81067. Exon 10 contains the sequence 82953~83162. Exon 11 contains the sequence 85777~85959. Exon 12 contains the sequence 88734~88931. Exon 13 contains the sequence 89318~89425. Exon 14 contains the sequence 89697~89767. Exon 15 contains the sequence 110536~110840. Exon 16 contains the sequence 112492~112555. Exon 17 contains the sequence 113451~113603. Exon 18 contains the sequence 113985~114051. Exon 19 contains the sequence 128574-128758. Exon 20 contains the sequence 129076-129208. Exon 21 contains the sequence 134601-134654. Exon 22 contains the sequence 141957-142102. Exon 23 contains the sequence 143060-143287. Exon 24 contains the sequence 145471-145639. Exon 25 contains the sequence 146476-146504. Intron 1 contains the sequence 533-43396. Intron 2 contains the sequence 43434-45098. Intron 3 contains the sequence 45159-46338. Intron 4 contains the sequence 46411-46885. Intron 5 contains the sequence 47037-73999. Intron 6 contains the sequence 74125-78342. Intron 7 contains the sequence 78435-79239.Intron 8 contains the sequence 79438~80888. Intron 9 contains the sequence 81068~82952. Intron 10 contains the sequence 83163~85776. Intron 11 contains the sequence 85960~88733. Intron 12 contains the sequence 88932~89317. Intron 13 contains the sequence 89426~89696. Intron 14 contains the sequence 89768~110535. Intron 15 contains the sequence 110841~112491. Intron 16 contains the sequence 112556~113450. Intron 17 contains the sequence 113604~113984. Intron 18 contains the sequence 114052~128573. Intron 19 contains the sequence 128759-129075. Intron 20 contains the sequence 129209-134600. Intron 21 contains the sequence 134655-141956. Intron 22 contains the sequence 142103-143059. Intron 23 contains the sequence 143288-145470. Intron 24 contains the sequence 145640-146475. An example of a pathogenic mutation in ATXN2 is the CAG trinucleotide expansion (32 or more CAG repeats) in exon 1. Examples of non-pathogenic mutations include ClinVar accession numbers VCV000522367, VCV000522368, VCV000522369, VCV000522370, VCV000128509, VCV000128508, VCV000128507, and VCV000218618.

[0055] The term "SNCA" gene refers to the gene that codes for the protein synuclein α. A representative sequence of the SNCA gene can be found in the NCBI reference sequence: NG_011851.1 and its corresponding sequence number 55. The boundaries between exons and introns can be defined by the sequences provided in sequence number 55. Specifically, exon 1 contains sequences 1-200. Exon 2 contains sequences 1470-1615. Exon 3 contains sequences 8978-9019. Exon 4 contains sequences 14774-14916. Exon 5 contains sequences 107885-107968. Exon 6 contains sequences 110502-113063. Intron 1 contains sequences 201-1469. Intron 2 contains sequences 1616-8977. Intron 3 contains the sequence 9020–14773. Intron 4 contains the sequence 14917–107884. Intron 5 contains the sequence 107969–110501. The start codon is located in intron 2. Examples of pathogenic mutations in SNCA include duplication or triplication of the gene, A53T, G51D, E46K, and A30P. Examples of non-pathogenic mutations include ClinVar accession numbers VCV000350063, VCV000350064, VCV000350086, and VCV000350093.

[0056] As defined herein, the SOD1 gene refers to the gene that produces the enzyme superoxide dismutase. A representative sequence of the SOD1 gene can be found in the NCBI reference sequence: NG_008689.1 and the corresponding sequence number 57. The boundaries between exons and introns can be defined by the sequences provided in sequence number 57. Specifically, exon 1 contains sequences 5001-5220. Exon 2 contains sequences 9169-9265. Exon 3 contains sequences 11828-11897. Exon 4 contains sequences 12637-12754. Exon 5 contains sequences 13850-14310. Intron 1 contains sequences 5221-9168. Intron 2 contains sequences 9170-11827. Intron 3 contains sequences 11898-12636. Intron 4 contains the sequence 12755–12849. The method described herein provides a transgene for integration into the SOD1 gene. The transgene may comprise a promoter, a partial SOD1 coding sequence, and a splice donor, and the integration site may be within introns 1, 2, 3, or 4 of the endogenous SOD1 gene. Furthermore, the transgene may include an RNAi cassette targeting the endogenous SOD1 transcript, a promoter, a partial SOD1 coding sequence (resistant to silencing by the RNAi cassette), and a splice donor. The transgene may be integrated into introns 1, 2, 3, or 4 of the endogenous SOD1 gene. The transgene may also comprise a splice acceptor, a partial SOD1 coding sequence (resistant to silencing by the RNAi cassette), a terminator, and an RNAi cassette targeting the endogenous SOD1 transcript. The transgene can be incorporated into introns 1, 2, 3, or 4 of the endogenous SOD1 gene. Examples of pathogenic mutations in SOD1 include A5V, C7F, G13R, G17S, E22K, G38R, L39V, G42S, F46C, H47R, G73S, H81R, L85V, G86R, G94R, E101G, I105F, and L107V. Examples of non-pathogenic mutations include ClinVar accessions VCV000440292, VCV000256202, VCV000586633, and VCV000395173.

[0057] As defined herein, the RHO gene refers to a gene that produces the protein rhodopsin. A representative sequence of the RHO gene can be found in the NCBI reference sequence: NC_000003.12 and the corresponding sequence number 58. The boundaries between exons and introns can be defined by the sequences shown in sequence number 58. Specifically, exon 1 contains sequences 1-456. Exon 2 contains sequences 2238-2406. Exon 3 contains sequences 3613-3778. Exon 4 contains sequences 3895-4134. Exon 5 contains sequences 4970-6706. Intron 1 contains sequences 457-2237. Intron 2 contains sequences 2407-3612. Intron 3 contains sequences 3779-3894. Intron 4 contains sequences 4135-4969. The methods described herein provide a transgene for integration into the RHO gene. The transgene may comprise a promoter, a partial RHO coding sequence, and a splice donor, and the integration site may be within introns 1, 2, 3, or 4 of the endogenous RHO gene. Furthermore, the transgene may include an RNAi cassette targeting the endogenous RHO transcript, a promoter, a partial RHO coding sequence (resistant to silencing by the RNAi cassette), and a splice donor. The transgene may be integrated within introns 1, 2, 3, or 4 of the endogenous RHO gene. The transgene may also comprise a splice acceptor, a partial RHO coding sequence (resistant to silencing by the RNAi cassette), a terminator, and an RNAi cassette targeting the endogenous RHO transcript. The transgene may be integrated within introns 1, 2, 3, or 4 of the endogenous RHO gene.Examples of pathogenic mutations in RHO include ClinVar accession numbers VCV000013039, VCV000013031, VCV000013017, VCV000013042, VCV000013018, VCV000625297, VCV000013055, VCV000013013, VCV000013019, VCV000013047, VCV000013016, VCV000013020, VCV000013021, and VCV000013 045, VCV000013054, VCV000625301, VCV000013038, VCV000013022, VCV000013035, VCV000013048, VCV000373094, VCV000013 028, VCV000279882, VCV000013024, VCV000013046, VCV000029875, VCV000013049, VCV000417867, VCV000013050, VCV000143 080, VCV000625303, VCV000013025, VCV000196282, VCV000013033, VCV000590911, VCV000143081, VCV000013023, VCV000013 026, VCV000013043, VCV000013027, VCV000013051, VCV000013034, VCV000013036, VCV000636084, VCV000013030, VCV000523 Examples include 376, VCV000013044, VCV000013029, VCV000419250, VCV000013056, VCV000013052, VCV000013015, VCV000013053, VCV000013032, VCV000013014, VCV000605502, VCV000605497, VCV000442401, VCV000442400, VCV000154258, and VCV000145614.Examples of non-pathogenic mutations include ClinVar accession numbers VCV000343272, VCV000256383, VCV000281512, VCV000256384, VCV000256382, VCV000343286, VCV000343290, VCV000343302, VCV000343303, VCV000343306, and VCV000606153.

[0058] As defined herein, the C9orf72 gene refers to a gene that produces a protein in various tissues and is associated with amyotrophic lateral sclerosis (ALS). A representative sequence of the C9orf72 gene can be found in the NCBI reference sequence: NG_031977.1 and the corresponding sequence number 59. The exon-intron boundaries can be defined by the sequence shown in sequence number 59. Specifically, exon 1 contains sequences 1-158; exon 2 contains sequences 6703-7190; exon 3 contains sequences 8277-8336; exon 4 contains sequences 11391-11486; exon 5 contains sequences 12218-12282; exon 6 contains sequences 13568-13640; and exon 7 contains sequences 15260-15376. Exon 8 contains the sequence 17071-17306. Exon 9 contains the sequence 23160-23217. Exon 10 contains the sequence 25201-25310. Exon 11 contains the sequence 25445-27321. Intron 1 contains the sequence 159-6702. Intron 2 contains the sequence 7191-8276. Intron 3 contains the sequence 8337-11390. Intron 4 contains the sequence 11487-12217. Intron 5 contains the sequence 12283-13567. Intron 6 contains the sequence 13641-15259. Intron 7 contains the sequence 15377-17070. Intron 8 contains the sequence 17307-23159. Intron 9 contains the sequence 23218–25200. Intron 10 contains the sequence 25311–25444. The method described herein provides a transgene for integration into the C9orf72 gene. The transgene may comprise a promoter, a partial C9orf72 coding sequence, and a splice donor, and the integration site may be within introns 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 of the endogenous C9orf72 gene. Furthermore, the transgene may include an RNAi cassette targeting the endogenous C9orf72 transcript, a promoter, a partial C9orf72 coding sequence (resistant to silencing by the RNAi cassette), and a splice donor.The transgene may be incorporated into introns 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 of the endogenous C9orf72 gene. The transgene may also contain a splice acceptor, a partial C9orf72 coding sequence (resistant to silencing by RNAi cassettes), a terminator, and an RNAi cassette targeting the endogenous C9orf72 transcript. Examples of pathogenic mutations in C9orf72 include duplication, triplication, or quadruplication of the C9or72 gene, or expansion of GGGGCC repeats. Examples of non-pathogenic mutations include ClinVar accession numbers VCV000366486, VCV000366521, VCV000366524, VCV000183033, and VCV000611705.

[0059] As defined herein, the CHRNA1 gene refers to a gene that produces the α1 subunit of the nicotinic choline receptor protein. A representative sequence of the CHRNA1 gene can be found in the NCBI reference sequence:NG_008172.1. As defined herein, the CHRND gene refers to a gene that produces the δ subunit of the nicotinic choline receptor protein. A representative sequence of the CHRND gene can be found in the NCBI reference sequence:NG_008028.1. As defined herein, the CHRNE gene refers to a gene that produces the ε subunit of the nicotinic choline receptor protein. A representative sequence of the CHRNE gene can be found in the NCBI reference sequence:NG_008029.2. As defined herein, the CHRNB1 gene refers to a gene that produces the β1 subunit of the nicotinic choline receptor protein. A representative sequence of the CHRNB1 gene can be found in the NCBI reference sequence:NG_008026.1. As defined herein, the PRPS1 gene refers to the gene that produces the protein phosphoribosyl pyrophosphate synthase 1. A representative sequence of the PRPS1 gene can be found in the NCBI reference sequence:NG_008407.1. As defined herein, the LRRK2 gene refers to the gene that produces the protein leucine-rich repeat kinase 2. A representative sequence of the LRRK2 gene can be found in the NCBI reference sequence:NG_011709.1. As defined herein, the STIM1 gene refers to the gene that produces protein-interstitial interaction molecule 1. A representative sequence of the STIM1 gene can be found in the NCBI reference sequence:NG_016277.1. As defined herein, the FGFR3 gene refers to the gene that produces the protein fibroblast growth factor receptor 3. A representative sequence of the FGFR3 gene can be found in the NCBI reference sequence:NG_012632.1. As defined herein, the MECP2 gene refers to the gene that produces protein methylated CpG-binding protein 2. A representative sequence of the MECP2 gene can be found in the NCBI reference sequence:NG_007107.2. As defined herein, the ATXN1 gene refers to the gene that produces protein attaxin 1.A representative sequence of the ATXN1 gene can be found in the NCBI reference sequence:NG_011571.1. As defined herein, the ATXN3 gene refers to the gene that produces the protein attaxin 3. A representative sequence of the ATXN3 gene can be found in the NCBI reference sequence:NG_008198.2. As defined herein, the CACNA1A gene refers to the gene that produces the protein voltage-gated calcium channel subunit α1A. A representative sequence of the CACNA1A gene can be found in the NCBI reference sequence:NG_011569.1. As defined herein, the ATXN7 gene refers to the gene that produces the protein attaxin 7. A representative sequence of the ATXN7 gene can be found in the NCBI reference sequence:NG_008227.1. As defined herein, the TBP gene refers to the gene that produces the protein TATA box-binding protein. A representative sequence of the TBP gene can be found in the NCBI reference sequence:NG_008165.1. As defined herein, the HTT gene refers to the gene that produces the protein huntingtin. A representative sequence of the HTT gene can be found in the NCBI reference sequence:NG_009378.1. As defined herein, the AR gene refers to the gene that produces the protein androgen receptor. A representative sequence of the AR gene can be found in the NCBI reference sequence:NG_009014.2. As defined herein, the FXN gene refers to the gene that produces the protein frataxin. A representative sequence of the FXN gene can be found in the NCBI reference sequence:NG_008845.2. As defined herein, the DMPK gene refers to the gene that produces the protein DM1 protein kinase. A representative sequence of the DMPK gene can be found in the NCBI reference sequence:NG_009784.1. As defined herein, the PABPN1 gene refers to the gene that produces the protein poly(A)-binding nuclear protein 1. A representative sequence of the PABPN1 gene can be found in the NCBI reference sequence:NG_008239.1. As defined herein, the ATXN8 gene refers to the gene that produces the protein attaxin 8.A representative sequence of the ATXN8 gene can be found at genomic coordinates (GRCh38):13:54,700,000-72,800,000.

[0060] As used herein, the term “silencing-resistant subcoding sequence” refers to a subcoding sequence that has mutations compared to a homologous sequence derived from the corresponding endogenous gene, the mutations being designed to block or reduce silencing by the corresponding RNAi cassette. The mutations may be nucleotide insertions, substitutions, or deletions within the DNA sequence encoding the target RNA sequence. The mutations may be sufficient to block or reduce the hybridization of short RNA molecules to the RNA transcript.

[0061] Where used herein, when referring to a silencing-resistant subcoding sequence, “lack of the sequence” refers to the deletion of one or more nucleotides within the corresponding RNAi target site. For example, if RNAi targets a transcript produced by the sequence GGTATCAAGACTACGAAC (within the exon of an endogenous gene), this sequence may also be located within the subcoding sequence of the transgene described herein. To prevent silencing of the modified gene, the RNAi target sequence within the subcoding sequence of the transgene can be modified. Specifically, the site may be mutated by insertion, substitution, or deletion of nucleotides within the site. If the mutation is a deletion, one or more nucleotides may be deleted. If nucleotides are deleted, it is preferable that the deletion be designed to be an in-frame deletion that does not eliminate protein function.

[0062] Where defined herein, “administration” may mean the delivery, provision, or introduction of an exogenous molecule into a cell. When a transgene or rare-cutting endonuclease is administered to a cell, the transgene or rare-cutting endonuclease is delivered, provided, or introduced into the cell. A rare-cutting endonuclease may be administered as a purified protein, nucleic acid, or a mixture of a purified protein and nucleic acid. The nucleic acid (i.e., RNA or DNA) may encode a rare-cutting endonuclease, or a portion of a rare-cutting endonuclease (e.g., gRNA). Administration may be achieved by methods such as lipid-mediated transport, electroporation, direct injection, cell fusion, particle injection, calcium phosphate coprecipitation, DEAE-dextran-mediated transport, viral vector-mediated transport, or any means suitable for delivering a purified protein or nucleic acid, or a mixture of a purified protein and nucleic acid, to a cell.

[0063] Percent sequence identity between a specific nucleic acid or amino acid sequence and a sequence referenced by a specific sequence identification number is determined as follows: First, the nucleic acid or amino acid sequence is compared to the sequence indicated by the specific sequence identification number using the BLAST2 sequencing (Bl2seq) program from the standalone version of BLASTZ, which includes BLASTN version 2.0.14 and BLASTP version 2.0.14. This standalone version of BLASTZ is available online from fr.com / blast or ncbi.nlm.nih.gov. Instructions for using the Bl2seq program can be found in the readme file included with BLASTZ. Bl2seq uses either the BLASTN or BLASTP algorithm to compare the two sequences. BLASTN is used to compare nucleic acid sequences, while BLASTP is used to compare amino acid sequences. To compare two nucleic acid sequences, the options are set as follows: -i is set to the file containing the first nucleic acid sequence to be compared (e.g., C:\seq1.txt), -j is set to the file containing the second nucleic acid sequence to be compared (e.g., C:\seq2.txt), -p is set to blastn, -o is set to any desired filename (e.g., C:\output.txt), -q is set to -1, -r is set to 2, and all other options remain at their default settings. For example, the following command can be used to generate an output file containing the comparison between the two sequences: C:\Bl2seq -ic:\seq1.txt -jc:\seq2.txt -p blastn -oc:\output.txt -q -1 -r 2. To compare two amino acid sequences, the Bl2seq options are set as follows: -i is set to the file containing the first amino acid sequence to be compared (e.g., C:\seq1.txt), -j is set to the file containing the second amino acid sequence to be compared (e.g., C:\seq2.txt), -p is set to blastp, -o is set to any desired filename (e.g., C:\output.txt), and all other options remain at their default settings. For example, the following command can be used to generate an output file containing a comparison between two amino acid sequences: C:\Bl2seq -ic:\seq1.txt -jc:\seq2.txt -p blastp -oc:\output.txt. If the two comparison sequences share homology, the specified output file will then present the homologous regions as aligned sequences. If the two comparison sequences do not share homology, the specified output file will not present aligned sequences.

[0064] When aligned, the number of matches is determined by counting the number of positions where the same nucleotide or amino acid residue is presented in both sequences. Percent sequence identity is determined by dividing the number of matches by either the length of the sequence shown in the identified sequence or the concatenated length (e.g., 100 consecutive nucleotides or amino acid residues from the sequence shown in the identified sequence), and then multiplying the resulting value by 100. The percentage sequence identity value is rounded to the nearest tenth.

[0065] Bidirectional gene repair system with promoter(s) In one embodiment, this document features a method for modifying the 5' end of a transgene and an endogenous gene. The transgene may include first and second promoters, the first promoter operably ligated to a first partial coding sequence, and the second promoter operably ligated to a second partial coding sequence. The first and second partial coding sequences may be operably ligated to first and second splice donor sequences, respectively (Figure 1). The first promoter, the first partial coding sequence, and the first splice donor may be arranged head-to-head with the second promoter, the second partial coding sequence, and the second splice donor. This transgene may be incorporated into the endogenous gene within an intron or at an exon-intron junction. In some embodiments, the transgene may be incorporated into the endogenous gene using a rare-cutting endonuclease or transposon. In one embodiment, the transgene, comprising first and second promoters, first and second partial coding sequences, and first and second splice donors, may be flanked by additional sequences, such as viral inverted terminal repeats (e.g., adeno-associated viral inverted repeats). These transgenes may be incorporated into the endogenous gene via targeted double-strand breaks using rare-cutting endonucleases.

[0066] In another embodiment, a transgene comprising first and second promoters, first and second partial coding sequences, and first and second splice donors can be adjacent to first and second rare-cutting endonuclease target sites. These transgenes can be incorporated into an endogenous gene via targeted double-strand breaks using one or more rare-cutting endonucleases, where one or more rare-cutting endonucleases cleave sequences within the endogenous gene and cleave adjacent target sites within the transgene.

[0067] In another embodiment, a transgene comprising first and second promoters, first and second partial coding sequences, and first and second splice donors may be adjacent to first and second homologous arms. These transgenes can be incorporated into an endogenous gene via targeted double-strand breaks using one or more rare-cutting endonucleases, which cleave the endogenous gene.

[0068] In another embodiment, a transgene comprising first and second promoters, first and second partial coding sequences, and first and second splice donors may be adjacent to first and second homologous arms and first and second rare-cutting endonuclease target sites. These transgenes can be incorporated into an endogenous gene via targeted double-strand breaks using one or more rare-cutting endonucleases, each of which cleaves a sequence in the endogenous gene and adjacent target sites in the transgene. The first and second target sites in the vector may be adjacent to the first and second homologous arms. Alternatively, the first or second target site, or the booths of the first and second target sites, may be located within the homologous arms.

[0069] In another embodiment, the transgene, comprising first and second promoters, first and second partial coding sequences, and first and second splice donors, can be adjacent to the left and right transposon ends. These transgenes can be incorporated into the endogenous gene through transposition using a transposase. As described herein, the transposase may be a CRISPR-associated transposase.

[0070] In some embodiments, the first and second promoters can be replaced with bidirectional promoters. In other embodiments, the transgene may further include first and second terminators positioned tail-to-tail orientation between the first and second promoters (Figure 1). Alternatively, the first and second terminators can be replaced with bidirectional terminators.

[0071] In one embodiment, this document features a method for modifying the 5' end of an endogenous gene, the endogenous gene having at least one intron between two coding exons. The intron may be any intron that is removed from precursor messenger RNA by a normal messenger RNA processing mechanism. The intron may be 20 bp to over 500 kb and include elements comprising a splice donor site, a branching sequence, and an acceptor site. The transgene disclosed herein for modification of the 5' end of an endogenous gene may include multiple functional elements, including a target site for a rare-cutting endonuclease, homologous arms, a splice acceptor sequence, a coding sequence, a promoter, and a transcription terminator (Figure 1).

[0072] In embodiments, the site for integration of the transgene may be an intron or an intron-exon junction. When targeting an intron, the partial coding sequence may include a sequence encoding a peptide produced by the exon preceding the intron in the endogenous gene. For example, if the transgene is designed to be integrated into intron 2 of an endogenous gene having 12 exons, the partial coding sequence may encode peptides produced by exons 1 and 2 of the endogenous gene. When targeting an exon-intron junction, the transgene may be integrated into the exon-intron junction in such a way that the intron sequence is preserved. In one embodiment, after integration, the intron sequence is preserved and the upstream exon sequence is preserved (i.e., a nucleotide derived from the transgene is added between the last nucleotide in the exon and the first nucleotide in the intron). Alternatively, in one embodiment, after integration, the intron sequence is preserved, but one or more nucleotides in the exon sequence are removed.

[0073] In one embodiment, the transgene contains two target sites for a rare-cutting endonuclease. The target sites may be sequences and chain lengths suitable for cleavage by the rare-cutting endonuclease. The target sites may be suitable for cleavage by the CRISPR system, TAL effector nuclease, zinc finger nuclease or meganuclease, or a combination of the CRISPR system, TALE nuclease, zinc finger nuclease or meganuclease, or any other rare-cutting endonuclease. The target sites may be positioned such that cleavage by the rare-cutting endonuclease results in the release of the transgene from the vector. The vector may include a viral vector (e.g., an adeno-associated vector) or a non-viral vector (e.g., a plasmid, a minicircle vector). If the transgene contains two target sites, the target sites may be the same sequence (i.e., targeted by the same rare-cutting endonuclease) or they may be different sequences (i.e., targeted by two or more different rare-cutting endonucleases).

[0074] In some embodiments, the transgenes provided herein may be incorporated by a transposase. The transposase may include a CRISPR transposase (Strecker et al., Science 10.1126 / science.aax9181, 2019; Klompe et al., Nature, 10.1038 / s41586-019-1323-z, 2019). The transposase can be used in combination with a transgene containing first and second splice acceptor sequences, first and second coding sequences, one bidirectional terminator, or first and second terminators (Figure 1), as well as the left and right ends of a transposon. The CRISPR transposase may include the Type V-U5, C2C5 CRISPR protein, and Cas12k, along with the proteins tnsB, tnsC, and tniQ. In some embodiments, Cas12k may be derived from Scytonema hofmanni (SEQ ID NO: 30) or Anabaena cylindrica (SEQ ID NO: 31). In one embodiment, the transgene described herein, including the left (SEQ ID NO: 32) and right (SEQ ID NO: 33) ends of the transposon, can be delivered to cells together with ShCas12k, tnsB, tnsC, TniQ, and gRNA (SEQ ID NO: 44). Alternatively, the CRISPR transposase may contain the Cas6 protein together with helper proteins including Cas7, Cas8, and TniQ. In one embodiment, the transgene described herein, including the left (SEQ ID NO: 41) and right (SEQ ID NO: 43) ends of a transposon, can be delivered to eukaryotic cells together with Cas6 (SEQ ID NO: 37), Cas7 (SEQ ID NO: 36), Cas8 (SEQ ID NO: 35), TniQ (SEQ ID NO: 34), TnsA (SEQ ID NO: 38), TnsB (SEQ ID NO: 39), TnsC (SEQ ID NO: 40), and gRNA (SEQ ID NO: 42). The proteins can be administered directly to cells as purified proteins or encoded in RNA or DNA. If encoded in RNA or DNA, the sequence can be optimized for codons for expression in eukaryotic cells. The gRNA (SEQ ID NO: 42) can be placed downstream of the RNApolIII promoter and terminated with a poly(T) terminator.

[0075] In one embodiment, the transgene includes first and second target sites, as well as first and second homologous arms. The first and second homologous arms may include sequences homologous to the genome sequence at or near the desired integration site. The homologous arms may have a suitable chain length for engaging in homologous recombination with the sequence at or near the desired integration site. The chain length of each homologous arm may be 50 nt to 10,000 nt (e.g., 50 nt, 100 nt, 200 nt, 300 nt, 400 nt, 500 nt, 600 nt, 700 nt, 800 nt, 900 nt, 1,000 nt, 2,000 nt, 3,000 nt, 4,000 nt, 5,000 nt, 6,000 nt, 7,000 nt, 8,000 nt, 9,000 nt, 10,000 nt). In one embodiment, the homologous arm may include a functional element containing a target site for a rare-cutting endonuclease. In one embodiment, the first homologous arm (e.g., the left homologous arm) may contain a sequence homologous to a target exon or intron, and the second homologous arm may contain a sequence homologous to a downstream genomic sequence of the first homologous arm. The first homologous arm must not have splice acceptor function with respect to the transcription direction from the promoter on the transgene. Several steps, including in silico analysis and experimental testing, can be performed to determine whether a sequence contains splice acceptor function. To determine whether there is potential for splice acceptor function, desired sequences for the second homologous arm can be searched for consensus branched sequences (e.g., YTRAC) and splice acceptor sites (e.g., Y-rich NCAGG). If branched or splice acceptor sequences are present, a single nucleotide polymorphism can be introduced to disrupt the function, or a different but adjacent sequence that does not contain such a sequence can be selected. To experimentally determine whether the first homologous arm possesses splice acceptor function, a synthetic construct containing the first homologous arm within an intron of the reporter gene can be constructed. The construct can then be administered to an appropriate cell type, and the reporter gene activity can be evaluated to monitor splicing function.

[0076] In one embodiment, the transgene comprises two splice donor sequences, hereafter referred to as the first and second splice donor sequences. The first and second splice donor sequences are located in the transgene in opposite directions (i.e., head-to-head orientation) and adjacent to each other (i.e., a partial coding sequence and a promoter). When the transgene is incorporated into an intron in a forward or reverse direction, the splice donor sequences facilitate the initiation of splicing of the intron within the corresponding premRNA. The first and second splice donor sequences may be the same sequence or different sequences. One or both splice donor sequences may be splice donor sequences of the intron into which the transgene is incorporated. One or both splice donor sequences may be synthetic splice donor sequences or splice donor sequences from introns derived from different genes.

[0077] In one embodiment, the transgene includes first and second coding sequences operably ligated to first and second splice donor sequences. The first and second coding sequences are positioned in opposite directions (i.e., head-to-head orientation) within the transgene. When the transgene is incorporated into the endogenous gene in a forward or reverse direction, the first and second coding sequences are transcribed into mRNA by a promoter located within the transgene. The coding sequences may be designed to correct a defective coding sequence, introduce a mutation, or introduce a novel peptide sequence. The first and second coding sequences may code for the same nucleic acid sequence and the same protein. Alternatively, the first and second coding sequences may be different nucleic acid sequences and code for the same protein (i.e., using codon degeneracy). The coding sequence can encode a purified tag (e.g., glutathione-S-transferase, poly(His), maltose-binding protein, Strep tag, Myc tag, AviTag, HA tag, or chitin-binding protein) or a reporter protein (e.g., GFP, RFP, lacZ, cat, luciferase, puro, neomycin).

[0078] In one embodiment, modification of the N-terminus of a protein encoded by an endogenous gene can be obtained by modifying the 5' end of the endogenous gene using the methods and compositions described herein. Modification of the 5' end of the coding sequence of an endogenous gene may include substitutions from the first coding exon to the exons between the first and last exons. For example, if the gene contains 12 exons, the modification may include substitutions of exon 1, or 1-2, or 1-3, or 1-4, or 1-5, or 1-6, or 1-7, or 1-8, or 1-9, or 1-10, or 1-11. In one embodiment, the endogenous exon to be substituted may be replaced with a similar sequence. For example, the first or second coding sequence of a transgene may include exon 1, or 1-2, or 1-3, or 1-4, or 1-5, or 1-6, or 1-7, or 1-8, or 1-9, or 1-10, or 1-11. The transgene can be integrated within the endogenous gene in the intron downstream of the last exon in the transgene's coding sequence (Figure 3). Alternatively, the transgene can be integrated within the exon corresponding to the last exon in the transgene's coding sequence (Figure 8). The transgene is designed to be 4.7 kb or less and can be incorporated into AAV vectors and particles and delivered to target cells in vivo.

[0079] In one embodiment, the transgene may include a bidirectional promoter, or a first and second promoter, operably ligated to the first and second coding sequences. The bidirectional promoter, or the first and second promoters, are positioned in opposite directions (i.e., head-to-head orientation) within the transgene. When the transgene is incorporated into the endogenous gene in a forward or reverse direction, the bidirectional promoter, or the first and second promoters, initiate transcription of the first and second coding sequences. The first and second promoters may be the same promoter or different promoters.

[0080] In one embodiment, the transgene may include a bidirectional promoter, or a first and second promoter, operably ligated to the first and second coding sequences. The bidirectional promoter, or the first and second promoters, are positioned in opposite directions (i.e., head-to-head orientation) within the transgene. When the transgene is integrated into the endogenous gene in a forward or reverse direction, the bidirectional promoter, or the first and second promoters, initiate transcription of the first and second coding sequences. The first and second promoters may be the same promoter or different promoters. The promoters may be selected from, for example, CMV, EF1α, SV40, PGK1, Ubc, human β-actin, CAG, or any promoter with sufficient activity to initiate transcription of the partial coding sequence. While not theoretically bound, a reverse-directed promoter may induce the generation of double-stranded RNA, thereby resulting in silencing of gene expression upstream of the integration site. Furthermore, a forward-directed promoter may not undergo the same silencing (e.g., due to codon degeneracy of the coding sequence) and may initiate RNA transcription. Methods for reducing potential RNAi from RNA produced in the reverse direction by a promoter are also described herein (Figure 5).

[0081] In one embodiment, the transgene may include a bidirectional terminator, or first and second terminators, between a first promoter and a second promoter (Figure 1). The bidirectional terminator, or first and second terminators, are positioned in opposite directions (i.e., tail-to-tail orientation) within the transgene. When the transgene is incorporated into the endogenous gene in a forward or reverse direction, the bidirectional terminator, or first and second terminators, terminate transcription from the promoter of the endogenous gene. The first and second terminators may be the same terminator or different terminators.

[0082] In one embodiment, this document provides a transgene comprising first and second rare-cutting endonuclease target sites, first and second splice donor sequences, first and second coding sequences, and one bidirectional promoter, or first and second promoters. The transgene can be incorporated into an endogenous gene via non-homologous-dependent methods, including non-homologous end joining and alternative non-homologous end joining, or by microhomology-mediated end joining. In one embodiment, the transgene is incorporated into an intron within the endogenous gene (Figure 2).

[0083] In another embodiment, this document provides a transgene comprising first and second homologous arms, first and second rare-cutting endonuclease target sites, first and second splice donor sequences, first and second coding sequences, and one bidirectional promoter, or first and second promoters. The transgene can be incorporated into an endogenous gene via both homology-dependent methods (e.g., synthesis-dependent chain annealing and microhomology-mediated end joining) and homology-independent methods (e.g., non-homologous end joining and alternative non-homologous end joining). In one embodiment, the transgene is incorporated into an intron within an endogenous gene (Figure 3). In another embodiment, the transgene is incorporated into an exon of an endogenous gene (Figure 8).

[0084] In another embodiment, this document provides a transgene comprising first and second homologous arms, first and second splice donor sequences, first and second coding sequences, and one bidirectional promoter, or first and second promoters (Figure 1).

[0085] In another embodiment, this document provides a transgene comprising first and second homologous arms, first and second coding sequences, first and second splice donor sequences, one bidirectional terminator, or first and second terminators, and first and second additional sequences (Figure 1). The additional sequences may be any additional sequences present on the transgene at the 5' and 3' ends, but the additional sequences should not include any elements that function as splice acceptors or splice donors. The additional sequences may be, for example, inverted end repeats of an adeno-associated virus genome, or left and right transposon ends.

[0086] In another embodiment, this document provides transgenes in viral vectors containing adeno-associated viruses and adenoviruses, the transgene comprising first and second splice donor sequences, first and second coding sequences, and one bidirectional terminator, or first and second terminators. Due to inverted terminal repeats of the viral vector, the transgene also comprises first and second additional sequences.

[0087] In another embodiment, this document provides transgenes in viral vectors containing adeno-associated viruses and adenoviruses, the transgene comprising first and second homologous arms, first and second splice donor sequences, first and second coding sequences, and one bidirectional promoter or first and second promoters. Due to inverted terminal repeats of the viral vector, the transgene also comprises first and second additional sequences.

[0088] In another embodiment, a transgene for integration may be designed to integrate through multiple repair pathways, producing the desired effect in each outcome. For example, a transgene may include first and second arm homologous arms, first and second rare-cutting endonuclease target sites, first and second coding sequences, and first and second promoters, and may be located within the AAV genome (i.e., adjacent to a 145-nucleotide inverted terminal repeat). After expression by rare-cutting endonuclease, the following outcomes may occur: 1) forward or reverse integration of the entire AAV genome at the target site by NHEJ, 2) forward or reverse integration of the sequence between the first and second rare-cutting endonuclease target sites at the target site by NHEJ, 3) integration by HR using the first and second homologous arms, or 4) any combination of the above outcomes. After integration with any of the above outcomes, the transgene described herein can modify or alter the protein sequence produced by the endogenous gene.

[0089] In some embodiments, the transgene described herein may have a combination of elements including a splice donor, a partial coding sequence, a promoter, homologous arms, left and right transposase ends, and a cleavage site by a rare-cutting endonuclease. In one embodiment, the combination may be 5' to 3'.

[0090] In some embodiments, the transgene described herein may have a combination of elements including a splice acceptor, a partial coding sequence, a terminator, homologous arms, left and right transposase ends, and cleavage sites by rare-cutting endonucleases.

[0091] In one embodiment, the combination may be [splice donor 1 RC]-[partial coding sequence 1 RC]-[promoter 1 RC]-[promoter 2]-[partial coding sequence 2]-[splice donor 2] from 5' to 3' (where RC represents reverse complementarity). This combination may be encapsulated on a linear DNA molecule or an AAV molecule and incorporated by NHEJ via targeted cleavage of the target gene.

[0092] In another embodiment, the combination from 5' to 3' may be [rare-cutting endonuclease cleavage site 1]-[splice donor 1 RC]-[partial coding sequence 1 RC]-[promoter 1 RC]-[promoter 2]-[partial coding sequence 2]-[splice donor 2]-[rare-cutting endonuclease cleavage site 2].

[0093] In another embodiment, the combination may be, from 5' to 3', [rare-cutting endonuclease cleavage site 1]-[homologous arm 1]-[splice donor 1 RC]-[partial coding sequence 1 RC]-[promoter 1 RC]-[promoter 2]-[partial coding sequence 2]-[splice donor 2]-[homologous arm 2]-[rare-cutting endonuclease cleavage site 2]. In this combination, one or more rare-cutting endonucleases can be used to promote HR and NHEJ. For example, a single rare-cutting nuclease can cleave a target gene (i.e., a desired intron), and the cleavage sites adjacent to the homologous arms can be designed to be the same target sequence within the intron.

[0094] In another embodiment, the combination may be 5' to 3', [homologous arm 1 + rare-cutting endonuclease cleavage site 1]-[splice donor 1 RC]-[partial coding sequence 1 RC]-[promoter 1 RC]-[promoter 2]-[partial coding sequence 2]-[splice donor 2]-[homologous arm 2]-[rare-cutting endonuclease cleavage site 2]. In this combination, one or more rare-cutting endonucleases can promote HR and NHEJ. For example, a single rare-cutting nuclease can cleave within homologous arm 1, downstream of homologous arm 2, and at genomic target sites (i.e., sites homologous to the sequence of homologous arm 1).

[0095] In another embodiment, the combination may be 5' to 3', [left end of transposase]-[splice donor 1 RC]-[partial coding sequence 1 RC]-[promoter 1 RC]-[promoter 2]-[partial coding sequence 2]-[splice donor 2]-[right end of transposase]. In all embodiments, splice donor 1 and splice donor 2 may be the same or different sequences, partial coding sequence 1 and partial coding sequence 2 may be the same or different sequences, and promoter 1 and promoter 2 may be the same or different sequences.

[0096] In embodiments, a transgene comprising the structure [rare-cutting endonuclease cleavage site 1]-[homologous arm 1]-[splice donor 1 RC]-[partial coding sequence 1]-[promoter 1 RC]-[promoter 2]-[partial coding sequence 2]-[splice donor 2]-[homologous arm 2]-[rare-cutting endonuclease cleavage site 2] can be incorporated into DNA through the delivery of one or more rare-cutting endonucleases. When one rare-cutting endonuclease is delivered, it can release the transgene by cleaving at rare-cutting endonuclease cleavage sites 1 and 2. Furthermore, the same rare-cutting endonuclease can generate cleavage within the target gene, simulating insertion via HR or NHEJ.

[0097] In other embodiments, a transgene comprising the structure [homologous arm 1 + rare-cutting endonuclease cleavage site 1]-[splice donor 1 RC]-[partial coding sequence 1]-[promoter 1 RC]-[promoter 2]-[partial coding sequence 2]-[splice donor 2]-[homologous arm 2]-[rare-cutting endonuclease cleavage site 1] can be incorporated into DNA by delivery of one or more rare-cutting endonucleases. When one rare-cutting endonuclease is delivered, it can release the transgene by cleavage at rare-cutting endonuclease cleavage sites 1 and 2. Furthermore, the same rare-cutting endonuclease can generate cleavage within the target gene, simulating insertion via HR or NHEJ. HR-mediated incorporation can occur when the cleavage is upstream of the incorporation site (i.e., within the homologous arm).

[0098] In embodiments, the subcoding sequences can be codon-tuned. Codon tuning may aim to 1) reduce double-stranded RNA pairing (Figure 5) and 2) optimize protein expression. An introduced gene containing first and second subcoding sequences operably linked to first and second promoters is incorporated into an endogenous gene, and if the first and second subcoding sequences are homologous to each other and the gene is endogenous, double-stranded RNA may be produced (Figure 5). The subcoding sequences can be codon-tuned to minimize RNA pairing. In one embodiment, codon optimization may be complete and different with respect to the first and second subcoding sequences. For example, subcoding sequence 1 may have a different nucleotide sequence from subcoding sequence 2, and both subcoding sequences 1 and 2 may be different sequences from the corresponding sequences in the endogenous gene of interest.

[0099] In another embodiment, codon optimization may be divided between first and second subcoding sequences. For example, the first subcoding sequence may have a mixture of an uncodon-adjusted sequence (i.e., homologous to the corresponding sequence in the endogenous gene of interest) and a codon-adjusted sequence. In this embodiment, the second subcoding sequence may have the opposite adjustment. For example, in 200-nucleotide subcoding sequences 1 and 2, nucleotides 1-100 of subcoding sequence 1 may be homologous to the sequence in the endogenous gene of interest, and nucleotides 101-200 may be codon-adjusted to have minimal sequence similarity to the endogenous gene of interest; and nucleotides 1-100 of subcoding sequence 2 may be codon-adjusted to have minimal sequence similarity to the endogenous gene of interest, and nucleotides 101-200 may be homologous to the sequence in the endogenous gene of interest.

[0100] In one embodiment, genome modification is the insertion of a transgene into the endogenous ATXN2 genome sequence. The transgene may contain a partial coding sequence for the ATXN2 protein. The partial coding sequence may be homologous to a coding sequence in the wild-type ATXN2 gene, a functional variant of the wild-type ATXN2 gene, a codon-modified version of the ATXN2 gene, or a mutant ATXN2 gene. In one embodiment, the transgene encoding a partial ATXN2 protein is inserted into intron 1 of the endogenous ATXN2 gene (Figures 3 and 4).

[0101] In one embodiment, the transgene provided herein comprises first and second partial coding sequences encoding peptides produced by exon 1 of the ATXN2 gene (Figure 7). The transgene may be incorporated into the endogenous ATXN2 gene within intron 1 or into the exon 1-intron 1 junction. This embodiment is particularly useful in cells containing extended trinucleotide repeats in exon 1 of ATXN2.

[0102] The methods and compositions provided herein can be used to modify genes encoding proteins within cells. Endogenous proteins include fibrinogen, prothrombin, tissue factor, factor V, factor VII, factor VIII, factor IX, factor X, factor XI, factor XII (Hagemann factor), factor XIII (fibrin stabilizing factor), von Willebrand factor, prekallikrein, high molecular weight kininogen (Fitzgerald factor), fibronectin, antithrombin factor III, heparin cofactor II, protein C, protein S, protein Z, protein Z-related protease inhibitors, and plasminogen. α2-antiplasmin, tissue plasminogen activator, urokinase, plasminogen activator inhibitor-1, plasminogen activator inhibitor-2, glucocerebrosidase (GBA), α-galactosidase A (GLA), iduronate sulfatase (IDS), iduronidase (IDUA), acid sphingomyelinase (SMPD1), MMAA, MMAB, MMACHC, MMADHC (C2orf25), MTRR, LMBRD1, MTR, propionyl-CoA carboxylase (PCC) (PCCA , and / or PCCB subunits), glucose-6-phosphate transporter (G6PT) protein or glucose-6-phosphatase (G6Pase), LDL receptor (LDLR), ApoB, LDLRAP-1, PCSK9, mitochondrial proteins (such as NAGS (N-acetylglutamate synthase), CPS1 (carbamoyl phosphate synthase I), and OTC (ornithine transcarbamylase)), ASS (argininosuccinate synthase), ASL (argininosuccinate lyase) ) and / or ARG1 (arginase), and / or solute carrier family 25 (SLC25A13, aspartic acid / glutamic acid carrier) protein, UGT1A1 or UDP glucuronsyltransferase polypeptide A1, fumarylacetoacetate hydrolyase (FAH), alaning glyoxylate aminotransferase (AGXT) protein, glyoxylate reductase / hydroxypyruvate reductase (GRHPR) protein, transthyretin gene (TTR) protein,Examples include ATP7B protein, phenylalanine hydroxylase (PAH) protein, USH2A protein, ATXN protein, and lipoprotein lyase (LPL) protein.

[0103] A transgene may contain a sequence for modifying an endogenous gene that carries a loss-of-function or gain-of-function mutation. Mutations may include those resulting in the following genetic disorders: achondroplasia, color blindness, acid maltase deficiency, adenosine deaminase deficiency, adrenoleukodystrophy, Aicardi syndrome, α1 antitrypsin deficiency, α-thalassemia, androgen insensitivity syndrome, Apert syndrome, arrhythmogenic right ventricular dysplasia, telangiectasia ataxia, Barth syndrome, β-thalassemia, blue rubber nipple nevus syndrome, Canavan disease, chronic granulomatous disease (CGD), cat cry syndrome, cystic fibrosis, Darkham's disease, ectodermal dysplasia, Fanconi anemia, fibrodysplasia ossificans progressive, fragile X syndrome, galactosemia, systemic gangliosidosis (e.g., GM1), hemochromatosis, and β-globin (HbC) 6. Hemoglobin C mutations in codons of the eye, hemophilia, Huntington's disease, hypophosphatasia, Klinefelter syndrome, Krabbe disease, Langer-Gideon syndrome, leukocyte adhesion deficiency, leukodystrophy, long QT syndrome, Marfan syndrome, Moebius syndrome, mucopolysaccharidosis (MPS), onychopatella syndrome, nephrogenic diabetes insipidus, neurofibromatosis, Niemann-Pick disease, osteogenesis imperfecta, porphyria, Prader-Willi syndrome, progeria, p Lotheus syndrome, retinoblastoma, Rett syndrome, Rubinstein-Taybe syndrome, Sanfilippo syndrome, severe combined immunodeficiency (SCID), Schwakman syndrome, sickle cell anemia, Smith-Magenis syndrome, Stickler syndrome, Tay-Sachs disease, thrombocytopenia-radial aplasia (TAR) syndrome, Treacher-Collins syndrome, trisomy, tuberous sclerosis, Turner syndrome, urea cycle disorders, von Hippel-Lindau disease, Waardenburg syndrome, Williams syndrome, Wilson's disease, Wiscott-Aldrich syndrome, X-linked lymphoproliferative syndrome, lysosomal storage disorders (e.g., Gaucher disease, GM1, Fabry disease, and Tay-Sachs disease), von Willebrand disease, Usher syndrome, polycystic kidney disease, spinocerebellar ataxia type 2, spinal and bulbar muscular atrophy, Friedreich's ataxia, and myotonic dystrophy type 2.

[0104] As described herein, the transgene may be contained within a viral or non-viral vector. The vector may be in the form of circular or linear double-stranded or single-stranded DNA. The donor molecule may be conjugated or associated with reagents that promote stability or cellular update. The reagents may be lipids, calcium phosphate, cationic polymers, DEAE-dextran, dendrimers, polyethylene glycol (PEG), cell membrane permeable peptides, gas-filled microbubbles, or magnetic beads. The donor molecule may be incorporated into a viral particle. The virus may be a retrovirus, adenovirus, adeno-associated vector (AAV), herpes simplex virus, poxvirus, hybrid adenovirus vector, Epstein-Barr virus, lentivirus, or herpes simplex virus.

[0105] Gene repair system using RNAi cassettes In another embodiment, the method described herein may be used to silence an endogenous gene while simultaneously replacing the RNA / protein lost due to silencing. In one embodiment, the method may include administering a transgene to cells, the transgene comprising two functional elements: 1) a silencing sequence, and 2) a complete coding sequence encoding a protein homologous to the silenced protein but resistant to silencing (Figure 9). The two functional elements may be on separate transgenes or on the same transgene. In another embodiment, the method may include administering a transgene to cells, the transgene comprising a partial or complete coding sequence for the repair of a mutant gene that is incorporated into the endogenous gene of interest and comprises 1) a silencing sequence, and 2) a silencing-resistant sequence (Figures 12-17).

[0106] A silencing sequence may include a promoter, a nucleic acid sequence that functions to silence a target nucleic acid, and a terminator. The nucleic acid sequence may be in a form that can induce gene silencing within a target nucleic acid (e.g., microRNA, hairpin RNA, antisense RNA). The nucleic acid sequence may target different regions of the mRNA of the target gene and may include the 5'UTR, coding sequence, or 3'UTR.

[0107] In one embodiment, this document describes a method for silencing and replacing the production of a target protein by administering the transgene shown in Figure 13 to cells and incorporating the transgene into the target endogenous gene. In one embodiment, the transgene may include a splice acceptor, a partial coding sequence (resistant to silencing), a terminator, and an RNAi cassette designed to silence the target endogenous gene. The splice acceptor may be operably ligated to a partial coding sequence which can be operably ligated to a terminator. The splice acceptor, partial coding sequence, terminator, and RNAi cassette may be adjacent to first and second homologous arms, or the left and right transposon ends. The transgene may be incorporated into an intron or an intron-exon junction within the target endogenous gene. The partial coding sequence may encode the remaining peptide sequence relative to the location where the transgene is incorporated. For example, if the transgene is incorporated into intron 3 of a gene containing five exons (Figure 13), the partial coding sequence can encode peptides produced by exons 4 and 5 of the endogenous gene. The RNAi cassettes within these transgenes can target sequences in exon 4 or 5 or in the 3' UTR. Thus, the corresponding target sites within the partial coding sequence in the transgene can be modified to block the silencing of the modified endogenous allele. In other embodiments, the transgene may include first and second splice acceptors, first and second partial coding sequences (both resistant to silencing), first and second terminators, and RNAi cassettes. These transgenes may be adjacent to additional sequences (e.g., viral ITRs), first and second rare-cutting endonuclease target sites, left and right transposon ends, or both first and second homologous arms and first and second rare-cutting endonuclease target sites.In one embodiment, the structure of the transgene may be, from 5' to 3', [homologous arm 1]-[splice acceptor]-[partial coding sequence]-[terminator]-[RNAi cassette]-[homologous arm 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [left end of transposase]-[splice acceptor]-[partial coding sequence]-[terminator]-[RNAi cassette]-[right end of transposase]. In yet another embodiment, the structure of the transgene may be, from 5' to 3', [additional sequence 1]-[splice acceptor 1]-[partial coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[partial coding sequence 2 RC]-[splice acceptor 2 RC]-[additional sequence 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[splice acceptor 1]-[partial coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[partial coding sequence 2 RC]-[splice acceptor 2 RC]-[rare-cutting endonuclease target site 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[homologous arm 1]-[splice acceptor 1]-[partial coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[partial coding sequence 2 RC]-[splice acceptor 2 RC]-[homologous arm 2]-[rare-cutting endonuclease target site 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [left end of transposase]-[splice acceptor 1]-[partial coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[partial coding sequence 2 RC]-[splice acceptor 2 RC]-[right end of transposase].

[0108] In one embodiment, this document describes a method for silencing and replacing the production of a target protein by administering the transgene shown in Figure 14 to cells and incorporating the transgene into the target endogenous gene. In one embodiment, the transgene may include a splice acceptor, a 2A sequence, a complete coding sequence (resistant to silencing), a terminator, and an RNAi cassette designed to silence the target endogenous gene. The splice acceptor may be operably ligated to the 2A sequence and operably ligated to the complete coding sequence, which may be operably ligated to the terminator. The splice acceptor, 2A sequence, complete coding sequence, terminator, and RNAi cassette may be adjacent to the first and second homologous arms, or the left and right transposon ends. The transgene may be incorporated into an intron or an intron-exon junction within the target endogenous gene (Figure 14). RNAi can be designed to silence the expression of an endogenous gene of interest, and the complete coding sequence within the transgene can be designed to be resistant to silencing. Thus, the corresponding target sites within the complete coding sequence within the transgene can be modified to block silencing. In other embodiments, the transgene may include first and second splice acceptors, first and second 2A sequences, first and second coding sequences (both resistant to silencing), first and second terminators, and an RNAi cassette. These transgenes may be adjacent to additional sequences (e.g., viral ITRs), first and second rare-cutting endonuclease target sites, left and right transposon ends, or both first and second homologous arms and first and second rare-cutting endonuclease target sites. In one embodiment, the structure of the transgene may be [homologous arm 1]-[splice acceptor]-[2A]-[coding sequence]-[terminator]-[RNAi cassette]-[homologous arm 2] from 5' to 3'.In another embodiment, the structure of the transgene may be, from 5' to 3', [left end of transposase]-[splice acceptor]-[2A]-[coding sequence]-[terminator]-[RNAi cassette]-[right end of transposase]. In another embodiment, the structure of the transgene may be, from 5' to 3', [additional sequence 1]-[splice acceptor 1]-[2A1]-[coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[coding sequence 2 RC]-[2A2 RC]-[splice acceptor 2 RC]-[additional sequence 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[splice acceptor 1]-[2A1]-[coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[coding sequence 2 RC]-[2A2 RC]-[splice acceptor 2 RC]-[rare-cutting endonuclease target site 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[homologous arm 1]-[splice acceptor 1]-[2A1]-[coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[coding sequence 2 RC]-[2A2 RC]-[splice acceptor 2 RC]-[homologous arm 2]-[rare-cutting endonuclease target site 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [left end of transposase]-[splice acceptor 1]-[2A1]-[coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[coding sequence 2 RC]-[2A2 RC]-[splice acceptor 2 RC]-[right end of transposase].

[0109] In one embodiment, this document describes a method for silencing and replacing the production of a target protein by administering the transgene shown in Figure 15 to cells and incorporating the transgene into the target endogenous gene. In one embodiment, the transgene may include a 2A sequence, a complete coding sequence (resistant to silencing), a terminator, and an RNAi cassette designed to silence the target endogenous gene. The 2A sequence may be operably ligated to a complete coding sequence which can be operably ligated to a terminator. The 2A sequence, complete coding sequence, terminator, and RNAi cassette may be adjacent to first and second homologous arms, or the left and right transposon ends. The transgene may be incorporated into an exon within the target endogenous gene (Figure 15). The RNAi may be designed to silence the expression of the target endogenous gene, and the complete coding sequence within the transgene may be designed to be resistant to silencing. Thus, the corresponding target site within the complete coding sequence within the transgene may be modified to block silencing. In other embodiments, the transgene may include first and second 2A sequences, first and second coding sequences (both resistant to silencing), first and second terminators, and an RNAi cassette. These transgenes may be adjacent to additional sequences (e.g., viral ITRs), first and second rare-cutting endonuclease target sites, left and right transposon ends, or both the first and second homologous arms and the first and second rare-cutting endonuclease target sites. In one embodiment, the structure of the transgene may be 5' to 3', [homologous arm 1]-[2A]-[coding sequence]-[terminator]-[RNAi cassette]-[homologous arm 2]. In another embodiment, the structure of the transgene may be 5' to 3', [left end of transposase]-[2A]-[coding sequence]-[terminator]-[RNAi cassette]-[right end of transposase].In another embodiment, the structure of the transgene may be, from 5' to 3', [additional sequence 1]-[2A1]-[coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[coding sequence 2 RC]-[2A2 RC]-[additional sequence 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[2A1]-[coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[coding sequence 2 RC]-[2A2 RC]-[rare-cutting endonuclease target site 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[homologous arm 1]-[2A1]-[coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[coding sequence 2 RC]-[2A2 RC]-[homologous arm 2]-[rare-cutting endonuclease target site 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [left end of transposase]-[2A1]-[coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[coding sequence 2 RC]-[2A2 RC]-[right end of transposase].

[0110] In one embodiment, this document describes a method for silencing and replacing the production of a target protein by administering the transgene shown in Figure 16 to cells and incorporating the transgene into the target endogenous gene. In one embodiment, the transgene may include a complete coding sequence (resistant to silencing and containing a start codon), a terminator, and an RNAi cassette designed to silence the target endogenous gene. The complete coding sequence may be operably linked to the terminator. The complete coding sequence, terminator, and RNAi cassette may be adjacent to the first and second homologous arms, or the left and right transposon ends. The integration site may be within the 5'UTR or before the start codon (Figure 16). Additional integration sites, if present, may be within an intron in the 5'UTR, but the transgene described in this embodiment must then include a splice acceptor sequence operably linked to the complete coding sequence(s). RNAi can be designed to silence the expression of an endogenous gene of interest, and the complete coding sequence within the transgene can be designed to be resistant to silencing. Thus, the corresponding target sites within the complete coding sequence within the transgene can be modified to block silencing. In other embodiments, the transgene may include first and second coding sequences (both resistant to silencing), first and second terminators, and an RNAi cassette. These transgenes may be adjacent to additional sequences (e.g., viral ITRs), first and second rare-cutting endonuclease target sites, left and right transposon ends, or both first and second homologous arms and first and second rare-cutting endonuclease target sites. In one embodiment, the structure of the transgene may be 5' to 3', [homologous arm 1]-[coding sequence]-[terminator]-[RNAi cassette]-[homologous arm 2]. In another embodiment, the structure of the transgene may be 5' to 3', [left end of transposase]-[coding sequence]-[terminator]-[RNAi cassette]-[right end of transposase].In another embodiment, the structure of the transgene may be, from 5' to 3', [additional sequence 1]-[coding sequence 1]-[terminator 1]-[RNAi cassette]-[terminator 2 RC]-[coding sequence 2 RC]-[additional sequence 2]. In yet another embodiment, the transgene may be designed to replace protein production without silencing the endogenous gene. In one embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[coding sequence 1]-[terminator 1]-[terminator 2 RC]-[coding sequence 2 RC]-[rare-cutting endonuclease target site 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[homologous arm 1]-[coding sequence 1]-[terminator 1]-[terminator 2 RC]-[coding sequence 2 RC]-[homologous arm 2]-[rare-cutting endonuclease target site 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [left end of transposase]-[coding sequence 1]-[terminator 1]-[terminator 2 RC]-[coding sequence 2 RC]-[right end of transposase]. In another embodiment, the structure of the transgene may be, from 5' to 3', [homologous arm 1]-[coding sequence]-[terminator]-[homologous arm 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [left end of transposase]-[coding sequence]-[terminator]-[right end of transposase]. In another embodiment, the structure of the transgene may be, from 5' to 3', [additional sequence 1]-[coding sequence 1]-[terminator 1]-[terminator 2 RC]-[coding sequence 2 RC]-[additional sequence 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[coding sequence 1]-[terminator 1]-[terminator 2 RC]-[coding sequence 2 RC]-[rare-cutting endonuclease target site 2].In another embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[homologous arm 1]-[coding sequence 1]-[terminator 1]-[terminator 2 RC]-[coding sequence 2 RC]-[homologous arm 2]-[rare-cutting endonuclease target site 2]. In yet another embodiment, the structure of the transgene may be, from 5' to 3', [left end of transposase]-[coding sequence 1]-[terminator 1]-[terminator 2 RC]-[coding sequence 2 RC]-[right end of transposase].

[0111] In one embodiment, this document describes a method for silencing and replacing the production of a target protein by administering the transgene shown in Figure 17 to cells and incorporating the transgene into the target endogenous gene. In one embodiment, the transgene may include an RNAi cassette, a promoter, a partial coding sequence (resistant to silencing), and a splice donor sequence designed to silence the endogenous gene. The promoter may be operably linked to a partial coding sequence that can be operably linked to a splice donor. The RNAi cassette, promoter, partial coding sequence, and splice donor may be adjacent to the first and second homologous arms or the left and right transposon ends. The transgene may be incorporated into an exon or intron within the target endogenous gene (Figure 17), but not into a site that disrupts the endogenous splice acceptor necessary for the production of the full-length protein. The RNAi may be designed to silence the expression of the target endogenous gene, and the partial coding sequence within the transgene may be designed to be resistant to silencing. Therefore, the corresponding target sites within the complete coding sequence of the transgene can be modified to prevent silencing. In other embodiments, the transgene may include first and second splice donor sequences, first and second partial coding sequences (both resistant to silencing), first and second promoters, and an RNAi cassette. These transgenes may be adjacent to additional sequences (e.g., viral ITRs), first and second rare-cutting endonuclease target sites, left and right transposon ends, or both the first and second homologous arms and the first and second rare-cutting endonuclease target sites. In one embodiment, the structure of the transgene may be 5' to 3', [homologous arm 1]-[RNAi cassette]-[promoter]-[partial coding sequence]-[splice donor]-[homologous arm 2]. In another embodiment, the structure of the transgene may be 5' to 3', [left end of transposon]-[RNAi cassette]-[promoter]-[partial coding sequence]-[splice donor]-[right end of transposon].In another embodiment, the structure of the transgene may be, from 5' to 3', [additional sequence 1]-[splice donor 1 RC]-[partial coding sequence 1 RC]-[promoter 1 RC]-[RNAi cassette]-[promoter 2]-[partial coding sequence 2]-[splice donor 2]-[additional sequence 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[splice donor 1 RC]-[partial coding sequence 1 RC]-[promoter 1 RC]-[RNAi cassette]-[promoter 2]-[partial coding sequence 2]-[splice donor 2]-[rare-cutting endonuclease target site 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [rare-cutting endonuclease target site 1]-[homologous arm 1]-[splice donor 1 RC]-[partial coding sequence 1 RC]-[promoter 1 RC]-[RNAi cassette]-[promoter 2]-[partial coding sequence 2]-[splice donor 2]-[rare-cutting endonuclease target site 2]. In another embodiment, the structure of the transgene may be, from 5' to 3', [left end of transposase]-[splice donor 1 RC]-[partial coding sequence 1 RC]-[promoter 1 RC]-[RNAi cassette]-[promoter 2]-[partial coding sequence 2]-[splice donor 2]-[right end of transposase]. The transgene can be used to modify the SNCA gene. Mutations in SNCA have been found to cause Parkinson's disease. The transgene described herein can be used to modify the gene expression of SNCA. In some cases, SNCAs are duplicated or tripled, leading to the overproduction of α-synuclein protein. In other cases, mutations such as Ala30Pro cause incorrect protein folding.The transgenes described herein provide a method for reducing endogenous SNCA expression (from gene duplication and intragenetic mutations) while replacing some or all of the expression of SNCA with SNCA isoforms (there are at least six SNCA transcripts, including full-length 140aa, 126aa, 112aa, 98aa, 67aa, and 115aa proteins). The SNCA gene contains six exons, with the start codon in exon 2. This document provides transgenes for integration into the SNCA gene. The transgenes may include an RNAi cassette targeting exon 1 or exon 2 of SNCA, a promoter, a partial coding sequence encoding a peptide produced by exon 2 of SNCA (this partial coding sequence is resistant to silencing by the RNAi cassette), and a splice donor.

[0112] In one embodiment, the method provided herein describes the delivery of a transgene having a fully functional silencing-resistant coding sequence and an RNAi silencing sequence (Figure 9). The functional coding sequence may include a promoter, a nucleic acid sequence that functions to produce an RNA or protein product, and a terminator. The nucleic acid sequence may be customized to avoid silencing by the silencing sequence (Figure 9). In one embodiment, the transgene may include a silencing sequence that targets the 5'UTR of the transcript. The functional coding sequence within the transgene may include a coding sequence for a silenced gene (either WT or codon-modified) with or without a 5'UTR, or with an alternative 5'UTR not derived from the target gene. In another embodiment, the transgene may include a silencing sequence that targets the 3'UTR of the transcript. The functional coding sequence within the transgene may include a coding sequence for a silenced gene (either WT or codon-modified) with or without a 3'UTR, or with an alternative 3'UTR not derived from the target gene. In yet another embodiment, the transgene may include a silencing sequence that targets the coding sequence of the gene. The functional coding sequence may include the coding sequence of a silencing gene, and either the entire coding sequence or a portion thereof may be modified to evade silencing by the silencing sequence. Modification may be achieved by methods such as codon optimization / codon adjustment, or by deletion of a target region. In one embodiment, the transgene described herein, comprising the silencing sequence and the functional coding sequence, may be transiently delivered to a cell (e.g., by a viral vector or plasmid DNA) or incorporated into the cell's genome. In some embodiments, the transgene may be delivered to a cell and may contain one or more genes having gain-of-function mutations (Figure 7).Examples of diseases involving gain-of-function mutations include HD (Huntington's disease), SBMA (spinal and bulbar muscular atrophy), SCA1 (spinocerebellar ataxia type 1), SCA2 (spinocerebellar ataxia type 2), SCA3 (spinocerebellar ataxia type 3 or Machado-Joseph disease), SCA6 (spinocerebellar ataxia type 6), SCA7 (spinocerebellar ataxia type 7), Fragile X syndrome, Fragile XE intellectual disability, Friedreich's ataxia, and myotonic dystrophy type 1. These include myotonic dystrophy type 2, spinocerebellar ataxia type 8, spinocerebellar ataxia type 12, spinal and bulbar muscular atrophy, JPH3, amyotrophic lateral sclerosis (ALS), hereditary motor-sensory neuropathy type IIC, postsynaptic slow channel congenital myasthenia gravis, PRPS1 hyperactivity, Parkinson's disease, tubular aggregate myopathy, achondroplasia, Rubus X-linked intellectual disability syndrome, and autosomal dominant retinitis pigmentosa.

[0113] In certain embodiments, the transgenes described herein, including silencing sequences and functional coding sequences, can be used to correct gain-of-function disorders by silencing specific genes and replacing gene expression. Genes may include SOD1, TRPV4, CHRNA1, CHRND, CHRNE, CHRNB1, PRPS1, LRRK2, STIM1, FGFR3, MECP2, SNCA, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, TBP, HTT, AR, FXN, DMPK, PABPN1, ATXN8, RHO, and C9orf72.

[0114] The transgenes described herein, including silencing sequences and functional coding sequences, can be delivered to cells using viral (e.g., AAV vectors) or nonviral methods. In certain embodiments, the AAV vectors described herein may be derived from any AAV. In certain embodiments, the AAV vectors are derived from defective and nonpathogenic adeno-associated virus type 2 of the family Parvoviridae. All such vectors are derived from plasmids that hold only the 145 bp inverted terminal repeat of the AAV adjacent to the transgene expression cassette. Efficient gene transport and stable transgene delivery by integration into the genome of transduced cells are key features of this vector system. (Wagner et al., Lancet 351:9117 1702-3, 1998; Kearns et al., Gene Ther. 9:748-55, 1996). Other AAV serotypes, including AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, and AAVrh.10, as well as any novel AAV serotypes, can also be used in accordance with the present invention. In some embodiments, chimeric AAVs are used, where the viral origin of the long-chain terminal repeat (LTR) sequence of the viral nucleic acid is heterologous to the viral origin of the capsid sequence. Non-limiting examples include chimeric viruses having an LTR derived from AAV2 and a capsid derived from AAV5, AAV6, AAV8, or AAV9 (i.e., AAV2 / 5, AAV2 / 6, AAV2 / 8, and AAV2 / 9, respectively).

[0115] Furthermore, the constructs described herein can also be incorporated into adenovirus vector systems. Adenovirus-based vectors enable very high transduction efficiency in many cell types and do not require cell division. Such vectors can be used to obtain high titers and high levels of expression.

[0116] The methods and compositions described herein are applicable to any eukaryotes in which modification of the organism through genome modification is desired. Examples of eukaryotes include plants, algae, animals, fungi, and protists. Eukaryotes may also include plant cells, algal cells, animal cells, fungal cells, and protist cells.

[0117] Examples of mammalian cells include oocytes, K562 cells, CHO (Chinese hamster ovary) cells, HEP-G2 cells, BaF-3 cells, Schneider cells, COS cells (monkey kidney cells expressing the SV40 T-antigen), CV-1 cells, HuTu80 cells, NTERA2 cells, NB4 cells, HL-60 cells, and HeLa cells, 293 cells (see, for example, Graham et al. (1977) J. Gen. Virol. 36:59), as well as myeloma cells such as SP2 or NS0 (e.g., Galfre Examples include, but are not limited to, Milstein (1981) Meth. Enzymol. 73(B):3 46. Peripheral blood mononuclear cells (PBMCs) or T cells can also be used, as can embryonic stem cells and adult stem cells. Examples of usable stem cells include embryonic stem cells (ES), induced pluripotent stem cells (iPSCs), mesenchymal stem cells, hematopoietic stem cells, liver stem cells, skin stem cells, and neural stem cells.

[0118] The methods and compositions of the present invention can be used for the production of modified organisms. Modified organisms may be small mammals, companion animals, livestock, and primates. Non-limited examples of rodents include mice, rats, hamsters, gerbils, and guinea pigs. Non-limited examples of companion animals include cats, dogs, rabbits, hedgehogs, and ferrets. Non-limited examples of livestock include horses, goats, sheep, pigs, llamas, alpacas, and cattle. Non-limited examples of primates include capuchin monkeys, chimpanzees, lemurs, macaques, marmosets, tamarins, spider monkeys, squirrel monkeys, and velvet monkeys. The methods and compositions of the present invention can be used on humans.

[0119] Exemplary plants and plant cells that can be modified using the methods described herein include, but are not limited to, monocotyledonous plants (e.g., wheat, maize, rice, millet, barley, sugarcane), dicotyledonous plants (e.g., soybeans, potatoes, tomatoes, alfalfa), fruit crops (e.g., tomatoes, apples, pears, strawberries, oranges), fodder crops (e.g., alfalfa), root vegetable crops (e.g., carrots, potatoes, sugar beets, yams), and leafy vegetable crops (e.g., lettuce, spinach). Examples include plants for consumption (e.g., soybeans and other legumes, pumpkins, peppers, eggplants, celery, etc.), flowering plants (e.g., petunias, roses, chrysanthemums), conifers and pines (e.g., pine, fir, spruce), poplar trees (e.g., P. tremula × P. alba), fiber crops (cotton, jute, flax, bamboo), plants used in phytomediation (e.g., heavy metal accumulating plants), oil crops (e.g., sunflowers, rapeseed), and plants used for experimental purposes (e.g., Arabidopsis thaliana). The methods disclosed herein can be used within the genera Asparagus, Avena, Brassica, Citrus, Melon, Capsicum, Pumpkin, Carrot, Artemisia, Glycine, Cotton, Barley, Lactuca, Rhus, Tomato, Apple, Cassava, Tobacco, Potato, Poria, Rice, Pear, Pea, Pear, Prunus, Radish, Rye, Solanum, Sorghum, Wheat, Vitis, Vigna, and Maize. The term plant cell includes isolated plant cells, as well as whole plants or parts of whole plants such as seeds, callus, leaves, and roots. The disclosure also includes seeds of the plants described above, and the seeds are modified using the compositions and / or methods described herein. This disclosure further encompasses the offspring, clones, cell lines, or cells of the transgenic plants described above, the offspring, clones, cell lines, or cells having a transgene or genetic construct.Exemplary algal species include microalgae, diatoms, Botryococcus braunii, Chlorella, Dunaliella tertiolecta, Gracileria, Pleurochrysis carterae, Sorghum, and Ulva.

[0120] The methods described herein may include the use of rare-cutting endonucleases to stimulate homologous or non-homologous integration of the transgene molecule into endogenous genes. Rare-cutting endonucleases may include CRISPR, TALEN, or zinc finger nucleases (ZFNs). CRISPR systems may include CRISPR / Cas9 or CRISPR / Cas12a (Cpf1). CRISPR systems may include variants exhibiting broad PAM capabilities (Hu et al., Nature 556, 57-63, 2018; Nishimasu et al., Science DOI:10.1126, 2018) or variants exhibiting higher on-target binding or cleavage activity (Kleinstiver et al., Nature 529:490-495, 2016). Gene editing reagents include nucleases (Mali et al., Science 339:823-826, 2013; Christian et al., Genetics 186:757-761, 2010) and nicasses (Cong et al., Science 339:819-823, 2013; Wu et al., Biochemical and Biophysical Research). Communications 1:261-266, 2014), CRISPR-FokI dimer (Tsai et al., Nature Biotechnology 32:569-576, 2014), or CRISPR-nickase pair (Ran et al.) It may be in the form of al., Cell 154:1380-1389, 2013).

[0121] The methods and compositions described herein can be used in situations where it is desired to modify the 5' end of the coding sequence of an endogenous gene. For example, a patient with SCA2 has an extended CAG repeat in exon 1. A patient with SCA2 can benefit from the replacement of exon 1. In another example, a patient with a hereditary disorder resulting from a loss-of-function mutation in the 5' end of an endogenous gene can benefit from the replacement of the first exon of that gene.

[0122] Furthermore, the methods and compositions described herein can be used in situations where it is desired to treat a gain-of-function genetic disorder while ensuring that wild-type protein is still produced. For example, a patient with retinitis pigmentosa who has a gain-of-function mutation in the RHO gene can benefit from a therapy containing a transgene that can silence the endogenous RHO gene and simultaneously produce wild-type RHO protein. A further advantage of this approach is the ability to select target sites for silencing that are not centered on the gain-of-function mutation site. This advantage allows for the design of effective silencing constructs (e.g., low off-targeting and highly effective on-targeting) and enables the design of monotherapies for patients with gain-of-function mutations in different regions of the RHO gene. Moreover, this method may be particularly useful in Parkinson's disease and gain-of-function disorders caused by genes that produce multiple isoforms, including SNCA. Cells with a gain-of-function mutation at the 5' end of the SNCA gene can benefit from the incorporation of an RNAi cassette targeting exon 2 with a transgene containing a promoter and a subcoding sequence resistant to RNAi silencing.

[0123] The present invention is further illustrated by the following embodiments, which do not limit the scope of the invention as described in the claims. [Examples]

[0124] Example 1: Targeted incorporation of DNA in the ATXN2 gene Three plasmids were constructed using transgenes designed to be integrated into the ATXN2 gene in human cells. All transgenes were designed to be integrated into intron 2 of the ATXN2 gene, and all transgenes were designed to insert a bidirectional partial coding sequence along with its individual promoter. The partial coding sequence encodes a peptide produced by exon 1 of the ATXN2 gene. The first plasmid (designated pBA1141) contained left and right homologous arms with sequences homologous to the start of intron 1 (i.e., successful gene targeting results in the insertion of the pBA1141 cargo into intron 1). Between the homologous arms, from 5' to 3', were a reverse-complementary orientation splice donor, a reverse-complementary orientation partial coding sequence 1 with codon adjustment (encoding a peptide produced by exon 1 of the ATXN2 gene), a reverse-complementary orientation EF1α promoter, a CMV promoter, a reverse-complementary orientation partial coding sequence 2 with codon adjustment (encoding a peptide produced by exon 1 of the ATXN2 gene), and a splice donor. The sequence of the pBA1141 transgene is shown in SEQ ID NO: 15 (Figure 6). Two nucleases were designed to facilitate the integration of pBA1141 into the genome: Cas9 with the target site (TGTGCAGGAGGGCCTGTTGGGGG, SEQ ID NO: 16), and Cas12a with the target site (TTTCCCTTGTGCCTCAAGTCCATCCGT, SEQ ID NO: 17). The target sites were also included in pBA1141 to facilitate the release of the donor molecule from the plasmid. The individual components within pBA1141 are shown in SEQ ID NOs. 18 is a sequence containing target sites for both Cas9 and Cas12a. SEQ ID NOs. 19 contains the sequence of the left homologous arm. SEQ ID NOs. 20 contains the reverse complementary codon-modifying partial coding sequence (exon 1) of the non-pathogenic ATXN2 gene. SEQ ID NOs. 21 contains the reverse complementary EF1α promoter. SEQ ID NOs. 22 contains the reverse complementary CMV promoter. SEQ ID NOs. 23 contains the codon-modifying partial coding sequence (exon 1) of the non-pathogenic ATXN2 gene. SEQ ID NOs. 24 contains the sequence of the right homologous arm. The second plasmid (referred to as pBA1142) contains the same cargo as pBA1135, but the homologous arm has been removed.The nuclease target site was preserved, promoting the release of the transgene from the plasmid. Successful cleavage of the plasmid was expected to release the transgene, making the sequence available for integration into the ATXN2 gene by NHEJ. The sequence of pBA1141 is shown in Sequence ID No. 25. The third plasmid (referred to as pBA1143) contained the same sequence as pBA1141, except that the sequence containing the nuclease target site (upstream of the left homologous arm) was removed and the right homologous arm was shortened to 600 bp.

[0125] HEK293T cells were used for transfection. HEK293T cells were maintained in DMEM supplemented with 10% fetal bovine serum (FBS) at 37°C and 5% CO2. HEK293T cells were transfected with 2 µg of donor, 2 µg of guide RNA (RNA form), and 2 µg of Cas9 (RNA form) or 2 µg of Cas12a plasmid (DNA form). Transfection was performed using electroporation. 72 hours after transfection, genomic DNA was isolated and integration events were evaluated. A list of primers used to detect integration or genomic DNA is shown in Table 1. [Table 1-1] [Table 1-2]

[0126] PCR was performed on genomic DNA to detect the integration of pBA1141, pBA1142, and pBA1143. For pBA1143, the transgene was designed to be precisely integrated via HR. Therefore, a band was detected by 3' junction PCR for both Cas9 and Cas12a transfected samples, indicating precise insertion into intron 1 (Figure 17, lanes 7-10). Expected band sizes were 1,225 bp (lanes 7 and 9) and 1,407 bp (lanes 8 and 10). Primers oNJB201+oNJB190 and oNJB202+oNJB191 were used for 3' junction PCR. For pBA1142, since no homologous arm was present, the transgene was predicted to be inserted via NHEJ insertion. NHEJ integration in the Cas9 transfected sample can be seen in lane 6 of Figure 17. The expected band size was 813 bp. Primers oNJB202+oNJB211 were used for PCR of the NHEJ insertion 3' junction. For pBA1141, both homologous arms and nuclease cleavage sites were present on the transgene (Figure 7). HR-mediated integration was observed in lanes 2-4 of Figure 17, and NHEJ-mediated integration was observed in lane 5 of Figure 17. The expected size for detecting HR-mediated insertion by PCR was 1594 bp (lane 2, primers oNJB201+oNJB190), 1775 bp (lane 3, primers oNJB202+oNJB191), and 1775 bp (lane 4, primers oNJB202+oNJB191). The expected size for detecting NHEJ-mediated insertion by PCR was 2067 bp (lane 5, primers oNJB202+oNJB211).

[0127] The results indicate that the described transgene, which contains a bidirectional partial coding sequence with a promoter, can be incorporated into genomic DNA via multiple different repair pathways.

[0128] HEK293T cells were used for transfection. HEK293T cells were maintained in DMEM supplemented with 10% fetal bovine serum (FBS) at 37°C and 5% CO2. HEK293T cells were transfected with 2 µg of donor, 2 µg of guide RNA (RNA form), and 2 µg of Cas9 (RNA form) or 2 µg of Cas12a plasmid (DNA form). Transfection was performed using electroporation. Single-cell clones containing the transfected cells were isolated, and RNA was extracted. New transcripts could be detected using RNA sequencing.

[0129] Example 2: Silencing of endogenous SOD1 gene expression and expression of substitution SOD1 protein This paper describes methods for using RNAi, RNAi resistance coding sequences, and gene editing to silence and replace the expression of endogenous genes. These methods are particularly useful for gain-of-function disorders, including amyotrophic lateral sclerosis with SOD1 gene mutations.

[0130] To validate gene silencing and substitution, a transgene was designed using an RNAi (shRNA) cassette target sequence within exon 2 of SOD1. The shRNA contained the sequence GGCCTGCATGGATTCCATGTTCAAGAGACATGGAATCCATGCAGGCC (SEQ ID NO: 49) and was positioned downstream of the U6 promoter. The transgene also included the SOD1 coding sequence downstream of the CMV promoter. Sequences within the coding sequence were modified to avoid shRNA silencing. The sequence of the transgene (referred to as pBA1148) is shown in SEQ ID NO: 10. A control vector was generated containing scrambled shRNA (referred to as pBA1147, SEQ ID NO: 53) and the coding sequence of WT SOD1 (referred to as pBA1149, SEQ ID NO: 54).

[0131] HEK293T cells were used for transfection. HEK293T cells were maintained in DMEM supplemented with 10% fetal bovine serum (FBS) at 37°C and 5% CO2. HEK293T cells were transfected with 2 ug of plasmid. Transfection was performed using electroporation. 48 hours after transfection, RNA was isolated and SOD1 mRNA levels were evaluated.

[0132] Two vectors are designed to be incorporated into intron 1 to silence the expression of the SOD1 gene using gene editing and produce a substitution SOD1 protein. The first vector contains, from 5' to 3', a left homologous arm, a splice acceptor, a partial coding sequence of SOD1 encoding the peptide produced by exons 2-5 (including mutations to avoid silencing by the RNAi cassette), a terminator, an RNAi cassette with the shRNA sequence shown in SEQ ID NO: 49, and a right homologous arm. The second vector contains, from 5' to 3', a nuclease target site, a splice acceptor, a partial coding sequence of SOD1 encoding the peptide produced by exons 2-5 (including mutations to avoid silencing by the RNAi cassette), a terminator, an RNAi cassette having the shRNA sequence shown in SEQ ID NO: 49, a reverse-complementary oriented second terminator, a reverse-complementary oriented second partial coding sequence of SOD1 encoding the peptide produced by exons 2-5 (including mutations to avoid silencing by the RNAi cassette), a reverse-complementary oriented second splice acceptor, and a second nuclease target site (Figure 12).

[0133] Two additional vectors are designed to be incorporated into intron 3 of the SOD1 gene. The first vector contains, from 5' to 3', a left homologous arm, an RNAi cassette with the shRNA sequence shown in SEQ ID NO: 49, a promoter, a partial coding sequence of SOD1 encoding the peptide produced by exons 1 and 2 (including mutations to avoid silencing by the RNAi cassette), a splice donor, and a right homologous arm. The second vector contains, from 5' to 3', a nuclease target site, a reverse-complementarily oriented splice donor, a reverse-complementarily oriented partial coding sequence of SOD1 encoding peptides produced by exons 1 and 2 (including mutations to avoid silencing by the RNAi cassette), a reverse-complementarily oriented promoter, an RNAi cassette having the shRNA sequence shown in SEQ ID NO: 49, a second promoter, a second partial coding sequence of SOD1 encoding peptides produced by exons 1 and 2 (including mutations to avoid silencing by the RNAi cassette), a splice donor, and a second nuclease target site (Figure 16).

[0134] HEK293T cells are used for transfection. HEK293T cells are maintained in DMEM supplemented with 10% fetal bovine serum (FBS) at 37°C and 5% CO2. HEK293T cells are transfected with 2 µg of plasmid, 2 µg of guide RNA (RNA form), and 2 µg of Cas9 (RNA form). Transfection is performed using electroporation. 72 hours after transfection, DNA is isolated and the integration of the transgene is evaluated. Clones containing the integration event are isolated and the mRNA levels of SOD1 (both from endogenous and modified genes) are evaluated.

[0135] Example 3: Silencing of endogenous SNCA gene expression and expression of two SNCA protein isoforms Mutations in SNCAs have been found to cause Parkinson's disease. The methods described herein can be used to modify the gene expression of SNCAs. In some cases, SNCAs are duplicated or tripled, resulting in the overproduction of α-synuclein protein. In other cases, mutations such as Ala30Pro cause incorrect protein folding. Methods for reducing the expression of endogenous SNCAs (from gene duplication and intragenetic mutations) while replacing some or all of SNCAs and SNCA isoforms (at least six present SNCA transcripts, including full-length 140aa, 126aa, 112aa, 98aa, 67aa, and 115aa proteins) are described herein.

[0136] The transgene was designed to contain shRNA to silence the expression of the endogenous SNCA gene. The transgene was also designed to replace two SNCA protein isoforms, each encoding two open reading frames for each isoform. The shRNA contains a 19nt hairpin sequence targeting the 3' end of the SNCA coding sequence (GGTATCAAGACTACGAAC, SEQ ID NO: 11). The two SNCA open reading frames within the transgene were designed to contain mutations at the shRNA target sites. SEQ ID NO: 12 shows the nucleic acid sequence of the transgene cloned into an expression plasmid (referred to as pBA1153). Two other transgenes were constructed: the first contains shRNA and two wild-type SNCA isoforms (without mutations that block shRNA silencing), and the second contains scrambled shRNA and two SNCA isoforms with mutations.

[0137] The transgene is transfected into HEK293 cells. HEK293 cells are maintained at 37°C and 5% CO2 in DMEM high-glucose medium without L-glutamine and sodium pyruvate, supplemented with 10% fetal bovine serum (FBS) and 1% penicillin streptomycin (PS) solution (×100). HEK293 cells are transfected with each of the plasmid constructs and their combinations using Lipofectamine 3000. 48 hours after transfection, RNA is extracted and the transcription level of SNCA is evaluated. A reduction in the expression of endogenous SNCA RNA and the expression of RNA from codon-modified SNCA sequences indicate the functionality of the transgene.

[0138] Two vectors are designed to be incorporated into the exon 2-intron 2 junction to silence SNCA gene expression using gene editing and produce a replacement SNCA protein while maintaining isoform production. The first vector contains, from 5' to 3', a left homologous arm, an RNAi cassette with an shRNA sequence targeting the exon 2 transcription sequence, a promoter (containing a 1,000 bp endogenous SNCA promoter), a start codon and a partial coding sequence encoding the peptide produced by exon 2 of the endogenous SNCA gene (including mutations to circumvent silencing by the RNAi cassette), a splice donor, and a right homologous arm. The splice donor and the right homologous arm are sequences derived from the 5' end of endogenous intron 2. The second vector contains, from 5' to 3', a nuclease target site, a reverse-complementarily oriented splice donor, a reverse-complementarily oriented partial coding sequence of the SNCA encoding the peptide produced by exon 2 (including mutations to avoid silencing by the RNAi cassette), a reverse-complementarily oriented promoter, an RNAi cassette with shRNA targeting exon 2, a second promoter, a second partial coding sequence of the SNCA encoding the peptide produced by exon 2 (including mutations to avoid silencing by the RNAi cassette), a splice donor, and a second nuclease target site (Figure 16). The splice donor sequence is a splice donor sequence derived from intron 2 of the SNCA gene. The nuclease is designed to facilitate the integration of the transgene into the exon 2-intron 2 junction.

[0139] The transgene and nuclease are transfected into HEK293 cells. HEK293 cells are maintained at 37°C and 5% CO2 in DMEM high-glucose medium, free of L-glutamine and sodium pyruvate, supplemented with 10% fetal bovine serum (FBS) and 1% penicillin streptomycin (PS) solution (×100). HEK293 cells are transfected with each of the plasmid constructs and their combinations using Lipofectamine 3000. Clones containing the integration event are isolated, and RNA is extracted. Reduced expression of endogenous SNCA RNA and increased expression of RNA from the modified SNCA gene indicate the functionality of the transgene.

[0140] Example 4: Silencing of endogenous RHO gene expression and expression of substitution RHO protein The transgene is designed to contain an shRNA that silences the expression of the endogenous RHO gene, and an open reading frame encoding the wild-type RHO protein. The sequence of the RHO protein is shown in SEQ ID NO: 13. The silencing sequence contains a hairpin sequence that targets the endogenous RHO transcript. The RHO open reading frame within the transgene is codon-tuned to contain minimal sequence homology at the shRNA target site.

[0141] The transgene is transfected into HEK293 cells. HEK293 cells are maintained at 37°C and 5% CO2 in DMEM high-glucose medium without L-glutamine and sodium pyruvate, supplemented with 10% fetal bovine serum (FBS) and 1% penicillin streptomycin (PS) solution (×100). HEK293 cells are transfected with each of the plasmid constructs and their combinations using Lipofectamine 3000. Three days after transfection, RNA is extracted from the cells and transcript levels are assessed. Reduced expression of endogenous RHO RNA and expression of RNA from codon-modified RHO sequences indicate the functionality of the transgene.

[0142] Example 5: Silencing of endogenous C9orf72 gene expression and expression of substituted C9orf72 protein The transgene is designed to contain an shRNA that silences the expression of the endogenous C9orf72 gene, and an open reading frame encoding the wild-type C9orf72 protein. The sequence of the C9orf72 protein is shown in SEQ ID NO: 14. The silencing sequence contains a hairpin sequence that targets the endogenous C9orf72 transcript. The open reading frame of C9orf72 within the transgene is codon-tuned to contain minimal sequence homology at the target site of the shRNA.

[0143] The transgene is transfected into HEK293 cells. HEK293 cells are maintained at 37°C and 5% CO2 in DMEM high-glucose medium without L-glutamine and sodium pyruvate, supplemented with 10% fetal bovine serum (FBS) and 1% penicillin streptomycin (PS) solution (×100). HEK293 cells are transfected with each of the plasmid constructs and their combinations using Lipofectamine 3000. Three days after transfection, RNA is extracted from the cells and transcript levels are assessed. Reduced expression of endogenous C9orf72 RNA and expression of codon-modified C9orf72 sequences indicate the functionality of the transgene.

[0144] Example 6: Targeted incorporation of DNA in the ATXN2 gene A transgene targeting ATXN2 is designed to replace the 5' end of the ATXN2 coding sequence. The plasmid, designated pBA1012-D1, is constructed with a transgene designed to incorporate the WT coding sequence into intron 1 of the ATXN2 gene (Figure 4). The transgene contains a first homologous arm homologous to the sequence, following the splice donor site in intron 1 (sequence number 2). Adjacent to the first homologous arm is the target site of the Cas9 nuclease. Following the first homologous arm are the reverse complementary splice donor sequence and exon 1 of the ATXN2 gene (non-extended CAG repeat sequence, sequence number 3). Following the first coding sequence is the EF1α promoter (sequence number 4). In head-to-head orientation, a second set of functional elements is present. The initiation of the elements in the second set includes a CMV promoter (sequence number 5) that drives the expression of the coding sequence of codon-modified exon 1 of the ATXN2 gene (sequence number 6). The coding sequence is followed by a splice donor site and a second homologous arm. The second homologous arm contains a rare-cutting endonuclease target site (SEQ ID NO: 8). The sequence of the transgene is shown in SEQ ID NO: 1.

[0145] The corresponding Cas9 nuclease is designed to generate three double-strand breaks: 1) within intron 1 of the endogenous ATXN2 gene, 2) adjacent to the first homologous arm in the pBA1012-D1 transgene, and 3) within the second homologous arm in the pBA1012-D1 transgene. The target sequence of the Cas9 nuclease is shown in Sequence ID No. 8.

[0146] Functional confirmation of the transgene and CRISPR vector is achieved by transfection of HEK293 cells. HEK293 cells are maintained at 37°C and 5% CO2 in DMEM high-glucose medium without L-glutamine and sodium pyruvate, supplemented with 10% fetal bovine serum (FBS) and 1% penicillin streptomycin (PS) solution (×100). HEK293 cells are transfected with each of the plasmid constructs and their combinations using Lipofectamine 3000. Two days after transfection, DNA is extracted and evaluated for mutations and targeted insertions in the ATXN2 gene. Nuclease activity is analyzed using the Cel-I assay or by deep sequencing of the amplicon containing the CRISPR / Cas9 target sequence. Successful transgene integration is analyzed using PCR.

[0147] Other Embodiments The present invention is described in conjunction with its detailed description, but it should be understood that the above description is intended to illustrate, not limit, the scope of the invention as defined by the appended claims. Other aspects, advantages, and modifications are within the following claims.

Claims

1. A combination comprising a transgene and at least one rare-cutting endonuclease that targets an intron in an endogenous gene, wherein the combination is for incorporating the transgene into the endogenous gene. The aforementioned introduced gene is located at 5' to 3': i. Splice acceptor sequence and ii. Partial code arrays and, iii. Terminator and, iv. One RNA interference cassette for silencing the expression of the endogenous gene and The introduced gene is incorporated into the endogenous gene, the partial coding sequence is operably linked to the promoter of the endogenous gene, and is resistant to silencing by the RNA interference cassette. A combination of items.

2. The combination according to claim 1, wherein the rare-cutting end nuclease is a CRISPR nuclease, a TAL effector nuclease, a zinc finger nuclease, or a meganuclease.

3. The combination according to claim 2, wherein the rare-cutting end nuclease is a CRISPR nuclease.

4. The combination according to claim 3, wherein the CRISPR nuclease is CRISPR / Cas12a nuclease or CRISPR / Cas9 nuclease.

5. The combination according to claim 1, wherein the endogenous gene is SOD1, TRPV4, CHRNA1, CHRND, CHRNE, CHRNB1, PRPS1, LRRK2, STIM1, FGFR3, MECP2, SNCA, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, TBP, HTT, AR, FXN, DMPK, PABPN1, ATXN8, RHO, or C9orf72.