CRISPR and AAV strategies for the treatment of X-linked juvenile retinoschisis

Nucleic acid constructs and AAV vectors are used to integrate and express retinoschisin in the RS1 locus, addressing the limitations of current gene therapies for XLRS by restoring functional retinoschisin protein expression and reducing mutant RS1 gene expression.

JP7756639B2Active Publication Date: 2025-10-20REGENERON PHARMACEUTICALS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022524995
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-08
Filing Date
2020-11-07
Publication Date
2025-10-20
Estimated Expiration
2040-11-07

AI Technical Summary

Technical Problem

Current gene therapy approaches for X-linked juvenile retinoschisis (XLRS) have not met their endpoints, highlighting the need for novel strategies to treat this condition effectively.

Method used

Nucleic acid constructs and compositions are developed to integrate a retinoschisin-encoding sequence into the endogenous RS1 locus, utilizing bidirectional constructs and vectors like AAV to express functional retinoschisin protein, potentially replacing nonfunctional protein and addressing mutations causing XLRS.

Benefits of technology

These constructs and compositions aim to restore functional retinoschisin expression, thereby reducing or eliminating expression of mutant RS1 genes, offering a potential therapeutic approach for XLRS.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007756639000087
    Figure 0007756639000087
  • Figure 0007756639000088
    Figure 0007756639000088
  • Figure 0007756639000089
    Figure 0007756639000089
Patent Text Reader

Abstract

Nucleic acid constructs and compositions are provided that allow for the insertion and / or expression of a retinoschisin coding sequence. Nuclease agents that target the RS1 locus are provided. Compositions and methods are also provided that use such constructs for integration into a target genomic locus and / or expression in cells. Methods for treating X-linked juvenile retinoschisis using nucleic acid constructs and compositions are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Application No. 62 / 932,608, filed November 8, 2019, which is incorporated herein by reference in its entirety for all purposes.

[0002] Reference to a sequence listing submitted as a text file via EFS WEB The sequence listing set forth in file 694232SEQLIST.txt is 1.22 MB, was created on November 6, 2020, and is incorporated herein by reference. [Background technology]

[0003] The RS1 gene encodes a highly conserved extracellular protein involved in retinal cellular organization. The RS1 gene is assembled and secreted from photoreceptors and bipolar cells as a homo-oligomeric protein complex. Over 200 mutations have been detected in RS1, many of which lead to early-onset macular degeneration due to nonfunctional protein or lack of protein secretion. Lack of functional Rs1 expression causes separation within the retinal layers, leading to the early and progressive vision loss associated with X-linked juvenile retinoschisis (XLRS). Clinical trials of gene therapy for XLRS are ongoing, but the trials have not met their endpoints. Novel strategies are needed to treat XLRS. Summary of the Invention [Means for solving the problem]

[0004] Nucleic acid constructs and compositions are provided that allow the insertion of a retinoschisin-encoding sequence into a target genomic locus, such as the endogenous RS1 locus, and / or the expression of a retinoschisin-encoding sequence.The nucleic acid constructs and compositions can be used in methods for integration into a target genomic locus and / or expression in cells, or in methods for treating X-linked juvenile retinoschisis.

[0005] In one embodiment, a bidirectional nucleic acid construct for integration into a target genomic locus is provided. Some such nucleic acid constructs include (a) a first segment comprising a first coding sequence of a first retinoschisin protein or a fragment thereof, and (b) a second segment comprising a reverse complement of a second coding sequence of a second retinoschisin protein or a fragment thereof. In some such constructs, the second segment is located 3' (i.e., downstream) of the first segment.

[0006] In some such constructs, the first retinoschisin protein or fragment thereof is a human retinoschisin protein or fragment thereof, the second retinoschisin protein or fragment thereof is a human retinoschisin protein or fragment thereof, or both the first retinoschisin protein or fragment thereof and the second retinoschisin protein or fragment thereof are human retinoschisin protein or fragment thereof. In some such constructs, the first coding sequence comprises, consists essentially of, or consists of complementary DNA (cDNA), the second coding sequence comprises, consists essentially of, or consists of cDNA, or both the first coding sequence and the second coding sequence comprise, consist essentially of, or consist of cDNA. In some such constructs, the first coding sequence comprises, consists essentially of, or consists of exons 2-6 of human RS1 or a degenerate variant thereof, the second coding sequence comprises, consists essentially of, or consists of exons 2-6 of human RS1 or a degenerate variant thereof, or both the first and second coding sequences comprise, consist essentially of, or consist of exons 2-6 of human RS1 or a degenerate variant thereof.

[0007] In some such constructs, the first segment comprises a fragment or portion of the first intron of human RS1 located 5' (i.e., upstream) of the first coding sequence, and / or the second segment comprises the reverse complement of a fragment or portion of the second intron of human RS1 located 3' (i.e., downstream) of the reverse complement of the second coding sequence.

[0008] In some such constructs, the first retinoschisin protein or fragment thereof is identical to the second retinoschisin protein or fragment thereof. In some such constructs, the second coding sequence employs codon usage that differs from the codon usage of the first coding sequence. In some such constructs, the second segment has at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, or at least about 99% complementarity to the first segment. In some such constructs, the second segment has less than about 30%, less than about 35%, less than about 40%, less than about 45%, less than about 50%, less than about 55%, less than about 60%, less than about 65%, less than about 70%, less than about 75%, less than about 80%, less than about 85%, less than about 90%, less than about 95%, less than about 97%, or less than about 99% complementarity to the first segment. In some such constructs, the reverse complement of the second coding sequence is (a) not substantially complementary to the first coding sequence, (b) not substantially complementary to a fragment of the first coding sequence, (c) highly complementary to the first coding sequence, (d) highly complementary to a fragment of the first coding sequence, (e) at least about 60%, at least about 70%, at least about 80%, or at least about 90% identical to the reverse complement of the first coding sequence, (f) about 50% to about 80% identical to the reverse complement of the first coding sequence, or (g) about 60% to about 100% identical to the reverse complement of the first coding sequence.

[0009] In some such constructs, the first segment is connected to the second segment by a linker. Optionally, the linker is about 5 to about 2000 nucleotides in length.

[0010] In some such constructs, the first segment comprises a first polyadenylation signal sequence located 3' of a first coding sequence, and the second segment comprises the reverse complement of a second polyadenylation signal sequence located 5' of the reverse complement of a second coding sequence. Optionally, the first polyadenylation signal sequence is different from the second polyadenylation signal sequence.

[0011] In some such constructs, the nucleic acid construct does not include a promoter driving expression of the first retinoschisin protein or fragment thereof or the second retinoschisin protein or fragment thereof. In some such constructs, the first segment includes a first splice acceptor site located 5' of the first coding sequence, and the second segment includes a reverse complement of a second splice acceptor site located 3' of the reverse complement of the second coding sequence. Optionally, the first splice acceptor site is derived from the RS1 gene, the second splice acceptor site is derived from the RS1 gene, or both the first and second splice acceptor sites are derived from the RS1 gene. Optionally, the first splice acceptor site is derived from intron 1 of human RS1, the second splice acceptor site is derived from intron 1 of human RS1, or both the first acceptor site and the second splice acceptor site are derived from intron 1 of human RS1.

[0012] In some such constructs, the nucleic acid construct does not comprise homology arms. In some such constructs, the nucleic acid construct comprises homology arms. In some such constructs, the nucleic acid construct is single-stranded. In some such constructs, the nucleic acid construct is double-stranded. In some such constructs, the nucleic acid construct comprises DNA.

[0013] In some such constructs, the first coding sequence is codon-optimized for expression in a host cell, the second coding sequence is codon-optimized for expression in a host cell, or both the first coding sequence and the second coding sequence are codon-optimized for expression in a host cell. In some such constructs, the nucleic acid construct comprises one or more of the following terminal structures: a hairpin, a loop, an inverted terminal repeat (ITR), or a toroid. Optionally, the nucleic acid construct comprises an ITR.

[0014] In some such constructs, the first retinoschisin protein or fragment thereof and / or the second retinoschisin protein or fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 5. In some such constructs, the first coding sequence and / or the second coding sequence comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 8 or 9. In some such constructs, the first coding sequence comprises, consists essentially of, or consists of a sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO:8, and the second coding sequence comprises, consists essentially of, or consists of a sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO:9. In some such constructs, the nucleic acid construct comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 46 or 47.

[0015] In some such constructs, the second segment is located 3' to the first segment, both the first retinoschisin protein or fragment thereof and the second retinoschisin protein or fragment thereof are human retinoschisin protein or fragment thereof, the first retinoschisin protein or fragment thereof is identical to the second retinoschisin protein or fragment thereof, both the first coding sequence and the second coding sequence comprise complementary DNA (cDNA) containing exons 2-6 of human RS1 or a degenerate variant thereof, the second coding sequence employs codon usage different from that of the first coding sequence, and the first segment is , a first polyadenylation signal sequence located 3' to the first coding sequence, and the second segment comprises the reverse complement of a second polyadenylation signal sequence located 5' to the reverse complement of the second coding sequence, the first segment comprises a first splice acceptor site located 5' to the first coding sequence, and the second segment comprises the reverse complement of a second splice acceptor site located 3' to the reverse complement of the second coding sequence, the nucleic acid construct does not comprise a promoter driving expression of the first retinoschisin protein or fragment thereof or the second retinoschisin protein or fragment thereof, and the nucleic acid construct does not comprise homology arms.

[0016] In another aspect, a vector is provided that includes any of the above bidirectional nucleic acid constructs. Some such vectors are viral vectors. Optionally, the vector is an adeno-associated virus (AAV) vector. Optionally, the AAV includes a single-stranded genome (ssAAV). Optionally, the AAV includes a self-complementary genome (scAAV). Optionally, the AAV is selected from the group consisting of AAV2, AAV5, AAV8, or AAV7m8.

[0017] Some such vectors do not include a promoter driving expression of the first retinoschisin protein or fragment thereof or the second retinoschisin protein or fragment thereof. Some such vectors do not include homology arms. Some such vectors include homology arms.

[0018] In another aspect, a lipid nanoparticle comprising any of the above bidirectional nucleic acid constructs is provided.

[0019] In another aspect, a cell is provided comprising any of the above bidirectional nucleic acid constructs. Some such cells are in vitro. Some such cells are in vivo. Some such cells are mammalian cells. Some such cells are human cells. Some such cells are retinal cells.

[0020] Some such cells express a first retinoschisin protein or fragment thereof or a second retinoschisin protein or fragment thereof. In some such cells, a nucleic acid construct is integrated into the genome at a target genomic locus. In some such cells, the target genomic locus is the endogenous RS1 locus. Optionally, the nucleic acid construct is integrated into the genome into intron 1 of the endogenous RS1 locus. Optionally, endogenous RS1 exon 1 splices into the first coding sequence or the second coding sequence of the nucleic acid construct. Optionally, the modified RS1 locus comprising the nucleic acid construct integrated into the genome encodes a protein comprising, consisting essentially of, or consisting of a sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 2 or 4.

[0021] In some such cells, integration of the nucleic acid construct into the endogenous RS1 locus prevents transcription of the endogenous RS1 gene downstream of the integration site. Optionally, integration of the nucleic acid construct into the endogenous RS1 locus reduces or eliminates expression of the endogenous retinoschisin protein and replaces it with expression of the retinoschisin protein or a fragment thereof encoded by the nucleic acid construct. Optionally, the endogenous RS1 locus contains a mutant RS1 gene containing a mutation that causes X-linked juvenile retinoschisis, and expression of the nucleic acid construct integrated into the genome reduces or eliminates expression of the mutant RS1 gene.

[0022] In another aspect, nucleic acid constructs are provided for homology-independent targeted integration into a target genomic locus. Some such nucleic acid constructs comprise a coding sequence for a retinoschisin protein or a fragment thereof flanked on each side by a nuclease target sequence for a nuclease agent. Nucleic acid constructs for homologous recombination with a target locus are also provided. Some such nucleic acid constructs comprise a coding sequence for a retinoschisin protein or a fragment thereof flanked on each side by homology arms, and optionally, the coding sequence and homology arms are further flanked on each side by a target sequence for the nuclease agent. Optionally, each homology arm is from about 25 nucleotides to about 2.5 kb in length.

[0023] In some such constructs for homology-independent targeted integration, the nuclease target sequence in the nucleic acid construct is identical to the nuclease target sequence for integration into the target genomic locus, and the nuclease target sequence in the target genomic locus is destroyed when the nucleic acid construct is inserted in the correct orientation, but is reformed when the nucleic acid construct is inserted into the target genomic locus in the opposite orientation.

[0024] In some such constructs, the retinoschisin protein or fragment thereof is a human retinoschisin protein or fragment thereof. In some such constructs, the coding sequence for the retinoschisin protein or fragment thereof comprises, consists essentially of, or consists of complementary DNA (cDNA). In some such constructs, the coding sequence for the retinoschisin protein or fragment thereof comprises, consists essentially of, or consists of exons 2-6 of human RS1 or a degenerate variant thereof.

[0025] In some such constructs, the nucleic acid construct comprises a fragment or portion of the first intron of human RS1 located 5' to the coding sequence. In some such constructs, the nucleic acid construct does not comprise a promoter driving expression of a retinoschisin protein or fragment thereof. In some such constructs, the nucleic acid construct comprises a polyadenylation signal sequence located 3' to the coding sequence. In some such constructs, the nucleic acid construct comprises a splice acceptor site located 5' to the coding sequence. Optionally, the splice acceptor site is derived from the RS1 gene. Optionally, the splice acceptor site is derived from intron 1 of human RS1.

[0026] Some such constructs are single-stranded. Some such constructs are double-stranded. Some such constructs comprise DNA. In some such constructs, the coding sequence is codon-optimized for expression in a host cell.

[0027] In some such constructs, the construct comprises one or more of the following terminal structures: a hairpin, a loop, an inverted terminal repeat (ITR), or a toroid. Optionally, the nucleic acid construct comprising the coding sequence and the nuclease target sequence is flanked by ITRs.

[0028] In some such constructs, the nuclease agent is a Cas protein and a guide RNA, and the nuclease target sequence is a guide RNA target sequence. Optionally, the guide RNA target sequence is a reverse guide RNA target sequence. Optionally, the Cas protein is Cas9.

[0029] In some such constructs, the retinoschisin protein or fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 5. In some such constructs, the coding sequence for the retinoschisin protein or fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 8 or 9. In some such constructs, the nucleic acid construct comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO:45.

[0030] In some such constructs, the nucleic acid construct is for homology-independent targeted integration into a target genomic locus, the retinoschisin protein or fragment thereof is human retinoschisin protein or fragment thereof, the coding sequence of the retinoschisin protein or fragment thereof comprises complementary DNA (cDNA) comprising exons 2 to 6 of human RS1 or a degenerate variant thereof, the nucleic acid construct does not comprise a promoter driving expression of the retinoschisin protein or fragment thereof, the nucleic acid construct comprises a polyadenylation signal sequence located 3' of the coding sequence, the nucleic acid construct comprises a splice acceptor site located 5' of the coding sequence, the nuclease target sequence in the nucleic acid construct is identical to the nuclease target sequence for integration into the target genomic locus, and the nuclease target sequence in the target genomic locus is destroyed when the nucleic acid construct is inserted in the correct orientation but is re-formed when the nucleic acid construct is inserted into the target genomic locus in the opposite orientation.

[0031] In some such constructs, the nucleic acid construct is for homologous recombination with a target genomic locus, the retinoschisin protein or fragment thereof is human retinoschisin protein or fragment thereof, the coding sequence of the retinoschisin protein or fragment thereof comprises complementary DNA (cDNA) comprising exons 2 to 6 of human RS1 or a degenerate variant thereof, the nucleic acid construct does not comprise a promoter driving expression of the retinoschisin protein or fragment thereof, the nucleic acid construct comprises a polyadenylation signal sequence located 3' of the coding sequence, and the nucleic acid construct comprises a splice acceptor site located 5' of the coding sequence, and each homology arm is from about 25 nucleotides to about 2.5 kb in length.

[0032] In another aspect, a vector is provided that includes any of the above nucleic acid constructs for homology-independent targeted integration. Some such vectors are viral vectors. Some such vectors are adeno-associated virus (AAV) vectors. Optionally, the AAV comprises a single-stranded genome (ssAAV). Optionally, the AAV comprises a self-complementary genome (scAAV). Optionally, the AAV is selected from the group consisting of AAV2, AAV5, AAV8, or AAV7m8.

[0033] In some such vectors, the vector does not comprise a promoter that drives expression of a retinoschisin protein or fragment thereof. In some such vectors, the vector does not comprise homology arms.

[0034] In another aspect, lipid nanoparticles are provided that include any of the above nucleic acid constructs for homology-independent targeted integration.

[0035] In another aspect, a cell is provided that contains any of the above-described nucleic acid constructs for homology-independent targeted integration. Some such cells are in vitro. Some such cells are in vivo. Some such cells are mammalian cells. Some such cells are human cells. Some such cells are retinal cells.

[0036] In some such cells, the cells express a retinoschisin protein or a fragment thereof. In some such cells, the nucleic acid construct is integrated into the genome at a target genomic locus. Optionally, the target genomic locus is the endogenous RS1 locus. Optionally, the nucleic acid construct is integrated into the genome into intron 1 of the endogenous RS1 locus. Optionally, endogenous RS1 exon 1 splices into the coding sequence of the retinoschisin protein or a fragment thereof in the nucleic acid construct. Optionally, the modified RS1 locus comprising the nucleic acid construct integrated into the genome encodes a protein comprising, consisting essentially of, or consisting of a sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 2 or 4.

[0037] In some such cells, integration of the nucleic acid construct into the endogenous RS1 locus prevents transcription of the endogenous RS1 gene downstream of the integration site. Optionally, integration of the nucleic acid construct into the endogenous RS1 locus reduces or eliminates expression of the endogenous retinoschisin protein and replaces it with expression of the retinoschisin protein or a fragment thereof encoded by the nucleic acid construct. Optionally, the endogenous RS1 locus contains a mutant RS1 gene containing a mutation that causes X-linked juvenile retinoschisis, and expression of the nucleic acid construct integrated into the genome reduces or eliminates expression of the mutant RS1 gene.

[0038] In another aspect, compositions are provided for use in expressing retinoschisin in a cell or for use in integrating the coding sequence of a retinoschisin protein or a fragment thereof into a target genomic locus in a cell. Some such compositions comprise (a) a nucleic acid construct comprising the coding sequence of a retinoschisin protein or a fragment thereof for integration into the target genomic locus, and (b) a nuclease agent or a nucleic acid encoding the nuclease agent, wherein the nuclease agent targets a nuclease target sequence in the target genomic locus.

[0039] Some such compositions include (a) any of the above-described nucleic acid constructs for homology-independent targeted integration, and (b) a nuclease agent or a nucleic acid encoding the nuclease agent, wherein the nuclease agent targets a nuclease target sequence at a target genomic locus. Optionally, the nuclease target sequence at the target genomic locus is identical to the nuclease target sequence of the nucleic acid construct. Optionally, the nuclease target sequence at the target genomic locus is destroyed when the nucleic acid construct is inserted in the correct orientation, but is reformed when the nucleic acid construct is inserted into the target genomic locus in the opposite orientation.

[0040] Some such compositions include (a) any of the bidirectional nucleic acid constructs described above; and (b) a nuclease agent or a nucleic acid encoding a nuclease agent, wherein the nuclease agent targets a nuclease target sequence in a target genomic locus.

[0041] In some such compositions, the target genomic locus is in the RS1 gene. Optionally, the nuclease target sequence of the target genomic locus is in the first intron of the RS1 gene. Optionally, integration of the nucleic acid construct into the endogenous RS1 locus prevents transcription of the endogenous RS1 gene downstream of the integration site. Optionally, integration of the nucleic acid construct into the endogenous RS1 locus in the cell reduces or eliminates expression of the endogenous retinoschisin protein and replaces it with expression of the retinoschisin protein or a fragment thereof encoded by the nucleic acid construct.

[0042] In some such compositions, the nuclease agent is a Cas protein and a guide RNA, and the nuclease target sequence is a guide RNA target sequence. Optionally, the Cas protein is Cas9. Optionally, the composition includes a guide RNA and a messenger RNA encoding the Cas protein. Optionally, the guide RNA and the messenger RNA encoding the Cas protein are in a lipid nanoparticle. Optionally, the composition includes DNA encoding the Cas protein and DNA encoding the guide RNA. Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in one or more viral vectors. Optionally, the one or more viral vectors are adeno-associated virus (AAV) viral vectors. Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in a single viral vector (e.g., a single AAV vector). Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in separate viral vectors (e.g., separate AAV vectors).

[0043] In some such compositions, the nucleic acid construct is in a viral vector. Optionally, the viral vector is an adeno-associated virus (AAV) viral vector. Optionally, the AAV is selected from the group consisting of AAV2, AAV5, AAV8, or AAV7m8.

[0044] Also provided is a composition comprising a guide RNA or DNA encoding a guide RNA, wherein the guide RNA comprises a DNA targeting segment that targets a guide RNA target sequence in the RS1 gene, and the guide RNA binds to a Cas protein and targets the Cas protein to the guide RNA target sequence in the RS1 gene.

[0045] In some such compositions or compositions for use, the composition further comprises a Cas protein or a nucleic acid encoding the Cas protein. Optionally, the Cas protein is a Cas9 protein. Optionally, the Cas protein is derived from the Streptococcus pyogenes Cas9 protein. In some such compositions or compositions for use, the composition comprises the Cas protein in protein form.

[0046] In some such compositions or compositions for use, the composition comprises a nucleic acid encoding a Cas protein, the nucleic acid comprising DNA encoding the Cas protein, and optionally the composition comprising DNA encoding a guide RNA. Optionally, the composition comprises a nucleic acid encoding a Cas protein, the nucleic acid comprising DNA encoding the Cas protein, the composition comprising DNA encoding a guide RNA, and the DNA encoding the Cas protein and the DNA encoding the guide RNA are in one or more viral vectors. Optionally, the one or more viral vectors are adeno-associated virus (AAV) viral vectors. Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in a single viral vector (e.g., a single AAV vector). Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in separate viral vectors (e.g., separate AAV vectors). Optionally, the AAV is selected from the group consisting of AAV2, AAV5, AAV8, or AAV7m8.

[0047] In some such compositions or compositions for use, the composition comprises a nucleic acid encoding a Cas protein, the nucleic acid comprising a messenger RNA encoding the Cas protein, and optionally the composition comprises a guide RNA in the form of RNA. In some such compositions or compositions for use, the composition comprises a nucleic acid encoding a Cas protein, the nucleic acid comprising a messenger RNA encoding the Cas protein, the composition comprises a guide RNA in the form of RNA, and the guide RNA and messenger RNA encoding the Cas protein are within a lipid nanoparticle.

[0048] In some such compositions or compositions for use, the messenger RNA encoding the Cas protein comprises at least one modification. Optionally, the messenger RNA encoding the Cas protein is modified to include modified uridines at one or more or all uridine positions. Optionally, the modified uridine is pseudouridine. Optionally, the messenger RNA encoding the Cas protein is completely substituted with pseudouridine. In some such compositions or compositions for use, the messenger RNA encoding the Cas protein comprises a 5' cap. In some such compositions or compositions for use, the messenger RNA encoding the Cas protein comprises a poly(A) tail. In some such compositions or compositions for use, the messenger RNA encoding the Cas protein comprises the sequence set forth in SEQ ID NO: 6243 or 6245.

[0049] In some such compositions or compositions for use, the nucleic acid encoding the Cas protein is codon-optimized for expression in mammalian or human cells. In some such compositions or compositions for use, the Cas protein comprises the sequence set forth in SEQ ID NO: 27, 6242, or 6246.

[0050] In some such compositions or compositions for use, the guide RNA target sequence is in an intron of the RS1 gene. Optionally, the intron is the first intron of the RS1 gene.

[0051] In some such compositions or compositions for use, the RS1 gene is a human RS1 gene.

[0052] In some such compositions or compositions for use, the DNA targeting segment comprises (a) at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 3148-6241; (b) at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 3148-4989; (c) at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351. (d) at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304; or (e) at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 4990 to 6241.

[0053] In some such compositions or compositions for use, the DNA targeting segment is (a) at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in any one of SEQ ID NOs: 3148-6241; (b) at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in any one of SEQ ID NOs: 3148-4989; or (c) at least 90% identical to the sequence set forth in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351. %, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304; or (e) at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in any one of SEQ ID NOs: 4990-6241.

[0054] In some such compositions or compositions for use, the DNA targeting segment comprises, consists essentially of, or consists of a sequence set forth in any one of SEQ ID NOs: 3148-6241, (b) any one of SEQ ID NOs: 3148-4989, (c) any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351, (d) any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304, or (e) any one of SEQ ID NOs: 4990-6241.

[0055] In some such compositions or compositions for use, the DNA targeting segment comprises at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304.

[0056] In some such compositions or compositions for use, the DNA targeting segment is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304.

[0057] In some such compositions or compositions for use, the DNA targeting segment comprises, consists essentially of, or consists of the sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304.

[0058] In some such compositions or compositions for use, the composition comprises a guide RNA in the form of RNA. In some such compositions or compositions for use, the composition comprises DNA encoding the guide RNA.

[0059] In some such compositions or compositions for use, the guide RNA comprises at least one modification. In some such compositions or compositions for use, the at least one modification comprises a 2'-O-methyl modified nucleotide. In some such compositions or compositions for use, the at least one modification comprises an internucleotide phosphorothioate bond. In some such compositions or compositions for use, the at least one modification comprises a modification in one or more of the first five nucleotides at the 5' end of the guide RNA. In some such compositions or compositions for use, the at least one modification comprises a modification in one or more of the last five nucleotides at the 3' end of the guide RNA. In some such compositions or compositions for use, the at least one modification comprises a phosphorothioate bond between the first four nucleotides at the 5' end of the guide RNA. In some such compositions or compositions for use, the at least one modification comprises a phosphorothioate bond between the last four nucleotides at the 3' end of the guide RNA. In some such compositions or compositions for use, the at least one modification comprises a 2'-O-methyl modified nucleotide in the first three nucleotides at the 5' end of the guide RNA. In some such compositions or compositions for use, the at least one modification comprises 2'-O-methyl modified nucleotides in the last three nucleotides of the 3' end of the guide RNA. In some such compositions or compositions for use, the at least one modification comprises (i) a phosphorothioate linkage between the first four nucleotides of the 5' end of the guide RNA, (ii) a phosphorothioate linkage between the last four nucleotides of the 3' end of the guide RNA, (iii) a 2'-O-methyl modified nucleotide in the first three nucleotides of the 5' end of the guide RNA, and (iv) a 2'-O-methyl modified nucleotide in the last three nucleotides of the 3' end of the guide RNA. In some such compositions or compositions for use, the guide RNA comprises a modified nucleotide of SEQ ID NO: 44.

[0060] In some such compositions or compositions for use, the guide RNA is a single guide RNA (sgRNA). Optionally, the guide RNA comprises, consists essentially of, or consists of a sequence set forth in any one of SEQ ID NOs: 33-39 and 53. In some such compositions or compositions for use, the guide RNA is a dual guide RNA (dgRNA) comprising two separate RNA molecules, including a CRISPR RNA (crRNA) and a trans-activating crRNA (tracrRNA). Optionally, the crRNA comprises a sequence set forth in any one of SEQ ID NOs: 29 and 52. Optionally, the tracrRNA comprises a sequence set forth in any one of SEQ ID NOs: 30-32.

[0061] In some such compositions or compositions for use, the composition is associated with lipid nanoparticles, and optionally, the composition includes a guide RNA. In some such compositions or compositions for use, the DNA encoding the guide RNA is in a viral vector. In some such compositions or compositions for use, the viral vector is an adeno-associated virus (AAV) viral vector. Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in a single viral vector (e.g., a single AAV vector). Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in separate viral vectors (e.g., separate AAV vectors). Optionally, the AAV is selected from the group consisting of AAV2, AAV5, AAV8, or AAV7m8.

[0062] In some such compositions or compositions for use, the composition is a pharmaceutical composition that includes a pharmaceutically acceptable carrier.

[0063] In some such compositions or compositions for use, the composition further comprises a second guide RNA or DNA encoding a second guide RNA, wherein the second guide RNA comprises a DNA-targeting segment that targets a second guide RNA target sequence in the RS1 gene, and the second guide RNA binds to a Cas protein and targets the Cas protein to the second guide RNA target sequence in the RS1 gene.

[0064] Also provided are cells comprising any of the above compositions or compositions for use. Optionally, the cells are in vitro. Optionally, the cells are in vivo. Some such cells are mammalian cells. Optionally, the cells are human cells. Optionally, the cells are retinal cells.

[0065] In some such cells, the cells express a first retinoschisin protein or fragment thereof, or a second retinoschisin protein or fragment thereof. In some such cells, the nucleic acid construct is integrated into the genome at a target genomic locus. Optionally, the target genomic locus is the endogenous RS1 locus. Optionally, the nucleic acid construct is integrated into the genome into intron 1 of the endogenous RS1 locus. Optionally, the endogenous RS1 exon 1 splices into the first coding sequence or the second coding sequence of the nucleic acid construct. Optionally, the modified RS1 locus comprising the nucleic acid construct integrated into the genome encodes a protein comprising, consisting essentially of, or consisting of a sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 2 or 4.

[0066] In some such cells, integration of the nucleic acid construct into the endogenous RS1 locus prevents transcription of the endogenous RS1 gene downstream of the integration site. Optionally, integration of the nucleic acid construct into the endogenous RS1 locus reduces or eliminates expression of the endogenous retinoschisin protein and replaces it with expression of the retinoschisin protein or a fragment thereof encoded by the nucleic acid construct. Optionally, the endogenous RS1 locus contains a mutant RS1 gene containing a mutation that causes X-linked juvenile retinoschisis, and expression of the nucleic acid construct integrated into the genome reduces or eliminates expression of the mutant RS1 gene.

[0067] In another aspect, methods are provided for integrating a coding sequence for a retinoschisin protein or a fragment thereof into a target genomic locus and expressing the retinoschisin protein or a fragment thereof in a cell. Some such methods comprise administering any of the nucleic acid constructs, vectors, lipid nanoparticles, or compositions described above to a cell, wherein the coding sequence is integrated into the target genomic locus and the retinoschisin protein or a fragment thereof is expressed in the cell. Optionally, the cell is a mammalian cell. Optionally, the cell is a human cell. Optionally, the cell is a retinal cell. Optionally, the cell is in vitro. Optionally, the cell is in vivo. Optionally, the cell is an in vivo retinal cell and the administration comprises subretinal injection or intravitreal injection.

[0068] In some such methods, the nucleic acid construct and the nuclease agent or nucleic acid encoding the nuclease agent are administered simultaneously. In some such methods, the nucleic acid construct and the nuclease agent or nucleic acid encoding the nuclease agent are administered sequentially in any order. Optionally, the nucleic acid construct is administered before the nuclease agent or nucleic acid encoding the nuclease agent. Optionally, the nucleic acid construct is administered after the nuclease agent or nucleic acid encoding the nuclease agent. Optionally, the time between sequential administrations is about 2 hours to about 48 hours. In some such methods, the nucleic acid construct and the nuclease agent or nucleic acid encoding the nuclease agent are administered in the same delivery vehicle. In some such methods, the nucleic acid construct and the nuclease agent or nucleic acid encoding the nuclease agent are administered in different delivery vehicles.

[0069] In some such methods, the target genomic locus is in the endogenous RS1 gene. Optionally, the nuclease target sequence of the target genomic locus is in the first intron of the endogenous RS1 gene. Optionally, the nucleic acid construct is integrated into the genome into intron 1 of the endogenous RS1 locus, and endogenous RS1 exon 1 splices to the coding sequence of a retinoschisin protein or a fragment thereof in the nucleic acid construct. Optionally, the modified RS1 locus comprising the nucleic acid construct integrated into the genome encodes a protein comprising, consisting essentially of, or consisting of a sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 2 or 4. In some such methods, integration of the nucleic acid construct into the endogenous RS1 locus prevents transcription of the endogenous RS1 gene downstream of the integration site. Optionally, integration of the nucleic acid construct into the endogenous RS1 locus reduces or eliminates expression of endogenous retinoschisin protein from the endogenous RS1 locus and replaces it with expression of retinoschisin protein or a fragment thereof encoded by the nucleic acid construct.

[0070] In another aspect, a method for treating a subject with X-linked juvenile retinoschisis is provided. Some such methods can include administering to a subject any of the above-mentioned nucleic acid constructs, vectors, lipid nanoparticles, or compositions, wherein the nucleic acid construct is integrated into and expressed from a target genomic locus in one or more retinal cells of the subject, thereby achieving a therapeutically effective level of retinoschisis expression in the subject. In some such methods, the subject is human. In some such methods, the subject has an endogenous RS1 gene containing at least one mutation associated with or causing X-linked juvenile retinoschisis. Optionally, the mutation is an R141C mutation. In some such methods, administration includes subretinal injection or intravitreal injection. In some such methods, the integration of the nucleic acid construct results in the restoration of retinal structure.

[0071] In some such methods, the nucleic acid construct and the nuclease agent or nucleic acid encoding the nuclease agent are administered simultaneously. In some such methods, the nucleic acid construct and the nuclease agent or nucleic acid encoding the nuclease agent are administered sequentially in any order. Optionally, the nucleic acid construct is administered before the nuclease agent or nucleic acid encoding the nuclease agent. Optionally, the nucleic acid construct is administered after the nuclease agent or nucleic acid encoding the nuclease agent. Optionally, the time between sequential administrations is about 2 hours to about 48 hours. In some such methods, the nucleic acid construct and the nuclease agent or nucleic acid encoding the nuclease agent are administered in the same delivery vehicle. In some such methods, the nucleic acid construct and the nuclease agent or nucleic acid encoding the nuclease agent are administered in different delivery vehicles.

[0072] In some such methods, the target genomic locus is in the endogenous RS1 gene. Optionally, the nuclease target sequence of the target genomic locus is in the first intron of the endogenous RS1 gene. Optionally, the nucleic acid construct is integrated into the genome into intron 1 of the endogenous RS1 locus, and endogenous RS1 exon 1 splices to the coding sequence of a retinoschisin protein or a fragment thereof in the nucleic acid construct. Optionally, the modified RS1 locus comprising the nucleic acid construct integrated into the genome encodes a protein comprising, consisting essentially of, or consisting of a sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 2 or 4. In some such methods, integration of the nucleic acid construct into the endogenous RS1 locus prevents transcription of the endogenous RS1 gene downstream of the integration site. Optionally, integration of the nucleic acid construct into the endogenous RS1 locus reduces or eliminates expression of endogenous retinoschisin protein from the endogenous RS1 locus and replaces it with expression of retinoschisin protein or a fragment thereof encoded by the nucleic acid construct.

[0073] In another aspect, a method for modifying the RS1 gene in a cell is provided. Some such methods include administering to a cell any of the above-described compositions comprising a guide RNA or DNA encoding the guide RNA and a Cas protein or a nucleic acid encoding the Cas protein, wherein the guide RNA binds to the Cas protein and targets the Cas protein to the guide RNA target sequence in the RS1 gene, and the Cas protein cleaves the guide RNA target sequence. In some such methods, the cell is a mammalian cell. Optionally, the cell is a human cell. Optionally, the cell is a retinal cell. Optionally, the cell is in vitro. Optionally, the cell is in vivo. In some such methods, the cell is a retinal cell, and the administration comprises subretinal injection or intravitreal injection. In some such methods, the guide RNA target sequence is in the first intron of the RS1 gene. [Brief explanation of the drawings]

[0074] [Figure 1] A schematic diagram (not to scale) of the mouse Rs1 locus is shown, including the location of the R141C mutation associated with X-linked juvenile retinoschisis (XLRS) and the insertion site of a nucleic acid construct containing exons 2-6 of human RS1. [Figure 2] Shown is an alignment of mouse retinoschisin, human retinoschisin, human retinoschisin with the R141C mutation, mouse retinoschisin with the R141C mutation, and a mouse / human retinoschisin hybrid expressed upon integration of a nucleic acid construct containing exons 2-6 of human RS1 into intron 1 of the mouse Rs1 locus. [Figure 3A]A schematic diagram (not to scale) of a bidirectional nucleic acid construct is shown, containing a first segment containing a splice acceptor (A), exons 2-6 of human RS1, and bovine growth hormone (bGH) polyA, and a second segment containing the reverse complement of SV40 polyA, the reverse complement of exons 2-6 of human RS1, and the reverse complement of the splice acceptor (A). The bidirectional construct also contains a U6 promoter operably linked to a sequence encoding a guide RNA targeting intron 1 of the mouse Rs1 locus between the two human RS1 segments. The horizontal arrow adjacent to the star represents the next-generation targeted resequencing amplicon designed for NGS. The bidirectional ssAAV construct is shown at the top, and the bidirectional scAAV construct is shown at the bottom. [Figure 3B] A schematic diagram (not to scale) of a homology-independent targeted integration nucleic acid construct is shown, containing a splice acceptor (A), exons 2-6 of human RS1, and a polyA sequence. The construct also contains a U6 promoter operably linked to a sequence encoding a guide RNA targeting intron 1 of the mouse Rs1 locus downstream of the human RS1 segment. The horizontal arrow adjacent to the star represents the next-generation targeted resequencing amplicon designed for NGS. [Figure 4] Scoring of retinal cavities shown in optical coherence tomography (OCT) scans of eyes from RosaCas9 / +, Rs1R141C / Y mice injected with RS1 viral vector version 1, RS1 viral vector version 2, or RS1 viral vector version 3 is shown. A score of 1 was assigned if at least one individual image contained 1–4 cavities. A score of 2 was assigned if at least one individual image contained 4 or more cavities but the cavities were not fused. A score of 3 was assigned if at least one individual image contained fused cavities. A score of 4 was assigned if at least one individual image contained fused cavities and the retina was elongated. The mean scores for each treatment group were compared to the pooled control group, including untreated eyes, by nonparametric Kruskal-Wallis one-way ANOVA with post-hoc Dunn's multiple comparison test. [Figure 5] NGS results from mouse retina samples from eyes of RosaCas9 / +, Rs1R141C / Y mice injected with RS1 viral vector version 1 (pssAAV mhRS1-sgu), RS1 viral vector version 2 (pscAAV rs1_tandem), or RS1 viral vector version 3 (pssAAV hRs1_HITI) are shown. Read counts for four expected sequence variants are shown: (1) WT mouse, mouse reference sequence without the R141C mutation; (2) mutant mouse, mouse reference sequence with the R141C mutation; (3) humanized transcript 1, human reference sequence; and (4) humanized transcript 2, mouse codon-optimized human reference sequence. Next-generation targeted resequencing amplicons were designed for the regions indicated in Figure 3 (horizontal arrows). mRNA from mouse retinas was used to generate cDNA, which served as a template for next-generation sequencing (NGS) amplification. [Figure 6A] NGS results from mouse retina samples from eyes of RosaCas9 / +, Rs1R141C / Y mice injected with RS1 viral vector version 1 (mhRS1-sgu), RS1 viral vector version 2 (pscAAV_rs1_tandem), or RS1 viral vector version 3 (hRs1_cDNA HITI) are shown. In these NGS results, separate amplicons were used to amplify the Rs1 intron 1 guide RNA target sequence. The number of reads that matched the mouse reference sequence or contained non-homologous end joining was quantified to assess the frequency of guide RNA cleavage without insertion. [Figure 6B]NGS results from mouse retina samples from eyes of RosaCas9 / +, Rs1R141C / Y mice injected with RS1 viral vector version 1 (mhRS1-sgu), RS1 viral vector version 2 (pscAAV_rs1_tandem), or RS1 viral vector version 3 (hRs1_cDNA HITI) are shown. In these NGS results, separate amplicons were used to amplify the Rs1 intron 1 guide RNA target sequence. The number of reads that matched the mouse reference sequence or contained non-homologous end joining was quantified to assess the frequency of guide RNA cleavage without insertion. [Figure 7A] NGS results from mouse retina samples from eyes of RosaCas9 / +, Rs1R141C / Y mice injected with RS1 viral vector version 1 (pssAAV mhRS1-sgu, Figure 7A), RS1 viral vector version 2 (pscAAV rs1_tandem, Figure 7B), or RS1 viral vector version 3 (pssAAV hRs1_HITI, Figure 7C) are shown. Read counts for four expected sequence variants are shown: (1) WT mouse, mouse reference sequence without the R141C mutation; (2) mutant mouse, mouse reference sequence with the R141C mutation; (3) humanized transcript 1, human reference sequence; and (4) humanized transcript 2, mouse codon-optimized human reference sequence. Next-generation targeted resequencing amplicons were designed for the regions indicated in Figure 3 (horizontal arrows). mRNA from mouse retinas was used to generate cDNA, which served as a template for next-generation sequencing (NGS) amplification. The number of reads that matched the mouse reference sequence or contained non-homologous end joining was quantified to assess the frequency of guide RNA truncation without insertion. [Figure 7B]NGS results from mouse retina samples from eyes of RosaCas9 / +, Rs1R141C / Y mice injected with RS1 viral vector version 1 (pssAAV mhRS1-sgu, Figure 7A), RS1 viral vector version 2 (pscAAV rs1_tandem, Figure 7B), or RS1 viral vector version 3 (pssAAV hRs1_HITI, Figure 7C) are shown. Read counts for four expected sequence variants are shown: (1) WT mouse, mouse reference sequence without the R141C mutation; (2) mutant mouse, mouse reference sequence with the R141C mutation; (3) humanized transcript 1, human reference sequence; and (4) humanized transcript 2, mouse codon-optimized human reference sequence. Next-generation targeted resequencing amplicons were designed for the regions indicated in Figure 3 (horizontal arrows). mRNA from mouse retinas was used to generate cDNA, which served as a template for next-generation sequencing (NGS) amplification. The number of reads that matched the mouse reference sequence or contained non-homologous end joining was quantified to assess the frequency of guide RNA truncation without insertion. [Figure 7C]NGS results from mouse retina samples from eyes of RosaCas9 / +, Rs1R141C / Y mice injected with RS1 viral vector version 1 (pssAAV mhRS1-sgu, Figure 7A), RS1 viral vector version 2 (pscAAV rs1_tandem, Figure 7B), or RS1 viral vector version 3 (pssAAV hRs1_HITI, Figure 7C) are shown. Read counts for four expected sequence variants are shown: (1) WT mouse, mouse reference sequence without the R141C mutation; (2) mutant mouse, mouse reference sequence with the R141C mutation; (3) humanized transcript 1, human reference sequence; and (4) humanized transcript 2, mouse codon-optimized human reference sequence. Next-generation targeted resequencing amplicons were designed for the regions indicated in Figure 3 (horizontal arrows). mRNA from mouse retinas was used to generate cDNA, which served as a template for next-generation sequencing (NGS) amplification. The number of reads that matched the mouse reference sequence or contained non-homologous end joining was quantified to assess the frequency of guide RNA truncation without insertion. [Figure 8A] Figure 8A shows RT-qPCR results from human retinoblastoma cells treated with RS1 viral vector version 1 (pssAAV mhRS1-sgu, Figure 8A) or RS1 viral vector version 2 (pscAAV rs1_tandem, Figure 8B) 2 hours prior to treatment with lipid nanoparticles formulated with Cas9 mRNA and one of six guide RNAs targeting human RS1 intron 1. Delta Ct values ​​are indicated (lower numbers indicate higher expression). "Ho" refers to the human reference sequence, and "Mo" refers to the human reference sequence codon-optimized for mouse expression. [Figure 8B]Figure 8A shows RT-qPCR results from human retinoblastoma cells treated with RS1 viral vector version 1 (pssAAV mhRS1-sgu, Figure 8A) or RS1 viral vector version 2 (pscAAV rs1_tandem, Figure 8B) 2 hours prior to treatment with lipid nanoparticles formulated with Cas9 mRNA and one of six guide RNAs targeting human RS1 intron 1. Delta Ct values ​​are indicated (lower numbers indicate higher expression). "Ho" refers to the human reference sequence, and "Mo" refers to the human reference sequence codon-optimized for mouse expression. [Figure 9A] Figure 9A shows RT-qPCR results from human retinoblastoma cells treated with RS1 viral vector version 1 (pssAAV mhRS1-sgu, Figure 9A) or RS1 viral vector version 2 (pscAAV rs1_tandem, Figure 9B) 2 hours after treatment with lipid nanoparticles formulated with Cas9 mRNA and one of six guide RNAs targeting human RS1 intron 1. Delta Ct values ​​are indicated (lower numbers indicate higher expression). "Ho" refers to the human reference sequence, and "Mo" refers to the human reference sequence codon-optimized for mouse expression. [Figure 9B] Figure 9A shows RT-qPCR results from human retinoblastoma cells treated with RS1 viral vector version 1 (pssAAV mhRS1-sgu, Figure 9A) or RS1 viral vector version 2 (pscAAV rs1_tandem, Figure 9B) 2 hours after treatment with lipid nanoparticles formulated with Cas9 mRNA and one of six guide RNAs targeting human RS1 intron 1. Delta Ct values ​​are indicated (lower numbers indicate higher expression). "Ho" refers to the human reference sequence, and "Mo" refers to the human reference sequence codon-optimized for mouse expression. [Figure 10]A schematic diagram of the nucleic acid construct for homologous recombination is shown, including a splice acceptor (A), exons 2-6 of human RS1, and a polyA sequence. The construct also contains a U6 promoter operably linked to a sequence encoding a guide RNA targeting intron 1 of the mouse Rs1 locus downstream of the human RS1 segment. The construct also contains upstream and downstream homology arms (HA). The horizontal arrow adjacent to the star represents the next-generation targeted resequencing amplicon designed for NGS. DETAILED DESCRIPTION OF THE INVENTION

[0075] definition The terms "protein," "polypeptide," and "peptide," used interchangeably herein, include polymeric forms of amino acids of any length, including coded and non-coded amino acids and amino acids that are chemically or biochemically modified or derivatized. These terms also include modified polymers, such as polypeptides with modified peptide backbones. The term "domain" refers to any portion of a protein or polypeptide having a specific function or structure.

[0076] The terms "nucleic acid" and "polynucleotide," used interchangeably herein, include polymeric forms of nucleotides of any length, containing ribonucleotides, deoxyribonucleotides, or analogs or modified versions thereof. These include single-, double-, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers that contain purine bases, pyrimidine bases, or other natural, chemically modified, biochemically modified, non-natural, or derivatized nucleotide bases.

[0077] The term "genomically integrated" refers to a nucleic acid that has been introduced into a cell such that the nucleotide sequence is integrated into the genome of the cell. Any protocol may be used for stable integration of a nucleic acid into the genome of a cell.

[0078] The terms "expression vector" or "expression construct" or "expression cassette" refer to a recombinant nucleic acid comprising a desired coding sequence operably linked to appropriate nucleic acid sequences necessary for expression of the operably linked coding sequence in a particular host cell or organism. Nucleic acid sequences necessary for expression in prokaryotes usually include a promoter, an operator (optional), a ribosome binding site, and other sequences. Eukaryotic cells are known to generally utilize promoters, enhancers, termination and polyadenylation signals, and some elements can be deleted and others added without sacrificing the required expression.

[0079] The term "viral vector" refers to a recombinant nucleic acid that contains at least one element of viral origin and contains elements sufficient for or that allow packaging into a viral vector particle. The vector and / or particle can be used to transfer DNA, RNA, or other nucleic acids into cells either ex vivo or in vivo. Many forms of viral vectors are known.

[0080] The term "isolated," with respect to proteins, nucleic acids, and cells, includes proteins, nucleic acids, and cells that are relatively purified with respect to other cellular or biological components that may normally be present in situ, up to and including substantially pure preparations of proteins, nucleic acids, or cells. The term "isolated" also includes proteins and nucleic acids that have no naturally occurring counterpart, or proteins or nucleic acids that are chemically synthesized and thus are substantially uncontaminated by other proteins or nucleic acids. The term "isolated" can also include proteins, nucleic acids, or cells that have been separated or purified from most other cellular or biological components with which they are naturally associated (e.g., other cellular proteins, nucleic acids, or cellular or extracellular components).

[0081] The term "wild-type" includes entities having a structure and / or activity as found in a normal state or context (as opposed to mutant, diseased, altered, etc.). Wild-type genes and polypeptides often exist in multiple alternative forms (e.g., alleles).

[0082] The term "endogenous sequence" refers to a nucleic acid sequence that is naturally occurring within a cell or animal. For example, an animal's endogenous RS1 sequence refers to the native RS1 sequence that is naturally occurring at the RS1 locus of the animal.

[0083] An "exogenous" molecule or sequence includes a molecule or sequence that is not normally present in a cell in that form. Normal presence includes presence in relation to a particular developmental stage and environmental conditions of the cell. An exogenous molecule or sequence may include, for example, a mutated version of a corresponding endogenous sequence in a cell, such as a humanized version of an endogenous sequence, or may include a sequence that corresponds to an endogenous sequence within a cell but in a different form (i.e., not within a chromosome). In contrast, an endogenous molecule or sequence includes a molecule or sequence that is normally present in that form, in a particular cell, at a particular developmental stage, and under particular environmental conditions.

[0084] The term "heterologous" when used in the context of a nucleic acid or protein indicates that the nucleic acid or protein contains at least two segments that do not naturally occur together in the same molecule. For example, the term "heterologous" when used with reference to a segment of a nucleic acid or a segment of a protein indicates that the nucleic acid or protein contains two or more subsequences that are not found in the same relationship to each other (e.g., linked together) in nature. As an example, a "heterologous" region of a nucleic acid vector is a segment of nucleic acid within or attached to another nucleic acid molecule that is not found in association with that other molecule in nature. For example, a heterologous region of a nucleic acid vector can include a coding sequence adjacent to a sequence not found in association with the coding sequence in nature. Similarly, a "heterologous" region of a protein is a segment of amino acids within or attached to another peptide molecule that is not found in association with other peptide molecules in nature (e.g., a fusion protein or a tagged protein). Similarly, a nucleic acid or protein can include a heterologous tag or a heterologous secretion or localization sequence.

[0085] "Codon optimization" refers to the process of modifying a nucleic acid sequence to enhance expression in a particular host cell by taking advantage of codon degeneracy, as indicated by the diversity of three-base pair codon combinations that specify amino acids, generally by replacing at least one codon of the native sequence with a codon more frequently or most frequently used in the host cell's genes while maintaining the native amino acid sequence. For example, a nucleic acid encoding a Cas9 protein can be modified to use an alternative codon that is more frequently used in a given prokaryotic or eukaryotic cell, including bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, or any other host cell, compared to the naturally occurring nucleic acid sequence. Codon usage tables are readily available, for example, in "codon usage databases." These tables can be adapted in various ways. See Nakamura et al. (2000) Nucleic Acids Research 28:292, incorporated herein by reference in its entirety for all purposes. Computer algorithms are also available for codon optimization of a particular sequence for expression in a particular host (see, eg, Gene Forge).

[0086] The term "locus" refers to the specific location of a gene (or key sequence), DNA sequence, polypeptide-encoding sequence, or position on a chromosome in the genome of an organism. For example, "RS1 locus" can refer to the specific location of the RS1 gene, RS1 DNA sequence, retinoschisin-encoding sequence, or the location of RS1 on a chromosome in the genome of an organism in which such a sequence is identified to reside. The "RS1 locus" can include regulatory elements of the RS1 gene, including, for example, an enhancer, promoter, 5' and / or 3' untranslated regions (UTRs), or a combination thereof.

[0087] The term "gene" refers to a DNA sequence in a chromosome that, when naturally occurring, may contain at least one coding region and at least one non-coding region. A DNA sequence in a chromosome that encodes a product (e.g., but not limited to, an RNA product and / or a polypeptide product) may include coding regions interrupted by non-coding introns and sequences located adjacent to the coding region at both the 5' and 3' ends, such that the gene corresponds to a full-length mRNA (including 5' and 3' untranslated sequences). Additionally, other non-coding sequences, including regulatory sequences (e.g., but not limited to, promoters, enhancers, and transcription factor binding sites), polyadenylation signals, internal ribosome entry sites, silencers, insulating sequences, and matrix attachment regions, may also be present in a gene. These sequences may be adjacent (e.g., within 10 kb) or distant from the coding region of the gene, and they affect the level or rate of gene transcription and translation.

[0088] The term "allele" refers to variant forms of a gene. Some genes have different forms that are located at the same position, or locus, on a chromosome. Diploid organisms have two alleles at each locus. Each pair of alleles represents a genotype at a particular locus. A genotype is described as homozygous if there are two identical alleles at a particular locus, and as heterozygous if the two alleles are different.

[0089] A "promoter" is a regulatory region of DNA that typically contains a TATA box that can direct RNA polymerase II to begin RNA synthesis at the appropriate transcription start site for a particular polynucleotide sequence. A promoter may further contain other regions that affect the rate of transcription initiation. The promoter sequences disclosed herein regulate the transcription of an operably linked polynucleotide. The promoter may be active in one or more cell types disclosed herein (e.g., eukaryotic cells, non-human mammalian cells, human cells, rodent cells, pluripotent cells, one-cell stage embryos, differentiated cells, or a combination thereof). The promoter may be, for example, a constitutively active promoter, a conditional promoter, an inducible promoter, a temporally restricted promoter (e.g., a developmentally regulated promoter), or a spatially restricted promoter (e.g., a cell-specific or tissue-specific promoter). Examples of promoters can be found, for example, in WO2013 / 176772, incorporated herein by reference in its entirety for all purposes.

[0090] Constitutive promoters are those that are active in all tissues or in specific tissues at all stages of development. Examples of constitutive promoters include human cytomegalovirus immediate early (hCMV), mouse cytomegalovirus immediate early (mCMV), human elongation factor 1 alpha (hEF1α), mouse elongation factor 1 alpha (mEF1α), mouse phosphoglycerate kinase (PGK), chicken beta actin hybrid (CAG or CBh), SV40 early, and beta 2 tubulin promoters.

[0091] Examples of inducible promoters include, for example, chemically regulated promoters and physically regulated promoters.Chemically regulated promoters include, for example, alcohol-regulated promoters (e.g., alcohol dehydrogenase (alcA) gene promoter), tetracycline-regulated promoters (e.g., tetracycline-responsive promoters, tetracycline operator sequence (tetO), tet-On promoters, or tet-Off promoters), steroid-regulated promoters (e.g., rat glucocorticoid receptor, estrogen receptor promoter, or ecdysone receptor promoter), or metal-regulated promoters (e.g., metalloprotein promoters).Physically regulated promoters include, for example, temperature-regulated promoters (e.g., heat shock promoters) and light-regulated promoters (e.g., light-inducible promoters or light-repressible promoters).

[0092] The tissue-specific promoter can be, for example, a neuron-specific promoter, a glial-specific promoter, a muscle cell-specific promoter, a cardiac cell-specific promoter, a kidney cell-specific promoter, a bone cell-specific promoter, an endothelial cell-specific promoter, or an immune cell-specific promoter (e.g., a B cell promoter or a T cell promoter).

[0093] Developmentally-regulated promoters include, for example, promoters that are active only during embryonic stages of development or only in adult cells.

[0094] "Operable linkage" or "operably linked" includes the juxtaposition of two or more components (e.g., a promoter and another sequence element) that allows for both components to function normally and for at least one of the components to mediate the function of at least one of the other components. For example, a promoter can be operably linked to a coding sequence if it controls the level of transcription of the coding sequence depending on the presence or absence of one or more transcriptional regulatory factors. Operable linkage can include proximity of such sequences to each other or acting in trans (e.g., regulatory sequences can act at a distance to control transcription of the coding sequence).

[0095] "Complementarity" of nucleic acids means that the nucleotide sequence of one strand of a nucleic acid will form hydrogen bonds with another sequence on the opposite strand due to the orientation of its nucleobase groups. In DNA, complementary bases are typically A and T and C and G. In RNA, they are typically C and G and U and A. Complementarity can be complete or substantial / sufficient. Complete complementarity between two nucleic acids means that the two nucleic acids can form a duplex, with all bases in the duplex binding to complementary bases through Watson-Crick pairing. "Substantial" or "sufficient" complementarity means that the sequence of one strand is not completely and / or perfectly complementary to the sequence of the opposite strand, but sufficient binding occurs between the bases of the two strands to form a stable hybrid complex under a set of hybridization conditions (e.g., salt concentration and temperature). Such conditions can be predicted by predicting the Tm (melting temperature) of the hybridized strands using the sequences and standard mathematical calculations, or by empirically determining the Tm using routine methods. Tm comprises the temperature at which the population of hybridization complexes formed between two nucleic acid strands is 50% denatured (i.e., the population of double-stranded nucleic acid molecules is half-dissociated into single strands). At temperatures below Tm, the formation of hybridization complexes is favored, while at temperatures above Tm, melting or separation of the strands in the hybridization complex is favored. While other known Tm calculations take into account the structural characteristics of nucleic acids, Tm can be estimated for a nucleic acid with a known G+C content in 1 M NaCl aqueous solution, for example, by using Tm = 81.5 + 0.41 (G+C%).

[0096] Hybridization requires that the two nucleic acids contain complementary sequences, although mismatches between bases are possible. Suitable conditions for hybridization between two nucleic acids depend on the length of the nucleic acids and the degree of complementarity, which are well-known variables. The greater the degree of complementarity between two nucleic acid sequences, the higher the melting temperature (Tm) of hybrids of nucleic acids having those sequences. For hybridization between nucleic acids having short stretches of complementarity (e.g., complementarity over 35 or fewer, 30 or fewer, 25 or fewer, 22 or fewer, 20 or fewer, or 18 or fewer nucleotides), the location of mismatches becomes important (see Sambrook et al., supra, 11.7-11.8). Typically, the length of a hybridizable nucleic acid is at least about 10 nucleotides. Exemplary minimum lengths for a hybridizable nucleic acid include at least about 15 nucleotides, at least about 20 nucleotides, at least about 22 nucleotides, at least about 25 nucleotides, and at least about 30 nucleotides. Additionally, the temperature and salt concentration of the wash solution may be adjusted as needed according to factors such as the length of the region of complementarity and the degree of complementarity.

[0097] The sequence of a polynucleotide does not need to be 100% complementary to the sequence of its target nucleic acid to specifically hybridize. Furthermore, a polynucleotide may hybridize across one or more segments, such that intervening or adjacent segments are not involved in the hybridization event (e.g., a loop structure or a hairpin structure). Polynucleotides (e.g., gRNAs) may contain at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence complementarity to the target region within the target nucleic acid sequence to which they are targeted. For example, a gRNA in which 18 of 20 nucleotides are complementary to the target region and therefore will specifically hybridize would exhibit 90% complementarity. In this example, the remaining non-complementary nucleotides may be clustered or interspersed with complementary nucleotides and need not be contiguous with each other or with complementary nucleotides.

[0098] The percent complementarity between nucleic acid sequences over a specific stretch within a nucleic acid can be routinely determined using the BLAST program (basic local alignment search tool) and PowerBLAST program (Altschul et al. (1990) J. Mol. Biol. 215:403-410; Zhang and Madden (1997) Genome Res. 7:649-656) or the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.) using default settings, which uses the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489).

[0099] The methods and compositions provided herein use a variety of different components. Some components throughout the description may have active variants and fragments. Such components include, for example, Cas proteins, CRISPR RNAs, tracrRNAs, and guide RNAs. The biological activities of each of these components are described elsewhere herein. The term "functionality" refers to the inherent ability of a protein or nucleic acid (or a fragment or variant thereof) to exhibit a biological activity or function. Such biological activity or function may include, for example, the ability of a Cas protein to bind to a guide RNA and a target DNA sequence. The biological function of a functional fragment or variant may be the same as or actually altered (e.g., in terms of their specificity, selectivity, or efficacy) compared to the original molecule, while retaining the basic biological function of the molecule.

[0100] The term "variant" refers to a nucleotide sequence (e.g., one nucleotide) that differs from the most common sequence in a population, or a protein sequence (e.g., one amino acid) that differs from the most common sequence in a population.

[0101] The term "fragment," when referring to a protein, refers to a protein that is shorter or has fewer amino acids than the full-length protein. The term "fragment," when referring to a nucleic acid, refers to a nucleic acid that is shorter or has fewer nucleotides than the full-length nucleic acid. A fragment, for example, when referring to a protein fragment, can be an N-terminal fragment (i.e., removal of a portion of the C-terminus of the protein), a C-terminal fragment (i.e., removal of a portion of the N-terminus of the protein), or an internal fragment (i.e., removal of a portion of an internal portion of the protein).

[0102] "Sequence identity" or "identity" in the context of two polynucleotide or polypeptide sequences refers to the residues of the two sequences that are the same when aligned for maximum correspondence over a specified comparison window. When using percentage sequence identity with respect to proteins, non-identical residue positions often differ by conservative amino acid substitutions, in which an amino acid residue is replaced with another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity) and therefore does not alter the functional properties of the molecule. When sequences differ by conservative substitutions, the percent sequence identity may be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." Means for making this adjustment are well known. Typically, this involves scoring conservative substitutions as partial rather than complete mismatches, thereby increasing the percentage sequence identity. Thus, for example, where identical amino acids are given a score of 1 and non-conservative substitutions are given a score of 0, conservative substitutions are given a score of 0 to 1. Scoring of conservative substitutions is calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, California).

[0103] "Percentage of sequence identity" includes a value determined by comparing two optimally aligned sequences (maximum number of perfectly matched residues) over a comparison window, where the portion of the polynucleotide sequence in the comparison window may contain additions or deletions (i.e., gaps) when compared to a reference sequence (not including additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions where the same nucleic acid base or amino acid residue occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Unless otherwise specified (e.g., the shorter sequence includes a concatenated non-homologous sequence), the comparison window is the full length of the shorter of the two sequences being compared.

[0104] Unless otherwise specified, sequence identity / similarity values ​​include values ​​obtained using GAP version 10 with the following parameters: % identity and % similarity for nucleotide sequences using a GAP weight of 50 and a length weight of 3, and the nwsgapdna.cmp scoring matrix; % identity and % similarity for amino acid sequences using a GAP weight of 8 and a length weight of 2, and the BLOSUM62 scoring matrix; or any equivalent program. "Equivalent program" includes any sequence comparison program that produces alignments with identical nucleotide or amino acid residue matches and identical percent sequence identity for any two sequences in question when compared to corresponding alignments produced by GAP version 10.

[0105] The term "conservative amino acid substitution" refers to the substitution of an amino acid normally present in a sequence with a different amino acid of similar size, charge, or polarity. Examples of conservative substitutions include the substitution of a non-polar (hydrophobic) residue such as isoleucine, valine, or leucine for another non-polar residue. Similarly, examples of conservative substitutions include the substitution of one polar (hydrophilic) residue for another, such as between arginine and lysine, between glutamine and asparagine, or between glycine and serine. Additionally, the substitution of a basic residue such as lysine, arginine, or histidine for another, or the substitution of one acidic residue such as aspartic acid or glutamic acid for another, are additional examples of conservative substitutions. Examples of non-conservative substitutions include the substitution of a non-polar (hydrophobic) amino acid residue, such as isoleucine, valine, leucine, alanine, or methionine, for a polar (hydrophilic) residue, such as cysteine, glutamine, glutamic acid, or lysine, and / or a polar residue for a non-polar residue. Typical amino acid classifications are summarized in Table 1 below. [Table 1]

[0106] A "homologous" sequence (e.g., a nucleic acid sequence) includes a sequence that is identical to or substantially similar to a known reference sequence, e.g., at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the known reference sequence. Homologous sequences can include, for example, orthologous and paralogous sequences. For example, homologous genes typically originate from a common ancestral DNA sequence through either speciation events (orthologous genes) or gene duplication events (paralogous genes). "Orthologous" genes include genes from different species that have evolved from a common ancestral gene through speciation. Orthologs typically retain the same function during evolution. "Paralogous" genes include genes that are related by duplication within a genome. Paralogs can evolve new functions during evolution.

[0107] The term "in vitro" includes an artificial environment and processes or reactions that occur within an artificial environment (e.g., a test tube or an isolated cell or cell line). The term "in vivo" includes a natural environment (e.g., a cell or organism or body) and processes or reactions that occur within a natural environment. The term "ex vivo" includes cells removed from an individual's body and processes or reactions that occur within such cells.

[0108] The repair of double-strand break (DSB) mainly occurs through two conserved DNA repair pathways: homologous recombination (HR) and non-homologous end joining (NHEJ).See Kasparek&Humphrey (2011)Seminars in Cell&Dev.Biol.22:886-897, and is incorporated herein by reference in its entirety for all purposes.Similarly, the repair of target nucleic acid mediated by exogenous donor nucleic acid can include any process of exchanging genetic information between two polynucleotides.

[0109] The term "recombination" includes any process of exchanging genetic information between two polynucleotides and can occur by any mechanism. Recombination can occur via homology-directed repair (HDR) or homologous recombination (HR). HDR or HR involves forms of nucleic acid repair that may require nucleotide sequence homology and use a "donor" molecule as a template for repair of a "target" molecule (i.e., a molecule that has experienced a double-strand break), resulting in the transfer of genetic information from the donor to the target. Without wishing to be bound by any particular theory, such transfer may involve mismatch correction of heteroduplex DNA formed between the broken target and the donor, and / or synthesis-dependent strand annealing, and / or related processes, where the donor resynthesizes the genetic information that will become part of the target. In some cases, the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide is integrated into the target DNA. See Wang et al. (2013) Cell 153:910-918, Mandalos et al. (2012) PLOS ONE 7:e45768:1-9, and Wang et al. (2013) Nat Biotechnol. 31:530-532, each of which is incorporated by reference in its entirety for all purposes.

[0110] Non-homologous end joining (NHEJ) involves the repair of double-strand breaks in nucleic acids by direct ligation of the cut ends to each other or to an exogenous sequence without the need for a homologous template. Ligation of non-adjacent sequences by NHEJ can often result in deletions, insertions, or translocations near the site of the double-strand break. For example, NHEJ can also result in targeted integration of an exogenous donor nucleic acid by directly ligating the cut ends to the ends of the exogenous donor nucleic acid (i.e., NHEJ-based capture). Such NHEJ-mediated targeted integration may be preferable for the insertion of an exogenous donor nucleic acid when the homology-directed repair (HDR) pathway is not readily available (e.g., non-dividing cells, primary cells, and cells that perform poorly in homology-based DNA repair). In addition, in contrast to homology-directed repair, knowledge of large regions of sequence identity adjacent to the break site is not required, which may be beneficial when attempting targeted insertion into an organism with a genome with limited knowledge of the genome sequence. Integration can be carried out by blunt-end ligation between exogenous donor nucleic acid and cleaved genome sequence, or by sticky-end ligation (i.e., having 5' or 3' overhang) using exogenous donor nucleic acid adjacent to the overhang that matches that generated by nuclease agent in cleaved genome sequence.See, for example, US2011 / 020722, WO2014 / 033644, WO2014 / 089290, and Maresca et al. (2013) Genome Res.23(3):539-546, each of which is incorporated herein by reference in its entirety for all purposes.When blunt-end ligation is carried out, the target and / or donor may need to be excised due to the generation of microhomology regions required for fragment binding, which may cause undesirable modifications to the target sequence.

[0111] A composition or method "comprising" or "including" one or more recited elements may include other elements not specifically recited. For example, a composition "comprises" or "includes" a protein may include the protein alone or in combination with other components. The transitional phrase "consisting essentially of" means that the claim scope shall be construed to include the specific elements recited in the claim as well as elements that do not materially affect the basic and novel characteristics of the claimed invention. Thus, the term "consisting essentially of," when used in the claims of the present invention, is not intended to be construed as equivalent to "comprising."

[0112] "Optional" or "optionally" means that the subsequently described event or circumstance may or may not occur, and that the description includes examples when the event or circumstance occurs and examples when it does not occur.

[0113] The specification of a range of values ​​includes all integers within or defining the range, and all subranges defined by integers within the range.

[0114] Unless otherwise clear from the context, the term "about" encompasses values ​​that are ±5 of the stated value.

[0115] The term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or").

[0116] The term "or" refers to any one member of a particular list and also includes any combination of members of that list.

[0117] The singular articles "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "protein" or "at least one protein" can include a plurality of proteins, including mixtures thereof.

[0118] Statistically significant means p≦0.05. (Mode for Carrying Out the Invention)

[0119] I. Overview X-linked juvenile retinoschisis (XLRS) is a form of early-onset macular degeneration caused by mutations in retinoschisin (RS1). The RS1 gene encodes a 24-kDa discoidin domain-containing protein secreted as a homo-oligomeric complex. Genetic mutations in RS1 cause either a nonfunctional protein or a lack of protein secretion, leading to fragmentation or division within the retinal layers and early, progressive vision loss. More than 200 different mutations in the RS1 gene are known to cause XLRS. Forty percent of disease-causing mutations are nonsense or frameshift mutations predicted to result in the absence of full-length retinoschisin protein. Fifty percent of disease-causing mutations are missense mutations that allow the production of full-length mutant proteins. Most of these are in the discoidin domain, resulting in misfolded proteins being retained in the ER.

[0120] Because XLRS is a recessive disorder caused by loss of retinoschisin function, gene replacement therapy is a potential treatment for this disease. Furthermore, because retinoschisin functions as an extracellular protein, beneficial treatments are not necessarily limited to transfected cells expressing the replacement gene but can encompass a wider range of tissues as the secreted protein spreads from the expression site.

[0121] Provided herein are nucleic acid constructs and compositions that allow the insertion of retinoschisin coding sequences into target genomic loci, such as endogenous RS1 loci, and / or the expression of retinoschisin coding sequences.The nucleic acid constructs and compositions can be used in methods for integration into target genomic loci and / or intracellular expression, or in methods for treating X-linked juvenile retinoschisis.Nucleic acid constructs (e.g., targeting endogenous RS1 loci) or nucleic acids encoding nuclease agents are also provided, which facilitate the integration of nucleic acid constructs into target genomic loci, such as endogenous RS1 loci.

[0122] Integration of a nucleic acid construct into the endogenous RS1 locus, such as intron 1 of RS1, can prevent transcription of the endogenous RS1 gene downstream of the integration site. Integration of a nucleic acid construct into the endogenous RS1 locus can reduce or eliminate expression of the endogenous retinoschisin protein (e.g., an endogenous retinoschisin protein having a mutation that causes XLRS) and replace it with expression of a retinoschisin protein encoded by the nucleic acid construct, or a fragment or variant thereof (e.g., a retinoschisin without the mutation that causes XLRS). In one example, integration of a nucleic acid construct into the endogenous RS1 locus reduces expression of the endogenous retinoschisin protein. In another example, integration of a nucleic acid construct into the endogenous RS1 locus eliminates expression of the endogenous retinoschisin protein. In this way, integration of the nucleic acid construct can simultaneously knock out the endogenous RS1 gene (e.g., an endogenous RS1 gene containing one or more mutations associated with or causing XLRS, such as R141C) and knock in a replacement retinoschisin coding sequence (e.g., a replacement retinoschisin coding sequence that does not contain a mutation associated with or causing XLRS).

[0123] II. Nucleic Acid Constructs Comprising Retinoschisin Coding Sequences for Integration into and Expression from Target Genomic Loci Provided herein is a nucleic acid construct (i.e., an exogenous donor nucleic acid) comprising a retinoschisin coding sequence (i.e., encoding a retinoschisin protein or a fragment or variant thereof) for integration into and expression from a target genomic locus. The nucleic acid construct may be an isolated nucleic acid construct.

[0124] Retinoschisin (X-linked juvenile retinoschisis protein) is a protein required for the normal structure and function of the retina. An exemplary human retinoschisin protein has been assigned UniProt accession number O15537 and has the sequence set forth in SEQ ID NO: 2. Orthologs in other species are also known. For example, an exemplary mouse retinoschisin protein has been assigned UniProt accession number Q9Z1L4 and has the sequence set forth in SEQ ID NO: 1. Retinoschisin is encoded by the RS1 gene (also known as XLRS1). The human RS1 gene contains six distinct exons spaced by five introns. The human RS1 gene has been assigned NCBI GeneID 6247. The mouse Rs1 gene has been assigned NCBI GeneID 20147. An exemplary coding sequence for human RS1 has been assigned CCDS ID CCDS14187.1 and is set forth in SEQ ID NO: 6. Mutations in retinoschisis cause X-linked juvenile retinoschisis (XLRS), a vitreoretinal dystrophy characterized by macular lesions and superficial retinal splitting. The nucleic acid constructs disclosed herein can be used in methods for treating XLRS, as described in more detail elsewhere herein.

[0125] The functional domains of RS1 are the signal peptide (SP), RS1, and discoidin domains. The signal sequence directs the translocation of nascent RS1 from the endoplasmic reticulum (site of synthesis) to the outer leaflet of the plasma membrane, during which the signal sequence is cleaved by signal peptidases to generate a mature protein with the characteristic RS1 and highly conserved discoidin domains. The distinct subdomains of the RS1 signal sequence are the amino-terminal positively charged N region that mediates translocation, the hydrophobic core (H) required for targeting and membrane insertion, and the polar "C" region that determines the site of recognition and cleavage by the signal peptidase. RS1 is prominently expressed by retinal photoreceptors and bipolar cells and is also present in the pineal gland.

[0126] The retinoschisin coding sequence contained in the nucleic acid construct disclosed herein can be the coding sequence for the full-length retinoschisin protein or a fragment or variant thereof. In one example, the retinoschisin coding sequence contained in the nucleic acid construct does not include the first exon of RS1. For example, the retinoschisin coding sequence contained in the nucleic acid construct can include exons 2-6 of the RS1 gene, or a variant or degenerate variant thereof. As an example, a cDNA fragment containing exons 2-6 of the RS1 gene can include the sequence set forth in SEQ ID NO: 8. The genetic code is degenerate (i.e., degenerate) because each of the 64 codons is specific for only one amino acid or stop signal, but a single amino acid may be coded for by more than one codon. Degenerate variants of a gene encode the same protein but use at least one different codon. The retinoschisin coding sequence in the nucleic acid construct can comprise complementary DNA (cDNA) without intervening introns, or the nucleic acid construct can include one or more introns separating the exons in the retinoschisin coding sequence. For example, a nucleic acid construct can include sequences corresponding to the RS1 genomic locus, including both exons and introns.

[0127] Retinoschisin coding sequence can be from any organism.For example, retinoschisin coding sequence can be mammal, non-human mammal, rodent, mouse, rat, or human or variant thereof.Alternatively, retinoschisin coding sequence can be chimeric (for example, part mouse and part human).In a specific example, retinoschisin coding sequence is human retinoschisin coding sequence.

[0128] The retinoschisin coding sequence can be codon-optimized for efficient translation into retinoschisin in a particular cell or organism. As an example, a codon-optimized version of exons 2-6 of human RS1 is set forth in SEQ ID NO: 9. For example, the nucleic acid can be modified to alternate codons more frequently used in human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest.

[0129] The retinoschisin coding sequence can encode a wild-type retinoschisin protein or a fragment or variant thereof. Similarly, the retinoschisin coding sequence can be a wild-type coding sequence or a variant thereof. In one example, the retinoschisin coding sequence does not contain a mutation associated with or causing X-linked juvenile retinoschisis. Alternatively, the retinoschisin coding sequence can contain one or more mutations associated with or causing X-linked juvenile retinoschisis (e.g., R141C).

[0130] The nucleic acid construct can further comprise one or more RS1 introns or fragments or variants thereof (e.g., one or more human RS1 introns or fragments or variants thereof). For example, the nucleic acid construct can comprise RS1 intron 1 or a fragment or variant thereof. The RS1 intron or a fragment or variant thereof can comprise a splice acceptor site or a fragment thereof. Examples of fragments of RS1 intron 1 are set forth in SEQ ID NOs: 15 and 16. In one specific example, the nucleic acid construct can comprise RS1 intron 1 or a fragment or variant thereof located 5' of exons 2-6 of RS1 (e.g., upstream of a cDNA sequence comprising, consisting essentially of, or consisting of exons 2-6 of RS1).

[0131] The nucleic acid construct can further comprise one or more splice acceptor sites. Examples of sequences (e.g., intron sequences) containing splice acceptor sites and their reverse complements are shown in SEQ ID NOs: 15-21. For example, the nucleic acid construct can comprise a splice acceptor site located 5' of the retinoschisin coding sequence. In a specific example, the retinoschisin coding sequence comprises, consists essentially of, or consists of exons 2-6 of RS1 (e.g., exons 2-6 of human RS1), and the splice acceptor site is a splice acceptor site from intron 1 of RS1 (e.g., human RS1) used to splice RS1 exon 1 to RS1 exon 2. The term splice acceptor site refers to a nucleic acid sequence at the 3' intron / exon boundary that can be recognized and bound by the splicing machinery.

[0132] The nucleic acid constructs disclosed herein can also include post-transcriptional regulatory elements, such as the woodchuck hepatitis virus post-transcriptional regulatory elements.

[0133] The nucleic acid construct can further include one or more polyadenylation signal sequences. Examples of polyadenylation signal sequences, or sequences containing polyadenylation signal sequences, or their reverse complements, are set forth in SEQ ID NOS: 22-25. For example, the nucleic acid construct can include a polyadenylation signal sequence located 3' of the retinoschisin coding sequence. Any suitable polyadenylation signal sequence can be used. The term polyadenylation signal sequence refers to any sequence that directs the termination of transcription and the addition of a poly(A) tail to an mRNA transcript. In eukaryotes, transcription terminators are recognized by protein factors, and termination is followed by polyadenylation, the process of adding a poly(A) tail to an mRNA transcript in the presence of poly(A) polymerase. Mammalian poly(A) signals typically consist of a core sequence approximately 45 nucleotides long, which may be flanked by various auxiliary sequences that help increase the efficiency of cleavage and polyadenylation. The core sequence, called the poly(A) recognition motif or poly(A) recognition sequence, consists of a highly conserved upstream element (AATAAA or AAUAAA) in mRNA that is recognized by the cleavage and polyadenylation specificity factor (CPSF), and a poorly defined downstream region (rich in U or G and U) that is bound by the cleavage stimulatory factor (CstF). Examples of transcription terminators that can be used include the human growth hormone (HGH) polyadenylation signal, the simian virus 40 (SV40) late polyadenylation signal, the rabbit beta globin polyadenylation signal, the bovine growth hormone (BGH) polyadenylation signal, the phosphoglycerate kinase (PGK) polyadenylation signal, the AOX1 transcription termination sequence, the CYC1 transcription termination sequence, or any transcription termination sequence known to be suitable for regulating gene expression in eukaryotic cells.

[0134] The nucleic acid construct can also include a polyadenylation signal sequence upstream of the retinoschisin coding sequence. The polyadenylation signal sequence upstream of the retinoschisin coding sequence can be adjacent to a recombinase recognition site recognized by a site-specific recombinase. In some constructs, the recombinase recognition site also flanks a selection cassette containing, for example, a coding sequence for a drug resistance protein. In some constructs, the recombinase recognition site does not flank the selection cassette. The polyadenylation signal sequence prevents transcription and expression of the protein or RNA encoded by the coding sequence. However, upon exposure to a site-specific recombinase, the polyadenylation signal sequence is excised, allowing the protein or RNA to be expressed.

[0135] Such a configuration can allow tissue-specific or developmental stage-specific expression in animals containing the retinoschisin coding sequence when the polyadenylation signal sequence is excised in a tissue-specific or developmental stage-specific manner.Excision of the polyadenylation signal sequence in a tissue-specific or developmental stage-specific manner can be achieved when the animal containing the nucleic acid construct further contains a coding sequence for a site-specific recombinase operably linked to a tissue-specific or developmental stage-specific promoter.The polyadenylation signal sequence is then excised only in those tissues or those developmental stages, allowing tissue-specific or developmental stage-specific expression.In one example, the retinoschisin or its fragment or variant encoded by the nucleic acid construct can be expressed in an eye-specific or retinal cell-specific manner.

[0136] Site-specific recombinases include enzymes that can promote recombination between recombinase recognition sites, where the two recombination sites are physically separated within a single nucleic acid or on separate nucleic acids. Examples of recombinases include Cre, Flp, and Dre recombinases. One example of a Cre recombinase gene is Crei, in which the two exons encoding Cre recombinase are separated by an intron, preventing expression in prokaryotic cells. Such recombinases can further contain a nuclear localization signal (e.g., NLS-Crei) to facilitate nuclear localization. Recombinase recognition sites include nucleotide sequences that can be recognized by a site-specific recombinase and serve as a substrate for a recombination event. Examples of recombinase recognition sites include FRT, FRT11, FRT71, attp, att, rox, and lox sites, such as loxP, lox511, lox2272, lox66, lox71, loxM2, and lox5171.

[0137] The nucleic acid construct can further include a promoter operably linked to the retinoschisin-encoding sequence. The retinoschisin-encoding sequence in the nucleic acid construct can be operably linked to any suitable promoter for expression in vivo in an animal or in vitro in isolated cells. The promoter can be a constitutively active promoter (e.g., a CAG promoter or a U6 promoter), a conditional promoter, an inducible promoter, a temporally restricted promoter (e.g., a developmentally regulated promoter), or a spatially restricted promoter (e.g., a cell-specific or tissue-specific promoter). Such promoters are well known and are discussed elsewhere herein. Promoters that can be used in the expression construct include, for example, promoters active in one or more of eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, rabbit cells, pluripotent cells, ocular cells, retinal cells, embryonic stem (ES) cells, or zygotes. In a specific example, the promoter is active in ocular or retinal cells.

[0138] Alternatively, some nucleic acid constructs do not include a promoter operably linked to the retinoschisin coding sequence (e.g., some nucleic acid constructs are promoterless constructs). Such nucleic acid constructs can be designed, for example, to be operably linked to an endogenous promoter of the target genomic locus (e.g., an endogenous RS1 promoter of an endogenous RS1 locus) upon integration into the target genomic locus.

[0139] Any target genomic locus capable of gene expression can be used, such as a safe harbor locus (safe harbor gene) or the endogenous RS1 locus. Interactions between the integrated exogenous DNA and the host genome can limit the reliability and safety of integration, potentially resulting in overt phenotypic effects due not to targeted gene modification but to unintended effects of integration on surrounding endogenous genes. For example, randomly inserted transgenes are susceptible to position effects and silencing, potentially making their expression unreliable and unpredictable. Similarly, integration of exogenous DNA into a chromosomal locus can affect surrounding endogenous genes and chromatin, thereby altering cellular behavior and phenotype. Safe harbor loci include chromosomal loci where a transgene or other exogenous nucleic acid insert can be stably and reliably expressed in all tissues of interest without overtly altering cellular behavior or phenotype (i.e., without adversely affecting the host cell). See, for example, Sadelain et al. (2012) Nat. Rev. Cancer 12:51-58, incorporated herein by reference in its entirety for all purposes. For example, a safe harbor locus can be a locus where the expression of an inserted gene sequence is not disrupted by read-through expression from adjacent genes. For example, a safe harbor locus can include a chromosomal locus where exogenous DNA can be integrated and function in a predictable manner without adversely affecting the structure or expression of endogenous genes. A safe harbor locus can include extragenic or intragenic regions, for example, loci within genes that are non-essential, unnecessary, or can be disrupted without obvious phenotypic consequences.

[0140] Such safe harbor loci can provide open chromatin structure in all tissues and can be ubiquitously expressed during embryonic development and in adults.For example, see Zambrowicz et al.(1997)Proc.Natl.Acad.Sci.USA94:3789-3794, the entire contents of which are incorporated herein by reference for all purposes.In addition, safe harbor loci can be targeted with high efficiency, and can be disrupted without obvious phenotypes.Examples of safe harbor loci include albumin, CCR5, HPRT, AAVS1 and Rosa26. See, for example, U.S. Patent Nos. 7,888,121, 7,972,854, 7,914,796, 7,951,925, 8,110,379, 8,409,861, and 8,586,526, as well as U.S. Patent Publication Nos. 2003 / 0232410, 2005 / 0208489, 2005 / 0208489, and 2006 / 0209410, each of which is incorporated by reference in its entirety for all purposes. See Patent Publications 2006 / 0063231, 2008 / 0159996, 2010 / 00218264, 2012 / 0017290, 2011 / 0265198, 2013 / 0137104, 2013 / 0122591, 2013 / 0177983, 2013 / 0177960, and 2013 / 0122591.

[0141] The target genomic locus can also be an endogenous RS1 locus, such as an endogenous RS1 locus containing one or more mutations associated with or causing XLRS (e.g., an R141C mutation in the encoded retinoschisin protein). Integration of a nucleic acid construct into the endogenous RS1 locus can, in some cases, disrupt transcription of the endogenous RS1 gene downstream of the integration site. Integration of a nucleic acid construct into the endogenous RS1 locus can reduce or eliminate expression of the endogenous retinoschisin protein and replace it with expression of a retinoschisin protein or a fragment or variant thereof encoded by the nucleic acid construct. In one example, integration of a nucleic acid construct into the endogenous RS1 locus reduces expression of the endogenous retinoschisin protein. In another example, integration of a nucleic acid construct into the endogenous RS1 locus eliminates expression of the endogenous retinoschisin protein. In this way, integration of the nucleic acid construct can simultaneously knock out the endogenous RS1 gene (e.g., an endogenous RS1 gene containing one or more mutations associated with or causing XLRS) and knock in a replacement retinoschisin coding sequence (e.g., a replacement retinoschisin coding sequence that does not contain a mutation associated with or causing XLRS).

[0142] The nucleic acid construct can be integrated into any part of the target genome locus. For example, the nucleic acid construct can be inserted into an intron or exon of the target genome locus, or can replace one or more introns and / or exons of the target genome locus. In a specific example, the nucleic acid construct can be integrated into an intron of the target genome locus, such as the first intron (e.g., RS1 intron 1) of the target genome locus. The expression cassette integrated into the target genome locus can be operably linked to an endogenous promoter (e.g., endogenous RS1 promoter) at the target genome locus, or can be operably linked to an exogenous promoter (e.g., CMV promoter) that is heterologous to the target genome locus.

[0143] Nucleic acid constructs can contain deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), can be single-stranded or double-stranded, and can be in linear or circular form. For example, the nucleic acid construct can be a single-stranded oligodeoxynucleotide (ssODN). See, for example, Yoshimi et al. (2016) Nat. Commun. 7:10431, incorporated herein by reference in its entirety for all purposes. The nucleic acid construct can be a naked nucleic acid or can be delivered by a vector such as an AAV vector. In a specific example, the nucleic acid construct can be delivered via AAV and inserted into the endogenous RS1 locus by non-homologous end joining (e.g., the nucleic acid construct can be one that does not contain homology arms). When introduced in linear form, the ends of the nucleic acid construct (e.g., donor sequence) can be protected (e.g., from exonuclease degradation) by well-known methods. For example, one or more dideoxynucleotide residues can be added to the 3'-end of linear molecule, and / or self-complementary oligonucleotides can be linked to one or both ends.See, for example, Chang et al.(1987) Proc.Natl.Acad.Sci.USA 84:4959-4963 and Nehls et al.(1996) Science 272:886-889, each of which is incorporated herein by reference in its entirety for all purposes.Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, adding terminal amino groups and using modified internucleotide linkages, such as phosphorothioate, phosphoramidate, and O-methylribose or deoxyribose residues.

[0144] Exemplary nucleic acid constructs are from about 50 nucleotides to about 5 kb in length, or from about 50 nucleotides to about 3 kb in length. Alternatively, the nucleic acid constructs can be from about 1 kb to about 1.5 kb, from about 1.5 kb to about 2 kb, from about 2 kb to about 2.5 kb, from about 2.5 kb to about 3 kb, from about 3 kb to about 3.5 kb, from about 3.5 kb to about 4 kb, from about 4 kb to about 4.5 kb, or from about 4.5 kb to about 5 kb in length. Alternatively, the nucleic acid constructs can be, for example, 5 kb, 4.5 kb, 4 kb, 3.5 kb, 3 kb, or 2.5 kb or less in length.

[0145] Integration of the nucleic acid construct at the target genomic locus can result in the addition of a nucleic acid sequence of interest to the target genomic locus, or the replacement (i.e., deletion and insertion) of a nucleic acid sequence of interest at the target genomic locus. Some nucleic acid constructs are designed for the insertion of a nucleic acid construct at the target genomic locus without a corresponding deletion at the target genomic locus. Other nucleic acid constructs are designed to delete a nucleic acid sequence of interest at the target genomic locus and replace it with the nucleic acid construct.

[0146] The nucleic acid construct or corresponding nucleic acid at the target genomic locus to be deleted and / or replaced can be of various lengths. Exemplary nucleic acid constructs or corresponding nucleic acids at the target genomic locus to be deleted and / or replaced are from about 1 nucleotide to about 5 kb in length, or from about 1 nucleotide to about 3 kb in length. For example, the nucleic acid construct or corresponding nucleic acid at the target genomic locus to be deleted and / or replaced can be from about 1 to about 100, from about 100 to about 200, from about 200 to about 300, from about 300 to about 400, from about 400 to about 500, from about 500 to about 600, from about 600 to about 700, from about 700 to about 800, from about 800 to about 900, or from about 900 to about 1,000 nucleotides in length. Similarly, the nucleic acid construct or corresponding nucleic acid at the target genomic locus to be deleted and / or replaced can be about 1 kb to 1.5 kb, about 1.5 kb to about 2 kb, about 2 kb to about 2.5 kb, about 2.5 kb to about 3 kb, about 3 kb to about 3.5 kb, about 3.5 kb to about 4 kb, about 4 kb to about 4.5 kb, about 4.5 kb to about 5 kb, or longer in length.

[0147] The nucleic acid construct or corresponding nucleic acid of the target genomic locus to be deleted and / or replaced can be a coding region such as an exon, an intron, an untranslated region, or a non-coding region such as a regulatory region (e.g., a promoter, enhancer, or transcriptional repressor binding element), or any combination thereof.

[0148] The nucleic acid construct optionally comprises one or more of the following terminal structures: a hairpin, a loop, an inverted terminal repeat (ITR), or a toroid. For example, the nucleic acid construct can comprise an ITR.

[0149] Some such nucleic acid constructs can modify a target genomic locus (e.g., but not limited to, the endogenous RS1 locus) following cleavage or nicking of the target genomic locus by a nuclease agent such as a Cas protein. The nucleic acid construct can be designed to repair the cleaved or nicked locus via ligation or homology-directed repair via non-homologous end joining (NHEJ). Optionally, repair by the nucleic acid construct removes or destroys the nuclease target sequence, such that the targeted allele cannot be retargeted by the nuclease agent.

[0150] Some nucleic acid constructs contain homology arms. The homology arms can be symmetrical (for example, each 40 nucleotides or each 60 nucleotides long), or asymmetrical (for example, one homology arm or complementary region is 36 nucleotides long, and one homology arm or complementary region is 91 nucleotides long). Other nucleic acid constructs do not contain homology arms.

[0151] Some nucleic acid constructs disclosed herein contain homology arms. The homology arms can flank the retinoschisin coding sequence. For ease of reference, the homology arms are referred to herein as 5' and 3' (i.e., upstream and downstream) homology arms. This term refers to the relative position of the homology arms to the nucleic acid insert (e.g., the retinoschisin coding sequence) within the nucleic acid construct. The 5' and 3' homology arms correspond to regions within the target genomic locus, and are referred to herein as the "5' target sequence" and the "3' target sequence," respectively.

[0152] A homology arm and a target sequence are "corresponding" or "corresponding" to each other if the two regions share a sufficient level of sequence identity to act as substrates for a homologous recombination reaction. The term "homology" includes DNA sequences that are identical to or share sequence identity with the corresponding sequence. The sequence identity between a given target sequence and the corresponding homology arm found in a nucleic acid construct can be any degree of sequence identity that allows homologous recombination to occur. For example, the amount of sequence identity shared by the homology arms of a nucleic acid construct (or fragment thereof) and the target sequence (or fragment thereof) can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity, allowing the sequences to undergo homologous recombination. Furthermore, the corresponding homologous regions between the homology arms and the corresponding target sequence can be any length sufficient to promote homologous recombination. Exemplary homology arms are about 25 nucleotides to about 2.5 kb in length, about 25 nucleotides to about 1.5 kb in length, or about 25 to about 500 nucleotides in length. For example, a given homology arm (or each of the homology arms) and / or corresponding target sequence can comprise a corresponding homology region that is about 25 to about 30, about 30 to about 40, about 40 to about 50, about 50 to about 60, about 60 to about 70, about 70 to about 80, about 80 to about 90, about 90 to about 100, about 100 to about 150, about 150 to about 200, about 200 to about 250, about 250 to about 300, about 300 to about 350, about 350 to about 400, about 400 to about 450, or about 450 to about 500 nucleotides in length, such that the homology arm has sufficient homology to undergo homologous recombination with the corresponding target sequence within the target nucleic acid. Alternatively, a given homology arm (or each homology arm) and / or corresponding target sequence can comprise a corresponding homology region that is about 0.5 kb to about 1 kb, about 1 kb to about 1.5 kb, about 1.5 kb to about 2 kb, or about 2 kb to about 2.5 kb in length. For example, the homology arms can each be about 750 nucleotides in length.In another example, the homology arms can each be about 150 to about 750, about 200 to about 700, about 250 to about 650, about 300 to about 600, about 350 to about 550, about 400 to about 500, about 150 to about 450, about 200 to about 450, about 250 to about 450, about 300 to about 450, about 350 to about 450, about 400 to about 450, about 450 to about 500, about 450 to about 550, about 450 to about 600, about 450 to about 650, about 450 to about 700, about 450 to about 750, or about 450 nucleotides in length. In another example, the homology arms each have a length of about 500 to about 1300, about 550 to about 1250, about 600 to about 1200, about 650 to about 1150, about 700 to about 1100, about 750 to about 1050, about 800 to about 1000, about 850 to about 950, about 500 to about 900, about 550 to about 900, about 600 to about 900, about 650 to about 900, about It can be 700 to about 900, about 750 to about 900, about 800 to about 900, about 850 to about 900, about 900 to about 950, about 900 to about 1000, about 900 to about 1050, about 900 to about 1100, about 900 to about 1150, about 900 to about 1200, about 900 to about 1250, about 900 to about 1300, or about 900 nucleotides in length. In another example, the homology arms can each be about 1500 to about 2100, about 1550 to about 2050, about 1600 to about 2000, about 1650 to about 1950, about 1700 to about 1900, about 1750 to about 1850, about 1500 to about 1800, about 1550 to about 1800, about 1600 to about 1800, about 1650 to about 1800, about 1700 to about 1800, about 1750 to about 1800, about 1800 to about 1850, about 1800 to about 1900, about 1800 to about 1950, about 1800 to about 2000, about 1800 to about 2050, about 1800 to about 2100, or about 1800 nucleotides. In another example, each homology arm is about 450 nucleotides or less, about 900 nucleotides or less, or about 1800 nucleotides or less. In another example, each homology arm is at least about 450 nucleotides, at least about 900 nucleotides, or at least about 1800 nucleotides. The homology arms can be symmetric (each about the same size in length) or asymmetric (one longer than the other).

[0153] When a CRISPR / Cas system or other nuclease agent is used in combination with the nucleic acid constructs disclosed herein, the 5' and 3' target sequences can be positioned sufficiently close to the nuclease cleavage site (e.g., sufficiently close to the guide RNA target sequence) to promote the occurrence of a homologous recombination event between the target sequence and the homology arm upon a single-strand break (nick) or double-strand break at the nuclease cleavage site. The term "nuclease cleavage site" includes the DNA sequence where a nick or double-strand break is created by a nuclease agent (e.g., a Cas9 protein complexed with a guide RNA). Target sequences within the targeted locus corresponding to the 5' and 3' homology arms of a nucleic acid construct are "sufficiently close" to a nuclease cleavage site if the distance is such that the distance promotes the occurrence of a homologous recombination event between the 5' and 3' target sequences and the homology arm upon a single-strand break or double-strand break at the nuclease cleavage site. Thus, the target sequence corresponding to the 5' and / or 3' homology arms of the nucleic acid construct can be, for example, within at least one nucleotide of a given nuclease cleavage site, or within at least 10 nucleotides to about 1,000 nucleotides of a given nuclease cleavage site. As an example, the nuclease cleavage site can be immediately adjacent to at least one or both of the target sequences.

[0154] The positional relationship between the target sequence corresponding to the homology arm of the nucleic acid construct and the nuclease cleavage site can vary. For example, the target sequence can be located 5' of the nuclease cleavage site, the target sequence can be located 3' of the nuclease cleavage site, or the target sequence can be adjacent to the nuclease cleavage site.

[0155] Other nucleic acid constructs do not contain any homology arms. Such nucleic acid constructs may be inserted by non-homologous end joining. For example, such nucleic acid constructs can be inserted into blunt-ended double-strand breaks after cleavage with a nuclease agent. In a specific example, the nucleic acid constructs can be delivered via AAV and inserted into target genomic loci by non-homologous end joining (for example, the nucleic acid constructs may not contain homology arms).

[0156] In a specific example, the nucleic acid construct can be inserted via homology-independent targeted integration. For example, the retinoschisin coding sequence of the nucleic acid construct can be flanked on each side by target sites for a nuclease agent (e.g., the same target site as the target genomic locus, and the same nuclease agent is used to cleave the target site in the target genomic locus). The nuclease agent can then cleave the target sites adjacent to the retinoschisin coding sequence. In a specific example, the nucleic acid construct is delivered using AAV-mediated delivery, and cleavage of the target sites adjacent to the retinoschisin coding sequence can remove the AAV inverted terminal repeats (ITRs). In some methods, when the retinoschisin coding sequence is inserted into the target genomic locus in the correct orientation, the target site (e.g., the gRNA target sequence including the adjacent protospacer adjacent motif) in the target genomic locus no longer exists, but is re-formed when the retinoschisin coding sequence is inserted into the target genomic locus in the opposite orientation. This can help ensure that the retinoschisin coding sequence is inserted in the correct orientation for expression.

[0157] In one exemplary nucleic acid construct for homology-independent targeted integration into a target genome locus, the retinoschisin protein or fragment thereof is a human retinoschisin protein or fragment thereof, the coding sequence of the retinoschisin protein or fragment thereof comprises a complementary DNA (cDNA) containing exons 2 to 6 of human RS1 or a degenerate variant thereof, the nucleic acid construct does not comprise a promoter driving expression of the retinoschisin protein or fragment thereof, the nucleic acid construct comprises a polyadenylation signal sequence located 3' of the coding sequence, the nucleic acid construct comprises a splice acceptor site located 5' of the coding sequence, the nuclease target sequence in the nucleic acid construct is identical to the nuclease target sequence for integration into the target genome locus, and the nuclease target sequence in the target genome locus is destroyed when the nucleic acid construct is inserted in the correct orientation but is re-formed when the nucleic acid construct is inserted into the target genome locus in the opposite orientation.

[0158] In one exemplary nucleic acid construct for homology-independent targeted integration into a target genomic locus, the retinoschisin protein or fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 2 or 5. In one exemplary nucleic acid construct for homology-independent targeted integration into a target genomic locus, the coding sequence of the retinoschisin protein or fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 6, 8, or 9, or a degenerate variant thereof. In one exemplary nucleic acid construct for homology-independent targeted integration into a target genomic locus, the nucleic acid construct comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO:45.

[0159] Other nucleic acid constructs have short single-stranded regions at the 5' and / or 3' ends that are complementary to one or more overhangs created by nuclease agent-mediated cleavage at the target genomic locus. For example, some nucleic acid constructs have short single-stranded regions at the 5' and / or 3' ends that are complementary to one or more overhangs created by nuclease-mediated cleavage at the 5' and / or 3' target sequence at the target genomic locus. Some such nucleic acid constructs have complementary regions only at the 5' end or only at the 3' end. For example, some such nucleic acid constructs have complementary regions only at the 5' end that are complementary to the overhang created at the 5' target sequence at the target genomic locus, or only at the 3' end that are complementary to the overhang created at the 3' target sequence at the target genomic locus. Other such nucleic acid constructs have complementary regions at both the 5' and 3' ends. For example, other such nucleic acid constructs have complementary regions (e.g., complementary to the first and second overhangs, respectively) at both the 5' and 3' ends generated by nuclease-mediated cleavage at the target genomic locus. For example, if the nucleic acid construct is double-stranded, a single-stranded complementary region can extend from the 5' end of the top strand of the donor nucleic acid and the 5' end of the bottom strand of the nucleic acid construct, creating 5' overhangs on both ends. Alternatively, a single-stranded complementary region can extend from the 3' end of the top strand of the nucleic acid construct and from the 3' end of the bottom strand of the template, creating 3' overhangs.

[0160] The complementary region can be of any length sufficient to promote ligation between the nucleic acid construct and the target nucleic acid. Exemplary complementary regions are about 1 to about 5 nucleotides in length, about 1 to about 25 nucleotides in length, or about 5 to about 150 nucleotides in length. For example, the complementary region can be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length. Alternatively, the complementary region may be about 5 to about 10, about 10 to about 20, about 20 to about 30, about 30 to about 40, about 40 to about 50, about 50 to about 60, about 60 to about 70, about 70 to about 80, about 80 to about 90, about 90 to about 100, about 100 to about 110, about 110 to about 120, about 120 to about 130, about 130 to about 140, about 140 to about 150 nucleotides in length, or longer.

[0161] Such complementary regions can complement the overhangs created by the two pairs of nickases. Using a first and second nickase to cleave opposite strands of DNA to create a first double-stranded break, and a third and fourth nickase to cleave opposite strands of DNA to create a second double-stranded break, two double-stranded breaks can be created, with staggered ends. For example, a Cas protein can be used to nick the first, second, third, and fourth guide RNA target sequences corresponding to the first, second, third, and fourth guide RNAs. The first and second guide RNA target sequences can be positioned to create a first cleavage site (i.e., the first cleavage site includes a nick in the first and second guide RNA target sequences) such that the nicks created by the first and second nickases on the first and second strands of DNA create double-stranded breaks. Similarly, the third and fourth guide RNA target sequences can be positioned to create a second cleavage site (i.e., the second cleavage site includes a nick within the third and fourth guide RNA target sequences) such that nicks created by the third and fourth nickases on the first and second strands of DNA create a double-stranded break. The nicks within the first and second guide RNA target sequences and / or the third and fourth guide RNA target sequences can be offset nicks that create an overhang. The offset window can be, for example, at least about 5 bp, 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, or more. See Ran et al. (2013) Cell 154:1380-1389, Mali et al. (2013) Nat. Biotechnol. 31:833-838, and Shen et al. (2014) Nat. Methods 11:399-404, each of which is incorporated by reference in its entirety for all purposes. In such cases, the double-stranded nucleic acid construct can be designed with single-stranded complementary regions complementary to the overhangs created by the nicks in the first and second guide RNA target sequences and the nicks in the third and fourth guide RNA target sequences.Such nucleic acid constructs can then be inserted by non-homologous end joining mediated ligation.

[0162] Some of the nucleic acid constructs disclosed herein are bidirectional constructs that can be inserted into and expressed from a target genome locus in either direction. Such nucleic acid constructs can include a first segment comprising a first coding sequence of a first retinoschisin protein or a fragment or variant thereof, and a second segment comprising a reverse complement of a second coding sequence of a second retinoschisin protein or a fragment or variant thereof. The second segment can be, for example, located 3' from the first segment of the nucleic acid construct.

[0163] The first and second segments can be linked directly together or can be linked by a linker, such as a peptide linker. The peptide linker can be of any suitable length. For example, the linker can be about 5 to about 2,000 nucleotides in length. By way of example, the linker sequence can be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 500, 1,000, 1,500, 2,000, or more nucleotides in length.

[0164] In some bidirectional constructs, the first retinoschisin protein, or fragment or variant thereof, is identical to the second retinoschisin protein, or fragment or variant thereof, In other bidirectional constructs, the first retinoschisin protein, or fragment or variant thereof, is different from the second retinoschisin protein, or fragment or variant thereof.

[0165] In some bidirectional constructs, the codon usage in the first coding sequence is the same as the codon usage in the second coding sequence. In some other bidirectional constructs, the second coding sequence employs a different codon usage than the codon usage of the first coding sequence to reduce hairpin formation. Such a reverse complement forms fewer base pairs than all nucleotides of the coding sequence of the first segment, but can optionally encode the same polypeptide.

[0166] The second segment can have any percentage of complementarity to the first segment. For example, the second segment sequence can have at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, or at least about 99% complementarity to the first segment. As another example, the second segment sequence can have less than about 30%, less than about 35%, less than about 40%, less than about 45%, less than about 50%, less than about 55%, less than about 60%, less than about 65%, less than about 70%, less than about 75%, less than about 80%, less than about 85%, less than about 90%, less than about 95%, less than about 97%, or less than about 99% complementarity to the first segment. In some nucleic acid constructs, the reverse complement of the second coding sequence may be substantially non-complementary to the first coding sequence (e.g., 70% or less complementary), not substantially complementary to a fragment of the first coding sequence, highly complementary to the first coding sequence (e.g., at least 90% complementary), highly complementary to a fragment of the first coding sequence, about 50% to about 80% identical to the reverse complement of the first coding sequence, or about 60% to about 100% identical to the reverse complement of the first coding sequence.

[0167] Bidirectional constructs can optionally include one or more (e.g., two) polyadenylation signal sequences. In some bidirectional constructs, the first segment can include a first polyadenylation signal sequence. In some bidirectional constructs, the first segment can include a second polyadenylation signal sequence. In some bidirectional constructs, the first segment can include a first polyadenylation signal sequence and the second segment can include a second polyadenylation signal sequence (e.g., the reverse complement of the polyadenylation signal sequence). In some bidirectional constructs, the first segment can include a first polyadenylation signal sequence located 3' of the first coding sequence. In some bidirectional constructs, the second segment can include the reverse complement of the second polyadenylation signal sequence located 5' of the reverse complement of the second coding sequence. In some bidirectional constructs, the first segment can include a first polyadenylation signal sequence located 3' of the first coding sequence, and the second segment can include the reverse complement of a second polyadenylation signal sequence located 5' of the reverse complement of the second coding sequence. The first and second polyadenylation signal sequences can be the same or different. In one example, the first and second polyadenylation signals are different.

[0168] Bidirectional constructs can optionally include one or more (e.g., two) splice acceptor sites. In some bidirectional constructs, the first segment can include a first splice acceptor site. In some bidirectional constructs, the first segment can include a second splice acceptor site. In some bidirectional constructs, the first segment can include a first splice acceptor site and the second segment can include a second splice acceptor site (e.g., the reverse complement of the splice acceptor site). In some bidirectional constructs, the first segment includes a first splice acceptor site located 5' of the first coding sequence. In some bidirectional constructs, the second segment includes the reverse complement of the second splice acceptor site located 3' of the reverse complement of the second coding sequence. In some bidirectional constructs, the first segment comprises a first splice acceptor site located 5' of the first coding sequence, and the second segment comprises a reverse complement of a second splice acceptor site located 3' of the reverse complement of the second coding sequence. The first and second splice acceptor sites can be the same or different. In one example, the first and second splice acceptor sites are different. The first and / or second splice acceptor sites can be derived from an RS1 gene, such as the human RS1 gene (e.g., from intron 1 of the RS1 gene).

[0169] Some bidirectional constructs can include a promoter driving expression of a first retinoschisin protein, or fragment or variant thereof, and / or the reverse complement of a promoter driving expression of a second retinoschisin protein, or fragment or variant thereof. Alternatively, a bidirectional construct can be a construct that does not include a promoter driving expression of the first retinoschisin protein, or fragment or variant thereof, or the second retinoschisin protein, or fragment or variant thereof (i.e., a promoterless construct).

[0170] One or both of the coding sequences may be codon-optimized for expression in a host cell. In some bidirectional constructs, only one of the coding sequences is codon-optimized. In some bidirectional constructs, the first coding sequence is codon-optimized. In some bidirectional constructs, the second coding sequence is codon-optimized. In some bidirectional constructs, both coding sequences are codon-optimized.

[0171] In an exemplary bidirectional construct, the second segment is located 3' of the first segment, both the first retinoschisin protein or fragment thereof and the second retinoschisin protein or fragment thereof are human retinoschisin protein or fragment thereof, the first retinoschisin protein or fragment thereof is identical to the second retinoschisin protein or fragment thereof, both the first coding sequence and the second coding sequence comprise complementary DNA (cDNA) containing exons 2-6 of human RS1 or a degenerate variant thereof, the second coding sequence employs codon usage that differs from the codon usage of the first coding sequence, and the first segment is the second segment comprises the reverse complement of a second polyadenylation signal sequence located 5' of the reverse complement of the second coding sequence; the first segment comprises a first splice acceptor site located 5' of the first coding sequence; and the second segment comprises the reverse complement of a second splice acceptor site located 3' of the reverse complement of the second coding sequence; the nucleic acid construct does not comprise a promoter driving expression of the first retinoschisin protein or fragment thereof or the second retinoschisin protein or fragment thereof; and optionally, the nucleic acid construct does not comprise a homology arm.

[0172] In exemplary bidirectional constructs, the first retinoschisin protein or fragment thereof and / or the second retinoschisin protein or fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 2 or 5. In exemplary bidirectional constructs, the first coding sequence and / or the second coding sequence comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 6, 8, or 9, or a degenerate variant thereof. In exemplary bidirectional constructs, the first coding sequence comprises, consists essentially of, or consists of a sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO:8, and the second coding sequence comprises, consists essentially of, or consists of a sequence at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO:9. In exemplary bidirectional constructs, the nucleic acid construct comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 46 or 47.

[0173] Nucleic acid constructs may contain modifications or sequences that provide additional desirable characteristics (e.g., modified or controlled stability, tracking or detection by fluorescent labels, binding sites for proteins or protein complexes, etc.). Nucleic acid constructs may contain one or more fluorescent labels, purification tags, epitope tags, or combinations thereof. For example, nucleic acid constructs may contain one or more fluorescent labels (e.g., fluorescent proteins or other fluorophores or dyes), such as at least one, at least two, at least three, at least four, or at least five fluorescent labels. Exemplary fluorescent labels include fluorophores such as fluorescein (e.g., 6-carboxyfluorescein (6-FAM)), Texas Red, HEX, Cy3, Cy5, Cy5.5, Pacific Blue, 5-(and -6)-carboxytetramethylrhodamine (TAMRA), and Cy7. A wide range of fluorescent dyes for labeling oligonucleotides are commercially available (e.g., from Integrated DNA Technologies). Such fluorescent labels (e.g., internal fluorescent labels) can be used, for example, to detect nucleic acid constructs directly incorporated into cleaved target nucleic acids with overhangs compatible with the ends of the nucleic acid construct. The label or tag can be at the 5' end, 3' end, or internal to the nucleic acid construct. For example, the nucleic acid construct can be conjugated to the 5' end with an IR700 fluorophore (5'IRDYE® 700) from Integrated DNA Technologies.

[0174] The nucleic acid construct can also contain a conditional allele. The conditional allele can be a multifunctional allele, as described in US2011 / 0104799, the entire contents of which are incorporated herein by reference for all purposes. For example, the conditional allele can include: (a) an actuating sequence in a sense orientation relative to transcription of the target gene; (b) a drug selection cassette (DSC) in a sense or antisense orientation; (c) a nucleotide sequence of interest (NSI) in an antisense orientation; and (d) a conditional inversion module (COIN, utilizing modules such as exon-splitting introns and invertible gene traps) in an inverted orientation. See, e.g., US2011 / 0104799. The conditional allele can further include a recombinable unit that recombines upon exposure to a first recombinase to form a conditional allele that (i) lacks the actuating sequence and DSC, and (ii) includes a NSI in a sense orientation and a COIN in an antisense orientation. See, e.g., US2011 / 0104799.

[0175] The nucleic acid construct may also include a polynucleotide encoding a selection marker. Alternatively, the nucleic acid construct may lack a polynucleotide encoding a selection marker. The selection marker may be included in a selection cassette. Optionally, the selection cassette may be a self-deleting cassette. See, for example, US8,697,851 and US2013 / 0312129, each of which is incorporated by reference in its entirety for all purposes. As an example, a self-deleting cassette may include a Crei gene (comprising two exons encoding Cre recombinase separated by an intron) operably linked to the mouse Prm1 promoter and a neomycin resistance gene operably linked to the human ubiquitin promoter. By using the Prm1 promoter, the self-deleting cassette can be specifically deleted in the male germ cells of F0 animals. Exemplary selection markers include neomycin phosphotransferase (neomycin phosphotransferase) and the neomycin resistant gene. r ), hygromycin B phosphotransferase (hyg r), puromycin-N-acetyltransferase (puro r ), blasticidin S deaminase (bsr r ), xanthine / guanine phosphoribosyltransferase (gpt), or herpes simplex virus thymidine kinase (HSV-k), or a combination thereof. The polynucleotide encoding the selection marker can be operably linked to a promoter active in the targeted cells. Examples of promoters are described elsewhere herein.

[0176] The nucleic acid construct can also contain a reporter gene. Exemplary reporter genes include those encoding luciferase, β-galactosidase, green fluorescent protein (GFP), enhanced green fluorescent protein (eGFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), blue fluorescent protein (BFP), enhanced blue fluorescent protein (eBFP), DsRed, ZsGreen, MmGFP, mPlum, mCherry, tdTomato, mStrawberry, J-Red, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, Cerulean, T-Sapphire, and alkaline phosphatase. Such reporter genes can be operably linked to a promoter active in the targeted cells. Examples of promoters are described elsewhere herein.

[0177] The nucleic acid construct may also contain one or more expression or deletion cassettes. A given cassette may contain one or more of a nucleotide sequence of interest, a polynucleotide encoding a selectable marker, and a reporter gene, along with various regulatory components that affect expression. Examples of selectable markers and reporter genes that may be included are discussed in detail elsewhere herein.

[0178] A nucleic acid construct can contain nucleic acids flanked by site-specific recombination target sequences. Alternatively, a nucleic acid construct can contain one or more site-specific recombination target sequences. The entire nucleic acid construct can be flanked by such site-specific recombination target sequences, but any region or individual polynucleotide of interest within the nucleic acid construct can also be flanked by such sites. Site-specific recombination target sequences that can flank a nucleic acid construct or any polynucleotide of interest within a nucleic acid construct can include, for example, loxP, lox511, lox2272, lox66, lox71, loxM2, lox5171, FRT, FRT11, FRT71, attp, att, FRT, rox, or combinations thereof. In one example, the site-specific recombination sites are flanked by polynucleotides encoding selectable markers and / or reporter genes contained within the nucleic acid construct. Following integration of the nucleic acid construct at the targeted locus, the sequences between the site-specific recombination sites can be removed.

[0179] Nucleic acid constructs can also contain one or more restriction sites for restriction endonucleases (i.e., restriction enzymes), including type I, type II, type III, and type IV endonucleases. Type I and type III restriction endonucleases recognize specific recognition sites but typically cleave at variable positions from the nuclease binding site, which can be several hundred base pairs away from the cleavage site (recognition site). In type II systems, restriction activity is independent of methylase activity, and cleavage typically occurs at specific sites within or near the binding site. Most type II enzymes cleave palindromic sequences, while type IIa enzymes recognize nonpalindromic recognition sites and cleave outside the recognition site, type IIb enzymes cleave a sequence twice, at both sites outside the recognition site, and type IIs enzymes recognize asymmetric recognition sites and cleave at a defined distance of approximately 1–20 nucleotides from the recognition site. Type IV restriction enzymes target methylated DNA. Restriction enzymes are further described and classified, for example, in the REBASE database (webpage at rebase.neb.com; Roberts et al., (2003) Nucleic Acids Res. 31:418-420, Roberts et al., (2003) Nucleic Acids Res. 31:1805-1812, and Belfort et al. (2002) in Mobile DNA II, pp. 761-783, Eds. Craigie et al., (ASM Press, Washington, DC)).

[0180] The nucleic acid constructs disclosed herein can also include additional coding sequences. For example, some nucleic acid constructs disclosed herein can include a sequence encoding a guide RNA that targets a target genomic locus (e.g., targets RS1, such as intron 1 of RS1). The sequence encoding the guide RNA can be operably linked to a promoter, such as a U6 promoter. In some nucleic acid constructs, the guide RNA expression cassette is located 3' (downstream) of the retinoschisin coding sequence. In some bidirectional nucleic acid constructs, the guide RNA expression cassette is located between the first segment and the second segment.

[0181] III. Vectors Containing Nucleic Acid Constructs Also provided herein are vectors comprising a nucleic acid construct (i.e., an exogenous donor nucleic acid) comprising a retinoschisin-encoding sequence (i.e., encoding a retinoschisin protein or a fragment or variant thereof) for integration into and expression from a target genome locus. Also provided herein are vectors comprising a nucleic acid encoding a nuclease agent (e.g., that targets an endogenous RS1 locus) disclosed elsewhere herein. Also provided herein are vectors (e.g., vectors comprising a nucleic acid construct and DNA encoding a guide RNA) comprising a nucleic acid construct and / or a nuclease agent (e.g., that targets an endogenous RS1 locus) disclosed elsewhere herein. Vectors can include additional sequences, such as, for example, an origin of replication, a promoter, and a gene encoding antibiotic resistance. Some such vectors include homology arms corresponding to the target site in the target genome locus. Other such vectors do not include any homology arms.

[0182] Some vectors can be circular. Alternatively, vectors can be linear. Vectors can be packaged to be delivered via lipid nanoparticles, liposomes, non-lipid nanoparticles, or viral capsids. Non-limiting exemplary vectors include plasmids, phagemids, cosmids, artificial chromosomes, mini-chromosomes, transposons, viral vectors, and expression vectors.

[0183] The vector may be a viral vector, such as an adeno-associated virus (AAV) vector. The AAV may be of any suitable serotype and may be single-stranded AAV (ssAAV) or self-complementary AAV (scAAV). Other exemplary viruses / viral vectors include retroviruses, adenoviruses, vaccinia viruses, poxviruses, and herpes simplex viruses. Viruses can infect dividing cells, non-dividing cells, or both dividing and non-dividing cells. Viruses may integrate into the host genome, or alternatively, may not integrate into the host genome. Such viruses may also be engineered to reduce immunity. Viruses may be replication-competent or replication-deficient (e.g., defective in one or more genes required for additional rounds of virion replication and / or packaging). Viruses may induce transient expression, long-term expression (e.g., for at least 1 week, 2 weeks, 1 month, 2 months, or 3 months), or persistent expression (e.g., of Cas9 and / or gRNA). Exemplary viral titers (e.g., AAV titers) are 10 12 , 10 13 , 10 14 , 10 15 , and 10 16 An exemplary viral titer (e.g., AAV titer) is about 10 12 , about 10 13 , about 10 14 , about 10 15 and about 10 16 vector genomes (vg) / mL, or approximately 10 12 ~about 10 16 , about 10 12 ~about 1015 , about 10 12 ~about 10 14 , about 10 12 ~about 10 13 , about 10 13 ~about 10 16 , about 10 14 ~about 10 16 , about 10 15 ~about 10 16 , or about 10 13 ~about 10 15 Other exemplary viral titers (e.g., AAV titers) include about 10 vg / mL. 12 , about 10 13 , about 10 14 , about 10 15 and about 10 16 of vector genome (vg) / kg body weight, or approximately 10 12 ~about 10 16 , about 10 12 ~about 10 15 , about 10 12 ~about 10 14 , about 10 12 ~about 10 13 , about 10 13 ~about 10 16 , about 10 14 ~about 10 16 , about 10 15 ~about 10 16 , or about 10 13 ~about 10 15 vg / kg body weight.

[0184] The ssDNA AAV genome consists of two open reading frames, Rep and Cap, flanked by two inverted terminal repeats that allow for the synthesis of complementary DNA strands. When constructing an AAV transfer plasmid, the transgene is placed between the two ITRs, and Rep and Cap are supplied in trans. In addition to Rep and Cap, AAV may require a helper plasmid containing genes from adenovirus. These genes (E4, E2a, and VA) mediate AAV replication. For example, the transfer plasmid, Rep / Cap, and helper plasmid can be transfected into HEK293 cells containing the adenovirus gene E1+ to produce infectious AAV particles. Alternatively, Rep, Cap, and adenovirus helper genes can be combined into a single plasmid. Similar packaging cells and methods can be used for other viruses, such as retroviruses.

[0185] Several AAV serotypes have been identified. These serotypes differ in the type of cells they infect (i.e., their tropism), allowing for preferential transduction of specific cell types. Serotypes targeting photoreceptor cells include AAV2, AAV5, and AAV8. Serotypes targeting retinal pigment epithelial tissue include AAV1, AAV2, AAV4, AAV5, and AAV8. In a specific example, the AAV vector containing the nucleic acid construct may be AAV2, AAV5, or AAV8.

[0186] Tropism can be further refined through pseudotyping, which involves mixing capsids and genomes from different viral serotypes. For example, AAV2 / 5 refers to a virus containing a serotype 2 genome packaged in a serotype 5 capsid. The use of pseudotyped viruses not only improves transduction efficiency but also alters tropism. Viral tropism can also be modified using hybrid capsids derived from different serotypes. For example, AAV-DJ contains hybrid capsids derived from eight serotypes and exhibits high infectivity across a wide range of cell types in vivo. AAV-DJ8 is another example that exhibits the properties of AAV-DJ but with enhanced brain uptake. AAV serotypes can also be modified through mutations. Examples of mutational modifications in AAV2 include Y444F, Y500F, Y730F, and S662V. Examples of mutational modifications in AAV3 include Y705F, Y731F, and T492V. Examples of AAV6 mutation modification include S663V and T492V.Other pseudotype / modified AAV variants include AAV2 / 1, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2.5, AAV8.2 and AAV / SASTG.In a specific example, AAV is AAV7m8, which is an AAV variant that mediates highly efficient delivery to all retinal layers and photoreceptors.For example, see Dalkara et al.(2013)Sci.Transl.Med.5:189ra76, the entire contents of which are incorporated herein by reference for all purposes.

[0187] To accelerate transgene expression, self-complementary AAV (scAAV) variants can be used. Because AAV relies on the cell's DNA replication machinery to synthesize the complementary strand of the AAV's single-stranded DNA genome, transgene expression can be delayed. To address this delay, scAAV can be used, which contain complementary sequences that can spontaneously anneal upon infection, eliminating the need for host cell DNA synthesis. However, single-stranded AAV (ssAAV) vectors can also be used.

[0188] To increase packaging capacity, longer transgenes can be split into two AAV transfer plasmids, one containing a 3' splice donor and the other a 5' splice acceptor. Upon co-infection of cells, these viruses form concatemers that are spliced ​​together to express the full-length transgene. This allows for expression of longer transgenes, but at a reduced efficiency. A similar method for increasing capacity utilizes homologous recombination. For example, the transgene can be split into two transfer plasmids, but with substantial sequence overlap, such that co-expression induces homologous recombination and expression of the full-length transgene.

[0189] IV. Lipid Nanoparticles Containing Nucleic Acid Constructs Also provided herein are lipid nanoparticles comprising a nucleic acid construct (i.e., an exogenous donor nucleic acid) comprising a retinoschisin coding sequence (i.e., encoding a retinoschisin protein or a fragment or variant thereof) for integration into and expression from a target genomic locus. Also provided herein are lipid nanoparticles comprising a nucleic acid encoding a nuclease agent (e.g., that targets the endogenous RS1 locus) as disclosed elsewhere herein. Also provided herein are lipid nanoparticles comprising a nucleic acid construct and a nucleic acid encoding a nuclease agent (e.g., that targets the endogenous RS1 locus) as disclosed elsewhere herein.

[0190] Lipid formulations can protect biomolecules from degradation while improving cellular uptake. Lipid nanoparticles are particles containing multiple lipid molecules physically bound to each other by intermolecular forces. These include microspheres (including unilamellar and multilamellar vesicles, e.g., liposomes), the dispersed phase in an emulsion, micelles, or the internal phase in a suspension. Such lipid nanoparticles can be used to encapsulate one or more nucleic acids or proteins for delivery. Formulations containing cationic lipids are useful for delivering polyanions such as nucleic acids. Other lipids that can be included are neutral lipids (i.e., uncharged or zwitterionic lipids), anionic lipids, helper lipids that enhance transfection, and stealth lipids that increase the time nanoparticles can survive in vivo. Examples of suitable cationic lipids, neutral lipids, anionic lipids, helper lipids, and stealth lipids can be found in WO2016 / 010840A1 and WO2017 / 173054A1, each of which is incorporated herein by reference in its entirety for all purposes. Exemplary lipid nanoparticles can include a cationic lipid and one or more other components. In one example, the other components can include a helper lipid such as cholesterol. In another example, the other components can include a helper lipid such as cholesterol and a neutral lipid such as DSPC. In another example, the other components can include a helper lipid such as cholesterol, any neutral lipid such as DSPC, and a stealth lipid such as S010, S024, S027, S031, or S033.

[0191] LNPs may include one or more or all of the following: (i) lipids for encapsulation and endosomal escape, (ii) neutral lipids for stabilization, (iii) helper lipids for stabilization, and (iv) stealth lipids. See, for example, Finn et al. (2018) Cell Rep. 22(9):2227-2235 and WO2017 / 173054A1, each of which is incorporated by reference in its entirety for all purposes. In certain LNPs, the cargo may further include a nuclease agent. In certain LNPs, the cargo may further include a guide RNA or a nucleic acid encoding the guide RNA. In certain LNPs, the cargo may further include an mRNA encoding a Cas nuclease such as Cas9, and a guide RNA or a nucleic acid encoding the guide RNA. In certain LNPs, the cargo may include an mRNA encoding a Cas nuclease such as Cas9, a guide RNA or a nucleic acid encoding the guide RNA, and a nucleic acid construct.

[0192] The lipid for encapsulation and endosomal escape can be a cationic lipid. The lipid can also be a biodegradable lipid, such as a biodegradable ionizable lipid. One example of a suitable lipid is lipid A or LP01, which is (9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate, also known as 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate. See, for example, Finn et al. (2018) Cell Rep. 22(9):2227-2235 and WO2017 / 173054A1, each of which is incorporated by reference in its entirety for all purposes. Another example of a suitable lipid is lipid B, which is ((5-((dimethylamino)methyl)-1,3-phenylene)bis(oxy))bis(octane-8,1-diyl)bis(decanoate), also known as ((5-((dimethylamino)methyl)-1,3-phenylene)bis(oxy))bis(octane-8,1-diyl)bis(decanoate). Another example of a suitable lipid is lipid C, which is 2-((4-(((3-(dimethylamino)propoxy)carbonyl)oxy)hexadecanoyl)oxy)propane-1,3-diyl(9Z,9'Z,12Z,12'Z)-bis(octadeca-9,12-dienoate). Another example of a suitable lipid is lipid D, which is 3-(((3-(dimethylamino)propoxy)carbonyl)oxy)-13-(octanoyloxy)tridecyl 3-octylundecanoate. Other suitable lipids include heptatriaconta-6,9,28,31-tetraen-19-yl 4-(dimethylamino)butanoate (also known as Dlin-MC3-DMA(MC3)).

[0193] Some such lipids suitable for use in the LNPs described herein are biodegradable in vivo. For example, LNPs comprising such lipids include those in which at least 75% of the lipid is cleared from plasma within 8, 10, 12, 24, or 48 hours, or within 3, 4, 5, 6, 7, or 10 days. As another example, at least 50% of the LNP is cleared from plasma within 8, 10, 12, 24, or 48 hours, or within 3, 4, 5, 6, 7, or 10 days.

[0194] Such lipids may be ionizable depending on the pH of the medium in which they are contained. For example, in a slightly acidic medium, the lipids may be protonated and thus positively charged. Conversely, in a slightly basic medium, such as blood, which has a pH of about 7.35, the lipids may not be protonated and therefore may not be charged. In some embodiments, the lipids may be protonated at a pH of at least about 9, 9.5, or 10. The ability of such lipids to be charged is related to their inherent pKa. For example, the lipids may independently have a pKa ranging from about 5.8 to about 6.2.

[0195] Neutral lipids function to stabilize the LNPs and improve their processing. Examples of suitable neutral lipids include various neutral, uncharged, or zwitterionic lipids. Examples of neutral phospholipids suitable for use in the present disclosure include 5-heptadecylbenzene-1,3-diol (resorcinol), dipalmitoylphosphatidylcholine (DPPC), disteroylphosphatidylcholine (DSPC), phosphocholine (DOPC), dimyristoylphosphatidylcholine (DMPC), phosphatidylcholine (PLPC), 1,2-distearoyl-sn-glycero-3-phosphocholine (DAPC), phosphatidylethanolamine (PE), egg phosphatidylcholine (EPC), dilauryloylphosphatidylcholine (DLPC), dimyristoylphosphatidylcholine (DMPC), 1-myristoyl-2-palmitoylphosphatidylcholine (MPPC), 1-palmitoyl-2-myristoylphosphatidylcholine (PMPC), 1-palmitoyl-2-stearoylphosphatidyl choline (PSPC), 1,2-diarachidoyl-sn-glycero-3-phosphocholine (DBPC), 1-stearoyl-2-palmitoylphosphatidylcholine (SPPC), 1,2-dieicosenoyl-sn-glycero-3-phosphocholine (DEPC), palmitoyloleoylphosphatidylcholine (POPC), lysophosphatidylcholine, dioleoylphosphatidylethanolamine (DOPE), dilinoleoylphosphatidylcholine distearoylphosphatidylethanolamine (DSPE), dimyristoylphosphatidylethanolamine (DMPE), dipalmitoylphosphatidylethanolamine (DPPE), palmitoyloleoylphosphatidylethanolamine (POPE), lysophosphatidylethanolamine and combinations thereof. For example, the neutral phospholipid may be selected from the group consisting of distearoylphosphatidylcholine (DSPC), and dimyristoylphosphatidylethanolamine (DMPE).

[0196] Helper lipids include lipids that enhance transfection. The mechanism by which helper lipids enhance transfection may include enhancing particle stability. In some cases, helper lipids can enhance membrane fusion. Helper lipids include steroids, sterols, and alkylresorcinols. Examples of suitable helper lipids include cholesterol, 5-heptadecylresorcinol, and cholesterol hemisuccinate. In one example, the helper lipid can be cholesterol or cholesterol hemisuccinate.

[0197] Stealth lipids include lipids that change the length of time that nanoparticles can exist in vivo. Stealth lipids are useful in the formulation process, for example, by reducing particle aggregation and controlling particle size. Stealth lipids can adjust the pharmacokinetic properties of LNPs. Suitable stealth lipids include lipids that have a hydrophilic head group linked to the lipid moiety.

[0198] The hydrophilic head group of the stealth lipid may comprise a polymer moiety selected from, for example, polymers based on PEG (often referred to as poly(ethylene oxide)), poly(oxazoline), poly(vinyl alcohol), poly(glycerol), poly(N-vinylpyrrolidone), polyamino acids, and poly(N-(2-hydroxypropyl)methacrylamide). The term PEG refers to polyethylene glycol or other polyalkylene ether polymers. In certain LNP formulations, the PEG is PEG-2K, also known as PEG2000, and has an average molecular weight of approximately 2,000 daltons. See, for example, WO2017 / 173054A1, which is incorporated herein by reference in its entirety for all purposes.

[0199] The lipid portion of the stealth lipid can be derived from, for example, a diacylglycerol or diacylglycamide, including those containing a dialkylglycerol or dialkylglycamide group having an alkyl chain length independently containing from about C4 to about C40 saturated or unsaturated carbon atoms, where the chain can contain one or more functional groups, such as, for example, an amide or ester. The dialkylglycerol or dialkylglycamide group can further contain one or more substituted alkyl groups.

[0200] By way of example, stealth lipids include PEG-dilaurylglycerol, PEG-dimyristoylglycerol (PEG-DMG), PEG-dipalmitoylglycerol, PEG-distearoylglycerol (PEG-DSPE), PEG-dilaurylglycamide, PEG-dimyristylglycamide, PEG-dipalmitoylglycamide, and PEG-distearoylglycamide, PEG-cholesterol (l-[8'-(cholest-5-ene-3[beta]-oxy)carboxamido-3',6'-dioxaotanyl]carbamoyl-[omega]-methyl-poly(ethylene glycol), PEG-DMB (3,4-ditetradecoxybenzyl-[omega]-methyl-poly(ethylene glycol)), and PEG-DMB (3,4-ditetradecoxybenzyl-[omega]-methyl-poly(ethylene glycol)). The stealth lipid may be selected from PEG2k-DMG, 1,2-dimyristoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DMG), 1,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DSPE), 1,2-distearoyl-sn-glycerol, methoxypolyethylene glycol (PEG2k-DSG), poly(ethylene glycol)-2000-dimethacrylate (PEG2k-DMA), and 1,2-distearyloxypropyl-3-amine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DSA). In one specific example, the stealth lipid may be PEG2k-DMG.

[0201] LNPs can contain different molar ratios of the component lipids in the formulation. The molar percentage of CCD lipids can be, for example, about 30 mol% to about 60 mol%, about 35 mol% to about 55 mol%, about 40 mol% to about 50 mol%, about 42 mol% to about 47 mol%, or about 45%. The molar percentage of helper lipids can be, for example, about 30 mol% to about 60 mol%, about 35 mol% to about 55 mol%, about 40 mol% to about 50 mol%, about 41 mol% to about 46 mol%, or about 44 mol%. The molar percentage of neutral lipids can be, for example, about 1 mol% to about 20 mol%, about 5 mol% to about 15 mol%, about 7 mol% to about 12 mol%, or about 9 mol%. The mole percent of the stealth lipid may be, for example, about 1 mole percent to about 10 mole percent, about 1 mole percent to about 5 mole percent, about 1 mole percent to about 3 mole percent, about 2 mole percent, or about 1 mole percent.

[0202] LNPs can have different ratios between the positively charged amine groups of the biodegradable lipids (N) and the negatively charged phosphate groups of the encapsulated nucleic acid (P). This can be mathematically represented by the equation N / P. For example, the N / P ratio can be about 0.5 to about 100, about 1 to about 50, about 1 to about 25, about 1 to about 10, about 1 to about 7, about 3 to about 5, about 4 to about 5, about 4, about 4.5, or about 5.

[0203] In some LNPs, the cargo can comprise a Cas mRNA and a gRNA. The Cas mRNA and gRNA can be in different ratios. For example, an LNP formulation can comprise a Cas mRNA to gRNA ratio ranging from about 25:1 to about 1:25, from about 10:1 to about 1:10, from about 5:1 to about 1:5, or about 1:1. Alternatively, an LNP formulation can comprise a Cas mRNA to gRNA nucleic acid ratio of about 1:1 to about 1:5, or about 10:1. Alternatively, an LNP formulation can comprise a Cas mRNA to gRNA nucleic acid ratio of about 1:10, 25:1, 10:1, 5:1, 3:1, 1:1, 1:3, 1:5, 1:10, or 1:25. Alternatively, an LNP formulation can comprise a Cas mRNA to gRNA nucleic acid ratio of about 1:1 to about 1:2. In specific examples, the ratio of Cas mRNA to gRNA may be about 1:1 or about 1:2.

[0204] Exemplary dosing of LNPs includes about 0.1, about 0.25, about 0.3, about 0.5, about 1, about 2, about 3, about 4, about 5, about 6, about 8, or about 10 mg / kg body weight (mpk) of total RNA (Cas9 mRNA and gRNA) cargo content, or about 0.1 to about 10, about 0.25 to about 10, about 0.3 to about 10, about 0.5 to about 10, about 1 to about 10, about 2 to about 10, about 3 to about 10, about 4 to about 10, about 5 to about 10, about 6 to about 10, about 8 to about 10, about 0.1 to about 8, about 0.1 to about 6, about 0.1 to about 5, about 0.1 to about 4, about 0.1 to about 3, about 0.1 to about 2, about 0.1 to about 1, about 0.1 to about 0.5, about 0.1 to about 0.3, about 0.1 to about 0.25, about 0.25 to about 8, about 0.3 to about 6, about 0.5 to about 5, about 1 to about 5, or about 2 to about 3 mg / kg body weight. Such LNP can be administered, for example, intravenously. In one example, an LNP dose of about 0.01 mg / kg to about 10 mg / kg, about 0.1 to about 10 mg / kg, or about 0.01 to about 0.3 mg / kg can be used. For example, LNP doses of about 0.01, about 0.03, about 0.1, about 0.3, about 1, about 3, or about 10 mg / kg can be used. Additional exemplary dosing of LNPs include about 0.1, about 0.25, about 0.3, about 0.5, about 1, about 2, about 3, about 4, about 5, about 6, about 8, or about 10 mg / kg (mpk) body weight relative to total RNA (Cas9 mRNA and gRNA) cargo content, or about 0.1 to about 10, about 0.25 to about 10, about 0.3 to about 10, about 0.5 to about 10, about 1 to about 10, about 2 to about 10, about 3 to about 10, about 4 to about 10, about 5 to about 10 , about 6 to about 10, about 8 to about 10, about 0.1 to about 8, about 0.1 to about 6, about 0.1 to about 5, about 0.1 to about 4, about 0.1 to about 3, about 0.1 to about 2, about 0.1 to about 1, about 0.1 to about 0.5, about 0.1 to about 0.3, about 0.1 to about 0.25, about 0.25 to about 8, about 0.3 to about 6, about 0.5 to about 5, about 1 to about 5, or about 2 to about 3 mg / kg body weight. Such LNP can be administered, for example, intravenously. In one example, an LNP dose of about 0.01 mg / kg to about 10 mg / kg, about 0.1 to about 10 mg / kg, or about 0.01 to about 0.3 mg / kg can be used.For example, an LNP dose of about 0.01, about 0.03, about 0.1, about 0.3, about 0.5, about 1, about 2, about 3, or about 10 mg / kg can be used. In another example, an LNP dose of about 0.5 to about 10, about 0.5 to about 5, about 0.5 to about 3, about 1 to about 10, about 1 to about 5, about 1 to about 3, or about 1 to about 2 mg / kg can be used.

[0205] V. NUCLEIC ACID CONSTRUCTS AND / OR COMPOSITIONS COMPRISING NUCLEASE AGENT OR NUCLEIC ACID ENCODING A NUCLEASE AGENT Also provided herein are compositions comprising a nucleic acid construct, vector, or lipid nanoparticle comprising a retinoschisin coding sequence (i.e., encoding a retinoschisin protein or a fragment or variant thereof) for integration into and expression from a target genomic locus as disclosed herein, as well as a nuclease agent or a nucleic acid encoding a nuclease agent. Also provided herein are compositions comprising a nucleic acid construct comprising a retinoschisin coding sequence (i.e., encoding a retinoschisin protein or a fragment or variant thereof), vector, or lipid nanoparticle for integration into and expression from a target genomic locus as disclosed herein. Also provided herein are compositions comprising a nuclease agent or a nucleic acid encoding a nuclease agent (e.g., a nuclease agent targeted to the RS1 gene or locus), or a vector or lipid nanoparticle comprising a nuclease agent or a nucleic acid encoding a nuclease agent. Such compositions can be, for example, for use in expressing retinoschisin in a cell or for use in integrating a coding sequence for a retinoschisin protein or a fragment or variant thereof into a target genomic locus in a cell. Such compositions may also be used to treat subjects with, for example, X-linked juvenile retinoschisis (XLRS). Such compositions may include a nucleic acid construct comprising a coding sequence of a retinoschisin protein or a fragment thereof for integration into a target genomic locus (or a vector or lipid nanoparticle comprising the nucleic acid construct), and a nuclease agent or a nucleic acid encoding the nuclease agent, wherein the nuclease agent targets a nuclease target sequence in the target genomic locus. The nuclease agent may be a CRISPR / Cas system (e.g., a Cas protein and a guide RNA) or any other suitable nuclease agent. Examples of suitable nuclease agents are provided below.

[0206] A. CRISPR / Cas system The methods and compositions disclosed herein can utilize a clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated (Cas) system or components of such a system to modify a genome (e.g., the RS1 locus) in a cell. A CRISPR / Cas system includes transcripts and other elements that are involved in the expression of or direct the activity of Cas genes. The CRISPR / Cas system can be, for example, a Type I, Type II, Type III, or Type V system (e.g., subtype VA or subtype VB). The methods and compositions disclosed herein can employ a CRISPR / Cas system by utilizing a CRISPR complex (including a guide RNA (gRNA) complexed with a Cas protein) for site-specific binding or cleavage of nucleic acids.

[0207] The CRISPR / Cas system used in the compositions and methods disclosed herein can be non-naturally occurring.A " non-naturally occurring " system includes any that shows the involvement of human hands, such as one or more components of the system are modified or mutated from their naturally occurring state, they are at least substantially free of at least one other component that is naturally associated, or they are associated with at least one other component that is not naturally associated.For example, some CRISPR / Cas systems use non-naturally occurring CRISPR complexes that include non-naturally occurring gRNA and Cas protein together, or use non-naturally occurring Cas protein, or use non-naturally occurring gRNA.

[0208] 1. Cas proteins Cas proteins generally contain at least one RNA recognition or binding domain capable of interacting with a guide RNA. Cas proteins may also contain a nuclease domain (e.g., a DNase or RNase domain), a DNA-binding domain, a helicase domain, a protein-protein interaction domain, a dimerization domain, and other domains. Some such domains (e.g., a DNase domain) may be derived from native Cas proteins. Other such domains can be added to create modified Cas proteins. The nuclease domain has catalytic activity for nucleic acid cleavage, including covalent cleavage of nucleic acid molecules. Cleavage can generate blunt or staggered ends, which can be single-stranded or double-stranded. For example, wild-type Cas9 proteins typically generate blunt cleavage products. Alternatively, wild-type Cpf1 proteins (e.g., FnCpf1) can result in cleavage products with a 5-nucleotide 5' overhang, with cleavage occurring 18 base pairs from the PAM sequence on the non-target strand and 23 bases on the target strand. A Cas protein can have full cleavage activity and create a double-stranded break at the target genomic locus (e.g., a double-stranded break with a blunt end), or it can be a nickase that creates a single-stranded break at the target genomic locus.

[0209] Examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csn1 or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (CasA), Cse2 (CasB), Cse3 (CasE), These include Cse4 (CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cu1966, and homologs or modified versions thereof.

[0210] Exemplary Cas protein is Cas9 protein or a protein derived from Cas9 protein.Cas9 protein is derived from type II CRISPR / Cas system, and typically shares four important motifs with conserved architecture.Modifications 1, 2, and 4 are RuvC-like motifs, and motif 3 is HNH motif. Exemplary Cas9 proteins are Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Neisseria meningitidis, or Campylobacter jejuni. Additional examples of Cas9 family members are described in WO2014 / 131833, which is incorporated by reference in its entirety for all purposes. S.Cas9 from S. pyogenes (SpCas9) (e.g., assigned UniProt accession number Q99ZW2) is an exemplary Cas9 protein. An exemplary SpCas9 protein sequence is set forth in SEQ ID NO:27 (encoded by the DNA sequence set forth in SEQ ID NO:26). An exemplary SpCas9 cDNA sequence is set forth in SEQ ID NO:28. Smaller Cas9 proteins (e.g., Cas9 proteins whose coding sequences are compatible with maximum AAV packaging capacity when combined with the guide RNA coding sequence and regulatory elements of Cas9 and guide RNA, such as SaCas9, CjCas9, and Nme2Cas9) are other exemplary Cas9 proteins. For example, Cas9 from S. aureus (SaCas9) (e.g., assigned UniProt accession number J7RUA5) is another exemplary Cas9 protein. Similarly, Cas9 from Campylobacter jejuni (CjCas9), for example (assigned UniProt accession number Q0P897), is another exemplary Cas9 protein. See, for example, Kim et al., (2017) Nat. Commun. 8:14500, the entire contents of which are incorporated herein by reference for all purposes. SaCas9 is smaller than SpCas9, and CjCas9 is smaller than both SaCas9 and SpCas9. Cas9 from Neisseria meningitidis (Nme2Cas9) is another exemplary Cas9 protein. See, for example, Edraki et al. (2019) Mol.See Cell 73(4):714-726, which is incorporated by reference in its entirety for all purposes. Other exemplary Cas9 proteins include Streptococcus thermophilus-derived Cas9 proteins (e.g., Streptococcus thermophilus LMD-9 Cas9 encoded by the CRISPR1 locus (St1Cas9) or Streptococcus thermophilus Cas9 encoded by the CRISPR3 locus (St3Cas9). Other exemplary Cas9 proteins include Francisella novicida-derived Cas9 (FnCas9) or the RHA Francisella novicida Cas9 variant (E1369R / E1449H / R1556A substitutions) that recognize alternative PAMs. These and other exemplary Cas9 proteins are described, for example, in Cebrian-Serrano and Davies (2017) Mamm.Genome 28(7):247-261, which are incorporated herein by reference in their entireties for all purposes. Examples of Cas9 coding sequences, Cas9 mRNA, and Cas9 protein sequences are provided in WO2013 / 176772, WO2014 / 065596, WO2016 / 106121, and WO2019 / 067910, each of which is incorporated herein by reference in its entirety for all purposes. Specific examples of ORFs and Cas9 amino acid sequences are provided in Table 30, paragraph

[0449] of WO2019 / 067910, and specific examples of Cas9 mRNAs and ORFs are provided in paragraphs

[0214] to

[0234] of WO2019 / 067910. By way of example, a Cas9 protein can comprise, consist essentially of, or consist of the sequence set forth in SEQ ID NO:6242. Such a Cas9 protein can be encoded by an mRNA comprising, consisting essentially of, or consisting of SEQ ID NO: 6243. As another example, a Cas9 protein can comprise, consist essentially of, or consist of the sequence set forth in SEQ ID NO: 6246. Such a Cas9 protein can be encoded by an mRNA comprising, consisting essentially of, or consisting of SEQ ID NO: 6245.

[0211] Another example of a Cas protein is the Cpf1 (CRISPR from Prevotella and Francisella1) protein. Cpf1 is a large protein (approximately 1300 amino acids) that contains a RuvC-like nuclease domain homologous to the corresponding domain in Cas9, along with a counterpart of Cas9's characteristic arginine-rich cluster. However, Cpf1 lacks the HNH nuclease domain present in the Cas9 protein, and in contrast to Cas9, which contains a long insertion containing an HNH domain, the RuvC-like domain is adjacent in the Cpf1 sequence. See, e.g., Zetsche et al. (2015) Cell 163(3):759-771, the entire contents of which are incorporated herein by reference for all purposes. Exemplary Cpf1 proteins are Francisella tularensis 1, Francisella tularensis subsp.novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp.SCADC, Acidaminococcus sp.BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis Cpf1 from Francisella novicida U112 (FnCpf1; assigned UniProt accession number A0Q7Q2) is an exemplary Cpf1 protein.

[0212] The Cas protein can be a wild-type protein (i.e., naturally occurring), a modified Cas protein (i.e., a Cas protein variant), or a fragment of a wild-type or modified Cas protein. The Cas protein can also be a variant or fragment that is active with respect to catalytic activity of the wild-type or modified Cas protein. A catalytically active variant or fragment can contain at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the wild-type or modified Cas protein or a portion thereof, and an active variant retains the ability to cleave at the desired cleavage site and thus retains nick-inducing or double-strand break-inducing activity. Assays for nick-inducing or double-strand break-inducing activity are known and generally measure the overall activity and specificity of a Cas protein on a DNA substrate containing the cleavage site.

[0213] Cas proteins can be modified to increase or decrease one or more of nucleic acid binding affinity, nucleic acid binding specificity, and enzymatic activity. Cas proteins can also be modified to alter other activities or properties of the protein, such as stability. For example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated, or the Cas protein can be truncated to remove domains that are not essential for protein function, or the activity or properties of the Cas protein can be optimized (e.g., enhanced or reduced).

[0214] One example of a modified Cas protein is the modified SpCas9-HF1 protein, which is a high-fidelity variant of Streptococcus pyogenes Cas9 (N497A / R661A / Q695A / Q926A) that contains modifications designed to reduce nonspecific DNA contact. See, e.g., Kleinstiver et al. (2016) Nature 529(7587):490-495, incorporated herein by reference in its entirety for all purposes. Another example of a modified Cas protein is the modified eSpCas9 variant (K848A / K1003A / R1060A) designed to reduce off-target effects. See, e.g., Slaymaker et al. (2016) Science 351(6268):84-88, incorporated herein by reference in its entirety for all purposes. Other SpCas9 variants include K855A and K810A / K1003A / R1060A. These and other modified Cas proteins are reviewed, for example, in Cebrian-Serrano and Davies (2017) Mamm. Genome 28(7):247-261, which is incorporated herein by reference in its entirety for all purposes. Another example of a modified Cas9 protein is xCas9, which is an SpCas9 variant that can recognize an expanded range of PAM sequences. See, for example, Hu et al. (2018) Nature 556:57-63, which is incorporated herein by reference in its entirety for all purposes.

[0215] Cas proteins can contain at least one nuclease domain, such as a DNase domain. For example, wild-type Cpf1 proteins generally contain a RuvC-like domain, likely in a dimeric conformation, that cleaves both strands of target DNA. Cas proteins can also contain at least two nuclease domains, such as a DNase domain. For example, wild-type Cas9 proteins generally contain a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC domain and the HNH domain can each cleave different strands of double-stranded DNA, creating a double-strand break in DNA. See, e.g., Jinek et al. (2012) Science 337:816-821, incorporated herein by reference in its entirety for all purposes.

[0216] Deleting or mutating one or more or all of the nuclease domains can render them non-functional or reduce their nuclease activity. For example, if one of the nuclease domains is deleted or mutated in a Cas9 protein, the resulting Cas9 protein, called a nickase, can generate single-strand breaks in double-stranded target DNA but not double-strand breaks (i.e., it can cleave either the complementary or non-complementary strand, but not both). If both nuclease domains are deleted or mutated, the resulting Cas protein (e.g., Cas9) will have reduced ability to cleave both strands of double-stranded DNA (e.g., a nuclease-null or nuclease-inactive Cas protein, or a Cas protein without catalytic activity (dCas)). If the nuclease domain is deleted or mutated in a Cas9 protein, the Cas9 protein retains double-strand break-inducing activity. An example of a mutation that converts Cas9 into a nickase is the D10A (aspartate to alanine at position 10 of Cas9) mutation in the RuvC domain of Cas9 from S. pyogenes. Similarly, H939A (histidine to alanine at amino acid position 839), H840A (histidine to alanine at amino acid position 840), or N863A (asparagine to alanine at amino acid position N863) in the HNH domain of Cas9 from S. pyogenes can convert Cas9 into a nickase. Other examples of mutations that convert Cas9 into a nickase include corresponding mutations in Cas9 from S. thermophilus. See, e.g., Sapranauskas et al. (2011) Nucleic Acids Res. 39(21):9275-9282 and WO2013 / 141680, each of which is incorporated by reference in its entirety for all purposes. Such mutations can be generated using methods such as site-directed mutagenesis, PCR-mediated mutagenesis, or total gene synthesis. Examples of other nickase-generating mutations can be found, for example, in WO2013 / 176772 and WO2013 / 142578, each of which is incorporated by reference in its entirety for all purposes.When all of the nuclease domains are deleted or mutated in a Cas protein (e.g., when both nuclease domains are deleted or mutated in a Cas9 protein), the resulting Cas protein (e.g., Cas9) has a reduced ability to cleave both strands of double-stranded DNA (e.g., a nuclease-null or nuclease-inactive Cas protein). One particular example is the D10A / H840A S. pyogenes Cas9 double mutant, or the corresponding double mutant of a Cas9 from another species when optimally aligned with S. pyogenes Cas9. Another particular example is the D10A / N863A S. pyogenes Cas9 double mutant, or the corresponding double mutant of a Cas9 from another species when optimally aligned with S. pyogenes Cas9.

[0217] Examples of inactivating mutations in the catalytic domain of xCas9 are the same as those described above for SpCas9. Examples of inactivating mutations in the catalytic domain of the Staphylococcus aureus Cas9 protein are also known. For example, the Staphylococcus aureus Cas9 enzyme (SaCas9) can contain a substitution at position N580 (e.g., an N580A substitution) and a substitution at position D10 (e.g., a D10A substitution), resulting in a nuclease-inactive Cas protein. See, e.g., WO2016 / 106236, incorporated herein by reference in its entirety for all purposes. Examples of inactivating mutations in the catalytic domain of Nme2Cas9 are also known (e.g., a combination of D16A and H588A). Examples of inactivating mutations in the catalytic domain of St1Cas9 are also known (e.g., a combination of D9A, D598A, H599A, and N622A). Examples of inactivating mutations in the catalytic domain of St3Cas9 are also known (e.g., the combination of D10A and N870A). Examples of inactivating mutations in the catalytic domain of CjCas9 are also known (e.g., the combination of D8A and H559A). Examples of inactivating mutations in the catalytic domain of FnCas9 and RHA FnCas9 are also known (e.g., N995A).

[0218] Examples of inactivating mutations in the catalytic domain of the Cpf1 protein are also known. For the Cpf1 proteins from Francisella novicida U112 (FnCpf1), Acidaminococcus sp. BV3L6 (AsCpf1), Lachnospiraceae bacterium ND2006 (LbCpf1), and Moraxella bovoculi 237 (MbCpf1 Cpf1), such mutations can include mutations at positions 908, 993, or 1263 of AsCpf1 or corresponding positions in Cpf1 orthologs, or at positions 832, 925, 947, or 1180 of LbCpf1 or corresponding positions in Cpf1 orthologs. Such mutations can include, for example, one or more of the mutations D908A, E993A, and D1263A in AsCpf1 or corresponding mutations in Cpf1 orthologs, or D832A, E925A, D947A, and D1180A in LbCpf1 or corresponding mutations in Cpf1 orthologs. See, e.g., US2016 / 0208243, incorporated herein by reference in its entirety for all purposes.

[0219] Cas protein can also be operably linked to heterologous polypeptide as a fusion protein. For example, Cas protein can be fused to a cleavage domain or epigenetic modification domain. See WO2014 / 089290, the entire contents of which are incorporated herein by reference for all purposes. Cas protein can also be fused to a heterologous polypeptide that provides increased or decreased stability. The fusion domain or heterologous polypeptide can be located at the N-terminus, C-terminus, or internal to the Cas protein.

[0220] For example, a Cas protein may be fused to one or more heterologous polypeptides that provide subcellular localization. Such heterologous polypeptides may include one or more nuclear localization signals (NLSs), such as a monopartite SV40 NLS and / or a bipartite alpha-importin NLS for targeting the nucleus, a mitochondrial localization signal for targeting mitochondria, an ER retention signal, or the like. See, for example, Lange et al. (2007) J. Biol. Chem. 282(8):5101-5105, incorporated herein by reference in its entirety for all purposes. Such subcellular localization signals may be located at the N-terminus, C-terminus, or anywhere within the Cas protein. The NLS may comprise a stretch of basic amino acids and may be a mono- or bi-knot sequence. Optionally, the Cas protein may include two or more NLSs, including an NLS at the N-terminus (e.g., an alpha-importin NLS or a mono-knot NLS) and an NLS at the C-terminus (e.g., an SV40 NLS or a bi-knot NLS). Cas proteins may also contain two or more NLSs at the N-terminus and / or two or more NLSs at the C-terminus.

[0221] The Cas protein may be fused to, for example, 1 to 10 NLSs (e.g., 1 to 5 NLSs), or to only one NLS. When one NLS is used, the NLS may be linked at the N-terminus or C-terminus of the Cas protein sequence. The NLS may also be inserted within the Cas protein sequence. Alternatively, the Cas protein may be fused to two or more NLSs. For example, the Cas protein may be fused to 2, 3, 4, or 5 NLSs. In a specific example, the Cas protein may be fused to two NLSs. In certain circumstances, the two NLSs may be the same (e.g., two SV40 NLSs) or different. For example, the Cas protein may be fused to two SV40 NLS sequences linked at the carboxy termini. Alternatively, the Cas protein may be fused to two NLSs, one linked at the N-terminus and one linked at the C-terminus. In other examples, the Cas protein may be fused to three NLSs or no NLS at all. The NLS may be, for example, an SV40 NLS. The NLS may be a single-part sequence, such as PKKKRKV (SEQ ID NO: 49) or PKKKRRV (SEQ ID NO: 50). The NLS may be a bipartite sequence, such as the nucleoplasmin NLS, KRPAATKKAGQAKKKK (SEQ ID NO: 51). In a specific example, a single PKKKRKV (SEQ ID NO: 49) NLS may be linked at the C-terminus of the Cas protein. One or more linkers are optionally included in the fusion site.

[0222] Cas proteins can also be operably linked to a cell penetration domain or protein transduction domain. For example, the cell penetration domain can be derived from the HIV-1 TAT protein, the TLM cell penetration motif from human hepatitis B virus, MPG, Pep-1, VP22, the cell penetration peptide from herpes simplex virus, or a polyarginine peptide sequence. See, for example, WO2014 / 089290 and WO2013 / 176772, each of which is incorporated herein by reference in its entirety for all purposes. The cell penetration domain can be located at the N-terminus, C-terminus, or anywhere within the Cas protein.

[0223] The Cas protein can also be operably linked to a heterologous polypeptide, such as a fluorescent protein, a purification tag, or an epitope tag, to facilitate tracking or purification. Examples of fluorescent proteins include green fluorescent protein (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, Emerald, Azami Green, Monomeric Azami). Green, CopGFP, AceGFP, ZsGreenl), yellow fluorescent proteins (e.g., YFP, eYFP, Citrine, Venus, YPet, PhiYFP, ZsYellowl), blue fluorescent proteins (e.g., eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., eCFP, Cerulean, CyPet, AmCyanl, Midoriishi-Cyan), red fluorescent proteins (e.g., mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred), orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, Monomeric Examples of tags include glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, hemagglutinin (HA), nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, histidine (His), biotin carboxyl carrier protein (BCCP), and calmodulin.

[0224] The Cas protein may be tethered to a labeled nucleic acid or donor sequence. Such tethering (i.e., physical association) can be achieved through covalent or non-covalent interactions, and tethering can be achieved directly (e.g., via direct fusion or chemical linkage, which can be achieved by modification of cysteine ​​or lysine residues on the protein or intein modifications) or through one or more intervening linker or adapter molecules, such as streptavidin or aptamers. See, for example, Pierce et al. (2005) Mini Rev. Med. Chem. 5(1):41-55, Duckworth et al. (2007) Angew. Chem. Int. Ed. Engl. 46(46):8819-8822, Schaeffer and Dixon (2009) Australian J. Chem. 62(10):1328-1332, Goodman et al. (2009) Chembiochem. 10(9):1551-1557, and Khatwani et al. (2012) Bioorg. Med. Chem. 20(14):4532-4539, each of which is incorporated by reference in its entirety for all purposes. Non-covalent strategies for synthesizing protein-nucleic acid conjugates include biotin-streptavidin and nickel-histidine methods. Covalent protein-nucleic acid conjugates can be synthesized by connecting appropriately functionalized nucleic acids and proteins using a variety of chemistries. Some of these chemistries involve direct attachment of oligonucleotides to amino acid residues on the protein surface (e.g., lysine amines or cysteine ​​thiols), while other, more complex schemes require post-translational modifications of the protein or the involvement of catalytic or reactive protein domains. Methods for covalently attaching proteins to nucleic acids can include, for example, chemical crosslinking of oligonucleotides to protein lysine or cysteine ​​residues, expressed protein ligation, chemoenzymatic methods, and the use of photoaptamers. The labeled nucleic acid or donor sequence can be tethered to the C-terminus, N-terminus, or internal region of the Cas protein.In one example, the labeled nucleic acid or donor sequence is tethered to the C-terminus or N-terminus of the Cas protein. Similarly, the Cas protein may be tethered to the 5'-end, 3'-end, or internal region of the labeled nucleic acid or donor sequence. That is, the labeled nucleic acid or donor sequence may be tethered in any direction and polarity. For example, the Cas protein may be tethered to the 5'-end or 3'-end of the labeled nucleic acid or donor sequence.

[0225] The Cas protein can be provided in any form. For example, the Cas protein can be provided in the form of a protein, such as a Cas protein complexed with a gRNA. Alternatively, the Cas protein can be provided in the form of a nucleic acid encoding the Cas protein, such as RNA (e.g., messenger RNA (mRNA)) or DNA. Optionally, the nucleic acid encoding the Cas protein can be codon-optimized for efficient translation into protein in a particular cell or organism. For example, the nucleic acid encoding the Cas protein can be modified to use alternative codons more frequently used in bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest, compared to the naturally occurring polynucleotide sequence. When the nucleic acid encoding the Cas protein is introduced into a cell, the Cas protein can be expressed transiently, conditionally, or constitutively within the cell.

[0226] Cas proteins provided as mRNA can be modified to improve stability and / or immunogenicity. Modifications can be made to one or more nucleosides within the mRNA. Examples of chemical modifications to mRNA nucleobases include pseudouridine, 1-methyl-pseudouridine, and 5-methyl-cytidine. For example, capped and polyadenylated Cas mRNA containing N1-methylpseudouridine can be used. Similarly, Cas mRNA can be modified by depleting uridines using synonymous codons.

[0227] The nucleic acid encoding the Cas protein can be stably integrated into the genome of the cell and operably linked to a promoter active in the cell. Alternatively, the nucleic acid encoding the Cas protein can be operably linked to a promoter in an expression construct. An expression construct includes any nucleic acid construct capable of directing the expression of a gene or other nucleic acid sequence of interest (e.g., a Cas gene) and introducing such a nucleic acid sequence of interest into a target cell. For example, the nucleic acid encoding the Cas protein can be present in a vector containing DNA encoding a gRNA. Alternatively, it can be a vector or plasmid separate from the vector containing DNA encoding the gRNA. Promoters that can be used in expression constructs include, for example, promoters active in one or more of eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, pluripotent cells, embryonic stem (ES) cells, adult stem cells, developmentally restricted progenitor cells, induced pluripotent stem (iPS) cells, or one-cell embryos. Such promoters can be, for example, conditional promoters, inducible promoters, constitutive promoters, or tissue-specific promoters. Optionally, the promoter may be a bidirectional promoter that drives expression of both the Cas protein in one direction and the guide RNA in the other direction. Such a bidirectional promoter may consist of (1) a complete conventional unidirectional Pol III promoter containing three external regulatory elements: a distal sequence element (DSE), a proximal sequence element (PSE), and a TATA box; and (2) a second basic Pol III promoter containing a PSE and a TATA box fused in reverse orientation to the 5' end of the DSE. For example, in the H1 promoter, the DSE is adjacent to the PSE and TATA box; the promoter can be made bidirectional by adding a PSE and a TATA box from the U6 promoter to create a hybrid promoter in which transcription in the reverse direction is controlled. See, for example, US2016 / 0074535, incorporated herein by reference in its entirety for all purposes.Bidirectional promoters can be used to simultaneously express genes encoding Cas proteins and guide RNAs, generating compact expression cassettes for easy delivery.

[0228] Various promoters can be used to drive Cas or Cas9 expression. In some methods, small promoters are used so that the Cas or Cas9 coding sequence can fit into an AAV construct. For example, Cas or Cas9 and one or more gRNAs (e.g., one gRNA, two gRNAs, three gRNAs, or four gRNAs) can be delivered via LNP-mediated delivery (e.g., in the form of RNA) or adeno-associated virus (AAV)-mediated delivery (e.g., AAV2-, AAV5-, AAV8-, or AAV7m8-mediated delivery). For example, the nuclease agent can be CRISPR / Cas9, and Cas9 mRNA and a gRNA targeting the endogenous RS1 locus (e.g., intron 1 of RS1) can be delivered via LNP-mediated delivery, or DNA encoding Cas9 and a gRNA targeting the endogenous RS1 locus (e.g., intron 1 of RS1) can be delivered via AAV-mediated delivery. Cas or Cas9 and gRNA can be delivered in a single AAV or via two separate AAVs. For example, the first AAV can carry a Cas or Cas9 expression cassette, and the second AAV can carry a gRNA expression cassette. Similarly, the first AAV can carry a Cas or Cas9 expression cassette, and the second AAV can carry two or more gRNA expression cassettes. Alternatively, a single AAV can carry a Cas or Cas9 expression cassette (e.g., a Cas or Cas9 coding sequence operably linked to a promoter) and a gRNA expression cassette (e.g., a gRNA coding sequence operably linked to a promoter). Similarly, a single AAV can carry a Cas or Cas9 expression cassette (e.g., a Cas or Cas9 coding sequence operably linked to a promoter) and two or more gRNA expression cassettes (e.g., a gRNA coding sequence operably linked to a promoter). Various promoters can be used to drive the expression of gRNA, such as the U6 promoter or the small tRNA Gln. Similarly, various promoters can be used to drive the expression of Cas9.For example, small promoters are used so that the Cas9 coding sequence can fit into the AAV construct, as well as small Cas9 proteins (e.g., SaCas9 or CjCas9 are used to maximize AAV packaging capacity).

[0229] Cas proteins provided as mRNA can be modified to improve stability and / or immunogenicity. Modifications can be made to one or more nucleosides within the mRNA. Examples of chemical modifications to mRNA nucleobases include pseudouridine, 1-methyl-pseudouridine, and 5-methyl-cytidine. mRNAs encoding Cas proteins can also be capped. The cap can be, for example, a cap 1 structure in which the +1 ribonucleotide is methylated at the 2'O position of the ribose. Capping can, for example, confer superior activity in vivo (e.g., by mimicking the native cap) or result in a native structure that reduces stimulation of the host's innate immune system (e.g., reducing activation of pattern recognition receptors of the innate immune system). The mRNA encoding the Cas protein can also be polyadenylated (to form a poly(A) tail). The mRNA encoding the Cas protein can also be modified to include pseudouridine (e.g., completely replaced with pseudouridine). As another example, capped and polyadenylated Cas mRNA containing N1-methylpseudouridine can be used. As another example, a Cas mRNA that is fully substituted with pseudouridine can be used (i.e., all standard uracil residues are replaced with pseudouridine, a uridine isomer in which uracil is attached by a carbon-carbon bond rather than a nitrogen-carbon bond). Similarly, a Cas mRNA can be modified by depleting uridines using synonymous codons. For example, a capped and polyadenylated Cas mRNA that is fully substituted with pseudouridine can be used.

[0230] The Cas mRNA can include modified uridines at at least one, more than one, or all uridine positions. The modified uridine can be a uridine modified at the 5-position (e.g., with a halogen, methyl, or ethyl). The modified uridine can be a pseudouridine modified at the 1-position (e.g., with a halogen, methyl, or ethyl). The modified uridine can be, for example, pseudouridine, N1-methyl-pseudouridine, 5-methoxyuridine, 5-iodouridine, or a combination thereof. In some examples, the modified uridine is 5-methoxyuridine. In some examples, the modified uridine is 5-iodouridine. In some examples, the modified uridine is pseudouridine. In some examples, the modified uridine is N1-methyl-pseudouridine. In some examples, the modified uridine is a combination of pseudouridine and N1-methyl-pseudouridine. In some examples, the modified uridine is a combination of pseudouridine and 5-methoxyuridine. In some examples, the modified uridine is a combination of N1-methylpseudouridine and 5-methoxyuridine. In some examples, the modified uridine is a combination of 5-iodouridine and N1-methyl-pseudouridine. In some examples, the modified uridine is a combination of pseudouridine and 5-iodouridine. In some examples, the modified uridine is a combination of 5-iodouridine and 5-methoxyuridine.

[0231] The Cas mRNAs disclosed herein can also include a 5' cap, such as Cap0, Cap1, or Cap2. The 5' cap is generally a 7-methylguanine ribonucleotide (which may be further modified, e.g., with ARCA) linked via a 5'-triphosphate to the 5' position of the first nucleotide of the 5' to 3' strand of the mRNA (i.e., the first cap-proximal nucleotide). In Cap0, the riboses of the first and second cap-proximal nucleotides of the mRNA both contain 2'-hydroxyl. In Cap1, the riboses of the first and second transcribed nucleotides of the mRNA contain 2'-methoxy and 2'-hydroxyl, respectively. In Cap2, the riboses of the first and second cap-proximal nucleotides of the mRNA both contain 2'-methoxy. See, for example, Katibah et al. (2014) Proc. Natl. Acad. Sci. USA 111(33):12025-30 and Abbas et al. (2017) Proc. Natl. Acad. Sci. USA 114(11):E2106-E2115, each of which is incorporated by reference in its entirety for all purposes. Most endogenous higher eukaryotic mRNAs, including mammalian mRNAs such as human mRNAs, contain Cap1 or Cap2. Cap0 and other cap structures distinct from Cap1 and Cap2 are recognized as non-self by components of the innate immune system, such as IFIT-1 and IFIT-5, and may be immunogenic in mammals, including humans, and may induce elevated levels of cytokines, including type I interferons. Components of the innate immune system, such as IFIT-1 and IFIT-5, may also compete with eIF4E for binding of mRNA to caps other than Cap1 or Cap2, inhibiting mRNA translation.

[0232] A cap can be incorporated co-transcriptionally. For example, ARCA (anti-reverse cap analog, Thermo Fisher Scientific catalog number AM8045) is a cap analog containing 7-methylguanine 3'-methoxy-5'-triphosphate linked to the 5' position of a guanine ribonucleotide, which can be incorporated into transcripts in vitro at the time of initiation. ARCA results in a Cap0 cap, in which the 2' position of the first cap-proximal nucleotide is hydroxyl. See, for example, Stepinski et al. (2001) RNA 7:1486-1495, the entire contents of which are incorporated herein by reference for all purposes.

[0233] CleanCap™ AG (m7G(5')ppp(5')(2'OMeA)pG, TriLink Biotechnologies catalog number N-7113) or CleanCap™ GG (m7G(5')ppp(5')(2'OMeG)pG, TriLink Biotechnologies catalog number N-7133) are used to co-transcriptionally provide the Cap1 structure. 3'-O-methylated versions of CleanCap™ AG and CleanCap™ GG are also commercially available from TriLink Biotechnologies as catalog numbers N-7413 and N-7433, respectively.

[0234] Alternatively, a cap can be added to RNA after transcription. For example, vaccinia capping enzyme is commercially available (New England Biolabs Catalog No. M2080S), which has RNA triphosphatase and guanylyltransferase activity provided by the D1 subunit and guanine methyltransferase activity provided by the D12 subunit. Thus, in the presence of S-adenosylmethionine and GTP, 7-methylguanine can be added to RNA to give Cap0. See, for example, Guo and Moss (1990) Proc. Natl. Acad. Sci. USA 87:4023-4027 and Mao and Shuman (1994) J. Biol. Chem. 269:24472-24479, each of which is incorporated herein by reference in its entirety for all purposes.

[0235] The Cas mRNA can further comprise a polyadenylation (polyA or poly(A) or polyadenine) tail. The polyA tail can comprise, for example, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 adenines, and optionally up to 300 adenines. For example, the polyA tail can comprise 95, 96, 97, 98, 99, or 100 adenine nucleotides.

[0236] 2. Guide RNA A "guide RNA" or "gRNA" is an RNA molecule that binds to a Cas protein (e.g., a Cas9 protein) and targets the Cas protein to a specific location within a target DNA. A guide RNA can be composed of two segments: a "DNA-targeting segment" (also called a guide sequence) and a "protein-binding segment." A "segment" includes a section or region of a molecule, such as a contiguous stretch of nucleotides in an RNA. Some gRNAs, such as those of Cas9, can contain two separate RNA molecules: an "activator RNA" (e.g., tracrRNA) and a "targeter RNA" (e.g., CRISPR RNA or crRNA). Other gRNAs are single RNA molecules (single RNA polynucleotides), which may also be referred to as "single-molecule gRNA," "single guide RNA," or "sgRNA." See, for example, WO2013 / 176772, WO2014 / 065596, WO2014 / 089290, WO2014 / 093622, WO2014 / 099750, WO2013 / 142578, and WO2014 / 131833, each of which is incorporated herein by reference in its entirety for all purposes. Guide RNA refers to either CRISPR RNA (crRNA) or a combination of crRNA and trans-activating CRISPR RNA (tracrRNA). The crRNA and tracrRNA can be associated as a single RNA molecule (single guide RNA or sgRNA) or as two separate RNA molecules (dual guide RNA or dgRNA). For example, in the case of Cas9, the single guide RNA can comprise a crRNA fused to a tracrRNA (e.g., via a linker). For example, in the case of Cpf1, only the crRNA is required to achieve binding to and / or cleavage of the target sequence. The terms "guide RNA" and "gRNA" include both double-molecule (i.e., modular) gRNAs and single-molecule gRNAs. In some methods and compositions disclosed herein, the gRNA is an S. pyogenes Cas9 gRNA or its equivalent.In some methods and compositions disclosed herein, the gRNA is a S. aureus Cas9 gRNA or its equivalent.

[0237] Exemplary bimolecular gRNAs include a crRNA-like ("CRISPR RNA" or "targeter RNA" or "crRNA" or "crRNA repeat") molecule and a corresponding tracrRNA-like ("trans-activating CRISPR RNA" or "activator RNA" or "tracrRNA") molecule. The crRNA contains both the DNA-targeting segment (single strand) of the gRNA and a stretch of nucleotides (i.e., the crRNA tail) that forms one half of the dsRNA duplex of the protein-binding segment of the gRNA. Examples of crRNA tails located downstream (3') of the DNA-targeting segment comprise, consist essentially of, or consist of GUUUUAGAGCUAUGCU (SEQ ID NO: 29) or GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 52). Any of the DNA-targeting segments disclosed herein can be attached to the 5' end of SEQ ID NO: 29 or 52 to form a crRNA.

[0238] The corresponding tracrRNA (activator RNA) contains a stretch of nucleotides that forms the other half of the dsRNA duplex of the protein-binding segment of the gRNA. The stretch of nucleotides in the crRNA is complementary to and hybridizes with the stretch of nucleotides in the tracrRNA, forming the dsRNA duplex of the protein-binding domain of the gRNA. Thus, each crRNA can be said to have a corresponding tracrRNA. Exemplary tracrRNA sequences comprise, consist essentially of, or consist of AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUU (SEQ ID NO: 30), AAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 31), or GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 32).

[0239] In systems where both crRNA and tracrRNA are required, the crRNA and the corresponding tracrRNA hybridize to form a gRNA. In systems where only crRNA is required, the crRNA may be the gRNA. The crRNA additionally provides a single-stranded DNA target segment that hybridizes to the complementary strand of the target DNA. When used for intracellular modification, the precise sequence of a given crRNA or tracrRNA molecule can be designed to be specific for the species in which the RNA molecule is used. See, for example, Mali et al. (2013) Science 339(6121):823-826, Jinek et al. (2012) Science 337(6096):816-821, Hwang et al. (2013) Nat. Biotechnol. 31(3):227-229, Jiang et al. (2013) Nat. Biotechnol. 31(3):233-239, and Cong et al. (2013) Science 339(6121):819-823, each of which is incorporated by reference in its entirety for all purposes.

[0240] The DNA-targeting segment (crRNA) of a given gRNA contains a nucleotide sequence that is complementary to a sequence on the complementary strand of the target DNA, as described in more detail below. The DNA-targeting segment of a gRNA interacts with the target DNA in a sequence-specific manner through hybridization (i.e., base pairing). Therefore, the nucleotide sequence of the DNA-targeting segment may vary and determines the location within the target DNA where the gRNA and target DNA interact. The DNA-targeting segment of a given gRNA can be modified to hybridize to any desired sequence within the target DNA. Naturally occurring crRNAs vary depending on the CRISPR / Cas system and organism, but often contain a targeting segment 21-72 nucleotides long flanked by two direct repeats (DRs) that are 21-46 nucleotides long (see, e.g., WO2014 / 131833, incorporated herein by reference in its entirety for all purposes). In S. pyogenes, the DRs are 36 nucleotides long and the targeting segment is 30 nucleotides long. The 3'-located DR is complementary to and hybridizes with the corresponding tracrRNA, which in turn binds to the Cas protein.

[0241] A DNA-targeting segment can have a length of, for example, at least about 12, at least about 15, at least about 17, at least about 18, at least about 19, at least about 20, at least about 25, at least about 30, at least about 35, or at least about 40 nucleotides. Such a DNA-targeting segment can have a length of, for example, about 12 to about 100, about 12 to about 80, about 12 to about 50, about 12 to about 40, about 12 to about 30, about 12 to about 25, or about 12 to about 20 nucleotides. For example, a DNA-targeting segment can be about 15 to about 25 nucleotides (e.g., about 17 to about 20 nucleotides, or about 17, 18, 19, or 20 nucleotides). See, e.g., US2016 / 0024523, incorporated herein by reference in its entirety for all purposes. For Cas9 derived from S. pyogenes, a typical DNA-targeting segment is 16 to 20 nucleotides in length, or 17 to 20 nucleotides in length. For Cas9 from S. aureus, a typical DNA-targeting segment is 21-23 nucleotides in length. For Cpf1, a typical DNA-targeting segment is at least 16 nucleotides in length or at least 18 nucleotides in length.

[0242] In one example, the DNA targeting segment can be about 20 nucleotides in length. However, shorter and longer sequences can also be used for the targeting segment (e.g., 15-25 nucleotides in length, such as 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length). The degree of identity between the DNA targeting segment and the corresponding guide RNA target sequence (or the degree of complementarity between the DNA targeting segment and the other strand of the guide RNA target sequence) can be, for example, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100%. The DNA targeting segment and the corresponding guide RNA target sequence can contain one or more mismatches. For example, the DNA-targeting segment of a guide RNA and the corresponding guide RNA target sequence can include 1 to 4, 1 to 3, 1 to 2, 1, 2, 3, or 4 mismatches (e.g., the total length of the guide RNA target sequence is at least 17, at least 18, at least 19, or at least 20 or more nucleotides). For example, the DNA-targeting segment of a guide RNA and the corresponding guide RNA target sequence can include 1 to 4, 1 to 3, 1 to 2, 1, 2, 3, or 4 mismatches where the total length of the guide RNA target sequence is 20 nucleotides.

[0243] As an example, a guide RNA targeting the RS1 gene can include a DNA-targeting segment (i.e., guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 3148-6241. Alternatively, a guide RNA targeting the RS1 gene can include a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 3148-6241. Alternatively, a guide RNA targeting the RS1 gene can include a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 3148-6241. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 3148-6241. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 3148-6241.Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 3148-6241. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from a sequence set forth in any one of SEQ ID NOs: 3148-6241 (DNA-targeting segment). Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of the sequence set forth in any one of SEQ ID NOs: 3148-6241. Examples of such guide sequences are shown in Tables 2 and 3.

[0244] The guide RNA can target the human RS1 gene. As an example, the guide RNA targeting the RS1 gene can include a DNA-targeting segment (i.e., a guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 3148-4989. Alternatively, the guide RNA targeting the RS1 gene can include a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 3148-4989. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 3148-4989. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 3148-4989. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 3148-4989.Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 3148-4989. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from a sequence set forth in any one of SEQ ID NOs: 3148-4989 (DNA-targeting segment). Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of the sequence set forth in any one of SEQ ID NOs: 3148-4989.

[0245] The guide RNA can be selected to target the human RS1 gene and avoid off-target effects. As an example, the guide RNA targeting the RS1 gene can include a DNA-targeting segment (i.e., a guide sequence) that includes, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351. Alternatively, the guide RNA targeting the RS1 gene can include a DNA-targeting segment that includes, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351. Alternatively, the guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351. Alternatively, the guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351.Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides (the DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from a sequence set forth in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351 (DNA-targeting segment). Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351 that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide.

[0246] The guide RNA can target the human RS1 gene. As an example, the guide RNA targeting the RS1 gene can comprise a DNA-targeting segment (i.e., a guide sequence) that comprises, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304. Alternatively, the guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304 (DNA-targeting segment). Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304 (the DNA-targeting segment).Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides (the DNA-targeting segment) of the sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (the DNA-targeting segment) of the sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304. Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304 (the DNA-targeting segment). Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of the sequence set forth in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304.

[0247] The guide RNA can target the mouse Rs1 gene. As an example, the guide RNA targeting the RS1 gene can include a DNA-targeting segment (i.e., guide sequence) that includes, consists essentially of, or consists of a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NO: 5477 or 5981). Alternatively, the guide RNA targeting the RS1 gene can include a DNA-targeting segment that includes, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides (DNA-targeting segment) of a sequence set forth in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NO: 5477 or 5981). Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NO: 5477 or 5981). Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a sequence (DNA-targeting segment) set forth in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NO: 5477 or 5981).Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (the DNA-targeting segment) of the sequence set forth in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NO: 5477 or 5981). Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) of the sequence set forth in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NO: 5477 or 5981). Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of a sequence that differs by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NO: 5477 or 5981). Alternatively, a guide RNA targeting the RS1 gene can comprise a DNA-targeting segment that comprises, consists essentially of, or consists of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides (DNA-targeting segment) that differ by no more than 3 nucleotides, no more than 2 nucleotides, or no more than 1 nucleotide from the sequence set forth in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NO: 5477 or 5981).

[0248] TracrRNA can be in any form (e.g., full-length tracrRNA or active partial tracrRNA) and of various lengths. They can include primary transcripts or processed forms. For example, tracrRNA (as part of a single guide RNA or as a separate molecule as part of a bimolecular gRNA) can comprise, consist essentially of, or consist of all or a portion of the wild-type tracrRNA sequence (e.g., about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides of the wild-type tracrRNA sequence). Examples of wild-type tracrRNA sequences from S. pyogenes include 171-nucleotide, 89-nucleotide, 75-nucleotide, and 65-nucleotide versions. See, for example, Deltcheva et al. (2011) Nature 471(7340):602-607, WO2014 / 093661, each of which is incorporated by reference in its entirety for all purposes. Examples of tracrRNAs within a single guide RNA (sgRNA) include the tracrRNA segments found within the +48, ​​+54, +67, and +85 versions of the sgRNA, where "+n" indicates that up to +n nucleotides of the wild-type tracrRNA are included in the sgRNA. See US8,697,359, which is incorporated by reference in its entirety for all purposes.

[0249] The percent complementarity between the DNA-targeting segment of the guide RNA and the complementary strand of the target DNA can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%). The percent complementarity between the DNA-targeting segment and the complementary strand of the target DNA can be at least 60% over approximately 20 consecutive nucleotides. As an example, the percent complementarity between the DNA-targeting segment and the complementary strand of the target DNA can be 100% over 14 consecutive nucleotides at the 5'-end of the complementary strand of the target DNA, and can be as low as 0% over the remainder. In such cases, the DNA-targeting segment can be considered to be 14 nucleotides in length. As another example, the percent complementarity between the DNA targeting segment and the complementary strand of the target DNA may be 100% over the seven consecutive nucleotides at the 5' end of the complementary strand of the target DNA, and as low as 0% over the remainder. In such cases, the DNA target segment can be considered to be 7 nucleotides long. In some guide RNAs, at least 17 nucleotides in the DNA target segment are complementary to the complementary strand of the target DNA. For example, the DNA target segment may be 20 nucleotides long and contain one, two, or three mismatches with the complementary strand of the target DNA. In one example, the mismatch is not adjacent to the region of the complementary strand that corresponds to the protospacer adjacent motif (PAM) sequence (i.e., the reverse complement of the PAM sequence) (e.g., the mismatch is at the 5' end of the DNA-targeting segment of the guide RNA, or the mismatch is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 base pairs away from the region of the complementary strand that corresponds to the PAM sequence).

[0250] The protein-binding segment of the gRNA may contain two stretches of nucleotides that are complementary to each other. The complementary nucleotides of the protein-binding segment hybridize to form a double-stranded RNA duplex (dsRNA). The protein-binding segment of the target gRNA interacts with the Cas protein, and the gRNA directs the bound Cas protein to a specific nucleotide sequence within the target DNA via the DNA-targeting segment.

[0251] A single guide RNA can include a DNA-targeting segment and a scaffold sequence (i.e., the protein-binding or Cas-binding sequence of the guide RNA). For example, such a guide RNA can have a 5' DNA-targeting segment linked to a 3' scaffold sequence. Exemplary scaffold sequences comprise, consist essentially of, or consist of the following: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCU (version 1, SEQ ID NO: 33), GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 2, SEQ ID NO: 34), GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 3, SEQ ID NO: 35), GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGU GC (version 4, SEQ ID NO: 36), GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU (version 5, SEQ ID NO: 37), GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (version 6, SEQ ID NO: 38), GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (version 7, SEQ ID NO: 39), or GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGGCACCGAGUCGGUGC (version 8, SEQ ID NO: 53). In some guide sgRNAs, the four terminal U residues of version 6 are absent.In some guide sgRNAs, only one, two, or three of the four terminal U residues in version 6 are present. A guide RNA targeting any of the guide RNA target sequences disclosed herein can include, for example, a DNA-targeting segment at the 5' end of the guide RNA fused to any of the exemplary guide RNA scaffold sequences at the 3' end of the guide RNA. That is, any of the DNA-targeting segments disclosed herein can be attached to the 5' end of any one of the above scaffold sequences to form a single guide RNA (chimeric guide RNA).

[0252] Guide RNAs may contain modifications or sequences that provide additional desirable characteristics (e.g., modified or modulated stability; intracellular targeting; tracking by fluorescent labeling; binding sites for proteins or protein complexes, etc.) That is, guide RNAs may contain one or more modified nucleosides or nucleotides, or one or more non-natural and / or naturally occurring components or structures used in place of or in addition to the standard A, G, C, and U residues. Examples of such modifications include, for example, a 5' cap (e.g., a 7-methylguanylate cap (m7G)), a 3' polyadenylation tail (i.e., a 3' poly(A) tail), a riboswitch sequence (e.g., to allow for regulated stability and / or regulated accessibility by proteins and / or protein complexes), a stability control sequence, a sequence that forms a dsRNA duplex (i.e., a hairpin), a modification or sequence that targets the RNA to a subcellular location (e.g., the nucleus, mitochondria, chloroplasts, etc.), a modification or sequence that provides tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows fluorescent detection, etc.), a modification or sequence that provides a binding site for a protein (e.g., a protein that acts on DNA, including a transcriptional activator, a transcriptional repressor, a DNA methyltransferase, a DNA demethylase, a histone acetyltransferase, a histone deacetylase, etc.), and combinations thereof. Other examples of modifications include an engineered stem-loop duplex, an engineered bulge region, an engineered hairpin 3' of a stem-loop duplex, or any combination thereof. See, e.g., US2015 / 0376586, incorporated herein by reference in its entirety for all purposes. A bulge can be an unpaired region of nucleotides within a duplex comprised of a crRNA-like region and a minimal tracrRNA-like region. A bulge can comprise an unpaired 5'-XXXY-3' on one side of the duplex, where X is any purine and Y can be a nucleotide that can form a wobble pair with a nucleotide on the opposite strand and with a region of unpaired nucleotides on the other side of the duplex.

[0253] Unmodified nucleic acids may be susceptible to degradation. Exogenous nucleic acids may also induce an innate immune response. Modifications may help introduce stability and reduce immunogenicity. Guide RNAs can include modified nucleosides and nucleotides, including, for example, one or more of the following: (1) modification or substitution of one or both of the non-linked phosphate oxygens and / or one or more of the linked phosphate oxygens in a phosphodiester backbone linkage (exemplary backbone modifications); (2) modification or substitution of a component of the ribose sugar, such as modification or substitution of the 2' hydroxyl of the ribose sugar (exemplary sugar modifications); (3) replacement of the phosphate moiety with a dephosphorylated linker (major substitution) (exemplary backbone modifications); (4) modification or substitution of a naturally occurring nucleobase, including non-standard nucleobases (exemplary base modifications); (5) substitution or modification of the ribose-phosphate backbone (exemplary backbone modifications); (6) modification of the 3' or 5' end of the oligonucleotide (e.g., removal, modification, or substitution of a terminal phosphate group, or attachment of a moiety, cap, or linker (3' or 5' cap modifications can include sugar and / or backbone modifications)); and (7) modification or substitution of a sugar (exemplary sugar modifications). Other possible guide RNA modifications include the modification or substitution of uracil or poly-uracil tract.For example, see WO2015 / 048577 and US2016 / 0237455, each of which is incorporated herein by reference in its entirety for all purposes.Similar modifications can be made to Cas-encoding nucleic acids, such as Cas mRNA.For example, Cas mRNA can be modified by using synonymous codons to deplete uridine.

[0254] Chemical modifications such as those described above can be combined to provide modified gRNAs and / or mRNAs containing residues (nucleosides and nucleotides) that can have two, three, four, or more modifications. For example, the modified residues can have modified sugars and modified nucleobases. In one example, all bases of the gRNA are modified (e.g., all bases have modified phosphate groups, such as phosphorothioate groups). For example, all or substantially all phosphate groups of the gRNA can be replaced with phosphorothioate groups. Alternatively or additionally, the modified gRNA can include at least one modified residue at or near the 5' end. Alternatively or additionally, the modified gRNA can include at least one modified residue at or near the 3' end.

[0255] Some gRNAs contain one, two, three, or more modified residues. For example, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% of the positions in the modified gRNA can be modified nucleosides or nucleotides.

[0256] Unmodified nucleic acids may be prone to degradation. Exogenous nucleic acids may also induce innate immune responses. Modifications may help to introduce stability and reduce immunogenicity. Some gRNAs described herein may contain one or more modified nucleosides or nucleotides to introduce stability against intracellular or serum-based nucleases. Some modified gRNAs described herein may exhibit reduced innate immune responses when introduced into a population of cells.

[0257] The gRNA disclosed herein can include backbone modifications, in which the phosphate group of the modified residue can be modified by replacing one or more oxygen atoms with different substituents.Modifications can include extensively replacing unmodified phosphate moieties with modified phosphate groups, as described herein.Backbone modifications of the phosphate backbone can also include changes that result in either an uncharged linker or a charged linker with asymmetric charge distribution.

[0258] Examples of modified phosphate groups include phosphorothioates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramidates, alkyl or aryl phosphonates, and phosphotriesters. The phosphorus atom of an unmodified phosphate group is achiral. However, replacing one of the non-bridging oxygens with one of the atoms or groups of atoms listed above can make the phosphorus atom chiral. The stereogenic phosphorus atom can have either the "R" configuration (Rp) or the "S" configuration (Sp). The backbone can also be modified by replacing the bridging oxygen (i.e., the oxygen connecting the phosphate to the nucleoside) with nitrogen (bridging phosphoramidates), sulfur (bridging phosphorothioates), and carbon (bridging methylene phosphonates). Substitutions can occur at both the bonded oxygen or the bonded oxygen.

[0259] The phosphate group can be substituted into a phosphorus-free connector with certain backbone modifications. In some embodiments, the charged phosphate group can be replaced with a neutral moiety. Examples of moieties that can replace the phosphate group include, but are not limited to, methylphosphonate, hydroxylamino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formate, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, and methyleneoxymethylimino.

[0260] Nucleic acid-mimicking scaffolds can also be constructed such that the phosphate linker and ribose sugar are replaced with nuclease-resistant nucleoside or nucleotide surrogates. Such modifications can include backbone and sugar modifications. In some embodiments, the nucleobases can be linked by a surrogate backbone. Examples include, but are not limited to, morpholino, cyclobutyl, pyrrolidine, and peptide nucleic acid (PNA) nucleoside surrogates.

[0261] Modified nucleosides and nucleotides can contain one or more modifications to the sugar group (sugar modifications). For example, the 2' hydroxyl group (OH) can be modified (e.g., substituted with several different oxy or deoxy substituents). Modifications to the 2' hydroxyl group can increase the stability of nucleic acids because the hydroxyl cannot be deprotonated to form a 2'-alkoxide ion.

[0262] Examples of 2' hydroxyl group modifications include alkoxy or aryloxy (or "R" can be, for example, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar), polyethylene glycol (PEG), O(CH2CHO) n Examples of suitable 2' hydroxyl group modifications include CH2CH2OR, where R can be, for example, H or an optionally substituted alkyl, and n can be an integer from 0 to 20 (e.g., 0 to 4, 0 to 8, 0 to 10, 0 to 16, 1 to 4, 1 to 8, 1 to 10, 1 to 16, 1 to 20, 2 to 4, 2 to 8, 2 to 10, 2 to 16, 2 to 20, 4 to 8, 4 to 10, 4 to 16, and 4 to 20). The 2' hydroxyl group modification can be 2'-O-Me. Similarly, the 2' hydroxyl group modification can be a 2'-fluoro modification, which replaces the 2' hydroxyl group with fluoride. The 2' hydroxyl group modification can be, for example, C 1-6 Alkylene or C 1-6They can include locked nucleic acids (LNAs) that can be connected to the 4' carbon of the same ribose sugar by a heteroalkylene bridge, exemplary bridges being methylene, propylene, ether or amino bridges; O-amino (amino can be, for example, NH; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine or polyamino) and aminoalkoxy, O(CH) n -amino, (amino can be, for example, NH; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino or diheteroarylamino, ethylenediamine, or polyamino). 2' hydroxyl group modifications can include unlocked nucleic acids (UNA), in which the ribose ring lacks a C2'-C3' bond. 2' hydroxyl group modifications can include methoxyethyl groups (MOE) (OCH2CH2OCH3, e.g., PEG derivatives).

[0263] Deoxy 2' modifications include hydrogen (i.e., deoxyribose sugars, e.g., in partial overhanging portions of dsRNA), halo (e.g., bromo, chloro, fluoro, or iodo), amino (amino can be, e.g., NH, alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or amino acid), NH(CHCHNH) n These may include CH2CH2-amino (amino, e.g., as described herein), -NHC(O)R (R may be, e.g., alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar), cyanomercapto, alkyl-thio-alkyl, thioalkoxy, and alkyl, cycloalkyl, aryl, alkenyl, and alkynyl, which may be optionally substituted, e.g., with amino, as described herein.

[0264] Sugar modifications can include sugar groups that can contain one or more carbons with the opposite stereochemical configuration to the corresponding carbon in ribose. Thus, modified nucleic acids can include nucleotides containing, for example, arabinose as the sugar. Modified nucleic acids can also include abasic sugars. These abasic sugars can also be further modified at one or more of the constituent sugar atoms. Modified nucleic acids can also include one or more sugars that are L-configured (e.g., L-nucleosides).

[0265] The modified nucleosides and modified nucleotides described herein that can be incorporated into modified nucleic acids can contain modified bases, also referred to as nucleobases. Examples of nucleobases include, but are not limited to, adenine (A), guanine (G), cytosine (C), and uracil (U). These nucleobases can be modified to provide modified residues that can be incorporated into modified nucleic acids, or can be completely replaced. The nucleobases of the nucleotides can be independently selected from purines, pyrimidines, purine analogs, or pyrimidine analogs. In some embodiments, the nucleobases can include, for example, naturally occurring synthetic derivatives of bases.

[0266] In dual guide RNAs, the crRNA and tracrRNA can each contain modifications. Such modifications can be at one or both ends of the crRNA and / or tracrRNA. In sgRNAs, one or more residues at one or both ends of the sgRNA can be chemically modified, and / or internal nucleosides can be modified, and / or the entire sgRNA can be chemically modified. Some gRNAs contain 5'-end modifications. Some gRNAs contain 3'-end modifications.

[0267] The guide RNAs disclosed herein may comprise one of the modification patterns disclosed in WO2018 / 107028A1, which is incorporated herein by reference in its entirety for all purposes. The guide RNAs disclosed herein may also comprise one of the structure / modification patterns disclosed in US2017 / 0114334, which is incorporated herein by reference in its entirety for all purposes. The guide RNAs disclosed herein may also comprise one of the structure / modification patterns disclosed in WO2017 / 136794, WO2017 / 004279, US2018 / 0187186, or US2019 / 0048338, each of which is incorporated herein by reference in its entirety for all purposes.

[0268] As one example, the nucleotides at the 5' or 3' end of the guide RNA can include phosphorothioate linkages (e.g., the base can have a modified phosphate group that is a phosphorothioate group). For example, the guide RNA can include phosphorothioate linkages between the two, three, or four terminal nucleotides at the 5' or 3' end of the guide RNA. As another example, the nucleotides at the 5' and / or 3' end of the guide RNA can have 2'-O-methyl modifications. For example, the guide RNA can include 2'-O-methyl modifications in the two, three, or four terminal nucleotides at the 5' and / or 3' end (e.g., the 5' end) of the guide RNA. See, for example, WO2017 / 173054A1 and Finn et al. (2018) Cell Rep. 22(9):2227-2235, each of which is incorporated herein by reference in its entirety for all purposes. Other possible modifications are described in more detail elsewhere herein. In one specific example, the guide RNA contains 2'-O-methyl analogs and 3' phosphorothioate internucleotide linkages in the first three 5' and 3' terminal RNA residues. In another specific example, all 2'OH groups that do not interact with the Cas9 protein are replaced with 2'-O-methyl analogs, and the tail region of the guide RNA that minimally interacts with Cas9 is modified with 5' and 3' phosphorothioate internucleotide linkages. Additionally, the DNA targeting segment can have 2'-fluoro modifications at some bases. See, for example, Yin et al. (2017) Nat. Biotech. 35(12):1179-1187, which is incorporated herein by reference in its entirety for all purposes. Other examples of modified guide RNAs are provided, for example, in WO2018 / 107028A1, which is incorporated herein by reference in its entirety for all purposes. Such chemical modifications may, for example, provide guide RNAs with greater stability and protection from exonucleases, allowing them to persist longer in cells than unmodified guide RNAs.Such chemical modifications can also protect against innate immune responses, which can, for example, aggressively degrade RNA or trigger immune cascades that result in cell death.

[0269] As an example, any of the guide RNAs described herein can include at least one modification. In one example, the at least one modification includes a 2'-O-methyl (2'-O-Me) modified nucleotide, a phosphorothioate (PS) internucleotide bond, a 2'-fluoro (2'-F) modified nucleotide, or a combination thereof. For example, the at least one modification can include a 2'-O-methyl (2'-O-Me) modified nucleotide. Alternatively or additionally, the at least one modification can include a phosphorothioate (PS) internucleotide bond. Alternatively or additionally, the at least one modification can include a 2'-fluoro (2'-F) modified nucleotide. In one example, the guide RNAs described herein include one or more 2'-O-methyl (2'-O-Me) modified nucleotides and one or more phosphorothioate (PS) internucleotide bonds.

[0270] The modification can occur anywhere in the guide RNA. For example, the guide RNA comprises a modification in one or more of the first five nucleotides at the 5' end of the guide RNA, and the guide RNA comprises a modification in one or more of the last five nucleotides at the 3' end of the guide RNA, or a combination thereof. For example, the guide RNA can comprise a phosphorothioate bond between the first four nucleotides of the guide RNA, a phosphorothioate bond between the last four nucleotides of the guide RNA, or a combination thereof. Alternatively or additionally, the guide RNA can comprise 2'-O-Me modified nucleotides in the first three nucleotides at the 5' end of the guide RNA, and 2'-O-Me modified nucleotides in the last three nucleotides at the 3' end of the guide RNA, or a combination thereof.

[0271] In one example, a modified gRNA can comprise the following sequence: mN*mN*mN*NNNNNNNNNNNNNNNNNGUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmU*mU*mU*mU (SEQ ID NO: 44), where "N" can be any natural or non-natural nucleotide, and the totality of the N residues can be any of the RS1 sequences described herein. The DNA-targeting segment (e.g., the sequence set forth in SEQ ID NO: 44, where the N residue is replaced with a DNA-targeting segment of any one of SEQ ID NOs: 3148-6241, or any one of SEQ ID NOs: 3148-4989, or any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351, or any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304, or any one of SEQ ID NOs: 4990-6241 (e.g., 5477 or 5981)) is used. The terms "mA," "mC," "mU," and "mG" refer to 2'-O-Me modified nucleotides (A, C, U, and G, respectively). The symbol "*" indicates a phosphorothioate modification. A phosphorothioate linkage or bond refers to a linkage in which sulfur replaces one non-bridging phosphate oxygen in a phosphodiester bond, e.g., a bond between nucleotide bases. When phosphorothioates are used to generate an oligonucleotide, the modified oligonucleotide may also be referred to as an S-oligo. The terms A*, C*, U*, or G* refer to a nucleotide linked to the next (e.g., 3') nucleotide by a phosphorothioate bond. The terms "mA*," "mC*," "mU*," and "mG*" refer to nucleotides (A, C, U, and G, respectively) that are 2'-O-Me substituted and linked to the next (e.g., 3') nucleotide by a phosphorothioate bond.

[0272] Another chemical modification that has been shown to affect the nucleotide sugar ring is halogen substitution. For example, 2'-fluoro (2'-F) substitution of the nucleotide sugar ring can increase oligonucleotide binding affinity and nuclease stability. An abasic nucleotide is a nucleotide lacking a nitrogenous base. An inverted base is one that has a linkage inverted from the normal 5' to 3' linkage (i.e., 5' to 5' or 3' to 3' linkage).

[0273] The abasic nucleotide can be linked by a reverse bond. For example, the abasic nucleotide can be linked to the terminal 5' nucleotide via a 5' to 5' bond, or the abasic nucleotide can be linked to the terminal 3' nucleotide via a 3' to 3' bond. The reverse abasic nucleotide at either the terminal 5' or 3' nucleotide can also be called a reverse abasic end cap.

[0274] In one example, one or more of the first 3, 4, or 5 nucleotides at the 5' end and one or more of the last 3, 4, or 5 nucleotides at the 3' end are modified. The modifications can be, for example, 2'-O-Me, 2'-F, reverse-basic nucleotides, phosphorothioate linkages, or other nucleotide modifications known to enhance stability and / or performance.

[0275] In another example, the first four nucleotides at the 5' end and the last four nucleotides at the 3' end can be linked by phosphorothioate bonds.

[0276] In another example, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end can comprise 2'-O-methyl (2'-O-Me) modified nucleotides. In another example, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end can comprise 2'-fluoro (2'-F) modified nucleotides. In another example, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end can comprise reverse-basic nucleotides.

[0277] The guide RNA can be provided in any form. For example, the gRNA can be provided in the form of RNA, either as two molecules (separate crRNA and tracrRNA) or as one molecule (sgRNA), and optionally in the form of a complex with a Cas protein. The gRNA can also be provided in the form of DNA encoding the gRNA. The DNA encoding the gRNA can encode a single RNA molecule (sgRNA) or separate RNA molecules (e.g., separate crRNA and tracrRNA). In the latter case, the DNA encoding the gRNA can be provided as one DNA molecule or as separate DNA molecules encoding the crRNA and tracrRNA, respectively.

[0278] When the gRNA is provided in the form of DNA, the gRNA can be expressed transiently, conditionally, or constitutively in the cell. The DNA encoding the gRNA can be stably integrated into the genome of the cell and operably linked to a promoter active in the cell. Alternatively, the DNA encoding the gRNA can be operably linked to a promoter in an expression construct. For example, the DNA encoding the gRNA can be within a vector containing a heterologous nucleic acid, such as a nucleic acid encoding a Cas protein. Alternatively, it can be a vector or plasmid separate from the vector containing the nucleic acid encoding the Cas protein. Promoters that can be used in such expression constructs include, for example, promoters active in one or more of eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, pluripotent cells, embryonic stem (ES) cells, adult stem cells, developmentally restricted progenitor cells, induced pluripotent stem (iPS) cells, or one-cell embryos. Such promoters can be, for example, conditional promoters, inducible promoters, constitutive promoters, or tissue-specific promoters. Such promoters can also be, for example, bidirectional promoters. Specific examples of suitable promoters include RNA polymerase III promoters, such as the human U6 promoter, rat U6 polymerase III promoter, or mouse U6 polymerase III promoter. In another example, small tRNA Gln may be used to drive expression of the guide RNA.

[0279] Alternatively, gRNA can be prepared by various other methods.For example, gRNA can be prepared by in vitro transcription using, for example, T7 RNA polymerase (see, for example, WO2014 / 089290 and WO2014 / 065596, each of which is incorporated herein by reference in its entirety for all purposes).Guide RNA can also be a synthetically produced molecule prepared by chemical synthesis.For example, guide RNA can be chemically synthesized to include 2'-O-methyl analogs and 3' phosphorothioate internucleotide linkages in the first three 5' and 3' terminal RNA residues.

[0280] The guide RNA (or nucleic acid encoding the guide RNA) can be in a composition comprising one or more guide RNAs (e.g., 1, 2, 3, 4, or more guide RNAs) and a carrier that increases the stability of the guide RNA (e.g., extends the period below a threshold such that degradation products remain below 0.5% by weight of the starting nucleic acid or protein under given storage conditions (e.g., -20°C, 4°C, or ambient temperature) or increases stability in vivo). Non-limiting examples of such carriers include poly(lactic acid) (PLA) microspheres, poly(D,L-lactic-coglycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipid cochleates, and lipid microtubules. Such compositions can further comprise a Cas protein, such as a Cas9 protein, or a nucleic acid encoding a Cas protein.

[0281] 3. Guide RNA target sequence The target DNA of a guide RNA includes a nucleic acid sequence present in the DNA to which the DNA target segment of the gRNA binds when conditions sufficient for binding exist. Suitable DNA / RNA binding conditions include physiological conditions normally present in cells. Other suitable DNA / RNA binding conditions (e.g., conditions in cell-free systems) are known in the art (see, e.g., Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001), incorporated herein by reference in its entirety for all purposes). The strand of the target DNA that is complementary to and hybridizes with the gRNA can be referred to as the "complementary strand," and the strand of the target DNA that is complementary to the "complementary strand" (and therefore not complementary to the Cas protein or gRNA) can be referred to as the "non-complementary strand" or "template strand."

[0282] The target DNA includes both the sequence on the complementary strand to which the guide RNA hybridizes and the corresponding sequence on the non-complementary strand (e.g., adjacent to the protospacer adjacent motif (PAM)). Unless otherwise specified, as used herein, the term "guide RNA target sequence" specifically refers to the sequence on the non-complementary strand that corresponds to (i.e., its reverse complement) the sequence to which the guide RNA hybridizes on the complementary strand. That is, the guide RNA target sequence refers to the sequence on the non-complementary strand adjacent to the PAM (e.g., upstream or 5' of the PAM in the case of Cas9). The guide RNA target sequence is equivalent to the DNA-targeting segment of the guide RNA, but contains thymine instead of uracil. As an example, the guide RNA target sequence of the SpCas9 enzyme may refer to the sequence upstream of the 5'-NGG-3' PAM on the non-complementary strand. Guide RNA is designed to have complementarity with the complementary strand of target DNA, and the hybridization between the DNA targeting segment of guide RNA and the complementary strand of target DNA promotes the formation of CRISPR complex.Complete complementarity is not necessarily required, as long as there is sufficient complementarity to cause hybridization and promote the formation of CRISPR complex.When guide RNA is referred to herein as targeting guide RNA target sequence, it means that guide RNA hybridizes with the complementary strand sequence of target DNA, which is the reverse complement of guide RNA target sequence on non-complementary strand.

[0283] The target DNA or guide RNA target sequence can comprise any polynucleotide, for example, it can be located in the nucleus or cytoplasm of a cell, or in a cell organelle such as mitochondria or chloroplast.The target DNA or guide RNA target sequence can be any nucleic acid sequence that is endogenous or exogenous to a cell.The guide RNA target sequence can be a sequence that encodes a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory sequence), or it can comprise both.

[0284] Site-specific binding and cleavage of target DNA by a Cas protein can occur at a location determined by both (i) base-pairing complementarity between the guide RNA and the complementary strand of the target DNA and (ii) a short motif, called a protospacer adjacent motif (PAM), on the non-complementary strand of the target DNA. The PAM can be adjacent to the guide RNA target sequence. Optionally, the guide RNA target sequence can be adjacent to the PAM at its 3' end (e.g., in the case of Cas9). Alternatively, the guide RNA target sequence can be adjacent to the PAM at its 5' end (e.g., in the case of Cpf1). For example, the cleavage site of the Cas protein can be about 1 to about 10 or about 2 to about 5 base pairs (e.g., 3 base pairs) upstream or downstream (e.g., within the guide RNA target sequence) of the PAM sequence. In the case of SpCas9, the PAM sequence (i.e., on the non-complementary strand) can be 5'-N1GG-3', where N1 is any DNA nucleotide, and PAM is immediately 3' to the guide RNA target sequence on the non-complementary strand of the target DNA. Thus, the sequence corresponding to PAM on the complementary strand (i.e., the reverse complement) is 5'-CCN2-3', where N2 is any DNA nucleotide, and is immediately 5' to the sequence where the DNA-targeting segment of the guide RNA hybridizes to the complementary strand of the target DNA. In some such cases, N1 and N2 can be complementary, and the N1-N2 base pair can be any base pair (e.g., N1=C and N2=G, N1=G and N2=C, N1=A and N2=T, or N1=T and N2=A). For Cas9 from S. aureus, the PAM can be NNGRRT or NNGRR, where N can be A, G, C, or T, and R can be G or A. For Cas9 from C. jejuni, the PAM can be, for example, NNNNACAC or NNNNRYAC, where N can be A, G, C, or T, and R can be G or A. In some cases (e.g., for FnCpfl), the PAM sequence is 5' upstream and can have the sequence 5'-TTN-3'.

[0285] An example of a guide RNA target sequence is a 20-nucleotide DNA sequence immediately preceding the NGG motif recognized by the SpCas9 protein. For example, two examples of a guide RNA target sequence + PAM are GN 19 NGG (SEQ ID NO: 40) or N 20 NGG (SEQ ID NO: 41). See, for example, WO2014 / 165825, the entire contents of which are incorporated herein by reference for all purposes. A guanine at the 5' end can promote transcription by RNA polymerase in cells. Other examples of guide RNA target sequences and PAMs may contain two guanine nucleotides at the 5' end (e.g., GGN ) to promote efficient transcription by T7 polymerase in vitro. 20 NGG, SEQ ID NO: 42). See, e.g., WO2014 / 065596, incorporated herein by reference in its entirety for all purposes. Other guide RNA target sequences and PAMs may have a length of 4-22 nucleotides of SEQ ID NOs: 40-42, including a 5' G or GG and a 3' GG or NGG. Still other guide RNA target sequences and PAMs may have a length of 14-20 nucleotides of SEQ ID NOs: 40-42.

[0286] A guide RNA targeting the RS1 gene can target, for example, the first intron of the RS1 gene or a sequence adjacent to the first intron of the RS1 gene (e.g., the first exon or second exon of the RS1 gene).

[0287] Formation of a CRISPR complex hybridized to the target DNA can result in cleavage of one or both strands of the target DNA within or near the region corresponding to the guide RNA target sequence (i.e., the guide RNA target sequence on the non-complementary strand of the target DNA and its reverse complement on the complementary strand to which the guide RNA hybridizes). For example, the cleavage site can be within the guide RNA target sequence (e.g., at a defined position relative to the PAM sequence). A "cleavage site" includes the location in the target DNA where the Cas protein generates a single-stranded or double-stranded break. The cleavage site can be single-stranded only (e.g., when a nickase is used) or on both strands of double-stranded DNA. The cleavage sites can be at the same position on both strands (generating blunt ends; e.g., Cas9) or at different sites on each strand (generating staggered ends (i.e., overhangs); e.g., Cpf1). Staggered ends can be generated, for example, by using two Cas proteins, each generating a single-stranded break at a different cleavage site on a different strand, thereby generating a double-stranded break. For example, a first nickase can create a single-stranded break on a first strand of double-stranded DNA (dsDNA), and a second nickase can create a single-stranded break on a second strand of the dsDNA such that an overhang sequence is created. In some cases, the guide RNA target sequence or cleavage site of the first strand nickase is separated from the guide RNA target sequence or cleavage site of the second strand nickase by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, or 1,000 base pairs.

[0288] A guide RNA targeting an RS1 gene, such as the human RS1 gene, can target any desired location in the RS1 gene. For example, a guide RNA targeting an RS1 gene can target the first intron of the RS1 gene or a sequence adjacent to the first intron of the RS1 gene (e.g., the first or second exon of the RS1 gene). For example, a guide RNA target sequence can include any continuous sequence of the RS1 gene. The term RS1 gene includes the genomic region containing the RS1 regulatory promoter and enhancer sequence and the coding sequence. A guide RNA target sequence can include a coding sequence, a non-coding sequence (e.g., a regulatory element such as a promoter or enhancer region), or a combination thereof. For example, a guide RNA target sequence can include any continuous coding sequence of an RS1 coding exon. For example, a guide RNA target sequence can be in exon 1 of the RS1 gene. For another example, a guide RNA target sequence can be in exon 2 of the RS1 gene. For another example, a guide RNA target sequence can be in exon 3 of the RS1 gene. As another example, the guide RNA target sequence may be in exon 4 of the RS1 gene. As another example, the guide RNA target sequence may be in exon 5 of the RS1 gene. As another example, the guide RNA target sequence may be in exon 6 of the RS1 gene. The guide RNA target sequence may also include a sequence contiguous with any of the RS1 introns. As one example, the guide RNA target sequence may be in intron 1 of the RS1 gene. As another example, the guide RNA target sequence may be in intron 2 of the RS1 gene. As another example, the guide RNA target sequence may be in intron 3 of the RS1 gene. As another example, the guide RNA target sequence may be in intron 4 of the RS1 gene. As another example, the guide RNA target sequence may be in intron 5 of the RS1 gene. As another example, the guide RNA target sequence may be in intron 6 of the RS1 gene.

[0289] Guide RNA target sequences can also be selected to minimize off-target modifications or avoid off-target effects (e.g., by avoiding no more than two mismatches with off-target genomic sequences).

[0290] As an example, a guide RNA targeting the RS1 gene can target a guide RNA target sequence set forth in any one of SEQ ID NOs: 54 to 3147. As another example, a guide RNA targeting the RS1 gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a guide RNA target sequence set forth in any one of SEQ ID NOs: 54 to 3147. Examples of such guide RNA target sequences are shown in Tables 2 and 3.

[0291] As one example, a guide RNA targeting the human RS1 gene can target a guide RNA target sequence set forth in any one of SEQ ID NOs: 54 to 1895. As another example, a guide RNA targeting the human RS1 gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a guide RNA target sequence set forth in any one of SEQ ID NOs: 54 to 1895.

[0292] As one example, a guide RNA targeting the human RS1 gene can target a guide RNA target sequence set forth in any one of SEQ ID NOs: 54-57, 60-92, 94-153, and 155-1257. As another example, a guide RNA targeting the human RS1 gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a guide RNA target sequence set forth in any one of SEQ ID NOs: 54-57, 60-92, 94-153, and 155-1257.

[0293] As an example, a guide RNA targeting the human RS1 gene may target a guide RNA target sequence set forth in any one of SEQ ID NOs: 56, 57, 65, 581, 1203, and 1210. As another example, a guide RNA targeting the human RS1 gene may target at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a guide RNA target sequence set forth in any one of SEQ ID NOs: 56, 57, 65, 581, 1203, and 1210.

[0294] As one example, a guide RNA targeting the mouse RS1 gene can target a guide RNA target sequence set forth in any one of SEQ ID NOs: 1896-3147 (e.g., SEQ ID NOs: 2383 or 2887). As another example, a guide RNA targeting the mouse RS1 gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of a guide RNA target sequence set forth in any one of SEQ ID NOs: 1896-3147 (e.g., SEQ ID NOs: 2383 or 2887). [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7]

Table 2-8

Table 2-9

Table 2-10

Table 2-11

Table 2-12

Table 2-13

Table 2-14

Table 2-15

Table 2-16

Table 2-17

Table 2-18

Table 2-19

Table 2-20

Table 2-21

Table 2-22

Table 2-23

Table 2-24

Table 2-25

Table 2-26

Table 2-27

Table 2-28

Table 2-29

Table 2-30

Table 2-31

Table 2-32

Table 2-33

Table 2-34

Table 2-35

Table 2-36

Table 2-37

Table 2-38

Table 2-39

Table 2-40

Table 2-41

Table 2-42

Table 2-43

Table 2-44

Table 2-45

Table 2-46

Table 2-47

Table 2-48

Table 2-49

Table 3-1

Table 3-2

Table 3-3

Table 3-4

Table 3-5

Table 3-6

Table 3-7

Table 3-8

Table 3-9

Table 3-10

Table 3-11

Table 3-12

Table 3-13

Table 3-14

Table 3-15

Table 3-16

Table 3-17

Table 3-18

Table 3-19

Table 3-20

Table 3-21

Table 3-22

Table 3-23

Table 3-24

Table 3-25

Table 3-26

Table 3-27

[0295] B. Other Nuclease Agents and Target Sequences of Nuclease Agents Any nuclease agent that induces a nick or double-strand break in a desired target sequence can be used in the methods and compositions disclosed herein. Naturally occurring or native nuclease agents can be used, as long as the nuclease agent induces a nick or double-strand break in the desired target sequence. Alternatively, modified or engineered nuclease agents can be used. "Engineered nuclease agents" include nucleases that have been engineered (modified or derived) from their native form to specifically recognize and induce a nick or double-strand break in a desired target sequence. Thus, engineered nuclease agents can be derived from native, naturally occurring nuclease agents or can be artificially created or synthesized. Engineered nucleases can, for example, induce a nick or double-strand break in a target sequence that is not a sequence that would be recognized by a natural (unengineered or unmodified) nuclease agent. The modification of the nuclease agent can be as little as one amino acid in a protein cleaving agent or one nucleotide in a nucleic acid cleaving agent. Generating a nick or double-stranded break in a target sequence or other DNA can be referred to herein as "cutting" or "cleaving" the target sequence or other DNA.

[0296] Also provided are active variants and fragments of exemplary target sequences.Such active variants can comprise at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with given target sequence, and active variants retain biological activity and can therefore be recognized and cut by nuclease agent in a sequence-specific manner.Measurement methods for measuring double-stranded cleavage of target sequence by nuclease agent are known.For example, see Frendewey et al.(2010)Methods in Enzymology476:295-307, the entire contents of which are incorporated herein by reference for all purposes.

[0297] The target sequence of the nuclease agent can be located anywhere in or near the target gene locus.The target sequence can be located in the coding region of the gene or in the regulatory region that affects the expression of the gene.The target sequence of the nuclease agent can be located in an intron, exon, promoter, enhancer, regulatory region, or any non-protein-coding region.

[0298] One type of nuclease agent is the transcription activator-like effector nuclease (TALEN). TAL effector nucleases are a class of sequence-specific nucleases that can be used to create double-strand breaks at specific target sequences within the genomes of prokaryotes or eukaryotes. TAL effector nucleases are created by fusing a natural or engineered transcription activator-like (TAL) effector, or a functional portion thereof, to the catalytic domain of an endonuclease, such as FokI. The unique modular TAL effector DNA-binding domain allows for the design of proteins with specific DNA recognition specificity. Therefore, the DNA-binding domain of a TAL effector nuclease can be engineered to recognize specific DNA target sites and can therefore be used to create double-strand breaks at desired target sequences. See WO2010 / 079430; Morbitzer et al. (2010) Proc. Natl. Acad. Sci. USA 107(50):21617-21622, Scholze & Boch (2010) Virulence 1:428-432, Christian et al. Genetics (2010) 186:757-761, Li et al. (2010) Nucleic Acids Res. (2010) doi:10.1093 / nar / gkq704, and Miller et al. (2011) Nat. Biotechnol. 29:143-148, each of which is incorporated by reference in its entirety for all purposes.

[0299] Examples of suitable TAL nucleases and methods for preparing suitable TAL nucleases are disclosed, for example, in US2011 / 0239315A1, US2011 / 0269234A1, US2011 / 0145940A1, US2003 / 0232410A1, US2005 / 0208489A1, US2005 / 0026157A1, US2005 / 0064474A1, US2006 / 0188987A1, and US2006 / 0063231A1, each of which is incorporated by reference in its entirety for all purposes. In various embodiments, for example, TAL effector nucleases are engineered that cleave within or near a target nucleic acid sequence at a locus or genomic locus of interest, where the target nucleic acid sequence is at or near a sequence to be modified by a targeting vector. TAL nucleases suitable for use in the various methods and compositions provided herein include those specifically designed to bind to or near the target nucleic acid sequence modified by the targeting vectors described herein.

[0300] In some TALENs, each TALEN monomer contains 33–35 TAL repeats that recognize a single base pair via two hypervariable residues. In some TALENs, the nuclease agent is a chimeric protein containing a TAL repeat-based DNA-binding domain operably linked to an independent nuclease, such as a FokI endonuclease. For example, a nuclease agent can include a first TAL repeat-based DNA-binding domain and a second TAL repeat-based DNA-binding domain, each of which is operably linked to a FokI nuclease. The first and second TAL repeat-based DNA-binding domains recognize two adjacent target DNA sequences on each strand of the target DNA sequence, separated by spacer sequences of various lengths (12–20 bp). The FokI nuclease subunits dimerize to create an active nuclease that makes a double-stranded break in the target sequence.

[0301] Nuclease agents used in the various methods and compositions disclosed herein can further include zinc finger nucleases (ZFNs). In some ZFNs, each monomer of the ZFN contains three or more zinc finger-based DNA-binding domains, each of which binds to a 3-bp subsite. In other ZFNs, the ZFN is a chimeric protein containing a zinc finger-based DNA-binding domain operably linked to a separate nuclease, such as a FokI endonuclease. For example, the nuclease agent can include a first ZFN and a second ZFN, each of which is operably linked to a FokI nuclease subunit, which recognize two adjacent target DNA sequences on each strand of the target DNA sequence, separated by an approximately 5-7 bp spacer. The FokI nuclease subunits dimerize to generate an active nuclease that makes a double-stranded cleavage. See, e.g., US2006 / 0246567, US2008 / 0182332, US2002 / 0081614, US2003 / 0021776, WO / 2002 / 057308A2, US2013 / 0123484, US2010 / 0291048, WO / 2011 / 017293A2, and Gaj et al. (2013) Trends Biotechnol., 31(7):397-405, each of which is incorporated by reference in its entirety for all purposes.

[0302] Active variants and fragments of nuclease agents (i.e., engineered nuclease agents) are also provided. Such active variants may have at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to the native nuclease agent, and the active variants retain the ability to cleave at the desired target sequence and thus retain nick or double-strand break-inducing activity. For example, any of the nuclease agents described herein may be modified from the native endonuclease sequence and designed to recognize and induce nicks or double-strand breaks at target sequences not recognized by the native nuclease agent. Thus, some engineered nucleases have the specificity to induce nicks or double-strand breaks at target sequences that are different from the corresponding native nuclease agent target sequence. Assays for nick or double-strand break-inducing activity are known and generally measure the overall activity and specificity of an endonuclease on a DNA substrate containing the target sequence.

[0303] Nuclease agents can be introduced into cells or animals by any known means. A polypeptide encoding a nuclease agent can be directly introduced into cells or animals. Alternatively, a polynucleotide encoding a nuclease agent can be introduced into cells or animals. Once a polynucleotide encoding a nuclease agent is introduced, the nuclease agent can be expressed transiently, conditionally, or constitutively in cells. The polynucleotide encoding a nuclease agent can be included in an expression cassette and can be operably linked to a conditional promoter, an inducible promoter, a constitutive promoter, or a tissue-specific promoter. Examples of promoters are discussed in more detail elsewhere herein. Alternatively, a nuclease agent can be introduced into cells as mRNA encoding the nuclease agent.

[0304] The polynucleotide encoding the nuclease agent can be stably integrated into the genome of the cell and operably linked to a promoter active within the cell. Alternatively, the polynucleotide encoding the nuclease agent can be in an expression vector or a targeting vector.

[0305] When nuclease agent is provided to cell by introducing polynucleotide that encodes nuclease agent, such polynucleotide that encodes nuclease agent can be modified to replace the codon that has higher frequency of use in target cell compared with the polynucleotide sequence that encodes naturally occurring nuclease agent.For example, compared with the polynucleotide sequence that encodes naturally occurring nuclease agent, the polynucleotide that encodes nuclease agent can be modified to the alternative codon that is more frequently used in the given eukaryotic cell of interest, including human cell, non-human cell, mammalian cell, rodent cell, mouse cell, rat cell or any other target host cell.

[0306] The term "nuclease agent target sequence" includes a DNA sequence in which a nick or double-strand break is induced by a nuclease agent. The nuclease agent target sequence can be endogenous (or natural) to the cell, or the target sequence can be exogenous to the cell. A target sequence that is exogenous to a cell does not naturally occur within the genome of the cell. The target sequence can also be exogenous to a polynucleotide of interest that is desired to be placed at the target locus. In some cases, the target sequence occurs only once in the genome of the host cell.

[0307] The length of the target sequence may vary, for example, comprising a target sequence of about 30-36 bp for a zinc finger nuclease (ZFN) pair (i.e., about 15-18 bp for each ZFN), about 36 bp for a transcription activator-like effector nuclease (TALEN), or about 20 bp for a CRISPR / Cas9 guide RNA.

[0308] VI. Cells or Animals or Genomes Comprising Nucleic Acid Constructs and / or Nuclease Agents or Nucleic Acids Encoding Nuclease Agents Genomes, cells, and animals produced by the methods disclosed herein are also provided. Genomes, cells, and animals containing a nucleic acid construct containing a retinoschisin-encoding sequence (i.e., encoding a retinoschisin protein or a fragment or variant thereof) for incorporation into and expression from a target genomic locus, vector, lipid nanoparticle, or composition described herein are also provided. Genomes, cells, and animals containing the described nuclease agents or nucleic acids encoding the nuclease agents (e.g., targeted to the endogenous RS1 locus) or vectors, lipid nanoparticles, or compositions described herein are also provided. The genomes, cells, or animals can contain a nucleic acid construct integrated into the genome at a target genomic locus (e.g., the RS1 locus) and express a retinoschisin protein or a fragment or variant thereof. Once integrated into the target genomic locus, the retinoschisin-encoding sequence can be operably linked to an endogenous promoter at the target genomic locus or can be operably linked to an exogenous promoter present in the nucleic acid construct. When the nucleic acid construct is a bidirectional nucleic acid construct disclosed herein, the genome, cell, or animal can express a first retinoschisin protein or its fragment or variant, or can express a second retinoschisin protein or its fragment or variant.In some genomes, cells, or animals, the target genomic locus is the RS1 locus.For example, the nucleic acid construct can be integrated into the genome into intron 1 of the endogenous RS1 locus.Then, the endogenous RS1 exon 1 can be spliced ​​to the coding sequence of the retinoschisin protein or its fragment or variant in the nucleic acid construct. In specific examples, the modified RS1 locus comprising the nucleic acid construct integrated into the genome comprises, consists essentially of, or encodes a protein consisting of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 2 or 4.In specific examples, the modified RS1 locus comprising the nucleic acid construct integrated into the genome comprises an RS1 coding sequence that comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 6, 11, or 12.

[0309] In some genomes, cells, or animals, integration of a nucleic acid construct into the endogenous RS1 locus prevents transcription of the endogenous RS1 gene downstream of the integration site. For example, integration of a nucleic acid construct into the endogenous RS1 locus can reduce or eliminate expression of the endogenous retinoschisin protein and replace it with expression of the retinoschisin protein or a fragment or variant thereof encoded by the nucleic acid construct. In one example, integration of a nucleic acid construct into the endogenous RS1 locus reduces expression of the endogenous retinoschisin protein. In another example, integration of a nucleic acid construct into the endogenous RS1 locus eliminates expression of the endogenous retinoschisin protein. In a specific example, the endogenous RS1 locus contains a mutant RS1 gene containing a mutation that causes X-linked juvenile retinoschisis, and expression of the nucleic acid construct integrated into the genome reduces or eliminates expression of the mutant RS1 gene.

[0310] The target genomic locus into which the nucleic acid construct is stably integrated can be heterozygous for the retinoschisin coding sequence from the nucleic acid construct, or homozygous for the retinoschisin coding sequence from the nucleic acid construct. Diploid organisms have two alleles at each locus. Each pair of alleles represents a genotype at a particular locus. A genotype is described as homozygous if there are two identical alleles at a particular locus, and as heterozygous if the two alleles are different. An animal comprising the nucleic acid construct described herein integrated into its genome can contain the nucleic acid construct at the target genomic locus in its germline.

[0311] The genomes, cells, or animals provided herein may be, for example, eukaryotes, including, for example, animals, mammals, non-human mammals, and humans. The term "animal" includes mammals, fish, and birds. A mammal may be, for example, a non-human mammal, a human, a rodent, a rat, a mouse, or a hamster. Other non-human mammals include, for example, non-human primates, monkeys, apes, cats, dogs, rabbits, horses, oxen, deer, bison, and livestock (e.g., bovine species such as cows and steers, ovine species such as sheep and goats, and porcine species such as pigs and wild boars). Birds include, for example, chickens, turkeys, ostriches, geese, ducks, and the like. Domestic and agricultural animals are also included. The term "non-human" excludes humans.

[0312] The cell may be an isolated cell (e.g., in vitro) or may be in vivo in an animal. The cell may also be in any type of undifferentiated or differentiated state. For example, the cell may be a totipotent cell, a pluripotent cell (e.g., a human pluripotent cell, or a non-human pluripotent cell such as a mouse embryonic stem (ES) cell or a rat ES cell), or a non-pluripotent cell. Totipotent cells include undifferentiated cells that can give rise to any cell type, and pluripotent cells include undifferentiated cells that have the ability to develop into two or more differentiated cell types.

[0313] The cells provided herein may also be germ cells (e.g., sperm or oocytes). The cells may be mitotically competent or mitotically inactive, meiotically competent or meiotically inactive. Similarly, the cells may also be primary somatic cells or cells that are not primary somatic cells. Somatic cells include any cell that is not a gamete, germ cell, gamete cell, or undifferentiated stem cell. For example, the cells may be hepatocytes, kidney cells, hematopoietic cells, endothelial cells, epithelial cells, fibroblasts, mesenchymal cells, keratinocytes, blood cells, melanocytes, monocytes, mononuclear cells, monocyte precursor cells, B cells, erythroblast-megakaryocytic cells, eosinophils, macrophages, T cells, pancreatic islet beta cells, exocrine cells, pancreatic precursor cells, endocrine precursor cells, adipocytes, preadipocytes, neurons, glial cells, neural stem cells, neurons, hepatoblasts, hepatocytes, cardiomyocytes, skeletal myoblasts, hepatocytes, leuk ... The cell may be a smooth muscle cell, a duct cell, an acinar cell, an alpha cell, a beta cell, a delta cell, a PP cell, a bile duct cell, a white or brown adipocyte, or an ocular cell (e.g., a trabecular meshwork cell, a retinal pigment epithelial cell, a retinal microvascular endothelial cell, a periretinal cell, a conjunctival epithelial cell, a conjunctival fibroblast, an iris pigment epithelial cell, a corneal cell, a lens epithelial cell, a non-pigmented ciliated epithelial cell, a choroidal fibroblast, a photoreceptor cell, a ganglion cell, a bipolar cell, a horizontal cell, or an amacrine cell). For example, the cell may be an ocular cell such as a retinal cell (e.g., a photoreceptor).

[0314] The cells provided herein may be normal, healthy cells, or may be diseased or mutant cells. For example, the cells may contain one or more mutations associated with or causing XLRS (e.g., encoding an R141C substitution in the retinoschisin protein).

[0315] The animals provided herein may be human or non-human. Non-human animals containing the nucleic acids or expression cassettes described herein can be produced by methods described elsewhere herein. The term "animal" includes mammals, fish, and birds. Mammals include, for example, humans, non-human primates, monkeys, apes, cats, dogs, horses, oxen, deer, bison, sheep, rabbits, rodents (e.g., mice, rats, hamsters, and guinea pigs), and livestock (e.g., bovine species such as cows and steers, ovine species such as sheep and goats, and porcine species such as pigs and wild boars). Birds include, for example, chickens, turkeys, ostriches, geese, and ducks. Domestic and agricultural animals are also included. The term "non-human animal" excludes humans. Specific examples of non-human animals include rodents such as mice and rats.

[0316] Non-human animals may be from any genetic background. For example, suitable mice may be from the 129 strain, the C57BL / 6 strain, a hybrid of 129 and C57BL / 6, the BALB / c strain, or the Swiss-Webster strain. Examples of 129 strains include 129P1, 129P2, 129P3, 129X1, 129S1 (e.g., 129S1 / SV, 129S1 / Svlm), 129S2, 129S4, 129S5, 129S9 / SvEvH, 129S6 (129 / SvEvTac), 129S7, 129S8, 129T1, and 129T2. See, for example, Festing et al. (1999) Mammalian Genome 10:836, incorporated herein by reference in its entirety for all purposes. Examples of C57BL strains include C57BL / A, C57BL / An, C57BL / GrFa, C57BL / Kal_wN, C57BL / 6, C57BL / 6J, C57BL / 6ByJ, C57BL / 6NJ, C57BL / 10, C57BL / 10ScSn, C57BL / 10Cr, and C57BL / Ola. Suitable mice can also be hybrids of the aforementioned 129 strain and the aforementioned C57BL / 6 strain (e.g., 50% 129 and 50% C57BL / 6). Similarly, suitable mice can be hybrids of the aforementioned 129 strain or hybrids of the aforementioned BL / 6 strain (e.g., 129S6 (129 / SvEvTac) strain).

[0317] Similarly, rats may be from, for example, the ACI rat strain, the Dark Agouti (DA) rat strain, the Wistar rat strain, the LEA rat strain, the Sprague Dawley (SD) rat strain, or a Fischer rat strain such as Fischer F344 or Fischer F6. Rats may also be obtained from strains derived from hybrids of two or more of the above-mentioned strains. For example, suitable rats may be from the DA strain or the ACI strain. The ACI rat strain has a white belly and paws, and an RT1 av1The Dark Agouti (DA) rat strain is characterized by the presence of the black agouti haplotype. Such strains are available from a variety of sources, including Harlan Laboratories. The Dark Agouti (DA) rat strain possesses the agouti coat and RT1 av1 The rat is characterized by having haplotype.Such rats can be obtained from various sources, including Charles River and Harlan Laboratories.In some cases, suitable rats can be derived from inbred rat strains.See, for example, US2014 / 0235933, the entirety of which is incorporated herein by reference for all purposes.

[0318] VII. METHODS FOR MODIFYING TARGETED GENETIC LOCUS, EXPRESSING RETINOSCHISIN IN CELLS, OR TREATING XLRS Also provided herein are methods for modifying a target genomic locus or for expressing retinoschisin in a cell using a nucleic acid construct comprising a retinoschisin coding sequence (i.e., encoding a retinoschisin protein or a fragment or variant thereof) for integration into and expression from a target genomic locus as disclosed herein. Also provided herein ...

Claims

1. 1. A composition for use in expressing retinoschisin in a cell or for use in integrating a coding sequence for a retinoschisin protein or a fragment thereof into a target genomic locus in a cell, comprising: (a) a nucleic acid construct for integration into the target genomic locus, wherein the nucleic acid construct is bidirectional; and (I) a first segment comprising a first coding sequence of a first retinoschisin protein or a fragment thereof, wherein the first coding sequence comprises exons 2-6 of human RS1 or a degenerate variant thereof; (II) a second segment comprising the reverse complement of a second coding sequence of a second retinoschisin protein or fragment thereof, wherein the second coding sequence comprises exons 2-6 of human RS1 or a degenerate variant thereof; and a nucleic acid construct comprising: (b) a nuclease agent or a nucleic acid encoding said nuclease agent, wherein said nuclease agent targets a nuclease target sequence at said target genomic locus; and Including, the target genomic locus is at an endogenous RS1 locus that comprises an endogenous RS1 gene; integration of the nucleic acid construct into the endogenous RS1 locus prevents transcription of the endogenous RS1 gene downstream of the integration site; said integration of said nucleic acid construct into said endogenous RS1 locus in said cell reduces or eliminates expression of an endogenous retinoschisin protein and replaces it with expression of said first retinoschisin protein or fragment thereof or said second retinoschisin protein or fragment thereof encoded by said nucleic acid construct; the nucleic acid construct does not contain homology arms, the nucleic acid construct does not comprise a promoter driving expression of the first retinoschisin protein or fragment thereof or the second retinoschisin protein or fragment thereof; the first retinoschisin protein or fragment thereof is identical to the second retinoschisin protein or fragment thereof, and the second coding sequence employs codon usage that differs from the codon usage of the first coding sequence; composition.

2. The composition for use of claim 1 , wherein the nucleic acid construct comprises a fragment or portion of the first intron of human RS1 located 5′ of the coding sequence.

3. The composition for use according to claim 1 or 2, wherein the nucleic acid construct comprises a polyadenylation signal sequence located 3' of the coding sequence.

4. The composition for use according to any one of claims 1 to 3, wherein the nucleic acid construct comprises a splice acceptor site located 5' to the coding sequence.

5. The composition for use according to claim 4, wherein the splice acceptor site is derived from intron 1 of human RS1.

6. the second segment is located 3' of the first segment; both the first retinoschisin protein or fragment thereof and the second retinoschisin protein or fragment thereof are human retinoschisin proteins or fragments thereof; both the first coding sequence and the second coding sequence comprise complementary DNA (cDNA) comprising exons 2-6 of human RS1 or a degenerate variant thereof; the first segment comprises a first polyadenylation signal sequence located 3' of the first coding sequence, and the second segment comprises the reverse complement of a second polyadenylation signal sequence located 5' of the reverse complement of the second coding sequence; the first segment comprises a first splice acceptor site located 5' of the first coding sequence, and the second segment comprises the reverse complement of a second splice acceptor site located 3' of the reverse complement of the second coding sequence; A composition for use according to any one of claims 1 to 5.

7. The composition for use according to any one of claims 1 to 6, wherein the nuclease target sequence of the target genomic locus is in the first intron of the endogenous RS1 gene of the endogenous RS1 locus.

8. The composition for use according to any one of claims 1 to 7, wherein the nucleic acid construct is in a viral vector.

9. The composition for use according to claim 8, wherein the viral vector is an adeno-associated viral (AAV) viral vector.

10. The composition for use according to claim 9, wherein the AAV is AAV2, AAV5, AAV8, or AAV7m8.

11. 11. The composition for use according to any one of claims 1 to 10, wherein the nuclease agent is a Cas protein and a guide RNA, and the nuclease target sequence is a guide RNA target sequence.

12. 12. The composition for use of claim 11, wherein the Cas protein is a Cas9 protein.

13. 13. The composition for use of claim 11 or 12, wherein the composition comprises the nucleic acid encoding the Cas protein, the nucleic acid comprising DNA encoding the Cas protein, and the composition comprising the DNA encoding the guide RNA.

14. 14. The composition for use of claim 13, wherein the DNA encoding the Cas protein and the DNA encoding the guide RNA are in one or more viral vectors.

15. 13. The composition for use of claim 11 or 12, wherein the composition comprises the nucleic acid encoding the Cas protein, the nucleic acid comprising messenger RNA encoding the Cas protein, and the composition comprises the guide RNA in the form of RNA.

16. 16. The composition for use of claim 15, wherein the composition comprises the guide RNA in the form of RNA, and the guide RNA and the messenger RNA encoding the Cas protein are within a lipid nanoparticle.

17. 17. The composition for use according to any one of claims 1 to 16, wherein the endogenous RS1 locus comprises a mutated RS1 gene comprising a mutation that causes X-linked juvenile retinoschisis.

18. The composition for use according to any one of claims 1 to 17, wherein the cells are human cells.

19. The composition for use according to any one of claims 1 to 18, wherein the cells are retinal cells.

20. The composition for use according to any one of claims 1 to 19, wherein the cell is in vivo in an animal.

21. 21. The composition for use according to claim 20, wherein integration of the nucleic acid construct results in restoration of retinal structure.

Citation Information

Patent Citations

  • Methods and compositions for the treatment of eye gene-related diseases

    JP2016510221A

  • Gene editing for autosomal dominant diseases

    WO2019183630A2