CRISPR and AAV strategies for the treatment of X-linked juvenile retinoschisis
The insertion or expression of retinal clexin-encoded sequences in retinal cells through CRISPR and AAV strategies solve the treatment problem of XLRS, and the expression of functional retinal clexin protein is achieved, resuming retinal function, and preventing the transcription of endogenous defective genes.
Patent Information
- Application Number
- CN202080077587.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-08
- Filing Date
- 2020-11-07
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2040-11-07
AI Technical Summary
Existing gene therapies fail to effectively treat X-linked adolescent retinoschisis (XLRS), resulting in vision loss, and new therapeutic strategies are needed.
Through CRISPR and AAV strategies, retinospermin coding sequences are inserted or expressed into target genomic loci, and bidirectional nucleic acid constructs and vectors such as AAV vectors or lipid nanoparticles are used to achieve the expression or integration of retinospermin proteins, replacing the expression of endogenous defective genes.
It is achieved effective expression of functional retinospertin protein in cells, replace defective genes, reduce or eliminate the expression of endogenous retinospertin protein, prevent the transcription of endogenous RS1 gene downstream of the integration site, and restore retinal function.
Smart Images

Figure CN114746125B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Application No. 62 / 932,608, filed November 8, 2019, which is incorporated herein by reference in its entirety for all purposes.
[0003] References to sequence listings submitted as text files via EFS WEB
[0004] The sequence listing written into file 694232SEQLIST.txt is 1.22 megabytes, was created on November 6, 2020, and is hereby incorporated by reference. Background Art
[0005] The RS1 gene encodes a highly conserved extracellular protein involved in the cellular organization of the retina. The gene is assembled and secreted as a homo-oligomeric protein complex by photoreceptors and bipolar cells. More than 200 mutations have been detected in RS1, many of which lead to the early onset of macular degeneration due to non-functional protein or absence of protein secretion. Lack of functional Rs1 expression can cause splitting within the retinal layers, leading to early and progressive vision loss associated with X-linked juvenile retinoschisis (XLRS). Although there have been clinical trials for gene therapy of XLRS, the trials have not reached their endpoints. New strategies are needed to treat XLRS. Summary of the Invention
[0006] Provided are nucleic acid constructs and compositions that allow for insertion of a retinoschisis coding sequence into a target genomic locus, such as the endogenous RS1 locus, and / or expression of a retinoschisis coding sequence. The nucleic acid constructs and compositions can be used in methods for integration into a target genomic locus and / or expression in cells, or in methods for treating X-linked juvenile retinoschisis.
[0007] In one aspect, bidirectional nucleic acid constructs for integration into a target genomic locus are provided. Some such nucleic acid constructs include: (a) a first segment comprising a first coding sequence for a first retinoschizin protein or a fragment thereof; and (b) a second segment comprising the reverse complement of a second coding sequence for a second retinoschizin protein or a fragment thereof. In some such constructs, the second segment is positioned 3' (i.e., downstream) of the first segment.
[0008] In some such constructs, the first retinoschizin protein or its fragment is a human retinoschizin protein or its fragment, the second retinoschizin protein or its fragment is a human retinoschizin protein or its fragment, or the first retinoschizin protein or its fragment and the second retinoschizin protein or its fragment are both human retinoschizin protein or its fragment. In some such constructs, the first coding sequence includes complementary DNA (cDNA), is essentially composed of it or is composed of it, the second coding sequence includes cDNA, is essentially composed of it or is composed of it, or the first coding sequence and the second coding sequence both include cDNA, is essentially composed of it or is composed of it. In some such constructs, the first coding sequence includes exons 2-6 of people RS1 or its degenerate variant, is essentially composed of it or is composed of it, the second coding sequence includes exons 2-6 of people RS1 or its degenerate variant, is essentially composed of it or is composed of it, or the first coding sequence and the second coding sequence both include exons 2-6 of people RS1 or its degenerate variant, is essentially composed of it or is composed of it.
[0009] In some such constructs, the first segment comprises a fragment or portion of a first intron of human RS1 located 5' (i.e., upstream) of the first coding sequence, and / or the second segment comprises the reverse complement of a fragment or portion of a second intron of human RS1 located 3' (i.e., downstream) of the reverse complement of the second coding sequence.
[0010] In some such constructs, the first retinoschizin protein or fragment thereof is identical to the second retinoschizin protein or fragment thereof. In some such constructs, the second coding sequence uses a codon usage that is different from the codon usage of the first coding sequence. In some such constructs, the second segment has at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, or at least about 99% complementarity to the first segment. In some such constructs, the second segment is less than about 30%, less than about 35%, less than about 40%, less than about 45%, less than about 50%, less than about 55%, less than about 60%, less than about 65%, less than about 70%, less than about 75%, less than about 80%, less than about 85%, less than about 90%, less than about 95%, less than about 97%, or less than about 99% complementary to the first segment. In some such constructs, the reverse complement of the second coding sequence: (a) is not significantly complementary to the first coding sequence; (b) is not significantly complementary to a fragment of the first coding sequence; (c) is highly complementary to the first coding sequence; (d) is highly complementary to the fragment of the first coding sequence; (e) is at least about 60%, at least about 70%, at least about 80%, or at least about 90% identical to the reverse complement of the first coding sequence; (f) is about 50% to about 80% identical to the reverse complement of the first coding sequence; or (g) is about 60% to about 100% identical to the reverse complement of the first coding sequence.
[0011] In some such constructs, the first segment is connected to the second segment by a linker. Optionally, the linker is about 5 to about 2000 nucleotides in length.
[0012] In some such constructs, the first segment includes a first polyadenylation signal sequence located 3' to the first coding sequence, and the second segment includes the reverse complement of a second polyadenylation signal sequence located 5' to the reverse complement of the second coding sequence. Optionally, the first polyadenylation signal sequence is different from the second polyadenylation signal sequence.
[0013] In some such constructs, the nucleic acid construct does not include a promoter that drives the expression of the first retinoschizine protein or its fragment or the second retinoschizine protein or its fragment. In some such constructs, the first segment includes a first splice acceptor site located at the 5' end of the first coding sequence, and the second segment includes the reverse complement of the second splice acceptor site located at the 3' end of the reverse complement of the second coding sequence. Optionally, the first splice acceptor site is from the RS1 gene, the second splice acceptor site is from the RS1 gene, or both the first splice acceptor site and the second splice acceptor site are from the RS1 gene. Optionally, the first splice acceptor site is from intron 1 of human RS1, the second splice acceptor site is from intron 1 of human RS1, or both the first acceptor site and the second splice acceptor site are from intron 1 of human RS1.
[0014] In some such constructs, the nucleic acid construct does not include homology arms. In some such constructs, the nucleic acid construct includes homology arms. In some such constructs, the nucleic acid construct is single-stranded. In some such constructs, the nucleic acid construct is double-stranded. In some such constructs, the nucleic acid construct includes DNA.
[0015] In some such constructs, the first coding sequence is codon optimized to express in a host cell, the second coding sequence is codon optimized to express in the host cell, or the first coding sequence and the second coding sequence are codon optimized to express in the host cell. In some such constructs, the nucleic acid construct includes one or more terminal structures in the following terminal structures: a hairpin, a loop, an inverted terminal repeat (ITR) or a toroid. Optionally, the nucleic acid construct includes ITR.
[0016] In some such constructs, the first retinoschizine protein or fragment thereof and / or the second retinoschizine protein or fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 5. In some such constructs, the first coding sequence and / or the second coding sequence comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 8 or 9. In some such constructs, the first coding sequence comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO:8, and the second coding sequence comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO:9. In some such constructs, the nucleic acid construct comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 46 or 47.
[0017] In some such constructs, the second segment is located 3' to the first segment, the first retinoschizine protein or fragment thereof and the second retinoschizine protein or fragment thereof are both human retinoschizine proteins or fragments thereof, the first retinoschizine protein or fragment thereof is identical to the second retinoschizine protein or fragment thereof, the first coding sequence and the second coding sequence both comprise complementary DNA (cDNA) comprising exons 2-6 of human RS1 or a degenerate variant thereof, the second coding sequence employs a codon usage that differs from the codon usage of the first coding sequence, wherein the first segment comprises a protein located 3' to the first segment. The first coding sequence comprises a first polyadenylation signal sequence located at the 3' end of the first coding sequence, and the second segment comprises the reverse complement of the second polyadenylation signal sequence located at the 5' end of the reverse complement of the second coding sequence, the first segment comprises a first splice acceptor site located at the 5' end of the first coding sequence, and the second segment comprises the reverse complement of the second splice acceptor site located at the 3' end of the reverse complement of the second coding sequence, the nucleic acid construct does not comprise a promoter driving expression of the first retinoschizine protein or fragment thereof or the second retinoschizine protein or fragment thereof, and the nucleic acid construct does not comprise homology arms.
[0018] On the other hand, a vector comprising any of the above-mentioned bidirectional nucleic acid constructs is provided. Some such vectors are viral vectors. Optionally, the vector is an adeno-associated virus (AAV) vector. Optionally, the AAV comprises a single-stranded genome (ssAAV). Optionally, the AAV comprises a self-complementary genome (scAAV). Optionally, the AAV is selected from the group consisting of: AAV2, AAV5, AAV8 or AAV7m8.
[0019] Some such vectors do not include a promoter driving expression of the first retinoschizine protein or fragment thereof or the second retinoschizine protein or fragment thereof. Some such vectors do not include homology arms. Some such vectors do include homology arms.
[0020] In another aspect, a lipid nanoparticle comprising any of the above-described bidirectional nucleic acid constructs is provided.
[0021] In another aspect, cells comprising any of the above-described bidirectional nucleic acid constructs are provided. Some such cells are in vitro. Some such cells are in vivo. Some such cells are mammalian cells. Some such cells are human cells. Some such cells are retinal cells.
[0022] Some such cells express the first retinoschizin protein or its fragment or the second retinoschizin protein or its fragment. In some such cells, the nucleic acid construct is integrated into the target genomic locus on the genome. In some such cells, the target genomic locus is an endogenous RS1 locus. Optionally, the nucleic acid construct is integrated into intron 1 of the endogenous RS1 locus on the genome. Optionally, endogenous RS1 exon 1 is spliced into the first coding sequence or the second coding sequence of the nucleic acid construct. Optionally, the modified RS1 locus encoding of the nucleic acid construct integrated on the genome includes a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 2 or 4, is essentially composed of or consists of a protein.
[0023] In some such cells, the nucleic acid construct is integrated into the endogenous RS1 locus to prevent the endogenous RS1 gene transcription downstream of the integration site. Optionally, the nucleic acid construct is integrated into the endogenous RS1 locus to reduce or eliminate the expression of endogenous retinoschizin protein, and the expression of the retinoschizin protein or its fragment encoded by the nucleic acid construct is substituted for the expression of the endogenous retinoschizin protein. Optionally, the endogenous RS1 locus includes the RS1 gene of mutation, the RS1 gene of mutation includes the mutation that causes X-linked juvenile retinoschizin, and the expression of the nucleic acid construct integrated on the genome reduces or eliminates the expression of the RS1 gene of mutation.
[0024] On the other hand, a nucleic acid construct for homology-independent targeting integration into the target genome locus is provided. Some such nucleic acid constructs include the coding sequence of the nuclease target sequence of the retinoschizin protein or its fragment on each side. A nucleic acid construct for homologous recombination with the target locus is also provided. Some such nucleic acid constructs include the coding sequence of the retinoschizin protein or its fragment on each side with a homology arm, optionally wherein the coding sequence and the homology arm are further flanked by the target sequence of the nuclease agent on each side. Optionally, the length of each homology arm is between about 25 nucleotides and about 2.5kb.
[0025] In some such constructs for homology-independent targeted integration, the nuclease target sequence in the nucleic acid construct is the same as the nuclease target sequence used for integration into the target genomic locus, wherein the nuclease target sequence in the target genomic locus is destroyed when the nucleic acid construct is inserted in the correct orientation, but is rearranged when the nucleic acid construct is inserted into the target genomic locus in the reverse orientation.
[0026] In some such constructs, the retinoschizin protein or fragment thereof is a human retinoschizin protein or fragment thereof. In some such constructs, the coding sequence for the retinoschizin protein or fragment thereof comprises, consists essentially of, or consists of complementary DNA (cDNA). In some such constructs, the coding sequence for the retinoschizin protein or fragment thereof comprises, consists essentially of, or consists of exons 2-6 of human RS1 or a degenerate variant thereof.
[0027] In some such constructs, the nucleic acid construct includes a fragment or part of the first intron of the human RS1 positioned at the 5' end of the coding sequence. In some such constructs, the nucleic acid construct does not include a promoter driving the expression of the retinoschizin protein or its fragment. In some such constructs, the nucleic acid construct includes a polyadenylation signal sequence positioned at the 3' end of the coding sequence. In some such constructs, the nucleic acid construct includes a splice acceptor site positioned at the 5' end of the coding sequence. Optionally, the splice acceptor site is from the RS1 gene. Optionally, the splice acceptor site is from intron 1 of human RS1.
[0028] Some such constructs are single-stranded. Some such constructs are double-stranded. Some such constructs comprise DNA. In some such constructs, the coding sequence is codon-optimized for expression in a host cell.
[0029] In some such constructs, the construct comprises one or more of the following terminal structures: a hairpin, a loop, an inverted terminal repeat (ITR), or a toroid. Optionally, the nucleic acid construct comprising the coding sequence and the nuclease target sequence is flanked by ITRs.
[0030] In some such constructs, the nuclease agent is a Cas protein and a guide RNA, and the nuclease target sequence is a guide RNA target sequence. Optionally, the guide RNA target sequence is a reverse guide RNA target sequence. Optionally, the Cas protein is Cas9.
[0031] In some such constructs, the retinoschizin protein or fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 5. In some such constructs, the coding sequence of the retinoschizin protein or fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 8 or 9. In some such constructs, the nucleic acid construct comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO:45.
[0032] In some such constructs, the nucleic acid construct is the nucleic acid construct for homology-independent targeted integration into the target genomic locus, the retinoschizine protein or fragment thereof is a human retinoschizine protein or fragment thereof, the coding sequence of the retinoschizine protein or fragment thereof comprises a complementary DNA (cDNA) comprising exons 2-6 of human RS1 or a degenerate variant thereof, the nucleic acid construct does not comprise a promoter driving expression of the retinoschizine protein or fragment thereof, the nucleic acid construct comprises a polyadenylation signal sequence located at the 3' end of the coding sequence, the nucleic acid construct comprises a splice acceptor site located at the 5' end of the coding sequence, and the nuclease target sequence in the nucleic acid construct is identical to the nuclease target sequence for integration into the target genomic locus, wherein the nuclease target sequence in the target genomic locus is destroyed when the nucleic acid construct is inserted in the correct orientation, but is rearranged when the nucleic acid construct is inserted into the target genomic locus in the reverse orientation.
[0033] In some such constructs, the nucleic acid construct is the nucleic acid construct for homologous recombination with the target genomic locus, the retinoschizine protein or fragment thereof is a human retinoschizine protein or fragment thereof, the coding sequence of the retinoschizine protein or fragment thereof comprises a complementary DNA (cDNA) comprising exons 2-6 of human RS1 or a degenerate variant thereof, the nucleic acid construct does not comprise a promoter driving expression of the retinoschizine protein or fragment thereof, the nucleic acid construct comprises a polyadenylation signal sequence located at the 3' end of the coding sequence, the nucleic acid construct comprises a splice acceptor site located at the 5' end of the coding sequence, and the length of each homology arm is between about 25 nucleotides and about 2.5 kb.
[0034] On the other hand, a vector comprising any of the above-mentioned nucleic acid constructs for homology-independent targeted integration is provided. Some such vectors are viral vectors. Some such vectors are adeno-associated virus (AAV) vectors. Optionally, the AAV comprises a single-stranded genome (ssAAV). Optionally, the AAV comprises a self-complementary genome (scAAV). Optionally, the AAV is selected from the group consisting of: AAV2, AAV5, AAV8 or AAV7m8.
[0035] In some such vectors, the vector does not include a promoter that drives expression of the retinoschizine protein or fragment thereof. In some such vectors, the vector does not include homology arms.
[0036] In another aspect, lipid nanoparticles comprising any of the above-described nucleic acid constructs for homology-independent targeted integration are provided.
[0037] In another aspect, cells comprising any of the above-described nucleic acid constructs for homology-independent targeted integration are provided. Some such cells are in vitro. Some such cells are in vivo. Some such cells are mammalian cells. Some such cells are human cells. Some such cells are retinal cells.
[0038] In some such cells, the cells express the retinoschizin protein or a fragment thereof. In some such cells, the nucleic acid construct is integrated at the target genomic locus on the genome. Optionally, the target genomic locus is an endogenous RS1 locus. Optionally, the nucleic acid construct is integrated in intron 1 of the endogenous RS1 locus on the genome. Optionally, endogenous RS1 exon 1 is spliced into the coding sequence of the retinoschizin protein or a fragment thereof in the nucleic acid construct. Optionally, the modified RS1 locus encoding of the nucleic acid construct integrated on the genome includes a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 2 or 4, is essentially composed of, or is composed of a protein.
[0039] In some such cells, the nucleic acid construct is integrated into the endogenous RS1 locus to prevent the endogenous RS1 gene transcription downstream of the integration site. Optionally, the nucleic acid construct is integrated into the endogenous RS1 locus to reduce or eliminate the expression of endogenous retinoschizin protein, and the expression of the retinoschizin protein or its fragment encoded by the nucleic acid construct is substituted for the expression of the endogenous retinoschizin protein. Optionally, the endogenous RS1 locus includes the RS1 gene of mutation, the RS1 gene of mutation includes the mutation that causes X-linked juvenile retinoschizin, and the expression of the nucleic acid construct integrated on the genome reduces or eliminates the expression of the RS1 gene of mutation.
[0040] In another aspect, compositions for expressing retinoschizine in a cell or for integrating a coding sequence for a retinoschizine protein or a fragment thereof into a target genomic locus in a cell are provided. Some such compositions comprise: (a) a nucleic acid construct comprising the coding sequence for the retinoschizine protein or a fragment thereof for integration into the target genomic locus; and (b) a nuclease agent or a nucleic acid encoding the nuclease agent, wherein the nuclease agent targets a nuclease target sequence in the target genomic locus.
[0041] Some such compositions include: (a) any of the above-mentioned nucleic acid constructs for homology-independent targeted integration; and (b) a nuclease agent or a nucleic acid encoding the nuclease agent, wherein the nuclease agent targets a nuclease target sequence in the target genomic locus. Optionally, the nuclease target sequence in the target genomic locus is identical to the nuclease target sequence in the nucleic acid construct. Optionally, the nuclease target sequence in the target genomic locus is destroyed when the nucleic acid construct is inserted in the correct orientation, but is reorganized when the nucleic acid construct is inserted into the target genomic locus in the opposite orientation.
[0042] Some such compositions include: (a) any of the above-described bidirectional nucleic acid constructs; and (b) a nuclease agent or a nucleic acid encoding the nuclease agent, wherein the nuclease agent targets a nuclease target sequence in the target genomic locus.
[0043] In some such compositions, the target genome locus is in the RS1 gene. Optionally, the nuclease target sequence in the target genome locus is in the first intron in the RS1 gene. Optionally, the nucleic acid construct is integrated into the endogenous RS1 locus to prevent the endogenous RS1 gene transcription downstream of the integration site. Optionally, the nucleic acid construct is integrated into the endogenous RS1 locus in the cell to reduce or eliminate the expression of endogenous retinoschizin protein, and the expression of the endogenous retinoschizin protein or its fragment encoded by the nucleic acid construct is substituted for the expression of the endogenous retinoschizin protein.
[0044] In some such compositions, the nuclease agent is a Cas protein and a guide RNA, and the nuclease target sequence is a guide RNA target sequence. Optionally, the Cas protein is Cas9. Optionally, the composition includes the guide RNA and the messenger RNA encoding the Cas protein. Optionally, the guide RNA and the messenger RNA encoding the Cas protein are in lipid nanoparticles. Optionally, the composition includes DNA encoding the Cas protein and DNA encoding the guide RNA. Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in one or more viral vectors. Optionally, the one or more viral vectors are adeno-associated virus (AAV) viral vectors. Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in a single viral vector (e.g., a single AAV vector). Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in separate viral vectors (e.g., separate AAV vectors).
[0045] In some such compositions, the nucleic acid construct is in a viral vector. Optionally, the viral vector is an adeno-associated virus (AAV) viral vector. Optionally, the AAV is selected from the group consisting of: AAV2, AAV5, AAV8, or AAV7m8.
[0046] Also provided are compositions comprising a guide RNA or a DNA encoding the guide RNA, wherein the guide RNA comprises a DNA targeting segment that targets a guide RNA target sequence in the RS1 gene, and wherein the guide RNA binds to a Cas protein and targets the Cas protein to the guide RNA target sequence in the RS1 gene.
[0047] In some such compositions or compositions used, the composition further includes the Cas protein or a nucleic acid encoding the Cas protein. Optionally, the Cas protein is a Cas9 protein. Optionally, the Cas protein is derived from Streptococcus pyogenes Cas9 protein. In some such compositions or compositions used, the composition includes a Cas protein in the form of a protein.
[0048] In some such compositions or compositions used, the composition includes the nucleic acid encoding the Cas protein, wherein the nucleic acid includes the DNA encoding the Cas protein, optionally wherein the composition includes the DNA encoding the guide RNA. Optionally, the composition includes the nucleic acid encoding the Cas protein, wherein the nucleic acid includes the DNA encoding the Cas protein, wherein the composition includes the DNA encoding the guide RNA, and wherein the DNA encoding the Cas protein and the DNA encoding the guide RNA are in one or more viral vectors. Optionally, the one or more viral vectors are adeno-associated virus (AAV) viral vectors. Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in a single viral vector (e.g., a single AAV vector). Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in separate viral vectors (e.g., separate AAV vectors). Optionally, the AAV is selected from the group consisting of: AAV2, AAV5, AAV8 or AAV7m8.
[0049] In some such compositions or compositions for use, the composition includes the nucleic acid encoding the Cas protein, wherein the nucleic acid includes a messenger RNA encoding the Cas protein, optionally wherein the composition includes the guide RNA in the form of RNA. In some such compositions or compositions for use, the composition includes the nucleic acid encoding the Cas protein, wherein the nucleic acid includes a messenger RNA encoding the Cas protein, wherein the composition includes the guide RNA in the form of RNA, and wherein the guide RNA and the messenger RNA encoding the Cas protein are in lipid nanoparticles.
[0050] In some such compositions or compositions used, the messenger RNA encoding the Cas protein includes at least one modification. Optionally, the messenger RNA encoding the Cas protein is modified to include a modified uridine at one or more or all uridine positions. Optionally, the modified uridine is a pseudouridine. Optionally, the messenger RNA encoding the Cas protein is completely replaced by pseudouridine. In some such compositions or compositions used, the messenger RNA encoding the Cas protein includes a 5' cap. In some such compositions or compositions used, the messenger RNA encoding the Cas protein includes a poly (A) tail. In some such compositions or compositions used, the messenger RNA encoding the Cas protein includes the sequence shown in SEQ ID NO:6243 or 6245.
[0051] In some such compositions or compositions used, the nucleic acid encoding the Cas protein is codon-optimized for expression in mammalian cells or human cells. In some such compositions or compositions used, the Cas protein comprises the sequence shown in SEQ ID NO: 27, 6242 or 6246.
[0052] In some such compositions or compositions for use, the guide RNA target sequence is in an intron of the RS1 gene. Optionally, the intron is the first intron of the RS1 gene.
[0053] In some such compositions or compositions for use, the RS1 gene is a human RS1 gene.
[0054] In some such compositions or compositions for use, the DNA targeting segment includes: (a) at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3148-6241; (b) at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3148-4989; (c) at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351; (d) at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304; or (e) at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: At least 17, at least 18, at least 19 or at least 20 consecutive nucleotides of the sequence set forth in any one of NOs: 4990-6241.
[0055] In some such compositions or compositions for use, the DNA targeting segment: (a) is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3148-6241; (b) is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3148-4989; (c) is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: (d) at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247 and 3249-4351; (d) at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297 and 4304; or (e) at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297 and 4304 The sequences set forth in any one of NOs: 4990-6241 are at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical.
[0056] In some such compositions or compositions for use, the DNA-targeting segment comprises, consists essentially of, or consists of the sequence shown in: (a) any one of SEQ ID NOs: 3148-6241; (b) any one of SEQ ID NOs: 3148-4989; (c) any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351; (d) any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304; or (e) any one of SEQ ID NOs: 4990-6241.
[0057] In some such compositions or compositions for use, the DNA-targeting segment includes at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304.
[0058] In some such compositions or compositions for use, the DNA targeting segment is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304.
[0059] In some such compositions or compositions for use, the DNA targeting segment comprises, consists essentially of, or consists of the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304.
[0060] In some such compositions or compositions for use, the composition includes the guide RNA in the form of RNA. In some such compositions or compositions for use, the composition includes the DNA encoding the guide RNA.
[0061] In some such compositions or compositions used, the guide RNA includes at least one modification. In some such compositions or compositions used, the at least one modification includes 2'-O-methyl modified nucleotides. In some such compositions or compositions used, the at least one modification includes phosphorothioate bonds between nucleotides. In some such compositions or compositions used, the at least one modification includes modification at one or more of the first five nucleotides at the 5' end of the guide RNA. In some such compositions or compositions used, the at least one modification includes modification at one or more of the last five nucleotides at the 3' end of the guide RNA. In some such compositions or compositions used, the at least one modification includes phosphorothioate bonds between the first four nucleotides at the 5' end of the guide RNA. In some such compositions or compositions used, the at least one modification includes phosphorothioate bonds between the last four nucleotides at the 3' end of the guide RNA. In some such compositions or compositions used, the at least one modification includes 2'-O-methyl modified nucleotides at the first three nucleotides at the 5' end of the guide RNA. In some such compositions or compositions used, the at least one modification includes 2'-O-methyl modified nucleotides at the last three nucleotides at the 3' end of the guide RNA. In some such compositions or compositions used, the at least one modification includes: (i) phosphorothioate bonds between the first four nucleotides at the 5' end of the guide RNA; (ii) phosphorothioate bonds between the last four nucleotides at the 3' end of the guide RNA; (iii) 2'-O-methyl modified nucleotides at the first three nucleotides at the 5' end of the guide RNA; and (iv) 2'-O-methyl modified nucleotides at the last three nucleotides at the 3' end of the guide RNA. In some such compositions or compositions used, the guide RNA includes modified nucleotides of SEQ ID NO: 44.
[0062] In some such compositions or compositions used, the guide RNA is a single guide RNA (sgRNA). Optionally, the guide RNA comprises, consists essentially of, or consists of the sequence shown in any one of SEQ ID NOs: 33-39 and 53. In some such compositions or compositions used, the guide RNA is a dual guide RNA (dgRNA) comprising two separate RNA molecules, the RNA molecules comprising CRISPR RNA (crRNA) and trans-activating crRNA (tracrRNA). Optionally, the crRNA comprises the sequence shown in any one of SEQ ID NOs: 29 and 52. Optionally, the tracrRNA comprises the sequence shown in any one of SEQ ID NOs: 30-32.
[0063] In some such compositions or compositions used, the composition is associated with lipid nanoparticles, optionally wherein the composition includes the guide RNA. In some such compositions or compositions used, the DNA encoding the guide RNA is in a viral vector. In some such compositions or compositions used, the viral vector is an adeno-associated virus (AAV) viral vector. Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in a single viral vector (e.g., a single AAV vector). Optionally, the DNA encoding the Cas protein and the DNA encoding the guide RNA are in separate viral vectors (e.g., separate AAV vectors). Optionally, the AAV is selected from the group consisting of: AAV2, AAV5, AAV8 or AAV7m8.
[0064] In some such compositions or compositions for use, the composition is a pharmaceutical composition that includes a pharmaceutically acceptable carrier.
[0065] In some such compositions or compositions for use, the composition further includes a second guide RNA or DNA encoding the second guide RNA, wherein the second guide RNA includes a DNA targeting segment that targets a second guide RNA target sequence in the RS1 gene, and wherein the second guide RNA binds to the Cas protein and targets the Cas protein to the second guide RNA target sequence in the RS1 gene.
[0066] Also provided are cells comprising any of the above compositions or compositions for use. Optionally, the cells are in vitro. Optionally, the cells are in vivo. Some such cells are mammalian cells. Optionally, the cells are human cells. Optionally, the cells are retinal cells.
[0067] In some such cells, the cells express the first retinoschizin protein or its fragment or the second retinoschizin protein or its fragment. In some such cells, the nucleic acid construct is integrated into the target genomic locus on the genome. Optionally, the target genomic locus is an endogenous RS1 locus. Optionally, the nucleic acid construct is integrated into intron 1 of the endogenous RS1 locus on the genome. Optionally, endogenous RS1 exon 1 is spliced into the first coding sequence or the second coding sequence of the nucleic acid construct. Optionally, the modified RS1 locus encoding of the nucleic acid construct included in the genome includes a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 2 or 4, is essentially composed of or consists of a protein.
[0068] In some such cells, the nucleic acid construct is integrated into the endogenous RS1 locus to prevent the endogenous RS1 gene transcription downstream of the integration site. Optionally, the nucleic acid construct is integrated into the endogenous RS1 locus to reduce or eliminate the expression of endogenous retinoschizin protein, and the expression of the retinoschizin protein or its fragment encoded by the nucleic acid construct is substituted for the expression of the endogenous retinoschizin protein. Optionally, the endogenous RS1 locus includes the RS1 gene of mutation, the RS1 gene of mutation includes the mutation that causes X-linked juvenile retinoschizin, and the expression of the nucleic acid construct integrated on the genome reduces or eliminates the expression of the RS1 gene of mutation.
[0069] In another aspect, methods are provided for integrating the coding sequence of a retinoschizine protein or a fragment thereof into a target genomic locus and expressing the retinoschizine protein or fragment thereof in a cell. Some such methods comprise administering any of the above-described nucleic acid constructs, vectors, lipid nanoparticles, or compositions to the cell, wherein the coding sequence is integrated into the target genomic locus and the retinoschizine protein or fragment thereof is expressed in the cell. Optionally, the cell is a mammalian cell. Optionally, the cell is a human cell. Optionally, the cell is a retinal cell. Optionally, the cell is in vitro. Optionally, the cell is in vivo. Optionally, the cell is a retinal cell in vivo, and the administration comprises subretinal injection or intravitreal injection.
[0070] In some such methods, the nucleic acid construct and the nuclease agent or the nucleic acid encoding the nuclease agent are administered simultaneously. In some such methods, the nucleic acid construct and the nuclease agent or the nucleic acid encoding the nuclease agent are administered sequentially in any order. Optionally, the nucleic acid construct is administered before the nuclease agent or the nucleic acid encoding the nuclease agent. Optionally, the nucleic acid construct is administered after the nuclease agent or the nucleic acid encoding the nuclease agent. Optionally, the time between consecutive administrations is from about 2 hours to about 48 hours. In some such methods, the nucleic acid construct and the nuclease agent or the nucleic acid encoding the nuclease agent are administered in the same delivery vehicle. In some such methods, the nucleic acid construct and the nuclease agent or the nucleic acid encoding the nuclease agent are administered in different delivery vehicles.
[0071] In some such compositions, the target genome locus is in endogenous RS1 gene. Optionally, the nuclease target sequence in the target genome locus is in the first intron in the endogenous RS1 gene. Optionally, the nucleic acid construct is integrated in the intron 1 of the endogenous RS1 locus on the genome, and wherein endogenous RS1 exon 1 is spliced into the retinoschizont protein in the nucleic acid construct or the coding sequence of its fragment. Optionally, the modified RS1 locus encoding included in the nucleic acid construct integrated on the genome includes and SEQ ID NO:2 or 4 at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical sequence, consisting essentially of it or consisting of a protein. In some such methods, the nucleic acid construct is integrated into the endogenous RS1 locus to prevent the transcription of the endogenous RS1 gene downstream of the integration site. Optionally, the nucleic acid construct is integrated into the endogenous RS1 locus to reduce or eliminate the expression of the endogenous retinoschizin protein from the endogenous RS1 locus, and the expression of the endogenous retinoschizin protein or its fragment encoded by the nucleic acid construct is substituted.
[0072] On the other hand, a method for treating an experimenter suffering from X-linked juvenile retinoschisis is provided. Some such methods can include administering any of the above-mentioned nucleic acid constructs, carriers, lipid nanoparticles or compositions to the experimenter, wherein the nucleic acid construct is integrated into the target genome locus in one or more retinal cells of the experimenter and expressed by the target genome locus, and wherein the expression of the retinoschisis inhibitor for the treatment of an effective level is achieved in the experimenter. In some such methods, the experimenter is a person. In some such methods, the experimenter has an endogenous RS1 gene, which includes at least one mutation associated with or causing X-linked juvenile retinoschisis. Optionally, the mutation is an R141C mutation. In some such methods, the administration includes subretinal injection or intravitreal injection. In some such methods, the integration of the nucleic acid construct causes retinal structure recovery.
[0073] In some such methods, the nucleic acid construct and the nuclease agent or the nucleic acid encoding the nuclease agent are administered simultaneously. In some such methods, the nucleic acid construct and the nuclease agent or the nucleic acid encoding the nuclease agent are administered sequentially in any order. Optionally, the nucleic acid construct is administered before the nuclease agent or the nucleic acid encoding the nuclease agent. Optionally, the nucleic acid construct is administered after the nuclease agent or the nucleic acid encoding the nuclease agent. Optionally, the time between consecutive administrations is from about 2 hours to about 48 hours. In some such methods, the nucleic acid construct and the nuclease agent or the nucleic acid encoding the nuclease agent are administered in the same delivery vehicle. In some such methods, the nucleic acid construct and the nuclease agent or the nucleic acid encoding the nuclease agent are administered in different delivery vehicles.
[0074] In some such compositions, the target genome locus is in endogenous RS1 gene. Optionally, the nuclease target sequence in the target genome locus is in the first intron in the endogenous RS1 gene. Optionally, the nucleic acid construct is integrated in the intron 1 of the endogenous RS1 locus on the genome, and wherein endogenous RS1 exon 1 is spliced into the retinoschizont protein in the nucleic acid construct or the coding sequence of its fragment. Optionally, the modified RS1 locus encoding included in the nucleic acid construct integrated on the genome includes and SEQ ID NO:2 or 4 at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical sequence, consisting essentially of it or consisting of a protein. In some such methods, the nucleic acid construct is integrated into the endogenous RS1 locus to prevent the transcription of the endogenous RS1 gene downstream of the integration site. Optionally, the nucleic acid construct is integrated into the endogenous RS1 locus to reduce or eliminate the expression of the endogenous retinoschizin protein from the endogenous RS1 locus, and the expression of the endogenous retinoschizin protein or its fragment encoded by the nucleic acid construct is substituted.
[0075] On the other hand, a method for modifying the RS1 gene in a cell is provided. Some such methods include administering any of the above-mentioned compositions to the cell, the composition including the guide RNA or the DNA encoding the guide RNA and the Cas protein or the nucleic acid encoding the Cas protein, wherein the guide RNA is combined with the Cas protein, and the Cas protein is targeted to the guide RNA target sequence in the RS1 gene, and the Cas protein cuts the guide RNA target sequence. In some such methods, the cell is a mammalian cell. Optionally, the cell is a human cell. Optionally, the cell is a retinal cell. Optionally, the cell is in vitro. Optionally, the cell is in vivo. In some such methods, the cell is a retinal cell, and the administration includes subretinal injection or intravitreal injection. In some such methods, the guide RNA target sequence is in the first intron in the RS1 gene. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 (Not to scale) Shown is a schematic diagram of the murine Rs1 locus including the location of the R141C mutation associated with X-linked juvenile retinoschisis (XLRS) and the insertion site for a nucleic acid construct encompassing exons 2-6 of human RS1.
[0077] Figure 2 Shown is an alignment of mouse retinoschizin, human retinoschizin, human retinoschizin with the R141C mutation, mouse retinoschizin with the R141C mutation, and a mouse / human retinoschizin hybrid expressed when a nucleic acid construct including exons 2-6 of human RS1 is integrated into intron 1 of the mouse Rs1 locus.
[0078] Figure 3A (Not to scale) A schematic diagram of a bidirectional nucleic acid construct is shown, comprising a first segment comprising a splice acceptor (A), exons 2-6 of human RS1, and bovine growth hormone (bGH) polyA, and a second segment comprising the reverse complement of SV40 polyA, the reverse complement of exons 2-6 of human RS1, and the reverse complement of the splice acceptor (A). The bidirectional construct also includes a U6 promoter operably linked to a sequence encoding a guide RNA targeting intron 1 of the murine Rs1 locus between two human RS1 segments. The horizontal arrows flanking the stars represent next-generation targeted resequencing amplicons designed for NGS. The bidirectional ssAAV structure is shown at the top, and the bidirectional scAAV structure is shown at the bottom.
[0079] Figure 3B (Not to scale) Schematic diagram of a homology-independent targeted integration nucleic acid construct comprising a splice acceptor (A), exons 2-6 of human RS1, and a polyA sequence. The construct also includes a U6 promoter operably linked to a sequence encoding a guide RNA targeting intron 1 of the murine Rs1 locus downstream of the human RS1 segment. The horizontal arrows flanking the stars represent next-generation targeted resequencing amplicons designed for NGS.
[0080] Figure 4 Shown from Injected with RS1 viral vector version 1, RS1 viral vector version 2, or RS1 viral vector version 3 Scoring of retinal cavities shown in optical coherence tomography (OCT) scans of mouse eyes. If there were 1-4 cavities on at least one individual image, a score of 1 was assigned. If there were ≥4 cavities on at least one individual image, but the cavities were not fused, a score of 2 was assigned. If there were fused cavities on at least one individual image, a score of 3 was assigned. If there were fused cavities on at least one individual image and the retina was stretched, a score of 4 was assigned. The mean score for each treatment group was compared with a control group containing pooled untreated eyes by a nonparametric Kruskal–Wallis one-way analysis of variance followed by a post hoc Dunn's multiple comparison test.
[0081] Figure 5 Shows the self Injected with RS1 viral vector version 1 (pssAAV mhRS1-sgu), RS1 viral vector version 2 (pscAAV rs1_tandem), or RS1 viral vector version 3 (pssAAV hRs1_HITI) NGS results of mouse retinal samples from mouse eyes. Read counts for four expected sequence variants are shown: (1) WT mouse, mouse reference sequence without the R141C mutant; (2) mutant mouse, mouse reference sequence with the R141C mutant; (3) humanized transcript 1, human reference sequence; and (4) humanized transcript 2, mouse codon-optimized human reference sequence. Next-generation targeted resequencing amplicons were designed for the regions depicted in Figure 3 (horizontal arrows). mRNA from mouse retina was used to generate cDNA to serve as template for next-generation sequencing (NGS) amplification.
[0082] Figure 6A and 6B Shows the self Injected with RS1 viral vector version 1 (mhRS1-sgu), RS1 viral vector version 2 (pscAAV_rs1_tandem), or RS1 viral vector version 3 (hRs1_cDNA HITI) NGS results for mouse retinal samples from mouse eyes. For these NGS results, separate amplicons were used to amplify the Rs1 intron 1 guide RNA target sequence. Reads that matched the mouse reference sequence or contained nonhomologous end joins were quantified to assess the frequency of guide RNA cleavage without insertion.
[0083] Figures 7A-7C Shows the self injected with RS1 viral vector version 1 (pssAAV mhRS1-sgu; Figure 7A ), RS1 viral vector version 2 (pscAAV rs1_tandem; Figure 7B ) or RS1 viral vector version 3 (pssAAV hRs1_HITI; Figure 7C )of NGS results of mouse retinal samples in the eyes of mice. Read counts of four expected sequence variants are shown: (1) WT mouse, mouse reference sequence without the R141C mutant; (2) mutant mouse, mouse reference sequence with the R141C mutant; (3) humanized transcript 1, human reference sequence; and (4) humanized transcript 2, mouse codon-optimized human reference sequence. Next-generation targeted resequencing amplicons were designed for the regions depicted in Figure 3 (horizontal arrows). mRNA from mouse retina was used to generate cDNA to serve as a template for next-generation sequencing (NGS) amplification. Reads that matched the mouse reference sequence or contained non-homologous end joins were quantified to assess the frequency of guide RNA cutting without insertion.
[0084] Figure 8A and 8B Shown are cells transfected with RS1 viral vector version 1 (pssAAV mhRS1-sgu; Figure 8A ) or RS1 viral vector version 2 (pscAAV rs1_tandem; Figure 8B ) treated human retinoblastoma cells. ΔCt values are shown (the lower the number, the higher the expression). "Ho" refers to the human reference sequence, and "Mo" refers to the human reference sequence codon-optimized for mouse expression.
[0085] Figure 9A and 9B Shown is the expression of RS1 viral vector version 1 (pssAAV mhRS1-sgu; Figure 9A ) or RS1 viral vector version 2 (pscAAV rs1_tandem; Figure 9B ) treated human retinoblastoma cells. ΔCt values are shown (the lower the number, the higher the expression). "Ho" refers to the human reference sequence, and "Mo" refers to the human reference sequence codon-optimized for mouse expression.
[0086] Figure 10Shown is a schematic diagram of a nucleic acid construct for homologous recombination, comprising a splice acceptor (A), exons 2-6 of people RS1, and a polyA sequence. Construct also includes a U6 promoter, which is operably connected to a sequence encoding a guide RNA targeting intron 1 of the murine Rs1 locus located downstream of the people RS1 segment. The construct also includes upstream and downstream homology arms (HA). The horizontal arrows flanking the stars represent next-generation targeted resequencing amplicons designed for NGS.
[0087] definition
[0088] The terms "protein," "polypeptide," and "peptide," used interchangeably herein, encompass polymeric forms of amino acids of any length, including coded and non-coded amino acids, as well as chemically or biochemically modified or chemically or biochemically derivatized amino acids. These terms also encompass polymers that have been modified, such as polypeptides with modified peptide backbones. The term "domain" refers to any portion of a protein or polypeptide that has a specific function or structure.
[0089] The terms "nucleic acid" and "polynucleotide," as used interchangeably herein, include polymeric forms of nucleotides of any length, including ribonucleotides, deoxyribonucleotides, or analogs or modified versions thereof, including single-, double-, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers comprising purine bases, pyrimidine bases, or other natural, chemically modified, biochemically modified, non-natural, or derivatized nucleotide bases.
[0090] The term "genomically integrated" refers to a nucleic acid that has been introduced into a cell such that the nucleotide sequence is integrated into the genome of the cell. Any protocol can be used for stably incorporating a nucleic acid into the genome of a cell.
[0091] The term "expression vector" or "expression construct" or "expression cassette" refers to a recombinant nucleic acid containing a desired coding sequence operably linked to appropriate nucleic acid sequences necessary for expression of the operably linked coding sequence in a particular host cell or organism. The nucleic acid sequences necessary for expression in prokaryotes typically include a promoter, an operator (optional), and a ribosome binding site, among other sequences. It is well known that eukaryotic cells utilize promoters, enhancers, as well as termination and polyadenylation signals, but some elements may be deleted and others added without sacrificing the necessary expression.
[0092] The term "viral vector" refers to a recombinant nucleic acid that contains at least one element of viral origin and contains elements sufficient or allowing packaging into viral vector particles. The vector and / or particle can be used to transfer DNA, RNA or other nucleic acids into cells in vitro or in vivo. Many forms of viral vectors are known.
[0093] The term "isolated" with respect to proteins, nucleic acids and cells comprises proteins, nucleic acids and cells that are relatively purified relative to other cells or organism components that may normally be present in situ, up to and including substantially pure preparations of proteins, nucleic acids or cells. The term "isolated" can comprise proteins and nucleic acids that do not have naturally occurring counterparts, or proteins or nucleic acids that have been chemically synthesized and are therefore substantially uncontaminated by other proteins or nucleic acids. The term "isolated" can comprise proteins, nucleic acids or cells that have been separated or purified from most other cellular components or organism components (e.g., but not limited to other cellular proteins, nucleic acids or cells or extracellular components) that naturally accompany proteins, nucleic acids or cells.
[0094] The term "wild-type" encompasses an entity having structure and / or activity as found in a normal (as compared to mutated, diseased, altered, etc.) state or condition. Wild-type genes and polypeptides typically exist in multiple different forms (e.g., alleles).
[0095] The term "endogenous sequence" refers to a nucleic acid sequence that naturally exists in a cell or animal. For example, an endogenous RS1 sequence of an animal refers to a natural RS1 sequence that naturally exists at the RS1 locus of the animal.
[0096] "Exogenous" molecules or sequences comprise molecules or sequences that are not usually present in the cell in the form described. Normal existence comprises the existence of the specific developmental stage and environmental conditions about the cell. For example, exogenous molecules or sequences can comprise a mutant version of the endogenous sequence corresponding in the cell (such as a humanized version of the endogenous sequence), or can comprise a sequence corresponding to the endogenous sequence in the cell but in different forms (that is, not in the chromosome). By contrast, endogenous molecules or sequences are included in molecules or sequences that exist usually in the form described in a specific cell at a specific developmental stage under specific environmental conditions.
[0097] When used in the context of a nucleic acid or protein, the term "heterologous" indicates that the nucleic acid or protein includes at least two segments that do not naturally occur together in the same molecule. For example, when used with respect to a segment of a nucleic acid or a segment of a protein, the term "heterologous" indicates that the nucleic acid or protein includes two or more subsequences that are not found in the same relationship (e.g., linked together) to each other in nature. As an example, a "heterologous" region of a nucleic acid vector is a nucleic acid segment that is not found in nature within or attached to another nucleic acid molecule associated with other molecules. For example, a heterologous region of a nucleic acid vector can include a coding sequence flanked by sequences that are not found in nature associated with a coding sequence. Similarly, a "heterologous" region of a protein is a segment of amino acids that is not found in nature within or attached to another peptide molecule (e.g., a fusion protein or a protein with a tag) associated with other peptide molecules. Similarly, a nucleic acid or protein can include a heterologous marker or a heterologous secretion or localization sequence.
[0098] "Codon optimization" utilizes the degeneracy of codons, as shown by the diversity of the three base pair codon combinations of the specified amino acids, and generally comprises the process of modifying the nucleic acid sequence to enhance expression in a specific host cell by replacing at least one codon of the native sequence with a codon more frequently or most frequently used in the genes of the host cell while maintaining the native amino acid sequence. For example, compared with naturally occurring nucleic acid sequences, the nucleic acid encoding Cas9 protein can be modified to replace codons with a higher frequency of use in a given prokaryotic or eukaryotic cell, including bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, hamster cells or any other host cells. Codon usage tables are easily available, for example, in "codon usage databases". These tables can be adjusted in a variety of ways. See Nakamura et al. (2000), " Nucleic Acids Research (Nucleic Acids Research) " 28:292, which is incorporated herein by reference in its entirety for all purposes. Computer algorithms are also available for codon optimization of a specific sequence for expression in a specific host (see, eg, Gene Forge).
[0099] The term " locus " refers to the specific location of the position on the chromosome of the genome of a gene (or significant sequence), a DNA sequence dna, a polypeptide encoding sequence or an organism. For example, the RS1 locus can refer to the specific location of the RS1 gene, the RS1 DNA sequence, the retinoschizine encoding sequence or the positioning of the RS1 on the chromosome of the genome of an organism, and the position has been identified as the position where this type of sequence resides." RS1 locus " can include the regulatory element of the RS1 gene, comprises for example enhancer, promoter, 5 ' and / or 3 ' untranslated region (UTR) or its combination.
[0100] The term "gene" refers to a DNA sequence in a chromosome that, if naturally occurring, may contain at least one coding region and at least one non-coding region. The DNA sequence encoding a product (such as, but not limited to, an RNA product and / or a polypeptide product) in the chromosome may include a coding region interrupted by non-coding introns and a sequence (including 5' and 3' non-translated sequences) positioned adjacent to the coding region at both the 5' and 3' ends such that the gene corresponds to a full-length mRNA. In addition, other non-coding sequences, including regulatory sequences (such as, but not limited to, promoters, enhancers, and transcription factor binding sites), polyadenylation signals, internal ribosome entry sites, silencers, insulating sequences, and matrix attachment regions may be present in a gene. These sequences may be close to the coding region of a gene (such as, but not limited to, within 10 kb) or located at a distant site, and these sequences may affect the transcription and translation levels or rates of the gene.
[0101] The term "allele" refers to a variant form of a gene. Some genes have multiple different forms that are located at the same position or genetic locus on a chromosome. Diploid organisms have two alleles at each locus. Each pair of alleles represents the genotype of a specific locus. If there are two identical alleles at a particular locus, the genotype is described as homozygous, and if the two alleles are different, the genotype is described as heterozygous.
[0102] " Promoter " is the regulatory region of DNA, which generally includes a TATA box that can guide RNA polymerase II to initiate RNA synthesis at the appropriate transcription start site of a specific polynucleotide sequence. Promoters may additionally include other regions that affect transcription initiation rate. Promoter sequences disclosed herein regulate the transcription of operably connected polynucleotides. Promoters can be active in one or more cell types in cell types disclosed herein (e.g., eukaryotic cells, non-human mammalian cells, human cells, rodent cells, pluripotent cells, single-cell embryos, differentiated cells, or a combination thereof). Promoters can be, for example, constitutively active promoters, conditional promoters, inducible promoters, time-limited promoters (e.g., promoters regulated by development) or spatially limited promoters (e.g., cell-specific or tissue-specific promoters). Examples of promoters can be found, for example, in WO 2013 / 176772, which is incorporated herein by reference in its entirety for all purposes.
[0103] A constitutive promoter is a promoter that is active in all tissues or specific tissues at all developmental stages. Examples of constitutive promoters include human cytomegalovirus immediate early (hCMV) promoter, mouse cytomegalovirus immediate early (mCMV) promoter, human elongation factor 1 alpha (hEF1α) promoter, mouse elongation factor 1 alpha (mEF1α) promoter, mouse phosphoglycerate kinase (PGK) promoter, chicken β-actin hybrid (CAG or CBh) promoter, SV40 early promoter, and β2 tubulin promoter.
[0104] Examples of inducible promoters include, for example, chemically regulated promoters and physically regulated promoters. Chemically regulated promoters include, for example, alcohol-regulated promoters (e.g., alcohol dehydrogenase (alcA) gene promoters), tetracycline-regulated promoters (e.g., tetracycline-responsive promoters, tetracycline operator sequences (tetO), tet-On promoters, or tet-Off promoters), steroid-regulated promoters (e.g., rat glucocorticoid receptor, estrogen receptor promoter, or ecdysone receptor promoter), or metal-regulated promoters (e.g., metalloprotein promoters). Physically regulated promoters include, for example, temperature-regulated promoters (e.g., heat shock promoters), and light-regulated promoters (e.g., light-inducible promoters or light-inhibitory promoters).
[0105] The tissue-specific promoter can be, for example, a neuron-specific promoter, a glial-specific promoter, a muscle cell-specific promoter, a cardiac cell-specific promoter, a kidney cell-specific promoter, a bone cell-specific promoter, an endothelial cell-specific promoter, or an immune cell-specific promoter (e.g., a B cell promoter or a T cell promoter).
[0106] Developmentally regulated promoters include, for example, promoters that are active only during embryonic development or only in adult cells.
[0107] "Operably linked" or "operably connected" encompasses the juxtaposition of two or more components (e.g., a promoter and another sequence element) such that both components function normally and at least one component is capable of mediating a function imposed on at least one other component. For example, a promoter may be operably linked to a coding sequence if it controls the level of transcription of the coding sequence in response to the presence or absence of one or more transcriptional regulatory factors. Operable linkage can include these sequences being adjacent to each other or acting in trans (e.g., regulatory sequences can act at a distance to control the transcription of a coding sequence).
[0108] The "complementarity" of nucleic acids means that due to the orientation of their core base groups, the nucleotide sequence in one nucleic acid chain forms hydrogen bonds with the sequence on another relative nucleic acid chain. The complementary bases in DNA are typically A and T and C and G. In RNA, they are typically C and G and U and A. Complementarity can be perfect or substantial / sufficient. Perfect complementarity between two nucleic acids means that the two nucleic acids can form a duplex, wherein each base in the duplex is bonded to a complementary base by Watson-Crick pairing. "Substantially" or "sufficiently" complementary means that the sequence in one chain is not completely and / or completely complementary to the sequence in the relative chain, but enough bonding occurs between the bases on the two chains to form a stable hybrid complex under a set of hybridization conditions (e.g., salt concentration and temperature). Such conditions can be predicted by using sequence and standard mathematical calculations to predict the Tm (melting temperature) of the hybrid chain, or by empirical determination of Tm using conventional methods. The temperature at which the Tm comprises the hybrid complex population formed between the two nucleic acid chains is denatured 50% (i.e., the double-stranded nucleic acid molecule population is half-dissociated into single strands). At temperatures below the Tm, the formation of hybridization complexes is favored, while at temperatures above the Tm, melting or separation of chains in the hybridization complex is favored. The Tm of a nucleic acid with a known G+C content in 1M NaCl aqueous solution can be estimated using, for example, Tm = 81.5 + 0.41 (% G+C), but other known Tm calculations take into account nucleic acid structural properties.
[0109] Hybridization requires that two nucleic acids contain complementary sequences, but there may be mismatches between the bases. The appropriate conditions for hybridization between two nucleic acids depend on the length and degree of complementarity of the nucleic acid, and these variables are well known. The greater the degree of complementarity between two nucleotide sequences, the greater the melting temperature (Tm) value of the nucleic acid hybrid with these sequences. For hybridization between nucleic acids with shorter complementary segments (for example, complementary on 35 or less, 30 or less, 25 or less, 22 or less, 20 or less or 18 or less nucleotides), the mismatch position becomes particularly important (see Sambrook et al., supra, 11.7-11.8). Generally, the length of hybridizable nucleic acids is at least about 10 nucleotides. The illustrative minimum length of hybridizable nucleic acids comprises at least about 15 nucleotides, at least about 20 nucleotides, at least about 22 nucleotides, at least about 25 nucleotides and at least about 30 nucleotides. In addition, the temperature and washing solution salt concentration can be adjusted as needed based on factors such as the length and degree of complementarity of the complementary region.
[0110] The polynucleotide sequence does not have to have 100% complementarity with its target nucleic acid to which it can specifically hybridize. In addition, the polynucleotide can hybridize on one or more segments so that the insertion or adjacent segment does not participate in the hybridization event (e.g., a loop structure or a hairpin structure). The polynucleotide (e.g., gRNA) can include at least 70%, at least 80%, at least 90%, at least 95%, at least 99% or 100% sequence complementarity to the target region within the target nucleic acid sequence it targets. For example, a gRNA in which 18 of the 20 nucleotides are complementary to the target region and therefore specifically hybridize will represent 90% complementarity. In this embodiment, the remaining non-complementary nucleotides can be clustered or interspersed with complementary nucleotides and do not need to be adjacent to each other or adjacent to complementary nucleotides.
[0111] The percent complementarity between specific nucleic acid sequence segments within a nucleic acid can be determined conventionally by using the BLAST program (Basic Local Alignment Search Tool) and the PowerBLAST program (Altschul et al. (1990) J. Mol. Biol. 215:403-410; Zhang and Madden (1997) Genome Res. 7:649-656) or by using the Gap program (University Research Park, Madison Wis., Genetics Computer Group, Unix version 8, Wisconsin Sequence Analysis Package) using default settings using the Smith and Waterman algorithm (Adv. Appl. Math., 1981, 2, 482-489).
[0112] The methods and compositions provided herein employ a variety of different components. Some components throughout the specification may have active variants and fragments. Such components include, for example, Cas proteins, CRISPR RNA, tracrRNA, and guide RNA. The biological activity of each of these components is described elsewhere herein. The term "functional" refers to the innate ability of a protein or nucleic acid (or its fragment or variant) to exhibit a biological activity or function. Such biological activity or function may include, for example, the ability of a Cas protein to bind to a guide RNA and a target DNA sequence. Compared to the original molecule, the biological function of a functional fragment or variant may be the same or may actually be altered (e.g., with respect to its specificity or selectivity or efficacy) but retain the basic biological function of the molecule.
[0113] The term "variant" refers to a nucleotide sequence that differs from the most prevalent sequence in a population (eg, by one nucleotide) or a protein sequence that differs from the most prevalent sequence in a population (eg, by one amino acid).
[0114] When referring to a protein, the term "fragment" means a protein that is shorter or has fewer amino acids than the full-length protein. When referring to a nucleic acid, the term "fragment" means a nucleic acid that is shorter or has fewer nucleotides than the full-length nucleic acid. When referring to a protein fragment, the fragment can be, for example, an N-terminal fragment (i.e., a portion of the C-terminus of the protein removed), a C-terminal fragment (i.e., a portion of the N-terminus of the protein removed), or an internal fragment (i.e., a portion of the internal portion of the protein removed).
[0115] In the context of two polynucleotides or polypeptide sequences, "sequence identity" or "identity" refers to the residues that are identical in the two sequences when aligned for maximum correspondence over a specified comparison window. When referring to the percentage of sequence identity of a protein, non-identical residue positions typically differ by conservative amino acid substitutions, in which an amino acid residue is substituted by another amino acid residue with similar chemical properties (e.g., charge or hydrophobicity), thereby not changing the functional properties of the molecule. When the conservative substitutions of a sequence are different, the percentage sequence identity can be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ due to such conservative substitutions are considered to have "sequence similarity" or "similarity." Methods for making such adjustments are well known. Typically, this involves counting conservative substitutions as partial mismatches rather than complete mismatches, thereby increasing the percentage sequence identity. Thus, for example, when the score for the identical amino acid is 1 and the score for the non-conservative substitution is zero, the score for the conservative substitution is between zero and 1. Scores for conservative substitutions are calculated, for example, by implementations in Project PC / GENE (Intelligenetics, Mountain View, California).
[0116] "Percentage of sequence identity" includes the value determined by comparing two optimally aligned sequences over a comparison window (the maximum number of fully matched residues), wherein the portion of the polynucleotide sequence in the comparison window may include additions or deletions (i.e., gaps) compared to the reference sequence (excluding additions or deletions) to achieve optimal alignment of the two sequences. The number of matching positions is obtained by calculating the percentage by measuring the number of positions at which the same nucleic acid base or amino acid residue appears in the two sequences, dividing the number of matching positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity. Unless otherwise specified (e.g., the shorter sequence comprises a linked heterologous sequence), the comparison window is the full length of the shorter of the two compared sequences.
[0117] Unless otherwise indicated, sequence identity / similarity values comprise values obtained using GAP version 10 using the following parameters: percent identity and percent similarity for nucleotide sequences using a GAP weight of 50 and a length weight of 3 and the nwsgapdna.cmp scoring matrix; percent identity and percent similarity for amino acid sequences using a GAP weight of 8 and a length weight of 2 and the BLOSUM62 scoring matrix; or any equivalent programs thereof. "Equivalent program" includes any sequence comparison program that produces an alignment having identical nucleotide or amino acid residue matches and the same percent sequence identity for any two sequences in question when compared to the corresponding alignment generated by GAP version 10.
[0118] The term "conservative amino acid substitution" refers to replacing an amino acid normally present in a sequence with a different amino acid of similar size, charge or polarity. Examples of conservative substitutions include replacing another non-polar residue with a non-polar (hydrophobic) residue (such as isoleucine, valine or leucine). Similarly, examples of conservative substitutions include replacing another polar residue with a polar (hydrophilic) residue, such as the polar residue between arginine and lysine, the polar residue between glutamine and asparagine, or the polar residue between glycine and serine. In addition, replacing another basic residue with a basic residue (such as lysine, arginine or histidine) or replacing another acidic residue with an acidic residue (such as aspartic acid or glutamic acid) is another example of conservative substitution. Examples of non-conservative substitutions include replacing polar (hydrophilic) residues such as cysteine, glutamine, glutamic acid or lysine with non-polar (hydrophobic) amino acid residues such as isoleucine, valine, leucine, alanine or methionine and / or replacing non-polar residues with polar residues. Typical amino acid classifications are summarized in Table 1 below.
[0119] Table 1. Amino acid classification.
[0120]
[0121] "Homologous" sequences (e.g., nucleic acid sequences) comprise sequences identical or substantially similar to known reference sequences, such that they, for example, have at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identity to known reference sequences. Homologous sequences can comprise for example orthologous sequences and paralogous sequences. For example, homologous genes typically descend from a common ancestral DNA sequence through a speciation event (orthologous genes) or a gene duplication event (paralogous genes). "Orthologous" genes comprise genes that evolved from a common ancestral gene through speciation in different species. Orthologs typically retain the same function during evolution. "Paralogous" genes comprise genes related to duplication within the genome. Paralogs can evolve new functions during evolution.
[0122] The term "in vitro" encompasses an artificial environment, and processes or reactions that occur within an artificial environment (e.g., a test tube or isolated cells or cell lines). The term "in vivo" encompasses a natural environment (e.g., a cell, an organism, or the body), and processes or reactions that occur within a natural environment. The term "ex vivo" encompasses cells that have been removed from an individual, and processes or reactions that occur within such cells.
[0123] Repair in response to double-strand breaks (DSBs) occurs primarily through two conserved DNA repair pathways: homologous recombination (HR) and non-homologous end joining (NHEJ). See Kasparek and Humphrey (2011) Seminars in Cell & Dev. Biol. 22:886-897, which is incorporated herein by reference in its entirety for all purposes. Similarly, target nucleic acid repair mediated by exogenous donor nucleic acids can include any process by which genetic information is exchanged between two polynucleotides.
[0124] The term "recombination" includes any process of genetic information exchange between the two polynucleotides, and can occur by any mechanism. Recombination can occur by homology directed repair (HDR) or homologous recombination (HR). HDR or HR include nucleic acid repair forms that may require nucleotide sequence homology, using "donor" molecule to repair "target" molecule (i.e., the molecule that experiences double-strand break) as a template, and cause genetic information to be transferred from the donor to the target. Without wishing to be bound by any particular theory, this transfer can relate to mismatch correction and / or synthesis-dependent strand annealing (synthesis-dependent strand annealing) of the heteroduplex DNA formed between the target and the donor of the fracture, wherein the donor is used to resynthesize the genetic information and / or related processes that will become a part of the target. In some cases, a part for a donor polynucleotide, a part for a copy of a donor polynucleotide, or a copy of a donor polynucleotide is integrated into the target DNA. See Wang et al. (2013) Cell 153:910-918; Mandalos et al. (2012) PLOS ONE 7:e45768:1-9; and Wang et al. (2013) Nat Biotechnol. 31:530-532, each of which is incorporated herein by reference in its entirety for all purposes.
[0125] Non-homologous end joining (NHEJ) is included in the case where homology template is not needed by directly connecting the broken ends to each other or to an exogenous sequence to repair the double-strand breaks in nucleic acids. Discontinuous sequences are connected by NHEJ and usually cause disappearance, insertion or translocation near the double-strand break site. For example, NHEJ can also cause the targeted integration of exogenous donor nucleic acids by the direct connection (that is, based on the capture of NHEJ) of the end of the broken end and the exogenous donor nucleic acid. When homology directed repair (HDR) approach is not easy to use (for example, in non-dividing cells, primary cells and execution based on homology DNA repair poor cells), the targeted integration of such NHEJ mediations can be preferably used for inserting exogenous donor nucleic acids. In addition, contrary to homology directed repair, there is no need for knowledge of the larger sequence identity region about the flanking cleavage site, which may be beneficial when attempting to target and insert into an organism with a limited genome knowledge of genomic sequence. Integration can be carried out by connecting the flat ends between the exogenous donor nucleic acid and the genomic sequence of the cutting, or by connecting sticky ends (that is, with 5' or 3' overhangs) using exogenous donor nucleic acids, which are flanked by overhangs compatible with the overhangs produced by the nuclease agent in the genomic sequence of the cutting. See, for example, US 2011 / 020722, WO 2014 / 033644, WO 2014 / 089290 and Maresca et al. (2013), Genome Research 23 (3): 539-546, each of which is incorporated herein by reference in its entirety for all purposes. If flat ends are connected, it may be necessary to excise the target and / or donor to produce the microhomology region required for fragment connection, which may produce undesirable changes in the target sequence.
[0126] A composition or method that "comprises" or "includes" one or more recited elements may include other elements not specifically recited. For example, a composition that "comprises" or "includes" a protein may contain the protein alone or in combination with other ingredients. The transition phrase "consisting essentially of" means that the scope of a claim should be interpreted to encompass the specified elements recited in the claim as well as those elements that do not materially affect the basic and novel characteristics of the claimed invention. Therefore, when used in the claims of the present invention, the term "consisting essentially of" should not be interpreted as equivalent to "comprising."
[0127] "Optional" or "optionally" means that the subsequently described event or circumstance can or cannot occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.
[0128] The specification of a numerical range includes all integers within or defining the range and all subranges defined by integers within the range.
[0129] Unless the context indicates otherwise, the term "about" encompasses values ±5 of the stated value.
[0130] The term "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ("or").
[0131] The term "or" refers to any one member of a particular list and also includes any combination of members of that list.
[0132] As used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "a protein" or "at least one protein" may include a plurality of proteins, including mixtures thereof.
[0133] Statistically significant means p≤0.05. DETAILED DESCRIPTION
[0134] I. Overview
[0135] X-linked juvenile retinoschisis (XLRS) is a juvenile macular degeneration caused by a mutation in the retinoschisis protein (RS1). The RS1 gene encodes a 24kDa protein containing a lectin domain that is secreted as a homo-oligomeric complex. Genetic mutations in RS1 result in non-functional protein or the absence of protein secretion, which causes splitting or division within the retinal layers, leading to early and progressive vision loss. More than 200 different mutations in the RS1 gene are known to cause XLRS. Forty percent of the pathogenic mutations are nonsense mutations or frameshift mutations that are expected to result in the loss of the full-length retinoschisis protein. Fifty percent of the pathogenic mutations are missense mutations that allow the production of full-length mutant proteins. Most of these are located in the lectin domain and result in misfolded proteins being retained in the ER.
[0136] Because XLRS is a recessive disease caused by loss of retinoschizine function, gene replacement therapy is a potential treatment for this disease. Furthermore, because retinoschizine is delivered as an extracellular protein, beneficial treatment is not necessarily limited to transfected cells expressing the replacement gene but can encompass a wider area due to the diffusion of the secreted protein from the site of expression.
[0137] Provided herein is a nucleic acid construct and composition allowing a retinoschizine coding sequence to be inserted into a target genome locus such as an endogenous RS1 locus and / or expressing a retinoschizine coding sequence. The nucleic acid construct and composition can be used for being incorporated into the target genome locus and / or expressing in a cell, or used in a method for treating X-linked juvenile retinoschizine. Nuclease agents (e.g., targeting endogenous RS1 locus) or nucleic acids encoding nuclease agents are also provided to promote the integration of nucleic acid constructs into the target genome locus such as an endogenous RS1 locus.
[0138] In some embodiments, the present invention provides the method for the invention of the present invention.Described nucleic acid construct is integrated into the endogenous RS1 locus such as the intron 1 of RS1 and can stop the endogenous RS1 gene transcription in integration site downstream.Described nucleic acid construct is integrated into the expression of endogenous retinoschizin protein (for example, there is the endogenous retinoschizin protein that causes XLRS sudden change), and with the expression of described retinoschizin protein or its fragment or variant (for example, not causing the sudden change retinoschizin of XLRS) encoded by described nucleic acid construct, substitute the expression of described endogenous retinoschizin protein.In one example, nucleic acid construct is integrated into the expression of endogenous retinoschizin protein that can reduce in endogenous RS1 locus.In another example, nucleic acid construct is integrated into the expression of endogenous retinoschizin protein that can eliminate in endogenous RS1 locus. In this way, integration of the nucleic acid construct can simultaneously knock out the endogenous RS1 gene (e.g., one or more endogenous RS1 genes including mutations associated with or causing XLRS, such as R141C) and knock in an alternative retinoschizine coding sequence (e.g., an alternative retinoschizine coding sequence that does not include mutations associated with or causing XLRS).
[0139] II. Nucleic Acid Constructs Comprising a Retinoschizine Coding Sequence for Integration into and Expression from a Target Genomic Locus
[0140] Provided herein are nucleic acid constructs (ie, exogenous donor nucleic acids) comprising a retinoschizine coding sequence (ie, encoding a retinoschizine protein or a fragment or variant thereof) for integration into and expression from a target genomic locus. The nucleic acid construct can be an isolated nucleic acid construct.
[0141] Retinoschizin (X-linked juvenile retinoschizin) is a protein required for the normal structure and function of the retina. An exemplary human retinoschizin protein is designated as UniProt accession number O15537 and has the sequence shown in SEQ ID NO:2. Orthologs of other species are also known. For example, an exemplary mouse retinoschizin protein is designated as UniProt accession number Q9Z1L4 and has the sequence shown in SEQ ID NO:1. Retinoschizin is encoded by the RS1 gene (also known as XLRS1). The human RS1 gene contains six independent exons spaced apart by five introns. The human RS1 gene is designated as NCBI GeneID 6247. The mouse Rs1 gene is designated as NCBI GeneID 20147. The exemplary coding sequence of human RS1 is designated as CCDS ID CCDS14187.1 and is shown in SEQ ID NO:6. Retinoschizin mutations cause X-linked juvenile retinoschisis (XLRS), a vitreoretinal dystrophy characterized by maculopathy and superficial retinal splitting. As described in more detail elsewhere herein, the nucleic acid constructs disclosed herein can be used in methods for treating XLRS.
[0142] The functional domains of RS1 are the signal peptide (SP), RS1, and the Dictyostelium lectin domain. The signal sequence directs the translocation of nascent RS1 from the endoplasmic reticulum (site of synthesis) to the outer leaflet of the plasma membrane, during which the signal sequence is cleaved by a signal peptidase to produce a mature protein with characteristic RS1 and highly conserved Dictyostelium lectin domains. The different subdomains of the RS1 signal sequence are the positively charged N region at the amino terminus that mediates translocation, the hydrophobic core (H) required for targeting and membrane insertion, and the polar "C" region that determines the site of recognition and cleavage by the signal peptidase. RS1 is primarily expressed by retinal photoreceptor cells and bipolar cells, and is also found in the pineal gland.
[0143] The retinoschizine coding sequence contained in the nucleic acid construct disclosed herein can be a coding sequence for a full-length retinoschizine protein or a fragment or variant thereof. In one example, the retinoschizine coding sequence contained in the nucleic acid construct does not include the first exon of RS1. For example, the retinoschizine coding sequence contained in the nucleic acid construct can include exons 2-6 of the RS1 gene or a variant or degenerate variant thereof. As an example, a cDNA fragment comprising exons 2-6 of the RS1 gene can include the sequence shown in SEQ ID NO: 8. Although each of the 64 codons is specific for only one amino acid or termination signal, the genetic code is degenerate (i.e., redundant) because a single amino acid can be encoded by more than one codon. Degenerate variants of a gene encode the same protein but use at least one different codon. The retinoschizine coding sequence in the nucleic acid construct can include a complementary DNA (cDNA) without an inserted intron, or the nucleic acid construct can include one or more introns separating the exons in the retinoschizine coding sequence. For example, a nucleic acid construct can include a sequence corresponding to the RS1 genomic locus having both exons and introns.
[0144] The retinoschizine coding sequence can be from any organism. For example, the retinoschizine coding sequence can be a mammal, a non-human mammal, a rodent, a mouse, a rat, or a human or a variant thereof. Alternatively, the retinoschizine coding sequence can be chimeric (e.g., part mouse and part human). In a specific example, the retinoschizine coding sequence is a human retinoschizine coding sequence.
[0145] The retinoschizine coding sequence can be codon-optimized for efficient translation into retinoschizine in a particular cell or organism. As an example, a codon-optimized version of exons 2-6 of human RS1 is shown in SEQ ID NO: 9. For example, the nucleic acid can be modified to replace codons with higher usage frequencies in human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest.
[0146] The retinoschizin coding sequence may encode a wild-type retinoschizin protein or a fragment or variant thereof. Similarly, the retinoschizin coding sequence may be a wild-type coding sequence or a variant thereof. In one example, the retinoschizin coding sequence does not include any mutation associated with or causing X-linked juvenile retinoschizin. Alternatively, the retinoschizin coding sequence may include one or more mutations associated with or causing X-linked juvenile retinoschizin (e.g., R141C).
[0147] The nucleic acid construct may further include one or more RS1 introns or fragments thereof or variants (e.g., one or more human RS1 introns or fragments thereof or variants). For example, the nucleic acid construct may include RS1 intron 1 or a fragment thereof or variant. RS1 introns or fragments thereof or variants may include a splice acceptor site or a fragment thereof. The example of a fragment of RS1 intron 1 is shown in SEQ ID NO: 15 and 16. In a specific example, the nucleic acid construct may include RS1 intron 1 or a fragment thereof or variant (e.g., upstream of a cDNA sequence comprising exon 2-6 of RS1, consisting essentially of or consisting of) located at the 5' end of exon 2-6 of RS1.
[0148] The nucleic acid construct can further include one or more splice acceptor sites. The example comprising the sequence (for example, intron sequence) of a splice acceptor site and its reverse complement is shown in SEQ ID NO:15-21. For example, the nucleic acid construct can include a splice acceptor site positioned at the 5' end of the retinoschizine coding sequence. In a specific example, the retinoschizine coding sequence includes exon 2-6 of RS1 (for example, exon 2-6 of people RS1), is essentially composed of or is composed of, and the splice acceptor site is a splice acceptor site from intron 1 of RS1 (for example, people RS1) for splicing RS1 exon 1 to RS1 exon 2. The term splice acceptor site refers to a nucleic acid sequence that can be identified and combined by a splicing mechanism at 3' intron / exon boundaries.
[0149] The nucleic acid constructs disclosed herein can also include post-transcriptional regulatory elements, such as woodchuck hepatitis virus post-transcriptional regulatory elements.
[0150] The nucleic acid construct may further include one or more polyadenylation signal sequences. Examples of polyadenylation signal sequences or sequences including polyadenylation signal sequences or their reverse complements are shown in SEQ ID NOs: 22-25. For example, a nucleic acid construct may include a polyadenylation signal sequence positioned at the 3' end of the retinoschizine coding sequence. Any suitable polyadenylation signal sequence may be used. The term polyadenylation signal sequence refers to any sequence that directs transcription termination and adds a poly-A tail to an mRNA transcript. In eukaryotes, transcription terminators are recognized by protein factors, and polyadenylation is the process of adding a poly(A) tail to an mRNA transcript in the presence of a poly(A) polymerase after termination. Mammalian poly(A) signals typically consist of a core sequence of approximately 45 nucleotides long, which may be flanked by different auxiliary sequences for enhancing cleavage and polyadenylation efficiency. The core sequence consists of the following: an upstream element (AATAAA or AAUAAA) that is highly conserved in mRNA, which is called the poly A recognition motif or poly A recognition sequence, which is recognized by the cleavage and polyadenylation specificity factor (CPSF); and a poorly defined downstream region (enriched in Us or Gs and Us) that is bound by the cleavage stimulating factor (CstF). Examples of transcription terminators that can be used include, for example, the human growth hormone (HGH) polyadenylation signal, the simian virus 40 (SV40) late polyadenylation signal, the rabbit β-globin polyadenylation signal, the bovine growth hormone (BGH) polyadenylation signal, the phosphoglycerate kinase (PGK) polyadenylation signal, the AOX1 transcription termination sequence, the CYC1 transcription termination sequence, or any transcription termination sequence known to be suitable for regulating gene expression in eukaryotic cells.
[0151] The nucleic acid construct can also include a polyadenylation signal sequence upstream of the retinoschizine coding sequence. The polyadenylation signal sequence upstream of the retinoschizine coding sequence can be flanked by a recombinase recognition site that is recognized by a site-specific recombinase. In some constructs, the recombinase recognition site also flanks a selection cassette that includes, for example, the coding sequence of a drug-resistant protein. In other constructs, the recombinase recognition site is not flanked by a selection cassette. The polyadenylation signal sequence prevents transcription and expression of the protein or RNA encoded by the coding sequence. However, after exposure to a site-specific recombinase, the polyadenylation signal sequence will be excised, and the protein or RNA can be expressed.
[0152] If the polyadenylation signal sequence is excised in a tissue-specific or developmental stage-specific manner, such configuration can achieve tissue-specific expression or developmental stage-specific expression in an animal comprising the retinoschizine coding sequence. If the animal comprising the nucleic acid construct further comprises the coding sequence of a site-specific recombinase operably linked to a tissue-specific or developmental stage-specific promoter, excision of the polyadenylation signal sequence in a tissue-specific or developmental stage-specific manner can be achieved. The polyadenylation signal sequence will then be excised only in those tissues or at those developmental stages, thereby achieving tissue-specific expression or developmental stage-specific expression. In one embodiment, the retinoschizine or a fragment or variant thereof encoded by the nucleic acid construct can be expressed in an eye-specific or retinal cell-specific manner.
[0153] Site-specific recombinase comprises the enzyme that can promote recombining between recombinase recognition site, and wherein two recombination sites are physically separated in single nucleic acid or on independent nucleic acid.The example of recombinase comprises Cre, Flp and Dre recombinase.An example of Cre recombinase gene is Crei, and wherein two exons of coding Cre recombinase are separated by intron, to prevent it from expressing in prokaryotic cell.This type of recombinase can further comprise the nuclear localization signal for promoting being positioned to core (for example, NLS-Crei).The recombinase recognition site comprises the nucleotide sequence that is recognized by site-specific recombinase and can be used as the substrate of recombination event.The example of recombinase recognition site comprises FRT, FRT11, FRT71, attp, att, rox and lox site, as loxP, lox511, lox2272, lox66, lox71, loxM2 and lox5171.
[0154] The nucleic acid construct can further include a promoter operably connected to the retinoschizine coding sequence. The retinoschizine coding sequence in the nucleic acid construct can be operably connected to any suitable promoter for expression in an animal or in vitro in an isolated cell. The promoter can be a constitutively active promoter (e.g., CAG promoter or U6 promoter), a conditional promoter, an inducible promoter, a time-limited promoter (e.g., a developmentally regulated promoter) or a spatially restricted promoter (e.g., a cell-specific or tissue-specific promoter). Such promoters are well-known and discussed elsewhere herein. The promoter that can be used in the expression construct includes, for example, a promoter with activity in one or more of eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, hamster cells, rabbit cells, pluripotent cells, eye cells, retinal cells, embryonic stem (ES) cells or fertilized eggs. In a specific example, the promoter is active in eye cells or retinal cells.
[0155] Alternatively, some nucleic acid constructs do not include the promoter operably connected with the retinoschizine coding sequence (for example, some nucleic acid constructs are promoterless constructs).This type of nucleic acid construct can be designed to, for example, when being integrated into the target genome locus, be operably connected with the endogenous promoter at the target genome locus (for example, the endogenous RS1 promoter at the endogenous RS1 locus).
[0156] Any target genome locus capable of expressing a gene can be used, such as safe harbor locus (safe harbor gene) or endogenous RS1 locus. The interaction between the exogenous DNA integrated and the host genome can limit the reliability and safety of integration, and may cause obvious phenotypic effects, which are not due to targeted gene modification but due to the unexpected effects of integration on surrounding endogenous genes. For example, the transgenic randomly inserted may be affected by position effect and silence, so that its expression is unreliable and unpredictable. Similarly, the integration of exogenous DNA into the chromosomal locus can affect the surrounding endogenous genes and chromatin, thus changing cell behavior and phenotype. The safe harbor locus includes a chromosomal locus, in which transgenic or other exogenous nucleic acid inserts can be stably and reliably expressed in all tissues of interest, without significantly changing cell behavior or phenotype (that is, without any deleterious effects on host cells). See, for example, Sadelain et al. (2012) " Cancer Nature Reviews (Nat. Rev. Cancer)" 12:51-58, which is incorporated herein by reference in its entirety for all purposes. For example, a safe harbor locus can be a locus where expression of an inserted gene sequence is not interfered with by any read-through expression from adjacent genes. For example, a safe harbor locus can comprise a chromosomal locus where exogenous DNA can integrate and function in a predictable manner without adversely affecting endogenous gene structure or expression. A safe harbor locus can comprise an extragenic region or an intragenic region, such as a locus within a gene that is non-essential, dispensable, or capable of disruption without overt phenotypic consequences.
[0157] Such safe harbor loci can provide an open chromatin configuration in all tissues and can be ubiquitously expressed during embryonic development and in adults. See, for example, Zambrowicz et al. (1997) Proc. Natl. Acad. Sci. USA 94:3789-3794, which is incorporated herein by reference in its entirety for all purposes. In addition, safe harbor loci can be efficiently targeted, and safe harbor loci can be broken without an obvious phenotype. Examples of safe harbor loci include albumin, CCR5, HPRT, AAVS1, and Rosa26. See, e.g., U.S. Patent Nos. 7,888,121; 7,972,854; 7,914,796; 7,951,925; 8,110,379; 8,409,861; 8,586,526; and U.S. Patent Publication Nos. 2003 / 0232410; 2005 / 0208489; 2005 / 0026157; 2006 / 0063231; 2 No. 008 / 0159996; No. 2010 / 00218264; No. 2012 / 0017290; No. 2011 / 0265198; No. 2013 / 0137104; No. 2013 / 0122591; No. 2013 / 0177983; No. 2013 / 0177960; and No. 2013 / 0122591, each of which is incorporated herein by reference in its entirety for all purposes.
[0158] In some embodiments, the target genome locus can be an endogenous RS1 locus, such as an endogenous RS1 locus including one or more mutations (such as, the R141C mutation in the encoded retinoschizine protein) relevant to XLRS or causing XLRS. In some cases, the endogenous RS1 gene transcription downstream of the integration site can be stopped by the nucleic acid construct being integrated into the endogenous RS1 locus. The nucleic acid construct can be integrated into the endogenous RS1 locus to reduce or eliminate the expression of endogenous retinoschizine protein, and the expression of the endogenous retinoschizine protein coded by the nucleic acid construct or its fragment or variant is substituted. In one example, the expression of endogenous retinoschizine protein can be reduced by the nucleic acid construct being integrated into the endogenous RS1 locus. In another example, the expression of endogenous retinoschizine protein can be eliminated by the nucleic acid construct being integrated into the endogenous RS1 locus. In this way, integration of the nucleic acid construct can simultaneously knock out an endogenous RS1 gene (e.g., one or more endogenous RS1 genes that include a mutation associated with or causing XLRS) and knock in an alternative retinoschizine coding sequence (e.g., an alternative retinoschizine coding sequence that does not include a mutation associated with or causing XLRS).
[0159] Nucleic acid construct can be integrated in any part of the target genome locus.For example, nucleic acid construct can be inserted in the intron or the exon of the target genome locus, or can replace one or more introns and / or the exon of the target genome locus.In a specific example, nucleic acid construct can be integrated in the intron of the target genome locus, such as the first intron (for example, RS1 intron 1) of the target genome locus.The expression cassette that is integrated in the target genome locus can be operably connected to the endogenous promoter (for example, endogenous RS1 promoter) at the target genome locus, or can be operably connected to the exogenous promoter (for example, CMV promoter) allogenic to the target genome locus.
[0160] Nucleic acid construct can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), which can be single-stranded or double-stranded, and can be linear or circular form.For example, nucleic acid construct can be single-stranded oligodeoxynucleotide (ssODN).See, for example, Yoshimi et al. (2016) " Nature Communications " 7:10431, the document is incorporated herein by reference as a whole for all purposes. Nucleic acid construct can be naked nucleic acid or can be delivered by carrier such as AAV vector. In a specific example, nucleic acid construct can be delivered by AAV and can be inserted into endogenous RS1 locus by non-homologous end joining (for example, nucleic acid construct can be a nucleic acid construct not including homology arms). If introduced in linear form, the end (for example, from exonuclease degradation) of nucleic acid construct (for example, donor sequence) can be protected by well-known methods. For example, one or more dideoxynucleotide residues can be added to the 3 ' end of linear molecule and / or self-complementary oligonucleotide can be connected to one or two ends. See, for example, Chang et al., (1987) PNAS 84:4959-4963 and Nehls et al., (1996) Science 272:886-889, each of which is incorporated herein by reference in its entirety for all purposes. Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, adding terminal amino groups and using modified internucleotide linkages, such as phosphorothioates, phosphoramidates, and O-methyl ribose or deoxyribose residues.
[0161] In one embodiment, the length of exemplary nucleic acid construct is about 50 Nucleotide to about 5kb or about 50 Nucleotide to about 3kb.Alternately, the length of nucleic acid construct can be about 1kb to about 1.5kb, about 1.5kb to about 2kb, about 2kb to about 2.5kb, about 2.5kb to about 3kb, about 3kb to about 3.5kb, about 3.5kb to about 4kb, about 4kb to about 4.5kb or about 4.5kb to about 5kb.Alternately, the length of nucleic acid construct can be for example no more than 5kb, 4.5kb, 4kb, 3.5kb, 3kb or 2.5kb.
[0162] The integration of nucleic acid construct at the target genome locus can cause the interpolation of the nucleic acid sequence being paid close attention to the target genome locus or the substitution (that is, deletion and insertion) of the nucleic acid sequence being paid close attention to at the target genome locus. Some nucleic acid constructs are designed to insert nucleic acid constructs at the target genome locus without any corresponding deletion at the target genome locus. Other nucleic acid constructs are designed to delete the nucleic acid sequence being paid close attention to at the target genome locus and replace it with nucleic acid construct.
[0163] The nucleic acid construct at the target genome locus that is deleted and / or substituted or corresponding nucleic acid can have various lengths.The exemplary nucleic acid construct at the target genome locus that is deleted and / or substituted or the length of corresponding nucleic acid are about 1 Nucleotide to about 5kb between or about 1 Nucleotide to about 3kb Nucleotide between.For example, the nucleic acid construct at the target genome locus that is deleted and / or substituted or the length of corresponding nucleic acid can be about 1 to about 100, about 100 to about 200, about 200 to about 300, about 300 to about 400, about 400 to about 500, about 500 to about 600, about 600 to about 700, about 700 to about 800, about 800 to about 900 or about 900 to about 1,000 Nucleotide between. Likewise, the length of the nucleic acid construct or corresponding nucleic acid at the target genomic locus that is deleted and / or replaced can be between about 1 kb to about 1.5 kb, about 1.5 kb to about 2 kb, about 2 kb to about 2.5 kb, about 2.5 kb to about 3 kb, about 3 kb to about 3.5 kb, about 3.5 kb to about 4 kb, about 4 kb to about 4.5 kb, about 4.5 kb to about 5 kb, or longer.
[0164] The nucleic acid construct or corresponding nucleic acid at the target genomic locus that is deleted and / or replaced can be a coding region such as an exon, a non-coding region such as an intron, an untranslated region, or a regulatory region (e.g., a promoter, enhancer, or transcription repressor binding element), or any combination thereof.
[0165] In some cases, the nucleic acid construct can include one or more of the following terminal structures: a hairpin, a loop, an inverted terminal repeat (ITR), or a toroid.For example, the nucleic acid construct can include an ITR.
[0166] Some such nucleic acid constructs can be modified after cutting the target genome locus with a nuclease agent such as Cas protein or making the target genome locus have a nick (such as but not limited to endogenous RS1 locus). Nucleic acid constructs can be designed to repair cutting or nicking locus by connection or homology directed repair mediated by non-homologous end joining (NHEJ). Optionally, repair constructs are removed or destroyed nuclease target sequences with nucleic acid constructs so that the allele targeted can not be re-targeted by the nuclease agent.
[0167] Some nucleic acid constructs include homology arms. Homology arms can be symmetrical (for example, a length of 40 nucleotides or 60 nucleotides respectively), or they can be asymmetrical (for example, a homology arm or a complementary region has a length of 36 nucleotides, and a homology arm or a complementary region has a length of 91 nucleotides). Other nucleic acid constructs do not include homology arms.
[0168] Some nucleic acid constructs disclosed herein include homology arms. The homology arms may be flanked by a retinoschizine coding sequence. For ease of reference, the homology arms are referred to herein as 5' and 3' (i.e., upstream and downstream) homology arms. This term relates to the relative position of the homology arms to the nucleic acid insert (e.g., retinoschizine coding sequence) within the nucleic acid construct. The 5' and 3' homology arms correspond to regions within the target genomic locus, which are referred to herein as "5' target sequence" and "3' target sequence," respectively.
[0169] In another embodiment, the homology arms and the target sequence share a sufficient level of sequence identity to each other, and the two regions are said to "correspond or correspond" to each other to act as substrates for homologous recombination reactions. The term "homology" includes DNA sequences that are identical or share sequence identity to corresponding sequences. The sequence identity between a given target sequence and the corresponding homology arms found in a nucleic acid construct can be any degree of sequence identity that allows homologous recombination to occur. For example, the amount of sequence identity shared by the homology arms of a nucleic acid construct (or its fragment) and the target sequence (or its fragment) can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% such that the sequence undergoes homologous recombination. In addition, the corresponding homology region between the homology arm and the corresponding target sequence can have any length sufficient to promote homologous recombination. The length of exemplary homology arms is between about 25 nucleotides and about 2.5 kb, between about 25 nucleotides and about 1.5 kb, or between about 25 and about 500 nucleotides. For example, a given homology arm (or each of the homology arms) and / or corresponding target sequence can include a corresponding homology region having a length of between about 25 to about 30, about 30 to about 40, about 40 to about 50, about 50 to about 60, about 60 to about 70, about 70 to about 80, about 80 to about 90, about 90 to about 100, about 100 to about 150, about 150 to about 200, about 200 to about 250, about 250 to about 300, about 300 to about 350, about 350 to about 400, about 400 to about 450, or about 450 to about 500 nucleotides, such that the homology arm has sufficient homology to undergo homologous recombination with the corresponding target sequence within the target nucleic acid. Alternatively, a given homology arm (or each homology arm) and / or corresponding target sequence can include a corresponding homology region having a length of about 0.5 kb to about 1 kb, about 1 kb to about 1.5 kb, about 1.5 kb to about 2 kb, or about 2 kb to about 2.5 kb. For example, each homology arm can be about 750 nucleotides in length.In another example, the length of the homology arms can each be about 150 to about 750, about 200 to about 700, about 250 to about 650, about 300 to about 600, about 350 to about 550, about 400 to about 500, about 150 to about 450, about 200 to about 450, about 250 to about 450, about 300 to about 450, about 350 to about 450, about 400 to about 450, about 450 to about 500, about 450 to about 550, about 450 to about 600, about 450 to about 650, about 450 to about 700, about 450 to about 750, or about 450 nucleotides. In another example, the lengths of the homology arms can each be about 500 to about 1300, about 550 to about 1250, about 600 to about 1200, about 650 to about 1150, about 700 to about 1100, about 750 to about 1050, about 800 to about 1000, about 850 to about 950, about 500 to about 900, about 550 to about 900, about 600 to about 900, about 650 to about 10 ... about 900, about 700 to about 900, about 750 to about 900, about 800 to about 900, about 850 to about 900, about 900 to about 950, about 900 to about 1000, about 900 to about 1050, about 900 to about 1100, about 900 to about 1150, about 900 to about 1200, about 900 to about 1250, about 900 to about 1300, or about 900 nucleotides. In another example, the homology arms can each be about 1500 to about 2100, about 1550 to about 2050, about 1600 to about 2000, about 1650 to about 1950, about 1700 to about 1900, about 1750 to about 1850, about 1500 to about 1800, about 1550 to about 1800, about 1600 to about 1800, about 1650 to about 1800, about 1700 to about 1800, about 1750 to about 1800, about 1800 to about 1850, about 1800 to about 1900, about 1800 to about 1950, about 1800 to about 2000, about 1800 to about 2050, about 1800 to about 2100, or about 1800 nucleotides. In another example, each homology arm is no more than about 450 nucleotides, no more than about 900 nucleotides, or no more than about 1800 nucleotides. In another example, each homology arm is at least about 450 nucleotides, at least about 900 nucleotides, or at least about 1800 nucleotides. The homology arms can be symmetrical (each length is about the same), or can be asymmetrical (one is longer than the other).
[0170] When the CRISPR / Cas system or other nuclease agent is used in combination with the nucleic acid constructs described herein, the 5' and 3' target sequences can be positioned sufficiently close to the nuclease cleavage site (e.g., within a position sufficiently close to the guide RNA target sequence) to promote the occurrence of homologous recombination events between the target sequence and the homology arms after a single-strand break (nick) or double-strand break at the nuclease cleavage site or the nuclease cleavage site. The term "nuclease cleavage site" includes a DNA sequence in which a nick or double-strand break is produced by a nuclease agent (e.g., a Cas9 protein complexed with a guide RNA). The target sequence corresponding to the 5' and 3' homology arms of the nucleic acid construct in the target locus is "positioned sufficiently close to" the nuclease cleavage site if such a distance is to promote the occurrence of homologous recombination events between the 5' and 3' target sequence and the homology arms after a single-strand break or double-strand break at the nuclease cleavage site. Thus, the target sequence corresponding to the 5' and / or 3' homology arm of the nucleic acid construct can be, for example, within at least 1 nucleotide of a given nuclease cleavage site, or within at least 10 nucleotides to about 1,000 nucleotides of a given nuclease cleavage site. As an example, the nuclease cleavage site can be immediately adjacent to at least one or both target sequences of the target sequence.
[0171] The spatial relationship of the target sequence corresponding to the homology arms of the nucleic acid construct and the nuclease cleavage site can vary. For example, the target sequence can be positioned 5' to the nuclease cleavage site, the target sequence can be positioned 3' to the nuclease cleavage site, or the target sequence can flank the nuclease cleavage site.
[0172] Other nucleic acid constructs do not include any homology arms. Such nucleic acid constructs can be inserted by non-homologous end joining. For example, such nucleic acid constructs can be inserted into a blunt-ended double-strand break after cutting with a nuclease agent. In a specific example, the nucleic acid construct can be delivered by AAV and can be inserted into the target genome locus by non-homologous end joining (for example, the nucleic acid construct can be a nucleic acid construct that does not include homology arms).
[0173] In a specific example, the nucleic acid construct can be inserted by homology-independent targeted integration. For example, the retinoschizine coding sequence in the nucleic acid construct can be flanked with a target site for a nuclease agent on each side (for example, the same target site as in the target genome locus, and the same nuclease agent is used to cut the target site in the target genome locus). The nuclease agent can then cut the target site of the flanking retinoschizine coding sequence. In a specific example, the nucleic acid construct is delivered by the delivery of AAV mediation, and the cutting of the target site of the flanking retinoschizine coding sequence can remove the inverted terminal repeats (ITR) of AAV. In some methods, the target site in the target genome locus (for example, the gRNA target sequence comprising the adjacent motif of the flanking protospacer) no longer exists when the retinoschizine coding sequence is inserted into the target genome locus with the correct direction, but is reorganized when the retinoschizine coding sequence is inserted into the target genome locus with the opposite direction. This can help the retinoschizine coding sequence to be inserted into the correct expression direction.
[0174] In an exemplary nucleic acid construct for homology-independent targeted integration into a target genomic locus, the retinoschizine protein or fragment thereof is a human retinoschizine protein or fragment thereof, the coding sequence of the retinoschizine protein or fragment thereof comprises a complementary DNA (cDNA) comprising exons 2-6 of human RS1 or a degenerate variant thereof, the nucleic acid construct does not comprise a promoter driving expression of the retinoschizine protein or fragment thereof, the nucleic acid construct comprises a polyadenylation signal sequence located at the 3' end of the coding sequence, the nucleic acid construct comprises a splice acceptor site located at the 5' end of the coding sequence, and the nuclease target sequence in the nucleic acid construct is the same as the nuclease target sequence used for integration into the target genomic locus, wherein the nuclease target sequence in the target genomic locus is destroyed when the nucleic acid construct is inserted in the correct orientation, but is rearranged when the nucleic acid construct is inserted into the target genomic locus in the reverse orientation.
[0175] In an exemplary nucleic acid construct for homology-independent targeted integration into a target genomic locus, the retinoschizin protein or a fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 2 or 5. In an exemplary nucleic acid construct for homology-independent targeted integration into a target genomic locus, the coding sequence of the retinoschizin protein or a fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 6, 8, or 9, or a degenerate variant thereof. In an exemplary nucleic acid construct for homology-independent targeted integration into a target genomic locus, the nucleic acid construct comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 45.
[0176] Other nucleic acid constructs can have a short single-stranded region complementary to one or more overhangs produced by the cutting mediated by the nuclease agent at the target genome locus at the 5' end and / or 3' end. For example, some nucleic acid constructs have a short single-stranded region complementary to one or more overhangs produced by the cutting mediated by the nuclease at the 5' and / or 3' target sequence of the target genome locus at the 5' end and / or 3' end. Some such nucleic acid constructs have a complementary region only at the 5' end or only at the 3' end. For example, some such nucleic acid constructs have a complementary region only at the 5' end complementary to the overhang produced at the 5' target sequence of the target genome locus or only at the 3' end complementary to the overhang produced at the 3' target sequence of the target genome locus. Other such nucleic acid constructs have a complementary region at both the 5' and 3' ends. For example, other such nucleic acid constructs have a complementary region (for example, respectively complementary to the first overhang and the second overhang) produced by the cutting mediated by the nuclease at both the 5' and 3' ends. For example, if the nucleic acid construct is double-stranded, the single-stranded complementary region can extend from the 5' end of the top strand of the nucleic acid construct and the 5' end of the bottom strand of the donor nucleic acid, thereby generating a 5' overhang at each end. Alternatively, the single-stranded complementary region can extend from the 3' end of the top strand of the nucleic acid construct and the 3' end of the bottom strand of the template, thereby generating a 3' overhang.
[0177] In some embodiments, the complementary region comprises at least one nucleotide sequence, a nucleic acid construct, a target nucleic acid, and a nucleic acid molecule. The complementary region can have any length that is enough to promote the connection between the nucleic acid construct and the target nucleic acid. The length of exemplary complementary region is between about 1 to about 5 nucleotide, between about 1 to about 25 nucleotide or between about 5 to about 150 nucleotide. For example, the length of complementary region can be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotide. Alternatively, the length of the complementary region can be about 5 to about 10, about 10 to about 20, about 20 to about 30, about 30 to about 40, about 40 to about 50, about 50 to about 60, about 60 to about 70, about 70 to about 80, about 80 to about 90, about 90 to about 100, about 100 to about 110, about 110 to about 120, about 120 to about 130, about 130 to about 140, about 140 to about 150 nucleotides, or longer.
[0178] Such complementary regions can be complementary to the overhangs produced by two pairs of nickases. By using the first nickase and the second nickase that cut opposite DNA chains to produce the first double-strand break and the third nickase and the fourth nickase that cut opposite DNA chains to produce the second double-strand break, two double-strand breaks with staggered ends can be produced. For example, Cas proteins can be used to cut the first, second, third and fourth guide RNA target sequences corresponding to the first, second, third and fourth guide RNAs. The first and second guide RNA target sequences can be positioned to produce a first cleavage site so that the nick produced by the first and second nickases on the first and second DNA chains produces a double-strand break (that is, the first cleavage site includes the nick in the first and second guide RNA target sequences). Similarly, the third and fourth guide RNA target sequences can be positioned to produce a second cleavage site so that the nick produced by the third and fourth nickases on the first and second DNA chains produces a double-strand break (that is, the second cleavage site includes the nick in the third and fourth guide RNA target sequences). The nick in the first and second guide RNA target sequences and / or the third and fourth guide RNA target sequences can be a biased nick that produces an overhang. The offset window can be, for example, at least about 5bp, 10bp, 20bp, 30bp, 40bp, 50bp, 60bp, 70bp, 80bp, 90bp, 100bp or more. See Ran et al. (2013), Cell 154: 1380-1389; Mali et al. (2013) Nature Biotechnology 31: 833-838; and Shen et al. (2014) Nat. Methods 11: 399-404, each of which is incorporated herein by reference as a whole for all purposes. In this case, the double-stranded nucleic acid construct can be designed to have a single-stranded complementary region that is complementary to the overhang produced by the nicks in the first and second guide RNA target sequences and the nicks in the third and fourth guide RNA target sequences. Such nucleic acid constructs can then be inserted by a connection mediated by non-homologous end joining.
[0179] Some of the nucleic acid constructs disclosed herein are bidirectional constructs that can be inserted into and expressed from a target genomic locus in either orientation. Such nucleic acid constructs can include a first segment comprising a first coding sequence for a first retinoschizin protein, or a fragment or variant thereof; and a second segment comprising the reverse complement of a second coding sequence for a second retinoschizin protein, or a fragment or variant thereof. For example, the second segment can be positioned 3' to the first segment in the nucleic acid construct.
[0180] The first segment and the second segment can be directly connected together, or can be connected by a linker such as a peptide linker. The peptide linker can be of any suitable length. For example, the length of the linker can be between about 5 and about 2000 nucleotides. For example, the length of the linker sequence can be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 500, 1000, 1500, 2000 or more nucleotides.
[0181] In some bidirectional constructs, the first retinoschizine protein, or fragment or variant thereof, is the same as the second retinoschizine protein, or fragment or variant thereof. In other bidirectional constructs, the first retinoschizine protein, or fragment or variant thereof, is different from the second first retinoschizine protein, or fragment or variant thereof.
[0182] In some bidirectional constructs, the codon usage in the first coding sequence is the same as the codon usage in the second coding sequence. In other bidirectional constructs, the second coding sequence uses a codon usage that is different from the codon usage of the first coding sequence to reduce hairpin formation. Such a reverse complement forms base pairs with all nucleotides of the coding sequence in the first segment, but it can optionally encode the same polypeptide.
[0183] The second segment can have any percentage of complementarity with the first segment. For example, the second segment sequence can have at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, or at least about 99% complementarity with the first segment. As another example, the second segment sequence can have less than about 30%, less than about 35%, less than about 40%, less than about 45%, less than about 50%, less than about 55%, less than about 60%, less than about 65%, less than about 70%, less than about 75%, less than about 80%, less than about 85%, less than about 90%, less than about 95%, less than about 97%, or less than about 99% complementarity with the first segment. In some nucleic acid constructs, the reverse complement of the second coding sequence can be not significantly complementary to the first coding sequence (e.g., no more than 70% complementary), not significantly complementary to a fragment of the first coding sequence, highly complementary to the first coding sequence (e.g., at least 90% complementary), highly complementary to a fragment of the first coding sequence, about 50% to about 80% identical to the reverse complement of the first coding sequence, or about 60% to about 100% identical to the reverse complement of the first coding sequence.
[0184] In some cases, bidirectional constructs can include one or more (for example, two) polyadenylation signal sequences. In some bidirectional constructs, the first section can include the first polyadenylation signal sequence. In some bidirectional constructs, the first section can include the second polyadenylation signal sequence. In some bidirectional constructs, the first section can include the first polyadenylation signal sequence, and the second section can include the second polyadenylation signal sequence (for example, the reverse complement of the polyadenylation signal sequence). In some bidirectional constructs, the first section can include the first polyadenylation signal sequence positioned at the 3 ' end of the first coding sequence. In some bidirectional constructs, the second section can include the reverse complement of the second polyadenylation signal sequence positioned at the 5 ' end of the reverse complement of the second coding sequence. In some bidirectional constructs, the first section can include the first polyadenylation signal sequence positioned at the 3 ' end of the first coding sequence, and the second section can include the reverse complement of the second polyadenylation signal sequence positioned at the 5 ' end of the reverse complement of the second coding sequence. The first polyadenylation signal sequence and the second polyadenylation signal sequence can be the same or different. In one example, the first polyadenylation signal and the second polyadenylation signal are different.
[0185] In some cases, a bidirectional construct may include one or more (e.g., two) splice acceptor sites. In some bidirectional constructs, the first segment may include the first splice acceptor site. In some bidirectional constructs, the first segment may include the second splice acceptor site. In some bidirectional constructs, the first segment may include the first splice acceptor site, and the second segment may include the second splice acceptor site (e.g., the reverse complement of the splice acceptor site). In some bidirectional constructs, the first segment includes the first splice acceptor site positioned at the 5' end of the first coding sequence. In some bidirectional constructs, the second segment includes the reverse complement of the second splice acceptor site positioned at the 3' end of the reverse complement of the second coding sequence. In some bidirectional constructs, the first segment includes the first splice acceptor site positioned at the 5' end of the first coding sequence, and the second segment includes the reverse complement of the second splice acceptor site positioned at the 3' end of the reverse complement of the second coding sequence. The first splice acceptor site and the second splice acceptor site may be identical or different. In one example, the first splice acceptor site and the second splice acceptor site are different. The first splice acceptor site and / or the second splice acceptor site can be from an RS1 gene (eg, from intron 1 of an RS1 gene), such as a human RS1 gene.
[0186] Some bidirectional constructs can include a promoter driving expression of a first retinoschizine protein, or a fragment or variant thereof, and / or the reverse complement of a promoter driving expression of a second retinoschizine protein, or a fragment or variant thereof. Alternatively, a bidirectional construct can be a construct that does not include a promoter driving expression of a first retinoschizine protein, or a fragment or variant thereof, or a second retinoschizine protein, or a fragment or variant thereof (i.e., a promoterless construct).
[0187] In some bidirectional constructs, one or both of the coding sequences can be codon optimized to express in a host cell. In some bidirectional constructs, only one of the coding sequences is codon optimized. In some bidirectional constructs, the first coding sequence is codon optimized. In some bidirectional constructs, the second coding sequence is codon optimized. In some bidirectional constructs, both coding sequences are codon optimized.
[0188] In an exemplary bidirectional construct, the second segment is located 3' to the first segment, the first retinoschizine protein or fragment thereof and the second retinoschizine protein or fragment thereof are both human retinoschizine proteins or fragments thereof, the first retinoschizine protein or fragment thereof is identical to the second retinoschizine protein or fragment thereof, the first coding sequence and the second coding sequence both comprise complementary DNA (cDNA) comprising exons 2-6 of human RS1 or a degenerate variant thereof, the second coding sequence employs a codon usage different from that of the first coding sequence, wherein the first segment comprises a protein located 3' to the first segment. The nucleic acid construct comprises a first polyadenylation signal sequence located at the 3' end of the first coding sequence, and the second segment comprises the reverse complement of the second polyadenylation signal sequence located at the 5' end of the reverse complement of the second coding sequence, the first segment comprises a first splice acceptor site located at the 5' end of the first coding sequence, and the second segment comprises the reverse complement of the second splice acceptor site located at the 3' end of the reverse complement of the second coding sequence, the nucleic acid construct does not comprise a promoter driving expression of the first retinoschizine protein or fragment thereof or the second retinoschizine protein or fragment thereof, and optionally the nucleic acid construct does not comprise a homology arm.
[0189] In an exemplary bidirectional construct, the first retinoschizin protein or fragment thereof and / or the second retinoschizin protein or fragment thereof comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 2 or 5. In an exemplary bidirectional construct, the first coding sequence and / or the second coding sequence comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% identical to SEQ ID NO: 6, 8, or 9, or a degenerate variant thereof. In an exemplary bidirectional construct, the first coding sequence comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO:8, and the second coding sequence comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO:9. In exemplary bidirectional constructs, the nucleic acid construct comprises, consists essentially of, or consists of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% identical to SEQ ID NO: 46 or 47.
[0190] The nucleic acid construct can include modifications or sequences that provide additional desired characteristics (e.g., modified or regulated stability; tracking or detection with fluorescent markers; binding sites for proteins or protein complexes, etc.). The nucleic acid construct can include one or more fluorescent markers, purification tags, epitope tags, or combinations thereof. For example, the nucleic acid construct can include one or more fluorescent markers (e.g., fluorescent proteins or other fluorophores or dyes), such as at least 1, at least 2, at least 3, at least 4, or at least 5 fluorescent markers. Exemplary fluorescent markers include fluorophores, such as fluorescein (e.g., 6-carboxyfluorescein (6-FAM)), Texas Red, HEX, Cy3, Cy5, Cy5.5, Pacific Blue, 5-(and-6)-carboxytetramethylrhodamine (TAMRA), and Cy7. A variety of fluorescent dyes are commercially available for labeling oligonucleotides (e.g., from Integrated DNA Technologies). Such fluorescent labels (e.g., internal fluorescent labels) can be used, for example, to detect nucleic acid constructs that have been directly integrated into a cleaved target nucleic acid having overhangs that are compatible with the ends of the nucleic acid construct. The label or marker can be located at the 5' end, the 3' end, or internal to the nucleic acid construct. For example, a nucleic acid construct can be attached at the 5' end to a fluorescent marker from Integrated DNA Technologies (5' 700) is conjugated with the IR700 fluorophore.
[0191] Nucleic acid construct can also include conditional alleles. Conditional alleles can be multifunctional alleles, as described in US2011 / 0104799, which is incorporated herein by reference in its entirety for all purposes. For example, conditional alleles can include: (a) a promoter sequence relative to gene transcription in a sense orientation; (b) a drug selection cassette (DSC) in a sense or antisense orientation; (c) a nucleotide sequence of interest (NSI) in an antisense orientation; and (d) a conditional reversal module (COIN, which utilizes exon splitting introns and reversible gene capture-like modules) in the opposite orientation. See, for example, US 2011 / 0104799. Conditional alleles can further include a recombinant unit, which is recombined to form a conditional allele after being exposed to the first recombinase, wherein the conditional allele (i) lacks a promoter sequence and a DSC; and (ii) contains a COIN in the NSI in the sense orientation and the antisense orientation. See, for example, US 2011 / 0104799.
[0192] The nucleic acid construct can also include the polynucleotides encoding a selection marker. Alternatively, the nucleic acid construct may lack the polynucleotides encoding a selection marker. The selection marker may be contained in a selection box. Optionally, the selection box may be a self-deletion box. See, for example, US 8,697,851 and US 2013 / 0312129, each document in the document is incorporated herein by reference in its entirety for all purposes. As an example, the self-deletion box can include a Crei gene (comprising two exons separated by introns encoding Cre recombinase) operably connected to the mouse Prm1 promoter and a neomycin resistance gene operably connected to the human ubiquitin promoter. By adopting the Prm1 promoter, the self-deletion box can be specifically deleted in the male germ cells of F0 animals. Exemplary selection markers include neomycin phosphotransferase (neo r ), hygromycin B phosphotransferase (hyg r ), puromycin-N-acetyltransferase (puromycin r ), blasticidin-S deaminase (bsr r ), xanthine / guanine phosphoribosyltransferase (gpt) or herpes simplex virus thymidine kinase (HSV-k) or a combination thereof. The polynucleotide encoding the selection marker can be operably linked to a promoter that is active in the targeted cell. Examples of promoters are described elsewhere herein.
[0193] Nucleic acid constructs can also include reporter genes.Exemplary reporter genes include genes encoding: luciferase, beta-galactosidase, green fluorescent protein (GFP), enhanced green fluorescent protein (eGFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), blue fluorescent protein (BFP), enhanced blue fluorescent protein (eBFP), DsRed, ZsGreen, MmGFP, mPlum, mCherry, tdTomato, mStrawberry, J-Red, mOrange, mKO, mCitrine, Venus, YPet, emerald, CyPet, Cerulean, T- sky blue and alkaline phosphatase. Such reporter genes can be operably linked to a promoter active in the targeted cell. The example of a promoter is described elsewhere herein.
[0194] The nucleic acid construct can also include one or more expression cassettes or deletion cassettes. A given cassette can include one or more of the nucleotide sequence of interest, the polynucleotide encoding the selected marker, and the reporter gene, as well as various regulatory components that affect expression. Examples of selectable markers and reporter genes that can be included are discussed in detail elsewhere herein.
[0195] Nucleic acid construct can comprise the nucleic acid that side joint has site-specific recombinant target sequence.Alternately, nucleic acid construct can comprise one or more site-specific recombinant target sequences.Although whole nucleic acid construct can side joint has this type of site-specific recombinant target sequence, any paid close attention to district or individual polynucleotide in nucleic acid construct also can side joint has this type of site.The site-specific recombinant target sequence of any paid close attention to polynucleotide in side joint nucleic acid construct or nucleic acid construct can comprise for example loxP, lox511, lox2272, lox66, lox71, loxM2, lox5171, FRT, FRT11, FRT71, attp, att, FRT, rox or its combination.In an example, the polynucleotide of the selective marker and / or reporter gene that comprises in the site-specific recombination site side joint encoding nucleic acid construct.After the targeting locus place integrates nucleic acid construct, the sequence between the site-specific recombination site can be removed.
[0196] Nucleic acid construct can also include the restriction site of one or more restriction endonucleases (i.e., restriction enzymes), and the restriction endonucleases comprise I type, II type, III type and IV type endonucleases. I type and III type restriction endonucleases identify specific recognition sites, but usually cut at the variable position from the nuclease binding site, and the nuclease binding site may be hundreds of base pairs away from the cleavage site (recognition site). In the II type system, restriction activity is independent of any methylase activity, and cutting usually occurs in or near the specific site of the binding site. Most of II type enzymes cut off palindromic sequences, but IIa type enzymes identify non-palindromic recognition sites and cut outside the recognition site, IIb type enzymes cut off sequence twice with two sites outside the recognition site, and IIs type enzymes identify asymmetric recognition sites and cut on one side and at a limited distance from the approximately 1-20 nucleotide of the recognition site. IV type restriction enzymes target methylated DNA. Restriction enzymes are further described and cataloged, for example, in the REBASE database (rebase.neb.com webpage; Roberts et al., (2003) Nucleic Acids Res. 31:418-420; Roberts et al., (2003) Nucleic Acids Res. 31:1805-1812; and Belfort et al. (2002) Mobile DNA II, pp. 761-783, Craigie et al., eds. (ASM Press, Washington, D.C.)).
[0197] Nucleic acid constructs disclosed herein can also include additional coding sequences. For example, some nucleic acid constructs disclosed herein can include sequences encoding guide RNAs targeting target genomic loci (e.g., targeting RS1, such as intron 1 of RS1). The sequences encoding guide RNAs can be operably connected to promoters such as U6 promoters. In some nucleic acid constructs, guide RNA expression cassettes are positioned at 3' (downstream) of the retinoschizine coding sequence. In some bidirectional nucleic acid constructs, guide RNA expression cassettes are positioned between the first section and the second section.
[0198] III. Vectors containing nucleic acid constructs
[0199] Also provided herein are vectors comprising nucleic acid constructs (i.e., exogenous donor nucleic acid), the nucleic acid constructs comprising retinoschizine coding sequences (i.e., encoding retinoschizine protein or its fragment or variant) for integration into and expressed from the target genome locus. Also provided herein are vectors comprising nucleic acids encoding nuclease agents disclosed elsewhere herein (e.g., targeting endogenous RS1 locus). Also provided herein are vectors comprising nucleic acids encoding nuclease agents disclosed elsewhere herein (e.g., targeting endogenous RS1 locus) (e.g., vectors comprising nucleic acid constructs and DNA encoding guide RNA). The vector can include other sequences, such as replication origin, promoter, and genes encoding antibiotic resistance. Some such vectors include homology arms corresponding to the target site in the target genome locus. Other such vectors do not include any homology arms.
[0200] Some vectors can be circular. Alternatively, the vector can be linear. The vector can be packaged for delivery via lipid nanoparticles, liposomes, non-lipid nanoparticles, or viral capsids. Non-limiting exemplary vectors include plasmids, phagemids, cosmids, artificial chromosomes, minichromosomes, transposons, viral vectors, and expression vectors.
[0201] The vector can be, for example, a viral vector, such as an adeno-associated virus (AAV) vector. AAV can be any suitable serotype and can be single-stranded AAV (ssAAV) or self-complementary AAV (scAAV). Other exemplary viruses / viral vectors include retroviruses, adenoviruses, vaccinia viruses, poxviruses, and herpes simplex viruses. The virus can infect dividing cells, non-dividing cells, or both dividing and non-dividing cells. The virus can be integrated into the host genome, or alternatively not integrated into the host genome. Such viruses can also be engineered to have reduced immunity. The virus may have replication ability or may have replication defects (for example, defects in one or more genes necessary for replication and / or packaging of additional rounds of virions). The virus can cause transient expression, long-term expression (for example, at least 1 week, 2 weeks, 1 month, 2 months, or 3 months) or permanent expression (for example, Cas9 and / or gRNA). Exemplary viral titers (for example, AAV titers) include 10 12 , 10 13 , 10 14 , 10 15 and 10 16 Vector genomes / mL. Exemplary viral titers (e.g., AAV titers) include about 10 12 , about 10 13 , about 10 14 , about 10 15 and about 10 16 vector genomes (vg) / mL, or between approximately 10 12 To about 10 16 Between, about 10 12 To about 10 15 Between, about 10 12 To about 10 14 Between, about 10 12 To about 10 13 Between, about 10 13 To about 10 16 Between, about 10 14 To about 10 16 Between, about 10 15 To about 10 16 or between about 10 13 To about 10 15 Other exemplary viral titers (e.g., AAV titers) include about 1012, about 1013, about 1014, about 1015, and about 1016 vector genomes (vg) / kg body weight, or between about 10 12 To about 10 16 Between, about 10 12 To about 10 15 Between, about 1012 To about 10 14 Between, about 10 12 To about 10 13 Between, about 10 13 To about 10 16 Between, about 10 14 To about 10 16 Between, about 10 15 To about 10 16 or between about 10 13 To about 10 15 Between vg / kg body weight.
[0202] The ssDNA AAV genome consists of two open reading frames, Rep and Cap, flanked by two inverted terminal repeats that allow synthesis of complementary DNA chains. When constructing the AAV transfer plasmid, the transgene is placed between the two ITRs, and Rep and Cap can be provided in trans. In addition to Rep and Cap, AAV may also require a helper plasmid containing adenoviral genes. These genes (E4, E2a, and VA) mediate AAV replication. For example, the transfer plasmid, Rep / Cap, and helper plasmid can be transfected into HEK293 cells containing the adenoviral gene E1+ to produce infectious AAV particles. Alternatively, Rep, Cap, and adenoviral helper genes can be combined into a single plasmid. Similar packaging cells and methods can be used for other viruses, such as retroviruses.
[0203] A variety of AAV serotypes have been identified. These serotypes differ in the cell types they infect (i.e., their tropism), allowing for preferential transduction of specific cell types. The serotypes for photoreceptor cells include AAV2, AAV5, and AAV8. The serotypes for retinal pigment epithelium tissue include AAV1, AAV2, AAV4, AAV5, and AAV8. In a specific example, the AAV vector comprising the nucleic acid construct can be AAV2, AAV5, or AAV8.
[0204] Tropism can be further refined by pseudotyping, which is a mixture of capsids and genomes from different viral serotypes. For example, AAV2 / 5 indicates a virus containing a serotype 2 genome packaged in a capsid from serotype 5. The use of pseudotyped viruses can improve transduction efficiency and alter tropism. Hybrid capsids derived from different serotypes can also be used to alter viral tropism. For example, AAV-DJ contains hybrid capsids from eight serotypes and exhibits high infectivity in a wide range of cell types in vivo. AAV-DJ8 is another example that exhibits the properties of AAV-DJ, but with enhanced brain uptake. AAV serotypes can also be modified by mutations. Examples of AAV2 mutation modifications include Y444F, Y500F, Y730F, and S662V. Examples of AAV3 mutation modifications include Y705F, Y731F, and T492V. Examples of AAV6 mutation modifications include S663V and T492V. Other pseudotyped / modified AAV variants include AAV2 / 1, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2.5, AAV8.2, and AAV / SASTG. In a specific example, the AAV is AAV7m8, an AAV variant that mediates efficient delivery to all retinal layers and photoreceptors. See, for example, Dalkara et al. (2013) Sci. Transl. Med. 5:189ra76, which is incorporated herein by reference in its entirety for all purposes.
[0205] To accelerate transgene expression, self-complementary AAV (scAAV) variants can be used. Because AAV relies on the cell's DNA replication machinery to synthesize the complementary strand of the AAV single-stranded DNA genome, transgene expression may be delayed. To address this delay, scAAV can be used that contains complementary sequences that can spontaneously anneal after infection, thereby eliminating the need for host cell DNA synthesis. However, single-stranded AAV (ssAAV) vectors can also be used.
[0206] To increase packaging capacity, a longer transgene can be split between two AAV transfer plasmids, the first with a 3' splice donor and the second with a 5' splice acceptor. After co-infection of the cells, these viruses form concatemers, splice together, and the full-length transgene can be expressed. While this allows for expression of longer transgenes, the expression efficiency is lower. A similar approach for increasing capacity utilizes homologous recombination. For example, a transgene can be split between two transfer plasmids but with a large amount of sequence overlap, so that co-expression induces homologous recombination and expression of the full-length transgene.
[0207] IV. Lipid Nanoparticles Comprising Nucleic Acid Constructs
[0208] Also provided herein are lipid nanoparticles comprising nucleic acid constructs (i.e., exogenous donor nucleic acids) comprising retinoschizine coding sequences (i.e., encoding retinoschizine proteins or fragments or variants thereof) for integration into and expressed from the target genomic loci. Also provided herein are lipid nanoparticles comprising nucleic acids encoding nuclease agents (e.g., targeting endogenous RS1 loci) disclosed elsewhere herein. Also provided herein are lipid nanoparticles comprising nucleic acids encoding nuclease agents (e.g., targeting endogenous RS1 loci) disclosed elsewhere herein.
[0209] Lipid formulations can protect biomolecules from degradation while improving their cellular uptake. Lipid nanoparticles are particles comprising multiple lipid molecules that are physically associated with each other by intermolecular forces. These particles include microspheres (including unilamellar and multilamellar vesicles, for example, liposomes), dispersed phases in emulsions, micelles, or internal phases in suspensions. Such lipid nanoparticles can be used to encapsulate one or more nucleic acids or proteins for delivery. Formulations containing cationic lipids can be used to deliver polyanions such as nucleic acids. Other lipids that can be included are neutral lipids (i.e., uncharged or zwitterionic lipids), anionic lipids, helper lipids that enhance transfection, and stealth lipids that increase the length of time that nanoparticles can exist in vivo. Examples of suitable cationic lipids, neutral lipids, anionic lipids, helper lipids, and stealth lipids can be found in WO 2016 / 010840 A1 and WO 2017 / 173054 A1, which are incorporated herein by reference for all purposes. Exemplary lipid nanoparticles can include cationic lipids and one or more other components. In one example, other components can include helper lipids such as cholesterol. In another example, other components can include helper lipids such as cholesterol and neutral lipids such as DSPC. In another example, other components can include helper lipids such as cholesterol, optional neutral lipids such as DSPC, and stealth lipids such as S010, S024, S027, S031, or S033.
[0210] LNP can contain one or more or all of the following: (i) lipids for encapsulation and for endosome escape; (ii) neutral lipids for stabilization; (iii) auxiliary lipids for stabilization; (iv) stealth lipids. See, for example, Finn et al. (2018) Cell Reports (Cell Rep.) 22 (9): 2227-2235 and WO 2017 / 173054 A1, each of which is incorporated herein by reference as a whole for all purposes. In some LNPs, the payload can further include a nuclease agent. In some LNPs, the payload can further include a guide RNA or a nucleic acid encoding a guide RNA. In some LNPs, the payload can further include mRNA encoding Cas nucleases such as Cas9 and guide RNA or nucleic acid encoding a guide RNA. In some LNPs, the payload can include mRNA encoding Cas nucleases such as Cas9, guide RNA or nucleic acid encoding a guide RNA and a nucleic acid construct.
[0211] The lipid for encapsulation and endosome escape can be a cationic lipid. The lipid can also be a biodegradable lipid, such as a biodegradable ionizable lipid. An example of a suitable lipid is lipid A or LPO1, i.e., (9Z, 12Z)-3-((4,4-bis(octyloxy)butyryl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadecane-9,12-dienoate, also known as 3-((4,4-bis(octyloxy)butyryl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl (9Z, 12Z)-octadecane-9,12-dienoate. See, e.g., Finn et al. (2018) Cell Rep. 22(9):2227-2235 and WO 2017 / 173054A1, each of which is incorporated herein by reference in its entirety for all purposes. Another example of a suitable lipid is lipid B, i.e., ((5-((dimethylamino)methyl)-1,3-phenylene)bis(oxy))bis(octane-8,1-diyl)bis(decanoate), also known as ((5-((dimethylamino)methyl)-1,3-phenylene)bis(oxy))bis(octane-8,1-diyl)bis(decanoate). Another example of a suitable lipid is lipid C, 2-((4-(((3-(dimethylamino)propoxy)carbonyl)oxy)hexadecanoyl)oxy)propane-1,3-diyl(9Z,9'Z,12Z,12'Z)-bis(octadec-9,12-dienoate). Another example of a suitable lipid is lipid D, 3-(((3-(dimethylamino)propoxy)carbonyl)oxy)-13-(octanoyloxy)tridecyl 3-octyl undecanoate. Other suitable lipids include heptathriacontac-6,9,28,31-tetraen-19-yl 4-(dimethylamino)butanoate (also known as Dlin-MC3-DMA (MC3)).
[0212] Some such lipids suitable for LNP described herein are biodegradable in vivo. For example, LNPs comprising such lipids are included in at least 75% of those that remove lipids from blood plasma within 8 hours, 10 hours, 12 hours, 24 hours, or 48 hours, or within 3 days, 4 days, 5 days, 6 days, 7 days, or 10 days. As another example, at least 50% of the LNPs are removed from blood plasma within 8 hours, 10 hours, 12 hours, 24 hours, or 48 hours, or within 3 days, 4 days, 5 days, 6 days, 7 days, or 10 days.
[0213] In some embodiments, lipid can be ionizable. For example, in slightly acidic medium, lipid can be protonated and therefore have positive charge. On the contrary, in weakly alkaline medium, for example, in blood having a pH of approximately 7.35, lipid may not be protonated and therefore have no charge. In certain embodiments, lipid can be protonated at a pH of at least about 9, 9.5 or 10. The charged ability of this lipid is relevant with its intrinsic pKa. For example, the pKa of lipid can be independently in the range of about 5.8 to about 6.2.
[0214] The role of neutral lipids is to stabilize and improve the processing of LNP. The example of suitable neutral lipids includes various neutral, uncharged or zwitterionic lipids. The example of neutral phospholipids suitable for the present disclosure includes but is not limited to 5-heptadecanediol-1,3-diol (resorcinol), dipalmitoylphosphatidylcholine (DPPC), distearoylphosphatidylcholine (DSPC), phosphorylcholine (DOPC), dimyristoylphosphatidylcholine (DMPC), phosphatidylcholine (PLPC), 1,2-distearoyl-sn-glycero-3-phosphocholine (DAPC), phosphatidylethanolamine (PE), egg phosphatidylcholine (EPC), dilauroylphosphatidylcholine (DLPC), dimyristoylphosphatidylcholine (DMPC), 1-myristoyl-2-palmitoylphosphatidylcholine (MPPC), 1-palmitoyl-2-myristoylphosphatidylcholine (PMPC). The neutral phospholipids include, but are not limited to, 1,2-diisopropylphosphatidylcholine (DPPC), 1-stearoyl-2-palmitoylphosphatidylcholine (SPPC), 1,2-diisopropylphosphatidylcholine (DEPC), palmitoyloleoylphosphatidylcholine (POPC), lysophosphatidylcholine, dioleoylphosphatidylethanolamine (DOPE), dilinoleoylphosphatidylcholine distearoylphosphatidylethanolamine (DSPE), dimyristoylphosphatidylethanolamine (DMPE), dipalmitoylphosphatidylethanolamine (DPPE), palmitoyloleoylphosphatidylethanolamine (POPE), lysophosphatidylethanolamine, and combinations thereof. For example, the neutral phospholipids may be selected from the group consisting of distearoylphosphatidylcholine (DSPC) and dimyristoylphosphatidylethanolamine (DMPE).
[0215] Helper lipids include lipids that enhance transfection. The mechanism by which helper lipids enhance transfection may include enhanced particle stability. In some cases, helper lipids may enhance membrane fusogenicity. Helper lipids include steroids, sterols, and alkylresorcinols. Examples of suitable helper lipids include cholesterol, 5-heptadecanylresorcinol, and cholesterol hemisuccinate. In one example, the helper lipid may be cholesterol or cholesterol hemisuccinate.
[0216] Stealth lipids include lipids that change the length of time a nanoparticle can exist in vivo. Stealth lipids can help the formulation process by, for example, reducing particle aggregation and controlling particle size. Stealth lipids can modulate the pharmacokinetic properties of LNPs. Suitable stealth lipids include lipids with a hydrophilic head group attached to the lipid moiety.
[0217] The hydrophilic head group of the stealth lipid can include, for example, a polymer moiety selected from polymers based on PEG (sometimes referred to as poly(ethylene oxide)), poly(oxazoline), poly(vinyl alcohol), poly(glycerol), poly(N-vinyl pyrrolidone), polyamino acids, and poly-N-(2-hydroxypropyl)methacrylamide. The term PEG means any polyethylene glycol or other polyalkylene ether polymer. In certain LNP formulations, the PEG is PEG-2K, also known as PEG 2000, which has an average molecular weight of approximately 2,000 Daltons. See, for example, WO 2017 / 173054 A1, which is incorporated herein by reference in its entirety for all purposes.
[0218] The lipid portion of the stealth lipid can be derived, for example, from a diacylglycerol or dialkylglycylamide, including those comprising a dialkylglycerol or dialkylglycylamide group having an alkyl chain length independently comprising from about C4 to about C40 saturated or unsaturated carbon atoms, wherein the chain can include one or more functional groups, such as amides or esters. The diacylglycerol or dialkylglycylamide group can further include one or more substituted alkyl groups.
[0219] As an example, the stealth lipid can be selected from PEG-dilaurin, PEG-dimyristoylglycerol (PEG-DMG), PEG-dipalmitoylglycerol, PEG-distearoylglycerol (PEG-DSPE), PEG-dilaurylamide, PEG-dimyristoylglyceramide, PEG-dipalmitoylglyceramide and PEG-distearoylglyceramide, PEG-cholesterol (1-[8'-(cholest-5-en-3[β]-oxy)formamido-3',6'-dioxaoctyl]carbamoyl-[ω]-methyl-poly(ethylene glycol), PEG-DMB (3,4-dioctylbenzyl-[ω]-methyl-poly(ethylene glycol) ether), 1,2-dimyristoylglycerol. In one embodiment, the stealth lipid may be PEG2k-DMG, 1,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000](PEG2k-DSPE), 1,2-distearoyl-sn-glycero, methoxypolyethylene glycol (PEG2k-DSG), poly(ethylene glycol)-2000-dimethacrylate (PEG2k-DMA), and 1,2-distearoyloxypropyl-3-amine-N-[methoxy(polyethylene glycol)-2000](PEG2k-DSA). In a specific example, the stealth lipid may be PEG2k-DMG.
[0220] LNP can comprise the component lipid of corresponding mol ratio in the composite.The mol-% of CCD lipid can be for example about 30mol-% to about 60mol-%, about 35mol-% to about 55mol-%, about 40mol-% to about 50mol-%, about 42mol-% to about 47mol-% or about 45%.The mol-% of helper lipid can be for example about 30mol-% to about 60mol-%, about 35mol-% to about 55mol-%, about 40mol-% to about 50mol-%, about 41mol-% to about 46mol-% or about 44mol-%.The mol-% of neutral lipid can be for example about 1mol-% to about 20mol-%, about 5mol-% to about 15mol-%, about 7mol-% to about 12mol-% or about 9mol-%. The mol-% of stealth lipids can be, for example, about 1 mol-% to about 10 mol-%, about 1 mol-% to about 5 mol-%, about 1 mol-% to about 3 mol-%, about 2 mol-%, or about 1 mol-%.
[0221] LNP can have different ratios between the positively charged amine groups of the biodegradable lipid (N) and the negatively charged phosphate groups (P) of the nucleic acid to be encapsulated. This can be mathematically represented by the equation N / P. For example, the N / P ratio can be about 0.5 to about 100, about 1 to about 50, about 1 to about 25, about 1 to about 10, about 1 to about 7, about 3 to about 5, about 4 to about 5, about 4, about 4.5 or about 5.
[0222] In some LNPs, the cargo can include Cas mRNA and gRNA. The ratio of Cas mRNA to gRNA can be different. For example, the ratio of Cas mRNA to gRNA nucleic acid of the LNP formulation can range from about 25:1 to about 1:25, about 10:1 to about 1:10, about 5:1 to about 1:5 or about 1:1. Alternatively, the ratio of Cas mRNA to gRNA nucleic acid of the LNP formulation can be about 1:1 to about 1:5 or about 10:1. Alternatively, the ratio of Cas mRNA to gRNA nucleic acid of the LNP formulation can be about 1:10, 25:1, 10:1, 5:1, 3:1, 1:1, 1:3, 1:5, 1:10 or 1:25. Alternatively, the ratio of Cas mRNA to gRNA nucleic acid of the LNP formulation can be about 1:1 to about 1:2. In a specific example, the ratio of Cas mRNA to gRNA can be about 1:1 or about 1:2.
[0223] Exemplary dosages of LNPs comprise about 0.1, about 0.25, about 0.3, about 0.5, about 1, about 2, about 3, about 4, about 5, about 6, about 8, or about 10 mg / kg body weight (mpk), or about 0.1 to about 10, about 0.25 to about 10, about 0.3 to about 10, about 0.5 to about 10, about 1 to about 10, about 2 to about 10, about 3 to about 10, about 4 to about 10, about 5 to about 10, relative to the total RNA (Cas9 mRNA and gRNA) loading content. , about 6 to about 10, about 8 to about 10, about 0.1 to about 8, about 0.1 to about 6, about 0.1 to about 5, about 0.1 to about 4, about 0.1 to about 3, about 0.1 to about 2, about 0.1 to about 1, about 0.1 to about 0.5, about 0.1 to about 0.3, about 0.1 to about 0.25, about 0.25 to about 8, about 0.3 to about 6, about 0.5 to about 5, about 1 to about 5, or about 2 to about 3 mg / kg body weight. Such LNPs can be administered, for example, intravenously. In one example, a LNP dosage of between about 0.01 mg / kg and about 10 mg / kg, between about 0.1 and about 10 mg / kg, or between about 0.01 and about 0.3 mg / kg can be used. For example, a LNP dosage of about 0.01, about 0.03, about 0.1, about 0.3, about 1, about 3, or about 10 mg / kg can be used. Additional exemplary dosages of LNPs include about 0.1, about 0.25, about 0.3, about 0.5, about 1, about 2, about 3, about 4, about 5, about 6, about 8, or about 10 mg / kg (mpk) body weight, or about 0.1 to about 10, about 0.25 to about 10, about 0.3 to about 10, about 0.5 to about 10, about 1 to about 10, about 2 to about 10, about 3 to about 10, about 4 to about 10, about 5 to about 10, or about 10 mg / kg (mpk) body weight relative to the total RNA (Cas9 mRNA and gRNA) loading content. , about 6 to about 10, about 8 to about 10, about 0.1 to about 8, about 0.1 to about 6, about 0.1 to about 5, about 0.1 to about 4, about 0.1 to about 3, about 0.1 to about 2, about 0.1 to about 1, about 0.1 to about 0.5, about 0.1 to about 0.3, about 0.1 to about 0.25, about 0.25 to about 8, about 0.3 to about 6, about 0.5 to about 5, about 1 to about 5, or about 2 to about 3 mg / kg body weight. Such LNPs can be administered, for example, intravenously. In one example, a LNP dosage of between about 0.01 mg / kg and about 10 mg / kg, between about 0.1 and about 10 mg / kg, or between about 0.01 and about 0.3 mg / kg can be used. For example, a dose of about 0.01, about 0.03, about 0.1, about 0.3, about 0.5, about 1, about 2, about 3, or about 10 mg / kg of LNP can be used.In another example, a LNP dosage of between about 0.5 and about 10, between about 0.5 and about 5, between about 0.5 and about 3, between about 1 and about 10, between about 1 and about 5, between about 1 and about 3, or between about 1 and about 2 mg / kg can be used.
[0224] V. Compositions Comprising Nucleic Acid Constructs and / or Nuclease Agents or Nucleic Acids Encoding Nuclease Agents
[0225] Also provided herein are compositions comprising nucleic acid constructs, comprising retinoschizine coding sequences (i.e., encoding retinoschizine proteins or fragments or variants thereof) and nuclease agents or nucleic acids encoding nuclease agents for integration into and expression from target genomic loci, vectors, or lipid nanoparticles disclosed herein. Also provided herein are compositions comprising nucleic acid constructs, comprising retinoschizine coding sequences (i.e., encoding retinoschizine proteins or fragments or variants thereof) for integration into and expression from target genomic loci, vectors, or lipid nanoparticles disclosed herein. Also provided herein are compositions comprising nuclease agents or nucleic acids encoding nuclease agents (e.g., wherein the nuclease agent targets RS1 gene or locus) or vectors or lipid nanoparticles comprising nuclease agents or nucleic acids encoding nuclease agents. Such compositions can, for example, be used to express retinoschizine in cells or for integrating the coding sequence of retinoschizine proteins or fragments or variants thereof into target genomic loci in cells. For example, such compositions can also be used to treat subjects with X-linked juvenile retinoschizine (XLRS). Such compositions may include a nucleic acid construct comprising a coding sequence for a retinoschizine protein or a fragment thereof for integration into a target genomic locus (or a carrier or lipid nanoparticle comprising a nucleic acid construct) and a nuclease agent or a nucleic acid encoding a nuclease agent, wherein the nuclease agent targets the nuclease target sequence in the target genomic locus. The nuclease agent may be a CRISPR / Cas system (e.g., Cas protein and guide RNA) or any other suitable nuclease agent. Examples of suitable nuclease agents are provided below.
[0226] A. CRISPR / Cas system
[0227] The methods and compositions disclosed herein can utilize clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated (Cas) systems or components of such systems to modify the genome (e.g., RS1 locus) within a cell. The CRISPR / Cas system comprises transcripts and other elements related to Cas gene expression or guiding its activity. The CRISPR / Cas system can be, for example, a type I, type II, type III system, or a type V system (e.g., VA subtype or VB subtype). The methods and compositions disclosed herein can employ a CRISPR / Cas system by utilizing a CRISPR complex (including a guide RNA (gRNA) complexed with a Cas protein) for site-directed binding or cutting of nucleic acids.
[0228] The CRISPR / Cas systems used in the compositions and methods disclosed herein may be non-naturally occurring. A "non-naturally occurring" system includes anything that indicates human involvement, such as one or more components of the system being altered or mutated from their naturally occurring state, being at least substantially free of at least one other component with which it is naturally associated in nature, or being associated with at least one other component with which it is not naturally associated. For example, some CRISPR / Cas systems employ non-naturally occurring CRISPR complexes that include non-naturally occurring gRNAs and Cas proteins, employ non-naturally occurring Cas proteins, or employ non-naturally occurring gRNAs.
[0229] 1. Cas proteins
[0230] Cas proteins typically include at least one RNA recognition or binding domain that can interact with guide RNA. Cas proteins can also include a nuclease domain (e.g., a DNase domain or an RNase domain), a DNA binding domain, a helicase domain, a protein-protein interaction domain, a dimerization domain, and other domains. Some such domains (e.g., a DNase domain) can be derived from natural Cas proteins. Other such domains can be added to prepare modified Cas proteins. The nuclease domain has catalytic activity for nucleic acid cleavage, which comprises the breaking of the covalent bonds of nucleic acid molecules. The cutting can produce flat ends or staggered ends, and it can be single-stranded or double-stranded. For example, wild-type Cas9 protein will typically produce blunt cleavage products. Alternatively, wild-type Cpf1 protein (e.g., FnCpf1) can produce cleavage products with 5 nucleotide 5' overhangs, wherein the cutting occurs after the 18th base pair from the PAM sequence on the non-targeting chain and after the 23rd base pair based on the targeting chain. The Cas protein can have full cleavage activity to generate double-strand breaks (e.g., double-strand breaks with blunt ends) at the target genomic locus, or it can be a nickase that generates single-strand breaks at the target genomic locus.
[0231] Examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csn1 or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (CasA), Cse2 (CasB), Cse3 (CasE), Cse4 (CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cu1966, and homologs or modified versions thereof.
[0232] An exemplary Cas protein is a Cas9 protein or a protein derived from a Cas9 protein. Cas9 proteins are from type II CRISPR / Cas systems and generally share four key motifs with conserved structures. Motifs 1, 2, and 4 are RuvC-like motifs, and motif 3 is an HNH motif. Exemplary Cas9 proteins are from Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, delbrueckii), Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp.), Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus), Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp.), Petrotoga mobilis, Thermosiphoafricanus, Acaryochloris marina, Neisseria meningitidis or Campylobacter jejuni. Additional examples of Cas9 family members are described in WO2014 / 131833, which is incorporated herein by reference in its entirety for all purposes. Cas9 (SpCas9) from Streptococcus pyogenes (e.g., assigned UniProt accession number Q99ZW2) is an exemplary Cas9 protein. An exemplary SpCas9 protein sequence is shown in SEQ ID NO: 27 (encoded by the DNA sequence shown in SEQ ID NO: 26). An exemplary SpCas9 cDNA sequence is shown in SEQ ID NO: 28. Smaller Cas9 proteins (e.g., Cas9 proteins whose coding sequences are compatible with maximum AAV packaging capacity when combined with guide RNA coding sequences and regulatory elements of Cas9 and guide RNA, such as SaCas9 and CjCas9 and Nme2Cas9) are other exemplary Cas9 proteins. For example, Cas9 from Staphylococcus aureus (SaCas9) (assigned UniProt accession number J7RUA5) is another exemplary Cas9 protein. Similarly, Cas9 from Campylobacter jejuni (CjCas9), for example (assigned UniProt accession number QOP897) is another exemplary Cas9 protein. See, for example, Kim et al. (2017), Nature Communications 8: 14500, which is incorporated herein by reference in its entirety for all purposes. SaCas9 is smaller than SpCas9, and CjCas9 is smaller than both SaCas9 and SpCas9. Cas9 from Neisseria meningitidis (Nme2Cas9) is another exemplary Cas9 protein. See, for example, Edraki et al. (2019) Mol.Cell, 73(4):714-726, which is incorporated herein by reference in its entirety for all purposes. Cas9 proteins from Streptococcus thermophilus (e.g., Streptococcus thermophilus LMD-9Cas9 (St1Cas9) encoded by the CRISPR1 locus or Streptococcus thermophilus Cas9 (St3Cas9) from the CRISPR3 locus) are other exemplary Cas9 proteins. Cas9 from Francisella novicida (FnCas9) or RHA Francisella novicida Cas9 variants that recognize alternative PAMs (E1369R / E1449H / R1556A substitutions) are other exemplary Cas9 proteins. These and other exemplary Cas9 proteins are reviewed, for example, in Cebrian-Serrano and Davies (2017), Mammal Genome, 28(7):247-261, which is incorporated herein by reference in its entirety for all purposes. Examples of Cas9 coding sequences, Cas9 mRNAs, and Cas9 protein sequences are provided in WO 2013 / 176772, WO 2014 / 065596, WO 2016 / 106121, and WO 2019 / 067910, each of which is incorporated herein by reference in its entirety for all purposes. Specific examples of ORFs and Cas9 amino acid sequences are provided in Table 30 of
[0449] of WO 2019 / 067910, and specific examples of Cas9 mRNAs and ORFs are provided in
[0214] -
[0234] of WO 2019 / 067910. As an example, Cas9 proteins include, are essentially composed of, or are composed of the sequence shown in SEQ ID NO:6242. Such Cas9 protein sequences can be encoded by, are essentially composed of, or are composed of mRNA encoding: SEQ ID NO:6243. As another example, the Cas9 protein can comprise, consist essentially of, or consist of the sequence shown in SEQ ID NO: 6246. Such a Cas9 protein sequence can be encoded by an mRNA comprising, consisting essentially of, or consisting of SEQ ID NO: 6245.
[0233] Another example of a Cas protein is Cpf1 (CRISPR from Prevotella and Francisella 1). Cpf1 is a large protein (approximately 1300 amino acids) that contains a RuvC-like nuclease domain homologous to the corresponding domain of Cas9, as well as a counterpart to the characteristic arginine-rich Cas9 cluster. However, Cpf1 lacks the HNH nuclease domain present in the Cas9 protein, and the RuvC-like domain is continuous in the Cpf1 sequence, while Cas9, on the contrary, contains a long insert containing the HNH domain. See, for example, Zetsche et al. (2015) Cell 163(3):759-771, which is incorporated herein by reference in its entirety for all purposes. Exemplary Cpf1 proteins are from Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrioproteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Shigella sp. SCADC, Acidaminococcus sp.) BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, and Porphyromonas macacae.Cpf1 from Francisella novicida U112 (FnCpf1; assigned UniProt accession number A0Q7Q2) is an exemplary Cpf1 protein.
[0234] The Cas protein can be a wild-type protein (i.e., those proteins present in nature), a modified Cas protein (i.e., a Cas protein variant), or a fragment of a wild-type or modified Cas protein. In terms of the catalytic activity of the wild-type or modified Cas protein, the Cas protein can also be an active variant or fragment. In terms of catalytic activity, the active variant or fragment may include at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the wild-type or modified Cas protein or a portion thereof, wherein the active variant retains the ability to cut at the desired cleavage site and thus retains nick induction or double-strand break induction activity. Assays for nick induction or double-strand break induction activity are known, and typically measure the overall activity and specificity of the Cas protein for DNA substrates containing cleavage sites.
[0235] The Cas protein can be modified to increase or decrease one or more of nucleic acid binding affinity, nucleic acid binding specificity, and enzymatic activity. The Cas protein can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or the Cas protein can be truncated to remove domains that are unnecessary for protein function or to optimize (e.g., enhance or reduce) the activity or property of the Cas protein.
[0236] An example of a modified Cas protein is a modified SpCas9-HF1 protein, which is a high-fidelity variant of Streptococcus pyogenes Cas9 with changes (N497A / R661A / Q695A / Q926A) designed to reduce non-specific DNA contacts. See, for example, Kleinstiver et al. (2016) Nature 529(7587):490-495, which is incorporated herein by reference in its entirety for all purposes. Another example of a modified Cas protein is a modified eSpCas9 variant (K848A / K1003A / R1060A) designed to reduce off-target effects. See, for example, Slaymaker et al. (2016) Science 351(6268):84-88, which is incorporated herein by reference in its entirety for all purposes. Other SpCas9 variants include K855A and K810A / K1003A / R1060A. These and other modified Cas proteins are reviewed, for example, in Cebrian-Serrano and Davies (2017) Mammalian Genomes 28(7): 247-261, which is incorporated herein by reference in its entirety for all purposes. Another example of a modified Cas9 protein is xCas9, which is a SpCas9 variant that can recognize an expanded range of PAM sequences. See, for example, Hu et al. (2018), Nature 556: 57-63, which is incorporated herein by reference in its entirety for all purposes.
[0237] Cas protein can include at least one nuclease domain, such as a DNase domain. For example, wild-type Cpf1 protein generally includes a RuvC-like domain that cuts two chains of target DNA, which may be in a dimer configuration. Cas protein can also include at least two nuclease domains, such as a DNase domain. For example, wild-type Cas9 protein generally includes a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC domain and the HNH domain can each cut different chains of double-stranded DNA to produce double-strand breaks in DNA. See, for example, Jinek et al. (2012), Science 337: 816-821, which is incorporated herein by reference in its entirety for all purposes.
[0238] One or more or all nuclease domains in the nuclease domain can be missing or mutated so that it no longer has function or has reduced nuclease activity.For example, if one of the nuclease domains in the Cas9 protein is missing or mutated, the resulting Cas9 protein can be referred to as a nickase, and single-strand breaks can be produced in double-stranded target DNA, but double-strand breaks will not be produced (that is, it can cut complementary strands or non-complementary strands, but can not cut both at the same time). If the two in the nuclease domain are missing or mutated, the ability of the two chains of the resulting Cas protein (for example, Cas9) to cut double-stranded DNA will be reduced (for example, nuclease is invalid or nuclease inactivation Cas protein, or catalytic death Cas protein (dCas)). If there is no nuclease domain in the Cas9 protein that is missing or mutated, the Cas9 protein will retain double-strand break inducing activity. The example of the mutation Cas9 is converted into a nickase is the D10A (aspartic acid is converted to alanine at position 10 of Cas9) mutation in the RuvC domain of the Cas9 from Streptococcus pyogenes. Similarly, H939A (histidine is converted to alanine at amino acid position 839), H840A (histidine is converted to alanine at amino acid position 840) or N863A (asparagine is converted to alanine at amino acid position N863) in the HNH domain of Cas9 from Streptococcus pyogenes can convert Cas9 into a nicking enzyme. Other examples of mutations that convert Cas9 into a nicking enzyme include corresponding mutations of Streptococcus thermophilus to Cas9. See, for example, Sapranauskas et al. (2011) Nucleic Acids Research 39 (21): 9275-9282 and WO 2013 / 141680, each of which is incorporated herein by reference in its entirety for all purposes. Such mutations can be produced using methods such as site-directed mutagenesis, PCR-mediated mutagenesis, or total gene synthesis. Other mutation examples producing nickase can be found in, for example, WO 2013 / 176772 and WO 2013 / 142578, and each document in the document is incorporated herein by reference as a whole for all purposes. If all nuclease domains in Cas protein are deleted or mutated (for example, two nuclease domains in Cas9 protein are deleted or mutated), the ability of the two chains of the double-stranded DNA cut by the gained Cas protein (for example, Cas9) will be reduced (for example, nuclease is invalid or nuclease inactivation Cas protein). An instantiation is D10A / H840A Streptococcus pyogenes Cas9 double mutant or the corresponding double mutant from the Cas9 of another species when compared with Streptococcus pyogenes Cas9 best. Another instantiation is D10A / N863A Streptococcus pyogenes Cas9 double mutant or the corresponding double mutant from the Cas9 of another species when compared with Streptococcus pyogenes Cas9 best.
[0239] The example of the inactivation mutation in the catalytic domain of xCas9 is the same as the mutation for SpCas9 described above. The example of the inactivation mutation in the catalytic domain of Staphylococcus aureus Cas9 protein is also known. For example, Staphylococcus aureus Cas9 enzyme (SaCas9) can include a substitution at position N580 (e.g., N580A substitution) and a substitution at position D10 (e.g., D10A substitution) for producing a nuclease-inactivated Cas protein. See, for example, WO 2016 / 106236, which is incorporated herein by reference in its entirety for all purposes. The example of the inactivation mutation in the catalytic domain of Nme2Cas9 is also known (e.g., a combination of D16A and H588A). The example of the inactivation mutation in the catalytic domain of St1Cas9 is also known (e.g., a combination of D9A, D598A, H599A and N622A). The example of the inactivation mutation in the catalytic domain of St3Cas9 is also known (e.g., a combination of D10A and N870A). Examples of inactivating mutations in the catalytic domain of CjCas9 are also known (e.g., a combination of D8A and H559A). Examples of inactivating mutations in the catalytic domain of FnCas9 and RHA FnCas9 are also known (e.g., N995A).
[0240] Examples of inactivating mutations in the catalytic domain of Cpf1 proteins are also known. With respect to Cpf1 proteins from Francisella novicida U112 (FnCpf1), Acidaminococcus sp. BV3L6 (AsCpf1), Lachnospiraceae ND2006 (LbCpf1), and Moraxella bovis 237 (MbCpf1 Cpf1), such mutations may comprise mutations at positions 908, 993, or 1263 of AsCpf1, or corresponding positions in Cpf1 orthologs, or positions 832, 925, 947, or 1180 of LbCpf1, or corresponding positions in Cpf1 orthologs. Such mutations may include, for example, one or more of the mutations D908A, E993A, and D1263A of AsCpf1 or corresponding mutations in Cpf1 orthologs, or D832A, E925A, D947A, and D1180A of LbCpf1 or corresponding mutations in Cpf1 orthologs. See, e.g., US 2016 / 0208243, which is incorporated herein by reference in its entirety for all purposes.
[0241] Cas proteins can also be operably linked to heterologous polypeptides as fusion proteins. For example, Cas proteins can be fused to cleavage domains or epigenetic modification domains. See WO 2014 / 089290, which is incorporated herein by reference in its entirety for all purposes. Cas proteins can also be fused to heterologous polypeptides to provide increased or decreased stability. The fusion domain or heterologous polypeptide can be located at the N-terminus, C-terminus, or inside the Cas protein.
[0242] As an example, the Cas protein can be fused to one or more heterologous polypeptides that provide subcellular localization. Such heterologous polypeptides may include, for example, one or more nuclear localization signals (NLS), such as single-component SV40 NLS and / or two-component α-import protein NLS for targeting the nucleus, mitochondrial localization signals for targeting mitochondria, ER retention signals, etc. See, for example, Lange et al. (2007) Journal of Biological Chemistry (J.Biol.Chem.) 282 (8): 5101-5105, which is incorporated herein by reference in its entirety for all purposes. Such subcellular localization signals can be located at any position within the N-terminus, C-terminus, or Cas protein. The NLS may include a basic amino acid segment and may be a monopartite sequence or a dipartite sequence. Optionally, the Cas protein may include two or more NLSs, including an NLS at the N-terminus (e.g., an α-import protein NLS or a monopartite NLS) and an NLS at the C-terminus (e.g., an SV40 NLS or a dipartite NLS). The Cas protein may also include two or more NLSs at the N-terminus and / or two or more NLSs at the C-terminus.
[0243] For example, the Cas protein can be fused to 1-10 NLSs (e.g., fused to 1-5 NLSs or fused to one NLS). When using one NLS, the NLS can be connected to the N-terminus or C-terminus of the Cas protein sequence. It can also be inserted into the Cas protein sequence. Alternatively, the Cas protein can be fused to more than one NLS. For example, the Cas protein can be fused to 2, 3, 4 or 5 NLSs. In a specific example, the Cas protein can be fused to two NLSs. In some cases, the two NLSs can be the same (e.g., two SV40 NLSs) or different. For example, the Cas protein can be fused to two SV40 NLS sequences connected at the carboxyl termini. Alternatively, the Cas protein can be fused to two NLSs, one connected to the N-terminus and one connected to the C-terminus. In other examples, the Cas protein can be fused to 3 NLSs or not fused to NLSs. NLS can be a single sequence, such as SV40 NLS, PKKKRKV (SEQ ID NO: 49) or PKKKRRV (SEQ ID NO: 50). The NLS can be a bipartite sequence, such as the NLS of the nucleoplasmin, KRPAATKKAGQAKKKK (SEQ ID NO: 51). In a specific example, a single PKKKRKV (SEQ ID NO: 49) NLS can be attached to the C-terminus of the Cas protein. One or more linkers are optionally included at the fusion site.
[0244] Cas protein can also be operably connected to a cell penetrating domain or a protein transduction domain. For example, the cell penetrating domain can be derived from HIV-1TAT protein, TLM cell penetration motif, MPG, Pep-1, VP22, cell penetrating peptides or polyarginine peptide sequences from human hepatitis B virus. See, for example, WO2014 / 089290 and WO 2013 / 176772, each document in the document is incorporated herein by reference as a whole for all purposes. The cell penetrating domain can be positioned at any position within the N-terminal, C-terminal or Cas protein.
[0245] Cas proteins can also be operably linked to heterologous polypeptides for tracking or purification, such as fluorescent proteins, purification tags or epitope tags. Examples of fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, emerald, Azami green, monomeric Azami green, CopGFP, AceGFP, ZsGreenl), yellow fluorescent proteins (e.g., YFP, eYFP, lemon yellow, Venus, YPet, PhiYFP, ZsYellowl), blue fluorescent proteins (e.g., eBFP, eBFP2, azurite, mKalamal, GFPuv, sky blue, T-sapphire), cyan fluorescent proteins (e.g., eCFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), red fluorescent proteins (e.g., mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-expressed, DsRed2, DsRed-monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRaspberry, mStrawberry, Jred), orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, monomeric Kusabira-Orange, mTangerine, tdTomato), and any other suitable fluorescent protein. Examples of tags include glutathione-S-transferase (GST), chitin binding protein (CBP), maltose binding protein, thioredoxin (TRX), poly (NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, hemagglutinin (HA), nus, Softag 1, Softag3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, histidine (His), biotin carboxyl carrier protein (BCCP), and calmodulin.
[0246] Cas proteins can also be tethered to labeled nucleic acids or donor sequences. This tethering (i.e., physical connection) can be achieved through covalent interactions or non-covalent interactions, and the tethering can be direct (e.g., by direct fusion or chemical conjugation, which can be achieved by modification of cysteine or lysine residues on the protein or intron modification), or can be achieved through one or more intermediate linkers or adapter molecules such as streptavidin or aptamers. See, e.g., Pierce et al. (2005), Mini Rev. Med. Chem. 5(1):41-55; Duckworth et al. (2007), Angew. Chem. Int. Ed. Engl. 46(46):8819-8822; Schaeffer and Dixon (2009), Australian J. Chem. 62(10):1328-1332; Goodman et al. (2009), Chembiochem. 10(9):1551-1557; and Khatwani et al. (2012), Bioorg. Med. Chem. 20(14):4532-4539, each of which is incorporated herein by reference in its entirety for all purposes. Non-covalent strategies for synthesizing protein-nucleic acid conjugates include biotin-streptavidin and nickel-histidine methods. Covalent protein-nucleic acid conjugates can be synthesized by connecting appropriately functionalized nucleic acids and proteins using a variety of chemical reactions. Some of these chemical reactions involve attaching oligonucleotides directly to amino acid residues (e.g., lysine amine or cysteine thiol) on the protein surface, while other more complex schemes require the participation of post-translational modification of proteins or catalytic or reactive protein domains. Methods for covalently linking proteins to nucleic acids can include, for example, chemical cross-linking of oligonucleotides to protein lysine or cysteine residues, expressed protein connections, chemoenzymatic methods, and the use of photoaptamers. The labeled nucleic acid or donor sequence can be tethered to the C-terminus, N-terminus, or internal region within the Cas protein. In one example, the labeled nucleic acid or donor sequence is tethered to the C-terminus or N-terminus of the Cas protein. Similarly, the Cas protein can be tethered to the 5' end, 3' end, or internal region within the labeled nucleic acid or donor sequence. That is, the labeled nucleic acid or donor sequence can be tethered in any orientation and polarity. For example, the Cas protein can be tethered to the 5' end or the 3' end of the labeled nucleic acid or donor sequence.
[0247] Cas protein can be provided in any form. For example, Cas protein can be provided in the form of protein, such as Cas protein complexed with gRNA. Alternatively, Cas protein can be provided in the form of nucleic acid encoding Cas protein, such as RNA (e.g., messenger RNA (mRNA)) or DNA. Optionally, the nucleic acid encoding Cas protein can be codon optimized to be effectively translated into protein in a specific cell or organism. For example, compared with naturally occurring polynucleotide sequences, the nucleic acid encoding Cas protein can be modified to replace codons with higher usage frequencies in bacterial cells, yeast cells, human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells or any other host cells of interest. When the nucleic acid encoding Cas protein is introduced into a cell, the Cas protein can be expressed transiently, conditionally or constitutively in the cell.
[0248] The Cas protein provided as mRNA can be modified to improve stability and / or immunogenic properties. One or more nucleosides within the mRNA can be modified. Examples of chemical modifications of mRNA nucleobases include pseudouridine, 1-methyl-pseudouridine, and 5-methyl-cytidine. For example, capped and polyadenylated Cas mRNAs containing N1-methylpseudouridine can be used. Similarly, Cas mRNAs can be modified by deleting uridine using synonymous codons.
[0249] The nucleic acid encoding the Cas protein can be stably integrated into the genome of the cell and operably connected to a promoter active in the cell. Alternatively, the nucleic acid encoding the Cas protein can be operably connected to a promoter in an expression construct. The expression construct comprises any nucleic acid construct capable of directing the expression of a gene or other nucleic acid sequence of interest (e.g., Cas gene) and can transfer such nucleic acid sequence of interest to a target cell. For example, the nucleic acid encoding the Cas protein can be in a vector comprising a DNA encoding gRNA. Alternatively, it can be in a vector or plasmid separated from a vector comprising a DNA encoding gRNA. The promoter that can be used for the expression construct is included in one or more cells in, for example, eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, pluripotent cells, embryonic stem (ES) cells, adult stem cells, developmentally restricted progenitor cells, inducible pluripotent stem (iPS) cells, or single-cell embryos. Such promoters can be, for example, conditional promoters, inducible promoters, constitutive promoters, or tissue-specific promoters. Optionally, the promoter can be a bidirectional promoter that drives expression of the Cas protein in one direction and expression of the guide RNA in the other direction. Such bidirectional promoters can be composed of: (1) a complete, conventional, unidirectional Pol III promoter containing three external control elements: a distal sequence element (DSE), a proximal sequence element (PSE), and a TATA box; (2) a second basic Pol III promoter containing a PSE and a TATA box fused to the 5' end of the DSE in the opposite orientation. For example, in the H1 promoter, the DSE is adjacent to the PSE and the TATA box, and the promoter can be bidirectionalized by generating a hybrid promoter in which reverse transcription is controlled by an additional PSE and a TATA box derived from the U6 promoter. See, for example, US 2016 / 0074535, which is incorporated herein by reference in its entirety for all purposes. The use of a bidirectional promoter to simultaneously express genes encoding Cas proteins and guide RNAs allows the generation of compact expression cassettes to facilitate delivery.
[0250] Different promoters can be used to drive Cas expression or Cas9 expression. In some methods, a small promoter is used so that the Cas or Cas9 coding sequence can be adapted to an AAV construct. For example, Cas or Cas9 and one or more gRNAs (for example, 1 gRNA or 2 gRNAs or 3 gRNAs or 4 gRNAs) can be delivered by LNP-mediated delivery (for example, in the form of RNA) or adeno-associated virus (AAV)-mediated delivery (for example, AAV2-mediated delivery, AAV5-mediated delivery, AAV8-mediated delivery or AAV7m8-mediated delivery). For example, the nuclease agent can be CRISPR / Cas9, and the Cas9 mRNA and gRNA targeting endogenous RS1 locus (for example, intron 1 of RS1) can be delivered by LNP-mediated delivery, or the DNA encoding Cas9 and the DNA encoding the gRNA targeting endogenous RS1 locus (for example, intron 1 of RS1) can be delivered by AAV-mediated delivery. Cas or Cas9 and gRNA can be delivered in a single AAV or by two separate AAVs. For example, the first AAV can carry a Cas or Cas9 expression cassette, and the second AAV can carry a gRNA expression cassette. Similarly, the first AAV can carry a Cas or Cas9 expression cassette, and the second AAV can carry two or more gRNA expression cassettes. Alternatively, a single AAV can carry a Cas or Cas9 expression cassette (for example, a Cas or Cas9 encoding sequence operably connected to a promoter) and a gRNA expression cassette (for example, a gRNA encoding sequence operably connected to a promoter). Similarly, a single AAV can carry a Cas or Cas9 expression cassette (for example, a Cas or Cas9 encoding sequence operably connected to a promoter) and two or more gRNA expression cassettes (for example, a gRNA encoding sequence operably connected to a promoter). Different promoters can be used to drive the expression of gRNA, such as U6 promoter or small tRNA Gln. Similarly, different promoters can be used to drive Cas9 expression. For example, a small promoter is used so that the Cas9 encoding sequence can be adapted to the AAV construct. Similarly, small Cas9 proteins (e.g., SaCas9 or CjCas9) are used to maximize AAV packaging capacity.
[0251] The Cas protein provided as mRNA can be modified to improve stability and / or immunogenic properties. One or more nucleosides in the mRNA can be modified. Examples of chemical modifications to mRNA core bases include pseudouridine, 1-methyl-pseudouridine and 5-methyl-cytidine. The mRNA encoding the Cas protein can also be capped. For example, the cap can be a cap 1 structure, in which +1 ribonucleotides are methylated at the 2'O position of the ribose. For example, capping can produce excellent activity in vivo (for example, by mimicking a natural cap), and can produce a natural structure that reduces the stimulation of the host's innate immune system (for example, the activation of pattern recognition receptors in the innate immune system can be reduced). The mRNA encoding the Cas protein can also be polyadenylated (to include a poly (A) tail). The mRNA encoding the Cas protein can also be modified to include pseudouridine (for example, it can be completely replaced with pseudouridine). As another example, capping and polyadenylation Cas mRNA containing N1-methylpseudouridine can be used. As another example, a Cas mRNA completely substituted with pseudouridine (i.e., all standard uracil residues are replaced with pseudouridine, which is a uridine isomer in which uracil is linked by carbon-carbon bonds rather than nitrogen-carbon bonds) can be used. Similarly, the Cas mRNA can be modified by deleting uridine using synonymous codons. For example, a capped and polyadenylated Cas mRNA completely substituted with pseudouridine can be used.
[0252] Cas mRNA can include modified uridine at at least one, multiple or all uridine positions. The modified uridine can be a uridine modified at the 5 position (e.g., a uridine modified with a halogen, methyl or ethyl). The modified uridine can be a pseudouridine modified at the 1 position (e.g., a pseudouridine modified with a halogen, methyl or ethyl). The modified uridine can be, for example, pseudouridine, N1-methylpseudouridine, 5-methoxyuridine, 5-iodouridine or a combination thereof. In some instances, the modified uridine is 5-methoxyuridine. In some instances, the modified uridine is 5-iodouridine. In some instances, the modified uridine is pseudouridine. In some instances, the modified uridine is N1-methylpseudouridine. In some instances, the modified uridine is a combination of pseudouridine and N1-methylpseudouridine. In some instances, the modified uridine is a combination of pseudouridine and 5-methoxyuridine. In some instances, the modified uridine is a combination of N1-methylpseudouridine and 5-methoxyuridine. In some instances, the modified uridine is a combination of 5-iodouridine and N1-methylpseudouridine. In some instances, the modified uridine is a combination of pseudouridine and 5-iodouridine. In some instances, the modified uridine is a combination of 5-iodouridine and 5-methoxyuridine.
[0253] The Cas mRNA disclosed herein may also include a 5' cap, such as Cap0, Cap1, or Cap2. The 5' cap is typically a 7-methylguanine ribonucleotide (which may be further modified, for example, with respect to ARCA) connected to the 5' position of the first nucleotide (that is, the first cap-proximal nucleotide) of the 5' to 3' chain of the mRNA via a 5'-triphosphate. In Cap 0, the ribose of the first cap-proximal nucleotide and the second cap-proximal nucleotide of the mRNA both include a 2'-hydroxyl group. In Cap 1, the ribose of the first transcribed nucleotide and the second transcribed nucleotide of the mRNA include a 2'-methoxy group and a 2'-hydroxyl group, respectively. In Cap 2, the ribose of the first cap-proximal nucleotide and the second cap-proximal nucleotide of the mRNA both include a 2'-methoxy group. See, for example, Katibah et al. (2014) Proceedings of the National Academy of Sciences of the United States of America 111(33):12025-30 and Abbas et al. (2017) Proceedings of the National Academy of Sciences of the United States of America 114(11):E2106-E2115, each of which is incorporated herein by reference in its entirety for all purposes. Most endogenous higher eukaryotic mRNAs (including mammalian mRNAs, such as human mRNAs) include cap 1 or cap 2. Cap 0 and other cap structures that are different from cap 1 and cap 2 may be immunogenic in mammals such as humans because components of the innate immune system (such as IFIT-1 and IFIT-5) recognize them as non-self, which may lead to elevated levels of cytokines including type I interferons. Components of the innate immune system, such as IFIT-1 and IFIT-5, may also compete with eIF4E to bind to mRNAs with caps other than cap 1 or cap 2, thereby potentially inhibiting translation of the mRNA.
[0254] The cap can be co-transcriptionally included. For example, ARCA (anti-reverse cap analog; ThermoFisher Scientific catalog number AM8045) is a cap analog comprising 7-methylguanine 3'-methoxy-5'-triphosphate linked to the 5' position of a guanine ribonucleotide that can be incorporated into the transcript in vitro at the beginning. ARCA produces a cap 0 cap in which the 2' position of the first cap-proximal nucleotide is a hydroxyl group. See, for example, Stepinski et al. (2001) RNA 7: 1486-1495, which is incorporated herein by reference in its entirety for all purposes.
[0255] CleanCap TM AG(m7G(5')ppp(5')(2'OMeA)pG; TriLink Biotechnology Co., Ltd. catalog number: N-7113 or CleanCap TMGG (m7G (5') ppp (5') (2'OMeG) pG; TriLink Biotechnology Co., Ltd. Catalog No.: N-7133) can be used to co-transcriptionally provide the cap 1 structure. TM AG and CleanCap TM 3'-O-methylated versions of GG can also be obtained from TriLink Biotechnology, Inc., catalog numbers N-7413 and N-7433, respectively.
[0256] Alternatively, a cap can be added to the RNA after transcription. For example, vaccinia capping enzyme is commercially available (New England Biolabs, catalog number: M2080S) and has RNA triphosphatase and uridyltransferase activity provided by its D1 subunit and guanine methyltransferase provided by its D12 subunit. In this way, the vaccinia capping enzyme can add 7-methylguanine to RNA in the presence of S-adenosylmethionine and GTP to provide cap O. See, for example, Guo and Moss, (1990) Proceedings of the National Academy of Sciences of the United States of America 87:4023-4027 and Mao and Shuman, (1994) Journal of Biological Chemistry 269:24472-24479, each of which is incorporated herein by reference in its entirety for all purposes.
[0257] The Cas mRNA may further include an adenylated (poly-A or poly(A) or poly-adenine) tail. For example, the poly-A tail may include at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, or at least 100 adenines, and optionally up to 300 adenines. For example, the poly-A tail may include 95, 96, 97, 98, 99, or 100 adenine nucleotides.
[0258] 2. Guide RNA
[0259] "Guide RNA" or "gRNA" is an RNA molecule that binds to a Cas protein (e.g., Cas9 protein) and targets the Cas protein to a specific location within the target DNA. A guide RNA can include two segments: a "DNA targeting segment" (also referred to as a "guide sequence") and a "protein binding segment." A "segment" comprises a portion or region of a molecule, such as a continuous stretch of nucleotides in an RNA. Some gRNAs, such as those of Cas9, can include two separate RNA molecules: an "activator RNA" (e.g., tracrRNA) and a "targeting factor RNA" (e.g., CRISPR RNA or crRNA). Other gRNAs are single RNA molecules (single RNA polynucleotides), which can also be referred to as "single molecule gRNA," "single guide RNA," or "sgRNA." See, for example, WO 2013 / 176772, WO 2014 / 065596, WO 2014 / 089290, WO 2014 / 093622, WO 2014 / 099750, WO 2013 / 142578, and WO 2014 / 131833, each of which is incorporated herein by reference in its entirety for the purposes described. Guide RNA can refer to a combination of CRISPR RNA (crRNA) or crRNA and transactivating CRISPR RNA (tracrRNA). CrRNA and tracrRNA can be associated as a single RNA molecule (single guide RNA or sgRNA) or in two separate RNA molecules (dual guide RNA or dgRNA). For example, for Cas9, a single guide RNA can include (for example, via a linker) a crRNA fused to tracrRNA. For example, for Cpf1, only one crRNA is required to achieve binding and / or cutting to the target sequence. The terms "guide RNA" and "gRNA" include both bimolecular (i.e., modular) gRNAs and unimolecular gRNAs. In some methods and compositions disclosed herein, the gRNA is a Streptococcus pyogenes Cas9 gRNA or its equivalent. In some methods and compositions disclosed herein, the gRNA is a Staphylococcus aureus Cas9 gRNA or its equivalent.
[0260] Exemplary bimolecular gRNAs include crRNA-like ("CRISPR RNA" or "targeting factor RNA" or "crRNA" or "crRNA repeats") molecules and corresponding tracrRNA-like ("trans-activating CRISPR RNA" or "activator RNA" or "tracrRNA") molecules. The crRNA includes both the DNA targeting segment (single strand) of the gRNA and the nucleotide segment (i.e., crRNA tail) that forms half of the dsRNA duplex of the protein binding segment of the gRNA. Examples of crRNA tails positioned downstream (3') of the DNA targeting segment include, consist essentially of, or consist of GUUUUAGAGCUAUGCU (SEQ ID NO: 29) or GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 52). Any of the DNA targeting segments disclosed herein can be linked to the 5' end of SEQ ID NO: 29 or 52 to form a crRNA.
[0261] The corresponding tracrRNA (activator RNA) includes a nucleotide segment of the other half of the dsRNA duplex forming the protein binding segment of the gRNA. The nucleotide segment of crRNA is complementary to the nucleotide segment of tracrRNA and hybridizes with it to form a dsRNA duplex of the protein binding domain of the gRNA. Therefore, each crRNA can be regarded as having a corresponding tracrRNA. Exemplary tracrRNA sequences include, are essentially composed of, or are composed of: AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUU (SEQ ID NO: 30), AAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 31) or GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 32).
[0262] In a system where both crRNA and tracrRNA are needed, crRNA and corresponding tracrRNA hybridize to form gRNA. In a system where only crRNA is needed, crRNA can be gRNA. CrRNA additionally provides a single-stranded DNA targeting segment that hybridizes with the complementary strand of the target DNA. If used for intracellular modification, the exact sequence of a given crRNA or tracrRNA molecule can be designed to be specific to the species in which the RNA molecule will be used. See, e.g., Mali et al. (2013) Science 339(6121):823-826; Jinek et al. (2012) Science 337(6096):816-821; Hwang et al. (2013) Nature Biotechnology 31(3):227-229; Jiang et al. (2013) Nature Biotechnology 31(3):233-239; and Cong et al. (2013) Science 339(6121):819-823, each of which is incorporated herein by reference in its entirety for all purposes.
[0263] The DNA targeting segment (crRNA) of a given gRNA includes a nucleotide sequence complementary to the sequence on the complementary strand of the target DNA, as described in more detail below. The DNA targeting segment of gRNA interacts with the target DNA in a sequence-specific manner by hybridization (i.e., base pairing). Therefore, the nucleotide sequence of the DNA targeting segment can be different, and the positioning in the target DNA with which the gRNA and target DNA interact is determined. The DNA targeting segment of the subject gRNA can be modified to hybridize with any desired sequence in the target DNA. Naturally occurring crRNA is different from CRISPR / Cas systems and organisms, but generally contains a length of 21 to 72 nucleotides flanked by two direct repeats (DR) of 21 to 46 nucleotides (see, for example, WO 2014 / 131833, which is incorporated herein by reference for all purposes). In the case of Streptococcus pyogenes, the length of DR is 36 nucleotides, and the length of the targeting segment is 30 nucleotides. The DR positioned at the 3' end is complementary and hybridized with the corresponding tracrRNA, which then binds to the Cas protein.
[0264] The length of the DNA targeting segment can be, for example, at least about 12, at least about 15, at least about 17, at least about 18, at least about 19, at least about 20, at least about 25, at least about 30, at least about 35, or at least about 40 nucleotides. Such DNA targeting segments can be, for example, about 12 to about 100, about 12 to about 80, about 12 to about 50, about 12 to about 40, about 12 to about 30, about 12 to about 25, or about 12 to about 20 nucleotides in length. For example, the DNA targeting segment can be about 15 to about 25 nucleotides (e.g., about 17 to about 20 nucleotides or about 17, 18, 19, or 20 nucleotides). See, for example, US 2016 / 0024523, which is incorporated herein by reference in its entirety for all purposes. For Cas9 from Streptococcus pyogenes, typical DNA targeting segments are between 16 and 20 nucleotides in length or between 17 and 20 nucleotides in length. For Cas9 from Staphylococcus aureus, typical DNA targeting segments are between 21 and 23 nucleotides in length. For Cpf1, typical DNA targeting segments are at least 16 nucleotides in length or at least 18 nucleotides in length.
[0265] In one example, the length of the DNA targeting segment can be about 20 nucleotides. However, shorter and longer sequences can also be used for the targeting segment (e.g., a length of 15-25 nucleotides, such as a length of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides). The degree of identity between the DNA targeting segment and the corresponding guide RNA target sequence (or the degree of complementarity between the DNA targeting segment and the other strand of the guide RNA target sequence) can be, for example, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100%. The DNA targeting segment and the corresponding guide RNA target sequence can include one or more mismatches. For example, the DNA-targeting segment of a guide RNA and the corresponding guide RNA target sequence can contain 1-4, 1-3, 1-2, 1, 2, 3, or 4 mismatches (e.g., wherein the total length of the guide RNA target sequence is at least 17, at least 18, at least 19, or at least 20 or more nucleotides). For example, the DNA-targeting segment of a guide RNA and the corresponding guide RNA target sequence can contain 1-4, 1-3, 1-2, 1, 2, 3, or 4 mismatches, wherein the total length of the guide RNA target sequence is 20 nucleotides.
[0266] As an example, a guide RNA targeting an RS1 gene can include a DNA targeting segment (i.e., a guide sequence) comprising, consisting essentially of, or consisting of a sequence as set forth in any one of SEQ ID NOs: 3148-6241 (DNA targeting segment). Alternatively, a guide RNA targeting an RS1 gene can include a DNA targeting segment comprising, consisting essentially of, or consisting of a sequence as set forth in any one of SEQ ID NOs: 3148-6241 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3148-6241 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3148-6241 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to at least 17, at least 18, at least 19 or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3148-6241 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to at least 17, at least 18, at least 19 or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3148-6241 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment comprising, consisting essentially of, or consisting of a sequence that differs from the sequence shown in any one of SEQ ID NOs: 3148-6241 (DNA targeting segment) by no more than 3, no more than 2, or no more than 1 nucleotides.Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3148-6241 (DNA targeting segment). Examples of such guide sequences are shown in Tables 2 and 3.
[0267] Guide RNA can target people RS1 gene.As an example, the guide RNA of targeting RS1 gene can include DNA targeting segment (that is, guide sequence), and described DNA targeting segment includes the sequence (DNA targeting segment) shown in any one of SEQ ID NO:3148-4989, is substantially composed of it or is composed of it.Alternately, the guide RNA of targeting RS1 gene can include DNA targeting segment, and described DNA targeting segment includes at least 17, at least 18, at least 19 or at least 20 continuous nucleotides of the sequence (DNA targeting segment) shown in any one of SEQ ID NO:3148-4989, is substantially composed of it or is composed of it. Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3148-4989 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3148-4989 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to at least 17, at least 18, at least 19 or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3148-4989 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to at least 17, at least 18, at least 19 or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3148-4989 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment comprising, consisting essentially of, or consisting of a sequence that differs from the sequence shown in any one of SEQ ID NOs: 3148-4989 (DNA targeting segment) by no more than 3, no more than 2, or no more than 1 nucleotides.Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment comprising, consisting essentially of, or consisting of a sequence that differs from at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3148-4989 (DNA targeting segment) by no more than 3, no more than 2, or no more than 1 nucleotide.
[0268] Guide RNA can target people RS1 gene and be selected to avoid missing target effect.As an example, the guide RNA of target RS1 gene can include DNA targeting section (that is, guide sequence), described DNA targeting section includes the sequence (DNA targeting section) shown in any one of SEQ ID NO:3148-3151,3154-3186,3188-3247 and 3249-4351, is substantially composed of it or is composed of it.Alternately, the guide RNA of target RS1 gene can include DNA targeting section, described DNA targeting section includes the sequence (DNA targeting section) shown in any one of SEQ ID NO:3148-3151,3154-3186,3188-3247 and 3249-4351 at least 17, at least 18, at least 19 or at least 20 continuous nucleotides, is substantially composed of it or is composed of it. Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to at least 17, at least 18, at least 19 or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247 and 3249-4351 (DNA targeting segment).Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to at least 17, at least 18, at least 19 or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247 and 3249-4351 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotides from the sequence shown in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotides from the sequence shown in any one of SEQ ID NOs: 3148-3151, 3154-3186, 3188-3247, and 3249-4351 (DNA targeting segment).
[0269] Guide RNA can target people RS1 gene.As an example, the guide RNA of targeting RS1 gene can include DNA targeting segment (that is, guide sequence), and described DNA targeting segment includes the sequence shown in any one of SEQ ID NO:3150,3151,3159,3675,4297 and 4304 (DNA targeting segment), is essentially composed of it or is composed of it.Alternately, the guide RNA of targeting RS1 gene can include DNA targeting segment, and described DNA targeting segment includes the sequence shown in any one of SEQ ID NO:3150,3151,3159,3675,4297 and 4304 (DNA targeting segment) at least 17, at least 18, at least 19 or at least 20 continuous nucleotides, is essentially composed of it or is composed of it. Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to at least 17, at least 18, at least 19 or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297 and 4304 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to at least 17, at least 18, at least 19 or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297 and 4304 (DNA targeting segment).Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotides from the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304 (DNA targeting segment). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotides from the sequence shown in any one of SEQ ID NOs: 3150, 3151, 3159, 3675, 4297, and 4304 (DNA targeting segment).
[0270] Guide RNA can target mouse Rs1 gene.As an example, the guide RNA to RS1 gene can include DNA targeting segment (that is, guide sequence), and described DNA targeting segment includes the sequence (DNA targeting segment) shown in any one of SEQ ID NO:4990-6241 (for example, SEQ ID NO:5477 or 5981), is substantially composed of it or is composed of it.Alternately, the guide RNA to target RS1 gene can include DNA targeting segment, and described DNA targeting segment includes the sequence (DNA targeting segment) shown in any one of SEQ ID NO:4990-6241 (for example, SEQ ID NO:5477 or 5981) at least 17, at least 18, at least 19 or at least 20 continuous nucleotides, is substantially composed of it or is composed of it. Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NOs: 5477 or 5981). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the sequence shown in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NOs: 5477 or 5981). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to at least 17, at least 18, at least 19 or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NO: 5477 or 5981). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment that is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to at least 17, at least 18, at least 19 or at least 20 consecutive nucleotides of the sequence shown in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NO: 5477 or 5981).Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotides from the sequence shown in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NOs: 5477 or 5981). Alternatively, the guide RNA targeting the RS1 gene can include a DNA targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotides from the sequence shown in any one of SEQ ID NOs: 4990-6241 (e.g., SEQ ID NOs: 5477 or 5981).
[0271] TracrRNA can be in any form (for example, full-length tracrRNA or active partial tracrRNA) and has different lengths.TracrRNA can include primary transcripts or processed forms.For example, tracrRNA (as a part of a single guide RNA or as a separate molecule as a part of a bimolecular gRNA) can include a wild-type tracrRNA sequence (for example, about or more than about 20, 26, 32, 45, 48, 54, 63, 67, 85 or more nucleotides of a wild-type tracrRNA sequence), consisting essentially of or consisting of. The example of the wild-type tracrRNA sequence from Streptococcus pyogenes includes 171 nucleotides, 89 nucleotides, 75 nucleotides and 65 nucleotide versions.See, for example, Deltcheva et al. (2011) Nature 471 (7340): 602-607; WO 2014 / 093661, each of which is incorporated herein by reference in its entirety for all purposes. Examples of tracrRNA within a single guide RNA (sgRNA) include the tracrRNA segments found in the +48, +54, +67, and +85 versions of the sgRNA, where "+n" indicates that up to +n nucleotides of the wild-type tracrRNA are included in the sgRNA. See US 8,697,359, which is incorporated herein by reference in its entirety for all purposes.
[0272] The complementarity percentage between the DNA targeting segment of the guide RNA and the complementary strand of the target DNA can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100%). The complementarity percentage between the DNA targeting segment and the complementary strand of the target DNA can be at least 60% over about 20 consecutive nucleotides. As an example, the complementarity percentage between the DNA targeting segment and the complementary strand of the target DNA can be 100% over 14 consecutive nucleotides at the 5' end of the complementary strand of the target DNA and be as low as 0% over the remainder. In such cases, the DNA targeting segment can be considered to be 14 nucleotides in length. As another example, the complementarity percentage between the DNA targeting segment and the complementary strand of the target DNA can be 100% over seven consecutive nucleotides at the 5' end of the complementary strand of the target DNA and be as low as 0% over the remainder. In such cases, the DNA targeting segment can be considered to be 7 nucleotides in length. In some guide RNAs, at least 17 nucleotides in the DNA targeting segment are complementary to the complementary strand of the target DNA. For example, the length of the DNA targeting segment can be 20 nucleotides and can include 1, 2 or 3 mismatches with the complementary strand of the target DNA. In one example, the mismatch is not adjacent to the district corresponding to the complementary strand of the protospacer adjacent motif (PAM) sequence (i.e., the reverse complement of the PAM sequence) (for example, the mismatch is located at the 5' end of the DNA targeting segment of the guide RNA, or the mismatch is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18 or 19 base pairs from the district corresponding to the complementary strand of the PAM sequence).
[0273] The protein binding fragment of the gRNA can include two nucleotide segments that are complementary to each other. The complementary nucleotides of the protein binding segment hybridize to form a double-stranded RNA duplex (dsRNA). The protein binding fragment of the subject gRNA interacts with the Cas protein, and the gRNA guides the bound Cas protein to a specific nucleotide sequence within the target DNA through the DNA targeting segment.
[0274] A single guide RNA can include a DNA targeting segment and a scaffold sequence (i.e., the protein binding or Cas binding sequence of the guide RNA). For example, such a guide RNA can have a 5' DNA targeting segment connected to a 3' scaffold sequence. Exemplary scaffold sequences include, consist essentially of, or consist of: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCU (version 1; SEQ ID NO: 33); GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 2; SEQ ID NO: 34); GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 3; SEQ ID NO: 35); GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 4; SEQ ID NO:36); GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU (version 5; SEQ ID NO:37); GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUU (version 6; SEQ ID NO:38); GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUU (version 7; SEQ ID NO:39); or GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGGCACCGAGUCGGUGC (version 8; SEQ ID NO:53). In some guide sgRNAs, the four terminal U residues of version 6 were absent. In some sgRNAs, only 1, 2, or 3 of the four terminal U residues of version 6 were present.A guide RNA targeting any of the guide RNA target sequences disclosed herein can comprise, for example, a DNA targeting segment on the 5' end of the guide RNA fused to any of the exemplary guide RNA scaffold sequences on the 3' end of the guide RNA. That is, any of the DNA targeting segments disclosed herein can be linked to the 5' end of any of the scaffold sequences described above to form a single guide RNA (chimeric guide RNA).
[0275] The guide RNA may comprise modifications or sequences that provide additional desired characteristics (e.g., modified or regulated stability; subcellular targeting; tracking with fluorescent markers; binding sites for proteins or protein complexes; etc.). That is, the guide RNA may comprise one or more modified nucleosides or nucleotides, or one or more non-natural and / or naturally occurring components or configurations that replace or are in addition to the canonical A, G, C, and U residues. Examples of such modifications include, for example, a 5' cap (e.g., a 7-methylguanylate cap (m7G)); a 3' polyadenylation tail (i.e., a 3' poly(A) tail); a riboswitch sequence (e.g., to allow regulation of stability and / or regulation of accessibility of a protein and / or protein complex); a stability control sequence; a sequence that forms a dsRNA duplex (i.e., a hairpin); a modification or sequence that targets the RNA to a subcellular location (e.g., the nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides for tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescence detection, a sequence that allows fluorescence detection, etc.); a modification or sequence that provides a binding site for a protein (e.g., a protein that acts on DNA, including transcription activators, transcription repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, etc.); and combinations thereof. Other examples of modification include engineered stem-loop duplex structures, engineered bulge regions, engineered hairpins 3' of stem-loop duplex structures, or any combination thereof. See, for example, US 2015 / 0376586, which is incorporated herein by reference in its entirety for all purposes. The bulge can be an unpaired region of nucleotides within the duplex consisting of a crRNA-like region and a minimal tracrRNA-like region. The bulge can include an unpaired 5'-XXXY-3' on one side of the duplex, wherein X is any purine, and Y can include nucleotides that can form a wobble pair with the nucleotides on the opposite chain; and include an unpaired nucleotide region on the other side of the duplex.
[0276] Unmodified nucleic acids may be susceptible to degradation. Exogenous nucleic acids may also induce an innate immune response. Modifications may help introduce stability and reduce immunogenicity. Guide RNAs may include modified nucleosides and modified nucleotides, including, for example, one or more of the following: (1) alterations or substitutions of one or both of the non-linked phosphate oxygens and / or one or more of the linked phosphate oxygens in the phosphodiester backbone linkage (exemplary backbone modifications); (2) alterations or substitutions of the composition of the ribose sugar, such as alterations or substitutions of the 2' hydroxyl group on the ribose sugar (exemplary sugar modifications); (3) substitutions (e.g., large-scale substitutions) of phosphate moieties with dephospholinkers (exemplary backbone modifications); (4) substitutions of naturally occurring nucleobases (including substitutions of non-canonical nucleobases with ... (Exemplary base modifications); (5) substitution or modification of the ribose-phosphate backbone (Exemplary backbone modifications); (6) removal, modification or substitution of the terminal phosphate group or a portion of the 3' or 5' end of the oligonucleotide (such 3' or 5' cap modifications may include sugar and / or backbone modifications); and (7) modification or substitution of a sugar (Exemplary sugar modifications). Other possible guide RNA modifications include modification or substitution of uracil or polyuracil tracts. See, for example, WO 2015 / 048577 and US 2016 / 0237455, each of which is incorporated herein by reference in its entirety for all purposes. Similar modifications can be made to Cas encoding nucleic acids, such as Cas mRNA. For example, Cas mRNA can be modified by depletion of uridine using synonymous codons.
[0277] Chemical modifications (such as those listed above) can be combined to provide modified gRNA and / or mRNA including residues (nucleosides and nucleotides) that can have two, three, four or more modifications. For example, the modified residue can have a modified sugar and a modified core base. In one example, each base of the modified gRNA (e.g., all bases have a modified phosphate group, such as a thiophosphate group). For example, all or substantially all phosphate groups of the gRNA can be replaced with a thiophosphate group. Alternatively or additionally, the modified gRNA can include at least one modified residue at or near the 5' end. Alternatively or additionally, the modified gRNA can include at least one modified residue at or near the 3' end.
[0278] Some gRNAs include one, two, three or more modified residues. For example, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or 100% of the positions in the modified gRNA can be modified nucleosides or nucleotides.
[0279] Unmodified nucleic acids may be susceptible to degradation. Exogenous nucleic acids can also induce an innate immune response. Modification can help introduce stability and reduce immunogenicity. Some gRNAs described herein may comprise one or more modified nucleosides or nucleotides to introduce stability to intracellular or serum-based nucleases. When introduced into a cell population, some modified gRNAs described herein may exhibit a reduced innate immune response.
[0280] The gRNA disclosed herein can include backbone modifications, wherein the phosphate groups of the modified residues can be modified by replacing one or more oxygens with different substituents. The modifications can include large-scale replacement of unmodified phosphate moieties with modified phosphate groups as described herein. The backbone modifications of the phosphate backbone can also include changes that result in uncharged linkers or charged linkers with asymmetric charge distribution.
[0281] The example of modified phosphate group comprises thiophosphate, selenophosphate, borane phosphoric acid (boranophosphate), borane phosphate ester (borano phosphate ester), hydrogen phosphate, phosphoramidate, alkyl or aryl phosphonate and phosphotriester.The phosphorus atom in unmodified phosphate group is achiral.However, one of the non-bridge oxygen can be replaced with one of the above-mentioned atoms or atomic groups to make the phosphorus atom chiral.The stereophosphorus atom can have " R " configuration (Rp) or " S " configuration (Sp).The skeleton can also be modified by replacing the bridge oxygen (that is, the oxygen connecting phosphate and nucleoside) with nitrogen (bridged phosphoramidate), sulfur (bridged thiophosphate) and carbon (bridged methylene phosphonate).Substitution can occur at the connection oxygen place or at two connection oxygen places in the connection oxygen.
[0282] In some backbone modifications, phosphate group can be replaced by a phosphorus-free joint. In certain embodiments, charged phosphate group can be replaced by a neutral moiety. Examples of parts that can replace phosphate group include but are not limited to, such as methylphosphonate, hydroxyamino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformaldehyde, methylaldehyde, oxime, methyleneimino, methyleneformimino, methylenehydrazono (methylenehydrazo), methylenedimethylhydrazono (methylenedimethylhydrazo) and methyleneoxyformimino.
[0283] It is also possible to construct a scaffold for mimicking nucleic acids in which the phosphate linker and ribose are replaced by nuclease-resistant nucleoside or nucleotide substitutes. Such modifications may include backbone modifications and sugar modifications. In certain embodiments, the nucleobase may be tethered to an alternative backbone. Examples may include, but are not limited to, morpholino, cyclobutyl, pyrrolidine, and peptide nucleic acid (PNA) nucleoside substitutes.
[0284] Modified nucleosides and modified nucleotides can include one or more modifications to the sugar group (sugar modifications). For example, the 2' hydroxyl (OH) group can be modified (e.g., replaced by a variety of different "oxy" or "deoxy" substituents). Modifications to the 2' hydroxyl group can enhance the stability of the nucleic acid because the hydroxyl group can no longer be deprotonated to form a 2'-alkoxide ion.
[0285] Examples of 2' hydroxyl modifications may include: alkoxy or aryloxy (OR, where "R" can be, for example, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar); polyethylene glycol (PEG), O(CH2CH2O); n CH2CH2OR, wherein R can be, for example, H or an optionally substituted alkyl group, and n can be an integer from 0 to 20 (e.g., 0 to 4, 0 to 8, 0 to 10, 0 to 16, 1 to 4, 1 to 8, 1 to 10, 1 to 16, 1 to 20, 2 to 4, 2 to 8, 2 to 10, 2 to 16, 2 to 20, 4 to 8, 4 to 10, 4 to 16, and 4 to 20). The 2' hydroxyl modification can be 2'-O-Me. Similarly, the 2' hydroxyl modification can be a 2'-fluoro modification, which replaces the 2' hydroxyl group with fluoride. The 2' hydroxyl modification can comprise a "locked" nucleic acid (LNA), wherein the 2' hydroxyl group can be, for example, C 1-6 Alkylene or C 1-6 A heteroalkylene bridge is attached to the 4' carbon of the same ribose, wherein exemplary bridges can include a methylene bridge, a propylene bridge, an ether bridge, or an amino bridge; an ortho-amino group (wherein the amino group can be, for example, NH2; an alkylamino group, a dialkylamino group, a heterocyclyl group, an arylamino group, a diarylamino group, a heteroarylamino group, or a diheteroarylamino group, ethylenediamine, or a polyamino group) and an aminoalkoxy group, O(CH2) n-amino (wherein the amino group can be, for example, NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, or polyamino). The 2' hydroxyl modification can comprise an unlocked nucleic acid (UNA), wherein the ribose ring lacks a C2'-C3' bond. The 2' hydroxyl modification can comprise a methoxyethyl (MOE), (OCH2CH2OCH3, such as a PEG derivative).
[0286] The deoxy 2' modification can comprise hydrogen (i.e., deoxyribose, such as part of an overhang of a dsRNA); halogen (e.g., bromine, chloride, fluorine, or iodine); amino (wherein the amino group can be, for example, NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or amino acid); NH(CH2CH2NH) n CH2CH2-amino (wherein amino can be, for example, as described herein), -NHC(O)R (wherein R can be, for example, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl or a sugar), cyano; thiol; alkyl-thio-alkyl; thioalkoxy; and alkyl, cycloalkyl, aryl, alkenyl and alkynyl, which can be optionally substituted, for example, with amino as described herein.
[0287] Sugar modification can include sugar groups that can also contain one or more carbons with a stereochemical configuration opposite to the stereochemical configuration of the corresponding carbon in ribose. Therefore, modified nucleic acids can include nucleotides containing, for example, arabinose as sugar. Modified nucleic acids can also include abasic sugars. These abasic sugars can also be further modified at one or more sugar atoms in the composition sugar atoms. Modified nucleic acids can also include one or more sugars in the form of L (e.g., L-nucleosides).
[0288] The modified nucleosides and modified nucleotides that can be incorporated into modified nucleic acids described herein can include modified bases, also referred to as core bases. Examples of core bases include but are not limited to adenine (A), guanine (G), cytosine (C) and uracil (U). These core bases can be modified or completely replaced to provide modified residues that can be incorporated into modified nucleic acids. The core bases of nucleotides can be independently selected from purine, pyrimidine, purine analogs or pyrimidine analogs. In certain embodiments, core bases can include, for example, naturally occurring and synthetic base derivatives.
[0289] In dual guide RNAs, each of the crRNA and tracrRNA can contain modifications. Such modifications can be at one or both ends of the crRNA and / or tracrRNA. In sgRNA, one or more residues at one or both ends of the sgRNA can be chemically modified, and / or internal nucleosides can be modified, and / or the entire sgRNA can be chemically modified. Some gRNAs include 5' end modifications. Some gRNAs include 3' end modifications.
[0290] The guide RNA disclosed herein may include one of the modification patterns disclosed in WO 2018 / 107028 A1, which is incorporated herein by reference in its entirety for all purposes. The guide RNA disclosed herein may also include one of the structures / modification patterns disclosed in US2017 / 0114334, which is incorporated herein by reference in its entirety for all purposes. The guide RNA disclosed herein may also include one of the structures / modification patterns disclosed in WO 2017 / 136794, WO 2017 / 004279, US 2018 / 0187186 or US 2019 / 0048338, each of which is incorporated herein by reference in its entirety for all purposes.
[0291] As an example, the nucleotides at the 5' or 3' ends of the guide RNA may include a phosphorothioate bond (e.g., a base may have a modified phosphate group, i.e., a phosphorothioate group). For example, the guide RNA may include a phosphorothioate bond between 2, 3, or 4 terminal nucleotides at the 5' or 3' ends of the guide RNA. As another example, the nucleotides at the 5' and / or 3' ends of the guide RNA may have a 2'-O-methyl modification. For example, the guide RNA may include 2'-O-methyl modifications at 2, 3, or 4 terminal nucleotides at the 5' and / or 3' ends (e.g., 5' ends) of the guide RNA. See, for example, WO 2017 / 173054A1 and Finn et al., (2018) Cell Reports 22 (9): 2227-2235, each of which is incorporated herein by reference in its entirety for all purposes. Other possible modifications are described in more detail elsewhere herein. In one embodiment, the guide RNA comprises a 2'-O-methyl analog and a 3' phosphorothioate internucleotide bond at the first three 5' and 3' end RNA residues. In another embodiment, the guide RNA is modified so that all 2'OH groups that do not interact with the Cas9 protein are replaced by 2'-O-methyl analogs, and the tail region of the guide RNA that interacts minimally with Cas9 is modified with a 5' and 3' phosphorothioate internucleotide bond. In addition, the DNA targeting segment may have a 2'-fluoro modification on certain bases. See, for example, Yin et al. (2017), Nature Biotechnology 35(12): 1179-1187, which is incorporated herein by reference in its entirety for all purposes. Other examples of modified guide RNAs are provided in, for example, WO 2018 / 107028 A1, which is incorporated herein by reference in its entirety for all purposes. For example, such chemical modifications can provide guide RNA with higher stability and protection from exonucleases, making its intracellular residence time longer than that of unmodified guide RNA. For example, such chemical modifications may also prevent innate intracellular immune responses that could actively degrade RNA or trigger immune cascades leading to cell death.
[0292] As an example, any guide RNA in the guide RNA described herein can include at least one modification. In one example, the at least one modification includes nucleotides modified with 2'-O-methyl (2'-O-Me), phosphorothioate (PS) bonds between nucleotides, nucleotides modified with 2'-fluoro (2'-F), or a combination thereof. For example, at least one modification can include nucleotides modified with 2'-O-methyl (2'-O-Me). Alternatively or additionally, the at least one modification can include phosphorothioate (PS) bonds between nucleotides. Alternatively or additionally, the at least one modification can include nucleotides modified with 2'-fluoro (2'-F). In one example, the guide RNA described herein includes one or more nucleotides modified with 2'-O-methyl (2'-O-Me) and one or more phosphorothioate (PS) bonds between nucleotides.
[0293] Modifications can occur anywhere in the guide RNA. As an example, the guide RNA includes modifications at one or more of the first five nucleotides at the 5' end of the guide RNA, the guide RNA includes modifications at one or more of the last five nucleotides at the 3' end of the guide RNA, or a combination thereof. For example, the guide RNA can include a phosphorothioate bond between the first four nucleotides of the guide RNA, a phosphorothioate bond between the last four nucleotides of the guide RNA, or a combination thereof. Alternatively or additionally, the guide RNA can include 2'-O-Me modified nucleotides at the first three nucleotides at the 5' end of the guide RNA, can include 2'-O-Me modified nucleotides at the last three nucleotides at the 3' end of the guide RNA, or a combination thereof.
[0294] In one example, the modified gRNA can include the following sequence: mN*mN*mN*NNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmU*mU*mU*mU (SEQ ID NO: 44), wherein "N" can be any natural or non-natural nucleotide, and wherein the sum of N residues includes an RS1 DNA-targeting segment as described herein (e.g., as shown in SEQ ID NO: 44), wherein the N residues are represented by any one of SEQ ID NOs: 3148-6241, or any one of SEQ ID NOs: 3148-4989, or SEQ ID NO: 44. NO: 3148-3151, 3154-3186, 3188-3247 and 3249-4351, or any one of SEQ ID NO: 3150, 3151, 3159, 3675, 4297 and 4304, or any one of SEQ ID NO: 4990-6241 (e.g., 5477 or 5981). The terms "mA", "mC", "mU" and "mG" refer to nucleotides that have been modified with 2'-O-Me (A, C, U and G, respectively). The symbol "*" depicts a phosphorothioate modification. A phosphorothioate linkage or bond refers to a bond in which sulfur replaces a non-bridging phosphate oxygen in a phosphodiester linkage, such as in a bond between nucleotide bases. When phosphorothioates are used to generate oligonucleotides, the modified oligonucleotides may also be referred to as sulfur oligonucleotides. The terms A*, C*, U*, or G* represent a nucleotide that is linked to the next (e.g., 3') nucleotide via a phosphorothioate bond. The terms "mA*," "mC*," "mU*," and "mG*" represent a nucleotide (A, C, U, and G, respectively) that has been substituted with 2'-O-Me and is linked to the next (e.g., 3') nucleotide via a phosphorothioate bond.
[0295] Another chemical modification that has been shown to affect the nucleotide sugar ring is halogen substitution. For example, 2'-fluorine (2'-F) substitutions on the nucleotide sugar ring can increase oligonucleotide binding affinity and nuclease stability. Abasic nucleotides refer to those nucleotides that lack a nitrogenous base. Reverse bases refer to those reverse bases that have connections that are reversed from the normal 5' to 3' connection (that is, 5' to 5' connection or 3' to 3' connection).
[0296] Abasic nucleotides can be connected to reverse linkages. For example, an abasic nucleotide can be connected to a terminal 5' nucleotide via a 5' to 5' linkage, or an abasic nucleotide can be connected to a terminal 3' nucleotide via a 3' to 3' linkage. Reverse abasic nucleotides at the terminal 5' or 3' nucleotides can also be referred to as reverse abasic end caps.
[0297] In one example, one or more of the first three, four, or five nucleotides at the 5' end and one or more of the last three, four, or five nucleotides at the 3' end are modified. The modifications can be, for example, 2'-O-Me, 2'-F, inverted abasic nucleotides, phosphorothioate linkages, or other well-known nucleotide modifications that increase stability and / or performance.
[0298] In another example, the first four nucleotides at the 5' end and the last four nucleotides at the 3' end can be linked with phosphorothioate bonds.
[0299] In another example, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end can include nucleotides modified with 2'-O-methyl (2'-O-Me). In another example, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end include nucleotides modified with 2'-fluoro (2'-F). In another example, the first three nucleotides at the 5' end and the last three nucleotides at the 3' end include reverse abasic nucleotides.
[0300] Guide RNA can be provided in any form. For example, gRNA can be provided in the form of RNA, as two molecules (individual crRNA and tracrRNA) or as one molecule (sgRNA), and optionally provided in the form of a complex with Cas protein. gRNA can also be provided in the form of DNA encoding gRNA. The DNA encoding gRNA can encode a single RNA molecule (sgRNA) or a separate RNA molecule (e.g., a separate crRNA and tracrRNA). In the latter case, the DNA encoding gRNA can be provided as a DNA molecule or as a separate DNA molecule encoding crRNA and tracrRNA respectively.
[0301] When gRNA is provided in the form of DNA, gRNA can be transient, conditional or constitutively expressed in cells. The DNA encoding gRNA can be stably integrated into the genome of the cell and operably connected to a promoter active in the cell. Alternatively, the DNA encoding gRNA can be operably connected to a promoter in an expression construct. For example, the DNA encoding gRNA can be in a vector comprising heterologous nucleic acids such as nucleic acids encoding Cas proteins. Alternatively, the DNA encoding gRNA can be in a vector or plasmid separated from a vector comprising nucleic acids encoding Cas proteins. The promoter that can be used for such expression constructs is included in one or more cells in, for example, eukaryotic cells, human cells, non-human cells, mammalian cells, non-human mammalian cells, rodent cells, mouse cells, rat cells, pluripotent cells, embryonic stem (ES) cells, adult stem cells, developmentally restricted progenitor cells, inducible pluripotent stem (iPS) cells or single cell stage embryos and has an active promoter. Such promoters can be, for example, conditional promoters, inducible promoters, constitutive promoters or tissue-specific promoters. Such promoters can also be, for example, bidirectional promoters. Specific examples of suitable promoters include RNA polymerase III promoters, such as human U6 promoter, rat U6 polymerase III promoter, or mouse U6 polymerase III promoter. In another example, small tRNA Gln can be used to drive expression of the guide RNA.
[0302] Alternatively, gRNA can be prepared by various other methods. For example, gRNA can be prepared by in vitro transcription using, for example, T7 RNA polymerase (see, for example, WO 2014 / 089290 and WO 2014 / 065596, each document in the document is incorporated herein by reference in its entirety for all purposes). Guide RNA can also be a synthetic molecule prepared by chemical synthesis. For example, guide RNA can be chemically synthesized to comprise 2'-O-methyl analogs and 3' phosphorothioate internucleotide bonds at the first three 5' and 3' end RNA residues.
[0303] The guide RNA (or nucleic acid encoding the guide RNA) can be in a composition comprising one or more guide RNAs (e.g., 1, 2, 3, 4 or more guide RNAs) and a carrier that increases the stability of the guide RNA (e.g., prolongs the time that degradation products remain below a threshold value, such as less than 0.5% of the weight of the starting nucleic acid or protein, under given storage conditions (e.g., -20°C, 4°C or ambient temperature); or increases in vivo stability). Non-limiting examples of such carriers include poly(lactic acid) (PLA) microspheres, poly(D,L-lactic-co-glycolic acid) (PLGA) microspheres, liposomes, micelles, reverse micelles, lipospires, and lipid microtubules. Such compositions can further include a Cas protein, such as a Cas9 protein, or a nucleic acid encoding a Cas protein.
[0304] 3. Guide RNA target sequence
[0305] The target DNA for guide RNA includes the nucleic acid sequence present in DNA, and the DNA targeting segment of gRNA will be combined with it, provided that there are enough binding conditions.Suitable DNA / RNA binding conditions include the physiological conditions normally present in the cell.Other suitable DNA / RNA binding conditions (for example, conditions in cell-free systems) are known in the art (see, for example, "Molecular Cloning: Laboratory Manual", 3rd edition (Sambrook et al., Harbor Laboratory Press (Harbor Laboratory Press) 2001), the document is incorporated herein by reference in its entirety for all purposes). The chain of the target DNA complementary to gRNA and hybridized can be referred to as "complementary chain", and the chain of the target DNA complementary to "complementary chain" (and therefore not complementary to Cas protein or gRNA) can be referred to as "non-complementary chain" or "template chain".
[0306] The target DNA comprises a sequence on the complementary strand to which the guide RNA hybridizes and a corresponding sequence on the non-complementary strand (e.g., adjacent to the protospacer sequence adjacent motif (PAM)). Unless otherwise indicated, as used herein, the term "guide RNA target sequence" specifically refers to a sequence on the non-complementary strand corresponding to the sequence to which the guide RNA hybridizes on the complementary strand (i.e., the reverse complement). That is, the guide RNA target sequence refers to a sequence on the non-complementary strand adjacent to the PAM (e.g., upstream or 5' of the PAM in the case of Cas9). The guide RNA target sequence is equivalent to the DNA targeting segment of the guide RNA, but has thymine instead of uracil. As an example, the guide RNA target sequence of the SpCas9 enzyme can refer to a sequence upstream of the 5'-NGG-3' PAM on the non-complementary strand. The guide RNA is designed to complement the complementary strand of the target DNA, wherein the hybridization between the DNA targeting segment of the guide RNA and the complementary strand of the target DNA promotes the formation of the CRISPR complex. Complete complementarity is not necessarily required, provided that there is sufficient complementarity to cause hybridization and promote the formation of the CRISPR complex. If a guide RNA is referred to herein as targeting a guide RNA target sequence, it means that the guide RNA hybridizes to a complementary strand sequence of the target DNA that is the reverse complement of the guide RNA target sequence on the non-complementary strand.
[0307] The target DNA or guide RNA target sequence can include any polynucleotide and can be located, for example, in the nucleus or cytoplasm of a cell or in an organelle of the cell, such as a mitochondria or chloroplast. The target DNA or guide RNA target sequence can be any nucleic acid sequence that is endogenous or exogenous to the cell. The guide RNA target sequence can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory sequence) or can comprise both.
[0308] Site-specific binding and cleavage of the target DNA by the Cas protein can occur at a location determined by both (i) base pairing complementarity between the guide RNA and the complementary strand of the target DNA and (ii) a short motif in the non-complementary strand of the target DNA, known as a protospacer adjacent motif (PAM). The PAM can flank the guide RNA target sequence. Optionally, the guide RNA target sequence can be flanked by a PAM on the 3' end (e.g., for Cas9). Alternatively, the guide RNA target sequence can be flanked by a PAM on the 5' end (e.g., for Cpf1). For example, the cleavage site of the Cas protein can be about 1 to about 10 or about 2 to about 5 base pairs (e.g., 3 base pairs) upstream or downstream of the PAM sequence (e.g., within the guide RNA target sequence). In the case of SpCas9, the PAM sequence (i.e., on the non-complementary strand) can be 5'-N1GG-3', where N1 is any DNA nucleotide, and where the PAM is immediately 3' of the guide RNA target sequence on the non-complementary strand of the target DNA. Thus, the sequence corresponding to the PAM on the complementary strand (i.e., the reverse complement) will be 5'-CCN2-3', where N2 is any DNA nucleotide and is immediately 5' of the sequence to which the DNA targeting segment of the guide RNA hybridizes on the complementary strand of the target DNA. In some such cases, N1 and N2 can be complementary, and the N1-N2 base pair can be any base pair (e.g., N1 = C and N2 = G; N1 = G and N2 = C; N1 = A and N2 = T; or N1 = T and N2 = A). In the case of Cas9 from Staphylococcus aureus, the PAM can be NNGRRT or NNGRR, where N can be A, G, C, or T, and R can be G or A. In the case of Cas9 from Campylobacter jejuni, the PAM can be, for example, NNNNACAC or NNNNRYAC, where N can be A, G, C, or T, and R can be G or A. In some cases (eg, for FnCpf1), the PAM sequence may be located upstream of the 5' end and have the sequence 5'-TTN-3'.
[0309] An example of a guide RNA target sequence is a 20-nucleotide DNA sequence immediately preceding the NGG motif recognized by the SpCas9 protein. For example, two examples of guide RNA target sequences plus a PAM are GN 19 NGG (SEQ ID NO: 40) or N 20 NGG (SEQ ID NO: 41). See, for example, WO 2014 / 165825, which is incorporated herein by reference in its entirety for all purposes. The guanine at the 5' end can promote transcription by RNA polymerase in cells. Other examples of guide RNA target sequences plus PAM can include two guanine nucleotides at the 5' end (e.g., GGN 20NGG; SEQ ID NO: 42) to promote efficient transcription by T7 polymerase in vitro. See, for example, WO 2014 / 065596, which is incorporated herein by reference in its entirety for all purposes. Other guide RNA target sequences plus PAMs can have SEQ ID NOs: 40-42 with a length between 4 and 22 nucleotides, including a 5' G or GG and a 3' GG or NGG. Still other guide RNA target sequences plus PAMs can have SEQ ID NOs: 40-42 with a length between 14 and 20 nucleotides.
[0310] The guide RNA targeting the RS1 gene can target the first intron of the RS1 gene or a sequence adjacent to the first intron of the RS1 gene (e.g., in the first exon or the second exon of the RS1 gene).
[0311] The formation of a CRISPR complex hybridized with the target DNA can result in cutting of one or both chains of the target DNA in or near a region corresponding to the guide RNA target sequence (i.e., the guide RNA target sequence on the non-complementary strand of the target DNA and the guide RNA on the complementary strand hybridizing therewith). For example, the cleavage site can be within the guide RNA target sequence (e.g., at a defined location relative to the PAM sequence). The "cleavage site" comprises the position of the target DNA where the Cas protein produces a single-strand break or a double-strand break. The cleavage site can be only on one strand of the double-stranded DNA (e.g., when using a nickase) or on both strands. The cleavage site can be at the same position on both chains (producing blunt ends; e.g., Cas9) or at different sites on each chain (producing staggered ends (i.e., overhangs); e.g., Cpf1). For example, staggered ends can be produced by using two Cas proteins, each of which produces a single-strand break at different cleavage sites on different chains, thereby producing a double-strand break. For example, a first nickase can produce a single-strand break on a first strand of double-stranded DNA (dsDNA), and a second nickase can produce a single-strand break on a second strand of the dsDNA, such that an overhang sequence is generated. In some cases, the guide RNA target sequence or cleavage site for the nickase on the first strand is separated from the guide RNA target sequence or cleavage site for the nickase on the second strand by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, or 1,000 base pairs.
[0312] The guide RNA of targeting RS1 gene such as people RS1 gene can target any desired positioning in RS1 gene.The guide RNA of targeting RS1 gene can target the first intron of RS1 gene or the sequence adjacent to the first intron of RS1 gene (for example, in the first exon or the second exon of RS1 gene).For example, the guide RNA target sequence can include any continuous sequence in RS1 gene.The term RS1 gene includes the genomic region that covers RS1 regulatory promoter and enhancer sequence and coding sequence.The guide RNA target sequence can include coding sequence, non-coding sequence (for example, regulatory element, such as promoter or enhancer region), or its combination.As an example, the guide RNA target sequence can include the continuous coding sequence in any RS1 coding exon.As an example, the guide RNA target sequence can be located in exon 1 of RS1 gene.As another example, the guide RNA target sequence can be located in exon 2 of RS1 gene.As another example, the guide RNA target sequence can be located in exon 3 of RS1 gene.As another example, the guide RNA target sequence can be located in exon 4 of RS1 gene. As another example, the guide RNA target sequence can be located in exon 5 of the RS1 gene. As another example, the guide RNA target sequence can be located in exon 6 of the RS1 gene. The guide RNA target sequence can also include a contiguous sequence in any RS1 intron. As an example, the guide RNA target sequence can be located in intron 1 of the RS1 gene. As another example, the guide RNA target sequence can be located in intron 2 of the RS1 gene. As another example, the guide RNA target sequence can be located in intron 3 of the RS1 gene. As another example, the guide RNA target sequence can be located in intron 4 of the RS1 gene. As another example, the guide RNA target sequence can be located in intron 5 of the RS1 gene. As another example, the guide RNA target sequence can be located in intron 6 of the RS1 gene.
[0313] Guide RNA target sequences can also be selected to minimize off-target modifications or avoid off-target effects (e.g., by avoiding two or fewer mismatches with off-target genomic sequences).
[0314] As an example, a guide RNA targeting the RS1 gene can target a guide RNA target sequence as set forth in any one of SEQ ID NOs: 54-3147. As another example, a guide RNA targeting the RS1 gene can target at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a guide RNA target sequence as set forth in any one of SEQ ID NOs: 54-3147. Examples of such guide RNA target sequences are shown in Tables 2 and 3.
[0315] As an example, a guide RNA targeting a human RS1 gene can target a guide RNA target sequence as set forth in any one of SEQ ID NOs: 54 to 1895. As another example, a guide RNA targeting a human RS1 gene can target at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a guide RNA target sequence as set forth in any one of SEQ ID NOs: 54 to 1895.
[0316] As an example, a guide RNA targeting a human RS1 gene can target a guide RNA target sequence set forth in any one of SEQ ID NOs: 54-57, 60-92, 94-153, and 155-1257. As another example, a guide RNA targeting a human RS1 gene can target at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of a guide RNA target sequence set forth in any one of SEQ ID NOs: 54-57, 60-92, 94-153, and 155-1257.
[0317] As an example, a guide RNA targeting the human RS1 gene can target the guide RNA target sequence shown in any one of SEQ ID NOs: 56, 57, 65, 581, 1203, and 1210. As another example, a guide RNA targeting the human RS1 gene can target at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the guide RNA target sequence shown in any one of SEQ ID NOs: 56, 57, 65, 581, 1203, and 1210.
[0318] As an example, a guide RNA targeting the mouse Rs1 gene can target the guide RNA target sequence set forth in any one of SEQ ID NOs: 1896-3147 (e.g., SEQ ID NOs: 2383 or 2887). As another example, a guide RNA targeting the mouse Rs1 gene can target at least 17, at least 18, at least 19, or at least 20 consecutive nucleotides of the guide RNA target sequence set forth in any one of SEQ ID NOs: 1896-3147 (e.g., SEQ ID NOs: 2383 or 2887).
[0319] Table 2. Human RS1 intron 1 guide RNA target and guide sequences.
[0320]
[0321]
[0322]
[0323]
[0324]
[0325]
[0326]
[0327]
[0328]
[0329]
[0330]
[0331]
[0332]
[0333]
[0334]
[0335]
[0336]
[0337]
[0338]
[0339]
[0340]
[0341]
[0342]
[0343]
[0344]
[0345]
[0346]
[0347]
[0348]
[0349]
[0350]
[0351]
[0352]
[0353]
[0354]
[0355]
[0356]
[0357]
[0358]
[0359]
[0360]
[0361]
[0362]
[0363] Table 3. Mouse Rs1 intron 1 guide RNA target and guide sequences.
[0364]
[0365]
[0366]
[0367]
[0368]
[0369]
[0370]
[0371]
[0372]
[0373]
[0374]
[0375]
[0376]
[0377]
[0378]
[0379]
[0380]
[0381]
[0382]
[0383]
[0384]
[0385]
[0386]
[0387]
[0388]
[0389]
[0390]
[0391]
[0392]
[0393]
[0394] B. Other Nuclease Agents and Target Sequences of Nuclease Agents
[0395] Any nuclease agent that induces nicks or double-strand breaks at the desired target sequence can be used in the methods and compositions disclosed herein. Naturally occurring or natural nuclease agents can be used, as long as the nuclease agent induces nicks or double-strand breaks at the desired target sequence. Alternatively, modified or engineered nuclease agents can be used. "Engineering nuclease agents" include nucleases that are engineered from their native form (modified or derived from the native form) to specifically recognize and induce nicks or double-strand breaks in the desired target sequence. Therefore, engineered nuclease agents can be derived from natural, naturally occurring nuclease agents, or can be artificially produced or synthesized. Engineered nucleases can induce nicks or double-strand breaks in, for example, a target sequence, wherein the target sequence is not a sequence recognized by a natural (non-engineered or unmodified) nuclease agent. The modification of a nuclease agent can be only an amino acid in a protein cleavage agent or a nucleotide in a nucleic acid cleavage agent. Creating a nick or double-strand break at a target sequence or other DNA may be referred to herein as "cutting" or "cleaving" the target sequence or other DNA.
[0396] Also provided are active variants and fragments of exemplary target sequences. Such active variants may include having at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with a given target sequence, wherein the active variant retains biological activity and therefore can be recognized and cut by a nuclease agent in a sequence-specific manner. The determination of target sequence double-strand breaks by a nuclease agent is well known. See, for example, Frendewey et al. (2010), Methods in Enzymology 476: 295-307, which is incorporated herein by reference in its entirety for all purposes.
[0397] The target sequence of a nuclease agent can be located anywhere in or near the target locus. The target sequence can be located within the coding region of a gene, or within a regulatory region that affects gene expression. The target sequence of a nuclease agent can be located within an intron, exon, promoter, enhancer, regulatory region, or any non-protein coding region.
[0398] One type of nuclease agent is a transcription activator-like effector nuclease (TALEN). TAL effector nucleases are a class of sequence-specific nucleases that can be used to make double-strand breaks at specific target sequences in prokaryotic or eukaryotic genomes. TAL effector nucleases are produced by fusing natural or engineered transcription activator-like (TAL) effectors or functional portions thereof with the catalytic domain of, for example, an endonuclease. The unique modular TAL effector DNA binding domain allows the design of proteins with potential recognition specificity for any given DNA. Therefore, the DNA binding domain of the TAL effector nuclease can be engineered to recognize specific DNA target sites and, therefore, be used to make double-strand breaks at the desired target sequence. See WO 2010 / 079430; Morbitzer et al. (2010) Proceedings of the National Academy of Sciences of the United States of America 107(50):21617-21622; Scholze and Boch (2010) Virulence 1:428-432; Christian et al., Genetics (2010) 186:757-761; Li et al. (2010) Nucleic Acids Res. (2010) doi:10.1093 / nar / gkq704; and Miller et al. (2011) Nature Biotechnology 29:143-148, each of which is incorporated herein by reference in its entirety for all purposes.
[0399] Examples of suitable TAL nucleases and methods for preparing suitable TAL nucleases are disclosed in, for example, US 2011 / 0239315 A1, US 2011 / 0269234 A1, US 2011 / 0145940 A1, US 2003 / 0232410 A1, US 2005 / 0208489 A1, US 2005 / 0026157 A1, US 2005 / 0064474 A1, US 2006 / 0188987 A1, and US 2006 / 0063231 A1, each of which is incorporated herein by reference in its entirety for all purposes. In various embodiments, a TAL effector nuclease is engineered that cuts in or near a target nucleic acid sequence, for example, in a locus of interest or a genomic locus of interest, wherein the target nucleic acid sequence is at or near a sequence to be modified by a targeting vector. TAL nucleases suitable for use with the various methods and compositions provided herein include those specifically designed to bind at or near a target nucleic acid sequence to be modified by a targeting vector as described herein.
[0400] In some TALENs, each monomer of the TALEN includes 33-35 TAL repeats that recognize a single base pair through two hypervariable residues. In some TALENs, the nuclease agent is a chimeric protein that includes a TAL repeat-based DNA binding domain operably linked to an independent nuclease such as the FokI endonuclease. For example, the nuclease agent can include a first TAL repeat-based DNA binding domain and a second TAL repeat-based DNA binding domain, wherein each of the first and second TAL repeat-based DNA binding domains is operably linked to a FokI nuclease, wherein the first and second TAL repeat-based DNA binding domains recognize two consecutive target DNA sequences in each chain of a target DNA sequence separated by spacer sequences of different lengths (12-20bp), and wherein the FokI nuclease subunits dimerize to produce an active nuclease that makes a double-strand break on the target sequence.
[0401] The nuclease agent used in the various methods and compositions disclosed herein may further include a zinc finger nuclease (ZFN). In some ZFNs, each monomer of the ZFN includes three or more zinc finger-based DNA binding domains, wherein each zinc finger-based DNA binding domain binds to a 3bp subsite. In other ZFNs, the ZFN is a chimeric protein including a zinc finger-based DNA binding domain that is operably linked to an independent nuclease such as a FokI endonuclease. For example, a nuclease agent may include a first ZFN and a second ZFN, wherein each of the first ZFN and the second ZFN is operably linked to a FokI nuclease subunit, wherein the first and second ZFNs recognize two consecutive target DNA sequences in each chain of a target DNA sequence separated by an approximately 5-7bp spacer, and wherein the FokI nuclease subunits dimerize to produce an active nuclease that breaks the double strand. See, e.g., US20060246567; US20080182332; US20020081614; US20030021776; WO / 2002 / 057308A2; US20130123484; US20100291048; WO / 2011 / 017293A2; and Gaj et al. (2013) Trends Biotechnol. 31(7):397-405, each of which is incorporated herein by reference in its entirety for all purposes.
[0402] Also provided are active variants and fragments of nuclease agents (i.e., engineered nuclease agents). Such active variants can include and have at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with natural nuclease agents, wherein active variants retain the ability to cut at the desired target sequence and therefore retain nick or double-strand break inducing activity. For example, any nuclease agent in the nuclease agent described herein can be modified from the natural endonuclease sequence and designed to recognize and induce nick or double-strand break at the target sequence that is not recognized by the natural nuclease agent. Therefore, some engineered nucleases have specificity to induce nick or double-strand break at the target sequence that is different from the corresponding natural nuclease agent target sequence. The determination of nick or double-strand break inducing activity is known, and the overall activity and specificity of endonuclease to the DNA substrate containing the target sequence are usually measured.
[0403] Nuclease agents can be introduced into cells or animals by any known means. Polypeptides encoding nuclease agents can be directly introduced into cells or animals. Alternatively, polynucleotides encoding nuclease agents can be introduced into cells or animals. When polynucleotides encoding nuclease agents are introduced, the nuclease agents can be expressed transiently, conditionally or constitutively in the cell. The polynucleotides encoding nuclease agents can be contained in an expression cassette and operably connected to a conditional promoter, an inducible promoter, a constitutive promoter or a tissue-specific promoter. Examples of promoters are discussed in further detail elsewhere herein. Alternatively, the nuclease agent can be introduced into the cell as an mRNA encoding the nuclease agent.
[0404] The polynucleotide encoding the nuclease agent can be stably integrated into the genome of the cell and operably linked to a promoter active in the cell. Alternatively, the polynucleotide encoding the nuclease agent can be in an expression vector or a targeting vector.
[0405] When providing a nuclease agent to a cell by introducing a polynucleotide encoding the nuclease agent, the polynucleotide encoding the nuclease agent can be modified to substitute codons that are used more frequently in the cell of interest compared to the naturally occurring polynucleotide sequence encoding the nuclease agent. For example, the polynucleotide encoding the nuclease agent can be modified to substitute codons that are used more frequently in a given eukaryotic cell of interest, including human cells, non-human cells, mammalian cells, rodent cells, mouse cells, rat cells, or any other host cell of interest, compared to the naturally occurring polynucleotide sequence.
[0406] The term "target sequence of a nuclease agent" comprises a DNA sequence in which a nuclease agent induces a nick or double-strand break. The target sequence of a nuclease agent can be endogenous (or natural) to the cell, or the target sequence can be exogenous to the cell. A target sequence exogenous to the cell is not naturally present in the genome of the cell. The target sequence can also be exogenous to the polynucleotide of interest that is desired to be located at the target locus. In some cases, the target sequence is present only once in the genome of the host cell.
[0407] The length of the target sequence can vary and include, for example, a target sequence of about 30-36 bp for a zinc finger nuclease (ZFN) pair (i.e., about 15-18 bp for each ZFN), about 36 bp for a transcription activator-like effector nuclease (TALEN), or about 20 bp for a CRISPR / Cas9 guide RNA.
[0408] VI. Cells or Animals or Genomes Comprising Nucleic Acid Constructs and / or Nuclease Agents or Nucleic Acids Encoding Nuclease Agents
[0409] Also provided are genomes, cells and animals produced by methods disclosed herein.Similarly, also provided are genomes, cells and animals comprising nucleic acid constructs, carriers, lipid nanoparticles or compositions as described herein, wherein the nucleic acid constructs include retinoschizine coding sequences (i.e., encoding retinoschizine protein or its fragment or variant) for being integrated into the target genome locus and expressed from the target genome locus.Similarly, also provided are nucleic acids (e.g., targeting endogenous RS1 locus) comprising described nuclease agents or encoding nuclease agents, or genomes, cells and animals comprising carriers, lipid nanoparticles or compositions as described herein. The genomes, cells or animals may include nucleic acid constructs that may be integrated into the target genome locus (e.g., RS1 locus) on the genome, and may express retinoschizine protein or its fragment or variant. When being integrated into the target genome locus, the retinoschizine coding sequence may be operably connected to the endogenous promoter at the target genome locus, or it may be operably connected to the exogenous promoter present in the nucleic acid construct. If the nucleic acid construct is a bidirectional nucleic acid construct disclosed herein, the genome, cell or animal can express a first retinoschizin protein or a fragment or variant thereof, or can express a second retinoschizin protein or a fragment or variant thereof. In some genomes, cells or animals, the target genomic locus is an RS1 locus. For example, the nucleic acid construct can be integrated into the intron 1 of the endogenous RS1 locus on the genome. Endogenous RS1 exon 1 can then be spliced into the coding sequence of the retinoschizin protein or its fragment or variant in the nucleic acid construct. In a specific example, the modified RS1 locus encoding of the nucleic acid construct integrated on the genome includes a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 2 or 4, is essentially composed of, or is composed of a protein. In specific examples, the modified RS1 locus comprising a genomically integrated nucleic acid construct comprises an RS1 coding sequence comprising, consisting essentially of, or consisting of a sequence that is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% identical to SEQ ID NO: 6, 11 or 12.
[0410] In some genomes, cells or animals, the nucleic acid construct is integrated into the endogenous RS1 locus to prevent the endogenous RS1 gene transcription downstream of the integration site. For example, the nucleic acid construct is integrated into the endogenous RS1 locus to reduce or eliminate the expression of endogenous retinoschizin protein, and the expression of the retinoschizin protein or its fragment or variant encoded by the nucleic acid construct is substituted for the expression of the endogenous retinoschizin protein. In one example, the nucleic acid construct is integrated into the endogenous RS1 locus to reduce the expression of endogenous retinoschizin protein. In another example, the nucleic acid construct is integrated into the endogenous RS1 locus to eliminate the expression of endogenous retinoschizin protein. In a specific example, the endogenous RS1 locus includes the RS1 gene of sudden change, the RS1 gene of sudden change includes the sudden change that causes X-linked juvenile retinoschizin, and the expression of the nucleic acid construct integrated on the genome reduces or eliminates the expression of the RS1 gene of the sudden change.
[0411] The target genomic locus that the nucleic acid construct stably integrates can be a heterozygote of the retinoschizine coding sequence from the nucleic acid construct or a homozygote of the retinoschizine coding sequence from the nucleic acid construct. Diploid organisms have two alleles at each locus. Every pair of alleles represents the genotype of a specific locus. If there are two identical alleles at a specific locus, the genotype is described as homozygous, and if the two alleles are different, the genotype is described as heterozygous. Animals comprising a nucleic acid construct integrated on the genome as described herein can include the nucleic acid construct in the target genomic locus of their germline.
[0412] The genome, cell or animal provided herein can be, for example, a eukaryotic organism, including, for example, animals, mammals, non-human mammals and humans. The term "animal" includes mammals, fish and birds. Mammals can be, for example, non-human mammals, humans, rodents, rats, mice or hamsters. Other non-human mammals include, for example, non-human primates, monkeys, apes, cats, dogs, rabbits, horses, bulls, deer, bison, livestock (for example, cattle species, such as cows, bulls, etc.; sheep species, such as sheep, goats, etc.; and pig species, such as pigs and wild boars). Birds include, for example, chickens, turkeys, ostriches, geese, ducks, etc. Domestic animals and agricultural animals are also included. The term "non-human" does not include humans.
[0413] The cell can be an isolated cell (e.g., in vitro) or can be in an animal body. The cell can also be in any type of undifferentiated or differentiated state. For example, the cell can be a totipotent cell, a pluripotent cell (e.g., a human pluripotent cell or a non-human pluripotent cell, such as a mouse embryonic stem (ES) cell or a rat ES cell) or a non-pluripotent cell. The totipotent cell comprises an undifferentiated cell that can produce any cell type, and the pluripotent cell comprises an undifferentiated cell with the ability to develop into more than one differentiated cell type.
[0414] The cell provided herein can also be a germ cell (for example, sperm or oocyte). The cell can be a mitotic competent cell or a mitotic inactive cell, a meiotic competent cell or a meiotic inactive cell. Similarly, the cell can also be a primary somatic cell or a cell that is not a primary somatic cell. Somatic cells comprise any cell that is not a gamete, germ cell, gametocyte or undifferentiated stem cell. For example, the cell can be a hepatocyte, a nephrocyte, a hematopoietic cell, an endothelial cell, an epithelial cell, a fibroblast, a mesenchymal cell, a keratinocyte, a hemocyte, a melanocyte, a monocyte, a mononuclear cell, a mononuclear cell precursor, a B cell, an erythrocyte-megakaryocyte, an eosinophil, a macrophage, a T cell, a pancreatic islet beta cell, an exocrine cell, a pancreatic progenitor cell, an endocrine progenitor cell, a fat cell, a preadipocyte, a neuron, a glial cell, a neural stem cell, a neuron, a hepatoblast, a hepatocyte, a cardiomyocyte, a bone marrow cell, a my ... The cell may be a skeletal muscle cell, a smooth muscle cell, a ductal cell, an alveolar cell, an α cell, a β cell, a δ cell, a PP cell, a bile duct cell, a white or brown fat cell, or an eye cell (e.g., a trabecular meshwork cell, a retinal pigment epithelial cell, a retinal microvascular endothelial cell, a retinal pericyte, a conjunctival epithelial cell, a conjunctival fibroblast, an iris pigment epithelial cell, a keratocyte, a lens epithelial cell, a non-pigmented ciliary epithelial cell, an eye choroidal fibroblast, a photoreceptor cell, a ganglion cell, a bipolar cell, a horizontal cell, or amacrine cell). For example, the cell may be an eye cell, such as a retinal cell (e.g., a photoreceptor).
[0415] The cell provided herein can be a normal, healthy cell, or can be a cell that is diseased or carries a mutant. For example, the cell can include one or more mutations associated with or causing XLRS (e.g., R141C substitutions in the encoding retinoschizin protein).
[0416] The animals provided herein can be humans or non-human animals. Non-human animals comprising nucleic acids or expression cassettes as described herein can be prepared by methods described elsewhere herein. The term "animal" includes mammals, fish, and birds. Mammals include, for example, humans, non-human primates, monkeys, apes, cats, dogs, horses, bulls, deer, bison, sheep, rabbits, rodents (e.g., mice, rats, hamsters, and guinea pigs) and livestock (e.g., cattle species, such as cows and bulls; sheep species, such as sheep and goats; and pig species, such as pigs and wild boars). Birds include, for example, chickens, turkeys, ostriches, geese, and ducks. Domestic animals and agricultural animals are also included. The term "non-human animal" does not include humans. Specific examples of non-human animals include rodents, such as mice and rats.
[0417] In some embodiments, non-human animals can be from any genetic background. For example, suitable mice can be from 129 strains, C57BL / 6 strains, a mixture of 129 and C57BL / 6, BALB / c strains or Swiss Webster strains. The example of 129 strains includes 129P1, 129P2, 129P3, 129X1, 129S1 (for example, 129S1 / SV, 129S1 / Svlm), 129S2, 129S4, 129S5, 129S9 / SvEvH, 129S6 (129 / SvEvTac), 129S7, 129S8, 129T1 and 129T2. See, for example, Festing et al. (1999), Mammalian Genome 10:836, which is incorporated herein by reference in its entirety for all purposes. Examples of C57BL strains include C57BL / A, C57BL / An, C57BL / GrFa, C57BL / Kal_wN, C57BL / 6, C57BL / 6J, C57BL / 6ByJ, C57BL / 6NJ, C57BL / 10, C57BL / 10ScSn, C57BL / 10Cr, and C57BL / Ola. Suitable mice can also be derived from a mixture of the above-mentioned 129 strain and the above-mentioned C57BL / 6 strain (e.g., 50% 129 and 50% C57BL / 6). Similarly, suitable mice can be derived from a mixture of the above-mentioned 129 strain or a mixture of the above-mentioned BL / 6 strain (e.g., 129S6 (129 / SvEvTac) strain).
[0418] Similarly, rats can be from any rat strain, including, for example, the ACI rat strain, the black spiny (DA) rat strain, the Wistar rat strain, the LEA rat strain, the Sprague Dawley (SD) rat strain, or the Fischer rat strain, such as the Fischer F344 or Fischer F6. Rats can also be obtained from mixed strains derived from two or more of the above strains. For example, suitable rats can be from the DA strain or the ACI strain. The AC...
Claims
1. A composition for expressing retinoschizine in a cell or for integrating a coding sequence for a retinoschizine protein or a fragment thereof into a target genomic locus in a cell, the composition comprising: (a) a nucleic acid construct comprising the coding sequence of the retinoschizine protein or a fragment thereof for integration into the target genomic locus, wherein the coding sequence comprises exons 2-6 of human RS1 or a degenerate variant thereof; and (b) a nuclease agent or a nucleic acid encoding the nuclease agent, wherein the nuclease agent targets a nuclease target sequence in the target genomic locus, wherein the target genomic locus is in the endogenous RS1 locus, wherein the nucleic acid construct is integrated into the endogenous RS1 locus to prevent transcription of the endogenous RS1 gene downstream of the integration site, wherein integrating the nucleic acid construct into the endogenous RS1 locus in the cell reduces or eliminates expression of endogenous retinoschizine protein and replaces expression of the endogenous retinoschizine protein with expression of the retinoschizine protein or a fragment thereof encoded by the nucleic acid construct, and wherein the nucleic acid construct is bidirectional and comprises: (I) a first segment comprising a first coding sequence for a first retinoschizine protein or a fragment thereof, a first polyadenylation signal sequence located at the 3' end of the first coding sequence, and a first splice acceptor site located at the 5' end of the first coding sequence; and (II) a second segment comprising the reverse complement of a second coding sequence for a second retinoschizine protein or fragment thereof, the reverse complement of a second polyadenylation signal sequence located 5' to the reverse complement of the second coding sequence, and the reverse complement of a second splice acceptor site located 3' to the reverse complement of the second coding sequence, The first splice acceptor site is from intron 1 of RS1 gene, The second splice acceptor site is from intron 1 of RS1 gene, wherein the second segment is located at the 3' end of the first segment, The two ends of the nucleic acid construct are flanked by inverted terminal repeats (ITRs), wherein the nucleic acid construct does not include homology arms, wherein the nucleic acid construct does not include a promoter driving expression of the first retinoschizine protein or a fragment thereof or the second retinoschizine protein or a fragment thereof, wherein the first retinoschizin protein or fragment thereof is identical to the second retinoschizin protein or fragment thereof, and wherein the second coding sequence employs a codon usage that is different from the codon usage of the first coding sequence, and The complementarity of the second coding sequence to the first coding sequence is less than 85% to reduce hairpin formation. 2 . The composition according to claim 1 , wherein the nucleic acid construct comprises a fragment or portion of the first intron of human RS1 located 5′ of the coding sequence.
3. The composition according to claim 1 or 2, wherein the first splice acceptor site is from intron 1 of human RS1.
4. The composition of claim 1 or 2, wherein the nuclease target sequence in the target genomic locus is in the first intron in the RS1 gene of the endogenous RS1 locus.
5. The composition of claim 1 or 2, wherein the nucleic acid construct is in a viral vector.
6. The composition of claim 5, wherein the viral vector is an adeno-associated virus (AAV) viral vector.
7. The composition of claim 6, wherein the AAV is selected from the group consisting of: AAV2, AAV5, AAV8, or AAV7m8.
8. The composition of any one of claims 1, 2, 6, and 7, wherein the nuclease agent is a Cas protein and a guide RNA, and the nuclease target sequence is a guide RNA target sequence.
9. The composition of claim 8, wherein the Cas protein is a Cas9 protein.
10. The composition of claim 9, wherein the composition comprises a nucleic acid encoding the Cas protein, wherein the nucleic acid comprises a DNA encoding the Cas protein, and wherein the composition comprises a DNA encoding the guide RNA.
11. The composition of claim 10, wherein the DNA encoding the Cas protein and the DNA encoding the guide RNA are in one or more viral vectors.
12. The composition of claim 9, wherein the composition comprises a nucleic acid encoding the Cas protein, wherein the nucleic acid comprises a messenger RNA encoding the Cas protein, and wherein the composition comprises the guide RNA in the form of RNA.
13. The composition of claim 12, wherein the composition comprises the guide RNA in the form of RNA, and wherein the guide RNA and the messenger RNA encoding the Cas protein are in lipid nanoparticles.
14. The composition of any one of claims 1, 2, 6, 7, and 9-13, wherein the endogenous RS1 locus comprises a mutated RS1 gene comprising a mutation that causes X-linked juvenile retinoschisis.
15. The composition of any one of claims 1, 2, 6, 7, and 9-13, wherein the cells are human cells.
16. The composition of claim 15, wherein the human cells are retinal cells.
17. The composition of any one of claims 1, 2, 6, 7, 9-13, and 16, wherein the cell is in an animal.
18. The composition of claim 17, wherein integration of the nucleic acid construct results in restoration of retinal structure.
Citation Information
Patent Citations
Functional genomics using zinc finger proteins
US20020081614A1
Regulation of angiogenesis with zinc finger proteins
US20030021776A1
Methods and compositions for using zinc finger endonucleases to enhance homologous recombination
US20030232410A1
Use of chimeric nucleases to stimulate gene targeting
US20050026157A1
Methods and compositions for targeted cleavage and recombination
US20050064474A1