Polynucleotides, compositions, and methods for genome editing
Patent Information
- Application Number
- JP2024525233
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-14
- Filing Date
- 2022-11-02
- Publication Date
- 2025-11-12
AI Technical Summary
Existing RNA-guided DNA binding agents like CRISPR-Cas systems face challenges in achieving robust expression levels and reducing immunogenicity, particularly in mammalian cells, leading to undesirable cytokine elevation.
The development of polynucleotides encoding NmeCas9 polypeptides with optimized coding sequences, including nuclear localization signals (NLS) and linker sequences, along with modified uridines, to enhance expression and reduce immunogenicity, specifically tailored for mammalian organs like the liver.
The optimized polynucleotides provide improved expression levels and reduced immunogenicity of NmeCas9, enhancing genome editing efficiency and specificity in mammalian cells.
Abstract
Description
[Technical field]
[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 275,425, filed November 3, 2021, and U.S. Provisional Patent Application No. 63 / 352,158, filed June 14, 2022, the contents of both applications being incorporated by reference in their entireties.
[0002] This application contains a Sequence Listing that has been submitted electronically in ASCII format, which is incorporated herein by reference in its entirety. The name of this ASCII copy, created on November 1, 2022, is 01155-0049-00PCT_ST26 and is 11,070,539 bytes in size.
[0003] The present disclosure relates to polynucleotides, compositions, and methods for genome editing, including RNA-guided DNA binding agents such as the CRISPR-Cas system and subunits thereof. [Background technology]
[0004] RNA-guided DNA binding agents, such as the CRISPR-Cas system, can be used for targeted genome editing, including in eukaryotic cells and in vivo. Such editing has been shown to be able to inactivate certain deleterious alleles or correct certain deleterious point mutations. For example, Neisseria meningitidis Cas9 (NmeCas9) advantageously has a low off-target cleavage rate. RNA-guided DNA binding agents can be produced in the system by contacting cells with polynucleotides, such as mRNA or expression constructs. However, existing approaches may provide less robust expression than desired in certain cell types or organisms, such as mammals, or may be undesirably immunogenic, e.g., inducing undesirable increases in cytokine levels.
[0005] Thus, there is a need for polynucleotides, compositions, and methods for the expression of polypeptides such as NmeCas9. The present disclosure aims to provide compositions and methods for polypeptide expression that provide one or more advantages, such as at least one of improved expression levels, increased activity of the encoded polypeptide, or reduced immunogenicity (e.g., reduced cytokine rise upon administration), or at least provide the public with a useful choice. In some embodiments, a polynucleotide encoding an RNA-guided DNA binding agent (e.g., NmeCas9) is provided, which differs from existing polynucleotides in one or more of its coding sequence, codon usage, non-coding sequence (e.g., UTR), or heterologous domain (e.g., NLS) in a manner disclosed herein. It has been found that such characteristics can provide advantages such as those described above. In some embodiments, the improved editing efficiency occurs in or is specific to a mammalian organ or cell type, such as the liver or hepatocytes. Summary of the Invention
[0006] The following embodiments are provided by the present disclosure.
[0007] In some embodiments, a polynucleotide is provided, the polynucleotide comprising an open reading frame (ORF), the ORF comprising a nucleotide sequence encoding a C-terminal Neisseria meningitidis (Nme) Cas9 polypeptide that is at least 90% identical to any one of SEQ ID NOs: 29, 32-41, 224-226, 231-233, 238-240, 245-247, 252-254, 259-261, 266-268, 273-275, 280-282, 287-289, 294-296, or 301-303, and 317-321, wherein the Nme Cas9 is Nme2 Cas9, Nme1 Cas9, or Nme3 Cas9, and comprises a nucleotide sequence encoding a first nuclear localization signal (NLS).
[0008] In some embodiments, the ORF further comprises a nucleotide sequence encoding a second NLS. In some embodiments, the first and second NLS are independently selected from SEQ ID NOs: 388 and 410-422. In some embodiments, the polynucleotide further comprises a polyA sequence or a polyadenylation signal sequence. In some embodiments, the ORF further comprises a nucleotide sequence encoding a linker sequence between the first NLS and the second NLS. In some embodiments, the ORF further comprises a nucleotide sequence encoding a linker spacer sequence between the NmeCas9 coding sequence and the NLS adjacent to the NmeCas9 coding sequence. In some embodiments, the ORF NmeCas9 has double-stranded endonuclease activity. In some embodiments, the ORF NmeCas9 has nickase activity. In some embodiments, the ORF NmeCas9 comprises a dCas9 DNA binding domain.
[0009] The following numbered embodiments provide additional support and explanation for the embodiments herein.
[0010] Embodiment 1 is a polynucleotide comprising an open reading frame (ORF), the ORF comprising a nucleotide sequence encoding a C-terminal Neisseria meningitidis (Nme) Cas9 polypeptide at least 90% identical to any one of SEQ ID NOs: 29, 32-41, 224-226, 231-233, 238-240, 245-247, 252-254, 259-261, 266-268, 273-275, 280-282, 287-289, 294-296, 301-303, or 316-321, wherein Nme Cas9 is Nme2 Cas9, Nme1 Cas9, or Nme3 Cas9, and a nucleotide sequence encoding a first nuclear localization signal (NLS).
[0011] Embodiment 2 is the polynucleotide of embodiment 1, wherein the ORF further comprises a nucleotide sequence encoding a second NLS.
[0012] Embodiment 3 is the polynucleotide of embodiment 1, wherein the first and second NLS are independently selected from SEQ ID NOs: 388 and 410-422.
[0013] Embodiment 4 is the polynucleotide according to any one of embodiments 1 to 3, wherein the polynucleotide further comprises a polyA sequence or a polyadenylation signal sequence.
[0014] Embodiment 5 is the polynucleotide of embodiment 4, wherein the polyA sequence comprises non-adenine nucleotides.
[0015] Embodiment 6 is the polynucleotide according to embodiment 4 or 5, wherein the polyA sequence comprises 100 to 400 nucleotides.
[0016] Embodiment 7 is the polynucleotide according to any one of embodiments 4 to 6, wherein the polyA sequence comprises the sequence of SEQ ID NO:409.
[0017] Embodiment 8 is a polynucleotide according to any one of embodiments 1 to 7, wherein the ORF further comprises a nucleotide sequence encoding a linker sequence between the first NLS and the second NLS.
[0018] Embodiment 9 is a polynucleotide according to any one of embodiments 1 to 8, wherein the ORF further comprises a nucleotide sequence encoding a linker spacer sequence between the Nme Cas9 coding sequence and the NLS adjacent to the Nme Cas9 coding sequence.
[0019] Embodiment 10 is the polynucleotide of embodiment 8 or 9, wherein the linker comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 amino acids.
[0020] Embodiment 11 is a polynucleotide according to any one of embodiments 8 to 10, wherein the linker sequence comprises GGG or GGGS, and optionally, the GGG or GGGS sequence is at the N-terminus of the spacer sequence.
[0021] Embodiment 12 is the polynucleotide according to any one of embodiments 8 to 11, wherein the linker sequence comprises any one of SEQ ID NOs: 61 to 122.
[0022] Embodiment 13 is a polynucleotide according to any one of embodiments 1 to 12, wherein the ORF further comprises one or more additional heterologous functional domains.
[0023] Embodiment 14 is a polynucleotide according to any one of embodiments 1 to 13, wherein Nme Cas9 has double-stranded endonuclease activity.
[0024] Embodiment 15 is a polynucleotide according to any one of embodiments 1 to 14, wherein Nme Cas9 has nickase activity.
[0025] Embodiment 16 is a polynucleotide according to any one of embodiments 1 to 14, wherein the Nme Cas9 comprises a dCas9 DNA binding domain.
[0026] Embodiment 17 is a polynucleotide described in any one of embodiments 1 to 16, wherein NmeCas9 comprises at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of any one of SEQ ID NOs: 1, 4 to 13, 220, 227, 234, 241, 248, 255, 262, 269, 276, 283, 290, 297, or 310 to 315.
[0027] Embodiment 18 is a polynucleotide according to any one of embodiments 1 to 17, wherein NmeCas9 comprises an amino acid sequence of SEQ ID NO: 1 and 4 to 13, 220, 227, 234, 241, 248, 255, 262, 269, 276, 283, 290, 297, or 310 to 315.
[0028] Embodiment 19 is a polynucleotide according to any one of embodiments 1 to 18, wherein the sequence encoding NmeCas9 comprises a nucleotide sequence having at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to any one of the nucleotide sequences of SEQ ID NOs: 15, 18-27, 29, 32-41, 221-226, 228-233, 235-240, 242-247, 249-254, 256-261, 263-268, 270-275, 277-282, 284-289, 291-296, 298-303, 304-309, or 316-321.
[0029] Embodiment 20 is a polynucleotide according to any one of embodiments 1 to 19, wherein the sequence encoding NmeCas9 comprises any one of the nucleotide sequences of SEQ ID NOs: 15, 18 to 27, 29, 32 to 41, 221 to 226, 228 to 233, 235 to 240, 242 to 247, 249 to 254, 256 to 261, 263 to 268, 270 to 275, 277 to 282, 284 to 289, 291 to 296, 298 to 303, 304 to 309, or 316 to 321.
[0030]
[0036] Embodiment 21 is a polynucleotide comprising an open reading frame (ORF) encoding a polypeptide comprising a cytidine deaminase, which is optionally an APOBEC3A deaminase, and a nucleotide sequence encoding a C-terminal meningococcal Cas9 nickase polypeptide at least 90% identical to any one of SEQ ID NOs: 1, 4-13, 220, 227, 234, 241, 248, 255, 262, 269, 276, 283, 290, or 297, wherein the Nme Cas9 nickase is an Nme2 Cas9 nickase, an Nme1 Cas9 nickase, or an Nme3 Cas9 nickase, and a nucleotide sequence encoding a first nuclear localization signal (NLS), wherein the polypeptide does not comprise a uracil glycosylase inhibitor (UGI).
[0031] Embodiment 22 is the polynucleotide of embodiment 21, wherein the ORF further comprises a nucleotide sequence encoding a second NLS.
[0032] Embodiment 23 is the polynucleotide of embodiment 21 or 22, wherein the deaminase is located N-terminal to the NLS in the polypeptide.
[0033] Embodiment 24 is the polynucleotide of any one of embodiments 21 to 23, wherein the cytidine deaminase is located N-terminal to the first and second NLSs in the polypeptide.
[0034] Embodiment 25 is the polynucleotide of embodiment 21 or 22, wherein the cytidine deaminase is located C-terminal to the NLS in the polypeptide.
[0035] Embodiment 26 is the polynucleotide of any one of embodiments 23 to 25, wherein the cytidine deaminase is located C-terminal to the first and second NLSs in the polypeptide.
[0036] Embodiment 27 is a polynucleotide according to any one of embodiments 21 to 26, wherein the ORF does not comprise a coding sequence for an NLS C-terminal to the ORF encoding Nme Cas9.
[0037] Embodiment 28 is a polynucleotide according to any one of embodiments 21 to 26, wherein the ORF does not contain a coding sequence C-terminal to the ORF encoding Nme Cas9.
[0038] Embodiment 29 is the polynucleotide of any one of embodiments 1 to 28, wherein the cytidine deaminase comprises an amino acid sequence having at least 87% identity to SEQ ID NO:151.
[0039] Embodiment 30 is the polynucleotide according to any one of embodiments 1 to 28, wherein the cytidine deaminase comprises an amino acid sequence having at least 80% identity to SEQ ID NOs: 152 to 216.
[0040] Embodiment 31 is the polynucleotide according to any one of embodiments 1 to 28, wherein the cytidine deaminase comprises an amino acid sequence having at least 80% identity to SEQ ID NO:14.
[0041] Embodiment 32 is the polynucleotide of any one of embodiments 1 to 31, wherein the ORF comprises a nucleotide sequence at least 80% identical to SEQ ID NO:42.
[0042] Embodiment 33 is the polynucleotide of any one of embodiments 1 to 32, wherein the polynucleotide comprises a 5'UTR having at least 90% identity to any one of SEQ ID NOs: 391 to 398.
[0043] Embodiment 34 is the polynucleotide according to any one of embodiments 1 to 33, wherein the polynucleotide comprises a 5'UTR comprising any one of SEQ ID NOs: 391 to 398.
[0044] Embodiment 35 is the polynucleotide of any one of embodiments 1 to 34, wherein the polynucleotide comprises a 3'UTR having at least 90% identity to any one of SEQ ID NOs: 399 to 406.
[0045] Embodiment 36 is the polynucleotide according to any one of embodiments 1 to 35, wherein the polynucleotide comprises a 3'UTR comprising any one of SEQ ID NOs: 399 to 306.
[0046] Embodiment 37 is the polynucleotide of any one of embodiments 1 to 36, wherein the polynucleotide comprises a 5'UTR and a 3'UTR from the same source.
[0047] Embodiment 38 is a polynucleotide according to any one of embodiments 1 to 37, wherein the polynucleotide comprises a 5' cap, optionally wherein the 5' cap is Cap0, Cap1, or Cap2.
[0048] Embodiment 39 is a polynucleotide according to any one of embodiments 1 to 38, wherein at least 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons in the ORF are minimal adenine codons or minimal uridine codons.
[0049] Embodiment 40 is a polynucleotide according to any one of embodiments 1 to 39, wherein the ORF comprises or consists of codons that increase translation of the mRNA in a mammal.
[0050] Embodiment 41 is a polynucleotide according to any one of embodiments 1 to 40, wherein the ORF comprises or consists of codons that increase translation of the mRNA in humans.
[0051] Embodiment 42 is the polynucleotide according to any one of embodiments 1 to 41, wherein the polynucleotide is an mRNA.
[0052] Embodiment 43 is a polynucleotide according to embodiment 42, wherein the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to any one of SEQ ID NOs: 29, 32-41, 224-226, 231-233, 238-240, 245-247, 252-254, 259-261, 266-268, 273-275, 280-282, 287-289, 294-296, 301-303, or 316-321.
[0053] Embodiment 44 is the polynucleotide of embodiment 42 or 43, wherein at least 10% of the uridines in the mRNA are substituted with modified uridines.
[0054] Embodiment 45 is the polynucleotide of embodiment 42 or 43, wherein less than 10% of the uridines in the mRNA are substituted with modified uridines.
[0055] Embodiment 46 is the polynucleotide of embodiment 45, wherein the modified uridine is one or more of N1-methyl-pseudouridine, pseudouridine, 5-methoxyuridine, or 5-iodouridine.
[0056] Embodiment 47 is the polynucleotide of embodiment 45, wherein the modified uridine is one or both of N1-methyl-pseudouridine or 5-methoxyuridine.
[0057] Embodiment 48 is the polynucleotide of any one of embodiments 45 to 47, wherein the modified uridine is N1-methyl-pseudouridine.
[0058] Embodiment 49 is the polynucleotide of any one of embodiments 45 to 47, wherein the modified uridine is 5-methoxyuridine.
[0059] Embodiment 50 is the polynucleotide according to any one of embodiments 44 and 46 to 49, wherein 15% to 45% of the uridines are substituted with modified uridines.
[0060] Embodiment 51 is the polynucleotide of embodiment 50, wherein at least 20% or at least 30% of the uridines are substituted with modified uridines.
[0061] Embodiment 52 is the polynucleotide of embodiment 51, wherein at least 80% or at least 90% of the uridines are substituted with modified uridines.
[0062] Embodiment 53 is the polynucleotide of embodiment 52, wherein 100% of the uridines are replaced with modified uridines.
[0063] Embodiment 54 is the polynucleotide of embodiment 42, wherein less than 10% of the nucleotides in the mRNA are substituted with modified nucleotides.
[0064] Embodiment 55 is a composition comprising the polynucleotide according to any one of embodiments 1 to 54 and at least one guide RNA (gRNA).
[0065]
[0023] Embodiment 56 is a composition comprising a first polynucleotide comprising a first open reading frame (ORF) encoding a polypeptide comprising a cytidine deaminase, optionally an APOBEC3A deaminase, and an NmeCas9 nickase, and a second polynucleotide comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), the second polynucleotide being different from the first polynucleotide, and optionally further comprising a guide RNA (gRNA).
[0066] Embodiment 57 is the composition of embodiment 55 or 56, wherein the gRNA is a single guide RNA.
[0067] Embodiment 58 is the composition of embodiment 55 or 56, wherein the gRNA is a dual guide RNA.
[0068] Embodiment 59 is a composition comprising the polynucleotide of any one of embodiments 1 to 57, further comprising a single guide RNA, wherein the single guide RNA comprises a guide region and a conserved region, (a) a truncated repeat / anti-repeat region, the truncated repeat / anti-repeat region lacking 2 to 24 nucleotides; (i) one or more of nucleotides 37-48 and 53-64 are deleted, and optionally one or more of nucleotides 37-64 are substituted with respect to SEQ ID NO:500; (ii) a truncated repeat / anti-repeat region in which nucleotide 36 is linked to nucleotide 65 by at least two nucleotides; or (b) a truncated hairpin 1 region, wherein truncated hairpin 1 lacks 2-10, optionally 2-8, nucleotides; (i) one or more of nucleotides 82-86 and 91-95 are deleted, and optionally one or more of positions 82-96 are substituted relative to SEQ ID NO:500; (ii) a truncated hairpin 1 region, in which nucleotide 81 is linked to nucleotide 96 by at least four nucleotides; or (c) a truncated hairpin 2 region, wherein the truncated hairpin 2 is missing 2-18, optionally 2-16, nucleotides; (i) one or more of nucleotides 113-121 and 126-134 are deleted, and optionally one or more of nucleotides 113-134 are substituted with respect to SEQ ID NO:500; (ii) a truncated hairpin 2 region, in which nucleotide 112 is linked to nucleotide 135 by at least four nucleotides; one or both nucleotides 144-145 are optionally deleted relative to SEQ ID NO:500, The composition wherein at least 10 nucleotides are modified nucleotides.
[0069] Embodiment 60 is a composition comprising the polynucleotide of any one of embodiments 1 to 57, further comprising a single guide RNA, wherein the single guide RNA comprises a guide region and a conserved region, (a) a truncated repeat / anti-repeat region, the truncated repeat / anti-repeat region lacking 2 to 24 nucleotides; (i) one or more of nucleotides 37 to 64 are deleted and optionally substituted relative to SEQ ID NO:500; (ii) nucleotide 36 is (i) a first internal linker that replaces four nucleotides, either alone or in combination with nucleotides, or (ii) a shortened repeat / anti-repeat region that is connected to nucleotide 65 by at least four nucleotides; or (b) a truncated hairpin 1 region, wherein truncated hairpin 1 lacks 2-10, optionally 2-8, nucleotides; (i) one or more of nucleotides 82 to 95 are deleted and optionally substituted relative to SEQ ID NO:500; (ii) nucleotide 81 is (i) a second internal linker that replaces four nucleotides, either alone or in combination with nucleotides, or (ii) a shortened hairpin 1 region that is linked to nucleotide 96 by at least four nucleotides; or (c) a truncated hairpin 2 region, wherein the truncated hairpin 2 is missing 2-18, optionally 2-16, nucleotides; (i) one or more of nucleotides 113 to 134 are deleted and optionally substituted relative to SEQ ID NO:500; (ii) nucleotide 112 includes one or more of: (i) a third internal linker, alone or in combination with nucleotides, substituting four nucleotides; or (ii) a truncated hairpin 2 region, connected to nucleotide 135 by at least four nucleotides; one or both nucleotides 144 to 145 are optionally deleted compared to SEQ ID NO:500, The composition, wherein the gRNA comprises at least one of a first internal linker, a second internal linker, and a third internal linker.
[0070] Embodiment 61 is a polypeptide encoded by the polynucleotide according to any one of embodiments 1 to 60.
[0071] Embodiment 62 is a vector comprising the polynucleotide according to any one of embodiments 1 to 60.
[0072] Embodiment 63 is an expression construct comprising a promoter operably linked to a sequence encoding the polynucleotide of any one of embodiments 1 to 60.
[0073] Embodiment 64 is an expression construct according to embodiment 63, wherein the promoter is an RNA polymerase promoter, optionally a bacterial RNA polymerase promoter.
[0074] Embodiment 65 is an expression construct according to embodiment 63 or 64, further comprising a polyA tail sequence or a polyadenylation signal sequence.
[0075] Embodiment 66 is an expression construct according to embodiment 65, wherein the poly A tail sequence is an encoded poly A tail sequence.
[0076] Embodiment 67 is a plasmid comprising an expression construct according to any one of embodiments 63 to 66.
[0077] Embodiment 68 is a host cell comprising the vector according to embodiment 62, the expression construct according to any one of embodiments 63 to 66, or the plasmid according to embodiment 67.
[0078] Embodiment 69 is a pharmaceutical composition comprising the polynucleotide, composition, or polypeptide according to any one of embodiments 1 to 61 and a pharma- ceutically acceptable carrier.
[0079] Embodiment 70 is a kit comprising the polynucleotide, composition, or polypeptide according to any one of embodiments 1 to 61.
[0080] Embodiment 71 is the use of a polynucleotide, a composition or a polypeptide according to any one of embodiments 1 to 61 for modifying a target gene in a cell.
[0081] Embodiment 72 is the use of a polynucleotide, a composition or a polypeptide according to any one of embodiments 1 to 61 for the manufacture of a medicament for modifying a target gene in a cell.
[0082] Embodiment 73 is a polynucleotide or composition according to any one of embodiments 1 to 60, wherein the polynucleotide or composition is formulated as a lipid nucleic acid assembly composition, optionally formulated as a lipid nanoparticle.
[0083] Embodiment 74 is a method for modifying a target gene, comprising delivering to a cell a polynucleotide, a polypeptide, or a composition according to any one of embodiments 1 to 61.
[0084] Embodiment 75 is a method of modifying a target gene comprising delivering to said cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising a polynucleotide according to any one of embodiments 1 to 60 and one or more guide RNAs.
[0085] Embodiment 76 is the method of embodiment 74 or 75, wherein at least one lipid nucleic acid assembly composition comprises a lipid nanoparticle (LNP), and optionally, all lipid nucleic acid assembly compositions comprise LNPs.
[0086] Embodiment 77 is the method of any one of embodiments 74-75, wherein at least one lipid nucleic acid assembly composition is a lipoplex composition.
[0087] Embodiment 78 is a composition or method according to any one of embodiments 75-77, wherein the lipid nucleic acid assembly composition comprises an ionizable lipid.
[0088] Embodiment 79 is a method for producing a polynucleotide according to any one of embodiments 1 to 54, comprising contacting an expression construct according to embodiments 63 to 66 with an RNA polymerase comprising at least one modified nucleotide and NTPs.
[0089] Embodiment 80 is the method of embodiment 79, wherein the NTP comprises one modified nucleotide.
[0090] Embodiment 81 is the method of embodiment 79 or 80, wherein the modified nucleotide comprises a modified uridine.
[0091] Embodiment 82 is the method of embodiment 81, wherein at least 80%, or at least 90%, or 100% of the uridine positions are modified uridines.
[0092] Embodiment 83 is the method of embodiment 81 or 82, wherein the modified uridine comprises or is a substituted uridine, pseudouridine, or a substituted pseudouridine, optionally N1-methyl-pseudouridine.
[0093] Embodiment 84 is the method of any one of embodiments 79 to 83, wherein the expression construct comprises an encoded polyA tail sequence. [Brief description of the drawings]
[0094] [Figure 1] FIG. 1 shows the average percent editing at the TTR locus in PMH with increasing doses of Nme2Cas9 mRNA and chemically modified sgRNA. [Figure 2A] FIG. 1 shows the average percent editing at the TTR locus in PMH using various ratios of sgRNA and Nme2Cas9 mRNA. [Figure 2B] FIG. 1 shows the average percent editing at the TTR locus in PMH using various ratios of pgRNA and Nme2Cas9 mRNA. [Diagram 3] FIG. 13 shows the average percent editing at the TTR locus in PMH for pgRNAs with Nme2Cas9 mRNA. [Figure 4] FIG. 1 shows the average editing percentage at the PCSK9 locus in PMH. [Figure 5A] FIG. 13 shows the average editing results at the VEGFA locus in HEK cells treated with mRNA C. [Figure 5B] FIG. 13 shows the average editing results at the VEGFA locus in HEK cells treated with mRNA I. [Figure 5C] FIG. 13 shows the average editing results at the VEGFA locus in HEK cells treated with mRNA J. [Figure 5D] FIG. 13 shows the average editing results at the VEGFA locus in PHH cells treated with mRNA C. [Figure 5E] FIG. 13 shows the average editing results at the VEGFA locus in PHH cells treated with mRNA I. [Figure 5F] FIG. 13 shows the average editing results at the VEGFA locus in PHH cells treated with mRNA J. [Figure 6] FIG. 1 shows the average percent editing at the mouse TTR locus in PMH cells treated with NmeCas9 constructs designed with one or two nuclear localization sequences. [Figure 7] FIG. 1 shows the average percent editing at the mouse TTR locus in PMH cells treated with pgRNA and various Nme2Cas9 mRNAs. [Figure 8] FIG. 1 shows the fold change in Nme2Cas9 protein expression compared to SpyCas9 protein expression in PMH, PRH, PCH and PHH cells. [Figure 9A] FIG. 1 shows the fold change in Nme2Cas9 protein expression compared to SpyCas9 protein expression in T cells from two donors assayed 24, 48, and 72 hours after treatment. [Figure 9B] FIG. 1 shows the fold change in Nme2Cas9 protein expression compared to SpyCas9 protein expression in T cells from two donors assayed 24, 48, and 72 hours after treatment. [Figure 9C] FIG. 1 shows the fold change in Nme2Cas9 protein expression compared to SpyCas9 protein expression in T cells from two donors assayed 24, 48, and 72 hours after treatment. [Figure 9D] FIG. 1 shows the fold change in Nme2Cas9 protein expression compared to SpyCas9 protein expression in T cells from two donors assayed 24, 48, and 72 hours after treatment. [Figure 9E] FIG. 1 shows the fold change in Nme2Cas9 protein expression compared to SpyCas9 protein expression in T cells from two donors assayed 24, 48, and 72 hours after treatment. [Figure 9F] FIG. 1 shows the fold change in Nme2Cas9 protein expression compared to SpyCas9 protein expression in T cells from two donors assayed 24, 48, and 72 hours after treatment. [Figure 10] FIG. 1 shows the average percent editing at the TTR locus in mouse liver treated with sgRNA and Nme2Cas9. [Figure 11A]FIG. 1 shows the average percent editing at the TTR locus in mouse liver following treatment with pgRNA and Nme2Cas9. [Figure 11B] FIG. 1 shows mean serum TTR protein after treatment with pgRNA and Nme2Cas9. [Figure 11C] FIG. 1 shows the average percent TTR knockdown following treatment with pgRNA and Nme2Cas9. [Figure 11D] FIG. 1 shows the average percent editing at the TTR locus in mouse liver following treatment with pgRNA and various Nme2Cas9s. [Figure 11E] FIG. 1 shows serum TTR protein knockdown following treatment with pgRNA and various Nme2Cas9s. [Figure 12] FIG. 1 shows the average percentage editing in mouse liver after treatment with various Nme2Cas9 constructs. [Figure 13] FIG. 1 shows the average percent editing in mouse liver following treatment with pgRNA and various Nme2Cas9s. [Figure 14] FIG. 1 shows the average percentage editing in mouse liver after treatment with various base editors. [Figure 15] FIG. 1 shows an exemplary schematic diagram of an Nme2 sgRNA in possible secondary structures, including a repeat / anti-repeat region and a hairpin region, including hairpin1 and hairpin2 regions, further showing the guide region (or target region) (shown in gray fill with dashed lines), bases that do not tolerate single or paired deletions (shown in gray fill with solid lines), and bases that tolerate single or paired deletions (open circles). [Figure 16] FIG. 1 shows the average percentage of CD3 negative T cells after TRAC editing with Nme1Cas9. [Figure 17] FIG. 1 shows the average percentage of CD3-negative T cells after TRAC editing with Nme3Cas9. [Figure 18] FIG. 1 shows expression of Nme-HiBiT constructs in T cells at 24 hours. [Figure 19] FIG. 13 shows CD3 negative cell population as a function of NmeCas9 mRNA amount. [Figure 20] FIG. 13 shows dose-response curves of selected gRNAs in PCH. [Figure 21] FIG. 1 shows dose response curves of LNP dilution series in PCH. [Figure 22] FIG. 1 shows serum TTR levels in mice. [Figure 23] FIG. 1 shows percent editing at the TTR locus in mouse liver samples. [Figure 24] FIG. 1 shows dose-response curves of selected gRNAs in PMH. [Diagram 25] FIG. 1 shows dose-response curves of selected gRNAs in PMH. [Figure 26] FIG. 1 shows the average percent editing at the PCSK9 locus in PMH with modified sgRNAs. [Figure 27] FIG. 1 shows the average percent editing in PMH of several Nme2Cas9 mRNAs with modified sgRNAs. [Figure 28] FIG. 1 shows percent editing at the TTR locus in primary mouse hepatocytes. [Figure 29] FIG. 1 shows serum TTR levels in mice. [Diagram 30] FIG. 1 shows percent editing at the TTR locus in mouse liver samples. [Diagram 31] FIG. 1 shows serum TTR measurements after treatment in mice. [Diagram 32] FIG. 1 shows percent editing at the TTR locus in mouse liver samples. [Diagram 33]FIG. 1 shows an exemplary sgRNA (G021536, SEQ ID NO: 139) in possible secondary structures. Methylation is indicated in bold and phosphorothioate bonds are indicated with "*". Watson-Crick base pairing is indicated by "-" between nucleotides of the double-stranded portion. Non-Watson-Crick base pairing is indicated by "●" between nucleotides of the double-stranded portion. [Diagram 34] FIG. 1 shows an exemplary sgRNA (G032572, SEQ ID NO:528) in possible secondary structures. Unmodified nucleotides are shown in bold, methylations are shown in light font, and phosphorothioate bonds are indicated with "*". Watson-Crick base pairing is indicated by a "-" between nucleotides of the double-stranded portion. Non-Watson-Crick base pairing is indicated by a "●" between nucleotides of the double-stranded portion. [Diagram 35] FIG. 1 shows an exemplary sgRNA (G031771, SEQ ID NO:529) in possible secondary structures. Unmodified nucleotides are shown in bold, methylations are shown in light font, and phosphorothioate bonds are indicated with "*". Watson-Crick base pairing is indicated by a "-" between nucleotides of the double-stranded portion. Non-Watson-Crick base pairing is indicated by a "●" between nucleotides of the double-stranded portion. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0095] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5]
[0096] The transcription sequence may generally include ARCA as the first three nucleotides for use with CleanCap™ or GGG as the first three nucleotides for use with AGG. Thus, the first three nucleotides may be modified for use with other capping approaches, such as vaccinia capping enzyme. Promoters and polyA sequences are not included in the transcription sequence. Promoters such as U6 promoter (SEQ ID NO: 389) or CMV promoter (SEQ ID NO: 390) and polyA sequences such as SEQ ID NO: 409 may be added to the 5' and 3' ends of the disclosed transcription sequences, respectively. Most nucleotide sequences are provided as DNA, but can be easily converted to RNA by changing T to U.
[0097] Reference will now be made in detail to certain specific embodiments of the invention, examples of which are illustrated in the accompanying drawings. While the invention will be described in conjunction with the illustrated embodiments, it will be understood that they are not intended to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents that may be included within the invention as defined by the appended claims.
[0098] Before describing the teachings of the present invention in detail, it is to be understood that the present disclosure is not limited to specific compositions or process steps, as such may vary. It should be noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, a reference to "a conjugate" includes a plurality of conjugates, a reference to "a cell" includes a plurality of cells, and so forth.
[0099] Numerical ranges include the numbers defining the range. Measurements and measurable values are understood to be approximations taking into account the number of significant digits and errors associated with the measurements. Accordingly, unless otherwise indicated, the numerical parameters set forth in the following specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained. Each numerical parameter should, at the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, be construed at least in light of the number of reported significant digits and by applying ordinary rounding techniques.
[0100] The term "about" is used herein to mean within a typical tolerance in the art. For example, "about" can be understood as about 2 standard deviations from the mean. In certain embodiments, about means +10%. In certain embodiments, about means +5%, +2%, or +1%. When about is present before a series of numbers or ranges, it is understood that "about" can modify each number in the series or range. Each numerical parameter should be interpreted at least in light of the number of reported significant digits and by applying ordinary rounding techniques, but not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims.
[0101] The use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" are not intended to be limiting. It is to be understood that the foregoing general description and detailed description are exemplary and explanatory only and are not intended to limit the present teachings.
[0102] The term "at least" preceding a number or a series of numbers is understood to include the number adjacent to the term "at least" and all subsequent numbers or integers that may be logically included, as is clear from the context. For example, the number of nucleotides in a nucleic acid molecule must be an integer. For example, "at least 17 nucleotides of a 20-nucleotide nucleic acid molecule" means that 17, 18, 19, or 20 nucleotides have the indicated property. When at least is present before a series of numbers or a range, it is understood that "at least" can modify each of the numbers in the series or range.
[0103] As used herein, "less than" or "less than" is understood to mean the expression and the logical lower zero or adjacent integer value that is logical from the context. For example, a duplex region of "2 or less nucleotide base pairs" has 2, 1, or 0 nucleotide base pairs. When "less than" or "less than" is present before a series of numbers or ranges, it is understood that each of the series of numbers or ranges is modified.
[0104] As used herein, ranges include both upper and lower limits.
[0105] As used herein, when the maximum amount of a value is expressed as 100% (e.g., 100% inhibition), it is understood that the value is interpreted in the context of the detection method. For example, 100% inhibition, etc. is understood as inhibition to a level below the detection level of the assay.
[0106] Unless otherwise stated in the specification above, embodiments herein that recite various components as "comprising" are also assumed to be "consisting of" or "consisting essentially of" the recited components. Embodiments herein that recite various components as "consisting of" are also assumed to be "comprising" or "consisting essentially of" the recited components. Embodiments herein that recite various components as "consisting essentially of" are also assumed to be "consisting of" or "comprising" the recited components (this interchangeability does not apply to the use of these terms in the claims).
[0107] The section headings used herein are for organizational purposes only and should not be construed as limiting the desired subject matter in any manner. In the event that any document incorporated by reference contradicts the express contents of this specification, including but not limited to definitions, the express contents of this specification shall prevail. Although the teachings of the present invention have been described in conjunction with various embodiments, it is not intended that the teachings of the present invention be limited to those embodiments. On the contrary, the teachings of the present invention encompass various alternatives, modifications, and equivalents, as will be apparent to those skilled in the art.
[0108] I. Definition Unless otherwise stated, the following terms and phrases used herein are intended to have the following meanings.
[0109] The term "or combinations thereof" as used herein refers to all permutations and combinations of the terms listed before this term. For example, "A, B, C, or combinations thereof" is intended to include at least one of A, B, C, AB, AC, BC, or ABC, and also BA, CA, CB, ACB, CBA, BCA, BAC, or CAB, if the order is important in a particular context. Following this example, combinations containing one or more repeats of an item or term, such as BB, AAA, AAB, BBC, AAABC, CBBA, BABB, etc., are expressly included. Those skilled in the art will understand that there is typically no limit to the number of items or terms in any combination, unless otherwise clear from the context.
[0110] As used herein, the term "kit" refers to a packaged set of related components (one or more polynucleotides or compositions) and one or more associated substances (such as a delivery device (e.g., a syringe), solvent, solution, buffer, instructions or desiccant).
[0111] "Or" is used in its inclusive sense, ie, equivalent to "and / or," unless the context requires otherwise.
[0112] "Polynucleotide" and "nucleic acid" are used herein to refer to polymeric compounds that contain nucleosides or nucleoside analogs with nitrogen-containing heterocyclic bases or base analogs linked together along the backbone, including polymers that are conventional RNA, DNA, mixed RNA-DNA, and analogs thereof. The nucleic acid "backbone" can be composed of a variety of linkages, including one or more of sugar phosphodiester linkages, peptide-nucleic acid linkages ("peptide nucleic acid" or PNA, PCT No. WO95 / 32305), phosphorothioate linkages, methylphosphonate linkages, or combinations thereof. The sugar moiety of the nucleic acid can be ribose, deoxyribose, or similar compounds with substitutions (e.g., 2' methoxy or 2' halide substitutions). The nitrogenous base may be a conventional base (A, G, C, T, U), an analog thereof (e.g., a modified uridine such as 5-methoxyuridine, pseudouridine, or N1-methylpseudouridine), an inosine, a purine, or a pyrimidine derivative (e.g., N4-methyldeoxyguanosine, deaza- or aza-purines, deaza- or aza-pyrimidines, pyrimidine bases having a substituent at the 5- or 6-position (e.g., 5-methylcytosine), purine bases having a substituent at the 2-, 6-, or 8-position, 2-amino-6-methylaminopurine, O6-methylguanine, 4-thio-pyrimidine, 4-amino-pyrimidine, 4-dimethylhydrazine-pyrimidine, and O4-alkyl-pyrimidines, U.S. Pat. No. 5,378,825 and PCT WO 93 / 13121). For a general discussion, see The Biochemistry of the Nucleic Acids 5-36, edited by Adams et al., 11th ed., 1992). Nucleic acids may contain one or more "abasic" residues in which the backbone does not contain a nitrogenous base at a position or positions in the polymer (U.S. Pat. No. 5,585,481). Nucleic acids may contain only conventional RNA or DNA sugars, bases, and linkages, or may contain both conventional building blocks and substitutions (e.g., polymers containing conventional bases with 2' methoxy linkages, or both conventional bases and one or more base analogs).Nucleic acids include "locked nucleic acids" (LNAs), which are analogs that contain one or more LNA nucleotide monomers that have a bicyclic furanose unit locked to an RNA-mimetic sugar structure, enhancing hybridization affinity to complementary RNA and DNA sequences (Vester and Wengel, 2004, Biochemistry 43(42):13233-41). RNA and DNA have different sugar moieties and can differ by the presence of uracil or its analogs in RNA and thymine or its analogs in DNA.
[0113] "Polypeptide" as used herein refers to a multimeric compound comprising amino acid residues capable of adopting a three-dimensional conformation. Polypeptides include, but are not limited to, enzymes, zymogen proteins, regulatory proteins, structural proteins, receptors, nucleic acid binding proteins, antibodies, and the like. Polypeptides may, but do not necessarily, include post-translational modifications, unnatural amino acids, prosthetic groups, and the like.
[0114] As used herein, "cytidine deaminase" means a polypeptide or polypeptide complex capable of cytidine deaminase activity that catalyzes the hydrolytic deamination of cytidine or deoxycytidine, typically to yield uridine or deoxyuridine. Cytidine deaminases include enzymes within the cytidine deaminase superfamily, in particular the APOBEC family of enzymes (APOBEC1, APOBEC2, APOBEC4, and APOBEC3 subgroups of enzymes), activation-induced cytidine deaminases (AID or AICDA), and CMP deaminases (see, e.g., Conticello et al., Mol. Biol. Evol. 22:367-77, 2005; Conticello, Genome Biol. 9:229, 2008; Muramatsu et al., J. Biol. Chem. 274:18470-6, 1999; Carrington et al., Cells 9:1690 (2020)). In some embodiments, variants of any known cytidine deaminase or APOBEC protein are included. Variants include proteins with sequences that differ from the wild-type protein by one or more mutations (i.e., substitutions, deletions, insertions), e.g., one or more single-point substitutions. For example, truncated sequences can be used, e.g., by deleting 1-4 amino acids at the N-terminus, C-terminus, or internal amino acids, preferably the C-terminus of the sequence. As used herein, the term "variant" refers to allelic variants, splicing variants, and natural or artificial mutants that are homologous to a reference sequence. Variants are "functional" in that they exhibit catalytic activity for DNA editing.
[0115] As used herein, the term "APOBEC3A" refers to a cytidine deaminase, such as the protein expressed by the human A3A gene. APOBEC3A may have catalytic DNA editing activity. The amino acid sequence of APOBEC3A has been described (UniPROT Accession ID: p31941) and is included herein as SEQ ID NO: 151. In some embodiments, the APOBEC3A protein is a human APOBEC3A protein or a wild-type protein. Variants include proteins having a sequence that differs from the wild-type APOBEC3A protein by one or more mutations (i.e., substitutions, deletions, insertions), e.g., one or more single-point substitutions. For example, truncated APOBEC3A sequences can be used, e.g., by deleting 1-4 amino acids at the N-terminus, C-terminus, or internal amino acids, preferably the C-terminus of the sequence. As used herein, the term "variant" refers to allelic variants, splicing variants, and natural or artificial mutants that are homologous to the APOBEC3A reference sequence. The variants are "functional" in that they exhibit catalytic activity for DNA editing. In some embodiments, APOBEC3A (e.g., human APOBEC3A) has a wild-type amino acid position 57 (numbered in the wild-type sequence). In some embodiments, APOBEC3A (e.g., human APOBEC3A) has an asparagine at amino acid position 57 (numbered in the wild-type sequence).
[0116] Several Cas9 orthologues have been obtained from Neisseria meningitidis (Esvelt et al., NAT. METHODS, vol. 10, 2013, 1116-1121; Hou et al., PNAS, vol. 110, 2013, pages 15644-15649) (Nme1Cas9, Nme2Cas9, and Nme3Cas9). The Nme2Cas9 orthologue functions efficiently in mammalian cells, recognizes the N4CC PAM, and can be used for in vivo editing (Ran et al., NATURE, vol. 520, 2015, pages 186-191; Kim et al., NAT. COMMUN., vol. 8, 2017, pages 14500). Nme2Cas9 has been shown to be naturally resistant to off-target editing (Lee et al., MOL. THER., vol. 24, 2016, pages 645-654; Kim et al., 2017). See also, e.g., WO / 2020081568 (e.g., pages 28 and 42), describing Nme2Cas9 D16A nickase, the contents of which are incorporated herein by reference in their entirety. Additionally, NmeCas9 variants are known in the art, see, e.g., Huang et al., Nature Biotech. 2022, doi.org / 10.1038 / s41587-022-01410-2, describing Cas9 mutants that target single nucleotide-pyrimidine PAMs. Throughout, "NmeCas9" or "Nme Cas9" is a generic term and encompasses any type of NmeCas9, including Nme1Cas9, Nme2Cas9, and Nme3Cas9.
[0117] As used herein, the term "fusion protein" refers to a hybrid polypeptide that includes polypeptides from at least two different proteins or sources. One polypeptide may be located at the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) portion of the fusion protein, thus forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein", respectively. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins that include peptide linkers. Methods of recombinant protein expression and purification are well known and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.
[0118] As used herein, the term "uracil glycosylase inhibitor," "uracil-DNA glycosylase inhibitor" or "UGI" refers to a protein capable of inhibiting the uracil-DNA glycosylase (UDG) base excision repair enzyme (e.g., UniProt ID: P14739, SEQ ID NO: 3).
[0119] The term "linker", as used herein, refers to a chemical group or molecule that connects two adjacent molecules or moieties. Typically, a linker is located between or sandwiched between two groups, molecules or other moieties and is linked to each other by a covalent bond. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). Exemplary peptide linkers are disclosed elsewhere herein.
[0120] As used herein, "modified uridine" refers to a nucleoside other than thymidine that has the same hydrogen bond acceptor as uridine and one or more structural differences from uridine. In some embodiments, the modified uridine is a substituted uridine, i.e., a uridine in which one or more aprotic substituents (e.g., alkoxy, such as methoxy) replace a proton. In some embodiments, the modified uridine is a pseudouridine. In some embodiments, the modified uridine is a substituted pseudouridine, i.e., a pseudouridine in which one or more aprotic substituents (e.g., alkyl, such as methyl) replace a proton. In some embodiments, the modified uridine is either a substituted uridine, a pseudouridine, or a substituted pseudouridine, e.g., N1-methyl-pseudouridine.
[0121] As used herein, a "uridine position" refers to a position in a polynucleotide that is occupied by a uridine or modified uridine. Thus, for example, a polynucleotide with "100% of the uridine positions being modified uridines" contains modified uridines at all positions that would be uridines in a conventional RNA of the same sequence (where all bases are standard A, U, C, or G bases). Unless otherwise specified, U in the polynucleotide sequences in this disclosure or in the sequence listing or sequence listings accompanying it can be a uridine or modified uridine.
[0122] As used herein, a first sequence is considered to "contain a sequence having at least X% identity to" a second sequence if alignment of the first sequence to the second sequence shows that X% or more of the positions of the second sequence match the first sequence overall. For example, the sequence AAGA contains a sequence having 100% identity to the sequence AAG, because the alignment gives 100% identity in that there is a match at all three positions of the second sequence. Differences between RNA and DNA (generally the replacement of thymidine with uridine or vice versa) and the presence of nucleoside analogs such as modified uridines do not contribute to differences in identity or complementarity between polynucleotides, as long as the relevant nucleotides (such as thymidine, uridine, or modified uridine) have the same complement (e.g., adenosine for all of thymidine, uridine, or modified uridine; another example is cytosine and 5-methylcytosine, both of which have guanosine as their complement). Thus, for example, in the sequence 5'-AXG, where X is any modified uridine, such as pseudouridine, N1-methylpseudouridine, or 5-methoxyuridine, it is considered to be 100% identical to AUG, since they are all perfectly complementary to the same sequence (5'-CAU). Exemplary alignment algorithms are the Smith-Waterman and Needleman-Wunsch algorithms, which are well known in the art. Those skilled in the art will understand what algorithm and parameter setting is appropriate for a given pair of sequences to be aligned. For sequences of generally similar length and predicted identity of more than 50% for amino acids or more than 75% for nucleotides, the Needleman-Wunsch algorithm, using the default settings of the Needleman-Wunsch algorithm interface provided by the EBI at the www.ebi.ac.uk web server, is generally appropriate.
[0123] "mRNA" is used herein to refer to a polynucleotide that is an RNA or modified RNA and contains an open reading frame that can be translated into a polypeptide (i.e., can serve as a substrate for translation by ribosomes and aminoacylated tRNAs). An mRNA can contain a phosphate sugar backbone that contains ribose residues or analogs thereof, e.g., 2'-methoxyribose residues. In some embodiments, the sugars of an mRNA phosphate-sugar backbone consist essentially of ribose residues, 2'-methoxyribose residues, or combinations thereof. Generally, an mRNA does not contain a significant amount of thymidine residues (e.g., 0 residues, or fewer than 30, 20, 10, 5, 4, 3, or 2 thymidine residues; or a thymidine content of less than 10%, 9%, 8%, 7%, 6%, 5%, 4%, 4%, 3%, 2%, 1%, 0.5%, 0.2%, or 0.1%). An mRNA can contain modified uridines at some or all of its uridine positions.
[0124] As used herein, "RNA-guided DNA binder" refers to a polypeptide or polypeptide complex having RNA-binding activity and DNA-binding activity, or a DNA-binding subunit of such a complex, whose DNA-binding activity is sequence-specific and depends on the sequence of its RNA. Exemplary RNA-guided DNA binders include Cas cleavase / Cas nickase and their inactivated forms ("dCas DNA binders"). "Cas nuclease", also referred to as "Cas protein" as used herein, includes Cas cleavase, Cas nickase, and dCas DNA binders. dCas DNA binders can be inactive nucleases that include non-functional nuclease domains (RuvC or HNH domains). In some embodiments, Cas cleavase or Cas nickase includes dCas DNA binders that have been modified to enable DNA cleavage, for example, via fusion with a FokI domain. Exemplary nucleotide and polypeptide sequences of Cas9 molecules are shown below. Methods for identifying alternative nucleotide sequences encoding Cas9 polypeptide sequences, including alternative naturally occurring variants, are known in the art. Sequences with at least 75%, 80%, 85%, preferably 90%, 95%, 96%, 97%, 98%, or 99% identity to any of the Cas9 nucleic acid sequences, amino acid sequences, or nucleic acid sequences encoding the amino acid sequences provided herein are also contemplated. Exemplary open reading frames of Cas9 are shown in Table 39A.
[0125] As used herein, the "minimal uridine codon(s)" for a given amino acid is the codon(s) with the fewest uridines (usually 0 or 1, except for the codon for phenylalanine, where the minimal uridine codon has 2 uridines). Modified uridine residues are considered equivalent to uridine for the purposes of assessing uridine content.
[0126] As used herein, the "uridine dinucleotide (UU) content" of an ORF can be expressed in absolute terms, based on a percentage, as a count of UU dinucleotides in the ORF, or as the percentage of positions occupied by uridine of the uridine dinucleotides (e.g., AUUAU has a uridine dinucleotide content of 40% because 2 of the 5 positions are occupied by uridine of the uridine dinucleotide). Modified uridine residues are considered equivalent to uridine for purposes of assessing uridine dinucleotide content.
[0127] As used herein, the "minimal adenine codon(s)" of a given amino acid is the codon(s) with the fewest adenines (typically 0 or 1, except for codons for lysine and asparagine, where the minimal adenine codon has 2 adenines). Modified adenine residues are considered equivalent to adenine for the purposes of assessing adenine content.
[0128] As used herein, the "adenine dinucleotide content" of an ORF can be expressed in absolute terms, based on a percentage, as a count of AA dinucleotides in the ORF, or as the percentage of positions occupied by adenine of adenine dinucleotides (e.g., UAAUA has an adenine dinucleotide content of 40% because 2 of the 5 positions are occupied by adenine of adenine dinucleotides). Modified adenine residues are considered equivalent to adenine for purposes of assessing adenine dinucleotide content.
[0129] As used herein, the "minimum repeat content" of a given open reading frame (ORF) is the smallest possible total occurrence of AA, CC, GG, and TT (or TU, UT, or UU) dinucleotides in an ORF that encodes the same amino acid sequence as the given ORF. Repeat content can be expressed in absolute terms, based on percentages, as the count of AA, CC, GG, and TT (or TU, UT, or UU) dinucleotides in an ORF, or as the count of AA, CC, GG, and TT (or TU, UT, or UU) dinucleotides in an ORF divided by the nucleotide length of the ORF (e.g., UAAUA has a repeat content of 20% because it occurs once in a sequence of five nucleotides). Modified adenine, guanine, cytosine, thymine, and uracil residues are considered equivalent to adenine, guanine, cytosine, thymine, and uracil residues for purposes of assessing minimum repeat content.
[0130] "Guide RNA", "gRNA", and "guide" are used interchangeably herein to refer to either crRNA (also known as CRISPR RNA) or a combination of crRNA and trRNA (also known as tracrRNA). crRNA and trRNA may associate as a single-stranded RNA molecule (single guide RNA, sgRNA) or in two separate RNA molecules (dual guide RNA, dgRNA). "Guide RNA" or "gRNA" refer to each type. trRNA may be a naturally occurring sequence or a trRNA sequence that has modifications or variations compared to a naturally occurring sequence. Guide RNA may include modified RNA as described herein. Unless otherwise clear from the context, guide RNAs described herein are suitable for use with Nme Cas9, e.g., Nme1, Nme2, or Nme3 Cas9. For example, FIG. 15 shows an exemplary schematic of Nme2 sgRNA in possible secondary structures.
[0131] As used herein, "guide sequence" or "guide region" or "spacer" or "spacer sequence" and the like refer to a sequence within a guide RNA that is complementary to a target sequence and functions to direct the guide RNA to the target sequence for binding or modification (e.g., cleavage) by NmeCas9. A guide sequence can be, for example, 20-25 nucleotides in length, for Nme Cas9 and related Cas9 homologs / orthologues. Shorter or longer sequences, for example, 20, 21, 22, 23, 24, or 25 nucleotides in length, can also be used as guides. For Nme Cas9, a guide sequence can be at least 22, 23, 24, or 25 nucleotides in length. A guide sequence can form a 22, 23, 24, or 25 contiguous base pair duplex, for example, a 24 contiguous base pair duplex, for Nme Cas9, with its target sequence.
[0132] Since the nucleic acid substrate of the Cas protein is a double-stranded nucleic acid, the target sequence of the Cas protein includes both the plus and minus strands of genomic DNA (i.e., the given sequence and the reverse complement of that sequence). Thus, when a guide sequence is described as being "complementary to a target sequence," it is understood that the guide sequence can direct the guide RNA to bind to the reverse complement of the target sequence. That is, in some embodiments, when a guide sequence binds to the reverse complement of a target sequence, the guide sequence is identical to a particular nucleotide of the target sequence (e.g., the target sequence without the PAM) except for the T to U substitution in the guide sequence.
[0133] As used herein, "indel" refers to an insertion / deletion mutation consisting of a number of nucleotides that are either inserted or deleted at the site of a double strand break (DSB) in a nucleic acid.
[0134] As used herein, "knockdown" refers to a reduction in the expression of a particular gene product (e.g., protein, mRNA, or both). Protein knockdown can be measured either by detecting the protein secreted by a tissue or cell aggregate (e.g., in serum or cell culture medium) or by detecting the total cellular amount of protein from a tissue or cell aggregate of interest. Methods for measuring mRNA knockdown are known and include sequencing of mRNA isolated from a tissue or cell population of interest. In some embodiments, "knockdown" can refer to some loss of expression of a particular gene product, for example, a reduction in the amount of transcribed mRNA, or a reduction in the amount of protein expressed or secreted by a cell population (including in vivo populations such as those found in tissues).
[0135] As used herein, "knockout" refers to the loss of expression of a particular protein in a cell. Knockout can be measured either by detecting the amount of protein secreted from a tissue or cell aggregate (e.g., in serum or cell medium) or by detecting the total cellular amount of the protein in a tissue or cell aggregate. In some embodiments, the disclosed method "knocks out" a target protein in one or more cells (e.g., in an aggregate of cells, including in vivo aggregates such as those found in tissues). In some embodiments, the knockout is a complete loss of expression of a target protein in a cell, rather than the formation of a mutant of the target protein, e.g., generated by indels.
[0136] As used herein, "ribonucleoprotein" (RNP) or "RNP complex" refers to a guide RNA comprising an RNA-guided DNA binding agent, such as a Cas cleavase, a Cas nickase, or a dCas DNA binding agent (e.g., Cas9). In some embodiments, the guide RNA guides an RNA-guided DNA binding agent, such as Cas9, to a target sequence, where the guide RNA hybridizes to the target sequence, the binding agent binds to the target sequence, and if the binding agent is a cleavase or nickase, can cleave or nick after binding.
[0137] As used herein, "target sequence" refers to a nucleic acid sequence in a target gene that has complementarity to a guide sequence of a gRNA. The interaction of the target sequence with the guide sequence induces an RNA-guided DNA binding agent to bind to the target sequence and potentially nick or cleave (depending on the activity of the binding agent) in the target sequence.
[0138] In some embodiments, the target sequence may be adjacent to the PAM. In some embodiments, the PAM may be adjacent to or within 1, 2, 3, or 4 nucleotides of the 3' end of the target sequence. The length and sequence of the PAM may depend on the Cas protein used. For example, the PAM may be selected from a consensus sequence or a specific PAM sequence for a particular Nme Cas9 protein or Nme Cas9 ortholog (Edraki et al., 2019). In some embodiments, the PAM may comprise 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length. Non-limiting exemplary PAM sequences include NCC, N4GAYW, N4GYTT, N4GTCT, NNNNCC(a), NNNNCAAA (N is defined as any nucleotide, W is defined as either A or T, R is defined as either A or G, and (a) is preferably, but not required to be, an A after the second C). In some embodiments, the PAM sequence may be NCC.
[0139] As used herein, "treatment" refers to any administration or application of therapy for a disease or disorder in a subject, including slowing or halting the progression or development of the disease, alleviating one or more signs or symptoms of the disease, curing the disease, or preventing the recurrence of one or more symptoms of the disease.
[0140] As used herein, the term "lipid nanoparticle" (LNP) refers to a particle that includes a plurality (i.e., two or more) of lipid molecules physically associated with each other by intermolecular forces. LNPs can be, for example, microspheres (including unilamellar and multilamellar vesicles, e.g., "liposomes," which are substantially spherical in some embodiments and can further include an aqueous core, e.g., containing a substantial portion of RNA molecules), the dispersed phase of an emulsion, a micelle, or the internal phase of a suspension. Emulsions, micelles, and suspensions can be suitable compositions for local or topical delivery. See, for example, WO2017173054A1, the contents of which are incorporated herein by reference in their entirety. Any LNP known to those of skill in the art to be capable of delivering nucleotides to a subject can be utilized with nucleic acids encoding guide RNAs and RNA-guided DNA binding agents as described herein.
[0141] As used herein, the term "nuclear localization signal" (NLS) or "nuclear localization sequence" refers to an amino acid sequence that directs the transport of a molecule that contains or is linked to such a sequence into the nucleus of a eukaryotic cell. The nuclear localization signal may form part of the molecule that is transported. In some embodiments, the NLS may be linked to the remainder of the molecule by covalent bonds, hydrogen bonds, or ionic interactions.
[0142] As used herein, "delivery" and "administration" are used interchangeably and include ex vivo and in vivo applications.
[0143] Concurrent administration, as used herein, means that two or more agents are administered sufficiently close in time that the agents act together. Concurrent administration includes administering the agents together in a single formulation, and administering the agents in separate formulations close enough in time that the agents act together.
[0144] As used herein, the phrase "pharmaceutical acceptable" means useful for preparing pharmaceutical compositions that are generally non-toxic, not biologically undesirable, and not otherwise unacceptable for pharmaceutical use. Pharmaceutically acceptable generally refers to a material that is not pyrogenic. Pharmaceutically acceptable may refer to a material that is sterile, particularly for pharmaceutical materials for injection or infusion.
[0145] II. Exemplary Polynucleotides and Compositions In some embodiments, a polynucleotide is provided, the polynucleotide comprising an open reading frame (ORF), the ORF comprising: and a nucleotide sequence encoding a C-terminal Neisseria meningitidis (Nme) Cas9 polypeptide that is at least 90% identical to any one of SEQ ID NOs: 29, 32-41, 224-226, 231-233, 238-240, 245-247, 252-254, 259-261, 266-268, 273-275, 280-282, 287-289, 294-296, 301-303, or 316-321; and a nucleotide sequence encoding a first nuclear localization signal (NLS). In some embodiments, the Nme Cas9 is Nme2 Cas9. In some embodiments, the Nme Cas9 is Nme1 Cas9. In some embodiments, the Nme Cas9 is Nme3 Cas9. In some embodiments, the ORF further comprises a nucleotide sequence encoding a second NLS. In some embodiments, the polynucleotide is mRNA.
[0146] In some embodiments, the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to any one of SEQ ID NOs: 29, 32-41, 224-226, 231-233, 238-240, 245-247, 252-254, 259-261, 266-268, 273-275, 280-282, 287-289, 294-296, 301-303, or 316-321. In some embodiments, the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to any one of SEQ ID NOs: 29 or 32-41. In some embodiments, the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to the sequence of SEQ ID NO: 32. In some embodiments, the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to the sequence of SEQ ID NO: 33. In some embodiments, the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to the sequence of SEQ ID NO: 34. In some embodiments, the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to the sequence of SEQ ID NO: 35. In some embodiments, the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to the sequence of SEQ ID NO: 36. In some embodiments, the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to the sequence of SEQ ID NO: 38. In some embodiments, the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to the sequence of SEQ ID NO: 39. In some embodiments, the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to the sequence of SEQ ID NO:41.
[0147] In some embodiments, the ORF comprises a sequence having at least 90%, 95%, 98% or 100% identity to the sequence of SEQ ID NO:38 or 41.
[0148] In some embodiments, a polynucleotide is provided, the polynucleotide comprising an ORF disclosed herein. In some embodiments, a polynucleotide is provided, the polynucleotide encoding an Nme Cas9 polypeptide at least 90% identical to any one of SEQ ID NOs: 29, 32-41, 224-226, 231-233, 238-240, 245-247, 252-254, 259-261, 266-268, 273-275, 280-282, 287-289, 294-296, 301-303, or 316-321, the Nme Cas9 being Nme2 Cas9, Nme3 Cas9, or Nme1 Cas9, a first nuclear localization signal (NLS), and a second NLS, the encoded first NLS and second NLS being N-terminal to the Nme Cas9 polypeptide.
[0149] In some embodiments, a polypeptide is provided, the polypeptide comprising an NmeCas9 polypeptide that is at least 90% identical to any one of at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% amino acid sequence of any one of SEQ ID NOs: 1 and 4-13, 220, 227, 234, 241, 248, 255, 262, 269, 276, 283, 290, or 297, or 310-315, wherein the NmeCas9 polypeptide is Nme2 Cas9, Nme3 Cas9, or Nme1 Cas9, a first nuclear localization signal (NLS), and a second NLS, wherein the encoded first NLS and second NLS are located N-terminal to the NmeCas9 polypeptide.
[0150] In some embodiments, a method of modifying a target gene is provided that comprises administering a composition described herein. In some embodiments, the method comprises delivering a polynucleotide comprising an open reading frame (ORF) to a cell, the ORF comprising a nucleotide sequence encoding a C-terminal Neisseria meningitidis (Nme) Cas9 polypeptide at least 90% identical to any one of SEQ ID NOs: 29, 32-41, 224-226, 231-233, 238-240, 245-247, 252-254, 259-261, 266-268, 273-275, 280-282, 287-289, 294-296, 301-303, or 316-321, wherein the Nme Cas9 is Nme2 Cas9 or Nme1 Cas9 or Nme3 Cas9, a nucleotide sequence encoding a first nuclear localization signal (NLS), and optionally a nucleotide sequence encoding a second NLS. In some embodiments, the polynucleotide is delivered to the cell in vitro. In some embodiments, the polynucleotide is delivered to a cell in vivo.
[0151] In some embodiments, the compositions described herein further comprise at least one gRNA. In some embodiments, a composition is provided comprising an mRNA described herein and at least one gRNA. In some embodiments, the gRNA is a single guide RNA (sgRNA). In some embodiments, the gRNA is a dual guide RNA (dgRNA).
[0152] In some embodiments, the composition is capable of effecting genome editing upon administration to a subject. In some embodiments, the subject is a human.
[0153] A. RNA-guided DNA-binding agent, NmeCas9 The RNA-guided DNA-binding agents described herein include Neisseria meningitidis Cas9 (NmeCas9) and modifications and variants thereof. In some embodiments, NmeCas9 is Nme2 Cas9. In some embodiments, NmeCas9 is Nme1 Cas9. In some embodiments, NmeCas9 is Nme3 Cas9.
[0154] Modified versions with one inactive catalytic domain, either RuvC or HNH, are referred to as "nickases". Nickases cut only one strand on the target DNA, thus generating a single-stranded break. A single-stranded break may also be known as a "nick". In some embodiments, the compositions and methods include a nickase. In some embodiments, the compositions and methods include a nickase RNA-guided DNA binding agent, such as a nickase Cas, e.g., a nickase Cas9, that induces a nick rather than a double-stranded break in the target DNA.
[0155] In some embodiments, the NmeCas9 nuclease can be modified to contain only one functional nuclease domain. For example, the RNA-guided DNA binding agent can be modified to have one of the nuclease domains mutated or completely or partially deleted to reduce nucleic acid cleavage activity.
[0156] In some embodiments, an NmeCas9 nickase with a RuvC domain that has reduced activity is used. In some embodiments, an NmeCas9 nickase with an inactive RuvC domain is used. In some embodiments, an NmeCas9 nickase with an HNH domain that has reduced activity is used. In some embodiments, an NmeCas9 nickase with an inactive HNH domain is used.
[0157] In some embodiments, conserved amino acids within the NmeCas9 nuclease domain are substituted to reduce or alter nuclease activity. Wild-type Cas9 has two nuclease domains: RuvC and HNH. The RuvC domain cleaves the non-target DNA strand, and the HNH domain cleaves the target strand of DNA. In some embodiments, the Cas9 nuclease comprises two or more RuvC domains or two or more HNH domains. In some embodiments, the Cas9 nuclease is wild-type Cas9. In some embodiments, the Cas9 can induce a double-stranded break in the target DNA. In certain embodiments, the Cas nuclease may cleave dsDNA, cleave one strand of dsDNA, or may not have DNA cleavase or nickase activity. In some embodiments, the NmeCas9 may comprise an amino acid substitution within the RuvC or RuvC-like nuclease domain. Exemplary amino acid substitutions in the RuvC or RuvC-like nuclease domain include H588A (based on the meningococcal Cas9 protein). In some embodiments, the Cas protein may include an amino acid substitution in the HNH or HNH-like nuclease domain. Exemplary amino acid substitutions in the HNH or HNH-like nuclease domain include D16A (based on the NmeCas9 protein).
[0158] In some embodiments, chimeric Cas proteins are used, where one domain or region of a protein is replaced by a portion of a different protein. In some embodiments, the NmeCas9 nuclease domain can be replaced with a domain from a different nuclease, such as Fok1. In some embodiments, the NmeCas9 protein can be a modified NmeCas9 nuclease.
[0159] In some embodiments, the nuclease can be modified to introduce a point mutation or a base change, for example, deamination.
[0160] In some embodiments, the Cas protein comprises a fusion protein comprising a Cas nuclease (e.g., NmeCas9), where the Cas nuclease is a nickase or is catalytically inactive and is linked to a heterologous functional domain. In some embodiments, the Cas protein comprises a fusion protein comprising a catalytically inactive Cas nuclease (e.g., NmeCas9) linked to a heterologous functional domain (see, e.g., WO2014152432). In some embodiments, the catalytically inactive Cas9 is from Neisseria meningitidis Cas9. In some embodiments, the catalytically inactive Cas comprises a mutation that inactivates the Cas.
[0161] In some embodiments, the heterologous functional domain is a domain that modifies gene expression, histone, or DNA.In some embodiments, the heterologous functional domain is a transcription activation domain or a transcription repression domain.In some embodiments, the nuclease is a catalytically inactive Cas nuclease, such as dCas9.
[0162] In some embodiments, the heterologous functional domain is a deaminase, such as cytidine deaminase or adenine deaminase. In certain embodiments, the heterologous functional domain is a C to T base conversion (cytidine deaminase), such as apolipoprotein B mRNA editing enzyme (APOBEC) deaminase. A heterologous functional domain, such as a deaminase, can be part of a fusion protein with a Cas nuclease with nickase activity, or a catalytically inactive Cas nuclease, as further described below.
[0163] In some embodiments, Nme Cas9 has double-stranded endonuclease activity.
[0164] In some embodiments, Nme Cas9 has nickase activity.
[0165] In some embodiments, Nme Cas9 comprises a dCas9 DNA binding domain.
[0166] In some embodiments, Nme Cas9 comprises an amino acid sequence that is at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 1 and 4-13, 220, 227, 234, 241, 248, 255, 262, 269, 276, 283, 290, 297, or 310-315 (as shown in Table 39A). In some embodiments, Nme Cas9 comprises an amino acid sequence of any one of SEQ ID NOs: 1 and 4-13, 220, 227, 234, 241, 248, 255, 262, 269, 276, 283, 290, 297, or 310-315 (as shown in Table 39A).
[0167] In some embodiments, the sequence encoding NmeCas9 comprises a nucleotide sequence at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 15, 18-27, 29, 32-41, 221-226, 228-233, 235-240, 242-247, 249-254, 256-261, 263-268, 270-275, 277-282, 284-289, 291-296, 298-303, 304-309, or 316-321 (as set forth in Table 39A). In some embodiments, the sequence encoding NmeCas9 comprises the nucleotide sequence of any one of SEQ ID NOs: 15, 18-27, 29, 32-41, 221-226, 228-233, 235-240, 242-247, 249-254, 256-261, 263-268, 270-275, 277-282, 284-289, 291-296, 298-303, 304-309, or 316-321 (as shown in Table 39A).
[0168] In some embodiments, any of the foregoing levels of identity is at least 95%, at least 98%, at least 99%, or 100%.
[0169] B. Exemplary Coding Sequences In any of the embodiments shown herein, the polynucleotide is an mRNA comprising an ORF encoding the RNA-guided DNA-binding agent disclosed above.In any of the embodiments described herein, the polynucleotide is an mRNA comprising an ORF encoding NmeCas9.In any of the embodiments described herein, the polynucleotide can be an expression construct comprising a promoter operably linked to an ORF encoding an RNA-guided DNA-binding agent (e.g., NmeCas9).
[0170] Certain ORFs are translated more efficiently in vivo than other ORFs in terms of polypeptide molecules produced per mRNA molecule.The codon pair usage of such efficiently translated ORFs can contribute to translation efficiency.Further description of ORF coding sequences, codon pair usage, and codon repeat content improvement is disclosed in WO2019 / 0067910 and WO2020 / 198641, the contents of each of which are incorporated herein by reference in their entirety.
[0171] For example, in some embodiments, at least 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons of the ORF are minimal adenine codons or minimal uridine codons. In some embodiments, the ORF comprises or consists of codons that increase translation of mRNA in mammals. In some embodiments, the ORF comprises or consists of codons that increase translation of mRNA in humans. Increased translation in a mammal, cell type, mammalian organ, human, human organ, etc. can be determined relative to the degree of translation of a wild-type sequence of the ORF, or relative to an ORF with a codon distribution that matches the codon distribution of the organism from which the ORF is derived or the organism that contains the most similar ORF at the amino acid level.
[0172] In some embodiments, the GC content of the ORF is 56% or more. In some embodiments, the GC content of the ORF is 56.5% or more. In some embodiments, the GC content of the ORF is 57% or more. In some embodiments, the GC content of the ORF is 57.5% or more. In some embodiments, the GC content of the ORF is 58% or more. In some embodiments, the GC content of the ORF is 58.5% or more. In some embodiments, the GC content of the ORF is 59% or more. In some embodiments, the GC content of the ORF is 63% or less. In some embodiments, the GC content of the ORF is 62.6% or less. In some embodiments, the GC content of the ORF is 62.1% or less. In some embodiments, the GC content of the ORF is 61.6% or less. In some embodiments, the GC content of the ORF is 61.1% or less. In some embodiments, the GC content of the ORF is 60.6% or less. In some embodiments, the GC content of the ORF is 60.1% or less.
[0173] In some embodiments, the ORF consists of a set of codons, where at least about 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons are codons listed in Table 1. [Table 2]
[0174] 1. Low uridine content ORF In some embodiments, an ORF encoding a polypeptide has a uridine content ranging from its minimum uridine content to about 150% of its minimum uridine content. In some embodiments, the uridine content of an ORF is about 145% or less, 140% or less, 135% or less, 130% or less, 125% or less, 120% or less, 115% or less, 110% or less, 105% or less, 104% or less, 103% or less, 102% or less, or 101% or less of its minimum uridine content. In some embodiments, an ORF has a uridine content equal to its minimum uridine content. In some embodiments, an ORF has a uridine content of about 150% or less of its minimum uridine content. In some embodiments, an ORF has a uridine content of about 145% or less of its minimum uridine content. In some embodiments, an ORF has a uridine content of about 140% or less of its minimum uridine content. In some embodiments, the ORF has a uridine content of about 135% or less of its minimum uridine content. In some embodiments, the ORF has a uridine content of about 130% or less of its minimum uridine content. In some embodiments, the ORF has a uridine content of about 125% or less of its minimum uridine content. In some embodiments, the ORF has a uridine content of about 120% or less of its minimum uridine content. In some embodiments, the ORF has a uridine content of about 115% or less of its minimum uridine content. In some embodiments, the ORF has a uridine content of about 110% or less of its minimum uridine content. In some embodiments, the ORF has a uridine content of about 105% or less of its minimum uridine content. In some embodiments, the ORF has a uridine content of about 104% or less of its minimum uridine content. In some embodiments, the ORF has a uridine content of about 103% or less of its minimum uridine content. In some embodiments, the ORF has a uridine content of about 102% or less of its minimum uridine content, hi some embodiments, the ORF has a uridine content of about 101% or less of its minimum uridine content.
[0175] In some embodiments, an ORF has a uridine dinucleotide content ranging from its minimum uridine dinucleotide content to 200% of its minimum uridine dinucleotide content. In some embodiments, the uridine dinucleotide content of an ORF is about 195% or less, 190% or less, 185% or less, 180% or less, 175% or less, 170% or less, 165% or less, 160% or less, 155% or less, 150% or less, 145% or less, 140% or less, 135% or less, 130% or less, 125% or less, 120% or less, 115% or less, 110% or less, 105% or less, 104% or less, 103% or less, 102% or less, or 101% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content equal to its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 200% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 195% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 190% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 185% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 180% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 175% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 170% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 165% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 160% or less of its minimum uridine dinucleotide content.In some embodiments, an ORF has a uridine dinucleotide content of about 155% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content equal to its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 150% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 145% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 140% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 135% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 130% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 125% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 120% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 115% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 110% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 105% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 104% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 103% or less of its minimum uridine dinucleotide content. In some embodiments, an ORF has a uridine dinucleotide content of about 102% or less of its minimum uridine dinucleotide content.In some embodiments, an ORF has a uridine dinucleotide content of less than or equal to about 101% of its minimum uridine dinucleotide content.
[0176] In some embodiments, an ORF has a uridine dinucleotide content ranging from its minimum uridine dinucleotide content to a uridine dinucleotide content that is 90% or less of the maximum uridine dinucleotide content of a reference sequence that encodes the same protein as the mRNA of interest. In some embodiments, the uridine dinucleotide content of an ORF is about 85% or less, 80% or less, 75% or less, 70% or less, 65% or less, 60% or less, 55% or less, 50% or less, 45% or less, 40% or less, 35% or less, 30% or less, 25% or less, 20% or less, 15% or less, 10% or less, or 5% or less of the maximum uridine dinucleotide content of a reference sequence that encodes the same protein as the mRNA of interest.
[0177] In some embodiments, the ORF has a uridine trinucleotide content ranging from 0 uridine trinucleotides to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, or 50 uridine trinucleotides (longer numbers of uridines are counted as the number of unique three uridine segments therein, e.g., a uridine tetranucleotide contains two uridine trinucleotides, a uridine pentanucleotide contains three uridine trinucleotides, etc.). In some embodiments, the ORF has a uridine trinucleotide content ranging from 0% uridine trinucleotides to 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 1.5%, or 2% uridine trinucleotides, where the uridine trinucleotide content percentage is calculated as the percentage of positions in the sequence that are occupied by uridines that form part of a uridine trinucleotide (or a longer number of uridines), such that the sequences UUUAAA and UUUUAAAA each have a uridine trinucleotide content of 50%. For example, in some embodiments, the ORF has a uridine trinucleotide content of 2% or less. For example, in some embodiments, the ORF has a uridine trinucleotide content of 1.5% or less. In some embodiments, the ORF has a uridine trinucleotide content of 1% or less. In some embodiments, the ORF has a uridine trinucleotide content of 0.9% or less. In some embodiments, the ORF has a uridine trinucleotide content of 0.8% or less. In some embodiments, the ORF has a uridine trinucleotide content of 0.7% or less. In some embodiments, the ORF has a uridine trinucleotide content of 0.6% or less. In some embodiments, the ORF has a uridine trinucleotide content of 0.5% or less. In some embodiments, the ORF has a uridine trinucleotide content of 0.4% or less. In some embodiments, the ORF has a uridine trinucleotide content of 0.3% or less.In some embodiments, the ORF has a uridine trinucleotide content of 0.2% or less. In some embodiments, the ORF has a uridine trinucleotide content of 0.1% or less. In some embodiments, the ORF has no uridine trinucleotides.
[0178] In some embodiments, an ORF has a uridine trinucleotide content ranging from its minimum uridine trinucleotide content to a uridine trinucleotide content that is 90% or less of the maximum uridine trinucleotide content of a reference sequence encoding the same protein as the polynucleotide of interest. In some embodiments, the uridine trinucleotide content of the ORF is about 85% or less, 80% or less, 75% or less, 70% or less, 65% or less, 60% or less, 55% or less, 50% or less, 45% or less, 40% or less, 35% or less, 30% or less, 25% or less, 20% or less, 15% or less, 10% or less, or 5% or less of the maximum uridine trinucleotide content of a reference sequence encoding the same protein as the polynucleotide of interest.
[0179] In some embodiments, the ORF has minimal nucleotide homopolymers, e.g., repeated strings of the same nucleotide. For example, in some embodiments, when selecting a minimal uridine codon from the codons listed in Table 2, the polynucleotide is constructed by selecting a minimal uridine codon that reduces the number and length of nucleotide homopolymers, e.g., selecting GCA instead of GCC for alanine, or selecting GGA instead of GGG for glycine, or selecting AAG instead of AAA for lysine.
[0180] A given ORF can be reduced in uridine content or uridine dinucleotide content or uridine trinucleotide content, for example, by using minimal uridine codons in a sufficient fraction of the ORF. For example, the amino acid sequences of the polypeptides encoded by the ORFs described herein can be back-translated into ORF sequences by converting the amino acids to codons, with some or all of the ORF using the exemplary minimal uridine codons shown below. In some embodiments, at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons in the ORF are codons listed in Table 2. [Table 3]
[0181] In some embodiments, the ORF consists of a set of codons, where at least about 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons are codons listed in Table 2.
[0182] 2. Low adenine content ORFs In some embodiments, an ORF has an adenine content ranging from its minimum adenine content to about 150% of its minimum adenine content. In some embodiments, an ORF's adenine content is about 145% or less, 140% or less, 135% or less, 130% or less, 125% or less, 120% or less, 115% or less, 110% or less, 105% or less, 104% or less, 103% or less, 102% or less, or 101% or less of its minimum adenine content. In some embodiments, an ORF has an adenine content equal to its minimum adenine content. In some embodiments, an ORF has an adenine content of about 150% or less of its minimum adenine content. In some embodiments, an ORF has an adenine content of about 145% or less of its minimum adenine content. In some embodiments, an ORF has an adenine content of about 140% or less of its minimum adenine content. In some embodiments, the ORF has an adenine content of about 135% or less of its minimum adenine content. In some embodiments, the ORF has an adenine content of about 130% or less of its minimum adenine content. In some embodiments, the ORF has an adenine content of about 125% or less of its minimum adenine content. In some embodiments, the ORF has an adenine content of about 120% or less of its minimum adenine content. In some embodiments, the ORF has an adenine content of about 115% or less of its minimum adenine content. In some embodiments, the ORF has an adenine content of about 110% or less of its minimum adenine content. In some embodiments, the ORF has an adenine content of about 105% or less of its minimum adenine content. In some embodiments, the ORF has an adenine content of about 104% or less of its minimum adenine content. In some embodiments, the ORF has an adenine content of about 103% or less of its minimum adenine content. In some embodiments, the ORF has an adenine content of about 102% or less of its minimum adenine content. In some embodiments, the ORF has an adenine content of about 101% or less of its minimum adenine content.
[0183] In some embodiments, an ORF has an adenine dinucleotide content ranging from its minimum adenine dinucleotide content to 200% of its minimum adenine dinucleotide content. In some embodiments, the adenine dinucleotide content of an ORF is about 195% or less, 190% or less, 185% or less, 180% or less, 175% or less, 170% or less, 165% or less, 160% or less, 155% or less, 150% or less, 145% or less, 140% or less, 135% or less, 130% or less, 125% or less, 120% or less, 115% or less, 110% or less, 105% or less, 104% or less, 103% or less, 102% or less, or 101% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content equal to its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 200% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 195% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 190% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 185% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 180% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 175% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 170% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 165% or less of its minimum adenine dinucleotide content, hi some embodiments, an ORF has an adenine dinucleotide content of about 160% or less of its minimum adenine dinucleotide content.In some embodiments, an ORF has an adenine dinucleotide content of about 155% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content equal to its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 150% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 145% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 140% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 135% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 130% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 125% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 120% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 115% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 110% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 105% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 104% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of about 103% or less of its minimum adenine dinucleotide content. In some embodiments, an ORF has an adenine dinucleotide content of less than or equal to about 102% of its minimum adenine dinucleotide content.In some embodiments, an ORF has an adenine dinucleotide content of about 101% or less of its minimum adenine dinucleotide content.
[0184] In some embodiments, the ORF has an adenine dinucleotide content ranging from its minimum adenine dinucleotide content to an adenine dinucleotide content that is 90% or less of the maximum adenine dinucleotide content of a reference sequence encoding the same protein as the polynucleotide of interest. In some embodiments, the adenine dinucleotide content of the ORF is about 85% or less, 80% or less, 75% or less, 70% or less, 65% or less, 60% or less, 55% or less, 50% or less, 45% or less, 40% or less, 35% or less, 30% or less, 25% or less, 20% or less, 15% or less, 10% or less, or 5% or less of the maximum adenine dinucleotide content of a reference sequence encoding the same protein as the polynucleotide of interest.
[0185] In some embodiments, an ORF has an adenine trinucleotide content ranging from 0 adenine trinucleotides to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, or 50 adenine trinucleotides (longer numbers of adenines count as the number of unique three-adenine segments therein, e.g., an adenine tetranucleotide contains two adenine trinucleotides, an adenine pentanucleotide contains three adenine trinucleotides, etc.). In some embodiments, the ORF has an adenine trinucleotide content ranging from 0% adenine trinucleotide to 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 1.5%, or 2% adenine trinucleotide, where the adenine trinucleotide content percentage is calculated as the percentage of positions in the sequence that are occupied by adenines that form part of an adenine trinucleotide (or a longer number of adenines), and thus the sequences UUUAAA and UUUUAAAA each have an adenine trinucleotide content of 50%. For example, in some embodiments, the ORF has an adenine trinucleotide content of 2% or less. For example, in some embodiments, the ORF has an adenine trinucleotide content of 1.5% or less. In some embodiments, the ORF has an adenine trinucleotide content of 1% or less. In some embodiments, the ORF has an adenine trinucleotide content of 0.9% or less. In some embodiments, the ORF has an adenine trinucleotide content of 0.8% or less. In some embodiments, the ORF has an adenine trinucleotide content of 0.7% or less. In some embodiments, the ORF has an adenine trinucleotide content of 0.6% or less. In some embodiments, the ORF has an adenine trinucleotide content of 0.5% or less. In some embodiments, the ORF has an adenine trinucleotide content of 0.4% or less. In some embodiments, the ORF has an adenine trinucleotide content of 0.3% or less.In some embodiments, the ORF has an adenine trinucleotide content of 0.2% or less. In some embodiments, the ORF has an adenine trinucleotide content of 0.1% or less. In some embodiments, the ORF has no adenine trinucleotides.
[0186] In some embodiments, the ORF has an adenine trinucleotide content ranging from its minimum adenine trinucleotide content to an adenine trinucleotide content that is 90% or less of the maximum adenine trinucleotide content of the reference sequence that encodes the same protein as the polynucleotide of interest.In some embodiments, the adenine trinucleotide content of the ORF is about 85% or less, 80% or less, 75% or less, 70% or less, 65% or less, 60% or less, 55% or less, 50% or less, 45% or less, 40% or less, 35% or less, 30% or less, 25% or less, 20% or less, 15% or less, 10% or less, or 5% or less of the maximum adenine trinucleotide content of the reference sequence that encodes the same protein as the polynucleotide of interest.In some embodiments, the ORF has a minimum nucleotide homopolymer, e.g., a repeated string of the same nucleotide. For example, in some embodiments, when selecting a minimal adenine codon from the codons listed in Table 3, a polynucleotide is constructed by selecting a minimal adenine codon that reduces the number and length of nucleotide homopolymers, for example, selecting GCA instead of GCC for alanine, or selecting GGA instead of GGG for glycine, or selecting AAG instead of AAA for lysine. A given ORF can be reduced in adenine content or adenine dinucleotide content or adenine trinucleotide content, for example, by using a minimal adenine codon in a sufficient fraction of the ORF. For example, the amino acid sequence of a polypeptide encoded by an ORF described herein can be back-translated into an ORF sequence by converting the amino acids to codons, with some or all of the ORF using the exemplary minimal adenine codons shown below. In some embodiments, at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons in the ORF are codons listed in Table 3. [Table 4]
[0187] In some embodiments, the ORF consists of a set of codons, where at least about 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons are codons listed in Table 3.
[0188] 3. ORFs with low adenine and uridine content To the extent practicable, any of the features described above in relation to low adenine content may be combined with any of the features described above in relation to low uridine content. For example, an ORF has a uridine content ranging from its minimum uridine content to about 150% of its minimum uridine content (e.g., the uridine content of an ORF is about 145% or less, 140% or less, 135% or less, 130% or less, 125% or less, 120% or less, 115% or less, 110% or less, 105% or less, 104% or less, 103% or less, 102% or less, or 101% or less of its minimum uridine content), and an adenine content ranging from its minimum adenine content to about 150% of its minimum adenine content (e.g., about 145% or less, 140% or less, 135% or less, 130% or less, 125% or less, 120% or less, 115% or less, 110% or less, 105% or less, 104% or less, 103% or less, 102% or less, or 101% or less of its minimum uridine content). The same applies to uridine and adenine dinucleotides. Similarly, the content of uridine nucleotides and adenine dinucleotides in the ORF may be as described above. Similarly, the content of uridine dinucleotides and adenine dinucleotides in the ORF may be as described above.
[0189] A given ORF can be reduced in uridine and adenine nucleotide or dinucleotide content, for example, by using minimal uridine and adenine codons in a sufficient fraction of the ORF. For example, the amino acid sequences of the polypeptides encoded by the ORFs described herein can be reverse translated into ORF sequences by converting the amino acids to codons, with some or all of the ORF using the exemplary minimal uridine and adenine codons shown below. In some embodiments, at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons in the ORF are codons listed in Table 4. [Table 5]
[0190] In some embodiments, the ORF is composed of a set of codons where at least about 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons are codons listed in Table 4. As can be seen from Table 4, each of the three serine codons listed contains either one A or one U. In some embodiments, uridine minimization is preferred by using the AGC codon for serine. In some embodiments, adenine minimization is preferred by using the UCC or UCG codon for serine.
[0191] 4. Codons that increase translation or correspond to highly expressed tRNAs, exemplary codon sets In some embodiments, the ORF has codons that increase translation in a mammal, such as a human. In further embodiments, the ORF has codons that increase translation in a mammal, such as a human organ, such as the liver. In further embodiments, the ORF has codons that increase translation in a mammal, such as a human cell type, such as a hepatocyte. The increase in translation in a mammal, cell type, mammalian organ, human, human organ, etc. can be determined relative to the degree of translation of a wild-type sequence of the ORF, or relative to an ORF that has a codon distribution that matches the codon distribution of the organism from which the ORF is derived or the organism that contains the most similar ORF at the amino acid level.
[0192] In some embodiments, the polypeptide encoded by the ORF is a Cas9 nuclease derived from a prokaryote, as described below, and the increase in translation in a mammal, cell type, mammalian organ, human, human organ, etc. can be determined relative to the degree of translation of a wild-type sequence of the ORF, or relative to an ORF of interest, such as an ORF encoding a human protein or transgene for expression in a human cell. For example, the ORF may be an ORF that has a codon distribution that matches that of the organism from which the ORF is derived or that contains the most similar ORF at the amino acid level, such as N. meningitidis, or all else is equal, including any applicable point mutations, heterologous domains, etc., compared to a translation of the Cas9 ORF contained in SEQ ID NOs: 29, 32-41, 224-226, 231-233, 238-240, 245-247, 252-254, 259-261, 266-268, 273-275, 280-282, 287-289, 294-296, 301-303, or 316-321. Codons useful for increasing expression in humans, including human liver and human hepatocytes, may be codons that correspond to highly expressed tRNAs in human liver / hepatocytes, as described in Dittmar KA, PLOS Genetics 2(12):e221 (2006). In some embodiments, at least about 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in the ORF are codons that correspond to highly expressed tRNAs (e.g., the most highly expressed tRNA for each amino acid) in a mammal, such as a human. In some embodiments, at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in the ORF are codons that correspond to highly expressed tRNAs (e.g., the most highly expressed tRNA for each amino acid) in a mammalian organ, such as a human organ.In some embodiments, at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in the ORF are codons that correspond to tRNAs that are highly expressed in mammalian liver, such as human liver (e.g., the most highly expressed tRNA for each amino acid). In some embodiments, at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in the ORF are codons that correspond to tRNAs that are highly expressed in mammalian liver cells, such as human hepatocytes (e.g., the most highly expressed tRNA for each amino acid).
[0193] Alternatively, codons corresponding to tRNAs that are highly expressed throughout organisms (eg, humans) may be used.
[0194] Any of the aforementioned approaches to codon selection can be combined with the selection of codons that contribute to lower repeat content, as set forth above, or the use of the codon sets of Table 1, as set forth above, e.g., the use of the minimal uridine or adenine codons of Tables 2, 3, or 4, and then, if more than one option is available, the use of the codon that corresponds to the more highly expressed tRNA, either in the organism in general (e.g., human), or in the organ or cell type of interest, e.g., liver or hepatocytes (e.g., human liver or human hepatocytes).
[0195] C. Nuclear localization signal (NLS) The nuclear localization signal (NLS) disclosed herein may facilitate the transport of the RNA-guided DNA binder to the cell nucleus. The first NLS disclosed herein, and if present, the second NLS, may be N-terminally linked to the RNA-guided DNA binder sequence, i.e., the RNA-guided DNA binder is a C-terminal domain in the encoded polypeptide. The first NLS disclosed herein, and if present, the second NLS, may be N-terminally linked to the NmeCas9 coding sequence. An additional NLS may be linked to the N-terminus of the NmeCas9 coding sequence. In some embodiments, the encoded polypeptide comprises three NLSs at the N-terminus of the NmeCas9 coding sequence. In some embodiments, at least one NLS is provided at the C-terminus of the RNA-guided DNA binder sequence (e.g., with or without an intervening spacer between the NLS and the preceding domain). In some embodiments, the first NLS and the second NLS are provided at the C-terminus of the RNA-guided DNA binder sequence (e.g., with or without an intervening spacer between the NLS and the preceding domain).
[0196] Thus, in some embodiments, an ORF encoding a polypeptide disclosed herein comprises a coding sequence for a first NLS and a coding sequence for a second NLS such that the encoded first and second NLSs are located N-terminal to the NmeCas9 polypeptide, hi some embodiments, the ORF further comprises a coding sequence for a third NLS C-terminal to the ORF encoding NmeCas9.
[0197] In some embodiments, the NLS may be a single part sequence, such as, for example, the SV40 NLS, PKKKRKV (SEQ ID NO: 388) or PKKKRRV (SEQ ID NO: 421). In some embodiments, the NLS may be a bipartite sequence, such as the NLS of nucleoplasmin, KRPAATKKAGQAKKKK (SEQ ID NO: 422). In some embodiments, the NLS sequence may include LAAKRSRTT (SEQ ID NO: 410), QAAKRSRTT (SEQ ID NO: 411), PAPAKRERTT (SEQ ID NO: 412), QAAKRPRTT (SEQ ID NO: 413), RAAKRPRTT (SEQ ID NO: 414), AAAKRSWSMAA (SEQ ID NO: 415), AAAKRVWSMAF (SEQ ID NO: 416), AAAKRSWSMAF (SEQ ID NO: 417), AAAKRKYFAA (SEQ ID NO: 418), RAAKRKAFAA (SEQ ID NO: 419), or RAAKRKYFAV (SEQ ID NO: 420). The NLS may be a snurportin-1 importin-β (IBB domain, e.g., the SPN1-impβ sequence; see Huber et al., 2002, J. Cell Bio., 156, 467-479. In certain embodiments, it is a single PKKKRKV (SEQ ID NO: 388). In some embodiments, the first and second NLSs are independently selected from an SV40 NLS, a nucleoplasmin NLS, a bipartite NLS, a c-myc-like NLS, and an NLS comprising the sequence KTRAD. In certain embodiments, the first and second NLSs may be the same (e.g., two SV40 NLSs). In certain embodiments, the first and second NLSs may be different.
[0198] In some embodiments, the first NLS is an SV40 NLS and the second NLS is a nucleoplasmin NLS.
[0199] In some embodiments, the SV40 NLS comprises the sequence of PKKKRKVE (SEQ ID NO: 383) or KKKRKVE (SEQ ID NO: 384). In some embodiments, the nucleoplasmin NLS comprises the sequence of KRPAATKKAGQAKKKK (SEQ ID NO: 422). In some embodiments, the bisecting NLS comprises the sequence of KRTADGSEFESPKKKRKVE (SEQ ID NO: 385). In some embodiments, the c-myc-like NLS comprises the sequence of PAAKKKKLD (SEQ ID NO: 386).
[0200] In some embodiments, one or more NLSs according to any of the preceding embodiments are present in an RNA-guided DNA binding agent in combination with one or more additional heterologous functional domains, such as any of the heterologous functional domains described below.
[0201] D. Other heterologous functional domains In some embodiments, the polypeptides (e.g., RNA-guided DNA binding agents) encoded by the ORFs described herein comprise (e.g., are or comprise) one or more additional heterologous functional domains. In some embodiments, the ORFs further comprise a nucleotide sequence encoding one or more additional heterologous functional domains.
[0202] In some embodiments, the heterologous functional domain may alter the intracellular half-life of the RNA-guided DNA binding agent. In some embodiments, the RNA-guided DNA binding agent may be increased in half-life. In some embodiments, the RNA-guided DNA binding agent may be decreased in half-life. In some embodiments, the heterologous functional domain may increase the stability of the RNA-guided DNA binding agent. In some embodiments, the heterologous functional domain may decrease the stability of the RNA-guided DNA binding agent. In some embodiments, the heterologous functional domain may act as a signal peptide for protein degradation. In some embodiments, the protein degradation may be mediated by proteolytic enzymes, such as, for example, proteasomes, lysosomal proteases, or calpain proteases. In some embodiments, the heterologous functional domain may include a PEST sequence. In some embodiments, the RNA-guided DNA binding agent may be modified by adding ubiquitin or polyubiquitin chains. In some embodiments, the ubiquitin may be a ubiquitin-like protein (UBL). Non-limiting examples of ubiquitin-like proteins include small ubiquitin-like modifiers (SUMO), ubiquitin cross-reactive protein (UCRP, also known as interferon-stimulated gene-15 (ISG15)), ubiquitin-related modifier 1 (URM1), neuronal-precursor-cell-expressed developmentally downregulated protein-8 (NEDD8, also known as Rub1 in S. cerevisiae), human leukocyte antigen F-related (FAT10), autophagy-8 (ATG8) and autophagy-12 (ATG12), Fau ubiquitin-like protein (FUB1), membrane-anchored UBL (MUB), ubiquitin fold modifier-1 (UFM1), and ubiquitin-like protein-5 (UBL5).
[0203] In some embodiments, the heterologous functional domain may be a marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, epitope tags, and reporter gene sequences. In some embodiments, the marker domain may be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, sfGFP, EGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalamal, GFPuv, Sapphire, T-sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyan1, Midoriishi-Cyan), red fluorescent proteins (e.g., mKate, mKate2, mPlum, DsRed), and the like. Monomeric, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRasberry, mStrawberry, Jred) and orange fluorescent protein (mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, tdTomato), or any other suitable fluorescent protein. In another embodiment, the marker domain can be a purification tag or an epitope tag.Non-limiting exemplary tags include glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein (MBP), thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, 8xHis, biotin carboxyl carrier protein (BCCP), poly-His, calmodulin, and HiBiT. Non-limiting exemplary reporter genes include glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, luciferase or fluorescent proteins.
[0204] In additional embodiments, the heterologous functional domain may target the RNA-guided DNA binding agent to a given organelle, cell type, tissue or organ.
[0205] In further embodiments, the heterologous functional domain can be an effector domain. When an RNA-guided DNA binding agent is guided to its target sequence, for example, when a Cas nuclease is guided to the target sequence by gRNA, the effector domain can modify or affect the target sequence. In some embodiments, the effector domain can be selected from a nucleic acid binding domain, a nuclease domain (e.g., a non-Cas nuclease domain), an epigenetic modification domain, a transcription activation domain, or a transcription repressor domain. In some embodiments, the heterologous functional domain is a nuclease, such as FokI nuclease. See, for example, U.S. Patent No. 9,023,649. In some embodiments, the heterologous functional domain is a transcription activator or a transcription repressor. See, for example, Qi et al., "Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression", Cell 152:1173-83 (2013); Perez-Pinera et al., "RNA-guided gene activation by CRISPR-Cas9-based transcription factors", Nat. Methods 10:973-6 (2013); Mali et al., "CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering", Nat. Biotechnol. 31:833-8 (2013); Gilbert et al., "CRISPR-mediated modular RNA-guided regulation of transcription in eukaryotes", Cell 154:442-51 (2013). Thus, RNA-guided DNA binding agents essentially become transcription factors that can be induced to bind to desired target sequences using guide RNAs.In certain embodiments, the DNA modifying domain is a methylation domain, such as a demethylation domain or a methyltransferase domain.In certain embodiments, the effector domain is a DNA modifying domain, such as a base editing domain.In certain embodiments, the DNA modifying domain is a nucleic acid editing domain, such as a deaminase domain, that introduces specific modifications into DNA, which will be described further below.
[0206] Linker In some embodiments, the ORF further comprises a nucleotide sequence encoding a linker sequence between the first NLS and the second NLS.
[0207] In some embodiments, the ORF further comprises a nucleotide sequence encoding a linker sequence between the Nme Cas9 coding sequence and the NLS adjacent to the Nme Cas9 coding sequence.
[0208] In some embodiments, the spacer comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more amino acids. In some embodiments, the spacer comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids.
[0209] In some embodiments, the peptide linker is a 16 residue "XTEN" linker, or a variant thereof (see, e.g., the Examples, and Schellenberger et al. A recombinant polypeptide extends the in vivo half-life of peptides and proteins in a tunable manner. Nat. Biotechnol. 27, 1186-1190 (2009)). In some embodiments, the XTEN linker comprises the sequence SGSETPGTSESATPES (SEQ ID NO: 58), SGSETPGTSESA (SEQ ID NO: 59), or SGSETPGTSESATPEGGSGGS (SEQ ID NO: 60).
[0210] In some embodiments, the peptide linker is (GGGGS) n (SEQ ID NO: 62), (G) n , (EAAAK) n (SEQ ID NO: 63), (GGS) n , (SEQ ID NO: 61), or the SGSETPGTSESATPES (SEQ ID NO: 58) motif (see, e.g., Guilinger JP, Thompson DB, Liu D R. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014;32(6):577-82, the entire contents of which are incorporated herein by reference), or (XP) n motif, or any combination thereof, wherein n is independently an integer from 1 to 30. See WO2015089406, e.g., paragraph
[0012] , the entire contents of which are incorporated herein by reference.
[0211] In some embodiments, the peptide linker comprises one or more sequences selected from SEQ ID NOs: 61-122.
[0212] E.UTR, Kozak sequence In some embodiments, the polynucleotide comprises at least one UTR from hydroxysteroid 17-beta dehydrogenase 4 (HSD17B4 or HSD), e.g., a 5'UTR from HSD. In some embodiments, the polynucleotide comprises at least one UTR from a globin mRNA, e.g., human alpha globin (HBA) mRNA, human beta globin (HBB) mRNA, or Xenopus laevis beta globin (XBG) mRNA. In some embodiments, the polynucleotide comprises a 5'UTR, a 3'UTR, or a 5' and a 3'UTR from a globin mRNA, e.g., HBA, HBB, or XBG. In some embodiments, the polynucleotide comprises a 5'UTR from bovine growth hormone, cytomegalovirus (CMV), mouse Hba-a1, HSD, albumin gene, HBA, HBB, or XBG. In some embodiments, the polynucleotide comprises a 3'UTR from bovine growth hormone, cytomegalovirus, mouse Hba-a1, HSD, albumin gene, HBA, HBB, or XBG. In some embodiments, the polynucleotide comprises a 5' and 3'UTR from bovine growth hormone, cytomegalovirus, mouse Hba-a1, HSD, albumin gene, HBA, HBB, XBG, heat shock protein 90 (Hsp90), glyceraldehyde 3-phosphate dehydrogenase (GAPDH), beta-actin, alpha-tubulin, tumor protein (p53), or epidermal growth factor receptor (EGFR).
[0213] In some embodiments, the polynucleotide comprises 5' and 3' UTRs from the same source, for example, a constitutively expressed mRNA such as actin, albumin, or a globin such as HBA, HBB, or XBG.
[0214] In some embodiments, the polynucleotides disclosed herein comprise a 5'UTR having at least 90% identity to any one of SEQ ID NOs: 391-398. In some embodiments, the polynucleotides disclosed herein comprise a 3'UTR having at least 90% identity to any one of SEQ ID NOs: 399-406. In some embodiments, any of the aforementioned levels of identity are at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, the mRNAs disclosed herein comprise a 5'UTR having the sequence of any one of SEQ ID NOs: 391-398. In some embodiments, the polynucleotides disclosed herein comprise a 3'UTR having the sequence of any one of SEQ ID NOs: 399-406.
[0215] In some embodiments, the polynucleotide does not include a 5'UTR, e.g., there are no additional nucleotides between the 5'cap and the start codon. In some embodiments, the mRNA includes a Kozak sequence (described below) between the 5'cap and the start codon, but does not have any additional 5'UTR. In some embodiments, the mRNA does not include a 3'UTR, e.g., there are no additional nucleotides between the stop codon and the polyA tail.
[0216] In some embodiments, the mRNA comprises a Kozak sequence. The Kozak sequence can affect translation initiation and the overall yield of polypeptides translated from the mRNA. The Kozak sequence includes a methionine codon that can function as an initiation codon. The minimal Kozak sequence is NNNRUGN, where at least one of the following is true: the first N is A or G, and the second N is G. In the context of a nucleotide sequence, R denotes a purine (A or G). In some embodiments, the Kozak sequence is RNNRUGN, NNNRUGG, RNNRUGG, RNNAUGN, NNNAUGG, or RNNAUGG. In some embodiments, the Kozak sequence is rccRUGg, with zero mismatches or with a maximum of one or two mismatches to the lowercase positions. In some embodiments, the Kozak sequence is rccAUGg, with zero mismatches or with a maximum of one or two mismatches to the lowercase positions. In some embodiments, the Kozak sequence is gccRccAUGG (nucleotides 4-13 of SEQ ID NO:408; SEQ ID NO:407) with zero mismatches or with a maximum of one, two, or three mismatches to the lowercase positions. In some embodiments, the Kozak sequence is gccAccAUG with zero mismatches or with a maximum of one, two, three, or four mismatches to the lowercase positions. In some embodiments, the Kozak sequence is GCCACCAUG. In some embodiments, the Kozak sequence is gccgccRccAUGG (SEQ ID NO:408) with zero mismatches or with a maximum of one, two, three, or four mismatches to the lowercase positions.
[0217] 5' Cap In some embodiments, a polynucleotide (eg, an mRNA) disclosed herein includes a 5' cap, e.g., Cap0, Cap1, or Cap2.
[0218] The 5' cap is generally a 7-methylguanine ribonucleotide (which may be further modified, e.g., as described below for ARCA) linked via a 5'-triphosphate to the 5' position of the first nucleotide (i.e., the first cap-proximal nucleotide) of the 5'-3' strand of the nucleic acid. In Cap0, the riboses of the first and second cap-proximal nucleotides of the mRNA both contain 2'-hydroxyl. In Cap1, the riboses of the first and second transcribed nucleotides of the mRNA both contain 2'-methoxy and 2'-hydroxyl, respectively. In Cap2, the riboses of the first and second cap-proximal nucleotides of the mRNA both contain 2'-methoxy. See, e.g., Katibah et al. (2014) Proc Natl Acad Sci USA 111(33):12025-30; Abbas et al. (2017) Proc Natl Acad Sci USA 114(11):E2106-E2115. Most endogenous higher eukaryotic nucleic acids, including mammalian nucleic acids such as human nucleic acids, contain Cap1 or Cap2. Cap0, as well as other cap structures distinct from Cap1 and Cap2, can be immunogenic in mammals, such as humans, because they are recognized as "non-self" by components of the innate immune system, such as IFIT-1 and IFIT-5, which can result in increased concentrations of cytokines, such as type I interferons. Components of the innate immune system, such as IFIT-1 and IFIT-5, can also compete with eIF4E for binding to nucleic acids with caps other than Cap1 or Cap2, potentially inhibiting translation of the nucleic acid.
[0219] A cap can be included co-transcriptionally. For example, ARCA (anti-reverse cap analog; Thermo Fisher Scientific catalog no. AM8045) is a cap analog containing 7-methylguanine 3'-methoxy-5'-triphosphate linked to the 5' position of a guanine ribonucleotide that can be incorporated into transcripts during transcription initiation in vitro. ARCA results in a Cap0 cap or a Cap0-like cap in which the 2' position of the first cap-proximal nucleotide is a hydroxyl. See, e.g., Stepinski et al., (2001) "Synthesis and properties of mRNAs containing the novel 'anti-reverse' cap analogs 7-methyl(3'-O-methyl)GpppG and 7-methyl(3'deoxy)GpppG," RNA 7:1486-1495. The structure of ARCA is shown below. [ka]
[0220] CleanCap™ AG (m7G(5')ppp(5')(2'OMeA)pG; TriLink Biotechnologies catalog number N-7113) or CleanCap™ GG (m7G(5')ppp(5')(2'OMeG)pG, TriLink Biotechnologies catalog number N-7133) can be used to provide the Cap1 structure co-transcriptionally. 3'-O-methylated versions of CleanCap™ AG and CleanCap™ GG are also available from TriLink Biotechnologies as catalog numbers N-7413 and N-7433, respectively. The structure of CleanCap™ AG is shown below. CleanCap™ structures are sometimes referred to herein using the last three digits of the above catalog numbers (e.g., "CleanCap™ 113" for TriLink Biotechnologies catalog number N-7113). [ka]
[0221] Alternatively, a cap can be added to RNA post-transcriptionally. For example, vaccinia capping enzyme is commercially available (New England Biolabs Catalog No. M2080S), which has RNA triphosphatase and guanylyltransferase activities provided by its D1 subunit, and guanine methyltransferase provided by its D12 subunit. Thus, in the presence of S-adenosylmethionine and GTP, 7-methylguanine can be added to RNA to give Cap0. See, e.g., Guo, P. and Moss, B. (1990) Proc. Natl. Acad. Sci. USA 87, 4023-4027; Mao, X. and Shuman, S. (1994) J. Biol. Chem. 269, 24472-24479. For additional discussion of caps and capping approaches, see, e.g., WO2017 / 053297 and Ishikawa et al., Nucl. Acids. Symp. Ser. (2009) No. 53, 129-130.
[0222] F. Poly A tail In some embodiments, the polynucleotide is an mRNA encoding a polypeptide disclosed herein that comprises an ORF, wherein the mRNA further comprises a poly-adenylation (polyA) tail.
[0223] In some embodiments, the polynucleotides disclosed herein further comprise a polyA tail sequence or a polyadenylation signal sequence. In some embodiments, the polyA tail sequence comprises 100 to 400 nucleotides.
[0224] In some embodiments, the polyA sequence comprises non-adenine nucleotides. In some cases, the polyA tail is "interrupted" with one or more non-adenine nucleotide "anchors" at one or more positions within the polyA tail. The polyA tail may comprise at least eight consecutive adenine nucleotides, but may also comprise one or more non-adenine nucleotides. As used herein, "non-adenine nucleotides" refers to any natural or non-natural nucleotide that does not contain adenine. Guanine, thymine, and cytosine nucleotides are exemplary non-adenine nucleotides. Thus, the polyA tail on an mRNA described herein may comprise consecutive adenine nucleotides located 3' to the nucleotides encoding the polypeptides disclosed herein. In some cases, the polyA tail of an mRNA comprises non-consecutive adenine nucleotides located 3' to the nucleotides encoding the RNA-guided DNA binding agent or sequence of interest, where the non-adenine nucleotides interrupt the adenine nucleotides at regular or irregular intervals.
[0225] In some embodiments, the polyA tail is encoded by the plasmid used for in vitro transcription of the mRNA and becomes part of the transcription product. The polyA sequence encoded by the plasmid, i.e., the number of consecutive adenine nucleotides in the polyA sequence, may not be exact, e.g., 100 polyA sequences in the plasmid may not result in exactly 100 polyA sequences in the transcribed mRNA. In some embodiments, the polyA tail is not encoded by the plasmid and is added by PCR tailing or enzymatic tailing, e.g., using E. coli poly(A) polymerase.
[0226] In some embodiments, one or more non-adenine nucleotides are positioned to interrupt consecutive adenine nucleotides such that poly(A) binding protein can bind to a series of consecutive adenine nucleotides. In some embodiments, one or more non-adenine nucleotides are positioned after at least 8, 9, 10, 11, or 12 consecutive adenine nucleotides. In some embodiments, one or more non-adenine nucleotides are positioned after at least 8-50 consecutive adenine nucleotides. In some embodiments, one or more non-adenine nucleotides are positioned after at least 8-100 consecutive adenine nucleotides. In some embodiments, the non-adenine nucleotides are positioned after 1, 2, 3, 4, 5, 6, or 7 adenine nucleotides followed by at least 8 consecutive adenine nucleotides.
[0227] A polyA tail of the present disclosure may comprise a sequence of consecutive adenine nucleotides, followed by one or more non-adenine nucleotides, optionally followed by an additional adenine nucleotide.
[0228] In some embodiments, the poly-A tail comprises or contains one non-adenine nucleotide or one contiguous string of 2-10 non-adenine nucleotides. In some embodiments, the non-adenine nucleotide(s) is / are positioned after at least 8, 9, 10, 11, or 12 contiguous adenine nucleotides. In some cases, the one or more non-adenine nucleotides is / are positioned after at least 8-50 contiguous adenine nucleotides. In some embodiments, the one or more non-adenine nucleotides are positioned after at least 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 consecutive adenine nucleotides.
[0229] In some embodiments, the non-adenine nucleotides are guanine, cytosine, or thymine. In some cases, the non-adenine nucleotides are guanine nucleotides. In some embodiments, the non-adenine nucleotides are cytosine nucleotides. In some embodiments, the non-adenine nucleotides are thymine nucleotides. In some cases, when there are more than one non-adenine nucleotides, the non-adenine nucleotides may be selected from a) guanine and thymine nucleotides, b) guanine and cytosine nucleotides, c) thymine and cytosine nucleotides, or d) guanine, thymine, and cytosine nucleotides. An exemplary polyA tail comprising non-adenine nucleotides is provided as SEQ ID NO:409.
[0230] In some embodiments, the polyA tail sequence comprises the sequence of SEQ ID NO:409.
[0231] G. Modified Nucleotides In some embodiments, a nucleic acid comprising an ORF encoding a polypeptide disclosed herein comprises modified uridines at some or all of the uridine positions.
[0232] In some embodiments, the modified uridine is a uridine modified at the 5-position, e.g., with a halogen or a C1-C3 alkoxy. In some embodiments, the modified uridine is a pseudouridine modified at the 1-position, e.g., with a C1-C3 alkyl. The modified uridine can be, for example, pseudouridine, N1-methyl-pseudouridine, 5-methoxyuridine, 5-iodouridine, or a combination thereof. In some embodiments, the modified uridine is 5-methoxyuridine. In some embodiments, the modified uridine is 5-iodouridine. In some embodiments, the modified uridine is pseudouridine. In some embodiments, the modified uridine is N1-methyl-pseudouridine. In some embodiments, the modified uridine is a combination of pseudouridine and N1-methyl-pseudouridine. In some embodiments, the modified uridine is a combination of pseudouridine and 5-methoxyuridine. In some embodiments, the modified uridine is a combination of N1-methylpseudouridine and 5-methoxyuridine. In some embodiments, the modified uridine is a combination of 5-iodouridine and N1-methyl-pseudouridine. In some embodiments, the modified uridine is a combination of pseudouridine and 5-iodouridine. In some embodiments, the modified uridine is a combination of 5-iodouridine and 5-methoxyuridine.
[0233] In some embodiments, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the uridine positions in a polynucleotide according to the disclosure are modified uridines. In some embodiments, at least 10% of the uridines at the uridine positions in a polynucleotide according to the disclosure are substituted with modified uridines. In some embodiments, at least 20% of the uridines at the uridine positions in a polynucleotide according to the disclosure are substituted with modified uridines. In some embodiments, at least 30% of the uridines at the uridine positions in a polynucleotide according to the disclosure are substituted with modified uridines. In some embodiments, at least 80% of the uridines at the uridine positions in a polynucleotide according to the disclosure are substituted with modified uridines. In some embodiments, at least 90% of the uridines at the uridine positions in a polynucleotide according to the disclosure are substituted with modified uridines.
[0234] In some embodiments, between 10% and 25%, between 15% and 25%, between 25% and 35%, between 35% and 45%, between 45% and 55%, between 55% and 65%, between 65% and 75%, between 75% and 85%, between 85% and 95%, or between 90% and 100% of the uridine positions in a polynucleotide of the disclosure are modified uridines. In some embodiments, between 15% and 45% of the uridines at the uridine positions in a polynucleotide of the disclosure are substituted with modified uridines.
[0235] In some embodiments, 100% of the uridines at the uridine positions in a polynucleotide according to the disclosure are substituted with modified uridines.
[0236] In some embodiments, the modified uridine is one or more of N1-methyl-pseudouridine, pseudouridine, 5-methoxyuridine, or 5-iodouridine, or a combination thereof. In some embodiments, the modified uridine is one or both of N1-methyl-pseudouridine or 5-methoxyuridine. In some embodiments, the modified uridine is N1-methyl-pseudouridine. In some embodiments, the modified uridine is 5-methoxyuridine.
[0237] In some embodiments, 10%-25%, 15%-25%, 25%-35%, 35%-45%, 45%-55%, 55%-65%, 65%-75%, 75%-85%, 85%-95%, or 90%-100% of the uridine positions in a polynucleotide of the disclosure are 5-methoxyuridine. In some embodiments, 10%-25%, 15%-25%, 25%-35%, 35%-45%, 45%-55%, 55%-65%, 65%-75%, 75%-85%, 85%-95%, or 90%-100% of the uridine positions in a polynucleotide of the disclosure are pseudouridine. In some embodiments, 10%-25%, 15%-25%, 25%-35%, 35%-45%, 45%-55%, 55%-65%, 65%-75%, 75%-85%, 85%-95%, or 90%-100% of the uridine positions in a polynucleotide of the disclosure are N1-methylpseudouridine. In some embodiments, 10%-25%, 15%-25%, 25%-35%, 35%-45%, 45%-55%, 55%-65%, 65%-75%, 75%-85%, 85%-95%, or 90%-100% of the uridine positions in a polynucleotide of the disclosure are 5-iodouridine. In some embodiments, 10%-25%, 15%-25%, 25%-35%, 35%-45%, 45%-55%, 55%-65%, 65%-75%, 75%-85%, 85%-95%, or 90%-100% of the uridine positions in a polynucleotide of the disclosure are 5-methoxyuridine and the remainder are N1-methylpseudouridine. In some embodiments, 10%-25%, 15%-25%, 25%-35%, 35%-45%, 45%-55%, 55%-65%, 65%-75%, 75%-85%, 85%-95%, or 90%-100% of the uridine positions in a polynucleotide of the disclosure are 5-iodouridine and the remainder are N1-methylpseudouridine. In some embodiments, between 15% and 45%, between 45% and 55%, between 55% and 65%, between 65% and 75%, between 75% and 85%, between 85% and 95%, or between 90% and 100% of the uridine positions in a polynucleotide according to the present disclosure are substituted with a modified uridine, optionally, the modified uridine is an N1-methyl-pseudouridine.In some embodiments, 15%-45%, 45%-55%, 55%-65%, 65%-75%, 75%-85%, 85%-95%, or 90%-100% of the uridine positions in a polynucleotide of the disclosure are substituted with N1-methyl-pseudouridine. In some embodiments, 85%, 90%, 95%, or 100% of the uridine positions in a polynucleotide of the disclosure are substituted with N1-methyl-pseudouridine. In some embodiments, 100% of the uridines are substituted with N1-methyl-pseudouridine. In some embodiments, 15%-45%, 45%-55%, 55%-65%, 65%-75%, 75%-85%, 85%-95%, or 90%-100% of the uridine positions in a polynucleotide according to the disclosure are substituted with a modified uridine, optionally the modified uridine is a pseudouridine. In some embodiments, 15%-45%, 45%-55%, 55%-65%, 65%-75%, 75%-85%, 85%-95%, or 90%-100% of the uridine positions in a polynucleotide according to the disclosure are substituted with a pseudouridine. In some embodiments, 85%, 90%, 95%, or 100% of the uridine positions in a polynucleotide of the disclosure are substituted with a pseudouridine. In some embodiments, 100% of the uridines are substituted with a pseudouridine.
[0238] II. Exemplary Polynucleotides and Compositions Comprising Deaminases and RNA-Guided Nickases The RNA-guided DNA binding agents disclosed herein may further comprise a base-editing domain that introduces specific modifications into a target nucleic acid, such as a deaminase domain.
[0239] In some embodiments, a nucleic acid is provided, the nucleic acid comprising an open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., A3A) and a C-terminal NmeCas9 nickase, and a first nuclear localization signal (NLS), wherein the polypeptide does not comprise a uracil glycosylase inhibitor (UGI).
[0240] In some embodiments, the second NLS is N-terminal to the Nme Cas9 nickase. In some embodiments, the deaminase is N-terminal to the NLS (i.e., the first NLS or the second NLS). In some embodiments, the deaminase is N-terminal to all NLSs in the polypeptide. In some embodiments, the polypeptide does not include a uracil glycosylase inhibitor (UGI).
[0241] In some embodiments, the polynucleotide is DNA or RNA. In some embodiments, the polynucleotide is mRNA. In some embodiments, a polypeptide encoded by the mRNA is provided.
[0242] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, a cytidine deaminase (e.g., APOBEC3A), an optional linker, and a D16A NmeCas9 nickase. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, a cytidine deaminase (e.g., APOBEC3A), an optional linker, and a D16A Nme2Cas9 nickase. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, a first and a second NLS, a cytidine deaminase (e.g., APOBEC3A), an optional linker, and a D16A NmeCas9 nickase. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, a first and a second NLS, a cytidine deaminase (e.g., APOBEC3A), an optional linker, and a D16A Nme2Cas9 nickase. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, a first NLS, a cytidine deaminase (e.g., APOBEC3A), a second NLS, an optional linker, a D16A NmeCas9 nickase. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, a first NLS, a cytidine deaminase (e.g., APOBEC3A), a second NLS, an optional linker, a D16A Nme2Cas9 nickase.
[0243] In some embodiments, a polypeptide comprising A3A and an RNA-guided nickase does not include a uracil glycosylase inhibitor (UGI).
[0244] In some embodiments, compositions are provided that include a first polypeptide or an mRNA encoding a first polypeptide comprising a cytidine deaminase, optionally an APOBEC3A deaminase (A3A), a C-terminal NmeCas9 nickase, a first nuclear localization signal (NLS), and optionally a second NLS, where the first NLS, and, if present, the second NLS, are located N-terminal to a sequence encoding the NmeCas9 nickase, and where the first polypeptide does not comprise a uracil glycosylase inhibitor (UGI), and an mRNA encoding a second polypeptide or a second polypeptide comprising a uracil glycosylase inhibitor (UGI), where the second polypeptide is different from the first polypeptide.
[0245] In some embodiments, a method for modifying a target gene is provided, comprising administering a composition described herein. In some embodiments, the method comprises delivering to a cell a first nucleic acid comprising a first open reading frame encoding a first polypeptide comprising a cytidine deaminase, optionally APOBEC3A deaminase (A3A), a C-terminal NmeCas9 nickase, a first nuclear localization signal (NLS), and optionally a second NLS, wherein the first NLS and, if present, the second NLS are located N-terminal to the sequence encoding the NmeCas9 nickase, and the first polypeptide does not comprise a uracil glycosylase inhibitor (UGI), and a second nucleic acid comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), the second nucleic acid being different from the first nucleic acid.
[0246] In some embodiments, the method includes delivering to a cell a polypeptide comprising a deaminase, optionally an APOBEC3A deaminase (A3A), a C-terminal NmeCas9 nickase, a first nuclear localization signal (NLS), and a second NLS, wherein the first NLS and the second NLS are located N-terminal to a sequence encoding the NmeCas9 nickase, and the first polypeptide does not comprise a uracil glycosylase inhibitor (UGI), or a nucleic acid encoding the polypeptide; and delivering to the cell a uracil glycosylase inhibitor (UGI), or a nucleic acid encoding a UGI.
[0247] In some embodiments, the molar ratio of mRNA encoding UGI to mRNA encoding APOBEC3A deaminase (A3A) and RNA-guided nickase is about 1:35 to about 30:1. In some embodiments, the molar ratio is about 1:25 to about 25:1. In some embodiments, the molar ratio is about 1:20 to about 25:1. In some embodiments, the molar ratio is about 1:10 to about 22:1. In some embodiments, the molar ratio is about 1:5 to about 25:1. In some embodiments, the molar ratio is about 1:1 to about 30:1. In some embodiments, the molar ratio is about 2:1 to about 10:1. In some embodiments, the molar ratio is about 5:1 to about 20:1. In some embodiments, the molar ratio is about 1:1 to about 25:1. In some embodiments, the molar ratio is about 1:35, 1:34, 1:33, 1:32, 1:31, 1:30, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:1 The molar ratio may be 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, or 30:1. In some embodiments, the molar ratio is about 1:1 or greater. In some embodiments, the molar ratio is about 1:1. In some embodiments, the molar ratio is about 2:1. In some embodiments, the molar ratio is about 3:1. In some embodiments, the molar ratio is about 4:1. In some embodiments, the molar ratio is about 5:1. In some embodiments, the molar ratio is about 6:1. In some embodiments, the molar ratio is about 7:1. In some embodiments, the molar ratio is about 8:1. In some embodiments, the molar ratio is about 9:1.In some embodiments, the molar ratio is about 10:1. In some embodiments, the molar ratio is about 11:1. In some embodiments, the molar ratio is about 12:1. In some embodiments, the molar ratio is about 13:1. In some embodiments, the molar ratio is about 14:1. In some embodiments, the molar ratio is about 15:1. In some embodiments, the molar ratio is about 16:1. In some embodiments, the molar ratio is about 17:1. In some embodiments, the molar ratio is about 18:1. In some embodiments, the molar ratio is about 19:1. In some embodiments, the molar ratio is about 20:1. In some embodiments, the molar ratio is about 21:1. In some embodiments, the molar ratio is about 22:1. In some embodiments, the molar ratio is about 23:1. In some embodiments, the molar ratio is about 24:1. In some embodiments, the molar ratio is about 25:1.
[0248] Similarly, in some embodiments, the molar ratios described above for the mRNA encoding the UGI protein to the mRNA encoding APOBEC3A deaminase (A3A) and RNA-guided nickase are similar when delivering proteins.
[0249] In some embodiments, the compositions described herein further comprise at least one gRNA. In some embodiments, a composition is provided comprising an mRNA described herein and at least one gRNA. In some embodiments, the gRNA is a single guide RNA (sgRNA). In some embodiments, the gRNA is a dual guide RNA (dgRNA).
[0250] In some embodiments, the composition is capable of effecting genome editing upon administration to a subject.
[0251] A. Cytidine deaminase, APOBEC3A deaminase Cytidine deaminases include enzymes within the cytidine deaminase superfamily, in particular the APOBEC family of enzymes (APOBEC1, APOBEC2, APOBEC4, and APOBEC3 subgroups of enzymes), activation-induced cytidine deaminases (AID or AICDA), and CMP deaminases (see, e.g., Conticello et al., Mol. Biol. Evol. 22:367-77, 2005; Conticello, Genome Biol. 9:229, 2008; Muramatsu et al., J. Biol. Chem. 274:18470-6, 1999; and Carrington et al., Cells 9:1690 (2020)).
[0252] In some embodiments, the cytidine deaminase disclosed herein is an enzyme of the APOBEC family. In some embodiments, the cytidine deaminase disclosed herein is an enzyme of the APOBEC1, APOBEC2, APOBEC4, and APOBEC3 subgroups. In some embodiments, the cytidine deaminase disclosed herein is an enzyme of the APOBEC3 subgroup. In some embodiments, the cytidine deaminase disclosed herein is an APOBEC3A deaminase (A3A). In some embodiments, the deaminase comprises an APOBEC3A deaminase.
[0253] In some embodiments, the APOBEC3A deaminase (A3A) disclosed herein is human A3A. In some embodiments, the APOBEC3A deaminase (A3A) disclosed herein is human A3A. In some embodiments, the A3A is wild-type A3A.
[0254] In some embodiments, A3A is an A3A variant. The A3A variant shares homology with wild-type A3A, or a fragment thereof. In some embodiments, the A3A variant has at least about 80% identity, at least about 85% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to wild-type A3A. In some embodiments, an A3A variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type A3A. In some embodiments, an A3A variant comprises a fragment of A3A such that the fragment has at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to a corresponding fragment of wild-type A3A.
[0255] In some embodiments, an A3A variant is a protein having a sequence that differs from the wild-type A3A protein by one or more mutations, e.g., substitutions, deletions, insertions, one or more single point substitutions. In some embodiments, a truncated A3A sequence can be used, e.g., by deleting N-terminal, C-terminal, or internal amino acids. In some embodiments, a truncated A3A sequence is used, with one to four amino acids deleted at the C-terminus of the sequence. In some embodiments, APOBEC3A (e.g., human APOBEC3A) has wild-type amino acid position 57 (numbered in the wild-type sequence). In some embodiments, APOBEC3A (e.g., human APOBEC3A) has an asparagine at amino acid position 57 (numbered in the wild-type sequence).
[0256] In some embodiments, the wild-type A3A is human A3A (UniProt Accession ID: p319411, SEQ ID NO: 151).
[0257] In some embodiments, A3A disclosed herein comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 151. In some embodiments, the level of identity is at least 85%, at least 87%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, A3A comprises an amino acid sequence having at least 87% identity to SEQ ID NO: 151. In some embodiments, A3A comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 151. In some embodiments, A3A comprises an amino acid sequence having at least 95% identity to SEQ ID NO: 151. In some embodiments, A3A comprises an amino acid sequence having at least 98% identity to SEQ ID NO: 151. In some embodiments, A3A comprises an amino acid sequence having at least 99% identity to SEQ ID NO: 151. In some embodiments, A3A comprises the amino acid sequence of SEQ ID NO: 151.
[0258] In some embodiments, the cytidine deaminase disclosed herein comprises an amino acid sequence having at least 80% identity to any one of SEQ ID NOs: 151-216. In some embodiments, the level of identity is at least 85%, at least 87%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, the cytidine deaminase comprises an amino acid sequence of any one of SEQ ID NOs: 151-216.
[0259] B. UGI Without being bound by any theory, providing a UGI with a polypeptide comprising a deaminase may be useful in the methods described herein by inhibiting cellular DNA repair mechanisms (e.g., UDG and downstream repair effectors) that recognize uracil in DNA as a form of DNA damage or otherwise excise or modify uracil or surrounding nucleotides. It is understood that the use of a UGI may increase the editing efficiency of enzymes that can deaminate C residues.
[0260] Suitable UGI protein and nucleotide sequences are provided herein, and additional suitable UGI sequences will be known to those of skill in the art, see, e.g., Wang et al., Uracil-DNA glycosylase inhibitor gene of bacteriophage PBS2 encodes a binding protein specific for uracil-DNA glycosylase. J. Biol. Chem. 264:1163-1171 (1989); Lundquist et al., Site-directed mutagenesis and characterization of uracil-DNA glycosylase inhibitor protein. Role of specific carboxylic amino acids in complex formation with Escherichia coli uracil-DNA glycosylase. J. Biol. Chem. 272:21408-21419 (1997); Ravishankar et al., X-ray analysis of a complex of Escherichia coli uracil DNA glycosylase (EcUDG) with a proteinaceous inhibitor. The structure elucidation of a prokaryotic Nucleic Acids Res. 26:4880-4887 (1998), and Putnam et al., Protein mimicry of DNA from crystal structures of the uracil-DNA glycosylase inhibitor protein and its complex with Escherichia coli uracil-DNA glycosylase. J. Mol. Biol. 287:331-346 (1999), the entire contents of each of which are incorporated herein by reference. It should be understood that any protein capable of inhibiting uracil-DNA glycosylase base excision repair enzyme is within the scope of the present disclosure.In addition, any protein that blocks or inhibits base excision repair is within the scope of the present disclosure. In some embodiments, the uracil glycosylase inhibitor is a protein that binds to uracil. In some embodiments, the uracil glycosylase inhibitor is a protein that binds to uracil in DNA. In some embodiments, the uracil glycosylase inhibitor is a single-stranded binding protein. In some embodiments, the uracil glycosylase inhibitor is a catalytically inactive uracil DNA glycosylase protein. In some embodiments, the uracil glycosylase inhibitor is a catalytically inactive uracil DNA glycosylase protein that does not excise uracil from DNA. In some embodiments, the uracil glycosylase inhibitor is a catalytically inactive UDG.
[0261] In some embodiments, a uracil glycosylase inhibitor (UGI) disclosed herein comprises an amino acid sequence that is at least 80% identical to SEQ ID NO:3. In some embodiments, any of the aforementioned levels of identity are at least 90%, at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, the UGI comprises an amino acid sequence that has at least 90% identity to SEQ ID NO:3. In some embodiments, the UGI comprises an amino acid sequence that has at least 95% identity to SEQ ID NO:3. In some embodiments, the UGI comprises an amino acid sequence that has at least 98% identity to SEQ ID NO:3. In some embodiments, the UGI comprises an amino acid sequence that has at least 99% identity to SEQ ID NO:3. In some embodiments, the UGI comprises the amino acid sequence of SEQ ID NO:3.
[0262] C. Linker In some embodiments, the polypeptide comprising the deaminase and the RNA-guided nickase described herein further comprises a linker connecting the deaminase and the RNA-guided nickase. In some embodiments, the linker is a peptide linker. In some embodiments, the nucleic acid encoding the polypeptide comprising the deaminase and the RNA-guided nickase further comprises a sequence encoding the peptide linker. In some embodiments, an mRNA encoding a deaminase-linker-RNA-guided nickase fusion protein is provided.
[0263] In some embodiments, the peptide linker is any amino acid region having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, or at least 50 or more amino acids.
[0264] In some embodiments, the peptide linker is a 16 residue "XTEN" linker, or a variant thereof (see, e.g., the Examples, and Schellenberger et al. A recombinant polypeptide extends the in vivo half-life of peptides and proteins in a tunable manner. Nat. Biotechnol. 27, 1186-1190 (2009)). In some embodiments, the XTEN linker comprises the sequence SGSETPGTSESATPES (SEQ ID NO: 58), SGSETPGTSESA (SEQ ID NO: 59), or SGSETPGTSESATPEGGSGGS (SEQ ID NO: 60).
[0265] In some embodiments, the peptide linker is (GGGGS) n (SEQ ID NO: 62), (G) n , (EAAAK) n (SEQ ID NO: 63), (GGS)n (SEQ ID NO: 61), or the SGSETPGTSESATPES (SEQ ID NO: 58) motif (see, e.g., Guilinger JP, Thompson DB, Liu D R. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014; 32(6):577-82, the entire contents of which are incorporated herein by reference), or (XP) n motif, or any combination thereof, wherein n is independently an integer from 1 to 30. See WO2015089406, e.g., paragraph
[0012] , the entire contents of which are incorporated herein by reference.
[0266] In some embodiments, the peptide linker comprises one or more sequences selected from SEQ ID NOs: 58 to 122. In some embodiments, the peptide linker comprises one or more sequences selected from SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 118, SEQ ID NO: 119, SEQ ID NO: 120, SEQ ID NO: 121, and SEQ ID NO: 122.
[0267] D. Compositions Comprising APOBEC3A Deaminase and RNA-Guided Nickase In some embodiments, an mRNA is provided that encodes a polypeptide comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase. In some embodiments, the polypeptide comprises a nucleotide sequence encoding a human deaminase (e.g., A3A) and a C-terminal RNA-guided nickase, and a first NLS and optionally a second NLS. In certain embodiments, the deaminase is N-terminal to the NLS. In certain embodiments, the deaminase is N-terminal to all NLSs.
[0268] In some embodiments, the polypeptide comprises a wild-type deaminase (e.g., A3A) and a C-terminal RNA-guided nickase. In some embodiments, the polypeptide comprises an A3A variant and an RNA-guided nickase. In some embodiments, the polypeptide comprises a deaminase (e.g., A3A) and a Cas9 nickase. In some embodiments, the polypeptide comprises a deaminase (e.g., A3A) and a D16A NmeCas9 nickase. In some embodiments, the polypeptide comprises a human deaminase (e.g., A3A) and a D16A NmeCas9 nickase. In some embodiments, the polypeptide comprises an A3A variant and a D16A NmeCas9 nickase. In some embodiments, the polypeptide lacks a UGI. In some embodiments, the deaminase (e.g., A3A) and the RNA-guided nickase are linked via a linker. In some embodiments, the polypeptide further comprises one or more additional heterologous functional domains. In some embodiments, the polypeptide further comprises a nuclear localization sequence (NLS) (described herein).
[0269] In some embodiments, the polypeptide comprises a human deaminase (e.g., A3A) and a C-terminal D16A NmeCas9 nickase, where the human deaminase (e.g., A3A) and the D16A NmeCas9 are fused via a linker. In some embodiments, the polypeptide comprises a human A3A and a C-terminal D16A NmeCas9 nickase, and an NLS at the N-terminus of the fusion polypeptide. In some embodiments, the polypeptide comprises a human A3A and a C-terminal D16A NmeCas9 nickase, where the human A3A and the D16A NmeCas9 are fused via a linker, and optionally comprises an NLS fused to the N-terminus of the human A3A, optionally via a linker.
[0270] The polypeptide may be organized in any number of ways to form a single chain. The first NLS and, if present, the second NLS are located N-terminal to the sequence encoding the Cas9 nickase. The additional NLS may be N-terminal to the Cas9 nickase. The A3A may be N-terminal or C-terminal relative to the NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, the first NLS, an optional second NLS, a deaminase, an optional linker, an RNA-guided nickase, and an optional NLS. In some embodiments, the linker is independently present between the first NLS and the second NLS, and between the NLS and the deaminase. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, a deaminase, a first NLS, an optional second NLS, and a C-terminal RNA-guided nickase. In some embodiments, linkers are present, independently, between the deaminase and the first NLS, between the first NLS and the second NLS, and between the NLS and the C-terminal nickase.
[0271] In any of the foregoing embodiments, the polypeptide may comprise an amino acid sequence having at least 80% identity to SEQ ID NO: 14. In some embodiments, any of the foregoing identity levels are at least 85%, 90%, 95%, 98%, or 99%, or 100% identical. In some embodiments, the polypeptides disclosed herein may comprise an amino acid sequence having at least 90% identity to SEQ ID NO: 14. In some embodiments, the polypeptides disclosed herein may comprise an amino acid sequence having at least 95% identity to SEQ ID NO: 3 or 6. In some embodiments, the polypeptides disclosed herein may comprise an amino acid sequence having at least 98% identity to SEQ ID NO: 14. In some embodiments, the polypeptides disclosed herein may comprise an amino acid sequence having at least 99% identity to SEQ ID NO: 14. In some embodiments, the polypeptides disclosed herein may comprise the amino acid sequence of SEQ ID NO: 14.
[0272] In any of the foregoing embodiments, a nucleic acid sequence comprising an open reading frame encoding a polypeptide disclosed herein may comprise a nucleic acid sequence having at least 80% identity to SEQ ID NO: 42. In some embodiments, any of the foregoing levels of identity are at least 85%, 90%, 95%, 98%, or 99%, or 100% identical.
[0273] In any of the foregoing embodiments, an mRNA sequence encoding a polypeptide disclosed herein may comprise a nucleic acid sequence having at least 80% identity to SEQ ID NO: 28. In some embodiments, any of the foregoing levels of identity are at least 85%, 90%, 95%, 98%, or 99%, or 100% identical.
[0274] In any of the foregoing embodiments, A3A may comprise an amino acid sequence having at least 80% identity to SEQ ID NO: 151. In some embodiments, the level of identity is at least 85%, 87%, 90%, 95%, 98%, or 99%, or 100% identical. In some embodiments, A3A comprises the amino acid sequence of SEQ ID NO: 151.
[0275] In any of the foregoing embodiments, the NmeCas9 nickase may comprise an amino acid sequence having at least 80%, 90%, 95%, 98%, or 99% identity to any one of SEQ ID NOs: 220, 248, or 276. In some embodiments, the level of identity is at least 85%, 87%, 90%, 95%, 98%, or 99%, or 100% identical. In some embodiments, the RNA-guided nickase comprises the amino acid sequence of SEQ ID NO: 220, 248, or 276. In some embodiments, the RNA-guided nickase comprises the amino acid sequence of SEQ ID NO: 220, 248, or 276. In some embodiments, the RNA-guided nickase comprises the amino acid sequence of SEQ ID NO: 220, 248, or 276.
[0276] III. Guide RNA In some embodiments, at least one guide RNA is provided in combination with the polynucleotide disclosed herein, such as the polynucleotide encoding the RNA-guided DNA binding agent.In some embodiments, guide RNA is provided as a molecule separate from the polynucleotide.In some embodiments, guide RNA is provided as a part of the polynucleotide disclosed herein, such as a part of UTR.
[0277] In some embodiments, a composition comprising a polynucleotide disclosed herein further comprises at least one guide RNA (or "gRNA").
[0278] In some embodiments, the gRNA is a single guide RNA (or "sgRNA").
[0279] In some embodiments, the gRNA is a dual guide RNA.
[0280] In some embodiments, the guide RNA comprises a modified sgRNA. The sgRNA may be modified to improve its in vivo stability.
[0281] In some embodiments, the gRNA described herein is a Neisseria meningitidis Cas9 (NmeCas9) gRNA that includes a conserved portion that includes the repeat / anti-repeat region, the hairpin 1 region, and the hairpin 2 region, where one or more of the repeat / anti-repeat region, the hairpin 1 region, and the hairpin 2 region are truncated. An exemplary wild-type NmeCas9 guide RNA is (N) 20~25 As used herein, (N) includes the sequence of: GUUGUAGCUCCCUUUCUCAUUUCGGAAACGAAAUGAGAACCGUUGCUACAAUAAGGCCGUCUGAAAAGAUGUGCCGCAACGCUCUGCCCCUUAAAGCUUCUGCUUUAAGGGGCAUCGUUUA (SEQ ID NO: 500). 20~25represents 20 to 25, i.e., 20, 21, 22, 23, 24, or 25 consecutive Ns. A, C, G, and U represent nucleotides having adenine, cytosine, guanine, and uracil bases, respectively. In some embodiments, (N) 20~25 has a length of 24 nucleotides. N is any natural or unnatural nucleotide, and the entirety of N comprises the guide sequence.
[0282] In some embodiments, the single guide RNA comprises a guide region and a conserved region, (a) a truncated repeat / anti-repeat region, the truncated repeat / anti-repeat region lacking 2 to 24 nucleotides; (i) one or more of nucleotides 37-48 and 53-64 are deleted, and optionally one or more of nucleotides 37-64 are substituted with respect to SEQ ID NO:500; (ii) a truncated repeat / anti-repeat region in which nucleotide 36 is linked to nucleotide 65 by at least two nucleotides; or (b) a truncated hairpin 1 region, wherein truncated hairpin 1 lacks 2-10, optionally 2-8, nucleotides; (i) one or more of nucleotides 82-86 and 91-95 are deleted, and optionally one or more of positions 82-96 are substituted relative to SEQ ID NO:500; (ii) a truncated hairpin 1 region, in which nucleotide 81 is linked to nucleotide 96 by at least four nucleotides; or (c) a truncated hairpin 2 region, wherein the truncated hairpin 2 is missing 2-18, optionally 2-16, nucleotides; (i) one or more of nucleotides 113-121 and 126-134 are deleted, and optionally one or more of nucleotides 113-134 are substituted with respect to SEQ ID NO:500; (ii) a truncated hairpin 2 region, in which nucleotide 112 is linked to nucleotide 135 by at least four nucleotides; One or both nucleotides 144-145 are optionally deleted relative to SEQ ID NO:500, and at least 10 nucleotides are modified nucleotides.
[0283] In some embodiments, the truncated repeat / anti-repeat region of the gRNA lacks 18 nucleotides. In some embodiments, the truncated repeat / anti-repeat region of the gRNA lacks 22 anti-repeat nucleotides.
[0284] In some embodiments, in the truncated repeat / anti-repeat region of the gRNA, nucleotide 36 is linked to nucleotide 65 by 6 nucleotides. In some embodiments, in the truncated repeat / anti-repeat region of the gRNA, nucleotide 36 is linked to nucleotide 65 by 7 nucleotides. In some embodiments, in the truncated repeat / anti-repeat region of the gRNA, nucleotide 36 is linked to nucleotide 65 by 8 nucleotides. In some embodiments, in the truncated repeat / anti-repeat region of the gRNA, nucleotide 36 is linked to nucleotide 65 by 9 nucleotides. In some embodiments, in the truncated repeat / anti-repeat region of the gRNA, nucleotide 36 is linked to nucleotide 65 by 10 nucleotides.
[0285] In some embodiments, in the truncated repeat / anti-repeat region of the gRNA, nucleotides 38-48 and 53-63 are deleted relative to SEQ ID NO: 500. In some embodiments, in the truncated repeat / anti-repeat region of the gRNA, nucleotides 38, 41-48, 53-60, and 63 are deleted relative to SEQ ID NO: 500.
[0286] In some embodiments, in the truncated repeat / anti-repeat region of the gRNA, nucleotide 36 is linked to nucleotide 65 by 6 nucleotides. In some embodiments, in the truncated repeat / anti-repeat region of the gRNA, nucleotides 38-48 and 53-63 are deleted relative to SEQ ID NO:500, and nucleotide 36 is linked to nucleotide 65 by nucleotides 37, 49-52, and 64.
[0287] In some embodiments, in the truncated repeat / anti-repeat region of the gRNA, nucleotide 36 is linked to nucleotide 65 by 10 nucleotides. In some embodiments, in the truncated repeat / anti-repeat region of the gRNA, nucleotides 38, 41-48, 53-60, and 63 are deleted relative to SEQ ID NO:500, and nucleotide 36 is linked to nucleotide 65 by nucleotides 37, 39, 40, 49-52, 61, 62, and 64.
[0288] In some embodiments, nucleotides 38-48 and all of nucleotides 53-63 of the upper stem of the truncated repeat / repeat region are deleted relative to SEQ ID NO:500.
[0289] In some embodiments, nucleotides 39-48 and all of nucleotides 53-62 of the upper stem of the truncated repeat / anti-repeat region are deleted relative to SEQ ID NO:500, and nucleotides 38 and 63 are replaced.
[0290] In some embodiments, the truncated repeat / anti-repeat region has 14 modified nucleotides. In some embodiments, the truncated repeat / anti-repeat region has 15 modified nucleotides. In some embodiments, the truncated repeat / anti-repeat region has 16 modified nucleotides. In some embodiments, the truncated repeat / anti-repeat region has 17 modified nucleotides. In some embodiments, the truncated repeat / anti-repeat region has 18 modified nucleotides. In some embodiments, the truncated repeat / anti-repeat region has 19 modified nucleotides. In some embodiments, the truncated repeat / anti-repeat region has 20 modified nucleotides.
[0291] In some embodiments, the truncated hairpin 1 region is missing 2 nucleotides. In some embodiments, the truncated hairpin 1 region is missing 21 nucleotides. In some embodiments, the truncated hairpin 1 region is missing 2 nucleotides, with nucleotides 86 and 91 being deleted relative to SEQ ID NO:500. In some embodiments, the truncated hairpin 1 region is missing 2 nucleotides, with nucleotides 85 and 92 being deleted relative to SEQ ID NO:500. In some embodiments, in the truncated hairpin 1 region, nucleotide 81 is linked to nucleotide 96 by 12 nucleotides. In some embodiments, in the truncated hairpin 1 region, nucleotides 86 and 91 are deleted relative to SEQ ID NO:500, with nucleotide 81 being linked to nucleotide 96 by nucleotides 82-85, 87-90, and 92-95. In some embodiments, in the shortened hairpin 1 region, nucleotides 85 and 92 are deleted relative to SEQ ID NO:500, and nucleotide 81 is joined to nucleotide 96 by nucleotides 82-84, 86-91, and 93-95.
[0292] In some embodiments, the shortened hairpin 1 region has a double-stranded portion 7 base paired nucleotides in length. In some embodiments, the shortened hairpin 1 region has a double-stranded portion 8 base paired nucleotides in length.
[0293] The stem of the truncated hairpin 1 region has a length of 7 base paired nucleotides, in some embodiments, nucleotides 85-86 and 91-92 of the truncated hairpin 1 region are deleted.
[0294] In some embodiments, the shortened hairpin 1 region has 13 modified nucleotides.
[0295] In some embodiments, the shortened hairpin 2 lacks 18 nucleotides. In some embodiments, the shortened hairpin 2 has 24 nucleotides. In some embodiments, nucleotides 113-121 and 126-134 of the shortened hairpin 2 are deleted relative to SEQ ID NO:500. In some embodiments, the shortened hairpin 2 lacks 18 nucleotides, and nucleotides 113-121 and 126-134 are deleted relative to SEQ ID NO:500. In some embodiments, in the shortened hairpin 2 region, nucleotide 112 is linked to nucleotide 135 by 4 nucleotides. In some embodiments, in the shortened hairpin 2 region, nucleotides 113-121 and 126-134 are deleted relative to SEQ ID NO:500, and nucleotide 112 is linked to nucleotide 135 by nucleotides 122-125.
[0296] In some embodiments, the truncated repeat / anti-repeat region has a length of 28 nucleotides. In some embodiments, the truncated repeat / anti-repeat region has a length of 32 nucleotides.
[0297] In some embodiments, the upper stem of the truncated repeat / anti-repeat region comprises 1 base pair or less. In some embodiments, the upper stem of the truncated repeat / anti-repeat region comprises 3 base pairs or less.
[0298] In some embodiments, the shortened hairpin 2 region has 8 modified nucleotides.
[0299] In some embodiments, the guide RNA (gRNA) comprises a guide region and a conserved region, (a) a truncated repeat / anti-repeat region, wherein the truncated repeat / anti-repeat region lacks 18 to 22 nucleotides relative to SEQ ID NO: 500; (i) nucleotides 38 to 48 and 53 to 63 are deleted; (ii) a truncated repeat / anti-repeat region in which nucleotide 36 is linked to nucleotide 65 by 6 to 10 nucleotides; (b) a truncated hairpin 1 region, wherein the truncated hairpin 1 is missing two nucleotides, nucleotides 86 and 91 being deleted or nucleotides 85 and 92 being deleted relative to SEQ ID NO:500; (c) a truncated hairpin 2 region, wherein truncated hairpin 2 lacks 18 nucleotides, wherein nucleotides 113-121 and 126-134 are deleted relative to SEQ ID NO:500, and nucleotides 144-145 are deleted relative to SEQ ID NO:500, and wherein at least 10 nucleotides are modified nucleotides.
[0300] In some embodiments, the guide RNA (gRNA) comprises a guide region and a conserved region, (a) a truncated repeat / anti-repeat region, wherein the truncated repeat / anti-repeat region lacks 18 to 22 nucleotides relative to SEQ ID NO: 500; (i) deletion of nucleotides 38, 41 to 48, 53 to 60, and 63; (ii) a truncated repeat / anti-repeat region in which nucleotide 36 is linked to nucleotide 65 by 6 to 10 nucleotides; (b) a truncated hairpin 1 region, wherein the truncated hairpin 1 is missing two nucleotides, nucleotides 86 and 91 being deleted or nucleotides 85 and 92 being deleted relative to SEQ ID NO:500; (c) a truncated hairpin 2 region, wherein the truncated hairpin 2 lacks 18 nucleotides, and nucleotides 113-121 and 126-134 are deleted relative to SEQ ID NO:500; Nucleotides 144 to 145 are deleted relative to SEQ ID NO: 500, and a shortened hairpin 2 region, in which at least 10 nucleotides are modified nucleotides.
[0301] In some embodiments, a guide RNA (gRNA) is provided, the gRNA comprising a guide region and a conserved region, the conserved region comprising: (a) a truncated repeat / anti-repeat region, wherein the truncated repeat / anti-repeat region lacks 18 to 22 nucleotides relative to SEQ ID NO: 500; (i) nucleotides 37 to 48 and 53 to 64 are deleted; (ii) a truncated repeat / anti-repeat region in which nucleotide 36 is linked to nucleotide 65 by between 6 and 10 nucleotides; or (b) a truncated hairpin 1 region, wherein the truncated hairpin 1 is missing two nucleotides, nucleotides 86 and 91 being deleted or nucleotides 85 and 92 being deleted relative to SEQ ID NO:500; or (c) a truncated hairpin 2 region, wherein the truncated hairpin 2 lacks 18 nucleotides, and nucleotides 113-121 and 126-134 are deleted relative to SEQ ID NO:500; Nucleotides 144 to 145 are deleted relative to SEQ ID NO: 500, It comprises one or more of a truncated hairpin 2 region, in which at least 10 nucleotides are modified nucleotides.
[0302] In a further embodiment, the truncated repeat / anti-repeat region lacks 22 nucleotides relative to SEQ ID NO: 500. In a further embodiment, nucleotide 36 is linked to nucleotide 65 by a sequence comprising the nucleotide sequence UGAAAC. In a further embodiment, nucleotide 36 is linked to nucleotide 65 by 10 nucleotides. In a further embodiment, nucleotide 36 is linked to nucleotide 65 by a sequence comprising the nucleotide sequence UUCGAAAGAC.
[0303] In some embodiments, the guide RNA (gRNA) of the previous embodiment comprises a guide region and a conserved region, (a) a truncated repeat / anti-repeat region, the truncated repeat / anti-repeat region lacking 18 to 22 nucleotides; (i) nucleotides 37 to 48 and 53 to 64 are deleted relative to SEQ ID NO: 500, (ii) a truncated repeat / anti-repeat region in which nucleotide 36 is linked to nucleotide 65 by 6 to 10 nucleotides; (b) a truncated hairpin 1 region, wherein the truncated hairpin 1 is missing two nucleotides relative to SEQ ID NO:500, nucleotides 86 and 91 are deleted, or nucleotides 85 and 92 are deleted; (c) a truncated hairpin 2 region, wherein the truncated hairpin 2 lacks 18 nucleotides, and nucleotides 113-121 and 126-134 are deleted relative to SEQ ID NO: 500; (d) nucleotides 144 to 145 are deleted relative to SEQ ID NO:500; At least 10 nucleotides are modified nucleotides.
[0304] In a further embodiment, the truncated repeat / anti-repeat region lacks 22 nucleotides relative to SEQ ID NO: 500. In a further embodiment, nucleotide 36 is linked to nucleotide 65 by a sequence comprising the nucleotide sequence UGAAAC. In a further embodiment, nucleotide 36 is linked to nucleotide 65 by 10 nucleotides. In a further embodiment, nucleotide 36 is linked to nucleotide 65 by a sequence comprising the nucleotide sequence UUCGAAAGAC.
[0305] Figures 33-35 show exemplary sgRNAs in possible secondary structures.
[0306] In some embodiments, the NmeCas9 short sgRNA comprises one of the following sequences in the 5' to 3' orientation: (N) 20~25 GUUGUAGCUCCCUGAAACCGUUGCUACAAUAAGGCCGUCGAAAGAUGU GCCGCAACGCUCUGCCUUCUGGCAUCGUU (SEQ ID NO: 501), (N) 20~25 GUUGUAGCUCCCUGAAACCGUUGCUACAAUAAGGCCGUCGAAAGAUGU GCCGCAACGCUCUGCCUUCUGGCAUCGUUUAUU (SEQ ID NO: 502), (N) 20~25 GUUGUAGCUCCCUGGAAACCCGUUGCUACAAUAAGGCCGUCGAAAGA UGUGCCGCAACGCUCUGCCUUCUGGCAUCGUUUAUU (SEQ ID NO: 503), or (N) 20~25 GUUGUAGCUCCCUUCGAAAGACCGUUGCUACAAUAAGGCCGUCGAAAGAUGUGCCGCAACGCUCUGCCUUCUGGCAUCGUU (sequence number 514). where N is a nucleotide encoding a guide sequence. In some embodiments, N is equal to 24. In some embodiments, N is equal to 25. N represents a nucleotide having any base, e.g., A, C, G, or U. (N) 20~25 represents a consecutive N from 20 to 25, i.e., 20, 21, 22, 23, 24, or 25.
[0307] In some embodiments, at least 10 nucleotides of the conserved portion of the NmeCas9 short sgRNA are modified nucleotides.
[0308] In some embodiments, the NmeCas9 short sgRNA comprises a conserved region that comprises one of the following sequences in the 5' to 3' orientation: mGUUGmUmAmGmCUCCCmUmGmAmAmAmCmCGUUmGmCUAmCAAU*AAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGCmAmAmCmGCUCUmGmCCmUmUmCmUGmGCmAmUC*mG*mU*mU (SEQ ID NO: 504), mGUUGmUmAmGmCUCCCmUmGmAmAmAmCmCGUUmGmCUAmCAAU*AAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGmCAAmCGCUCUmGmCCmUmUmCmUGGCAUCG*mU*mU (SEQ ID NO: 505), or Additional examples of NmeCas9 short gRNAs (e.g., SEQ ID NOs: 512-530) are shown in Table 39B.
[0309] In some embodiments, the NmeCas9 short gRNA comprises one of the following sequences in the 5' to 3' orientation: mN*mN*mN*mNmNNmNmNNmNNmNNNNmNNNNmNNNNmNNNmGUUGmUmAmGmCUCCCmUmGmAmAmAmCmCGUUmGmCUAmCAAU*AAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGmCAAmCGCUCUmGmCCmUmUmCmUGGCAUCG*mU*mU (SEQ ID NO: 512), mN*mN*mN*mNmNNmNmNNmNNmNNNNmNNNNmNNNNmNNNmGUUGmUmAmGmCUCCCmUmGmAmAmAmCmCGUUmGmCUAmCAAUAAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGmCAAmCGCUCUmGmCCmUmUmCmUGGCAUCG*mU*mU (SEQ ID NO: 525), mN*mN*mN*mNmNNmNmNNmNNmNNNNmNNNNmNNNNmNNNmGUUGmUmAmGmCUCCCmUmGmAmAmAmCmCGUUmGmCUAmCAAU*AAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGmCAAmCGmCmUmCmUmGmCCmUmUmCmUGGCAUCG*mU*mU (SEQ ID NO: 526), mN*mN*mN*mNmNNmNmNNmNNmNNNNmNNNNmNNNNmNNNmGUUGmUmAmGmCUCCCmUmGmAmAmAmCmCGUUmGmCUAmCAAUAAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGmCAAmCGmCmUmCmUmGmCCmUmUmCmUGGCAUCG*mU*mU (SEQ ID NO: 527), mN*mN*mN*mNmNNmNmNNmNNmNNNNmNNNNmNNNNmNNNmGUUGmUmAmGmCUCCCmUmUmCmGmAmAmAmGmAmCmCGUUmGmCUAmCAAU*AAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGmCAAmCGCUCUmGmCCmUmUmCmUGGCAUCG*mU*mU (SEQ ID NO: 516), mN*mN*mN*mNmNNmNmNNmNNmNNNNmNNNNmNNNNmNNNmGUUGmUmAmGmCUCCCmUmUmCmGmAmAmAmGmAmCmCGUUmGmCUAmCAAUAAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGmCAAmCGCUCUmGmCCmUmUmCmUGGCAUCG*mU*mU (SEQ ID NO: 520), mN*mN*mN*mNmNNmNmNNmNNmNNNNmNNNNmNNNNmNNNNmGUUGmUmAmGmCUCCCmUmUmCmGmAmAmAmGmAmCmCGUUmGmCUAmCAAU*AAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGmCAAmCGmCmUmCmUmGmCCmUmUmCmUGGCAUCG*mU*mU (SEQ ID NO: 521), or mN*mN*mN*mNmNNmNmNNmNNmNNNNmNNNNmNNNNmNNNmGUUGmUmAmGmCUCCCmUmUmCmGmAmAmAmGmAmCmCGUUmGmCUAmCAAUAAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGmCAAmCGmCmUmCmUmGmCCmUmUmCmUGGCAUCG*mU*mU (SEQ ID NO: 522), (N) 20~25 mGUUGmUmAmGmCUCCCmUmGmAmAmAmCmCGUUmGmCUAmCAAUAAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGCmAmAmCmGmCmUmCmUmGmCCmUmUmCmUGmGCmAmUC*mG*mU*mU (SEQ ID NO: 523), or (N) 20~25 mGUUGmUmAmGmCUCCCmUmUmCmGmAmAmAmGmAmCmCGUUmGmCUAmCAAUAAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGmCAAmCGmCmUmCmUmGmCCmUmUmCmUGGCAUCG*mU*mU (sequence number 524). In the formula, N represents any base, for example, A, C, G, or U. (mN*) 3represents three consecutive nucleotides, each having an arbitrary base, a 2'-OMe, and a 3'PS linkage to the adjacent nucleotide. The nucleotide modifications are: m is a 2'-OMe modification; * is indicated as a PS linkage. In the context of modified nucleotide sequences, in certain embodiments, N, A, C, G, and U are unmodified RNA nucleotides, i.e., a 2'-OH and a phosphodiesterase linkage to the 3' nucleotide.
[0310] The truncated NmeCas9 gRNA may include an internal linker as disclosed herein.
[0311] As used herein, "internal linker" describes a non-nucleotide segment that connects two nucleotides in a guide RNA. If the gRNA contains a spacer region, the internal linker is located outside the spacer region (e.g., within the scaffold or conserved region of the gRNA). In the case of a V-shaped guide, it is understood that the last hairpin is the only hairpin in the structure, i.e., the repeat-repeat prevention region. In some embodiments, the internal linker comprises a PEG linker disclosed herein. In some embodiments, the internal linker comprises a PEG linker disclosed herein.
[0312] In some embodiments, the single guide RNA comprises a guide region and a conserved region, (a) a truncated repeat / anti-repeat region, the truncated repeat / anti-repeat region lacking 2 to 24 nucleotides; (i) one or more of nucleotides 37 to 64 are deleted and optionally substituted relative to SEQ ID NO:500; (ii) nucleotide 36 is (i) a first internal linker that replaces four nucleotides, either alone or in combination with nucleotides, or (ii) a shortened repeat / anti-repeat region that is connected to nucleotide 65 by at least four nucleotides; or (b) a truncated hairpin 1 region, wherein truncated hairpin 1 lacks 2-10, optionally 2-8, nucleotides; (i) one or more of nucleotides 82 to 95 are deleted and optionally substituted relative to SEQ ID NO:500; (ii) nucleotide 81 is (i) a second internal linker that replaces four nucleotides, either alone or in combination with nucleotides, or (ii) a shortened hairpin 1 region that is linked to nucleotide 96 by at least four nucleotides; or (c) a truncated hairpin 2 region, wherein the truncated hairpin 2 is missing 2-18, optionally 2-16, nucleotides; (i) one or more of nucleotides 113 to 134 are deleted and optionally substituted relative to SEQ ID NO:500; (ii) nucleotide 112 includes one or more of: (i) a third internal linker, alone or in combination with nucleotides, substituting four nucleotides; or (ii) a truncated hairpin 2 region, connected to nucleotide 135 by at least four nucleotides; one or both nucleotides 144 to 145 are optionally deleted compared to SEQ ID NO:500, The gRNA comprises at least one of a first internal linker, a second internal linker, and a third internal linker.
[0313] Exemplary locations of the linker are as shown below. (N) 20-25 GUUGUAGCUCCCCUUC(L1)GACCGUUGCUACAAUAAGGCCGUC(L1)GAUGU GCCGCAACGCUCUGCC(L1)GGCAUCGUU (SEQ ID NO: 506). As used herein, (L1) refers to an internal linker having a bridge length of approximately 15 to 21 atoms.
[0314] In some embodiments, the truncated NmeCas9 guide RNAs containing internal linkers may be chemically modified. Exemplary modifications include the modification pattern of the following sequence: mN*mN*mN*mNNNmNmNNmNNNNmNNNNNNmNNNNmNNNNmGUUGmUmAmGmCUCCCmUmUmC(L1)mGmAmCmCGUUmGmCUAmCAAU*AAGmGmCCmGmUmC(L1)mGmAmUGUGCmCGmCAAmCGCUCUmGmCC(L1)GGCAUCG*mU*mU (SEQ ID NO: 507).
[0315] In some embodiments, the sgRNA comprises the modification pattern set forth in SEQ ID NOs: 141 and 143-150 (Nme PEG guide), where N is any natural or unnatural nucleotide and the entirety of N comprises the guide sequence.
[0316] IV. Delivery In some embodiments, the polynucleotides or compositions disclosed herein are formulated in or administered via lipid nanoparticles. See, e.g., WO2017173054, the entirety of which is incorporated herein by reference.
[0317] Lipids Disclosed herein are various embodiments that use lipid-nucleic acid assembly compositions that include the nucleic acid(s) or composition(s) described herein. In some embodiments, the lipid-nucleic acid assembly compositions include a nucleic acid (e.g., an mRNA) that includes an open reading frame (ORF) that encodes a polynucleotide that includes an open reading frame, the ORF including a nucleotide sequence that encodes a C-terminal Neisseria meningitidis (Nme) Cas9 polypeptide disclosed herein and a nucleotide sequence that encodes a first nuclear localization signal (NLS). In some embodiments, the NmeCas9 is Nme2Cas9, Nme1Cas9, or Nme3Cas9.
[0318] As used herein, "lipid-nucleic acid assembly composition" refers to lipid-based delivery compositions, including lipid nanoparticles (LNPs) and lipoplexes. LNPs refer to lipid nanoparticles less than 100 nM. LNPs are formed by precisely mixing the lipid components (e.g., in ethanol) with the aqueous nucleic acid components, and LNPs are uniform in size. Lipoplexes are particles formed by bulk mixing of lipid and nucleic acid components, and are about 100 nm to 1 micron in size. In certain embodiments, the lipid-nucleic acid assembly is an LNP. As used herein, a "lipid-nucleic acid assembly" comprises a plurality (i.e., two or more) lipid molecules physically associated with each other by intermolecular forces. The lipid-nucleic acid assembly may comprise bioavailable lipids with pKa values less than 7.5 or less than 7. The lipid-nucleic acid assembly is formed by mixing a nucleic acid-containing aqueous solution with an organic solvent-based lipid solution, e.g., 100% ethanol. Suitable solutions or solvents may include or contain water, PBS, Tris buffer, NaCl, citrate buffer, ethanol, chloroform, diethyl ether, cyclohexane, tetrahydrofuran, methanol, isopropanol. Optionally, a pharmaceutically acceptable buffer may be included in the pharmaceutical formulation that includes lipid nucleic acid assembly, for example, for ex vivo therapy. In some embodiments, the aqueous solution includes RNA, such as mRNA or gRNA. In some embodiments, the aqueous solution includes mRNA that codes for an RNA-guided DNA-binding agent, such as Cas9.
[0319] As used herein, lipid nanoparticles (LNPs) refer to particles that include a plurality (i.e., two or more) of lipid molecules physically associated with each other by intermolecular forces. LNPs can be, for example, microspheres (including unilamellar and multilamellar vesicles, e.g., lamellar phase lipid bilayers referred to as "liposomes," which in some embodiments are substantially spherical and, in certain embodiments, can include an aqueous core, e.g., containing a substantial portion of RNA molecules), the dispersed phase of an emulsion, a micelle, or the internal phase of a suspension. Emulsions, micelles, and suspensions can be suitable compositions for local and / or topical delivery. See, for example, WO2017173054A1, the contents of which are incorporated herein by reference in their entirety. Any LNP known to one of skill in the art to be capable of delivering nucleotides to a subject can be utilized with the guide RNA and nucleic acids encoding NmeCas9 and NLS described herein.
[0320] In some embodiments, the aqueous solution comprises a nucleic acid encoding a polypeptide comprising A3A and an RNA-guided nickase. Pharmaceutical formulations comprising lipid-nucleic acid assembly compositions can optionally include a pharma- ceutical acceptable buffer.
[0321] In some embodiments, lipid nucleic acid assembly compositions include "amine lipids" (sometimes referred to herein or elsewhere as "ionizable lipids" or "biodegradable lipids") along with optional "helper lipids," "neutral lipids," and stealth lipids such as PEG lipids. In some embodiments, the amine lipids or ionizable lipids are cationic, depending on the pH.
[0322] 1. Amine lipids In some embodiments, lipid nucleic acid assembly compositions comprise an "amine lipid," which is an ionizable lipid, such as lipid A or an equivalent thereof, including, for example, the acetal analog of lipid A.
[0323] In some embodiments, the amine lipid is lipid A, which is (9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate, alternatively 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl(9Z,12Z)-octadeca-9,12-dienoate. Lipid A is [ka] It can be shown as:
[0324] Lipid A can be synthesized according to WO2015 / 095340 (e.g., pages 84-86). In some embodiments, the amine lipid is an equivalent of lipid A.
[0325] In some embodiments, the amine lipid is an analog of lipid A. In some embodiments, the lipid A analog is an acetal analog of lipid A. In certain lipid nucleic acid assembly compositions, the acetal analog is a C4-C12 acetal analog. In some embodiments, the acetal analog is a C5-C12 acetal analog. In additional embodiments, the acetal analog is a C5-C10 acetal analog. In further embodiments, the acetal analog is selected from C4, C5, C6, C7, C9, C10, C11, and C12 acetal analogs.
[0326] Amine lipids and other "biodegradable lipids" suitable for use in the lipid nucleic acid assemblies described herein are biodegradable in vivo or ex vivo. Amine lipids have low toxicity (e.g., tolerated in animal models without side effects at doses of 10 mg / kg or greater). In some embodiments, lipid nucleic acid assemblies comprising amine lipids include those in which at least 75% of the amine lipid is cleared from plasma or engineered cells within 8, 10, 12, 24, or 48 hours, or within 3, 4, 5, 6, 7, or 10 days. In some embodiments, lipid nucleic acid assemblies comprising amine lipids include those in which at least 50% of the nucleic acid, e.g., mRNA or gRNA, is cleared from plasma within 8, 10, 12, 24, or 48 hours, or within 3, 4, 5, 6, 7, or 10 days. In some embodiments, lipid nucleic acid assemblies comprising amine lipids include those in which at least 50% of the lipid nucleic acid assembly is cleared from plasma within 8, 10, 12, 24, or 48 hours, or within 3, 4, 5, 6, 7, or 10 days, e.g., by measuring lipid (e.g., amine lipid), nucleic acid, e.g., RNA / mRNA, or other components. In some embodiments, lipid encapsulated versus free lipid, RNA, or nucleic acid components of the lipid nucleic acid assembly are measured.
[0327] Biodegradable lipids include, for example, the biodegradable lipids of WO / 2020 / 219876, WO / 2020 / 118041, WO / 2020 / 072605, WO / 2019 / 067992, WO / 2017 / 173054, WO2015 / 095340, and WO2014 / 136086, and LNPs include the LNP compositions described therein, which lipids and compositions are incorporated herein by reference.
[0328] Lipid clearance can be measured as described in the literature. See Maier, MA, et al. Biodegradable Lipids Enabling Rapidly Eliminated Lipid Nanoparticles for Systemic Delivery of RNAi Therapeutics. Mol. Ther. 2013, 21(8), 1570-78 ("Maier"). For example, in Maier, an LNP-siRNA system containing siRNA targeting luciferase was administered at 0.3 mg / kg by intravenous bolus injection via the lateral tail vein to 6-8 week old male C57Bl / 6 mice. Blood, liver, and spleen samples were collected at 0.083, 0.25, 0.5, 1, 2, 4, 8, 24, 48, 96, and 168 hours after administration. Mice were perfused with saline prior to tissue collection and blood samples were processed to obtain plasma. All samples were processed and analyzed by LC-MS. Furthermore, Maier describes procedures for evaluating toxicity after administration of LNP-siRNA formulations. For example, siRNA targeting luciferase was administered to male Sprague-Dawley rats at 0, 1, 3, 5, and 10 mg / kg (5 animals / group) by a single intravenous bolus injection at a dose volume of 5 mL / kg. After 24 hours, approximately 1 mL of blood was obtained from the jugular vein of conscious animals and serum was isolated. At 72 hours after administration, all animals were euthanized for necropsy. Evaluations of clinical signs, body weight, serum chemistry, organ weight, and histopathology were performed. Although Maier describes methods for evaluating siRNA-LNP formulations, these methods may be applied to evaluate clearance, pharmacokinetics, and toxicity of administration of lipid nucleic acid assembly compositions of the present disclosure.
[0329] Ionizable lipid and bioavailable lipid for LNP delivery of nucleic acid known in the art are suitable.Lipid may be ionizable depending on the pH of the medium in which it is contained.For example, in a slightly acidic medium, lipid such as amine lipid may be protonated and therefore positively charged.On the contrary, in a slightly basic medium such as blood, which has a pH of about 7.35, lipid such as amine lipid may not be protonated and therefore uncharged.
[0330] The ability of a lipid to carry a charge is related to its inherent pKa. In some embodiments, the amine lipids of the present disclosure may each independently have a pKa in the range of about 5.1 to about 7.4. In some embodiments, the bioavailable lipids of the present disclosure may each independently have a pKa in the range of about 5.1 to about 7.4, e.g., about 5.5 to about 6.6, about 5.6 to about 6.4, about 5.8 to about 6.2, or about 5.8 to about 6.5. For example, the amine lipids of the present disclosure may each independently have a pKa in the range of about 5.8 to about 6.5. Lipids having a pKa in the range of about 5.1 to about 7.4 are effective in delivering cargo in vivo, e.g., to the liver. Additionally, lipids having a pKa in the range of about 5.3 to about 6.4 have been found to be effective in delivering cargo in vivo, e.g., to tumors. See, e.g., WO2014 / 136086.
[0331] 2. Additional lipids "Neutral lipids" suitable for use in the lipid nucleic acid assembly compositions of the present disclosure include, for example, a variety of neutral, uncharged, or zwitterionic lipids. Examples of neutral phospholipids suitable for use in the present disclosure include, but are not limited to, 5-heptadecylbenzene-1,3-diol (resorcinol), dipalmitoylphosphatidylcholine (DPPC), distearoylphosphatidylcholine, such as 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC), posphocholine (DOPC), dimyristoylphosphatidylcholine (DMPC), phosphatidylcholine (PLPC), 1,2-distearoyl-sn-glycero-3-phosphocholine (DAPC), phosphatidylethanolamine (PE), egg phosphatidylcholine (EPC), dilauroylphosphatidylcholine (DLPC), dimyristoylphosphatidylcholine (DMPC), 1-myristoyl-2-palmitoylphosphatidylcholine (MPPC), 1-palmitoyl-2-myristoylphosphatidylcholine (PMPC), 1-palmitoyl-2-stearoylphosphatidylcholine (PSPC), 1,2-arachidoyl-sn-glycero-3-phosphocholine (DBPC), 1-stearoyl-2-palmitoylphosphatidylcholine (SPPC), 1,2-dieicosenoyl-sn-glycero-3-phosphocholine (DEPC), palmitoyloleoylphosphatidylcholine (POPC), lysophosphatidylcholine, dioleoylphosphatidylethanolamine (DOPE), dilinoleoylphosphatidylcholine distearoylphosphatidylethanolamine (DSPE), dimyristoylphosphatidylethanolamine (DMPE), dipalmitoylphosphatidylethanolamine (DPPE), palmitoyloleoylphosphatidylethanolamine (POPE), lysophosphatidylethanolamine and combinations thereof. In one embodiment, the neutral phospholipid may be selected from the group consisting of distearoylphosphatidylcholine (DSPC) and dimyristoylphosphatidylethanolamine (DMPE). In another embodiment, the neutral phospholipid may be distearoylphosphatidylcholine (DSPC).
[0332] "Helper lipids" include steroids, sterols, and alkylresorcinols. Helper lipids suitable for use in the present disclosure include, but are not limited to, cholesterol, 5-heptadecylresorcinol, and cholesterol hemisuccinate. In one embodiment, the helper lipid can be cholesterol. In one embodiment, the helper lipid can be cholesterol hemisuccinate.
[0333] A "stealth lipid" is a lipid that alters the length of time that a nanoparticle can reside in vivo (e.g., in a blood vessel). Stealth lipids can aid in the formulation process, for example, by reducing particle aggregation and controlling particle size. Stealth lipids as used herein can modulate the pharmacokinetic properties of lipid nucleic acid assemblies or aid in the stability of nanoparticles ex vivo. Stealth lipids suitable for use in the lipid nucleic acid assembly compositions of the present disclosure include, but are not limited to, stealth lipids that have a hydrophilic head group linked to the lipid moiety. Information regarding stealth lipids suitable for use in the lipid nucleic acid assembly compositions of the present disclosure and the biochemistry of such lipids can be found in Romberg et al., Pharmaceutical Research, Vol. 25, No. 1, 2008, pg. 55-71 and Hoekstra et al., Biochimica et Biophysica Acta 1660 (2004) 41-52. Additional suitable PEG lipids are disclosed in WO2006 / 007712.
[0334] In one embodiment, the hydrophilic head group of the stealth lipid comprises a polymer moiety selected from PEG-based polymers. The stealth lipid may comprise a lipid moiety. In some embodiments, the stealth lipid is a PEG lipid.
[0335] In one embodiment, the stealth lipid comprises a polymer moiety selected from PEG (sometimes called poly(ethylene oxide)), poly(oxazoline), poly(vinyl alcohol), poly(glycerol), poly(N-vinylpyrrolidone), polyamino acids, and poly[N-(2-hydroxypropyl)methacrylamide]-based polymers.
[0336] In one embodiment, the PEG lipid comprises a PEG-based polymer moiety, sometimes referred to as poly(ethylene oxide).
[0337] The PEG lipid further comprises a lipid moiety. In some embodiments, the lipid moiety may be derived from a diacylglycerol or diacylglycamide, such as one that comprises, independently, a dialkylglycerol or dialkylglycamide group having an alkyl chain length of about C4 to about C40 saturated or unsaturated carbon atoms, and the chain may comprise one or more functional groups, such as, for example, an amide or ester. In some embodiments, the alkyl chain length comprises about C10 to C20. The dialkylglycerol or dialkylglycamide group may further comprise one or more substituted alkyl groups. The chain length may be symmetrical or asymmetrical.
[0338] Unless otherwise specified, the term "PEG" as used herein means any polyethylene glycol or other polyalkylene ether polymer. In one embodiment, PEG is an optionally substituted linear or branched polymer of ethylene glycol or ethylene oxide. In one embodiment, PEG is unsubstituted. In one embodiment, PEG is substituted, for example, by one or more alkyl, alkoxy, acyl, hydroxy, or aryl groups. In one embodiment, the term includes PEG copolymers, such as PEG-polyurethane or PEG-polypropylene (see, e.g., J. Milton Harris, Poly(ethylene glycol) chemistry: biotechnical and biomedical applications (1992)). In another embodiment, the term does not include PEG copolymers. In one embodiment, the PEG has a molecular weight of about 130 to about 50,000, in a subembodiment, about 150 to about 30,000, in a subembodiment, about 150 to about 20,000, in a subembodiment, about 150 to about 15,000, in a subembodiment, about 150 to about 10,000, in a subembodiment, about 150 to about 6,000, in a subembodiment, about 150 to about 5,000, in a subembodiment, about 150 to about 4,000, in a subembodiment, about 150 to about 3,000, in a subembodiment, about 300 to about 3,000, in a subembodiment, about 1,000 to about 3,000, and in a subembodiment, about 1,500 to about 2,500.
[0339] In some embodiments, the PEG (e.g., conjugated to a lipid moiety or lipid, such as a stealth lipid) is "PEG-2K," also referred to as "PEG2000," which has an average molecular weight of about 2,000 daltons. PEG-2K is represented herein by the following formula (I): [ka] However, other embodiments of PEG known in the art may be used, for example, having a number average degree of polymerization of about 23 subunits (n=23) and / or 68 subunits (n=68). In some embodiments, n may range from about 30 to about 60. In some embodiments, n may range from about 35 to about 55. In some embodiments, n may range from about 40 to about 50. In some embodiments, n may range from about 42 to about 48. In some embodiments, n may be 45. In some embodiments, R may be selected from H, substituted alkyl, and unsubstituted alkyl. In some embodiments, R may be unsubstituted alkyl. In some embodiments, R may be methyl.
[0340] In any of the embodiments described herein, the PEG lipid may be PEG-dilaurylglycerol, PEG-dimyristoylglycerol (e.g., 1,2-dimyristoyl-rac-glycero-3-methylpolyoxyethylene glycol 2000 (PEG2k-DMG) or PEG-DMG (Cat. No. GM-020, NOF Corp., Tokyo, Japan), PEG-dipalmitoylglycerol, PEG-distearoylglycerol (PEG-DSPE) (Cat. No. DSPE-020CN, NOF Corp., Tokyo, Japan), PEG-dilaurylglycamide, PEG-dimyristylglycamide, PEG-dipalmitoylglycerol (e.g., 1,2-dimyristoyl-rac-glycero-3-methylpolyoxyethylene glycol 2000 (PEG2k-DMG) or PEG-DMG (Cat. No. GM-020 ... Mitoylglycamide, and PEG-distearoylglycamide, PEG-cholesterol (1-[8'-(cholest-5-ene-3[beta]-oxy)carboxamido-3',6'-dioxaoctanyl]carbamoyl-[omega]-methyl-poly(ethylene glycol), PEG-DMB (3,4-ditetradecoxylbenzyl-[omega]-methyl-poly(ethylene glycol) ether), 1,2-dimyristoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DMG) (catalog no. 880150P, Avanti Polar Lipids, Alabaster, AL, USA), 1,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DSPE) (catalog no. 880120C, Avanti Polar Lipids, Alabaster, AL, USA), Lipids, Alabaster, Alabama, USA), 1,2-distearoyl-sn-glycerol, methoxypolyethylene glycol (PEG2k-DSG; GS-020, NOF Corp., Tokyo, Japan), poly(ethylene glycol)-2000-dimethacrylate (PEG2k-DMA), and 1,2-distearyloxypropyl-3-amine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DSA). In one embodiment, the PEG lipid can be 1,2-dimyristoyl-rac-glycero-3-methylpolyoxyethylene glycol 2000 (PEG2k-DMG). In one embodiment, the PEG lipid can be PEG2k-DMG.In one embodiment, the PEG lipid may be PEG2k-DMG. In some embodiments, the PEG lipid may be PEG2k-DSG. In one embodiment, the PEG lipid may be PEG2k-DSPE. In one embodiment, the PEG lipid may be PEG2k-DMA. In one embodiment, the PEG lipid may be PEG2k-C-DMA. In one embodiment, the PEG lipid may be compound S027 disclosed in WO2016 / 010840 (paragraphs
[0240] to
[0244] ). In one embodiment, the PEG lipid may be PEG2k-DSA. In one embodiment, the PEG lipid may be PEG2k-C11. In some embodiments, the PEG lipid may be PEG2k-C14. In some embodiments, the PEG lipid may be PEG2k-C16. In some embodiments, the PEG lipid may be PEG2k-C18.
[0341] In a preferred embodiment, the PEG lipid comprises a glycerol group. In a preferred embodiment, the PEG lipid comprises a dimyristoyl glycerol (DMG) group. In a preferred embodiment, the PEG lipid comprises PEG-2k. In a preferred embodiment, the PEG lipid is PEG-DMG. In a preferred embodiment, the PEG lipid is PEG-2k-DMG. In a preferred embodiment, the PEG lipid is 1,2-dimyristoyl-rac-glycero-3-methoxypolyethylene glycol-2000. In a preferred embodiment, the PEG-2k-DMG is 1,2-dimyristoyl-rac-glycero-3-methoxypolyethylene glycol-2000.
[0342] LNP Lipid nanoparticles (LNPs) are well-known vehicles for the delivery of nucleotide and protein cargoes and may be used to deliver the polynucleotides, compositions, or pharmaceutical formulations disclosed herein. In some embodiments, the LNPs deliver nucleic acids, proteins, or nucleic acids and proteins.
[0343] As used herein, lipid nanoparticles (LNPs) refer to particles that contain a plurality (i.e., two or more) of lipid molecules physically associated with each other by intermolecular forces. LNPs can be, for example, microspheres (including unilamellar and multilamellar vesicles, e.g., lamellar phase lipid bilayers referred to as "liposomes," which in some embodiments are substantially spherical, and in certain embodiments, can contain an aqueous core, e.g., containing a substantial portion of RNA molecules), the dispersed phase of an emulsion, a micelle, or the internal phase of a suspension (see, for example, WO2017173054, the contents of which are incorporated herein by reference in their entirety). Any LNP known to those skilled in the art to be capable of delivering nucleotides to a subject can be utilized.
[0344] In some embodiments, the LNPs comprise a cationic lipid. In some embodiments, the LNPs comprise (9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate, also known as 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl(9Z,12Z)-octadeca-9,12-dienoate) (referred to herein as lipid A). In some embodiments, the LNPs comprise a molar ratio of cationic lipid amine to RNA phosphate (N:P) of about 4.5. In some embodiments, the LNPs comprise nonyl 8-((7,7-bis(octyloxy)heptyl)(2-hydroxyethyl)amino)octanoate. In some embodiments, the LNPs comprise a molar ratio of cationic lipid amines to RNA phosphate (N:P) of about 4.5 to 6.5. In some embodiments, the LNPs comprise a molar ratio of cationic lipid amines to RNA phosphate (N:P) of about 4.5. In some embodiments, the LNPs comprise a molar ratio of cationic lipid amines to RNA phosphate (N:P) of about 6.0.
[0345] In some embodiments, the disclosure includes a method for delivering a polynucleotide or a composition disclosed herein to a subject, where the polynucleotide is associated with a LNP. In some embodiments, the disclosure includes a method for delivering a first polynucleotide and a second polynucleotide to a subject, or a composition for delivering a first polynucleotide and a second polynucleotide to a subject, where the first polynucleotide and the second polynucleotide are associated with the same LNP, e.g., co-formulated with the same LNP. In some embodiments, the disclosure includes a method for delivering a first polynucleotide and a second polynucleotide to a subject, or a composition for delivering a first polynucleotide and a second polynucleotide to a subject, where the first polynucleotide and the second polynucleotide are each associated with a separate LNP, e.g., each polynucleotide is associated with a separate LNP for administration to a subject, or for use together, e.g., for simultaneous administration. In some embodiments, the first polynucleotide and the second polynucleotide encode NmeCas9 nickase and UGI. In some embodiments, the composition further comprises one or more guide RNAs. In some embodiments, the method further comprises delivering one or more guide RNAs.
[0346] In some embodiments, provided herein is a method of delivering any of the polynucleotides or compositions described herein to a cell or cell population or subject, including a cell or cell population of a subject in vivo, and any one or more of the components are associated with LNP.In some embodiments, the composition further comprises one or more guide RNAs.In some embodiments, the method further comprises delivering one or more guide RNAs.
[0347] In some embodiments, the present specification provides a composition comprising any of the polynucleotides or compositions described herein or the donor constructs disclosed herein, alone or in combination, and LNP.In some embodiments, the composition further comprises one or more guide RNAs.In some embodiments, the method further comprises delivering one or more guide RNAs.
[0348] In some embodiments, the LNPs associated with the polynucleotides or compositions disclosed herein are for use in the preparation of a medicament for treating a disease or disorder.
[0349] In some embodiments, a method of modifying a target gene is provided, the method comprising delivering to a cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles comprising a polynucleotide as disclosed herein, and one or more guide RNAs.
[0350] In some embodiments, at least one lipid-nucleic acid assembly composition comprises a lipid nanoparticle (LNP), and optionally, all lipid-nucleic acid assembly compositions comprise LNP. In some embodiments, at least one lipid-nucleic acid assembly composition is a lipoplex composition. In some embodiments, the lipid-nucleic acid assembly composition comprises an ionized lipid.
[0351] Electroporation is a well-known means for the delivery of cargo, and any electroporation methodology can be used for the delivery of polynucleotide or composition disclosed herein.In some embodiments, electroporation can be used to deliver any one of the polynucleotides or compositions disclosed herein.
[0352] In some embodiments, the present disclosure includes a method of delivering a polynucleotide, polypeptide, or composition disclosed herein to an ex vivo cell, wherein the polynucleotide or composition is associated with an LNP or is not associated with an LNP. In some embodiments, the LNP is also associated with one or more guide RNAs. See, for example, PCT / US2021 / 029446, which is incorporated herein by reference.
[0353] In some embodiments, a kit is provided that includes a polynucleotide, polypeptide, or composition disclosed herein.
[0354] In some embodiments, a pharmaceutical formulation is provided that comprises the polynucleotide, polypeptide, or composition disclosed herein. The pharmaceutical formulation may further comprise a pharma- ceutically acceptable carrier, such as water or a buffer. The pharmaceutical formulation may further comprise one or more pharma- ceutically acceptable excipients, such as stabilizers, preservatives, bulking agents, etc. The pharmaceutical formulation may further comprise one or more pharma- ceutically acceptable salts, such as sodium chloride. In some embodiments, the pharmaceutical formulation is formulated for intravenous administration. In some embodiments, the pharmaceutical formulation is non-pyrogenic. In some embodiments, the pharmaceutical formulation is sterile, particularly for pharmaceutical formulations for injection or infusion.
[0355] V. Exemplary Uses, Methods, and Treatments In some embodiments, the polynucleotides, expression constructs, compositions, lipid nanoparticles (LNPs), or pharmaceutical compositions disclosed herein are for use in, for example, gene therapy of target genes.
[0356] In some embodiments, there is provided the use of a polynucleotide, composition, or polypeptide disclosed herein in modifying a target gene in a cell.
[0357] In some embodiments, there is provided a use of a polynucleotide, composition, or polypeptide disclosed herein in the manufacture of a medicament for modifying a target gene in a cell.
[0358] In some embodiments, the polynucleotide or composition is formulated as a lipid-nucleic acid assembly composition, optionally a lipid nanoparticle.
[0359] In some embodiments, a method of modifying a target gene is provided, the method comprising delivering a polynucleotide, a polypeptide, or a composition disclosed herein to a cell.
[0360] In some embodiments, the polynucleotide, expression construct, composition, lipid nanoparticle (LNP), or pharmaceutical composition is for use in genome editing, e.g., editing a target gene, where the polynucleotide encodes an RNA-guided DNA-binding agent (e.g., NmeCas9).
[0361] In some embodiments, a polynucleotide, expression construct, composition, lipid nanoparticle (LNP), or pharmaceutical composition disclosed herein encoding a polypeptide disclosed herein is for use in expressing the polypeptide in a heterologous cell (e.g., a human or mouse cell).
[0362] In some embodiments, the polynucleotide, expression construct, composition, lipid nanoparticle (LNP), or pharmaceutical composition is for use in modifying a target gene, e.g., altering its sequence or epigenetic state, where the polynucleotide encodes an RNA-guided DNA-binding agent (e.g., NmeCas9).
[0363] In some embodiments, the polynucleotide, expression construct, composition, lipid nanoparticle (LNP), or pharmaceutical composition is for use in inducing a double-strand break (DSB) in a target gene. In some embodiments, the polynucleotide, expression construct, composition, lipid nanoparticle (LNP), or pharmaceutical composition is for use in inducing an indel in a target gene. In some embodiments, the use of the polynucleotide, expression construct, composition, lipid nanoparticle (LNP), or pharmaceutical composition disclosed herein is provided for the preparation of a medicament for genome editing, for example, editing a target gene, where the polynucleotide encodes an RNA-guided DNA binding agent (e.g., NmeCas9). In some embodiments, the use of the polynucleotide, expression construct, composition, lipid nanoparticle (LNP), or pharmaceutical composition disclosed herein that encodes a polypeptide disclosed herein is provided for the preparation of a medicament for expressing or increasing the expression of a polypeptide in a heterologous cell (e.g., a human cell or a mouse cell). In some embodiments, the use of the polynucleotides, expression constructs, compositions, lipid nanoparticles (LNPs), or pharmaceutical compositions disclosed herein is provided for the preparation of a medicament for modifying a target gene, e.g., altering its sequence or epigenetic state. In some embodiments, the use of the polynucleotides, expression constructs, compositions, lipid nanoparticles (LNPs), or pharmaceutical compositions disclosed herein is provided for the preparation of a medicament for inducing a double-strand break (DSB) in a target gene. In some embodiments, the use of the polynucleotides, expression constructs, compositions, lipid nanoparticles (LNPs), or pharmaceutical compositions disclosed herein is provided for the preparation of a medicament for inducing an indel in a target gene.
[0364] In some embodiments, the target gene is a transgene. In some embodiments, the target gene is an endogenous gene. The target gene can be a target gene in a subject, e.g., a mammal (e.g., a human). In some embodiments, the target gene is a target gene in an organ, e.g., a liver, e.g., a mammalian liver, e.g., a human liver. In some embodiments, the target gene is a target gene in a liver cell, e.g., a mammalian liver cell, e.g., a human liver cell. In some embodiments, the target gene is a target gene in a hepatocyte, a mammalian hepatocyte, e.g., a human hepatocyte. In some embodiments, the liver cell or hepatocyte is in situ. In some embodiments, the liver cell or hepatocyte is isolated in culture, e.g., a primary culture. In some embodiments, the target cell is a peripheral blood mononuclear cell (PBMC), e.g., a mammalian PBMC, e.g., a human PBMC. In some embodiments, the PBMC is an immune cell, e.g., a T cell, a B cell, a NK cell. In some embodiments, the cells are pluripotent cells, e.g., mammalian pluripotent cells, e.g., human pluripotent cells. In some embodiments, the target cells are stem cells, e.g., mammalian stem cells, e.g., human stem cells. In some embodiments, the stem cells are present in bone marrow. In some embodiments, the stem cells are induced pluripotent stem cells (iPCS). In some embodiments, the cells are isolated, e.g., by ex vivo culture.
[0365] Also provided are methods corresponding to the uses disclosed herein, comprising administering to a subject a polynucleotide, expression construct, composition, lipid nanoparticle (LNP), or pharmaceutical composition disclosed herein, or contacting a cell, such as those described above, with a polynucleotide, LNP, or pharmaceutical composition disclosed herein, e.g., to express a polypeptide disclosed herein or increase expression of a polypeptide disclosed herein in a heterologous cell (e.g., a human or mouse cell).
[0366] In any of the foregoing embodiments involving a subject, the subject can be a mammal. In any of the foregoing embodiments involving a subject, the subject can be a human.
[0367] In some embodiments, the polynucleotides, expression constructs, compositions, lipid nanoparticles (LNPs), or pharmaceutical compositions disclosed herein are administered intravenously or are for intravenous administration.
[0368] In some embodiments, a single administration of the polynucleotide, LNP, or pharmaceutical composition disclosed herein is sufficient to knock down the expression of the target gene product. In some embodiments, a single administration of the polynucleotide, LNP, or pharmaceutical composition disclosed herein is sufficient to knock out the expression of the target gene product. In other embodiments, two or more administrations of the polynucleotide, LNP, or pharmaceutical composition disclosed herein may be beneficial to maximize the cumulative effect of editing, modification, indel formation, DSB formation, etc.
[0369] VI. Exemplary DNA Molecules, Vectors, Expression Constructs, Host Cells, and Methods of Production In certain embodiments, the present disclosure provides a DNA molecule comprising an ORF sequence that encodes the polypeptide disclosed herein.In some embodiments, in addition to the ORF sequence, the DNA molecule further comprises a nucleic acid that does not encode the polypeptide disclosed herein.The nucleic acid that does not encode the polypeptide includes, but is not limited to, the nucleic acid that encodes promoter, enhancer, control sequence, and guide RNA.
[0370] In some embodiments, the DNA molecule further comprises a nucleotide sequence encoding a crRNA, a trRNA, or a crRNA and a trRNA. In some embodiments, the nucleotide sequence encoding the crRNA, the trRNA, or the crRNA and a trRNA comprises or consists of a guide sequence flanked by all or part of repeat sequences from a naturally occurring CRISPR / Cas system. The nucleic acid comprising or consisting of the crRNA, the trRNA, or the crRNA and a trRNA may further comprise a vector sequence comprising or consisting of a nucleic acid not found in nature with the crRNA, the trRNA, or the crRNA and a trRNA. In some embodiments, the crRNA and the trRNA are encoded by non-contiguous nucleic acids within one vector. In other embodiments, the crRNA and the trRNA may be encoded by contiguous nucleic acids. In some embodiments, the crRNA and the trRNA are encoded by opposite strands of a single nucleic acid. In other embodiments, the crRNA and the trRNA are encoded by the same strand of a single nucleic acid.
[0371] In some embodiments, the DNA molecule further comprises a promoter operably linked to a sequence encoding any of the ORFs encoding the polypeptides disclosed herein. In some embodiments, the DNA molecule is an expression construct suitable for expression in mammalian cells, such as human cells or mouse cells, such as human hepatocytes or rodent (e.g., mouse) hepatocytes. In some embodiments, the DNA molecule is an expression construct suitable for expression in cells of a mammalian organ, such as human liver or rodent (e.g., mouse) liver. In some embodiments, the DNA molecule is a plasmid or episome. In some embodiments, the DNA molecule is contained in a host cell, such as a bacterium or a cultured eukaryotic cell. Exemplary bacteria include proteobacteria, such as E. coli. Exemplary cultured eukaryotic cells include primary hepatocytes, including hepatocytes of rodent (e.g., mouse) or human origin, hepatocyte cell lines, including hepatocytes of rodent (e.g., mouse) or human origin, human cell lines, rodent (e.g., mouse) cell lines, CHO cells, fungi, e.g., fission yeast or budding yeast, e.g., Saccharomyces, e.g., S. cerevisiae, and insect cells.
[0372] In some embodiments, a method for producing the mRNA disclosed herein is provided. In some embodiments, such a method comprises contacting a DNA molecule described herein with an RNA polymerase under conditions that allow transcription. In some embodiments, the contacting is performed in vitro, for example, in a cell-free system. In some embodiments, the RNA polymerase is an RNA polymerase of bacteriophage origin, such as T7 RNA polymerase. In some embodiments, a NTP is provided that comprises at least one modified nucleotide as described above. In some embodiments, the NTP comprises at least one modified nucleotide as described above and does not comprise UTP.
[0373] In some embodiments, methods of producing a polynucleotide disclosed herein are provided. In some embodiments, such methods include contacting an expression construct disclosed herein with an RNA polymerase comprising at least one modified nucleotide and an NTP. In some embodiments, the modified nucleotide comprises a modified uridine. In further embodiments, at least 80% of the uridine positions are modified uridines. In further embodiments, at least 90% of the uridine positions are modified uridines. In further embodiments, 100% of the uridine positions are modified uridines. In further embodiments, the modified uridine comprises or is a substituted uridine, pseudouridine, or substituted pseudouridine. In further embodiments, the modified uridine comprises or is an N1-methyl-pseudouridine. In some embodiments, the expression construct comprises an encoded poly-A tail sequence.
[0374] In some embodiments, the polynucleotides disclosed herein may be contained within or delivered by a vector system of one or more vectors. In some embodiments, one or more of the vectors, or all of the vectors, may be DNA vectors. In some embodiments, one or more of the vectors, or all of the vectors, may be RNA vectors. In some embodiments, one or more of the vectors, or all of the vectors, may be circular. In other embodiments, one or more of the vectors, or all of the vectors, may be linear. In some embodiments, one or more of the vectors, or all of the vectors, may be encapsulated in a lipid nanoparticle, a liposome, a non-lipid nanoparticle, or a viral capsid. Non-limiting examples of vectors include plasmids, phagemids, cosmids, artificial chromosomes, mini-chromosomes, transposons, viral vectors, and expression vectors.
[0375] Non-limiting examples of viral vectors include adeno-associated virus (AAV) vectors, lentivirus vectors, adenovirus vectors, helper-dependent adenovirus vectors (HDAd), herpes simplex virus (HSV-1) vectors, bacteriophage T4, baculovirus vectors, and retrovirus vectors. In some embodiments, the viral vector can be an AAV vector. In other embodiments, the viral vector can be a lentivirus vector. In some embodiments, the lentivirus can be non-integrating. In some embodiments, the viral vector can be an adenovirus vector. In some embodiments, the adenovirus can be a high cloning capacity, or "gutless," adenovirus, in which all coding viral regions are separated from the 5' and 3' inverted terminal repeats (ITRs) and the packaging signal ("I") has been deleted from the virus to increase its packaging capacity. In yet other embodiments, the viral vector can be an HSV-1 vector. In some embodiments, the HSV-1 based vector is helper-dependent, while in other embodiments, it is not vector-dependent. For example, an amplicon vector that retains only the packaging sequence requires a helper virus with structural components for packaging, while a 30 kb deleted HSV-1 vector that removes non-essential viral functions does not require a helper virus. In additional embodiments, the viral vector can be bacteriophage T4. In some embodiments, bacteriophage T4 can package any linear or circular DNA or RNA molecule when the viral head is empty. In further embodiments, the viral vector can be a baculovirus vector. In yet further embodiments, the viral vector can be a retrovirus vector. In embodiments that use AAV or lentivirus vectors with smaller cloning capacity, it may be necessary to use two or more vectors to deliver all the components of the vector system as disclosed herein.For example, one AAV vector can include a sequence encoding a Cas protein, and a second AAV vector can include one or more guide sequences.
[0376] In some embodiments, the vector may be capable of driving expression of one or more coding sequences in a cell, for example, coding sequences of mRNAs disclosed herein. In some embodiments, the cell may be a prokaryotic cell, such as, for example, a bacterial cell. In some embodiments, the cell may be a eukaryotic cell, such as, for example, a yeast, plant, insect, or mammalian cell. In some embodiments, the eukaryotic cell may be a mammalian cell. In some embodiments, the eukaryotic cell may be a rodent cell. In some embodiments, the eukaryotic cell may be a human cell. Suitable promoters for driving expression in different types of cells are known in the art. In some embodiments, the promoter may be wild type. In other embodiments, the promoter may be modified for more efficient or effective expression. In still other embodiments, the promoter may be truncated but still maintain its function. For example, the promoter may have a normal size or a reduced size suitable for proper packaging of the vector into a virus.
[0377] In some embodiments, the vector system may contain one copy of the nucleotide sequence comprising the ORF encoding the polypeptide disclosed herein. In other embodiments, the vector system may contain two or more copies of the nucleotide sequence encoding the polypeptide disclosed herein. In some embodiments, the nucleotide sequence encoding the polypeptide disclosed herein may be operably linked to at least one transcriptional or translational control sequence. In some embodiments, the nucleotide sequence encoding the nuclease may be operably linked to at least one promoter.
[0378] In some embodiments, the promoter may be constitutive, inducible, or tissue-specific. In some embodiments, the promoter may be a constitutive promoter. Non-limiting exemplary constitutive promoters include cytomegalovirus immediate early promoter (CMV), simian virus (SV40) promoter, adenovirus major late (MLP) promoter, Rous sarcoma virus (RSV) promoter, mouse mammary tumor virus (MMTV) promoter, phosphoglyserate kinase (PGK) promoter, elongation factor alpha (EF1a) promoter, ubiquitin promoter, actin promoter, tubulin promoter, immunoglobulin promoter, functional fragments thereof, or any combination of the above. In some embodiments, the promoter may be a CMV promoter. In some embodiments, the promoter may be a truncated CMV promoter. In other embodiments, the promoter may be an EF1a promoter. In some embodiments, the promoter may be an inducible promoter. Non-limiting exemplary inducible promoters include those that are inducible by heat shock, light, chemicals, peptides, metals, steroids, antibiotics, or alcohol. In some embodiments, the inducible promoter can be one that has a low basal (uninduced) expression level, such as the Tet-On® promoter (Clontech).
[0379] In some embodiments, the promoter may be a tissue-specific promoter, for example a promoter specific for expression in the liver.
[0380] The vector may further comprise a nucleotide sequence encoding at least one guide RNA. In some embodiments, the vector comprises one copy of the guide RNA. In other embodiments, the vector comprises two or more copies of the guide RNA. In embodiments with two or more guide RNAs, the guide RNAs may not be identical so as to target different target sequences, or may be identical in that they target the same target sequence. In some embodiments where the vector comprises two or more guide RNAs, each guide RNA may have other distinct properties, such as activity or stability in a ribonucleoprotein complex with an RNA-guided DNA-binding agent (e.g., NmeCas9). In some embodiments, the nucleotide sequence encoding the guide RNA may be operably linked to at least one transcriptional or translational control sequence (e.g., promoter, 3'UTR, or 5'UTR). In one embodiment, the promoter is a tRNA promoter, e.g., ... Lys3, or tRNA chimeras. See Mefferd et al., RNA. 2015 21:1683-9; Scherer et al., Nucleic Acids Res. 2007 35:2620-2628. In some embodiments, the promoter may be recognized by RNA polymerase III (Pol III). Non-limiting examples of Pol III promoters include U6 and H1 promoters. In some embodiments, the nucleotide sequence encoding the guide RNA may be operably linked to a mouse or human U6 promoter. In other embodiments, the nucleotide sequence encoding the guide RNA may be operably linked to a mouse or human H1 promoter. In embodiments with two or more guide RNAs, the promoters used to drive expression may be the same or different. In some embodiments, the nucleotides encoding the crRNA of the guide RNA and the nucleotides encoding the trRNA of the guide RNA may be provided on the same vector. In some embodiments, the nucleotides encoding the crRNA and the trRNA may be driven by the same promoter. In some embodiments, the crRNA and the trRNA may be transcribed into a single transcript. For example, the crRNA and trRNA can be processed from a single transcript to form a dual molecule guide RNA. Alternatively, the crRNA and trRNA can be transcribed into a single molecule guide RNA. In other embodiments, the crRNA and trRNA can be driven by corresponding promoters on the same vector. In yet other embodiments, the crRNA and trRNA can be encoded by different vectors.
[0381] In some embodiments, the composition comprises a vector system, the system comprising two or more vectors. In some embodiments, the vector system can comprise one single vector. In other embodiments, the vector system can comprise two vectors. In additional embodiments, the vector system can comprise three vectors. When different polynucleotides are used for multiplexing, or when multiple copies of a polynucleotide are used, the vector system can comprise four or more vectors.
[0382] In some embodiments, a host cell is provided, the host cell comprising a vector, expression construct, or plasmid disclosed herein.
[0383] In some embodiments, the vector system can include an inducible promoter to initiate expression only after delivery to a target cell. Non-limiting exemplary inducible promoters include those that can be induced by heat shock, light, chemicals, peptides, metals, steroids, antibiotics, or alcohol. In some embodiments, the inducible promoter can be one that has a low basal (uninduced) expression level, such as the Tet-On® promoter (Clontech).
[0384] In additional embodiments, the vector system can include tissue-specific promoters to initiate expression only after delivery into a particular tissue.
[0385] VII. Determining ORF Validity The efficacy of a polynucleotide comprising an ORF encoding a polypeptide disclosed herein can be determined, for example, using any of the art-recognized methods for detecting the presence, expression level, or activity of a particular polypeptide when the polypeptide is expressed together with other components for a target function or system, such as enzyme-linked immunosorbent assay (ELISA), other immunological methods, Western blot, liquid chromatography-mass spectrometry (LC-MS), FACS analysis, HiBiT peptide assay (Promega), or other assays described herein, or by methods for determining enzyme activity levels in biological samples (e.g., cells, cell lysates or extracts, conditioned medium, whole blood, serum, plasma, urine, or tissues), such as in vitro activity assays. Exemplary assays for the activity of various encoded polypeptides, such as RNA-guided DNA binders, described herein include assays for indel formation, deamination, or mRNA or protein expression. In some embodiments, the efficacy of a polynucleotide comprising an ORF encoding a polypeptide disclosed herein is determined based on an in vitro model.
[0386] 1. Determining the availability of ORFs encoding RNA-guided DNA-binding agents In some embodiments, the efficacy of the mRNA is determined when expressed together with other components of the RNP, e.g., at least one gRNA, e.g., a gRNA targeting TTR.
[0387] RNA-guided DNA binding agents with cleavase activity (e.g., NmeCas9) can result in double-strand breaks in DNA. Non-homologous end joining (NHEJ) is a process by which double-strand breaks (DSBs) in DNA are repaired by religation of the broken ends, which can result in errors in the form of insertion / deletion (indel) mutations. The DNA ends of DSBs frequently undergo enzymatic processing, resulting in the addition or removal of nucleotides in one or both strands before the ends are rejoined. These additions or removals prior to rejoining result in the presence of insertion or deletion (indel) mutations in the DNA sequence at the NHEJ repair site. Many mutations resulting from indels alter the reading frame or introduce premature stop codons, thus producing non-functional proteins.
[0388] In some embodiments, the effectiveness of the mRNA encoding nuclease is determined based on an in vitro model.In some embodiments, the in vitro model is HEK293 cell.In some embodiments, the in vitro model is HUH7 human liver cancer cell.In some embodiments, the in vitro model is primary hepatocyte, for example primary human or mouse hepatocyte.
[0389] In some embodiments, detecting gene editing events, such as the formation of insertion / deletion ("indel") mutations, utilizes linear amplification with tagged primers and isolating the tagged amplification products (hereinafter referred to as "LAM-PCR" or "linear amplification (LA)") methods as described in WO2018 / 067447 or Schmidt et al., Nature Methods 4:1051-1057 (2007), or next generation sequencing ("NGS"; e.g., using the Illumina NGS platform), as described below, or other methods known in the art for detecting indel mutations.
[0390] For example, to quantitatively determine the efficiency of editing at a target location in a genome, in NGS methods, genomic DNA is isolated and deep sequencing is used to identify the presence of insertions and deletions introduced by gene editing. PCR primers around the target site (e.g., TTR) are designed to amplify the genomic region of interest. An additional PCR is performed according to the manufacturer's (Illumina) protocol, adding the necessary chemistry for sequencing. The amplicons are sequenced on an Illumina MiSeq instrument. The reads are aligned to a reference genome (e.g., mm10) after filtering out those with low quality scores. The resulting file containing the reads is mapped against a reference genome (BAM file), reads that overlap with the target sequence of interest are selected, and the number of wild-type reads and the number of reads containing insertions, substitutions, or deletions are calculated. The editing ratio (e.g., "editing efficiency" or "editing rate") is defined as the total number of sequence reads containing insertions or deletions relative to the total number of sequence reads containing wild-type. EXAMPLES
[0391] The following examples are provided to illustrate certain disclosed embodiments and should not be construed in any way as limiting the scope of the disclosure.
[0392] Example 1. Materials and Methods In vitro transcription ("IVT") of nuclease mRNA Capped and polyadenylated mRNA containing N1-methylpseudo-U was generated by in vitro transcription using routine methods. For example, plasmid DNA containing the T7 promoter, sequence for transcription, and polyadenylation region was linearized with XbaI according to the manufacturer's protocol. XbaI was inactivated by heating. The linearized plasmid was purified from enzymes and buffer salts. IVT reactions to generate modified mRNA were performed by incubating at 37°C. The following components were added to the reaction mixture: 50 ng / μL linearized plasmid, 2-5 mM, 10-25 mM ARCA (Trilink), 5 U / μL T7 RNA polymerase, 1 U / μL mouse RNase inhibitor (NEB), 0.004 U / μL inorganic E. coli pyrophosphatase (NEB), and 1x reaction buffer. TURBO DNase (Thermo Fisher) was added to a final concentration of 0.01 U / μL, and the reaction was incubated at 37°C to remove the DNA template.
[0393] The mRNA was purified using MegaClear transcription clean-up kit (Thermo Fisher) or RNeasy Maxi kit (Qiagen) according to the manufacturer's protocol. Alternatively, the mRNA was purified using a precipitation protocol (in some cases followed by HPLC-based purification). Briefly, after DNase digestion, the mRNA was purified using LiCl precipitation, ammonium acetate precipitation, and sodium acetate precipitation. For HPLC-purified mRNA, after LiCl precipitation and reconstitution, the mRNA was purified by RP-IP HPLC (see, e.g., Kariko, et al. Nucleic Acids Research, 2011, Vol. 39, No. 21 el42). The fractions selected for pooling were combined and desalted by sodium acetate / ethanol precipitation as described above. In a further alternative method, the mRNA was purified by LiCl precipitation and then further purified by tangential flow filtration. RNA concentrations were determined by measuring light absorbance at 260 nm (Nanodrop) and transcripts were analyzed by capillary electrophoresis on a Bioanalyzer (Agilent).
[0394] When the sequences cited in this paragraph are referred to below in terms of RNA, it is understood that T should be replaced by U (which may be a modified nucleoside as described above). The messenger RNA used in the examples includes a 5' cap and a 3' polyadenylation sequence, for example, up to 100 nt. Guide RNAs were chemically synthesized by commercial vendors or using standard in vitro synthesis techniques with modified nucleotides.
[0395] Streptococcus pyogenes ("Spy") Cas9 mRNA was generated from plasmid DNA encoding an open reading frame according to SEQ ID NOs: 43-47 and 49 (see sequences in Table 39A).
[0396] Hepatocyte preparation Primary mouse hepatocytes (PMH), primary rat hepatocytes (PRH), primary human hepatocytes (PHH), and primary cynomolgus monkey hepatocytes (PCH) were prepared as follows: PMH (Gibco, MCM837 unless otherwise specified), PRH (Gibco, Rs977 unless otherwise specified), PCH (In Vitro ADMET Laboratories, 10136011 unless otherwise specified), and PHH (Gibco, Hu8284 unless otherwise specified) were thawed and resuspended in 50 mL of cryopreserved hepatocyte recovery medium (CHRM) (Invitrogen, CM7000), followed by centrifugation. Cells were resuspended in Hepatocyte medium with seeding supplement, Williams E medium seeding supplement with FBS (Gibco, catalog number A13450). Cells were pelleted by centrifugation, resuspended in medium, and seeded at 20,000 cells / well for PMH and 30,000 for PHH onto Bio-coat collagen I-coated 96-well plates (Corning #354407). The seeded cells were incubated at 37°C and 5% CO 2 The cells were left to adhere for 4-6 hours in a tissue culture incubator at ambient temperature. After incubation, the cells were examined for monolayer formation, washed once, and plated in 100 μL of Hepatocyte Maintenance Medium, Williams E Medium (Gibco, Cat. No. A12176-01) and Supplement Pack (Gibco, Cat. No. CM3000).
[0397] HEK cell preparation HEK-293 cells (ATCC, CRL-1573 unless otherwise specified) were thawed and resuspended in serum-free Dulbecco's Modified Eagle Medium (Corning #10-013-CV) with 10% FBS content (Gibco #A31605-02) and 1% penicillin-streptomycin (Gibco #15070063). Cells were counted and seeded onto 96-well tissue culture plates (Falcon, #353072) in Dulbecco's Modified Eagle Medium (Corning #10-013-CV) with 10% FBS content (Gibco #A31605-02). The seeded cells were incubated at 37°C and 5% CO 2The plates were left at ambient temperature in a tissue culture incubator for 18 hours to allow adhesion.
[0398] Preparation of LNP formulations containing sgRNA and Cas9 mRNA In general, lipid nanoparticle components were dissolved in 100% ethanol at various molar ratios. RNA cargo (e.g., Cas9 mRNA and sgRNA) was dissolved in 25 mM citric acid, 100 mM NaCl, pH 5.0, resulting in a concentration of RNA cargo of approximately 0.45 mg / mL. The LNP used was the ionized lipid ((9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate, also known as 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl The lipid mixture contained (9Z,12Z)-octadeca-9,12-dienoate (also referred to herein as lipid A), cholesterol, disteroylphosphatidylcholine (DSPC), and 1,2-dimyristoyl-rac-glycero-3-methylpolyoxyethylene glycol 2000 (PEG2k-DMG) in a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. The LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of about 6. The LNPs used contain a single RNA species, such as Cas9 mRNA or sgRNA. LNPs are similarly prepared with a mixture of Cas9 mRNA and guide RNA.
[0399] LNPs were prepared by impingement jet mixing of lipids in ethanol with two volumes of RNA solution and one volume of water using cross-flow technology. First, lipids in ethanol were mixed with two volumes of RNA solution by cross mixing. A fourth water stream was then mixed with the cross outlet stream through a series of tees (see FIG. 2 of WO2016010840). The LNPs were kept at room temperature for 1 hour and further diluted with water (approximately 1:1 v / v). The diluted LNPs were buffer exchanged into 50 mM Tris, 45 mM NaCl, 5% (w / v) sucrose, pH 7.5 (TSS) and concentrated as required by methods known in the art. The resulting mixture was then filtered using a 0.2 μm sterile filter. Characterization of the final LNPs was performed to determine encapsulation efficiency, polydispersity index, and average particle size. The final LNPs were stored at 4° C. or −80° C. until further use.
[0400] sgRNA and Cas9 mRNA lipofection Lipofection of Cas9 mRNA and gRNA used a premixed lipid formulation. The lipofection reagent contained ionizable lipid A, cholesterol, DSPC, and PEG2k-DMG in a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. The mixture was reconstituted in 100% ethanol and then mixed with RNA (e.g., Cas9 mRNA and gRNA) at a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6.0.
[0401] Next generation sequencing ("NGS") and editing efficiency analysis Genomic DNA was extracted using a commercially available kit, for example, QuickExtract™ DNA Extraction Solution (Lucigen, Cat. No. QE09050), according to the manufacturer's protocol. To quantitatively determine the efficiency of editing at the target location in the genome, deep sequencing was utilized to identify the presence of insertions and deletions introduced by gene editing. PCR primers were designed based on the target site in the gene of interest (e.g., TRAC) to amplify the genomic segment of interest. The design of primer sequences was performed according to standards in the field.
[0402] An additional PCR was performed according to the manufacturer's (Illumina) protocol, and sequencing chemistry was added. Amplicons were sequenced on an Illumina MiSeq instrument. After removing reads with low quality scores, reads were aligned to the human reference genome (e.g., hg38). Reads overlapping the target region of interest were realigned to the local genomic sequence to improve alignment. The number of wild-type reads relative to the number of reads containing CT mutations, CA / G mutations, or indels was then calculated. Insertions and deletions were scored in a 20 bp region centered around the predicted Cas9 cleavage site. The indel percentage is defined as the total number of sequencing reads with one or more bases inserted or deleted within the 20 bp score region divided by the total number of sequencing reads containing wild type. CT or CA / G mutations were scored in a 40 bp region including 10 bp upstream and 10 bp downstream of the 20 bp sgRNA target sequence. The CT editing percentage is defined as the total number of sequencing reads with one or more CT mutations within a 40 bp region divided by the total number of sequencing reads containing the wild type. The percentage of CA / G mutations is calculated similarly.
[0403] Example 2. In vitro editing with selected guides in primary mouse hepatocytes (PMH) A modified sgRNA screen was performed to evaluate the editing efficiency of 95 different sgRNAs targeting various sites within the mouse TTR gene. Based on that study, two sgRNAs (G021320 and G021256) were selected for evaluation in a dose-response assay. These two test guides were compared to the mouse TTR SpyCas9 guide (G000502) with a 20 nucleotide guide sequence. The tested NmeCas9 sgRNA targeting the mouse TTR gene comprises a 24 nucleotide guide sequence (represented by N) and a guide scaffold as follows: mN*mNNNNNNNNmNNNNNNNNNNNNNNNNNmGUUGmUmAmGmCUCCCmUmGmAmAmAmCmCGUUmGmCUAmCAAU*AAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGCmAmAmCmGCUCUmGmCCmUmUmCmUGmGCmAmUC*mG*mU*mU (SEQ ID NO:508), where A, C, G, U, and N are adenine, cytosine, guanine, uracil, and optional ribonucleotides, respectively, unless otherwise specified. m indicates a 2'O-methyl modification and * indicates a phosphorothioate linkage between nucleotides. Unmodified and modified versions of the guide are provided in Table 39B.
[0404] Guide and Cas9 mRNA were lipofected into primary mouse hepatocytes (PMH) as described below. PMH (In Vitro ADMET Laboratories MCM114) were prepared as described in Example 1. Lipofection was performed as described in Example 1 using a dose response of sgRNA and mRNA. Briefly, cells were incubated at 37° C., 5% CO 2After incubation at 37°C for 24 hours, the cells were treated with lipoplexes. The lipoplexes were incubated in maintenance medium containing 10% fetal bovine serum (FBS) for 10 minutes at 37°C. After incubation, the lipoplexes were added to mouse hepatocytes in an 8-point 3-fold dose-response assay starting with the highest dose of 300ng Cas9 mRNA and 50nM sgRNA. Only gRNA doses are listed in Table 5, but messenger RNA doses increase and decrease with gRNA dose in each condition. After 72 hours of treatment, cells were lysed and NGS analysis was performed as described in Example 1.
[0405] A dose response of editing efficiency for each guide concentration was performed in triplicate samples. Table 5 shows the calculated EC 50 The mean percent edited and standard deviation (SD) of the values are shown. The means and standard deviations (SD) are shown in FIG. 1. [Table 6]
[0406] Example 3. sgRNA:mRNA ratios for sgRNA or pgRNA using LNPs A study was conducted to evaluate the editing efficiency of sgRNA designs containing PEG linkers (pgRNA). This study compared two gRNAs targeting TTR with the same guide sequence, one of which contained three PEG linkers within the constant region of the guide (pgRNA, G021846) and one of which did not (G021845), as shown in Table 39B. Guide and mRNA were formulated in separate LNPs and mixed in the desired ratio for delivery to primary mouse hepatocytes (PMH) via lipid nanoparticles (LNPs).
[0407] Unless otherwise noted, PMH cells were prepared, treated, and analyzed as described in Example 1. PMH cells from In Vitro ADMET Laboratories (Lot No. MCM114) were seeded at a density of 15,000 cells / well. Cells were treated with LNPs as described below. LNPs were generally prepared as described in Example 1. LNPs were prepared with a lipid composition of 50 / 9 / 38 / 3, expressed as the molar ratio of ionized lipid A / cholesterol / DSPC / PEG, respectively. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of about 6. LNPs encapsulated a single RNA species, either gRNA G021845, gRNA G021846, or mRNA (mRNA M), as described in Example 1.
[0408] PMH cells were treated with various amounts of LNPs at gRNA to mRNA weight ratios of 1:4, 1:2, 1:1, 2:1, 4:1, or 8:1 of RNA cargo. Duplicate samples were included in each assay. Guides were assayed in an eight-point three-fold dose-response curve starting at a total RNA concentration of 1 ng / μL, as shown in Table 6. The average percent editing results are shown in Table 6. Figure 2A shows the average percent editing of sgRNA G021845, and Figure 2B shows the average percent editing of sgRNA G021846. "ND" in the table represents a value that was not detectable due to experimental failure. [Table 7-1] [Table 7-2]
[0409] Example 3.1. In vitro editing of modified pegylated guide (pgRNA) in PMH using LNPs Modified pgRNAs with the same target site in the mouse TTR gene were assayed to assess editing efficiency in PMH cells.
[0410] Unless otherwise noted, PMH cells were prepared, treated, and analyzed as described in Example 1. PMH cells from In Vitro ADMET Laboratories (Lot No. MC148) were used and seeded at a density of 15,000 cells / well. LNP formulations were prepared as described in Example 1. LNPs were prepared with a lipid composition of 50 / 9 / 38 / 3, expressed as the molar ratio of ionized lipid A / cholesterol / DSPC / PEG, respectively. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of about 6, and gRNA or mRNA as shown in Table 7.
[0411] PMH in 100 μl of medium were treated with LNP for 30 ng total mRNA (mRNA P) by weight, and the amount of LNP for gRNA is shown in Table 7. Samples were run in duplicate. The average editing results of PMH are shown in Table 7 and Figure 3. [Table 8-1] [Table 8-2]
[0412] Example 4. Nme2-mRNA studies Example 4.1 - In vitro editing in primary mouse hepatocytes Messenger mRNAs encoding Nme2Cas9 ORFs with different NLS arrangements were assayed for editing efficiency in primary mouse hepatocytes (PMH).
[0413] PMH was prepared as described in Example 1. Lipofection was performed using Lipofectamine MessengerMAX transfection reagent (Invitrogen LMRNA001) according to the manufacturer's protocol to transfect cells with 100 nM sgRNA G020361 targeting mouse PCSK9 and mRNA at the concentrations listed in Table 8. Triplicate samples were included in the assay. After 72 hours of incubation at 37°C in maintenance medium, cells were harvested and NGS analysis was performed as described in Example 1. The average editing results with standard deviation (SD) are shown in Table 8 and Figure 4. [Table 9-1] [Table 9-2]
[0414] Example 4.2 - Dose response of Nme2 ORF mutants and guides with chemical modification mutations Messenger mRNAs encoding Nme2Cas9 ORFs with different NLS configurations were assayed for editing efficiency in primary human hepatocytes (PHH) and HEK-293 cells. The assays were performed using gRNAs with identical guide sequences targeting the VEGFA locus TS47, with gRNAs of various lengths and chemical modification patterns. PHH cells were prepared as described in Example 1. HEK293 cells were thawed and seeded at a density of 30,000 cells / well in 96-well plates in DMEM (Corning, 10-013-CV) containing 10% FBS and incubated for 24 hours. Lipofection was performed using Lipofectamine MessengerMAX transfection reagent (Invitrogen LMRNA001) according to the manufacturer's protocol. Cells were transfected with gRNA at the concentrations listed in Tables 9-10 using a dose-response 1:3 dilution series starting with the highest dose of 100 nM gRNA and 1 ng / μL mRNA. Multiple samples were included in the assay. After 72 hours of incubation at 37° C., cells were harvested and subjected to NGS analysis as described in Example 1. The average editing results with standard deviation (SD) are shown in Table 9 and Figures 5A-5C for HEK cells and in Table 10 and Figures 5D-5F for PHH cells. [Table 10] [Table 11]
[0415] Example 4.3 - Dose response of Nme2 NLS variants with LNP in PMH Messenger mRNAs encoding Nme2Cas9 ORFs with different NLS arrangements were assayed for editing efficiency in primary mouse hepatocytes (PMH). The assay tested guides targeting the mouse TTR locus and included both sgRNA and pgRNA designs.
[0416] PMH was prepared as in Example 1. LNP was prepared generally as described in Example 1, using a single RNA species as shown in Table 11 as cargo. LNP was prepared with a lipid composition of 50 / 9 / 38 / 3, expressed as molar ratios of ionizable lipid A / cholesterol / DSPC / PEG, respectively. LNP was formulated with a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6.
[0417] Cells were treated with gRNA-containing LNPs at 60 ng / 100 μl by RNA weight, and mRNA-containing LNPs as shown in Table 11. Cells were incubated for 72 hours at 37° C. in Williams E medium (Gibco, A1217601) with maintenance supplements and 10% fetal bovine serum. After 72 hours of incubation at 37° C., cells were harvested and editing was assessed by NGS as described in Example 1. The average percent editing data is shown in Table 11 and FIG. 6. [Table 12-1] [Table 12-2]
[0418] Example 4.4 - Dose response of Nme2 NLS variants with LNP in PMH Messenger mRNAs encoding Nme2Cas9 ORFs with different NLS arrangements were assayed for editing efficiency in primary mouse hepatocytes (PMH).
[0419] PMH (Gibco, MC148) was prepared as described in Example 1. LNPs were prepared generally as described in Example 1, using a single RNA species as cargo. LNPs were prepared with a lipid composition of 50 / 9 / 38 / 3, expressed as molar ratios of ionizable lipid A / cholesterol / DSPC / PEG, respectively. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6.
[0420] Cells were treated with 30ng gRNA by RNA weight / 100μl gRNA G021844-containing LNPs and mRNA-containing LNPs as shown in Table LS4. Cells were incubated for 24 hours in Williams E medium (Gibco, A1217601) with maintenance supplements and 10% fetal bovine serum. After 72 hours of incubation, cells were harvested and editing was assessed by NGS as described in Example 1. The average editing percentage data is shown in Table 12 and Figure 7. [Table 13-1] [Table 13-2] [Table 13-3]
[0421] Example 5 - NmeCas9 Protein Expression Example 5.1 Protein Expression in Primary Human Hepatocytes To quantify the expression of each mRNA construct, mRNA and protein expression levels were measured following LNP delivery of mRNA encoding either SpyCas9 or NmeCas9 into primary human hepatocytes.
[0422] PHH cells were prepared as described in Example 1. LNPs were prepared generally as described in Example 1, using a single RNA species as cargo. LNPs were made with a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6.
[0423] Cells were administered one LNP containing mRNA (mRNA only) or two LNPs containing either mRNA or gRNA. Each LNP was applied to cells at 16.7ng total RNA cargo / 100μl. After LNP treatment, cells were incubated for 24 hours in Williams E medium (Gibco, A1217601) with maintenance supplements and 10% fetal bovine serum. After 24 hours of incubation at 37°C, cells were harvested and expression was quantified via Nano-Glo HiBiT lysis detection system (Promega, N3030) according to the manufacturer's instructions. Raw luminescence was normalized to a standard curve using HiBiT control protein (Promega, N3010). Protein expression of the different Cas9 variants shown in Table 13 and Figure 8 was normalized to the expression of SpyCas9 measured in the corresponding hepatocytes delivered with SpyCas9 mRNA only. Consistent with the data shown in Table 13, protein expression from these same constructs was higher for the NmeCas9 constructs than for the SpyCas9 constructs, as detected by Western blot with anti-HiBiT antibody from PHH cell extracts or measured by HiBiT detection in PMH, PCH, PHH, and PRH cells. [Table 14]
[0424] Example 5.2: Protein Expression in T Cells To quantify the expression of each mRNA construct, protein expression levels were measured following LNP delivery of mRNA encoding either SpyCas9 or Nme2Cas9 into T cells.
[0425] Healthy human donor apheresis was obtained commercially (Hemacare). T cells from two donors (W106 and W864) were isolated by negative selection using the EasySep Human T Cell Isolation Kit (Stem Cell Technology, Cat. 17951) on a MultiMACS Cell24 Separator Plus instrument according to the manufacturer's instructions. Isolated T cells were cryopreserved in CS10 freezing medium (Cryostor, Cat. No. 07930) for future use.
[0426] Upon thawing, T cells were cultured in complete T cell growth medium consisting of CTS OpTmizer base medium (CTS OpTmizer medium (Gibco, A1048501) containing 1x GlutaMAX, 10 mM HEPES buffer, 1% penicillin / streptomycin) supplemented with cytokines (200 IU / ml IL2, 5 ng / ml IL7 and 5 ng / ml IL15) and 2.5% human serum (Gemini, 100-512). After overnight incubation at 37°C, T cells at a density of 1e6 / mL were activated with T cell TransAct reagent (1:100 dilution, Miltenyi) and incubated for 48 hours in a tissue culture incubator.
[0427] Activated T cells were treated with LNPs delivering mRNA encoding Nme2-mRNA or Spy mRNA with HiBiT tag. LNPs were prepared generally as in Example 1. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of about 6. LNPs encapsulating Nme2Cas9 mRNA used lipid A, cholesterol, DSPC, and PEG2k-DMG in a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs encapsulating SpyCas9 mRNA used lipid A, cholesterol, DSPC, and PEG2k-DMG in a molar ratio of 50% lipid A, 38.5% cholesterol, 10% DSPC, and 1.5% PEG2k-DMG.
[0428] Just prior to LNP treatment of T cells, LNPs were incubated with 10 ug / mL ApoE3 (Peprotech, Cat. No. 200-02), 5 ng / ml IL7 (Peprotech, Cat. No. 200-07), and 5 ng / ml IL15 (Peprotech, Cat. No. 200-15) in complete T cell medium supplemented with cytokines (200 IU / ml IL2 (Peprotech, Cat. No. 200-02), 5 ng / ml IL7 (Peprotech, Cat. No. 200-07), and 5 ng / ml IL15 (Peprotech, Cat. No. 200-15) and 2.5% human serum (Gemini, 100-512). The cells were pre-incubated for 5 minutes at 37°C with a LNP concentration of 13.33ug / ml total RNA containing the ApoE ELISA kit (catalog no. 350-02). After incubation, the LNPs were then mixed 1:1 by volume with T cells in complete T cell media containing the cytokines used for ApoE incubation. T cells were harvested for protein expression analysis 24, 48 and 72 hours after LNP treatment. T cells were lysed by Nano-Glo® HiBiT Lytic Assay (Promega) and Cas9 protein levels were quantified by Nano-Glo® Nano-Glo HiBiT Extracellular Detection System (Promega, Catalog no. N2420) according to the manufacturer's instructions. Luminescence was measured using a Biotek Neo2 plate reader. A linear regression was plotted on GraphPad using protein counts and luminescence readouts from standard controls, forcing the line through X=0, Y=0. To calculate the number of proteins per lysate, the formula Y=ax+0 was used.
[0429] Samples were normalized to the average of SpyCas9 at 0.83ug / ml LNP dose. Tables 14A-14B and Figures 9A-9F show the relative Cas9 protein expression in activated cells for mRNA at 24, 48, and 72 hours after LNP treatment in donor 1 or donor 2. Cas9 was expressed in a dose-dependent manner in activated T cells. Protein expression was higher from Nme2Cas9 samples compared to SpyCas9 samples in activated T cells. [Table 15] [Table 16]
[0430] Example 6. In vivo editing in mouse liver using lipid nanoparticles (LNPs) The LNPs used in all in vivo studies were formulated as described in Example 1. Deviations from the protocol are noted in each example. The transport and storage solution (TSS) used for LNP preparation was administered in the experiment as a vehicle-only negative control.
[0431] In vivo editing in mouse models Selected guide designs were tested for editing efficiency in vivo. In each study involving mice, CD-1 female mice ranging from 6 to 10 weeks of age were used. Animals were weighed prior to administration. LNPs were formulated generally as described in Example 1. LNPs contained a molar ratio of 50% ionizable lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. Lipid-nucleic acid assemblies were formulated at a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6.
[0432] LNPs were administered via a lateral tail vein in a volume of 0.2 mL per animal (approximately 10 mL per kilogram of body weight). Body weight was measured 24 hours after administration. Approximately 6-7 days after LNP delivery, animals were euthanized by exsanguination under isoflurane anesthesia after administration. Blood was collected into serum separator tubes by cardiac puncture. For studies using in vivo editing, liver tissue was collected from the left middle lobe from each animal for DNA extraction and analysis.
[0433] For in vivo studies, genomic DNA was extracted from tissues using a bead-based extraction kit, e.g., Zymo Quick-DNA96 kit (Zymo Research, Catalog No. D3010), following the manufacturer's protocol. NGS analysis was performed as described in Example 1.
[0434] Transthyretin (TTR) ELISA Assay for Use in Animal Studies Blood was collected and serum was isolated as described above. Total TTR serum levels were determined using the Mouse Prealbumin (Transthyretin) ELISA kit (Aviva Systems Biology, Cat. No. OKIA00111). Kit reagents and standards were prepared according to the manufacturer's protocol. Mouse serum was diluted to a final dilution of 10,000 times with 1x assay diluent. Both standard curve dilutions (100 μL each) and diluted serum samples were added to each well of the ELISA plate pre-coated with capture antibody. The plate was incubated at room temperature for 30 minutes and then washed. Enzyme-antibody conjugate (100 μL per well) was added for a 20 minute incubation. Unbound antibody conjugate was removed and the plate was washed again before adding color substrate solution. The plate was incubated for 10 minutes before adding 100 μL of stop solution, e.g., sulfuric acid (approximately 0.3 M). The plate was read on a Clariostar at 450 nm absorbance. Serum TTR levels were calculated by SoftMax Pro software version 6.4.2 or Mars software version 3.31 using a 4-parameter logistic curve fit from the standard curve. Final serum values were adjusted for assay dilution. Unless otherwise specified, percent protein knockdown (%KD) values were determined relative to controls, typically animals sham-treated with vehicle (TSS). TSS percent was calculated by dividing each sample TTR value by the mean value of the TSS group, then adjusted to a percentage value.
[0435] Example 6.1 In vivo editing using co-formulated LNPs The editing efficiency of the modified sgRNAs tested in Example 4.2 was further evaluated in a mouse model. Guide RNA designs with identical guide sequences targeting mouse PCSK9 but with conserved regions of different lengths were tested by preparing LNPs as described in Example 1. LNPs were prepared using ionized lipid A, cholesterol, DSPC, and PEG2k-DMG in a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of about 6. As shown in Table 15, gRNA targeting the PCSK9 gene and mRNA C were co-formulated in LNPs at a gRNA to mRNA weight ratio of 1:2. LNPs were administered to female CD-1 mice (n=5) at a dose of 1 mg / kg total RNA as described above. Mice were euthanized 7 days after administration. The editing efficiencies of LNPs containing the indicated sgRNAs are shown in Table 15 and illustrated in FIG. 10. [Table 17]
[0436] Example 6.2. In vivo editing using pgRNA and mRNA LNPs The editing efficiency of the modified pgRNA was evaluated in vivo. In addition to the guide modifications identified in the previous study in Example 6.1, four nucleotides in each of the repeat / anti-repeat regions, hairpin 1, and hairpin 2 loops were replaced with a spacer-18 PEG linker.
[0437] Using a single RNA species as cargo, LNPs were prepared generally as described in Example 1. LNPs contained ionizable lipid A, cholesterol, DSPC, and PEG2k-DMG in a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6.
[0438] LNPs containing gRNAs targeting the TTR gene shown in Table 16 were administered to female CD-1 mice (n=5) at a dose of 0.1 mg / kg or 0.3 mg / kg total RNA as described above. LNPs containing mRNA (mRNA M SEQ ID NO:23) and LNPs containing pgRNA (G021846 or G021844), respectively, were delivered simultaneously at a ratio of 1:2 by RNA weight. Mice were euthanized 7 days after administration.
[0439] The editing efficiency, serum TTR knockdown, and percent TSS of LNPs containing the indicated pgRNAs are shown in Table 16 and illustrated in Figures 11A-11C, respectively. [Table 18]
[0440] The pgRNA (G021844) from the above study was evaluated in mice with alternative mRNA at various dose levels. Using a single RNA species as cargo, LNPs were prepared generally as described in Example 1. LNPs containing pgRNA (G21844) or mRNA (mRNA P or mRNA M) were formulated as described in Example 1. The LNPs used were prepared using ionizable lipid A, cholesterol, DSPC, and PEG2k-DMG in a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of about 6. Both G000502 and G021844 target exon 3 of the mouse TTR gene. LNPs containing pgRNA and LNPs containing mRNA were administered simultaneously based on total RNA weight, with a guide:mRNA ratio of 2:1 by RNA weight, respectively. Additional LNPs were co-formulated with G000502 and SpyCas9 mRNA at a weight ratio of 1:2, respectively, to provide the preferred SpyCas9 guide:mRNA ratio.
[0441] The LNPs shown in Table 17 were administered to female CD-1 mice (n=4) at a dose of 0.1 mg / kg or 0.03 mg / kg of total RNA. The editing efficiency of LNPs containing the indicated gRNAs is shown in Table 17 and illustrated in Figures 11D and 11E. [Table 19]
[0442] Example 6.3. In vivo editing using sgRNA and mRNA LNPs LNPs were prepared generally as described in Example 1 using a single RNA species as cargo. The LNPs used were prepared using lipid A, cholesterol, DSPC, and PEG2k-DMG in a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6. sgRNAs were designed to target the pcsk9 gene (G020361) or the Rosa26 gene (G020848).
[0443] LNPs containing sgRNA or mRNA were administered to female CD-1 mice (n=5) at a dose of 1 mg / kg of total RNA. The tested mRNAs (mRNA C, mRNA J, mRNA Q, mRNA N) were designed with various numbers and sequences of NLSs. LNPs were administered simultaneously based on the total weight of the RNA cargo at a 1:1 ratio of gRNA:mRNA by RNA weight. The average editing percentages are shown in Table 18 and illustrated in FIG. 12. [Table 20]
[0444] Example 7. In vivo editing with NmeCas9 and either sgRNA or pgRNA The editing efficiency of the modified pgRNAs tested with Nme2Cas9 was tested in a mouse model. All Nme sgRNAs tested contained the same 24-nt guide sequence targeting mTTR.
[0445] LNPs were prepared generally as described in Example 1 using a single RNA species as cargo. The LNPs used were prepared using lipid A, cholesterol, DSPC, and PEG2k-DMG in a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of about 6. LNPs were mixed with gRNA and mRNA cargo in a weight ratio of 2:1. Doses are calculated based on the total RNA mass of gRNA and mRNA. The transport and storage solution (TSS) used for LNP preparation was administered in the experiment as a vehicle-only negative control.
[0446] In each study involving mice, 6-10 week old CD-1 female mice were used (n=5 per group except for TSS control n=4). Formulations were administered intravenously via tail vein injection according to the doses listed in Table 19. Animals were observed periodically for adverse effects for at least 24 hours after dosing. Six days after treatment, animals were euthanized by cardiac puncture under isoflurane anesthesia and liver tissue was collected for downstream analysis. Liver punches weighing 5mg to 15mg were collected for isolation of genomic DNA and total RNA. Genomic DNA samples were analyzed using NGS sequencing as described in Example 1. Editing efficiencies of LNPs containing the indicated mRNAs and gRNAs are shown in Table 19 and illustrated in Figure 13. [Table 21]
[0447] Example 8. In vivo base editing with Nme2Cas9 gRNA The editing efficiency of modified gRNA with different mRNA was tested with Nme base editor constructs in mouse model. This experiment was performed in parallel with Example 7, and the same control sample was used. Using a single RNA species as cargo, LNPs were generally prepared as described in Example 1. The LNPs used were prepared using lipid A, cholesterol, DSPC, and PEG2k-DMG in a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. The LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of about 6. The LNPs used were formulated as described in Example 1, except that each component, guide RNA, or mRNA, was formulated into LNPs separately, and the LNPs were mixed before administration as described in Table 20. For Nme2Cas9 and Nme2Cas9 base editor samples, LNPs were mixed with gRNA and editor mRNA cargo in a weight ratio of 2:1. For SpyCas9 base editor samples, LNPs were mixed with gRNA and editor mRNA cargo at a weight ratio of 1:2. As shown in Table 20 and Figure 14, the dose is calculated based on the total RNA weight of gRNA and editor mRNA. Base editor samples were treated with an additional 0.03mpk of UGI mRNA. Transport and storage solution (TSS) used for LNP preparation was administered in the experiment as a vehicle-only negative control.
[0448] In each study involving mice, 6-10 week old CD-1 female mice were used (n=5 per group except for TSS control n=4). Formulations were administered intravenously via tail vein injection according to the doses listed in Table 20. Animals were observed regularly for adverse effects for at least 24 hours after dosing. Six days after treatment, animals were euthanized by cardiac puncture under isoflurane anesthesia and liver tissue was collected for downstream analysis. Liver punches weighing 5mg-15mg were collected for isolation of genomic DNA and total RNA. Genomic DNA was extracted using a DNA isolation kit (ZymoResearch, D3010) and samples were analyzed using NGS sequencing as described in Example 1. Editing efficiencies of LNPs containing the indicated gRNAs are shown in Table 20 and illustrated in Figure 14. [Table 22]
[0449] Example 9. Guided screening with Nme1Cas9 and Nme3Cas9 mRNA in T cells The editing efficiency of one modified gRNA scaffold was tested in T cells harboring Nme1Cas9 or Nme3Cas9 mRNA using guides with nine different target sequences within the TRAC locus.
[0450] Healthy human donor apheresis was obtained commercially (Hemacare, Donor 3786), cells were washed and resuspended in CliniMACS® PBS / EDTA buffer (Miltenyi Biotec catalog 130-070-525) and processed in a MultiMACS™ Cell 24 Separator Plus device (Miltenyi Biotec). T cells were isolated via positive selection using the Straight from Leukopak® CD4 / CD8 MicroBead kit, human (Miltenyi Biotec catalog 130-122-352). T cells were aliquoted and cryopreserved for use in a Cryostor® CS10 (StemCell Technologies, catalog 07930). Upon thawing, cells were cultured at 1.0 x 10^ in T cell proliferation medium (TCGM) consisting of CTS OpTmizer T cell proliferation SFM and T cell proliferation supplement (Thermo Fisher, Cat. A1048501), 5% human AB serum (GeminiBio, Cat. 100-512), 1x penicillin-streptomycin, 1x Glutamax, 10 mM HEPES, 200 U / mL recombinant human interleukin-2 (Peprotech, Cat. 200-02), 5 ng / mL recombinant human interleukin-7 (Peprotech, Cat. 200-07), and 5 ng / mL recombinant human interleukin-15 (Peprotech, Cat. 200-15). 6T cells were seeded at a density of 10000 cells / mL. T cells were left in this medium for 24 hours, at which point they were activated with T cell TransAct™ Human Reagent (Miltenyi, catalog 130-111-160) added at a volume ratio of 1:100.
[0451] For Nme1Cas9 guide screening, a solution containing mRNA encoding Nme1Cas9 (mRNA AB) was prepared in P3 buffer. Guide RNAs targeting various sites within the TRAC locus were denatured at 95°C for 2 min and incubated at room temperature for 5 min. 48 h after activation, T cells were harvested, centrifuged, and 12.5 × 10^ cells were electroporated in P3 electroporation buffer (Lonza). 6 Resuspended at a concentration of 1 x 10^ cells / mL for each well to be electroporated. 5 The cells were mixed with 600ng of Nme1Cas9 mRNA and 5μM of gRNA in a final volume of 20μL of P3 electroporation buffer. The mixture was transferred in duplicate to a 96-well Nucleofector™ plate and electroporated using the manufacturer's pulse code. The electroporated T cells were immediately placed in CTS OpTmizer T cell growth medium without cytokines for 15 minutes and then transferred to a new flat-bottom 96-well plate containing additional CTS OpTmizer T cell growth medium supplemented with cytokines. The resulting plate was incubated at 37°C for 3 days. Three days after electroporation, the cells were split 1:2 into two U-bottom plates.
[0452] Seven days after electroporation, seeded T cells were assayed by flow cytometry to determine surface expression of T cell receptors. Briefly, T cells were incubated with antibodies against CD3 (BioLegend, Cat. No. 317336), CD4 (BioLegend, Cat. No. 317434), CD8 (BioLegend, Cat. No. 301046), and Viakrome (Beckman Coulter, Cat. No. C36628). Cells were then washed, resuspended in cell staining buffer, and processed on a Cytoflex flow cytometer (Beckman Coulter). Flow cytometry data was analyzed using the FlowJo software package. T cells were gated based on size, shape, viability, and expression of CD8 and CD3. Samples were run in duplicate.
[0453] CD3 is a cell surface component of the T cell receptor complex, and its presence on the cell surface is used as a surrogate marker for TRAC protein expression. The CD3 negative cell populations for each of the indicated gRNAs, and the corresponding standard deviations (SD), are shown in Table 21 and illustrated in FIG. [Table 23]
[0454] For screening of guides with Nme3Cas9 mRNA, T cells were prepared as described in this example. Solutions containing mRNA encoding Nme3Cas9 (mRNA Z), as well as controls for Nme1Cas9 (mRNA AB) and Nme2Cas9 (mRNA O) were prepared in P3 buffer. Electroporation of NmeCas9 (e.g., Nme1Cas9, Nme2Cas9, or Nme3Cas9) gRNA and mRNA was performed as described above. Samples were electroporated in triplicate. Three days after electroporation, cells were assayed via flow cytometry as described above.
[0455] The CD3 negative cell populations for each of the indicated gRNAs, and the corresponding standard deviations (SD), are shown in Table 22 and illustrated in FIG. 17. [Table 24]
[0456] Example 10. Expression of codon-optimized NmeCas9 mRNA To quantify the expression of each mRNA construct, T cells were electroporated with mRNA encoding Nme1Cas9, Nme2Cas9, or Nme3Cas9, and then protein expression levels were measured. All of the NmeCas9 mRNA constructs have the same general structure, with consecutive SV40 and nucleoplasmin nuclear localization signal coding sequences N-terminal to the NmeCas9 open reading frame. The constructs contain a coding sequence for a HiBiT tag C-terminal to the NmeCas9 open reading frame. The components are linked by linkers, and the specific sequences are provided herein.
[0457] Healthy human donor apheresis was obtained commercially (Hemacare, Donor 3786), cells were washed and resuspended in CliniMACS® PBS / EDTA buffer (Miltenyi Biotec catalog 130-070-525) and processed in a MultiMACS™ Cell 24 Separator Plus device (Miltenyi Biotec). T cells were isolated via positive selection using the Straight from Leukopak® CD4 / CD8 MicroBead Kit, human (Miltenyi Biotec catalog 130-122-352). T cells were aliquoted and resuspended in Cryostor® CS10 (StemCell Technologies, catalog 07930). Upon thawing, cells were cultured at 1.0 x 10^ in T cell proliferation medium (TCGM) consisting of CTS OpTmizer T cell proliferation SFM and T cell proliferation supplement (ThermoFisher, Cat. A1048501), 5% human AB serum (GeminiBio, Cat. 100-512), 1x penicillin-streptomycin, 1x Glutamax, 10 mM HEPES, 200 U / mL recombinant human interleukin-2 (Peprotech, Cat. 200-02), 5 ng / mL recombinant human interleukin-7 (Peprotech, Cat. 200-07), and 5 ng / mL recombinant human interleukin-15 (Peprotech, Cat. 200-15). 6 T cells were seeded at a density of 10000 cells / mL. T cells were left in this medium for 24 hours, at which point they were activated with T cell TransAct™ Human Reagent (Miltenyi, catalog 130-111-160) added at a volume ratio of 1:100.
[0458] A solution containing mRNA encoding NmeCas9 was prepared in P3 buffer. Guide RNA targeting the TRAC locus was removed from storage, denatured at 95°C for 2 min, and incubated at room temperature for 5 min. 48 h after activation, T cells were harvested, centrifuged, and diluted to 12.5 × 10^ in P3 electroporation buffer (Lonza). 6Each well to be electroporated was resuspended at a concentration of 1 x 10^ cells / mL in a final volume of 20 μL of P3 electroporation buffer. 5 The T cells contained NmeCas9 mRNA as specified in Table 23, and 1 μM of gRNA as specified in Table 23 (G028853 for Nme1Cas9, G021469 for Nme2Cas9, G028848 for Nme3Cas9). NmeCas9 mRNA was tested using a 3-fold 5-point serial dilution starting with 600 ng of mRNA. The appropriate gRNA and mRNA mixtures were transferred in triplicate to 96-well Nucleofector™ plates and electroporated using the manufacturer's pulse code. Electroporated T cells were immediately placed in CTS OpTmizer T cell growth medium without cytokines for 15 minutes before being transferred to a new flat-bottom 96-well plate containing additional CTS OpTmizer T cell growth medium supplemented with cytokines. The resulting plates were incubated at 37° C. for 24 hours prior to HiBiT luminescence assay or 96 hours prior to flow cytometry.
[0459] T cells were harvested for protein expression analysis 24 hours after electroporation. T cells were lysed by Nano-Glo® HiBiT Lytic Assay (Promega). Luminescence was measured using a Biotek Neo2 plate reader. Table 23 and Figure 18 show Cas9 protein expression in activated cells as relative luminescence units (RLU) and the corresponding standard deviation (SD). [Table 25]
[0460] Four days after editing, T cells were assayed by flow cytometry to determine surface protein expression. Briefly, T cells were incubated with a mixture of antibodies diluted in cell staining buffer (BioLegend, Cat. No. 420201) for 30 minutes at 4°C. Antibodies against CD3 (BioLegend, Cat. No. 317336), CD4 (BioLegend, Cat. No. 317434), CD8 (BioLegend, Cat. No. 301046), and Viakrome (Beckman Coulter, Cat. No. C36628) were diluted 1:100. Cells were then washed, resuspended in 100 μL of cell staining buffer, and processed on a Cytoflex flow cytometer (Beckman Coulter). Flow cytometry data were analyzed using the FlowJo software package. T cells were gated based on size, shape, viability, CD8, and CD3. Samples were run in triplicate. The CD3 negative cell populations for each of the indicated gRNAs, and the corresponding standard deviations (SD), are shown in Table 24 and illustrated in FIG. [Table 26]
[0461] Example 11. In vitro editing with selected guides in primary cynomolgus monkey hepatocytes (PCH) Three NmeCas9 sgRNAs (G024739, G024741, and G024743) were selected for evaluation in a dose-response assay. The tested NmeCas9 sgRNAs targeting the cynomolgus TTR gene contain a 24-nucleotide guide sequence.
[0462] Unmodified and modified versions of the guide are provided in Table 25. [Table 27]
[0463] gRNA and Cas9 mRNA were lipofected into primary cynomolgus monkey hepatocytes (PCH) as described below. PCH (In Vitro ADMET Laboratories 10136011) were prepared as described in Example 1. PCH were seeded at a density of 40,000 cells / well. LNP formulations were prepared as described in Example 1. LNPs were prepared with a lipid composition at a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of about 6 and gRNA as shown in Table 25. PCH in 100 μL of media were treated with an 8-point, 4-fold dilution series of LNPs containing sgRNA, starting at 70 ng, and a fixed 30 ng dose of LNPs encapsulating mRNA O by mRNA weight. The sgRNA concentration in each well is shown in Table 26. 72 hours after treatment, cells were lysed and NGS analysis was performed as described in Example 1. The dose response of editing efficiency to guide concentration was measured in triplicate samples. Table 26 and Figure 20 show the average editing percentage and standard deviation (SD) at each guide concentration. [Table 28]
[0464] Example 12. In vitro editing of LNPs using mRNA dilution series in PCH Modified sgRNAs with the same target site in the cynomolgus TTR gene were assayed to evaluate the editing efficiency in PCH with different mRNAs (mRNA O, mRNA AA) and formulation ratios. Unless otherwise noted, PCH (In Vitro ADMET Laboratories, 10136011) were prepared, treated, and analyzed as described in this example as in Example 1. PCH were used and seeded at a density of 50,000 cells / well. LNP formulations were prepared as described in Example 1. LNPs were prepared with a lipid composition having a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of about 6 and gRNA as shown in Table 27. PCH in 100 μL of media were treated with eight-point, three-fold serial dilutions of mixed (separately formulated) or co-formulated LNPs with various ratios of gRNA:mRNA. The highest dose was 3 ng / μL total RNA by weight, and the gRNA:mRNA ratios for the dilution series were as shown in Table 27. Samples were run in triplicate. The mean percent editing, standard deviation (SD), and calculated EC50 are shown in Table 27 and Figure 21. [Table 29]
[0465] Example 13. In vivo editing with NmeCas9 gRNA The editing efficiency of the modified gRNAs was tested with the Nme2Cas9 construct in mice. All Nme sgRNAs tested contained the same 24-nt guide sequence targeting the mouse TTR gene (mTTR).
[0466] LNPs were prepared generally as described in Example 1, with a cargo of 1:2 by weight of gRNA to mRNA O. The LNPs used were prepared with a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6. Dose was calculated based on the total RNA weight of gRNA and mRNA. The transport and storage solution (TSS) used for LNP preparation was administered in the experiment as a vehicle-only negative control.
[0467] In each study involving mice, CD-1 female mice approximately 6-8 weeks of age were used. Animals were fed a regular diet with standard maintenance. Animals were weighed before dose administration. TSS and LNP formulations were administered intravenously via tail vein injection at a dose of 0.03 mpk. Animals were observed periodically for adverse effects for at least 24 hours after dosing. After 14 days of treatment, animals were euthanized by cardiac exsanguination under isoflurane anesthesia, and blood for serum preparation and liver tissue were collected for downstream analysis.
[0468] Serum TTR levels shown in Table 28 and FIG. 22 were generated for all experimental groups using a Serum TTR ELISA-Prealbumin ELISA (Aviva Systems, Cat. No. OKIA00111) according to the manufacturer's protocol and compared to a negative control (TSS). [Table 30]
[0469] Liver biopsy punches weighing 5-15 mg were collected for isolation of genomic DNA. Genomic DNA was extracted using a DNA isolation kit (ZymoResearch, D3012) and samples were analyzed using NGS sequencing as described in Example 1. The editing efficiencies of LNPs containing the indicated gRNAs are shown in Table 29 and illustrated in Figure 23. [Table 31]
[0470] Example 14. Dose-response curve of NmeCas9 gRNA in PMH with Nme2Cas9 The editing efficiency of the modified gRNAs was tested with the Nme2Cas9 construct in primary mouse hepatocytes (PMH). All Nme sgRNAs tested contained the same 24-nt guide sequence targeting the mouse TTR gene (mTTR).
[0471] PMH (Gibco, Lot No. MC931) were thawed and resuspended in hepatocyte thawing medium followed by centrifugation. The supernatant was discarded and the pelleted cells were resuspended in hepatocyte seeding medium (Williams E medium (Gibco, Catalog No. A12176-01)) with seeding supplement dexamethasone + cocktail supplement (Gibco, Catalog No. A15563, Lot No. 2459010) and FBS content (Gibco, Catalog No. A13450, Lot No. 2486425). Cells were counted and seeded on Biocoat Collagen I coated 96-well plates (Corning, Ref356407, Lot08722018) at a concentration of 15,000 cells / well. Seeded cells were incubated at 37°C and 5% CO 2 The cells were left in a tissue culture incubator at room temperature for 4-6 hours to allow for attachment. After incubation, the cells were examined for monolayer formation and washed once with Hepatocyte Maintenance Medium (Williams E Medium) with Seeding Medium Supplement (Gibco, Cat. No. A15564, Lot No. 2459014).
[0472] LNPs were prepared generally as described in Example 1, with a cargo of 1:2 by weight of gRNA to mRNA O. The LNPs used were prepared with a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. The LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of 6. Each LNP was applied to cells using 8-point 3-fold serial dilutions starting with 450 ng total cargo per 100 μl well at the highest dose (300 ng mRNA O and 46.5 nM gRNA (approximately 150 ng gRNA)) as shown in Table 30. After treatment with LNPs, cells were incubated at 37° C. for 24 hours in Williams E medium with seeding medium supplement (Gibco, Cat. No. A15564, Lot No. 2459014) and 3% fetal bovine serum. After 72 hours, cells were harvested and analyzed by NGS as described in Example 1.
[0473] The editing efficiencies of LNPs containing the indicated gRNAs, and their corresponding EC50s, are shown in Table 30 and illustrated in Figure 24. [Table 32]
[0474] Example 15. Dose-response curve of NmeCas9 gRNA in PMH with Nme2Cas9 The editing efficiency of the modified gRNAs was tested with the Nme2Cas9 construct in primary mouse hepatocytes (PMH). All Nme sgRNAs tested contained the same 24-nt guide sequence targeting the mouse TTR gene (mTTR).
[0475] PMH (Gibco, Lot No. MC931) were thawed and resuspended in Hepatocyte Thawing Medium followed by centrifugation. The supernatant was discarded and the pelleted cells were resuspended in Hepatocyte Seeding Medium (Williams E Medium (Gibco, Catalog No. A12176-01)) and Seeding Supplement Dexamethasone + Cocktail Supplement (Gibco, Catalog No. A15563, Lot No. 2459010) and FBS content (Gibco, Catalog No. A13450, Lot No. 2486425). Cells were counted and seeded on Biocoat Collagen I coated 96-well plates (Corning, Ref356407, Lot08722018) at a concentration of 15,000 cells / well. Seeded cells were incubated at 37°C and 5% CO 2 The cells were left in a tissue culture incubator at room temperature for 4-6 hours to allow for attachment. After incubation, the cells were examined for monolayer formation and washed once with Hepatocyte Maintenance Medium (Williams E Medium) with Seeding Medium Supplement (Gibco, Cat. No. A15564, Lot No. 2459014).
[0476] LNPs were prepared generally as described in Example 1, with a cargo of 1:2 by weight of gRNA to mRNA O. The LNPs used were prepared with a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. The LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of 6. Each LNP was applied to cells using 8-point 3-fold serial dilutions starting with 450 ng total cargo per 100 μl well at the highest dose (300 ng mRNA O and 46.5 nM gRNA (i.e., 150 ng gRNA)) as shown in Table 31. After treatment with LNPs, cells were incubated at 37° C. for 24 hours in Williams E medium with seeding medium supplement (Gibco, Cat. No. A15564, Lot No. 2459014) and 3% fetal bovine serum. Samples were run in triplicate. After 72 hours, cells were harvested and analyzed by NGS as described in Example 1.
[0477] The editing efficiencies of the LNPs containing gRNAs, and their corresponding EC50s, are shown in Table 31 and illustrated in Figure 25. [Table 33-1] [Table 33-2]
[0478] Example 16. In vitro editing in primary mouse hepatocytes (PMH) by dilution curve A. Example 16.1. Modified sgRNA Evaluation Using Dilution Series Modified sgRNAs with various scaffold structures targeting previously published sites in all mouse pcsk9 genes (see WO2019094791) were designed as shown in Tables 1-2 and tested for editing efficiency using primary mouse hepatocytes (PMH). Cells were prepared as described in Example 1 using PMH cells (In Vitro ADMET Laboratories) and seeded at a density of 20,000 cells / well. Cells were transfected using MessengerMax (Invitrogen) according to the manufacturer's protocol with 1 ng / μl Nme2 Cas9 mRNA (mRNA U) and sgRNA at concentrations as shown in Table 32. Duplicate samples were included in the assay. Cells were harvested 72 hours after transfection and analyzed by NGS as described in Example 1. The average editing percentage with standard deviation is shown in Table 32 and Figure 26. [Table 34]
[0479] B. Example 16.2. Evaluation of mRNA polyA tail modifications and cargo ratios sgRNAs targeting mouse psck9 gene were selected from Table 32 to evaluate guided editing efficiency resulting from specific combinations of polyA tail modifications and sgRNA:mRNA ratios. Unless otherwise noted, PMH cells used were prepared, treated, and analyzed as described in Example 1. PMH (Gibco) were seeded at a density of 15,000 cells / well.
[0480] LNPs were prepared generally as described in Example 1. LNPs were prepared with a lipid composition of 50 / 9 / 38 / 3, expressed as a molar ratio of ionized lipid A / cholesterol / DSPC / PEG, respectively. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6. As shown in Table 33, LNPs encapsulated gRNA G017566 or one of three mRNAs that encode the same Nme2Cas9 open reading frame (ORF) but with different encoded polyA tails. Preliminary experiments keeping the application of sgRNA constant and varying the amount of mRNA applied showed that a weight ratio of 1:1 sgRNA:mRNA resulted in the highest editing percentage. In this example, increasing doses of mRNA LNPs and gRNA LNPs were applied to cells in 100ul medium while maintaining a weight ratio of 1:1 sgRNA:mRNA, as described in Table 33. Table 33 and Figure 27 show the mean edit percentages and standard deviations (SD). [Table 35]
[0481] Example 17. Dose-response curve of NmeCas9 gRNA in PMH with Nme2Cas9 The editing efficiency of the modified gRNAs was tested with the Nme2Cas9 construct in primary mouse hepatocytes (PMH). All Nme sgRNAs tested contained the same 24-nt guide sequence targeting the mouse TTR gene (mTTR).
[0482] PMH (Gibco, lot MC931) were thawed and resuspended in stem cell thawing medium containing seeding supplement (Williams E medium (Gibco, catalog no. A12176-01)) and dexamethasone + cocktail supplement (Gibco, catalog no. A15563, lot no. 2019842), as well as seeding supplement and FBS content (Gibco, catalog no. A13450, lot no. 1970698), followed by centrifugation. The supernatant was discarded and the pelleted cells were resuspended in Hepatocyte Seeding Medium and Supplement Pack (Invitrogen, catalog no. A1217601 and Gibco, catalog no. CM3000). Cells were counted and seeded at a density of 15,000 cells / well on Biocoat Collagen I-coated 96-well plates (Thermo Fisher, catalog no. 877272). Seeded cells were incubated at 37°C and 5% CO 2 The cells were allowed to adhere for 4-6 hours in a tissue culture incubator at ambient temperature. After incubation, the cells were examined for monolayer formation and washed once with Hepatocyte Maintenance Medium (Invitrogen, Cat. No. A1217601 and Gibco, Cat. No. CM4000).
[0483] LNPs were prepared generally as described in Example 1, with a cargo of 1:2 by weight of gRNA to mRNA O. The LNPs used were prepared with a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. The LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of 6. Each LNP was applied to cells using an 8-point 4-fold serial dilution starting with 300 ng total RNA per 100 μl well (approximately 32.25 nM gRNA concentration per well), as shown in Table 32. After treatment with LNPs, cells were incubated at 37° C. for 24 hours in Williams E medium (Gibco, A1217601) with maintenance supplements and 3% fetal bovine serum. Samples were run in triplicate. After 72 hours, cells were harvested and analyzed by NGS as described in Example 1.
[0484] The editing efficiencies of LNPs containing the indicated gRNAs, and their corresponding EC50s, are shown in Table 34 and illustrated in Figure 28. [Table 36]
[0485] Example 18. In vivo editing with NmeCas9 gRNA The editing efficiency of the modified gRNAs was tested with the Nme2Cas9 construct in mice. All Nme sgRNAs tested contained the same 24-nt guide sequence targeting the mouse TTR gene (mTTR).
[0486] LNPs were prepared generally as described in Example 1, with a cargo of 1:2 by weight of gRNA to mRNA O. The LNPs used were prepared with a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6. Dose was calculated based on the total RNA weight of gRNA and mRNA. The transport and storage solution (TSS) used for LNP preparation was administered in the experiment as a vehicle-only negative control.
[0487] In each study involving mice, CD-1 female mice approximately 6-8 weeks of age were used (n=5 for all groups). Animals were fed regular chow with standard maintenance. Animals were weighed before dose administration. TSS and LNP formulations were administered intravenously via tail vein injection at a dose of 0.03 mpk. Animals were observed periodically for adverse effects for at least 24 hours after dosing. Seven days after treatment, animals were euthanized by cardiac exsanguination under isoflurane anesthesia, and blood for serum preparation and liver tissue were collected for downstream analysis.
[0488] Serum TTR levels shown in Table 35 and Figure 29 were generated using a Serum TTR ELISA-Prealbumin ELISA (Aviva Systems, Catalog No. OKIA00111) according to the manufacturer's protocol. Serum TTR levels are significantly lower in all experimental groups compared to the negative control (TSS). [Table 37]
[0489] Liver biopsy punches weighing 5-15 mg were collected for isolation of genomic DNA. Genomic DNA was extracted using a DNA isolation kit (ZymoResearch, D3012) and samples were analyzed using NGS sequencing (n=5 for all groups) as described in Example 1. The editing efficiencies of LNPs containing the indicated gRNAs are shown in Table 36 and illustrated in Figure 30. [Table 38]
[0490] Example 19. In vivo editing with NmeCas9 gRNA The editing efficiency of the modified gRNAs was tested with Nme2Cas9 mRNA in mice. All Nme sgRNAs tested contained the same 24-nt guide sequence targeting mTTR.
[0491] LNPs were prepared generally as described in Example 1, with a cargo of 1:2 by weight of gRNA to mRNA O. The LNPs used were prepared with a molar ratio of 50% lipid A, 38% cholesterol, 9% DSPC, and 3% PEG2k-DMG. LNPs were formulated with a lipid amine to RNA phosphate (N:P) molar ratio of approximately 6. Dose was calculated based on the total RNA weight of gRNA and mRNA. The transport and storage solution (TSS) used for LNP preparation was administered in the experiment as a vehicle-only negative control.
[0492] In each study involving mice, CD-1 female mice approximately 6 weeks of age were used (n=5 for all groups). Animals were weighed before dose administration for dose calculation and 24 hours after administration for monitoring. TSS and LNP formulations were administered intravenously via tail vein injection at a dose of 0.01mpk or 0.03mpk. Animals were observed periodically for adverse effects for at least 24 hours after dosing. Seven days after treatment, animals were euthanized by cardiac exsanguination under isoflurane anesthesia. Blood was collected by cardiac puncture for serum TTR ELISA and liver tissue was collected for downstream analysis.
[0493] Serum TTR results prepared using Serum TTR ELISA-Prealbumin ELISA (Aviva Systems, Cat. No. OKIA00111) according to the manufacturer's protocol are shown in FIG. 31 and Table 37. [Table 39]
[0494] Liver biopsy punches weighing approximately 5 mg to 15 mg were collected for isolation of genomic DNA and total RNA. Genomic DNA was extracted using a DNA isolation kit (ZymoResearch, D3012) and samples were analyzed using NGS sequencing (n=5 for all groups) as described in Example 1. The editing efficiency of LNPs containing the indicated gRNAs is shown in Table 38 and illustrated in Figure 32. [Table 40]
[0495] Example 20. Additional embodiments The following numbered items provide additional support and explanation for the embodiments herein.
[0496] Item 1 is a polynucleotide comprising an open reading frame (ORF), the ORF comprising a nucleotide sequence encoding a C-terminal Neisseria meningitidis (Nme) Cas9 polypeptide that is at least 90% identical to any one of SEQ ID NOs: 29, 32-41, 224-226, 231-233, 238-240, 245-247, 252-254, 259-261, 266-268, 273-275, 280-282, 287-289, 294-296, 301-303, or 316-321, wherein the Nme Cas9 is Nme2 Cas9, Nme1 Cas9, or Nme3 Cas9, and a nucleotide sequence encoding a first nuclear localization signal (NLS).
[0497] Item 2 is the polynucleotide of item 1, wherein the ORF further comprises a nucleotide sequence encoding a second NLS.
[0498] Item 3 is the polynucleotide according to Item 1, wherein the first and second NLSs are independently selected from SEQ ID NOs: 388 and 410 to 422.
[0499] Item 4 is a polynucleotide according to any one of the preceding items, wherein the polynucleotide further comprises a polyA sequence or a polyadenylation signal sequence.
[0500] Item 5 is the polynucleotide according to item 4, wherein the polyA sequence includes non-adenine nucleotides.
[0501] Item 6 is the polynucleotide according to item 4 or 5, wherein the polyA sequence contains 100 to 400 nucleotides.
[0502] Item 7 is the polynucleotide according to any one of items 4 to 6, wherein the polyA sequence comprises the sequence of SEQ ID NO:409.
[0503] Item 8 is the polynucleotide of any one of the preceding items, wherein the ORF further comprises a nucleotide sequence encoding a linker sequence between the first NLS and the second NLS.
[0504] Item 9 is the polynucleotide of any one of the preceding items, wherein the ORF further comprises a nucleotide sequence encoding a linker spacer sequence between the Nme Cas9 coding sequence and the NLS adjacent to the Nme Cas9 coding sequence.
[0505] Item 10 is the polynucleotide of item 8 or 9, wherein the linker comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 amino acids.
[0506] Item 11 is the polynucleotide according to any one of items 8 to 10, wherein the linker sequence comprises GGG or GGGS, and optionally the GGG or GGGS sequence is at the N-terminus of the spacer sequence.
[0507] Item 12 is the polynucleotide according to any one of Items 8 to 11, wherein the linker sequence includes any one of SEQ ID NOs: 61 to 122.
[0508] Item 13 is the polynucleotide of any one of the preceding items, wherein the ORF further comprises one or more additional heterologous functional domains.
[0509] Item 14 is a polynucleotide according to any one of the preceding items, wherein Nme Cas9 has double-stranded endonuclease activity.
[0510] Item 15 is the polynucleotide according to any one of items 1 to 14, wherein Nme Cas9 has nickase activity.
[0511] Item 16 is the polynucleotide according to any one of items 1 to 14, wherein Nme Cas9 comprises a dCas9 DNA binding domain.
[0512] Item 17 is a polynucleotide according to any one of the preceding items, wherein NmeCas9 comprises at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence of any one of SEQ ID NOs: 1, 4-13, 220, 227, 234, 241, 248, 255, 262, 269, 276, 283, 290, 297, or 310-315.
[0513] Item 18 is the polynucleotide of any one of the preceding items, wherein NmeCas9 comprises the amino acid sequence of SEQ ID NO: 1 and 4 to 13, 220, 227, 234, 241, 248, 255, 262, 269, 276, 283, 290, 297, or 310 to 315.
[0514] Item 19 is the polynucleotide of any one of the preceding items, wherein the sequence encoding NmeCas9 comprises a nucleotide sequence having at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the nucleotide sequence of any one of SEQ ID NOs: 15, 18-27, 29, 32-41, 221-226, 228-233, 235-240, 242-247, 249-254, 256-261, 263-268, 270-275, 277-282, 284-289, 291-296, 298-303, 304-309, or 316-321.
[0515] Item 20 is a polynucleotide according to any one of the preceding items, wherein the sequence encoding NmeCas9 comprises any one of the nucleotide sequences of SEQ ID NOs: 15, 18-27, 29, 32-41, 221-226, 228-233, 235-240, 242-247, 249-254, 256-261, 263-268, 270-275, 277-282, 284-289, 291-296, 298-303, 304-309, or 316-321.
[0516] Item 21 is a polynucleotide comprising an open reading frame (ORF) encoding a polypeptide comprising a cytidine deaminase, which is optionally an APOBEC3A deaminase, and a nucleotide sequence encoding a C-terminal Neisseria meningitidis (Nme) Cas9 nickase polypeptide that is at least 90% identical to any one of SEQ ID NOs: 1, 4-13, 220, 227, 234, 241, 248, 255, 262, 269, 276, 283, 290, or 297, where the Nme Cas9 nickase is an Nme2 Cas9 nickase, an Nme1 Cas9 nickase, or an Nme3 Cas9 nickase, and a nucleotide sequence encoding a first nuclear localization signal (NLS), wherein the polypeptide does not comprise a uracil glycosylase inhibitor (UGI). 【...
Claims
[Claim 1] The invention described in the specification.