Polypeptides, polynucleotides, compositions, and methods for genome editing involving chromatin remodelers
Fusing HMGB1 with CRISPR-associated nucleases enhances genome editing efficiency by improving chromatin accessibility, addressing the challenge of modifying targets within compacted chromatin regions.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- INTELLIA THERAPEUTICS INC
- Filing Date
- 2025-11-03
- Publication Date
- 2026-05-07
AI Technical Summary
Existing genome editing technologies face challenges in accessing and modifying targets within highly structured chromatin regions due to reduced efficiency of CRISPR-associated nucleases like Cas9, particularly when protospacer adjacent motifs (PAMs) are located within nucleosomes.
Fusion of a chromatin remodeler, such as the High Mobility Group Box 1 (HMGB1) protein, with a programmable DNA-binding protein (e.g., CRISPR-associated nuclease) to enhance chromatin accessibility and editing efficiency at compacted chromatin loci.
The fusion of HMGB1 with CRISPR-associated nucleases increases genome editing efficiency by altering DNA-protein interactions, allowing for effective modifications in diverse chromatin environments.
Smart Images

Figure US2025053779_07052026_PF_FP_ABST
Abstract
Description
POLYPEPTIDES, POLYNUCLEOTIDES, COMPOSITIONS, AND METHODS FOR GENOME EDITING INVOLVING CHROMATIN REMODELERSCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to US Provisional Application No. 63 / 716,226, filed November 4, 2024, the content of which is herein incorporated by reference in its entirety.REFERENCE TO ELECTRONIC SEQUENCE LISTING
[0002] This application contains a sequence listing, which has been submitted electronically in XML file format and is hereby incorporated by reference in its entirety. Said XML file, created on November 2, 2025, is named “01155-0073-00PCT.xml” and is 4,963,446 bytes in size.FIELD
[0003] The present disclosure relates to polypeptides, polynucleotides, compositions, methods, and systems for genome editing.BACKGROUND
[0004] The ability to introduce modifications into a genomic sequence of a cell is of interest for gene editing and clinical therapeutic applications. However, in order to enact gene editing, a programmable DNA-binding protein (e.g., a CRISPR-associated nuclease) must be able to access the highly structured chromosomal DNA. The effect of local chromatin accessibility on editing efficiency can potentially create challenges for certain targets. Target sequences located in chromatin-dense regions, in particular, may be less susceptible to genome editing. For example, the CRISPR-associated nuclease Cas9 exhibits reduced editing of targets whose protospacer adjacent motifs (PAMs) are located within nucleosomes. The ability to introduce genetic modifications at loci within diverse chromatin environments is thus of great interest to the field of genetic engineering.SUMMARY
[0005] The present disclosure provides polypeptides, polynucleotides, compositions, methods, and systems for genome editing involving a chromatin remodeler. For example, the High Mobility Group Box 1 (HMGB1) protein is a nucleosome-interacting protein that is capable of altering the interaction between DNA and DNA binding proteins, thereby increasing chromatin accessibility. Such chromatin remodelers may be fused with aprogrammable DNA-binding protein (e.g., a CRISPR-associated nuclease), conferring substantial advantages for genome editing applications. For example, such fusion proteins may show increased editing efficiency at loci that are positioned within compacted chromatin.
[0006] Accordingly, in some embodiments, the present disclosure provides, inter alia, an HMGB1 polypeptide comprising an HMGB1 Box B domain and at least 7 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain.
[0007] In some embodiments, the present disclosure provides, inter alia, an HMGB 1 polypeptide comprising a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to SEQ ID NO: 7 or that is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to SEQ ID NO: 17.
[0008] In some embodiments, the present disclosure provides, inter alia, a system comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a Class II Cas nuclease; wherein the HMGB1 polypeptide is operably linked to the Class II Cas nuclease.
[0009] In some embodiments, the present disclosure provides, inter alia, a fusion protein comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide domain comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a Class II Cas nuclease.
[0010] In some embodiments, the present disclosure provides, inter alia, a system comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) an NmeCas9 nuclease; wherein the HMGB1 polypeptide is operably linked to the NmeCas9 nuclease.
[0011] In some embodiments, the present disclosure provides, inter alia, a fusion protein comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) an NmeCas9 nuclease.
[0012] In some embodiments, the present disclosure provides, inter alia, a system comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; (2) a SpyCas9 nuclease; wherein the HMGB1 polypeptide is operably linked to the SpyCas9 nuclease.
[0013] In some embodiments, the present disclosure provides, inter aha, a fusion protein comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a SpyCas9 nuclease.
[0014] In some embodiments, the present disclosure provides, inter aha, a system comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; (2) a SpyCas9 nuclease; wherein the HMGB1 polypeptide is operably linked to the SpyCas9 nuclease.
[0015] In some embodiments, the present disclosure provides, inter aha, a fusion protein comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a SpyCas9 nuclease.
[0016] In some embodiments, the present disclosure provides, inter aha, an mRNA comprising an open reading frame (ORF) encoding an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain.
[0017] In some embodiments, the present disclosure provides, inter aha, an mRNA comprising an open reading frame (ORF) encoding a fusion protein comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a Class II Cas nuclease.
[0018] In some embodiments, the present disclosure provides, inter alia, an mRNA comprising an open reading frame (ORF) encoding a fusion protein comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) an NmeCas9 nuclease.
[0019] In some embodiments, the present disclosure provides, inter alia, an mRNA comprising an open reading frame (ORF) encoding a fusion protein comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a SpyCas9 nuclease.
[0020] In some embodiments, the present disclosure provides, inter alia, a composition comprising an mRNA comprising an open reading frame (ORF) encoding: (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; (2) a fusion protein comprising (i) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (ii) a Class II Cas nuclease; (3) a fusion protein comprising (i) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (ii) an NmeCas9 nuclease; or (4) a fusion protein comprising (i) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (ii) a SpyCas9 nuclease; optionally wherein the composition comprises a guide RNA (gRNA).
[0021] In some embodiments, the present disclosure provides, inter alia, a lipid nanoparticle (LNP) composition comprising an mRNA comprising an open reading frame (ORF) encoding: (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic taildomain; (2) a fusion protein comprising (i) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (ii) a Class II Cas nuclease; (3) a fusion protein comprising (i) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (ii) an NmeCas9 nuclease; or (4) a fusion protein comprising (i) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (ii) a SpyCas9 nuclease; optionally wherein the composition comprises a guide RNA (gRNA).
[0022] In some embodiments, the present disclosure provides, inter alia, a method of producing a modification in a genomic sequence of a target cell, the method comprising contacting the cell with: (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; (2) a fusion protein comprising (i) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (ii) a Class II Cas nuclease; (3) a fusion protein comprising (i) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (ii) an NmeCas9 nuclease; (4) a fusion protein comprising (i) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (ii) a SpyCas9 nuclease; or (5) an mRNA comprising an open reading frame (ORF) encoding any of (l)-(4).BRIEF DESCRIPTION OF THE DRAWINGS
[0023] FIG. 1 shows the percent indel frequency at the PCSK9 locus in cells treated with the indicated mRNA (N=3).
[0024] FIG. 2 shows the percent indel frequency at the PCSK9 locus in cells treated with the indicated mRNA (N=3 unless otherwise indicated; *N=2).
[0025] FIG. 3 shows the percent indel frequency at the PCSK9 locus in cells treated with the indicated mRNA (N=3).
[0026] FIG. 4 shows the percent indel frequency at the TTR locus in cells treated with the indicated mRNA (N=3).
[0027] FIGs. 5A and B show the percent indel frequency at the SCAP locus in cells treated with (FIG. 5A) the indicated Nme2Cas9 mRNA or (FIG. 5B) the indicated SpyCas9 mRNA.
[0028] FIG. 6 shows the percent indel frequency at the SCAP locus in cells treated with the indicated mRNAs encoding Nme2Cas9 and an HMGB1 polypeptide.
[0029] FIG. 7 shows the percent indel frequency in cells treated with the indicated mRNAs and guide RNAs.
[0030] FIG. 8 shows a box-and- whisker plot from a one-way AN OVA analysis of the indel percentage for each of mRNA A, mRNA B and mRNA H, across all loci targeted by the guide RNAs indicated in FIG. 7.
[0031] FIG. 9 shows the rate of luciferase reporter insertion at the albumin locus in primary mouse hepatocyte (PMH) cells treated with G023805 (SEQ ID NO: 1006), the indicated mRNA, and AAV-DJ with a luciferase reporter sequence.
[0032] FIG. 10 shows the percent indel frequency at the albumin locus in primary mouse hepatocyte (PMH) cells treated with G023805 (SEQ ID NO: 1006), the indicated mRNA, and AAV-DJ with a luciferase reporter sequence.
[0033] FIG. 11 shows the percent indel frequency at the SCAP locus in cells treated with the indicated mRNAs encoding Nme2Cas9 or an Nme2Cas9-HMGBl fusion protein.
[0034] FIG. 12 shows the percent indel frequency at the TTR locus in cells treated with the indicated mRNAs encoding Nme2Cas9 or a Nme2Cas9-HMGBl fusion protein.
[0035] FIG. 13 shows the percent indel frequency at the PCSK9 locus in cells treated with the indicated mRNAs encoding Nme2Cas9 or a Nme2Cas9-HMGB 1 fusion protein.
[0036] FIG. 14 shows the percent C-to-T base editing at the PCSK9 locus in cells treated with the indicated mRNAs encoding Nme2Cas9-BC22 or a Nme2Cas9-BC22-HMGBl fusion protein.
[0037] FIG. 15 shows the percent indel frequency at the PCSK9 locus in the livers of mice treated with the indicated mRNAs encoding Nme2Cas9 or a Nme2Cas9-HMGBl fusion protein.
[0038] FIG. 16 shows the frequency of C-to-T base editing, and indels at the PCSK9 locus in the livers of mice treated with an mRNA encoding UGI and the indicated mRNAs encoding Nme2Cas9-BC22 or a Nme2Cas9-BC22-HMGBl fusion protein.
[0039] FIG. 17 shows the percent indel frequency at the SCAP locus in cells treated with the indicated mRNAs encoding Nme2Cas9 or a Nme2Cas9-HMGB 1 fusion protein and a SCAP targeting guide RNA.
[0040] FIG. 18 shows the percent indel frequency at the TRAC locus in cells treated with the indicated mRNAs encoding Nme2Cas9 or a Nme2Cas9-HMGB 1 fusion protein and a TRAC targeting guide RNA.
[0041] FIGs. 19A-19C show scatter plots showing a comparison of the upregulation and downregulation of transcriptome expression in primary human hepatocyte (PHH) cells treated with mRNA A vs. untreated cells (no LNP) (FIG. 19A), treated with mRNA B vs. untreated cells (no LNP) (FIG. 19B), and treated with mRNA A vs. treated with mRNA B (FIG. 19C). NS=not statist ical ly significant.
[0042] FIGs. 20A-C show scatter plots showing a comparison of the upregulation and downregulation of transcriptome expression in T-cells treated with mRNA A vs. untreated cells (no LNP) (FIG. 20A), treated with mRNA B vs. untreated cells (no LNP) (FIG. 20B), or treated with mRNA A vs. treated with mRNA B (FIG. 20C). NS=not statistically significant.
[0043] FIG. 21 shows the percent C-to-T base editing at the SCAP or TTR locus in cells treated with the indicated guide RNA targeting the SCAP (G026304; SEQ ID NO: 1007) or TTR (G021884) locus and the indicated mRNA encoding Nme2Cas9-BC22 or a Nme2Cas9-BC22-HMGBl fusion protein.
[0044] FIG. 22 shows a structure of wildtype HMGB1, including various domains therein.
[0045] FIG. 23 shows domains of exemplary fusion proteins comprising an HMGB1 polypeptide and a programmable DNA-binding protein domain. In the context of FIG. 23, “DNA binding domain” refers to a programmable DNA-binding domain disclosed herein and “NLS” refers to a heterologous NLS disclosed herein.BRIEF DESCRIPTION OF DISCLOSED SEQUENCESDETAILED DESCRIPTION
[0046] The section headings used herein are for organizational purposes only and are not to be construed as limiting the desired subject matter in any way. In the event that any material incorporated by reference contradicts any term defined in this specification or any other express content of this specification, this specification controls. While the present teachings are described in conjunction with various embodiments, it is not intended that the present teachings be limited to such embodiments. On the contrary, the present teachings encompass various alternatives, modifications, and equivalents, as will be appreciated by those of skill in the art.Definitions
[0047] Unless stated otherwise, the following terms and phrases as used herein are intended to have the following meanings:
[0048] “Polynucleotide” and “nucleic acid” are used herein to refer to a multi meric compound comprising nucleosides or nucleoside analogs which have nitrogenous heterocyclic bases or base analogs linked together along a backbone, including conventional RNA, DNA, mixed RNA-DNA, and polymers that are analogs thereof. A nucleic acid “backbone” can be made up of a variety of linkages, including one or more of sugar-phosphodiester linkages, peptide-nucleic acid bonds (“peptide nucleic acids” or PNA; PCT No. WO 95 / 32305), phosphorothioate linkages, methylphosphonate linkages, or combinations thereof. Sugar moieties of a nucleic acid can be ribose, deoxyribose, or similar compounds with subslilulions, e.g., 2’ methoxy, 2’ halide, or 2’-O-(2-methoxyethyl) (2’-O-moe) subslilulions. Nitrogenous bases can be conventional bases (A, G, C, T, U), analogs thereof (e.g., modified uridines such as 5 -methoxyuridine, pseudouridine, or N1 -methylpseudouridine, or others); inosine; derivatives of purines or pyrimidines (e.g., N4-methyl deoxyguanosine, deaza- or aza-purines, deaza- or aza-pyrimidines, pyrimidine bases with substituent groups at the 5 or 6 position (e.g., 5-methylcytosine), purine bases with a substituent at the 2, 6, or 8 positions, 2-amino-6-methylaminopurine, O6-methylguanine, 4-thio-pyrimidines, 4-amino-pyrimidines, 4-dimethylhydrazine-pyrimidines, and O4-alkyl-pyrimidines; US Pat. No. 5,378,825 and PCT No. WO 93 / 13121). For general discussion see The Biochemistry of the Nucleic Acids 5-36, Adams et al., ed., 11thed., 1992). Nucleic acids can include one or more “abasic” residues where the backbone includes no nitrogenous base for position(s) of the polymer (US Pat. No.5,585,481)- A nucleic acid can comprise only conventional RNA or DNA sugars, bases and linkages, or can include both conventional components and substitutions (e.g., conventional bases with 2’ methoxy linkages, or polymers containing both conventional bases and one or more base analogs). Nucleic acid includes “locked nucleic acid” (LNA), an analogue containing one or more LNA nucleotide monomers with a bicyclic furanose unit locked in an RNA mimicking sugar conformation, which enhance hybridization affinity toward complementary RNA and DNA sequences (Vester and Wengel, 2004, Biochemistry 43(42): 13233-41). Nucleic acid includes “unlocked nucleic acid” enables the modulation of the thermodynamic stability and also provides nuclease stability. RNA and DNA have different sugar moieties and can differ by the presence of uracil or analogs thereof in RNA and thymine or analogs thereof in DNA. It is understood that, depending on the function of the polynucleotide, certain nucleosides or nucleoside analogs may be preferred, or certain nucleosides or nucleoside analogs may not be tolerated.
[0049] “Polypeptide” as used herein refers to a multimeric compound comprising amino acid residues that can adopt a three-dimensional conformation. Polypeptides include but are not limited to enzymes, enzyme precursor proteins, regulatory proteins, structural proteins, receptors, nucleic acid binding proteins, antibodies, etc. Polypeptides may, but do not necessarily, comprise post-translational modifications, non-natural amino acids, prosthetic groups, and the like. In some embodiments, a polypeptide may comprise an amino acid sequence provided by a SEQ ID NO listed herein (e.g., in Table 37) that includes an N-terminal methionine. In some embodiments, such a polypeptide may comprise the amino acid sequence of the SEQ ID NO comprising the N-terminal methionine. In other embodiments, the polypeptide may comprise the amino acid sequence of the SEQ ID NO: as listed, but without the N-terminal methionine. For example, a SEQ ID NO provided herein may, in some embodiments, be operably linked to another moiety at the N-terminus (e.g., a nuclear localization signal, a second polypeptide, etc.). In such embodiments, a person of skill in the art will recognize that the N-terminal methionine is optional. Accordingly, in SEQ ID NOs provided herein that contain an N-terminal methionine residue, the N-terminal methionine residue should be considered an optional feature. Likewise, for nucleic acid sequences provided herein that encode a protein or polypeptide that comprises an N-terminal methionine residue, the nucleotides encoding the N-terminal methionine residue should be considered an optional feature.
[0050] As used herein, “ribonucleoprotein” (RNP) or “RNP complex” refers to a guide RNA together with an RNA-guided DNA binding agent, such as a Cas nuclease, e.g., a Cascleavase, Cas nickase, or dCas DNA binding agent (e.g., Cas9). In some embodiments, the guide RNA guides the RNA-guided DNA binding agent such as Cas9 to a target sequence, and the guide RNA hybridizes with the target sequence and the agent binds to the target sequence; in cases where the agent is a cleavase or nickase, binding can be followed by cleaving or nicking.
[0051] “Programmable DNA-binding protein” as used herein refers to a protein capable of binding to a specific sequence of DNA. The DNA-binding domain of the programmable DNA-binding protein is programmable, meaning that it can be designed or engineered to recognize and bind different DNA sequences. In some embodiments, for example, DNA binding is mediated by interactions between the programmable DNA-binding protein and the target DNA. In such embodiments, the DNA-binding domain can be programmed to bind a DNA sequence of interest by protein engineering. In other embodiments, for example, DNA binding is mediated by a guide RNA that interacts with the programmable DNA-binding protein and the target DNA. In such instances, the programmable DNA-binding protein can be targeted to a DNA sequence of interest by designing the appropriate guide RNA.
[0052] In some embodiments, the programmable DNA-binding protein further comprises a domain for modifying the DNA. In some embodiments, the modification has nuclease activity and can cleave one or both strands of a double-stranded DNA sequence. The DNA break can then be repaired by a cellular DNA repair process such as non-homologous end joining (NHEJ) or homology-directed repair (HDR), such that the DNA sequence can be modified by a deletion, insertion, or substitution of at least one base pair. Exemplary programmable DNA-binding proteins having nuclease activity include, but are not limited to, CRISPR nucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases, meganucleases, and programmable DNA binding domains linked to nuclease domains.
[0053] The programmable DNA-binding protein can comprise wildtype or naturally occurring programmable DNA-binding domains or DNA-modification domains, modified versions of naturally occurring programmable DNA-binding domains or DNA-modification domains, synthetic or artificial DNA-binding domains or DNA -modification domains, and combinations thereof.
[0054] “Cas nuclease,” also called “Cas protein” as used herein, encompasses Cas cleavases, Cas nickases, and dCas DNA binding agents. Cas cleavases / nickases and dCas DNA binding agents include a Csm or Cmr complex of a type III CRISPR system, the CaslO,Csml, or Cmr2 subunit thereof, a Cascade complex of a type I CRISPR system, the Cas3 subunit thereof, and Class 2 Cas nucleases. As used herein, a “Class 2 Cas nuclease” is a single-chain polypepride with RNA-guided DNA binding activity. Class 2 Cas nucleases include Class 2 Cas cleavases, Class 2 Cas nickases (e.g., H840A, D10A, or N863A variants), which further have RNA-guided DNA cleavases or nickase activity, and Class 2 dCas DNA binding agents, in which cleavase / nickase activity is inactivated. Class 2 Cas nucleases include, for example, Cas9, Cpfl, C2cl, C2c2, C2c3, HF Cas9 (e.g., N497A, R661A, Q695A, Q926A variants), HypaCas9 (e.g., N692A, M694A, Q695A, H698A variants), eSPCas9(1.0) (e.g., K810A, K1003A, R1060A variants), and eSPCas9(l.l) (e.g., K848A, K1003A, R1060A variants) proteins and modifications thereof. Cpfl protein, Zetsche et al., Cell, 163: 1-13 (2015), is homologous to Cas9, and contains a RuvC-like nuclease domain. Cpfl sequences of Zetsche are incorporated by reference in their entirety. See, e.g., Zetsche, Tables SI and S3. See, e.g., Makarova et al., Nat Rev Microbiol, 13(11): 722-36 (2015); Shmakov et al., Molecular Cell, 60:385-397 (2015). In certain embodiments, the Cas9 is a Streptococcus pyogenes Cas9 (SpyCas9). In certain embodiments, the Cas9 is <\ Neisseria meningitidis Cas9 (NmeCas9), e.g., an Nme2Cas9.
[0055] The term “High Mobility Group Box 1” or “HMGB1,” as used herein in the context of the HMGB 1 protein, refers to a non-histone, nuclear DNA-binding protein belonging to the High Mobility Group-Box superfamily, or a portion thereof. The wildtype HMGB1 protein is composed of 215 amino acids in three structural domains: a HMGB1 Box A DNA-binding domain (alternatively referred elsewhere herein as a HMGB1 Box A domain), a Box B DNA-binding domain (alternatively referred elsewhere herein as an HMGB1 Box B domain), and an acidic tail domain. The wildtype HMGB1 protein comprising SEQ ID NO: 1 may be understood to comprise, from N-terminus to C-terminus, a HMGB1 Box A DNA-binding domain comprising the sequence of amino acid residues 2-78 of SEQ ID NO: 1, an HMGB1 Box B DNA-binding domain comprising the sequence of amino acid residues 89-162 of SEQ ID NO: 1, a cryptic nuclear localization signal (NLS) comprising the sequence of amino acid residues 179-185 of SEQ ID NO: 1, and an acidic tail corresponding to amino acid residues 186-215 of SEQ ID NO: 1. FIG. 22 shows an exemplary schematic of the wildtype HMGB1 protein. The term “HMGB1 polypeptide” as used herein refers to a polypeptide comprising the amino acid sequence of HMGB 1, or a portion or fragment thereof. For example, in some embodiments, an HMGB1 polypeptide can comprise one or more HMGB1 domains (such as an HMGB1 Box B domain, an HMGB1 Box A domain, and an acidic tail domain). It is understood that, when an HMGB1 polypeptide is present as part of afusion protein, the HMGB1 polypeptide may be referred as “HMGB1” (e.g., “a fusion protein comprising, from N-terminus to C-terminus, deaminase-first linker-DNA-binding domain-heterologous NLS-second linker-HMGBl”). In some embodiments, the acidic tail may serve as a transcription stimulatory domain. “HMGB1” as used herein in the context of nucleic acids refers to a nucleic acid (e.g., DNA or mRNA) encoding an HMGB1 polypeptide. The human HMGB1 gene has accession number NC_000013.11 (30456704..30617597). It is understood that, when an HMGB1 polypeptide (e.g., an HMGB1 protein of SEQ ID NO: 1) is present as part of a fusion protein, the N-terminal methionine residue of HMGB 1 may be omitted.
[0056] As used herein, in the context of the wildtype HMGB 1 protein, the term “Box A” refers to an amino acid sequence comprising or consisting of SEQ ID NO: 3. In some embodiments, the HMGB1 Box A domain may comprise or consist of amino acid residues 2-78 of SEQ ID NO: 1. In some embodiments, the HMGB1 Box A domain may have DNA binding activity.
[0057] As used herein, in the context of the wildtype HMGB 1 protein, the term “Box B” refers to an amino acid sequence comprising or consisting of SEQ ID NO: 4. In some embodiments, the HMGB1 Box B domain may comprise or consist of amino acid residues 89-162 of SEQ ID NO: 1. In some embodiments, the HMGB1 Box B domain may have DNA binding activity.
[0058] As used herein, in the context of the wildtype HMGB 1 protein, the term “cryptic nuclear localization signal” or “cryptic NLS” refers to an amino acid sequence comprising or consisting of EKSKKKK (SEQ ID NO: 5) or a variant thereof. In some embodiments, the cryptic NLS may comprise amino acid residues 179-185 of SEQ ID NO: 1. In some embodiments, the cryptic NLS may comprise additional amino acid residues. In some embodiments, the cryptic NLS may comprise a part of or all of amino acid residues 166-185 of SEQ ID NO: 1. In some embodiments, the cryptic NLS disclosed herein comprises a variant of EKSKKKK (SEQ ID NO: 5) wherein one amino acid residue of SEQ ID NO: 5 comprises a conservative substitution thereof. As used herein, a cryptic NLS, as provided by SEQ ID NO: 5, or as a variant thereof, is not a heterologous NLS described herein.
[0059] In some embodiments, the cryptic NLS may have constitutive NLS activity (that is, act as a signal fragment that mediates the nuclear import of the HMGB 1 protein). In some embodiments, the cryptic NLS may not have constitutive NLS activity.
[0060] As used herein, in the context of the wildtype HMGB 1 protein, the term “receptor for advanced glycation end-products (RAGE) binding domain” or “RAGE binding domain” refers to an amino acid sequence comprising amino acid residues 150-183 of SEQ IDNO:1. See also FIG. 22. In some embodiments, the HMGB1 polypeptide described herein comprises a RAGE binding domain, or portion thereof, C-terminal to the HMGB1 Box B domain. In some embodiments, the RAGE binding domain comprises the amino acid sequence of SEQ ID NO: 21 C-terminal to the HMGB1 Box B domain.
[0061] As used herein, in the context of the wildtype HMGB 1 protein, the term “acidic tail domain” of an amino acid sequence comprising or consisting of SEQ ID NO: 6. In some embodiments, the acidic tail domain may comprise or consist of amino acid residues 186-215 of SEQ ID NO: 1. In some embodiments, the acidic tail domain may have transcriptional stimulatory function.
[0062] As used herein, the term “operably linked” refers to a juxtaposition of at least two components, e.g., polypeptide chain components or polynucleotide chain components, in a manner to allow one component to exert an effect on the other, or to allow the components when operably linked to have an effect that the two components could not exert individually or as a mixture of the components. In certain embodiments, two components may be operably linked by a covalent linkage, e.g., a peptide bond between two polypeptide chains encoded by a single open reading frame. In certain embodiments, the linkage can be mediated by a linker, e.g., a peptide linker. In certain embodiments, the linker is covalently attached to at least one of the components. In certain embodiments, the linker is covalently attached to both components. In certain embodiments, e.g., polynucleotides, the components can be joined by one or more phosphodiester or phosphorothioate bonds, e.g., a single phosphodiester (PO) or phosphorothioate (PS) bond, or a polynucleotide linker sequence, or by other types of linkages, or any combination thereof.
[0063] In certain embodiments, components are joined directly by a non-covalent linkage. For example, each of the components can include a domain, e.g., a heterodimerization domain (or heterodimer domain, which is interchangeably used herein with a heterodimerization domain), that mediates the binding of the components to each other. In some embodiments, each of the components is a polypeptide chain and each polypeptide chain comprises a fusion protein comprising the heterodimer domain. In some embodiments, the heterodimer domains comprise leucine zipper domains, PDZ domains, streptavidin and streptavidin binding protein domains, foldon domains, hydrophobic polypeptides, an antibody or one or more binding fragments thereof (e.g., scFv, VHH domain) on one component that binds an epitope, either naturally occurring or inserted, in the other component. The position of the heterodimer domain is selected independently for each component.
[0064] As used herein, the term “system” refers to at least two operably linked proteins that are not part of the same fusion protein. In some embodiments, the system may comprise an HMGB1 polypeptide. In some embodiments, the system may further comprise a programmable DNA-binding protein. In some embodiments, the system may further comprise a polynucleotide.
[0065] As used herein, the term “fusion protein” refers to a hybrid polypeptide which comprises polypeptides from at least two different proteins or sources. One polypeptide may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxyterminal (C- terminal) protein thus forming an “amino-terminal fusion protein” or a “carboxyterminal fusion protein,” respectively. It is understood that, if one or more of the polypeptides includes an N-terminal methionine corresponding to a start codon, the methionine may be omitted from the fusion protein. For example, a polypeptide that is not located at the N-terminal portion in some cases may lack an N-terminal methionine and its coding sequence may similarly lack a start codon in its open reading frame that ordinarily would encode an N-terminal methionine. Any of the proteins provided herein may be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N. Y. (2012)), the entire contents of which are incorporated herein by reference.
[0066] The term “domain,” as used herein in the context of a protein, refers to a distinct physical, structural, or functional region of the protein, or an amino acid sequence in a protein that is associated with a particular function. In the context of a polypeptide such as a wildtype HMGB 1 protein, a domain may comprise a structural or functional region such as Box A DNA binding domain, Box B DNA-binding domain, the cryptic NLS, or the acidic tail domain, as shown in FIG. 22. In the context of a fusion protein, which comprises polypeptides from at least two different proteins or sources, a domain may comprise one of those proteins or a physical, structural, or functional region of one of those proteins. For example, in the context of an NmeCas9-HMGBl fusion protein, a domain may comprise the full HMGB1 protein or one or more of Box A, Box B, the cryptic NLS, and the acidic tail. Certain polypeptides, systems, and fusion proteins disclosed herein comprise an HMGB1 polypeptide, for example, which may comprise one or more of an HMGB1 Box A DNA binding domain, an HMGB1 Box B DNA binding domain, a cryptic NLS, an acidic tail domain, or RAGEbinding domain, but which may also lack one or more domains from wildtype HMGB 1 such as an acidic tail domain or an HMGB1 Box A DNA binding domain.
[0067] As used herein, the term “intervening peptide sequence” refers to a plurality of amino acid residues that can be used to link two moieties, e.g., two polypeptide sequences, to provide a fusion protein. In certain embodiments, the intervening peptide sequence provides the desired distance between the two polypeptide sequences, where the exact identity of the amino acid residues is not essential to all functions of the fusion protein. In certain embodiments, the intervening peptide sequence can include a linker sequence. In certain embodiments, the intervening peptide sequence can include an NLS sequence, e.g., a heterologous NLS sequence. In certain embodiments, the intervening peptide sequence can include both a linker sequence and an NLS sequence, e.g., a heterologous NLS sequence. In certain embodiments, the NLS sequence of the intervening peptide sequence can function to promote transport of the fusion protein into the nucleus, and also function to provide the desired distance between the two polypeptides in the fusion protein.
[0068] The term “linker,” as used herein, refers to a chemical group or a molecule linking two adjacent molecules or moieties. Typically, the linker is positioned between, or flanked by, two groups, molecules, or other moieties and connected to each one via a covalent bond. In some embodiments, the linker is a peptide linker comprising an amino acid or a plurality of amino acids (e.g., a peptide or protein) such as a 16-amino acid residue “XTEN” linker, or a variant thereof (See, e.g., the Examples; and Schellenberger et al. A recombinant polypeptide extends the in vivo half-life of peptides and proteins in a tunable manner. Nat. Biotechnol. 27, 1186-1190 (2009) or a GS linker. As used herein, the term “GS linker” refers to a linker sequence that is rich in glycine and serine, e.g., at least 50%, 60%, 70%, 80%, 90%, or 100% glycine or serine amino acid residues. In some embodiments, the linker comprises one or more sequences selected from SEQ ID NOs: 301-365 and 425-435.
[0069] As used herein, the terms “nuclear localization signal” (NLS) or “nuclear localization sequence” refers to an amino acid sequence which induces transport of molecules comprising such sequences or linked to such sequences into the nucleus of eukaryotic cells. The nuclear localization signal may form part of the molecule to be transported. In some embodiments, the NLS may be fused to the molecule by a covalent bond, hydrogen bonds or ionic interactions. In some embodiments, the NLS may be fused to the molecule via a linker.
[0070] As used herein, “open reading frame” or “ORF” of a gene refers to a sequence consisting of a series of codons that specify the amino acid sequence of the protein that the gene codes for. The ORF generally begins with a start codon (e.g., ATG in DNA or AUG inRNA) and ends with a stop codon, e.g., TAA, TAG or TGA in DNA or UAA, UAG, or UGA in RNA. The ORF sequences described herein may or may not include a start codon encoding an N-terminal methionine. Thus, in some cases, an ORF described in a SEQ ID NO herein comprises a start codon and the ORF includes the start codon. In other cases, an ORF comprises the sequence of the listed SEQ ID NO but without the start codon provided in the sequence of the listed SEQ ID NO. In some cases, an ORF described in a SEQ ID NO herein comprises a stop codon and the ORF includes the stop codon. In other cases, an ORF comprises the sequence of the listed SEQ ID NO but without the stop codon provided in the sequence of the listed SEQ ID NO.
[0071] “Messenger RNA” or “mRNA” is used herein to refer to a polynucleotide that is not DNA and comprises an open reading frame that can be translated into a polypeptide (i.e., can serve as a substrate for translation by a ribosome and amino-acylated tRNAs). mRNA can comprise one or more modifications, e.g., as provided below. In general, mRNAs do not contain a substantial quantity of thymidine residues (e.g., 0 residues or fewer than 30, 20, 10, 5, 4, 3, or 2 thymidine residues; or less than 10%, 9%, 8%, 7%, 6%, 5%, 4%, 4%, 3%, 2%, 1%, 0.5%, 0.2%, or 0.1% thymidine content). An mRNA can comprise one or more chemically modified nucleosides such as 5-methyl-cytidine (5mC), 2-thio-uridine (2sU), Nl-methylpseudouridine (ml \| / U) and pseudo-uridine (yU), or a modified cap structure as provided below. For example, an mRNA can contain modified uridines at some or all of its uridine positions.
[0072] “Modified uridine” is used herein to refer to a nucleoside other than thymidine with the same hydrogen bond acceptors as uridine and one or more structural differences from uridine. In some embodiments, a modified uridine is a substituted uridine, i.e., a uridine in which one or more non-proton substituents (e.g., alkoxy, such as methoxy) takes the place of a proton. In some embodiments, a modified uridine is pseudouridine. In some embodiments, a modified uridine is a substituted pseudouridine, i.e., a pseudouridine in which one or more non-proton substituents (e.g., alkyl, such as methyl) takes the place of a proton. In some embodiments, a modified uridine is any of a substituted uridine, pseudouridine, or a substituted pseudouridine.
[0073] “Uridine position” as used herein refers to a position in a polynucleotide occupied by a uridine or a modified uridine. Thus, for example, a polynucleotide in which “100% of the uridine positions are modified uridines” contains a modified uridine at every position that would be a uridine in a conventional RNA (where all bases are standard A, U, C, or G bases) of the same sequence. Unless otherwise indicated, a U in a polynucleotidesequence of a sequence table or sequence listing in or accompanying this disclosure can be a uridine or a modified uridine.
[0074] As used herein, “HiBiT” or “HiBiT tag” refers to a 11 -amino acid tag sequence (e.g., SEQ ID NO: 437) typically used in vitro for observing protein expression. A person of skill in the art would recognize that a HiBiT tag can be useful for research purposes, but that any polypeptide or fusion protein constructs described herein containing a tag, or polynucleotides encoding same, can also be used without the tag, without further optimization. It is to be understood that, if the sequence includes a tag, the tag is not included in the amino acid sequence within the scope of the claimed invention. It is to be understood that, if the amino acid sequence includes a tag (e.g., HiBiT tag (e.g., SEQ ID NO: 437) or HiBiT tag with preceding linker (e.g., SEQ ID NO: 439), the tag (and preceding linker) is not included in the amino acid sequence within the scope of the claimed invention. It is to be understood that, if the nucleic acid sequence includes a sequence encoding tag (e.g., HiBiT tag (e.g., SEQ ID NO: 438) or a HiBiT tag with preceding linker (e.g., SEQ ID NO: 440), the sequence encoding the tag (and preceding linker) is not included in the nucleic acid sequence within the scope of the claimed invention.
[0075] “Guide RNA”, “gRNA”, and “guide” are used herein interchangeably to refer to either a crRNA (also known as CRISPR RNA), or the combination of a crRNA and a trRNA (also known as tracrRNA). The crRNA and trRNA may be associated as a single RNA molecule (single guide RNA, sgRNA) or in two separate RNA molecules (dual guide RNA, dgRNA). “Guide RNA” or “gRNA” refers to each type. The trRNA may be a naturally occurring sequence, or a trRNA sequence with modifications or variations compared to naturally occurring sequences.
[0076] As used herein, the term “genomic locus,” when used in the context of a genomic locus being targeted by a guide RNA, includes one or more parts of a genomic sequence, the targeting of which affects the expression of the gene that is associated with the locus. For example, a genomic locus may include a coding sequence of a gene, an intron sequence of a gene, a regulatory sequence, a transcriptional control sequence of a gene, a translational control sequence of a gene, a splicing site, or a non-coding sequence between genes (e.g., intergenic space).
[0077] As used herein, a “guide sequence” or “guide region” or “targeting sequence” or “spacer” or “spacer sequence” and the like refers to a sequence within a guide RNA that is complementary to a target sequence and functions to direct a guide RNA to a target sequence for binding or modification (e.g., cleavage) by an RNA-guided nickase. A guide sequence can be 20 nucleotides in length, e.g., in the case of Streptococcus pyogenes (i.e., Spy Cas9 (also referred to as SpCas9)) and related Cas9 homologs / orthologs. Shorter or longer sequences can also be used as guides, e.g., 17-, 18-, 19-, 21-, 22-, 23-, 24-, or 25 -nucleotides in length. Aguide sequence can be 20-25 nucleotides in length, e.g., in the case of NmeCas9, e.g., 20-, 21-, 22-, 23-, 24-or 25 -nucleotides in length. For example, a guide sequence of 24 nucleotides in length can be used with NmeCas9, e.g., Nme2Cas9.
[0078] In some embodiments, the target sequence is in a genomic locus or on a chromosome, for example, and is complementary to the guide sequence. In some embodiments, the degree of complementarity or identity between a guide sequence and its corresponding target sequence may be about 80%, 85%, preferably about 90%, 95%, or 100%. In some embodiments, the guide sequence and the target sequence may be 100% complementary or identical. In other embodiments, the guide sequence and the target sequence may contain at least one mismatch. For example, the guide sequence and the target sequence may contain 1, 2, 3, or 4 mismatches, where the total length of the duplex formed between the guide sequence and the target sequence in a genomic sequence is at least 20 base pairs. In some embodiments, the degree of complementarity or identity between a guide sequence and its corresponding target sequence is at least 80%, 85%, preferably at least 90%, or 95%, for example when, the guide sequence comprises a sequence 20 contiguous nucleotides. In other embodiments, the guide sequence and the target sequence may contain at least one mismatch, i.e., one nucleotide that is not identical or not complementary, depending on the reference sequence. For example, the guide sequence and the target sequence may contain 1-2, preferably no more than 1 mismatch, where the total length of the target sequence is 19, 20, 21, 22, 23, or 24, nucleotides, or more. In some embodiments, the guide sequence and the target region may contain 1-2 mismatches where the guide sequence comprises at least 24 nucleotides, or more. In some embodiments, the guide sequence and the target region may contain 1-2 mismatches where the guide sequence comprises 24 nucleotides.
[0079] As used herein, a “target sequence” or “genomic target sequence” refers to a sequence of nucleic acid in a target genomic locus, in either the positive or the negative strand, that has complementarity to the guide sequence of the guide RNA, i.e., that is sufficiently complementary to the guide sequence of the guide RNA to permit specific binding of the guide to the target sequence. The interaction of the target sequence and the guide scaffold sequence directs an RNA-guided DNA binding agent, e.g., a SpyCas9 nuclease, to bind, and potentially nick or cleave (depending on the activity of the agent), within the target sequence. The specific length of the target sequence and the number of mismatches possible between the target sequence and the guide sequence depend, for example, on the identity of the Cas9 nuclease being directed by the guide RNA. Target sequences for Cas proteins include both the positive and negative strands of genomic DNA (i.e., the sequence given and the sequence’sreverse complement), as a nucleic acid substrate for a Cas protein is a double stranded nucleic acid. Accordingly, where a guide sequence is said to be “complementary to a target sequence,” it is to be understood that the guide sequence may direct an RNA-guided DNA binding agent (e.g., SpyCas9 cleavase or nickase) to bind to the reverse complement of a target sequence. Thus, in some embodiments, where the guide sequence binds the reverse complement of a target sequence, the guide sequence is identical to certain nucleotides of the target sequence (e.g., the target sequence not including the PAM) except for the substitution of U for T in the guide sequence.
[0080] As used herein, a first sequence is considered to “comprise a sequence that is at least X% identical to” a second sequence if an alignment of the first sequence to the second sequence shows that X% or more of the positions of the second sequence in its entirety are matched by the first sequence. For example, the sequence AAGA comprises a sequence with 100% identity to the sequence AAG because an alignment would give 100% identity in that there are matches to all three positions of the second sequence. The differences between RNA and DNA (generally the exchange of uridine for thymidine or vice versa) and the presence of nucleoside analogs such as modified uridines do not contribute to differences in identity or complementarity among polynucleotides as long as the relevant nucleotides (such as thymidine, uridine, or modified uridine) have the same complement (e.g., adenosine for all of thymidine, uridine, or modified uridine; another example is cytosine and 5 -methylcytosine, both of which have guanosine as a complement). Thus, for example, the sequence 5’-AXG where X is any modified uridine, such as pseudouridine, N1 -methyl pseudouridine, or 5-methoxyuridine, is considered 100% identical to AUG in that both are perfectly complementary to the same sequence (5 ’-GAU). Exemplary alignment algorithms are the Smith-Waterman and Needleman-Wunsch algorithms, which are well-known in the art. One skilled in the art will understand what choice of algorithm and parameter settings are appropriate for a given pair of sequences to be aligned; for sequences of generally similar length and expected identity >50% for amino acids or >75% for nucleotides, the Needleman-Wunsch algorithm with default settings of the Needleman-Wunsch algorithm interface provided by the EBI at the www.ebi.ac.uk web server is generally appropriate.
[0081] As used herein, the term “internal linker”, in the context of a guide RNA, describes a non-nucleotide segment joining two nucleotides within a guide RNA. If the guide RNA contains a guide region, the internal linker is located outside of the spacer region (e.g., in the scaffold or conserved region of the guide RNA). In some embodiments, the internal linker comprises a polyethylene glycol (PEG) linker disclosed herein. The guide RNAs comprisingan internal linker disclosed herein comprise one of the structures / modification patterns disclosed in WO2022 / 261292, the contents of which are hereby incorporated by reference in its entirety.
[0082] As used herein, the term “contact” refers to providing at least one component so that the component physically contacts a cell, including physically contacting the cell surface, cytosol, or nucleus of the cell. “Contacting” a cell with a polypeptide encompasses, for example, contacting the cell with a nucleic acid that encodes the polypeptide and allowing the cell to express the polypeptide.
[0083] As used herein, “indel” refers to an insertion or deletion mutation consisting of a number of nucleotides that are either inserted, deleted, or inserted and deleted, e.g., at the site of double-strand breaks (DSBs), in a target nucleic acid. As used herein, when indel formation results in an insertion, the insertion is a random insertion at the site of a DSB and is not generally directed by or based on a template sequence.
[0084] As used herein, a “conservative substitution” refers to a substitution of an original amino acid residue in a protein with a different amino acid residue with similar biochemical properties, e.g., size, charge, and hydrophobicity. Thus, although a protein with a conservative substitution has an amino acid sequence that differs from the amino acid sequence of the original protein, it may retain some or all of its structure and function. The biochemical properties of different amino acid residues are known in the art, and those skilled in the art will understand how to select an appropriate substation for an original amino acid residue based on its specific properties. Conservative substitution tables providing functionally similar amino acids are well known in the art. Non-limiting examples of conservative substitutions include amino acids listed in one of the following eight groups, each of which list examples of amino acids that are conservative substitutions for one another: Group 1:Alanine (A), Glycine (G); Group 2: Aspartic acid (D), Glutamic acid (E); Group 3:Asparagine (N), Glutamine (Q); Group 4: Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); Group 6: Phenylalanine (E), Tyrosine (Y), Tryptophan (W); Group 7: Serine (S), Threonine (T); and Group 8: Cysteine (C).
[0085] As used herein, “increased” (or “increases”) expression of a protein on a cell refers to an increase in expression of the protein relative to an unmodified cell. In some embodiments, the surface expression of a protein on a cell is measured by flow cytometry and has “increased” surface expression relative to an unmodified cell as evidenced by an increase in fluorescence signal upon staining with the same antibody against the protein. The “increase” of protein expression can be measured by other known techniques in the field withappropriate controls known to those skilled in the art, e.g., western blot, ELISA, immunohistochemistry. In certain embodiments, expression is assayed in a subject sample, e.g., a tissue or body fluid, e.g., urine, blood, or serum or plasma derived therefrom. In certain embodiments, increase of expression can be assessed in the same subject sample that may not represent increase in all sites of expression systemically. For example, a protein may be expressed in multiple tissues, but treatment results in an increase of expression predominantly in a single tissue, e.g., in liver. In certain embodiments, increase can be assessed by an activity assay, or by the level of a surrogate marker present in a subject sample, e.g., urine or blood, that correlates with an increase of protein expression. In certain embodiments, increase can include a greater increase of expression of a protein a wild-type protein as compared to expression of the protein containing a mutation as in a subject heterozygous at a particular genomic sequence. In certain embodiments, increase can include increase of expression of a wild-type protein relative to expression of the protein containing a mutation. The ratio may be changed, for example, by preferentially increasing the expression of the wild-type protein by altering the genomic sequence encoding the protein containing a mutation to the wild-type sequence.
[0086] As used herein, “knockdown” or “reduction” refers to a partial decrease or loss in expression of a particular gene product (e.g., protein, mRNA, or both) relative to an unmodified cell. Knockdown of a protein can be measured either by detecting protein secreted by tissue or population of cells (e.g., in serum or cell media) or by detecting total cellular amount of the protein from a tissue or cell population of interest. Methods for measuring knockdown of mRNA are known and include sequencing of mRNA isolated from a tissue or cell population of interest. In some embodiments, “knockdown” may refer to some loss of expression of a particular gene product, for example a decrease in the amount of mRNA transcribed or a decrease in the amount of protein expressed or secreted by a population of cells (including in vivo populations such as those found in tissues).
[0087] As used herein, a “control” is understood as an appropriate matched sample or subject for comparison. For example, a control can be a cell population treated in the same manner as the test population except that the treatment used for the control population lacks at least one active agent, e.g., a guide RNA, an mRNA encoding a nuclease, an insertion construct, a lipid formulation. In certain embodiments, a control may be an internal control, e.g., a cell population or subject prior to treatment.
[0088] As used herein, a “population of cells comprising edited cells” (or “population of cells comprising engineered cells”) or the like refers to a cell population that comprisesedited cells (or engineered cells), however not all cells in the population must be edited. A cell population comprising edited cells may also include non-edited cells. The percentage of edited cells within a cell population comprising edited cells may be determined by counting the number of cells within the population that are edited in the population as determined by standard cell counting methods. For example, in some embodiments, a cell population comprising edited cells comprising a single genome edit will have at least 20%, 30%, 40%, preferably at least 50%, 60%, 70%, 80%, 90%, or 95% of the cells in the population with the single edit.
[0089] As used herein, “treatment” refers to any administration or application of a therapeutic for disease or disorder in a subject, and includes inhibiting the disease, arresting its development, relieving one or more symptoms of the disease, curing the disease, or preventing one or more symptoms of the disease, including reoccurrence of the symptom.
[0090] As used herein, “delivering” and “administering” are used interchangeably, and include ex vivo and in vivo applications.
[0091] Co-administration, as used herein, means that a plurality of substances are administered sufficiently close together in time so that the agents act together. Coadministration encompasses administering substances together in a single formulation and administering substances in separate formulations close enough in lime so that the agents act together.
[0092] As used herein, the phrase “pharmaceutically acceptable” means that which is useful in preparing a pharmaceutical composition that is generally non-toxic and is not biologically undesirable and that are not otherwise unacceptable for pharmaceutical use. Pharmaceutically acceptable generally refers to substances that are non-pyrogenic.Pharmaceutically acceptable can refer to substances that are sterile, especially for pharmaceutical substances that are for injection or infusion.
[0093] As used herein, a “subject” refers to any member of the animal kingdom. In some embodiments, “subject” refers to humans. In some embodiments, “subject” refers to non-human animals. In some embodiments, “subject” refers to primates. In some embodiments, a subject may be a transgenic animal, genetically engineered animal, or a clone. In certain embodiments of the present invention the subject is an adult, an adolescent, or an infant. In some embodiments, terms “individual” or “patient” are used and are intended to be interchangeable with “subject”.
[0094] The term “about” or “approximately” means an acceptable error for a particular value as determined by one of ordinary skill in the art, which depends in part on how the valueis measured or determined, or a degree of variation that does not substantially affect the properties of the described subject matter, or within the tolerances accepted in the art, e.g.. within 10%, 5%, 2%, or 1% or within two standard deviations of a set of values. Accordingly, unless indicated to the contrary, the numerical parameters set forth in the following specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. When “about” is present before the first value of a series, it is understood to modify each value in the series.
[0095] Before describing the present teachings in detail, it is to be understood that the disclosure is not limited to specific compositions or process steps, as such may vary. It should be noted that, as used in this specification and the appended claims, the singular form “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. Thus, for example, reference to “a conjugate” includes a plurality of conjugates and reference to “a cell” includes a plurality of cells and the like.
[0096] Numeric ranges are inclusive of the numbers defining the range. Measured and measurable values are understood to be approximate, taking into account significant digits and the error associated with the measurement. Also, the use of “comprise”, “comprises”, “comprising”, “contain”, “contains”, “containing”, “include”, “includes”, and “including” are not intended to be limiting. It is to be understood that both the foregoing general description and detailed description are exemplary and explanatory only and are not restrictive of the teachings.
[0097] Unless specifically noted in the specification, embodiments in the specification that recite “comprising” various components are also contemplated as “consisting of’ or “consi sling essentially of’ the recited components; embodiments in the specification that recite “consisting of’ various components are also contemplated as “comprising” or “consi sling essentially of’ the recited components; and embodiments in the specification that recite “consisting cssenlially of’ various components are also contemplated as “consisting of’ or “comprising” the recited components (this interchangeability does not apply to the use of these terms in the claims).
[0098] The term “or” is used in an inclusive sense, i.e., equivalent to “and / or,” unless the context clearly indicates otherwise.
[0099] Ranges are understood to include the numbers at the end of the range and all logical values therebetween. For example, 5-10 nucleotides is understood as 5, 6, 7, 8, 9, or 10 nucleotides, whereas 5-10% is understood to contain 5% and all possible values through 10%.
[0100] At least 17 nucleotides of a 20-nucleotide sequence is understood to include 17, 18, 19, or 20 nucleotides of the sequence provided, thereby providing a upper limit even if one is not specifically provided as it would be clearly understood. Similarly, up to 3 nucleotides would be understood to encompass 0, 1, 2, or 3 nucleotides, providing a lower limit even if one is not specifically provided. When “at least,” “up to,” or other similar language modifies a number, it is understood to modify each number in the series.
[0101] As used herein, “no more than” or “less than” is understood as the value adjacent to the phrase and logical lower values or integers, as logical from context, to zero. For example, a duplex region of “no more than 2 nucleotide base pairs” has a 2, 1, or 0 nucleotide base pairs. When “no more than” or “less than” is present before a series of numbers or a range, it is understood that each of the numbers in the series or range is modified.
[0102] In the event of a conflict between a sequence in the application and an indicated accession number or position in an accession number, the sequence in the application predominates.
[0103] As used herein, “detecting an analyte” and the like is understood as performing an assay in which the analyte can be detected, if present, wherein the analyte is present in an amount above the level of detection of the assay.
[0104] As used herein, it is understood that when the maximum amount of a value is represented by 100% (e.g., 100% inhibition or 100% encapsulation) that the value is limited by the method of detection. For example, 100% inhibition is understood as inhibition to a level below the level of detection of the assay, and 100% encapsulation is understood as no material intended for encapsulation can be detected outside the vesicles.
[0105] Reference will now be made in detail to certain embodiments of the invention, examples of which are illustrated in the accompanying drawings. While the invention is described in conjunction with the illustrated embodiments, it will be understood that they are not intended to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents, which may be included within the invention as defined by the appended claims and included embodiments.HMGB1
[0106] Chromatin structure may impact the efficacy of various genome editing techniques, and target sequences located in chromatin-dense regions, in particular, may be less amenable to genome editing. For example, the CRISPR-associated nuclease Cas9 can exhibit insufficient editing of certain targets whose protospacer adjacent motifs (PAMs) are located within nucleosomes. Chromatin modulating peptides fused with Cas9 nucleases have been reported to improve editing efficiency at refractory target sites. See, e.g., Ding X, Seebeck T, Feng Y, Jiang Y, Davis GD, Chen F. Improving CRISPR-Cas9 Genome Editing Efficiency by Fusion with Chromatin-Modulating Peptides. CRISPR J. 2019 Feb;2:51-63.
[0107] The wildtype HMGB 1 protein is a non-sequence-specific DNA-binding protein that interacts with DNA through its DNA-binding domains. HMGB1 binding to DNA introduces local DNA distortions, which may affect the structure and stability of nearby DNA-protein complexes. In particular, HMGB 1 binding within the linker region of nucleosomal DNA may destabilize the nucleosome structure, thereby increasing the accessibility of the surrounding chromatin. See, e.g., Starkova TY, Polyanichko AM, Artamonova TO, Tsimokha AS, Tomilin AN, Chikhirzhina EV. Structural Characteristics of High-Mobility Group Proteins HMGB1 and HMGB2 and Their Interaction with DNA. Int J Mol Sci. 2023 Feb 10;24(4):3577 and Lange SS, Vasquez KM. HMGB1: the jack-of-all-trades protein is a master DNA repair mechanic. Mol Carcinog. 2009 Jul;48(7):571-80.
[0108] The wildtype HMGB1 protein comprises, from N-terminus to C-terminus, a Box A DNA-binding domain, a Box B DNA-binding domain, a cryptic nuclear localization signal (NLS), and an acidic tail. The wildtype HMGB1 protein further comprises a receptor for advanced glycation end-products (RAGE) binding domain.
[0109] In the context of the wildtype protein, the acidic tail interacts with the Box A and Box B domains, causing the HMGB1 protein to adopt an auto-inhibited conformation that is not competent for DNA binding. The acidic tail therefore suppresses the chromatin remodeling activity of the wildtype HMGB 1 protein. In contrast, a truncated HMGB 1 protein lacking the acidic tail may remain in the active conformation and thereby exhibit increased chromatin remodeling activity.
[0110] While certain truncated HMGB1 proteins lacking an acidic tail may have increased chromatin remodeling activity relative to the wildtype HMGB1 protein, other elements within the HMGB1 protein may also promote chromatin remodeling. For example, truncated HMGB1 proteins comprising a cryptic NLS, which may or may not haveconstitutive NLS activity, may exhibit increased chromatin remodeling activity compared to truncated HMGB1 proteins that lack the cryptic NLS.
[0111] The present disclosure provides for an HMGB 1 polypeptide comprising certain domains of HMGB 1, or nucleic acids encoding the same. In some embodiments, the HMGB1 polypeptide comprises an HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises at least 7 contiguous HMGB1 amino acid residues C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 contiguous HMGB1 amino acid residues C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises at least 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 contiguous HMGB1 amino acid residues C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises 7-10, 7-12, 7-14, 7-16, 7-18, 7-20, 8-10, 8-12, 8-14, 8-16, 8-18, 8-18, 8-20, 10-12, 10-14, 10-16, 10-18, 10-20, 12-14, 12-16, 12-18, 12-20, 14-16, 14-18, 14-20, 16-18, 16-20, or 18-20 contiguous HMGB1 amino acid residues C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises 7-9, 7-11, 7-13, 7-15, 7-17, 7-19, 7-20, 9-11, 9-13, 9-15, 9-17, 9-19, 9-20, 11-13, 11-15, 11-17, 11-19, 11-20, 13-15, 13-17, 13-19, 13-20, 15-17, 15-19, 15-20, 17-19, 17-20, or 19-20 contiguous HMGB1 amino acid residues C-terminal to the HMGB1 Box B domain.
[0112] In some embodiments, the HMGB 1 polypeptide lacks all or part of an HMGB 1 acidic tail domain. In some embodiments, the HMGB 1 polypeptide lacks all of an acidic tail domain. In some embodiments, the HMGB1 polypeptide lacks part of an acidic tail domain. In some embodiments, the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain. In some embodiments, the HMGB1 polypeptide comprises 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises at least 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 contiguous amino acid residues C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises 7-10, 7-12, 7-14, 7-16, 7-18, 7-20, 8-10, 8-12, 8-14, 8-16, 8-18, 8-18, 8-20, 10-12, 10-14, 10-16, 10-18, 10-20, 12-14, 12-16, 12-18, 12-20, 14-16, 14-18, 14-20, 16-18, 16-20, or 18-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises 7-9, 7-11, 7-13, 7-15, 7-17, 7-19, 7-20, 9-11, 9-13, 9-15, 9-17, 9-19, 9-20, 11-13, 11-15, 11-17, 11-19, 11-20, 13-15, 13-17, 13-19, 13-20, 15-17, 15-19, 15-20, 17-19, 17-20, or 19-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain.
[0113] In some embodiments, the HMGB 1 polypeptide comprises a cryptic nuclear localization signal (NLS). In some embodiments, the cryptic NLS comprises the sequence of EKSKKKK (SEQ ID NO: 5) or a variant thereof. In some embodiments, the HMGB1 polypeptide comprises a receptor for advanced glycation end-products (RAGE) binding domain. In some embodiments, the RAGE binding domain comprises the sequence of SEQ ID NO: 21. In some embodiments, the HMGB1 polypeptide comprises 7-20 amino acid residues C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises 7-20 amino acid residues immediately C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises the sequence of SEQ ID NO: 22.
[0114] In some embodiments, the HMGB 1 polypeptide lacks an acidic tail domain. In some embodiments, the HMGB1 polypeptide lacks a sequence at least 50%, 60%, 70%, 80%, 90%, 94%, or 97% identical to SEQ ID NO: 6. In some embodiments, the HMGB1 polypeptide lacks the sequence of SEQ ID NO: 6. In some embodiments, the HMGB1 polypeptide lacks transcriptional stimulatory function.
[0115] In some embodiments, the HMGB 1 polypeptide comprises a cryptic NLS that is located C-terminal to the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises a cryptic NLS that is located C-terminal to the HMGB1 Box B domain, optionally wherein the cryptic NLS is comprised within the at least 7 contiguous amino acid residues. In some embodiments, the cryptic NLS immediately follows the C-terminal end of the HMGB1 Box B domain. In some embodiments, the HMGB1 polypeptide comprises a cryptic NLS that is located N-terminal to the HMGB1 Box B domain. In some embodiments, the cryptic NLS immediately precedes the N-terminal end of the HMGB1 Box B domain.
[0116] In some embodiments, the HMGB1 polypeptide lacks an HMGB1 Box A domain. In some embodiments, the HMGB1 polypeptide comprises two HMGB1 Box B domains. In some embodiments, the cryptic NLS is located C-terminal to the C-terminal end of the two HMGB1 Box B domains. In some embodiments, the HMGB1 polypeptide comprises an HMGB1 Box A domain. In some embodiments, the HMGB1 Box B domain is located C-terminal to the HMGB1 Box A domain.
[0117] In some embodiments, the HMGB1 polypeptide comprises, from N-terminus to C-terminus: (1) HMGB1 Box A domain-HMGB1 Box B domain-cryptic NLS; (2) HMGB1 Box B domain-HMGB1 Box B domain-cryptic NLS; or (3) HMGB1 Box B domain-cryptic NLS. In some embodiments, the HMGB1 polypeptide comprises a sequence having at least80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 2-5,7-10, and 21-22, or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 12-20. In some embodiments, the HMGB1 polypeptide comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, and 21-22. In some embodiments, the HMGB1 polypeptide comprises a sequence that is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 12-20.
[0118] In some embodiments, the HMGB 1 Box A domain comprises amino acid residues 2-78 of SEQ ID NO: 1. In some embodiments, the HMGB1 Box A domain comprises amino acid residues 2-84 of SEQ ID NO: 1. In some embodiments, the HMGB1 Box A domain comprises amino acid residues 2-88 of SEQ ID NO: 1. In some embodiments, the HMGB1 Box B domain comprises amino acid residues 89-162 of SEQ ID NO: 1. In some embodiments, the HMGB1 Box B domain comprises amino acid residues 85-185 of SEQ ID NO: 1, including the cryptic nuclear localization signal. In some embodiments, the HMGB1 Box B domain consists of amino acid residues 85-185, including the cryptic nuclear localization signal.
[0119] In some embodiments, the present disclosure provides for an HMGB1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB 1 Box B domain and at least 7 amino acid residues C-terminal to the HMGB1 Box B domain, and wherein the HMGB 1 polypeptide lacks all or part of an acidic tail domain. In some embodiments, the acidic tail domain comprises amino acid residues 186-215 relative to SEQ ID NO: 1. In some embodiments, the cryptic NLS comprises the sequence of EKSKKKK (SEQ ID NO: 5) or a variant thereof. In some embodiments, the present disclosure provides for an HMGB1 polypeptide comprising an HMGB1 Box B domain and lacking an acidic tail domain, further wherein the HMGB1 polypeptide comprises a cryptic NLS comprising amino acid residues 179-185 of SEQ ID NO: 1 or a variant thereof. In some embodiments, the present disclosure provides for an HMGB1 polypeptide comprising an HMGB1 polypeptide comprising an HMGB 1 Box B domain and lacking an acidic tail domain, further wherein the HMGB 1 polypeptide comprises amino acid residues 166-185 of SEQ ID NO: 1, which comprises the cryptic NLS (SEQ ID NO: 5), or a variant of amino acid residues 166-185 of SEQ ID NO: 1.
[0120] In some embodiments, the cryptic NLS is located C-terminal to the HMGB 1 Box B domain. In some embodiments, the cryptic NLS immediately follows the C-terminal end of the HMGB1 Box B domain. In some embodiments, the cryptic NLS is located N-terminal to the HMGB1 Box B domain. In some embodiments, the cryptic NLS immediately precedes the C-terminal end of the HMGB1 Box B domain.
[0121] In some embodiments, the sequence of EKSKKKK (SEQ ID NO: 5) is located C-terminal to the HMGB1 Box B domain. In some embodiments, the sequence of EKSKKKK (SEQ ID NO: 5) immediately follows the C-terminal end of the HMGB1 Box B domain. In some embodiments, the sequence of EKSKKKK (SEQ ID NO: 5) is located N-terminal to the HMGB1 Box B domain. In some embodiments, the sequence of EKSKKKK (SEQ ID NO: 5) immediately precedes the C-terminal end of the HMGB1 Box B domain.
[0122] In some embodiments, the HMGB1 polypeptide comprises the sequence of amino acid residues 166-185 relative to SEQ ID NO: 1, which comprises the cryptic NLS (SEQ ID NO: 5). In some embodiments, the HMGB1 polypeptide of amino acid residues 166-185 relative to SEQ ID NO: 1 has 2, 1, or 0 amino acid substitutions. In some embodiments, the amino acid substitutions are conservative amino acid substitutions.
[0123] In some embodiments, the HMGB1 polypeptide lacks an HMGB1 Box A domain. In some embodiments, the HMGB1 polypeptide comprises two HMGB1 Box B domains. In some embodiments, wherein the HMGB1 polypeptide comprises two HMGB1 Box B domains, the cryptic NLS is located C-terminal to the most C-terminal HMGB1 Box B domain. In some embodiments, wherein the HMGB1 polypeptide comprises two HMGB1 Box B domains, the sequence of EKSKKKK (SEQ ID NO: 5) is located C-terminal to the most C-terminal Box B domain. In some embodiments, the HMGB1 polypeptide comprises an HMGB1 Box A domain. In some embodiments, wherein the HMGB1 polypeptide comprises an HMGB1 Box A domain, the HMGB1 Box B domain is located C-terminal to the HMGB1 Box A domain.
[0124] In some embodiments, wherein the HMGB1 polypeptide of amino acid residues 166-185 relative to SEQ ID NO: 1 is present, the HMGB1 polypeptide is located C-terminal to the most C-terminal Box B domain. When only one Box B domain is present, it is understood to be the most C-terminal Box B domain.
[0125] In some embodiments, the HMGB1 polypeptide comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 2-5,7-10, and 21-22 or that is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 12-20.
[0126] In some embodiments, the HMGB 1 polypeptide comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to SEQ ID NO: 7 or that is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to SEQ ID NO: 17.
[0127] In some embodiments, the present disclosure provide for a polynucleotide encoding certain domains of HMGB1. In some embodiments, the polynucleotide comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 12-20. In some embodiments, the polynucleotide comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of SEQ ID NO: 17.Systems and Fusion Proteins Comprising HMGB1 and a Programmable DNA-Binding Protein
[0128] One aspect of the present disclosure provides for systems and fusion proteins comprising an HMGB1 polypeptide and a programmable DNA-binding protein. See, e.g., FIG. 23, which shows certain non-limiting, exemplary HMGB1 -programmable DNA-binding protein fusion proteins.
[0129] The efficacy of genetic tools comprising programmable DNA-binding proteins is predicated on the accessibility of target loci. For example, loci positioned within compacted chromatin may be inaccessible or less accessible to programmable DNA-binding proteins. As shown in the working examples herein, systems and fusion proteins comprising a HMGB 1 polypeptide may exhibit an increased ability to bind and edit target loci in chromatin dense regions of the genome relative to systems and fusion proteins that lack an HMGB 1 polypeptide.
[0130] In some embodiments, the present disclosure provides for a system or fusion protein comprising a programmable DNA-binding protein and an HMGB 1 polypeptide comprising an HMGB1 Box B domain and lacking an acid tail domain. In some embodiments, the HMGB1 polypeptide comprises a cryptic NLS. In some embodiments, the cryptic NLS comprises the sequence of EKSKKKK (SEQ ID NO: 5). In some embodiments, the programmable DNA-binding protein is located N-terminal to the HMGB1 polypeptide. In some embodiments, the programmable DNA-binding protein is located C-terminal to the HMGB1 polypeptide. In some embodiments, the system or fusion protein comprises at least one heterologous nuclear localization signal (NLS).Programmable DNA-Binding Proteins
[0131] In some embodiments, a programmable DNA-binding protein is a protein capable of binding to a specific sequence of DNA. The DNA-binding domain of theprogrammable DNA-binding protein is programmable, meaning that it can be designed or engineered to recognize and bind different DNA sequences. In some embodiments, for example, DNA binding is mediated by interactions between the programmable DNA-binding protein and the target DNA. In such embodiments, the DNA binding domain can be programmed to bind a DNA sequence of interest by protein engineering. In other embodiments, for example, DNA binding is mediated by a guide RNA that interacts with the programmable DNA-binding protein and the target DNA. In such instances, the programmable DNA-binding protein can be targeted to a DNA sequence of interest by designing the appropriate guide RNA.
[0132] In some embodiments, the programmable DNA-binding protein further comprises a domain for modifying the DNA. In some embodiments, the modification has nuclease act i vity and can cleave one or both strands of a double-stranded DNA sequence. The DNA break can then be repaired by a cellular DNA repair process such as non-homologous end joining (NHEJ) or homology-directed repair (HDR), such that the DNA sequence can be modified by a deletion, insertion, or substitution of at least one base pair. Exemplary programmable DNA-binding proteins having nuclease activity include, but are not limited to, CRISPR nucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases, meganucleases, and programmable DNA binding domains linked to nuclease domains.
[0133] The programmable DNA-binding protein can comprise wildtype or naturally-occurring DNA-binding or DNA-modification domains, modified versions of naturally occurring DNA-binding or DNA-modification domain, synthetic or artificial DNA binding or modification domains, and combinations thereof.
[0134] In some embodiments, the programmable DNA-binding protein is selected from a clustered regularly interspaced short palindromic repeats (CRISPR) nuclease, a zinc finger nuclease (ZFN), a transcription activator-like effector (TALE), a transcription activator-like effector nuclease (TALEN), a meganuclease, or a chimeric protein comprising a programmable DNA-binding domain linked to a nuclease domain.Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) Nucleases
[0135] In some embodiments, the programmable DNA-binding protein disclosed herein is a clustered regularly interspaced short palindromic repeat (CRISPR) nuclease. In some embodiments, the CRISPR nuclease is a CRISPR-associated (Cas) nuclease.
[0136] In some embodiments, the programmable DNA-binding protein is a Class II Cas nuclease. In some embodiments, the programmable DNA-binding protein is a modified Class II Cas protein or derived from a Class II Cas protein. In some embodiments, the programmable DNA-binding protein is modified or derived from a Cas protein, such as a Class II Cas nuclease (which may be, e.g., a Cas nuclease of Type II, V, or VI). Class II Cas nuclease include, for example, Cas9, Cpfl, C2cl, C2c2, and C2c3 proteins and modifications thereof. In some embodiments, the Class II Cas nuclease is selected from Cas9, Cpfl, C2cl, C2c2, C2c3, HF Cas9, HypaCas9, eSPCas9 (1.0), eSPCas9 (1.1), and variants thereof.Examples of Cas9 nucleases include those of the type II CRISPR systems of 5. pyogenes, S. aureus, and other prokaryotes (see, e.g., the list in the next paragraph), and modified (e.g., engineered or mutant) versions thereof. See, e.g., US2016 / 0312198 Al; US 2016 / 0312199 Al, which is incorporated by reference in its entirety. Other examples of Cas nucleases include a Csm or Cmr complex of a type III CRISPR system or the CaslO, Csml, or Cmr2 subunit thereof; and a Cascade complex of a type I CRISPR system, or the Cas3 subunit thereof. In some embodiments, the Cas nuclease may be from a Type-IIA, Type-IIB, or Type-IIC system. For discussion of various CRISPR systems and Cas nucleases, see, e.g., Makarova et al., Nat. Rev. Microbiol. 9:467-477 (2011); Makarova et al., Nat. Rev.Microbiol, 13: 722-36 (2015); Shmakov et al., Molecular Cell, 60:385-397 (2015).
[0137] A Cas nuclease described herein may be a cleavase, nickase, or deactivated Cas9 (dCas9) form of a Cas nuclease from the species including, but not limited to, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Listeria innocua, Lactobacillus gasseri, Francisella novicida, Wolinella succinogenes, Sutterella wadsworthensis, Gammaproteobacterium, Neisseria meningitidis, Campylobacter jejuni, Pasteurella multocida, Fibrobacter succinogene, Rhodospirillum rubrum, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Lactobacillus buchneri, Treponema denticola, Microscilla marina, Burkholderiales bacterium, Polar omonas naphthalenivorans, Polar omonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus,Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalter omonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Streptococcus pasteurianus, Neisseria cinerea, Campylobacter lari, Parvibaculum lavamentivorans, Corynebacterium diphtheria, Acidothermus cellulolyticus, Rhodopseudomonas palustris, Actinomyces naeslundii, Acidaminococcus sp., Lachnospiraceae bacterium ND2006, or Acaryochloris marina.
[0138] In some embodiments, the Class II Cas nuclease is a Cas9 nuclease. In some embodiments, the Cas9 nuclease is a cleavase, nickase, or dCas9 form of the Cas9 nuclease from Streptococcus pyogenes. In some embodiments, the Cas9 nuclease is a cleavase, nickase, or dCas9 form of the Cas9 nuclease from Streptococcus thermophilus. In some embodiments, the Cas9 nuclease is a cleavase, nickase, or dCas9 form of the Cas9 nuclease from Neisseria meningitidis. In some embodiments, the Cas9 nuclease is a cleavase, nickase, or dCas9 form of the Cas9 nuclease from Staphylococcus aureus. In some embodiments, the Cas9 nuclease is a cleavase, nickase, or dCas9 form of the Cas9 nuclease from C. diphtheriae. In some embodiments, the Cas9 nuclease is a cleavase, nickase, or dCas9 form of the Cas9 nuclease from A. cellulolyticus. In some embodiments, the Cas9 nuclease is a cleavase, nickase, or dCas9 form of the Cas9 nuclease from C. jejuni. In some embodiments, the Cas9 nuclease is a cleavase, nickase, or dCas9 form of the Cas9 nuclease from R. palustris. In some embodiments, the Cas9 nuclease is a cleavase, nickase, or dCas9 form of the Cas9 nuclease from R. rubrum. In some embodiments, the Cas9 nuclease is a cleavase, nickase, or dCas9 form of the Cas9 nuclease from A. naeslundii. In some embodiments, the Cas9 nuclease is a cleavase, nickase, or dCas9 form of the Cas9 nuclease from Francisella novicida.
[0139] In some embodiments, the Cas9 is an NmelCas9, an Nme2Cas9, an Nme3Cas9, or SpyCas9.
[0140] In some embodiments, the Cas nuclease is a Class V Cas nuclease. In some embodiments, the Cas nuclease is a Casl2. In some embodiments, the Casl2 is Lachnospiraceae bacterium Casl2a (LbCasl2a) or the Casl2 is Acidaminococcus sp. Casl2a (AsCasl2a). In some embodiments, the Cas nuclease is an Eubacterium siraeum Casl3d (EsCasl3d).
[0141] In some embodiments, the Cas nuclease is a cleavase, nickase, or dCas form of the Cpfl nuclease from Francisella novicida. In some embodiments, the Cas nuclease is a cleavase, nickase, or dCas form of the Cpfl nuclease from Acidaminococcus sp. In some embodiments, the Cas nuclease is a cleavase, nickase, or dCas form of the Cpfl nuclease from Lachnospiraceae bacterium ND2006. In further embodiments, the Cas nuclease is a cleavase, nickase, or dCas form of the Cpfl nuclease from Francisella tularensis, Lachnospiraceae bacterium, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium, Parcubacteria bacterium, Smithella, Acidaminococcus, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi, Leptospira inadai, Porphyromonas crevioricanis, Prevotella disiens, or Porphyromonas macacae. In certain embodiments, the Cas nuclease is a cleavase, nickase, or dCas form of a Cpfl nuclease from an Acidaminococcus or Lachnospiraceae. A nickase may be derived from (i.e. related to) a specific Cas nuclease in that the nickase is a form of the nuclease in which one of its two catalytic domains is inactivated, e.g., by mutating an active site residue essential for nucleolysis, such as DIO, H840, or N863 in SpyCas9, or D16, D587, H588, N611, or K614 in Nme2Cas9. In some embodiments, the nickase may comprise inactivating mutations at two or more amino acid residues in one of its two catalytic domains, such as D587 and H588; H588 and N611; H588 and K614; D587 and N611; D587 and K614; or D587, H588, and N611 in Nme2Cas9. A dCas may be derived from (i.e., related to) a specific Cas nuclease in that the dCas is a form of the nuclease in which both of its two catalytic domains are inactivated, e.g., by mutating active site residues essential for nucleolysis, such as DIO, H840, or N863 in SpyCas9 or D16, D587, H588, N611, or K614 in NmeCas9. In some embodiments, the dCas may comprise inactivating mutations at two or more amino acid residues in one or both of its two catalytic domains, such as D587 and H588; H588 and N611; H588 and K614; D587 and N611; D587 and K614; or D587, H588, and N611 in NmeCas9. One skilled in the art will be familiar with techniques for easily identifying corresponding residues in other Cas proteins, such as sequence alignment and structural alignment, which is discussed in detail below.
[0142] In other embodiments, the Cas nuclease may relate to a Type-I CRISPR / Cas system. In some embodiments, the Cas nuclease may be a component of the Cascade complex of a Type-I CRISPR / Cas system. In some embodiments, the Cas nuclease may be a Cas3 protein. In some embodiments, the Cas nuclease may be from a Type-Ill CRISPR / Cas system.
[0143] In some embodiments, the Cas9 nuclease is a cleavase or a nickase.
[0144] In some embodiments, the Cas9 nuclease is a nickase. In some embodiments, a Cas nuclease is a nickase form of a Cas nuclease or a modified Cas nuclease in which an endonucleolytic active site is inactivated, e.g., by one or more alterations (e.g., point mutations) in a catalytic domain. See, e.g., US Pat. No. 8,889,356 for discussion of Cas nickases and exemplary catalytic domain alterations.
[0145] In some embodiments, the Cas9 nuclease is a dCas9. In some embodiments, a Cas nuclease is a dCas9 form of a Cas nuclease or a modified Cas nuclease in which an endonucleolytic active site is inactivated, e.g., by one or more alterations (e.g., point mutations) in each catalytic domain.Streptococcus pyogenes Cas9 (SpyCas9)
[0146] Programmable DNA-binding proteins described herein encompass Streptococcus pyogenes Cas9 (SpyCas9) and modified versions and variants thereof.
[0147] Wild type S. pyogenes Cas9 has two catalytic domains: RuvC and HNH. The RuvC domain cleaves the non-target DNA strand, and the HNH domain cleaves the target strand of DNA. In some embodiments, a Cas nuclease may comprise an amino acid substitution in the RuvC or RuvC-like nuclease domain. Exemplary amino acid substitutions in the RuvC or RuvC-like nuclease domain include D10A (based on the S. pyogenes Cas9 protein). See, e.g., Zetsche et al. (2015) Cell Oct 22:163(3): 759-771. In some embodiments, the Cas nuclease may comprise an amino acid substitution in the HNH or HNH-like nuclease domain. Exemplary amino acid substitutions in the HNH or HNH-like nuclease domain include E762A, H840A, N863A, H983A, and D986A (based on the S. pyogenes Cas9 protein). See, e.g., Zetsche et al. (2015). Further exemplary amino acid substitutions include D917A, E1006A, and D1255A (based on the Francisella novicida U112 Cpfl (FnCpfl) sequence (UniProtKB - A0Q7Q2 (CPF1_FRATN)).
[0148] In some embodiments, a Cas nuclease such as a Cas9 nickase or dCas9 has an inactivated RuvC or HNH domain. In some embodiments, a nuclease is used having a RuvC domain with reduced activity. In some embodiments, a nuclease is used having an inactive RuvC domain. In some embodiments, a nuclease is used having an HNH domain with reduced activity. In some embodiments, a nuclease is used having an inactive HNH domain. See, W02019067910, the entire content of which is incorporated herein by reference.
[0149] In some embodiments, the SpyCas9 comprises an amino acid sequence at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 702, 704-705, and 707-708. In some embodiments, the SpyCas9 comprises an amino acidsequence of any one of SEQ ID NOs: 702, 704-705, and 707-708. In some embodiments, the amino acid sequence of the SpyCas9 lacks an N-terminal methionine, where the SpyCas9 is comprised within a fusion protein. In some embodiments, the amino acid sequence of the SpyCas9 lacks the N-terminal methionine of any one of SEQ ID NOs: 702, 704-705, and 707-708, for instance where the SpyCas9 is comprised within a fusion protein. In other embodiments, the SpyCas9 comprises any one of SEQ ID NOs: 702, 704-705, and 707-708, including or lacking the N-terminal methionine.
[0150] In some embodiments, the sequence encoding the SpyCas9 comprises a nucleotide sequence at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 699-701, 703, and 706. In some embodiments, the sequence encoding the SpyCas9 comprises a nucleotide sequence of any one of SEQ ID NOs: 699-701, 703, and 706. In some embodiments, the nucleotide sequence encoding the SpyCas9 lacks a start codon or stop codon, for instance, where the SpyCas9 is comprised within a fusion protein. In some embodiments, the nucleotide sequence encoding the SpyCas9 lacks the start codon or stop codon of any one of SEQ ID NOs: 699-701, 703, and 706, for instance, where the SpyCas9 is comprised within a fusion protein. In other embodiments, the nucleotide sequence of the SpyCas9 comprises any one of SEQ ID NOs: 699-701, 703, and 706, including or lacking the start codon or stop codon.
[0151] In some embodiments, any of the foregoing levels of identity is at least 95%, at least 98%, at least 99%, or 100%.Neisseria meningitidis Cas9 (NmeCas9)
[0152] Programmable DNA-binding proteins described herein encompass Neisseria meningitidis Cas9 (NmeCas9) and modified versions and variants thereof. See, e.g., WQ2023081689, the entire content of which is incorporated herein by reference. In some embodiments, the programmable DNA binding protein is an NmeCas9. In some embodiments, the NmeCas9 is Nme2Cas9. In some embodiments, the NmeCas9 is NmelCas9. In some embodiments, the NmeCas9 is Nme3Cas9.
[0153] In some embodiments, the NmeCas9 is a cleavase or a nickase.
[0154] In some embodiments, the NmeCas9 is a cleavase. Cleavases cut both strands of the target DNA, thus creating a double-strand break. In some embodiments, the compositions and methods comprise cleavases. In some embodiments, the compositions and methods comprise a cleavase programmable DNA-binding protein, such as a cleavase Cas, e.g., a cleavase Cas9, that induces a double-strand break.
[0155] In some embodiments, the NmeCas9 is a nickase. Modified versions having one catalytic domain, either RuvC or HNH, that is inactive are termed “nickases.” Nickases cut only one strand on the target DNA, thus creating a single-strand break. A single-strand break may also be known as a “nick.” In some embodiments, the compositions and methods comprise nickases. In some embodiments, the compositions and methods comprise a nickase programmable DNA-binding protein, such as a nickase Cas, e.g., a nickase Cas9, that induces a nick rather than a double strand break in the target DNA.
[0156] In some embodiments, the NmeCas9 is a dCas9. Modified versions in which both catalytic domains are inactive are termed “dCas.” dCas binds to the target DNA without creating a single- or double-strand break. In some embodiments, the compositions and methods comprise a dCas programmable DNA-binding protein, such as a dCas9, that does not produce a single- or double-strand break in the target DNA.
[0157] In some embodiments, the NmeCas9 nuclease may be modified to contain only one functional nuclease domain. For example, the programmable DNA-binding protein may be modified such that one of the nuclease domains is mutated or fully or partially deleted to reduce its nucleic acid cleavage activity.
[0158] In some embodiments, the NmeCas9 may be modified to contain no functional nuclease domain. For example, the programmable DNA-binding protein may be modified such that both of the nuclease domains are mutated or fully or partially deleted to reduce their nucleic acid cleavage activity.
[0159] In some embodiments, a NmeCas9 nickase is used having a RuvC domain with reduced activity. In some embodiments, a NmeCas9 nickase is used having an inactive RuvC domain. In some embodiments, a NmeCas9 nickase is used having an HNH domain with reduced activity. In some embodiments, a NmeCas9 nickase is used having an inactive HNH domain.
[0160] In some embodiments, a NmeCas9 dCas9 is used having an inactive RuvC domain and an inactive HNH domain.
[0161] In some embodiments, a conserved amino acid within a NmeCas9 nuclease domain is substituted to reduce or alter nuclease activity. Wild type Cas9 has two nuclease domains: RuvC and HNH. The RuvC domain cleaves the non-target DNA strand, and the HNH domain cleaves the target strand of DNA. In some embodiments, the Cas9 nuclease comprises more than one RuvC domain or more than one HNH domain. In some embodiments, the Cas9 nuclease is a wild type Cas9. In some embodiments, the Cas9 is capable of inducing a double strand break in target DNA. In certain embodiments, the Casnuclease may cleave dsDNA, it may cleave one strand of dsDNA, or it may not have DNA cleavase or nickase activity. In some embodiments, a NmeCas9 may comprise an amino acid substitution in the RuvC or RuvC-like nuclease domain. Exemplary amino acid substitutions in the RuvC or RuvC-like nuclease domain include H588A (based on the N. meningitidis Cas9 protein). In some embodiments, the Cas protein may comprise an amino acid substitution in the HNH or HNH-like nuclease domain. Exemplary amino acid substitutions in the HNH or HNH-like nuclease domain include D16A (based on the NmeCas9 protein).
[0162] In some embodiments, chimeric Cas proteins are used, where one domain or region of the protein is replaced by a portion of a different protein. In some embodiments, an NmeCas9 may include a PAM interacting domain from another Cas9 protein. In some embodiments, a NmeCas9 nuclease domain may be replaced with a domain from a different nuclease such as Fokl. In some embodiments, a NmeCas9 protein may be a modified NmeCas9 nuclease.
[0163] In some embodiments, the nuclease may be modified to induce a point mutation or base change, e.g., a deamination.
[0164] In some embodiments, the Cas protein comprises a fusion protein comprising a Cas nuclease (e.g., NmeCas9), which is a nickase or is catalytically inactive, linked to a heterologous functional domain. In some embodiments, the Cas protein comprises a fusion protein comprising a catalytically inactive Cas nuclease (e.g., NmeCas9) linked to a heterologous functional domain (see, e.g., WO2014152432). In some embodiments, the catalytically inactive Cas9 is from the N. meningitidis Cas9. In some embodiments, the catalytically inactive Cas comprises mutations that inactivate the Cas.
[0165] In some embodiments, the heterologous functional domain is a domain that modifies gene expression, histones, or DNA. In some embodiments, the heterologous functional domain is a transcriptional activation domain or a transcriptional repressor domain. In some embodiments, the nuclease is a catalytically inactive Cas nuclease, such as dCas9.
[0166] In some embodiments, the heterologous functional domain is a deaminase, such as a cytidine deaminase or an adenosine deaminase. In certain embodiments, the heterologous functional domain is a C to T base converter (cytidine deaminase), such as an apolipoprotein B mRNA editing enzyme (APOBEC) deaminase. A heterologous functional domain such as a deaminase may be part of a fusion protein with a Cas nuclease having nickase activity or a Cas nuclease that is catalytically inactive discussed further below.
[0167] In some embodiments, the NmeCas9 has double stranded endonuclease activity.
[0168] In some embodiments, the NmeCas9 has nickase activity.
[0169] In some embodiments, the NmeCas9 comprises a dCas9 DNA binding domain.
[0170] In some embodiments, the NmeCas9 comprises an amino acid sequence at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 712-713, 715-716, and 719-721. In some embodiments, the NmeCas9 comprises an amino acid sequence of any one of SEQ ID NOs: 712-713, 715-716, and 719-721. In some embodiments, the amino acid sequence of the NmeCas9 lacks an N-terminal methionine, where the NmeCas9 is comprised within a fusion protein. In some embodiments, the amino acid sequence of the NmeCas9 lacks the N-terminal methionine of any one of SEQ ID NOs: 712-713, 715-716, and 719-721, for instance where the NmeCas9 is comprised within a fusion protein. In other embodiments, the NmeCas9 comprises any one of SEQ ID NOs: 712-713, 715-716, and 719-721, including or lacking the N-terminal methionine.
[0171] In some embodiments, the sequence encoding the NmeCas9 comprises a nucleotide sequence at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 711, 714, 717-718, and 733-736. In some embodiments, the sequence encoding the NmeCas9 comprises a nucleotide sequence of any one of SEQ ID NOs: 711, 714, 717-718, and 733-736. In some embodiments, the nucleotide sequence encoding the NmeCas9 lacks a start codon or stop codon, for instance, where the NmeCas9 is comprised within a fusion protein. In some embodiments, the nucleotide sequence encoding the NmeCas9 lacks the start codon or stop codon of any one of SEQ ID NOs: 711, 714, 717-718, and 733-736, for instance, where the NmeCas9 is comprised within a fusion protein. In other embodiments, the nucleotide sequence of the NmeCas9 comprises any one of SEQ ID NOs: 711, 714, 717-718, and 733-736, including or lacking the start codon or stop codon.
[0172] In some embodiments, any of the foregoing levels of identity is at least 95%, at least 98%, at least 99%, or 100%.
[0173] In some embodiments, the NmeCas9 is an Nme2Cas9. In some embodiments, the Nme2Cas9 comprises an amino acid sequence at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 712, 715-716, and 719-720. In some embodiments, the Nme2Cas9 comprises an amino acid sequence of any one of SEQ ID NOs: 712, 715-716, and 719-720. In some embodiments, the amino acid sequence of the Nme2Cas9 lacks an N-terminal methionine, where the Nme2Cas9 is comprised within a fusion protein. In some embodiments, the amino acid sequence of the Nme2Cas9 lacks the N-terminal methionine of any one of SEQ ID NOs: 712, 715-716, and 719-720, for instance where the Nme2Cas9 is comprised within a fusion protein. In other embodiments, theNme2Cas9 comprises any one of SEQ ID NOs: 712, 715-716, and 719-720, including or lacking the N-terminal methionine.
[0174] In some embodiments, the sequence encoding the Nme2Cas9 comprises a nucleotide sequence at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 711, 714, 717-718, and 734-735. In some embodiments, the sequence encoding the Nme2Cas9 comprises a nucleotide sequence of any one of SEQ ID NOs: 711, 714, 717-718, and 734-735. In some embodiments, the nucleotide sequence encoding the Nme2Cas9 lacks a start codon or stop codon, for instance, where the Nme2Cas9 is comprised within a fusion protein. In some embodiments, the nucleotide sequence encoding the Nme2Cas9 lacks the start codon or stop codon of any one of SEQ ID NOs: 711, 714, 717-718, and 734-735, for instance, where the Nme2Cas9 is comprised within a fusion protein. In other embodiments, the nucleotide sequence of the Nme2Cas9 comprises any one of SEQ ID NOs: 711, 714, 717-718, and 734-735, including or lacking the start codon or stop codon.
[0175] In some embodiments, any of the foregoing levels of identity is at least 95%, at least 98%, at least 99%, or 100%.Deaminases
[0176] In some embodiments, the HMGB1 polypeptide, system, or fusion protein described herein further comprises a deaminase. In some embodiments, the deaminase is an adenosine deaminase. In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, wherein the deaminase is a cytidine deaminase, the cytidine deaminase is an APOBEC3A deaminase (A3A).
[0177] In some embodiments, the deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 728 and 729.
[0178] In some embodiments, the amino acid sequence of the deaminase lacks an N-terminal methionine, for instance, where the deaminase is comprised within a fusion protein. In some embodiments, the amino acid sequence of the deaminase lacks the N-terminal methionine of any one of SEQ ID NO: 728 and 729, for instance, where the deaminase is comprised within a fusion protein. In other embodiments, the amino acid sequence of the deaminase comprises any one of SEQ ID NOs: 728 and 729, including or lacking the N-terminal methionine.
[0179] In some embodiments, the deaminase comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 737 and 738.
[0180] In some embodiments, the nucleotide sequence encoding the deaminase lacks a start codon or stop codon, for instance, where the deaminase is comprised within a fusion protein. In some embodiments, the nucleotide sequence encoding the deaminase lacks the start codon or stop codon of any one of SEQ ID NOs: 737 and 738, for instance, where the deaminase is comprised within a fusion protein. In other embodiments, the nucleotide sequence of the deaminase comprises any one of SEQ ID NOs: 737 and 738, including or lacking the start codon or stop codon.Cytidine Deaminases
[0181] In some embodiments, the HMGB1 polypeptide, system, or fusion protein described herein further comprises a cytidine deaminase. Cytidine deaminases encompass enzymes in the cytidine deaminase superfamily, and in particular, enzymes of the APOBEC family (APOBEC 1, APOBEC2, APOBEC4, and APOBEC3 subgroups of enzymes), activation-induced cytidine deaminase (AID or AICDA) and CMP deaminases (see, e.g., Conticello et al., Mol. Biol. Evol. 22:367-77, 2005; Conticello, Genome Biol. 9:229, 2008; Muramatsu et al., J. Biol. Chem. 274: 18470-6, 1999); and Carrington et al., Cells 9:1690 (2020)).
[0182] In some embodiments, the cytidine deaminase disclosed herein is an enzyme of APOBEC family. In some embodiments, the cytidine deaminase disclosed herein is an enzyme of APOBEC 1, APOBEC2, APOBEC4, and APOBEC3 subgroups. In some embodiments, the cytidine deaminase disclosed herein is an enzyme of APOBEC3 subgroup. In some embodiments, the cytidine deaminase disclosed herein is an APOBEC3A deaminase (A3A).
[0183] In some embodiments, the cytidine deaminase is a cytidine deaminase comprising an amino acid sequence having at least 80%, 85% 87%, 90%, 95%, 98%, 99%, or 100% identity to any one of SEQ ID NOs: 728 and 729.
[0184] In some embodiments, the amino acid sequence of the cytidine deaminase lacks an N-terminal methionine, where the cytidine deaminase is comprised within a fusion protein. In some embodiments, the amino acid sequence of the cytidine deaminase lacks the N-terminal methionine of any one of SEQ ID NOs: 728 and 729, for instance, where the cytidine deaminase is comprised within a fusion protein. In other embodiments, the amino acidsequence of the cytidine deaminase comprises any one of SEQ ID NOs: 728 and 729, including or lacking the N-terminal methionine.
[0185] In some embodiments, the cytidine deaminase comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 737 and 738.
[0186] In some embodiments, the nucleotide sequence encoding the cytidine deaminase lacks a start codon or stop codon, for instance, where the cytidine deaminase is comprised within a fusion protein. In some embodiments, the nucleotide sequence encoding the cytidine deaminase lacks the start codon or stop codon of any one of SEQ ID NOs: 737 and 738, for instance, where the cytidine deaminase is comprised within a fusion protein. In other embodiments, the nucleotide sequence of the cytidine deaminase comprises any one of SEQ ID NOs: 737 and 738, including or lacking the start codon or stop codon.APOBEC3A Deaminases
[0187] In some embodiments, the HMGB1 polypeptide, system, or fusion protein described herein further comprises an APOBEC3A deaminase (A3A). In some embodiments, an APOBEC3A deaminase (A3A) disclosed herein is a human A3A. In some embodiments, the A3A is a wild-type A3A.
[0188] In some embodiment, the A3A is an A3A variant. A3 A variants share homology to wild-type A3A, or a fragment thereof. In some embodiments, an A3A variant has at least about 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to a wild type A3A. In some embodiments, the A3A variant comprises a fragment of an A3A.
[0189] In some embodiments, an A3A variant is a protein having a sequence that differs from a wild-type A3A protein by one or several mutations, such as substitutions, deletions, insertions, one or several single point substitutions. In some embodiments, a shortened A3A sequence could be used, e.g., by deleting N-terminal, C-terminal, or internal amino acids. In some embodiments, a shortened A3A sequence is used where one to four amino acids at the C-terminus of the sequence is deleted. In some embodiments, an APOBEC3A (such as a human APOBEC3A) has a wild-type amino acid position 57 (as numbered in the wild-type sequence). In some embodiments, an APOBEC3A (such as a human APOBEC3A) has an asparagine at amino acid position 57 (as numbered in the wildtype sequence).
[0190] In some embodiments, the wild-type A3A is a human A3A (UniPROT accession ID: P31941, SEQ ID NO: 728).
[0191] In some embodiments, the A3A disclosed herein comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 728. In some embodiments, the level of identity is at least 85%, 87%, 90%, 95%, 98%, or 99%; or 100% identical.
[0192] In some embodiments, the amino acid sequence of the A3A lacks an N-terminal methionine, where the A3A is comprised within a fusion protein. In some embodiments, the amino acid sequence of the A3A lacks the N-terminal methionine of SEQ ID NOs: 728, for instance, where the A3A is comprised within a fusion protein. In other embodiments, the amino acid sequence of the A3A comprises SEQ ID NO: 728, including or lacking the N-terminal methionine.
[0193] In some embodiments, the A3A comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to SEQ ID NO: 737.
[0194] In some embodiments, the nucleotide sequence encoding the A3A lacks a start codon or stop codon, for instance, where the A3A is comprised within a fusion protein. In some embodiments, the nucleotide sequence encoding the A3A lacks the start codon or stop codon of SEQ ID NO: 737, for instance, where the A3A is comprised within a fusion protein. In other embodiments, the nucleotide sequence of the A3A comprises SEQ ID NO: 737, including or lacking the start codon or stop codon.Zinc Finger Nucleases (ZFNs)
[0195] In still other embodiments, the programmable DNA-binding protein comprises a pair of zinc finger nucleases (ZFNs). A ZFN comprises a DNA-binding zinc finger and a nuclease domain. The zinc finger region can comprise from about two to seven zinc fingers, for example, about four to six zinc fingers, wherein each zinc finger binds three consecutive base pairs. The zinc finger region can be engineered to recognize and bind to any DNA sequence. Zinc finger design tools or algorithms are available on the internet or from commercial sources. The zinc fingers can be linked together using suitable linker sequences.
[0196] A ZFN also comprises a nuclease domain, which can be obtained from any endonuclease or exonuclease. Non-limiting examples of endonucleases from which a nuclease domain can be derived include, but are not limited to, restriction endonucleases and homing endonucleases. In some embodiments, the nuclease domain can be derived from a type II-S restriction endonuclease. Type II-S endonucleases cleave DNA at sites that are typically several base pairs away from the recognition / binding site and, as such, have separable binding and cleavage domains. These enzymes generally are monomers thattransiently associate to form dimers to cleave each strand of DNA at staggered locations. Non-limiting examples of suitable type II-S endonucleases include Bfil, Bpml, Bsal, Bsgl, BsmBI, BsmI, BspMI, Fokl, MboII, and Sapl. In some embodiments, the nuclease domain can be a Fokl nuclease domain or a derivative thereof. The type II-S nuclease domain can be modified to facilitate dimerization of two different nuclease domains. For example, the cleavage domain of Fokl can be modified by mutating certain amino acid residues. By way of non-limiting example, amino acid residues at positions 446, 447, 479, 483, 484, 486, 487, 490, 491, 496, 498, 499, 500, 531, 534, 537, and 538 of Fokl nuclease domains are targets for modification. In specific embodiments, the Fokl nuclease domain can comprise a first Fokl half- domain comprising Q486E, I499L, or N496D mutations, and a second Fokl halfdomain comprising E490K, I538K, or H537R mutations. In some embodiments, the ZFN has double-stranded cleavage activity. In other embodiments, the ZFN has nickase activity (i.e., one of the nuclease domains has been inactivated).Transcription Activator-Like Effectors (TALEs) or TALE Nucleases (TALENs)
[0197] In some embodiments, the programmable DNA-binding protein can be a transcription activator-like effector (TALE) or a TALE nuclease (TALEN).
[0198] In some embodiments, the programmable DNA-binding protein comprises a TALE. TALEs are proteins secreted by the plant pathogen Xanthomonas to alter transcription of genes in host plant cells. TALE repeat arrays can be engineered via modular protein design to target any DNA sequence of interest.
[0199] In some embodiments, the programmable DNA-binding protein comprises a TALEN. TALENs comprise a DNA-binding domain composed of highly conserved repeats derived from TALEs linked to a nuclease domain. The nuclease domain of TALENs can be any nuclease domain as described above in the subsection describing ZFNs. In specific embodiments, the nuclease domain is derived from Fokl (Sanjana et al., 2012, Nat Protoc, 7(1): 171-192). The TALEN can have cleavase or nickase activity.Meganucleases and Rare-Cutting Endonucleases
[0200] In some embodiments, the programmable DNA-binding protein comprises a meganuclease or derivative thereof. Meganucleases are endodeoxyribonucleases characterized by long recognition sequences, i.e., the recognition sequence generally ranges from about 12 base pairs to about 45 base pairs. As a consequence of this requirement, the recognitionsequence generally occurs only once in any given genome. Among meganucleases, the family of homing endonucleases named LAGLID ADG has become a valuable tool for the study of genomes and genome engineering. In some embodiments, the meganuclease can be I-Scel, I-TevI, or variants thereof. A meganuclease can be targeted to a specific chromosomal sequence by modifying its recognition sequence using techniques well known to those skilled in the art.
[0201] In alternative embodiments, the programmable DNA-binding protein comprises a rare-culling endonuclease or derivative thereof. Rare-cutting endonucleases are site-specific endonucleases whose recognition sequence occurs rarely in a genome, preferably only once in a genome. The rare-cutting endonuclease may recognize a 7-nucleotide sequence, an 8-nucleotide sequence, or longer recognition sequence. Non-limiting examples of rare-cutting endonucleases include Notl, Asci, Pad, AsiSI, Sbfl, and Fsel.Chimeric Proteins Comprising a Programmable DNA-Binding Domain Linked to a Nuclease Domain
[0202] In some embodiments, the programmable DNA-binding protein comprises a chimeric protein comprising a programmable DNA-binding domain linked to a nuclease domain. The nuclease domain can be any of those described above in the subsection describing ZFNs (e.g., the nuclease domain can be a Fokl nuclease domain), a nuclease domain derived from a CRISPR nuclease (e.g., RuvC or HNH nuclease domains of Cas9), or a nuclease domain derived from a meganuclease or rare-cutting endonuclease.
[0203] The programmable DNA-binding domain of the chimeric protein can be any programmable DNA-binding protein domain, such as, e.g., a zinc finger protein or a TALE. Alternatively, the programmable DNA-binding domain can be a catalytically inactive (dead) CRISPR protein that was modified by deletion or mutation to lack all nuclease activity. For example, the catalytically inactive CRISPR protein can be a catalytically inactive (dead) SpyCas9 (dCas9) in which the RuvC domain comprises a D10A, D8A, E762A, or D986A mutation and the HNH domain comprises a H840A, H559A, N854A, N865A, or N863A mutation. Alternatively, the catalytically inactive CRISPR protein can be a catalytically inactive (dead) NmeCas9, e.g., Nme2Cas9 in which the RuvC domain comprises a D16A mutation and the HNH domain comprises a H588A mutation. Alternatively, the catalytically inactive CRISPR protein can be a catalytically inactive (dead) Cpfl protein comprising comparable mutations in the nuclease domains. In still other embodiments, the programmable DNA-binding domain can be a catalytically inactive meganuclease in which nuclease activitywas eliminated by mutation or deletion, e.g., the catalytically inactive meganuclease can comprise a C-terminal truncation.Intervening Peptide Sequences
[0204] In some embodiments, the fusion protein described herein further comprises an intervening peptide sequence between the HMGB1 polypeptide and the programmable DNA- binding protein. In some embodiments, the programmable DNA-binding protein is attached to the HMGB 1 polypeptide by an intervening peptide sequence of a specific length as defined by a number of amino acid residues. In some embodiments, the Class II Cas nuclease is attached to the HMGB1 polypeptide by an intervening peptide sequence. In some embodiments, the nucleic acid encoding the HMGB1 polypeptide and the programmable DNA-binding protein further comprises a sequence encoding the intervening peptide sequence.
[0205] In some embodiments, the intervening peptide sequence is at least 10 amino acid residues in length, optionally at least 12 amino acid residues in length. For example, in some embodiments, the intervening peptide sequence may be at least 15, 20, 25, 30, 35, 40, 45, 50, or more amino acid residues in length. In some embodiments, the intervening peptide sequence is 10-50 amino acid residues in length, optionally 12-42 amino acid residues in length, further optionally 12, 15, 20, 33, 41, or 42 amino acid residues in length. In some embodiments, the intervening peptide sequence is at least 10, 15, 20, 25, 30, 35, 40, 45, 50, or 55 amino acid residues in length. In some embodiments, the intervening peptide sequence is at least 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, or 54 amino acid residues in length. In some embodiments, the intervening peptide sequence is 10-15, 10-20, 10-25, 10-30, 10-35, 10-40, 10-45, 10-50, 10-55, 15-20, 15-25, 15-30, 15-35, 15-40, 15-45, 15-50, 15-55, 20-25, 20-30, 20-35, 20-40, 20-45, 20-50, 20-55, 25-30, 25-35, 25-40, 25-45, 25-50, 25-55, 30-35, 30-40, 30-45, 30-50, 30-55, 35-40, 35-45, 35-50, 35-55, 40-45, 40-50, 40-55, 45-50, 45-55, or 50-55 amino acid residues in length. In some embodiments, the intervening peptide sequence is 12-14, 12-16, 12-18, 12-20, 12-22, 12-24, 12-26, 12-28, 12-30, 12-32, 12-34, 12-36, 12-38, 12-40, 12-42, 14-16, 14-18, 14-20, 14-22, 14-24, 14-26, 14-28, 14-30, 14-32, 14-34, 14-36, 14-38, 14-40, 14-42, 16-18, 16-20, 16-22, 16-24, 16-26, 16-28, 16-30, 16-32, 16-34, 16-36, 16-38, 16-40, 16-42, 18-20, 18-22, 18-24, 18-26, 18-28, 18-30, 18-32, 18-34, 18-36, 18-38, 18-40, 18-42, 20-22, 20-24, 20-26, 20-28, 20-30, 20-32, 20-34, 20-36, 20-38, 20-40, 20-42, 22-24, 22-26, 22-28, 22-30, 22-32, 22-34,22-36, 22-38, 22-40, 22-42, 24-26, 24-28, 24-30, 24-32, 24-34, 24-36, 24-38, 24-40, 24-42, 26-28, 26-30, 26-32, 26-34, 26-36, 26-38, 26-40, 26-42, 28-30, 28-32, 28-34, 28-36, 28-40, 28-42, 30-32, 30-34, 30-36, 30-38, 30-40, 30-42, 32-34, 32-36, 32-38, 32-40, 32-42, 34-36, 34-38, 34-40, 34-42, 36-38, 36-40, 36-42, 38-40, 38-42, 40-42 amino acid residues in length. In some embodiments, the intervening peptide sequence is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, or 55 amino acid residues in length.
[0206] In some embodiments, the intervening peptide sequence comprises a linker or a nuclear localization signal (NLS). In some embodiments, the intervening peptide sequence comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435. In some embodiments, the intervening peptide sequence comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 366-369 and 370-384.Peptide linkers
[0207] Other types of intervening peptide sequences may be used herein. In some embodiments, the intervening peptide sequence is a linker. For example, the intervening peptide sequence may be the 16 residue “XTEN” linker, or a variant thereof (See, e.g., Schellenberger et al. A recombinant polypeptide extends the in vivo half-life of peptides and proteins in a tunable manner. Nat. Biotechnol. 27, 1186-1190 (2009)). In some embodiments, the XTEN linker comprises a sequence that is any one of SGSETPGTSESATPES (SEQ ID NO: 301), SGSETPGTSESA (SEQ ID NO: 302), or SGSETPGTSESATPEGGSGGS (SEQ ID NO: 303). In some embodiments, the XTEN linker consists of the sequence SGSETPGTSESATPES (SEQ ID NO: 301), SGSETPGTSESA (SEQ ID NO: 302), or SGSETPGTSESATPEGGSGGS (SEQ ID NO: 303).
[0208] In some embodiments, the linker comprises a (GGGGS)n(e.g., SEQ ID NOs: 305, 309-311, 314-318, 320-331, or 333-359), a (G)n, an (EAAAK)n(e.g„ SEQ ID NOs: 306, 310-312, 315-318, 320-331, or 334-360), a (GGS)n, an SGSETPGTSESATPES (SEQ ID NO: 301) motif see, e.g., Guilinger J P, Thompson D B, Liu D R. Fusion of catalytically inactive Cas9 to Fokl nuclease improves the specificity of genome modification. Nat.Biotechnol. 2014; 32(6): 577-82; the entire contents are incorporated herein by reference), or an (XP)nmotif, or a combination of any of these, wherein n is independently an integer between 1 and 30. See, W02015089406, e.g., paragraph
[0012] , the entire content of which is incorporated herein by reference. As used herein, the term GS linker refers to a linkersequence that is rich in glycine and serine, e.g., at least 50% glycine or serine amino acid residues. Exemplary GS linkers include, but are not limited to, GGSGG (SEQ ID NO: 425), GAPESATESGGTSTESEGSAGTSTESEGSAGSAGSTSGSSS (SEQ ID NO: 426), GGGS (SEQ ID NO: 427), LEGGGGS (SEQ ID NO: 428), SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 429), GGSGGSGGSGGSGGSGGSGG (SEQ ID NO: 430), GGSGGSGGSGGSGGS (SEQ ID NO: 431), GGSGGSGGSGGS (SEQ ID NO: 432), GGSGGG (SEQ ID NO: 433), GSGGGS (SEQ ID NO: 434), and SGGSSGGS (SEQ ID NO: 435).
[0209] In some embodiments, the linker is 12 amino acid residues in length (or referred as “12 AA linker”). Non-limiting examples of the 12 AA linker is the sequence of SEQ ID NO: 432. In some embodiments, the linker is 15 amino acid residues in length (referred elsewhere herein as “15 AA linker,” “Linkerl5,” or “L15”). Non-limiting examples of the 15 AA linker is the sequence of SEQ ID NO: 431. In some embodiments, the linker is 20 amino acid residues in length (referred elsewhere herein as “20 AA linker,” “Linker20,” “GS20 linker,” or “L20”). Non-limiting examples of the 20 AA linker is the sequence of SEQ ID NO: 430. In some embodiments, the linker is 33 amino acid residues in length (referred elsewhere herein as “33 AA linker,” “Linker33,” or “L33”). Non-limiting examples of the 33 AA linker is the sequence of SEQ ID NO: 429. In some embodiments, the linker is 41 amino acid residues in length (referred elsewhere herein as “41 AA linker,” “Linker41,” or “L41”). Non-limiting examples of the 41 AA linker is the sequence of SEQ ID NO: 426.
[0210] In some embodiments, the linker comprises one or more sequences selected from SEQ ID NO: 301, SEQ ID NO: 302, SEQ ID NO: 303, SEQ ID NO: 361, SEQ ID NO: 362, SEQ ID NO: 363. SEQ ID NO: 364 and SEQ ID NO: 365. In some embodiments, the peptide linker comprises a sequence of SEQ ID NO: 361 (alternatively referred elsewhere herein as “GH5 linker”). In some embodiments, the peptide linker comprises three consecutive copies of SEQ ID NO: 361 (referred to herein as “3xGH5 linker”).Heterologous Nuclear Localization Signals (NLS)
[0211] In some embodiments, a heterologous functional domain may facilitate transport of the HMGB 1 polypeptide, system, or fusion protein disclosed herein into the nucleus of a cell. For example, the heterologous functional domain may be a heterologous nuclear localization signal (NLS). In some embodiments, the HMGB1 polypeptide, system,or fusion protein comprises a heterologous nuclear localization signal (NLS), i.e., an NLS sequence not present in the HMGB1 protein.
[0212] In some embodiments, the HMGB 1 polypepride, system, or fusion protein disclosed herein comprises at least one heterologous nuclear localization signal (NLS). In some embodiments, the HMGB1 polypeptide, system, or fusion protein may comprise 1, 2, 3, 4 or 5 heterologous NLS(s). In some embodiments, the HMGB1 polypeptide, system, or fusion protein may comprise two heterologous NLSs. In some embodiments, the HMGB1 polypeptide, system, or fusion protein may comprise three heterologous NLSs. In some embodiments, the HMGB1 polypeptide, system, or fusion protein may comprise four heterologous NLSs.
[0213] Where one heterologous NLS is used in a polypeptide or fusion protein, the heterologous NLS may be present at the N-terminus or the C-terminus of the HMGB1 polypeptide or fusion protein. In some embodiments, the HMGB1 polypeptide or fusion protein disclosed herein may comprise C-terminally at least one heterologous NLS. A heterologous NLS may also be inserted within the HMGB 1 polypeptide or fusion protein. In other embodiments, the HMGB1 polypeptide, system, or fusion protein may comprise more than one heterologous NLS.
[0214] In some embodiments, the HMGB 1 polypeptide, system, or fusion protein comprises 2, 3, 4, or 5 heterologous NLSs. In certain circumstances, the heterologous NLSs may be the same (e.g., two SV40 NLSs). In certain circumstances, the heterologous NLSs may be different, e.g., at least one of one sequence and at least one of another sequence, or different (e.g., an SV40 NLS and a nucleoplasmin NLS). In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises two SV40 NLS sequences at the C-terminus. In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises two heterologous NLSs. In some embodiments, the first heterologous NLS is located at the N-terminus and the second heterologous NLS is present at the C-terminus. In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises an SV40 NLS and a nucleoplasmin NLS. In some embodiments, the SV40 NLS is located N-terminal to the nucleoplasmin NLS. In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises 3 heterologous NLSs.
[0215] In some embodiments, the heterologous NLS may be a monopartite sequence, such as, e.g., the SV40 NLS, for example, PKKKRKVE (SEQ ID NO: 366), KKKRKVE (SEQ ID NO: 367), PKKKRKV (SEQ ID NO: 371) or PKKKRRV (SEQ ID NO: 383). In some embodiments, the heterologous NLS may be a bipartite sequence, such as the NLS ofnucleoplasmin, KRPAATKKAGQAKKKK (SEQ ID NO: 384). In a specific embodiment, a single PKKKRKV (SEQ ID NO: 371). One or more linkers or intervening peptide sequences are optionally included at the fusion site (e.g., between fusion protein disclosed herein and heterologous NLS).
[0216] In some embodiments, one or more heterologous NLS(s) according to any of the foregoing embodiments are present in the HMGB 1 polypeptide, system, or fusion protein in combination with one or more additional heterologous functional domains, such as any of the heterologous functional domains described below.
[0217] In some embodiments, the HMGB 1 polypeptide, system, or fusion protein comprises a heterologous nuclear localization signal (NLS) and the heterologous NLS is present at the C-terminus of the HMGB1 polypeptide, system, or fusion protein. In some embodiments, the HMGB1 polypeptide or fusion protein comprises a heterologous nuclear localization signal (NLS) and the heterologous NLS is present at the N-terminus of the HMGB1 polypeptide or fusion protein. In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises a heterologous nuclear localization signal (NLS) and the heterologous NLS is present at both the N-terminus and C-terminus of the HMGB1 polypeptide, system, or fusion protein. In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises a heterologous nuclear localization signal (NLS) and the heterologous NLS is present between the HMGB 1 polypeptide or programmable DNA-binding protein and the linker sequence. In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises a heterologous nuclear localization signal (NLS) and the heterologous NLS is present within the linker sequence. In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises a heterologous nuclear localization signal (NLS) and the heterologous NLS is present between two linker sequences. In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises a heterologous nuclear localization signal (NLS) and the heterologous NLS is present between the C-terminus of the HMGB1 polypeptide disclosed herein and the linker sequence. In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises a nuclear localization signal (NLS) and the heterologous NLS is present between the C-terminus of the programmable DNA-binding protein disclosed herein and the linker sequence. In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises a nuclear localization signal (NLS) and the heterologous NLS is present between the N-terminus of the HMGB1 polypeptide disclosed herein and the linker sequence. In some embodiments, the HMGB 1 polypeptide, system, or fusion protein comprises a nuclear localization signal (NLS)and the heterologous NLS is present between N-terminus of the programmable DNA-binding protein disclosed herein and the linker sequence.
[0218] In some embodiments, the at least one heterologous NLS is an SV40 NLS or a nucleoplasmin NLS.
[0219] In some embodiments, the at least one heterologous NLS comprises a sequence having at least 80%, 85%, 90%, 95%, or 98% identity to any one of SEQ ID NOs: 366-369 and 371-384. In some embodiments, the HMGB1 polypeptide, system, or fusion protein comprises one, two, or three nuclear localization signals (NLSs) independently selected from SEQ ID NOs: 366-369 and 371-384.
[0220] In some embodiments, the at least one heterologous NLS is encoded by a nucleic acid sequence having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 370 and 385-397.
[0221] In some embodiments, the HMGB 1 polypeptide, system, or fusion protein may not comprise a heterologous NLS. For example, in some embodiments, the HMGB1 polypeptide, system, or fusion protein may comprise a cryptic NLS that allows for nuclear localization.Exemplary fusion proteins comprising an HMGB1 Polypeptide and a Programmable DNA-Binding Protein
[0222] In some embodiments, the present disclosure provides for a system or fusion protein comprising a programmable DNA-binding protein and an HMGB 1 polypeptide comprising an HMGB1 Box B domain and lacking an acid tail domain. In some embodiments, the HMGB1 polypeptide comprises a cryptic NLS. In some embodiments, the cryptic NLS comprises the sequence of EKSKKKK (SEQ ID NO: 5). In some embodiments, the HMGB1 polypeptide comprises an HMGB1 Box A domain and an HMGB1 Box B domain, and lacks an acidic tail domain. In some embodiments, the HMGB1 polypeptide comprises the sequence of SEQ ID NO: 7. In some embodiments, the programmable DNA-binding protein is located N-terminal to the HMGB1 polypeptide. In some embodiments, the programmable DNA-binding protein is located C-terminal to the HMGB 1 polypeptide. In some embodiments, the system or fusion protein comprises at least one heterologous nuclear localization signal (NLS). See, e.g., FIG. 22, which shows certain non-limiting, exemplary HMGB 1 -programmable DNA-binding protein fusion proteins.
[0223] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a programmable DNA binding domain; a heterologous NLS; a linker; and anHMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a programmable DNA-binding protein; a linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a programmable DNA-binding protein; a first linker; a heterologous NLS; a second linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a first heterologous NLS; a programmable DNA-binding protein; a first linker; a second heterologous NLS; a second linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a heterologous NLS; programmable DNA-binding protein; a linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a first heterologous NLS; a second heterologous NLS; a programmable DNA-binding protein; a linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an HMGB1 polypeptide; a first heterologous NLS; a second heterologous NLS; and a programmable DNA-binding protein. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a first heterologous NLS; a second heterologous NLS; a deaminase; a first linker; a DNA-binding domain; a second linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a deaminase; a first linker; a DNA-binding domain; a heterologous NLS; a second linker; and an HMGB1 binding domain. In some embodiments, wherein the fusion protein comprises a first and second heterologous NLS, the first and second heterologous NLS are the same. In some embodiments, wherein the fusion protein comprises a first and second heterologous NLS, the first and second heterologous NLS are different. In some embodiments, wherein the fusion protein comprises a first and second linker, the first and second linkers are the same. In some embodiments, wherein the fusion protein comprises a first and second linker, the first and second linkers are different. In any one of the aforementioned fusion proteins, the programmable DNA binding domain may comprise an amino acid sequence at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704, 705, 707, 708, 712, 713, 715, 716, and 719-721. In any one of the aforementioned fusion proteins, the heterologous NLS may comprise an amino acid sequence having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384. In any one of the aforementioned fusion proteins, the linker may comprise an amino acid sequence having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435. In some any one of the aforementioned fusion proteins, the HMGB1 polypeptide may comprise an amino acid sequence having at least 90%, 95%, 99%, or 100%identity to any one of SEQ ID NOs: 2-5, 7-10, 21, and 22. In any one of the aforementioned fusion proteins, the deaminase may comprise an amino acid sequence having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 728 and 729; In some embodiments where the fusion protein comprises a first and second heterologous NLS, the first and second heterologous NLS are the same. In some embodiments where the fusion protein comprises a first and second heterologous NLS, the first and second heterologous NLS are different. In some embodiments where the fusion protein comprises a first and second linker, the first and second linkers are the same. In some embodiments where the fusion protein comprises a first and second linker, the first and second linkers are different.
[0224] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a programmable DNA-binding protein; a 41 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a programmable DNA-binding protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a programmable DNA-binding protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a programmable DNA-binding protein; a GS20 linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a programmable DNA-binding protein; a GS15 linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a programmable DNA-binding protein; a GS12 linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a programmable DNA-binding protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a nucleoplasmin NLS; a programmable DNA-binding protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a nucleoplasmin NLS; a programmable DNA-binding protein; a first GS linker; an SV40 NLS; a second GS linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a programmable DNA-binding protein; a GS linker; a nucleoplasmin NLS; a GS linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an HMGB1 polypeptide; an SV40 NLS; a nucleoplasmin NLS; anda programmable DNA-binding protein. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a deaminase; a 3xGH5 linker; a programmable DNA-binding protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a deaminase; a GH5 linker; a programmable DNA-binding protein; an SV40 NLS; a 41 amino acid residue linker; and an HMGB1 polypeptide. In any one of the aforementioned fusion proteins, the programmable DNA binding domain may comprise an amino acid sequence having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704, 705, 707, 708, 712, 713, 715, 716, and 719-721. In any one of the aforementioned fusion proteins, the HMGB1 polypeptide may comprise an amino acid sequence having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, 21, and 22. In some any one of the aforementioned fusion proteins, the deaminase may comprise an amino acid sequence having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 728 and 729;
[0225] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a programmable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704, 705, 707, 708, 712, 713, 715, 716, and 719-721; a linker having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435; and an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, 21, and 22.
[0226] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a programmable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704, 705, 707, 708, 712, 713, 715, 716, and 719-721; a first linker having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435; a heterologous NLS having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384; a second linker having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435; and an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, 21, and 22.
[0227] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a first heterologous NLS having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384; a programmable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704, 705, 707, 708, 712, 713, 715, 716, and 719-721; a first linker having at least 90%, 95%, 99%, or 100% identity to anyone of SEQ ID NOs: 301-365 and 425-435; a second heterologous NLS having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384; a second linker having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435; and an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, 21, and 22.
[0228] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a heterologous NLS having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384; programmable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704, 705, 707, 708, 712, 713, 715, 716, and 719-721; a linker having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435; and an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, 21, and 22.
[0229] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a first heterologous NLS having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384; a second heterologous NLS having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384; a programmable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704, 705, 707, 708, 712, 713, 715, 716, and 719-721; a linker having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435; and an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, 21, and 22.
[0230] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, 21, and 22; a first heterologous NLS having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384; a second heterologous NLS having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384; and a programmable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704, 705, 707, 708, 712, 713, 715, 716, and 719-721.
[0231] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a first heterologous NLS having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384; a second heterologous NLS having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384; a deaminase having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 728 and 729; a first linker having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435; aprogrammable DNA-binding domain having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704, 705, 707, 708, 712, 713, 715, 716, and 719-721; a second linker having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435; and an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, 21, and 22.
[0232] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a deaminase having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 728 and 729; a first linker having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435; a DNA-binding domain having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704, 705, 707, 708, 712, 713, 715, 716, and 719-721; a heterologous NLS having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 366-384; a second linker having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 301-365 and 425-435; and an HMGB1 binding domain having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, 21, and 22.
[0233] In some embodiments, for example, as shown above, the fusion protein comprises, from N-terminus to C -terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433 or 429; 5) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;6) a linker having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NO: 426 or 429; and7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7 or 10.In some embodiments, a nucleic acid (e.g., mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0234] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 426; and 7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0235] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429; and 7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0236] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433; 5) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429; and 7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 10.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0237] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433; 5) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429; and 7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0238] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;3) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429; and 4) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0239] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433; 5) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 430; 7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0240] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 431; 7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0241] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433; 5) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 432; 7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0242] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 371; 3) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433; 4) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;5) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429;6) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0243] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;3) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;4) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;5) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429;6) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0244] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;3) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;4) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;5) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 435;6) a bipartite NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 436; 7) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 435;8) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0245] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) a programmable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;2) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433; 5) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0246] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7;2) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434; 3) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 4) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;5) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433; 6) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0247] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366;3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) aprogrammable DNA-binding protein having at least 90%, 95%, 99%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719- 721;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: SEQ ID NO: 429;7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0248] In some embodiments, the fusion protein comprises a sequence having at least at least 80%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%. 97%, 98%, 99%, or 100% identity to any one of SEQ ID NOs: 515, 524, 527, 530, 533, 536, 539, 542, 549, 552, 555, 567, 573, 576, 588, 591, 597, 603, 624, 627, 630, 633, 636, 639, 642, 649, 652, 655, 664, 667, or 673, or is encoded by a nucleic acid having at least at least 80%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%. 97%, 98%, 99%, or 100% identity to the sequence of any one of SEQ ID NOs: 513, 514, 522, 523, 525, 526, 528, 529, 531, 532, 534, 535, 537, 538, 540, 541, 547, 548, 550, 551, 553, 554, 565, 566, 571, 572, 574, 575, 586, 587, 589, 590, 595, 596, 601, 602, 622, 623, 625, 626, 628, 629, 631, 632, 634, 635, 637, 638, 640, 641, 647, 648, 650, 651, 653, 654, 662, 663, 665, 666, 671, or 672.
[0249] In some embodiments, the fusion protein comprises a sequence having at least 80%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%. 97%, 98%, 99%, or 100% to any one of SEQ ID NOs: 515, 567, 591, 597, 603, 624, 627, 630, 633, 636, 639, 642, 649, 652, 655, 664, 667, or 673, or is encoded by a nucleic acid having at least 80%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%. 97%, 98%, 99%, or 100% identity to the sequence of any one of SEQ ID NOs: 513, 514, 565, 566, 589, 590, 595, 596, 601, 602, 622, 623, 625, 626, 628, 629, 631, 632, 634, 635, 637, 638, 640, 641, 647, 648, 650, 651, 653, 654, 662, 663, 665, 666, 671, or 672. In some embodiments, a nucleic acid (e.g., mRNA) encoding the fusion protein is provided, wherein the polynucleotide comprises a nucleotide sequence having at least 80%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%. 97%, 98%, 99%, or 100% identity to the sequence of any one of SEQ ID NOs: 513, 514, 565, 566, 589, 590, 595, 596,601, 602, 622, 623, 625, 626, 628, 629, 631, 632, 634, 635, 637, 638, 640, 641, 647, 648, 650, 651, 653, 654, 662, 663, 665, 666, 671, or 672.
[0250] In some embodiments, the fusion protein comprises a sequence having at least 80%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%. 97%, 98%, 99%, or 100% identity to any one of SEQ ID NOs: 515, 567, 591, 597, 603, 624, 627, 630, 633, 636, 639, 642, 649, 652, 655, or is encoded by a nucleic acid having at least 80%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%. 97%, 98%, 99%, or 100% to the sequence of any one of SEQ ID NOs: 513, 514, 565, 566, 589, 590, 595, 596, 601, 602, 622, 623, 625, 626, 628, 629, 631, 632, 634, 635, 637, 638, 640, 641, 647, 648, 650, 651, 653, or 654. In some embodiments, a nucleic acid (e.g., mRNA) encoding the fusion protein is provided, wherein the polynucleotide comprises a nucleotide sequence having at least 80%, 85%, 88%, 90%, 91%, 92%, 93%, 94%, 95%, 96%.97%, 98%, 99%, or 100% identity to the sequence of any one of SEQ ID NOs: 513, 514, 565, 566, 589, 590, 595, 596, 601, 602, 622, 623, 625, 626, 628, 629, 631, 632, 634, 635, 637, 638, 640, 641, 647, 648, 650, 651, 653, or 654.
[0251] In some embodiments, the nucleic acid is an mRNA sequence. In some embodiments, the nucleic acid is a DNA sequence. It is to be understood that, if a sequence is provided as a DNA sequence (i.e., comprising thymidine nucleotides instead of uridine nucleotides), the corresponding RNA sequence would comprise uridine nucleotides instead of thymidine nucleotides. Similarly, if a sequence is provided as an RNA sequence (i.e., comprising uridine nucleotides instead of thymidine nucleotides), the corresponding DNA sequence would comprise thymidine nucleotides instead of uridine nucleotides.
[0252] In some embodiments, the present disclosure provides for a fusion protein comprising a Class II Cas nuclease and an HMGB1 polypeptide comprising an HMGB1 Box B domain and lacking amino acid residues 186-215 relative to SEQ ID NO: 1, further wherein the HMGB1 polypeptide comprises a cryptic NLS. In some embodiments, the Class II Cas nuclease is located N-terminal to the HMGB1 polypeptide. In some embodiments, the Class II Cas nuclease is located C-terminal to the HMGB1 polypeptide. In some embodiments, the fusion protein comprises at least one heterologous nuclear localization signal (NLS).
[0253] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a Cas protein; a heterologous NLS; a linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a Cas protein; a linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a Cas protein; a first linker; a heterologous NLS; a second linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a first heterologous NLS; a Cas protein; a first linker; a second heterologous NLS; a second linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a heterologous NLS; a Cas protein; a linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a first heterologous NLS; a second heterologous NLS; a Cas protein; a linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an HMGB1 polypeptide; a first heterologous NLS; a second heterologous NLS; and a Cas protein. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a first heterologous NLS; a second heterologous NLS; a deaminase; a first linker; a DNA-binding domain; a second linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a deaminase; a first linker; a DNA-binding domain; a heterologous NLS; a second linker; and an HMGB1 polypeptide. In some embodiments, wherein the fusion protein comprises a first and second heterologous NLS, the first and second heterologous NLS are the same. In some embodiments, wherein the fusion protein comprises a first and second heterologous NLS, the first and second heterologous NLS are different. In some embodiments, wherein the fusion protein comprises a first and second linker, the first and second linkers are the same. In some embodiments, wherein the fusion protein comprises a first and second linker, the first and second linkers are different.
[0254] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a Cas9 protein; a 41 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a Cas9 protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a Cas9 protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a Cas9 protein; a GS20 linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a Cas9 protein; a GS15 linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a Cas9 protein; a GS12 linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a Cas9 protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus toC-terminus: a nucleoplasmin NLS; a Cas9 protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a nucleoplasmin NLS; a Cas9 protein; a first GS linker; an SV40 NLS; a second GS linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a Cas9 protein; a GS linker; a nucleoplasmin NLS; a GS linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an HMGB1 polypeptide; an SV40 NLS; a nucleoplasmin NLS; and a Cas9 protein. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a deaminase; a 3xGH5 linker; a Cas9 nickase; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a deaminase; a GH5 linker; a Cas9 nickase; an SV40 NLS; a 41 amino acid residue linker; and an HMGB1 polypeptide.
[0255] In some embodiments, the Cas9 is a SpyCas9 or an NmeCas9. In some embodiments, wherein the Cas9 is an NmeCas9, the NmeCas9 is an Nme2Cas9.
[0256] In some embodiments, the Cas9 comprises a mutation in a RuvC domain or an HNH domain. In some embodiments, wherein the Cas9 is a SpyCas9, the SpyCas9 comprises a point mutation in H840, DIO, or N863. In some embodiments, wherein the Cas9 is a SpyCas9, the SpyCas9 comprises an H840A, D10A, or N863 mutation. In some embodiments, wherein the Cas9 is an NmeCas9, the NmeCas9 comprises a point mutation in D16 or H588. In some embodiments, wherein the Cas9 is an NmeCas9, the NmeCas9 comprises a D16A or H588A mutation.
[0257] In some embodiments, the fusion protein comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 515, 524, 527, 530, 533, 536, 539, 542, 549, 552, 555, 567, 573, 576, 588, 591, 597, 603, 624, 627, 630, 633, 636, 639, 642, 649, 652, 655, 664, 667, or 673, or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 513, 514, 522, 523, 525, 526, 528, 529, 531, 532, 534, 535, 537, 538, 540, 541, 547, 548, 550, 551, 553, 554, 565, 566, 571, 572, 574, 575, 586, 587, 589, 590, 595, 596, 601, 602, 622, 623, 625, 626, 628, 629, 631, 632, 634, 635, 637, 638, 640, 641, 647, 648, 650, 651, 653, 654, 662, 663, 665, 666, 671, or 672.
[0258] In some embodiments, the fusion protein comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 515, 567, 591, 597, 603, 624, 627, 630, 633, 636, 639, 642, 649, 652, 655, 664, 667, or 673, or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any oneof SEQ ID NOs: 513, 514, 565, 566, 589, 590, 595, 596, 601, 602, 622, 623, 625, 626, 628, 629, 631, 632, 634, 635, 637, 638, 640, 641, 647, 648, 650, 651, 653, 654, 662, 663, 665, 666, 671, or 672. In some embodiments, the nucleic acid is an mRNA sequence. In some embodiments, the nucleic acid is a DNA sequence. It is to be understood that, if a sequence is provided as a DNA sequence (i.e., comprising thymidine nucleotides instead of uridine nucleotides), the corresponding RNA sequence would comprise uridine nucleotides instead of thymidine nucleotides. Similarly, if a sequence is provided as an RNA sequence (i.e., comprising uridine nucleotides instead of thymidine nucleotides), the corresponding DNA sequence would comprise thymidine nucleotides instead of uridine nucleotides.
[0259] In some embodiments, the present disclosure provides for a fusion protein comprising an NmeCas9 nuclease and an HMGB1 polypeptide comprising an HMGB1 Box B domain and lacking amino acid residues 186-215 relative to SEQ ID NO: 1, further wherein the HMGB1 polypeptide comprises a cryptic NLS comprising the sequence of EKSKKKK (SEQ ID NO: 5). In some embodiments, the NmeCas9 is located N-terminal to the HMGB1 polypeptide. In some embodiments, the NmeCas9 is located C-terminal to the HMGB1 polypeptide. In some embodiments, the fusion protein comprises at least one heterologous nuclear localization signal (NLS).
[0260] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; an NmeCas9 protein; a 41 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; an NmeCas9 protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an NmeCas9 protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; an NmeCas9 protein; a GS20 linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; an NmeCas9 protein; a GS15 linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; an NmeCas9 protein; a GS12 linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; an NmeCas9 protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a nucleoplasmin NLS; an NmeCas9 protein; a 33 amino acid residue linker; and an HMGB1polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a nucleoplasmin NLS; an NmeCas9 protein; a first GS linker; an SV40 NLS; a second GS linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an NmeCas9 protein; a GS linker; a nucleoplasmin NLS; a GS linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an HMGB1 polypeptide; an SV40 NLS; a nucleoplasmin NLS; and an NmeCas9 protein. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a deaminase; a 3xGH5 linker; an NmeCas9 nickase; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a SpyCas9 protein; an SV40 NLS; a 41 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; a SpyCas9 protein; a 33 amino acid residue linker; and an HMGB1 polypeptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a deaminase; a GH5 linker; a SpyCas9 nickase; a 41 amino acid residue linker; and an HMGB 1 polypeptide.
[0261] In some embodiments, the fusion protein comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 712-713, 715-716, and 719-721 or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 711, 714, 717-718, and 733-736.
[0262] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: an SV40 NLS; a nucleoplasmin NLS; SpyCas9; a 33 amino acid residue linker; and an HMGB 1 polypeptide.
[0263] In some embodiments, the fusion protein comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 702 and 704-708 or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 702 and 704-708.
[0264] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 426; and 7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0265] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429; and 7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0266] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429; and7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 10.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0267] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429; and 7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0268] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;3) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429; and 4) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0269] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 430;7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0270] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 431;7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0271] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 432;7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0272] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 371; 3) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;4) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;5) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429;6) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0273] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;3) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;4) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;5) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429;6) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0274] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;3) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;4) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;5) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 435;6) a bipartite NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 436; 7) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 435;8) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0275] In some embodiments, the fusion protein comprises, from N-terminus to C- terminus:1) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;2) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0276] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7;2) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;3) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366;4) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;5) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;6) an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0277] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus:1) optionally a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434;2) an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; 3) a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;4) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433;5) a SpyCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 704;6) a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: SEQ ID NO: 429;7) an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.In some embodiments, a nucleic acid (e.g., a mRNA) comprising an open reading frame (ORF) encoding the aforementioned fusion protein is provided.
[0278] In some embodiments, the fusion protein comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 567 and 591 or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 565-566 and 589-590.Additional Features
[0279] The following section provides additional features of the HMGB 1 polypeptide, system, fusion protein, the nucleic acid or nucleic acids encoding the same, guide RNAs, and compositions disclosed herein. In any of the embodiments set forth herein, the nucleic acid or nucleic acids may be one or more expressions construct comprising a promoter operablylinked to an ORF encoding the HMGB 1 polypeptide or fusion protein disclosed herein.Additional Heterologous functional domains
[0280] In some embodiments, the heterologous functional domain may be capable of modifying the intracellular half-life of the HMGB1 polypeptide, system, or fusion protein disclosed herein. In some embodiments, the half-life of the HMGB1 polypeptide, system, or fusion protein disclosed herein may be increased. In some embodiments, the half-life of the HMGB1 polypeptide, system, or fusion protein disclosed herein may be reduced. In some embodiments, the heterologous functional domain may be capable of increasing the stability of the HMGB1 polypeptide, system, or fusion protein disclosed herein. In some embodiments, the heterologous functional domain may be capable of reducing the stability of the HMGB1 polypeptide, system, or fusion protein disclosed herein. In some embodiments, the heterologous functional domain may act as a signal peptide for protein degradation. In some embodiments, the protein degradation may be mediated by proteolytic enzymes, such as, for example, proteasomes, lysosomal proteases, or calpain proteases. In some embodiments, the heterologous functional domain may comprise a PEST sequence. In some embodiments, the HMGB1 polypeptide may be modified by addition of ubiquitin or a polyubiquitin chain. In some embodiments, the ubiquitin may be a ubiquitin-like protein (UBL). Non-limiting examples of ubiquitin-like proteins include small ubiquitin-like modifier (SUMO), ubiquitin cross-reactive protein (UCRP, also known as interferon-stimulated gene-15 (ISG15)), ubiquitin-related modifier-1 (URM1), neuronal-precursor-cell-expressed developmentally downregulated protein-8 (NEDD8, also called Rub1 in S. cerevisiae), human leukocyte antigen F-associated (FAT10), autophagy-8 (ATG8) and -12 (ATG12), Fau ubiquitin-like protein (FUB1), membrane-anchored UBE (MUB), ubiquitin fold-modifier- 1 (UFM1), and ubiquitin-like protein-5 (UBE5).
[0281] In some embodiments, the heterologous functional domain may be a marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, epitope tags, and reporter gene sequences. In some embodiments, the marker domain may be a fluorescent protein. Any known fluorescent proteins may be used as the marker domain such as GFP, YFP, EBFP, ECFP, DsRed or any other suitable fluorescent protein. In some embodiments, the marker domain may be a purification tag or an epitope tag. Nonlimiting exemplary tags include glutathione-S-transferase (GST), chitin binding protein (GBP), maltose binding protein (MBP), thioredoxin (TRX), poly(NANP), tandem affinitypurification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, SI, T7, V5, VSV-G, 6xHis, 8xHis, biotin carboxyl carrier protein (BCCP), poly-His, and calmodulin. In some embodiments, the marker domain may be a reporter gene. Non-limiting exemplary reporter genes include glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, or fluorescent proteins.
[0282] In additional embodiments, the heterologous functional domain may target the HMGB1 polypeptide, system, or fusion protein disclosed herein to a specific organelle, cell type, tissue, or organ. In some embodiments, the heterologous functional domain may target the HMGB1 polypeptide, system, or fusion protein disclosed herein to mitochondria.UTRs; Kozak sequences
[0283] In some embodiments, the nucleic acid (e.g., mRNA) disclosed herein comprises a 5’ UTR, 3’ UTR, or 5’ and 3’ UTRs from Hydroxysteroid 17-Beta Dehydrogenase 4 (HSD17B4 or HSD) or globin such as human alpha globin (HBA), human beta globin (HBB), Xenopus laevis beta globin (XBG), bovine growth hormone, cytomegalovirus (CMV), mouse Hba-al, heat shock protein 90 (Hsp90), glyceraldehyde 3-phosphate dehydrogenase (GAPDH), beta-actin, alpha-tubulin, tumor protein (p53), or epidermal growth factor receptor (EGFR).
[0284] In some embodiments, the nucleic acid described herein does not comprise a 5’ UTR, e.g., there are no additional nucleotides between the 5’ cap and the start codon. In some embodiments, the nucleic acid comprises a Kozak sequence (described below) between the 5’ cap and the start codon, but does not have any additional 5’ UTR. In some embodiments, the nucleic acid does not comprise a 3’ UTR, e.g., there are no additional nucleotides between the stop codon and the poly- A tail.
[0285] In some embodiments, the polynucleotide comprises a 5’ UTR with at least 85%, 90%, or 95% identity to any one of SEQ ID NOs: 398-405. In some embodiments, the polynucleotide comprises a 3’ UTR with at least 85%, 90%, or 95% identity to any one of SEQ ID NOs: 406-413. In some embodiments, the polynucleotide comprises a 5’ UTR and 3’ UTR from the same source.
[0286] In some embodiments, the nucleic acid herein comprises a Kozak sequence. The Kozak sequence can affect translation initiation and the overall yield of a polypeptide translated from an mRNA. A Kozak sequence includes a methionine codon that can function as the start codon. A minimal Kozak sequence is NNNRUGN (SEQ ID NO: 416) wherein atleast one of the following is true: the first N is A or G and the second N is G. In the context of a nucleotide sequence, R means a purine (A or G). In some embodiments, the Kozak sequence is RNNRUGN (SEQ ID NO: 417), NNNRUGG (SEQ ID NO: 418), RNNRUGG (SEQ ID NO: 419), RNNAUGN (SEQ ID NO: 420), NNNAUGG (SEQ ID NO: 421), RNNAUGG (SEQ ID NO: 422), or GCCACCAUG (SEQ ID NO: 423).Poly-A tail
[0287] In some embodiments, the nucleic acid disclosed herein further comprises a poly-adenylated (poly-A) tail. The poly-A tails may comprise at least 8 consecutive adenine nucleotides, but also comprise one or more non-adenine nucleotide. As used herein, “nonadenine nucleotides” refers to any natural or non-natural nucleotides that do not comprise adenine. Guanine, thymine, and cytosine nucleotides are exemplary non-adenine nucleotides. Thus, the poly-A tails on the nucleic acid described herein may comprise consecutive adenine nucleotides located 3’ to nucleotides encoding a polypeptide of interest. In some instances, the poly-A tails on the nucleic acid comprise non-consecutive adenine nucleotides located 3’ to nucleotides encoding the HMGB 1 polypeptide, wherein non-adenine nucleotides interrupt the adenine nucleotides at regular or irregularly spaced intervals.
[0288] In some embodiments, the poly-A tail is encoded in a plasmid used for in vitro transcription of an mRNA and becomes part of the transcript. The poly-A sequence encoded in the plasmid, i.e., the number of consecutive adenine nucleotides in the poly-A sequence, may not be exact, e.g., a 100 poly-A sequence in the plasmid may not result in a precisely 100 poly-A sequence in the transcribed mRNA. In some embodiments, the poly-A tail is not encoded in the plasmid, and is added by PCR tailing or enzymatic tailing, e.g., using E. coli poly(A) polymerase.
[0289] In some embodiments, the one or more non-adenine nucleotides are positioned to interrupt the consecutive adenine nucleotides so that a poly(A) binding protein can bind to a stretch of consecutive adenine nucleotides. In some embodiments, one or more non-adenine nucleotide(s) is located after at least 8, 9, 10, 11, or 12 consecutive adenine nucleotides. In some embodiments, the one or more non-adenine nucleotide is located after 8-50 consecutive adenine nucleotides. In some embodiments, the one or more non-adenine nucleotide is located after 8-100 consecutive adenine nucleotides.
[0290] In some embodiments, the poly-A tail comprises or contains one non-adenine nucleotide or one consecutive stretch of 2-10 non-adenine nucleotides.
[0291] In some embodiments, the non-adenine nucleotide is guanine, cytosine, or thymine. In some instances, where more than one non-adenine nucleotide is present, the non-adenine nucleotide may be selected from: a) guanine and thymine nucleotides; b) guanine and cytosine nucleotides; c) thymine and cytosine nucleotides; or d) guanine, thymine and cytosine nucleotides.Modified nucleotides
[0292] In some embodiments, the ORF or mRNA disclosed herein comprises a modified uridine at some or all uridine positions. In some embodiments, the modified uridine is a uridine modified at the 5 position, e.g., with a halogen or C 1 -C3 alkoxy. In some embodiments, the modified uridine is a pseudouridine modified at the 1 position, e.g., with a C1-C3 alkyl. The modified uridine can be, for example, pseudouridine, N1 -methylpseudouridine, 5-methoxyuridine, 5 -iodouridine, or a combination thereof.
[0293] In some embodiments, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the uridine positions in the nucleic acid disclosed herein are modified uridines. In some embodiments, 10%-25%, 15-25%, 25-35%, 35-45%, 45-55%, 55-65%, 65-75%, 75-85%, 85-95%, or 90-100% of the uridine positions in an mRNA disclosed herein are modified uridines, e.g., 5-methoxyuridine, 5 -iodouridine, N1 -methyl pseudouridine, pseudouridine, or a combination thereof. In some embodiments, 80-95% or 80-100% of the uridine positions in an mRNA disclosed herein are modified uridines, e.g., 5-methoxyuridine, 5 -iodouridine, Nl-methyl pseudouridine, pseudouridine, or a combination thereof.
[0294] In some embodiments, at least 50% of the uridine is substituted with a modified uridine. In some embodiments, 15% to 45% of the uridine is substituted with the modified uridine. In some embodiments, at least 70%, 80%, 85%, 90%, 95%, or 100% of the uridine is substituted with the modified uridine.
[0295] In some embodiments, at least 85% of the uridine is substituted with modified uridine. In some embodiments, at least 95% of the uridine is substituted with modified uridine.
[0296] In some embodiments, the modified uridine is one or more of N1 -methyl -pseudouridine, pseudouridine or 5 -iodouridine. In some embodiments, the modified uridine is Nl-methyl -pseudouridine. In some embodiments, the modified uridine is pseudouridine. In some embodiments, the modified uridine is 5-iodouridine. In some embodiments, at least 85% of the uridine is substituted with the modified uridine. In some embodiments, 100% uridine is substituted with the modified uridine.5’ Cap
[0297] In some embodiments, the nucleic acid disclosed herein comprises a 5’ cap, such as a CapO, Capl, or Cap2. A 5’ cap is generally a 7-methylguanine ribonucleotide (which may be further modified, as discussed below e.g., with respect to ARCA) linked through a 5’-triphosphate to the 5’ position of the first nucleotide of the 5’-to-3’ chain of the nucleic acid, i.e., the first cap-proximal nucleotide. In CapO, the riboses of the first and second cap-proximal nucleotides of the mRNA both comprise a 2’-hydroxyl. In Capl, the riboses of the first and second transcribed nucleotides of the nucleic acid comprise a 2’-methoxy and a 2’-hydroxyl, respectively. In Cap2, the riboses of the first and second cap-proximal nucleotides of the nucleic acid both comprise a 2’-methoxy. See, e.g., Katibah et al. (2014) Proc Natl Acad Sci USA 111(33): 12025-30; Abbas et al. (2017) Proc Natl Acad Sci USA 114(11): E2106-E2115. Most endogenous higher eukaryotic nucleic acids, including mammalian nucleic acids such as human nucleic acids, comprise Capl or Cap2. CapO and other cap structures differing from Capl and Cap2 may be immunogenic in mammals, such as humans, due to recognition as “non-self” by components of the innate immune system such as IFIT-1 and IFIT-5, which can result in elevated cytokine levels including type I interferon. Components of the innate immune system such as IFIT-1 and IFIT-5 may also compete with eIF4E for binding of a nucleic acids with a cap other than Capl or Cap2, potentially inhibiting translation of the nucleic acid.
[0298] A 5’ cap can be included co-transcriptionally. For example, ARCA (antireverse cap analog; Thermo Fisher Scientific Cat. No. AM8045) is a cap analog comprising a 7-methylguanine 3 ’-methoxy-5’ -triphosphate linked to the 5’ position of a guanine ribonucleotide which can be incorporated in vitro into a transcript at initiation. ARCA results in a CapO cap or a CapO-like cap in which the 2’ position of the first cap-proximal nucleotide is hydroxyl. See, e.g., Stepinski et al., (2001) “Synthesis and properties of mRNAs containing the novel ‘anti-reverse’ cap analogs 7-methyl(3'-O-methyl)GpppG and 7-methyl(3'deoxy)GpppG,” RNA 7: 1486-1495. The ARCA structure is shown below.
[0299] CleanCap™ AG (m7G(5')ppp(5')(2'OMeA)pG; TriLink Biotechnologies Cat. No. N-7113) or CleanCap™ GG (m7G(5')ppp(5')(2'OMeG)pG; TriLink Biotechnologies Cat. No. N-7133) can be used to provide a Capl structure co-transcriptionally. 3’-O-methylated versions of CleanCap™ AG and CleanCap™ GG are also available from TriLink Biotechnologies as Cat. Nos. N-7413 and N-7433, respectively. The CleanCap™ AG structure is shown below. CleanCap™ structures are sometimes referred to herein using the last three digits of the catalog numbers listed above (e.g., “CleanCap™ 113” for TriLink Biotechnologies Cat. No. N-7113).
[0300] Alternatively, a cap can be added to an RNA post-transcriptionally. For example, Vaccinia capping enzyme is commercially available (New England Biolabs Cat. No. M2080S) and has RNA triphosphatase and guanylyltransferase activities, provided by its DI subunit, and guanine methyltransferase, provided by its D12 subunit. As such, it can add a 7- methylguanine to an RNA, so as to give CapO, in the presence of S-adenosyl methionine and GTP. See, e.g., Guo, P. and Moss, B. (1990) Proc. Natl. Acad. Sci. USA 87, 4023-4027; Mao, X. and Shuman, S. (1994) J. Biol. Chem. 269, 24472-24479. For additional discussion of caps and capping approaches, see, e.g., WO2017 / 053297 and Ishikawa et al., Nucl. Acids. Symp. Ser. (2009) No. 53, 129-130.Guide RNAs
[0301] In some embodiments, the systems, methods and compositions of the present disclosure further comprises a guide RNA or use thereof. As used herein, a guide RNA, or gRNA, is understood to include at least a guide / spacer (targeting) sequence that is complementary to a genomic locus; and a scaffold sequence for binding a SpyCas9 nuclease. The guide / spacer and scaffold can be in a single polynucleotide, i.e., a single guide RNA(sgRNA), or in two polynucleotides, i.e., a dual guide RNA (dgRNA) in which the crRNA and tracrRNA are separate polynucleotides.Target / Guide Sequences and Genes
[0302] In some embodiments, the methods and compositions of the present disclosure utilize a polypeptide, system, or fusion protein comprising a Class II nuclease.
[0303] For example, a target sequence may be recognized and cleaved by a Class II Cas nuclease. A target sequence for a Class II Cas nuclease is located near the nuclease’s cognate PAM sequence. In some embodiments, a Class II Cas nuclease may be directed by a guide RNA to a target sequence of a gene, where the guide RNA hybridizes with and the Class II Cas protein cleaves, nicks, or binds the target sequence. In some embodiments, the guide RNA hybridizes with a Class II Cas nuclease cleaves, nicks, or binds the target sequence adjacent to or comprising its cognate PAM. The target sequence may be complementary or have identity to a targeting sequence of the guide RNA. In some embodiments, the degree of complementarity between a targeting sequence of a guide RNA and the portion of the corresponding target sequence that hybridizes to the guide RNA may be about 80%, 85%, preferably about 90%, 95%, or 100%. In some embodiments, the percent identity between a targeting sequence of a guide RNA and the portion of the corresponding target sequence that hybridizes to the guide RNA may be about 80%, 85%, preferably about 90%, 95%, or 100%. The homology region of the target is adjacent to a cognate PAM sequence. In some embodiments, the target sequence may comprise a sequence 100% complementary or 100% identical with the targeting sequence of the guide RNA. In other embodiments, the target sequence may comprise at least one mismatch, deletion, or insertion, as compared to the targeting sequence of the guide RNA.
[0304] The length of the target sequence may depend on the nuclease system used. For example, the targeting sequence of a guide RNA for a CRISPR / Cas system may comprise 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30in length and the target sequence is a corresponding length, optionally adjacent to a PAM sequence. In some embodiments, the target sequence may comprise 15-24, optionally 16-22, nucleotides in length. In some embodiments, the target sequence may comprise 17-21, optionally 18-20, nucleotides in length. In some embodiments, the target sequence is 20 nucleotides in length. In some embodiments, the target sequence may comprise 24 nucleotides in length.
[0305] The target nucleic acid molecule may be any DNA molecule that is endogenous or exogenous to a cell. In some embodiments, the target nucleic acid molecule may be anepisomal DNA, a plasmid, a genomic DNA, viral genome, or chromosomal DNA. In some embodiments, the target sequence of the gene may be a genomic sequence from a cell or in a cell, including a human cell.
[0306] In further embodiments, the target sequence may be a viral sequence. In further embodiments, the target sequence may be a pathogen sequence. In yet other embodiments, the target sequence may be a synthesized sequence. In further embodiments, the target sequence may be a genomic sequence. In certain embodiments, the target sequence may comprise a translocation junction, e.g., a translocation associated with a cancer. In some embodiments, the target sequence may be on a eukaryotic chromosome, such as a human chromosome.
[0307] In some embodiments, the target sequence may be located in a genomic locus; for example, the target sequence may be located in a coding sequence of a gene, an intron sequence of a gene, a regulatory sequence, a transcriptional control sequence of a gene, a translational control sequence of a gene, a splicing site, or a non-coding sequence between genes (e.g., intergenic space). In some embodiments, the gene may be a protein coding gene. In other embodiments, the genomic locus may be at a non-coding locus. In some embodiments, the target sequence may be in a disease-associated gene. In some embodiments, the target sequence may be located in a non-genic functional site in a genomic sequence, for example a site that controls aspects of chromatin organization, such as a scaffold site or locus control region.
[0308] In some embodiments involving a Cas nuclease, such as a Class II Cas nuclease, the target sequence may be adjacent to a protospacer adjacent motif (“PAM”). In some embodiments, the PAM may be adjacent to or within 1, 2, 3, or 4, nucleotides of the 3' end of the target sequence. The length and the sequence of the PAM may depend on the Cas protein used. For example, the PAM may be selected from a consensus or a particular PAM sequence for a specific Spy Cas9 protein or Spy Cas9 ortholog, including those disclosed in FIG. 1 of Ran et al., Nature, 520: 186-191 (2015), and FIG. S5 of Zetsche 2015, the relevant disclosure of each of which is incorporated herein by reference. In some embodiments, the PAM may be 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length. Non-limiting exemplary PAM sequences include NGG, NGGNG, NG, NAAAAN, NNAAAAW, NNNNACA, GNNNCNNA, TTN, and NNNNGATT (wherein N is defined as any nucleotide, and W is defined as either A or T). In some embodiments, the PAM sequence may be NGG. In some embodiments, the PAM sequence may be NGGNG. In some embodiments, the PAM sequence may be TTN. In some embodiments, the PAM sequence may be NNAAAAW. It is understood that Cas9 nucleases can be modified to alter PAM recognition. It is understood that the use ofa SpyCas9 nuclease or NmeCas9 nuclease with altered PAM recognition is within the scope of the disclosure provided herein.
[0309] In some embodiments, the PAM may be selected from a consensus or a particular PAM sequence for a specific Nme Cas9 protein or Nme Cas9 ortholog (Edraki et al., 2019). In some embodiments, the Nme Cas9 PAM may comprise 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length. Non-limiting exemplary PAM sequences include NCC, N4GAYW, N4GYTT, N4GTCT, NNNNCC(a), NNNNCAAA (wherein N is defined as any nucleotide, W is defined as either A or T, and R is defined as either A or G; and (a) is a preferred, but not required, A after the second C)). In some embodiments, the PAM sequence may be NCC.
[0310] In one embodiment, the PAM may be selected from a consensus or a particular PAM sequence for other Class II-C Cas9 orthologs. In some embodiments, the SmuCas9 PAM may comprise one to four required nucleotides selected from the group consisting of N4CN3, N4CT, N4CCN, N4CCA, and N4GNT3. In one embodiment, the one to four required nucleotides are selected from the group consisting of C, CT, CCN, CCA, CN3and GNT2. In one embodiment, Type II-C Cas9 is bound to a truncated sgRNA.
[0311] In some embodiments, the guide RNA is a single guide RNA (sgRNA). In some embodiments, the gRNA is a short-single guide RNA (short-sgRNA) comprising a conserved portion of an sgRNA comprising a hairpin region, wherein the hairpin region lacks at least 5-10 nucleotides and wherein the short-sgRNA comprises a 5’ end modification or a 3’ end modification or both.Guide RNA Scaffolds
[0312] In some embodiments, the guide RNA is a SpyCas9 guide RNA. In the case of a SpyCas9 single guide RNA (sgRNA), the above guide sequences may further comprise additional nucleotides to form a sgRNA, e.g., with the following exemplary nucleotide sequence following the 3’ end of the guide sequence:GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 102) in 5’ to 3’ orientation.
[0313] In the case of a sgRNA, the above guide sequences may further comprise additional nucleotides to form a sgRNA, e.g., with any one of the following exemplary nucleotide sequence following the 3’ end of the guide sequence:GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 101) in 5’ to 3’ orientation; orGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAAAAUG GCACCGAGUCGGUGCU (SEQ ID NO: 106) in 5’ to 3’ orientation.
[0314] In the case of a sgRNA, the guide sequences may be integrated into the following modified motif:mN*mN*mN*NNNNNNNNNNNNNNNNNGUUUUAGAmGmCmUmAmGmAmAmAmU mAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmA mAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmU*mU*mU*mU (SEQ ID NO: 172) where “N” may be any natural or non-natural nucleotide, preferably an RNA nucleotide; sugar moieties of the nucleotide can be ribose, deoxyribose, or similar compounds with substitutions; m is a 2’-O-methyl modified nucleotide, and * is a phosphorothioate linkage to the adjacent nucleotide residue; and wherein the N’s are collectively the nucleotide sequence of a guide sequence. In the context of a modified sequence, unless otherwise indicated, A, C, G, N, and U are an unmodified RNA nucleotide, i.e., a 2’-OH sugar moiety with a phosphodiesterase linkage to the adjacent nucleotide residue, or a 5 ’-terminal PO4.
[0315] In the case of a sgRNA, the guide sequences may further comprise a SpyCas9 sgRNA sequence. An example of a SpyCas9 sgRNA sequence is shown in Table 38 (SEQ ID NO: 102:GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGC - “Exemplary SpyCas9 sgRNA- 1”), included at the 3’ end of the guide sequence, and provided with the domains as shown in Table 38 below. LS is lower stem. B is bulge. US is upper stem. Hl and H2 are hairpin 1 and hairpin 2, respectively. Collectively Hl and H2 are referred to as the hairpin region. A model of the structure is provided in FIG. 10A of WO2019237069 which is incorporated herein by reference.
[0316] The nucleotide sequence of Exemplary SpyCas9 sgRNA- 1 may serve as a template sequence for specific chemical modifications, sequence substitutions and truncations.
[0317] In certain embodiments, the guide RNA is an sgRNA or a dgRNA, for example, and it optionally comprises a chemical modification. In some embodiments, the modified sgRNA comprises a guide sequence and a SpyCas9 sgRNA sequence, e.g., Exemplary SpyCas9 sgRNA- 1. A guide RNA, such as an sgRNA, may include modifications on the 5’ end of the guide sequence or on the 3’ end of the SpyCas9 sgRNA sequence, such as, e.g., Exemplary SpyCas9 sgRNA-1 at one or more of the terminal nucleotides, e.g., at 1, 2, 3, or 4 of the nucleotides at the 3’ end or at the 5’ end. In certain embodiments, the modifiednucleotide is selected from a 2’-O-methyl (2’-0Me) modified nucleotide, a 2’-O-(2-methoxyethyl) (2’-O-moe) modified nucleotide, a 2 ’-fluoro (2’-F) modified nucleotide, a phosphorothioate (PS) linkage between nucleotides, or an inverted abasic modified nucleotide; or a combination thereof. In certain embodiments, the modified nucleotide includes a 2’-0Me modified nucleotide. In certain embodiments, the modified nucleotide includes a PS linkage. In certain embodiments, the modified nucleotide includes a 2’-0Me modified nucleotide and a PS linkage.
[0318] In certain embodiments, using SEQ ID NO: 102 (“Exemplary SpyCas9 sgRNA-1”) as an example, the Exemplary SpyCas9 sgRNA-1 further includes one or more of: (A) a shortened hairpin 1 region, or a substituted and optionally shortened hairpin 1 region, wherein (1) at least one of the following pairs of nucleotides are substituted in hairpin 1 with Watson-Crick pairing nucleotides: Hl-1 and Hl-12, Hl-2 and Hl-11, Hl-3 and Hl-10, or Hl-4 and Hl -9, and the hairpin 1 region optionally lacks (a) any one or two of Hl -5 through Hl-8, (b) one, two, or three of the following pairs of nucleotides: Hl-1 and Hl-12, Hl-2 and Hill, Hl-3 and Hl-10, and Hl-4 and Hl-9, or (c) 1-8 nucleotides of hairpin 1 region; or (2) the shortened hairpin 1 region lacks 4-8 nucleotides, preferably 4-6 nucleotides, and (a) one or more of positions Hl-1, Hl-2, or Hl-3 is deleted or substituted relative to Exemplary SpyCas9 sgRNA-1 (SEQ ID NO: 102), or (b) one or more of positions Hl-6 through Hl-10 is substituted relative to Exemplary SpyCas9 sgRNA-l(SEQ ID NO: 102); or (3) the shortened hairpin 1 region lacks 5-10 nucleotides, preferably 5-6 nucleotides, and one or more of positions N18, Hl-12, or n is substituted relative to Exemplary SpyCas9 sgRNA-1 (SEQ ID NO: 102); or (B) a shortened upper stem region, wherein the shortened upper stem region lacks 1-6 nucleotides and wherein the 6, 7, 8, 9, 10, or 11 nucleotides of the shortened upper stem region include less than or equal to 4 substitutions relative to Exemplary SpyCas9 sgRNA-1 (SEQ ID NO: 102); or (C) a substitution relative to Exemplary SpyCas9 sgRNA-1 (SEQ ID NO: 102) at any one or more of LS6, LS7, US3, US10, B3, N7, N15, N17, H2-2 and H2-14, wherein the substituent nucleotide is neither a pyrimidine that is followed by an adenine, nor an adenine that is preceded by a pyrimidine; or (D) an Exemplary SpyCas9 sgRNA-1 (SEQ ID NO: 102) with an upper stem region, wherein the upper stem modification comprises a modification to any one or more of US 1 -US 12 in the upper stem region, wherein (1) the modified nucleotide is optionally selected from a 2’-O-methyl (2’-OMe) modified nucleotide, a 2’-O-(2-methoxyethyl) (2’-O-moe) modified nucleotide, a 2’-fluoro (2’-F) modified nucleotide, a phosphorothioate (PS) linkage between nucleotides, an inverted abasicmodified nucleotide, or a combination thereof; or (2) the modified nucleotide optionally includes a 2’-0Me modified nucleotide.
[0319] In some embodiments, the unmodified sgRNA comprises the following sequence:(N)2OGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCACG AAAGGGCACCGAGUCGGUGC (SEQ ID NO: 126); or (N oGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCACG AAAGGGCACCGAGUCGGUGCU (SEQ ID NO: 127).
[0320] In some embodiments, the sgRNA comprises a modified motif disclosed herein, including any modified motif shown in Tables 1C, ID, 2C, and 2D, where a guide RNA, or “N” may be any natural or non-natural nucleotide, preferably an RNA nucleotide; sugar moieties of the nucleotide can be ribose, deoxyribose, or similar compounds with substitutions; m is a 2’-O-methyl modified nucleotide, and * is a phosphoro thioate linkage to the adjacent nucleotide residue; and wherein the N’s are collectively the nucleotide sequence of a guide sequence.
[0321] In the context of a modified sequence, unless otherwise indicated, A, C, G, N, and U are an unmodified RNA nucleotide, i.e., a 2’-OH sugar moiety with a phosphodiester linkage to the adjacent nucleotide residue, or a 5 ’-terminal PO4.
[0322] In some embodiments, the guide RNA that directs the cleavase to a genomic locus is a SpyCas9 guide RNA. In some embodiments, the SpyCas9 guide RNA is a single guide RNA comprising: a conserved portion of an sgRNA comprising an upper stem and hairpin region, wherein every nucleotide in the upper stem region is modified with 2’-O-Me, and every nucleotide in the hairpin region is modified with 2’-O-Me; a 3’ end modification comprising 2’-O-Me modified nucleotides at the last three nucleotides of the 3’ end and phosphorothioate (PS) bonds between the last four nucleotides of the 3’ end; and 5’ end modification comprising 2’-O-Me modified nucleotides at the first three nucleotides of the 5’ end; and phosphorothioate (PS) bonds between the first four nucleotides of the 5’ end.
[0323] In some embodiments, the SpyCas9 guide RNA is a short-single guide RNA (short-sgRNA) comprising a conserved portion of an sgRNA comprising a hairpin region, wherein the hairpin region lacks at least 5-10 nucleotides and wherein the short-sgRNA comprises (i) a 5’ end modification or (ii) a 3’ end modification.
[0324] In some embodiments, the guide RNA is a SpyCas9 guide RNA that is a single guide RNA comprising a nucleotide sequence selected from SEQ ID NOs: 83-110, 112-138,140-159, and 167-237, or a nucleotide sequence that is at least 85%, 90%, or 95% identical to SEQ ID NOs: 83-110, 112-138, 140-159, and 167-237.
[0325] In some embodiments, the single guide RNA comprises, from 5’ to 3’, (1) the guide sequence; and (2) GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUG (SEQ ID NO: 102);GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCACGAAA GGGCACCGAGUCGGUGC (SEQ ID NO: 107); or GUUUUAGACGUAGAAAUACGAAGUUAAAAUAAGGCUAGUCCGUUAUCACGAAA GGGCACCGAGUCGGUGC (SEQ ID NO: 109).
[0326] In some embodiments, the sgRNA comprises Exemplary SpyCas9 sgRNA-1 or the modified versions thereof provided herein, or a version as provided in Table IB, where the totality of the N’s comprise a guide sequence that directs a nuclease to a target sequence. Each N is independently modified or unmodified. In certain embodiments, in the absence of an indication of a modification, the nucleotide is an unmodified RNA nucleotide residue, i.e., a ribose sugar and a phosphodiester backbone.TABLE 1A - Exemplary Unmodified SpyCas9 Scaffold SequencesTable IB: Exemplary Unmodified SpyCas9 Guide RNA Sequenceswherein the Ns collectively are a guide sequence provided herein. Within the table, in the context of an unmodified sequence, A, C, G, U, and N are, independently, any natural or nonnatural adenine, cytosine, guanine, uracil, and any nucleotide (e.g., A, C, G, or U), respectively.wherein “m” indicates a 2’-0-Me modification, “f” indicates a 2’-fluoro modification, a indicates a phosphorothioate linkage between nucleotides, and no modification in the context of a modified sequence indicates an RNA (2 ’-OH) and phosphodiesterase linkage to the 3’ nucleotide when one is present.
[0327] In certain embodiments, the guide sequence is a chemically modified sequence. In certain embodiments, the chemically modified guide sequence is (mN*)3(N)13-17. In certain embodiments, the guide sequence is (mN*)3(N)17, i.e.,mN*mN*mN*NNNNNNNNNNNNNNNNN. In certain embodiments, each N of the (N)13-17 or the (N)17 is unmodified, i.e., an RNA (2’-OH and phosphodiesterase linkage to the 3’ nucleotide when one is present). In certain embodiments, each N in the (N)13-17 or the (N)17 is independently modified, e.g., independently modified with a 2’-O-methyl modification.
[0328] In some embodiments, the sgRNA disclosed herein may be modified as shown herein or in the sequence mN*mN*mN*NNNNNNNNNNNNNNNNN GUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGU CCGUUAUCACGAAAGGGCACCGAGUCGG*mU*mG*mC (SEQ ID NO: 170).
[0329] In some embodiments, the sgRNA disclosed herein may be modified as shown herein or in the sequence mN*mN*mN*NNNNNNNNNNNNNNNNN GUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGU CCGUUAUCACGAAAGGGCACCGAGUCGGmU*mG*mC*mU (SEQ ID NO: 171).
[0330] In the case of a sgRNA, the guide sequences may further comprise a SpyCas9 sgRNA scaffold sequence. An example of a SpyCas9 sgRNA scaffold sequence is shown in the Table 38 below (SEQ ID NO: 102:GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGC, or SEQ ID NO: 107 or 109 - “Exemplary SpyCas9 sgRNA- 1”), included at the 3’ end of the guide sequence, and provided with the domains as shown in the table below. LS is lower stem. B is bulge. US is upper stem. Hl and H2 are hairpin 1 and hairpin 2, respectively. Collectively Hl and H2 are referred to as the hairpin region. A model of the structure containing both a guide sequence and a scaffold sequence is provided in FIG. 10A of WO2019237069, which is incorporated herein by reference.
[0331] The nucleotide sequence of Exemplary SpyCas9 sgRNA-1 may serve as a template sequence for specific chemical modifications, sequence substitutions and truncations.
[0332] In certain embodiments, the guide RNA is an sgRNA or a dgRNA, for example, and it optionally comprises a chemical modification. In some embodiments, the modified sgRNA comprises a guide sequence and a SpyCas9 sgRNA sequence, e.g., Exemplary SpyCas9 sgRNA-1. A guide RNA, such as an sgRNA, may include modifications on the 5’ end of the guide sequence or on the 3’ end of the SpyCas9 sgRNA sequence, such as, e.g., Exemplary SpyCas9 sgRNA-1 at one or more of the terminal nucleotides, e.g., at 1, 2, 3, or 4 of the nucleotides at the 3’ end or at the 5’ end. In certain embodiments, the modified nucleotide is selected from a 2’-O-methyl (2’-OMe) modified nucleotide, a 2’-O-(2-methoxyethyl) (2’-O-moe) modified nucleotide, a 2 ’-fluoro (2’-F) modified nucleotide, a phosphorothioate (PS) linkage between nucleotides, or an inverted abasic modified nucleotide;or a combination thereof. In certain embodiments, the modified nucleotide includes a 2’-0Me modified nucleotide. In certain embodiments, the modified nucleotide includes a PS linkage. In certain embodiments, the modified nucleotide includes a 2’-0Me modified nucleotide and a PS linkage.
[0333] In certain embodiments, using SEQ ID NO: 102:GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGC “Exemplary SpyCas9 sgRNA-1,” see WO2019237069, the contents of which are incorporated herein by reference). The portions of the Exemplary SpyCas9 sgRNA-1 and position numbering scheme are set forth in Table 38 below.
[0334] As an example, the Exemplary SpyCas9 sgRNA-1 further includes one or more of:A. a shortened hairpin 1 region, or a substituted and optionally shortened hairpin 1 region, wherein1. at least one of the following pairs of nucleotides are substituted in hairpin 1 with Watson-Crick pairing nucleotides: Hl-1 and Hl-12, Hl-2 and Hl-11, Hl-3 and Hl-10, or Hl-4 and Hl-9, and the hairpin 1 region optionally lacksa. any one or two of H 1 -5 through H 1 - 8,b. one, two, or three of the following pairs of nucleotides: Hl-1 and Hl-12, Hl-2 and Hl-11, Hl-3 and Hl-10, and Hl-4 and Hl-9, orc. 1-8 nucleotides of hairpin 1 region; or2. the shortened hairpin 1 region lacks 4-8 nucleotides, preferably 4-6 nucleotides; anda. one or more of positions Hl-1, Hl-2, or Hl-3 is deleted or substituted relative to Exemplary SpyCas9 sgRNA-1 (SEQ ID NO: 102) orb. one or more of positions Hl-6 through Hl-10 is substituted relative to Exemplary SpyCas9 sgRNA-l(SEQ ID NO: 102); or 3. the shortened hairpin 1 region lacks 5-10 nucleotides, preferably 5-6 nucleotides, and one or more of positions N18, Hl-12, or n is substituted relative to Exemplary SpyCas9 sgRNA-1 (SEQ ID NO: 102); or B. a shortened upper stem region, wherein the shortened upper stem region lacks 1-6 nucleotides and wherein the 6, 7, 8, 9, 10, or 11 nucleotides of the shortened upperstem region include less than or equal to 4 substitutions relative to Exemplary SpyCas9 sgRNA-1 (SEQ ID NO: 102); orC. a substitution relative to Exemplary SpyCas9 sgRNA-1 (SEQ ID NO: 102) at any one or more of LS6, LS7, US3, US10, B3, N7, N15, N17, H2-2 and H2-14, wherein the substituent nucleotide is neither a pyrimidine that is followed by an adenine, nor an adenine that is preceded by a pyrimidine; orD. an Exemplary SpyCas9 sgRNA-1 (SEQ ID NO: 102) with an upper stem region, wherein the upper stem modification comprises a modification to any one or more of US1-US12 in the upper stem region, wherein1. the modified nucleotide is optionally selected from a 2’-O-methyl (2’-OMe) modified nucleotide, a 2’-O-(2-methoxyethyl) (2’-O-moe) modified nucleotide, a 2’-fluoro (2’-F) modified nucleotide, a phosphorothioate (PS) linkage between nucleotides, an inverted abasic modified nucleotide, or a combination thereof; or2. the modified nucleotide optionally includes a 2’-OMe modified nucleotide. Other SpyCas9 spacer and scaffold sequences and chemical modification patterns are provided, for example in WQ2018107028, WO2019237069, and WO 2021119275, each of which is incorporated herein by reference.
[0335] In some embodiments, the guide RNA disclosed herein is a NmeCas9 guide RNA.
[0336] In some embodiments, the guide RNA is a NmeCas9 guide RNA that is a single guide RNA comprising a nucleotide sequence selected from SEQ ID NOs: 238-245, 248-255, 258-278, and 281-290, or a nucleotide sequence that is at least 85%, 90%, or 95% identical to SEQ ID NOs: 238-245, 248-255, 258-278, and 281-290. In some embodiments, the second guide comprises one or more internal polyethylene glycol (PEG) linker. In some embodiments, the NmeCas9 guide RNA comprising internal linkers comprise a sequence selected from SEQ ID NOs: 281-290.
[0337] In certain embodiments, using SEQ ID NO: 253 (“Exemplary NmeCas9 sgRNA-1,” as shown in Table 39) as an example, the Exemplary NmeCas9 sgRNA-1 includes: (A) A guide RNA (gRNA) comprising a guide region and a conserved region, the conserved region comprising one or more of: (a) a shortened repeat / anti-repeat region, wherein the shortened repeat / anti-repeat region lacks 2-24 nucleotides, wherein (i) one or more of nucleotides 37-48 and 53-64 is deleted and optionally one or more of nucleotides 37-64 is substituted relative to SEQ ID NO: 253; and (ii) nucleotide 36 is linked to nucleotide 65by at least 2 nucleotides; or (b) a shortened hairpin 1 region, wherein the shortened hairpin 1 lacks 2-10, optionally 2-8 nucleotides, wherein (i) one or more of nucleotides 82-86 and 91-95 is deleted and optionally one or more of positions 82-96 is substituted relative to SEQ ID NO: 253; and (ii) nucleotide 81 is linked to nucleotide 96 by at least 4 nucleotides; or (c) a shortened hairpin 2 region, wherein the shortened hairpin 2 lacks 2-18, optionally 2-16 nucleotides, wherein (i) one or more of nucleotides 113-121 and 126-134 is deleted and optionally one or more of nucleotides 113-134 is substituted relative to SEQ ID NO: 253; and (ii) nucleotide 112 is linked to nucleotide 135 by at least 4 nucle li des; wherein one or both nucleotides 144-145 are optionally deleted relative to SEQ ID NO: 253; wherein optionally at least 10 nucleotides are modified nucleotides.
[0338] Exemplary unmodified conserved portion nucleotide sequences are provided in Table 2A.
[0339] In the case of a sgRNA, the guide sequences may be integrated into one of the following exemplary modified conserved portion motifs as shown in Table 2B.
[0340] In certain embodiments, the guide sequence is 20-25 nucleotides in length ((N)20-25), wherein each nucleotide may be independently modified. In certain embodiments, each of nucleotides 1-3 of the 5’ end of the guide is independently modified. In certain embodiments, each of nucleotides 1-3 of the 5’ end of the guide is independently modified with a 2’-0Me modification. In certain embodiments, each of nucleotides 1-3 of the 5’ end of the guide is independently modified with a phosphorothioate linkage to the adjacent nucleotide residue. In certain embodiments, each of nucleotides 1-3 of the 5’ end of the guide is independently modified with a 2’-0Me modification and a phosphorothioate linkage to the adjacent nucleotide residue.
[0341] In the case of a sgRNA, modified guide sequences may be integrated into one of the following exemplary modified conserved portion motifs as shown in Table 2B.
[0342] In some embodiments, the guide RNA comprises a sgRNA comprising a guide region and a conserved portion of an sgRNA, for example, the conserved portion of sgRNA shown as Exemplary NmeCas9 sgRNA- 1 or the conserved portions of the guide RNAs shown in Table 2A-3B and throughout the specification.
[0343] In some embodiments, the sgRNA comprises Exemplary NmeCas9 sgRNA- 1 or the modified versions thereof provided herein, or a version as provided in Table 2B. Each N is independently modified or unmodified. In certain embodiments, in the absence of an indication of a modification, the nucleotide is an unmodified RNA nucleotide residue, i.e., a ribose sugar and a phosphodiester backbone.Table 2A: Exemplary Unmodified NmeCas9 Scaffold SequencesTable 2B: Exemplary Unmodified NmeCas9 Guide RNA SequencesTable 2C: Exemplary Modified NmeCas9 Scaffold Sequenceswherein “m” indicates a 2’-0-Me modification, and a indicates a phosphorothioate linkage between nucleotides, and no modification in the context of a modified sequence indicates an RNA (2’-OH) and a phosphorothioate linkage.Table 2D. Exemplary Modified NmeCas9 Guide RNA Sequences
[0344] In certain embodiments, the guide sequence is a chemically modified sequence. In certain embodiments, the chemically modified guide sequence is (mN*)3(N) 17-22. In certain embodiments, the guide sequence is (mN*)3(N)21, i.e., mN*mN*mN*NNNNNNNNNNNNNNNNNNNNN. In certain embodiments, each N of the (N)17-22 or the (N)21 is unmodified. In certain embodiments, the each N in the (N)18-21 or the (N)21 is independently modified, e.g. independently modified with a 2’-O-methyl modification.Linker containing guide RNAs
[0345] In certain embodiments, the guide RNA comprises one or more internal linkers. As used herein, “internal linker” describes a non-nucleotide segment joining two nucleotides within a guide RNA. If the guide RNA contains a spacer region, the internal linker is located outside of the spacer region (e.g., in the scaffold or conserved region of the guide RNA). For Type V guides, it is understood that the last hairpin is the only hairpin in the structure, i.e., the repeal -anti -repeat region. The length of an internal linker may be dependent on, for example, the number of nucleotides replaced by the linker and the position of the linker in the guide RNA. Internal linkers and their use in the context of guide RNA are provided in WO2022261292., the contents of which are hereby incorporated by reference in its entirety.
[0346] Guide RNAs disclosed herein may comprise an internal linker. In general, any internal linker compatible with the function of the guide RNA may be used. It may be desirable for the linker to have a degree of flexibility. In some embodiments, the internal linker comprises at least two, three, four, five, six, or more on-pathway single bonds. A bond is on-pathway if it is part of the shortest path of bonds between the two nucleotides whose 5’ and 3’ positions are connected to the linker.
[0347] As used herein the length of the internal linker can be defined by its bridging length. The “bridging length” of an internal linker as used herein refers to the distance or number of atoms in the shortest chain of atoms on the pathway from the first atom of the linker (bound to a 3’ substituent, such as an oxygen or phosphate, of the preceding nucleotide to the last atom of the linker (bound to a 5’ substituent, such as an oxygen or phosphate) of the following nucleotide) (e.g., from ~ to # in the structure of Formula (I) described below).Approximate predicted bridging lengths for various linkers are provided in a table below.
[0348] Exemplary predicted linker lengths by number of atoms, number of ethylene glycol units, approximate linker length in Angstroms on the assumption that an ethylene glycol monomer is about 3.7 Angstroms, and suitable location for substitution of at least the entire loop portion of a hairpin structure are provided in Table 3 below. Substitution of two nucleotides requires a linker length of at least about 11 Angstroms. Substitution of at least 3 nucleotides requires a linker length of at least about 16 Angstroms.
[0349] In some embodiments, the internal linker has a bridging length of about 3 Angstroms to about 37 Angstroms. In some embodiments, the internal linker has a bridging length of about 6 Angstroms to about 37 Angstroms. In some embodiments, the internal linker has a bridging length of about 7 Angstroms to about 22 Angstroms.
[0350] In some embodiments, the internal linker comprises 1-10 ethylene glycol subunits covalently linked to each other. In some embodiments, the internal linker comprises at least two ethylene glycol subunits covalently linked to each other. In some embodiments, the internal linker comprises at least six ethylene glycol subunits covalently linked to each other.
[0351] In some embodiments, the internal linker comprises a polyethylene glycol (PEG) linker. In some embodiments, the internal linker comprises a PEG linker having from 1 to 10 ethylene glycol units. In some embodiments, the internal linker comprises a PEG linker having from 3 to 6 ethylene glycol units. In some embodiments, the internal linker comprises a PEG linker having 3 ethylene glycol units. In some embodiments, the internal linker comprises a PEG linker having 6 ethylene glycol units.Table 3
[0352] In some embodiments, the internal linker comprises a structure of formula (I):—L0-L1-L2-#(I)wherein:~ indicates a bond to a 3’ substituent of the preceding nucleotide;# indicates a bond to a 5’ substituent of the following nucleotide;L0 is null or C1-3 aliphatic;LI is -[E1-(R1)]m-, whereeach R1is independently a C1-5 aliphatic group, optionally substituted with 1 or 2 E2, each E1and E2are independently a hydrogen bond acceptor, or are each independently chosen from cyclic hydrocarbons, and heterocyclic hydrocarbons, andeach m is 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10; andL2 is null, Ci-3 aliphatic, or is a hydrogen bond acceptor.
[0353] In some embodiments, LI comprises one or more -CH2CH2O-, -CH2OCH2-, or -OCH2CH2- units (“ethylene glycol subunits”). In some embodiments, the number of -CH2CH2O-, -CH2OCH2-, or -OCH2CH2- units is in the range of 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10.
[0354] In some embodiments, m is 1, 2, 3, 4 or 5. In some embodiments, m is 1, 2, or 3. In some embodiments, m is 6, 7, 8, 9, or 10.
[0355] In some embodiments, L0 is null. In some embodiments, L0 is -CH2- or -CH2CH2-.
[0356] In some embodiments, L2 is null. In some embodiments, L2 is -O-, -S-, or C1-3 aliphatic. In some embodiments, L2 is -O-. In some embodiments, L2 is -S-. In some embodiments, L2 is -CH2- or -CH2CH2-.
[0357] In the tables herein, LI and L2, are optionally, C9 and C18, respectively as follows:
[0358] In certain embodiments, the internal linker has a bridging length of about 3-30 atoms, optionally 12-21 atoms, and the linker substitutes for at least 2 nucleotides of the guide RNA. In certain embodiments, the internal linker has a bridging length of about 6-18 atoms, optionally about 6-12 atoms, and the linker substitutes for at least 2 nucleotides of the guide RNA. In certain embodiments, the internal linker substitutes for 2-12 nucleotides.
[0359] In some embodiments, the guide RNA comprises a nucleic acid sequence of SEQ ID NO: 102, including modifications disclosed elsewhere herein. Table 4 shows various embodiments of the guide RNA structures and species with possible number of internal linkers and positions.Table 4.
[0360] In certain embodiments, the internal linker is in a repeat-anti-repeat region of the guide RNA. In certain embodiments, the internal linker substitutes for at least 4 nucleotides of the repeat-anti-repeat region of the guide RNA. In certain embodiments, the internal linker substitutes for up to 28 nucleotides in the repeat / anti-repeat region of the guide RNA. In certain embodiments, the internal linker is in a hairpin between a first portion and a second portion of the repeat / anti-repeat region, wherein the first portion and the second portion together form a duplex portion. In certain embodiments, the internal linker substitutes for the loop in the repeat-anti-repeat region of a Spy Cas9 guide RNA, corresponding to nucleotides 13-16 in SEQ ID NO: 102.
[0361] In certain embodiments, the internal linker is in a nexus region of the guide RNA. In certain embodiments, the internal linker substitutes for 2, 3, or 4 nucleotides of the nexus region of the guide RNA. In certain embodiments, the internal linker substitutes for the loop in the nexus region of a Spy Cas9 guide RNA corresponding to nucleotides 33-36 of SEQ ID NO: 102.
[0362] In certain embodiments, the internal linker is in a hairpin region of the guide RNA. In certain embodiments, the internal linker is in a hairpin region between a first portion of the guide RNA and a second portion of the guide RNA, wherein the first portion and the second portion together form a duplex region. In certain embodiments, the internal linkersubstitutes for a hairpin of the guide RNA. In certain embodiments, the internal linker substitutes for at least 4 nucleotides of the hairpin region of the guide RNA.
[0363] In certain embodiments, the internal linker is in a hairpin 1 region. In certain embodiments, the internal linker substitutes for the loop in the hairpin 1 region of a Spy Cas9 guide RNA, corresponding to nucleotides 53-56 in SEQ ID NO: 102 (see Table 38).
[0364] In certain embodiments, wherein the guide RNA is a single guide RNA (sgRNA) comprising a guide region and a conserved portion 3’ to the guide region, the conserved portion comprises a repeat / anti-repeat region, a nexus region, a hairpin 1 region, and a hairpin 2 region, and comprises at least one of: (1) a first internal linker substituting for at least 2 nucleotides of an upper stem region of the repeat / anti-repeat region; (2) a second internal linker substituting for 1 or 2 nucleotides of the nexus region; and (3) a third internal linker substituting for at least 2 nucleotides of the hairpin 1 region. In certain embodiments, the first internal linker has a bridging length of about 9-30 atoms, optionally about 15-21 atoms. In certain embodiments, the first internal linker substitutes for a loop, or part thereof, of the upper stem region. In certain embodiments, the first internal linker substitutes for a loop, or part thereof, of the upper stem region. In certain embodiments, the first internal linker subslilules for the loop and the stem, or part thereof, of the upper stem region. In certain embodiments, the first internal linker subslilules for all of the nucleotides consliluling the loop of the upper stem region. In certain embodiments, the first internal linker substitutes for all of the nucleotides constituting the loop and the stem of the upper stem region. In certain embodiments, the second internal linker has a bridging length of about 6-18 atoms, optionally about 6-12 atoms. In certain embodiments, the third internal linker has a bridging length of about 9-30, optionally about 12-21 atoms. In certain embodiments, the third internal linker substitutes for a loop, or part thereof, of the hairpin 1. In certain embodiments, the third internal linker substitutes for the loop and the stem, or part thereof, of the hairpin 1. In certain embodiments, the third internal linker substitutes for all of the nucleotides consliluling the loop of the hairpin 1. In certain embodiments, the third internal linker substitutes for all of the nucleotides constituting the loop and the stem of the hairpin 1. In certain embodiments, the hairpin 2 region of the sgRNA does not contain any internal linker.
[0365] In some embodiments, wherein the guide RNA is a SpyCas9 guide RNA, the SpyCas9 guide RNA scaffold sequence comprises an internal linker. In some embodiments, the SpyCas9 guide RNA scaffold sequence comprises (i) an internal linker in or substitutes for the upper stem region; (ii) an internal linker in a nexus region; or (iii) an internal linker in the hairpin 1 region. In certain embodiments, the internal linker in the upper stem regionsubstitutes for US1 through US 12 relative to SEQ ID NO: 102. In certain embodiments, the internal linker in the hairpin 1 substitutes for Hl-5 through Hl-8 relative to SEQ ID NO: 102. In certain embodiments, nucleotides 29-40 (US1 through US12 of the upper stem region) are substituted for the first internal linker relative to SpyCas9 guide RNA reference sequence; and nucleotides 70-78 (HP1-HP10 of the hairpin 1 region) are substituted for the second internal linker relative to SpyCas9 guide RNA reference sequence. In certain embodiments, the guide RNA comprises the sequence of GUUUUAGA(L3)AAGUUAAAAUAAGGCUAGUCCGUUAUCAC(L1)GGGCACCGAGU CGGUGCU (SEQ ID NO: 220), wherein LI denotes that the linker is S18, and L3 denotes that the linker is S6. In certain embodiments, “(LI)” denotes an S18 internal linker, “(L2)” denotes an S9 internal linker, “(L3)” denotes an S6 internal linker, “(L4)” denotes an S3 internal linker, a “(dS)” indicates an abasic site having 1’, 2 ’-dideoxyribose modification (e.g., dSpacer from IDT). As used herein, “S3” indicates a 5’-O-CH2-CH2-CH2-O-PO3-3’ internal linker, “S6” indicates a 5’-O-(CH2O)2-PO3- 3’ internal linker, “S9” indicates a 5’-O-(CH2O)3-PO3- 3’ internal linker, “SI 8” indicates a 5’-O-(CH2O)6-PO3- 3’ internal linker, and no modification in the context of a modified sequence indicates an RNA (2’-OH) and phosphodiesterase linkage to the 3’ nucleotide when one is present.Table 5. Exemplary SpyCas9 Scaffold Sequences and Guide RNAs Comprising Linkersfollows: wherein “m” indicates a 2’-0-Me modification, a indicates a phosphorothioate linkage between nucleotides, and within the individually indicated nucleotides, no modification indicates an RNA (2’ -OH) with a phosphodiesterase backbone.
[0367] The shortened NmeCas9 guide RNA may comprise internal linkers disclosed herein. In some embodiments, the internal linker comprises a polyethylene glycol (PEG) linker. The guide RNAs comprising an internal linker disclosed herein comprise one of the structures / modification patterns disclosed in WO2022 / 261292, the contents of which are hereby incorporated by reference in its entirety. Further exemplary NmeCas9 guide RNAs comprising linkers are provided in Table 6.
[0368] In some embodiments, the shortened NmeCas9 guide RNA comprising internal linkers may be chemically modified as shown in Table 6. In certain embodiments, the guide sequence is a chemically modified sequence as shown in Table 6.Table 6: Exemplary Modified NmeCas9 Scaffold and Guide RNA Sequences00369] In certain embodiments, an sgRNA comprises an Exemplary SpyCas9 sgRNA- 1 (e.g., SEQ ID NO: 102, 107, or 109 shown in Table 38). In certain embodiments, the Exemplary SpyCas9 sgRNA- 1 does not include a 3’ tail. In certain embodiments, the Exemplary SpyCas9 sgRNA-1 may include a 3’ tail, e.g., a 3’ tail of 1, 2, 3, 4nucleotides, e.g., a tail of 1 nucleotide, e.g., a tail of 1 uridine nucleotide. It is understood that a 3’ tail is not a 3’ extension, which includes a template and a DRS, is distinct in structure and function from a 3’ tail. The nucleotides and nucleotide modifications discussed in relation to the Exemplary SpyCas9 sgRNA-1 spacer and scaffold may or may not be tolerated in a 3’ extension as discussed further below. In certain embodiments, the tail includes one or more modified nucleotides. In certain embodiments, the modified nucleotide is selected from a 2’-O-methyl (2’-OMe) modified nucleotide, a2’-O-(2-methoxyethyl) (2’-O-moe) modified nucleotide, a 2’-fluoro (2’-F) modified nucleotide, a phosphorothioate (PS) linkage between nucleotides; or a combination thereof. In certain embodiments, the modified nucleotide includes a 2’-OMemodified nucleotide. In certain embodiments, the modified nucleotide includes a PS linkage between nucleotides. In certain embodiments, the modified nucleotide includes a 2’-0Me modified nucleotide and a PS linkage between nucleotides.
[0370] In certain embodiments, the hairpin region includes one or more modified nucleotides. In certain embodiments, the modified nucleotide is selected from a 2’-O-methyl (2’-OMe) modified nucleotide, a2’-O-(2-methoxyethyl) (2’-O-moe) modified nucleotide, a 2 ’-fluoro (2’-F) modified nucleotide, a phosphorothioate (PS) linkage between nucleotides; or a combination thereof. In certain embodiments, the modified nucleotide includes a 2’-OMe modified nucleotide.
[0371] In certain embodiments, the upper stem region includes one or more modified nucleotides. In certain embodiments, the modified nucleotide selected from a 2’-O-methyl (2’-OMe) modified nucleotide, a 2’-O-(2-methoxyethyl) (2’-O-moe) modified nucleotide, a 2’-fluoro (2’-F) modified nucleotide, a phosphorothioate (PS) linkage between nucleotides; or a combination thereof. In certain embodiments, the modified nucleotide includes a 2’-OMe modified nucleotide.
[0372] In certain embodiments, the Exemplary SpyCas9 sgRNA-1 comprises one or more YA dinucleotides, wherein Y is a pyrimidine, wherein the YA dinucleotide includes a modified nucleotide. In certain embodiments, the modified nucleotide selected from a 2’-O-methyl (2’-OMe) modified nucleotide, a2’-O-(2-methoxyethyl) (2’-O-moe) modified nucleotide, a 2’ -fluoro (2’-F) modified nucleotide, a phosphorothioate (PS) linkage between nucleotides,; or a combination thereof. In certain embodiments, the modified nucleotide includes a 2’-OMe modified nucleotide.
[0373] In certain embodiments, the Exemplary SpyCas9 sgRNA-1 comprises one or more YA dinucleotides, wherein Y is a pyrimidine, wherein the YA dinucleotide includes a sequence substituted nucleotide, wherein the pyrimidine is substituted for a purine. In certain embodiments, when the pyrimidine forms a Watson-Crick base pair in the single guide, the Watson-Crick based nucleotide of the sequence substituted pyrimidine nucleotide is substituted to maintain Watson-Crick base pairing.
[0374] In some embodiments, the Exemplary SpyCas9 sgRNA-1 is chemically modified. An Exemplary SpyCas9 sgRNA-1, SpyCas9 spacer and scaffold, or more simply guide RNA, comprising one or more modified nucleosides or nucleotides is called a “modified” guide RNA or “chemically modified” guide RNA, to describe the presence of one or more non-naturally or naturally occurring components or configurations that are used instead of or in addition to the canonical A, G, C, and U residues. In some embodiments, amodified guide RNA is synthesized with a non-canonical nucleoside or nucleotide, is here called “modified.” Modified nucleosides and nucleotides can include one or more of: (i) alteration, e.g., replacement, of one or both of the non-linking phosphate oxygens or of one or more of the linking phosphate oxygens in the phosphodiester backbone linkage (an exemplary backbone modification); (ii) alteration, e.g., replacement, of a constituent of the ribose sugar, e.g., of the 2' hydroxyl on the ribose sugar (an exemplary sugar modification); (iii) modification or replacement of a naturally occurring nucleobase, including with a non-canonical nucleobase (an exemplary base modification); and (iv) modification of the 3' end or 5' end of the oligonucleotide to provide exonuclease stability, e.g., with 2’ O-me, 2’ halide, or 2’ deoxy substituted ribose; or replacement of phosphodiester with phosphorothioate.
[0375] Chemical modifications such as those listed above can be combined to provide modified Exemplary SpyCas9 sgRNA-1 comprising nucleosides and nucleotides (collectively “residues”) that can have two, three, four, or more modifications. For example, a modified residue can have a modified sugar and a modified nucleobase. In some embodiments, modified guide RNAs comprise at least one modified residue at or near the 5' end of the RNA. In some embodiments, modified guide RNAs comprise at least one modified residue at or near the 3' end of the RNA.
[0376] In some embodiments, the Exemplary SpyCas9 sgRNA-1 comprises one, two, three or more modified residues. In some embodiments, at least 10% (e.g., at least 10%, 15%, preferably at least 20%, 25%, 30%, 35%, 40%, 45%, or 50%) of the positions in a modified guide RNA are modified nucleosides or nucleotides. In some embodiments, at least 10% of the positions in the modified guide RNA are modified nucleotides or nucleosides. In some embodiments, at least 10% of the positions in the modified guide RNA are modified nucleotides or nucleosides. In some embodiments at least 15% of the positions in the modified guide RNA are modified nucleotides or nucleosides. In some embodiments preferably at least 20% of the positions in the modified guide RNA are modified nucleotides or nucleosides. In some embodiments, no more than 65% of the positions in the modified guide RNA are modified nucleotides. In some embodiments, no more than 55% of the positions in the modified guide RNA are modified nucleotides. In some embodiments, no more than 50% of the positions in the modified guide RNA are modified nucleotides. In some embodiments, 10-80% of the positions in the modified guide RNA are modified nucleotides. In some embodiments, 20-70% of the positions in the modified guide RNA are modified nucleotides. In some embodiments, 20-50% of the positions in the modified guide RNA are modified nucleotides and the nuclease is a SpyCas9 nuclease.
[0377] Unmodified nucleic acids can be prone to degradation by, e.g., intracellular nucleases or those found in serum. For example, nucleases can hydrolyze nucleic acid phosphodiester bonds. Accordingly, in one aspect the guide RNAs described herein can contain one or more modified nucleosides or nucleotides, e.g., to introduce stability toward intracellular or serum-based nucleases. In some embodiments, the modified guide RNA molecules described herein can exhibit a reduced innate immune response when introduced into a population of cells, both in vivo and ex vivo. The term “innate immune response” includes a cellular response to exogenous nucleic acids, including single stranded nucleic acids, which involves the induction of cytokine expression and release, particularly the interferons, and cell death.
[0378] In some embodiments of a backbone modification, the phosphate group of a modified residue can be modified by replacing one or more of the oxygens with a different substituent. Further, the modified residue, e.g., modified residue present in a modified nucleic acid, can include the replacement of an unmodified phosphate moiety with a modified phosphate group as described herein. In some embodiments, the backbone modification of the phosphate backbone can include alterations that result in either an uncharged linker or a charged linker with unsymmetrical charge distribution.
[0379] Examples of modified phosphate groups include, phosphorothioate, borano phosphate esters, methyl phosphonates, phosphoroamidates, phosphodithioate, alkyl or aryl phosphonates and phosphotriesters. In certain embodiments, the modified phosphate group is a phosphorothioate group. The phosphorous atom in an unmodified phosphate group is achiral. However, replacement of one of the non-bridging oxygens with one of the above atoms or groups of atoms can render the phosphorous atom chiral. The stereogenic phosphorous atom can possess either the “R” configuration (herein Rp) or the “S” configuration (herein Sp). The backbone can also be modified by replacement of a bridging oxygen, (i.e., the oxygen that links the phosphate to the nucleoside), with nitrogen (bridged phosphoroamidates), sulfur (bridged phosphorothioates) and carbon (bridged methylenephosphonates). The replacement can occur at either linking oxygen or at both of the linking oxygens.
[0380] The phosphate group can be replaced by non-phosphorus containing connectors in certain backbone modifications, e.g., an amide linkage. In some embodiments, the charged phosphate group can be replaced by a neutral moiety. Examples of moieties which can replace the phosphate group can include, without limitation, e.g., methyl phosphonate, carboxymethyl, carbamate, amide, thioether. Further examples of moieties which can replace the phosphate group can include, without limitation, e.g., ethylene oxide linker, sulfonate, sulfonamide,thioformacetal, formacetal, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo and methyleneoxymethylimino.
[0381] Scaffolds that can mimic nucleic acids can also be constructed wherein the phosphate linker and ribose sugar are replaced by nuclease resistant nucleoside or nucleotide surrogates. Such modifications may comprise backbone and sugar modifications. In some embodiments, the nucleobases can be tethered by a surrogate backbone. Examples can include, without limitation, the morpholino, cyclobutyl, pyrrolidine and peptide nucleic acid (PNA) nucleoside surrogates.
[0382] The modified nucleosides and modified nucleotides can include one or more modifications to the sugar group, i.e. at sugar modification. For example, the 2' hydroxyl group (OH) can be modified, e.g. replaced with a number of different “oxy” or “deoxy” substituents. In some embodiments, modifications to the 2' hydroxyl group can enhance the stability of the nucleic acid since the hydroxyl can no longer be deprotonated to form a 2'-alkoxide ion.
[0383] Examples of 2' hydroxyl group modifications can include alkoxy or aryloxy (OR, wherein “R” can be, e.g., alkyl, cycloalkyl, aryl, aralkyl, heteroaryl or a sugar); polyethyleneglycols (PEG), O(CH2CH2O)nCH2CH2OR wherein R can be, e.g., H or optionally substituted alkyl, and n can be an integer from 0 to 20 (e.g., from 0 to 4, from 0 to 8, from 0 to 10, from 0 to 16, from 1 to 4, from 1 to 8, from 1 to 10, from 1 to 16, from 1 to 20, from 2 to 4, from 2 to 8, from 2 to 10, from 2 to 16, from 2 to 20, from 4 to 8, from 4 to 10, from 4 to 16, and from 4 to 20). In some embodiments, the 2' hydroxyl group modification can be 2'-O-Me. In some embodiments, the 2' hydroxyl group modification can be a 2'-fluoro modification, which replaces the 2' hydroxyl group with a fluoride. In some embodiments, the 2' hydroxyl group modification can include “locked” nucleic acids (LNA) in which the 2' hydroxyl can be connected, e.g., by a Cl-6 alkylene or Cl-6 heteroalkylene bridge, to the 4' carbon of the same ribose sugar, where exemplary bridges can include methylene, propylene, ether, or amino bridges; 0-amino (wherein amino can be, e.g., NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, or polyamino) and aminoalkoxy, O(CH2)n-amino, (wherein amino can be, e.g., NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, or polyamino). In some embodiments, the 2' hydroxyl group modification can include "unlocked" nucleic acids (UNA) in which the ribose ring lacks the C2'-C3' bond. In some embodiments, the 2' hydroxyl group modification can include the methoxyethyl group (MOE), (OCH2CH2OCH3, e.g., a PEG derivative). 2' modifications can include hydrogen (i.e.deoxyribose sugars); halo (e.g., bromo, chloro, fluoro, or iodo); amino (wherein amino can be, e.g., NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or amino acid); NH(CH2CH2NH)nCH2CH2- amino (wherein amino can be, e.g., as described herein), -NHC(O)R (wherein R can be, e.g., alkyl, cycloalkyl, aryl, aralkyl, heteroaryl or sugar), cyano; mercapto; alkyl-thio-alkyl; thioalkoxy; and alkyl, cycloalkyl, aryl, alkenyl and alkynyl, which may be optionally substituted with e.g., an amino as described herein.
[0384] The sugar modification can comprise a sugar group which may also contain one or more carbons that possess the opposite stereochemical configuration than that of the corresponding carbon in ribose. Thus, a modified nucleic acid can include nucleotides containing e.g., arabinose, as the sugar. The modified nucleic acids can also include abasic sugars. These abasic sugars can also be further modified at one or more of the constituent sugar atoms. The modified nucleic acids can also include one or more sugars that are in the L form, e.g., L- nucleosides. As used herein, a single abasic sugar is not understood to result in a discontinuity of a duplex.
[0385] In certain embodiments, 2’ modifications, include, for example, modifications include 2’-OMe, 2’-F, 2’-H, LNA, optionally 2’-O-Me.
[0386] The modified nucleosides and modified nucleotides described herein, which can be incorporated into a modified nucleic acid, can include a modified base, also called a nucleobase. Examples of nucleobases include, but are not limited to, adenine (A), guanine (G), cytosine (C), and uracil (U). These nucleobases can be modified or wholly replaced to provide modified residues that can be incorporated into modified nucleic acids. The nucleobase of the nucleotide can be independently selected from a purine, a pyrimidine, a purine analog, or pyrimidine analog. In some embodiments, the nucleobase can include, for example, naturally occurring and synthetic derivatives of a base.
[0387] In some embodiments, the guide RNAs, spacer and scaffold disclosed herein comprise one or more modifications at a YA site (also referred to as a YA modification). In some embodiments, the guide RNAs, spacer, and scaffold disclosed herein comprise one or more of the YA modifications disclosed in WO2019 / 237069, the contents of which are incorporated by reference herein.
[0388] In embodiments employing a dual guide RNA, each of the crRNA and the tracr RNA can contain modifications. Such modifications may be at one or both ends of the crRNA or tracr RNA. In embodiments comprising an sgRNA, one or more residues at one or both ends of the sgRNA may be chemically modified, or internal nucleosides may be modified, orthe sgRNA may be chemically modified throughout. Certain embodiments comprise a 5' end modification. Certain embodiments comprise a 3' end modification. Certain embodiments comprise a 5’ end modification and a 3’ end modification.
[0389] In some embodiments, the guide RNAs, spacer and scaffold, disclosed herein comprise one of the modification patterns disclosed in WO2018 / 107028, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, the guide RNAs disclosed herein comprise one of the structures / modification patterns disclosed in US20170114334, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, the guide RNAs disclosed herein comprise one of the structures / modification patterns disclosed in WO2017 / 136794, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, the guide RNAs disclosed herein comprise one of the structures / modification patterns disclosed in WO2019 / 237069, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, the guide RNAs disclosed herein comprise one of the structures / modification patterns disclosed in WO2021 / 119275, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, the guide RNAs disclosed herein comprise one of the structures / modification patterns disclosed in WO2023081687A1, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, the guide RNAs disclosed herein comprise one of the structures / modification patterns disclosed in WO2022 / 261292, the contents of which are hereby incorporated by reference in their entirety.Lipid-based Delivery
[0390] The following section provides additional features of lipid-based delivery compositions, including lipid nanoparticles (LNPs) and lipoplexes, for the nucleic acid or nucleic acids encoding the same. In some embodiments, the nucleic acid or nucleic acids encoding the same is delivered to the cell via at least one lipid nanoparticle (LNP).
[0391] Lipid nanoparticles (LNPs) are a well- known means for delivery of nucleotide and protein cargo, and may be used for delivery of the guide RNAs (e.g., sgRNAs, dgRNAs, or crRNAs), compositions, or pharmaceutical formulations disclosed herein. In some embodiments, the LNPs deliver nucleic acid, protein, or nucleic acid together with protein. As used herein, lipid nanoparticle (LNP) refers to a particle that comprises a plurality of (i.e., more than one) lipid molecules physically associated with each other by intermolecular forces.Any LNP known to those of skill in the art to be capable of delivering nucleotides to subjects may be utilized..
[0392] In certain embodiments, an LNP has a diameter of about 50-175 nm, or a population of the LNP with an average diameter of about 50-175 nm as measured by dynamic light scatering. In preferred embodiments, an LNP composition has a diameter of 75-150 nm.
[0393] In embodiments, the average diameter is a Z-average diameter. In certain embodiments, the Z-average diameter is measured by dynamic light scattering (DLS) using methods known in the art. For example, average particle size and polydispersity can be measured by dynamic light scatering (DLS) using a Malvern Zetasizer DLS instrument. LNP samples are diluted with PBS buffer prior to being measured by DLS. Z-average diameter and number average diameter along with a polydispersity index (pdi) can be determined. The Z average is the intensity weighted mean hydrodynamic size of the ensemble collection of particles. The number average is the particle number weighted mean hydrodynamic size of the ensemble collection of particles. A Malvern Zetasizer instrument can also be used to measure the zeta potential of the LNP using methods known in the art.
[0394] In some embodiments, Dynamic Light Scattering (“DLS”) may be used to characterize the polydispersity index (PDI) and size of the LNPs of the present disclosure. DLS measures the scattering of light that results from subjecting a sample to a light source. PDI, as determined from DLS measurements, represents the distribution of particle size (around the mean particle size) in a population, with a perfectly uniform population having a PDI of zero.
[0395] In some embodiments, the LNPs disclosed herein have a PDI from about 0.005 to about 0.75. In some embodiments, the LNPs disclosed herein have a PDI from about 0.005 to about 0.1. In some embodiments, the LNPs disclosed herein have a PDI from about 0.005 to about 0.09, about 0.005 to about 0.08, about 0.005 to about 0.07, or about 0.006 to about 0.05. In some embodiments, the LNP have a PDI from about 0.01 to about 0.5. In some embodiments, the LNP have a PDI from about zero to about 0.4. In some embodiments, the LNP have a PDI from about zero to about 0.35. In some embodiments, the LNP PDI may range from about zero to about 0.3. In some embodiments, the LNP have a PDI that may range from about zero to about 0.25. In some embodiments, the LNP PDI may range from about zero to about 0.2. In some embodiments, the LNP have a PDI from about zero to about 0.05. In some embodiments, the LNP have a PDI from about zero to about 0.01. In some embodiments, the LNP have a PDI less than about 0.01, 0.02, 0.05, 0.08, 0.1, 0.15, 0.2, or 0.4.
[0396] LNP size may be measured by various analytical methods known in the art. In some embodiments, LNP size may be measured using Asymmetric-Flow Field Flow Fractionation - Multi-Angle Light Scattering (AF4-MALS). In certain embodiments, LNP size may be measured by separating particles in the composition by hydrodynamic radius, followed by measuring the molecular weights, hydrodynamic radii and root mean square radii of the fractionated particles. In some embodiments, LNP size and particle concentration may be measured by nanoparticle tracking analysis (NTA, Malvern Nanosight). In certain embodiments, LNP samples are diluted appropriately and injected onto a microscope slide. A camera records the scattered light as the particles are slowly infused through field of view. After the movie is captured, the Nanoparticle Tracking Analysis processes the movie by tracking pixels and calculating a diffusion coefficient. This diffusion coefficient can be translated into the hydrodynamic radius of the particle. Such methods may also count the number of individual particles to give particle concentration. In some embodiments, LNP size, morphology, and structural characteristics may be determined by cryo-electron microscopy (“cryo-EM”).
[0397] The LNPs disclosed herein can have a size (e.g. Z-average diameter or numberaverage diameter) of about 1 to about 250 nm. In some embodiments, the LNPs have a size of about 10 to about 200 nm. In further embodiments, the LNPs have a size of about 20 to about 150 nm. In some embodiments, the LNPs have a size of about 50 to about 150 nm or about 70 to 130 nm. In some embodiments, the LNPs have a size of about 50 to about 100 nm. In some embodiments, the LNPs have a size of about 50 to about 120 nm. In some embodiments, the LNPs have a size of about 60 to about 100 nm. In some embodiments, the LNPs have a size of about 75 to about 150 nm. In some embodiments, the LNPs have a size of about 75 to about 120 nm. In some embodiments, the LNPs have a size of about 75 to about 100 nm. In some embodiments, the LNPs have a size of about 50 to about 145 nm, about 50 to about 120 nm, about 50 to about 120 nm, about 50 to about 115 nm, about 50 to about 100 nm, about 60 to about 145 nm, about 60 to about 120 nm, about 60 to about 115 nm, or about 60 to about 100 nm. In some embodiments, the LNPs have a size of less than about 145 nm, less than about 120 nm, less than about 115 nm, less than about 100 nm, or less than about 80 nm. In some embodiments, the LNPs have a size of greater than about 50 nm or greater than about 60 nm. In some embodiments, the particle size is a Z-average particle size. In some embodiments, the particle size is a number-average particle size. In some embodiments, the particle size is the size of an individual LNP. Unless indicated otherwise, all sizes referred to herein are the average sizes (diameters) of the fully formed nanoparticles, as measured by dynamic lightscattering on a Malvern Zetasizer or Wyatt NanoStar. The nanoparticle sample is diluted in phosphate buffered saline (PBS) so that the count rate is approximately 200-400 kcps.
[0398] LNPs are formed by precise mixing a lipid component (e.g., in ethanol) with an aqueous nucleic acid component and LNPs are uniform in size. Lipoplexes are particles formed by bulk mixing the lipid and nucleic acid components and are between about 100nm and 1 micron in size. In certain embodiments the lipid nucleic acid assemblies are LNPs. As used herein, a “lipid nucleic acid assembly” comprises a plurality of (i.e., more than one) lipid molecules physically associated with each other by intermolecular forces. A lipid nucleic acid assembly may comprise a bioavailable lipid having a pKa value of < 7.5 or < 7. The lipid nucleic acid assemblies are formed by mixing an aqueous nucleic acid-containing solution with an organic solvent-based lipid solution, e.g., 100% ethanol. Suitable solutions or solvents include or may contain: water, PBS, Tris buffer, NaCl, citrate buffer, ethanol, chloroform, diethylether, cyclohexane, tetrahydrofuran, methanol, isopropanol. A pharmaceutically acceptable buffer may optionally be comprised in a pharmaceutical formulation comprising the lipid nucleic acid assemblies, e.g., for an ex vivo ACT therapy. In some embodiments, the aqueous solution comprises an RNA, such as an mRNA or a guide RNA. In some embodiments, the aqueous solution comprises an mRNA encoding an RNA-guided DNA binding agent, such as Cas9. In some embodiments, the aqueous solution comprises an mRNA encoding a polypeptide.
[0399] In some embodiments, the lipid nucleic acid assembly formulations include an “amine lipid” (sometimes herein or elsewhere described as an “ionizable lipid” or a “biodegradable lipid”), together with an optional “helper lipid”, a “neutral lipid”, and a stealth lipid such as a PEG lipid. In some embodiments, the amine lipids or ionizable lipids are cationic depending on the pH.lonizable / Amine Lipids
[0400] In some embodiments, LNPs comprise an ionizable lipid such as those provided herein.
[0401] In some embodiments, the ionizable lipid is represented by Formula (I),whereinX1is O, NR1, or a direct bond,X2is C2-5 alkylene,X3is C(=O) or a direct bond,R1is H or Me,R3is C1-3 alkyl,R2is C1-3 alkyl, orR2taken together with the nitrogen atom to which it is attached and 1-3 carbon atoms of X2form a 4-, 5-, or 6-membered ring, orX1is NR1, R1and R2taken together with the nitrogen atoms to which they are attached form a 5- or 6-membered ring, orR2taken together with R3and the nitrogen atom to which they are attached form a 5-, 6-, or 7-membered ring,Y1is C2-12 alkylene,Y2is selected from(in either orientation),(in either orientation),(in either orientation),n is 0 to 3,R4is Ci-15 alkyl,Z1is C2-6 alkylene or a direct bond,(in either orientation) or absent, provided that if Z1is a direct bond, Z2is absent;R5is C5-9 alkyl or Ce-io alkoxy,R6is C5-9 alkyl or Ce-io alkoxy,W is methylene or a direct bond, andR7is H or Me,or a salt thereof,provided that if R3and R2are C2 alkyls, X1is O, X2is linear C3 alkylene, X3isC(=O), Y1is linear Ce alkylene, (Y2)n-R4is, R4is linear C5 alkyl, Z1is C2 alkylene, Z2is absent, W is methylene, and R7is H, then R5and R6are not Cs alkoxy,or a salt thereof.
[0402] In some embodiments, the ionizable lipid is Lipid A, which is (9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate, also called 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl (9Z, 12Z)-octadeca-9, 12-dienoate. Lipid A can be depicted as:salt thereof.
[0403] Lipid A may be synthesized according to W02015 / 095340 (e.g., pp. 84-86). In some embodiments, the ionizable lipid is Lipid A, or an ionizable lipid provided in WO2020 / 219876, which is hereby incorporated by reference. In some embodiments, the ionizable lipid is represented by structural Formula (II),or a salt thereof,wherein,A is O, NH, or a direct bond,X1is a Ci-5 alkylene,R1and R2are each independently a C1-3 alkyl, orR1taken together with the nitrogen atom to which it is attached and 1-3 carbon atoms of X1form a 4-, 5-, or 6-membered ring, orR1taken together with R2and the nitrogen atom to which they are attached form a 5-, 6-, or 7-membered ring, andR3is H or C1-3 alkyl,Z1and Z2are each independently a C1-5 alkylene,Z3and Z4are each independently a -C(=O)O- in either direction,Z5and Z6are each independently a direct bond or a C1-3 alkylene,Y1is selected from H, a C1-10 alkyl, C3-10 alkenyl, and C3-10 alkynyl,Y2, Y3, and Y4are each independently selected from a C3-10 alkyl, C3-10 alkenyl, or C3-10 alkynyl, andn is 0 or 1,or a salt thereof.
[0404] In some embodiments, the ionizable lipid is Lipid B, which is O, O'-(2-((((2-(diethylamino)ethyl)carbamoyl)oxy)methyl)propane- 1, 3 -diyl) di(heptadecan-9-yl) diglutarate. Lipid B can be depicted as:
[0405] Lipid B may be synthesized as follows: to a solution of di(heptadecan-9-yl) 0,0'-(2-((((4-nitrophenoxy)carbonyl)oxy)methyl)propane-l,3-diyl) diglutarate in MeCN (0.05 -0.25 M) was added N', N'-diethylethane-l,2-diamine (1.0 - 3.0 equiv.), pyridine (1.0 - 2.0 equiv.) and DMAP (0.1 - 1.0 equiv.). Then the mixture was stirred at 15-25 °C for at least 12 h under N2 atmosphere. The reaction mixture was concentrated under reduced pressure to remove solvent. The residue was diluted with EtOAc and washed 2-5x with IN NaHCOs and 3x with H2O. The organic layer was dried over Na2SO4, filtered and the filtrate was concentrated under reduced pressure to give a residue. The residue was purified by silica gel chromatography to afford product as a colorless oil.JH NMR (400 MHz, CDCh) 55.35 (s, 1H), 4.85 (p, J = 6.3 Hz, 2H), 4.12 (t, J = 6.0 Hz, 6H), 3.22 (q, J = 5.8 Hz, 2H), 2.53 (d, J = 6.8 Hz, 6H), 2.36 (dt, J = 14.7, 7.4 Hz, 9H), 1.93 (p, J = 7.5 Hz, 4H), 1.49 (q, J = 5.9 Hz, 8H), 1.24 (s, 48H), 1.00 (t, J = 7.1 Hz, 6H), 0.87 (t, J = 6.7 Hz, 12H). MS: 953.6 m / z [M+H],
[0406] In some embodiments, the ionizable lipid is Lipid C, which is 3-(((2-(azepan-l-yl)ethyl)carbamoyl)oxy)-2-((((9Z,12Z)-octadeca-9,12-dienoyl)oxy)methyl)propyl heptadecan-9-yl glutarate. Lipid C can be depicted as:
[0407] Lipid C may be synthesized from heptadecan-9 -yl (3-(((4-nitrophenoxy)carbonyl)oxy)-2-((((9Z,12Z)-octadeca-9,12-dienoyl)oxy)methyl)propyl) glutarate and 2-(azepan-l-yl)ethan-l -amine using the same method employed for Lipid B. ’ll NMR (400 MHz, CDCh) 55.44 - 5.26 (m, 5H), 4.86 (p, J = 6.3 Hz, 1H), 4.13 (t, J = 5.8 Hz, 6H), 3.21 (t, J = 6.0 Hz, 2H), 2.81 - 2.73 (m, 2H), 2.66 - 2.56 (m, 5H), 2.42 - 2.26 (m, 7H), 2.05 (q, J = 6.8 Hz, 4H), 1.94 (p, J = 7.5 Hz, 2H), 1.62 (d, J = 16.8 Hz, 10H), 1.50 (d, J = 6.2 Hz, 4H), 1.41 - 1.26 (m, 24H), 1.25 (s, 16H), 0.88 (td, J = 6.9, 5.0 Hz, 9H). MS: 890.5 m / z [M+H],
[0408] In some embodiments, the ionizable lipid is represented by structural Formula (III),wherein, independently for each occurrence,X1is Ci-3 alkyleneX2is selected from O, NH, NMe, and a bond, provided that when X2is O, R2taken together with the nitrogen atom and either R1or a carbon atom of X3form a 4-membered, 5 -membered, or 6-membered ring;X3is C2-4 alkylene,X4is Ci alkylene or a bond,X5is Ci alkylene or a bond,R1is C1-3 alkyl,R2is C1-3 alkyl, orR2taken together with the nitrogen atom and either R1or a carbon atom of X3form a 4-membered, 5-membered, or 6-membered ring,Y1is selected from a bond, -CH=CH-, -(C=O)O-, and -O(C=O)-,Y2is selected from -CH2-CH=CH- and C3-C4 alkylene,R3is selected from H, C5-7 cycloalkyl, Cs-Cio alkenyl, and C3-18 alkyl, andR4is C4-8 alkyl,or a salt thereof.
[0409] In some embodiments, the ionizable lipid is Lipid D, which is 2-((4-(((3- (pyrrolidin- 1 -yl)propoxy)carbonyl)oxy)decanoyl)oxy)propane- 1,3-diyl (9Z,9'Z, 12Z, 12'Z)-bis(octadeca-9,12-dienoate). Lipid D can be depicted as:or a salt thereof.
[0410] Lipid D may be synthesized according to WO 2020 / 118041, which is incorporated by reference in its entirety. In some embodiments, the ionizable lipid is Lipid D, or an ionizable lipid provided in WO 2020 / 118041, which is hereby incorporated by reference.
[0411] In some embodiments, the ionizable lipid is represented by structural Formula (IV),wherein, independently for each occurrence,X1is C5-11 alkylene,Y1is C3-11 alkylene,wherein ai is a bond to Y1, and a2 is a bond to R1,Z1is C2-4 alkylene,Z2is selected from -OH, -NH2, -OC(=O)R3, -OC(=O)NHR3, -NHC(=O)NHR3, and -NHS(=O)2R3,R1is C4-12 alkyl or C3-12 alkenyl,each R2is independently C4-12 alkyl, andR3is C1-3 alkyl,or a salt thereof.
[0412] In some embodiments, the ionizable lipid is Lipid E, which is 8-((8,8-bis(octyloxy)octyl)(2-hydroxyethyl)amino)octyl nonanoate. Lipid E can be depicted as:, or a salt thereof.
[0413] In some embodiments, the ionizable lipid is Lipid F, which is nonyl 8-((7,7-bis(octyloxy)heptyl)(2-hydroxyethyl)amino)octanoate. Lipid F can be depicted as:
[0414] Lipids E and F may be synthesized according to W02020072605 and Mol. Ther. 2018, 26(6), 1509-1519 (“Sahms”), which are incorporated by reference in their entireties. In some embodiments, the ionizable lipid Lipids E and F, or an ionizable lipid provided in W02020072605, which is hereby incorporated by reference.
[0415] In other embodiments, ionizable lipids can include, for example the ionizable lipids of WO 2020 / 219876 (e.g., at pp. 13-33, 66-87), WO 2020 / 118041, WO 2020 / 072605 (e.g., at pp. 5-12, 21-29, 61-68, WO 2019 / 067992, WO 2017 / 173054, WO 2015 / 095340, and WO 2014 / 136086, the entire contents of which, are hereby incorporated by reference.
[0416] Ionizable lipids and other “biodegradable lipids” suitable for use in the lipid nucleic acid assemblies described herein are biodegradable in vivo or ex vivo. The ionizable lipids have low toxicity (e.g., are tolerated in animal models without adverse effect in amounts of greater than or equal to 10 mg / kg). In some embodiments, lipid nucleic acid assemblies comprising an ionizable lipid include those where at least 75% of the ionizable lipid is cleared from the plasma or the engineered cell within 8, 10, 12, 24, or 48 hours, or 3, 4, 5, 6, 7, or 10 days. In some embodiments, lipid nucleic acid assemblies comprising an ionizable lipid include those where at least 50% of the nucleic acid, e.g., mRNA or guide RNA, is cleared from the plasma within 8, 10, 12, 24, or 48 hours, or 3, 4, 5, 6, 7, or 10 days. In some embodiments, lipid nucleic acid assemblies comprising an ionizable lipid include those where at least 50% of the lipid nucleic acid assembly is cleared from the plasma within 8, 10, 12, 24, or 48 hours, or 3, 4, 5, 6, 7, or 10 days, for example by measuring a lipid (e.g., an ionizable lipid), nucleic acid, e.g.,RNA / mRNA, or other component. In some embodiments, lipid-encapsulated versus free lipid, RNA, or nucleic acid component of the lipid nucleic acid assembly is measured.
[0417] Lipid clearance may be measured as described in literature. See Maier, M. A., et al. Biodegradable Lipids Enabling Rapidly Eliminated Lipid Nanoparticles for Systemic Delivery of RNAi Therapeutics. Mol. Ther. 2013, 21(8), 1570-78 “Maier"').
[0418] Ionizable and bioavailable lipids for LNP delivery of nucleic acids known in the art are suitable. The ionizable lipids of the present disclosure may form salts depending upon the pH of the medium they are in. For example, in a slightly acidic medium, the ionizable lipids may be protonated and thus bear a positive charge. Conversely, in a slightly basic medium, such as, for example, blood where pH is approximately 7.35, the ionizable lipids may not be protonated and thus bear no charge. In some embodiments, the ionizable lipids of the present disclosure may be predominantly protonated at a pH of at least about 9. In some embodiments, the ionizable lipids of the present disclosure may be predominantly protonated at a pH of at least about 10.
[0419] The pH at which the ionizable lipids is predominantly protonated is related to its intrinsic pKa. In some embodiments, an ionizable lipid of the present disclosure has a pKa in the range of from about 5.1 to about 8.0, even more preferably from about 5.5 to about 7.6. In some embodiments, an ionizable lipid of the present disclosure has a pKa in the range of from about 5.7 to about 8, from about 5.7 to about 7.6, from about 6 to about 8, from about 6 to about 7.5, from about 6 to about 7, from about 6 to about 6.9, from about 6 to about 6.5, from about 6.1 to about 6.9, or from about 6 to about 6.85. In some embodiments, an ionizable lipid of the present disclosure has a pKa of about 6.0, about 6.1, about 6.1, about 6.2, about 6.3, about 6.4, about 6.6, about 6.7, about 6.8, or about 6.9. Alternatively, an ionizable lipid of the present disclosure has a pKa in the range of from about 6 to about 8. The pKa can be an important consideration in formulating LNPs, as it has been found that LNPs formulated with certain lipids having a pKa ranging from about 5.5 to about 7.0 are effective for delivery of cargo in vivo. Further, it has been found that LNPs formulated with certain lipids having a pKa ranging from about 5.3 to about 6.4 are effective for delivery in vivo, e.g. to tumors. See, e.g., WO 2014 / 136086. In some embodiments, the ionizable lipids are positively charged at an acidic pH but neutral in the blood.
[0420] Ionizable lipids and other “biodegradable lipids” suitable for use in the lipid nucleic acid assemblies described herein can be biodegradable in vivo or ex vivo. The amine lipids have low toxicity (e.g., are tolerated in animal models without adverse effect in amounts of greater than or equal to 10 mg / kg). In some embodiments, lipid nucleic acid assembliescomprising an amine lipid include those where at least 75% of the amine lipid is cleared from the plasma or the engineered cell within 8, 10, 12, 24, or 48 hours, or 3, 4, 5, 6, 7, or 10 days. In some embodiments, lipid nucleic acid assemblies comprising an amine lipid include those where at least 50% of the nucleic acid, e.g., mRNA or guide RNA, is cleared from the plasma within 8, 10, 12, 24, or 48 hours, or 3, 4, 5, 6, 7, or 10 days. In some embodiments, lipid nucleic acid assemblies comprising an amine lipid include those where at least 50% of the lipid nucleic acid assembly is cleared from the plasma within 8, 10, 12, 24, or 48 hours, or 3, 4, 5, 6, 7, or 10 days, for example by measuring a lipid (e.g., an amine lipid), nucleic acid, e.g., RNA / mRNA, or other component. In some embodiments, lipid-encapsulated versus free lipid, RNA, or nucleic acid component of the lipid nucleic acid assembly is measured.
[0421] Biodegradable lipids include, for example the biodegradable lipids of WO 2020 / 219876 (e.g., atpp. 13-33, 66-87), WO 2020 / 118041, WO 2020 / 072605 (e.g., atpp.5-12, 21-29, 61-68, WO 2019 / 067992, WO 2017 / 173054, WO 2015 / 095340, andWO 2014 / 136086, and LNPs include LNP compositions described therein, the lipids and compositions of which are hereby incorporated by reference.
[0422] Lipid clearance may be measured as described in literature. See Maier, M. A., et al. Biodegradable Lipids Enabling Rapidly Eliminated Lipid Nanoparticles for Systemic Delivery of RNAi Therapeutics. Mol. Ther. 2013, 21(8), 1570-78 (“Maier"').
[0423] Ionizable and bioavailable lipids for LNP delivery of nucleic acids known in the art are suitable. Lipids may be ionizable depending upon the pH of the medium they are in. For example, in a slightly acidic medium, the lipid, such as an amine lipid, may be protonated and thus bear a positive charge. Conversely, in a slightly basic medium, such as, for example, blood where pH is approximately 7.35, the lipid, such as an amine lipid, may not be protonated and thus bear no charge.
[0424] The ability of a lipid to bear a charge is related to its intrinsic pKa. In some embodiments, the amine lipids of the present disclosure may each, independently, have a pKa in the range of about 5.1-7.4.Additional Lipids
[0425] “Neutral lipids” suitable for use in a lipid composition of the disclosure include, for example, a variety of neutral, uncharged or zwitterionic lipids. Examples of neutral phospholipids suitable for use in the present disclosure include, but are not limited to, 5-heptadecylbenzene-l,3-diol (resorcinol), dipalmitoylphosphatidylcholine (DPPC), distearoylphosphatidylcholine (DSPC), l,2-dioleoyl-sn-glycero-3-phosphocoline (DOPC),dimyristoylphosphatidylcholine (DMPC), phosphatidylcholine (PLPC), 1,2-distearoyl-sn-glycero-3-phosphocholine (DAPC), phosphatidylethanolamine (PE), egg phosphatidylcholine (EPC), dilauryl oylphosphatidylcholine (DLPC), dimyristoylphosphatidylcholine (DMPC), 1-myristoyl-2-palmitoyl phosphatidylcholine (MPPC), l-palmitoyl-2-myristoyl phosphatidylcholine (PMPC), l-palmitoyl-2-stearoyl phosphatidylcholine (PSPC), 1,2-diarachidoyl-sn-glycero-3-phosphocholine (DBPC), l-stearoyl-2-palmitoyl phosphatidylcholine (SPPC), l,2-dieicosenoyl-sn-glycero-3-phosphocholine (DEPC), palmitoyloleoyl phosphatidylcholine (POPC), lysophosphatidyl choline, dioleoyl phosphatidylethanolamine (DOPE), dilinoleoylphosphatidylcholine distearoylphosphatidylethanolamine (DSPE), dimyristoyl phosphatidylethanolamine (DMPE), dipalmitoyl phosphatidylethanolamine (DPPE), palmitoyloleoyl phosphatidylethanolamine (POPE), lysophosphatidylethanolamine and combinations thereof. In certain embodiments, the neutral phospholipid may be selected from distearoylphosphatidylcholine (DSPC) and dimyristoyl phosphatidyl ethanolamine (DMPE). In other embodiments, the neutral phospholipid may be distearoylphosphatidylcholine (DSPC).
[0426] “Helper lipids” include steroids, sterols, and alkyl resorcinols. Helper lipids suitable for use in the present disclosure include, but are not limited to, cholesterol, 5-heptadecylresorcinol, and cholesterol hemisuccinate. In certain embodiments, the helper lipid may be cholesterol or a deri vali ve thereof. In one embodiment, the helper lipid may be cholesterol hemisuccinate.
[0427] “Stealth lipids” are lipids that alter the length of time the nanoparticles can exist in vivo (e.g., in the blood). Stealth lipids may assist in the formulation process by, for example, reducing particle aggregation and controlling particle size. Stealth lipids used herein may modulate pharmacokinetic properties of the lipid nucleic acid assembly or aid in stability of the nanoparticle ex vivo. Stealth lipids suitable for use in a lipid composition of the disclosure include, but are not limited to, stealth lipids having a hydrophilic head group linked to a lipid moiety. Stealth lipids suitable for use in a lipid composition of the present disclosure and information about the biochemistry of such lipids can be found in Romberg et al., Pharmaceutical Research, Vol. 25, No. 1, 2008, pg. 55-71 and Hoekstra et al., Biochimica et Biophysica Acta 1660 (2004) 41-52. Additional suitable PEG lipids are disclosed, e.g., in WO 2006 / 007712.
[0428] In one embodiment, the hydrophilic head group of stealth lipid comprises a polymer moiety selected from polymers based on PEG. Stealth lipids may comprise a lipid moiety. In some embodiments, the stealth lipid is a PEG lipid.
[0429] In one embodiment, a stealth lipid comprises a polymer moiety selected from polymers based on PEG (sometimes referred to as polyethylene oxide)), poly(oxazoline), poly(vinyl alcohol), poly(glycerol), poly(N- vinylpyrrolidone), polyaminoacids and poly[N-(2-hydroxypropyl)methacrylamide].
[0430] In one embodiment, the PEG lipid comprises a polymer moiety based on PEG (sometimes referred to as poly (ethylene oxide)).
[0431] The PEG lipid further comprises a lipid moiety. In some embodiments, the lipid moiety may be derived from diacylglycerol or diacylglycamide, including those comprising a dialkylglycerol or dialkylglycamide group having alkyl chain length independently comprising from about C4 to about C40 saturated or unsaturated carbon atoms, wherein the chain may comprise one or more functional groups such as, for example, an amide or ester. In some embodiments, the alkyl chain length comprises about CIO to C20. The dialkylglycerol or dialkylglycamide group can further comprise one or more substituted alkyl groups. The chain lengths may be symmetrical or asymmetrical.
[0432] Unless otherwise indicated, the term “PEG” as used herein means any polyethylene glycol or other polyalkylene ether polymer. In one embodiment, PEG is an optionally substituted linear or branched polymer of ethylene glycol or ethylene oxide. In one embodiment, PEG is unsubstituted. In one embodiment, the PEG is substituted, e.g., by one or more alkyl, alkoxy, acyl, hydroxy, or aryl groups. In one embodiment, the term includes PEG copolymers such as PEG-polyurethane or PEG-polypropylene (see, e.g., J. Milton Harris, Polyethylene glycol) chemistry: biotechnical and biomedical applications (1992)); in another embodiment, the term does not include PEG copolymers.
[0433] In some embodiments, the PEG (e.g., conjugated to a lipid moiety or lipid, such as a stealth lipid), is a “PEG-2K,” also termed “PEG 2000,” which has an average molecular weight of about 2,000 Daltons. PEG-2K is represented herein by the following formula (IV), wherein n is 45, meaning that the number averaged degree of polymerization comprises about45 subunits
[0434] However, other PEG embodiments known in the art may be used, including, e.g., those where the number-averaged degree of polymerization comprises about 23 subunits (n=23), or 68 subunits (n=68). In some embodiments, R may be selected from H, substituted alkyl, and unsubstituted alkyl. In some embodiments, R may be unsubstituted alkyl. In some embodiments, R may be methyl.
[0435] In any of the embodiments described herein, the PEG lipid may be selected from PEG-dilauroylglycerol, PEG-dimyristoylglycerol (PEG-DMG catalog # GM-020 from NOF, Tokyo, Japan), such as e.g., l,2-dimyristoyl-rac-glycero-3-methylpolyoxyethylene glycol 2000 (PEG2k-DMG), PEG-dipalmitoylglycerol, PEG-distearoylglycerol (PEG-DSPE) (catalog # DSPE-020CN, NOF, Tokyo, Japan), PEG-dilaurylglycamide, PEG-dimyristylglycamide, PEG-dipalmitoylglycamide, and PEG-distearoylglycamide, PEG-cholesterol ( 1 - [8 ’ -(Cholest-5 -en-3 [beta] -oxy)carboxamido-3 ’,6 ’ -dioxaoctanyl]carbamoyl-[omega] -methyl -poly(ethylene glycol), PEG-DMB (3,4-ditetradecoxylbenzyl-[omega]-methyl-poly(ethylene glycol)ether), l,2-dimyristoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DMPE) (cat. #88O15OP from Avanti Polar Lipids, Alabaster, Alabama, USA), l,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DSPE) (cat. #880120C from Avanti Polar Lipids, Alabaster, Alabama, USA), 1,2-distearoyl-sn-glycerol, methoxypolyethylene glycol (PEG2k-DSG; GS-020, NOF Tokyo, Japan), polyethylene glycol)-2000-dimethacrylate (PEG2k-DMA), and l,2-distearyloxypropyl-3-amine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DSA), and methoxy-PEG2000-carbamoyl-l,2-tridecyoxypropylamine) (Cl 3 Ether), and methoxy-PEG2000-carbamoyl-l,2-tetradecyoxypropylamine (C14 Ether). In one embodiment, the PEG lipid may be l,2-dimyristoyl-rac-glycero-3-methylpolyoxyethylene glycol 2000. In one embodiment, the PEG lipid may be PEG2k-DMG. In some embodiments, the PEG lipid may be PEG2k-DSG. In one embodiment, the PEG lipid may be PEG2k-DSPE. In one embodiment, the PEG lipid may be PEG2k-DMA. In one embodiment, the PEG lipid may be PEG2k-C-DMA. In one embodiment, the PEG lipid may be compound S027, disclosed in W02016 / 010840 (paragraphs
[0240] to
[0244] ). In one embodiment, the PEG lipid may be PEG2k-DSA. In one embodiment, the PEG lipid may be PEG2k-Cl 1. In some embodiments, the PEG lipid may be PEG2k-C14. In some embodiments, the PEG lipid may be PEG2k-C16. In some embodiments, the PEG lipid may be PEG2k-C18.
[0436] In some embodiments, the PEG lipid includes a glycerol group. In some embodiments, the PEG lipid includes a dimyristoylglycerol (DMG) group. In some embodiments, the PEG lipid comprises PEG-2k. In some embodiments, the PEG lipid is a PEG-DMG. In some embodiments, the PEG lipid is a PEG-2k-DMG. In some embodiments, the PEG lipid is l,2-dimyristoyl-rac-glycero-3-methoxypolyethylene glycol2000. In some embodiments, the PEG-2k-DMG is l,2-dimyristoyl-rac-glycero-3-methoxypolyethylene glycol-2000.Lipid Nanoparticles (LNPs)
[0437] A LNP may contain (i) an ionizable (biodegradable) lipid, (ii) an optional neutral lipid, (iii) a helper lipid, and (iv) a stealth lipid, such as a PEG lipid. The lipid nucleic acid assembly may contain a biodegradable lipid and one or more of a neutral lipid, a helper lipid, and a stealth lipid, such as a PEG lipid.
[0438] The lipid nucleic acid assembly may contain (i) an ionizable lipid for encapsulation and for endosomal escape, (ii) a neutral lipid for stabilization, (iii) a helper lipid, also for stabilization, and (iv) a stealth lipid, such as a PEG lipid. The lipid nucleic acid assembly may contain an ionizable lipid and one or more of a neutral lipid, a helper lipid, also for stabilization, and a stealth lipid, such as a PEG lipid.
[0439] An LNP may comprise a nucleic acid, e.g., an RNA, component that includes one or more of an RNA-guided DNA-binding agent, a Gas nuclease mRNA, a Class II Cas nuclease mRNA, a Cas9 mRNA, and a guide RNA. In some embodiments, a LNP may include a Class II Cas nuclease and a guide RNA as the RNA component. In some embodiments, an LNP may comprise the RNA component, an ionizable lipid, a helper lipid, a neutral lipid, and a stealth lipid. In certain LNPs, the helper lipid is cholesterol. In other compositions, the neutral lipid is DSPC. In additional embodiments, the stealth lipid is PEG2k-DMG or PEG2k-C11. In some embodiments, the LNP comprises Lipid A-E or an equivalent of Lipid A-E; a helper lipid; a neutral lipid; a stealth lipid; and an RNA such as a guide RNA. In some embodiments, the LNP comprises Lipid A-E or an equivalent of Lipid A-E; a helper lipid; a stealth lipid; and an RNA such as a guide RNA. In some compositions, the ionizable lipid is Lipid A-E. In some compositions, the ionizable lipid is Lipid A-E or an acetal analog thereof; the helper lipid is cholesterol; the neutral lipid is DSPC; and the stealth lipid is PEG2k-DMG. In some compositions, the amine lipid is Lipid A-F or an acetal analog thereof; the helper lipid is cholesterol; the neutral lipid is DSPC; and the stealth lipid is C13 Ether.
[0440] In some embodiments, LNP compositions are described according to the respective molar ratios of the component lipids in the formulation. Embodiments of the present disclosure provide LNP compositions described according to the respective molar ratios of the component lipids in the formulation. In one embodiment, the mol % of the amine lipid may be from about 30 mol % to about 60 mol %. All mol % numbers are given as a fraction of the lipid component of the LNPs.
[0441] Embodiments of the present disclosure provide LNPs described according to the respective molar ratios of the component lipids in the composition. All mol % numbers aregiven as a fraction of the lipid component of the lipid composition or, more specifically, the LNP compositions. In some embodiments, the lipid mol % of a lipid relative to the lipid component will be ±30%, ±25%, ±20%, ±15%, ±10%, ±5%, or ±2.5% of the specified, nominal, or actual mol % of the lipid. In some embodiments, the lipid mol % of a lipid relative to the lipid component will be ±4 mol %, ±3 mol %, ±2 mol %, ±1.5 mol %, ±1 mol %, ±0.5 mol %, ±0.25 mol %, or ±0.05 mol % of the specified, nominal, or actual mol % of the lipid component. In certain embodiments, the lipid mol % will vary by less than 15%, 10%, 5%, 1%, or 0.5% from the specified, nominal, or actual mol % of the lipid. In some embodiments, the mol % numbers are based on nominal concentration. As used herein, “nominal concentration” refers to concentration based on the input amounts of substances combined to form a resulting composition. For example, if 100 mg of solute is added to 1 L water, the nominal concentration is 100 mg / L. In some embodiments, the mol % numbers are based on actual concentration, e.g., concentration determined by an analytic method. In some embodiments, actual concentration of the lipids of the lipid component may be determined, for example, from chromatography, such as liquid chromatography, followed by a detection method, such as charged aerosol detection. In some embodiments, actual concentration of the lipids of the lipid component may be characterized by lipid analysis, AF4-MALS, NTA, or cryo-EM. All mol % numbers are given as a percentage of the lipids of the lipid component.
[0442] In one embodiment, the mol % of the ionizable lipid may be from about 30 mol % to about 60 mol %. In one embodiment, the mol % of the neutral lipid may be from about 5 mol % to about 15 mol %.
[0443] In one embodiment, the mol % of the helper lipid may be from about 20 mol % to about 60 mol %. In one embodiment, the mol % of the helper lipid is adjusted based on the ionizable lipid, neutral lipid, and PEG lipid concentrations to bring the lipid component to 100 mol %.
[0444] In one embodiment, the mol % of the PEG lipid may be from about 1 mol % to about 10 mol %.
[0445] Embodiments of the present disclosure provide LNP compositions, for example, LNP compositions comprising an ionizable lipid, a helper lipid, a helper lipid, and a PEG lipid, described according to the respective molar ratios of the component lipids in the formulation. In certain embodiments, the amount of the ionizable lipid is from about 25 mol % to about 45 mol %; the amount of the neutral lipid is from about 10 mol % to about 30 mol %; the amount of the helper lipid is from about 25 mol % to about 65 mol %; and the amount of the PEG lipid is from about 1.5 mol % to about 3.5 mol %.
[0446] Other embodiments of the present disclosure provide LNP compositions, for example, LNP compositions comprising an ionizable lipid, a helper lipid, a helper lipid, and a PEG lipid, described according to the respective molar ratios of the component lipids in the formulation. In certain embodiments, the amount of the ionizable lipid is from about 25 mol % to about 50 mol %; the amount of the neutral lipid is from about 7 mol % to about 25 mol %; the amount of the helper lipid is from about 39 mol % to about 65 mol %; and the amount of the PEG lipid is from about 0.5 mol % to about 1.8 mol %.
[0447] In certain embodiments, the amount of the ionizable lipid is about 20-55 mol %, about 20-45 mol %. In certain embodiments, the amount of the neutral lipid is about 7-25 mol %. In certain embodiments, the amount of the helper lipid is about 39-65 mol %, about 39-59 mol %,. In certain embodiments, the amount of the PEG lipid is about 0.5-1.8 mol %.
[0448] In some embodiments, the cargo includes an mRNA encoding an RNA-guided DNA-binding agent (e.g., a Gas nuclease, a Class II Cas nuclease, or Cas9), or a guide RNA or a nucleic acid encoding a guide RNA, or a combination of mRNA and guide RNA. In one embodiment, a LNP comprising this cargo may contain an ionizable lipid, e.g., an ionizable lipid provided herein. In some embodiments, a LNP comprising this cargo may comprise Lipid A. In some embodiments, a LNP comprising this cargo may comprise Lipid B. In some embodiments, a LNP comprising this cargo may comprise Lipid C. In some embodiments, a LNP comprising this cargo may comprise Lipid D. In some embodiments, a LNP comprising this cargo may comprise Lipid E. In some embodiments, a LNP comprising this cargo may comprise Lipid E In various embodiments, a LNP comprises an amine lipid, a neutral lipid, a helper lipid, and a PEG lipid. In some embodiments, the helper lipid is cholesterol. In some embodiments, the neutral lipid is DSPC. In specific embodiments, PEG lipid is PEG2k-DMG. In some embodiments, a LNP may comprise an amine lipid (e.g., Lipid A, Lipid B, Lipid C, Lipid D, Lipid E, or Lipid E), a helper lipid, a neutral lipid, and a PEG lipid. In some embodiments, a LNP comprises an amine lipid (e.g., Lipid A, Lipid B, Lipid C, Lipid D, Lipid E, or Lipid E), DSPC, cholesterol, and a PEG lipid. In some embodiments, the LNP comprises a PEG lipid comprising DMG. In some embodiments, the amine lipid is selected from Lipid A, and an equivalent of Lipid A, including an acetal analog of Lipid A, or an amine lipid provided in WO2020 / 219876; or Lipid D or an amine lipid provided in W02020 / 072605. In additional embodiments, a LNP comprises Lipid A, cholesterol, DSPC, and PEG2k-DMG. In additional embodiments, a LNP comprises Lipid B, cholesterol, DSPC, and PEG2k-DMG. In ...
Claims
1. We claim:
1. An HMGB1 polypeptide comprising an HMGB1 Box B domain and at least 7 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain.
2. The HMGB 1 polypeptide of claim 1, wherein the HMGB 1 polypeptide comprises a cryptic nuclear localization signal (NLS).
3. The HMGB 1 polypeptide of claim 1 or 2, wherein the cryptic NLS comprises the sequence of EKSKKKK (SEQ ID NO: 5).
4. The HMGB1 polypeptide of any one of claims 1-3, wherein the HMGB1 polypeptide comprises a receptor for advanced glycation end-products (RAGE) binding domain.
5. The HMGB1 polypeptide of claim 4, wherein the RAGE binding domain comprises the sequence of SEQ ID NO: 21.
6. The HMGB1 polypeptide of any one of claims 1-5, wherein the HMGB1 polypeptide comprises 7-20 amino acid residues C-terminal to the HMGB1 Box B domain.
7. The HMGB1 polypeptide of any one of claims 1-6, wherein the HMGB1 polypeptide comprises the sequence of SEQ ID NO: 22.
8. The HMGB 1 polypeptide of any one of claims 1-7, wherein the HMGB 1 polypeptide lacks a sequence at least 50%, 60%, 70%, 80%, 90%, 94%, or 97% identical to SEQ ID NO: 6 or lacks the sequence of SEQ ID NO: 6.
9. The HMGB 1 polypeptide of any one of claims 1-8, wherein the HMGB 1 polypeptide lacks transcriptional stimulatory function.
10. The HMGB1 polypeptide of any one of claims 1-9, wherein the HMGB1 polypeptide comprises a cryptic NLS that is located C-terminal to the HMGB 1 Box B domain.
11. The HMGB 1 polypeptide of any one of claims 6-10, wherein the at least 7-20 contiguous amino acid residues C-terminal end of the HMGB 1 Box B domain comprise the cryptic NLS.
12. The HMGB1 polypeptide of any one of claims 1-9, wherein the HMGB1 polypeptide comprises a cryptic NLS that is located N-terminal to the HMGB1 Box B domain.
13. The HMGB1 polypeptide of claim 12, wherein the cryptic NLS immediately precedes the N-terminal end of the HMGB 1 Box B domain.
14. The HMGB1 polypeptide of any one of claims 1-13, wherein the HMGB1 polypeptide lacks an HMGB1 Box A domain.
15. The HMGB1 polypeptide of any one of claims 1-14, wherein the HMGB1 polypeptide comprises two HMGB1 Box B domains.
16. The HMGB1 polypeptide of claim 15, wherein the cryptic NLS is located C-terminal to the C-terminal end of the two HMGB1 Box B domains.
17. The HMGB1 polypeptide of any one of claims 1-13 or 15-16, wherein the HMGB1 polypeptide comprises an HMGB1 Box A domain.
18. The HMGB1 polypeptide of claim 17, wherein the HMGB1 Box B domain is located C-terminal to the HMGB 1 Box A domain.
19. The HMGB1 polypeptide of any one of claims 1-18, wherein the HMGB1 polypeptide comprises, from N-terminus to C-terminus:19.1 ) HMGB 1 B ox A domain-HMGB 1 B ox B domain-cryptic NLS;20.2) HMGB1 Box B domain-HMGB 1 Box B domain-cryptic NLS; or21.3) HMGB 1 Box B domain-cryptic NLS.
20. The HMGB1 polypeptide of any one of claims 1-19, wherein the HMGB1 polypeptide comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, and 21-22, or that is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 12-20.
21. An HMGB1 polypeptide comprising a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to SEQ ID NO: 7 or that is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to SEQ ID NO: 17.
22. A system comprising the HMGB1 polypeptide of any one of claims 1-21 and a programmable DNA-binding protein.
23. The system of claim 22, wherein the HMGB1 polypeptide is operably linked to a programmable DNA-binding protein.
24. A fusion protein comprising the HMGB1 polypeptide of any one of claims 1-21, further comprising a programmable DNA-binding protein.
25. The system or fusion protein of any one of claims 22-24, wherein the programmable DNA-binding protein is selected from a clustered regularly interspaced short palindromic repeats (CRISPR) nuclease, a zinc finger nuclease (ZFN), a transcription activator-like effector (TALE), a transcription activator-like effector nuclease (TALEN), a meganuclease, and a chimeric protein comprising a programmable DNA-binding domain linked to a nuclease domain.
26. The system or fusion protein of claim 25, wherein the programmable DNA-binding protein is a Class II Cas nuclease.
27. The system or fusion protein of claim 26, wherein the Class II Cas nuclease is selected from Cas9, Cpfl, C2cl, C2c2, C2c3, HF Cas9, HypaCas9, eSPCas9(1.0), eSPCas9(l.l).
28. The system or fusion protein of claim 27, wherein the Class II Cas nuclease is a Cas9 nuclease.
29. The system or fusion protein of claim 28, wherein the Cas9 nuclease is a cleavase or a nickase.
30. The system or fusion protein of claim 28, wherein the Cas9 nuclease is a dCas9.
31. The system or fusion protein of any one of claims 28-30, wherein the Cas9 nuclease comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719-721, optionally wherein the Cas9 nuclease includes or lacks an N-terminal methionine; or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 699-701, 703, 706, 711, 714, 717-718, and 733-736, optionally wherein the nucleic acid includes or lacks a start or stop codon.
32. The system or fusion protein of claim 25, wherein the programmable DNA-binding protein is a Class V Cas nuclease.
33. The system or fusion protein of claim 32, wherein the Class V Cas nuclease is a Casl2 nuclease, optionally a Casl2a (Cpfl) or a Casl2e (CasX), optionally wherein the Casl2 is Lachnospiraceae bacterium Casl2a (LbCasl2a) or Acidaminococcus sp. Casl2a (AsCasl2a).
34. The system or fusion protein of any one of claims 22 and 24-33, further comprising a deaminase.
35. The system or fusion protein of claim 34, wherein the deaminase is an adenosine deaminase or a cytidine deaminase.
36. The system or fusion protein of claim 35, wherein the deaminase is an APOBEC3A deaminase (A3A).
37. The system or fusion protein of any one of claims 35 or 36, wherein the deaminase comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to SEQ ID NO: 728.
38. The system or fusion protein of any one of claims 22 and 24-37, wherein the programmable DNA-binding protein is located N-terminal to the HMGB 1 polypeptide.
39. The system or fusion protein of any one of claims 22 and 24-38, wherein the programmable DNA-binding protein is located C-terminal to the HMGB 1 polypeptide.
40. The system or fusion protein of any one of claims 22 and 24-39, further comprising a heterologous nuclear localization signal (NLS).
41. The system or fusion protein of claim 40, wherein the heterologous NLS comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 366-369 and 370-384.
42. The system or fusion protein of claim 41, wherein the heterologous NLS comprises an SV40 NLS or a nucleoplasmin NLS.
43. The system or fusion protein of claim 42, wherein the heterologous NLS comprises a sequence having at least 80%, 85%, 90%, 95%, or 100% identity to any one of SEQ ID NOs: 366-367, 371, 383-384 or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, or 100% identity to the sequence of any one of SEQ ID NOs: 370, 385, and 397.
44. The system or fusion protein of any one of claims 22 and 24-43, comprising two heterologous NLSs.
45. The system or fusion protein of any one of claims 22 and 24-44, comprising an SV40 NLS and a nucleoplasmin NLS, optionally wherein the SV40 NLS is located N-terminal to the nucleoplasmin NLS.
46. The system or fusion protein of any one of claims 22 and 24-45, wherein at least one heterologous NLS is present at the N-terminus of the HMGB1 polypeptide or the fusion protein.
47. The fusion protein of any one of claims 24-46, further comprising an intervening peptide sequence, wherein the programmable DNA-binding protein is linked to the HMGB 1 polypeptide by the intervening peptide sequence.
48. The fusion protein of claim 47, wherein the intervening peptide sequence is at least 10 amino acid residues in length, optionally at least 12 amino acid residues in length.
49. The fusion protein of claim 47 or 48, wherein the intervening peptide sequence is 10-50 amino acid residues in length, optionally 12-50 amino acid residues in length, further optionally 12, 15, 20, 33, 41, or 42 amino acid residues in length.
50. The fusion protein of any one of claims 47-49, wherein the intervening peptide sequence comprises a linker or a heterologous NLS.
51. The fusion protein of claim 50, wherein the intervening peptide sequence comprises, from N-terminus to C-terminus:50.i.linker;51.ii.linker -heterologous NLS;52.iii.heterologous NLS -linker;53.iv.first linker -heterologous NLS -second linker; or54.v.first linker -first heterologous NLS - second heterologous NLS- second linker.
52. The fusion protein of any one of claims 24-51, wherein the fusion protein comprises, from N-terminus to C-terminus:55.1) DNA-binding domain-heterologous NLS-linker-HMGB 1 polypeptide; 2) DNA-binding domain-linker-HMGB 1 polypeptide;56.3) DNA-binding domain-first linker-heterologous NLS-second linker- HMGB1 polypeptide;57.4) heterologous NLS-DNA-binding domain-first linker- heterologous NLS- second linker- HMGB 1 polypeptide;58.5) heterologous NLS-DNA-binding domain-linker-HMGB 1 polypeptide; 6) heterologous NLS-heterologous NLS-DNA-binding domain-linker- HMGB 1 polypeptide;59.7) HMGB1 polypeptide -first linker-first heterologous NLS-second heterologous NLS-second linker-DNA-binding domain;60.8) First heterologous NLS-second heterologous NLS-deaminase-first linker- DNA-binding domain-second linker-HMGBl polypeptide; or61.9) Deaminase-first linker-DNA-binding domain-heterologous NLS-second linker-HMGB 1 polypeptide.
53. The fusion protein of any one of claims 50-52, wherein the linker comprises the amino acid sequence of any one of SEQ ID NOs: 301-365 and 425-435; or wherein the first linker and the second linker are independently selected from any one of SEQ ID NOs: 301-365 and 425-435.
54. The fusion protein of any one of claims 50-53, wherein the heterologous NLS comprises the amino acid sequence of any one of SEQ ID NOs: 366-369 and 370-384; or wherein the firstheterologous NLS and the second heterologous NLS are independently selected from SEQ ID NOs: 366-369 and 370-384.
55. The fusion protein of any one of claims 24-54, wherein the fusion protein comprises a sequence having at least 80%, 90%, 95%, 98%, 99%, or 100% identity to any one of SEQ ID NOs: 515, 567, 591, 597, 603, 624, 627, 630, 633, 636, 639, 642, 649, 652, 655, 664, 667, or 673, or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98%, 99%, or 100% identity to the sequence of any one of SEQ ID NOs: 513, 514, 565, 566, 589, 590, 595, 596, 601, 602, 622, 623, 625, 626, 628, 629, 631, 632, 634, 635, 637, 638, 640, 641, 647, 648, 650, 651, 653, 654, 662, 663, 665, 666, 671, or 672.
56. A system comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB 1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB 1 Box B domain, wherein the HMGB 1 polypeptide lacks all or part of an acidic tail domain; and (2) a Class II Cas nuclease; wherein the HMGB 1 polypeptide is operably linked to the Class II Cas nuclease.
57. The system of claim 56, wherein the HMGB1 polypeptide and Class II Cas nuclease are covalently linked in a fusion protein.
58. A fusion protein comprising (1) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a Class II Cas nuclease.
59. The system or fusion protein of any one of claims 56-58, wherein the HMGB 1 polypeptide comprises a cryptic nuclear localization signal (NLS).
60. The system or fusion protein of claim 59, wherein the cryptic NLS comprises the sequence of EKSKKKK (SEQ ID NO: 5).
61. The system or fusion protein of any one of claims 56-60, wherein the HMGB 1 polypeptide comprises a receptor for advanced glycation end-products (RAGE) binding domain.
62. The system or fusion protein of claim 61, wherein the RAGE binding domain comprises the sequence of SEQ ID NO: 21.
63. The system or fusion protein of any one of claims 56-62, wherein the HMGB 1 polypeptide comprises 7-20 amino acid residues C-terminal to the HMGB1 Box B domain.
64. The system or fusion protein of any one of claims 56-63, wherein the HMGB 1 polypeptide comprises the sequence of SEQ ID NO: 22.
65. The system or fusion protein of any one of claims 56-64, wherein the HMGB 1 polypeptide lacks a sequence at least 50%, 60%, 70%, 80%, 90%, 94%, or 97% identical to SEQ ID NO: 6 or lacks the sequence of SEQ ID NO: 6.
66. The system or fusion protein of any one of claims 56-65, wherein the HMGB 1 polypeptide lacks transcriptional stimulatory function.
67. The system or fusion protein of any one of claims 56-66, wherein the HMGB 1 polypeptide comprises a cryptic NLS that is located C-terminal to the HMGB 1 Box B domain.
68. The system or fusion protein of claim 67, wherein the at least 7-20 contiguous amino acid residues at the C-terminal end of the HMGB 1 Box B domain comprise the cryptic NLS.
69. The system or fusion protein of any one of claims 56-66, wherein the HMGB 1 polypeptide comprises a cryptic NLS that is located N-terminal to the HMGB1 Box B domain.
70. The system or fusion protein of claim 69, wherein the cryptic NLS immediately precedes the N-terminal end of the HMGB 1 Box B domain.
71. The system or fusion protein of any one of claims 56-70, wherein the HMGB 1 polypeptide lacks an HMGB1 Box A domain.
72. The system or fusion protein of any one of claims 56-71, wherein the HMGB 1 polypeptide comprises two HMGB1 Box B domains.
73. The system or fusion protein of claim 72, wherein the cryptic NLS is located C-terminal to the C-terminal end of the two HMGB1 Box B domain.
74. The system or fusion protein of any one of claims 56-70 or 72-73, wherein the HMGB1 polypeptide comprises an HMGB1 Box A domain.
75. The system or fusion protein of claim 74, wherein the HMGB1 Box B domain is located C-terminal to the HMGB 1 Box A domain.
76. The system or fusion protein of any one of claims 56-75, wherein the HMGB 1 polypeptide comprises, from N-terminus to C-terminus:85.1 ) HMGB 1 B ox A domain-HMGB 1 B ox B domain-cryptic NLS;86.2) HMGB1 Box B domain-HMGB 1 Box B domain-cryptic NLS; or87.3) HMGB 1 Box B domain-cryptic NLS.
77. The system or fusion protein of any one of claims 56-76, wherein the HMGB 1 polypeptide comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, and 21-22, or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 12-20.
78. The system or fusion protein of any one of claims 56-77, wherein the Class II Cas nuclease is selected from Cas9, Cpfl, C2cl, C2c2, C2c3, HF Cas9, HypaCas9, eSPCas9(1.0), eSPCas9(l.l).
79. The system or fusion protein of any one of claims 56-78, wherein the Class II Cas nuclease is a Cas9 nuclease.
80. The system or fusion protein of claim 79, wherein the Cas9 is a cleavase or a nickase.
81. The system or fusion protein of claim 79, wherein the Cas9 is a dCas9.
82. The system or fusion protein of any one of claims 79-81, wherein the Cas9 nuclease comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, 712-713, 715-716, and 719-721, optionally wherein the Cas9 nuclease includes or lacks an N-terminal methionine; or that is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 699-701, 703, 706, 711, 714, 717-718, and 733-736, optionally wherein the nucleic acid includes or lacks a start or stop codon.
83. The system or fusion protein of any one of claims 56-82, further comprising a deaminase.
84. The system or fusion protein of claim 83, wherein the deaminase is an adenosine deaminase or a cytidine deaminase.
85. The system or fusion protein of claim 83 or 84, wherein the deaminase is an APOBEC3A deaminase (A3A).
86. The system or fusion protein of any one of claims 83-85, wherein the deaminase comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to SEQ ID NO:
728.
87. The system or fusion protein of any one of claims 56-86, wherein the Class II Cas nuclease is located N-terminal to the HMGB1 polypeptide.
88. The system or fusion protein of any one of claims 56-86, wherein the Class II Cas nuclease is located C-terminal to the HMGB 1 polypeptide.
89. The system or fusion protein of any one of claims 56-88, further comprising a heterologous nuclear localization signal (NLS).
90. The system or fusion protein of claim 89, wherein the heterologous NLS comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 366-369 and 370-384.
91. The system or fusion protein of claim 89 or 90, wherein the heterologous NLS comprises an SV40 NLS or a nucleoplasmin NLS.
92. The system or fusion protein of any one of claims 89-91, wherein the heterologous NLS comprises a sequence having at least 80%, 85%, 90%, 95%, or 100% identity to any one of SEQ ID NOs: 366-367, 371, 383-384 or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, or 100% identity to the sequence of any one of SEQ ID NOs: 370, 385, and 397.
93. The system or fusion protein of any one of claims 56-92, comprising two heterologous NLSs.
94. The system or fusion protein of any one of claims 56-93, further comprising an SV40 NLS and a nucleoplasmin NLS, optionally wherein the SV40 NLS is located N-terminal to the nucleoplasmin NLS.
95. The system or fusion protein of any one of claims 89-94, wherein at least one heterologous NLS is present at the N-terminus of the HMGB 1 polypeptide or the fusion protein.
96. The fusion protein of any one of claims 58-95, further comprising an intervening peptide sequence, wherein the Class II Cas nuclease is linked to the HMGB1 polypeptide by the intervening peptide sequence.
97. The fusion protein of claim 96, wherein the intervening peptide sequence is at least 10 amino acid residues in length, optionally at least 12 amino acid residues in length.
98. The fusion protein of claim 96 or 97, wherein the intervening peptide sequence is 10-50 amino acid residues in length, optionally 12-50 amino acid residues in length, further optionally 12, 15, 20, 33, 41, or 42 amino acid residues in length.
99. The fusion protein of any one of claims 95-98, wherein the intervening peptide sequence comprises a linker or a heterologous NLS.
100. The fusion protein of claim 99, wherein the intervening peptide sequence comprises, from N-terminal to C-terminal:108.i. linker;109.ii. linker -heterologous NLS;110.iii. heterologous NLS -linker; iv. first linker -heterologous NLS -second linker; or111.v. first linker -first heterologous NLS - second heterologous NLS- second linker.
101. The fusion protein of any one of claims 99-100, comprising, from N-terminus to C-terminus:113.1) Cas- heterologous NLS- linker- HMGB1 polypeptide;114.2) Cas- linker-HMGB1 polypeptide;115.3) Cas- linker- heterologous NLS- linker- HMGB 1 polypeptide;116.4) heterologous NLS- Cas- linker- heterologous NLS- linker- HMGB1 polypeptide;117.5) heterologous NLS- Cas- linker- HMGB1 polypeptide;118.6) heterologous NLS- heterologous NLS- Cas- linker- HMGB 1 polypeptide; or 7) HMGB 1 polypeptide-first linker-first heterologous NLS-second heterologous NLS -second linker- Cas119.8) First heterologous NLS-second heterologous NLS-deaminase-first linker-Cas nickase-second linker-HMGB 1 polypeptide; or120.9) Deaminase-first linker-Cas nickase -NLS-linker-HMGB 1 polypeptide.
102. The fusion protein of any one of claims 99-101, wherein the linker is a GS linker, comprising at least 80% glycine (G) or serine (S).
103. The fusion protein of any one of claims 99-102, comprising, from N-terminus to C-terminus:123.1) SV40 NLS-nucleoplasmin NLS-Cas9-41 AA linker-HMGB1 polypeptide; 2) SV40 NLS-nucleoplasmin NLS-Cas9-33 AA linker-HMGB 1 polypeptide; 3) Cas9-33 AA linker-HMGB1 polypeptide;124.4) SV40 NLS-nucleoplasmin NLS-Cas9-GS20 linker-HMGB 1 polypeptide; 5) SV40 NLS-nucleoplasmin NLS-Cas9-GS15 linker-HMGB1 polypeptide; 6) SV40 NLS-nucleoplasmin NLS-Cas9-GS12 linker-HMGB1 polypeptide; 7) SV40 NLS-Cas9-33 AA linker-HMGB 1 polypeptide;125.8) nucleoplasmin NLS-Cas9-33 AA linker-HMGB1 polypeptide;126.9) nucleoplasmin NLS-Cas9-GS linker-SV40 NLS-GS linker-HMGB1 polypeptide;127.10) Cas9-GS linker-nucleoplasmin NLS-GS linker-HMGB1 polypeptide; or 11) HMGB1 polypeptide-GS linker-SV40 NLS -nucleoplasmin NLS-GS linker- Cas9128.12) SV40 NLS -nucleoplasmin NLS-deaminase-3xGH5 linker-Cas9 nickase -33 AA linker-HMGB1 polypeptide; or129.13) deaminase-GH5 linker-Cas9 nickase -SV40 NLS-41 AA linker- HMGB1 polypeptide.
104. The fusion protein of any one of claims 99-103, wherein the linker comprises the amino acid sequence of any one of SEQ ID NOs: 301-365 and 425-435; or the first linker and the second linker are independently selected from any one of SEQ ID NOs: 301-365 and 425-435.
105. The fusion protein of any one of claims 99-104, wherein the heterologous NLS comprises the amino acid sequence of any one of SEQ ID NOs: 366-369 and 370-384; or wherein the first heterologous NLS and the second heterologous NLS are independently selected from SEQ ID NOs: 366-369 and 370-384.
106. The system or fusion protein of any one of claim 56-105, wherein the Cas9 is a SpyCas9 or an NmeCas9, optionally wherein the NmeCas9 is an Nme2Cas9.
107. The system or fusion protein of any one of claims 56-106, wherein the Cas9 comprises a mutation in a RuvC domain or an HNH domain.
108. The system or fusion protein of claim 106 or 107, wherein the SpyCas9 comprises a substitution at amino acid position H840, D10, or N863, optionally wherein the SpyCas9 comprises a substitution comprising H840A, D10A, or N863A.
109. The system or fusion protein of claim 106 or 107, wherein NmeCas9 comprises a substitution at D16 or H588, optionally wherein the NmeCas9 comprises a substitution comprising H840A, D16A, or H588A.
110. The fusion protein of any one of claims 58-109, wherein the fusion protein comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 515, 567, 591, 597, 603, 624, 627, 630, 633, 636, 639, 642, 649, 652, 655, 664, 667, or 673, or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 513, 514, 565, 566, 589, 590, 595, 596, 601, 602, 622, 623, 625, 626, 628, 629, 631, 632, 634, 635, 637, 638, 640, 641, 647, 648, 650, 651, 653, 654, 662, 663, 665, 666, 671, or 672.
111. A system comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) an NmeCas9 nuclease; wherein the HMGB1 polypeptide is operably linked to the NmeCas9 nuclease.
112. The system of claim 111, wherein the HMGB1 polypeptide and NmeCas9 nuclease are covalently linked in a fusion protein.
113. A fusion protein comprising ( 1 ) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) an NmeCas9 nuclease.
114. The system or fusion protein of any one of claims 111-113, wherein the HMGB1 polypeptide comprises a cryptic nuclear localization signal (NLS).
115. The system or fusion protein of claim 114, wherein the cryptic NLS has the sequence of EKSKKKK (SEQ ID NO: 5).
116. The system or fusion protein of any one of claims 111-115, wherein the HMGB1 polypeptide comprises a receptor for advanced glycation end-products (RAGE) binding domain.
117. The system or fusion protein of claim 116, wherein the RAGE binding domain comprises the sequence of SEQ ID NO: 21.
118. The system or fusion protein of any one of claims 111-117, wherein the HMGB1 polypeptide comprises 7-20 amino acid residues C-terminal to the HMGB1 Box B domain.
119. The system or fusion protein of any one of claims 111-118, wherein the HMGB1 polypeptide comprises the sequence of SEQ ID NO: 22.
120. The system or fusion protein of any one of claims 111-119, wherein the HMGB1 polypeptide lacks a sequence at least 50%, 60%, 70%, 80%, 90%, 94%, or 97% identical to SEQ ID NO: 6 or lacks the sequence of SEQ ID NO: 6.
121. The system or fusion protein of any one of claims 111-120, wherein the HMGB1 polypeptide lacking all or part of an acidic tail domain lacks transcriptional stimulatory function.
122. The system or fusion protein of any one of claims 111-121, wherein the HMGB1 polypeptide comprises a cryptic NLS that is located C-terminal to the HMGB1 Box B domain.
123. The system or fusion protein of claim 122, wherein the at least 7-20 contiguous amino acid residues C-terminal end of the HMGB 1 Box B domain comprise the cryptic NLS.
124. The system or fusion protein of any one of claims 111-123, wherein the HMGB1 polypeptide comprises a cryptic NLS that is located N-terminal to the HMGB1 Box B domain.
125. The system or fusion protein of claim 124, wherein the cryptic NLS immediately precedes the N-terminal end of the HMGB 1 Box B domain.
126. The system or fusion protein of any one of claims 111-125, wherein the HMGB1 polypeptide lacks an HMGB1 Box A domain.
127. The system or fusion protein of any one of claims 111-126, wherein the HMGB1 polypeptide comprises two Box B domains.
128. The system or fusion protein of claim 127, wherein the cryptic NLS is located C-terminal to the C-terminal end of the two HMGB1 Box B domain.
129. The system or fusion protein of any one of claims 111-125 or 127-128, wherein the HMGB1 polypeptide comprises an HMGB1 Box A domain.
130. The system or fusion protein of claim 129, wherein the HMGB 1 Box B domain is located C-terminal to the HMGB 1 Box A domain.
131. The system or fusion protein of any one of claims 111-130, comprising, from N-terminus to C-terminus:155.i. HMGB 1 Box A domain-HMGB 1 Box B domain-cryptic NLS;156.ii. HMGB1 Box B domain-HMGB 1 Box B domain-cryptic NLS; or iii. HMGB 1 Box B domain-cryptic NLS.
132. The system or fusion protein of any one of claims 111-131, wherein the HMGB1 polypeptide comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, and 21-22, or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 12-20.
133. The system or fusion protein of any one of claims 111-132, wherein the NmeCas9 is a cleavase or a nickase.
134. The system or fusion protein of any one of claims 111-132, wherein the NmeCas9 is a dCas9.
135. The system or fusion protein of any one of claims 111-134, wherein the NmeCas9 is an Nme2Cas9.
136. The system or fusion protein of any one of claims 111-135, wherein the NmeCas9 nuclease comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 712-713, 715-716, and 719-721, optionally wherein the Cas9 nuclease includes or lacks an N-terminal methionine; or that is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 711, 714, 717-718, and 733-736, optionally wherein the nucleic acid includes or lacks a start or stop codon.
137. The system or fusion protein of any one of claims 111-136, further comprising a deaminase.
138. The system or fusion protein of claim 137, wherein the deaminase is an adenosine deaminase or a cytidine deaminase.
139. The system or fusion protein of claim 137 or 138, wherein the deaminase is an APOBEC3A deaminase (A3A).
140. The system or fusion protein of any one of claims 137-139, wherein the deaminase comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to SEQ ID NO: 728.
141. The system or fusion protein of any one of claims 111-140, wherein the NmeCas9 is located N-terminal to the HMGB 1 polypeptide.
142. The system or fusion protein of any one of claims 111-140, wherein the NmeCas9 is located C-terminal to the HMGB 1 polypeptide.
143. The system or fusion protein of any one of claims 111-140, further comprising a heterologous nuclear localization signal (NLS).
144. The system or fusion protein of claim 143, wherein the heterologous NLS comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 366-369 and 370-384.
145. The system or fusion protein of claim 143 or 144, wherein the heterologous NLS comprises an SV40 NLS or a nucleoplasmin NLS.
146. The system or fusion protein of any one of claims 143-145, wherein the heterologous NLS comprises a sequence having at least 80%, 85%, 90%, 95%, or 100% identity to any one of SEQ ID NOs: 366-367, 371, 383-384 or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, or 100% identity to the sequence of any one of SEQ ID NOs: 370, 385, and 397.
147. The system or fusion protein of any one of claims 111-146, comprising two heterologous NLSs.
148. The system or fusion protein of any one of claims 111-147, further comprising an SV40 NLS and a nucleoplasmin NLS, optionally wherein the SV40 NLS is located N-terminal to the nucleoplasmin NLS.
149. The system or fusion protein of any one of claims 111-148, wherein the at least one heterologous NLS is present at the N-terminus of the HMGB1 polypeptide or the fusion protein.
150. The fusion protein of any one of claims 111-149, further comprising an intervening peptide sequence, wherein the NmeCas9 nuclease is linked to the HMGB 1 polypeptide by the intervening peptide sequence.
151. The fusion protein of claim 150, wherein the intervening peptide sequence is at least 10 amino acid residues in length, optionally at least 12 amino acid residues in length.
152. The fusion protein of claim 150 or 151, wherein the intervening peptide sequence is 10-50 amino acid residues in length, optionally 12-50 amino acid residues in length, further optionally 12, 15, 20, 33, 41, or 42 amino acid residues in length.
153. The fusion protein of any one of claims 150-152, wherein the at least one intervening peptide sequence comprises a linker or a heterologous NLS.
154. The fusion protein of claim 153, wherein the intervening peptide sequence comprises, from N-terminus to C-terminus,178.1) linker;179.2) linker -heterologous NLS;180.3) heterologous NLS -linker;181.4) first linker -heterologous NLS -second linker; or182.5) first linker -first heterologous NLS - second heterologous NLS- second linker.
155. The fusion protein of any one of claims 153-154, wherein the linker is a GS linker comprising at least 80% glycine (G) or serine (S).
156. The fusion protein of any one of claims 153-155, comprising, from N-terminus to C-terminus:185.1) SV40 NLS-nucleoplasmin NLS-Nme2Cas9-41 AA linker-HMGB1 polypeptide; 2) S V40 NLS-nucleoplasmin NLS-Nme2Cas9-33 AA linker-HMGB 1 polypeptide; 3) Nme2Cas9-33 AA linker-HMGB1 polypeptide;186.4) S V40 NLS-nucleoplasmin NLS-Nme2Cas9-GS20 linker-HMGB 1 polypeptide; 5) SV40 NLS-nucleoplasmin NLS-Nme2Cas9-GS 15 linker-HMGB 1 polypeptide; 6) S V40 NLS-nucleoplasmin NLS-Nme2Cas9-GS 12 linker-HMGB 1 polypeptide; 7) S V40 NLS-Nme2Cas9-33 AA linker-HMGB 1 polypeptide;187.8) nucleoplasmin NLS-Nme2Cas9-33 AA linker-HMGB 1 polypeptide;188.9) nucleoplasmin NLS-Nme2Cas9-GS linker-SV40 NLS-GS linker- HMGB 1 polypeptide;189.10) Nme2Cas9-GS linker-nucleoplasmin NLS-GS linker-HMGB 1 polypeptide;190.11) HMGB1 polypeptide-GS linker-SV40 NLS-nucleoplasmin NLS-GS linker- Nme2Cas9; or191.12) SV40 NLS-nucleoplasmin NLS-deaminase-3xGH5 linker-Nme2Cas9-33 AA linker-HMGB 1 polypeptide.
157. The fusion protein of any one of claims 153-156, wherein the linker comprises the amino acid sequence of any one of SEQ ID NOs: 301-365 and 425-435; or the first linker and the second linker are independently selected from any one of SEQ ID NOs: 301-365 and 425-435.
158. The fusion protein of any one of claims 153-157, wherein the heterologous NLS comprises the amino acid sequence of any one of SEQ ID NOs: 366-369 and 370-384; or wherein the first heterologous NLS and the second heterologous NLS are independently selected from SEQ ID NOs: 366-369 and 370-384.
159. The fusion protein of any one of claims 153-158, wherein the fusion protein comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 515, 567, 591, 597, 603, 624, 627, 630, 633, 636, 639, 642, 649, 652, 655, or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 513, 514, 565, 566, 589, 590, 595, 596, 601, 602, 622, 623, 625, 626, 628, 629, 631, 632, 634, 635, 637, 638, 640, 641, 647, 648, 650, 651, 653, or 654.
160. The fusion protein of any one of claims 153-159, wherein the fusion protein comprises, from N-terminus to C-terminus:196.i. a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434 ii. an SV40 NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 366; iii. a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO:197.384;198.iv. a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433; v. an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 715;199.vi. a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429; and vii. an HMGB1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.
161. The fusion protein of any one of claims 153-159, wherein the fusion protein comprises, from N-terminus to C-terminus:201.i. a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 434 ii. a nucleoplasmin NLS having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 384;202.iii. a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 433; iv. an NmeCas9 protein having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO:203.715;204.v. a linker having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 429; vi. an HMGB 1 polypeptide having at least 90%, 95%, 99%, or 100% identity to SEQ ID NO: 7.
162. The fusion protein of any one of claims 153-160, wherein the fusion protein comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 567 and 591 or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 565-566 and 589-590.
163. A system comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; (2) a SpyCas9 nuclease; wherein the HMGB1 polypeptide is operably linked to the SpyCas9 nuclease.
164. The system of claim 163, wherein the HMGB1 polypeptide and SpyCas9 are covalently linked in a fusion protein.
165. A fusion protein comprising ( 1 ) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a SpyCas9 nuclease.
166. The system or fusion protein of any one of claims 163-165, wherein the HMGB1 polypeptide comprises a cryptic nuclear localization signal (NLS).
167. The system or fusion protein of claim 166, wherein the cryptic NLS comprises the sequence of EKSKKKK (SEQ ID NO: 5).
168. The system or fusion protein of any one of claims 163-167, wherein the HMGB1 polypeptide comprises a receptor for advanced glycation end-products (RAGE) binding domain.
169. The system or fusion protein of claim 168, wherein the RAGE binding domain comprises the sequence of SEQ ID NO: 21.
170. The system or fusion protein of any one of claims 163-169, wherein the HMGB1 polypeptide comprises 7-20 amino acid residues C-terminal to the HMGB1 Box B domain.
171. The system or fusion protein of any one of claims 163-170, wherein the HMGB1 polypeptide comprises the sequence of SEQ ID NO: 22.
172. The system or fusion protein of any one of claims 163-171, wherein the HMGB1 polypeptide lacks a sequence at least 50%, 60%, 70%, 80%, 90%, 94%, or 97% identical to SEQ ID NO: 6 or lacks the sequence of SEQ ID NO: 6.
173. The system or fusion protein of any one of claims 163-172, wherein the HMGB1 polypeptide lacking all or part of an acidic tail domain lacks transcriptional stimulatory function.
174. The system or fusion protein of any one of claims 163-173, wherein the HMGB1 polypeptide comprises a cryptic NLS that is located C-terminal to the HMGB1 Box B domain.
175. The system or fusion protein of claim 174, wherein the at least 7-20 contiguous amino acid residues at the C-terminal end of the HMGB 1 Box B domain comprise the cryptic NLS.
176. The system or fusion protein of any one of claims 163-175, wherein the HMGB1 polypeptide comprises a cryptic NLS that is located N-terminal to the HMGB1 Box B domain.
177. The system or fusion protein of claim 176, wherein the cryptic NLS immediately precedes the N-terminal end of the HMGB 1 Box B domain.
178. The system or fusion protein of any one of claims 163-177, wherein the HMGB1 polypeptide lacks an HMGB1 Box A domain.
179. The system or fusion protein of any one of claims 163-178, wherein the HMGB1 polypeptide comprises two Box B domains.
180. The system or fusion protein of claim 179, wherein the cryptic NLS is located C-terminal to the C-terminal end of the two HMGB1 Box B domain.
181. The system or fusion protein of any one of claims 163-177 or 179-180, wherein the HMGB1 polypeptide comprises an HMGB1 Box A domain.
182. The system or fusion protein of claim 181, wherein the HMGB 1 Box B domain is located C-terminal to the HMGB 1 Box A domain.
183. The system or fusion protein of any one of claims 163-182, comprising, from N-terminus to C-terminus:224.i. HMGB 1 Box A domain-HMGB 1 Box B domain-cryptic NLS;225.ii. HMGB1 Box B domain-HMGB 1 Box B domain-cryptic NLS; or226.iii. HMGB1 Box B domain-cryptic NLS.
184. The system or fusion protein of any one of claims 163-183, wherein the HMGB1 polypeptide comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 2-5, 7-10, and 21-22 or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 12-20.
185. The system or fusion protein of any one of claims 163-184, wherein the SpyCas9 is a cleavase or a nickase.
186. The system or fusion protein of any one of claims 163-184, wherein the SpyCas9 is a dCas9.
187. The fusion protein of any one of claims 165-186, wherein the SpyCas9 nuclease comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 702, 704-705, 707-708, optionally wherein the SpyCas9 nuclease includes or lacks an N-terminal methionine or that is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 699-701, 703, and 706, optionally wherein the nucleic acid includes or lacks a start or stop codon.
188. The system or fusion protein of any one of claims 163-187, wherein the fusion protein further comprises a deaminase.
189. The system or fusion protein of claim 188, wherein the deaminase is an adenosine deaminase or a cytidine deaminase.
190. The system or fusion protein of claim 188 or 189, wherein the deaminase is an APOBEC3A deaminase (A3A).
191. The system or fusion protein of any one of claims 188-190, wherein the deaminase comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to SEQ ID NO: 728.
192. The system or fusion protein of any one of claims 163-191, wherein the SpyCas9 is located N-terminal to the HMGB 1 polypeptide.
193. The system or fusion protein of any one of claims 163-191, wherein the SpyCas9 is located C-terminal to the HMGB 1 polypeptide.
194. The system or fusion protein of any one of claims 163-193, further comprising a heterologous nuclear localization signal (NLS).
195. The system or fusion protein of claim 194, wherein the heterologous NLS comprises a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 366-369 and 370-384.
196. The system or fusion protein of claim 194 or 195, wherein the heterologous NLS comprises an SV40 NLS or a nucleoplasmin NLS.
197. The system or fusion protein of any one of claims 194-196, wherein the at least one heterologous NLS comprises a sequence having at least 80%, 85%, 90%, 95%, or 100% identity to any one of SEQ ID NOs: 366-367, 371, 383-384 or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, or 100% identity to the sequence of any one of SEQ ID NOs: 370, 385, and 397.
198. The system or fusion protein of any one of claims 163-197, comprising two heterologous NLSs.
199. The system or fusion protein of any one of claims 163-198, comprising an SV40 NLS and a nucleoplasmin NLS, optionally wherein the SV40 NLS is located N-terminal to the nucleoplasmin NLS.
200. The system or fusion protein of any one of claims 163-199, wherein at least one heterologous NLS is present at the N-terminus of the HMGB1 polypeptide or the fusion protein.
201. The fusion protein of any one of claims 165-200, further comprising an intervening peptide sequence, wherein the SpyCas9 nuclease is linked to the HMGB 1 polypeptide by the intervening peptide sequence.
202. The fusion protein of claim 201, wherein the intervening peptide sequence is at least 10 amino acid residues in length, optionally at least 12 amino acid residues in length.
203. The fusion protein of claim 201 or 202, wherein the intervening peptide sequence is 10-50 amino acid residues in length, optionally 12-50 amino acid residues in length, further optionally 12, 15, 20, 33, 41, or 42 amino acid residues in length.
204. The fusion protein of any one of claims 201-203, wherein the intervening peptide sequence comprises a linker or a heterologous NLS.
205. The fusion protein of any one of claims 201-204, wherein the intervening peptide sequence comprises, from N-terminal to C-terminal:247.i. linker;248.ii. linker -heterologous NLS;249.iii. heterologous NLS -linker;250.iv. first linker -heterologous NLS -second linker; or251.v. first linker -first heterologous NLS - second heterologous NLS- second linker.
206. The fusion protein of any one of claims 204-205, comprising, from N-terminus to C-terminus:252.1) SpyCas9-SV40 NLS-41 AA linker-HMGBl polypeptide;253.2) S V40 NLS-nucleoplasmin NLS-SpyCas9-33 AA linker-HMGB 1 polypeptide; or 3) Deaminase-GH5 linker-SpyCas9 nickase-SV40NLS-41 AA linker-HMGB 1 polypeptide.
207. The fusion protein of any one of claims 204-206, wherein the linker comprises the amino acid sequence of any one of SEQ ID NOs: 301-365 and 425-435; or the first linker and the second linker are independently selected from any one of SEQ ID NOs: 301-365 and 425-435.
208. The fusion protein of any one of claims 204-207, wherein the heterologous NLS comprises the amino acid sequence of any one of SEQ ID NOs: 366-369 and 370-384; or wherein the first heterologous NLS and the second heterologous NLS are independently selected from SEQ ID NOs: 366-369 and 370-384.
209. The fusion protein of any one of claims 204-208, wherein the fusion protein comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to SEQ ID NO: 588 or is encoded by a nucleic acid having at least 80%, 85%, 90%, 95%, 98% or 100% identity to the sequence of any one of SEQ ID NOs: 586 and 587.
210. An mRNA comprising an open reading frame (ORF) encoding an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain.
211. An mRNA comprising an open reading frame (ORF) encoding a fusion protein comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB 1 Box B domain, wherein the HMGB 1 polypeptide lacks all or part of an acidic tail domain; and (2) a Class II Cas nuclease.
212. An mRNA comprising an open reading frame (ORF) encoding a fusion protein comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB 1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) an NmeCas9 nuclease.
213. An mRNA comprising an open reading frame (ORF) encoding a fusion protein comprising (1) an HMGB1 polypeptide, wherein the HMGB1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB 1 Box B domain, wherein the HMGB 1 polypeptide lacks all or part of an acidic tail domain; and (2) a SpyCas9 nuclease.
214. An mRNA comprising an open reading frame encoding the HMGB 1 polypeptide or fusion protein of any one of claims 1-209.
215. The mRNA of any one of claims 210-214, wherein at least 10% of the uridine in the mRNA is substituted with a modified uridine.
216. The mRNA of claim 215, wherein the modified uridine is one or more of Nl-methyl-pseudouridine, pseudouridine, 5 -methoxyuridine, or 5 -iodouridine.
217. The mRNA of claim 215 or 216, wherein at least 20%, 30%, 80%, 90%, or 95% of the uridine is substituted with the modified uridine.
218. The mRNA of any one of claims 215-217, wherein 100% of the uridine is substituted with the modified uridine.
219. The mRNA of any one of claims 210-218, wherein the mRNA comprises a 5’ UTR having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 398-405.
220. The mRNA of any one of claims 210-219, wherein the mRNA comprises a 3’ UTR having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 406-413.
221. The mRNA of any one of claims 210-220, wherein the mRNA further comprises a 5’ cap, optionally wherein the cap is selected from CapO, Capl, or Cap2.
222. The mRNA of any one of claims 210-221, wherein the mRNA further comprises a polyadenylated (poly- A) tail, optionally wherein the poly-A tail is an encoded poly-A tail comprising a sequence of SEQ ID NO: 424.
223. The mRNA of any one of claims 210-222, wherein the mRNA comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 12-20.
224. The mRNA of any one of claims 210-223, wherein the mRNA comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to SEQ ID NO: 17.
225. The mRNA of any one of claims 210-224, wherein the mRNA comprises a sequence having at least 80%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 513, 522, 525, 528, 531, 534, 537, 540, 547, 550, 553, 565, 571, 574, 586, 589, 595, and 601.
226. A composition comprising an mRNA comprising an open reading frame (ORF) encoding:272.1 ) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB 1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain;273.2) a fusion protein comprising (1) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a Class II Cas nuclease; 3) a fusion protein comprising (1) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) an NmeCas9 nuclease; or 4) a fusion protein comprising (1) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a SpyCas9 nuclease; optionally wherein the composition comprises a guide RNA (gRNA).
227. A composition comprising the mRNA of any one of claims 210-225, optionally wherein the composition comprises a guide RNA (gRNA).
228. The composition of claim 226 or 227, wherein the mRNA or the guide RNA is associated with one or more lipid nanoparticles (LNP).
229. The composition of any one of claims 226-228, wherein the mRNA and guide RNA are associated with the same lipid nanoparticle (LNP).
230. The composition of any one of claims 226-229, wherein the mRNA and guide RNA are each associated with a separate lipid nanoparticle (LNP).
231. A lipid nanoparticle (LNP) composition comprising an mRNA comprising an open reading frame (ORF) encoding:278.1 ) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB 1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain;279.2) a fusion protein comprising (1) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a Class II Cas nuclease; 3) a fusion protein comprising (1) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) an NmeCas9 nuclease; or 4) a fusion protein comprising (1) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a SpyCas9 nuclease; optionally further comprising a guide RNA (gRNA).
232. A lipid nanoparticle (LNP) composition comprising the mRNA of any one of claims 210-225, optionally further comprising a guide RNA (gRNA).
233. The composition or LNP composition of claim 232, wherein the guide RNA is a single guide RNA (sgRNA).
234. The composition or LNP composition of claim 233, wherein the sgRNA is a SpyCas9 guide RNA.
235. The composition or LNP composition of claim 233, wherein the sgRNA is an NmeCas9 guide RNA.
236. The composition or LNP composition of claim 234, wherein the SpyCas9 guide RNA comprises a modified nucleotide sequence having at least 90% identity to any one of SEQ ID NOs: 140-159 and 167-237.
237. The composition or LNP composition of claim 234 or 236, wherein the SpyCas9 guide RNA comprises a modified nucleotide sequence of any one of SEQ ID NOs: 140-159 and 167- 237.
238. The LNP composition of any one of claims 231-237, wherein the LNP comprises (i) an ionizable lipid; (ii) a helper lipid; (iii) a stealth lipid; (iv) a neutral lipid; or combinations of one or more of (i)-(iv).
239. The composition or LNP composition of any one of claims 226-238, wherein the fusion protein further comprises a cytidine deaminase.
240. The composition or LNP composition of claim 239, further comprising a uracil glycosylase inhibitor (UGI) or an mRNA encoding a UGI.
241. A vector comprising a sequence encoding the mRNA of any one of claims 210-225 or an expression construct comprising a promoter operably linked to a sequence encoding the mRNA of any one of claims 210-225, optionally wherein the expression construct is a plasmid.
242. A host cell comprising the vector or expression construct of claim 241.
243. A pharmaceutical composition comprising the HMGB1 polypeptide, fusion protein, mRNA, composition, or LNP composition of any one of claims 1-240, and a pharmaceutically acceptable carrier.
244. A method of producing a modification in a genomic sequence of a target cell, the method comprising contacting the cell with:292.1 ) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB 1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain;293.2) a fusion protein comprising (1) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a Class II Cas nuclease; 3) a fusion protein comprising (1) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) an NmeCas9 nuclease; 4) a fusion protein comprising (1) an HMGB 1 polypeptide, wherein the HMGB 1 polypeptide comprises an HMGB1 Box B domain and at least 7-20 contiguous amino acid residues C-terminal to the HMGB1 Box B domain, wherein the HMGB1 polypeptide lacks all or part of an acidic tail domain; and (2) a SpyCas9 nuclease; or 5) an mRNA comprising an open reading frame (ORF) encoding any one of (l)-(4).
245. A method of producing a modification in a genomic sequence of a target cell, the method comprising contacting the cell with the HMGB1 polypeptide, fusion protein, mRNA, composition, or LNP composition of any one of claims 1-243.
246. The method of claim 244 or 245, wherein the modification comprises an insertion, deletion, or substitution in the genomic DNA.
247. The method of any one of claims 244-246, wherein the modification results in a change in an amino acid sequence encoded by a locus in which the genomic sequence is present.
248. The method of claim 247, wherein the change in the amino acid sequence encoded by a locus in which the genomic sequence is present does not require indel formation.
249. The method of any one of claims 244-248, wherein the cell is in a subject.
250. The method of any one of claims 244-249, wherein the method produces an increased level of modifications in target cells relative to cells not contacted with a DNA-binding protein.
251. The method of any one of claims 244-250, wherein the modification does not result in upregulation or downregulation of the transcriptome of the target cell relative to a cell contacted with a DNA-binding protein.
252. The method of any one of claims 244-251, wherein the HMGB1 polypeptide, fusion protein, mRNA encoding the fusion protein, and optionally a guide RNA are delivered to the target cell via electroporation.
253. The method of any one of claims 244-252, wherein the HMGB1 polypeptide, fusion protein, mRNA encoding the fusion protein, and optionally a guide RNA are delivered to the target cell via a lipid nanoparticle.
254. The method of any one of claims 249-252, wherein the modification is in vivo.
255. The method of any one of claims 249-252, wherein the modification is in vitro.
256. An engineered cell or population of engineered cells altered by the method of any one of claims 244-252.
257. The population of engineered cells of claim 256, wherein contacting the cells with the HMGB 1 polypeptide, fusion protein, mRNA, composition, or LNP composition results in a higher percentage of modified cells than contacting a second cell population not contacted with the composition comprising the HMGB 1 polypeptide.
258. The engineered cell or population of engineered cells of claim 256, wherein contacting the cell or population of cells with the HMGB 1 polypeptide, fusion protein, mRNA composition, or LNP composition does not result in upregulation or downregulation of the transcriptome of the engineered cell or population of engineered cells relative to a second cell or population of cells not contacted with the composition comprising an HMGB 1 polypeptide.
259. A kit comprising the HMGB1 polypeptide, fusion protein, mRNA, composition, or LNP composition of any one of claims 1-243.
260. The HMGB1 polypeptide, fusion protein, mRNA, composition, or LNP composition of any one of claims 1-243 for use in producing a modification in a genomic sequence of a target cell.
261. Use of the HMGB 1 polypeptide, fusion protein, mRNA, composition, or LNP composition of any one of claims 1-243 for producing a modification in a genomic sequence of a target cell.
262. Use of the HMGB 1 polypeptide, fusion protein, mRNA, composition, or LNP composition of any one of claims 1-243 for the manufacture of a medicament for producing a modification in a genomic sequence of a target cell.
Citation Information
Patent Citations
Engineered CRISPR-CAS9 NUCLEASES WITH ALTERED PAM SPECIFICITY
US20160312198A1
Engineered CRISPR-CAS9 Nucleases with Altered PAM Specificity
US20160312199A1
RNA Modification to Engineer Cas9 Activity
US20170114334A1
Backbone modified oligonucleotide analogs
US5378825A
Linking reagents for nucleotide probes
US5585481A