Polynucleotides, compositions, and methods for genome editing involving deamination
Patent Information
- Application Number
- JP2023535332
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-03
- Filing Date
- 2021-12-10
- Publication Date
- 2025-09-04
AI Technical Summary
Current genome editing techniques using DNA base editors, such as BE3 and APOBEC3A, suffer from high off-target editing and bystander mutations, necessitating improved compositions and methods for targeted C-to-T base editing with reduced bystander mutations.
A method involving mRNA encoding a cytidine deaminase, such as APOBEC3A, and an RNA-guided nickase, optionally with a uracil glycosylase inhibitor (UGI), is used to enhance the specificity and fidelity of C-to-T conversions in target nucleotides, minimizing bystander mutations.
The method achieves higher purity and reduced bystander mutations in C-to-T edits, improving the accuracy and efficiency of genome editing.
Abstract
Description
[Technical Field]
[0001] This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Patent Application No. 63 / 124,060, filed December 11, 2020, U.S. Provisional Patent Application No. 63 / 130,104, filed December 23, 2020, U.S. Provisional Patent Application No. 63 / 165,636, filed March 24, 2021, and U.S. Provisional Patent Application No. 63 / 275,424, filed November 3, 2021, each of which is incorporated by reference in its entirety.
[0002] This application is filed with an electronic Sequence Listing, which is provided as a file entitled "2021-12-08_01155-0016-00PCT_ST25.txt," created on December 8, 2021, and is 1,557,107 bytes in size. The information in the electronic format of this Sequence Listing is incorporated herein by reference in its entirety. Summary of the Invention
[0003] The present disclosure relates to polynucleotides, compositions, and methods for genome editing involving deamination.
[0004] Base editing is a genome editing technique that directly generates point mutations in specific regions of genomic DNA without causing double-strand breaks (DSBs). DNA base editors (BEs) contain fusions of catalytically impaired Cas nucleases with base-modifying enzymes. Currently, the effector for cytidine-to-thymidine (C-to-T) editing fuses cytidine deaminase with a nickase and a uracil glycosylase inhibitor (UGI). For example, base editor 3 (BE3) consists of a Cas9 nuclease fused to APOBEC1 (apolipoprotein mRNA editing enzyme, catalytic polypeptide 1) deaminase and UGI, mutated to convert it into a nickase (nCas9) (e.g., Wang et al., Cell Research 27:1289-1292 (2017)). It was reported that the nCas9-fused UGI domain remains important for achieving high fidelity base editing, even in the presence of high levels of free UGI. As another example, an engineered human APOBEC3A (A3A or apolipoprotein mRNA editing enzyme, catalytic polypeptide 3A) deaminase was investigated as a replacement for the rat APOBEC1 deaminase (RAPO1) in the original BE3 (Gehrke et al., Nature Biotechnology, 36:977-982 (2018)). However, it was noted that "the ability of a base editor to edit all Cs within its editing window may have deleterious effects" and that "mutation of the N57 residue of human A3A deaminase was important to restore the accuracy of the original target sequence in the context of the base editor and further reduce its off-target editing activity." In fact, an APOBEC3A-class 2 Cas nickase (D10A) base editor has been reported to be "highly mutagenic" and exhibit Cas9-independent off-target base editing (Doman et al., Nature Biotechnology, 38:620-628 (2020)). Thus, there is a need for improved compositions and methods for targeted C-to-T base editing using cytidine deaminase (e.g., APOBEC3A deaminase) and RNA-guided nickase.
[0005] Thus, the present disclosure provides polynucleotides, compositions, and methods for genome editing involving a cytidine deaminase (e.g., an APOBEC3A deaminase) and an RNA-guided nickase that can induce C-to-T conversions at targeted nucleotides with greater fidelity and minimize bystander mutations. The present disclosure is based, in part, on the discovery that pairing a cytidine deaminase (e.g., an APOBEC deaminase) and an RNA-guided nickase system with a UGI in trans (e.g., as a separate mRNA) can reduce the amount of other base edits (C-to-A / G conversions, insertions, or deletions) and increase the purity of the C-to-T edit.
[0006] Therefore, the following embodiments are provided.
[0007] In some embodiments, a composition is provided comprising a first mRNA comprising a first open reading frame encoding a polypeptide comprising cytidine deaminase and RNA-guided nickase, and a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), the second mRNA being different from the first mRNA. In some embodiments, the first open reading frame does not comprise a sequence encoding UGI. In some embodiments, the composition comprises a first composition and a second composition, the first composition comprising a first mRNA comprising a first open reading frame encoding a polypeptide comprising cytidine deaminase and RNA-guided nickase, but not a uracil glycosylase inhibitor (UGI), and the second composition comprising a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), the second mRNA being different from the first mRNA. In some embodiments, the composition comprises lipid nanoparticles.
[0008] In some embodiments, methods are provided for modifying a target gene, the method comprising delivering to a cell a first mRNA comprising a first open reading frame encoding a first polypeptide comprising a cytidine deaminase and an RNA-guided nickase, a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), the second mRNA being distinct from the first mRNA, and at least one guide RNA (gRNA).
[0009] In some embodiments, a method for modifying at least one cytidine in a target gene in a cell is provided, the method comprising expressing in a cell or contacting the cell with (i) a first polypeptide comprising a cytidine deaminase and an RNA-guided nickase, the first polypeptide not comprising a uracil glycosylase inhibitor (UGI), (ii) a UGI polypeptide, and (iii) at least one guide RNA (gRNA), wherein the first polypeptide and the gRNA form a complex with the target gene and modify at least one cytidine in the target gene. In some embodiments, the ratio of the UGI polypeptide to the first polypeptide is 10:1 to 50:1.
[0010] In some embodiments, provided is an mRNA comprising an open reading frame (ORF) encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, the polypeptide not comprising a uracil glycosylase inhibitor (UGI). Polypeptides encoded by the mRNA are also provided. In some embodiments, provided is a method of modifying a target gene, the method comprising delivering to a cell an mRNA or a polypeptide described herein.
[0011] In some embodiments, a composition comprises two different mRNAs, where a first mRNA comprises an ORF encoding a cytidine deaminase (e.g., A3A) and an RNA-guided nickase, and a second mRNA comprises an ORF encoding a uracil glycosylase inhibitor (UGI). In some embodiments, the first mRNA in the composition does not comprise an ORF encoding a UGI. In some embodiments, the molar ratio of the second mRNA to the first mRNA is 1:1 to 30:1, 2:1 to 30:1, 7:1 to 22:1. In some embodiments, the molar ratio of the second mRNA to the first mRNA is 22:1, 7:1, 2:1, or 1:1, 1:4, 1:11, or 1:33.
[0012] Further embodiments are provided and described through the claims and drawings. [Brief explanation of the drawings]
[0013] [Figure 1A] Figure 1 shows C-to-T converting active deaminase editors profiled against five different guide target sequences. [Figure 1B] Figure 1 shows C-to-T converting active deaminase editors profiled against five different guide target sequences. [Figure 1C] Figure 1 shows C-to-T converting active deaminase editors profiled against five different guide target sequences. [Figure 1D] Figure 1 shows C-to-T converting active deaminase editors profiled against five different guide target sequences. [Figure 1E] Figure 1 shows C-to-T converting active deaminase editors profiled against five different guide target sequences. [Figure 2A] For deaminase editors profiled using sg000296, the percentage of edited reads in which the four targeted cytosines were converted to thymidines is shown. [Figure 2B]For deaminase editors profiled using sg0001373, the percentage of edited reads in which the five targeted cytosines were converted to thymidines is shown. [Figure 2C] For deaminase editors profiled using sg001400, the percentage of edited reads in which the four targeted cytosines were converted to thymidines is shown. [Figure 2D] For deaminase editors profiled using sg003018, the percentage of edited reads in which the six targeted cytosines were converted to thymidines is shown. [Figure 2E] For deaminase editors profiled using sg005883, the percentage of edited reads in which at least six of the eight targeted cytosines were converted to thymidines is shown. [Figure 3A] The percentage of C-to-T conversions at each base position for each of the five different guide target sequences is shown. [Figure 3B] The percentage of C-to-T conversions at each base position for each of the five different guide target sequences is shown. [Figure 3C] The percentage of C-to-T conversions at each base position for each of the five different guide target sequences is shown. [Figure 3D] The percentage of C-to-T conversions at each base position for each of the five different guide target sequences is shown. [Figure 3E] The percentage of C-to-T conversions at each base position for each of the five different guide target sequences is shown. [Figure 4] Editing profiles are shown as percentages of total reads in U-2OS cells (A) and HuH-7 cells (B). [Figure 5] Editing profiles are shown as the percentage of edited reads in U-2OS cells (A) and HuH-7 cells (B). [Figure 6] Editing profiles are expressed as % of total reads in titrations of UGI mRNA (SEQ ID NOs: 25 and 34). [Figure 7]Figure 1 shows TTR editing levels in CD-1 mice treated with different LNP combinations with mRNA constructs and sgRNA. The editing rates (%) of C-to-T conversions, C-to-A / G conversions, and indels in DNA sequences (NGS sequencing) extracted from liver tissue samples collected from CD-1 mice are shown. [Figure 8] Shows TTR editing levels in liver tissue taken from CD-1 mice treated with different LNP combinations with mRNA constructs and sgRNAs when the UGI sequence was delivered in trans (as a separate mRNA). [Figure 9A] Scatter plots depicting statistically significant (adjusted p-value < 0.05) differential gene expression events (black dots) in liver samples from mice treated with sgRNA G000282 and BC22n mRNA in the absence of UGI mRNA in trans are shown. [Figure 9B] Scatter plots depicting statistically significant (adjusted p-value < 0.05) differential gene expression events (black dots) in liver samples from mice treated with sgRNA G000282 and BC22n mRNA in the presence of UGI mRNA in trans are shown. [Figure 9C] Scatter plots depicting statistically significant (adjusted p-value < 0.05) differential gene expression events (black dots) in liver samples from mice treated with Cas9 mRNA in the presence of UGI mRNA in trans are shown. [Figure 10] 1 shows the editing profile in T cells after treatment with different mRNA constructs and CIITA-targeting sgRNA. [Figure 11] MHC class II negative cells assessed by flow cytometry analysis of T cells treated with different mRNA constructs and CIITA guide RNA are shown. [Figure 12] Scatter plots showing statistically significant (* = adjusted p-value < 0.05) differential gene expression events (black dots) in T cells treated with sgRNA G018076, UGI mRNA, and Cas9 mRNA (A) or sgRNA G018076, UGI mRNA, and BC22n mRNA (B) are shown. [Figure 13]Scatter plots showing statistically significant (* = adjusted p-value < 0.05) differential gene expression events (black dots) in T cells treated with sgRNA G018117, UGI mRNA, and Cas9 mRNA (A) or sgRNA G018117, UGI mRNA, and BC22n mRNA (B) are shown. [Figure 14] Figure 1 shows the protein-protein interaction networks enriched in the differentially expressed gene list in T cells treated with sgRNA G018076, UGI mRNA, and Cas9 mRNA (A) or sgRNA G018076, UGI mRNA, and BC22n mRNA (B). [Figure 15] Figure 1 shows the protein-protein interaction networks enriched in the differentially expressed gene list in T cells treated with sgRNA G018117, UGI mRNA, and Cas9 mRNA (A) or sgRNA G018117, UGI mRNA, and BC22 mRNA (B). [Figure 16A] T cell editing profiles are shown. The editing profile of the on-target TRAC locus (A) is shown in addition to the editing profiles of 10 loci (B-C) previously described as mutational hotspots in APOBEC-positive tumors. [Figure 16B] T cell editing profiles are shown. The editing profile of the on-target TRAC locus (A) is shown in addition to the editing profiles of 10 loci (B-C) previously described as mutational hotspots in APOBEC-positive tumors. [Figure 16C] T cell editing profiles are shown. The editing profile of the on-target TRAC locus (A) is shown in addition to the editing profiles of 10 loci (B-C) previously described as mutational hotspots in APOBEC-positive tumors. [Figure 17A] The editing profile of T cells treated with various levels of BC22n mRNA and Cas9 mRNA is shown. Cells were edited using sgRNA G015995. [Figure 17B]The editing profile of T cells treated with various levels of BC22n mRNA and Cas9 mRNA is shown. Cells were edited using sgRNA G016017. [Figure 17C] The editing profile of T cells treated with various levels of BC22n mRNA and Cas9 mRNA is shown. Cells were edited using sgRNA G016206. [Figure 17D] The editing profile of T cells treated with various levels of BC22n mRNA and Cas9 mRNA is shown. Cells were edited using sgRNA G018117. [Figure 17E] The editing profile of T cells treated with various levels of BC22n mRNA and Cas9 mRNA is shown. Cells were edited using sgRNA G016086. [Figure 18A] Editing profiles of T cells simultaneously edited with four guides using varying levels of BC22n mRNA or Cas9 mRNA are shown. The editing profile at each edited locus is represented separately: G015995 (A), G016017 (B), G016206 (C), and G018117 (D). [Figure 18B] Editing profiles of T cells simultaneously edited with four guides using varying levels of BC22n mRNA or Cas9 mRNA are shown. The editing profile at each edited locus is represented separately: G015995 (A), G016017 (B), G016206 (C), and G018117 (D). [Figure 18C] Editing profiles of T cells simultaneously edited with four guides using varying levels of BC22n mRNA or Cas9 mRNA are shown. The editing profile at each edited locus is represented separately: G015995 (A), G016017 (B), G016206 (C), and G018117 (D). [Figure 18D]Editing profiles of T cells simultaneously edited with four guides using varying levels of BC22n mRNA or Cas9 mRNA are shown. The editing profile at each edited locus is represented separately: G015995 (A), G016017 (B), G016206 (C), and G018117 (D). [Figure 19A] Phenotypic analysis results as the percentage of cells negative for antibody binding with increasing total RNA for BC22n and Cas9 samples are shown. The percentage of B2M-negative cells is shown when B2M guide G015995 was used for editing. [Figure 19B] Phenotyping results as the percentage of cells negative for antibody binding with increasing total RNA are shown for BC22n and Cas9 samples. The percentage of B2M-negative cells is shown when multiple guides are used for editing. [Figure 19C] Phenotypic analysis of BC22n and Cas9 samples as a percentage of cells negative for antibody binding with increasing total RNA is shown. The percentage of CD3-negative cells is shown when TRAC guide G016017 was used for editing. [Figure 19D] The results of phenotypic analysis of BC22n and Cas9 samples as a percentage of cells negative for antibody binding with increasing total RNA are shown. The percentage of CD3-negative cells is shown when TRBC guide G016206 was used for editing. [Figure 19E] Phenotyping results as the percentage of cells negative for antibody binding with increasing total RNA are shown for BC22n and Cas9 samples. The percentage of CD3-negative cells is shown when multiple guides are used for editing. [Figure 19F] Phenotypic analysis results as the percentage of cells negative for antibody binding with increasing total RNA for BC22n and Cas9 samples are shown. The percentage of MHC class II-negative cells is shown when CIITA guide G018117 was used for editing. [Figure 19G]Phenotyping results as the percentage of cells negative for antibody binding with increasing total RNA are shown for BC22n and Cas9 samples. The percentage of MHC class II-negative cells is shown when multiple guides are used for editing. [Figure 19H] Phenotyping results as the percentage of cells negative for antibody binding with increasing total RNA are shown for BC22n and Cas9 samples. The percentage of triple (B2M, CD3, MHC II) negative cells is shown when multiple guides are used for editing. [Figure 19I] Phenotypic analysis results as the percentage of cells negative for antibody binding with increasing total RNA are shown for BC22n and Cas9 samples. The percentage of MHC class II-negative T cells is shown when CIITA guide G016086 was used for editing. [Figure 20A] Figure 1 shows the editing profile of T cells using various levels of trans UGI mRNA and various editor mRNAs. The percentage of editing with 27.3 nM BC22n mRNA is shown. [Figure 20B] Figure 1 shows the editing profile of T cells using various levels of trans UGI mRNA and various editor mRNAs. The percent editing rate with 24.7 nM BC22-2×UGI mRNA is shown. [Figure 20C] Figure 1 shows the editing profile of T cells using various levels of trans UGI mRNA and various editor mRNAs. The percentage of editing with 24.0 nM BE4Max mRNA is shown. [Figure 21] C-to-T conversion rates (%) as a percentage of edited reads are shown using various levels of UGI mRNA in trans and various editor mRNAs (BC22n at 27.3 nM, BC22-2×UGI at 24.7 nM, BE4Max at 24.0 nM). [Figure 22]For several guides using Cas9 or BC22, the average percentage of T cells negative for cell surface expression of MHC II ("% MHC II negative") is shown in relation to the distance between the cleavage site boundary nucleotides, shown as base pairs ("bp"). Positive numbers indicate the splice site boundary nucleotide 3' to the cleavage site, and negative numbers indicate the splice site boundary nucleotide 5' to the cleavage site. [Figure 23A] An exemplary sgRNA (SEQ ID NO: 141, methylation not shown) in possible secondary structures is shown with labels designating individual nucleotides of the conserved regions of the sgRNA, including the lower stem, bulge, upper stem, nexus (the nucleotides of which may be referred to as N1 through N18, respectively, in the 5' to 3' direction), and hairpin region (including hairpin 1 and hairpin 2 regions). The nucleotides between hairpin 1 and hairpin 2 are labeled n. A guide region may also be present on the sgRNA, and is denoted in this figure as "(N)x" before the conserved region of the sgRNA. [Figure 23B] The ten conserved YA sites in the exemplary sgRNA sequence (SEQ ID NO: 141, methylation not shown) are labeled 1 through 10. The numbers 25, 45, 50, 56, 64, 67, and 83 represent the pyrimidine positions of YA sites 1, 5, 6, 7, 8, 9, and 10 in an sgRNA with a guide region denoted as (N)x (e.g., x is optionally 20). [Figure 24] Figure 1 shows the results of the efficacy of three CIITA guides (G016086, G016092, and G016067) for editing T cells with BC22. A shows the C-to-T conversion rate (%). B shows the percentage of MHC class II-negative T cells (%). [Figure 25] The percentage of B2M-negative T cells after editing with different mRNA combinations is shown. EP stands for electroporation. [Figure 26] The percentage of single nucleotide variants (SNVs) that are C-to-U transitions in the transcriptomes of T cells edited with B2M using different mRNA combinations is shown (ns = not significant). [Figure 27] The percentage of B2M-negative T cells after treatment with different mRNA combinations is shown. [Figure 28] 1 shows the editing profile at the B2M locus in T cells after treatment with different mRNA combinations. [Figure 29] The percentage of SNVs that are C-to-T transitions in amplified genomic DNA from single T cells edited with B2M using different mRNA combinations is shown (ns = not significant). [Figure 30] The percentage of B2M-negative eHap1 cells after editing with different mRNA combinations is shown. [Figure 31] 1 shows the editing profile of the B2M locus in eHap1 cells after treatment with different mRNA combinations. [Figure 32] The percentage of SNVs that are C-to-T transitions in clonally expanded eHap1 cells edited with B2M using different mRNA combinations is shown (ns = not significant). [Figure 33] Figure 1 shows cell viability compared to untreated cells after electroporation or LNP delivery of BC22n or Cas9 editors and single or multiple guides. [Figure 34] Shown is the total γH2AX spot intensity per nucleus after electroporation or LNP delivery of BC22n or Cas9 editor and single or multiple guides. [Figure 35] Shown are the % editing rates at the locus of interest after LNP delivery of the BC22n or Cas9 editor and single or multiple guides. [Figure 36] The percentage of cells negative for the indicated surface proteins after LNP delivery of the BC22n or Cas9 editor and single or multiple guides is shown. [Figure 37] Shown is the percentage of unique interchromosomal translocations across molecules after LNP delivery of BC22n or Cas9 editors and multiple guides. [Figure 38]Average % editing rates in mouse liver after treatment with various base editors are shown. [Figure 39] Average % C-to-T conversion activity for base editor constructs designed with various deaminase domains is shown. [Figure 40A] The percentage of total reads containing at least one C to T conversion is shown for base-edited constructs designed with various linkers. [Figure 40B] The percentage of total reads containing at least one C to T conversion is shown for base-edited constructs designed with various linkers. [Figure 40C] The percentage of total reads containing at least one C to T conversion is shown for base-edited constructs designed with various linkers. [Figure 40D] The percentage of total reads containing at least one C to T conversion is shown for base-edited constructs designed with various linkers. [Figure 40E] The percentage of total reads containing at least one C to T conversion is shown for base-edited constructs designed with various linkers. [Figure 40F] The percentage of total reads containing at least one C to T conversion is shown for base-edited constructs designed with various linkers. [Figure 40G] The percentage of total reads containing at least one C to T conversion is shown for base-edited constructs designed with various linkers. [Figure 40H] The percentage of total reads containing at least one C to T conversion is shown for base-edited constructs designed with various linkers. [Figure 40I] The percentage of total reads containing at least one C to T conversion is shown for base-edited constructs designed with various linkers. [Figure 40J] The percentage of total reads containing at least one C to T conversion is shown for base-edited constructs designed with various linkers. [Figure 40K] The percentage of total reads containing at least one C to T conversion is shown for base-edited constructs designed with various linkers. [Figure 41] Shown is the average % editing rate in SERPINA1 in Huh-7 after treatment with base editor constructs designed with various linkers. [Figure 42] The EC90 values for base-edited mRNAs designed with various linkers are shown. The 95% confidence interval (CI) for each EC90 value is also shown. [Figure 43] Average percent editing at the ANAPC5 locus in PHHs using base-editing constructs designed with various linkers is shown. [Figure 44] The EC95 values for base-edited mRNAs designed with various linkers are shown. The 95% confidence interval (CI) for each EC95 value is also shown. [Figure 45] Average percent editing at the TRAC locus in PHHs using base-editing constructs designed with various linkers is shown. [Figure 46] The masses of base editor mRNAs designed with various linkers that result in 90% of the maximal knockdown of CD3 (EC90) are shown. The 95% confidence interval for each EC50 value is also shown. [Figure 47] The mass of BC22n mRNA designed with various linkers that resulted in 90% of the maximal knockdown (EC90) of CD3, HLA-A3, and HLA-DR, DP, DQ (EC50) is shown. The 95% confidence interval (CI) of each EC50 value is also shown. [Figure 48] The average percent editing at the TRAC locus in T cells treated with sgRNAs in 100-mer or 91-mer formats is shown. [Figure 49A] The average percent editing rate at the TRBC1 locus in T cells treated with sgRNAs in 100-mer or 91-mer formats is shown. [Figure 49B]The average percent editing rate at the TRBC2 locus in T cells treated with sgRNAs in 100-mer or 91-mer formats is shown. [Figure 50] The average percent editing at the CIITA locus in T cells treated with sgRNAs in 100-mer or 91-mer formats is shown. [Figure 51] The average percent editing at the B2M locus in T cells treated with sgRNAs in 100-mer or 91-mer formats is shown. [Figure 52] The average percent editing at the CD38 locus in T cells treated with sgRNAs in 100-mer or 91-mer formats is shown. [Figure 53A] The mean percentage of CD8+ T cells negative for the CD3 surface receptor is shown after treatment with sgRNA targeting TRAC in 100-mer or 91-mer format. [Figure 53B] The mean percentage of CD8+ T cells negative for the CD3 surface receptor is shown after treatment with sgRNAs in 100-mer or 91-mer formats targeting TRBCs. [Figure 54A] The mean percentage of CD8+ T cells negative for HLA-DR, DP, and DQ surface receptors is shown after treatment with sgRNAs targeting CIITA in 100-mer or 91-mer formats. [Figure 54B] The mean percentage of CD8+ T cells negative for HLA-A surface receptors is shown after treatment with sgRNAs targeting HLA-A in 100-mer or 91-mer formats. [Figure 55A] The mean percentage of CD8+ T cells negative for the B2M surface receptor is shown after treatment with sgRNAs targeting B2M in 100-mer or 91-mer formats. [Figure 55B]The mean percentage of CD8+ T cells negative for the CD38 surface receptor is shown after treatment with sgRNAs targeting CD38 in 100-mer or 91-mer formats. [Figure 56] Average percent editing at the ANAPC5 locus in mouse liver with increasing amounts of UGI mRNA is shown. [Figure 57] Shows the % C to T editing purity at the ANAPC5 locus in mouse liver after editing with increasing amounts of UGI mRNA. [Figure 58A] Figure 1 shows total editing at different BC22n-HiBiT mRNA concentrations in B2M in PHH cells. [Figure 58B] 1 shows C-to-T purity at different UGI-HiBiT mRNA concentrations in B2M in PHH cells. [Figure 58C] Figure 1 shows total editing at different BC22-2xUGI-HiBiT mRNA concentrations in B2M in PHH cells. [Figure 58D] Figure 1 shows C-to-T purity at different BC22-2xUGI-HiBiT mRNA concentrations in B2M in PHH cells. [Figure 59A] Figure 1 shows total editing at different BC22n-HiBiT mRNA concentrations in T cells. [Figure 59B] C-to-T purity at different UGI-HiBiT mRNA concentrations in T cells is shown. [Figure 59C] Figure 1 shows total editing at different BC22-2xUGI-HiBiT mRNA concentrations in T cells. [Figure 59D] C-to-T purity at different BC22-2×UGI-HiBiT mRNA concentrations in T cells is shown. [Figure 60] Shows editing in liver tissue taken from CD-1 mice treated with LNPs harboring a fixed dose of sgRNA and base editor mRNA and different doses of UGI mRNA. [Figure 61]C-to-T purity in liver tissue collected from CD-1 mice treated with LNPs carrying a fixed dose of sgRNA and base editor mRNA and different doses of UGI mRNA. [Figure 62] Shown is the lysis rate (%) of T cells targeted by NK cells at different effector:target (E:T) ratios treated with sgRNA and base editor and UGI mRNA. [Figure 63] The conversion rate of each guide nucleotide position of the highly active guide edited by the Spy base editor is shown. [Figure 64] The conversion rate of each guide nucleotide position of the highly active guide edited by the Nme2 base editor is shown. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8] [Table 1-9] [Table 1-10]
[0014] For the sequences themselves, see the sequence listing below. The transcribed sequences generally may contain GGG as the first three nucleotides for use with ARCA or AGG as the first three nucleotides for use with CleanCap™. Thus, the first three nucleotides can be modified for use with other capping approaches, such as Vaccinia capping enzyme. Promoters and polyA sequences are not included in the transcribed sequences. Promoters such as the U6 promoter (SEQ ID NO: 67) or CMV promoter (SEQ ID NO: 68) and polyA sequences such as SEQ ID NO: 109 can be added to the disclosed transcribed sequences at the 5' and 3' ends, respectively. Most nucleotide sequences are provided as DNA, but can be easily converted to RNA by changing T to U. DETAILED DESCRIPTION OF THE INVENTION
[0015] Reference will now be made in detail to certain embodiments of the invention, examples of which are illustrated in the accompanying drawings. While the invention will be described in conjunction with the illustrated embodiments, it will be understood that they are not intended to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents, which may be included within the scope of the invention as defined by the appended claims.
[0016] Before describing the present teachings in detail, it is to be understood that the present disclosure is not limited to particular compositions or process steps, as such may vary. It should be noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly indicates otherwise. Thus, for example, a reference to "a conjugate" includes a plurality of conjugates, a reference to "a cell" includes a plurality of cells, and so forth.
[0017] Numerical ranges are inclusive of the numbers defining the range. It is understood that measured and measurable values are approximations taking into account significant digits and error associated with the measurements. Also, the use of "comprise," "comprises," "comprising," "contain," "contains," "containing," "include," "includes," and "including" is not intended to be limiting. It is understood that both the foregoing general description and the detailed description are exemplary and explanatory only and are not limiting of the present teachings.
[0018] The term "about" or "approximately" refers to an acceptable error for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined and on the degree of variation that does not materially affect the properties of the described subject matter (i.e., within the acceptable error tolerance accepted in the art, e.g., 10%, 5%, 2%, or 1%). Accordingly, unless otherwise indicated, the numerical parameters set forth in the following specification and appended claims are approximations that may vary depending upon the desired properties sought to be obtained. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.
[0019] Unless otherwise stated in the specification above, embodiments herein that recite various components as "comprising" are also assumed to "consist of" or "consist essentially of" the recited components; embodiments herein that recite various components as "consisting of" are also assumed to "comprise" or "consist essentially of" the recited components; and embodiments herein that recite various components as "consisting essentially of" are also assumed to "consist of" or "comprise" the recited components (this interchangeability does not apply to the use of these terms in the claims).
[0020] The section headings used herein are for organizational purposes only and should not be construed as limiting the desired subject matter in any way. In the event that a document incorporated by reference contradicts the express contents of this specification, including but not limited to definitions, the express contents of this specification shall control. While the present teachings have been described in conjunction with various embodiments, it is not intended that the present teachings be limited to such embodiments. Rather, the present teachings encompass various alternatives, modifications, and equivalents, as will be appreciated by those skilled in the art.
[0021] I. Definition Unless otherwise stated, the following terms and phrases used herein are intended to have the following meanings:
[0022] As used herein, the term "or combinations thereof" refers to all permutations and combinations of the terms listed before it. For example, "A, B, C, or combinations thereof" is intended to include at least one of A, B, C, AB, AC, BC, or ABC, and, where order is important in a particular context, also BA, CA, CB, ACB, CBA, BCA, BAC, or CAB. Following this example, combinations containing one or more repeats of an item or term, such as BB, AAA, AAB, BBC, CBBA, CABA, etc., are expressly included. Those skilled in the art will understand that there is typically no limit to the number of items or terms in any combination unless otherwise clear from the context.
[0023] As used herein, the term "kit" refers to a packaged set of one or more polynucleotides or compositions and one or more related materials, e.g., related components such as a delivery device (e.g., a syringe), solvent, solution, buffer, instructions, or desiccant.
[0024] "Or" is used in its inclusive sense, ie, equivalent to "and / or," unless the context requires otherwise.
[0025] "Polynucleotide" and "nucleic acid" are used herein to refer to polymeric compounds containing nucleosides or nucleoside analogs having nitrogen-containing heterocyclic bases or base analogs linked together along a backbone, and include polymers of conventional RNA, DNA, mixed RNA-DNA, and analogs thereof. The nucleic acid "backbone" can be composed of a variety of linkages, including one or more of sugar-phosphodiester linkages, peptide-nucleic acid linkages ("peptide nucleic acids" or PNA, PCT Publication No. WO 95 / 32305), phosphorothioate linkages, methylphosphonate linkages, or combinations thereof. The sugar moiety of the nucleic acid can be ribose, deoxyribose, or similar compounds with substitutions, such as 2' methoxy or 2' halide substitutions. Nitrogenous bases include the common bases (A, G, C, T, U), their analogs (e.g., modified uridines, such as 5-methoxyuridine, pseudouridine, or N1-methylpseudouridine); inosine; derivatives of purines or pyrimidines (e.g., N 4 -methyldeoxyguanosine, deaza or azapurines, deaza or azapyrimidines, pyrimidine bases with substituents at the 5 or 6 positions (e.g., 5-methylcytosine), purine bases with substituents at the 2, 6, or 8 positions, 2-amino-6-methylaminopurine, O 6 -methylguanine, 4-thio-pyrimidine, 4-amino-pyrimidine, 4-dimethylhydrazine-pyrimidine, and O 4 -alkyl-pyrimidines; U.S. Pat. No. 5,378,825 and PCT Publication No. WO 93 / 13121). For a general discussion, see The Biochemistry of the Nucleic Acids, vol. 5-36, Adams et al., ed., 11 thed., 1992. Nucleic acids can contain one or more "abasic" residues, where the backbone does not contain a nitrogenous base at one or more positions in the polymer (U.S. Patent No. 5,585,481). Nucleic acids can contain only conventional RNA or DNA sugars, bases, and linkages, or they can contain both conventional components and substitutions (e.g., polymers containing conventional bases with 2' methoxy linkages, or conventional bases and one or more base analogs). Nucleic acids include "locked nucleic acids" (LNAs), which are analogs containing one or more LNA nucleotide monomers with a bicyclic furanose unit locked to an RNA-mimetic sugar structure, enhancing hybridization affinity for complementary RNA and DNA sequences (Vester and Wengel, 2004, Biochemistry 43(42):13233-41). RNA and DNA differ in their sugar moieties, which can differ by the presence of uracil or its analogs in RNA and thymine or its analogs in DNA.
[0026] As used herein, "polypeptide" refers to a multimeric compound comprising amino acid residues capable of adopting a three-dimensional conformation. Polypeptides include, but are not limited to, enzymes, proenzyme proteins, regulatory proteins, structural proteins, receptors, nucleic acid binding proteins, antibodies, and the like. Polypeptides may, but do not necessarily, include post-translational modifications, unnatural amino acids, artificial groups, and the like.
[0027] As used herein, "cytidine deaminase" means a polypeptide or complex of polypeptides capable of cytidine deaminase activity that catalyzes the hydrolytic deamination of cytidine or deoxycytidine, typically to yield uridine or deoxyuridine. Cytidine deaminases include enzymes of the cytidine deaminase superfamily, particularly enzymes of the APOBEC family (APOBEC1, APOBEC2, APOBEC4, and APOBEC3 subgroups of enzymes), activation-induced cytidine deaminase (AID or AICDA) and CMP deaminase (see, e.g., Conticello et al., Mol. Biol. Evol. 22:367-77, 2005; Conticello, Genome Biol. 9:229, 2008; Muramatsu et al., J. Biol. Chem. 274:18470-6, 1999; Carrington et al., Cells 9:1690 (2020)). In some embodiments, variants of any known cytidine deaminase or APOBEC protein are included. Variants include proteins having a sequence that differs from that of the wild-type protein by one or several mutations (i.e., substitutions, deletions, insertions), e.g., one or several single-point substitutions. For example, truncated sequences can be used, e.g., by deleting 1 to 4 amino acids at the N-terminus, C-terminus, or internal amino acids of the sequence, preferably the C-terminus. As used herein, the term "variant" refers to allelic variants, splicing variants, and natural or artificial mutants that are homologous to a reference sequence. Variants are "functional" in that they exhibit catalytic activity for DNA editing.
[0028] As used herein, the term "APOBEC3A" refers to a cytidine deaminase, e.g., a protein expressed by the human A3A gene. APOBEC3A can have catalytic DNA editing activity. The amino acid sequence of APOBEC3A has been described (UniPROT Accession ID: p31941) and is included herein as SEQ ID NO: 40. In some embodiments, the APOBEC3A protein is a human APOBEC3A protein and / or a wild-type protein. Variants include proteins having a sequence that differs from the wild-type APOBEC3A protein by one or several mutations (i.e., substitutions, deletions, insertions), e.g., one or several single-point substitutions. For example, truncated APOBEC3A sequences can be used, e.g., by deleting 1 to 4 amino acids at the N-terminus, C-terminus, or internal amino acids of the sequence, preferably the C-terminus. As used herein, the term "variant" refers to allelic variants, splicing variants, and natural or artificial mutants homologous to the APOBEC3A reference sequence. The variants are "functional" in that they exhibit catalytic activity for DNA editing. In some embodiments, the APOBEC3A (such as human APOBEC3A) has a wild-type amino acid at position 57 (as numbered in the wild-type sequence). In some embodiments, the APOBEC3A (such as human APOBEC3A) has an asparagine at amino acid position 57 (as numbered in the wild-type sequence).
[0029] As used herein, a "nickase" is an enzyme that generates a single-strand break (also known as a "nick") in double-stranded DNA, i.e., cleaves one strand of the DNA double helix but not the other. As used herein, an "RNA-guided nickase" refers to a polypeptide or complex of polypeptides with DNA nickase activity, where the DNA nickase activity is sequence-specific and dependent on the sequence of the RNA. Exemplary RNA-guided nickases include Cas nickases. Cas nickases include, but are not limited to, the Csm or Cmr complex of a type III CRISPR system and its Cas10, Csm1, or Cmr2 subunit, the Cascade complex of a type I CRISPR system and its Cas3 subunit, and nickase forms of class 2 Cas nucleases. Class 2 Cas nickases include class 2 Cas nuclease variants in which only one of the two catalytic domains is inactivated and which have RNA-guided DNA nickase activity. Class 2 Cas nickases include, for example, Cas9 (e.g., H840A, D10A, or N863A variants of SpyCas9), Cpf1, C2c1, C2c2, C2c3, HF Cas9 (e.g., N497A, R661A, Q695A, Q926A variants), HypaCas9 (e.g., N692A, M694A, Q695A, H698A variants), eSPCas9(1.0) (e.g., K810A, K1003A, R1060A variants), and eSPCas9(1.1) (e.g., K848A, K1003A, R1060A variants) proteins and modified forms thereof. The Cpf1 protein (Zetsche et al., Cell, 163:1-13 (2015)) is homologous to Cas9 and contains a RuvC-like protein domain. The Cpf1 sequence of Zetsche is incorporated by reference in its entirety. See, e.g., Tables S1 and S3 of Zetsche. "Cas9" encompasses S. pyogenes (Spy) Cas9, variants of Cas9 listed herein, and their equivalents.See, for example, Makarova et al., Nat Rev Microbiol, 13(11):722-36 (2015); Shmakov et al., Molecular Cell, 60:385-397 (2015).
[0030] Several Cas9 orthologs have been isolated from N. meningitidis (Esvelt et al., NAT. METHODS, Vol. 10, 2013, pp. 1116-1121; Hou et al., PNAS, Vol. 110, 2013, pp. 15644-15649; Edraki et al., Mol. Cell 73:714-726, 2019) (Nme1Cas9, Nme2Cas9, and Nme3Cas9). The Nme2Cas9 ortholog functions efficiently in mammalian cells, recognizes the N4CC PAM, and can be used for in vivo editing (Ran et al., NATURE, Vol. 520, 2015, pp. 186-191; Kim et al., NAT. COMMUN., Vol. 8, 2017, pp. 14500). Nme2Cas9 has been shown to be naturally resistant to off-target editing (Lee et al., MOL. THER., vol. 24, 2016, pp. 645-654; Kim et al., 2017). See also, e.g., WO / 2020081568 (e.g., pp. 28 and 42), the contents of which are incorporated herein by reference in their entirety, describing the Nme2Cas9 D16A nickase. Throughout, "NmeCas9" is a generic term and encompasses any type of NmeCas9, including Nme1Cas9, Nme2Cas9, and Nme3Cas9.
[0031] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising polypeptides from at least two different proteins or sources. One polypeptide can be located at the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) portion of the fusion protein, thus forming an "amino-terminal fusion protein" or a "carboxy-terminal fusion protein," respectively. Any of the proteins provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced via recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Methods for recombinant protein expression and purification are well known and include those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)), the entire contents of which are incorporated herein by reference.
[0032] As used herein, the term "linker" refers to a chemical group or molecule that connects two adjacent molecules or moieties. Typically, a linker is positioned between or sandwiched by two groups, molecules, or other moieties, connecting them to each other via a covalent bond. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein), such as a 16-amino acid residue "XTEN" linker, or variants thereof (see, e.g., the Examples; and Schellenberger et al., "A recombinant polypeptide extends the in vivo half-life of peptides and proteins in a tunable manner." Nat. Biotechnol. 27, 1186-1190 (2009)). In some embodiments, the XTEN linker comprises the sequence SGSETPGTSESATPES (SEQ ID NO: 46), SGSETPGTSESA (SEQ ID NO: 47), or SGSETPGTSESATPEGGSGGS (SEQ ID NO: 48). In some embodiments, the linker comprises one or more sequences selected from SEQ ID NOs: 46-59, 61, and 211-272.
[0033] As used herein, the terms "uracil glycosylase inhibitor," "uracil-DNA glycosylase inhibitor," or "UGI" refer to proteins that can inhibit the uracil-DNA glycosylase (UDG) base excision repair enzyme (e.g., UniPROT ID: P14739; SEQ ID NO: 27; SEQ ID NO: 43).
[0034] As used herein, the "open reading frame" or "ORF" of a gene refers to a sequence of codons that specifies the amino acid sequence of the protein encoded by the gene. An ORF generally begins with a start codon (e.g., ATG in DNA or AUG in RNA) and ends with a stop codon (e.g., TAA, TAG, or TGA in DNA or UAA, UAG, or UGA in RNA).
[0035] "Guide RNA," "gRNA," and "guide" are used interchangeably herein to refer to either crRNA (also known as CRISPR RNA) or the combination of crRNA and trRNA (also known as tracrRNA). The crRNA and trRNA can associate as a single RNA molecule (single guide RNA, sgRNA) or as two separate RNA molecules (dual guide RNA, dgRNA). "Guide RNA" or "gRNA" refers to each type. The trRNA may be a naturally occurring sequence or a trRNA sequence with modifications or mutations compared to the naturally occurring sequence.
[0036] As used herein, "guide sequence" or "guide region" or "spacer" or "spacer sequence" refers to a sequence within a gRNA that is complementary to a target sequence and functions to direct the gRNA to the target sequence for binding or modification (e.g., cleavage) by an RNA-guided nickase. A guide sequence can be 20 nucleotides in length, for example, in the case of Streptococcus pyogenes (i.e., Spy Cas9 (also known as SpCas9)) and related Cas9 homologs / orthologs. Shorter or longer sequences, e.g., 15, 16, 17, 18, 19, 21, 22, 23, 24, or 25 nucleotides in length, can also be used as guides. A guide sequence can be 20-25 nucleotides in length, e.g., 20, 21, 22, 23, 24, or 25 nucleotides in length, for example, in the case of Nme Cas9. For example, a 24-nucleotide guide sequence can be used with Nme Cas9, e.g., Nme2 Cas9.
[0037] In some embodiments, the target sequence is intragenic or chromosomal, e.g., complementary to the guide sequence. In some embodiments, the degree of complementarity or identity between the guide sequence and its corresponding target sequence can be about 75%, 80%, 85%, 90%, 95%, or 100%. In some embodiments, the guide sequence and target region can be 100% complementary or identical. In other embodiments, the guide sequence and target region can contain at least one mismatch. For example, the guide sequence and target sequence can contain 1, 2, 3, or 4 mismatches, in which case the total length of the target sequence is at least 17, 18, 19, 20, or more base pairs. In some embodiments, the guide sequence and target region can contain 1 to 4 mismatches, in which case the guide sequence contains at least 17, 18, 19, 20, or more nucleotides. In some embodiments, the guide sequence and target region can contain 1, 2, 3, or 4 mismatches, in which case the guide sequence contains 20 nucleotides.
[0038] As used herein, "target sequence" or "genomic target sequence" refers to a sequence of nucleic acid within a target gene that has complementarity to the guide sequence of a gRNA. The interaction of the target sequence and guide sequence directs an RNA-guided DNA-binding agent to bind and potentially nick or cleave (depending on the activity of the agent) within the target sequence. The target sequence of a Cas protein includes both the plus and minus strands of genomic DNA (i.e., a given sequence and its reverse complement) because the nucleic acid substrate of a Cas protein is a double-stranded nucleic acid. Thus, when a guide sequence is said to be "complementary to a target sequence," it is understood that the guide sequence can direct an RNA-guided DNA-binding agent (e.g., dCas9 or impaired Cas9) to bind to the reverse complement of the target sequence. Thus, in some embodiments in which the guide sequence binds to the reverse complement of the target sequence, the guide sequence is identical to certain nucleotides of the target sequence (e.g., a target sequence without a PAM), except that T is replaced with U in the guide sequence.
[0039] As used herein, a first sequence is considered to "contain a sequence having at least X% identity" to a second sequence if alignment of the first and second sequences shows that X% or more positions throughout the second sequence match the first sequence. For example, the sequence AAGA contains a sequence having 100% identity to the sequence AAG, because alignment would result in 100% identity in that all three positions in the second sequence match. Differences between RNA and DNA (generally, the exchange of uridine with thymidine or vice versa) and the presence of nucleoside analogs such as modified uridines do not contribute to differences in identity or complementarity between polynucleotides, as long as the related nucleotide (e.g., thymidine, uridine, or modified uridine) has the same complement (e.g., adenosine for all thymidine, uridine, or modified uridine; another example is cytosine and 5-methylcytosine, both of which have guanosine as their complement). Thus, for example, the sequence 5'-AXG (where X is any modified uridine, e.g., pseudouridine, N1-methylpseudouridine, or 5-methoxyuridine) is considered 100% identical to AUG, in that both are perfectly complementary to the same sequence (5'-CAU). Exemplary alignment algorithms are the Smith-Waterman and Needleman-Wunsch algorithms, which are well known in the art. Those skilled in the art will understand which algorithm selection and parameter settings are appropriate for a given pair of sequences to be aligned. For sequences of generally similar length and predicted greater than 50% amino acid identity and greater than 75% nucleotide identity, the Needleman-Wunsch algorithm, with default settings in the Needleman-Wunsch algorithm interface provided by EBI on its web server at www.ebi.ac.uk, is generally appropriate.
[0040] "MRNA" is used herein to refer to a polynucleotide, other than DNA, that contains an open reading frame that can be translated into a polypeptide (i.e., can serve as a substrate for translation by ribosomes and aminoacylated tRNAs). mRNA can include a phosphate-sugar backbone that includes ribose residues or analogs thereof, such as 2'-methoxyribose residues. In some embodiments, the sugars of the phosphate-sugar backbone of an mRNA consist essentially of ribose residues, 2'-methoxyribose residues, or combinations thereof. Generally, mRNA does not contain a substantial amount of thymidine residues (e.g., 0 or less than 30, 20, 10, 5, 4, 3, or 2 thymidine residues; or a thymidine content of less than 10%, 9%, 8%, 7%, 6%, 5%, 4%, 4%, 3%, 2%, 1%, 0.5%, 0.2%, or 0.1%). mRNA can include modified uridines at some or all of its uridine positions.
[0041] "Modified uridine" is used herein to refer to a nucleoside other than thymidine that has the same hydrogen bond acceptor as uridine and has one or more structural differences from uridine. In some embodiments, the modified uridine is a substituted uridine, i.e., a uridine in which one or more aprotic substituents (e.g., alkoxy, e.g., methoxy) replace the protons. In some embodiments, the modified uridine is a substituted pseudouridine. In some embodiments, the modified uridine is a substituted pseudouridine, i.e., a pseudouridine in which one or more aprotic substituents (e.g., alkyl, e.g., methyl) replace the protons. In some embodiments, the modified uridine is either a substituted uridine, a pseudouridine, or a substituted pseudouridine.
[0042] As used herein, a "uridine position" refers to a position in a polynucleotide occupied by a uridine or modified uridine. Thus, for example, a polynucleotide in which "100% of the uridine positions are modified uridines" contains a modified uridine at every position that would be a uridine in normal RNA of the same sequence (where all bases are standard A, U, C, or G bases). Unless otherwise indicated, the U in the polynucleotide sequences in this disclosure or in the sequence listing or sequence listings attached to this disclosure can be a uridine or a modified uridine.
[0043] As used herein, the "minimal uridine codon(s)" of a given amino acid are the codon(s) with the fewest uridines (usually 0 or 1, except for the codon for phenylalanine, which has two uridines as its minimal uridine codon). In assessing uridine content, modified uridine residues are considered equivalent to uridine.
[0044] As used herein, the "uridine dinucleotide (UU) content" of an ORF can be expressed in absolute terms as the number of UU dinucleotides in the ORF, or on a percentage basis as the percentage of positions occupied by uridine dinucleotides (e.g., AUUAU has a 40% uridine dinucleotide content because 2 of the 5 positions are occupied by uridine dinucleotides). In assessing uridine dinucleotide content, modified uridine residues are considered equivalent to uridine.
[0045] As used herein, the "minimal adenine codon(s)" of a given amino acid are the codon(s) with the fewest adenines (usually 0 or 1, except for the codons for lysine and asparagine, which have two adenines as their minimal adenine codon). In assessing adenine content, modified adenine residues are considered equivalent to adenine.
[0046] As used herein, the "adenine dinucleotide content" of an ORF can be expressed in absolute terms as the number of AA dinucleotides in the ORF, or on a percentage basis as the percentage of positions occupied by adenine dinucleotides (e.g., UAAUA has an adenine dinucleotide content of 40% because 2 of 5 positions are occupied by adenine dinucleotides). In assessing adenine dinucleotide content, modified adenine residues are considered equivalent to adenine.
[0047] As used herein, "indel" refers to an insertion / deletion mutation consisting of several nucleotides inserted or deleted, for example, at a double-strand break (DSB) site in a target nucleic acid.
[0048] As used herein, "knockdown" refers to a reduction in the expression of a particular gene product (e.g., protein, mRNA, or both). Protein knockdown can be measured by detecting the protein secreted by a tissue or population of cells (e.g., in serum or cell culture medium), or by detecting the total cellular amount of protein from the tissue or cell population of interest. Methods for measuring mRNA knockdown are known and include sequencing mRNA isolated from the tissue or cell population of interest. In some embodiments, "knockdown" can refer to some loss of expression of a particular gene product, for example, a reduction in the amount of mRNA transcribed, or a reduction in the amount of protein expressed or secreted by a population of cells (including in vivo populations such as those found in tissues).
[0049] As used herein, "knockout" refers to the loss of expression of a particular protein in a cell. Knockout can be measured by detecting the amount of protein secreted from a tissue or population of cells (e.g., in serum or cell culture medium), or by detecting the total cellular amount of the protein in a tissue or population of cells. In some embodiments, the methods of the present disclosure "knock out" a target protein in one or more cells (e.g., in a population of cells, including in vivo populations such as those found in tissues). In some embodiments, the knockout is not the formation of a mutant of the target protein, e.g., created by an indel, but rather a complete loss of expression of the target protein in the cell (i.e., a reduction in expression below the detection level of the assay used).
[0050] As used herein, the term "nuclear localization signal" (NLS) or "nuclear localization sequence" refers to an amino acid sequence that directs the transport of a molecule containing or linked to such a sequence into the nucleus of a eukaryotic cell. The nuclear localization signal may form part of the molecule to be transported. In some embodiments, the NLS may be fused to the molecule by a covalent bond, hydrogen bond, or ionic interaction. In some embodiments, the NLS may be fused to the molecule via a linker.
[0051] As used herein, "β2M" or "B2M" refers to the nucleic acid or protein sequence of "β-2 microglobulin," the human gene having reference GRCh38.p13, accession number NC_000015 (range 44711492..44718877). The B2M protein associates with MHC class I molecules as a heterodimer on the surface of nucleated cells and is required for expression of MHC class I proteins.
[0052] As used herein, "CIITA" or "CIITA" or "C2TA" refers to the nucleic acid or protein sequence of the "class II major histocompatibility complex transactivator," the human gene having the accession number NC_000016.10 (range 10866208..10941562) with reference to GRCh38.p13. The nuclear CIITA protein acts as a positive regulator of MHC class II gene transcription and is required for the expression of MHC class II proteins.
[0053] As used herein, "MHC" or "MHC molecule(s)" or "MHC protein" or "MHC complex(es)" refers to major histocompatibility complex molecule(s), including, for example, MHC class I and MHC class II molecules. In humans, MHC molecules are referred to as "human leukocyte antigen" complexes or "HLA molecules" or "HLA proteins." The use of the terms "MHC" and "HLA" is not meant to be limiting, and as used herein, the term "MHC" can be used to refer to human MHC molecules, i.e., HLA molecules. Thus, the terms "MHC" and "HLA" are used interchangeably herein.
[0054] The term "HLA-A," as used herein in the context of an HLA-A protein, refers to an MHC class I protein molecule that is a heterodimer consisting of a heavy chain (encoded by the HLA-A gene) and a light chain (i.e., beta-2 microglobulin). The term "HLA-A" or "HLA-A gene," as used herein in the context of a nucleic acid, refers to the gene that encodes the heavy chain of the HLA-A protein molecule. The HLA-A gene is also referred to as "HLA class I histocompatibility, A alpha chain," and the human gene has the accession number NC_000006.12 (29942532..29945870). The HLA-A gene is known to exist in thousands of different versions (also referred to as "alleles") across the population (and an individual can receive two different alleles of the HLA-A gene). A public database of HLA-A alleles, including sequence information, can be accessed at IPD-IMGT / HLA: https: / / www.ebi.ac.uk / ipd / imgt / hla / . All alleles of HLA-A are encompassed by the terms "HLA-A" and "HLA-A gene."
[0055] As used herein in the context of nucleic acids, "HLA-B" refers to the gene encoding the heavy chain of the HLA-B protein molecule. HLA-B is also referred to as "HLA class I histocompatibility, B alpha chain," and the human gene has the accession number NC_000006.12 (31353875..31357179).
[0056] The term "HLA-C" as used herein in the context of nucleic acids refers to the gene encoding the heavy chain of the HLA-C protein molecule. HLA-C is also referred to as "HLA class I histocompatibility, C alpha chain," and the human gene has the accession number NC_000006.12 (31268749..31272092).
[0057] As used herein in the context of nucleic acids, "TRBC1" and "TRBC2" refer to two homologous genes encoding the T cell receptor β chain. "TRBC" or "TRBC1 / 2" are used herein to refer to TRBC1 and TRBC2. The human wild-type TRBC1 sequence is available at NCBI Gene ID: 28639; Ensembl: ENSG00000211751. T cell receptor beta constant region, V_segment translation product, BV05S1J2.2, TCRBC1, and TCRB are gene synonyms for TRBC1. The human wild-type TRBC2 sequence is available at NCBI Gene ID: 28638; Ensembl: ENSG00000211772. T cell receptor beta constant region, V_segment translation product, and TCRBC2 are gene synonyms for TRBC2.
[0058] As used herein in the context of nucleic acids, "TRAC" refers to the gene encoding the T cell receptor alpha chain. The human wild-type TRAC sequence is available at NCBI Gene ID: 28755; Ensembl: ENSG00000277734. T cell receptor alpha constant region, TCRA, IMD7, TRCA, and TRA are gene synonyms of TRAC.
[0059] As used herein, the term "homozygous" refers to having two identical alleles of a particular gene.
[0060] As used herein, "treatment" refers to any administration or application of a therapeutic agent to a disease or disorder in a subject, including inhibiting the disease, preventing its onset, alleviating one or more symptoms of the disease, curing the disease, or preventing one or more symptoms of the disease (including the recurrence of symptoms).
[0061] As used herein, "delivery" and "administration" are used interchangeably and include ex vivo and in vivo applications.
[0062] As used herein, "co-administration" means that two or more agents are administered sufficiently close in time that the agents act together. Co-administration includes administering the agents together in a single formulation, and administering the agents in separate formulations sufficiently close in time that the agents act together.
[0063] As used herein, the phrase "pharmaceutically acceptable" means useful for preparing pharmaceutical compositions that are generally non-toxic, not biologically undesirable, and not otherwise unacceptable for pharmaceutical use. Pharmaceutically acceptable generally refers to a substance that is non-pyrogenic. Pharmaceutically acceptable can refer to a substance that is sterile, particularly a pharmaceutical substance for injection or infusion.
[0064] As used herein, "subject" refers to any member of the animal kingdom. In some embodiments, "subject" refers to a human. In some embodiments, "subject" refers to a non-human animal. In some embodiments, "subject" refers to a primate. In some embodiments, subjects include, but are not limited to, mammals, birds, reptiles, amphibians, fish, insects, and / or worms. In certain embodiments, the non-human subject is a mammal (e.g., a rodent, mouse, rat, rabbit, monkey, dog, cat, sheep, cow, primate, and / or pig). In some embodiments, the subject may be a transgenic animal, a genetically engineered animal, and / or a clone. In certain embodiments of the invention, the subject is an adult, an adolescent, or an infant. In some embodiments, the terms "individual" or "patient" are used and are intended to be interchangeable with "subject."
[0065] As used herein, "reduced or eliminated" expression of a protein on a cell refers to a partial or complete loss of expression of the protein compared to unmodified cells. In some embodiments, surface expression of a protein on a cell is measured by flow cytometry, and the cell has "reduced or eliminated" surface expression compared to unmodified cells, as evidenced by a reduced fluorescent signal when stained with the same antibody against the protein. Cells that have "reduced or eliminated" surface expression of a protein by flow cytometry compared to unmodified cells may be referred to as "negative" for the expression of that protein, as evidenced by a fluorescent signal similar to that of cells stained with an isotype control antibody. "Reduced or eliminated" protein expression may also be measured by other techniques known in the art, using appropriate controls known to those skilled in the art.
[0066] II. Exemplary Compositions and Methods In some embodiments, a nucleic acid is provided that includes an open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase, the polypeptide not including a uracil glycosylase inhibitor (UGI). In some embodiments, the nucleic acid is DNA or RNA. In some embodiments, the nucleic acid is mRNA. In some embodiments, a polypeptide encoded by an mRNA is provided.
[0067] In some embodiments, a polypeptide or mRNA encoding the polypeptide is provided, the polypeptide comprising a cytidine deaminase and an RNA-guided nickase, the polypeptide not comprising a UGI. In some embodiments, the cytidine deaminase is A3A. In some embodiments, the RNA-guided nickase does not comprise a uracil glycosylase inhibitor (UGI). In some embodiments, a composition is provided comprising a first polypeptide or mRNA encoding the first polypeptide comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase, and a second polypeptide or mRNA encoding the second polypeptide that comprises a uracil glycosylase inhibitor (UGI), the second polypeptide being different from the first polypeptide.
[0068] In some embodiments, compositions are provided that include a first nucleic acid comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase, and a second nucleic acid, different from the first nucleic acid, comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI). In some embodiments, the first nucleic acid encodes a polypeptide that does not comprise a UGI.
[0069] In some embodiments, methods are provided for modifying a target gene, comprising administering a composition described herein. In some embodiments, the methods comprise delivering to a cell a first nucleic acid comprising a first open reading frame encoding a first polypeptide comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase, and a second nucleic acid, different from the first nucleic acid, comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI).
[0070] In some embodiments, the method includes delivering to the cell a polypeptide comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase, or a nucleic acid encoding the polypeptide, and separately (e.g., not via the same nucleic acid construct) delivering to the cell a uracil glycosylase inhibitor (UGI), or a nucleic acid encoding a UGI.
[0071] In some embodiments, the molar ratio of mRNA encoding UGI to mRNA encoding cytidine deaminase (e.g., A3A) and RNA-guided nickase is about 1:35 to about 30:1. In some embodiments, the molar ratio is about 1:25 to about 25:1. In some embodiments, the molar ratio is about 1:20 to about 25:1. In some embodiments, the molar ratio is about 1:10 to about 22:1. In some embodiments, the molar ratio is about 1:5 to about 25:1. In some embodiments, the molar ratio is about 1:1 to about 30:1. In some embodiments, the molar ratio is about 2:1 to about 10:1. In some embodiments, the molar ratio is about 5:1 to about 20:1. In some embodiments, the molar ratio is about 1:1 to about 25:1. In some embodiments, the molar ratio is about 1:35, 1:34, 1:33, 1:32, 1:31, 1:30, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:9, 1:8, 1:19, 1:20, 1:21, 1:22, 1:23, 1:24, 1:25, 1:26, 1:27, 1:28, 1:29, 1:30, 1:32, 1:31, 1:30, 1:29, 1:28, 1:27, 1:26, 1:25, 1:24, 1:23, 1:22, 1:21, 1:20, 1:19, 1:18, 1:17, 1:16, 1:15, 1:14, 1:13, 1:12, 1:11, 1:10, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:21, 1:22, 1:23, 1:24, 1:25, 1:26, 1:27, 1:28, 1:29, 1:30, 1:31, 1:32 The molar ratio may be 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, or 30:1. In some embodiments, the molar ratio is equal to or greater than about 1:1. In some embodiments, the molar ratio is about 1:1. In some embodiments, the molar ratio is about 2:1. In some embodiments, the molar ratio is about 3:1. In some embodiments, the molar ratio is about 4:1. In some embodiments, the molar ratio is about 5:1. In some embodiments, the molar ratio is about 6:1. In some embodiments, the molar ratio is about 7:1. In some embodiments, the molar ratio is about 8:1. In some embodiments, the molar ratio is about 9:1. In some embodiments, the molar ratio is about 10:1. In some embodiments, the molar ratio is about 11:1.In some embodiments, the molar ratio is about 12:1. In some embodiments, the molar ratio is about 13:1. In some embodiments, the molar ratio is about 14:1. In some embodiments, the molar ratio is about 15:1. In some embodiments, the molar ratio is about 16:1. In some embodiments, the molar ratio is about 17:1. In some embodiments, the molar ratio is about 18:1. In some embodiments, the molar ratio is about 19:1. In some embodiments, the molar ratio is about 20:1. In some embodiments, the molar ratio is about 21:1. In some embodiments, the molar ratio is about 22:1. In some embodiments, the molar ratio is about 23:1. In some embodiments, the molar ratio is about 24:1. In some embodiments, the molar ratio is about 25:1.
[0072] Similarly, in some embodiments, the molar ratios described above for the mRNA encoding a UGI protein and the mRNA encoding a cytidine deaminase (e.g., A3A) and an RNA-guided nickase are similar when delivering proteins.
[0073] For example, in some embodiments, the molar ratio of the delivered UGI protein to the delivered cytidine deaminase (e.g., A3A) and RNA-guided nickase is about 1:35 to about 30:1, and in some embodiments, the molar ratio is about 1:1 to about 30:1.
[0074] In some embodiments, the molar ratio of the UGI peptide to the cytidine deaminase (eg, A3A) and RNA-guided nickase is about 10:1 to about 50:1. In some embodiments, the molar ratio can be about 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, or 50:1. In some embodiments, the molar ratio is from about 10:1 to about 40:1. In some embodiments, the molar ratio is about 10:1 to about 30:1. In some embodiments, the molar ratio is about 2:1. In some embodiments, the molar ratio is about 10:1 to about 20:1. In some embodiments, the molar ratio is about 10:1 to about 15:1. In some embodiments, the molar ratio is about 15:1 to about 50:1. In some embodiments, the molar ratio is about 6:1. In some embodiments, the molar ratio is about 20:1 to about 50:1. In some embodiments, the molar ratio is about 8:1. In some embodiments, the molar ratio is about 30:1 to about 50:1. In some embodiments, the molar ratio is about 30:1 to about 40:1. In some embodiments, the molar ratio is about 11:1. In some embodiments, the molar ratio is about 20:1 to about 30:1.
[0075] In some embodiments, the compositions described herein further comprise at least one gRNA. In some embodiments, compositions are provided comprising an mRNA described herein and at least one gRNA. In some embodiments, the gRNA is a single guide RNA (sgRNA). In some embodiments, the gRNA is a dual guide RNA (dgRNA).
[0076] In some embodiments, the composition is capable of effectively effecting genome editing upon administration to a subject.
[0077] A.UGI Without being bound by any theory, providing a UGI in conjunction with a polypeptide comprising a deaminase may be useful in the methods described herein by recognizing uracil in DNA as a form of DNA damage and inhibiting cellular DNA repair mechanisms (e.g., UDG and downstream repair effectors) that would otherwise excise or modify the uracil and / or surrounding nucleotides. It is understood that the use of a UGI may enhance the editing efficiency of enzymes capable of deaminating C residues.
[0078] Suitable UGI protein and nucleotide sequences are provided herein; additional suitable UGI sequences will be known to those of skill in the art, see, e.g., Wang et al., Uracil-DNA glycosylase inhibitor gene of bacteriophage PBS2 encodes a binding protein specific for uracil-DNA glycosylase. J. Biol. Chem. 264:1163-1171 (1989); Lundquist et al., Site-directed mutagenesis and characterization of uracil-DNA glycosylase inhibitor protein. Role of specific carboxylic amino acids in complex formation with Escherichia coli uracil-DNA glycosylase. J. Biol. Chem. 272:21408-21419 (1997); Ravishankar et al., X-ray analysis of a complex of Escherichia coli uracil DNA glycosylase (EcUDG) with a proteinaceous inhibitor. The structure elucidation of a prokaryotic UDG. Nucleic Acids Res. 26:4880-4887 (1998); and Putnam et al., Protein mimicry of DNA from crystal structures of the uracil-DNA glycosylase inhibitor protein and its complex with Escherichia coli uracil-DNA glycosylase. J. Mol. Biol. 287:331-346 (1999), the entire contents of each of which are incorporated herein by reference. It should be understood that any protein capable of inhibiting uracil-DNA glycosylase base excision repair enzymes is within the scope of the present disclosure.Furthermore, any protein that blocks or inhibits base excision repair is also within the scope of the present disclosure. In some embodiments, the uracil glycosylase inhibitor is a protein that binds uracil. In some embodiments, the uracil glycosylase inhibitor is a protein that binds uracil in DNA. In some embodiments, the uracil glycosylase inhibitor is a single-stranded binding protein. In some embodiments, the uracil glycosylase inhibitor is a catalytically inactive uracil DNA-glycosylase protein. In some embodiments, the uracil glycosylase inhibitor is a catalytically inactive uracil DNA-glycosylase protein that does not excise uracil from DNA. In some embodiments, the uracil glycosylase inhibitor is a catalytically inactive UDG.
[0079] In some embodiments, a uracil glycosylase inhibitor (UGI) disclosed herein comprises an amino acid sequence at least 80% identical to SEQ ID NO: 27 or 43. In some embodiments, any of the foregoing levels of identity is at least 90%, at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, a UGI comprises an amino acid sequence at least 90% identical to SEQ ID NO: 27 or 43. In some embodiments, a UGI comprises an amino acid sequence at least 95% identical to SEQ ID NO: 27 or 43. In some embodiments, a UGI comprises an amino acid sequence at least 98% identical to SEQ ID NO: 27 or 43. In some embodiments, a UGI comprises an amino acid sequence at least 99% identical to SEQ ID NO: 27 or 43. In some embodiments, a UGI comprises the amino acid sequence of SEQ ID NO: 27 or 43.
[0080] B. Cytidine deaminase Cytidine deaminases include enzymes of the cytidine deaminase superfamily, in particular enzymes of the APOBEC family (APOBEC1, APOBEC2, APOBEC4, and APOBEC3 subgroups of enzymes), activation-induced cytidine deaminase (AID or AICDA) and CMP deaminase (see, e.g., Conticello et al., Mol. Biol. Evol. 22:367-77, 2005; Conticello, Genome Biol. 9:229, 2008; Muramatsu et al., J. Biol. Chem. 274:18470-6, 1999; and Carrington et al., Cells 9:1690 (2020)).
[0081] In some embodiments, the cytidine deaminase disclosed herein is an enzyme of the APOBEC family. In some embodiments, the cytidine deaminase disclosed herein is an enzyme of the APOBEC1, APOBEC2, APOBEC4, and APOBEC3 subgroups. In some embodiments, the cytidine deaminase disclosed herein is an enzyme of the APOBEC3 subgroup. In some embodiments, the cytidine deaminase disclosed herein is an APOBEC3A deaminase (A3A).
[0082] In some embodiments, the cytidine deaminase is (i) an enzyme of the APOBEC family, optionally an enzyme of the APOBEC3 subgroup; (ii) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 41, and 960-1023; (iii) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 41, and 960-1013; (iv) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009; or (v) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 976, 981, 984, 986, and 1014-1023.
[0083] In some embodiments, the cytidine deaminase is a cytidine deaminase comprising an amino acid sequence at least 80%, 85%, 87%, 90%, 95%, 98%, 99%, or 100% identical to SEQ ID NOs: 40, 41, and 960-1023. In some embodiments, the cytidine deaminase is a cytidine deaminase comprising an amino acid sequence that is at least 80%, 85%, 87%, 90%, 95%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 40, 41, and 960-1013. In some embodiments, the cytidine deaminase is a cytidine deaminase comprising an amino acid sequence that is at least 80%, 85%, 87%, 90%, 95%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the cytidine deaminase is a cytidine deaminase comprising an amino acid sequence that is at least 80%, 85%, 87%, 90%, 95%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 40, 976, 981, 984, 986, and 1014-1023. In some embodiments, the cytidine deaminase is a cytidine deaminase comprising an amino acid sequence that is at least 80%, 85%, 87%, 90%, 95%, 98%, 99%, or 100% identical to SEQ ID NOs: 976, 977, 993-1006, and 1009.
[0084] 1. APOBEC3A deaminase In some embodiments, the APOBEC3A deaminase (A3A) disclosed herein is human A3A. In some embodiments, the A3A is wild-type A3A.
[0085] In some embodiments, the A3A is an A3A variant. The A3A variant shares homology with wild-type A3A or a fragment thereof. In some embodiments, the A3A variant has at least about 80% identity, at least about 85% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to wild-type A3A. In some embodiments, an A3A variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes compared to wild-type A3A. In some embodiments, an A3A variant comprises a fragment of A3A such that the fragment has at least about 80% identity, at least about 90% identity, at least about 95% identity, at least about 96% identity, at least about 97% identity, at least about 98% identity, at least about 99% identity, at least about 99.5% identity, or at least about 99.9% identity to the corresponding fragment of wild-type A3A.
[0086] In some embodiments, an A3A variant is a protein having a sequence that differs from that of a wild-type A3A protein by one or more mutations, e.g., substitutions, deletions, insertions, or one or more single-point substitutions. In some embodiments, a truncated A3A sequence may be used, e.g., by deleting N-terminal, C-terminal, or internal amino acids. In some embodiments, a truncated A3A sequence is used in which 1 to 4 amino acids at the C-terminus of the sequence are deleted. In some embodiments, the APOBEC3A (e.g., human APOBEC3A) has a wild-type amino acid at position 57 (as numbered in the wild-type sequence). In some embodiments, the APOBEC3A (e.g., human APOBEC3A) has an asparagine at amino acid position 57 (as numbered in the wild-type sequence).
[0087] In some embodiments, the wild-type A3A is human A3A (UniPROT Accession ID: p319411, SEQ ID NO: 40).
[0088] In some embodiments, A3A disclosed herein comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 40. In some embodiments, the level of identity is at least 85%, at least 87%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, A3A comprises an amino acid sequence having at least 87% identity to SEQ ID NO: 40. In some embodiments, A3A comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 40. In some embodiments, A3A comprises an amino acid sequence having at least 95% identity to SEQ ID NO: 40. In some embodiments, A3A comprises an amino acid sequence having at least 98% identity to SEQ ID NO: 40. In some embodiments, A3A comprises an amino acid sequence having at least 99% identity to A3A No. 40. In some embodiments, A3A comprises the amino acid sequence of SEQ ID NO: 40.
[0089] C. Linker In some embodiments, a polypeptide comprising A3A and an RNA-guided nickase described herein further comprises a linker connecting the A3A and the RNA-guided nickase. In some embodiments, the linker is an organic molecule, a polymer, or a chemical moiety. In some embodiments, the linker is a peptide linker. In some embodiments, a nucleic acid encoding a polypeptide comprising A3A and an RNA-guided nickase further comprises a sequence encoding the peptide linker. An mRNA encoding an A3A-linker-RNA-guided nickase fusion protein is provided.
[0090] In some embodiments, the peptide linker is any stretch of amino acids having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, or more amino acids.
[0091] In some embodiments, the peptide linker is a 16-residue "XTEN" linker, or a variant thereof (see, e.g., the Examples; and Schellenberger et al. A recombinant polypeptide extends the in vivo half-life of peptides and proteins in a tunable manner. Nat. Biotechnol. 27, 1186-1190 (2009)). In some embodiments, the XTEN linker comprises a sequence that is any one of SGSETPGTSESATPES (SEQ ID NO: 46), SGSETPGTSESA (SEQ ID NO: 47), or SGSETPGTSESATPEGGSGGS (SEQ ID NO: 48). In some embodiments, the XTEN linker consists of the sequence SGSETPGTSESATPES (SEQ ID NO: 46), SGSETPGTSESA (SEQ ID NO: 47), or SGSETPGTSESATPEGGSGGS (SEQ ID NO: 48).
[0092] In some embodiments, the peptide linker is (GGGGS) n (e.g., SEQ ID NOs: 212, 216, 221, 240), (G) n , (EAAAK) n (e.g., SEQ ID NOs: 213, 219, 267), (GGS) n , SGSETPGTSESATPES (SEQ ID NO: 46) motif (see, e.g., Guilinger JP, Thompson DB, Liu D R. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014;32(6):577-82; the entire contents of which are incorporated herein by reference), or (XP) n motif, or any combination thereof, wherein n is independently an integer between 1 and 30. See WO2015089406, e.g., paragraph
[0012] (the entire contents of which are incorporated herein by reference).
[0093] In some embodiments, the peptide linker comprises one or more sequences selected from SEQ ID NOs: 46-59, 61, and 211-272. In some embodiments, the peptide linker comprises one or more sequences selected from SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 268, SEQ ID NO: 269, SEQ ID NO: 270, SEQ ID NO: 271, and SEQ ID NO: 272. In some embodiments, the peptide linker comprises the sequence of SEQ ID NO: 268.
[0094] D. RNA-guided nickase In some embodiments, the RNA-guided nickase disclosed herein is a Cas nickase. In some embodiments, the RNA-guided nickase is from a specific Cas nuclease whose catalytic domain(s) have been inactivated. In some embodiments, the RNA-guided nickase is a Class 2 Cas nickase, such as Cas9 nickase or Cpf1 nickase. In some embodiments, the RNA-guided nickase is S. pyogenes Cas9 nickase. In some embodiments, the RNA-guided nickase is Neisseria meningitidis Cas9 nickase.
[0095] In some embodiments, the RNA-guided nickase is a modified Class 2 Cas protein or is derived from a Class 2 Cas protein. In some embodiments, the RNA-guided nickase is modified or derived from a Cas protein, such as a Class 2 Cas nuclease (which may be, for example, a Type II, Type V, or Type VI Cas nuclease). Class 2 Cas nucleases include, for example, Cas9, Cpfl, C2cl, C2c2, and C2c3 proteins, and variants thereof. Examples of Cas9 nucleases include those in Type II CRISPR systems of S. pyogenes, S. aureus, and other prokaryotes (see, for example, the list in the next paragraph), and modified (e.g., engineered or mutated) versions thereof. See, for example, US2016 / 0312198 A1; US2016 / 0312199 A1, which are incorporated by reference in their entireties. Other examples of Cas nucleases include the Csm or Cmr complexes of type III CRISPR systems, or the Cas10, Csm1, or Cmr2 subunits thereof; and the Cascade complex of type I CRISPR systems, or the Cas3 subunit thereof. In some embodiments, the Cas nuclease may be from a type IIA, type IIB, or type IIC system. For a discussion of various CRISPR systems and Cas nucleases, see, e.g., Makarova et al., Nat. Rev. Microbiol. 9:467-477 (2011); Makarova et al., Nat. Rev. Microbiol, 13:722-36 (2015); Shmakov et al., Molecular Cell, 60:385-397 (2015).
[0096] Casnickases described herein include, but are not limited to, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus species, Staphylococcus aureus, Listeria innocua, Lactobacillus gasseri, Francisella novicida, Wolinella succinogenes, Sutterella wadsworthensis, Gammaproteobacterium, Neisseria meningitidis, Campylobacter jejuni, Pasteurella multocida, Fibrobacter succinogene, Rhodospirillum rubrum, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Lactobacillus buchneri, Treponema denticola, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas species, Crocosphaera watsonii, Cyanothece species, Microcystis aeruginosa, Synechococcus species, Acetohalobium arabaticum, Ammonifex degensii, Caldicellosiruptor becscii, Candidatus Desulforudis, Clostridiumbotulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter species, Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohlobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc species, Arthrospira maxima, Arthrospira platensis, Arthrospira species, Lyngbya species, Microcoleus chthonoplastes, Oscillatoria species, Petrotoga mobilis, Thermosipho africanus, Streptococcus pasteurianus, Neisseria cinerea, Campylobacter lari, Parvibaculum lavamentivorans, Corynebacterium diphtheria, Acidaminococcus species, Lachnospiraceae bacterium ND2006, or Acaryochloris marina.
[0097] In some embodiments, the Cas nickase is a nickase form of Cas9 nuclease from Streptococcus pyogenes. In some embodiments, the Cas nickase is a nickase form of Cas9 nuclease from Streptococcus thermophilus. In some embodiments, the Cas nickase is a nickase form of Cas9 nuclease from Neisseria meningitidis. See, e.g., WO / 2020081568, which describes the Nme2Cas9 D16A nickase. In some embodiments, the Cas nickase is a nickase form of Cas9 nuclease from Staphylococcus aureus. In some embodiments, the Cas nickase is a nickase form of Cpf1 nuclease from Francisella novicida. In some embodiments, the Cas nickase is a nickase form of Cpf1 nuclease from Acidaminococcus species. In some embodiments, the Cas nickase is a nickase form of Cpf1 nuclease from Lachnospiraceae bacterium ND2006. In further embodiments, the Cas nickase is a nickase form of Cpf1 nuclease from Francisella tularensis, Lachnospiraceae bacterium, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium, Parcubacteria bacterium, Smithella, Acidaminococcus, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi, Leptospira inadai, Porphyromonas crevioricanis, Prevotella disiens, or Porphyromonas macacae. In certain embodiments, the Cas nickase is a nickase form of Cpf1 nuclease from Acidaminococcus or Lachnospiraceae.As discussed elsewhere, a nickase may be derived from (i.e., related to) a particular Cas nuclease in that the nickase is a form of a nuclease in which one of its two catalytic domains has been inactivated by mutating active site residues essential for nucleolysis, such as D10, H840, or N863 of Spy Cas 9. Those skilled in the art will be familiar with techniques for readily identifying corresponding residues in other Cas proteins, such as sequence alignment and structural alignment, which are described in more detail below.
[0098] In other embodiments, the Cas nickase may be associated with a type I CRISPR / Cas system. In some embodiments, the Cas nickase may be a component of a Cascade complex of a type I CRISPR / Cas system. In some embodiments, the Cas nickase may be a Cas3 protein. In some embodiments, the Cas nickase may be from a type III CRISPR / Cas system.
[0099] In some embodiments, the Cas nickase is a nickase form of a Cas nuclease or modified Cas nuclease in which the endonucleolytic active site has been inactivated, for example, by one or more modifications (e.g., point mutations) in the catalytic domain (see, e.g., U.S. Patent No. 8,889,356 for a discussion of Cas nickases and exemplary catalytic domain modifications).
[0100] Wild-type S. pyogenes Cas9 has two catalytic domains: RuvC and HNH. The RuvC domain cleaves the non-target DNA strand, and the HNH domain cleaves the target strand of DNA. In some embodiments, the Cas nuclease can include an amino acid substitution in the RuvC or RuvC-like nuclease domain. Exemplary amino acid substitutions in the RuvC or RuvC-like nuclease domain include D10A (based on the S. pyogenes Cas9 protein). See, e.g., Zetsche et al. (2015) Cell Oct 22:163(3):759-771. In some embodiments, the Cas nuclease can include an amino acid substitution in the HNH or HNH-like nuclease domain. Exemplary amino acid substitutions in the HNH or HNH-like nuclease domain include E762A, H840A, N863A, H983A, and D986A (based on the S. pyogenes Cas9 protein). See, e.g., Zetsche et al. (2015). Additional exemplary amino acid substitutions include D917A, E1006A, and D1255A (based on the Francisella novicida U112 Cpf1 (FnCpf1) sequence (UniProtKB-A0Q7Q2(CPF1_FRATN))).
[0101] In some embodiments, a Cas nickase, such as a Cas9 nickase, has an inactivated RuvC domain or HNH domain. In some embodiments, a nickase with a RuvC domain that has reduced activity is used. In some embodiments, a nickase with an inactive RuvC domain is used. In some embodiments, a nickase with an HNH domain that has reduced activity is used. In some embodiments, a nickase with an inactive HNH domain is used.
[0102] In some embodiments, the Cas9 nickase has an active HNH nuclease domain, capable of cleaving the non-target strand of DNA, i.e., the strand bound by the gRNA, and an inactive RuvC nuclease domain, unable to cleave the target strand of DNA, i.e., the strand where base editing by the deaminase is desired.
[0103] An exemplary Cas9 nickase amino acid sequence is provided as SEQ ID NO: 70. An exemplary Cas9 nickase mRNA ORF sequence, including a start and stop codon, is provided as SEQ ID NO: 71. An exemplary Cas9 nickase mRNA coding sequence suitable for incorporation into a fusion protein is provided as SEQ ID NO: 72.
[0104] In some embodiments, the RNA-guided nickase is a Class 2 Cas nickase described herein. In some embodiments, the RNA-guided nickase is a Cas9 nickase described herein.
[0105] In some embodiments, the RNA-guided nickase is a S. pyogenes Cas9 nickase described herein.
[0106] In some embodiments, the RNA-guided nickase is a D10A SpyCas9 nickase described herein. In some embodiments, the RNA-guided nickase comprises an amino acid sequence having at least 80%, 90%, 95%, 98%, or 99% identity to any one of SEQ ID NOs: 70, 73, or 76. In some embodiments, the RNA-guided nickase comprises the amino acid sequence of SEQ ID NO: 70.
[0107] In some embodiments, the mRNA ORF sequence encoding the RNA-guided nickase, including the start codon and stop codon, comprises a nucleotide sequence having at least 80%, 90%, 95%, 98%, 99%, or 100% identity to the nucleotide sequence of any one of SEQ ID NOs: 71, 74, or 77. In some embodiments, the mRNA sequence encoding the RNA-guided nickase comprises a nucleotide sequence having at least 80%, 90%, 95%, 98%, 99%, or 100% identity to the nucleotide sequence of any one of SEQ ID NOs: 72, 75, or 78. In some embodiments, the level of identity is at least 90%. In some embodiments, the level of identity is at least 95%. In some embodiments, the level of identity is at least 98%. In some embodiments, the level of identity is at least 99%. In some embodiments, the level of identity is at least 100%. In some embodiments, the sequence encoding the RNA-guided nickase comprises the nucleotide sequence of any one of SEQ ID NOs: 71, 72, 74, 75, 77, or 78.
[0108] In some embodiments, the RNA-guided nickase is a Neisseria meningitidis (Nme) Cas9 nickase described herein.
[0109] In some embodiments, the RNA-guided nickase is a D16A NmeCas9 nickase described herein. In some embodiments, the D16A NmeCas9 nickase is a D16A Nme2Cas9 nickase. In some embodiments, the D16A Nme2Cas9 nickase comprises an amino acid sequence at least 80%, 90%, 95%, 98%, 99%, or 100% identical to SEQ ID NO: 387. In some embodiments, the sequence encoding D16A Nme2Cas9 comprises a nucleotide sequence at least 80%, 90%, 95%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 388-393.
[0110] E. Compositions Comprising Cytidine Deaminase and RNA-Guided Nickase In some embodiments, mRNA encoding a polypeptide comprising a cytidine deaminase and an RNA-guided nickase, wherein the polypeptide is free of a uracil glycosylase inhibitor (UGI), is provided.
[0111] 1. Exemplary Compositions As described herein, compositions, methods, and uses are provided that include an mRNA comprising an open reading frame encoding a polypeptide comprising cytidine deaminase and RNA-guided nickase, the polypeptide being free of a uracil glycosylase inhibitor (UGI). For each exemplary composition described below, the mRNA does not include a UGI.
[0112] In some embodiments, mRNAs are provided that encode polypeptides comprising cytidine deaminase and RNA-guided nickase. In some embodiments, APOBEC family enzymes and RNA-guided nickases are provided. In some embodiments, the polypeptides comprise APOBEC1 subgroup enzymes and RNA-guided nickase. In some embodiments, the polypeptides comprise APOBEC2 subgroup enzymes and RNA-guided nickase. In some embodiments, the polypeptides comprise APOBEC4 subgroup enzymes and RNA-guided nickase. In some embodiments, the polypeptides comprise APOBEC3 subgroup enzymes and RNA-guided nickase.
[0113] In some embodiments, mRNA is provided that encodes a polypeptide comprising a cytidine deaminase and an RNA-guided nickase. In some embodiments, an enzyme of the APOBEC family and a D10A SpyCas9 nickase is provided. In some embodiments, the polypeptide comprises an enzyme of the APOBEC1 subgroup and a D10A SpyCas9 nickase. In some embodiments, the polypeptide comprises an enzyme of the APOBEC2 subgroup and a D10A SpyCas9 nickase. In some embodiments, the polypeptide comprises an enzyme of the APOBEC4 subgroup and a D10A SpyCas9 nickase. In some embodiments, the polypeptide comprises an enzyme of the APOBEC3 subgroup and a D10A SpyCas9 nickase.
[0114] In some embodiments, mRNA is provided that encodes a polypeptide comprising a cytidine deaminase and an RNA-guided nickase. In some embodiments, an enzyme of the APOBEC family and a D16A NmeCas9 nickase is provided. In some embodiments, an enzyme of the APOBEC family and a D16A Nme2Cas9 nickase is provided. In some embodiments, the polypeptide comprises an enzyme of the APOBEC1 subgroup and a D16A Nme2Cas9 nickase. In some embodiments, the polypeptide comprises an enzyme of the APOBEC2 subgroup and a D16A Nme2Cas9 nickase. In some embodiments, the polypeptide comprises an enzyme of the APOBEC4 subgroup and a D16A Nme2Cas9 nickase. In some embodiments, the polypeptide comprises an enzyme of the APOBEC3 subgroup and a D16A Nme2Cas9 nickase.
[0115] In some embodiments, the polypeptide lacks a UGI.
[0116] In some embodiments, the cytidine deaminase and the RNA-guided nickase are linked via a linker. In some embodiments, the cytidine deaminase and the RNA-guided nickase are linked via a peptide linker. In some embodiments, the peptide linker comprises one or more sequences selected from SEQ ID NOs: 46-59, 61, and 211-272.
[0117] In some embodiments, the polypeptide further comprises one or more additional heterologous functional domains. In some embodiments, the polypeptide further comprises one or more nuclear localization sequences (NLS) (described herein) at the C-terminus of the polypeptide or the N-terminus of the polypeptide.
[0118] In some embodiments, mRNAs are provided that encode polypeptides comprising cytidine deaminase and RNA-guided nickase. In some embodiments, APOBEC family enzymes and RNA-guided nickases are provided. In some embodiments, the polypeptides comprise APOBEC1 subgroup enzymes and RNA-guided nickase. In some embodiments, the polypeptides comprise APOBEC2 subgroup enzymes and RNA-guided nickase. In some embodiments, the polypeptides comprise APOBEC4 subgroup enzymes and RNA-guided nickase. In some embodiments, the polypeptides comprise APOBEC3 subgroup enzymes and RNA-guided nickase.
[0119] In some embodiments, an mRNA is provided that encodes a polypeptide comprising a cytidine deaminase and an RNA-guided nickase. In some embodiments, the polypeptide comprises an APOBEC family enzyme and a D10A SpyCas9 nickase, wherein the APOBEC family enzyme and the D10A SpyCas9 nickase are fused via a linker. In some embodiments, the polypeptide comprises an APOBEC family enzyme and a D10A SpyCas9 nickase, and a nuclear localization sequence (NLS) at the C-terminus of the fusion polypeptide. In some embodiments, the polypeptide comprises an APOBEC family enzyme and a D10A SpyCas9 nickase, and an NLS at the N-terminus of the fusion polypeptide. In some embodiments, the polypeptide comprises an APOBEC family enzyme and a D10A SpyCas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D10A SpyCas9 nickase, optionally via a linker. In some embodiments, the polypeptide comprises an APOBEC family enzyme and a D10A SpyCas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D10A SpyCas9 nickase, optionally via a linker.
[0120] In some embodiments, the polypeptide comprises an APOBEC family enzyme and a D16A Nme2Cas9 nickase fused via a linker. In some embodiments, the polypeptide comprises an APOBEC family enzyme and a D16A Nme2Cas9 nickase fused via a linker. In some embodiments, the polypeptide comprises an APOBEC family enzyme and a D16A Nme2Cas9 nickase, and a nuclear localization sequence (NLS) at the C-terminus of the fusion polypeptide. In some embodiments, the polypeptide comprises an APOBEC family enzyme and a D16A Nme2Cas9 nickase, and an NLS at the N-terminus of the fusion polypeptide. In some embodiments, the polypeptide comprises an APOBEC family enzyme and a D16A Nme2Cas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D16A Nme2Cas9 nickase, optionally via a linker. In some embodiments, the polypeptide comprises an APOBEC family enzyme and a D16A Nme2Cas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D16A Nme2Cas9 nickase, optionally via a linker.
[0121] In some embodiments, the polypeptide comprises an APOBEC1 subgroup enzyme and a D10A SpyCas9 nickase, wherein the APOBEC subgroup enzyme and the D10A SpyCas9 nickase are fused via a linker. In some embodiments, the polypeptide comprises an APOBEC1 subgroup enzyme and a D10A SpyCas9 nickase, and a nuclear localization sequence (NLS) at the C-terminus of the fusion polypeptide. In some embodiments, the polypeptide comprises an APOBEC1 subgroup enzyme and a D10A SpyCas9 nickase, and an NLS at the N-terminus of the fusion polypeptide. In some embodiments, the polypeptide comprises an APOBEC1 subgroup enzyme and a D10A SpyCas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D10A SpyCas9 nickase, optionally via a linker. In some embodiments, the polypeptide comprises an APOBEC1 subgroup enzyme and a D10A SpyCas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D10A SpyCas9 nickase, optionally via a linker.
[0122] In some embodiments, a polypeptide comprises an APOBEC1 subgroup enzyme and a D16A Nme2Cas9 nickase, wherein the APOBEC1 subgroup enzyme and the D16A Nme2Cas9 nickase are fused via a linker. In some embodiments, a polypeptide comprises an APOBEC1 subgroup enzyme and a D16A Nme2Cas9 nickase, wherein the APOBEC1 subgroup enzyme and the D16A Nme2Cas9 nickase are fused via a linker. In some embodiments, a polypeptide comprises an APOBEC1 subgroup enzyme and a D16A Nme2Cas9 nickase, and a nuclear localization sequence (NLS) at the C-terminus of the fusion polypeptide. In some embodiments, a polypeptide comprises an APOBEC1 subgroup enzyme and a D16A Nme2Cas9 nickase, and an NLS at the N-terminus of the fusion polypeptide. In some embodiments, a polypeptide comprises an APOBEC1 subgroup enzyme and a D16A Nme2Cas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D16A Nme2Cas9 nickase, optionally via a linker. In some embodiments, a polypeptide comprises an APOBEC1 subgroup enzyme and a D16A Nme2Cas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D16A Nme2Cas9 nickase, optionally via a linker.
[0123] In some embodiments, the polypeptide comprises an APOBEC3 subgroup enzyme and a D10A SpyCas9 nickase, wherein the APOBEC3 subgroup enzyme and the D10A SpyCas9 nickase are fused via a linker. In some embodiments, the polypeptide comprises an APOBEC3 subgroup enzyme and a D10A SpyCas9 nickase, and a nuclear localization sequence (NLS) at the C-terminus of the fusion polypeptide. In some embodiments, the polypeptide comprises an APOBEC3 subgroup enzyme and a D10A SpyCas9 nickase, and an NLS at the N-terminus of the fusion polypeptide. In some embodiments, the polypeptide comprises an APOBEC3 subgroup enzyme and a D10A SpyCas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D10A SpyCas9 nickase, optionally via a linker. In some embodiments, the polypeptide comprises an APOBEC3 subgroup enzyme and a D10A SpyCas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D10A SpyCas9 nickase, optionally via a linker.
[0124] In some embodiments, a polypeptide comprises an APOBEC3 subgroup enzyme and a D16A Nme2Cas9 nickase, wherein the APOBEC3 subgroup enzyme and the D16A Nme2Cas9 nickase are fused via a linker. In some embodiments, a polypeptide comprises an APOBEC3 subgroup enzyme and a D16A Nme2Cas9 nickase, wherein the APOBEC3 subgroup enzyme and the D16A Nme2Cas9 nickase are fused via a linker. In some embodiments, a polypeptide comprises an APOBEC3 subgroup enzyme and a D16A Nme2Cas9 nickase, and a nuclear localization sequence (NLS) at the C-terminus of the fusion polypeptide. In some embodiments, a polypeptide comprises an APOBEC3 subgroup enzyme and a D16A Nme2Cas9 nickase, and an NLS at the N-terminus of the fusion polypeptide. In some embodiments, a polypeptide comprises an APOBEC3 subgroup enzyme and a D16A Nme2Cas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D16A Nme2Cas9 nickase, optionally via a linker. In some embodiments, a polypeptide comprises an APOBEC3 subgroup enzyme and a D16A Nme2Cas9 nickase fused via a linker, and an NLS fused to the C-terminus of the D16A Nme2Cas9 nickase, optionally via a linker.
[0125] In some embodiments, the polypeptide comprises a D10A SpyCas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:268, and a cytidine deaminase comprising an amino acid sequence at least 85% identical to any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D10A SpyCas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:269, and a cytidine deaminase comprising an amino acid sequence at least 85% identical to any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D10A SpyCas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:270, and a cytidine deaminase comprising an amino acid sequence at least 85% identical to any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D10A SpyCas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:271, and a cytidine deaminase comprising an amino acid sequence at least 85% identical to any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D10A SpyCas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO: 272, and a cytidine deaminase comprising an amino acid sequence at least 85% identical to any one of SEQ ID NOs: 40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In any of the foregoing embodiments, the D10A SpyCas9 nickase can comprise an amino acid sequence at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to any one of SEQ ID NOs: 70, 73, or 76.
[0126] In some embodiments, the polypeptide comprises a D16A Nme2Cas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:268, and a cytidine deaminase comprising an amino acid sequence at least 85% identical to any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D16A Nme2Cas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:269, and a cytidine deaminase comprising an amino acid sequence at least 85% identical to any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D16A Nme2Cas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:270, and a cytidine deaminase comprising an amino acid sequence at least 85% identical to any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D16A Nme2Cas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:271, and a cytidine deaminase comprising an amino acid sequence at least 85% identical to any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D16A Nme2Cas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:272, and a cytidine deaminase comprising an amino acid sequence at least 85% identical to any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In any of the foregoing embodiments, the D16A Nme2Cas9 nickase can comprise an amino acid sequence at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:387.
[0127] In some embodiments, the polypeptide comprises a D10A SpyCas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO: 268, and a cytidine deaminase comprising an amino acid sequence selected from any one of SEQ ID NOs: 40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D10A SpyCas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO: 269, and a cytidine deaminase comprising an amino acid sequence selected from any one of SEQ ID NOs: 40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D10A SpyCas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO: 270, and a cytidine deaminase comprising an amino acid sequence selected from any one of SEQ ID NOs: 40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D10A SpyCas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO: 271, and a cytidine deaminase comprising an amino acid sequence selected from any one of SEQ ID NOs: 40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D10A SpyCas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO: 272, and a cytidine deaminase comprising an amino acid sequence selected from any one of SEQ ID NOs: 40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In any of the foregoing embodiments, the D10A SpyCas9 comprises an amino acid sequence at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to any one of SEQ ID NOs: 70, 73, or 76.
[0128] In some embodiments, the polypeptide comprises a D16A Nme2Cas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:268, and a cytidine deaminase comprising an amino acid sequence selected from any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D16A Nme2Cas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:269, and a cytidine deaminase comprising an amino acid sequence selected from any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D16A Nme2Cas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO: 270, and a cytidine deaminase comprising an amino acid sequence selected from any one of SEQ ID NOs: 40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D16A Nme2Cas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO: 271, and a cytidine deaminase comprising an amino acid sequence selected from any one of SEQ ID NOs: 40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In some embodiments, the polypeptide comprises a D16A Nme2Cas9 nickase, a linker comprising the amino acid sequence of SEQ ID NO:272, and a cytidine deaminase comprising an amino acid sequence selected from any one of SEQ ID NOs:40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009. In any of the foregoing embodiments, the D16A Nme2Cas9 nickase comprises an amino acid sequence at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identical to SEQ ID NO:387.
[0129] Polypeptides can be organized in any number of ways to form a single chain. The NLS can be N- or C-terminal, or both N- and C-terminal, and the cytidine deaminase can be N- or C-terminal to the RNA-guided nickase. In some embodiments, a polypeptide comprises, from N- to C-terminus, a cytidine deaminase, an optional linker, an RNA-guided nickase, and an optional NLS. In some embodiments, a polypeptide comprises, from N- to C-terminus, an RNA-guided nickase, an optional linker, a cytidine deaminase, and an optional NLS. In some embodiments, a polypeptide comprises, from N- to C-terminus, an optional NLS, an RNA-guided nickase, an optional linker, and a cytidine deaminase. In some embodiments, a polypeptide comprises, from N- to C-terminus, an optional NLS, an RNA-guided nickase, an optional linker, and a cytidine deaminase.
[0130] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, an APOBEC family enzyme, an optional linker, an RNA-guided nickase, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, an RNA-guided nickase, an optional linker, an APOBEC family enzyme, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, an RNA-guided nickase, an optional linker, an APOBEC family enzyme, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, an RNA-guided nickase, an optional linker, an APOBEC family enzyme, and an optional NLS.
[0131] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, an enzyme of the APOBEC3 subgroup, an optional linker, an RNA-guided nickase, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, an RNA-guided nickase, an optional linker, an enzyme of the APOBEC3 subgroup, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, an RNA-guided nickase, an optional linker, an enzyme of the APOBEC3 subgroup, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, an RNA-guided nickase, an optional linker, an enzyme of the APOBEC3 subgroup, and an optional NLS.
[0132] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, an APOBEC family enzyme, an optional linker, a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, an optional linker, an APOBEC family enzyme, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, an optional linker, an APOBEC family enzyme, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, an optional linker, and an APOBEC family enzyme, and an optional NLS.
[0133] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, an enzyme of the APOBEC3 subgroup, an optional linker, a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, an optional linker, an enzyme of the APOBEC3 subgroup, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, an optional linker, an enzyme of the APOBEC3 subgroup, and an optional NLS. In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, an optional linker, and an enzyme of the APOBEC3 subgroup, and an optional NLS.
[0134] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, an optional NLS, an enzyme of the APOBEC3 subgroup, an optional linker, and a D16A Nme2Cas9 nickase.
[0135] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, (i) an optional NLS, (ii) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 41, and 960-1023, (iii) a linker comprising one or more sequences selected from SEQ ID NOs: 46-59, 61, and 211-272, (iv) a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, and (v) an optional NLS.
[0136] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, (i) an optional NLS, (ii) a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, (iii) a linker comprising one or more sequences selected from SEQ ID NOs: 46-59, 61, and 211-272, (iv) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 41, and 960-1023, and (v) an optional NLS.
[0137] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, (i) an optional NLS, (ii) a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, (iii) a linker comprising one or more sequences selected from SEQ ID NOs: 46-59, 61, and 211-272, (iv) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 41, and 960-1023, and (v) an optional NLS.
[0138] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, (i) an optional NLS, (ii) a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, (iii) a linker comprising one or more sequences selected from SEQ ID NOs: 46-59, 61, and 211-272, (iv) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 41, and 960-1023, and (v) an optional NLS.
[0139] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, (i) an optional NLS, (ii) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 41, and 960-1023, (iii) a linker comprising one or more sequences selected from SEQ ID NOs: 46-59, 61, and 211-272, (iv) a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, and (v) an optional NLS.
[0140] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, (i) an optional NLS, (ii) a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, (iii) a linker comprising one or more sequences selected from SEQ ID NOs: 46-59, 61, and 211-272, (iv) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 41, and 960-1023, and (v) an optional NLS.
[0141] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, (i) an optional NLS, (ii) a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, (iii) a linker comprising one or more sequences selected from SEQ ID NOs: 46-59, 61, and 211-272, (iv) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 41, and 960-1023, and (v) an optional NLS.
[0142] In some embodiments, the polypeptide comprises, from N-terminus to C-terminus, (i) an optional NLS, (ii) a D10A SpyCas9 nickase or a D16A Nme2Cas9 nickase, (iii) a linker comprising one or more sequences selected from SEQ ID NOs: 46-59, 61, and 211-272, and (iv) a cytidine deaminase comprising an amino acid sequence at least 80% identical to any one of SEQ ID NOs: 40, 41, and 960-1023, and (v) an optional NLS.
[0143] 2. Composition Comprising APOBEC3A Deaminase and RNA-Guided Nickase In some embodiments, an mRNA is provided that encodes a polypeptide comprising APOBEC3A deaminase (A3A) and an RNA-guided nickase. In some embodiments, the polypeptide comprises human A3A and an RNA-guided nickase. In some embodiments, the polypeptide comprises wild-type A3A and an RNA-guided nickase. In some embodiments, the polypeptide comprises an A3A variant and an RNA-guided nickase. In some embodiments, the polypeptide comprises A3A and a Cas9 nickase. In some embodiments, the polypeptide comprises A3A and a D10A SpyCas9 nickase. In some embodiments, the polypeptide comprises human A3A and a D10A SpyCas9 nickase. In some embodiments, the polypeptide comprises an A3A variant and a D10A SpyCas9 nickase. In some embodiments, the polypeptide lacks UGI. In some embodiments, the A3A and the RNA-guided nickase are linked via a linker. In some embodiments, the polypeptide further comprises one or more additional heterologous functional domains. In some embodiments, the polypeptide further comprises a nuclear localization sequence (NLS) (described herein) at the C-terminus of the polypeptide or the N-terminus of the polypeptide.
[0144] In some embodiments, the polypeptide comprises human A3A and D10A SpyCas9 nickase, wherein the human A3A and the D10A SpyCas9 nickase are fused via a linker. In some embodiments, the polypeptide comprises human A3A and D10A SpyCas9 nickase and a nuclear localization sequence (NLS) at the C-terminus of the fusion polypeptide. In some embodiments, the polypeptide comprises human A3A and D10A SpyCas9 nickase and an NLS at the N-terminus of the fusion polypeptide. In some embodiments, the polypeptide comprises human A3A and D10A SpyCas9 nickase fused via a linker and an NLS fused to the C-terminus of the D10A SpyCas9 nickase, optionally via a linker. In some embodiments, the polypeptide comprises human A3A and D10A SpyCas9 nickases fused via a linker, and an NLS fused to the C-terminus of the D10A SpyCas9 nickase, optionally via a linker.
[0145] The polypeptides can be organized in any number of ways to form a single chain. The NLS can be N- or C-terminal, or both N- and C-terminal, and A3A can be N- or C-terminal to the RNA-guided nickase. In some embodiments, the polypeptide comprises, from N- to C-terminus, A3A, an optional linker, RNA-guided nickase, and an optional NLS. In some embodiments, the polypeptide comprises, from N- to C-terminus, RNA-guided nickase, an optional linker, A3A, and an optional NLS. In some embodiments, the polypeptide comprises, from N- to C-terminus, an optional NLS, RNA-guided nickase, an optional linker, and A3A. In some embodiments, the polypeptide comprises, from N- to C-terminus, an optional NLS, RNA-guided nickase, an optional linker, A3A, and an optional NLS.
[0146] In any of the foregoing embodiments, the polypeptide may comprise an amino acid sequence having at least 80% identity to SEQ ID NO: 3 or 6. In some embodiments, any of the foregoing levels of identity is at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, the polypeptides disclosed herein may comprise an amino acid sequence having at least 90% identity to SEQ ID NO: 3 or 6. In some embodiments, the polypeptides disclosed herein may comprise an amino acid sequence having at least 95% identity to SEQ ID NO: 3 or 6. In some embodiments, the polypeptides disclosed herein may comprise an amino acid sequence having at least 98% identity to SEQ ID NO: 3 or 6. In some embodiments, the polypeptides disclosed herein may comprise an amino acid sequence having at least 99% identity to SEQ ID NO: 3 or 6. In some embodiments, the polypeptides disclosed herein may comprise the amino acid sequence of SEQ ID NO: 3 or 6.
[0147] In any of the foregoing embodiments, a nucleic acid sequence comprising an open reading frame encoding a polypeptide disclosed herein may comprise a nucleic acid sequence having at least 80% identity to SEQ ID NO: 2 or 5. In some embodiments, any of the foregoing levels of identity is at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%.
[0148] In any of the foregoing embodiments, an mRNA sequence encoding a polypeptide disclosed herein may comprise a nucleic acid sequence having at least 80% identity to SEQ ID NO: 1 or 4. In some embodiments, any of the foregoing levels of identity is at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%.
[0149] In any of the foregoing embodiments, the polypeptide may comprise an amino acid sequence having at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 303, 306, 309, or 312. In some embodiments, a polypeptide disclosed herein may comprise the amino acid sequence of SEQ ID NO: 303, 306, 309, or 312. In any of the foregoing embodiments, a nucleic acid sequence comprising an open reading frame encoding a polypeptide disclosed herein may comprise a nucleic acid sequence having at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 302, 305, 308, or 311. In some embodiments, a nucleic acid sequence comprising an open reading frame encoding a polypeptide disclosed herein comprises the nucleic acid sequence of SEQ ID NO: 302, 305, 308, or 311. In any of the foregoing embodiments, an mRNA sequence encoding a polypeptide disclosed herein may comprise a nucleic acid sequence having at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% identity to SEQ ID NO: 301, 304, 307, or 310. In any of the foregoing embodiments, an mRNA sequence encoding a polypeptide disclosed herein may comprise the nucleic acid sequence of SEQ ID NO: 301, 304, 307, or 310.
[0150] In any of the foregoing embodiments, A3A may comprise an amino acid sequence having at least 80% identity to SEQ ID NO: 40. In some embodiments, the level of identity is at least 85%, at least 87%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, A3A comprises the amino acid sequence of SEQ ID NO: 40.
[0151] In any of the foregoing embodiments, the RNA-guided nickase may comprise an amino acid sequence having at least 80%, 90%, 95%, 98%, or 99% identity to any one of SEQ ID NOs: 70, 73, or 76. In some embodiments, the level of identity is at least 85%, at least 87%, at least 90%, at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, the RNA-guided nickase comprises the amino acid sequence of SEQ ID NO: 70. In some embodiments, the RNA-guided nickase comprises the amino acid sequence of SEQ ID NO: 73. In some embodiments, the RNA-guided nickase comprises the amino acid sequence of SEQ ID NO: 76.
[0152] In any of the foregoing embodiments, A3A may comprise an amino acid sequence having at least 80% identity to SEQ ID NO:40, and the RNA-guided nickase may comprise an amino acid sequence having at least 80%, 90%, 95%, 98%, or 99% identity to any one of SEQ ID NOs:70, 73, or 76. In some embodiments, A3A comprises the amino acid sequence of SEQ ID NO:40, and the RNA-guided nickase comprises the amino acid sequence of SEQ ID NO:70.
[0153] F. Additional Features 1. Codon optimization In some embodiments, the polypeptide comprising UGI, or a cytidine deaminase (e.g., APOBEC3A deaminase) and an RNA-guided nickase, is encoded by an open reading frame (ORF) comprising a codon-optimized nucleic acid sequence. In some embodiments, the codon-optimized nucleic acid sequence comprises minimal adenine codons and / or minimal uridine codons.
[0154] A given ORF can have a reduced uridine content or uridine dinucleotide content, for example, by using minimal uridine codons in a sufficient portion of the ORF. For example, the amino acid sequences of the polypeptides described herein can be reverse-translated into ORF sequences by converting amino acids to codons, with some or all of the ORF using the exemplary minimal uridine codons shown below. In some embodiments, at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons in the ORF are codons listed in Table 1. [Table 2]
[0155] In some embodiments, an ORF can consist of a set of codons in which at least about 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons are codons listed in Table 1.
[0156] A given ORF can have a reduced adenine or adenine dinucleotide content, for example, by using minimal adenine codons in a sufficient portion of the ORF. For example, the amino acid sequences of the polypeptides described herein can be reverse-translated into ORF sequences by converting amino acids to codons, with some or all of the ORF using the exemplary minimal adenine codons shown below. In some embodiments, at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons in the ORF are codons listed in Table 2. [Table 3]
[0157] In some embodiments, an ORF can consist of a set of codons in which at least about 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons are codons listed in Table 2.
[0158] To the extent feasible, any of the features described above for low adenine content can be combined with any of the features described above for low uridine content. The same applies to uridine and adenine dinucleotides. Similarly, the content of uridine nucleotides and adenine dinucleotides in the ORF can be as described above. Similarly, the content of uridine dinucleotides and adenine nucleotides in the ORF can be as described above.
[0159] A given ORF can have a reduced uridine and adenine nucleotide and / or dinucleotide content, for example, by using minimal uridine and adenine codons in a sufficient portion of the ORF. For example, the amino acid sequences of the polypeptides described herein can be reverse-translated into ORF sequences by converting amino acids to codons, with some or all of the ORF using the exemplary minimal uridine and adenine codons shown below. In some embodiments, at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons in the ORF are codons listed in Table 3. [Table 4]
[0160] In some embodiments, an ORF can consist of a set of codons in which at least about 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the codons are codons listed in Table 3. As seen in Table 3, each of the three serine codons listed contains either one A or one U. In some embodiments, uridine minimization is prioritized by using the AGC codon for serine. In some embodiments, adenine minimization is prioritized by using the UCC and / or UCG codons for serine.
[0161] In some embodiments, the ORF can have codons that increase translation in a mammal, e.g., a human. In further embodiments, the mRNA includes an ORF with codons that increase translation in a mammal, e.g., a human organ, such as the liver. In further embodiments, the ORF can have codons that increase translation in a mammal, e.g., a cell type, such as a human hepatocyte. Increased translation in a mammal, cell type, mammalian organ, human, human organ, etc. can be determined relative to the degree of translation of the wild-type sequence of the ORF, or relative to an ORF with a codon distribution that matches the codon distribution of the organism from which the ORF is derived or the organism containing the most similar ORF at the amino acid level. Alternatively, in some embodiments, increased translation of a Cas9 sequence in a mammal, cell type, mammalian organ, human, human organ, etc. is determined relative to the translation of an ORF having the sequence of SEQ ID NO: 2 or 5, all other things being equal (including any applicable point mutations, heterologous domains, etc.). In some embodiments, at least about 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in an ORF are codons that correspond to highly expressed tRNAs in a mammal, e.g., a human (e.g., the most highly expressed tRNA for each amino acid). In some embodiments, at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in an ORF are codons that correspond to highly expressed tRNAs in a mammalian organ, e.g., a human organ (e.g., the most highly expressed tRNA for each amino acid).
[0162] Alternatively, codons corresponding to tRNAs that are highly expressed in common organisms (eg, humans) may be used.
[0163] Any of the foregoing approaches to codon selection can be combined with the minimal uridine and / or adenine codons shown above, for example, by starting with a codon from Table 1, 2, or 3 and then, if multiple choices are available, using the codon that corresponds to the more highly expressed tRNA, either in the organism in general (e.g., humans) or in the organ or cell type of interest (e.g., human liver or human hepatocytes).
[0164] In some embodiments, at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in an ORF are from a codon set shown in Table 4 (e.g., a low U1, low A, or low A / U codon set). Codons in the low U1, low G, low A, and low A / U sets use codons that minimize the indicated nucleotides, while using codons that correspond to highly expressed tRNAs when multiple choices are available. In some embodiments, at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in an ORF are from a low U1 codon set shown in Table 4. In some embodiments, at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in an ORF are from the low A codon set shown in Table 4. In some embodiments, at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the codons in an ORF are from the low A / U codon set shown in Table 4. [Table 5]
[0165] 2. Heterologous functional domains; nuclear localization signals (NLS) In some embodiments, the polypeptide comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase further comprises one or more additional heterologous functional domains (e.g., is or comprises a ternary or higher order fusion polypeptide).
[0166] In some embodiments, the heterologous functional domain can facilitate transport of the polypeptide to the nucleus of the cell. For example, the heterologous functional domain can be a nuclear localization signal (NLS). In some embodiments, the polypeptide can be fused to 1 to 10 NLS(s). In some embodiments, the polypeptide can be fused to 1 to 5 NLS(s). In some embodiments, the polypeptide can be fused to one NLS. When one NLS is used, the NLS can be fused to the N-terminus or C-terminus of the polypeptide sequence. In some embodiments, the polypeptide can be fused to the C-terminus of at least one NLS. An NLS can also be inserted within the polypeptide sequence. In other embodiments, the polypeptide can be fused to two or more NLSs. In some embodiments, the polypeptide can be fused to two, three, four, or five NLSs. In some embodiments, the polypeptide can be fused to two NLSs. In certain situations, the two NLSs can be the same (e.g., two SV40 NLSs) or different. In some embodiments, the polypeptide is fused to two SV40 NLS sequences at the carboxy terminus. In some embodiments, a polypeptide may be fused to two NLSs (one at the N-terminus and one at the C-terminus). In some embodiments, a polypeptide may be fused to three NLSs. In some embodiments, a polypeptide may not be fused to an NLS. In some embodiments, an NLS may be a monokaryotic sequence, such as the SV40 NLS, PKKKRKV (SEQ ID NO: 63) or PKKKRRV (SEQ ID NO: 121). In some embodiments, an NLS may be a bikaryotic sequence, such as the nucleoplasmin NLS, KRPAATKKAGQAKKKK (SEQ ID NO: 122). In certain embodiments, a single PKKKRKV (SEQ ID NO: 63) NLS may be fused to the C-terminus of a polypeptide. One or more linkers are optionally included at the fusion site (e.g., between the polypeptide and the NLS). In some embodiments, one or more NLS(s) according to any of the foregoing embodiments are present in a polypeptide in combination with one or more additional heterologous functional domains, for example, any of the heterologous functional domains described below.
[0167] In some embodiments of the mRNA disclosed herein, a cytidine deaminase (e.g., A3A) is located at the N-terminus of the RNA-guided nickase in the polypeptide. In some embodiments of the mRNA disclosed herein, the encoded RNA-guided nickase comprises a nuclear localization signal (NLS). In some embodiments, the NLS is fused to the C-terminus of the RNA-guided nickase. In some embodiments, the NLS is fused to the C-terminus of the RNA-guided nickase via a linker. In some embodiments, the NLS is fused to the N-terminus of the RNA-guided nickase. In some embodiments, the NLS is fused to the N-terminus of the RNA-guided nickase via a linker (e.g., SEQ ID NO: 61). In some embodiments, the NLS comprises a sequence having at least 80%, 85%, 90%, or 95% identity to any one of SEQ ID NOs: 63 and 110-122. In some embodiments, the NLS comprises the sequence of any one of SEQ ID NOs: 63 and 110-122. In some embodiments, the NLS is encoded by a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 63 and 110-122.
[0168] In some embodiments, the heterologous functional domain may be capable of altering the intracellular half-life of A3A and / or RNA-guided nickase in a polypeptide. In some embodiments, the half-life of A3A and / or RNA-guided nickase in a polypeptide may be increased. In some embodiments, the half-life of A3A and / or RNA-guided nickase in a polypeptide may be decreased. In some embodiments, the heterologous functional domain may be capable of increasing the stability of A3A and / or RNA-guided nickase in a polypeptide. In some embodiments, the heterologous functional domain may be capable of decreasing the stability of A3A and / or RNA-guided nickase in a polypeptide. In some embodiments, the heterologous functional domain may function as a signal peptide for protein degradation. In some embodiments, protein degradation may be mediated by proteolytic enzymes such as, for example, proteasomes, lysosomal proteases, or calpain proteases. In some embodiments, the heterologous functional domain may comprise a PEST sequence. In some embodiments, the polypeptide may be modified by the addition of ubiquitin or polyubiquitin chains. In some embodiments, the ubiquitin can be a ubiquitin-like protein (UBL). Non-limiting examples of ubiquitin-like proteins include small ubiquitin-like modifier (SUMO), ubiquitin cross-reactive protein (UCRP, also known as interferon-stimulated gene 15 (ISG15)), ubiquitin-related modifier 1 (URM1), neural progenitor cell-expressed developmentally downregulated protein 8 (NEDD8, also known as Rub1 in S. cerevisiae), human leukocyte antigen F-related (FAT10), autophagy 8 (ATG8) and 12 (ATG12), Fau ubiquitin-like protein (FUB1), membrane-anchored UBL (MUB), ubiquitin-fold modifier 1 (UFM1), and ubiquitin-like protein 5 (UBL5).
[0169] In some embodiments, the heterologous functional domain can be a marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, epitope tags, and reporter gene sequences. In some embodiments, the marker domain can be a fluorescent protein. Any known fluorescent protein, such as GFP, YFP, EBFP, ECFP, DsRed, or any other suitable fluorescent protein, can be used as the marker domain. In some embodiments, the marker domain can be a purification tag and / or an epitope tag. Non-limiting exemplary tags include glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein (MBP), thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, 8xHis, biotin carboxyl carrier protein (BCCP), poly-His, and calmodulin. In some embodiments, the marker domain can be a reporter gene. Non-limiting exemplary reporter genes include glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, or fluorescent protein.
[0170] In additional embodiments, the heterologous functional domain may target the polypeptide to a particular organelle, cell type, tissue, or organ, hi some embodiments, the heterologous functional domain may target the polypeptide to mitochondria.
[0171] 3.UTR; Kozak sequence In some embodiments, a nucleic acid (e.g., mRNA) disclosed herein comprises the 5' UTR, 3' UTR, or 5' and 3' UTRs from hydroxysteroid 17 beta dehydrogenase 4 (HSD17B4 or HSD), or a globin, such as human alpha globin (HBA), human beta globin (HBB), Xenopus laevis beta globin (XBG), bovine growth hormone, cytomegalovirus (CMV), mouse Hba-a1, heat shock protein 90 (Hsp90), glyceraldehyde-3-phosphate dehydrogenase (GAPDH), beta actin, alpha tubulin, tumor protein (p53), or epidermal growth factor receptor (EGFR).
[0172] In some embodiments, the nucleic acids disclosed herein comprise a 5' UTR from an HSD and a 3' UTR from a human albumin gene. In some embodiments, the mRNAs disclosed herein comprise a 5' UTR having at least 90% identity to any one of SEQ ID NOs: 93 and a 3' UTR having at least 90% identity to any one of SEQ ID NOs: 69.
[0173] In some embodiments, a nucleic acid disclosed herein comprises a 5' UTR having at least 90% identity to any one of SEQ ID NOs: 91-98. In some embodiments, an mRNA disclosed herein comprises a 3' UTR having at least 90% identity to any one of SEQ ID NOs: 69, 99-106. In some embodiments, any of the foregoing levels of identity is at least 95%, at least 98%, at least 99%, or 100%. In some embodiments, an mRNA disclosed herein comprises a 5' UTR having the sequence of any one of SEQ ID NOs: 91-98. In some embodiments, an mRNA disclosed herein comprises a 3' UTR having the sequence of any one of SEQ ID NOs: 69, 99-106. In some embodiments, an mRNA comprises a 5' UTR and a 3' UTR from the same source.
[0174] In some embodiments, the nucleic acids described herein do not include a 5'UTR, e.g., there are no additional nucleotides between the 5' cap and the start codon. In some embodiments, the mRNA includes a Kozak sequence (described below) between the 5' cap and the start codon, but does not have an additional 5'UTR. In some embodiments, the mRNA does not include a 3'UTR, e.g., there are no additional nucleotides between the stop codon and the polyA tail.
[0175] In some embodiments, the nucleic acids herein comprise a Kozak sequence. Kozak sequences can affect translation initiation and the overall yield of polypeptides translated from mRNA. Kozak sequences include a methionine codon, which can function as a start codon. A minimal Kozak sequence is NNNRUGN, where at least one of the following is true: the first N is A or G and the second N is G. In the context of the nucleotide sequence, R denotes a purine (A or G). In some embodiments, the Kozak sequence is RNNRUGN, NNNRUGG, RNNRUGG, RNNAUGN, NNNAUGG, RNNAUGG, or GCCACCAUG. In some embodiments, the Kozak sequence is rccRUGg, rccAUGg, gccAccAUG, gccRccAUGG (SEQ ID NO: 107), or gccgccRccAUGG (SEQ ID NO: 108), with zero mismatches or up to one or two mismatches at the lowercase positions.
[0176] 4. PolyA tail In some embodiments, the nucleic acids disclosed herein further comprise a polyadenylation (polyA) tail. The polyA tail can comprise at least eight consecutive adenine nucleotides, but may also comprise one or more non-adenine nucleotides. As used herein, "non-adenine nucleotide" refers to any natural or non-natural nucleotide that does not contain adenine. Guanine, thymine, and cytosine nucleotides are exemplary non-adenine nucleotides. Thus, the polyA tail on a nucleic acid described herein can comprise consecutive adenine nucleotides located 3' to the nucleotide encoding a polypeptide of interest. In some examples, the polyA tail on an mRNA comprises non-consecutive adenine nucleotides located 3' to the nucleotide encoding a polypeptide comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase or a sequence of interest, with the non-adenine nucleotides interrupting the adenine nucleotides at regular or irregular intervals.
[0177] In some embodiments, the polyA tail is encoded by a plasmid used for in vitro transcription of mRNA and becomes part of the transcript. The polyA sequence encoded by the plasmid, i.e., the number of consecutive adenine nucleotides in the polyA sequence, may not be precise; for example, 100 polyA sequences in the plasmid may not result in exactly 100 polyA sequences in the transcribed mRNA. In some embodiments, the polyA tail is not encoded by the plasmid but is added by PCR tailing or enzymatic tailing, e.g., using E. coli poly(A) polymerase.
[0178] In some embodiments, one or more non-adenine nucleotides are positioned to interrupt consecutive adenine nucleotides, thereby allowing poly(A) binding proteins to bind to the stretch of consecutive adenine nucleotides. In some embodiments, the one or more non-adenine nucleotide(s) are positioned after at least 8, 9, 10, 11, or 12 consecutive adenine nucleotides. In some embodiments, the one or more non-adenine nucleotides are positioned after 8 to 50 consecutive adenine nucleotides. In some embodiments, the one or more non-adenine nucleotides are positioned after 8 to 100 consecutive adenine nucleotides.
[0179] In some embodiments, the poly-A tail comprises or contains one non-adenine nucleotide or one contiguous stretch of 2-10 non-adenine nucleotides.
[0180] In some embodiments, the non-adenine nucleotides are guanine, cytosine, or thymine. In some cases where more than one non-adenine nucleotide is present, the non-adenine nucleotides may be selected from the following: a) guanine and thymine nucleotides, b) guanine and cytosine nucleotides, c) thymine and cytosine nucleotides, or d) guanine, thymine, and cytosine nucleotides. An exemplary poly-A tail comprising non-adenine nucleotides is provided as SEQ ID NO: 109.
[0181] 5. Modified Nucleotides In some embodiments, the nucleic acids disclosed herein contain modified uridines at some or all uridine positions. In some embodiments, the modified uridine is a uridine modified at the 5-position, for example, with a halogen or a C1-C3 alkoxy. In some embodiments, the modified uridine is a pseudouridine modified at the 1-position, for example, with a C1-C3 alkyl. The modified uridine can be, for example, pseudouridine, N1-methyl-pseudouridine, 5-methoxyuridine, 5-iodouridine, or a combination thereof.
[0182] In some embodiments, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% of the uridine positions in a nucleic acid disclosed herein are modified uridines. In some embodiments, 10% to 25%, 15% to 25%, 25% to 35%, 35% to 45%, 45% to 55%, 55% to 65%, 65% to 75%, 75% to 85%, 85% to 95%, or 90% to 100% of the uridine positions in an mRNA disclosed herein are modified uridines, such as 5-methoxyuridine, 5-iodouridine, N1-methylpseudouridine, pseudouridine, or a combination thereof.
[0183] In some embodiments, at least 10% of the uridines are substituted with modified uridines. In some embodiments, 15% to 45% of the uridines are substituted with modified uridines. In some embodiments, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100% of the uridines are substituted with modified uridines.
[0184] 6.5' Cap In some embodiments, the nucleic acids disclosed herein include a 5' cap, e.g., Cap0, Cap1, or Cap2. A 5' cap is generally a 7-methylguanine ribonucleotide (which may be further modified, e.g., as described below for ARCA) linked via a 5'-triphosphate to the 5' position of the first nucleotide of the 5' to 3' strand of the nucleic acid, i.e., the first cap-proximal nucleotide. In Cap0, the riboses of the first and second cap-proximal nucleotides of the mRNA both include a 2'-hydroxyl. In Cap1, the riboses of the first and second transcribed nucleotides of the mRNA include a 2'-methoxy and a 2'-hydroxyl, respectively. In Cap2, the riboses of the first and second cap-proximal nucleotides of the mRNA both include a 2'-methoxy. See, for example, Katibah et al. (2014) Proc Natl Acad Sci USA 111(33):12025-30; Abbas et al. (2017) Proc Natl Acad Sci USA 114(11):E2106-E2115. Most endogenous higher eukaryotic nucleic acids, including mammalian nucleic acids, e.g., human nucleic acids, contain Cap1 or Cap2. Cap0 and other cap structures distinct from Cap1 and Cap2 can be immunogenic in mammals, e.g., humans, because they are recognized as "non-self" by components of the innate immune system, such as IFIT-1 and IFIT-5, and can lead to elevated levels of cytokines, including type I interferons. Components of the innate immune system, such as IFIT-1 and IFIT-5, may also compete with eIF4E to bind to nucleic acids with caps other than Cap1 or Cap2, potentially inhibiting nucleic acid translation.
[0185] A cap can also be included co-transcriptionally. For example, ARCA (anti-reverse cap analog; Thermo Fisher Scientific catalog number AM8045) is a cap analog containing 7-methylguanine 3'-methoxy-5'-triphosphate linked to the 5' position of a guanine ribonucleotide and can be incorporated into transcripts in vitro at initiation. ARCA results in a Cap0 cap or Cap0-like cap in which the 2' position of the first cap-proximal nucleotide is hydroxyl. See, for example, Stepinski et al. (2001) "Synthesis and properties of mRNAs containing the novel 'anti-reverse' cap analogs 7-methyl(3'-O-methyl)GpppG and 7-methyl(3'deoxy)GpppG," RNA 7:1486-1495. The structure of ARCA is shown below. [ka]
[0186] CleanCap™ AG (m7G(5')ppp(5')(2'OMeA)pG; TriLink Biotechnologies catalog number N-7113) or CleanCap™ GG (m7G(5')ppp(5')(2'OMeG)pG; TriLink Biotechnologies catalog number N-7133) can be used to co-transcriptionally provide the Cap1 structure. 3'-O-methylated versions of CleanCap™ AG and CleanCap™ GG are also available from TriLink Biotechnologies as catalog numbers N-7413 and N-7433, respectively. The structure of CleanCap™ AG is shown below. CleanCap™ structures may be referred to herein using the last three digits of the above catalog number (e.g., "CleanCap™ 113" for TriLink Biotechnologies catalog number N-7113). [ka]
[0187] Alternatively, a cap can be added to RNA post-transcriptionally. For example, Vaccinia capping enzyme is commercially available (New England Biolabs, catalog number M2080S), which possesses RNA triphosphatase and guanylyltransferase activities provided by its D1 subunit and guanine methyltransferase activity provided by its D12 subunit. Thus, in the presence of S-adenosylmethionine and GTP, it can add 7-methylguanine to RNA to give Cap0. See, e.g., Guo, P. and Moss, B. (1990) Proc. Natl. Acad. Sci. USA 87, 4023-4027; Mao, X. and Shuman, S. (1994) J. Biol. Chem. 269, 24472-24479. For further discussion of caps and capping approaches, see, e.g., WO2017 / 053297 and Ishikawa et al., Nucl. Acids. Symp. Ser. (2009) No. 53, 129-130.
[0188] G. Guide RNA (gRNA) In some embodiments, the composition comprises at least one guide RNA (gRNA), and the method comprises delivering the at least one gRNA, wherein the gRNA directs the editor to a desired genomic location. In some embodiments, the composition comprises an mRNA described herein and at least one gRNA. In some embodiments, the composition comprises a polypeptide described herein and at least one gRNA. In some embodiments, the gRNA is a single guide RNA (sgRNA). In some embodiments, the gRNA is a dual guide RNA (dgRNA).
[0189] The gRNAs disclosed herein can include a guide sequence that directs a polypeptide, including a cytidine deaminase (e.g., APOBEC3A deaminase) and an RNA-guided nickase, to a cytosine (C) located in any region of a gene (e.g., within the coding region of a gene) for conversion of the cytosine (C) to a thymine (T) (a "C-to-T conversion").
[0190] In some embodiments, the C-to-T conversion modifies a DNA sequence, e.g., a human gene sequence. In some embodiments, the C-to-T conversion modifies the coding sequence of a gene. In some embodiments, the C-to-T conversion generates a stop codon, e.g., a premature stop codon, within the coding region of a gene. In some embodiments, the C-to-T conversion removes a stop codon. In some embodiments, the C-to-T conversion modifies a regulatory sequence of a gene (e.g., a gene promoter or a gene repressor). In some embodiments, the C-to-T conversion modifies the splicing of a gene. In some embodiments, the C-to-T conversion corrects a genetic defect associated with a disease or disorder.
[0191] In some embodiments, a guide RNA (gRNA) comprises a guide sequence that directs a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase) and an RNA-guided nickase to a splice donor or acceptor site within a gene. In some embodiments, the splice donor or acceptor is a splice donor site. In some embodiments, the splice donor or acceptor site is a splice acceptor site.
[0192] In some embodiments, the guide RNA (gRNA) comprises a guide sequence that directs a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase) and an RNA-guided nickase to an acceptor splice site boundary. In some embodiments, the guide RNA (gRNA) comprises a guide sequence that directs a polypeptide comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase to a donor splice site boundary.
[0193] In some embodiments, a guide RNA (gRNA) comprises a guide sequence that directs a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase) and an RNA-guided nickase to make a single-stranded cleavage in a gene at a cleavage site 3' to the acceptor splice site boundary or 5' to the acceptor splice site boundary. In this and the following description, 3' and 5' refer to the direction of the cleaved strand.
[0194] In some embodiments, a guide RNA (gRNA) disclosed herein comprises a guide sequence that directs a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase) and an RNA-guided nickase to make a single-stranded break in a gene at a cleavage site that is 3' to the donor splice site boundary or 5' to the donor splice site boundary.
[0195] As used herein, a "splice site" refers to the three nucleotides that make up an acceptor splice site or donor splice site (defined below), or any other nucleotide known in the art to be part of a splice site. See, e.g., Burset et al., Nucleic Acids Research 28(21):4364-4375(2000) (describing canonical and non-canonical splice sites in mammalian genomes). The three nucleotides that make up an "acceptor splice site" are two conserved residues at the 3' end of the intron (e.g., AG in humans) and a boundary nucleotide (i.e., the first nucleotide of the exon 3' to AG). The three nucleotides that make up a "donor splice site" are two conserved residues at the 5' end of the intron (e.g., GT (in genes) or GU (in RNAs such as pre-mRNA) in humans) and a boundary nucleotide (i.e., the first nucleotide of the exon 5' to GT).
[0196] In some embodiments, a composition comprising at least one gRNA is provided in combination with a nucleic acid (e.g., mRNA) disclosed herein. In some embodiments, one or more gRNAs are provided as separate molecules from the nucleic acid (e.g., mRNA) disclosed herein. In some embodiments, the gRNA is provided as part of a nucleic acid disclosed herein, for example, as part of a UTR.
[0197] In some embodiments, a composition is provided comprising a polypeptide comprising a cytidine deaminase and an RNA-guided nickase, and a gRNA. In some embodiments, a ribonucleoprotein complex (RNP) is provided, the RNP comprising a polypeptide comprising a cytidine deaminase and an RNA-guided nickase, and a gRNA. In some embodiments, the polypeptide does not comprise UGI.
[0198] The gRNA comprises a guide sequence that targets a specific gene or gene sequence. In some embodiments, the gRNA is a Cas nickase guide. In some embodiments, the gRNA is a class 2 Cas nickase guide. In further embodiments, the gRNA is a Cpf1 or Cas9 guide. In some embodiments, the gRNA is an Nme nickase guide. In some embodiments, the Nme nickase is an Nme1, Nme2, or Nme3 nickase. In some embodiments, the gRNA comprises a 5' guide sequence of RNA that forms two or more hairpin or stem-loop structures. CRISPR / Cas gRNA structures are known in the art and vary depending on their cognate Cas nuclease. Generally, a gRNA used with any particular Cas9 or Nme nickase described herein must function with that nickase. For example, if a polypeptide disclosed herein comprises a SpyCas9 nickase, the provided gRNA is a SpyCas9 guide RNA (as described herein). Where the polypeptide disclosed herein comprises an NmeCas9 nickase, the guide RNA is an NmeCas9 guide RNA (as described herein).
[0199] In some embodiments, the gRNA comprises a guide sequence that directs an RNA-guided nickase (e.g., Cas9 nickase) to a target DNA sequence within a target locus, such as a target gene. Targets and exemplary target sequences for each gene are exemplified herein and include, but are not limited to, the targets and guide sequences disclosed in WO2017185054 (against the trinucleotide repeat of transcription factor 4 (TCF4)); WO2018119182 A1 (targeting SERPINA1); WO2019 / 067872 (targeting transthyretin (TTR)); and WO2020 / 028327 A1 (targeting hydroxyacid oxidase 1 (HAO1)), the contents of each of which are incorporated herein by reference in their entirety. Those of skill in the art will be familiar with suitable guide sequences for targeting other genes or loci of interest.
[0200] The gRNA may comprise a crRNA that includes 17, 18, 19, 20, 21, 22, 23, 24, or 25 consecutive nucleotides of a guide sequence. The gRNA may further comprise a trRNA. In each composition and method embodiment described herein, the crRNA and trRNA may be associated as a single RNA (sgRNA) or may be on separate RNAs (dgRNA). In the context of an sgRNA, the crRNA and trRNA components may be covalently linked, for example, via a phosphodiester bond or other covalent bond.
[0201] In each of the composition, use, and method embodiments described herein, the gRNA may comprise two RNA molecules as a "dual guide RNA" or "dgRNA." The dgRNA comprises a first RNA molecule comprising a crRNA that includes a guide sequence, and a second RNA molecule comprising a trRNA. The first and second RNA molecules may not be covalently linked, but may form an RNA duplex through base pairing between portions of the crRNA and the trRNA.
[0202] In each of the embodiments of the compositions, uses, and methods described herein, the gRNA may comprise a single RNA molecule, referred to as a "single guide RNA" or "sgRNA." The sgRNA may comprise a crRNA (or a portion thereof) comprising a guide sequence covalently linked to a trRNA, for example. The sgRNA may comprise 17, 18, 19, 20, 21, 22, 23, 24, or 25 consecutive nucleotides of the guide sequence. In some embodiments, the crRNA and trRNA are covalently linked via a linker. In some embodiments, the sgRNA forms a stem-loop structure through base pairing between portions of the crRNA and the trRNA. In some embodiments, the crRNA and trRNA are covalently linked via one or more bonds that are not phosphodiester bonds.
[0203] In some embodiments, the trRNA may comprise all or part of a trRNA sequence derived from a naturally occurring CRISPR / Cas system. In some embodiments, the trRNA comprises a truncated or modified wild-type trRNA. The length of the trRNA varies depending on the CRISPR / Cas system used. In some embodiments, the trRNA comprises or consists of 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, or more than 100 nucleotides. In some embodiments, the trRNA may comprise a specific secondary structure, such as, for example, one or more hairpin or stem-loop structures, or one or more bulges.
[0204] The gRNAs provided herein can be useful for recognizing (e.g., hybridizing to) target sequences within a gene. In some embodiments, the selection of one or more gRNAs is determined based on the target sequence within a gene. In some embodiments, a gRNA that is complementary to or has complementarity to a target sequence within a target locus is used to direct polypeptides comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase to a specific location within the locus. The target locus can be recognized and nicked by the Cas nickase-containing gRNA.
[0205] In some embodiments, the guide sequence is at least 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, or 90% identical to the target sequence. In some embodiments, the target sequence may be complementary to the guide sequence of a gRNA. In some embodiments, the degree of complementarity or identity between the guide sequence of a gRNA and its corresponding target sequence may be about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the target sequence and the guide sequence of a gRNA may be 100% complementary or identical. In other embodiments, the target sequence and the guide sequence of a gRNA may contain at least one mismatch. For example, the target sequence and the guide sequence of the gRNA may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 mismatches, in which case the total length of the target sequence is at least about 17, 18, 19, 20, or more base pairs. In some embodiments, the target sequence and the guide sequence of the gRNA may contain 1, 2, 3, 4, 5, or 6 mismatches, in which case the guide sequence is 20 nucleotides.
[0206] The gRNA may include a guide sequence that ligates to additional nucleotides to form the crRNA, for example, having the following exemplary nucleotide sequence following the guide sequence at its 3' end: GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 139).
[0207] In the case of an sgRNA, the guide sequence may be linked to additional nucleotides to form the sgRNA, for example, having the following exemplary nucleotide sequence following the 3' end of the guide sequence: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 140) from 5' to 3'.
[0208] In some embodiments, the sgRNA comprises the modification pattern set forth below in SEQ ID NO: 141, where N is any natural or unnatural nucleotide, and the entire plurality of Ns constitutes a guide sequence described herein, and the modified sgRNA comprises the following sequence: mN*mN*mN*NNNNNNNNNNNNNNNNNNNGUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmCmAmCmCmGmAmGmUmCmGmUmGmCmU*mU*mU*mU (SEQ ID NO: 141), where "N" can be any natural or unnatural nucleotide. For example, SEQ ID NO: 141 is encompassed herein, where the plurality of Ns is replaced with any of the guide sequences disclosed herein. The modification remains as set forth in SEQ ID NO: 141, even when the plurality of Ns is replaced with a guide nucleotide. That is, even though "N's" are substituted by guide nucleotides, the first three nucleotides are 2'OMe modified and there are phosphorothioate linkages between the first and second nucleotides, the second and third nucleotides, and the third and fourth nucleotides.
[0209] Figure 23A shows an exemplary sgRNA (SEQ ID NO: 141; methylation not shown) in a possible secondary structure, with labels designating individual nucleotides of the conserved regions of the sgRNA, including the lower stem, bulge, upper stem, nexus (the nucleotides of which may be referred to as N1 through N18, respectively, in the 5' to 3' direction), and hairpin region (including hairpin 1 and hairpin 2 regions). The nucleotides between hairpin 1 and hairpin 2 are labeled n. A guide region may be present on the sgRNA, and is represented in this figure as "(N)x" before the conserved region of the sgRNA. In some embodiments, the sgRNA may further include one or more nucleotides between the lower stem and bulge region, between the bulge and upper stem region, between the upper stem and nexus, or between the nexus and hairpin 1 region, or between the hairpin 1 region and hairpin 2 region.
[0210] In some embodiments, the conserved portion of the sgRNA is a conserved region of spyCas9 or a spyCas9 equivalent. In some embodiments, the conserved portion of the sgRNA is not derived from S. pyogenes Cas9, such as Staphylococcus aureus Cas9 ("saCas9"). Further description of exemplary sgRNA regions is provided in WO2019 / 237069, published December 12, 2019, the entire contents of which are incorporated herein by reference.
[0211] The SpyCas9 gRNA may comprise an internal linker. In some embodiments, the internal linker may have a bridge length of about 3-30, optionally 12-21 atoms, where the linker replaces at least two nucleotides of the gRNA. In some embodiments, the internal linker has a bridge length of about 6-18, optionally 6-12 atoms, where the linker replaces at least two nucleotides of the gRNA. In some embodiments, the internal linker comprises at least two ethylene glycol subunits covalently linked to one another, hi some embodiments, the internal linker comprises a PEG-linker.
[0212] In some embodiments, the internal linker comprises a PEG-linker having 1 to 10 ethylene glycol units. In some embodiments, the internal linker comprises a PEG-linker having 3 to 6 ethylene glycol units. In some embodiments, the internal linker comprises a PEG-linker having 3 ethylene glycol units. In some embodiments, the internal linker comprises a PEG-linker having 6 ethylene glycol units.
[0213] In some embodiments, the conserved portion of the spyCas9 guide RNA comprises a repeat-antirepeat region, a hairpin 1 region, and a hairpin 2 region, and further comprises at least one of the following: 1) a first internal linker that replaces at least two nucleotides in the upper stem region of the repeat-antirepeat region of the sgRNA; 2) a second internal linker that replaces one or two nucleotides in hairpin 1 of the sgRNA; or 3) a third internal linker that replaces at least two nucleotides in hairpin 2 of the sgRNA.
[0214] Exemplary locations of linkers in spyCas9 guide RNAs are shown below: [ka] (In the sequence, N is the nucleotide that encodes the guide sequence).
[0215] As used herein, "Linker 1" or "L1" refers to an internal linker having a bridge length of about 15 to 21 atoms. As used herein, "Linker 2" or "L2" refers to an internal linker having a bridge length of about 6 to 12 atoms.
[0216] In some embodiments, spyCas9 guide RNAs containing internal linkers can be chemically modified. Exemplary modifications include modification patterns of the following sequences: mA*mC*mG*CAAAUAUCAGUCCAGCGGUUUUAGAmGmCmUmA(L1)mUmAmGmCAAGUUAAAAUAAGGC(L2)GUCCGUUAUCAC(L1)GGGCACCGAGUCGG*mU*mG*mC (SEQ ID NO: 523).
[0217] In some embodiments, the gRNA comprises a 3' tail. In some embodiments, the 3' tail consists of nucleotides containing uracil or modified uracil. In some embodiments, the 3' terminal nucleotide is a modified nucleotide. In some embodiments, the 3' tail comprises a modification of any one or more nucleotides present in the 3' tail. In further embodiments, the modification of the 3' tail is one or more 2'-O-methyl (2'-OMe) modified nucleotides, and a phosphorothioate (PS) linkage between the nucleotides at the terminal nucleotide.
[0218] 1. Short single guide RNA (short sgRNA) In some embodiments, the sgRNAs provided herein are short single guide RNAs (short sgRNAs), e.g., short sgRNAs that include a conserved portion of an sgRNA, including a hairpin region, where the hairpin region lacks at least 5-10 nucleotides or 6-10 nucleotides. In some embodiments, the 5-10 nucleotides or 6-10 nucleotides are contiguous.
[0219] In some embodiments, the short sgRNA lacks at least nucleotides 54-58 (AAAAA) of the conserved portion of the spyCas9 sgRNA. In some embodiments, the short sgRNA is a non-spyCas9 sgRNA that lacks nucleotides corresponding to nucleotides 54-58 (AAAAA) of the conserved portion of spyCas9, as determined, for example, by pairwise or structural alignment.
[0220] Structural alignment is useful when molecules share a common, similar structure despite significant sequence differences. Structural alignment involves identifying corresponding residues between two (or more) sequences by (i) modeling the structure of a first sequence with the known structure of a second sequence, or (ii) comparing the structures of the first and second sequences, if both are known, and identifying the residue in the first sequence that is most similar to the residue of interest in the second sequence. Corresponding residues are identified in some algorithms based on minimizing the distance (e.g., which set of paired positions minimizes the root-mean-square deviation of the alignment) at a given position in the superimposed structures (e.g., the 1-position of a nucleobase or the 1' carbon of a pentose ring in a polynucleotide, or the alpha carbon in a polypeptide). When identifying positions in a non-spyCas9 gRNA that correspond to positions described for a spyCas9 gRNA, the spyCas9 gRNA can be the "second" sequence. If a non-spyCas9 gRNA of interest does not have a known structure available but is more closely related to another non-spyCas9 gRNA with a known structure, it may be most effective to model the non-spyCas9 gRNA of interest using the known structure of the closely related non-spyCas9 gRNA and then compare the model to the spyCas9 gRNA structure to identify the desired corresponding residues in the non-spyCas9 gRNA of interest. There is an extensive literature on protein structural modeling and alignment, with representative disclosures including those cited in US6859736; US8738343; and Aslam et al., Electronic Journal of Biotechnology 20 (2016) 9-13. For a discussion of modeling structures based on known related structure(s), see, for example, Bordoli et al., Nature Protocols 4 (2009) 1-13 and the references cited therein. For nucleic acid alignment, see also Figure 2(F) in Nishimasu et al., Cell 162(5):1113-1126 (2015).
[0221] In some embodiments, the short sgRNAs described herein comprise a conserved portion that includes a hairpin region, wherein the hairpin region lacks 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides. In some embodiments, the missing nucleotides are 5-10 missing nucleotides or 6-10 missing nucleotides. In some embodiments, the missing nucleotides are contiguous. In some embodiments, the missing nucleotides span at least a portion of Hairpin 1 and a portion of Hairpin 2. In some embodiments, the 5-10 missing nucleotides comprise or consist of nucleotides 54-58, 54-61, or 53-60 of SEQ ID NO:140.
[0222] In some embodiments, the short sgRNAs described herein further comprise a nexus region, wherein the nexus region lacks at least one nucleotide (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of the nexus region). In some embodiments, the short sgRNA lacks every nucleotide of the nexus region.
[0223] In some embodiments, the SpyCas9 short sgRNA described herein comprises: In some embodiments, the short sgRNA described herein comprises the sequence of SEQ ID NO: 520: mN*mN*mN*GUUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCACGAAAGGGCACCGAGUCGGmUmGmC*mU (SEQ ID NO: 520), where A, C, G, U, and N are adenine, cytosine, guanine, uracil, and any ribonucleotide, respectively, unless otherwise indicated, m indicates a 2'O-methyl modification, and * indicates a phosphorothioate internucleotide linkage.
[0224] In some embodiments, the gRNAs described herein are N. meningitidis Cas9 (NmeCas9) gRNAs that contain a conserved portion comprising the repeat / anti-repeat region, hairpin 1 region, and hairpin 2 region, where one or more of the repeat / anti-repeat region, hairpin 1 region, and hairpin 2 region are truncated. An exemplary wild-type NmeCas9 guide RNA is (N) 20-25 GUUGUAGCUCCCUUUCUCAUUUCGGAAACGAAAUGAGAACCGUUGCUACAAUAAGGCCGUCUGAAAAGAUGUGCCGCAACGCUCUGCCCCUUAAAGCUUCUGCUUUAAGGGGCAUCGUUUA (SEQ ID NO: 512). 20-25 represents 20 to 25, i.e., 20, 21, 22, 23, 24, or 25 consecutive Ns. A, C, G, and U represent nucleotides having adenine, cytosine, guanine, and uracil bases, respectively. In some embodiments, (N) 20-25 has a length of 24 nucleotides, N is any natural or unnatural nucleotide, and the entire set of Ns constitutes the guide sequence.
[0225] In some embodiments, the conserved portion of the NmeCas9 short gRNA is (a) a truncated repeat / anti-repeat region lacking 2 to 24 nucleotides, (i) one or more of nucleotides 37-48 and 53-64 are deleted, and optionally one or more of nucleotides 37-64 are substituted, relative to SEQ ID NO: 512; and (ii) a truncated repeat / anti-repeat region in which nucleotide 36 is connected to nucleotide 65 by at least two nucleotides; or b) a truncated Hairpin 1 region lacking 2 to 10, optionally 2 to 8, nucleotides; (i) one or more of nucleotides 82-86 and 91-95 are deleted, and optionally one or more of positions 82-96 are substituted, relative to SEQ ID NO: 512; and (ii) a truncated hairpin 1 region in which nucleotide 81 is connected to nucleotide 96 by at least four nucleotides; or (c) a truncated hairpin 2 region lacking 2 to 18, optionally 2 to 16, nucleotides; (i) one or more of nucleotides 113-121 and 126-134 are deleted, and optionally one or more of nucleotides 113-134 are substituted, relative to SEQ ID NO: 512; and (ii) comprises a truncated hairpin 2 region in which nucleotide 112 is connected to nucleotide 135 by at least four nucleotides; One or both of nucleotides 144-145 are optionally deleted relative to SEQ ID NO: 512, and at least 10 nucleotides are modified nucleotides.
[0226] In some embodiments, the NmeCas9 short gRNA comprises the following sequence in 5' to 3' direction: (N) 20-25 GUUGUAGCUCCCUGAAACCGUUGCUACAAUAAGGCCGUCGAAAGAUGUGCCGCAACGCUCUGCCUUCUGGCAUCGUU (SEQ ID NO: 513); (N) 20-25 GUUGUAGCUCCCUGAAACCGUUGCUACAAUAAGGCCGUCGAAAGAUGUGCCGCAACGCUCUGCCUUCUGGCAUCGUUUAUU (SEQ ID NO: 514); (N) 20-25 GUUGUAGCUCCCUGGAAACCCGUUGCUACAAUAAGGCCGUCGAAAGA UGUGCCGCAACGCUCUGCCUUCUGGCAUCGUUUAUU (SEQ ID NO: 515), Contains one of the following:
[0227] In some embodiments, at least 10 nucleotides of the conserved portion of the NmeCas9 short sgRNA are modified nucleotides.
[0228] In some embodiments, the NmeCas9 short sgRNA has the following sequence in 5' to 3' direction: GUUGmUmAmGmCUCCCmUmGmAmAmAmCmCGUUmGmCUAmCAAU*AAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGCmAmAmCmGCUCUmGmCCmUmUmCmUGmGCmAmUC*mG*mU*mU (SEQ ID NO: 516) or GUUGmUmAmGmCUCCCmUmGmAmAmAmCmCGUUmGmCUAmCAAU*AAGmGmCCmGmUmCmGmAmAmAmGmAmUGUGCmCGmCAAmCGCUCUmGmCCmUmUmCmUGGCAUCG*mU*mU (SEQ ID NO: 517), The conserved region includes one of:
[0229] The truncated NmeCas9 gRNA may include an internal linker as disclosed herein.
[0230] As used herein, the term "internal linker" refers to a non-nucleotide segment that connects two nucleotides within a guide RNA. When a gRNA contains a spacer region, the internal linker is located outside the spacer region (e.g., in the scaffold or conserved region of the gRNA). In the case of a V-shaped guide, it is understood that the last hairpin is the only hairpin in the structure, i.e., the repeat-anti-repeat region. In some embodiments, the internal linker comprises a PEG-linker as disclosed herein.
[0231] Exemplary locations of the linker are shown below: [ka] As used herein, (L1) refers to an internal linker having a bridge length of about 15 to 21 atoms.
[0232] In some embodiments, the truncated NmeCas9 guide RNAs containing internal linkers can be chemically modified. Exemplary modifications include modification patterns of the following sequences: mN*mN*mN*mNmNNmNNNmNNmNNNNmNNNNmNNNNmNNNmGUUGmUmAmGmCUCCCmUmUmC(L1)mGmAmCmCGUUmGmCUAmCAAU*AAGmGmCCmGmUmC(L1)mGmAmUGUGCmCGmCAAmCGCUCUmGmCC(L1)GGCAUCG*mU*mU (SEQ ID NO: 519).
[0233] 2. Qualification In some embodiments, gRNAs (e.g., sgRNAs, short sgRNAs, dgRNAs, or crRNAs) are modified. The terms "modified" or "modification" in the context of gRNAs described herein include the modifications described above, such as (a) terminal modifications, e.g., 5'- or 3'-terminal modifications (including 5'- or 3'-protected terminal modifications), (b) nucleobase (or "base") modifications (including base substitutions or removals), (c) sugar modifications (including modifications at the 2', 3', and / or 4' positions), (d) internucleoside linkage modifications, and (e) backbone modifications (which may include modifications or replacements of phosphodiester linkages and / or ribose sugars). Modifications of a nucleotide at a given position include modifications or replacements of the phosphodiester linkage immediately 3' to the sugar of the nucleotide. Thus, for example, a nucleic acid containing a phosphorothioate between the first and second sugars from the 5'-end is considered to contain a modification at position 1. The term "modified gRNA" generally refers to a gRNA having modifications to the chemical structure of one or more of the phosphodiester linkages or backbone moieties, including the base, sugar, and nucleotide phosphate, all as detailed and exemplified herein (see, e.g., the modification patterns shown in SEQ ID NOs: 142-145, 181-185, and 191-203).
[0234] Further description and exemplary patterns of modifications are provided in Table 1 of WO2019 / 237069, published December 12, 2019, the entire contents of which are incorporated herein by reference.
[0235] In some embodiments, the gRNA comprises a modification at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more YA sites. In some embodiments, the pyrimidine at the YA site comprises a modification (including a modification that alters the internucleoside linkage immediately 3' to the sugar of the pyrimidine). In some embodiments, the adenine at the YA site comprises a modification (including a modification that alters the internucleoside linkage immediately 3' to the sugar of the adenine). In some embodiments, the pyrimidine and adenine at the YA site comprise a modification, e.g., a sugar, base, or internucleoside linkage modification. The YA modification can be any of the types of modifications described herein. In some embodiments, the YA modification comprises one or more of phosphorothioate, 2'-OMe, or 2'-fluoro. In some embodiments, the YA modification comprises a pyrimidine modification comprising one or more of phosphorothioate, 2'-OMe, 2'-H, inosine, or 2'-fluoro. In some embodiments, the YA modification comprises a bicyclic ribose analog (e.g., LNA, BNA, or ENA) within an RNA duplex region comprising one or more YA sites. In some embodiments, the YA modification comprises a bicyclic ribose analog (e.g., LNA, BNA, or ENA) within an RNA duplex region comprising a YA site, wherein the YA modification is distal to the YA site.
[0236] In some embodiments, the guide sequence (or guide region) of a gRNA comprises one, two, three, four, five, or more YA sites ("guide region YA sites") that may comprise a YA modification. In some embodiments, one or more YA sites located 5 to the 5' end, 6 to the 7 to the 8 to the 9 to the 10 from the 5' end of the 5' terminus ("5 to the 5" and the like refers to positions 5 to the 3' end of the guide region, i.e., to the 3'-most nucleotide within the guide region) comprise a YA modification. Modified guide region YA sites comprise a YA modification.
[0237] In some embodiments, the modified guide region YA site is within 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, or 9 nucleotides of the 3'-terminal nucleotide of the guide region. For example, if the modified guide region YA site is within 10 nucleotides of the 3'-terminal nucleotide of the guide region and the guide region is 20 nucleotides long, the modified nucleotide of the modified guide region YA site is located anywhere from 11 to 20. In some embodiments, the modified guide region YA site is at or after the 4th, 5th, 6th, 7th, 8th, 9th, 10th, or 11th nucleotide from the 5' end of the 5' terminus.
[0238] In some embodiments, the modified guide region YA site is other than a 5' end modification. For example, an sgRNA can include a 5' end modification described herein and can further include a modified guide region YA site. Alternatively, an sgRNA can include an unmodified 5' end and a modified guide region YA site. Alternatively, a short sgRNA can include a modified 5' end and an unmodified guide region YA site.
[0239] In some embodiments, the modified guide region YA site contains a modification that is absent in at least one nucleotide located 5' to the guide region YA site. For example, if nucleotides 1-3 contain phosphorothioate, nucleotide 4 contains only 2'-OMe modifications, and nucleotide 5 is a pyrimidine at the YA site and contains phosphorothioate, the modified guide region YA site contains a modification (phosphorothioate) that is absent in at least one nucleotide located 5' to the guide region YA site (nucleotide 4). In another example, if nucleotides 1-3 contain phosphorothioate, and nucleotide 4 is a pyrimidine at the YA site and contains 2'-OMe, the modified guide region YA site contains a modification (2'-OMe) that is absent in at least one nucleotide located 5' to the guide region YA site (any of nucleotides 1-3). This condition is always met, even when an unmodified nucleotide is located 5' to the modified guide region YA site.
[0240] In some embodiments, the modified guide region YA site comprises the modifications described for YA sites above. The guide region of a gRNA can be modified according to any embodiment, including a modified guide region, described herein.
[0241] Conserved region YA sites 1-10 are illustrated in Figure 23B. In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 conserved region YA sites comprise a modification. In some embodiments, conserved region YA sites 1, 8, or 1 and 8 comprise a YA modification. In some embodiments, conserved region YA sites 1, 2, 3, 4, and 10 comprise a YA modification. In some embodiments, YA sites 2, 3, 4, 8, and 10 comprise a YA modification. In some embodiments, conserved region YA sites 1, 2, 3, and 10 comprise a YA modification. In some embodiments, YA sites 2, 3, 8, and 10 comprise a YA modification. In some embodiments, YA sites 1, 2, 3, 4, 8, and 10 comprise a YA modification. In some embodiments, YA sites 1, 2, 3, 4, 8, and 10 comprise a YA modification. In some embodiments, 1, 2, 3, 4, 5, 6, 7, or 8 additional conserved region YA sites comprise a YA modification.
[0242] In some embodiments, the modified conserved region YA site comprises the modifications described for the YA site above. Any of the embodiments described elsewhere in this disclosure may be combined, to the extent feasible, with any of the foregoing embodiments.
[0243] In some embodiments, the 5' and / or 3' terminal regions of the gRNA are modified.
[0244] In some embodiments, the terminal (i.e., last) 1, 2, 3, 4, 5, 6, or 7 nucleotides of the 3'-terminal region are modified. Throughout, this modification may be referred to as a "3'-end modification." In some embodiments, the terminal (i.e., last) 1, 2, 3, 4, 5, 6, or 7 nucleotides of the 3'-terminal region comprise two or more modifications. In some embodiments, the 3'-end modification comprises, or further comprises, any one or more of the following: modified nucleotides selected from 2'-O-methyl (2'-O-Me) modified nucleotides, 2'-O-(2-methoxyethyl) (2'-O-moe) modified nucleotides, 2'-fluoro (2'-F) modified nucleotides, internucleotide phosphorothioate (PS) linkages, inverted abasic modified nucleotides, or combinations thereof. In some embodiments, the 3'-end modification comprises, or further comprises, modifications of 1, 2, 3, 4, 5, 6, or 7 nucleotides at the 3'-end of the gRNA. In some embodiments, the 3'-end modification comprises or further comprises one PS linkage, which is located between the last and penultimate nucleotide. In some embodiments, the 3'-end modification comprises or further comprises two PS linkages between the last three nucleotides. In some embodiments, the 3'-end modification comprises or further comprises four PS linkages between the last four nucleotides. In some embodiments, the 3'-end modification comprises or further comprises a PS linkage between any one or more of the last 2, 3, 4, 5, 6, or 7 nucleotides. In some embodiments, the gRNA comprising a 3'-end modification comprises or further comprises a 3' tail, which comprises a modification of any one or more nucleotides present in the 3' tail. In some embodiments, the 3' tail is fully modified. In some embodiments, the 3' tail comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, or 1-10 nucleotides, optionally any one or more of which nucleotides are modified.In some embodiments, a gRNA is provided that comprises a 3' end modification, wherein the 3' end modification comprises a 3' end modification set forth in any one of SEQ ID NOs: 141-145. In some embodiments, a gRNA is provided that comprises a 3' protected end modification. In some embodiments, the 3' tail comprises 1 to about 20 nucleotides, 1 to about 15 nucleotides, 1 to about 10 nucleotides, 1 to about 5 nucleotides, 1 to about 4 nucleotides, 1 to about 3 nucleotides, and 1 to about 2 nucleotides. In some embodiments, the gRNA does not comprise a 3' tail.
[0245] In some embodiments, the 5' terminal region is modified, e.g., the first 1, 2, 3, 4, 5, 6, or 7 nucleotides of the gRNA are modified. Throughout, this modification may be referred to as a "5' end modification." In some embodiments, the first 1, 2, 3, 4, 5, 6, or 7 nucleotides of the 5' terminal region comprise two or more modifications. In some embodiments, at least one of the terminal (i.e., first) 1, 2, 3, 4, 5, 6, or 7 nucleotides at the 5' end is modified. In some embodiments, both the 5' and 3' terminal regions (e.g., ends) of the gRNA are modified. In some embodiments, only the 5' terminal region of the gRNA is modified. In some embodiments, only the 3' terminal region (plus or minus the 3' tail) of the conserved portion of the gRNA is modified. In some embodiments, the gRNA comprises modifications in 1, 2, 3, 4, 5, 6, or 7 of the first 7 nucleotides in the 5' terminal region of the gRNA. In some embodiments, the gRNA comprises modifications at 1, 2, 3, 4, 5, 6, or 7 of the 7 terminal nucleotides in the 3'-terminal region. In some embodiments, 2, 3, or 4 of the first 4 nucleotides in the 5'-terminal region and / or 2, 3, or 4 of the last 4 nucleotides in the 3'-terminal region are modified. In some embodiments, 2, 3, or 4 of the first 4 nucleotides in the 5'-terminal region are linked with phosphorothioate (PS) linkages. In some embodiments, modifications to the 5'- and / or 3'-terminus comprise 2'-O-methyl (2'-O-Me) or 2'-O-(2-methoxyethyl) (2'-O-moe) modifications. In some embodiments, modifications comprise 2'-fluoro (2'-F) modifications to nucleotides. In some embodiments, modifications comprise internucleotide phosphorothioate (PS) linkages. In some embodiments, modifications comprise inverted abasic nucleotides. In some embodiments, modifications comprise protected end modifications. In some embodiments, the modifications include two or more modifications selected from a blocked end modification, 2'-O-Me, 2'-O-moe, 2'-fluoro (2'-F), a phosphorothioate internucleotide (PS) linkage, and an inverted abasic nucleotide.In some embodiments, equivalent modifications are encompassed. In some embodiments, a gRNA is provided that comprises a 5'-end modification, wherein the 5'-end modification comprises any one of SEQ ID NOs: 141-145.
[0246] In some embodiments, gRNAs are provided that include 5' and 3' end modifications, hi some embodiments, the gRNAs include modified nucleotides at neither the 5' nor the 3' end.
[0247] In some embodiments, an sgRNA is provided that includes an upper stem modification, the upper stem modification comprising a modification to any one or more of US1 through US12 in the upper stem region. In some embodiments, an sgRNA is provided that includes an upper stem modification comprising a modification to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or all 12 nucleotides in the upper stem region. In some embodiments, an sgRNA is provided that includes an upper stem modification comprising 1, 2, 3, 4, or 5 YA modifications at YA sites. In some embodiments, the upper stem modification comprises 2'-OMe-modified nucleotides, 2'-O-moe-modified nucleotides, 2'-F-modified nucleotides, and / or combinations thereof. Other modifications described herein, such as 5'-end modifications and / or 3'-end modifications, may be combined with the upper stem modification.
[0248] In some embodiments, the sgRNA comprises a modification in the hairpin region. In some embodiments, the hairpin region modification comprises at least one modified nucleotide selected from 2'-O-methyl (2'-OMe) modified nucleotides, 2'-fluoro (2'-F) modified nucleotides, and / or combinations thereof. In some embodiments, the hairpin region modification is present in the hairpin 1 region. In some embodiments, the hairpin region modification is present in the hairpin 2 region. In some embodiments, the hairpin modification comprises one, two, or three YA modifications at the YA sites. In some embodiments, the hairpin modification comprises at least one, two, three, four, five, or six YA modifications. Other modifications described herein, such as upper stem modifications, 5'-end modifications, and / or 3'-end modifications, may be combined with modifications in the hairpin region.
[0249] In some embodiments, the gRNA comprises a replaced, optionally shortened, Hairpin 1 region, in which at least one of the following nucleotide pairs: H1-1 and H1-12, H1-2 and H1-11, H1-3 and H1-10, and / or H1-4 and H1-9, is replaced with a Watson-Crick pairing nucleotide. "Watson-Crick pairing nucleotides" includes any pair capable of forming Watson-Crick base pairing, including AT, AU, TA, UA, CG, and GC pairs, as well as pairs containing modified versions of any of the foregoing nucleotides with the same base-pairing preference. In some embodiments, the Hairpin 1 region lacks any one or two of H1-5 through H1-8. In some embodiments, the Hairpin 1 region lacks one, two, or three of the following nucleotide pairs: H1-1 and H1-12, H1-2 and H1-11, H1-3 and H1-10, and / or H1-4 and H1-9. In some embodiments, the Hairpin 1 region lacks 1 to 8 nucleotides of the Hairpin 1 region. In any of the foregoing embodiments, the missing nucleotides can be such that one or more nucleotide pairs (H1-1 and H1-12, H1-2 and H1-11, H1-3 and H1-10, and / or H1-4 and H1-9) are replaced with Watson-Crick paired nucleotides to form base pairs within the gRNA.
[0250] In some embodiments, the gRNA further comprises an upper stem region lacking at least one nucleotide, e.g., any of the truncated upper stem regions shown in Table 7 of U.S. Application No. 62 / 946,905 (the contents of which are incorporated by reference herein in their entirety) or described elsewhere herein, and can be combined with any shortened or substituted hairpin 1 region described herein.
[0251] In some embodiments, the gRNAs described herein further comprise a nexus region lacking at least one nucleotide.
[0252] 3. Chemical Modification of gRNA In some embodiments, the gRNA is chemically modified. A gRNA that includes one or more modified nucleosides or nucleotides is referred to as a "modified" gRNA or a "chemically modified" gRNA, denoting the presence of one or more non-natural and / or naturally occurring components or structures used in place of or in addition to the standard A, G, C, and U residues. Modified nucleosides and nucleotides can include one or more of the following: (i) an alteration, e.g., substitution, of one or both of the non-linked phosphate oxygens in the phosphodiester backbone linkages and / or one or more linking phosphate oxygens (exemplary backbone modifications); (ii) an alteration, e.g., substitution, of a component of the ribose sugar, e.g., the 2' hydroxyl on the ribose sugar (exemplary sugar modifications); (iii) a wholesale replacement of the phosphate moiety with a "dephospho" linker (exemplary backbone modifications); (iv) a modification or substitution of a naturally occurring nucleobase, including with a non-standard nucleobase (exemplary base modifications); (v) a substitution or modification of the ribose-phosphate backbone (exemplary backbone modifications); (vi) a modification of the 3' or 5' end of the oligonucleotide, e.g., removal, modification, or substitution of the terminal phosphate group, or conjugation of a moiety, cap, or linker (such 3' or 5' cap modifications can include sugar and / or backbone modifications); and (vii) a modification or substitution of the sugar (exemplary sugar modifications).
[0253] The chemical modifications listed above can be combined to provide modified gRNAs comprising nucleosides and nucleotides (collectively "residues") that can have two, three, four, or more modifications. For example, modified residues can have modified sugars and modified nucleobases. In some embodiments, all bases of the gRNA are modified, e.g., all bases have modified phosphate groups, e.g., phosphorothioate groups. In certain embodiments, all, or substantially all, of the phosphate groups of a gRNA molecule are replaced with phosphorothioate groups. In some embodiments, modified gRNAs comprise at least one modified residue at or near the 5' end of the RNA. In some embodiments, modified gRNAs comprise at least one modified residue at or near the 3' end of the RNA.
[0254] In some embodiments, the gRNA comprises one, two, three, or more modified residues. In some embodiments, at least 5% (e.g., at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100%) of the positions in the modified gRNA are modified nucleosides or nucleotides.
[0255] In some embodiments of backbone modifications, the phosphate group of the modified residue can be modified by replacing one or more oxygen atoms with different substituents. Furthermore, modified residues, such as those present in modified nucleic acids, can include the total replacement of unmodified phosphate moieties with modified phosphate groups as described herein. In some embodiments, backbone modifications of the phosphate backbone can include modifications that result in either an uncharged linker or a charged linker with an asymmetric charge distribution.
[0256] Examples of modified phosphate groups include phosphorothioates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphoramidates, alkyl or aryl phosphonates, and phosphotriesters.
[0257] Nucleic acid-mimicking scaffolds can also be constructed, in which the phosphate linker and ribose sugar are replaced with nuclease-resistant nucleoside or nucleotide substitutes. Such modifications can include backbone modifications and sugar modifications. In some embodiments, the nucleobases can be linked by alternative backbones. Examples include, but are not limited to, morpholino, cyclobutyl, pyrrolidine, and peptide nucleic acid (PNA) nucleoside substitutes.
[0258] Modified nucleosides and nucleotides can include one or more modifications to the sugar group, i.e., sugar modifications. For example, the 2' hydroxyl group (OH) can be modified, e.g., replaced with a number of different "oxy" or "deoxy" substituents. In some embodiments, modifying the 2' hydroxyl group can improve the stability of the nucleic acid by preventing the hydroxyl from being deprotonated and forming a 2'-alkoxide ion. Examples of 2' hydroxyl group modifications include alkoxy or aryloxy (OR, where "R" can be, e.g., alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or a sugar); polyethylene glycol (PEG), O(CHCHO) n Examples of 2' hydroxyl group modifications include CH2CH2OR (where R can be, for example, H or optionally substituted alkyl, and n can be an integer from 0 to 20). In some embodiments, the 2' hydroxyl group modification can be 2'-O-Me. In some embodiments, the 2' hydroxyl group modification can be a 2'-fluoro modification, which replaces the 2' hydroxyl group with fluoride. In some embodiments, the 2' hydroxyl group modification can include "locked" nucleic acids (LNAs), where the 2' hydroxyl is replaced with, for example, a C1-6 alkylene or C 1-6 The 4' carbon of the same ribose sugar can be connected by a heteroalkylene bridge, and exemplary bridges can include methylene, propylene, ether, or amino bridges. In some embodiments, the 2' hydroxyl group modification can include an "unlocked" nucleic acid (UNA), in which the ribose ring lacks a C2'-C3' bond. In some embodiments, the 2' hydroxyl group modification can include a methoxyethyl group (MOE), (OCH2CH2OCH3, e.g., a PEG derivative).
[0259] "Deoxy" 2' modifications include hydrogen (i.e., deoxyribose sugars, such as those in partial overhanging portions of dsRNA); halo (e.g., bromo, chloro, fluoro, or iodo); amino (amino can be, for example, NH; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or amino acid); NH(CHCHNH) n These may include CH2CH2-amino (wherein amino may be, for example, as described herein), -NHC(O)R (wherein R may be, for example, alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or sugar), cyano; mercapto; alkyl-thio-alkyl; thioalkoxy; and alkyl, cycloalkyl, aryl, alkenyl, and alkynyl (which may optionally be substituted, for example, with amino, as described herein).
[0260] Sugar modifications can include sugar groups that can contain one or more carbons with a stereochemical configuration opposite that of the corresponding carbon in ribose. Thus, modified nucleic acids can include nucleotides containing, for example, arabinose as the sugar. Modified nucleic acids can also include abasic sugars. These abasic sugars can be further modified at one or more of the constituent sugar atoms. Modified nucleic acids can also include one or more sugars that are in the L-configuration, such as L-nucleosides.
[0261] The modified nucleosides and modified nucleotides described herein that can be incorporated into modified nucleic acids can contain modified bases, also referred to as nucleobases. Examples of nucleobases include, but are not limited to, adenine (A), guanine (G), cytosine (C), and uracil (U). These nucleobases can be modified or completely replaced to provide modified residues that can be incorporated into modified nucleic acids. The nucleobases of the nucleotides can be independently selected from purines, pyrimidines, purine analogs, or pyrimidine analogs. In some embodiments, the nucleobases can include, for example, naturally occurring and synthetic base derivatives.
[0262] In embodiments using dual guide RNAs, each of the crRNA and tracr RNA can include modifications. Such modifications can be present at one or both ends of the crRNA and / or tracr RNA. In embodiments including an sgRNA, one or more residues at one or both ends of the sgRNA can be chemically modified, or the entire sgRNA can be chemically modified. Certain embodiments include 5'-end modifications. Certain embodiments include 3'-end modifications. In certain embodiments, one or more or all of the nucleotides in the single-stranded overhang of the gRNA molecule are deoxynucleotides.
[0263] In some embodiments, the gRNAs disclosed herein comprise one of the modification patterns disclosed in WO2018 / 107028 A1, published June 14, 2018, the contents of which are incorporated herein by reference in their entirety.
[0264] The terms "mA," "mC," "mU," or "mG" may be used to indicate a nucleotide modified with 2'-O-Me. The terms "fA," "fC," "fU," or "fG" may be used to indicate a nucleotide substituted with 2'-F. "*" may be used to represent a PS modification. The terms A*, C*, U*, or G* may be used to indicate a nucleotide that is linked to the next (e.g., 3') nucleotide by a PS bond. The terms "mA*," "mC*," "mU*," or "mG*" may be used to indicate a nucleotide substituted with 2'-O-Me and linked to the next (e.g., 3') nucleotide by a PS bond.
[0265] H. Lipids; Formulations; Delivery Disclosed herein are various embodiments that utilize lipid-nucleic acid assembly compositions comprising the nucleic acid(s) or composition(s) described herein. In some embodiments, the lipid-nucleic acid assembly composition comprises a nucleic acid (e.g., mRNA) comprising an open reading frame encoding a first polypeptide comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase. In some embodiments, the lipid-nucleic acid assembly composition comprises a first nucleic acid comprising an open reading frame encoding a first polypeptide comprising a cytidine deaminase (e.g., A3A) and an RNA-guided nickase, and a second nucleic acid encoding UGI.
[0266] As used herein, "lipid-nucleic acid assembly composition" refers to lipid-based delivery compositions, including lipid nanoparticles (LNPs) and lipoplexes. LNPs refer to lipid nanoparticles of less than 100 nM. LNPs are formed by precisely mixing a lipid component (e.g., in ethanol) with an aqueous nucleic acid component, and LNPs are uniform in size. Lipoplexes are particles formed by bulk mixing of a lipid component and a nucleic acid component, and are approximately 100 nm to 1 micron in size. In certain embodiments, the lipid-nucleic acid assembly is an LNP. As used herein, a "lipid-nucleic acid assembly" comprises multiple (i.e., two or more) lipid molecules physically associated with each other through intermolecular forces. Lipid-nucleic acid assemblies may comprise bioavailable lipids with pKa values less than 7.5 or less than 7. Lipid-nucleic acid assemblies are formed by mixing an aqueous nucleic acid-containing solution with an organic solvent-based lipid solution, such as 100% ethanol. Suitable solutions or solvents may include or contain water, PBS, Tris buffer, NaCl, citrate buffer, ethanol, chloroform, diethyl ether, cyclohexane, tetrahydrofuran, methanol, isopropanol. Pharmaceutically acceptable buffers may optionally be included in pharmaceutical formulations containing lipid nucleic acid assemblies, for example, for ex vivo treatment. In some embodiments, the aqueous solution contains RNA, such as mRNA or gRNA. In some embodiments, the aqueous solution contains mRNA encoding an RNA-guided DNA-binding agent, such as Cas9.
[0267] As used herein, lipid nanoparticles (LNPs) refer to particles comprising a plurality (i.e., two or more) of lipid molecules physically associated with each other by intermolecular forces. LNPs can be, for example, microspheres (including unilamellar and multilamellar vesicles, e.g., "liposomes," i.e., lamellar phase lipid bilayers that are substantially spherical in some embodiments and, in more specific embodiments, can include an aqueous core containing, for example, a substantial portion of RNA molecules), the dispersed phase in an emulsion, a micelle, or the internal phase in a suspension. Emulsions, micelles, and suspensions can be compositions suitable for local and / or topical delivery. See, e.g., WO2017173054A1 (the contents of which are incorporated herein by reference in their entirety). Any LNP known to those skilled in the art to be capable of delivering nucleotides to a subject can be utilized with the guide RNAs described herein, as well as nucleic acids encoding RNA-guided nickase and cytidine deaminase.
[0268] In some embodiments, the aqueous solution comprises a nucleic acid encoding a polypeptide comprising A3A and an RNA-guided nickase. Pharmaceutical formulations comprising lipid-nucleic acid assembly compositions may optionally include a pharmaceutically acceptable buffer.
[0269] In some embodiments, lipid nucleic acid assembly compositions comprise "amine lipids" (sometimes referred to herein or elsewhere as "ionizable lipids" or "biodegradable lipids"), along with optional "helper lipids," "neutral lipids," and stealth lipids, such as PEG lipids. In some embodiments, the amine lipids or ionizable lipids are cationic depending on the pH.
[0270] 1. Amine lipids In some embodiments, the lipid nucleic acid assembly composition comprises an "amine lipid," which is an ionizable lipid such as, for example, lipid A or its equivalent, eg, an acetal analog of lipid A.
[0271] In some embodiments, the amine lipid is lipid A, which is (9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate, also referred to as 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl(9Z,12Z)-octadeca-9,12-dienoate. Lipid A can be represented as follows:
[0272] [ka]
[0273] Lipid A can be synthesized according to WO2015 / 095340 (e.g., pages 84-86). In some embodiments, the amine lipid is an equivalent of lipid A.
[0274] In some embodiments, the amine lipid is an analog of lipid A. In some embodiments, the lipid A analog is an acetal analog of lipid A. In certain lipid nucleic acid assembly compositions, the acetal analog is a C4-C12 acetal analog. In some embodiments, the acetal analog is a C5-C12 acetal analog. In additional embodiments, the acetal analog is a C5-C10 acetal analog. In further embodiments, the acetal analog is selected from C4, C5, C6, C7, C9, C10, C11, and C12 acetal analogs.
[0275] Amine lipids and other "biodegradable lipids" suitable for use in the lipid nucleic acid assemblies described herein are biodegradable in vivo or ex vivo. Amine lipids have low toxicity (e.g., tolerated without side effects in animal models at doses of 10 mg / kg or greater). In some embodiments, lipid nucleic acid assemblies comprising amine lipids include those in which at least 75% of the amine lipid is cleared from plasma or engineered cells within 8, 10, 12, 24, or 48 hours, or within 3, 4, 5, 6, 7, or 10 days. In some embodiments, lipid nucleic acid assemblies comprising amine lipids include those in which at least 50% of the nucleic acid, e.g., mRNA or gRNA, is cleared from plasma within 8, 10, 12, 24, or 48 hours, or within 3, 4, 5, 6, 7, or 10 days. In some embodiments, lipid nucleic acid assemblies comprising amine lipids include those in which at least 50% of the lipid nucleic acid assemblies are cleared from plasma within 8, 10, 12, 24, or 48 hours, or within 3, 4, 5, 6, 7, or 10 days, as determined by measuring lipids (e.g., amine lipids), nucleic acids, e.g., RNA / mRNA, or other components. In some embodiments, the lipid-encapsulated lipid, RNA, or nucleic acid components of the lipid nucleic acid assemblies are measured relative to the free lipid, RNA, or nucleic acid components.
[0276] Biodegradable lipids include, for example, those of WO / 2020 / 219876, WO / 2020 / 118041, WO / 2020 / 072605, WO / 2019 / 067992, WO / 2017 / 173054, WO2015 / 095340, and WO2014 / 136086, and LNPs include the LNP compositions described therein, which lipids and compositions are incorporated herein by reference.
[0277] Lipid clearance can be measured as described in the literature. See Maier, MA, et al. Biodegradable Lipids Enabling Rapidly Eliminated Lipid Nanoparticles for Systemic Delivery of RNAi Therapeutics. Mol. Ther. 2013, 21(8), 1570-78 ("Maier"). For example, in Maier, an LNP-siRNA system containing luciferase-targeting siRNA was administered to 6-8 week-old male C57Bl / 6 mice via an intravenous bolus injection at 0.3 mg / kg via the lateral tail vein. Blood, liver, and spleen samples were collected at 0.083, 0.25, 0.5, 1, 2, 4, 8, 24, 48, 96, and 168 hours post-administration. Mice were perfused with saline, followed by tissue collection, and blood samples were processed to obtain plasma. All samples were processed and analyzed by LC-MS. Furthermore, Maier describes a method for evaluating toxicity after administration of LNP-siRNA formulations. For example, luciferase-targeting siRNA was administered to male Sprague-Dawley rats at 0, 1, 3, 5, and 10 mg / kg (5 rats / group) via a single intravenous bolus injection in a 5 mL / kg dose volume. After 24 hours, approximately 1 mL of blood was obtained from the jugular vein of awake animals, and serum was isolated. Seventy-two hours after administration, all animals were euthanized for necropsy. Clinical signs, body weight, serum chemistry, organ weight, and histopathology were evaluated. Maier describes methods for evaluating siRNA-LNP formulations, which may be applied to assessing clearance, pharmacokinetics, and toxicity associated with the administration of lipid-nucleic acid assembly compositions of the present disclosure.
[0278] Suitable for the LNP delivery of nucleic acid known in the art is ionizable and bioavailable lipid.Lipid can be ionizable depending on the pH of the medium it is contained in.For example, in a slightly acidic medium, lipid such as amine lipid can be protonated and therefore carry positive charge.On the other hand, in a slightly basic medium such as blood, whose pH is about 7.35, lipid such as amine lipid can not be protonated and therefore carry no charge.
[0279] The ability of a lipid to carry a charge is related to its inherent pKa. In some embodiments, the amine lipids of the present disclosure can each independently have a pKa in the range of about 5.1 to about 7.4. In some embodiments, the bioavailable lipids of the present disclosure can each independently have a pKa in the range of about 5.1 to about 7.4, e.g., about 5.5 to about 6.6, about 5.6 to about 6.4, about 5.8 to about 6.2, or about 5.8 to about 6.5. For example, the amine lipids of the present disclosure can each independently have a pKa in the range of about 5.8 to about 6.5. Lipids with a pKa in the range of about 5.1 to about 7.4 are effective for delivering cargo in vivo, for example, to the liver. Furthermore, lipids with a pKa in the range of about 5.3 to about 6.4 have been found to be effective for delivering cargo in vivo, for example, to tumors. See, e.g., WO 2014 / 136086.
[0280] 2. Additional fats "Neutral lipids" suitable for use in the lipid-nucleic acid assembly compositions of the present disclosure include, for example, various neutral, uncharged, or zwitterionic lipids. Examples of neutral phospholipids suitable for use in the present disclosure include, but are not limited to, 5-heptadecylbenzene-1,3-diol (resorcinol), dipalmitoylphosphatidylcholine (DPPC), distearoylphosphatidylcholine (DSPC), phosphocholine (DOPC), dimyristoylphosphatidylcholine (DMPC), phosphatidylcholine (PLPC), 1,2-distearoyl-sn-glycero-3-phosphocholine (DAPC), phosphatidylethanolamine (PE), egg phosphatidylcholine (EPC), dilauryloylphosphatidylcholine (DLPC), dimyristoylphosphatidylcholine (DLPC), and phosphocholine (phosphocholine- ... (DMPC), 1-myristoyl-2-palmitoylphosphatidylcholine (MPPC), 1-palmitoyl-2-myristoylphosphatidylcholine (PMPC), 1-palmitoyl-2-stearoylphosphatidylcholine (PSPC), 1,2-diarachidoyl-sn-glycero-3-phosphocholine (DBPC), 1-stearoyl-2-palmitoylphosphatidylcholine (SPPC), 1,2-dieicosenoyl-sn-glycero-3-phosphocholine (DEPC), palmitoyloleoylphosphatidylcholine (POPC), lysophosphatidylcholine, dioleoylphosphatidylethanolamine (DOPE), dilinoleoylphosphatidylcholine Examples of the neutral phospholipid include distearoylphosphatidylethanolamine (DSPE), dimyristoylphosphatidylethanolamine (DMPE), dipalmitoylphosphatidylethanolamine (DPPE), palmitoyloleoylphosphatidylethanolamine (POPE), lysophosphatidylethanolamine, and combinations thereof. In one embodiment, the neutral phospholipid may be selected from the group consisting of distearoylphosphatidylcholine (DSPC) and dimyristoylphosphatidylethanolamine (DMPE). In another embodiment, the neutral phospholipid may be distearoylphosphatidylcholine (DSPC).
[0281] "Helper lipids" include steroids, sterols, and alkylresorcinols. Helper lipids suitable for use in the present disclosure include, but are not limited to, cholesterol, 5-heptadecylresorcinol, and cholesterol hemisuccinate. In one embodiment, the helper lipid can be cholesterol. In one embodiment, the helper lipid can be cholesterol hemisuccinate.
[0282] A "stealth lipid" is a lipid that alters the length of time that nanoparticles can persist in vivo (e.g., in the blood). Stealth lipids can aid in the formulation process, for example, by reducing particle aggregation and controlling particle size. Stealth lipids used herein can modulate the pharmacokinetic properties of lipid nucleic acid assemblies or aid in the ex vivo stability of nanoparticles. Stealth lipids suitable for use in the lipid nucleic acid assembly compositions of the present disclosure include, but are not limited to, stealth lipids with a hydrophilic head group linked to the lipid moiety. Information regarding stealth lipids suitable for use in the lipid nucleic acid assembly compositions of the present disclosure, and the biochemistry of such lipids, can be found in Romberg et al., Pharmaceutical Research, Vol. 25, No. 1, 2008, pp. 55-71 and Hoekstra et al., Biochimica et Biophysica Acta 1660 (2004) 41-52. Additional suitable PEG lipids are disclosed, for example, in WO 2006 / 007712.
[0283] In one embodiment, the hydrophilic head group of the stealth lipid comprises a polymer moiety selected from PEG-based polymers. The stealth lipid may comprise a lipid moiety. In some embodiments, the stealth lipid is a PEG-lipid.
[0284] In one embodiment, the stealth lipid comprises a polymer moiety selected from PEG-based polymers (sometimes referred to as poly(ethylene oxide)), poly(oxazoline), poly(vinyl alcohol), poly(glycerol), poly(N-vinylpyrrolidone), polyamino acids, and poly[N-(2-hydroxypropyl)methacrylamide].
[0285] In one embodiment, the PEG lipid comprises a PEG-based polymer moiety (sometimes referred to as poly(ethylene oxide)).
[0286] The PEG-lipid further comprises a lipid moiety. In some embodiments, the lipid moiety may be derived from a diacylglycerol or diacylglycamide, including those comprising a dialkylglycerol or dialkylglycamide group having an alkyl chain length independently containing from about C4 to about C40 saturated or unsaturated carbon atoms, and the chain may contain one or more functional groups, such as amides or esters. In some embodiments, the alkyl chain length comprises from about C10 to C20. The dialkylglycerol or dialkylglycamide group may further comprise one or more substituted alkyl groups. The chain length may be symmetric or asymmetric.
[0287] Unless otherwise indicated, the term "PEG," as used herein, refers to any polyethylene glycol or other polyalkylene ether polymer. In one embodiment, PEG is a linear or branched polymer of optionally substituted ethylene glycol or ethylene oxide. In one embodiment, PEG is unsubstituted. In one embodiment, PEG is substituted, for example, with one or more alkyl, alkoxy, acyl, hydroxy, or aryl groups. In one embodiment, the term includes PEG copolymers, such as PEG-polyurethane or PEG-polypropylene (see, e.g., J. Milton Harris, Poly(ethylene glycol) chemistry: biotechnical and biomedical applications (1992)). In another embodiment, the term does not include PEG copolymers. In one embodiment, the PEG has a molecular weight of about 130 to about 50,000, in a subembodiment, about 150 to about 30,000, in a subembodiment, about 150 to about 20,000, in a subembodiment, about 150 to about 15,000, in a subembodiment, about 150 to about 10,000, in a subembodiment, about 150 to about 6,000, in a subembodiment, about 150 to about 5,000, in a subembodiment, about 150 to about 4,000, in a subembodiment, about 150 to about 3,000, in a subembodiment, about 300 to about 3,000, in a subembodiment, about 1,000 to about 3,000, and in a subembodiment, about 1,500 to about 2,500.
[0288] In some embodiments, the PEG (e.g., conjugated to a lipid moiety or lipid, such as a stealth lipid) is "PEG-2K," also known as "PEG2000," which has an average molecular weight of about 2,000 daltons. PEG-2K is represented herein by the following formula (I): where n is 45, meaning that the number-average degree of polymerization comprises about 45 subunits. [ka] However, other embodiments of PEG known in the art may be used, including, for example, those having a number average degree of polymerization of about 23 subunits (n=23) and / or 68 subunits (n=68). In some embodiments, n may range from about 30 to about 60. In some embodiments, n may range from about 35 to about 55. In some embodiments, n may range from about 40 to about 50. In some embodiments, n may range from about 42 to about 48. In some embodiments, n may be 45. In some embodiments, R may be selected from H, substituted alkyl, and unsubstituted alkyl. In some embodiments, R may be unsubstituted alkyl. In some embodiments, R may be methyl.
[0289] In any of the embodiments described herein, the PEG lipid may be PEG-dilaurylglycerol, PEG-dimyristoylglycerol (PEG-DMG) (Cat. No. GM-020, NOF, Tokyo, Japan), PEG-dipalmitoylglycerol, PEG-distearoylglycerol (PEG-DSPE) (Cat. No. DSPE-020CN, NOF, Tokyo, Japan), PEG-dilaurylglycamide, PEG-dimyristylglycamide, PEG-dipalmitoylglycamide, and PEG-distearoylglycerol. Ricamide, PEG-cholesterol (1-[8'-(cholest-5-ene-3[beta]-oxy)carboxamido-3',6'-dioxaotanyl]carbamoyl-[omega]-methyl-poly(ethylene glycol), PEG-DMB (3,4-ditetradecoxylbenzyl-[omega]-methyl-poly(ethylene glycol) ether), 1,2-dimyristoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DMG) (catalog no. 880150P, Avanti Polar Lipids, Alabaster, Alabama, USA), 1,2-distearoyl-sn-glycero-3-phosphoethanolamine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DSPE) (catalog no. 880120C, Avanti Polar Lipids, Alabaster, Alabama, USA), 1,2-distearoyl-sn-glycerol, methoxypolyethylene glycol (PEG2k-DSG; GS-020, NOF Tokyo, Japan), poly(ethylene glycol)-2000-dimethacrylate (PEG2k-DMA), and 1,2-distearyloxypropyl-3-amine-N-[methoxy(polyethylene glycol)-2000] (PEG2k-DSA). In one embodiment, the PEG lipid can be PEG2k-DMG. In some embodiments, the PEG lipid can be PEG2k-DSG. In one embodiment, the PEG lipid can be PEG2k-DSPE. In one embodiment, the PEG lipid can be PEG2k-DMA. In one embodiment, the PEG lipid can be PEG2k-C-DMA.In one embodiment, the PEG lipid can be compound S027, which is disclosed in WO2016 / 010840 (paragraphs
[0240] to
[0244] ). In one embodiment, the PEG lipid can be PEG2k-DSA. In one embodiment, the PEG lipid can be PEG2k-C11. In some embodiments, the PEG lipid can be PEG2k-C14. In some embodiments, the PEG lipid can be PEG2k-C16. In some embodiments, the PEG lipid can be PEG2k-C18.
[0290] 3. Formulation The lipid-nucleic acid assembly may contain (i) a biodegradable lipid, (ii) an optional neutral lipid, (iii) a helper lipid, and (iv) a stealth lipid, such as a PEG lipid. The lipid-nucleic acid assembly may contain a biodegradable lipid and one or more of a neutral lipid, a helper lipid, and a stealth lipid, such as a PEG lipid.
[0291] The lipid nucleic acid assembly may contain (i) an amine lipid for encapsulation and endosomal escape, (ii) a neutral lipid for stabilization, (iii) a helper lipid, also for stabilization, and (iv) a stealth lipid, such as a PEG lipid. The lipid nucleic acid assembly may contain an amine lipid and one or more of a neutral lipid, a helper lipid for stabilization, and a stealth lipid, such as a PEG lipid.
[0292] The mRNA necessary to achieve the functional effects described herein can be delivered to cells in one or more lipid-nucleic acid assembly compositions. For example, one lipid-nucleic acid assembly composition can be formulated for delivery and include an mRNA encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, and additional mRNAs encoding, for example, one or more UGIs and one or more gRNAs. Alternatively, the mRNA encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, the mRNA encoding one or more UGIs, and the mRNA encoding one or more gRNAs can be formulated into separate lipid-nucleic acid assembly compositions. Thus, one or more lipid-nucleic acid assembly compositions can be delivered to cells in vitro or in vivo.
[0293] In some embodiments, there is provided a method of modifying a target gene in a cell, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs; and optionally lipid nanoparticles, to a cell.
[0294] In some embodiments, portions (a) and (b) are present in separate lipid nucleic acid assembly compositions. In some embodiments, portions (a) and (b) are present in the same lipid nucleic acid assembly composition. In some embodiments, portions (a) and (c) are present in separate lipid nucleic acid assembly compositions. In some embodiments, portions (a) and (c) are present in the same lipid nucleic acid assembly composition. In some embodiments, portions (b) and (c) are present in separate lipid nucleic acid assembly compositions. In some embodiments, portions (a) and (c) are present in the same lipid nucleic acid assembly composition, and portion (b) is present in a separate lipid nucleic acid assembly composition. In some embodiments, portions (a), (b), and (c) are each present in a separate lipid nucleic acid assembly composition. In some embodiments, portions (a), (b), and (c) are present in the same lipid nucleic acid assembly composition. In some embodiments, one or more guide RNAs are each present in a separate lipid nucleic acid assembly composition.
[0295] In some embodiments, the methods further comprise delivering one or more guide RNAs in one or more lipid nucleic acid assembly compositions separate from the lipid nucleic acid assembly composition comprising A3A and UGI.
[0296] In some embodiments, at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 lipid-nucleic acid assembly compositions are delivered to a cell. In some embodiments, at least one lipid-nucleic acid assembly composition comprises a lipid nanoparticle (LNP). In some embodiments, all lipid-nucleic acid assembly compositions comprise LNP. In some embodiments, at least one lipid-nucleic acid assembly composition is a lipoplex composition.
[0297] In some embodiments, lipid nucleic acid assembly compositions, e.g., LNP compositions, comprise an mRNA encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, as described herein. In some embodiments, lipid nucleic acid assembly compositions, e.g., LNP compositions, comprise an mRNA encoding a polypeptide comprising an APOBEC3A deaminase (A3A) and an RNA-guided nickase, and a gRNA.
[0298] In some embodiments, the lipid nucleic acid assembly composition comprises a first lipid nucleic acid assembly composition comprising an mRNA encoding a first polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase. In some embodiments, the lipid nucleic acid assembly composition further comprises a second lipid nucleic acid assembly composition comprising a gRNA.
[0299] In some embodiments, the lipid nucleic acid assembly composition comprises a first composition comprising an mRNA encoding a first polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, and one or more second mRNAs encoding a uracil glycosylase inhibitor (UGI). In some embodiments, the lipid nucleic acid assembly composition further comprises one or more gRNAs.
[0300] In some embodiments, the lipid nucleic acid assembly composition further comprises a second lipid nucleic acid assembly composition comprising a gRNA. In some embodiments, the lipid nucleic acid assembly composition comprises first and second lipid nucleic acid assembly compositions, wherein the first composition comprises mRNA encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, and the second composition comprises one or more mRNAs encoding a uracil glycosylase inhibitor (UGI). In some embodiments, the first lipid nucleic acid assembly composition or the second lipid nucleic acid assembly composition further comprises one or more gRNAs. In some embodiments, the lipid nucleic acid assembly composition further comprises a third lipid nucleic acid assembly composition comprising one or more gRNAs.
[0301] In some embodiments, the lipid nucleic acid assembly composition comprises a first lipid nucleic acid assembly composition comprising an mRNA comprising an open reading frame encoding a first polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, and a second lipid nucleic acid assembly composition comprising one or more guide RNAs (gRNAs).
[0302] In some embodiments, the lipid nucleic acid assembly composition comprises a first composition comprising a first mRNA comprising a first open reading frame encoding a first polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, and a second mRNA comprising one or more second open reading frames encoding a uracil glycosylase inhibitor (UGI). In some embodiments, the lipid nucleic acid assembly composition further comprises one or more gRNAs. In some embodiments, the lipid nucleic acid assembly composition further comprises a second lipid nucleic acid assembly composition comprising one or more gRNAs.
[0303] In some embodiments, the lipid nucleic acid assembly composition comprises first and second lipid nucleic acid assembly compositions, wherein the first composition comprises a first mRNA comprising a first open reading frame encoding a first polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, and the second composition comprises a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI). In some embodiments, the first lipid nucleic acid assembly composition or the second lipid nucleic acid assembly composition further comprises a gRNA. In some embodiments, the lipid nucleic acid assembly composition further comprises a third lipid nucleic acid assembly composition comprising a gRNA.
[0304] In some embodiments, the lipid nucleic acid assembly composition comprises a first composition comprising a first mRNA comprising a first open reading frame encoding a first polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, and the second composition comprises a uracil glycosylase inhibitor (UGI). In some embodiments, the first lipid nucleic acid assembly composition or the second lipid nucleic acid assembly composition further comprises a gRNA. In some embodiments, the lipid nucleic acid assembly composition further comprises a third lipid nucleic acid assembly composition comprising a gRNA.
[0305] In some embodiments, the lipid nucleic acid assembly composition comprises a first and a second lipid nucleic acid assembly composition, wherein the first composition comprises a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, and the second composition comprises a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI). In some embodiments, the first lipid nucleic acid assembly composition or the second lipid nucleic acid assembly composition further comprises a gRNA. In some embodiments, the lipid nucleic acid assembly composition further comprises a third lipid nucleic acid assembly composition comprising a gRNA.
[0306] In certain embodiments, the lipid nucleic acid assembly composition can include mRNA, optionally gRNA, an amine lipid, a helper lipid, a neutral lipid, and a stealth lipid. In certain lipid nucleic acid assembly compositions, the helper lipid is cholesterol. In other lipid nucleic acid assembly compositions, the neutral lipid is DSPC. In additional embodiments, the stealth lipid is PEG2k-DMG or PEG2k-C11. In certain embodiments, the lipid nucleic acid assembly composition includes lipid A or an equivalent of lipid A, a helper lipid, a neutral lipid, and a stealth lipid. In certain compositions, the amine lipid is lipid A. In certain compositions, the amine lipid is lipid A or an acetal analog thereof, the helper lipid is cholesterol, the neutral lipid is DSPC, and the stealth lipid is PEG2k-DMG.
[0307] Embodiments of the present disclosure also provide lipid-nucleic acid assembly compositions, which are expressed according to the molar ratio of the positively charged amine group (N) of the amine lipid to the negatively charged phosphate group (P) of the nucleic acid to be encapsulated. This can be mathematically represented by the formula N / P. In some embodiments, the lipid-nucleic acid assembly composition can include a lipid component comprising an amine lipid, a helper lipid, a neutral lipid, and a PEG lipid, and a nucleic acid component, with an N / P ratio of about 3 to 10. In some embodiments, the lipid-nucleic acid assembly composition can include a lipid component comprising an amine lipid, a helper lipid, and a PEG lipid, and a nucleic acid component, with an N / P ratio of about 3 to 10. In some embodiments, the lipid-nucleic acid assembly composition can include a lipid component comprising an amine lipid, a helper lipid, a neutral lipid, and a helper lipid, and an RNA component, with an N / P ratio of about 3 to 10. In some embodiments, the lipid-nucleic acid assembly composition can include a lipid component comprising an amine lipid, a helper lipid, and a PEG lipid, and an RNA component, with an N / P ratio of about 3 to 10. In one embodiment, the N / P ratio can be about 5-7. In one embodiment, the N / P ratio can be about 3-7. In one embodiment, the N / P ratio can be about 4.5-8. In one embodiment, the N / P ratio can be about 6. In one embodiment, the N / P ratio can be 6±1. In one embodiment, the N / P ratio can be 6±0.5. In some embodiments, the N / P ratio will be ±30%, ±25%, ±20%, ±15%, ±10%, ±5%, or ±2.5% of the target N / P ratio. In certain embodiments, the lot-to-lot variation of the LNP will be less than 15%, less than 10%, or less than 5%.
[0308] In some embodiments, the lipid-nucleic acid assembly composition is formed by mixing an aqueous RNA solution with an organic solvent-based lipid solution, such as 100% ethanol. Suitable solutions or solvents may include or contain water, PBS, Tris buffer, NaCl, citrate buffer, ethanol, chloroform, diethyl ether, cyclohexane, tetrahydrofuran, methanol, or isopropanol. For example, a pharmaceutically acceptable buffer for in vivo administration of the lipid-nucleic acid assembly composition may be used. In certain embodiments, a buffer is used to maintain the pH of a composition comprising the lipid-nucleic acid assembly composition at pH 6.5 or above. In certain embodiments, a buffer is used to maintain the pH of a composition comprising LNPs at pH 7.0 or above. In certain embodiments, the composition has a pH in the range of about 7.2 to about 7.7. In additional embodiments, the composition has a pH in the range of about 7.3 to about 7.7, or about 7.4 to about 7.6. In further embodiments, the composition has a pH of about 7.2, 7.3, 7.4, 7.5, 7.6, or 7.7. The pH of the composition can be measured using a micro pH probe. In certain embodiments, a cryoprotectant is included in the composition. Non-limiting examples of cryoprotectants include sucrose, trehalose, glycerol, DMSO, and ethylene glycol. Exemplary compositions can include up to 10% cryoprotectant, such as sucrose. In certain embodiments, lipid nucleic acid assembly compositions can include about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10% cryoprotectant. In certain embodiments, lipid nucleic acid assembly compositions can include about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10% sucrose. In some embodiments, lipid nucleic acid assembly compositions can include a buffer. In some embodiments, the buffer can include phosphate buffer (PBS), Tris buffer, citrate buffer, and mixtures thereof. In certain exemplary embodiments, the buffer includes NaCl. In certain embodiments, NaCl is omitted. Exemplary amounts of NaCl can range from about 20 mM to about 45 mM. Exemplary amounts of NaCl can range from about 40 mM to about 50 mM, hi some embodiments, the amount of NaCl is about 45 mM.In some embodiments, the buffer is a Tris buffer. Exemplary amounts of Tris can range from about 20 mM to about 60 mM. Exemplary amounts of Tris can range from about 40 mM to about 60 mM. In some embodiments, the amount of Tris is about 50 mM. In some embodiments, the buffer comprises NaCl and Tris. A specific exemplary embodiment of a lipid-nucleic acid assembly composition contains 5% sucrose and 45 mM NaCl in a Tris buffer. In other exemplary embodiments, the composition contains about 5% w / v sucrose, about 45 mM NaCl, and about 50 mM Tris at pH 7.5. The amounts of salt, buffer, and cryoprotectant may be varied to maintain the osmolality of the entire formulation. For example, the final osmolality may be maintained below 450 mOsm / L. In further embodiments, the osmolality is between 350 and 250 mOsm / L. Certain embodiments have a final osmolality of 300 + / - 20 mOsm / L.
[0309] In some embodiments, microfluidic mixing, T-mixing, or cross-mixing is used. In certain aspects, the flow rate, junction size, junction geometry, junction shape, tubing diameter, solution, and / or nucleic acid and lipid concentrations may be varied. The lipid-nucleic acid assembly composition may be concentrated or purified, for example, via dialysis, tangential flow filtration, or chromatography. The lipid-nucleic acid assembly composition may be stored, for example, as a suspension, emulsion, or lyophilized powder. In some embodiments, the lipid-nucleic acid assembly composition is stored at 2-8°C, and in certain aspects, the LNP composition is stored at room temperature. In additional embodiments, the lipid-nucleic acid assembly composition is stored frozen, for example, at -20°C or -80°C. In other embodiments, the lipid-nucleic acid assembly composition is stored at a temperature ranging from about 0°C to about -80°C. The frozen lipid-nucleic acid assembly composition may be thawed, for example, on ice, at room temperature, or at 25°C, prior to use.
[0310] The lipid nucleic acid assembly composition can be, for example, a microsphere, the dispersed phase in an emulsion, a micelle, or the internal phase in a suspension.
[0311] Furthermore, in some embodiments, lipid nucleic acid assembly compositions are biodegradable in that they do not accumulate to cytotoxic levels in vivo at therapeutically effective doses. In some embodiments, lipid nucleic acid assembly compositions do not elicit an innate immune response that results in substantial side effects at therapeutic dose levels. In some embodiments, lipid nucleic acid assembly compositions provided herein do not elicit toxicity at therapeutic dose levels.
[0312] The LNPs disclosed herein can have a size (e.g., Z-average diameter) of about 1 to about 150 nm. In some embodiments, the LNPs are about 10 to about 200 nm in size. In some embodiments, the LNPs are about 50 to about 100 nm in size. In some embodiments, the LNPs are about 60 to about 100 nm in size. In some embodiments, the LNPs are about 75 to about 100 nm in size. In some embodiments, the LNP composition comprises a population of LNPs having an average diameter of about 20 to 100 nm. In some embodiments, the LNP composition comprises a population of LNPs having an average diameter of about 50 to 100 nm. In some embodiments, the LNP composition comprises a population of LNPs having an average diameter of about 60 to 100 nm. In some embodiments, the LNP composition comprises a population of LNPs having an average diameter of about 75 to 100 nm. Unless otherwise indicated, all sizes referred to herein are the average size (diameter) of fully formed nanoparticles as measured by dynamic light scattering on a Malvern Zetasizer. Nanoparticle samples are diluted in phosphate-buffered saline (PBS) to achieve a count rate of approximately 200-400 kcps. Data are presented as a weighted average of intensity measurements (Z-average diameter).
[0313] In some embodiments, LNPs are formed with an average encapsulation efficiency ranging from about 50% to about 100%. In some embodiments, LNPs are formed with an average encapsulation efficiency ranging from about 50% to about 70%. In some embodiments, LNPs are formed with an average encapsulation efficiency ranging from about 70% to about 90%. In some embodiments, LNPs are formed with an average encapsulation efficiency ranging from about 90% to about 100%. In some embodiments, LNPs are formed with an average encapsulation efficiency ranging from about 75% to about 95%.
[0314] In some embodiments, LNPs are formed with an average molecular weight ranging from about 1.00E+05 g / mol to about 1.00E+10 g / mol. In some embodiments, LNPs are formed with an average molecular weight ranging from about 5.00E+05 g / mol to about 7.00E+07 g / mol. In some embodiments, LNPs are formed with an average molecular weight ranging from about 1.00E+06 g / mol to about 1.00E+10 g / mol. In some embodiments, LNPs are formed with an average molecular weight ranging from about 1.00E+07 g / mol to about 1.00E+09 g / mol. In some embodiments, LNPs are formed with an average molecular weight ranging from about 5.00E+06 g / mol to about 5.00E+09 g / mol.
[0315] In some embodiments, the polydispersity (Mw / Mn; the ratio of weight-average molar mass (Mw) to number-average molar mass (Mn)) may range from about 1.000 to about 2.000. In some embodiments, Mw / Mn may range from about 1.00 to about 1.500. In some embodiments, Mw / Mn may range from about 1.020 to about 1.400. In some embodiments, Mw / Mn may range from about 1.010 to about 1.100. In some embodiments, Mw / Mn may range from about 1.100 to about 1.350.
[0316] Dynamic light scattering ("DLS") can be used to characterize the polydispersity index ("pdi") and size of the LNPs of the present disclosure. DLS measures the scattering of light resulting from exposing a sample to a light source. The PDI determined from DLS measurements represents the distribution of particle sizes (near the mean particle size) within an aggregate, with a completely uniform aggregate having a PDI of zero. In some embodiments, the pdi can range from 0.005 to 0.75. In some embodiments, the pdi can range from 0.01 to 0.5. In some embodiments, the pdi can range from 0.02 to 0.4. In some embodiments, the pdi can range from 0.03 to 0.35. In some embodiments, the pdi can range from 0.1 to 0.35. In some embodiments, the pdi can range from about 0 to about 0.4, e.g., from about 0 to about 0.35. In some embodiments, the pdi can range from about 0 to about 0.35, from about 0 to about 0.3, from about 0 to about 0.25, or from about 0 to about 0.2. In some embodiments, the pdi is less than about 0.08, 0.1, 0.15, 0.2, or 0.4.
[0317] In some embodiments, the LNPs disclosed herein are 1-250 nm in size. In some embodiments, the LNPs are 10-200 nm in size. In further embodiments, the LNPs are 20-150 nm in size. In some embodiments, the LNPs are 50-150 nm in size. In some embodiments, the LNPs are 50-100 nm in size. In some embodiments, the LNPs are 50-120 nm in size. In some embodiments, the LNPs are 75-150 nm in size. In some embodiments, the LNPs are 30-200 nm in size. Unless otherwise indicated, all sizes referred to herein are the average size (diameter) of fully formed nanoparticles as measured by dynamic light scattering on a Malvern Zetasizer. Nanoparticle samples are diluted in phosphate-buffered saline (PBS) to achieve a count rate of approximately 200-400 kcals. Data are presented as a weighted average of intensity measurements. In some embodiments, LNPs are formed with an average encapsulation efficiency ranging from 50% to 100%. In some embodiments, LNPs are formed with an average encapsulation efficiency ranging from 50% to 70%. In some embodiments, LNPs are formed with an average encapsulation efficiency ranging from 70% to 90%. In some embodiments, LNPs are formed with an average encapsulation efficiency ranging from 90% to 100%. In some embodiments, LNPs are formed with an average encapsulation efficiency ranging from 75% to 95%.
[0318] Electroporation is also a well-known means for delivery of cargo, and any electroporation methodology can be used for delivery of any one of the RNAs disclosed herein.
[0319] In some embodiments, the methods include methods for delivering a composition comprising an mRNA disclosed herein to an ex vivo cell, wherein the mRNA is encapsulated in LNPs. In some embodiments, the composition comprises an mRNA disclosed herein and one or more additional RNAs encapsulated in LNPs.
[0320] In some embodiments, the lipid nucleic acid assembly composition comprises lipid components, the lipid components comprising an amine lipid, a neutral lipid, a helper lipid, and a stealth lipid, and an N / P ratio of about 1-10.
[0321] In some examples, the lipid components comprise lipid A or its acetal analog, cholesterol, DSPC, and PEG-DMG, and the N / P ratio is about 1 to 10. In some embodiments, the lipid components comprise about 40 to 60 mol% amine lipids, about 5 to 15 mol% neutral lipids, and about 1.5 to 10 mol% PEG lipids, with the remainder of the lipid components being helper lipids, and the N / P ratio of the lipid-nucleic acid assembly composition is about 3 to 10. In some embodiments, the lipid components comprise about 50 to 60 mol% amine lipids, about 8 to 10 mol% neutral lipids, and about 2.5 to 4 mol% PEG lipids, with the remainder of the lipid components being helper lipids, and the N / P ratio of the lipid-nucleic acid assembly composition is about 3 to 8. In some examples, the lipid components comprise about 50-60 mol% amine lipid, about 5-15 mol% DSPC, and about 2.5-4 mol% PEG lipid, with the remainder of the lipid components being cholesterol, and the N / P ratio of the lipid nucleic acid assembly composition is about 3 to 8. In some examples, the lipid components comprise 48-53 mol% lipid A, about 8-10 mol% DSPC, and 1.5-10 mol% PEG lipid, with the remainder of the lipid components being cholesterol, and the N / P ratio of the lipid nucleic acid assembly composition is 3-8 ± 0.2.
[0322] In some embodiments, the lipid components comprise about 50-60 mol% amine lipids such as lipid A, about 8-10 mol% neutral lipids, and about 2.5-4 mol% stealth lipids (e.g., PEG lipids), with the remainder of the lipid components being helper lipids, and the N / P ratio of the lipid-nucleic acid assembly composition is about 6. In some embodiments, the lipid components comprise about 50-60 mol% amine lipids such as lipid A, about 27-39.5 mol% helper lipids, about 8-10 mol% neutral lipids, and about 2.5-4 mol% stealth lipids (e.g., PEG lipids), and the N / P ratio of the lipid-nucleic acid assembly composition is about 5-7 (e.g., about 6). In some embodiments, the lipid components comprise about 50-60 mol% amine lipids such as lipid A, about 5-15 mol% neutral lipids, and about 2.5-4 mol% stealth lipids (e.g., PEG lipids), with the remainder of the lipid components being helper lipids, and the N / P ratio of the lipid-nucleic acid assembly composition is about 3-10. In some embodiments, the lipid components comprise about 40-60 mol% amine lipids such as lipid A, about 5-15 mol% neutral lipids, and about 2.5-4 mol% stealth lipids (e.g., PEG lipids), with the remainder of the lipid components being helper lipids, and the N / P ratio of the lipid-nucleic acid assembly composition is about 6. In some embodiments, the lipid components comprise about 50-60 mol% amine lipids such as lipid A, about 5-15 mol% neutral lipids, and about 1.5-10 mol% stealth lipids (e.g., PEG lipids), with the remainder of the lipid components being helper lipids, and the N / P ratio of the lipid-nucleic acid assembly composition is about 6. In some embodiments, the lipid components comprise about 40-60 mol% amine lipids such as lipid A, about 0-10 mol% neutral lipids, and about 1.5-10 mol% stealth lipids (e.g., PEG lipids), with the remainder of the lipid components being helper lipids, and the N / P ratio of the lipid-nucleic acid assembly composition is about 3-10. In some embodiments, the lipid components comprise about 40-60 mol% amine lipids such as lipid A, less than about 1 mol% neutral lipids, and about 1.5-10 mol% stealth lipids (e.g., PEG lipids), with the remainder of the lipid components being helper lipids, and the N / P ratio of the lipid-nucleic acid assembly composition is about 3-10.In some embodiments, the lipid components comprise about 40-60 mol% amine lipids, such as lipid A, and about 1.5-10 mol% stealth lipids (e.g., PEG lipids), with the remainder of the lipid components being helper lipids, the N / P ratio of the lipid-nucleic acid assembly composition being about 3-10, and the lipid-nucleic acid assembly composition being essentially free of or free of neutral phospholipids. In some embodiments, the lipid components comprise about 50-60 mol% amine lipids, such as lipid A, about 8-10 mol% neutral lipids, and about 2.5-4 mol% stealth lipids (e.g., PEG lipids), with the remainder of the lipid components being helper lipids, and the N / P ratio of the lipid-nucleic acid assembly composition being about 3-7.
[0323] In some embodiments, the amine lipid is present at about 50 mol%. In some embodiments, the neutral lipid is present at about 9 mol%. In some embodiments, the stealth lipid is present at about 3 mol%. In some embodiments, the helper lipid is present at about 38 mol%.
[0324] In some embodiments, the lipid components comprise, consist essentially of, or consist of about 50 mol% amine lipids such as lipid A, about 9 mol% neutral lipids such as DSPC, about 3 mol% stealth lipids such as PEG lipids, e.g., PEG2k-DMG, with the remainder of the lipid components being helper lipids such as cholesterol, and the N / P ratio of the lipid-nucleic acid assembly composition is about 6. In some embodiments, the amine lipid is lipid A. In some embodiments, the neutral lipid is DSPC. In some embodiments, the stealth lipid is a PEG lipid. In some embodiments, the stealth lipid is PEG2k-DMG. In some embodiments, the helper lipid is cholesterol. In some embodiments, the lipid comprises lipid components comprising about 50 mol% lipid A, about 9 mol% DSPC, and about 3 mol% PEG2k-DMG, with the remainder of the lipid components being cholesterol, and the N / P ratio of the lipid-nucleic acid assembly composition is about 6.
[0325] In some embodiments, the lipid components comprise, consist essentially of, or consist of about 25-45 mol% amine lipids such as lipid A, about 10-30 mol% neutral lipids such as DSPC, about 1.5-3.5 mol% stealth lipids such as PEG lipids, e.g., PEG2k-DMG, and about 25-65 mol% helper lipids such as cholesterol, and the N / P ratio of the lipid-nucleic acid assembly composition is about 6. In some embodiments, the lipid components comprise, consist essentially of, or consist of about 35 mol% amine lipids such as lipid A, about 15 mol% neutral lipids such as DSPC, and about 2.5 mol% stealth lipids such as PEG lipids, e.g., PEG2k-DMG, with the remainder of the lipid components being helper lipids such as cholesterol, and the N / P ratio of the lipid-nucleic acid assembly composition is about 6. In some embodiments, the amine lipid is lipid A. In some embodiments, the neutral lipid is DSPC. In some embodiments, the stealth lipid is a PEG lipid. In some embodiments, the stealth lipid is PEG2k-DMG. In some embodiments, the helper lipid is cholesterol. In some embodiments, the lipid comprises a lipid component comprising about 35 mol% lipid A, about 15 mol% DSPC, and about 2.5 mol% PEG2k-DMG, with the remainder of the lipid component being cholesterol, and the N / P ratio of the lipid-nucleic acid assembly composition is about 6.
[0326] I. Exemplary Uses, Methods, and Treatments In some embodiments, the nucleic acids (e.g., mRNA), polypeptides, compositions, or lipid-nucleic acid assembly compositions disclosed herein are for use in genome editing, e.g., editing or modifying a target gene. In some embodiments, the nucleic acids (e.g., mRNA), polypeptides, compositions, or lipid-nucleic acid assembly compositions disclosed herein are for use in modifying a target gene, e.g., modifying its sequence or epigenetic state. In some embodiments, the nucleic acids (e.g., mRNA), polypeptides, compositions, or lipid-nucleic acid assembly compositions disclosed herein are for use in the manufacture of a medicament for genome editing or modifying a target gene.
[0327] In some embodiments, the use of a nucleic acid (e.g., mRNA), polypeptide, composition, or lipid nucleic acid assembly composition disclosed herein is provided for the preparation of a medicament for genome editing, e.g., editing a target gene. In some embodiments, the use of a nucleic acid (e.g., mRNA), polypeptide, composition, or lipid nucleic acid assembly composition is provided for the preparation of a medicament for modifying a target gene, e.g., modifying its sequence or epigenetic state. In some embodiments, the use of a nucleic acid (e.g., mRNA), polypeptide, composition, or lipid nucleic acid assembly composition disclosed herein is provided for the preparation of a medicament for causing a C-to-T conversion in a target gene.
[0328] In some embodiments, methods of genome editing or modifying a target gene are provided, comprising delivering to a cell an mRNA, composition, or lipid nanoparticle(s) described herein.
[0329] In some embodiments, the method generates a cytosine (C) to thymine (T) conversion in the target gene.
[0330] In some embodiments, the method results in at least 50% C-to-T conversions relative to the total edited target sequence. As used herein, "total edited target sequence" refers to the sum of each read with an indel or at least one conversion, where an indel may contain two or more nucleotides. Indels are calculated as the total number of sequencing reads with one or more base insertions or deletions within a 20-bp scoring region divided by the total number of sequencing reads containing the wild-type sequence. C-to-T conversions or C-to-A / G conversions were scored in a 40-bp region including 10 bp upstream and 10 bp downstream of the 20-bp sgRNA target sequence. Any sequencing method (e.g., NGS) that allows for the reading of sequences that deviate from the wild-type alignment can be used. In some embodiments, the method results in at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% C-to-T conversions relative to total editing of the target sequence.
[0331] In some embodiments, the ratio of C-to-T conversion to unintended editing is greater than 1:1. As used herein, "unintended editing" refers to any editing in the target region that is not a C-to-T conversion. In some embodiments, the ratio of C-to-T conversion to unintended editing is greater than 2:1, greater than 3:1, greater than 4:1, greater than 5:1, greater than 6:1, greater than 7:1, or greater than 8:1. In some embodiments, the ratio of C-to-T conversion to unintended editing is between 2:1 and 99:1. In some embodiments, the ratio of C-to-T conversion to unintended editing is 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, or 8:1.
[0332] In some embodiments, the method causes A3A to make a base edit corresponding to any one of positions −1 to 10 relative to the 5′ end of the guide sequence.
[0333] In some embodiments, the method causes A3A to perform a base edit at 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotide positions from the 5' end of the guide sequence.
[0334] In some embodiments, the nickase is SpyCas9 nickase, and the method involves a cytidine deaminase making a base edit at a cytidine that is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 nucleotides from the 5' end of the guide sequence.
[0335] In some embodiments, the nickase is an NmeCas9 nickase, and the method involves a cytidine deaminase making a base edit at a cytidine located 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides from the 5' end of the guide sequence.
[0336] In some embodiments, the composition comprises a first mRNA comprising a first open reading frame encoding a polypeptide comprising A3A and an RNA-guided nickase, a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), and a gRNA, wherein the first mRNA, second mRNA, and, if present, the gRNA are delivered in a ratio of about 6:2:3 (w:w:w). In some embodiments, the target gene is present in a subject, e.g., a mammal, e.g., a human.
[0337] In some embodiments, methods are provided for modifying a target gene, the methods comprising delivering to a cell a first mRNA comprising a first open reading frame encoding a first polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), the second mRNA being distinct from the first mRNA, and at least one guide RNA (gRNA).
[0338] In some embodiments, methods are provided for modifying a target gene in a cell, the methods comprising delivering to the cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs. In some embodiments, the one or more guide RNAs are each present in a separate lipid-nucleic acid assembly composition.
[0339] In some embodiments, methods are provided for modifying a target gene in a cell, the method comprising delivering to a cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, the method comprising one gRNA targeting a gene that reduces or eliminates MHC class I expression on the cell surface, and / or one gRNA targeting a gene that reduces or eliminates MHC class II expression on the cell surface, and / or one gRNA targeting a gene that reduces or eliminates endogenous TCR expression.
[0340] In some embodiments, methods are provided for modifying a target gene in a cell, the method comprising delivering to a cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, the method comprising at least two gRNAs selected from one gRNA targeting a gene that reduces or eliminates MHC class I expression on the cell surface, one gRNA targeting a gene that reduces or eliminates MHC class II expression on the cell surface, and one gRNA targeting a gene that reduces or eliminates endogenous TCR expression.
[0341] In some embodiments, methods are provided for modifying target genes in cells, comprising delivering to a cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises one gRNA targeting a gene that reduces or eliminates MHC class I expression on the cell surface, one gRNA targeting a gene that reduces or eliminates MHC class II expression on the cell surface, and one gRNA targeting a gene that reduces or eliminates endogenous TCR expression.
[0342] In some embodiments, methods are provided for modifying a target gene in a cell, the method comprising delivering to a cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, the method comprising one gRNA targeting a gene that reduces or eliminates HLA-A expression on the cell surface, and / or one gRNA targeting a gene that reduces or eliminates MHC class II expression on the cell surface, and / or one gRNA targeting a gene that reduces or eliminates endogenous TCR expression.
[0343] In some embodiments, methods are provided for modifying a target gene in a cell, the method comprising delivering to a cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises at least two gRNAs selected from one gRNA targeting a gene that reduces or eliminates expression of HLA-A on the cell surface, one gRNA targeting a gene that reduces or eliminates MHC class II expression on the cell surface, and one gRNA targeting a gene that reduces or eliminates endogenous TCR expression.
[0344] In some embodiments, methods are provided for modifying target genes in cells, comprising delivering to a cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises one gRNA targeting a gene that reduces or eliminates expression of HLA-A on the cell surface, one gRNA targeting a gene that reduces or eliminates MHC class II expression on the cell surface, and one gRNA targeting a gene that reduces or eliminates endogenous TCR expression.
[0345] In some embodiments, methods are provided for modifying a target gene in a cell, the method comprising delivering to a cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises one gRNA selected from a gRNA targeting TRAC, TRBC, B2M, HLA-A, or CIITA. In some embodiments, one gRNA targets TRAC. In some embodiments, one gRNA targets TRBC. In some embodiments, one gRNA targets B2M. In some embodiments, one gRNA targets HLA-A. In some embodiments, one gRNA targets CIITA.
[0346] In some embodiments, a method of modifying a target gene in a cell is provided, comprising delivering to the cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises at least two gRNAs selected from gRNAs targeting TRAC, TRBC, or B2M, and wherein the two guide RNAs do not target the same gene. In some embodiments, each gRNA is present in a separate lipid-nucleic acid assembly composition.
[0347] In some embodiments, a method of modifying a target gene in a cell is provided, comprising delivering to the cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises at least two gRNAs selected from gRNAs targeting TRAC, TRBC, or HLA-A, and wherein the two guide RNAs do not target the same gene. In some embodiments, each gRNA is present in a separate lipid-nucleic acid assembly composition.
[0348] In some embodiments, a method of modifying a target gene in a cell is provided, comprising delivering to the cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises at least two gRNAs selected from gRNAs targeting TRAC, TRBC, and HLA-A, and wherein the two guide RNAs do not target the same gene. In some embodiments, each gRNA is present in a separate lipid-nucleic acid assembly composition.
[0349] In some embodiments, a method for modifying a target gene in a cell is provided, comprising delivering to the cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises one guide RNA targeting TRAC and one gRNA targeting TRBC. In some embodiments, the gRNAs are each present in separate lipid-nucleic acid assembly compositions.
[0350] In some embodiments, methods are provided for modifying a target gene in a cell, the methods comprising delivering to the cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises one guide RNA targeting B2M and one gRNA targeting CIITA. In some embodiments, the gRNAs are each present in separate lipid-nucleic acid assembly compositions.
[0351] In some embodiments, a method of modifying a target gene in a cell is provided, comprising delivering to the cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises one guide RNA targeting HLA-A and one gRNA targeting CIITA. In some embodiments, the gRNAs are each present in separate lipid-nucleic acid assembly compositions. In some embodiments, the cells are homozygous for HLA-B and homozygous for HLA-C.
[0352] In some embodiments, a method of modifying a target gene in a cell is provided, comprising delivering to the cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises one guide RNA targeting TRAC, one gRNA targeting TRBC, and one gRNA targeting B2M. In some embodiments, the gRNAs are each present in separate lipid-nucleic acid assembly compositions.
[0353] In some embodiments, a method of modifying a target gene in a cell is provided, comprising delivering to the cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises one guide RNA targeting TRAC, one gRNA targeting TRBC, and one gRNA targeting HLA-A. In some embodiments, the gRNAs are each present in separate lipid-nucleic acid assembly compositions. In some embodiments, the cells are homozygous for HLA-B and homozygous for HLA-C.
[0354] In some embodiments, a method of modifying a target gene in a cell is provided, comprising delivering to the cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises one guide RNA targeting TRAC, one gRNA targeting TRBC, one gRNA targeting B2M, and one gRNA targeting CIITA. In some embodiments, the gRNAs are each present in a separate lipid-nucleic acid assembly composition.
[0355] In some embodiments, a method of modifying a target gene in a cell is provided, comprising delivering to the cell one or more lipid-nucleic acid assembly compositions, optionally lipid nanoparticles, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs, wherein the method comprises one guide RNA targeting TRAC, one gRNA targeting TRBC, one gRNA targeting HLA-A, and one gRNA targeting CIITA. In some embodiments, the gRNAs are each present in separate lipid-nucleic acid assembly compositions. In some embodiments, the cells are homozygous for HLA-B and homozygous for HLA-C.
[0356] In some embodiments, a cell is provided that comprises a composition comprising a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase, and a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), the second mRNA being different from the first mRNA.
[0357] In some embodiments, engineered cells are provided that include at least one base edit and / or indel, where the base edit and / or indel is effected by contacting the cell with a composition comprising: a first mRNA that includes a first open reading frame encoding a polypeptide comprising a cytidine deaminase (e.g., APOBEC3A deaminase (A3A)) and an RNA-guided nickase; and a second mRNA that includes a second open reading frame encoding a uracil glycosylase inhibitor (UGI), where the second mRNA is different from the first mRNA.
[0358] In some embodiments, the cell is a human cell. In some embodiments, a genetically modified cell is referred to as an engineered cell. An engineered cell refers to a cell (or a progeny of a cell) that contains an engineered genetic modification, e.g., that has been contacted with and genetically modified by a gene editing system. The terms "engineered cell" and "genetically modified cell" are used interchangeably throughout. An engineered cell can be any exemplary cell type disclosed herein. In some embodiments, the cell is an allogeneic cell.
[0359] In some embodiments, the cell is an immune cell. As used herein, "immune cell" refers to a cell of the immune system, including, for example, lymphocytes (e.g., T cells, B cells, natural killer cells ("NK cells," and NKT cells, or iNKT cells)), monocytes, macrophages, mast cells, dendritic cells, or granulocytes (e.g., neutrophils, eosinophils, and basophils). In some embodiments, the cell is a primary immune cell. In some embodiments, the immune system cell may be selected from CD3+, CD4+, and CD8+ T cells, regulatory T cells (Tregs), B cells, NK cells, and dendritic cells (DCs). In some embodiments, the immune cell is allogeneic. In some embodiments, the cell is a lymphocyte. In some embodiments, the cell is an adaptive immune cell. In some embodiments, the cell is a T cell. In some embodiments, the cell is a B cell. In some embodiments, the cell is an NK cell. In some embodiments, the lymphocyte is allogeneic.
[0360] In some embodiments, genome editing or modification of the target gene is performed in vivo, hi some embodiments, genome editing or modification of the target gene is in isolated or cultured cells.
[0361] In some embodiments, the target gene is present in an organ, such as the liver, for example, a mammalian liver, for example, a human liver. In some embodiments, the target gene is present in a liver cell, for example, a mammalian liver cell, for example, a human liver cell. In some embodiments, the target gene is present in a hepatocyte, for example, a mammalian hepatocyte, for example, a human hepatocyte. In some embodiments, the liver cell or hepatocyte is present in situ. In some embodiments, the liver cell or hepatocyte is isolated in culture, for example, a primary culture.
[0362] In some embodiments, the genome editing or modification of the target gene inactivates a splice donor or splice acceptor site.
[0363] Also provided are methods corresponding to the uses disclosed herein, which include administering to a subject a nucleic acid (e.g., mRNA), polypeptide, composition, or lipid-nucleic acid assembly composition disclosed herein, or contacting such a cell with a nucleic acid (e.g., mRNA), polypeptide, composition, or lipid-nucleic acid assembly composition disclosed herein.
[0364] In some embodiments, a nucleic acid (e.g., mRNA), polypeptide, composition, or lipid-nucleic acid assembly composition disclosed herein is administered intravenously for any of the above-mentioned uses for an organism, organ, or cell in situ.
[0365] In any of the foregoing embodiments involving a subject, the subject can be a mammal. In any of the foregoing embodiments involving a subject, the subject can be a human. In any of the foregoing embodiments involving a subject, the subject can be a cow, pig, monkey, sheep, dog, cat, fish, or poultry.
[0366] In some embodiments, the nucleic acids (e.g., mRNA), polypeptides, compositions, or lipid-nucleic acid assembly compositions disclosed herein are administered intravenously or are for intravenous administration.
[0367] In some embodiments, genome editing or modification of the target gene knocks down the expression of the target gene. In some embodiments, genome editing or modification of the target gene knocks down the expression of the target gene by at least 50%, 55%, 60%, 65%, 70%, 75%, or 80%. In some embodiments, genome editing or modification of the target gene creates a missense mutation in the gene.
[0368] In some embodiments, a single administration of a nucleic acid (e.g., mRNA), polypeptide, composition, or lipid nucleic acid assembly composition disclosed herein is sufficient to knock down expression of a target gene product. In some embodiments, a single administration of a nucleic acid (e.g., mRNA), polypeptide, composition, or lipid nucleic acid assembly composition disclosed herein is sufficient to knock out expression of a target gene product. In other embodiments, multiple administrations of a nucleic acid (e.g., mRNA), polypeptide, composition, or lipid nucleic acid assembly composition disclosed herein may be beneficial to maximize editing through a cumulative effect.
[0369] In some embodiments, the efficacy of treatment with a nucleic acid (e.g., mRNA), polypeptide, composition, or lipid nucleic acid assembly composition disclosed herein is seen 1 year, 2 years, 3 years, 4 years, 5 years, or 10 years after delivery.
[0370] In some embodiments, treatment slows or stops the progression of the disease.
[0371] In some embodiments, treatment results in improvement, stabilization, or slowing of changes in organ function or symptoms of organ disease.
[0372] In some embodiments, the efficacy of treatment is measured by extending the subject's survival time.
[0373] 1. Exemplary Guide RNAs, Compositions, Methods, and Engineered Cells for TRAC and TRBC Editing The present disclosure provides guide RNAs that target TRAC. Guide sequences that target the TRAC gene are shown in SEQ ID NOs: 706-721 in Table 5A.
[0374] The present disclosure provides guide RNAs that target TRBC. Guide sequences that target the TRBC gene are shown in SEQ ID NOs: 618-669 in Table 5B.
[0375] In some embodiments, the guide sequence is complementary to the corresponding genomic region shown in the table below, according to coordinates from the human reference genome hg38. The guide sequence of further embodiments may be complementary to a sequence near a genomic coordinate listed in either Table 5A or Table 5B. For example, the guide sequence of further embodiments may be complementary to a sequence comprising 15 contiguous nucleotides ± 10 nucleotides of a genomic coordinate listed in either Table 5A or 5B.
[0376] As explained in the preceding sections, each of the guide sequences shown in Tables 5A and 5B may further comprise additional nucleotides to form a crRNA, for example, the following exemplary nucleotide sequence following the guide sequence at its 3' end: GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 139) in the 5' to 3' direction. In the case of an sgRNA, the guide sequence may further comprise additional nucleotides to form an sgRNA, for example, the following exemplary nucleotide sequence following the 3' end: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 140). The guide sequence may further comprise additional nucleotides to form an sgRNA, for example.
[0377] In some embodiments, the sgRNA comprises the modification pattern shown below in SEQ ID NO: 141, where N is any natural or unnatural nucleotide, and the entire plurality of Ns constitutes a guide sequence described herein, and the modified sgRNA comprises the following sequence: mN*mN*mN*NNNNNNNNNNNNNNNNNNNGUUUAGAmGmCmUmAmGmAmAmAmUmAmGmCAAGUUAAAAUAAGGCUAGUCCGUUAUCAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmU*mU*mU*mU (SEQ ID NO: 141), where "N" can be any natural or unnatural nucleotide. For example, SEQ ID NO: 141 is encompassed herein, where the plurality of Ns is replaced with any of the guide sequences disclosed herein. The modification remains as shown in SEQ ID NO: 141, even when the plurality of Ns is replaced with a guide nucleotide. That is, even though "N's" are substituted by guide nucleotides, the first three nucleotides are 2'OMe modified and there are phosphorothioate linkages between the first and second nucleotides, the second and third nucleotides, and the third and fourth nucleotides.
[0378] In some embodiments, a gRNA targeting TRAC comprises a guide sequence selected from: i) SEQ ID NOs: 706-721; ii) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from SEQ ID NOs: 706-721; iii) a guide sequence at least 95%, 90%, or 85% identical to a sequence selected from SEQ ID NOs: 706-721; iv) a sequence comprising 10 contiguous nucleotides ± 10 nucleotides of the genomic coordinates listed in Table 5A; v) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence from (iv); or vi) a guide sequence that is at least 95%, 90%, or 85% identical to a sequence selected from (v).
[0379] [Table 6-1] [Table 6-2] [Table 6-3] [Table 6-4] [Table 6-5]
[0380] In some embodiments, the guide sequence comprises SEQ ID NO: 706. In some embodiments, the guide sequence comprises SEQ ID NO: 707. In some embodiments, the guide sequence comprises SEQ ID NO: 708. In some embodiments, the guide sequence comprises SEQ ID NO: 709. In some embodiments, the guide sequence comprises SEQ ID NO: 710. In some embodiments, the guide sequence comprises SEQ ID NO: 711. In some embodiments, the guide sequence comprises SEQ ID NO: 712. In some embodiments, the guide sequence comprises SEQ ID NO: 713. In some embodiments, the guide sequence comprises SEQ ID NO: 714. In some embodiments, the guide sequence comprises SEQ ID NO: 715. In some embodiments, the guide sequence comprises SEQ ID NO: 716. In some embodiments, the guide sequence comprises SEQ ID NO: 717. In some embodiments, the guide sequence comprises SEQ ID NO: 718. In some embodiments, the guide sequence comprises SEQ ID NO: 719. In some embodiments, the guide sequence comprises SEQ ID NO: 720. In some embodiments, the guide sequence comprises SEQ ID NO: 721.
[0381] In some embodiments, a gRNA targeting a TRBC comprises a guide sequence selected from: i) SEQ ID NOs: 618-669; ii) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from SEQ ID NOs: 618-669; iii) a guide sequence at least 95%, 90%, or 85% identical to a sequence selected from SEQ ID NOs: 618-669; iv) a sequence comprising 10 contiguous nucleotides ± 10 nucleotides of the genomic coordinates listed in Table 5B; v) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence from (iv); or vi) a guide sequence at least 95%, 90%, or 85% identical to a sequence selected from (v).
[0382] [Table 7-1] [Table 7-2] [Table 7-3] [Table 7-4] [Table 7-5] [Table 7-6] [Table 7-7] [Table 7-8] [Table 7-9]
Table 7-10
Table 7-11
Table 7-12
Table 7-13
Table 7-14
Table 7-15
Table 7-16
Table 7-17
[0383] In some embodiments, the guide sequence comprises SEQ ID NO: 618. In some embodiments, the guide sequence comprises SEQ ID NO: 619. In some embodiments, the guide sequence comprises SEQ ID NO: 620. In some embodiments, the guide sequence comprises SEQ ID NO: 621. In some embodiments, the guide sequence comprises SEQ ID NO: 622. In some embodiments, the guide sequence comprises SEQ ID NO: 623. In some embodiments, the guide sequence comprises SEQ ID NO: 624. In some embodiments, the guide sequence comprises SEQ ID NO: 625. In some embodiments, the guide sequence comprises SEQ ID NO: 626. In some embodiments, the guide sequence comprises SEQ ID NO: 627. In some embodiments, the guide sequence comprises SEQ ID NO: 628. In some embodiments, the guide sequence comprises SEQ ID NO: 629. In some embodiments, the guide sequence comprises SEQ ID NO: 630. In some embodiments, the guide sequence comprises SEQ ID NO: 631. In some embodiments, the guide sequence comprises SEQ ID NO: 632. In some embodiments, the guide sequence comprises SEQ ID NO: 633. In some embodiments, the guide sequence comprises SEQ ID NO: 634. In some embodiments, the guide sequence comprises SEQ ID NO: 635. In some embodiments, the guide sequence comprises SEQ ID NO: 636. In some embodiments, the guide sequence comprises SEQ ID NO: 637. In some embodiments, the guide sequence comprises SEQ ID NO: 638. In some embodiments, the guide sequence comprises SEQ ID NO: 639. In some embodiments, the guide sequence comprises SEQ ID NO: 640. In some embodiments, the guide sequence comprises SEQ ID NO: 641. In some embodiments, the guide sequence comprises SEQ ID NO: 642. In some embodiments, the guide sequence comprises SEQ ID NO: 643. In some embodiments, the guide sequence comprises SEQ ID NO: 644. In some embodiments, the guide sequence comprises SEQ ID NO: 645. In some embodiments, the guide sequence comprises SEQ ID NO: 646. In some embodiments, the guide sequence comprises SEQ ID NO: 647. In some embodiments, the guide sequence comprises SEQ ID NO: 648. In some embodiments, the guide sequence comprises SEQ ID NO: 649. In some embodiments, the guide sequence comprises SEQ ID NO: 650. In some embodiments, the guide sequence comprises SEQ ID NO: 651.In some embodiments, the guide sequence comprises SEQ ID NO: 652. In some embodiments, the guide sequence comprises SEQ ID NO: 653. In some embodiments, the guide sequence comprises SEQ ID NO: 654. In some embodiments, the guide sequence comprises SEQ ID NO: 655. In some embodiments, the guide sequence comprises SEQ ID NO: 656. In some embodiments, the guide sequence comprises SEQ ID NO: 657. In some embodiments, the guide sequence comprises SEQ ID NO: 658. In some embodiments, the guide sequence comprises SEQ ID NO: 659. In some embodiments, the guide sequence comprises SEQ ID NO: 660. In some embodiments, the guide sequence comprises SEQ ID NO: 661. In some embodiments, the guide sequence comprises SEQ ID NO: 662. In some embodiments, the guide sequence comprises SEQ ID NO: 663. In some embodiments, the guide sequence comprises SEQ ID NO: 664. In some embodiments, the guide sequence comprises SEQ ID NO: 665. In some embodiments, the guide sequence comprises SEQ ID NO: 666. In some embodiments, the guide sequence comprises SEQ ID NO: 667. In some embodiments, the guide sequence comprises SEQ ID NO: 668. In some embodiments, the guide sequence comprises SEQ ID NO: 669.
[0384] In some embodiments, the present disclosure provides a method of modifying a DNA sequence in a TRAC gene, comprising delivering to a cell a composition disclosed herein. a gRNA comprising a guide sequence selected from: ai) SEQ ID NOs: 706-721; ii) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from SEQ ID NOs: 706-721; iii) a guide sequence at least 95%, 90%, or 85% identical to a sequence selected from SEQ ID NOs: 706-721; iv) a sequence comprising 10 contiguous nucleotides ± 10 nucleotides of a genomic coordinate listed in Table 5A; v) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence from (iv); or vi) a guide sequence that is at least 95%, 90%, or 85% identical to a sequence selected from (v); b. A nucleic acid encoding the gRNA of (a). may include:
[0385] In some embodiments, the present disclosure provides a method of reducing expression of a TRAC gene, comprising delivering to a cell a composition disclosed herein. a gRNA comprising a guide sequence selected from: ai) SEQ ID NOs: 706-721; ii) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from SEQ ID NOs: 706-721; iii) a guide sequence at least 95%, 90%, or 85% identical to a sequence selected from SEQ ID NOs: 706-721; iv) a sequence comprising 10 contiguous nucleotides ± 10 nucleotides of a genomic coordinate listed in Table 5A; v) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence from (iv); or vi) a guide sequence that is at least 95%, 90%, or 85% identical to a sequence selected from (v); b. A nucleic acid encoding the gRNA of (a). may include:
[0386] In some embodiments, the present disclosure provides methods of immunotherapy comprising administering to a subject, their autologous cells, and / or allogeneic cells, a composition disclosed herein. a gRNA comprising a guide sequence selected from: ai) SEQ ID NOs: 706-721; ii) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from SEQ ID NOs: 706-721; iii) a guide sequence at least 95%, 90%, or 85% identical to a sequence selected from SEQ ID NOs: 706-721; iv) a sequence comprising 10 contiguous nucleotides ± 10 nucleotides of a genomic coordinate listed in Table 5A; v) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence from (iv); or vi) a guide sequence that is at least 95%, 90%, or 85% identical to a sequence selected from (v); b. A nucleic acid encoding the gRNA of (a). may include:
[0387] In some embodiments, cells modified by the methods disclosed herein are provided. The cells may be modified ex vivo. The cells include T cells, CD4 + , or CD8 + The cells can be mammalian, primate, or human cells. The cells can be used for immunotherapy of a subject.
[0388] In some embodiments, a composition is provided comprising a gRNA comprising a guide sequence selected from: i) SEQ ID NOs:706-721, ii) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from SEQ ID NOs:706-721, iii) a guide sequence at least 95%, 90%, or 85% identical to a sequence selected from SEQ ID NOs:706-721, iv) a sequence comprising 10 contiguous nucleotides ± 10 nucleotides of a genomic coordinate listed in Table 5A, v) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence from (iv), or vi) a guide sequence that is at least 95%, 90%, or 85% identical to a sequence selected from (v). The composition may optionally further comprise any one of the nucleic acids (e.g., mRNA), polypeptides, compositions, or lipid-nucleic acid assembly compositions disclosed herein.
[0389] In certain embodiments, the compositions disclosed herein are used to modify the DNA sequence in the TRAC gene of a cell. In certain embodiments, the compositions disclosed herein are used to reduce the expression of the TRAC gene of a cell. In some embodiments, the compositions disclosed herein are used for immunotherapy of a subject.
[0390] In some embodiments, the present disclosure provides a method for modifying a DNA sequence within a TRBC1 and / or TRBC2 gene, comprising delivering to a cell a composition disclosed herein. a gRNA comprising a guide sequence selected from: ai) SEQ ID NOs: 618-669; ii) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from SEQ ID NOs: 618-669; iii) a guide sequence at least 95%, 90%, or 85% identical to a sequence selected from SEQ ID NOs: 618-669; iv) a sequence comprising 10 contiguous nucleotides ± 10 nucleotides of a genomic coordinate listed in Table 5B; v) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence from (iv); or vi) a guide sequence that is at least 95%, 90%, or 85% identical to a sequence selected from (v); or b. A nucleic acid encoding the gRNA of (a). may include:
[0391] In some embodiments, the present disclosure provides a method for reducing expression of the TRBC1 and / or TRBC2 genes, comprising delivering to a cell a composition disclosed herein. a gRNA comprising a guide sequence selected from: ai) SEQ ID NOs: 618-669; ii) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from SEQ ID NOs: 618-669; iii) a guide sequence at least 95%, 90%, or 85% identical to a sequence selected from SEQ ID NOs: 618-669; iv) a sequence comprising 10 contiguous nucleotides ± 10 nucleotides of a genomic coordinate listed in Table 5C; v) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence from (iv); or vi) a guide sequence that is at least 95%, 90%, or 85% identical to a sequence selected from (v); b. A nucleic acid encoding the gRNA of (a). may include:
[0392] In some embodiments, the present disclosure provides methods of immunotherapy comprising administering to a subject, their autologous cells, and / or allogeneic cells, a composition disclosed herein. a gRNA comprising a guide sequence selected from: ai) SEQ ID NOs: 618-669; ii) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from SEQ ID NOs: 618-669; iii) a guide sequence at least 95%, 90%, or 85% identical to a sequence selected from SEQ ID NOs: 618-669; iv) a sequence comprising 10 contiguous nucleotides ± 10 nucleotides of a genomic coordinate listed in Table 5B; v) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence from (iv); or vi) a guide sequence that is at least 95%, 90%, or 85% identical to a sequence selected from (v); b. A nucleic acid encoding the gRNA of (a). may include:
[0393] In some embodiments, cells modified by the methods disclosed herein may be provided. The cells may be modified ex vivo. The cells may include T cells, CD4 + , or CD8 + The cells are cells. The cells can be mammalian, primate, or human cells. The cells can be used for immunotherapy of a subject.
[0394] In some embodiments, a composition is provided comprising a gRNA comprising a guide sequence selected from: i) SEQ ID NOs: 618-669, ii) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence selected from SEQ ID NOs: 618-669, iii) a guide sequence at least 95%, 90%, or 85% identical to a sequence selected from SEQ ID NOs: 618-669, iv) a sequence comprising 10 contiguous nucleotides ± 10 nucleotides of a genomic coordinate listed in Table 5B, v) at least 17, 18, 19, or 20 contiguous nucleotides of a sequence from (iv), or vi) a guide sequence that is at least 95%, 90%, or 85% identical to a sequence selected from (v). The composition may optionally further comprise any one of the nucleic acids (e.g., mRNA), polypeptides, compositions, or lipid-nucleic acid assembly compositions disclosed herein.
[0395] In certain embodiments, the compositions disclosed herein are used to modify DNA sequences within the TRBC1 and / or TRBC2 genes in cells. In certain embodiments, the compositions disclosed herein are used to reduce expression of the TRBC1 and / or TRBC2 genes in cells. In some embodiments, the compositions disclosed herein are used for immunotherapy of a subject.
[0396] J. Exemplary DNA Molecules, Vectors, Expression Constructs, Host Cells, and Production Methods In certain embodiments, the present disclosure provides a DNA molecule comprising a sequence encoding a polypeptide described herein. In some embodiments, the DNA molecule further comprises a nucleic acid that does not encode a polypeptide. The nucleic acid that does not encode a polypeptide disclosed herein includes, but is not limited to, a promoter, an enhancer, a regulatory sequence, and a nucleic acid that encodes a gRNA.
[0397] In some embodiments, the DNA molecule further comprises a nucleotide sequence encoding a crRNA, a trRNA, or a crRNA and a trRNA. In some embodiments, the nucleotide sequence encoding the crRNA, the trRNA, or the crRNA and a trRNA comprises or consists of a guide sequence flanked by all or part of repeat sequences from a naturally occurring CRISPR / Cas system. The nucleic acid comprising or consisting of the crRNA, the trRNA, or the crRNA and a trRNA may further comprise a vector sequence, which comprises or consists of a nucleic acid not found in nature with the crRNA, the trRNA, or the crRNA and a trRNA. In some embodiments, the crRNA and the trRNA are encoded by non-contiguous nucleic acids within a single vector. In other embodiments, the crRNA and the trRNA may be encoded by contiguous nucleic acids. In some embodiments, the crRNA and the trRNA are encoded by opposite strands of a single nucleic acid. In other embodiments, the crRNA and the trRNA are encoded by the same strand of a single nucleic acid.
[0398] In some embodiments, the DNA molecule further comprises a promoter operably linked to a sequence encoding any of the mRNAs encoding the polypeptides described herein. In some embodiments, the DNA molecule is an expression construct suitable for expression in mammalian cells, such as human cells or mouse cells, such as human hepatocytes or rodent (e.g., mouse) hepatocytes. In some embodiments, the DNA molecule is an expression construct suitable for expression in cells of a mammalian organ, such as the human liver or rodent (e.g., mouse) liver. In some embodiments, the DNA molecule is a plasmid or episome. In some embodiments, the DNA molecule is contained in a host cell, such as a bacterium or a cultured eukaryotic cell. Exemplary bacteria include proteobacteria, such as E. coli. Exemplary cultured eukaryotic cells include primary hepatocytes, including hepatocytes of rodent (e.g., mouse) or human origin; hepatocyte cell lines, including hepatocytes of rodent (e.g., mouse) or human origin; human cell lines; rodent (e.g., mouse) cell lines; CHO cells; microbial fungi, such as fission yeast or budding yeast, such as Saccharomyces, e.g., S. cerevisiae; and insect cells.
[0399] In some embodiments, methods for producing mRNA disclosed herein are provided. In some embodiments, such methods include contacting a DNA molecule described herein with an RNA polymerase under conditions that allow transcription. In some embodiments, the contacting is performed in vitro, e.g., in a cell-free system. In some embodiments, the RNA polymerase is a bacteriophage-derived RNA polymerase, e.g., T7 RNA polymerase. In some embodiments, NTPs are provided that include at least one modified nucleotide as described above. In some embodiments, the NTPs include at least one modified nucleotide as described above and do not include UTP.
[0400] In some embodiments, the mRNA disclosed herein, alone or together with one or more gRNAs, can be contained within or delivered by a vector system of one or more vectors. In some embodiments, one or more or all of the vectors can be DNA vectors. In some embodiments, one or more or all of the vectors can be RNA vectors. In some embodiments, one or more or all of the vectors can be circular. In other embodiments, one or more or all of the vectors can be linear. In some embodiments, one or more or all of the vectors can be encapsulated in lipid nanoparticles, liposomes, non-lipid nanoparticles, or viral capsids. Non-limiting exemplary vectors include plasmids, phagemids, cosmids, artificial chromosomes, minichromosomes, transposons, viral vectors, and expression vectors.
[0401] Non-limiting exemplary viral vectors include adeno-associated viral (AAV) vectors, lentiviral vectors, adenoviral vectors, helper-dependent adenoviral vectors (HDAd), herpes simplex viral (HSV-1) vectors, bacteriophage T4, baculoviral vectors, and retroviral vectors. In some embodiments, the viral vector can be an AAV vector. In other embodiments, the viral vector can be a lentiviral vector. In some embodiments, the lentivirus can be non-integrating. In some embodiments, the viral vector can be an adenoviral vector. In some embodiments, the adenovirus can be a high-cloning-capacity or "gutless" adenovirus, in which all coding viral regions except the 5' and 3' inverted terminal repeats (ITRs) and packaging signal ("I") have been deleted from the virus to enhance its packaging ability. In yet other embodiments, the viral vector can be an HSV-1 vector. In some embodiments, the HSV-1-based vector is helper-dependent, while in other embodiments it is helper-independent. For example, an amplicon vector that retains only the packaging sequence requires a helper virus with structural elements for packaging, whereas a 30 kb deletion HSV-1 vector that removes non-essential viral functions does not. In additional embodiments, the viral vector can be bacteriophage T4. In some embodiments, bacteriophage T4 can package any linear or circular DNA or RNA molecule when the viral head is empty. In further embodiments, the viral vector can be a baculovirus vector. In yet further embodiments, the viral vector can be a retrovirus vector. In embodiments using AAV or lentivirus vectors with limited cloning capacity, it may be necessary to use multiple vectors to deliver all components of the vector system disclosed herein. For example, one AAV vector can contain a sequence encoding a Cas protein, while a second AAV vector can contain one or more guide sequences.
[0402] In some embodiments, the vector may be capable of driving expression of one or more coding sequences in a cell, such as the coding sequences of the mRNAs disclosed herein. In some embodiments, the cell may be a prokaryotic cell, such as a bacterial cell. In some embodiments, the cell may be a eukaryotic cell, such as a yeast, plant, insect, or mammalian cell. In some embodiments, the eukaryotic cell may be a mammalian cell. In some embodiments, the eukaryotic cell may be a rodent cell. In some embodiments, the eukaryotic cell may be a human cell. Suitable promoters for driving expression in different types of cells are known in the art. In some embodiments, the promoter may be wild-type. In other embodiments, the promoter may be modified for more efficient or effective expression. In still other embodiments, the promoter may be truncated but retain its function. For example, the promoter may have a normal size or a reduced size suitable for proper packaging of the vector into a virus.
[0403] In some embodiments, the vector system may contain one copy of the nucleotide sequence encoding the polypeptide disclosed herein. In other embodiments, the vector system may contain multiple copies of the nucleotide sequence encoding the polypeptide disclosed herein. In some embodiments, the nucleotide sequence encoding the polypeptide disclosed herein may be operably linked to at least one transcriptional or translational control sequence. In some embodiments, the nucleotide sequence encoding the protein may be operably linked to at least one promoter.
[0404] In some embodiments, the promoter can be constitutive, inducible, or tissue-specific. In some embodiments, the promoter can be a constitutive promoter. Non-limiting exemplary constitutive promoters include the cytomegalovirus immediate early promoter (CMV), the simian virus (SV40) promoter, the adenovirus major late (MLP) promoter, the Rous sarcoma virus (RSV) promoter, the mouse mammary tumor virus (MMTV) promoter, the phosphoglycerate kinase (PGK) promoter, the elongation factor-alpha (EF1a) promoter, the ubiquitin promoter, the actin promoter, the tubulin promoter, the immunoglobulin promoter, functional fragments thereof, or any combination of the above. In some embodiments, the promoter can be a CMV promoter. In some embodiments, the promoter can be a truncated CMV promoter. In other embodiments, the promoter can be an EF1a promoter. In some embodiments, the promoter can be an inducible promoter. Non-limiting exemplary inducible promoters include those inducible by heat shock, light, chemicals, peptides, metals, steroids, antibiotics, or alcohol. In some embodiments, the inducible promoter can be one that has a low basal (uninduced) expression level, such as the Tet-On® promoter (Clontech).
[0405] In some embodim...
Claims
1. A composition comprising a first mRNA having a first open reading frame encoding a polypeptide comprising cytidine deaminase and RNA-induced nickase, and a second mRNA having a second open reading frame encoding a uracil glycosylase inhibitor (UGI), the second mRNA being different from the first mRNA, and optionally, the composition comprising lipid nanoparticles.
2. The composition described in claim 1, wherein the first open reading frame does not contain a sequence encoding UGI.
3. A composition comprising a first lipid nanoparticle (LNP) and a second LNP, wherein the first LNP comprises a first mRNA comprising a first open reading frame encoding a polypeptide comprising cytidine deaminase and RNA-guided nickase, and wherein the first open reading frame does not comprise a sequence encoding a uracil glycosylase inhibitor (UGI), and the second LNP comprises a second mRNA comprising a second open reading frame encoding a UGI, and wherein the second mRNA is different from the first mRNA.
4. An in vitro or ex vivo method for modifying a target gene, the method comprising delivering to a cell a first mRNA comprising a first open reading frame encoding a first polypeptide comprising cytidine deaminase and RNA-guided nickase, and a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI), the second mRNA being different from the first mRNA, and at least one guide RNA (gRNA).
5. The method described in claim 4, wherein the first open reading frame does not contain a sequence encoding UGI.
6. A composition or method described in any one of claims 1 to 5, wherein the molar ratio of the second mRNA to the first mRNA is 1:1 to 30:1, 2:1 to 30:1, or 7:1 to 22:
1.
7. An mRNA comprising an open reading frame encoding a polypeptide comprising cytidine deaminase and RNA-induced nickase, the polypeptide not comprising uracil glycosylase inhibitor (UGI).
8. A cell, (i) the cell has been introduced with the mRNA or composition of any one of claims 1 to 3, 6 and 7, and the cell has been modified after said introduction, or (ii) the cell is an engineered cell that has been modified by the method of any one of claims 4 to 6; The cells.
9. The cytidine deaminase (i) an enzyme of the APOBEC family, optionally an enzyme of the APOBEC3 subgroup; (ii) a cytidine deaminase comprising an amino acid sequence that is at least 80%, 85%, 87%, 90%, 95%, 98%, 99% or 100% identical to any one of SEQ ID NOs: 40, 41, and 960-1023; (iii) a cytidine deaminase comprising an amino acid sequence that is at least 80%, 85%, 87%, 90%, 95%, 98%, 99% or 100% identical to any one of SEQ ID NOs: 40, 41, and 960-1013; (iv) a cytidine deaminase comprising an amino acid sequence that is at least 80%, 85%, 87%, 90%, 95%, 98%, 99% or 100% identical to any one of SEQ ID NOs: 40, 41, 976, 977, 979, 980, 984-987, 993-1006, and 1009; or (v) a cytidine deaminase comprising an amino acid sequence that is at least 80%, 85%, 87%, 90%, 95%, 98%, 99% or 100% identical to any one of SEQ ID NOs: 40, 976, 981, 984, 986, and 1014 to 1023. The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 8, wherein 10. The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 9, wherein the cytidine deaminase is APOBEC3A deaminase (A3A), optionally human A3A, and further optionally wild-type A3A, and optionally the A3A comprises the amino acid sequence of any one of SEQ ID NOs: 40, 976, 977, 993-1006, and 1009, or an amino acid sequence having at least 87%, at least 90%, at least 95%, at least 98%, or at least 99% identity to the amino acid sequence of any one of SEQ ID NOs: 40, 976, 977, 993-1006, and 1009.
11. The mRNA, composition, method, cell, or engineered cell described in any one of claims 1 to 10, wherein the UGI comprises the amino acid sequence of SEQ ID NO:27, or an amino acid sequence having at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% identity to SEQ ID NO:
27.
12. The RNA-guided nickase (i) S. pyogenes (Spy) Cas9 nickase, optionally a D10A SpyCas9 nickase; or (ii) N. meningitidis (Nme) Cas9 nickase, optionally a D16A NmeCas9 nickase, and further optionally a D16A Nme2Cas9; The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 11.
13. The RNA-guided nickase comprises the amino acid sequence of SEQ ID NO: 70, 73, or 76, or an amino acid sequence having at least 80%, 90%, 95%, 98%, or 99% identity to the amino acid sequence of SEQ ID NO: 70, 73, or 76; or the sequence encoding the RNA-guided nickase comprises the nucleotide sequence of any one of SEQ ID NOs: 71, 72, 74, 75, or 77-90, or a nucleotide sequence having at least 90%, 95%, 98%, or 99% identity to the nucleotide sequence of any one of SEQ ID NOs: 71, 72, 74, 75, or 77-90; The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 12.
14. The mRNA, a) a 5'UTR having at least 90% identity to any one of SEQ ID NOs: 91-98; b) a 3'UTR having at least 90% identity to any one of SEQ ID NOs: 99-106 c) a 5' cap selected from Cap0, Cap1, Cap2, and a co-transcriptionally or post-transcriptionally added cap, optionally wherein the co-transcriptionally added cap is selected from an anti-reverse cap analog (ARCA), AG(m7G(5')ppp(5')(2'OmeA)pG, or GG(m7G(5')ppp(5')(2'OmeG)pG, a post-transcriptionally added cap; and d) a polyadenylation (polyA) tail, optionally added to the mRNA by PCR tailing or enzymatic tailing, further optionally wherein the polyA tail is the sequence of SEQ ID NO:
109.
14. The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 13, comprising one or more of:
15. The mRNA, composition, method, cell, or engineered cell described in any one of claims 1 to 14, wherein the cytidine deaminase is located at the N-terminus of the RNA-guided nickase in the polypeptide.
16. The mRNA, composition, method, or cell described in any one of claims 1 to 15, wherein the polypeptide comprising the cytidine deaminase and the RNA-guided nickase further comprises a peptide linker between the cytidine deaminase and the RNA-guided nickase, and the peptide linker comprises one or more sequences selected from SEQ ID NOs: 268, 46-59, 61, 211-267, and 269-272.
17. The polypeptide comprising the cytidine deaminase and the RNA-guided nickase further comprises a nuclear localization signal (NLS); the NLS comprises any one of SEQ ID NOs: 63 and 110-122, or a sequence having at least 80%, 85%, 90%, or 95% identity to any one of SEQ ID NOs: 63 and 110-122, or the NLS is encoded by any one of SEQ ID NOs: 123-135, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 100% identity to any one of SEQ ID NOs: 123-135; The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 16. (i) the NLS is at the C-terminus of the RNA-guided nickase; (ii) the NLS is at the N-terminus of the RNA-guided nickase; (iii) the NLS is fused to both the N-terminus and C-terminus of the RNA-guided nickase; or (iv) the cytidine deaminase is located at the N-terminus of the NLS in the polypeptide, and the RNA-guided nickase is located at the N-terminus of the NLS in the polypeptide; 18. The mRNA, composition, method, cell, or engineered cell of claim 17.
19. The mRNA, composition, method, cell, or engineered cell of claim 17 or 18, wherein the polypeptide comprising the cytidine deaminase and the RNA-guided nickase further comprises two NLSs localized at the N-terminus of the polypeptide, optionally an SV40 NSL and a plasmin NSL at the N-terminus of the polypeptide, wherein the SV40 NSL comprises the sequence of SEQ ID NO: 63, and the plasmin NSL comprises the sequence of SEQ ID NO:
122.
20. The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 19, the polypeptide comprising the cytidine deaminase and the RNA-guided nickase, a) an APOBEC cytidine deaminase, optionally an APOBEC3A deaminase, wherein optionally the cytidine deaminase comprises an amino acid sequence having at least 87% identity to SEQ ID NO: 40 or at least 80% identity to SEQ ID NO: 41; b) the nickase is an N. meningitidis (Nme) Nme2Cas9 nickase, an Nme1Cas9 nickase, or an Nme3Cas9 nickase, optionally wherein the nickase comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 387; and c) a first nuclear localization signal (NLS), optionally comprising an amino acid sequence having at least 90% identity to SEQ ID NO: 63, and free of a uracil glycosylase inhibitor (UGI); and a second NLS, optionally comprising an amino acid sequence having at least 90% identity to SEQ ID NO:
122. Including, Optionally, the cytidine deaminase is located C-terminal to the first NLS and the second NLS in the polypeptide. The mRNA, composition, method, cell, or engineered cell.
21. The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 20, The open reading frame (ORF) encoding the polypeptide comprising the cytidine deaminase and the RNA-guided nickase comprises, in a 5' to 3' direction: a) a nucleotide sequence encoding a first nuclear localization signal (NLS), optionally wherein the first NLS comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 63, and optionally a nucleotide sequence encoding a second NLS, optionally wherein the second NLS comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 122; b) a nucleotide sequence encoding a cytidine deaminase comprising an amino acid sequence having at least 87% identity to SEQ ID NO: 40 or at least 80% identity to SEQ ID NO: 41; and c) (i) a nucleotide sequence encoding a N. meningitidis (Nme) Cas9 nickase polypeptide comprising an amino acid sequence having at least 95% identity to SEQ ID NO:387, optionally an amino acid sequence having at least 99% identity to SEQ ID NO:387, or (ii) a nucleotide sequence having at least 95% identity to SEQ ID NO:389, optionally having at least 99% identity to SEQ ID NO:
389. Including, Optionally, the ORF does not include a sequence encoding an NLS at the C-terminus of the sequence encoding NmeCas9, and optionally, the ORF does not include a coding sequence at the C-terminus of the ORF encoding NmeCas9. The mRNA, composition, method, cell, or engineered cell.
22. The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 21, (i) the polypeptide comprising cytidine deaminase and RNA-guided nickase comprises the amino acid sequence of any one of SEQ ID NOs: 303, 306, 309, 312, 313, and 321-337, or comprises an amino acid sequence having at least 80%, 85%, 88%, 90%, 95%, 98%, or 99% identity to any one of SEQ ID NOs: 303, 306, 309, 312, 313, and 321-337; or the open reading frame encoding the polypeptide comprising cytidine deaminase and RNA-guided nickase comprises the sequence of any one of SEQ ID NOs: 302, 305, 308, 311 and 361-377, or a sequence having at least 80%, 85%, 88%, 90%, 95%, 98% or 99% identity to any one of SEQ ID NOs: 302, 305, 308, 311 and 361-377; or (ii) the polypeptide comprising cytidine deaminase and RNA-guided nickase comprises the amino acid sequence of any one of SEQ ID NOs: 3, 6, 15, 21, and 33, or comprises an amino acid sequence having at least 80%, 85%, 88%, 90%, 95%, 98%, or 99% identity to any one of SEQ ID NOs: 3, 6, 15, 21, and 33; or the open reading frame encoding the polypeptide comprising cytidine deaminase and RNA-guided nickase comprises the sequence of any one of SEQ ID NOs: 2, 5, 14, 20, and 32, or a sequence having at least 80%, 85%, 88%, 90%, 95%, 98%, or 99% identity to any one of SEQ ID NOs: 2, 5, 14, 20, and 32; The mRNA, composition, method, cell, or engineered cell.
23. The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 22, the polypeptide comprising cytidine deaminase and RNA-guided nickase comprises the amino acid sequence of SEQ ID NO: 312, or comprises an amino acid sequence having at least 80%, 85%, 88%, 90%, 95%, 98% or 99% identity to SEQ ID NO: 312; or The open reading frame encoding the polypeptide comprising cytidine deaminase and RNA-guided nickase comprises the sequence of SEQ ID NO: 311 or a sequence having at least 80%, 90%, 95%, 98% or 99% identity to SEQ ID NO: 311; The mRNA, composition, method, cell, or engineered cell.
24. The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 22, The polypeptide comprising cytidine deaminase and RNA-guided nickase comprises the amino acid sequence of SEQ ID NO: 313 or an amino acid sequence having at least 80%, 90%, 95%, 98% or 99% identity to SEQ ID NO: 313; The mRNA, composition, method, cell, or engineered cell.
25. The mRNA, composition, method, cell, or engineered cell of any one of claims 1 to 20, wherein at least 10%, 20%, 30%, 80%, 90% or 100% of the uridines in the mRNA are substituted with modified uridines, and optionally the modified uridines are one or more of N1-methyl-pseudouridine, pseudouridine, 5-methoxyuridine, or 5-iodouridine.
26. A polypeptide encoded by any one of the mRNAs described in any one of claims 1 to 10 and 12 to 21.
27. A polypeptide comprising, from the N-terminus to the C-terminus: a) a first nuclear localization signal (NLS) comprising an amino acid sequence having at least 90% identity to, or comprising, SEQ ID NO: 63; b) a second NLS comprising an amino acid sequence having at least 90% identity to, or comprising, SEQ ID NO: 122; c) a cytidine deaminase, optionally an APOBEC3A deaminase, comprising an amino acid sequence having at least 87% identity to SEQ ID NO: 40, or the amino acid sequence of SEQ ID NO: 40, excluding the N-terminal methionine, and d) An Nme2Cas9 nickase comprising an amino acid sequence having at least 90% identity to SEQ ID NO: 387, or the amino acid sequence of SEQ ID NO: 387, excluding the N-terminal methionine. The polypeptide comprising: (i) a first peptide linker between the second NLS and the cytidine deaminase, optionally the first peptide linker comprising GGG or GGS; and (ii) a second peptide linker between the cytidine deaminase and the Nme2Cas9 nickase, optionally wherein the second peptide linker comprises the amino acid sequence of SEQ ID NO:
268.
28. The polypeptide of claim 27, further comprising one or both of:
29. A polypeptide comprising an amino acid sequence having at least 95%, 98% or 99% identity to SEQ ID NO: 313, or the amino acid sequence of SEQ ID NO:
313.
30. A polynucleotide encoding a polypeptide described in any one of claims 27 to 29.
31. A polynucleotide comprising an open reading frame (ORF), The ORF, from 5' to 3', a) a nucleotide sequence encoding a first nuclear localization signal (NLS), (i) the first NLS comprises an amino acid sequence having at least 90% identity to, or the amino acid sequence of, SEQ ID NO: 63; or (ii) the nucleotide sequence encoding the first NLS comprises a nucleotide sequence having at least 90% identity to SEQ ID NO: 62; a nucleotide sequence encoding said first NLS; b) a nucleotide sequence encoding a second NLS, wherein said second NLS comprises an amino acid sequence having at least 90% identity with, or the amino acid sequence of, SEQ ID NO: 122; c) a nucleotide sequence encoding a cytidine deaminase, wherein the cytidine deaminase comprises an amino acid sequence having at least 87% identity to SEQ ID NO: 40, or the amino acid sequence of SEQ ID NO: 40 excluding the N-terminal methionine; and d) a nucleotide sequence encoding an Nme2Cas9 nickase, (i) the Nme2Cas9 nickase comprises an amino acid sequence having at least 90% identity to, or the amino acid sequence of, SEQ ID NO: 387, excluding the N-terminal methionine, or (ii) the nucleotide sequence encoding the Nme2Cas9 nickase has at least 95% identity with SEQ ID NO: 389; Nucleotide sequence encoding the Nme2Cas9 nickase The polynucleotide comprising: (i) a nucleotide sequence encoding a first peptide linker between the second NLS and the cytidine deaminase, optionally wherein the first peptide linker comprises GGG or GGS; and (ii) a nucleotide sequence encoding a second peptide linker between the cytidine deaminase and the Nme2Cas9 nickase, optionally wherein the second peptide linker comprises the amino acid sequence of SEQ ID NO:
268.
32. The polynucleotide of claim 31, further comprising one or both of:
33. A polynucleotide comprising a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO:311, or an open reading frame (ORF) comprising the nucleotide sequence of SEQ ID NO:
311.
34. A ribonucleoprotein complex (RNP) comprising: (i) a polypeptide encoded by any one of the mRNAs or polynucleotides described in any one of claims 1 to 7, 9 to 25, and 30 to 33, or a polypeptide described in any one of claims 26 to 29; and (ii) a guide RNA.
35. A vector comprising a nucleotide sequence encoding the mRNA or polynucleotide described in any one of claims 1 to 3, 6, 7, 9 to 25 and 30 to 33.
36. An expression construct, or a plasmid containing said expression construct, comprising a promoter operably linked to a sequence encoding any one of the mRNAs or polynucleotides described in any one of claims 1 to 3, 6, 7, 9 to 25 and 30 to 33.
37. A host cell comprising the vector of claim 35, or the expression construct or plasmid of claim 36.
38. The mRNA, composition, polypeptide or polynucleotide described in any one of claims 1 to 3, 6, 7 and 9 to 33, wherein the mRNA or composition is formulated as a lipid nanoparticle (LNP) composition.
39. A lipid nanoparticle (LNP) composition comprising mRNA including an open reading frame (ORF) encoding a polypeptide comprising cytidine deaminase and RNA-guided nickase, wherein the polypeptide does not comprise a uracil glycosylase inhibitor (UGI).
40. One or more lipid nanoparticle (LNP) compositions, comprising: a) a first mRNA comprising a first open reading frame (ORF) encoding a polypeptide comprising a cytidine deaminase and an RNA-guided nickase; b) a second mRNA comprising a second ORF encoding a uracil glycosylase inhibitor (UGI); and c) one or more guide RNAs The LNP composition comprising:
41. An in vivo or ex vivo method for modifying a target gene in a cell, comprising: (a) a first mRNA comprising a first open reading frame encoding a polypeptide comprising a cytidine deaminase and an RNA-guided nickase; (b) a second mRNA comprising a second open reading frame encoding a uracil glycosylase inhibitor (UGI); and (c) one or more guide RNAs; The method comprises delivering to the cells one or more lipid nanoparticle (LNP) compositions comprising:
42. The LNP composition or method according to any one of claims 39 to 41, (i) portions (a), (b), and (c) are each present in a separate LNP; (ii) moieties (a), (b), and (c) are present in the same LNP; or (iii) moieties (a) and (c) are present in the same LNP and moiety (b) is present in a separate LNP; Further, optionally, each of the one or more guide RNAs is present in a separate LNP. The LNP composition or method.
43. The LNP composition comprising: (i) an amine lipid; (ii) a helper lipid; (iii) a stealth lipid; (iv) a neutral lipid; or a combination of two or more of (i) to (iv); Optionally, (i) the amine lipid is lipid A ((9Z,12Z)-3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl octadeca-9,12-dienoate (also referred to as 3-((4,4-bis(octyloxy)butanoyl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl(9Z,12Z)-octadeca-9,12-dienoate)), (ii) the helper lipid is cholesterol, (iii) the stealth lipid is PEG2k-DMG, (iv) the neutral lipid is DSPC, or a combination of two or more of (i)-(iv). The LNP composition or method of any one of claims 39 to 42.
44. The LNP composition or method of any one of claims 39 to 43, wherein the LNP composition comprises about 35 mol% of an amine lipid such as lipid A, about 15 mol% of a neutral lipid such as DSPC, about 2.5 mol% of a stealth lipid such as a PEG lipid, e.g., PEG2k-DMG, with the remainder of the lipid components being helper lipids such as cholesterol, and the N / P ratio is about 6.
45. A polypeptide comprising cytidine deaminase and RNA-induced nickase, wherein the polypeptide does not contain uracil glycosylase inhibitor (UGI).
46. A composition comprising: a first polypeptide comprising cytidine deaminase and RNA-induced nickase, wherein the first polypeptide does not comprise a uracil glycosylase inhibitor (UGI); and a second polypeptide comprising a UGI, wherein the second polypeptide is different from the first polypeptide.
47. A pharmaceutical composition or kit comprising an mRNA, RNP, composition, polypeptide, polynucleotide, or cell described in any one of claims 1 to 3, 6, 7, 9 to 34, 38 to 40, and 42 to 46, and optionally a pharmaceutically acceptable carrier.
48. An mRNA, RNP, composition, polypeptide, polynucleotide, or cell described in any one of claims 1 to 3, 6, 7, 9 to 34, 38 to 40, and 42 to 46 for use in modifying a target gene.
49. An mRNA, RNP, composition, polypeptide, polynucleotide, or cell described in any one of claims 1 to 3, 6, 7, 9 to 34, 38 to 40, and 42 to 46 for use as a pharmaceutical.