High-efficiency low-miss gene editing tool
Patent Information
- Application Number
- CN202480002331.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-25
- Filing Date
- 2024-10-25
- Publication Date
- 2025-06-27
Smart Images

Figure 00000079_0000 
Figure 00000079_0001 
Figure 00000080_0000
Abstract
Description
A gene editing tool with high efficiency and low off-target effects Technical Field
[0001] The present application belongs to the field of biomedicine technology, and specifically relates to a gene editing tool with high efficiency and low off-target effects. Background Art
[0002] CRISPR-Cas9 is the third-generation gene editing technology, following the introduction of ZFNs, TALENs, and other gene editing technologies. Since Zhang Feng first reported the high editing efficiency of Cas9 in mammalian cells in 2013, CRISPR-Cas9 has rapidly developed, becoming one of the most efficient, simplest, lowest-cost, and most versatile technologies for gene editing and modification, and the most mainstream gene editing system today.
[0003] CRISPR (Clustered Regularly Interspersed Short Palindromic Repeats) is a natural immune system in prokaryotes, found in 40% of sequenced bacteria and 90% of sequenced archaea. After being invaded by a virus, some bacteria can "store" a small fragment of the viral gene within their own DNA. Upon encountering another viral attack, the bacteria can recognize the virus based on this "stored" fragment and cut the viral DNA, rendering it ineffective.
[0004] Currently, the classic CRISPR-Cas9 system consists of two parts: sgRNA and Cas9 protein. The sgRNA (single guide RNA) is composed of crRNA and tracrRNA. Currently, a linker is often used to connect the 3' end of the crRNA to the 5' end of the tracrRNA to form a single sgRNA. The Cas9 protein binds to the sgRNA, forming an RNA-protein complex (RNP). It locates and recognizes the PAM and target sequence, unwinding the DNA double strands to form an R-loop. The sgRNA hybridizes with the complementary strand, while the other strand remains free. Subsequently, the Cas9 protein precisely cuts the strand at the blunt-end site three nucleotides upstream of the PAM, forming a blunt-ended product. The HNH domain of the Cas9 protein is responsible for cleaving the DNA strand that is complementary to the crRNA, while the RuvC domain is responsible for cleaving the other, non-complementary DNA strand. Ultimately, the action of Cas9 creates a DNA double-strand break (DSB), and gene knockout is achieved through frameshift mutations during the genome repair process.
[0005] Although the CRISPR-Cas9 system can efficiently edit target genes, the wild-type Cas9 protein has a high off-target rate, with an average of approximately 100 off-target sites. This high off-target frequency can lead to serious risks of genomic mutations and chromosomal translocations, hindering the development of gene editing.
[0006] Studies have shown that the specificity of the CRISPR / Cas9 system mainly depends on the recognition sequence of the sgRNA. Since the designed sgRNA may form mismatches with non-target DNA sequences, it may lead to unexpected gene mutations, which are called off-target effects. Since off-target mutations may cause genomic instability and disrupt the functions of other normal genes, researchers are committed to studying the factors affecting the off-target effect of the CRISPR / Cas9 system. Specifically, the factors affecting off-target effect include: (1) the effect of sgRNA and DNA pairing sequence on off-target effect: there may be base mismatches between the sgRNA sequence and the DNA binding region; (2) the effect of PAM on off-target effect: generally, different PAM sequences have different off-target effects; generally speaking, the more nucleotides required, the lower the off-target effect; (3) the influence of Cas9 and other factors: by modifying the wild-type Cas9 to improve the specificity of the CRISPR / Cas9 system, the use of mutant nucleases can induce efficient gene editing in human cells and greatly reduce the level of off-target mutations.
[0007] Regarding off-target effects, current CRISPR / Cas9 gene editing technology has the following problems: 1) The off-target rate of the Cas9 protein is high, resulting in a higher risk of genome damage and chromosomal aberrations; 2) The current mainstream gene editing delivery vector is AVV, which allows for prolonged expression of the editing tool, increasing the risk of off-target effects; 3) Viral vector delivery is immunogenic and unsuitable for repeated administration. Because the Cas9 nuclease has been modified to act as a precise DNA-binding protein under the guidance of gRNA, it can guide other functional proteins connected to it to the target site for nucleotide modification operations, thus forming a series of gene editing tools, including but not limited to base editors and epigenetic editors. However, these systems also have off-target effects for the same reasons.
[0008] Summary of the Invention
[0009] This application modifies and transforms the Cas9 nuclease to obtain a nuclease with significantly reduced off-target effects. This modification is applicable to all known gene editing systems based on the DNA targeting function of Cas9, including but not limited to Cas9 double-stranded nucleic acid cleavage enzyme, nCas9 (Cas9 nickase), or dCas9 (catalytically dead Cas9), or their fusion proteins, such as base editing systems such as ABE8e and epigenetic editors such as CRISPRoff-EE.
[0010] Specifically, this application involves the following technical solutions:
[0011] Item 1. An isolated modified Cas9 nuclease or DNA-binding fragment thereof, comprising a mutation at one or more amino acid residue positions selected from the group consisting of: K526, N692, Q695, H698, N497, Y450, Q926, K377, E387, D397, R400, D406, A421, L423, R424, Q426, Y430, K442, P449, V452, A456, R457, W464, M465, K468, E470, T474, P475, W476, F478, K484, S487, A488, T496, F498, L502, N504, K506, P509, F518, N522, E523, L540, S541, I548, D550, F553, V561, K562, E573, A589, L598, D605, L607, N609, N612, E617, D6 18. D628, R629, R635, K637, L651, K652, R654, T657, G658, L666, K673, S675, I679, L680, L683, N690, R691, F693, S701, F704, Q712, G715, Q716, H723, I724, L727, I733, L738, Q739, N803, Q805, Q807, K810, Y812, D8 29. N831, R832, S834, D835, Q844, S845, K848, R859, K862, R864, K866, K890, T893, Q894, D898, N899, K902, K913, K918, Q920, T924, R925, T928, K929, H930, S960, K961, S964, K968, R976, H982, H983, Y1013, K1031, T1033, SI106, K1107, S1109, Y1237, Y1242, K1244 and K1246, wherein the amino acid residue positions correspond to the amino acid numbering in the Streptococcus pyogenes Cas9 (SpCas9) protein amino acid sequence SEQ ID NO: 68 or are defined with reference to the amino acid numbering in the Streptococcus pyogenes Cas9 (SpCas9) protein amino acid sequence SEQ ID NO: 68.
[0012] Item 2. The modified Cas9 nuclease or DNA-binding fragment thereof according to Item 1, comprising one or more mutations selected from the group consisting of K526X, Q695X, H698X, and R691X, wherein X is glycine (G), alanine (A), valine (V), isoleucine (I), leucine (L), aspartic acid (D), glutamic acid (E), asparagine (N), glutamine (Q), serine (S), threonine (T), lysine (K), arginine (R), phenylalanine (F), or tyrosine (Y).
[0013] Item 3. The modified Cas9 nuclease or DNA-binding fragment thereof according to Item 1 or 2, wherein each X is independently selected from any one of the following: glycine (G), alanine (A), aspartic acid (D) and glutamic acid (E).
[0014] Item 4. The modified Cas9 nuclease or DNA-binding fragment thereof according to Item 3, comprising any combination of mutations selected from the group consisting of:
[0015] K526A+R691A+Q695A+H698A, K526A+R691A+N692A+Q695A+H698A, K526G+R691G+Q695G+H698G, K526D+R691A+Q695A+H698A, K526A+R69 1D+Q695A+H698A, K526A+R691A+Q695A+H698D, K526E+R691A+Q695A+H698A, K526A+R691E+Q695A+H698A and K526A+R691A+Q695A+H698E.
[0016] Item 5. The modified Cas9 nuclease or DNA-binding fragment thereof according to any one of Items 1 to 4, which is a Cas9 nucleic acid double-strand nickase, nCas9 (Cas9 nickase), or dCas9 (Catalytically dead Cas9).
[0017] Item 6. The modified Cas9 nuclease or DNA-binding fragment thereof according to any one of Items 1-5, further comprising one or more mutations in the RuvC domain and / or the HNH domain, optionally comprising a mutation from aspartic acid at position 10 to alanine (D10A) in the RuvC domain and / or a mutation from histidine at position 840 to alanine (H840A) in the HNH domain.
[0018] Item 7. A modified Cas9 nuclease or DNA-binding fragment thereof according to any one of Items 1 to 6, wherein the amino acid sequence of the Cas9 nuclease is as shown in SEQ ID NO.69, SEQ ID NO.78 or SEQ ID NO.81, or comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the amino acid sequence shown in SEQ ID NO.69, SEQ ID NO.78 or SEQ ID NO.81.
[0019] Item 8. The Cas9 nuclease or DNA-binding fragment thereof according to any one of Items 1 to 7, wherein the DNA-binding fragment does not comprise one or more amino acid segments selected from the group consisting of:
[0020] The amino acid segments at positions 494-501, 179-296, 503-708, 792-897, and 1010-1081, wherein the amino acid positions correspond to the amino acid numbers in the SpCas9 protein amino acid sequence SEQ ID NO: 68,
[0021] Optionally, the DNA binding fragment is a DNA binding fragment contained in SEQ ID NO.69, SEQ ID NO.78 or SEQ ID NO.81, wherein the DNA binding fragment does not contain one or more amino acid segments selected from the following group in these sequences:
[0022] Amino acid segments 494-501, 179-296, 503-708, 792-897 and 1010-1081.
[0023] Item 9. A fusion protein comprising the Cas9 nuclease or DNA binding fragment thereof according to any one of Items 1-8.
[0024] Item 10. The fusion protein according to Item 9, further comprising a nuclear localization signal peptide on the N-terminal side and / or the C-terminal side of the Cas9 nuclease or its DNA-binding fragment.
[0025] Item 11. The fusion protein according to Item 9 or 10, further comprising a cytosine deaminase, an adenine deaminase, an oxidase, a glycosidase, an alkyltransferase, a DNA synthetase, an RNA synthetase, a uracil glycosylase inhibitor (UGI), a transcription activator, a transcription repressor, a methylase, or a demethylase fused to the Cas9 nuclease or its DNA binding fragment, a Gam protein derived from bacteriophage Mu, and / or a fluorescent protein; optionally, the transcription activator includes VP64, and optionally, the transcription repressor includes a KRAB protein.
[0026] Item 12. The fusion protein according to Item 11, which comprises or is, from N-terminus to C-terminus:
[0027] 1) Nuclear localization signal peptide-Cas9 nucleic acid double-strand cleavage enzyme, nCas9 or dCas9 or its DNA binding fragment-nuclear localization signal peptide;
[0028] 2) Nuclear localization signal peptide-TadA* enzyme-nCas9-nuclear localization signal peptide;
[0029] 3) TadA* enzyme-nCas9;
[0030] 4)Dnmt3A-Dnmt3L-dCas9-KRAB; or
[0031] 5) Dnmt3A-Dnmt3L-dCas9-nuclear localization signal peptide-KRAB;
[0032] Wherein “-” indicates connection through a peptide bond or linker.
[0033] Item 13. The fusion protein according to Item 12, wherein:
[0034] Each of the linkers independently comprises or is any one or more amino acid sequences selected from the group consisting of SEQ ID NO.77, SEQ ID NO.85, SEQ ID NOs.88-91, or an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto;
[0035] The amino acid sequence of the nuclear localization signal peptide independently comprises or is any one selected from the following: SEQ ID NO.70, SEQ ID NO.71, SEQ ID NO.74, SEQ ID NO.75, optionally, the amino acid sequence as shown in SEQ ID NO.71 or 74 is located on the N-terminal side, and / or the amino acid sequence as shown in SEQ ID NO.70, 75 or 85 is located on the C-terminal side;
[0036] The amino acid sequence of the TadA* enzyme comprises or is the amino acid sequence shown in SEQ ID NO.76;
[0037] The amino acid sequence of KRAB comprises or is the amino acid sequence shown in SEQ ID NO.87;
[0038] The amino acid sequence of the DNMT3A comprises or is the amino acid sequence shown in SEQ ID NO.87; and / or
[0039] The amino acid sequence of DNMT3L comprises or is the amino acid sequence shown in SEQ ID NO.87.
[0040] Item 14. A fusion protein according to any one of Items 11-13, whose amino acid sequence is as shown in SEQ ID NO.66, SEQ ID NO.73 or SEQ ID NO.80, or comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the amino acid sequence shown in SEQ ID NO.66, SEQ ID NO.73 or SEQ ID NO.80.
[0041] Item 15. An engineered nucleic acid molecule comprising a nucleotide sequence encoding a modified Cas9 nuclease or DNA-binding fragment thereof according to any one of items 1-8 or a fusion protein according to any one of items 9-14.
[0042] Item 16. The nucleic acid molecule according to Item 15, comprising a 5' non-coding region sequence (5'UTR) and a 3' non-coding region sequence (3'UTR); wherein the 5'UTR comprises or is the 5'UTR encoding the tobacco phagocytic virus gene, and / or the 3'UTR comprises or is the 3'UTR encoding the human hemoglobin alpha 1 (hHBA1) gene.
[0043] Item 17. A nucleic acid molecule according to Item 15 or 16, wherein the nucleotide sequence of the 5'UTR is as shown in SEQ ID NO.10 or 11, or comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% nucleotide sequence identity with the nucleotide sequence as shown in SEQ ID NO.10 or 11, and / or the nucleotide sequence of the 3'UTR is as shown in SEQ ID NO.13, or comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the nucleotide sequence as shown in SEQ ID NO.13.
[0044] Item 18. The nucleic acid molecule according to any one of Items 15-17, further comprising a Poly (A) tail sequence or a tailing signal sequence, preferably, the Poly (A) tail sequence comprises a nucleotide sequence as shown in SEQ ID NO: 12 or comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with SEQ ID NO. 12.
[0045] Item 19. The nucleic acid molecule according to any one of Items 15 to 18, which is an mRNA and comprises a 5' cap structure, optionally wherein the 5' cap structure is m7G(5')ppp(5')(2'OMeA)pG.
[0046] Item 20. A nucleic acid molecule according to any one of Items 15-19, whose nucleotide sequence is as shown in SEQ ID NO: 2, 6 or 8, or comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the nucleotide sequence as shown in SEQ ID NO: 2, 6 or 8.
[0047] Item 21. A nucleic acid molecule according to any one of Items 15-20, which is an mRNA molecule comprising one or more modified bases of uridine (U), optionally wherein the modified base U is 1-methylpseudouridine; optionally, each U in the nucleic acid molecule encoding the Cas9 nuclease or an enzymatically active fragment thereof is 1-methylpseudouridine.
[0048] Item 22. A composition, ribonucleoprotein complex or protein lipid complex comprising a modified Cas9 nuclease or DNA binding fragment thereof according to any one of items 1-8, or a fusion protein according to any one of items 9-14, or a nucleic acid molecule according to any one of items 15-21.
[0049] Item 23. The composition, ribonucleoprotein complex or protein lipid complex according to Item 22, further comprising a gRNA targeting a target gene, a nucleic acid molecule encoding the gRNA, or a construct comprising the gRNA.
[0050] Item 24. A gene editing composition according to Item 23, wherein the target gene is selected from any one or more of the following: hepatitis B (HBV) gene, PCSK9, EMX1 and VEGFA3.
[0051] Item 25. A gene editing method, comprising introducing into a host cell the modified Cas9 nuclease or its DNA binding fragment described in any one of Items 1-8, or the fusion protein described in any one of Items 9-14, or the nucleic acid molecule described in any one of Items 15-21, or the composition, ribonucleoprotein complex or protein lipid complex described in any one of Items 22-24.
[0052] Item 26. Use of the modified Cas9 nuclease or its DNA binding fragment described in any of Items 1-8, or the fusion protein described in any of Items 9-14, or the nucleic acid molecule described in any of Items 15-21, or the composition, ribonucleoprotein complex or protein lipid complex described in any of Items 22-24 in the preparation of a medicament for treating a disease or condition in need thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1: Bioanalyzer analysis of the integrity of the mRNA encoding the Cas9 protein.
[0054] Figure 2: Results of the normal translation test of Cas9 mRNA in HEK293T cells.
[0055] Figure 3: HPLC-RP analysis of sgTTR-0 purity.
[0056] Figure 4: T7E1 enzyme digestion reaction (A) and ELISA (B) respectively detect the editing efficiency and protein knockdown effect of different gRNAs targeting the TTR gene. A shows the cleavage effect of T7E1 enzyme digestion; B shows the ELISA detection of TTR protein content in the supernatant of HepG2 cells.
[0057] Figure 5: GUIDE-seq analysis of off-target rates for sgTTR-0 and sgTTR-1. A shows the number of off-targets detected for sgTTR-0 and sgTTR-1 in combination with SpCas9 mRNA. B shows the distribution of chromosomal cleavage sites for the sgTTR-0 and Cas9 mRNA combinations. C shows the distribution of chromosomal cleavage sites for the sgTTR-1 and Cas9 mRNA combinations.
[0058] Figure 6: T7E1 enzyme digestion reaction to detect the cutting efficiency of TTR gene with gRNAs of different modifications and mutations.
[0059] Figure 7: Off-target rate (safety) test of Cas9-mut5 using GUIDE-seq. A shows the specific off-target sites for different Cas9 mutants and sgTTR-0 combinations as measured by GUIDE-seq; B shows the off-target frequency (off-target reads / on-target reads) for different Cas9 mutants and sgTTR-0 combinations; C shows the potential off-target sites for Cas9-mut5 and Cas9-WT across the entire genome.
[0060] Figure 8: Gene editing effects of different SpCas9 mutants combined with sgTTR-0. A shows the reduction in TTR protein in mouse serum one week after administration (1 mg / kg) as determined by ELISA; B shows the TTR gene editing efficiency in mouse liver one month after administration (1 mg / kg) as determined by amplicon sequencing; C shows the knockdown effect of sgTTR-0+SpCas9-mut5 on TTR protein in humanized mouse serum at different doses.
[0061] Figure 9: Gene editing efficiency of sgTTR-0 and SpCas9 mRNA using different LNP delivery formulations. A shows the changes in TTR protein levels in mouse serum measured by ELISA one week after dosing (0.3 mg / kg); B shows the TTR gene editing efficiency in the liver measured one month after dosing (0.3 mg / kg); C shows the changes in TTR protein levels in mouse serum measured by ELISA one week after dosing (1 mg / kg) using different delivery formulations.
[0062] Figure 10: Gene editing effects of sgTTR-0 and SpCas9 mRNA combinations at different mass ratios. A shows TTR protein levels in mouse serum measured by ELISA one week after dosing; B shows TTR gene editing efficiency in mouse liver measured by amplicon sequencing one month after dosing.
[0063] Figure 11: GUIDE-seq analysis of safety studies at different gRNA to mRNA mass ratios. A shows the number of off-target sites at different gRNA to mRNA mass ratios as measured by GUIDE-seq; B shows the statistically determined off-target frequency (number of off-target reads / number of on-target reads) at different gRNA to mRNA mass ratios.
[0064] Figure 12: Editing efficiency of different Cas9 mutants at the TTR locus in HEK293T cells and HepG2 cells. Panel A shows the editing efficiency of different Cas9 mutants at the TTR locus in HEK293T cells; Panel B shows the editing efficiency of different Cas9 mutants at the TTR locus in HepG2 cells; Panel C shows the knockdown efficiency of different Cas9 mutants at the TTR protein in HepG2 cells.
[0065] FIG13 shows the editing efficiency of SpCas9-WT and its mutants at four different off-target sites (off-target site 1 / 2 / 3 / 4, shown in A to C, respectively) under the guidance of sgTTR-0.
[0066] FIG14 shows the editing efficiency of SpCas9-WT and its mutants at four target sites of the TRAC gene (guided by sgRNA-1 / 2 / 3 / 4, respectively, shown in A to D).
[0067] Figure 15: Editing efficiency of SpCas9-WT and its mutants at two target sites (A: EMX1; B: VEGFA3), and off-target rates at corresponding high-frequency off-target sites.
[0068] Figure 16: Schematic diagram of the structures of ABE8e and ABE8e-Mut5. A is the ABE8e reported in the literature (i.e., ABE8e-WT in the examples of this application), and B is ABE8e-Mut5.
[0069] Figure 17: Next-generation sequencing identification of the on-target editing efficiency of ABE8e-mut5 and ABE8e-WT for 7 gRNAs (A: EMX1; B: VEGFA3; C: TTR-ABE-1; D: HEK site 1; E: HEK site 2; F: HEK site 3; G: HEK site 4) and the editing efficiency of the corresponding high-frequency off-target sites
[0070] Figure 18: Next-generation sequencing identified the editing efficiency of ABE8e-mut5 and ABE8e-WT within and outside the editing window of six sites on the genome (A: EMX1; B: VEGFA3; C: HEK293-2; D: PCSK9; E: F-site2; F: ATGg2).
[0071] Figure 19: Schematic diagram of the structures of EE-WT (A: CRISPRoff-EE-WT) and EE-mut5 (B: CRISPRoff-EE-Mut5).
[0072] Figure 20: Inhibitory efficiency of EE-mut5 and EE-WT on PCSK9 protein in different liver cancer cells (A: HepG2, B: Huh7).
[0073] Figure 21: Inhibitory effects of EE-mut5 and EE-WT on HBV DNA (A) and related proteins (B: HBsAg, C: HBeAg) in HepG2.2.15 cells, as well as their effects on cell viability (D).
[0074] Figure 22: The number of differentially methylated sites detected between EE-mut5 and EE-WT in the genome of HepG2.2.15 cells. DETAILED DESCRIPTION
[0075] definition:
[0076] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as understood by one of ordinary skill in the art.The abbreviations for amino acid residues are the standard three-letter and / or one-letter codes used in the art to refer to one of the 20 common L-amino acids.
[0077] As used herein, the term "Cas9 nuclease" refers to CRISPR associated protein 9, which is the same as understood by those skilled in the art, and can bind to and cut double-stranded DNA target sites under the guidance of gRNA. Cas9 nucleases can be modified, such as by amino acid mutations, deletions, or additions, to have higher editing efficiency, higher editing specificity, lower off-target rates, loss of the ability to cut one or both strands of DNA, and the like. In some embodiments of the present application, a "modified Cas9 nuclease" is a double-stranded cutting enzyme with higher editing specificity or lower off-target efficiency. In some embodiments, a "modified Cas9 nuclease" is a nickase (Cas9 nickase or nCas9) that cuts one strand of DNA. In some embodiments, a "modified Cas9 nuclease" loses the enzyme's cleavage activity on the DNA chain, but still retains the ability to target and bind to DNA under the guidance of gRNA. In this case, it is referred to as catalytically dead Cas9, dead Cas9, or dCas9.
[0078] As used herein, "coding sequence" may refer to a ribonucleotide sequence in a mature mRNA that can be translated into a protein, or may refer to the complementary sequence of a deoxyribonucleotide (DNA) sequence that serves as a template for transcribing the ribonucleotide (RNA) sequence. Furthermore, the "coding sequence" of the present application may further include polynucleotide sequences encoding functional nucleic acids, such as miRNA, shRNA, dsRNA, and the like.
[0079] In this application, "N-terminal side" is used to describe the relative positional relationship between two sequences, an amino acid and a sequence, or two amino acids in the same amino acid sequence. Among them, "N-terminal" refers to the end of the amino acid sequence that contains a free amino group. For example, "the N-terminal side of the Cas9 nuclease or its DNA-binding fragment further contains a nuclear localization signal peptide", which means that the "nuclear localization signal peptide" is closer to the N-terminus of the amino acid sequence in which they are co-located relative to "as9 nuclease or its DNA-binding fragment". Similarly, "C-terminal side" is also used to describe the relative positional relationship between two sequences, an amino acid and a sequence, or two amino acids in the same amino acid sequence. Among them, "C-terminal" refers to the end of the amino acid sequence that contains a free carboxyl group. For example, "the C-terminal side of the Cas9 nuclease or its DNA-binding fragment further contains a nuclear localization signal peptide", which means that the "nuclear localization signal peptide" is closer to the C-terminus of the amino acid sequence in which they are co-located relative to "Cas9 nuclease or its DNA-binding fragment". The sequence or amino acid located on the N-terminal side or C-terminal side of a sequence or an amino acid can be directly connected to the sequence or an amino acid, or separated by one or more amino acid residues.
[0080] Although the numerical ranges and parameter approximations shown in the broad scope of this application, the numerical values shown in the specific examples are recorded as accurately as possible. However, any numerical value is necessarily contained in a certain error, which is caused by the standard deviation present in their respective measurements. In addition, all ranges disclosed herein should be understood to cover any and all sub-ranges contained therein. For example, a range of "1 to 10" should be considered to include any and all sub-ranges between a minimum of 1 and a maximum of 10 (including endpoints); that is, all sub-ranges starting with a minimum of 1 or greater, such as 1 to 6.1, and sub-ranges ending with a maximum of 10 or less, such as 5.5 to 10. In addition, any reference referred to as "incorporated herein" should be understood to be incorporated in its entirety.
[0081] Those skilled in the art will appreciate that, due to the degeneracy of the genetic code, many different polynucleotides can encode the same polypeptide. It will also be understood that the skilled artisan can use conventional techniques to make nucleotide substitutions that do not affect the polypeptide sequence encoded by the nucleic acid molecule to reflect the codon usage of any particular host organism in which the polypeptide is expressed. Therefore, unless otherwise indicated, "polynucleotides encoding the protein or immunogenic fragment of the present application" include all polynucleotide sequences that are degenerate to each other and encode the same amino acid sequence.
[0082] The term "ribonucleoprotein" (RNP) or "RNP complex" refers to a guide RNA and an RNA-guided DNA binder, such as a Cas nuclease, e.g., a Cas cleavage enzyme, a Cas nickase, or a dCas DNA binder (e.g., Cas9). In some embodiments, the guide RNA guides the RNA-guided DNA binder, such as Cas9, to a target sequence, and the guide RNA hybridizes to the target sequence and the agent binds to the target sequence; where the agent is a cleavage enzyme or nickase, binding may be followed by cleavage or nicking.
[0083] The term "TTR", or transthyretin, refers to transthyretin, the gene product of the TTR gene, also known as vitamin A-binding protein. It is an important component of plasma proteins and is widely distributed in various cells, plasma, and tissue fluids. As a carrier protein, TTR is mainly synthesized in the liver and the choroid plexus in the brain, secreted into the blood and cerebrospinal fluid, and carries thyroxine and retinol (vitamin A) to various tissues and cells throughout the body. As a carrier protein, the function of TTR can often be replaced by thyroxine-binding globulin and albumin in plasma. Under physiological conditions, TTR is a stable protein. When TTR breaks down into monomers, it can cause amyloidosis. In recent years, more and more TTR-targeted drugs are gradually entering the clinic for use in neurological diseases, endocrine and metabolic diseases, and so on.
[0084] The terms "nuclear localization signal," "NLS," "nuclear localization signal peptide," or "nuclear localization sequence" refer to an amino acid sequence or peptide that induces transport of a molecule comprising or linked to such a sequence into the nucleus of a eukaryotic cell. The nuclear localization signal can form part of the molecule to be transported. In some embodiments, the NLS can be attached to the molecule via a covalent bond, a hydrogen bond, or an ionic interaction.
[0085] Nucleic Acids
[0086] The term "nucleic acid" or "nucleic acid molecule" will be recognized and understood by those of ordinary skill in the art. As used herein, the term "nucleic acid" or "nucleic acid molecule" preferably refers to a DNA (molecule) or RNA (molecule). It is preferably used synonymously with the term polynucleotide. Preferably, a nucleic acid or nucleic acid molecule is a polymer comprising or consisting of nucleotide monomers, which are covalently linked to each other via phosphodiester bonds of a sugar / phosphate backbone. The term "nucleic acid molecule" also includes modified nucleic acid molecules, such as base-modified, sugar-modified, or backbone-modified DNA or RNA molecules as defined herein.
[0087] Unless otherwise specified, "nucleotide" herein refers not only to naturally occurring ribonucleotides or deoxyribonucleotide monomers, but also to their related structural variants, including derivatives and analogs, which are functionally equivalent in the specific context in which the nucleotide is used, unless the context clearly indicates otherwise. For example, "nucleotide" refers to a deoxyribonucleotide or a ribonucleotide. Nucleotides can be standard nucleotides (i.e., adenosine (A), guanosine (G), cytidine (C), thymidine (T) and uridine (U), nucleotide isomers or nucleotide analogs, for example, U or T in the nucleotide sequence used to represent mRNA can represent natural uridine and pseudouridine, such as 1-methylpseudouridine. Nucleotide analogs refer to nucleotides with modified purine or pyrimidine bases or modified ribose moieties. Nucleotide analogs can be naturally occurring nucleotides (e.g., inosine, pseudouridine, etc.) or non-naturally occurring nucleotides. Non-limiting examples of modifications on the sugar or base portion of the nucleotide include the addition (or removal) of acetyl, amino, carboxyl, carboxymethyl, hydroxyl, methyl, phosphoryl and thiol groups, and substitution of the carbon and nitrogen atoms of the base with other atoms (e.g., 7-deazapurine). Nucleotide analogs also include dideoxynucleotides, 2'-O-methyl nucleotides, locked nucleic acids (LNA), peptide nucleic acids (PNA) and morpholino oligonucleotides.
[0088] In some embodiments, the nucleic acid comprises at least one heterologous untranslated region (UTR). The term "untranslated region" or "UTR" or "UTR element" will be recognized and understood by those of ordinary skill in the art to mean a portion of a nucleic acid molecule, typically located 5' or 3' to a coding sequence. The 5' end is referred to as a 5'UTR, and the 3' end is referred to as a 3'UTR. Generally speaking, UTRs are not translated into proteins; UTRs can be part of a nucleic acid, such as DNA or RNA. UTRs can comprise elements for controlling gene expression, also referred to as regulatory elements. Such regulatory elements can be ribosome binding sites, miRNA binding sites, etc.; RNA (e.g., mRNA) can further comprise a 5'UTR, a 3'UTR, a 3'-poly(A) and / or a 5' cap analog.
[0089] In some embodiments, the 5'UTR is a heterologous UTR, i.e., a UTR found in nature that is associated with a different ORF; in another embodiment, the 5'UTR is a synthetic UTR; the 5'UTR is a region of the mRNA that is located upstream (5') of the start codon (the first codon of the mRNA transcript translated by the ribosome). The 5'UTR does not encode a protein. The natural 5'UTR has characteristics that play a role in translation initiation, such as the Kozak sequence, which has a consensus CCR(A / G)CCAUGG; exemplary 5'UTRs also include tobacco etch virus, African clawed frog or human α-globin or β-globin, human cytochrome b-245a polypeptide, hydroxysteroid (17b) dehydrogenase, and alpha-1-globin 5'UTR, etc.
[0090] In some embodiments, the 3'UTR can be heterologous or synthetic; for example: HBA1 (human Hemoglobin Subunit Alpha 1), globin UTR, including African clawed frog β-globin UTR and human β-globin UTR; other 3'UTRs can also be CYBA (cytochrome b-245alpha chain), rabbit β-globin, hepatitis B virus (HBV), α-globin 3'UTR and VEEV (Venezuelan equine encephalitis virus) virus 3'UTR sequences. In some embodiments, rps9 (Ribosomal Protein S9) 3'UTR, FIG4 (FIG4 Phosphoinositide 5-Phosphatase), gp130, DH143 and human albumin hHBB (human hemoglobin subunit beta) 3'UTR can also be used.
[0091] In some embodiments, the 3'-poly (A) tail is also called a poly (A) tail; the poly (A) tail is located downstream of the 3' UTR, for example, the mRNA region immediately downstream (i.e., 3'), which contains multiple consecutive adenosine monophosphates. The poly (A) tail may contain 10 to 300 adenosine monophosphates, and may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 210, 220, 230, 240, 250, 260, 270, 280, 290 or 300 adenosine monophosphates. In some preferred embodiments, the poly (A) tail contains 50 to 250 adenosine monophosphates, more preferably 50-100 adenosine monophosphates; most preferably 100 adenosine monophosphates; in relevant biological environments (e.g., in cells, in vivo), the function of the 3'-poly (A) tail is to protect the mRNA from enzymatic degradation, for example in the cytoplasm, and to facilitate transcription termination and / or export of the mRNA from the nucleus and translation.
[0092] In some embodiments, the RNA (e.g., mRNA) further comprises a 5' guanosine cap; the 5' guanosine cap is a eukaryotic mRNA transcript, and the 5' cap is composed of an inverted 7-methylguanosine, connected to the rest of the eukaryotic mRNA via a 5'-5' triphosphate bridge, namely cap 0 (cap 0), which primarily serves as a quality control for correct mRNA processing and helps stabilize the eukaryotic mRNA; based on cap 0, the first nucleotide is methylated with 2'-OH, referred to as cap 1 (cap 1); in addition to cap 0 and cap 1, the second nucleotide can also be further methylated, referred to as cap 2; generally speaking, the 5'-cap can be synthesized by different synthetic routes of 5'-capped mRNA based on enzymatic, chemical, or chemoenzymatic methods;
[0093] In some embodiments, during in vitro transcription, a cap analog is directly added to the in vitro transcription (IVT) system, and the 5' cap analog includes but is not limited to: m 7 Gppp(2'OMeA)pG、m 7 GpppApA、m 7 GpppApC、m 7 GpppApG、m 7 GpppApU、m 7 GpppCpA、m 7 GpppCpC、m 7 GpppCpG、m 7 GpppCpU、m 7 GpppGpA、m 7 GpppGpC、m7 GpppGpG、m 7 GpppGpU、m 7 GpppUpA、m 7 GpppUpC、m 7 GpppUpG、m 7 GpppUpU、m 7 Gpppm 6 ApG, m 7 G 3’Ome pppApA、m 7 G 3’Ome pppApC、m 7 G 3’Ome pppApU、m 7 G 3’Ome pppApG、m 7 G 3’Ome pppCpA、m 7 G 3’Ome pppCpC、m 7 G 3’Ome pppCpG、m 7 G 3’Ome pppCpU、m 7 G 3’Ome pppUpA、m 7 G 3’Ome pppUpC、m 7 G 3’Ome pppUpG、m 7 G 3’Ome pppUpU、m 7 G 3’Ome pppA 2’Ome pG、m 7 G 3’Ome pppA 2’Ome pC、 m 7 G 3’Ome pppA 2’Ome pU、m 7 G 3’Ome pppA 2’Ome pA、m 7 G 3’Ome pppC 2’Ome pA、m 7 G 3’Ome pppC 2’Ome pU、m 7 G 3’Ome pppC 2’Ome pG、m 7 G 3’Ome pppC 2’Ome pC、m 7 G 3’OmepppG 2’Ome pA, m 7 G 3’Ome pppG 2’Ome pU、m 7 G 3’Ome pppG 2’Ome pG、m 7 G 3’Ome pppG 2’Ome pC、m 7 G 3’Ome pppU 2’Ome pA, m 7 G 3’Ome pppU 2’Ome pU、m 7 G 3’Ome pppU 2’Ome pG、m 7 G 3’Ome pppU 2’Ome pC etc.
[0094] In some embodiments, the capped analogs may also be other structures, such as tetramers, pentamers, hexamers, heptamers, octamers, nonamers, or decamers, etc. The specific sequence thereof may be determined according to the conditions of the template.
[0095] As used herein, "mRNA" (messenger RNA) is any RNA of a naturally occurring, non-naturally occurring, or modified amino acid polymer that encodes at least one protein and can be translated to produce the encoded protein in vitro, in vivo, in situ, or ex vivo. It will be appreciated by those skilled in the art that, unless otherwise indicated, the polynucleotide sequences described herein may use "T" to refer to thymine when representing a DNA sequence, but when the polynucleotide sequence represents RNA (e.g., mRNA), the "T" will be replaced by "U" (uracil). Thus, any DNA disclosed and identified by a particular sequence number (SEQ ID NO) herein also discloses an RNA (e.g., mRNA) sequence that is complementary or corresponding to the DNA, wherein each "T" of the DNA sequence is replaced by a "U."
[0096] open reading frame
[0097] An open reading frame (ORF) is a continuous stretch of DNA or RNA that begins with a start codon (ATG or AUG, which will be translated into, for example, methionine) and ends with a stop codon (e.g., TAA, TAG, or TGA, or UAA, UAG, or UGA). Generally speaking, an ORF typically encodes a protein. It will be understood that the sequences disclosed herein may also contain additional elements, such as 5' and 3' UTRs, but unlike ORFs, these elements are not necessarily present in the RNA polynucleotides of the present application.
[0098] In some embodiments, the composition comprises RNA (eg, mRNA) comprising a nucleotide sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100% identical to SEQ ID NO. 2.
[0099] In some embodiments, the open reading frame is preferably at least partially codon optimized. Codon optimization is based on such discovery: translation efficiency can be determined by the different frequencies of transfer RNA (tRNA) occurring in the cell. Therefore, if there is so-called "rare codon" of increasing degree in the coding region of the nucleic acid of the application defined herein, the translation of the corresponding modified nucleotide sequence is less efficient than when there is a codon encoding relatively "common" tRNA. Those skilled in the art can carry out codon optimization for sequence to be translated with the characteristics of its in vitro expression system.
[0100] Chemically modified or unmodified nucleotides
[0101] In some embodiments, RNA (e.g., mRNA) is not chemically modified, but rather comprises standard ribonucleotides consisting of adenosine, guanosine, cytosine, and uridine. In some embodiments, the nucleotides and nucleosides disclosed herein comprise standard nucleoside residues, such as those present in transcribed RNA (e.g., A, G, C, or U). In some embodiments, the nucleotides and nucleosides disclosed herein include standard deoxyribonucleosides, such as those present in DNA (e.g., dA, dG, dC, or dT);
[0102] In some embodiments, the nucleotides and nucleosides of the present application include modified nucleotides or nucleosides. Such modified nucleotides and nucleosides can be naturally occurring modified nucleotides and nucleosides or non-naturally occurring modified nucleotides and nucleosides. Such modifications can include the sugar of nucleotides and / or nucleosides well known in the art, the modification of the backbone or core base moiety.
[0103] In some embodiments, the modified nucleic acid base in the nucleic acid (e.g., RNA nucleic acid, e.g., mRNA nucleic acid) includes 1-methyl-pseudouridine, 1-ethyl-pseudouridine, 5-methoxy-uridine, 5-methyl-cytidine and / or pseudouridine, pseudouridine.
[0104] In vitro transcription system (IVT)
[0105] In vitro transcription is the process of using DNA as a template in an in vitro cell-free system containing components such as RNA polymerase and NTP to mimic the in vivo transcription process to generate mRNA. Generally speaking, the capped RNA synthesized in the in vitro transcription reaction can be used for subsequent experiments such as microinjection, in vitro translation, and transfection. The in vitro transcription system usually includes a transcription buffer, nucleotide triphosphates (NTPs), an RNase inhibitor, and a polymerase. NTPs can be synthesized by oneself or selected from a supplier. NTPs can be natural or non-natural NTPs. Optional polymerases include, but are not limited to, phage RNA polymerases, such as T7 RNA polymerase, T3 RNA polymerase, SP6 RNA polymerase, and / or polymerase mutants thereof, such as, but not limited to, polymerases capable of incorporating modified nucleic acids and / or modified nucleotides, including chemically modified nucleic acids and / or nucleotides. Some embodiments exclude the use of DNA enzymes. In some embodiments, the RNA contains a 5' guanosine cap.
[0106] In addition to synthesis by in vitro transcription systems, chemical synthesis methods can also be used, including solid-phase chemical synthesis and liquid-phase chemical synthesis; with respect to solid-phase chemical synthesis, the nucleic acids disclosed in this application can be prepared in whole or in part using solid-phase technology; solid-phase chemical synthesis of nucleic acids is an automated method in which molecules are fixed on a solid support and synthesized stepwise in a reactant solution. Solid-phase synthesis can be used for site-specific introduction of chemical modifications in nucleic acid sequences; with respect to liquid-phase chemical synthesis, the nucleic acids of this application can be synthesized in liquid phase by sequentially adding monomer constructs. In addition, the above-mentioned synthesis methods can also be used in combination, because the synthesis methods discussed above each have their own advantages and limitations, and attempts can be made to combine these methods to overcome the above-mentioned limitations. Combinations of these methods are within the scope of this application.
[0107] The term "identity" refers to the relationship between the sequences of two or more polypeptides (e.g., antigens) or polynucleotides (nucleic acids) determined by comparing sequences. Identity also refers to the degree of sequence relatedness between or among sequences determined by the number of matches between strings of two or more amino acid residues or nucleic acid residues. Identity measures the percentage of identical matches between the smaller of two or more sequences, where gap comparisons (if any) are solved by a specific mathematical model or computer program (e.g., an "algorithm"). The identities of the related antigens or nucleic acids can be easily calculated by known methods. "Percentage (%) identity" for polypeptide or polynucleotide sequences is defined as the percentage of residues (amino acid residues or nucleic acid residues) in a candidate amino acid or nucleic acid sequence that are identical to the residues in the amino acid sequence or the nucleic acid sequence of a second sequence after aligning the sequences and introducing gaps, if necessary, to obtain maximum percentage identity. Methods and computer programs for comparison are well known in the art. It is understood that identity depends on the calculation of percentage identity, but its value may vary due to gaps and penalties introduced in the calculation. Typically, variants of a particular polynucleotide or polypeptide have 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to a particular reference polynucleotide or polypeptide as determined by the sequence alignment programs and parameters described herein and known to those of skill in the art.
[0108] Lipid nanoparticles (LNPs)
[0109] The RNA of the present application (for example, mRNA, gRNA etc.) can be formulated in lipid nanoparticles (LNP).Lipid nanoparticles generally include ionizable cationic lipids, helper lipids, cholesterol and PEG lipid components and nucleic acid of interest. The lipid nanoparticles of the present application can use components, compositions and methods generally known in the art to generate.
[0110] pharmaceutical preparations
[0111] Provided herein are compositions (e.g., pharmaceutical compositions), methods, kits, and reagents for performing gene modification or editing in human and other mammalian cells.
[0112] The term "pharmaceutical composition" refers to the combination of an active agent and an inert or active carrier such that the composition is particularly suitable for in vivo or in vitro diagnostic or therapeutic use. A "pharmaceutically acceptable carrier" does not cause undesirable physiological effects upon administration to or after administration to a subject. The carrier in a pharmaceutical composition must be "acceptable" in the sense that it is compatible with the active ingredient and capable of stabilizing it. One or more solubilizing agents may be used as pharmaceutical carriers for delivering the active agent. Examples of pharmaceutically acceptable carriers include, but are not limited to, biocompatible carriers, adjuvants, additives, and diluents to obtain a composition that can be used as a dosage form. Examples of other carriers include colloidal silicon oxide, magnesium stearate, cellulose, and sodium lauryl sulfate. Other suitable pharmaceutical carriers and diluents, as well as pharmaceutical necessities for them, are described in Remington's Pharmaceutical Sciences.
[0113] sgRNA
[0114] In each of the embodiments of the compositions, uses, and methods described herein, the guide RNA may comprise a single RNA molecule as a "single guide RNA" or "sgRNA". The sgRNA may comprise a crRNA or portion thereof containing a guide sequence covalently linked to the tracrRNA. In some embodiments, the crRNA and tracrRNA are covalently linked via a linker. In some embodiments, the sgRNA forms a stem-loop structure via base pairing between the crRNA and the respective portions of the tracrRNA. In some embodiments, the crRNA and tracrRNA are covalently linked via one or more bonds that are not phosphodiester bonds. In some embodiments, the approximately 20 bases at the 5' end of the sgRNA are sequences that are complementary to the genome (referred to as spacers), and the bases of 21-100nt are sequences that interact with Cas9.
[0115] Modified Cas9 nuclease
[0116] The present application provides a technical solution to reduce the off-target efficiency of the CRISPR-Cas9 system, comprising introducing four sites into the Cas9 nuclease: 526, 691, 695, and 698 to form a modified Cas9 nuclease in which positions 526, 691, 695, and 698 are alanine (A) or its conservatively substituted amino acids, wherein the amino acid sites are numbered with reference to SEQ ID NO. 68. It should be understood that some Cas9 nucleases known in the prior art, such as nCas9, dCas9, or Cas9 double-stranded cutting enzymes, are A or its conservatively substituted amino acids at the four sites. Therefore, the solution of the present application should include the case where only one, two, or three amino acid mutations are introduced into the Cas9 nuclease before modification to form the modified Cas9 nuclease. It should be understood that conservatively substituted amino acids for A include G (glycine), D (aspartic acid), and E (glutamic acid). This application demonstrates through a large number of examples that the four mutations can reduce the off-target efficiency of Cas9 nuclease or other gene editing systems based on Cas9 nuclease (e.g., nCas9, dCas9, base editing system, epigenetic editing system) without reducing the gene editing efficiency, so that those skilled in the art can expect that introducing the above four mutations into a known functional Cas9 nuclease or a modified Cas9 nuclease or its fusion protein will still achieve its function and reduce the off-target efficiency at the same time. Based on this, the present application provides at least a modified Cas9 nuclease or a DNA-binding fragment thereof, and a fusion protein having a DNA modification function comprising the modified Cas9 nuclease or its DNA-binding fragment, or at least specifically recognizing a DNA target site under the guidance of gRNA, wherein the modified Cas9 nuclease or its DNA-binding fragment and the fusion protein are A or its conservatively substituted amino acids at positions 526, 691, 695 and 698 relative to the reference sequence SEQ ID NO.68, and the other positions are consistent with the amino acid sequence of the Cas9 nuclease (including dCas9) or its DNA-binding fragment before modification by the four mutations known in the prior art. In some embodiments, the modified Cas9 nuclease or DNA-binding fragment thereof and the fusion protein are all A at positions 526, 691, 695, and 698 relative to the reference sequence SEQ ID NO. 68. In some embodiments, the modified Cas9 nuclease is SpCas9.
[0117] Examples of Cas9 nucleases before the four mutation modifications include, for example, dCas9 and nCas9 disclosed in WO2019126709A1. Examples of DNA-binding fragments of Cas9 nucleases before the four mutation modifications include, for example, CN110241098A and WO2020005980A1. Examples of fusion proteins comprising Cas9 before the four mutation modifications include base editors, such as WO2018165629A1, WO2018213708A1, WO2018213726A1, WO2020102659A1, WO2020181195A1, WO2020181193A1, WO2020181180A1, WO2020181178A1, WO2021030666A1 , WO2017070632A3, WO2018027078A8 and WO2020181202A1; epigenetic editors, such as those disclosed in articles PMID: 34942274, 36864020, 35234927, 38418872, 38760566, 37116617 or 34580310, and WO2020014261A1; and other developed Cas9 nucleases in gene editing systems, such as those disclosed in WO2021183783A1, WO2020210751A1, WO2020191248A1, WO2014150624A1, WO2020041751A1, WO2021025750A1, WO2020182941A1, WO2020176389A1, WO2014093661A2, WO2016094874 A1, WO2016094867A1, AU2015101792A4, WO2016205613A1, WO2016205759A1, WO2018035387A1, WO2018209320A8, US11155803B2, WO2021158921A, WO2015089427A1 and WO2021183807A1, WO2018039438A1.
[0118] Example
[0119] The embodiments of the present application will be described in detail below with reference to the examples, but it will be understood by those skilled in the art that the following examples are merely illustrative of the present application and should not be considered as limiting the scope of the present application. In the examples, if no specific conditions are specified, the conditions according to conventional conditions or manufacturer recommendations are used. If the manufacturer is not specified for the reagents or instruments used, they are all conventional products that can be obtained commercially.
[0120] The SpCas9 mutants used in the following Examples 1-7 (excluding the nuclear localization signal and signal peptide portion), the amino acid mutation sites of the nCas9 and dCas9 mutants of Example 8 relative to the wild-type SpCas9 protein (SEQ ID NO.68), and the mRNA element structure encoding the mutants are shown in Table 1 below. All U (uridine) in the mRNA used in the examples of this application are 1-methyl-pseudouridine. The mRNA sequences of SpCas9-WT (SEQ ID NO.2), SpCas9-Mut5 (SEQ ID NO.2), SpCas9-HF1 (SEQ ID NO.3) and HiFi-Cas9 (SEQ ID NO.4) used in the following examples are shown in the sequence table. In the mRNA sequences of other SpCas9 cleavage enzyme variants (SpCas9-Mut1 to Mut7, and SpCas9-HF4), the codons corresponding to the amino acid mutation sites are changed accordingly. The other partial sequences are the same as the mRNA sequence of SpCas9-HF1 (SEQ ID NO.3).
[0121] Table 1: Cas9 cleavage enzyme mutant mRNA structures
[0122] It should be understood that the protein and nucleic acid sequences used in the following examples are as shown in the sequence listing, or can be obtained by introducing the mutations described in Table 1 on the basis of the sequences shown in the sequence listing.
[0123] Example 1: Cas9 mRNA quality detection
[0124] 1.1 In vitro transcription (IVT) of Cas9 mRNA
[0125] 1. According to the instructions of the IVT kit (E131, Novoprotein), the IVT reaction system was prepared by mixing 10% Transcription Buffer, ATP, GTP, CTP, 1-N-Me-Pseudo UTP (Cat. No. WA0992, Zhaowei Technology), 5' cap analog m7G(5')ppp(5')(2'OMeA)pG (Cat. No. GAGNH23C2L1B, Zhaowei Technology), water for injection, linearized plasmid template containing T7 promoter and DNA sequence encoding Cas9 mRNA (GenScript Biotech Co., Ltd.), and Enzyme Mix.
[0126] 2. The mixed reaction system was placed at 37°C for 40 minutes;
[0127] 3. Add the corresponding proportion of DNase I to terminate the reaction.
[0128] Cas9 protein mRNA was synthesized in vitro and then purified by hydrophobic chromatography and ultrafiltration. The purity of the mRNA encoding the Cas9 protein was analyzed by Bioanalyzer, confirming that high-purity mRNA (>95%) was obtained. The purity results of spCas9-WT mRNA are shown in Figure 1.
[0129] 1.2 Western blot analysis of Cas9 protein expression
[0130] HEK293T and HepG2 cells were transfected with mRNA encoding Cas9 as follows:
[0131] (1) Plating: HEK293T or HepG2 cells were digested and resuspended, and counted at 1.5×10 per well. 5 The cells were plated into 24-well plates.
[0132] (2) Transfection: 24 hours after plating, prepare two 1.5 ml EP tubes, add 250 μl Opti-MEM to each tube, add 500 ng Cas9 mRNA to EP tube No. 1 and mix well, add 2.5 μl Lipo2000 to EP tube No. 2 and mix well, then add the liquid in EP tube No. 2 to EP tube No. 1, mix well, centrifuge and let stand for 15 minutes before adding to the wells.
[0133] (3) Medium change: After 6 hours, remove the cell supernatant and add complete culture medium.
[0134] (4) Cell collection: 24 hours after mRNA transfection into HEK29T cells, the cells were collected and the expression of Cas9 protein was detected by western blot.
[0135] The results are shown in Figure 2, which shows that the Cas9 protein mRNA synthesized by the IVT step of Example 1 can be effectively expressed in HEK29T cells. It should be understood that the mRNA used in subsequent examples must undergo steps 1.1 and 1.2 before subsequent experiments can be carried out.
[0136] 1.3 gRNA chemical synthesis and purity testing
[0137] Step 1: Solid-Phase Synthesis: gRNA synthesis was performed using an automated synthesizer. Using the UnyLinker vector (Biocomma, 1000A CPG), synthesis was performed from the 3′ end to the 5′ end by sequentially coupling phosphoramidite nucleoside monomers. After the last monomer was coupled, the dimethoxytriphenyl (DMT) protecting group at the 5′ end was removed, and the protecting group (cyanoethyl) of the phosphate backbone was deprotected to obtain a solid-phase support intermediate.
[0138] Step 2: Cleavage / Aminolysis Deprotection: The solid support intermediate is cleaved from the solid support by aminolysis reaction to remove the protecting group, and the crude product is obtained after concentration and desilylation reaction.
[0139] Step 3: Chromatographic purification and ultrafiltration desalting: The crude product is purified by hydrophobic chromatography-acid hydrolysis-reverse phase chromatography to remove short chain sequences and other impurities. The collected qualified components are ultrafiltration desalted and pyrogen-free to obtain a gRNA aqueous solution without further freeze-drying.
[0140] Step 4: Aseptic filling: After the gRNA solution is sterile filtered twice, it is filled in the isolator to obtain the gRNA stock solution product.
[0141] The gRNA-related sequence structure involved in this application is as follows:
[0142] Table 2: gRNAs used in the implementation of this application
[0143] The purity of the relevant gRNA meets the requirements of subsequent experiments. The exemplary results are shown in Figure 3, which are the purity test results of sgTTR-0.
[0144] Example 2 Optimization of gRNA
[0145] 2.1 Editing efficiency test of different sRNAs on TTR gene and protein
[0146] HEK293T cells were transfected with different gRNAs (sgTTR-0, sgTTR-1, sgTTR-2, sgTTR-3, and sgTTR-7) in combination with SpCas9-WT mRNA (the mass ratio of gRNA to mRNA was 1:1). After 24 hours, the cells were lysed, and a DNA fragment of approximately 1 kb near the corresponding target site was amplified. The targeted cutting efficiency was detected by T7E1 enzyme digestion (the kit used was GeneArt™ Genomic Shearing Detection Kit, A24372, ThermoFisher).
[0147] The specific steps for detecting the targeted cutting efficiency by T7E1 enzyme cutting are as follows:
[0148] (1) Lyse the edited cells and PCR amplify the DNA bands approximately 1000 bp upstream and downstream of the editing site (i.e., design upstream primers for approximately 1000 bp upstream of the editing site, and design downstream primers for approximately 1000 bp downstream of the editing site);
[0149] (2) PCR products were purified by column and the concentration was determined by Nanodrop;
[0150] (3) Take 100 ng of purified DNA product and prepare the enzyme digestion reaction system according to the instructions. After annealing, add 1 μl of T7E1 enzyme and react at 37°C for 1 h.
[0151] (4) Agarose gel electrophoresis was used to identify the enzyme digestion results and detect the editing efficiency.
[0152] In addition, HepG2 cells were transfected with different gRNAs in combination with Cas9 mRNA (the mass ratio of gRNA to mRNA was 1:1). After 72 hours, the cell supernatant was collected and the TTR protein content in the cell supernatant was detected by ELISA. The ELISA detection steps are as follows:
[0153] (1) Sample preparation and standard preparation: The sample to be tested (cell supernatant or animal serum) is diluted in an appropriate ratio, and standard solutions with gradient concentrations are prepared according to the instructions.
[0154] (2) Sample addition: 50 μL / well was added to a 96-well ELISA plate, sealed with a sealing film, and incubated at room temperature for 1 h.
[0155] (3) Add detection antibody: After washing the plate three times with PBST, add biotinylated detection antibody at 50 μl / well, seal the plate with sealing film, and incubate at room temperature for 1 h.
[0156] (4) Add secondary antibody: After washing the plate 3 times with PBST, add secondary antibody, 50ul / well, seal the plate with sealing film, and incubate at room temperature for 30min
[0157] (5) Color development: After washing the plate three times with PBST, add TMB color development solution (50 μl / well) and develop the color at room temperature in the dark for 10-15 minutes.
[0158] (6) Termination: Add 50 μl of stop solution to each well to stop color development.
[0159] (7) Plate reading: Read the OD value using a microplate reader at 450 nm.
[0160] The above test results (Figure 4) show that sgTTR-0, sgTTR-1, sgTTR-2, sgTTR-3, and sgTTR-7 all have certain cleavage activity, among which sgTTR-0 and sgTTR-1 have better effects. Further tests will be conducted on the above two sgTTRs in the future.
[0161] 2.2 GUIDE-seq analysis of off-target rates of sgTTR-0 and sgTTR-1
[0162] The examples of this application use the GUIDE-seq method to detect off-target rates. For details, see, for example (Tsai, Zheng et al. 2015, Malinin, Lee et al. 2021).
[0163] GUIDE-seq is a method for detecting off-target effects of gene editing. Its basic principle is to use a short double-stranded oligodeoxynucleotide tag (dsODN) to mark the breaks induced by CRISPR-Cas9 (that is, after the Cas9 enzyme in the CRISPR system cuts the genome to produce DNA double-strand breaks, there is a certain probability that the dsODN will be integrated into the genome during genome repair). The sequence of the gene region where the dsODN tag is located is then amplified (that is, upstream primers upstream of the genomic break site and downstream primers targeting the genomic break site are set up for sequence amplification) to construct a library and perform high-throughput sequencing. Finally, bioinformatics analysis is used to determine the location and mutation frequency of off-target mutations.
[0164] The steps for off-target rate detection are as follows:
[0165] (1) Cell transfection. gRNA (sg-TTR-0 and sgTTR-1), wild-type SpCas9 mRNA, and double-stranded oligodeoxynucleotides (dsODN) tags were simultaneously transfected into HepG2 cells by electroporation.
[0166] (2) Genomic DNA extraction: 72 hours after transfection, genomic DNA of cells was extracted.
[0167] (3) ODN integration rate detection: T7E1 enzyme digestion was used to detect DNA double-strand cutting efficiency, and Ndel enzyme digestion was used to detect dsODN tag insertion efficiency. Subsequent detection was performed after the Ndel / T7E1 integration efficiency reached 30%.
[0168] (4) Take appropriate quality genomic DNA to construct a library and use the MGI-2000 system for sequencing; (5) Bioinformatics analysis. The analysis results are shown in Figure 5, where:
[0169] (A) Shows the off-target amounts measured when sgTTR-0 and sgTTR-1 were combined with SpCas9-WT mRNA;
[0170] (B) shows the distribution of cleavage sites on chromosomes by the combination of sgTTR-0 and Cas9 mRNA;
[0171] (C) Shows the distribution of chromosomal cleavage sites of the sgTTR-1 and Cas9 mRNA combinations.
[0172] It can be seen that different gRNAs show different off-target efficiencies, among which the off-target probability of sgTTR-0 is significantly lower than that of sgTTR-1.
[0173] 2.3 Further optimization exploration of sgTTR-0
[0174] 2.3.1 Western blot detection of TTR target protein
[0175] HEK293T cells were co-transfected with different sgTTR-0 variants (sgTTR-0-v02, sgTTR-0-v04, TTR-0-del (82 nt); and cr+tr, a combination of the crRNA and tracrRNA parts of sgTTR-0) and Cas9-HF1 mRNA. After 24 hours, the cells were lysed, and a DNA fragment of approximately 1 kb near the sgTTR-0 target site was amplified, and the cutting efficiency was detected by T7E1 enzyme digestion.
[0176] Among them, Figure 6 shows the results of the T7E1 enzyme digestion experiment, which shows that sgTTR-0-v02, sgTTR-0-v04, or splitting the sgRNA into crRNA and tracrRNA can maintain the editing effect on the TTR gene, while truncating the sgRNA to 82nt loses the editing effect on the TTR gene.
[0177] Example 3 Screening of Cas9 mutants with low off-target rates
[0178] 3.1 GUIDE-seq analysis of off-target rates (safety) of wild-type Cas9 (SpCas9-WT) and its mutants
[0179] After co-transfection of HepG2 cells with mRNA encoding SpCas9-WT and its mutants (SpCas9-WT, SpCas9-HF1, SpCas9-Mut1, SpCas9-Mut2, SpCas9-Mut3, SpCas9-Mut4, SpCas9-Mut5, and SpCas9-Mut6) and sgTTR-0, off-target efficiency was detected. The specific steps are as follows:
[0180] (1) Cell transfection. sgTTR-0, wild-type or mutant SpCas9 mRNA and double-stranded oligodeoxynucleotides (dsODN) tags were simultaneously transfected into HepG2 cells by electroporation.
[0181] (2) Genomic DNA extraction: 72 hours after transfection, genomic DNA of cells was extracted.
[0182] (3) ODN integration rate detection: T7E1 enzyme digestion was used to detect DNA double-strand cutting efficiency, and Ndel enzyme digestion was used to detect dsODN tag insertion efficiency. Subsequent detection was performed when the Ndel / T7E1 integration efficiency reached 30%.
[0183] (4) Take appropriate quality genomic DNA to construct a library and sequence it using the MGI-2000 system;
[0184] (5) Bioinformatics analysis.
[0185] The results are shown in Figure 7, where:
[0186] Part A of Figure 7 shows the specific number of off-target sites for different Cas9 mutants and sgTTR-0 combinations measured by GUIDE-seq. Among them, the combination of wild-type Cas9 and sgTTR-0 measured a total of 41 off-target sites, the reported low off-target mutant Cas9-HF1 and sgTTR-0 measured 11 off-target sites, and the combination of Cas9-mut5 and sgTTR-0 reduced the number of off-target sites to 6; Part B counted the off-target frequency (number of off-target reads / number of on-target reads) of different Cas9 mutants and sgTTR-0 combinations. The results of parts A and B of Figure 7 both show that the safety of Cas9-mut5 is not only much higher than that of wild-type Cas9, but also better than that of reported low off-target mutants and other mutants (such as Cas9-Mut1, Cas9-Mut2, Cas9-Mut3, Cas9-Mut4, and Cas9-Mut6).
[0187] Part C of Figure 7 shows the potential off-target sites of Cas9-mut5 and Cas9-WT across the entire genome.
[0188] 3.2 Detection of gene editing efficiency of different SpCas9 mutants combined with sgTTR-0
[0189] To test whether the editing efficiency of the aforementioned mutants with reduced off-target efficiency (e.g., SpCas9-Mut5) was affected by the mutation, the gene editing efficiency was tested at the TTR target site.
[0190] The mice used in this test are humanized TTR mouse models, purchased from Jiangsu Jicui Yaokang Biotechnology Co., Ltd., with strain number T055186. This model mouse model utilizes gene editing technology to replace the mouse TTR gene coding region and regulatory sequences with corresponding human TTR gene fragments. It is commonly used in transthyretin amyloidosis disease research and drug screening.
[0191] The specific test steps are as follows:
[0192] (1) Preparation of test drugs. sgTTR-0 was encapsulated in LNPs with the mRNAs of SpCas9-WT and various mutants (SpCas9-WT, SpCas9-HF1, SpCas9-HF4, SpCas9-Mut1, SpCas9-Mut2, Cas9-Mut3, Cas9-Mut4, SpCas9-Mut5, and SpCas9-Mut6) at a mass ratio of 1:1. The encapsulation method is described in Example 4 below.
[0193] (2) Animal Grouping. 6-8 week old TTR humanized mice (either sex) were used. One week before dosing, submandibular blood was collected from the mice. TTR protein expression in the serum was determined using an ELISA kit (ab231920, abcam). Mice were grouped according to the initial level of TTR expression before dosing, with 4-6 mice per group, to ensure that the initial level of TTR expression in each group was comparable.
[0194] (3) Administration. The drug was administered once via tail vein injection. The day of administration was designated as day 0. The dosage was 1 mpk (1 mg / kg), 0.3 mg / kg, or 3 mg / kg, where mg represents the total mass of the nucleic acid in the drug.
[0195] (4) Regular blood sampling. Blood was collected from the submandibular area of mice on days 4, 7, 14, 21, and 28 after administration, and serum was collected. TTR protein expression in the serum was detected using an ELISA kit (ab231920, abcam) to determine the knockdown level of TTR protein.
[0196] (5) Collection of samples from mice. 28 days after administration, all mice were killed, and the livers were quickly frozen in liquid nitrogen and stored at -80°C.
[0197] (6) Determination of liver editing efficiency. Liver genomic DNA was extracted, and PCR amplification was performed by designing primers near the target gene region. The PCR products were then subjected to high-throughput sequencing to obtain information on the mutation frequency of the target region. That is, the editing efficiency was determined by amplicon sequencing.
[0198] The experimental results are shown in Figure 8. Wild-type SpCas9 and various SpCas9 mutants combined with sgTTR-0 all showed some editing efficacy on the TTR gene in humanized TTR mice. Using the degree of TTR protein knockdown in the serum of humanized TTR mice as an indicator, the reported low-off-target mutant HF1 improved safety at the expense of efficacy, while Cas9-mut5 achieved even greater safety while maintaining comparable editing efficacy to wild-type Cas9.
[0199] Example 4 Optimization of LNP Delivery Formula
[0200] 4.1 Gene Editing Trials with sgTTR-0 and SpCas9 mRNA in Different LNP Delivery Formulations
[0201] To obtain a more optimal LNP delivery formulation for in vivo genome editing, this example tested various LNP formulations as shown in Table 3.
[0202] The main steps are as follows:
[0203] (1) Preparation of test drugs. sgTTR-0 and SpCas9 mRNA (SpCas9-HF1 mRNA or SpCas9-Mut5 mRNA) were encapsulated in LNPs with different formulations.
[0204] (2) Animal Grouping. One week before dosing, serum TTR protein expression was measured using an ELISA kit (ab231920, Abcam). Mice were divided into groups of 4-6 mice per group based on the initial level of TTR expression before dosing, ensuring that the initial TTR expression levels of mice in each group were comparable. Mice were 6-8 weeks old, TTR humanized mice of either sex.
[0205] Table 3
[0206] (3) Administration. Injection was performed via tail vein at a dose of 1 mpk or 0.3 mpk. The day of administration was recorded as day 0.
[0207] (4) One week after administration, the expression of TTR protein in serum was detected by ELISA to determine the knockdown level of TTR protein.
[0208] (5) Collection of samples from mice. 28 days after administration, the mice were killed and their livers were collected, quickly frozen in liquid nitrogen, and stored at -80°C.
[0209] (6) Determination of liver editing efficiency. Liver genomic DNA was extracted, the target gene region was amplified by PCR, and the PCR products were subjected to high-throughput sequencing to determine the editing efficiency.
[0210] The results showed (see Figure 9 ) that LNP-01 is a more suitable LNP formulation for delivery, and the combination of its encapsulated gRNA and Cas9 mRNA (whether SpCas9-HF1 in AC or SpCas9-MUT5 in D in Figure 9 ) has the best gene editing effect.
[0211] Therefore, the LNP formulation used in all other embodiments of this application is LNP-01. Therefore, taking LNP-01 as an example, the packaging method of LNP in the embodiments of this application is as follows:
[0212] (1) Accurately weigh a certain amount of SM-102, 10% DSPC, 38.5% cholesterol, and 1.5% DMG-PEG2000 lipids with a molar mass ratio of 50%, add appropriate amount of anhydrous ethanol to dissolve, and prepare a lipid working solution for use (final concentration of lipid working solution 20 mg / mL).
[0213] (2) Prepare citric acid buffer solution (10 mM, pH 4.0) containing 130 mM sodium chloride, Tris-NaOAc buffer solution (20 mM, 10.7 mM, pH 7.5), and Tris-NaOAc buffer solution (20 mM, 10.7 mM, pH 7.5) containing 60% sucrose respectively.
[0214] (3) Take an appropriate amount of mRNA stock solution and dilute it with the sodium chloride-citrate buffer solution prepared above to adjust the final concentration of the mRNA working solution to 0.18 mg / mL.
[0215] (4) Using a microfluidic instrument and a matching chip, the lipid working solution and the mRNA working solution were mixed in a volume ratio of 1:3 to prepare an mRNA-loaded LNP solution.
[0216] (5) Add 9 times the volume of Tris-NaOAc buffer solution to dilute the prepared LNP solution, and use TFF concentration and purification to remove the ethanol solution in the system.
[0217] (6) The mRNA content in the LNP solution was detected by ultraviolet spectrometry and an appropriate amount of Tris-NaOAc buffer solution (20 mM, 10.7 mM, pH 7.5) containing 60% sucrose was added to adjust the final mRNA concentration in the finished LNP solution to 100 μg / mL and the sucrose content in the external aqueous phase system to 8.7%.
[0218] The test results show that LNP encapsulation efficiency, particle size and other parameters meet the requirements of subsequent tests and can be used for lipid prescription screening.
[0219] 4.2 Detection of gene editing efficiency of sgTTR-0 and SpCas9 mRNA combinations at different mass ratios
[0220] To further optimize the editing efficiency of the gene editing system of this application, the ratio of gRNA and Cas9 mRNA in LNP liposomes was optimized based on the editing efficiency. The specific experimental steps are as follows:
[0221] (1) Preparation of test drugs. sgTTR-0 and SpCas9-HF1 mRNA were encapsulated in LNPs at different mass ratios, such as 1:1, 1:4, 1:8, 1:16, and 1:32. The encapsulation method is described in Section 4.1.
[0222] (2) Animal grouping. 6-8 week old TTR humanized mice (either sex) were used. One week before administration, serum TTR protein expression was measured using an ELISA kit (ab231920, abcam). Mice were then divided into groups based on the initial level of TTR expression before administration, with 4-6 mice per group to ensure that the initial TTR expression levels of mice in each group were comparable.
[0223] (3) Administration: The drug was injected into the tail vein at a dose of 1 mpk. The day of administration was designated as day 0.
[0224] (4) Regularly collect blood and detect the TTR protein expression in the serum to determine the knockdown level of TTR protein.
[0225] (5) Collection of samples from mice. 28 days after administration, the mice were killed and their livers were collected, quickly frozen in liquid nitrogen, and stored at -80°C.
[0226] (6) Determination of liver editing efficiency. Liver genomic DNA was extracted, the target gene region was amplified by PCR, and the PCR products were subjected to high-throughput sequencing to determine the editing efficiency.
[0227] The results, as shown in Figure 10, show that combinations of sgTTR-0 and SpCas9 mRNA at varying mass ratios all achieved a certain degree of editing efficiency on the TTR gene. The optimal editing efficiency was achieved at gRNA to SpCas9 mRNA ratios of 1:1 and 1:4. This ratio, starting at 1:8, reduced editing efficiency to a certain extent, and at a ratio of 1:32, the TTR editing efficiency was essentially halved.
[0228] 4.3 Off-target rate detection of sgTTR-0 and SpCas9 mRNA combinations at different mass ratios
[0229] To further optimize the editing efficiency of the gene editing system of this application, the ratio of gRNA and Cas9 mRNA in LNP liposomes was optimized based on off-target efficiency. The specific experimental steps are as follows:
[0230] (1) Cell transfection. sgTTR-0, SpCas9-WT mRNA and double-stranded oligodeoxynucleotides (dsODN) tags were simultaneously transfected into HepG2 cells by electroporation. The total transfection volume was kept at 1 μg, and sgTTR-0 and SpCas9 mRNA were transfected into HepG2 cells at different mass ratios.
[0231] (2) Genomic DNA extraction: 72 hours after transfection, genomic DNA of cells was extracted.
[0232] (3) ODN integration rate detection: T7E1 enzyme digestion was used to detect DNA double-strand cutting efficiency, and Ndel enzyme digestion was used to detect dsODN tag insertion efficiency. Subsequent detection was performed after the Ndel / T7E1 integration efficiency reached 30%.
[0233] (4) Take appropriate quality genomic DNA to construct a library and sequence it using the MGI-2000 system;
[0234] (5) Bioinformatics analysis.
[0235] The results are shown in Figure 11, indicating that the gRNA to mRNA mass ratio of 1:1 is safer.
[0236] By using the transthyretin gene (TTR) as a target through Examples 1-4, and by optimizing the gRNA, Cas9, and LNP delivery system, as well as the ratio of gRNA to Cas9 mRNA, a CRISPR-Cas9 cleavage enzyme system suitable for TTR gene editing with lower off-target efficiency was obtained. In particular, the aforementioned examples screened the low off-target Cas9 enzyme mutant SpCas9-Mut5 through a large number of experiments. While achieving a highly efficient TTR gene knockout effect, it also achieved the effect of reducing off-target effects to the greatest extent, reduced the side effects of the application of the system, and achieved unexpected technical effects.
[0237] In subsequent examples, the gene editing efficiency and low off-target rate of SpCas9-Mut5 will be verified in multiple cell lines and more target genes. At the same time, through subsequent examples, it was also found that after SpCas9-Mut5 was transformed into dead Cas9 (dCas9) and Cas9 nickase (Cas9 nikase), it still retained the characteristics of high editing efficiency and low off-target efficiency, and this feature can also be introduced into a variety of modified editing systems (such as Cas9nikase-based base editing systems and dCas9-based epigenetic editing systems).
[0238] Example 5: Verification of the editing efficiency of SpCas9 mutants in different cells
[0239] We further verified whether SpCas9-Mut5 editing efficiency was superior to wild-type Cas9 and other mutants in HEK293T and HepG2 cells. The specific methods are as follows:
[0240] (1) sgTTR-0 was combined with different Cas9 mutants (SpCas9-Mut1 to Mut7) to transfect HEK293T cells and HepG2 cells;
[0241] (2) After 72 h, cells were harvested and genomic DNA was extracted;
[0242] (3) Amplify a DNA fragment of approximately 200 bp near the sgTTR-0 target site and perform amplicon sequencing analysis;
[0243] (4) At the same time, the culture supernatant was collected and the TTR protein content in the cell supernatant was detected by ELISA.
[0244] The results are shown in Figure 12.
[0245] The results in Figure A show that in HEK293T cells, under the guidance of sgTTR-0, mut1-7 showed a high editing efficiency (genomic level) for the TTR gene, ranging from 40% to 60%, which is basically equivalent to the editing efficiency of the control group HF1 (about 55%).
[0246] The results in Figure B show that in HepG2 cells, under the guidance of sgTTR-0, the editing efficiency of mut1-7 on the TTR gene (genomic level) is basically equivalent to the editing efficiency of the control group HF1 (about 66%).
[0247] The results in Figure C show that in HepG2 cells, under the guidance of sgTTR-0, mut1-7 has a strong inhibitory effect on TTR protein (protein level), and its knockdown level of TTR protein (downregulation of approximately 60%) is higher than that of the control SpCas9-HF1 (downregulation of approximately 49.5%).
[0248] Example 6 Detection of SpCas9 mutant off-target rates at high-frequency off-target sites
[0249] After using GUIDE-Seq technology to predict the high-frequency off-target sites of SpCas9-WT, it was found that the off-target rate of SpCas9-Mut5 was lower among the four major high-frequency off-target sites. The specific method is as follows:
[0250] (1) sgTTR-0 was combined with different Cas9 mutants to transfect HepG2 cells;
[0251] (2) After 72 hours, cells were harvested and genomes were extracted;
[0252] (3) Amplify a DNA fragment of approximately 200 bp near the high-frequency off-target site;
[0253] (4) Perform amplicon sequencing on the amplified DNA fragments.
[0254] The editing efficiency of the wild-type SpCas9 and its mutants at four different off-target sites (off-target site 1 / 2 / 3 / 4) was analyzed. The results are shown in FIG13 . Specifically,
[0255] The results in Figure A show that under the guidance of sgTTR-0, spCas9-WT had an off-target editing efficiency of nearly 20% at off-target site 1, while mut2 had an off-target editing efficiency of approximately 4%. The remaining Cas9 mutants (including mut5) reduced off-target editing at this site to a level equivalent to that of the ctrl group (<1%). In terms of editing efficiency at this off-target site, mut5 showed a significant difference from WT (P<0.0001), mut5 showed a significant difference from mut2 (P=0.0015), and there was no statistical difference between mut5 and mut1 / 3 / 4 / 6 / 7.
[0256] The results in Figure B show that under the guidance of sgTTR-0, Cas9-WT had an off-target editing efficiency of over 10% at the off-target site 2, mut2 had an off-target editing efficiency of nearly 10%, and mut4 had an off-target editing efficiency of approximately 2%. The remaining Cas9 mutants (including mut5) reduced the off-target editing of this site to a level equivalent to that of the ctrl group (<0.5%). In terms of editing at this off-target site, mut5 showed a significant difference from WT (P<0.0001), mut5 showed a significant difference from mut2 (P<0.0001), mut5 showed a significant difference from mut4 (P<0.0001), and mut5 showed a significant difference from mut1 / 3 / 6 / 7.
[0257] The results in Figure C show that under the guidance of sgTTR-0, Cas9-WT had an off-target editing efficiency of nearly 10% at off-target site 3, and mut2 had an off-target editing efficiency of nearly 5%. The other Cas9 mutants (including mut5) reduced the off-target editing efficiency at this site to the level of the ctrl group (approximately 1%). In terms of editing efficiency at this off-target site, mut5 showed a significant difference from WT (P<0.0001), mut5 showed a significant difference from mut2 (P=0.0015), and there was no statistical difference between mut5 and mut1 / 3 / 4 / 6 / 7.
[0258] The results in Figure D show that under the guidance of sgTTR-0, Cas9-WT measured an off-target editing efficiency of approximately 7% at the off-target site 4 site, and mut5 significantly reduced the editing of this off-target site (reduced to approximately 4%, p value = 0.003).
[0259] Combined with the results of Example 3 (e.g., Figure 8), it can be seen that SpCas9-Mut5 has fewer off-target sites and a lower overall off-target rate compared to wild-type Cas9 and other mutants, and also has a lower off-target rate at each high-frequency off-target site than the wild type or other mutants.
[0260] Example 7: Testing the editing effect of Cas9-mut5 at more sites in the genome
[0261] 7.1 Editing Efficiency Detection
[0262] To investigate whether the editing efficiency of Cas9-mut5 at sites other than TTR is comparable to that of wild-type Cas9 and reported high-efficiency, low-off-target editors, we selected four gRNAs (sgRNA-1 / 2 / 3 / 4) and combined them with Cas9WT / HF1 / HiFiCas9 / Mut5 mRNA to transfect HepG2 cells. After 72 hours, the cells were harvested and DNA fragments of approximately 200 bp near the corresponding target sites were amplified and analyzed by amplicon sequencing. The results are shown in Figure 14. Specifically:
[0263] A shows that under the guidance of sgRNA-1 (AGAGTCTCTCAGCTGGTACA, target TRAC), the editing efficiency of Cas9-mut5 is comparable to that of Cas9 WT / HF1 / HiFiCas9, and their on-target editing efficiencies are all around 60%.
[0264] B shows that under the guidance of sgRNA-2 (TCAGGGTTCTGGATATCTGT, target gene is TRAC), the editing efficiency of Cas9-mut5 is comparable to that of Cas9 WT / HF1 / HiFiCas9, and their on-target editing efficiencies are all around 55%.
[0265] C shows that under the guidance of sgRNA-3 (CTGGATATCTGTGGGACAAG, target gene is TRAC), the editing efficiency of Cas9-mut5 is comparable to that of Cas9 WT / HF1 / HiFiCas9, and their on-target editing efficiencies are all around 60%.
[0266] D shows that under the guidance of sgRNA-4 (ACGACGCGTGGGTGGCAAGC, target gene is REGNASE-1), the editing efficiency of Cas9-mut5 is comparable to that of Cas9 WT / HF1 / HiFiCas9, and their on-target editing efficiencies are all around 40%.
[0267] The above results indicate that Mut5's editing ability at sites other than TTR is not inferior to WT or reported high-efficiency, low-off-target editing tools.
[0268] 7.2 Off-target rate detection
[0269] To investigate the off-target effects of Cas9-mut5 when combined with additional gRNAs, HEK293T cells were transfected with Cas9 WT, HF1, and HiFiCas9 / mut5, respectively, along with EMX1 and VEGFA3 gRNAs. The control group was transfected with gRNA only, and cells were harvested 72 hours later. A DNA fragment approximately 200 bp near the target site was amplified and analyzed by next-generation sequencing. The results are shown in Figure 15, where:
[0270] The results in Figure A show that under the guidance of EMX1 gRNA, the on-target editing efficiency of WT / HF1 / HiFiCas9 / mut5 for the EMX1 gene exceeded 70%, and the on-target editing efficiency of SpCas9-Mut5 was better than that of SpCas9-WT; at the same time, like other high-efficiency, low-off-target editors, mut5 had a significantly reduced off-target editing efficiency at the Sp-Cas9-WT high-frequency off-target site OT1 (from 18.5% to 0.9%, P < 0.0001), almost dropping to the level of the ctrl group, and SpCas9-Mut5 had the most significant effect in reducing the off-target rate.
[0271] The results in Figure B show that under the guidance of VEGFA3 gRNA, the on-target editing efficiency of SpCas9-Mut5 for the VEGFA3 gene exceeds 70%, which is better than SpCas9-WT; at the same time, the off-target rate of SpCas9-Mut5 at the SpCas9-WT high-frequency off-target site OT2 is lower than that of SpCas9-WT and other mutants, and it can significantly reduce the off-target editing of the SpCas9-WT high-frequency off-target site OT1 (from 11.2% to 6.0%, P < 0.0001).
[0272] Overall, the off-target rate of Mut5 in genes other than TTR is relatively lower, and its gene editing efficiency is often better than SpCas9-WT and other low-off-target Cas9 variants.
[0273] Example 8 Editing Effect Testing of SpCas9-Mut5 Nikase (nCas9-Mut5) and Dead SpCas9-Mut5 (dCas9-Mut5)
[0274] Wild-type Cas9 can induce double-strand breaks in DNA because it has two nuclease domains: RuvC and HNH. The RuvC domain cleaves the non-target DNA strand, while the HNH domain cleaves the target DNA strand. If a mutation is introduced into the RuvC nuclease active region to inactivate it, the mutated nuclease can only cleave one strand of the dsDNA. This mutant form of Cas9 nickase is called Cas9 nickase (nCas9). If both the RuvC and HNH nuclease active regions are mutated simultaneously, the nuclease loses the ability to cleave DNA and only retains the ability to enter the genome guided by the guide RNA. This mutant form of Cas9 is called dead Cas9 (dCas9).
[0275] As mentioned above, SpCas9-Mut5, obtained by introducing four mutations (K526A / R691A / Q695A / H698A) into the wild-type Cas9 nuclease, can show editing efficiency comparable to that of the wild type on multiple targets including TTR, indicating that the mutation sites carried by SpCas9-Mut5 do not affect the ability of the nuclease to cut the target DNA. The anti-off-target property of Mut5 is that the Mut5 quadruple mutation reduces the electrostatic and hydrophobic interactions between the Cas9 protein and the genomic DNA phosphate backbone, and because the interaction is independent of the DNA base sequence, the anti-off-target ability of the Mut5 quadruple mutation does not have sequence specificity.
[0276] To explore whether the introduction of the mut5 quadruple mutation into nCas9 and dCas9 still has the effect of high editing efficiency and low off-target, the examples of this application respectively tested the on-target and off-target editing efficiency of the nCas9 fusion protein (taking the single-base editor ABE8e as an example) and the dCas9 fusion protein (taking the epigenetic editor as an example) after the introduction of the Mut5 quadruple mutation.
[0277] 8.1 Mut5 quadruple mutations were applied to nCas9 fusion protein (single-base editor) to reduce off-target levels while maintaining the original on-target editing efficiency
[0278] The Cas9 nicking enzyme nCas9 is fused with the APOBEC family cytosine deaminase or adenine deaminase TadA to obtain a cytosine base editor (CBE) that realizes C to T base conversion, or an adenine base editor (ABE) that realizes A to G base conversion.
[0279] Taking the ABE adenine base editor as an example, the core components of this fusion protein are nCas9 (Cas9nickase, with a D10A mutation relative to the wild-type Cas9 protein shown in SEQ ID NO. 68) and the adenine deaminase TadA. When the fusion protein targets genomic DNA under the guidance of gRNA, the adenine deaminase can bind to the single-stranded DNA and deaminate adenine (A) within a certain range to inosine (I). Inosine is treated as a G base during DNA replication, ultimately achieving a direct conversion of AT base pairs to GC base pairs.
[0280] The editing efficiency of the first-generation ABE editors (such as ABE7.10 and ABEmax) is low. To improve the editing efficiency, David Liu's group developed a new ABE variant, ABE8e, through molecular evolution of the eTadA monomer (the structural diagram is shown in Figure 16A).
[0281] ABE8e exhibits high editing efficiency, with activity 3 to 11 times higher than that of ABE7.10. In this example, four mutations of mut5 were introduced into the nCas9 element of ABE8e to generate ABE8e-mut5 (schematic structure shown in Figure 16B).
[0282] In this example, the bpNLS on the left side of the structure in Figures A and B of Figure 16 is also called the N-terminal signal peptide, and the bpNLS on the right side is also called the C-terminal signal peptide; the aforementioned signal peptides and TadA*, 32-aa linker, nCas9-Mut5, and the amino acid sequences of ABE8e-WT (i.e., ABE8e in the figure) and ABE8e-Mut5 are all shown in the sequence listing.
[0283] In order to study whether the introduction of the mut5 quadruple mutation enables nCas9 to reduce the off-target level while maintaining the original on-target editing efficiency, this example detects the nCas9 introduced with the mut5 quadruple mutation in the ABE8e-mut5 system. In this experiment, mRNA encoding ABE8e-WT (amino acid sequence SEQ ID NO.71, nucleotide sequence SEQ ID NO.5) or ABE8e-mut5 (amino acid sequence SEQ ID NO.72, nucleotide sequence SEQ ID NO.6) was transfected into HEK293T cells with different gRNA combinations by transfection reagent. After 72 hours, the cells were collected, and DNA fragments of about 200bp near the target site (Target) and the related off-target site (OT) were amplified for second-generation sequencing analysis. In this experiment, the ctrl group was only transfected with gRNA without mRNA. The results are shown in Figure 17. The 7 gRNAs selected for the experiment are all sgRNAs reported in literature or patents. The related off-target sites are selected from the high-frequency mutation sites in GUIDE-seq data reported in literature or patents (Tsai, Zheng et al. 2015, Liang, Xie et al. 2019, Richter, Zhao et al. 2020 and WO2022246266A1) (except HEK site2 OT2 / OT3, which are predicted by the off-target prediction online tool, the website of the online tool is http: / / www.rgenome.net / cas-offinder / ).
[0284] Figure 17A shows that after ABE8e and ABE8e-mut5 were combined with EMX1 gRNA (sgRNA targeting the EMX1 gene), their on-target editing efficiencies were comparable (both exceeding 40%), with no significant difference between the two groups. In terms of off-target effects, ABE8e-mut5 significantly reduced the off-target level of the OT1 site (from 14.2% to 2.4%, P < 0.0001), but had no significant effect on the off-target level of the OT2 site.
[0285] Figure 17B shows that after ABE8e and ABE8e-mut5 were combined with VEGFA3 gRNA (sgRNA targeting the VEGFA3 gene), the on-target editing efficiency of ABE8e was 62.6%, and the editing efficiency of ABE8e-mut5 was 68.2%, that is, the on-target editing efficiency of ABE8e-mut5 at this site was significantly higher than that of ABE8e (P=0.0022); in terms of off-target, ABE8e-mut5 can significantly reduce the off-target level of the OT1 site (from 50.1% to 14.7%, P<0.0001) and the off-target level of the OT2 site (from 27% to 2%, P<0.0001)
[0286] Figure 17C shows that after ABE8e and ABE8e-mut5 were combined with TTR-ABE-1gRNA (sgRNA targeting the TTR gene), their on-target editing efficiencies were comparable (both exceeding 68%), and there was no significant difference between the two groups; in terms of off-target, ABE8e-mut5 could significantly reduce the off-target level of the OT2 site (from 23.5% to 4.4%, P < 0.0001) and the off-target level of OT3 (from 5.0% to 3.3%, P < 0.0001), but had no significant effect on the off-target level of OT1.
[0287] Figure 17D shows that when ABE8e and ABE8e-mut5 were combined with gRNA targeting HEK site 1 (sgRNA targeting HEK site 1), their on-target editing efficiencies were comparable (both around 62%). In terms of off-target effects, ABE8e-mut5 significantly reduced the off-target levels of OT1 (from 52.3% to 30.9%, P < 0.0001), OT2 (from 49.4% to 5.3%, P < 0.0001), and OT3 (from 25.1% to 5.2%, P < 0.0001).
[0288] Figure 17E shows that after ABE8e and ABE8e-mut5 were combined with HEK site 2 (sgRNA targeting HEK site 2), their target editing efficiencies were comparable (both exceeding 74%). In terms of off-target, ABE8e-mut5 significantly reduced the off-target level of the OT1 site (from 8.1% to 3.9%, P < 0.0001), but had no significant effect on the off-target levels of OT2 and OT3.
[0289] Figure 17F shows that after ABE8e and ABE8e-mut5 were combined with HEK site 3 (sgRNA targeting HEK site 3), ABE8e-mut5 had a significant improvement in on-target editing efficiency compared with ABE8e (from 65.9% to 74%, P = 0.001); in terms of off-target, ABE8e-mut5 can significantly reduce the off-target level of the OT1 site (from 14.9% to 8.0%, P < 0.0001) and the off-target level of OT2 (from 4.4% to 2.2%, P < 0.0001), but has no significant effect on the off-target level of OT3.
[0290] Figure 17G shows that after ABE8e and ABE8e-mut5 were combined with HEK site 4 (sgRNA targeting HEK site 4), ABE8e-mut5 had a significantly improved on-target editing efficiency compared with ABE8e (from 66.6% to 69.5%, P = 0.0022); in terms of off-target, ABE8e-mut5 can significantly reduce the off-target level of the OT1 site (from 59.3% to 17.3%, P < 0.0001) and the off-target level of OT2 (from 58.6% to 16.8%, P < 0.0001), but has no significant effect on the off-target level of OT3.
[0291] In summary, the results in Figure 17 demonstrate that ABE8e-mut5 exhibits on-target editing efficiencies comparable to or even exceeding those of ABE8e at multiple target sites, while simultaneously reducing editing levels at corresponding off-target sites. This demonstrates that introducing the mut5 quadruple mutation into the nCas9 fusion protein can maintain on-target editing efficiency while reducing off-target levels.
[0292] 8.2 Mut5 quadruple mutation applied to nCas9 fusion protein single base editor to reduce bystander editing effects
[0293] Compared with the first-generation ABE editors (such as ABE7.10, ABEmax), ABE8e has improved editing activity while further widening the editing window (the position range that is prone to deamination) (A4-A8, i.e., the A base from the 4th to the 8th base). This will cause non-target base changes and cause a serious bystander editing effect. To explore whether the introduction of the mut5 quadruple mutation can reduce the bystander editing effect of ABE8e, this example used a transfection reagent to transfect HEK293T cells with mRNA encoding ABE8e / ABE8e-mut5 and different gRNA combinations. After 72 hours, the cells were harvested and a DNA fragment of about 200 bp near the target site was amplified. After next-generation sequencing, the gRNA pairing region (as used herein, the "gRNA pairing region" refers to the double-stranded DNA segment in the genome that is 100% complementary to the gRNA spacer sequence. When describing the modification in the gRNA pairing region, it refers to the modification of the portion of the gRNA pairing region that is identical to the gRNA spacer sequence) was analyzed for the conversion rate of A to G at different positions within 20 bases. The results are shown in Figure 18, where
[0294] A shows that ABE8e-mut5 significantly improved the efficiency of converting the 8th base A to G in the EMX1 gRNA (sgRNA targeting the EMX1 gene) pairing region compared with ABE8e (from 24.2% to 34.4%, P value = 0.0001), and significantly reduced the efficiency of converting the 11th base A to G (from 23.9% to 11.1%, P value < 0.0001).
[0295] B shows that ABE8e-mut5 tends to improve the efficiency of converting the 5th base A to G in the VEGFA3 gRNA (sgRNA targeting the VEGFA3 gene) pairing region compared with ABE8e (from 61.6% to 67%, P value = 0.024), and tends to reduce the editing efficiency of the 9th base A (from 19.3% to 17.6%, P value = 0.0174).
[0296] C shows that ABE8e-mut5 and ABE8e have similar editing efficiencies for the A at the 5th base in the gRNA pairing region of HEK293_2sgRNA (see reference (Tsai, Zheng et al. 2015)) (both around 70%, with no significant difference). The editing efficiency of ABE8e-mut5 for the A at the 7th base is slightly lower than that of ABE8e (from 71.3% to 68.9%, p value = 0.0066). In addition, the A to G conversion rates of ABE8e-mut5 at other positions in the gRNA pairing region are significantly lower than those of ABE8e, including the A to G conversion rates at the 3rd base (ABE8e: 12.1%, ABE8e-mut5: 4.8%, P value < 0.0001) and the A to G conversion rates at the 8th base (ABE8e: 24.5%, ABE8e-mut5: 13.8%, P value < 0.0001). value < 0.0001), the A to G conversion rate at base 9 (ABE8e: 7.9%, ABE8e-mut5: 4.7%, P value < 0.0001), the A to G conversion rate at base 12 (ABE8e: 6.3%, ABE8e-mut5: 2.4%, P value < 0.0001), and the A to G conversion rate at base 14 (ABE8e: 1.8%, ABE8e-mut5: 1.5%, P value = 0.0187). At positions far from the editing window, such as A at positions 2 and 16, ABE8e-mut5 and ABE8e had almost no editing effect on them.
[0297] D shows that ABE8e-mut5 and ABE8e have similar editing efficiencies for the 6th base A in the pairing region of PCSK9 gRNA (i.e., sgRNA targeting the PCSK9 gene) (both greater than 60%, with no significant difference), and their editing efficiencies for the 16th base A are both less than 1%.
[0298] E showed that ABE8e-mut5 and ABE8e both had a certain degree of improvement in the editing efficiency of the 5th base A and the 8th base A in the F-site2 gRNA (targeting the non-coding region) pairing region, with the A to G conversion rate at the 5th base (ABE8e: 63.7%, ABE8e-mut5: 67.1%, P value = 0.0002) and the A to G conversion rate at the 8th base (ABE8e: 57.3%, ABE8e-mut5: 58.9%, P value = 0.0126). However, the A to G conversion rates at other positions in the gRNA pairing region were significantly lower in ABE8e-mut5 than in ABE8e, with the A to G conversion rate at the 2nd base (ABE8e: 7.7%, ABE8e-mut5: 4.9%, P value = 0.0126). The A-to-G conversion rates at position 12 (ABE8e: 14.6%, ABE8e-mut5: 6.4%, P value < 0.0001) and 14 (ABE8e: 3.9%, ABE8e-mut5: 3.3%, P value = 0.0002) were significantly higher than those at position 13 (P value < 0.0001). The editing efficiencies of ABE8e-mut5 and ABE8e for base 16 A were both less than 1%.
[0299] F shows that ABE8e-mut5 and ABE8e have similar editing efficiencies for the A at the 4th base in the ATGg2 gRNA (i.e., sgRNA targeting the ATGg2 gene) pairing region (both greater than 66%, with no significant difference). However, the A-to-G conversion rates of ABE8e-mut5 at other positions in the gRNA pairing region are significantly lower than those of ABE8e, including the A-to-G conversion rate at the 12th base (ABE8e: 6.6%, ABE8e-mut5: 2.0%, Pvalue < 0.0001) and the A-to-G conversion rate at the 13th base (ABE8e: 1.1%, ABE8e-mut5: 0.3%, Pvalue < 0.0001). In regions far from the editing window, such as A at positions 15, 16, and 19, the editing efficiencies of ABE8e-mut5 and ABE8e are both less than 1%.
[0300] In summary, the results in Figure 18 show that within the core region of the editing window (A4-A8), ABE8e-mut5 and ABE8e have comparable A-to-G conversion efficiencies. However, at the edges of the editing window or outside the editing window, ABE8e-mut5's A-to-G editing efficiency is lower than or comparable to that of ABE8e. This suggests that ABE8e-mut5, by reducing its binding to non-target sites, can reduce non-target base changes within the editing window, effectively reducing bystander editing while maintaining on-target editing efficiency.
[0301] 8.3 Mut5 quadruple mutation applied to dCas9 fusion protein (epigenetic editor EE) can maintain the original on-target editing efficiency
[0302] dCas9 is a Cas9 protein that has lost its cutting activity. It has lost the ability to cut DNA, but under the guidance of gRNA, it can still target and bind to DNA with the same precision. CRISPR-dCas9 combined with epigenetic modification enzymes (methyltransferases, acetyltransferases, etc.) can form a dCas9-epigenetic modification system. Currently, the common epigenetic modification enzymes that activate the expression of target genes are mainly histone acetyltransferases and DNA demethyltransferases, such as p300 and Tet1; while the epigenetic modification enzymes that inhibit the expression of target genes are mainly DNA methyltransferase DNMT3a, histone methylase LSD1 and histone deacetylase HDAC3.
[0303] In this example, the epigenetic editor EE-WT, formed by the fusion of dCas9 and DNA methyltransferase DNMT3a, is used as an example (the dCas9 protein has mutations D10A and H840A relative to the wild-type Cas9 protein, and the structural schematic is shown in part A of Figure 19). Four mutations of mut5 are introduced into its dCas9 element to obtain EE-mut5 (the structural schematic is shown in part B of Figure 19).
[0304] Among them, the amino acid sequences of dCas9-WT, Dnmt3A+3L, KRAB, dCas9-mut5, EE-WT and EE-Mut5 used in this example are all shown in the sequence listing.
[0305] To investigate whether the introduction of the mut5 quadruple mutation affects the editing efficiency of dCas9 and its fusion protein, EE-WT mRNA / EE-mut5 mRNA and a gRNA targeting the PCSK9 gene (EE-PCSK9) were encapsulated in LNPs, respectively, and the liver cancer cell lines HepG2 and Huh7 were treated with both LNPs at a dose of 500ng / 2.5E5 cells. The medium was changed or passaged every 2-4 days (cells were counted during passage to ensure consistent cell numbers in different groups). 12 days after transfection, the cell supernatant was collected and the PCSK9 protein content in the cell supernatant was measured using an ELISA kit. The results showed that both EE-mut5 and EE-WT could downregulate PCSK9 protein expression by more than 99% in different liver cancer cells, indicating that both EE-mut5 and EE-WT could effectively inhibit PCSK9 protein expression, and the degree of inhibition was comparable (the results are shown in Figure 20).
[0306] To further validate the efficacy of EE-mut5, HBV-targeting gRNAs (sgHBV-1 and sgHBV-2) were encapsulated with EE-WT mRNA and EE-mut5 mRNA, respectively, in LNPs. HepG2.2.15 cells were treated with LNPs at a dose of 40 ng / 2.25E4 cells. Two days after transfection, the medium was changed. Five days after transfection, the supernatant was collected and analyzed for HBV DNA, HBsAg (HBV s antigen), and HBeAg (HBV e antigen). Cell viability was also assessed using Cell-Titer Glo. The results are shown in Figure 21, where:
[0307] A shows that under the guidance of sgHBV-1 (i.e., EE-sgHBV-1), the inhibition rates of EE-mut5 and EE-WT on HBV DNA were comparable (both around 50%, with no significant difference); under the guidance of sgHBV-2 (i.e., EE-sgHBV-2), the inhibition rates of EE-mut5 and EE-WT on HBV DNA were also comparable (both exceeding 60%, with no significant difference).
[0308] B shows that under the guidance of sgHBV-1, the inhibition rates of EE-mut5 and EE-WT on HBsAg are comparable (both exceeding 98%, with no significant difference); under the guidance of sgHBV-2, the inhibition rates of EE-mut5 and EE-WT on HBsAg are also comparable (both around 90%, with no significant difference).
[0309] C shows that under the guidance of sgHBV-1, the inhibition rates of EE-mut5 and EE-WT on HBeAg are comparable (both at 97%, with no significant difference); under the guidance of sgHBV-2, the inhibition rates of EE-mut5 and EE-WT on HBeAg are also comparable (both exceeding 96%, with no significant difference).
[0310] D shows that there is no significant difference in the effect of treating cells with EE-mut5 and EE-WT in combination with different gRNAs on cell viability.
[0311] 8.4 Mut5 quadruple mutation applied to dCas9 fusion protein (epigenetic editor EE) to achieve lower off-target levels
[0312] To investigate whether the application of the mut5 quadruple mutation to epigenetic editors would reduce their off-target effects, HBV-targeting gRNA (sgHBV-1) was encapsulated in LNPs along with EE-WT mRNA and EE-mut5 mRNA, respectively. The LNPs in the control group contained only sgRNA. All three groups of LNPs were treated with HepG2.2.15 cells at a dose of 1.5 μg / 8E5 cells, with three replicates per group. After 7 days of treatment, cells were harvested and whole-genome epigenetic sequencing was performed. The methylation levels of CpG sites on the genomes of each group were analyzed, and differentially expressed genes between the EE-WT and EE-mut5 groups relative to the control group were counted. The results are shown in Figure 22. The results showed that compared to the control group, EE-WT had 141 sites with significantly increased methylation levels, while only 58 sites with significantly increased methylation were counted for EE-mut5 (37 of these differentially expressed sites overlapped between the EE-WT and EE-mut5 groups). This shows that EE-mut5 can significantly reduce the methylation level on the genome outside the target site compared with EE-WT, that is, the application of Mut5 quadruple mutation to the dCas9 fusion protein epigenetic editor EE can make it have a lower off-target level.
[0313] The relevant sequences involved in this application are shown in the following sequence table:
[0314] Sequence Listing
[0315] Note: When SEQ ID NO.1-SEQ ID NO.13 are used as mRNA sequences, all thymidine (T) represents uridine (U). When SEQ ID NO.1-SEQ ID NO.13 are used in the examples, all U represents 1-methylpseudouridine.
[0316] When SEQ ID NOs. 14-46 are used as gRNA sequences, thymidine (T) represents uridine (U). When gRNA sequences such as SEQ ID NOs. 14-23 are used in the examples, the sequences of SEQ ID NOs. 14-23 all have the following modifications: the first three nucleotides and the last three nucleotides from the 5' end are 2'-O-methylated, and the first three nucleotide linkages from the 5' end and the last three nucleotide linkages are phosphorothioate linkages.
[0317] The above describes the implementation methods of the present application. However, the present application is not limited to the above implementation methods. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. An isolated modified Cas9 nuclease or DNA binding fragment thereof, comprising a mutation at one or more amino acid residue positions selected from the group consisting of K526, N692, Q695, H698, N497, Y450, Q926, K377, E387, D397, R400, D406, A421, L423, R424, Q426, Y430, K442, P449, V452, A456, R457, W464, M465, K468, E470, T474, P475, W476 76. F478, K484, S487, A488, T496, F498, L502, N504, K506, P509, F518, N522, E523, L540, S541, I548, D550, F553, V561 , K562, E573, A589, L598, D605, L607, N609, N612, E617, D618, D628, R629, R635, K637, L651, K652, R654, T657, G658, L6 66. K673, S675, I679, L680, L683, N690, R691, F693, S701, F704, Q712, G715, Q716, H723, I724, L727, I733, L738, Q739 , N803, Q805, Q807, K810, Y812, D829, N831, R832, S834, D835, Q844, S845, K848, R859, K862, R864, K866, K890, T893, Q pyogenes Cas9 (SpCas9) protein amino acid sequence: SEQ ID NO: 68 or the amino acid numbering definition in reference to the amino acid numbering of the Streptococcus pyogenes Cas9 (SpCas9) protein amino acid sequence: SEQ ID NO:
68.
2. The modified Cas9 nuclease or DNA binding fragment thereof according to claim 1, comprising one or more mutations selected from the group consisting of K526X, Q695X, H698X and R691X, wherein X is glycine (G), alanine (A), valine (V), isoleucine (I), leucine (L), aspartic acid (D), glutamic acid (E), asparagine (N), glutamine (Q), serine (S), threonine (T), lysine (K), arginine (R), phenylalanine (F) or tyrosine (Y).
3. The modified Cas9 nuclease or DNA binding fragment thereof according to claim 1 or 2, wherein each X is independently selected from any one of the following: glycine (G), alanine (A), aspartic acid (D) and glutamic acid (E).
4. The modified Cas9 nuclease or DNA binding fragment thereof according to claim 3, comprising any combination of mutations selected from the group consisting of K526A+R691A+Q695A+H698A, K526A+R691A+N692A+Q695A+H698A, K526G+R691G+Q695G+H698G, K526D+R691A+Q695A+H698A, K526A+R691D+Q695A+H698A, K526A+R691A+Q695A+H698D, K526E+R691A+ Q695A+H698A, K526A+R691E+Q695A+H698A and K526A+R691A+Q695A+H698E.
5. The modified Cas9 nuclease or DNA binding fragment thereof according to any one of claims 1 to 4, which is Cas9 nucleic acid double-strand cleavage enzyme, nCas9 (Cas9 nickase), or dCas9 (Catalytically dead Cas9).
6. The modified Cas9 nuclease or DNA binding fragment thereof according to any one of claims 1 to 5, further comprising one or more mutations in the RuvC domain and / or the HNH domain, optionally comprising a mutation from aspartic acid at position 10 to alanine (D10A) in the RuvC domain and / or a mutation from histidine at position 840 to alanine (H840A) in the HNH domain.
7. The modified Cas9 nuclease or DNA binding fragment thereof according to any one of claims 1 to 6, wherein the amino acid sequence of the Cas9 nuclease is as shown in SEQ ID NO.69, SEQ ID NO.78 or SEQ ID NO.81, or comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the amino acid sequence shown in SEQ ID NO.69, SEQ ID NO.78 or SEQ ID NO.
81.
8. The Cas9 nuclease or DNA binding fragment thereof according to any one of claims 1 to 7, wherein the DNA binding fragment does not comprise one or more amino acid segments selected from the group consisting of: The amino acid segments at positions 494-501, 179-296, 503-708, 792-897 and 1010-1081, wherein the amino acid positions correspond to the amino acid numbers in the SpCas9 protein amino acid sequence SEQ ID NO: 68, Optionally, the DNA binding fragment is a DNA binding fragment contained in SEQ ID NO.69, SEQ ID NO.78 or SEQ ID NO.81, wherein the DNA binding fragment does not contain one or more amino acid segments selected from the following group in these sequences: Amino acid segments at positions 494-501, 179-296, 503-708, 792-897 and 1010-1081.
9. A fusion protein comprising the Cas9 nuclease or DNA binding fragment thereof according to any one of claims 1 to 8.
10. The fusion protein according to claim 9, further comprising a nuclear localization signal peptide on the N-terminal side and / or the C-terminal side of the Cas9 nuclease or its DNA binding fragment.
11. The fusion protein according to claim 9 or 10, further comprising a cytosine deaminase, an adenine deaminase, an oxidase, a glycosidase, an alkyltransferase, a DNA synthetase, an RNA synthetase, a uracil glycosylase inhibitor (UGI), a transcription activator, a transcription repressor, a methylase, or a demethylase fused to a Cas9 nuclease or a DNA binding fragment thereof, a Gam protein derived from bacteriophage Mu, and / or a fluorescent protein; optionally, the transcription activator comprises VP64, and optionally, the transcription repressor comprises a KRAB protein.
12. The fusion protein according to claim 11, which comprises or is, from N-terminus to C-terminus: 1) Nuclear localization signal peptide-Cas9 nucleic acid double-strand cleavage enzyme, nCas9 or dCas9 or its DNA binding fragment-nuclear localization signal peptide; 2) Nuclear localization signal peptide-TadA* enzyme-nCas9-nuclear localization signal peptide; 3) TadA* enzyme-nCas9; 4)Dnmt3A-Dnmt3L-dCas9-KRAB; or 5) Dnmt3A-Dnmt3L-dCas9-nuclear localization signal peptide-KRAB; Wherein "-" indicates connection through a peptide bond or a linker.
13. The fusion protein according to claim 12, wherein: Each of the linkers independently comprises or is any one or more amino acid sequences selected from the group consisting of SEQ ID NO.77, SEQ ID NO.85, SEQ ID NO.88-91, or an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto; The amino acid sequence of the nuclear localization signal peptide independently comprises or is any one selected from the following: SEQ ID NO.70, SEQ ID NO.71, SEQ ID NO.74, SEQ ID NO.75, optionally, the amino acid sequence as shown in SEQ ID NO.71 or 74 is located on the N-terminal side, and / or the amino acid sequence as shown in SEQ ID NO.70, 75 or 85 is located on the C-terminal side; The amino acid sequence of the TadA* enzyme comprises or is the amino acid sequence shown in SEQ ID NO.76; The amino acid sequence of KRAB comprises or is the amino acid sequence shown in SEQ ID NO.87; The amino acid sequence of DNMT3A comprises or is the amino acid sequence shown in SEQ ID NO.87; and / or The amino acid sequence of DNMT3L comprises or is the amino acid sequence shown in SEQ ID NO.
87.
14. The fusion protein according to any one of claims 11 to 13, whose amino acid sequence is as shown in SEQ ID NO.66, SEQ ID NO.73 or SEQ ID NO.80, or comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the amino acid sequence as shown in SEQ ID NO.66, SEQ ID NO.73 or SEQ ID NO.
80.
15. An engineered nucleic acid molecule comprising a nucleotide sequence encoding a modified Cas9 nuclease or a DNA binding fragment thereof according to any one of claims 1-8 or a fusion protein according to any one of claims 9-14.
16. The nucleic acid molecule according to claim 15, comprising a 5' non-coding region sequence (5'UTR) and a 3' non-coding region sequence (3'UTR); wherein, The 5'UTR comprises or is a 5'UTR encoding a gene of tobacco phagocytosis virus, and / or the 3'UTR comprises or is a 3'UTR encoding a gene of human hemoglobin protein alpha 1 (hHBA1).
17. A nucleic acid molecule according to claim 15 or 16, wherein the nucleotide sequence of the 5'UTR is as shown in SEQ ID NO.10 or 11, or comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% nucleotide sequence identity with the nucleotide sequence as shown in SEQ ID NO.10 or 11, and / or the nucleotide sequence of the 3'UTR is as shown in SEQ ID NO.13, or comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the nucleotide sequence as shown in SEQ ID NO.
13.
18. A nucleic acid molecule according to any one of claims 15 to 17, further comprising a Poly(A) tail sequence or a tailing signal sequence, preferably the Poly(A) tail sequence comprises the nucleotide sequence shown in SEQ ID NO: 12 or comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with SEQ ID NO.
12.
19. The nucleic acid molecule according to any one of claims 15 to 18, which is mRNA and comprises a 5' cap structure, optionally, the 5' cap structure is m7G(5')ppp(5')(2'OMeA)pG.
20. A nucleic acid molecule according to any one of claims 15 to 19, wherein the nucleotide sequence is as shown in SEQ ID NO: 2, 6 or 8, or comprises a nucleotide sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with the nucleotide sequence as shown in SEQ ID NO: 2, 6 or 8.
21. The nucleic acid molecule according to any one of claims 15-20, which is an mRNA molecule, comprising one or more modified uridine (U), optionally, the modified U is 1-methylpseudouridine; optionally, each U in the nucleic acid molecule encoding the Cas9 nuclease or an enzymatically active fragment thereof is 1-methylpseudouridine.
22. A composition, a ribonucleoprotein complex or a protein lipid complex comprising a modified Cas9 nuclease or a DNA binding fragment thereof according to any one of claims 1 to 8, or a fusion protein according to any one of claims 9 to 14, or a nucleic acid molecule according to any one of claims 15 to 21.
23. The composition, ribonucleoprotein complex or protein lipid complex according to claim 22, further comprising a gRNA targeting a target gene, a nucleic acid molecule encoding the gRNA, or a construct comprising the gRNA.
24. The gene editing composition according to claim 23, wherein the target gene is selected from any one or more of the following: hepatitis B (HBV) gene, PCSK9, EMX1 and VEGFA3.
25. A gene editing method, comprising introducing into a host cell the modified Cas9 nuclease or its DNA binding fragment described in any one of claims 1-8, or the fusion protein described in any one of claims 9-14, or the nucleic acid molecule described in any one of claims 15-21, or the composition, ribonucleoprotein complex or protein lipid complex described in any one of claims 22-24.
26. Use of the modified Cas9 nuclease or DNA binding fragment thereof according to any one of claims 1-8, or the fusion protein according to any one of claims 9-14, or the nucleic acid molecule according to any one of claims 15-21, or the composition, ribonucleoprotein complex or protein lipid complex according to any one of claims 22-24 in the preparation of a medicament for treating a disease or condition in need thereof.