Epigenetic editing system
By delivering a fusion protein of catalytically inactivated Cas protein and DNA methyltransferase via lipid nanoparticles, the challenge of large-size delivery of CRISPR editors has been solved, enabling durable epigenetic changes and safe in vivo application.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
- Filing Date
- 2026-01-19
- Publication Date
- 2026-05-12
AI Technical Summary
Existing CRISPR-based epigenetic editors are large in size, making efficient delivery via size-limited vectors such as adeno-associated virus (AAV) difficult, thus limiting their in vivo application and posing risks of viral vector DNA integration or prolonged plasmid expression.
Lipid nanoparticles (LNPs) are used to deliver engineered CRISPR-OFF-epigenome editors. By using catalytically inactivated Cas proteins such as dSpCas9 or dSF01 to fusion proteins with DNA methyltransferase DNMT3A and KRAB transcriptional repression domains, transient protein expression is achieved through mRNA delivery, enabling efficient and site-specific DNA methylation.
It enables persistent epigenetic changes without altering the genome sequence, avoids the risk of viral vector integration, and provides a safer and more flexible in vivo application method.
Smart Images

Figure CN122012459A_ABST
Abstract
Description
[0001] This application claims priority to Chinese patent application CN2025115522676, filed on October 28, 2025, entitled "An Epigenetic Editing System". The entire contents of the aforementioned Chinese patent application are incorporated herein by reference. Technical Field
[0002] This invention relates to the field of gene editing, particularly to the field of regularly clustered short palindromic repeats (CRISPR) technology. Specifically, this invention relates to an epigenetic editing system. Background Technology
[0003] Precisely programmable gene expression regulation through targeted epigenetic editing represents a revolutionary approach to developing novel therapeutics, offering the potential for long-term physiological changes without altering the underlying genome sequence. CRISPR-based epigenetic editors enable site-specific modifications to the epigenome, providing a powerful pathway for persistently silencing pathogenic genes or activating therapeutic genes. This innovative strategy works by inducing heritable chemical modifications to chromatin structure, such as DNA methylation or histone modifications, thereby altering gene expression patterns. However, a major obstacle to the widespread clinical application of many current CRISPR-based epigenome editors is their large size, primarily due to the multi-domain nature of these fusion proteins (Cas protein, guide RNA, and one or more effector domains). This large payload size presents a significant challenge for efficient in vivo delivery using size-constrained vectors such as adeno-associated virus (AAV).
[0004] This invention seeks to address these key limitations by delivering engineered mRNA for epigenetic repression systems via lipid nanoparticles (LNPs). The invention develops and systematically optimizes CRISPR OFF-Epigenome Editors (CRISPR OFF-EE) based on two different Cas platforms: the widely used Streptococcus pyogenes Cas9 (SpCas9) and the smaller Cas12i3 variant Cas-SF01. These systems involve fusing catalytically inactivated Cas proteins (dSpCas9 or dSF01) with potent DNA methyltransferase effector domains (DNMT3A and DNMT3L) and a KRAB transcriptional repression domain to promote efficient, site-specific DNA methylation, thereby inducing persistent gene silencing. By utilizing an mRNA delivery editor, the goal of this invention is to achieve transient protein expression, thereby programming persistent epigenetic changes, while mitigating the risks associated with viral vector DNA integration or long-term expression in plasmid-based systems. This approach provides a potentially safer and more flexible model for the in vivo application of these powerful epigenetic tools. Summary of the Invention
[0005] Fusion protein
[0006] On one hand, the present invention provides a fusion protein comprising a DNA methyltransferase, a Cas protein, and a transcriptional repressor domain.
[0007] In one embodiment, the DNA methyltransferase comprises DNMT3A and DNMT3L; the transcriptional repression domain is KRAB.
[0008] In one embodiment, the Cas protein is selected from Cas9 or Cas12 proteins.
[0009] In one embodiment, the Cas protein is selected from the Cas12i protein.
[0010] In one embodiment, the Cas protein is a Cas protein with nuclease activity inactivated.
[0011] In one embodiment, the Cas9 is dCas9 (Cas9 with nuclease activity inactivated).
[0012] In one embodiment, the Cas12i protein is Cas-SF01.
[0013] Specifically, CN116004573B discloses a Cas protein BC26312 with an amino acid mutation, which is referred to as Cas-SF01 in this invention.
[0014] In one embodiment, the Cas-SF01 is a Cas protein with nuclease inactivation. For example, the Cas-SF01 can be obtained by mutating E at position 844 to A or by mutating D at position 619 to A.
[0015] In one embodiment, the fusion protein comprises, from the N-terminus to the C-terminus, a DNA methyltransferase, a Cas protein, and a transcriptional repressor domain.
[0016] In one embodiment, the fusion protein comprises, from N-terminus to C-terminus, DNMT3A, DNMT3L, Cas protein, and a transcriptional repressor domain.
[0017] In one embodiment, the elements of the fusion protein can be directly connected or connected via a connector (e.g., an XTEN connector).
[0018] In one embodiment, the fusion protein further includes a nuclear localization sequence (NLS).
[0019] In one embodiment, the fusion protein optionally comprises one or two amino acid sequences as shown in SEQ ID NO: 1-SEQ ID NO: 6, and / or, comprises one or two amino acid sequences that are at least 95% homologous to the amino acid sequences shown in SEQ ID NO: 1-SEQ ID NO: 6.
[0020] In a preferred embodiment, the amino acid sequence of the fusion protein is shown in SEQ ID NO: 1.
[0021] In another preferred embodiment, the fusion protein comprises one or two amino acid sequences as shown in SEQ ID NO: 2 and SEQ ID NO: 3.
[0022] In another preferred embodiment, the fusion protein comprises one or two amino acid sequences as shown in SEQ ID NO: 4 and SEQ ID NO: 5.
[0023] In another preferred embodiment, the amino acid sequence of the fusion protein is shown in SEQ ID NO: 6.
[0024] In this invention, the amino acid site is an amino acid site starting from the N-terminus.
[0025] Those skilled in the art will understand that the structure of a protein can be altered without adversely affecting its activity and function. For example, one or more conserved amino acid substitutions can be introduced into the amino acid sequence of a protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule. Examples and implementations of conserved amino acid substitutions are familiar to those skilled in the art. Specifically, an amino acid residue can be substituted with another amino acid residue belonging to the same group as the site to be substituted, i.e., replacing another nonpolar amino acid residue with a nonpolar amino acid residue, replacing another polar uncharged amino acid residue with a polar uncharged amino acid residue, replacing another basic amino acid residue with a basic amino acid residue, and replacing another acidic amino acid residue with an acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. Conservative substitutions, where an amino acid is replaced by another amino acid belonging to the same group, fall within the scope of this invention, provided that the substitution does not lead to inactivation of the protein's biological activity. Therefore, the proteins of this invention can contain one or more conserved substitutions in their amino acid sequences, preferably generated by substitutions according to Table 1. Furthermore, this invention also covers proteins that also contain one or more other nonconservative substitutions, provided that such nonconservative substitutions do not significantly affect the desired function and biological activity of the proteins of this invention.
[0026] Conservative amino acid substitutions can occur at one or more predicted non-essential amino acid residues. “Non-essential” amino acid residues are those that can be altered (deleted, substituted, or replaced) without changing biological activity, while “essential” amino acid residues are required for biological activity. A “conservative amino acid substitution” is a substitution in which an amino acid residue is replaced by an amino acid residue with a similar side chain. Amino acid substitutions can occur in the non-conservative regions of the aforementioned Cas mutant protein. Generally, such substitutions are not performed on conserved amino acid residues, or on amino acid residues located within conserved motifs, where such residues are required for protein activity. However, those skilled in the art will understand that functional variants may have fewer conserved or non-conserved alterations in conserved regions.
[0027] Table 1
[0028]
[0029] As is well known in the art, one or more amino acid residues can be altered (replaced, deleted, truncated, or inserted) from the N and / or C ends of a protein while retaining its functional activity. Therefore, proteins that have one or more amino acid residues altered from their N and / or C ends while retaining their desired functional activity are also within the scope of this invention. These alterations can include those introduced by modern molecular methods such as PCR, which includes PCR amplification that alters or lengthens the protein-coding sequence by means of oligonucleotides containing amino acid-coding sequences used in the PCR amplification.
[0030] It should be recognized that proteins can be altered in various ways, including amino acid substitutions, deletions, truncations, and insertions, and methods for such operations are generally known in the art. For example, amino acid sequence variants of the aforementioned proteins can be prepared by mutating DNA. This can also be accomplished through other forms of mutagenesis and / or directed evolution, for example, using known mutagenesis, recombination, and / or shuffling methods, combined with relevant screening methods, to perform single or multiple amino acid substitutions, deletions, and / or insertions.
[0031] Those skilled in the art will understand that these minor amino acid changes in the Cas protein of this invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using r-DNA technology) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be altered, but the polypeptide may retain its activity. If the mutations are not located near the catalytic domain, active site, or other functional domains, a smaller impact can be expected.
[0032] Those skilled in the art can identify the essential amino acids of the Cas mutant protein of the present invention using methods known in the art, such as localized mutagenesis, protein evolution, or bioinformatics analysis. The catalytic domains, active sites, or other functional domains of the protein can also be determined through physical structural analysis, such as by techniques like nuclear magnetic resonance, crystallography, electron diffraction, or photoaffinity labeling, combined with mutations in presumed key site amino acids.
[0033] In this invention, amino acid residues can be represented by a single letter or by three letters, for example: alanine (Ala, A), valine (Val, V), glycine (Gly, G), leucine (Leu, L), glutamic acid (Gln, Q), phenylalanine (Phe, F), tryptophan (Trp, W), tyrosine (Tyr, Y), aspartic acid (Asp, D), asparagine (Asn, N), glutamic acid (Glu, E), lysine (Lys, K), methionine (Met, M), serine (Ser, S), threonine (Thr, T), cysteine (Cys, C), proline (Pro, P), isoleucine (Ile, I), histidine (His, H), and arginine (Arg, R).
[0034] The term "AxxB" indicates that amino acid A at position xx is changed to amino acid B. For example, D293K means that D at position 293 is mutated to K. When multiple amino acid sites are mutated simultaneously, it can be expressed in forms such as D293K_E494R, D293K / E494RK, and D293K+E494R. For example, D293K_E494R represents that D at position 293 is mutated to K while E at position 494 is mutated to R.
[0035] The specific amino acid positions (numbers) within the protein described in this invention are determined using standard sequence alignment tools by comparing the amino acid sequence of the target protein with SEQ ID NO.1. For example, the Smith-Waterman algorithm or the CLUSTALW2 algorithm can be used to align two sequences, with the sequence considered aligned when the alignment score is the highest. The alignment score can be calculated according to the method described in Wilbur, W.J. and Lipman, D.J. (1983) Rapid similarity searches of nucleic acid and protein data banks. Proc. Natl. Acad. Sci. USA, 80:726-730. In the ClustalW2 (1.82) algorithm, the default parameters are preferably used: protein gap opening penalty = 10.0; protein gap extension penalty = 0.2; protein matrix = Gonnet; protein / DNA end gap = -1; protein / DNA GAPDIST = 4. Preferably, the AlignX program (part of the vectorNTI group) is used with default parameters suitable for multiple alignments (gap opening penalty: 10, gap extension penalty: 0.05) to determine the position of a specific amino acid in the protein of the present invention by comparing the amino acid sequence of the protein with SEQ ID NO.1. Those skilled in the art can use commonly used software, such as Clustal Omega, to perform sequence identity comparison and alignment of the amino acid sequence of any parental Cas protein with SEQ ID NO.1, thereby obtaining the amino acid sites in the parental Cas protein corresponding to the amino acid sites defined in SEQ ID NO.1 as described in this application.
[0036] In one embodiment, the fusion protein may further include other modified portions; the modified portions are selected from other proteins or peptides, detectable markers, or any combination thereof.
[0037] In one embodiment, the modified portion is selected from epitope tags and reporter gene sequences.
[0038] The epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art can choose other suitable epitope tags (e.g., for purification, detection or tracing).
[0039] The reporter gene sequences are well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0040] In one embodiment, the fusion protein of the present invention contains a detectable marker, such as a fluorescent dye, such as FITC or DAPI.
[0041] In one embodiment, the Cas protein of the present invention is optionally coupled, conjugated, or fused to the modified portion via a linker.
[0042] The fusion protein of the present invention is not limited by its production method. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.
[0043] Nucleic acid encoding fusion protein
[0044] On the other hand, the present invention provides a nucleic acid comprising:
[0045] (a) The polynucleotide sequence encoding the fusion protein of the present invention;
[0046] Alternatively, a polynucleotide complementary to the polynucleotide described in (a).
[0047] In one embodiment, the nucleotide sequence is codon-optimized for expression in prokaryotic cells. In another embodiment, the nucleotide sequence is codon-optimized for expression in eukaryotic cells.
[0048] In one embodiment, the cell is an animal cell, such as a mammalian cell.
[0049] In one embodiment, the cell is a human cell.
[0050] In one embodiment, the cell is a plant cell, such as the cell of a cultivated plant (e.g., cassava, corn, sorghum, wheat, or rice), algae, tree, or vegetable.
[0051] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.
[0052] In one embodiment, the nucleic acid is mRNA.
[0053] In a preferred embodiment, the nucleic acid comprises a polynucleotide sequence encoding a fusion protein as shown in any of the amino acid sequences of SEQ ID NO: 1-SEQ ID NO: 6.
[0054] In a preferred embodiment, the nucleic acid comprises a polynucleotide sequence as shown in any of SEQ ID NO: 136, SEQ ID NO: 137 and SEQ ID NO: 147-150.
[0055] In a preferred embodiment, the nucleic acid has a polynucleotide sequence as shown in any of SEQ ID NO: 136, SEQ ID NO: 137 and SEQ ID NO: 147-150.
[0056] In a preferred embodiment, the nucleic acid has a polynucleotide sequence as shown in SEQ ID NO: 136.
[0057] In a preferred embodiment, the nucleic acid has a polynucleotide sequence as shown in SEQ ID NO: 137.
[0058] CRISPR epigenetic editing system and protein-nucleic acid complexes / compositions
[0059] This invention provides an engineered, non-naturally occurring vector system, or a CRISPR-Cas system, or an epigenetic editing system, comprising the aforementioned fusion protein or a nucleic acid sequence encoding the fusion protein and a nucleic acid encoding one or more guide RNAs (gRNAs).
[0060] The gRNA includes a region that binds to the Cas protein and a region that binds to the target sequence.
[0061] In one embodiment, the nucleic acid sequence encoding the fusion protein and the nucleic acid encoding one or more guide RNAs are artificially synthesized.
[0062] The guide RNA can form a complex with the fusion protein.
[0063] This guide RNA targets one or more target sequences.
[0064] The one or more target sequences hybridize with the genomic loci of the DNA molecule encoding one or more gene products and guide the Cas protein to the genomic locus of the DNA molecule of the one or more gene products, thereby altering or modifying the expression of the one or more gene products.
[0065] In one embodiment, the target sequence is a cell-derived target sequence; for example, a prokaryotic cell or a eukaryotic cell.
[0066] The cells of this invention include one or more of animals, plants, or microorganisms.
[0067] In some embodiments, the Cas protein is codon-optimized for expression in cells.
[0068] In one embodiment, the gRNA of the present invention includes a region that binds to the Cas9 protein and a region that binds to a target sequence. The gRNA is capable of targeting a region upstream of the PCSK9 gene or its transcription start site, and the region that binds to the target sequence is selected from sgRNA1 to sgRNA10, preferably sgRNA3 or sgRNA4. The present invention also provides a combination comprising two gRNAs, the combination of which is selected from a combination of sgRNA3 and Na-sgRNA, or a combination of sgRNA4 and Na-sgRNA.
[0069] In one embodiment, the gRNA of the present invention includes a region that binds to a Cas12i protein (e.g., Cas-SF01) and a region that binds to a target sequence. The gRNA is capable of targeting a region upstream of the PCSK9 gene or its transcription start site. The region that binds to the target sequence is selected from AS-cr1 to AS-cr18, S-cr1 to S-cr23, and crRNA1 to crRNA60, preferably from AS-cr7, S-cr14, crRNA6, crRNA7, crRNA9, crRNA13, crRNA16, crRNA25, and crRNA59, and more preferably from crRNA6, crRNA7, crRNA9, crRNA13, crRNA16, crRNA25, and crRNA59. The present invention also provides a combination comprising two gRNAs, wherein the regions of the two gRNAs that bind to the target sequence are respectively selected from AS-cr7 and S-cr14.
[0070] In some preferred embodiments, the epigenetic editing system comprises:
[0071] (A-1) (i) a nucleic acid sequence encoding a fusion protein as shown in any of the amino acid sequences of SEQ ID NO: 1-5, and (ii) a gRNA molecule as shown in any of SEQ ID NO: 66-75, or a combination of a gRNA molecule as shown in any of SEQ ID NO: 66-75 and a gRNA molecule as shown in SEQ ID NO: 146; or
[0072] (B-1) (i) a nucleic acid sequence encoding the fusion protein as shown in the amino acid sequence of SEQ ID NO: 6, and (ii) a crRNA molecule as shown in any of SEQ ID NO: 25-65 and SEQ ID NO: 76-135, or a combination of two of the crRNA molecules.
[0073] In some preferred embodiments, the epigenetic editing system comprises:
[0074] (A-2) (i) comprises any of the polynucleotide sequences shown in SEQ ID NO: 136, SEQ ID NO: 147-150, or a combination of any two of the polynucleotide sequences shown therein; and (ii) is selected from one of the gRNA molecules shown in SEQ ID NO: 68 and SEQ ID NO: 69, or a combination of one of them with a gRNA molecule shown in SEQ ID NO: 146; or
[0075] (B-2) contains the polynucleotide sequence shown in SEQ ID NO: 137 and (ii) is selected from one or a combination of two of the gRNA molecules shown in SEQ ID NO: 31, SEQ ID NO: 56, SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 84, SEQ ID NO: 88, SEQ ID NO: 91, SEQ ID NO: 100 and SEQ ID NO: 134.
[0076] In some other preferred embodiments, the epigenetic editing system comprises:
[0077] (A-3) (i) comprises the polynucleotide sequence shown in SEQ ID NO: 136, and (ii) is selected from one of the gRNA molecules shown in SEQ ID NO: 68 and SEQ ID NO: 69, or a combination of one of them with a gRNA molecule shown in SEQ ID NO: 146; or
[0078] (B-3) (ii) comprises the polynucleotide sequence shown in SEQ ID NO: 137 and (ii) is selected from gRNA molecules shown in SEQ ID NO: 81, SEQ ID NO: 82, SEQ ID NO: 84, SEQ ID NO: 88, SEQ ID NO: 91, SEQ ID NO: 100 and SEQ ID NO: 134.
[0079] The present invention also provides an engineered, non-naturally occurring carrier system, which may include one or more carriers, the one or more carriers comprising:
[0080] a) A first regulatory element, which is operatively linked to the gRNA.
[0081] b) A second regulatory element operatively linked to the fusion protein;
[0082] Components (a) and (b) are located on the same or different carriers in the system.
[0083] The first and second regulatory elements include promoters (e.g., constitutive or inducible promoters), enhancers (e.g., 35S promoters or 35S enhanced promoters), internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and polyU sequences).
[0084] In some embodiments, the vector in the system is a viral vector (e.g., a retroviral vector, lentiviral vector, adenovirus vector, adeno-associated vector, and herpes simplex vector), or it can be a plasmid, virus, granule, bacteriophage, or other type known to those skilled in the art.
[0085] In some embodiments, the system provided herein is a delivery system. In some embodiments, the delivery system is a nanoparticle, liposome, exosome, microbubble, or gene gun.
[0086] In one embodiment, the target sequence is a DNA or RNA sequence derived from prokaryotic or eukaryotic cells. In another embodiment, the target sequence is a non-naturally occurring DNA or RNA sequence.
[0087] In one embodiment, the target sequence is present within the cell. In another embodiment, the target sequence is present in the cell nucleus or cytoplasm (e.g., organelles). In one embodiment, the cell is a eukaryotic cell. In other embodiments, the cell is a prokaryotic cell.
[0088] On the other hand, the present invention provides a complex or composition comprising:
[0089] (i) Protein components selected from: the aforementioned fusion protein; and
[0090] (ii) A nucleic acid component comprising (a) a guide sequence capable of hybridizing with a target sequence; and (b) a region capable of binding to the Cas protein in the fusion protein of the present invention.
[0091] The protein components and nucleic acid components combine to form a complex.
[0092] In one embodiment, the nucleic acid component is a guide RNA in a CRISPR-Cas system.
[0093] In one embodiment, the complex or composition is non-natural or modified. In one embodiment, at least one component of the complex or composition is non-natural or modified. In one embodiment, the first component is non-natural or modified; and / or, the second component is non-natural or modified.
[0094] In one embodiment, the target sequence bound by the gRNA is a sequence derived from the PCSK9 gene or a sequence upstream of the transcription start site (TSS) of the PCSK9 gene. For example, a sequence within 1000 bp upstream of the PCSK9 gene transcription start site (TSS) (e.g., within 20 bp, 50 bp, 100 bp, 150 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, or 900 bp).
[0095] In one embodiment, the region of the gRNA that binds to the target sequence has 15 to 30 consecutive nucleotides, for example, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, or 29 nucleotides.
[0096] In one embodiment, the gRNA includes terminal chemical modifications to enhance stability. For example, the region binding to the target sequence includes 2'-O-methylated nucleosides at positions 1-3 and the last three positions, linked by 3'-thiophosphate nucleosides.
[0097] On the other hand, the present invention also provides the use of the above-mentioned CRISPR system or protein-nucleic acid complex / composition in inhibiting the transcription or initiation of the PCSK9 gene; or, in the preparation of reagents or kits for inhibiting the transcription or initiation of the PCSK9 gene.
[0098] In a preferred embodiment, the CRISPR system or protein-nucleic acid complex / composition can increase the methylation level of the PCSK9 gene promoter region, thereby inhibiting the transcription or initiation of the PCSK9 gene.
[0099] In this invention, the inhibition results in the repression or silencing of the PCSK9 gene transcription.
[0100] Guide RNA (gRNA)
[0101] On the other hand, the present invention provides a gRNA comprising a first segment and a second segment; the first segment is also referred to as a "backbone region", "protein binding region" or "protein binding sequence"; the second segment is also referred to as a "target sequence for targeting nucleic acids" or "targeting region for targeting nucleic acids" or "guide sequence for targeting target sequences".
[0102] The first segment of the gRNA can interact with the Cas protein of the present invention, thereby enabling the Cas protein and gRNA to form a complex.
[0103] The target sequence or target region of the nucleic acid targeted by this invention comprises a nucleotide sequence complementary to a sequence in the target nucleic acid. In other words, the target sequence or target region of the nucleic acid targeted by this invention interacts with the target nucleic acid in a sequence-specific manner through hybridization (i.e., base pairing). Therefore, the target sequence or target region of the nucleic acid can be altered or modified to hybridize with any desired sequence within the target nucleic acid. The nucleic acid is selected from DNA or RNA.
[0104] The percentage of complementarity between the target sequence or target region of the target nucleic acid and the target sequence of the target nucleic acid may be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%).
[0105] The "backbone region," "protein-binding region," and "protein-binding sequence" of the gRNA of this invention can interact with CRISPR proteins (or Cas proteins). The gRNA of this invention guides the interacting Cas protein to a specific nucleotide sequence within the target nucleic acid through the targeting sequence of the target nucleic acid.
[0106] The gRNA of this invention can form a complex with Cas protein.
[0107] In one embodiment, the gRNA of the present invention includes a region that binds to the Cas9 protein and a region that binds to a target sequence. The gRNA is capable of targeting a region upstream of the PCSK9 gene or its transcription start site, and the region that binds to the target sequence is selected from sgRNA3 or sgRNA4. The present invention also provides a combination comprising two gRNAs, wherein the regions of the two gRNAs that bind to the target sequence are selected from sgRNA3 and Na-sgRNA, respectively.
[0108] In one embodiment, the gRNA of the present invention includes a region that binds to a Cas12i protein (e.g., Cas-SF01) and a region that binds to a target sequence. The gRNA is capable of targeting a region upstream of the PCSK9 gene or its transcription start site, and the region binding to the target sequence is selected from AS-cr7 or S-cr14. The present invention also provides a combination comprising two gRNAs, wherein the regions of the two gRNAs that bind to the target sequence are selected from AS-cr7 and S-cr14, respectively.
[0109] On the other hand, the present invention also provides the use of a system or composition comprising the above-described gRNA in inhibiting the transcription or initiation of the PCSK9 gene; or, in the preparation of reagents or kits for inhibiting the transcription or initiation of the PCSK9 gene. In a preferred embodiment, the system or composition can increase the methylation level of the PCSK9 gene promoter region, thereby inhibiting the transcription or initiation of the PCSK9 gene. In the present invention, the inhibition results in the repression or silencing of the transcription of the PCSK9 gene.
[0110] In one embodiment, the system or composition further includes a Cas9 protein, and the gRNA includes a region that binds to the Cas9 protein and a region that binds to a target sequence; the gRNA is capable of targeting a region upstream of the PCSK9 gene or its transcription start site, and the region that binds to the target sequence is selected from sgRNA3 or sgRNA4. Optionally, the gRNA is a combination of two gRNAs, and the regions of the two gRNAs that bind to the target sequence are selected from sgRNA3 and Na-sgRNA, respectively.
[0111] In one embodiment, the system or composition further includes a Cas12i protein, and the gRNA includes a region that binds to a Cas12i protein (e.g., Cas-SF01) and a region that binds to a target sequence; the gRNA is capable of targeting a region upstream of the PCSK9 gene or its transcription start site, and the region that binds to the target sequence is selected from AS-cr7 or S-cr14. Optionally, the gRNA is a combination of two gRNAs, and the regions of the two gRNAs that bind to the target sequence are selected from AS-cr7 and S-cr14, respectively.
[0112] mRNA sequence optimization
[0113] The fusion protein methylation editing tool of this invention can be optimized using artificial intelligence algorithms. Protein sequences from epigenetic editors (EEs) are converted (translated) into codon-optimized DNA sequences using codon optimization tools provided by Integrated DNA Technologies (IDT, https: / / www.idtdna.com / CodonOpt), Twist Biosciences (https: / / ecommerce.twistdna.com / app), and GENEWIZ (https: / / clims4.genewiz.com / Toolbox / CodonOptimization). These tools select codons based on species-specific codon usage preferences to improve translation efficiency in mammalian cells.
[0114] The obtained DNA sequence was then transcribed into an mRNA sequence and further optimized using the LinearDesign algorithm (https: / / github.com / LinearDesignSoftware / LinearDesign, http: / / rna.baidu.com / ). LinearDesign balances two key computational metrics: the minimum free energy (MFE), representing the stability of the mRNA secondary structure, and the codon fitness index (CAI), representing translation efficiency. We tuned the LAMBDA hyperparameter to optimize the trade-off between MFE and CAI to ensure both mRNA sequence stability and high expression levels.
[0115] This embodiment uses RiboGraphViz (www.github.com / DasLab / RiboGraphViz) to predict the secondary structure of the optimized mRNA sequence to calculate the MFE-based folding conformation. We evaluated the sequence's folding free energy, CAI score, and the presence of immunogenic motifs such as Toll-like receptor (TLR) activation sequences. The mRNA's lifespan was predicted using DegScore (https: / / github.com / eternagame / DegScore).
[0116] This invention ultimately selected the optimal combination of low folding free energy, high naturalness, high CAI score, extended linear secondary structure, and minimal immunogenic motifs to generate novel mRNA sequences with higher translational capacity and stability for downstream in vitro and in vivo experiments. Both 573-Split-SpCas9-OFF-EE and 713-Split-SpCas9-OFF-EE are derived from SpCas9-OFF-EE V2 (SEQ ID NO: 136). Both SpCas9-OFF-EE V2 (SEQ ID NO: 136) and SF01-OFF-EE V2 mRNA (SEQ ID NO: 137) underwent AI-assisted optimization.
[0117] Delivery and delivery compositions (drug-LNP compositions)
[0118] The fusion proteins, gRNAs, nucleic acids, CRISPR systems, and protein-nucleic acid complexes / compositions of the present invention can be delivered by any method known in the art. Such methods include, but are not limited to, electroporation, lipid transfection, nuclear transfection, microinjection, acoustic pore effect, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, nuclear transfection, magnetic transfection, lipid transfection, puncture transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial viruses, etc.
[0119] In one embodiment, the present invention provides lipid nanoparticles (LNPs) for delivering the CRISPR system or protein-nucleic acid complex / composition or mRNA of the present invention into cells.
[0120] Based on this, the present invention also provides an LNP comprising a nucleic acid encoding the above-mentioned fusion protein and gRNA.
[0121] In one embodiment, the nucleic acid encoding the fusion protein is mRNA.
[0122] The present invention provides an LNP (drug-LNP composition) comprising one or more nucleic acids, said one or more nucleic acids comprising: (a) mRNA and / or gRNA encoding the above-mentioned fusion protein, said gRNA including a region binding to the Cas protein in the fusion protein and a region binding to a target sequence; (b) a cationic lipid or a salt thereof, said cationic lipid or a salt thereof; (c) a mixture of phospholipids and cholesterol or derivatives thereof; and (d) a PEG-lipid conjugate (e.g., DMG-PEG2000).
[0123] This invention delivers nucleic acids containing fusion proteins and gRNA into cells or subjects via the aforementioned LNP.
[0124] In one embodiment, the drug-LNP composition comprises the above-described epigenetic editing system and an LNP for delivering the epigenetic editing system, wherein the LNP is composed of cationic lipids, helper phospholipids (DSPC), cholesterol, and polyethylene glycol lipids (DMG-PEG). In one embodiment, the LNP of the present invention can be prepared by the following manner:
[0125] The ratio of amine to RNA phosphate (N:P) is 3:1–10:1 (e.g., 4:1, 5:1, 6:1, 7:1, 8:1, or 9:1). Lipids, including ionizable cationic lipids, cholesterol, DSPC, and DMG-PEG2000, are dissolved in ethanol at a molar ratio (e.g., 50:10:38.5:1.5), and the mRNA and gRNA containing the fusion protein (preferably, the weight ratio of mRNA to gRNA encoding the fusion protein is 1:1) are dissolved in acetate buffer. During mixing, the ratio of aqueous phase to organic solvent is maintained at approximately 3:1, and the flow rate is 12 ml / min. After mixing, the LNPs are diluted with PBS and dialyzed against a 10 kDa filter at 4°C for 12 hours for buffer exchange. The LNPs are concentrated using an Amicon® Ultra centrifuge filter and stored at 4°C for subsequent use.
[0126] host cells
[0127] The present invention also relates to an in vitro, ex vivo, or in vivo cell or cell line or its progeny, said cell or cell line or its progeny comprising: the fusion protein of the present invention, nucleic acid molecule, CRISPR-Cas system, protein-nucleic acid complex, vector, system or composition comprising the gRNA of the present invention, or delivery composition of the present invention.
[0128] In some implementations, the cell is a prokaryotic cell.
[0129] In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a non-human mammalian cell, such as cells of non-human primates, cattle, sheep, pigs, dogs, monkeys, rabbits, or rodents (such as rats or mice). In some embodiments, the cell is a non-mammalian eukaryotic cell, such as cells of poultry (such as chickens), fish, or crustaceans (such as clams or shrimp).
[0130] In some implementations, the cell is a stem cell or stem cell line.
[0131] In some cases, the host cells of the present invention contain genetic or genomic modifications that are not present in their wild type.
[0132] Methods and Applications
[0133] The fusion proteins, nucleic acid molecules, CRISPR-Cas systems, protein-nucleic acid complexes, vectors, systems or compositions containing the gRNA of the present invention, delivery compositions, or the host cells described above can be used for any or more of the following purposes: gene or genome editing, editing target sequences in target loci to modify organisms, and reducing or inhibiting PCSK9 gene expression or transcription. In other embodiments, they can also be used to prepare reagents or kits for any or more of the above purposes.
[0134] The present invention also provides a method for editing a target nucleic acid, the method comprising contacting the target nucleic acid with the aforementioned fusion protein, nucleic acid molecule, CRISPR-Cas system, protein-nucleic acid complex, vector, system or composition comprising the gRNA of the present invention, or delivery composition. In one embodiment, the method is for editing the target nucleic acid intracellularly or extracellularly.
[0135] The gene editing of the present invention can be performed in cells (e.g., prokaryotic cells and / or eukaryotic cells).
[0136] In one embodiment, the edited target nucleic acid is a sequence derived from the PCSK9 gene or a sequence upstream of the transcription start site (TSS) of the PCSK9 gene. For example, a sequence derived within 1000 bp upstream of the PCSK9 gene transcription start site (TSS) (e.g., within 20 bp, 50 bp, 100 bp, 150 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, or 900 bp).
[0137] In one implementation, the editing results in the suppression or reduction of the transcriptional level of the PCSK9 gene; for example, the transcription is repressed or silenced.
[0138] In one embodiment, the editing can increase the methylation level of the PCSK9 gene promoter region, thereby reducing or inhibiting the transcription or initiation of the PCSK9 gene.
[0139] In one embodiment, the above-mentioned gene or genome editing is performed in vivo or in vitro.
[0140] On the other hand, the present invention also provides a method for reducing or inhibiting PCSK9 gene expression or transcription, the method comprising the step of contacting the above-mentioned fusion protein, nucleic acid molecule, CRISPR-Cas system, protein-nucleic acid complex, vector, system or composition containing gRNA of the present invention or delivery composition with a target nucleic acid in a cell; wherein the target nucleic acid is a sequence derived from the PCSK9 gene or a sequence derived from upstream of the transcription start site (TSS) of the PCSK9 gene.
[0141] In one embodiment, the inhibition of PCSK9 gene expression or transcription occurs in vivo or in vitro.
[0142] In some embodiments, the cells of the present invention are eukaryotic cells. In some embodiments, the cells are mammalian cells. In some embodiments, the cells are human cells. In some embodiments, the cells are non-human mammalian cells, such as cells of non-human primates, cattle, sheep, pigs, dogs, monkeys, rabbits, or rodents (such as rats or mice). In some embodiments, the cells are non-mammalian eukaryotic cells, such as cells of poultry (such as chickens), fish, or crustaceans (such as clams or shrimp).
[0143] In some implementations, the cell is a stem cell or stem cell line.
[0144] On the other hand, the present invention also provides a method for treating PCSK9-related diseases in subjects, the method comprising the steps of reducing or inhibiting the expression or transcription of the PCSK9 gene in subjects using the above-mentioned fusion protein, nucleic acid molecule, CRISPR-Cas system, protein-nucleic acid complex, vector, system or composition containing gRNA of the present invention or delivery composition of the present invention.
[0145] On the other hand, the present invention also provides the use of the above-mentioned fusion protein, nucleic acid molecule, CRISPR-Cas system, protein-nucleic acid complex, vector, system or composition containing the gRNA of the present invention, or delivery composition of the present invention in the preparation of reagents or pharmaceutical compositions for treating PCSK9-related diseases in subjects.
[0146] The PCSK9-related diseases include autosomal dominant hypercholesterolemia (ADH), hypercholesterolemia, elevated total cholesterol levels, elevated low-density lipoprotein (LDL) levels, decreased high-density lipoprotein (HDL) levels, hepatic steatosis, coronary heart disease, ischemic stroke, peripheral vascular disease, thrombosis, type 2 diabetes, hypertension, obesity, Alzheimer's disease, neurodegeneration, age-related macular degeneration (AMD), or a combination thereof.
[0147] CRISPR system
[0148] As used herein, the terms “regularly clustered short palindromic repeats (CRISPR)-CRISPR-associated (Cas) (CRISPR-Cas) system” or “CRISPR system” are used interchangeably and have the meaning commonly understood by those skilled in the art, which typically includes transcripts or other elements associated with the expression of CRISPR-associated (“Cas”) genes, or transcripts or other elements capable of directing the activity of said Cas genes. The Cas protein in this invention is a Crisprassociated protein.
[0149] CRISPR / Cas complex
[0150] As used herein, the term “CRISPR / Cas complex” refers to a complex formed by the binding of guide RNA or mature crRNA to the Cas protein, which contains a guide sequence that hybridizes to the target sequence and a homologous repeat sequence that binds to the Cas protein. This complex is capable of recognizing and cleaving polynucleotides that hybridize with the guide RNA or mature crRNA.
[0151] target sequence
[0152] A "target sequence" refers to a polynucleotide targeted by a guide sequence in the gRNA, such as a sequence complementary to that guide sequence, where hybridization between the target and guide sequences will promote the formation of a CRISPR / Cas complex (including the Cas protein and gRNA). Perfect complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of a CRISPR / Cas complex.
[0153] The target sequence can contain any polynucleotide, such as DNA or RNA. In some cases, the target sequence is located inside or outside the cell. In some cases, the target sequence is located in the cell nucleus or cytoplasm. In some cases, the target sequence may be located in an organelle of a eukaryotic cell, such as a mitochondrion or chloroplast. The sequence or template that can be used for recombination into a target locus containing the target sequence is referred to as an "edit template," "edit polynucleotide," or "edit sequence." In one embodiment, the edit template is a foreign nucleic acid. In one embodiment, the recombination is homologous recombination.
[0154] In this invention, the "target sequence," "target polynucleotide," or "target nucleic acid" can be any endogenous or exogenous polynucleotide for a cell (e.g., a eukaryotic cell). For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or useless DNA). In some cases, the target sequence should be associated with a protospacer adjacent motif (PAM).
[0155] wild type
[0156] As used herein, the term “wildtype” has the meaning commonly understood by those skilled in the art as referring to the typical form of an organism, strain, or gene, or the characteristic that distinguishes it from mutant or variant forms when it exists in nature, is separable from its natural source and has not been intentionally modified by humans.
[0157] Derivatization
[0158] As used herein, the term "derivation" refers to the chemical modification of an amino acid, polypeptide, or protein in which one or more substituents are covalently linked to the amino acid, polypeptide, or protein. Substituents may also be referred to as side chains.
[0159] A derivatized protein is a derivative of the original protein. Generally, the derivatization of a protein does not adversely affect its desired activity (e.g., activity to bind to guide RNA, endonuclease activity, activity to bind to and cleave a target sequence at a specific site under the guidance of guide RNA). In other words, the derivative of a protein has the same activity as the original protein.
[0160] Derivatized proteins
[0161] Also known as "protein derivatives," these are modified forms of proteins, where one or more amino acids of the protein may be deleted, inserted, modified, and / or substituted.
[0162] Not naturally occurring
[0163] As used herein, the terms “non-naturally occurring” or “engineered” are used interchangeably and indicate artificial involvement. When these terms are used to describe nucleic acid molecules or peptides, they indicate that the nucleic acid molecule or peptide is at least substantially free from at least one other component bound to it, either naturally occurring or found in nature.
[0164] Orthologue (ortholog)
[0165] As used herein, the term "orthologue" has the meaning commonly understood by those skilled in the art. As further guidance, an "orthologue" of a protein, as described herein, refers to a protein belonging to a different species that performs the same or similar function as the protein that is its orthologue.
[0166] identity
[0167] As used herein, the term "identity" refers to the sequence matching between two polypeptides or two nucleic acids. Two compared sequences are identical at a position when the same base or amino acid monomeric subunit occupies the same location (e.g., a position in each of two DNA molecules is occupied by adenine, or a position in each of two polypeptides is occupied by lysine). The "percentage identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared × 100. For example, if six out of ten positions in two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT have 50% identity (three out of six positions match). Typically, two sequences are compared to produce the maximum identity. Such comparisons can be made using methods readily available, for example, computer programs such as the Align program (DNAstar, Inc.) Needleman et al. (1970) J. Mol. Biol. 48:443-453. The percentage identity between two amino acid sequences can also be determined using the algorithm of E. Meyers and W. Miller (Comput. Appl Biosci., 4:11-17 (1988)) integrated into the ALIGN program (version 2.0), which uses a PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4. Alternatively, the percentage identity between two amino acid sequences can be determined using the Needleman and Wunsch algorithm (J MoI Biol. 48:444-453 (1970)) in the GAP program integrated into the GCG software package (available at www.gcg.com), which uses a Blossum 62 matrix or a PAM250 matrix, along with gap weights of 16, 14, 12, 10, 8, 6, or 4, and length weights of 1, 2, 3, 4, 5, or 6.
[0168] carrier
[0169] The term "vector" refers to a nucleic acid molecule capable of delivering another nucleic acid molecule linked to it. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, or without free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and a wide variety of other polynucleotides known in the art. A vector can be introduced into a host cell through transformation, transduction, or transfection, thereby enabling the expression of its carried genetic material elements in the host cell. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc., as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector may contain a variety of elements controlling expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. Additionally, the vector may contain a replication initiation site.
[0170] One type of vector is a "plasmid," which is a circular double-stranded DNA loop into which another DNA fragment can be inserted, for example, using standard molecular cloning techniques.
[0171] Another type of vector is the viral vector, in which a virus-derived DNA or RNA sequence is present in a vector used to package the virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). Viral vectors also contain polynucleotides carried by the virus used for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and episodic mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.
[0172] Other vectors (e.g., non-attachment mammalian vectors) integrate into the host cell's genome upon introduction and thereby replicate along with the host genome. Furthermore, some vectors are capable of directing the expression of genes they are operatively linked to. Such vectors are referred to herein as "expression vectors."
[0173] host cells
[0174] As used herein, the term “host cell” refers to a cell that can be used to introduce a vector, including but not limited to prokaryotic cells such as Escherichia coli or Bacillus subtilis, and eukaryotic cells such as microbial cells, fungal cells, animal cells, and plant cells.
[0175] Those skilled in the art will understand that the design of expression vectors can depend on factors such as the selection of host cells to be transformed and the desired expression level.
[0176] Control element
[0177] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences), for which detailed description can be found in Goeddel, *Gene Expression Technology: Methods in Enzymology*, 185, Academic Press, San Diego, California (1990). In some cases, regulatory elements include those sequences that direct constitutive expression of a nucleotide sequence in many types of host cells and those sequences that direct expression of that nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can primarily direct expression in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, specific organs (e.g., liver, pancreas), or specific cell types (e.g., lymphocytes). In some cases, regulatory elements can also be directed to express in a time-dependent manner (such as in a cell cycle-dependent or developmental stage-dependent manner), which may or may not be tissue or cell type specific. In some cases, the term "regulatory element" covers enhancer elements such as WPRE; CMV enhancer; R-U5' fragment in the LTR of HTLV-I (Mol. Cell. Biol., Vol. 8(1), pp. 466-472, 1988); SV40 enhancer; and intron sequence between exons 2 and 3 of rabbit β-globin (Proc. Natl. Acad. Sci. USA., Vol. 78(3), pp. 1527-31, 1981).
[0178] promoter
[0179] As used herein, the term "promoter" has the meaning known to those skilled in the art, referring to a non-coding nucleotide sequence located upstream of a gene that initiates the expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell under most or all physiological conditions of the cell. An inducible promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when an inducer corresponding to the promoter is present in the cell. A tissue-specific promoter is a nucleotide sequence that, when operably linked to a polynucleotide encoding or defining a gene product, results in the production of the gene product in the cell substantially only when the cell is a cell of the tissue type corresponding to that promoter.
[0180] NLS
[0181] A “nuclear localization signal” or “nuclear localization sequence” (NLS) is an amino acid sequence that “tags” a protein to allow it to be transported to the nucleus via nuclear transport; that is, a protein with an NLS is transported to the nucleus. Typically, an NLS contains positively charged Lys or Arg residues exposed on the protein surface. Exemplary nuclear localization sequences include, but are not limited to, NLS from the following: SV40 large T antigen, EGL-13, c-Myc, and TUS protein. In some embodiments, the NLS contains the PKKKRKV sequence. In some embodiments, the NLS contains the AVKRPAATKKAGQAKKKKLD sequence. In some embodiments, the NLS contains the PAAKRVKLD sequence. In some embodiments, the NLS contains the MSRRRKANPTKLSENAKKLAKEVEN sequence. In some embodiments, the NLS contains the KLKIKRPVK sequence. Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP A1, the KIPIK sequence in the yeast transcriptional repressor Matα2, and PY-NLS.
[0182] Operable connection
[0183] As used herein, the term “operably linked” is intended to mean that the nucleotide sequence of interest is linked to one or more regulatory elements in a manner that allows the expression of that nucleotide sequence (e.g., in an in vitro transcription / translation system or in the host cell when the vector is introduced into the host cell).
[0184] Complementarity
[0185] As used herein, the term "complementarity" refers to the ability of a nucleic acid to form one or more hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. The percentage of complementarity indicates the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, 10 out of 10 are 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Complete complementarity" means that all consecutive residues in a nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in a second nucleic acid sequence. As used herein, “substantially complementary” refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region having 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.
[0186] Strict conditions
[0187] As used herein, “strict conditions” for hybridization refer to conditions under which a nucleic acid complementary to the target sequence hybridizes primarily with the target sequence and substantially does not hybridize to non-target sequences. Strict conditions are typically sequence-dependent and vary depending on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence.
[0188] Hybridization
[0189] The terms “hybridization” or “complementary” or “substantially complementary” refer to nucleic acids (such as RNA, DNA) containing nucleotide sequences that enable them to bind non-covalently, that is, to form base pairs and / or G / U base pairs with another nucleic acid in a sequence-specific, antiparallel manner (i.e., nucleic acid-specific binding of complementary nucleic acids), also known as “annealing” or “hybridization”.
[0190] Hybridization requires two nucleic acids to contain complementary sequences, although mismatches between bases are possible. Suitable conditions for hybridization between two nucleic acids depend on their length and degree of complementarity, variables well known in the art. Typically, hybridizable nucleic acids are 8 nucleotides or longer (e.g., 10 nucleotides or longer, 12 nucleotides or longer, 15 nucleotides or longer, 20 nucleotides or longer, 22 nucleotides or longer, 25 nucleotides or longer, or 30 nucleotides or longer).
[0191] It should be understood that the sequence of a polynucleotide does not need to be 100% complementary to the sequence of its target nucleic acid for specific hybridization. The polynucleotide may contain 60% or higher, 65% or higher, 70% or higher, 75% or higher, 80% or higher, 85% or higher, 90% or higher, 95% or higher, 98% or higher, 99% or higher, 99.5% or higher, or have 100% sequence complementarity with the target region of the target nucleic acid sequence it hybridizes with.
[0192] Hybridization of the target sequence with gRNA means that at least 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the nucleic acid sequences of the target sequence and gRNA can hybridize to form a complex; or it means that at least 12, 15, 16, 17, 18, 19, 20, 21, 22, or more bases of the nucleic acid sequences of the target sequence and gRNA can complement each other to form a complex.
[0193] Express
[0194] As used herein, the term "expression" refers to the process by which a DNA template is transcribed into polynucleotides (such as mRNA or other RNA transcripts) and / or the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotides are derived from genomic DNA, expression can include the splicing of mRNA in eukaryotic cells.
[0195] connector
[0196] As used herein, the term "linker" refers to a linear polypeptide formed by the linkage of multiple amino acid residues via peptide bonds. The linkers of this invention can be synthetically produced amino acid sequences or naturally occurring polypeptide sequences, such as polypeptides with hinge region functions. Such linker polypeptides are well known in the art (see, for example, Holliger, P. et al. (1993) Proc. Natl. Acad. Sci. USA 90:6444-6448; Poljak, RJ et al. (1994) Structure2:1121-1123).
[0197] treat
[0198] As used in this article, the term "treatment" means to treat or cure a disease, to delay the onset of symptoms of a disease, and / or to slow the progression of a disease.
[0199] Subjects
[0200] As used herein, the term “subject” includes, but is not limited to, various animals, plants and microorganisms.
[0201] animal
[0202] For example, mammals, such as bovids, equines, sheep, suidae, canines, felines, lagos, rodents (e.g., mice or rats), non-human primates (e.g., macaques or cynomolgus monkeys), or humans. In some embodiments, the subject (e.g., a human) suffers from a condition (e.g., a condition caused by a disease-related gene defect).
[0203] Beneficial effects of the invention
[0204] In this invention, the applicant utilizes Cas9 or a smaller Cas-SF01 (a Cas12i3 variant) to rationally design and engineer a compact, mRNA-delivered epigenetic repressor. Combined with an optimized mRNA structure and a lipid nanoparticle (LNP) carrier, a single intravenous injection of optimized OFF-EE-V2 mRNA and a selected guide RNA (gRNA) targeting mouse PCSK9 resulted in a reduction of approximately 90% in circulating PCSK9 protein and a corresponding reduction of approximately 55% in LDL-C levels, with effects lasting at least 180 days. The Cas-SF01-based editor exhibited higher specificity and fewer off-target methylation events compared to its Cas9-based counterpart. The optimized LNP formulation also demonstrated good safety.
[0205] This invention involves fusing the DNMT3A / DNMT3L methyltransferase domain and the transcriptional repressor domain KRAB with a nuclease-inactivated Cas protein (especially the small-sized Cas-SF01), and then optimizing the fusion protein and epigenetic editing system of this invention. The fusion protein and epigenetic editing system of this invention have the following beneficial effects:
[0206] 1. Compact structure and high delivery efficiency: Compared with traditional large Cas9 fusion proteins, the Cas-SF01 variant used in this invention significantly reduces the overall size of the fusion protein, which is more conducive to in vivo delivery via LNP-mRNA vector.
[0207] 2. Long-lasting gene silencing effect: The synergistic effect of DNMT3A / 3L and KRAB can induce high-level DNA methylation in the promoter region of the target gene. The optimized fusion protein and the epigenetic editing system OFF-EE-V2 can achieve long-lasting epigenetic silencing without changing the DNA sequence (for example, the silencing effect on PCSK9 can last for at least 180 days).
[0208] 3. High safety: Persistent epigenetic memory can be established through transient mRNA expression, avoiding the risk of viral vector integration into the genome, and Cas-SF01 exhibits lower off-target effects than SpCas9.
[0209] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and examples. However, those skilled in the art will understand that the following drawings and examples are for illustrative purposes only and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art from the following detailed description of the drawings and preferred embodiments. Attached Figure Description
[0210] Figure 1 Epigenetic silencing of the Gapdh-Snrpn-GFP reporter gene and endogenous CD151 was achieved in HEK293T cells using the SpCas9 / SF01-OFF-EE plasmid. Specifically:
[0211] a. Schematic diagram of the SpCas9 / SF01 epigenetic editor (SpCas9 / SF01-OFF-EE). The Snrpn-GFP reporter gene is targeted by catalytically inactivated Cas9 (dCas9) or SF01 (dSF01) (dCas9 / dSF01), a protein fused with an epigenetic editor (EE) domain. This complex is guided by a specific single sgRNA or crRNA to the Snrpn promoter region, resulting in DNA methylation and subsequent silencing of GFP expression.
[0212] b, Representative flow cytometry plot showing Snrpn-GFP expression in HEK293T cells 30 days after transfection with SpCas9 / SF01-OFF-EE plasmid and corresponding single sgRNA or crRNA.
[0213] c. Quantitative analysis of GFP silencing in Snrpn-GFP reporter cells 30 days after transfection. Values represent the mean ± standard deviation (SD) of three independent biological replicates.
[0214] Methylation levels of individual CpGs in the d, Snrpn promoter region and adjacent Gapdh genomic sites. Data are presented as mean percentage ± SD of three independent biological replicates.
[0215] e. Quantitative analysis of CD151 silencing in HEK293T cells 30 days after transfection with the SpCas9 / SF01-OFF-EE plasmid and a pool of three sgRNAs or three crRNAs targeting the CD151 promoter. Values represent the mean ± SD of three independent biological replicates.
[0216] methylation levels of individual CpGs within the f, CD151 promoter region. Data are presented as mean percentage ± SD for three independent biological replicates.
[0217] Figure 2 Optimize SpCas9-OFF and SF01-OFF epigenome editor (EE) mRNAs to achieve efficient and durable gene silencing. Specifically:
[0218] a. Schematic diagram of the SpCas9-OFF / SF01-OFF-EE mRNA construct. The V1 construct represents the unoptimized mRNA, whose coding sequence (CDS) is amplified from the previously described mammalian expression plasmid. The V2 construct incorporates several optimizations: modified 5' and 3' untranslated regions (UTRs), codon optimization, altered linker sequences, and revised nuclear localization signals (NLS).
[0219] b, c, Gene silencing efficiency comparison analysis using V1 and V2 SpCas9-OFF-EE mRNA with homologous sgRNA (b) or SF01-OFF-EE mRNA with homologous crRNA (c). Silencing of GFP and endogenous genes was assessed in Snrpn-GFP reporter cells and HEK293T cells 30 days after electroporation. Experiments included single sgRNA / crRNA targeting GFP, or three mixed sgRNA / crRNAs targeting multiple endogenous sites (CD29, CD81, and CD151).
[0220] d, Hepa 1-6 PCSK9 IRES-GFP Schematic diagram of the reporter gene cell line. An internal ribosome entry site (IRES)-EGFP box is integrated between the last exon and the 3' UTR of the endogenous PCSK9 gene.
[0221] e. Using Hepa 1-6 PCSK9 IRES-GFP The reporter gene cell lines were selected for sgRNA and crRNA screening using either SpCas9-OFF-EE or SF01-OFF-EE. The upper dashed box indicates the SF01 crRNA binding site. The lower dashed box indicates the SpCas9 sgRNA binding site. The x-axis position of each line represents the position of the target site relative to the PCSK9 transcription start site (TSS) (in nucleotides). Light blue line: crRNA targeting the positive strand; red line: crRNA targeting the negative strand. Light green shading: Two regions targeted by crRNAs that are strongly silenced by SF01-OFF-EE. Purple line: na-sgRNA. Orange line: 10 sgRNAs designed around the na-sgRNA binding site.
[0222] f, g. Using 10 selected sgRNAs (f) and 41 crRNAs (g) from (e), Hepa1-6 PCSK9 was evaluated. IRES-GFP The effect and persistence of EGFP silencing in cells lasting up to one week.
[0223] h, i. Quantitative analysis of CpG methylation status within a 700 bp genomic region containing the PCSK9 CpG island was performed using targeted amplicon bisulfite sequencing.
[0224] Figure 3 In vitro screening and multi-omics analysis were used to determine the optimal sgRNA / crRNA configuration for epigenetically silent PCSK9 in mouse hepatocytes. Specifically:
[0225] a, b, in Hepa 1-6 PCSK9 IRES-GFP The silencing effect and persistence of EGFP for up to 7 days were evaluated in cells. Cells were treated with SpCas9-OFF-EE mRNA combined with the top three single sgRNAs or combinations of two or three of them (from...). Figure 2 Cells were co-delivered with SF01-OFF-EE mRNA (selected from 10 sgRNAs in f) (a). Cells were co-delivered with the top five single crRNAs or combinations of two or three of them (selected from 41 crRNAs in Fig. 2g) (b).
[0226] c and d represent the treatments in (a) and (b), respectively, quantifying the CpG methylation status within a 700 bp genomic region containing the PCSK9 CpG island (CGI). Methylation was analyzed by targeted amplicon bisulfite sequencing. Data represent the mean ± SD of three independent biological replicates.
[0227] e, CpG methylation level at PCSK9 site.
[0228] f, The bar chart shows the global CpG methylation level of the specified sample, determined by whole-genome bisulfite sequencing (WGBS) analysis (n=3 for each experimental condition).
[0229] g, CpG methylation profiles within a ±5 kb genomic region centered at the PCSK9 transcription start site (TSS). Hepa 1-6 mouse hepatocytes were treated with different epigenetic editors: SpCas9-OFF-EE, 573-Split-SpCas9-OFF-EE, or 713-Split-SpCas9-OFF-EE mRNA were co-delivered with Na-sgRNA and SgRNA-4; or SF01-OFF-EE mRNA was co-delivered with AS-cr7 and S-cr14 crRNA. Samples were collected 7 days post-delivery for WGBS. SpCas9-OFF-EE mRNA was delivered alone as a mock control. Genomic regions containing CGI are marked with yellow rectangles.
[0230] Manhattan plots of genome-wide methylation changes determined by h, i, and WGBS are used to compare cells (g) treated with SpCas9-OFF-EE / sgRNA and the effector-only control, and cells (h) treated with SF01-OFF-EE / crRNA and the effector-only control. Benjamini-Hochberg false discovery rate (FDR) corrected p-values (DSS Wald test, two-tailed) for each CpG are plotted against its genomic coordinates. Differentially methylated CpGs (DMCs) within the differentially methylated regions (DMRs) of PCSK9 are shown in red.
[0231] j. Volcano plot of RNA-seq analysis, showing differential gene expression between simulated-treatment cells and cells treated with SpCas9-EE-OFF / sgRNA (left) or SF01-EE-OFF / crRNA (right) (n=3 for each experimental condition). P-values were derived from Wald's test for a binomial distribution, corrected for multiple tests using the Benjamini-Hochberg method. Horizontal dashed lines represent the threshold for corrected P-values (FDR ≤ 0.05), and vertical dashed lines represent the threshold for |log2 FC| ≥ 1. Upregulated genes are marked in red, downregulated genes in light blue, and non-differentially expressed genes in gray. The PCSK9 gene is highlighted.
[0232] The k, l, scatter plot correlates the DMR methylation differential of WGBS (y-axis) with the gene expression log2FC of RNA-seq (x-axis), showing the status of all genes within each DMR ±20 kb range. Comparisons are made between SpCas9-EE and effector-only controls (j), and between SF01-EE and effector-only controls (k). The PCSK9 gene is highlighted in blue. Thresholds (grey dashed lines) are set for methylation (β value) differential >0.2 or < −0.2, and RNA-seq log2FC >1 or < −1. DEG represents differentially expressed genes.
[0233] Figure 4 Epigenetic silencing of PCSK9 in mouse liver following LNP-mediated delivery of epigenome editor mRNA. Specifically:
[0234] a, b, Dose-dependent effects of a single injection of an LNP formulation containing SpCas9-OFF-EE-V2 or SF01-OFF-EE-V2 mRNA and homologous PCSK9-targeting sgRNA or crRNA. Circulating PCSK9 protein levels (a) and plasma LDL-C levels (b) were measured in C57BL / 6 mice (n=6 per group) on day 7 post-injection.
[0235] c, d, Comparison of the effects of six different editors on silencing PCSK9. Circulating PCSK9 protein levels (c) and plasma LDL-C levels (d) in C57BL / 6 mice (n=6 per group) were measured on day 7 after injection of LNPs encapsulating various editor payloads (SpCas9 nuclease, SpCas9-OFF-EE-V2, 573-Split-SpCas9-OFF-EE, 713-Split-SpCas9-OFF-EE, SF01 nuclease, or SF01-OFF-EE-V2) and homologous PCSK9-targeting sgRNA or crRNA.
[0236] e. Volcano plots from RNA-seq analysis show differential gene expression in liver tissues of C57BL / 6 mice (n=6 per group) treated with LNP vector (control) versus LNP-delivered SpCas9-OFF-EE mRNA plus PCSK9-targeting sgRNA (left) or SF01-OFF-EE mRNA plus PCSK9-targeting crRNA (right). P-values were determined using a Wald test for a binomial distribution, with multiple correction (Benjamini-Hochberg method). Horizontal dashed lines represent the threshold for corrected P-values (FDR ≤ 0.05), and vertical dashed lines represent the fold change (FC) threshold of |log2 FC| ≥ 1. Upregulated genes are marked in red, downregulated genes in light blue, and non-differentially expressed genes in gray. The PCSK9 gene is also highlighted.
[0237] f. Gene Ontology (GO) biological process pathway enrichment analysis, targeting differentially expressed genes (DEGs) identified by RNA-seq in mouse livers treated with SpCas9-OFF-EE / PCSK9-sgRNA (left) or SF01-OFF-EE / PCSK9-crRNA (right).
[0238] g, The bar chart illustrates the global CpG methylation level in the liver tissue of treated mice, determined by whole-genome bisulfite sequencing (WGBS) (n=3 for each experimental condition).
[0239] Figure 5 Persistence and specificity analysis of PCSK9 silencing in vivo. Among them:
[0240] a. In vivo experimental flowchart. Lipid nanoparticles (LNPs) were formulated with mRNA encoding an epigenome editor or nuclease and sgRNA or crRNA homologously targeting PCSK9. A single intravenous (IV) injection was administered to C57BL / 6 mice. Blood samples were collected at specified time points to measure plasma PCSK9 and low-density lipoprotein cholesterol (LDL-C) levels. Liver tissue was collected on days 30 and 120 post-injection for RNA sequencing (RNA-seq) and whole-genome bisulfite sequencing (WGBS).
[0241] b, c, Time-series analysis of circulating PCSK9 protein levels (b) and plasma LDL-cholesterol (LDL-c) levels (c) for up to 180 days after a single administration of an LNP formulation that delivers nuclease or EE mRNA and corresponding PCSK9-targeting sgRNA / crRNA (n=6 mice per group).
[0242] Representative CpG methylation profiles of a ±5kb genomic region centered at the PCSK9 transcription start site (TSS) in liver tissues of C57BL / 6 mice (d, e, and c). Mice were treated with various epigenetic editors: SpCas9-OFF-EE, 573-Split-SpCas9-OFF-EE, or 713-Split-SpCas9-OFF-EE mRNA co-delivered with Na-sgRNA and SgRNA-4; or SF01-OFF-EE mRNA co-delivered with AS-cr7 and S-cr14 crRNA. Liver tissues were collected at 30 and 120 days post-delivery for WGBS. LNP vector alone was administered as a mock control. Genomic regions containing CGI are marked with yellow rectangles.
[0243] f, g, Manhattan plots depicting genome-wide methylation changes in liver tissue, identified by WGBS. Comparisons show mice treated with SpCas9-OFF-EE / PCSK9-sgRNA versus LNP-only control (f), or SF01-OFF-EE / PCSK9-crRNA versus LNP-only control (g). Benjamini-Hochberg false discovery rate (FDR) corrected p-values (DSS Wald test, two-sided) for each CpG are plotted against its genomic coordinates. Differentially methylated CpGs (DMCs) within differentially methylated regions (DMRs) of PCSK9 are shown in red.
[0244] The h, i, scatter plots correlate WGBS DMR methylation differentials (y-axis) with RNA-seq log2 fold change (FC) of gene expression (x-axis), showing the status of all genes within each DMR ±20 kb range in liver tissue. Comparisons are made between SpCas9-OFF-EE and LNP-only vector controls (h), and SF01-OFF-EE and LNP-only vector controls (i). The PCSK9 gene is highlighted in blue. Thresholds (grey dashed lines) are set for methylation (β value) differential > 0.2 or < −0.2, and RNA-seq log2FC > 1 or < −1. DEG represents differentially expressed genes.
[0245] Figure 6 Endogenous CD151 was epigenetically silenced in HEK293T cells using SpCas9-OFF and SF01-OFF epigenetic editor plasmids. Specifically:
[0246] a. Representative flow cytometry histogram of unstained control HEK293T cells (without anti-CD151 antibody) at 30 days. Percentage indicates cells gated as CD151 negative.
[0247] b, Representative flow cytometry histogram of CD151 expression in untreated HEK293T cells stained at 30 days, showing the baseline positive population. Percentage indicates cells gated for CD151 staining positivity.
[0248] c, d, Representative flow cytometry histograms of CD151 expression 30 days after HEK293T cells transfected with SpCas9-OFF epigenetic editor (EE) plasmids and non-targeted (NT) sgRNA (c) or CD151-targeting sgRNA (d). Percentages represent the proportion of cells within the unsilenced and potentially silenced populations as defined by CD151 expression levels.
[0249] e, f, Representative flow cytometry histograms of CD151 expression 30 days after HEK293T cells transfected with SF01-OFF-EE plasmid and non-targeted (NT) crRNA (e) or CD151-targeting crRNA (f). Percentages represent the proportion of cells within the defined unsilenced and silenced populations. All experiments were performed in triplicate, and representative histograms are shown. Quantitative data derived from these replicates are presented in Figure 1e with standard deviation.
[0250] Figure 7Flow cytometry analysis of Gapdh-Snrpn-GFP reporter gene silencing in HEK293T cells using unoptimized SpCas9-OFF-EE and SF01-OFF-EE mRNA (V1). Among them:
[0251] a. A representative flow cytometry histogram of untreated HEK293T cells (cultured for 30 days) expressing the Snrpn-GFP reporter gene, serving as a mock control. Percentages represent the proportion of cells that were GFP-positive (unsilenced) and GFP-negative based on baseline fluorescence gating.
[0252] Representative flow cytometry histograms of Snrpn-GFP expression 30 days after electroporation of HEK293T cells with SpCas9-OFF-EE mRNA and non-targeted (NT) sgRNA (b) or sgRNA targeting the Snrpn site (c). Percentages represent the proportion of cells gated as GFP-positive (unsilenced) and GFP-negative (silenced).
[0253] d, Representative flow cytometry histogram of untreated HEK293T cells (cultured for 30 days) expressing the Snrpn-GFP reporter gene, serving as a control for the SF01-OFF-EE experiment. Percentages represent the proportion of cells that were GFP-positive (unsilenced) and GFP-negative based on baseline fluorescence gating.
[0254] Representative flow cytometry histograms of Snrpn-GFP expression 30 days after electroporation of SF01-OFF-EE mRNA and non-targeted (NT) crRNA (e) or Snrpn-targeted crRNA (f) in HEK293T cells. Percentages represent the proportion of cells gated as GFP-positive (unsilenced) and GFP-negative (silenced). Representative histograms are shown. All experiments were performed in triplicate. Quantitative data derived from these replicates are shown in Figures 2b and 2c with standard deviation.
[0255] Figure 8 Epigenetic silencing of endogenous CD151 expression was achieved in HEK293T cells using unoptimized SpCas9-OFF-EE and SF01-OFF-EE mRNA (V1). Specifically:
[0256] a, d, Representative flow cytometry histograms of unstained control HEK293T cells (without anti-CD151 antibody) cultured for 30 days. Percentages indicate cells gated as antibody-negative.
[0257] b, c, HEK293T cells electroporated with SpCas9-OFF-EE V1 mRNA and non-targeted (NT) sgRNA (b) or a mixed pool of three chemically synthesized sgRNAs targeting CD151 (c) Flow cytometry histograms representing CD151 expression 30 days later. Percentages represent the proportion of cells within the unsilenced and potentially silenced populations as defined by CD151 expression levels.
[0258] e, f, HEK293T cells electroporated with SF01-OFF-EE V1 mRNA and non-targeted (NT) crRNA (e) or a mixed pool of three chemically synthesized crRNAs targeting CD151 (f). Representative flow cytometry histograms of CD151 expression 30 days after electroporation. Percentages represent the proportion of cells within the defined unsilenced and silenced populations. Representative histograms are shown. All experiments were performed in triplicate. Quantitative data derived from these replicates are shown in Figures 2b and 2c with standard deviation.
[0259] Figure 9 Delivery of unoptimized SpCas9 / SF01-OFF-EE V1 mRNA via lipid nanoparticles (LNPs) did not alter serum PCSK9 or LDL-cholesterol levels in C57 mice.
[0260] Serum PCSK9 protein expression levels on day 7 after intravenous injection of unoptimized SpCas9 / SF01-OFF-V1 mRNA encapsulated in LNP in C57 mice. No significant changes in serum PCSK9 levels were observed compared to baseline or vector control groups.
[0261] Serum low-density lipoprotein cholesterol (LDL-C) levels on day 7 after intravenous injection of unoptimized SpCas9 / SF01-OFF-V1 mRNA encapsulated in LNP in C57 mice. No significant decrease in LDL-C levels was detected. Results are expressed as mean ± standard deviation (sd) (n=4 mice per group).
[0262] Figure 10 In HEK293T cells, optimized mRNA resulted in enhanced nuclear localization of SpCas9-OFF-EE. Specifically:
[0263] a. Representative fluorescence microscopy images of HEK293T cells 48 hours after transient co-transfection with mCherry mRNA (transfection control) and unoptimized (V1) or optimized (V2) SpCas9-OFF-EE-GFP mRNA. For ease of observation, GFP was inserted downstream of the KRAB domain in the SpCas9-OFF-EE V1 and V2 constructs (as shown in Figure 2a). Scale bar = 10 µm.
[0264] b. Quantification of the percentage of nuclear GFP signal. Each data point represents the ratio of the mean nuclear GFP signal intensity to the total cellular GFP intensity in all GFP-positive cells within a single microscopic field of view. Data are expressed as mean ± SD of three independent experiments. Statistical significance between groups V1 and V2 was determined using Student's t-test.
[0265] Figure 11 Optimized SpCas9-OFF-EE V2 mRNA exhibited improved translation efficiency and stability in HEK293T cells. Specifically:
[0266] a. Representative flow cytometry histograms of HEK293T cells expressing mCherry and EE mRNA. Percentages represent the proportion of cells that were mCherry or GFP positive or negative based on baseline fluorescence gating. HEK293T cells were co-transfected with mCherry mRNA (transfection control) and unoptimized (V1) or optimized (V2) SpCas9-OFF-EE-GFP mRNA. GFP and mCherry expression were analyzed by flow cytometry at 24, 96, and 168 hours post-electroporation.
[0267] b. Quantitative analysis of mCherry expression in cells co-transfected with mRNA at specified time points. No significant differences in mCherry expression (percentage of mCherry-positive cells) were observed between time points or between cells transfected with V1 or V2 SpCas9-OFF-EE-GFP mRNA. Data are presented as mean ± SD (n = 3 independent experiments).
[0268] c. Comparison of expression of V1 (unoptimized) and V2 (optimized) SpCas9-OFF-EE-GFP mRNA constructs in HEK293T cells at 24, 96, and 168 hours post-electroporation. The optimized SpCas9-OFF-EE-GFP V2 mRNA showed significantly higher expression levels (percentage of GFP-positive cells) than V1 mRNA at all evaluation time points. Data are presented as mean ± SD (n = 3 independent experiments). Two-way ANOVA was used to compare the V1 and V2 groups, with V2 used as the control column for multiple comparisons (Supplementary Table); **** indicates P ≤ 0.0001.
[0269] Figure 12 Compared to unoptimized EE V1 mRNA, optimized epigenome editor (EE) V2 mRNA showed higher purity and lower dsRNA content. Specifically:
[0270] a, b, capillary electrophoresis analysis of the integrity and purity of unoptimized EE V1 (a) and EE V2 (b) mRNA.
[0271] c. Statistical analysis of dsRNA content after in vitro transcription before and after mRNA optimization. The optimized mRNA vector produced significantly less dsRNA after in vitro transcription than the unoptimized vector. Quantitative comparison of double-stranded RNA (dsRNA) byproduct levels in in vitro transcribed EE V1 (unoptimized) and EE V2 (optimized) mRNA. The optimized EE V2 mRNA formulation showed significantly lower dsRNA content than the unoptimized EE V1 mRNA. Data are expressed as mean ± SD (n = 3 independent in vitro transcription reactions). Statistical significance between the V1 and V2 groups was determined using Student's t-test. ***P ≤ 0.001.
[0272] Figure 13 Comparison of the physicochemical properties of LNPs packaged in SpCas9-OFF-EE-V1 and SpCas9-OFF-EE-V2. Among them:
[0273] a. Intensity size distribution and polydispersity index (PDI) of LNPs encapsulating SpCas9-OFF-EE-V1 mRNA.
[0274] b. Intensity size distribution and polydispersity index (PDI) of LNPs encapsulating SpCas9-OFF-EE-V2 mRNA.
[0275] Figure 14 Optimized SpCas9-OFF-EE V2 (SEQ ID NO: 136) and SF01-OFF-EE V2 mRNA (SEQ ID NO: 137) resulted in dose-dependent GFP silencing at 7 and 21 days after delivery to Gapdh-Snrpn-GFP reporter cells. Among them:
[0276] a, b, dose-response curves illustrating GFP silencing following delivery of increasing concentrations of SpCas9-OFF-EE V2 mRNA (sgRNA with homologous Snrpn targeting) or SF01-OFF-EE V2 mRNA (crRNA with homologous Snrpn targeting) to the Gapdh-Snrpn-GFP reporter gene. GFP expression was assessed at 7 days (a) and 21 days (b) post-mRNA delivery. Data points represent the mean ± SD of three independent biological replicates.
[0277] c, d, Calculate the half-maximal effect concentration (EC50) of silencing GFP for each epigenome editor (EE) platform on day 7 (c) and day 21 (d). 50 By fitting the data to a four-parameter logistic model (R² for all conditions) 2 > 0.97) Determine EC 50 value.
[0278] Figure 15CpG methylation status at Snrpn sites after electroporation with optimized (V2) or unoptimized (V1) SpCas9-OFF-EE and SF01-OFF-EE mRNA. CpG methylation across a 370 bp region of the Snrpn gene was quantified one week (top) and three weeks (bottom) after electroporation of Gapdh-Snrpn-GFP reporter cells. Cells were treated with electroporation with SpCas9-OFF-EE (V1 or V2) and sgRNA targeting Snrpn, or SF01-OFF-EE (V1 or V2) and crRNA targeting Snrpn (Supplementary Table). The heatmap shows the percentage of methylated CpG dinucleotides at individual sites, with each square representing a specific CpG. Short blue lines indicate sgRNA binding sites, and short purple lines indicate crRNA binding sites. The x-axis position of each line represents the distance (in nucleotides) of the target site relative to the Snrpn transcription start site (TSS). The data represent the average percentage of CpG methylation from three independent experiments.
[0279] Figure 16 CpG methylation status of CD29 CpG islands (CGIs) after electroporation with optimized (V2) or unoptimized (V1) SpCas9-OFF-EE and SF01-OFF-EE mRNA. Quantification of CpG methylation across the CD29 gene CGI region at one week (top) and three weeks (bottom) after electroporation of HEK293T cells. Cells were treated with either electroporation with SpCas9-OFF-EE (V1 or V2) and a pool of three homologous CD29-targeting sgRNAs, or SF01-OFF-EE (V1 or V2) and a pool of three homologous CD29-targeting crRNAs (Supplementary Table). The heatmap shows the percentage of methylated CpG dinucleotides at individual sites, with each square representing a specific CpG. Short blue lines indicate sgRNA binding sites, and short purple lines indicate crRNA binding sites. The x-axis position of each line represents the distance (in nucleotides) of the target site relative to the CD29 transcription start site (TSS). The data represent the average percentage of CpG methylation from three independent experiments.
[0280] Figure 17CpG methylation status of CD81 CpG islands (CGIs) after electroporation with optimized (V2) or unoptimized (V1) SpCas9-OFF-EE and SF01-OFF-EE mRNA. Quantification of CpG methylation across the CD81 gene CGI region at one week (top) and three weeks (bottom) after electroporation of HEK293T cells. Cells were treated with either electroporation with SpCas9-OFF-EE (V1 or V2) and a pool of three homologous CD81-targeting sgRNAs, or SF01-OFF-EE (V1 or V2) and a pool of three homologous CD81-targeting crRNAs (Supplementary Table). The heatmap shows the percentage of methylated CpG dinucleotides at individual sites, with each square representing a specific CpG. Short blue lines indicate sgRNA binding sites, and short purple lines indicate crRNA binding sites. The x-axis position of each line represents the distance (in nucleotides) of the target site relative to the CD81 transcription start site (TSS). The data represent the average percentage of CpG methylation from three independent experiments.
[0281] Figure 18 CpG methylation status of CD151 CpG islands (CGIs) after electroporation with optimized (V2) or unoptimized (V1) SpCas9-OFF-EE and SF01-OFF-EE mRNA. Quantification of CpG methylation across the CD151 gene CGI region at one week (top) and three weeks (bottom) after electroporation of HEK293T cells. Cells were treated with either electroporation with SpCas9-OFF-EE (V1 or V2) and a pool of three homologous CD151-targeting sgRNAs, or SF01-OFF-EE (V1 or V2) and a pool of three homologous CD151-targeting crRNAs (Supplementary Table). The heatmap shows the percentage of methylated CpG dinucleotides at individual sites, with each square representing a specific CpG. Short blue lines indicate sgRNA binding sites, and short purple lines indicate crRNA binding sites. Squares represent the percentage of methylated CpG dinucleotides. The x-axis position of each line represents the distance (in nucleotides) of the target site relative to the CD151 transcription start site (TSS). The data represent the average percentage of CpG methylation from three independent experiments.
[0282] Figure 19 Hepa 1-6 PCSK9 IRES-GFP Functional validation of reporter gene cell lines. Among them:
[0283] a, b. Explanation of Hepa 1-6 PCSK9 IRES-GFP Dose-response curves for EGFP silencing in reporter genes. Cells were treated with increasing concentrations of SpCas9-OFF-EE V2 mRNA (sgRNA with homologous target for PCSK9) or SF01-OFF-EE V2 mRNA (crRNA with homologous target for PCSK9). EGFP protein expression was assessed by flow cytometry at 7 days (a) and 21 days (b) after mRNA delivery. Data points represent the mean ± SD of three independent biological replicates.
[0284] c, Hepa 1-6 PCSK9 IRES-GFP Pearson correlation analysis was performed on the EGFP silencing efficiency in reporter genes in cells and different doses (1, 2, and 4 µg) of SpCas9-OFF-EE V2 mRNA (with homologous sgRNA) or SF01-OFF-EE V2 mRNA (with homologous crRNA). EGFP signal was measured 7 days after mRNA delivery.
[0285] d, Hepa 1-6 PCSK9 IRES-GFP qRT-PCR analysis of PCSK9 mRNA expression in reporter genes. Cells were nuclear transfected with specified doses (1, 2, 4, and 8 µg) of SpCas9-OFF-EE V2 mRNA (with homologous sgRNA) or SF01-OFF-EE V2 mRNA (with homologous crRNA). PCSK9 mRNA levels were quantified 7 days after mRNA delivery.
[0286] e, Hepa 1-6 PCSK9 IRES-GFP Pearson correlation analysis of EGFP mRNA levels and PCSK9 mRNA expression in reporter gene cells. Cells were nuclear transfected with specified doses (1, 2, 4, and 8 µg) of SpCas9-OFF-EE V2 mRNA (with homologous sgRNA) or SF01-OFF-EE V2 mRNA (with homologous crRNA). Measurements were performed 7 days after mRNA delivery.
[0287] f, Hepa 1-6 PCSK9 IRES-GFP Pearson correlation analysis of the reporter gene EGFP mRNA level in cells and the CpG methylation status of PCSK9 CpG islands (CGI). Hepa 1-6 PCSK9IRES-GFP Reporter cells were nuclearly transfected with SpCas9-OFF-EE V2 mRNA (with homologous sgRNA) or SF01-OFF-EE V2 mRNA (with homologous crRNA) at specified doses (1, 2, 4, and 8 µg). Analysis was performed at 7 and 21 days post-mRNA delivery.
[0288] Figure 20 .SpCas9-OFF-EE in Hepa 1-6 PCSK9 IRES-GFP Persistent EGFP silencing can be achieved in reporter gene cells through different sgRNA configurations. Among them:
[0289] a, b in Hepa 1-6 PCSK9 IRES-GFP The effectiveness and persistence of EGFP silencing for up to 21 days were evaluated in cells. Cells were treated with SpCas9-OFF-EE mRNA co-delivered with 10 single sgRNAs (a) or combinations of two or three sgRNAs (b). Individual data points and mean values are shown (n=3 biological replicates).
[0290] c and d correspond to the treatments in (a) and (b), respectively, quantifying the CpG methylation status within a 700 bp genomic region containing the PCSK9 CpG island. Methylation was analyzed by targeted amplicon bisulfite sequencing. Data represent the mean ± SD of three independent biological replicates.
[0291] Figure 21 SF01-OFF-EE in Hepa 1-6 PCSK9 IRES-GFP Persistent EGFP silencing can be achieved in reporter gene cells through different crRNA configurations. Among them:
[0292] a, b in Hepa 1-6 PCSK9 IRES-GFP The effectiveness and persistence of EGFP silencing for up to 21 days were evaluated in cells. Cells were treated with SpCas9 / SF01-OFF-EE mRNA co-delivered with 41 single sgRNAs (a) or combinations of two or three sgRNAs (b). Individual data points and mean values are shown (n=3 biological replicates).
[0293] c and d correspond to the treatments in (a) and (b), respectively, quantifying the CpG methylation status within a 700 bp genomic region containing the PCSK9 CpG island. Methylation was analyzed by targeted amplicon bisulfite sequencing. Data represent the mean ± SD of three independent biological replicates.
[0294] Figure 22 The target-specific transcription of PCSK9 was downregulated after epigenetic silencing of PCSK9 using Split-SpCas9-OFF-EE. Specifically:
[0295] Manhattan plots of genome-wide methylation changes identified by WGBS, a, b, comparing cells treated with 573-Split-SpCas9-OFF-EE / sgRNA versus an effector-only control (a), or 713-Split-SpCas9-OFF-EE / sgRNA versus an effector-only control (b). Benjamini-Hochberg false discovery rate (FDR) corrected p-values (DSS Wald test, two-tailed) for each CpG are plotted against its genomic coordinates. Differentially methylated CpGs (DMCs) within PCSK9 differentially methylated regions (DMRs) are shown in red.
[0296] c. Volcano plot of RNA-seq analysis, showing differential gene expression between cells treated with simulated RNA and those treated with 573-Split-SpCas9-OFF-EE / sgRNA (left) or 713-Split-SpCas9-OFF-EE / sgRNA (right) (n=3 for each experimental condition). P-values were derived from Wald's test for a binomial distribution, corrected for multiple tests using the Benjamini-Hochberg method. Horizontal dashed lines represent the threshold for corrected P-values (FDR ≤ 0.05), and vertical dashed lines represent the threshold for |log2FC| ≥ 1. Upregulated genes are marked in red, downregulated genes in light blue, and non-differentially expressed genes in gray. PCSK9 is highlighted.
[0297] Scatter plots d and e correlate WGBS DMR methylation differentials (y-axis) with RNA-seq gene expression log2FC (x-axis), showing all genes within each DMR ±20 kb range. Comparisons are made between 573-Split-SpCas9-OFF-EE and effector-only controls (d), and 713-Split-SpCas9-OFF-EE and effector-only controls (e). The PCSK9 gene is highlighted in blue. Thresholds (grey dashed lines) are set for methylation (β value) differential >0.2 or < −0.2, and RNA-seq log2FC >1 or < −1. DEG represents differentially expressed genes.
[0298] Figure 23 Physicochemical properties of LNPs encapsulated in SpCas9-OFF-EE or SF01-OFF-EE. Among them:
[0299] a. Intensity size distribution of LNPs. b. Size and polydispersity index (PDI) of LNPs (n=3). c. Encapsulation efficiency of LNPs as measured using the RiboGreen RNA Detection Kit (n=3). Data represent mean ± sd.
[0300] Figure 24 Analysis of the efficacy and specificity of in vivo Split-SpCas9-OFF-EE silencing of PCSK9. Among them:
[0301] a, b Manhattan plots depicting genome-wide methylation changes in liver tissue, identified by WGBS. Comparisons show mice treated with 573-Split-SpCas9-OFF-EE / PCSK9-sgRNA versus LNP-only control (a), or 713-Split-SpCas9-OFF-EE / Pcsk-sgRNA versus LNP-only control (b). Benjamini-Hochberg false discovery rate (FDR) corrected p-values (DSS Wald test, two-sided) for each CpG are plotted against its genomic coordinates. Differentially methylated CpGs (DMCs) within differentially methylated regions (DMRs) of PCSK9 are shown in red.
[0302] c. Volcano plot of RNA-seq analysis, showing differential gene expression in liver tissues of C57BL / 6 mice (n=6 per group) treated with LNP vector (control) and LNP delivery of 573-Split-SpCas9-OFF-EE mRNA plus PCSK9-targeting sgRNA (left) or 713-Split-SpCas9-OFF-EE mRNA plus PCSK9-targeting sgRNA (right). P-values were determined using a Wald test for a binomial distribution, with multiple correction (Benjamini-Hochberg method). Horizontal dashed lines represent the threshold for corrected P-values (FDR ≤ 0.05), and vertical dashed lines represent the fold change (FC) threshold of |log2 FC| ≥ 1. Upregulated genes are marked in red, downregulated genes in light blue, and non-differentially expressed genes in gray. PCSK9 gene is marked.
[0303] Scatter plots d and e correlate WGBS DMR methylation differentials (y-axis) with RNA-seq log2 fold change (FC) of gene expression (x-axis), showing the status of all genes within each DMR ±20 kb range in liver tissue. Comparisons are made between 573-Split-SpCas9-OFF-EE and the LNP-only vector control (d), and 713-Split-SpCas9-OFF-EE and the LNP-only vector control (e). The PCSK9 gene is highlighted in blue. Thresholds (grey dashed lines) are set for methylation (β value) differential > 0.2 or < −0.2, and RNA-seq log2FC > 1 or < −1. DEG, differentially expressed genes.
[0304] Figure 25 LNP-mediated in vivo effects and biodistribution after delivery to various editor systems. Among them:
[0305] a. The bar chart shows the mean percentage of PCSK9 promoter CpG methylation in the specified organs. Mice were analyzed on day 7 after injection of a specified LNP formulation containing a combination of various epigenetic editor mRNAs and sgRNA or crRNA targeting the PCSK9 promoter (n=6 per group).
[0306] b. The bar chart shows the percentage of edited alleles (e.g., insertions and deletions) in the designated organ on day 7 post-injection. Mice received a designated LNP formulation (n=6 per group) encapsulating either Cas9 or SF01 nuclease-encoded mRNA and homologous sgRNA or crRNA targeting PCSK9 exon 1.
[0307] Figure 26 Histological analysis of various organs following LNP-mediated delivery of epigenetic editors. Hematoxylin and eosin (H&E) staining of liver, heart, lung, spleen, and kidney sections from C57BL / 6 mice. Tissues were collected on day 7 after injection of LNP formulations encapsulating various epigenetic editors at a dose of 3 mg / kg. Scale bar = 100 µm.
[0308] Figure 27 Assessment of potential systemic toxicity following LNP-mediated delivery of various editor systems. Plasma levels of (a) alanine aminotransferase (ALT), (b) aspartate aminotransferase (AST), (c) creatinine (CR), and (d) urea were measured in mice (n=5 per group) after LNP-mediated delivery of mRNAs encoding different editor payloads.
[0309] Figure 28 Screening of crRNAs targeting human PCSK9 using the SF01-OFF-EE epigenome editor. The image shows the results of quantification of PCSK9 mRNA extracted by qPCR 21 days post-transfection from 60 crRNAs targeting human PCSK9. Detailed Implementation
[0310] The following examples are for illustrative purposes only and are not intended to limit the invention. Unless otherwise specified, the experiments and methods described in the examples are generally performed according to conventional methods well known in the art and described in various references. For example, conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA used in this invention can be found in Sambrook, Fritsch, and Maniatis, *Molecular Cloning: A Laboratory Manual*, 2nd edition (1989); *Current Protocols in Molecular Biology* (edited by FM. Ausubel et al., (1987)); and the *Methods in Enzymology* series (academic publishing company): *PCR 2: A PRACTICAL*. APPROACH (edited by MJ MacPherson, BD Hames and GR Taylor (1995)), Harlow and Lane (1988) Antibodies, A Laboratory Manual, and Animal Cell Culture (edited by R.R. Freshney (1987)).
[0311] Furthermore, unless specific conditions are specified in the examples, conventional conditions or conditions recommended by the manufacturer should be followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products. Those skilled in the art will understand that the examples are described by way of illustration and are not intended to limit the scope of protection claimed by the invention. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety.
[0312] Example 1. Materials and Methods
[0313] General methods and molecular cloning:
[0314] Expression plasmids for sgRNA or crRNA were constructed based on a previously described backbone (Liang, SQ et al. Genome-wide profiling of prime editor off-target sites in vitro and in vivo using PE-tag. Nat Methods 20, 898-907 (2023).). Briefly, gblocks (ordered from Genscript) containing spacer-scaffold sequences (for sgRNA) or 3' direct repeat (DR)-spacer sequences (for crRNA) were cloned into U6 expression vectors via the Gibson assembly method. These vectors were linearized by digestion with BstXI and BlpI. These guide RNA expression plasmids also co-expressed the T2A-mCherry tag from a single promoter to assess transfection efficiency.
[0315] The mammalian expression plasmid SpCas9-OFF-EE V1 was obtained from Addgene (Addgene, #167981). To generate the Cas-SF01-OFF-EE V1 mammalian expression vector, a dSF01 (E844A) catalytic inactivation domain was created from a wild-type Cas-SF01 nuclease template via site-directed mutagenesis, which was then used to replace the dSpCas9 fragment in the SpCas9-OFF-EE V1 vector via Gibson assembly. The lentiviral plasmid used to generate the Gapdh-Snrpn-GFP reporter cell line was also obtained from Addgene (Addgene, #70148). All plasmids used in transient transfection experiments were purified using a kit including an endotoxin removal step (ZymoPURE Plasmid Miniprep Kit, Zymo Research).
[0316] The amino acid sequence of the Cas-SF01 wild-type protein is shown below:
[0317]
[0318] Construction of OFF-EE V1 and V2 mRNA in vitro transcription (IVT) vectors:
[0319] For the construction of the OFF-EE V1 mRNA IVT vector, the coding sequences of the dSpCas9 / dSF01, Dnmt3A-3L, and Krab modules were assembled into the IVT plasmid backbone digested with BamHI / NotI via Gibson assembly. This vector contained the 5′ and 3′ untranslated regions (UTRs) from human β-globin, the Kozak consensus sequence, and a 100-bp polyadenylate signal. For the optimized OFF-EE V2 mRNA IVT vector, codon-optimized OFF-EE coding sequences (dSpCas9 / dSF01, 573-Split-SpCas9, 713-Split-SpCas9, Dnmt3A-3L, and Krab) were cloned into the PEmax mRNA plasmid backbone digested with EcoRI via Gibson assembly. To facilitate visualization experiments, the GFP coding sequence was inserted between the Krab domain and the C-terminal nuclear localization signal (NLS) in the V1 and V2 SpCas9-OFF-EE IVT vectors via Gibson assembly, creating the SpCas9-OFF-EE-GFP construct.
[0320] RNA production and purification:
[0321] For the guide RNA required for in vitro experiments, sgRNA and crRNA are chemically synthesized by a commercial supplier (Genscript) and contain end chemical modifications to enhance stability, including 2′-O-methylated nucleosides at the first three and last three positions, linked by 3′ thiophosphate nucleosides.
[0322] Epigenetic editor mRNA was generated via in vitro transcription (IVT). The plasmid DNA template was first linearized with a PmeI restriction enzyme, which cleaves downstream of the polyadenylated tail sequence. The linearized DNA (500 ng) was then used as a template for the transcription reaction using the HiScribe T7 High-Yield RNA Synthesis Kit (NEB), which included co-transcriptional capping (using CleanCap® AG, TriLink Biotechnologies) and complete UTP replacement with N1-methylpseudouridine-5′-triphosphate (m1ΨTP, TriLink Biotechnologies) to reduce immunogenicity and enhance translation. After 1 hour of incubation, the DNA template was digested with DNase I, and the resulting mRNA was purified using the Monarch® RNA Cleanup Kit (NEB) and eluted in nuclease-free water. The mRNA concentration was quantified using a NanoDrop One UV-Vis spectrophotometer (Thermo Fisher Scientific), and integrity was assessed by capillary gel electrophoresis. The purified mRNA was stored at -80°C. For all in vivo mouse experiments, an additional cellulose purification step was performed to effectively remove dsRNA contaminants.
[0323] Cell culture:
[0324] Hepa1-6 (ATCC CRL-1830) and HEK293T (ATCC CRL-3216) cell lines were maintained in Dulbecco modified Eagle medium (DMEM, Thermo Fisher Scientific) supplemented with 10% fetal bovine serum (FBS) and 1% penicillin-streptomycin. Cells were cultured in a humidified incubator at 37°C and 5% CO2.
[0325] Plasmid and mRNA delivery in mammalian cells:
[0326] Plasmid delivery: Cells were seeded in 12-well plates and transfected using Lipofectamine 2000 (Thermo Fisher Scientific) according to the manufacturer's instructions when 60-70% confluence was achieved. A total of 1 µg of SpCas9-OFF-EE or SF01-OFF-EE plasmid and 500 ng of the corresponding sgRNA / crRNA expression vector were transfected. In the Gapdh-Snrpn-GFP reporter cell assay, cells co-expressing BFP (from the EE plasmid) and mCherry (from the guide plasmid) were sorted by FACS three days post-transfection to enrich successfully transfected cells, and cells were continuously cultured to monitor changes in GFP expression over time.
[0327] mRNA delivery: For imaging experiments, Lipofectamine MessengerMAX (Thermo Fisher Scientific) was used for delivery; for functional silencing assays, electroporation was used. For imaging experiments, 1x10... 5 HEK293T cells were seeded in 24-well plates on poly-L-lysine-coated coverslips. Cells were co-transfected with 4 µg of SpCas9-OFF-EE-GFP (V1 or V2) mRNA and 4 µg of mCherry mRNA. After 48 hours, cells were fixed with 4% paraformaldehyde (PFA), stained with DAPI, and imaged using a Zeiss LSM 980 confocal microscope. For time-dependent analysis of protein expression, transfected cells were analyzed by flow cytometry at 24, 96, and 7 days post-transfection. For functional silencing assays, mRNA was delivered to HEK293T, Hepa1-6, or reporter cell lines using a Neon™ electroporation system (100 µL kit, Thermo Fisher Scientific) with cell type-specific parameters (e.g., HEK293T: 1200 V, 20 ms, 2 pulses; Hepa1-6: 1350 V, 20 ms, 2 pulses). For each reaction, 1 x 102 cells were used. 6 Cell pellets (300 xg, 5 min) were resuspended in 110 µL Neon™ Buffer R and the specified reagents were added. For dose-dependent GFP silencing, Gapdh-Snrpn-GFP or Hepa1-6 PCSK9IRES-EGFP reporter cells were electroporated with different amounts of SpCas9-OFF-EE V2 or SF01-OFF-EE V2 mRNA and the corresponding guide RNA. To compare the efficacy of the mRNA editor, HEK293T or Hepa1-6 cells were electroporated with 4 µg of mRNA and 200 pmol of synthetic sgRNA(s) or crRNA(s). Endogenous sites targeted in HEK293T cells included CD29, CD81, and CD151, and multiple sgRNAs (SEQ ID NO: 7-15) or crRNAs (SEQ ID NO: 16-24) were designed. Seven or 21 days after electroporation, genomic DNA was isolated from all groups and stored at -80°C for subsequent library preparation and target amplicon deep sequencing.
[0328] Generation of reporter cell lines:
[0329] The Gapdh-Snrpn-GFP reporter cell line was generated via lentiviral transfection. Lentiviral virus was generated by co-transfecting HEK293T cells with the Gapdh-Snrpn-GFP lentiviral plasmid (Addgene, #70148) and packaging plasmids (psPAX.2 and psMD2.G). The viral supernatant was harvested, filtered, and used to infect HEK293T cells. Stably transduced cells were selected with puromycin, and a GFP-positive single-cell clone was sorted by FACS to establish the cell line.
[0330] Hepa1-6 PCSK9 IRES-EGFP Reporter cell line: generated via CRISPR-Cas9-mediated homologous recombination. Hepa1-6 cells were co-transfected with a Cas9 expression plasmid, an sgRNA plasmid targeting the 3' UTR of the endogenous mouse PCSK9 gene, and a donor plasmid containing the IRES-EGFP-T2A-BSD-polyA box (flanked by the homologous arm). Stable edited cells were selected with blastomycin, and a GFP-positive single-cell clone was sorted by FACS to establish the cell line.
[0331] RNA extraction and RT-qPCR:
[0332] Total RNA was extracted from cells using the RNeasy Kit (Takara Bio), and cDNA was synthesized using the PrimeScript™ RTMaster Mix (Takara Bio). Quantitative real-time PCR (qRT-PCR) was performed on a QuantStudio 5 Real-Time PCR system (Thermo Fisher Scientific) using TB Green Premix Ex Taq II (Takara Bio). Gene expression was quantified using specific primers for mouse PCSK9 and GAPDH (housekeeping genes). Relative fold changes were calculated using the ΔΔCT method.
[0333] Flow cytometry:
[0334] The expression of CD29, CD81, and CD151 on the surface of live cells was quantified using fluorescein-conjugated monoclonal antibodies. Briefly, single-cell suspensions were prepared in ice-cold FACS buffer (PBS pH 7.4, supplemented with 10% [v / v] heat-inactivated FBS and 2 mM EDTA). Cells were pre-blocked for 10 min at 4°C with Human TruStain FcX (BioLegend, 422302; 1:100 dilution), and then stained with predetermined optimal concentrations of APC-conjugated anti-CD29 (BioLegend, 303008), FITC-conjugated anti-CD81 (BioLegend, 349504), APC-conjugated anti-CD151 (BioLegend, 350406), and corresponding isotype controls (all from BioLegend). After incubation in the dark at 4°C for 30 min, cells were washed three times with 2 mL of FACS buffer (centrifuged at 300 × g for 5 min). Cells were resuspended in 200 µl of fresh FACS buffer and analyzed immediately on a Beckman CytoFLEX flow cytometer equipped with 488 nm and 638 nm lasers. GFP expression dynamics were quantified by flow cytometry for all engineered reporter cell lines. Briefly, untreated wild-type cell lines were used as negative controls, and Gapdh-Snrpn-GFP and Hepa1-6 PCSK9IRES-EGFP cell lines were used as positive controls. The percentage of GFP-negative cells after epitope editing tool treatment was analyzed.
[0335] ELISA and blood biochemistry tests:
[0336] To track PCSK9 levels in mouse serum, blood was collected from the orbital vein using a serum separation tube. Serum was separated by centrifugation at 2000 g for 120 minutes. PCSK9 levels were measured using the Mouse PCSK9 ELISA Kit (Proteintech, KE10050) according to the manufacturer's instructions. Briefly, serum from each mouse was diluted 1:200 with sample dilution buffer. After adding the diluted mouse serum and incubating for 2 hours, antibodies binding to the protein were detected using HRP-conjugated secondary antibodies. TMB substrate solution was added, and the reaction was terminated 15 minutes after adding TMB substrate stop solution. Absorbance was measured at 450 nm. LDL-C, ALT, AST, urea, and creatinine in mouse serum were detected using a fully automated biochemical analyzer (Medicalsystem Biotechnology, MS480). The catalog numbers for the test reagents are as follows: LDL-C (201SJTZ207A, MedicalsystemBiotechnology), ALT (201SJTZ001, Medicalsystem Biotechnology), AST (201SJTZ002, Medicalsystem Biotechnology), urea (201SJTZ106, Medicalsystem Biotechnology), and creatinine (201SJTZ105A, Medicalsystem Biotechnology).
[0337] LNP preparation:
[0338] LNPs were formed by microfluidic mixing of lipids and RNA solution at an amine to RNA phosphate (N:P) ratio of 6. The lipids, including ionizable cationic lipids, cholesterol, DSPC, and DMG-PEG2000, were dissolved in ethanol at a molar ratio of 50:10:38.5:1.5. The RNA cargo (mRNA weight ratio 1:1) was dissolved in 25 mM acetate buffer (pH 5.0). During mixing, the aqueous phase to organic solvent ratio was maintained at approximately 3:1 at a flow rate of 12 ml / min. After mixing, the LNPs were diluted with PBS and dialyzed against a 10 kDa filter at 4°C for 12 hours for buffer exchange. The LNPs were concentrated using an Amicon® Ultra centrifuge filter and stored at 4°C for subsequent use. The size and polydispersity of the LNPs were measured using dynamic light scattering (DLS) with a ZetaSizer Nano ZS (Malvern, UK). The encapsulation efficiency (EE) of the LNPs was determined using the Quanti-it™ RiboGreen RNA Assay Kit according to the manufacturer's instructions.
[0339] Biological distribution of LNPs in vivo:
[0340] LNPs encapsulated with firefly luciferase mRNA (mRNA dose: 0.25 mg / kg) were injected intravenously into 8-week-old Balb / c mice. Six hours post-injection, each mouse was intraperitoneally injected with 200 µL of D-luciferin (15 mg / mL dissolved in PBS), followed by whole-body imaging using an IVIS-spectrum (Perkin Elmer) (exposure time set to 0.5 seconds). Major organs were also collected for subsequent imaging.
[0341] In vivo delivery of epigenetic editor to silence PCSK9 in mice:
[0342] Eight-week-old female C57BL / 6 mice received either vector controls or LNPs encapsulated with epigenetic editors (EE mRNA and corresponding PCSK9-targeting sgRNA or crRNA) via tail vein injection at total RNA doses of 0.75, 1.5, or 3.0 mg / kg. Anticoagulated blood samples were collected from the orbital sinus at specific time points post-injection (days 7, 30, 60, 90, 120, 150, and 180) using EDTA-coated tubes. Serum was separated by centrifugation (2,000 × g, 4°C, 15 min), and PCSK9 protein, LDL-C levels, AST, ALT, creatinine, and urea were determined by ELISA. Mice were euthanized by cervical dislocation following CO2 asphyxia. Major organs (heart, liver, spleen, lung, and kidney) were immediately collected; one leaf was rapidly frozen in liquid nitrogen for subsequent DNA / RNA extraction (stored at -80°C), and the other leaf was fixed with 4% paraformaldehyde for histological examination.
[0343] Histology and staining:
[0344] Freshly harvested tissues were immersed in 4% PFA at 4°C for 24 hours, followed by dehydration with a series of ethanol solutions (70%, 95%, and 100%) to remove residual moisture, and then embedded in paraffin. Paraffin blocks were cut into 5 µm thick sections, dewaxed with xylene, and rehydrated with water. The sections were stained with hematoxylin and eosin (Sigma-Aldrich, no. H3136 and Thermo Fisher Scientific, no. 6766008) and examined for histopathological changes.
[0345] Targeted amplicon bisulfite sequencing:
[0346] Using the Genomic DNA Extraction Kit (QIAGEN), follow the manufacturer's instructions from 1×10 6Genomic DNA was extracted from cells. For each sample, 1 µg of DNA was bisulfite converted and purified using the EZ DNA Methylation-Gold Kit (Zymo Research). The modified DNA was amplified using nested PCR. The first round of nested PCR conditions were as follows: 95°C for 5 min, 65°C for 30 s, repeat step 2 30 times (decreasing by 0.5°C per cycle), 68°C for 5 min, and hold at 12°C. Primers used for bisulfite sequencing are listed in Table S4. Sequencing connectors were added to both ends of the second round of nested PCR primers for subsequent next-generation sequencing. Amplicons were generated using KOD DNA polymerase (TOYOBO), and the amplicons from each condition were mixed so that each sample contained all relevant amplicons and used as input for NEBNext Ultra II DNA library preparation, amplified using a unique dual index to allow multiplex sequencing. The libraries were sequenced using a BGI Genomics DNBSEQ-G99 (2 × 150 bp paired end reads) with a read depth of at least 10,000 read pairs per sample. The average percentage of CpG methylation under all conditions was calculated using R (version 4.3.1) within a 500 bp region centered on the gRNA binding site.
[0347] Targeted amplicon deep sequencing to assess edit rate:
[0348] Genomic DNA was isolated from cultured cells or mouse liver for editing analysis. Genomic loci across each target site were amplified by PCR using site-specific primers carrying tails complementary to the Truseq adapter. A first round of PCR was performed on 200 ng of genomic DNA using Phusionmaster mix (Thermo) and site-specific primers containing i5 and i7 adapter tails. A second round of PCR was performed on the first round PCR products using i5 and i7 primers to complete the adapter and include i5 and i7 indices. All primers used for amplicon sequencing are listed in the supplementary table. PCR products were purified using Ampure magnetic beads (0.9X reaction volume), eluted with 25 µl TE buffer, and quantified using Qubit. Each amplicon was mixed in an equimolar ratio and sequenced using Illumina Miniseq (Illumina Miniseq Control software (3.1)). Amplicon sequencing data were analyzed using CRISPResso2 (https: / / crispresso.pinellolab.partners.org / ). In summary, demultiplexing and base calling were performed using bcl2fastq Conversion Software v2.19 (Illumina, Inc.), allowing for one barcode mismatch, with a minimum trimmed read length of 95 (TrimGalore v. 0.6.2). Sequencing reads were aligned to each amplicon sequence using CRISPResso2. Since many epigenetic edited samples contain multiple base changes or insertions / deletions, C-to-T editing efficiency was analyzed using CRISPResso2 in standard batch mode with the following parameters: '-q 30', '--discard_indel_reads TRUE', and '-qwc (or --quantification_window_coordinates)' and '--expected_hdr_amplicon_seq' to provide the expected edited amplicon sequence for each target site.
[0349] Sequence alignment and cytosine methylation level detection:
[0350] The cleaned reads were mapped back to the reference genome using BSMAP software version 2.90. The parameters used were set to "-n 0 -v 0.08 -g 1 -p 48". The methylation ratio was extracted from the BSMAP output (SAM) using the Python script (methratio.py) distributed with the BSMAP package. In short, the methylation level is calculated based on the percentage of methylated cytosine (mC) in the entire genome, with site methylation level = 100 × (number of sequences with methylated cytosine / total number of valid sequences).
[0351] DMC and DMR testing:
[0352] Differentially methylated cytosine (DMCs) were detected at CpG sites (at least 5-fold coverage) using Radmeth (v1.0). Detected DMCs were filtered according to the following criteria: (1) Q value must be less than 0.05; (2) methylation level difference must be greater than 0.1. Differentially methylated regions (DMRs) were detected at CpG sites (at least 5-fold coverage) using metilene in de-novo mode. The parameters used were set to "--mincpgs 3 --minMethDiff 0 --mode 1 -mtc 2". Detected DMRs were then filtered according to the following criteria: (1) corrected MWU test P value must be less than 0.05; (2) methylation level difference must be greater than 0.1; (3) the number of CpGs contained in the DMR must be greater than 5; (4) the length of the DMR must be greater than 50 bp. Relevant elements and genes of DMCs and DMRs were located in the genome using genomic gff files. The genome was annotated with GO and KEGG based on the EGGNOG database using emapper. Then, GO and KEGG enrichment analyses were performed on the relevant genes using Allenricher (v1.0).
[0353] RNA-seq experiments and data analysis:
[0354] Total RNA was extracted from mouse hepatocytes or mouse hepatocyte cell lines. 1 µg of RNA was used as input material for RNA sample preparation for each sample. Sequencing libraries were generated according to the manufacturer's recommendations using the NEBNext® Ultra™ RNA Library Preparation Kit (Illumina®) (NEB, USA), and index codes were added to each sample for attribution sequences. Briefly, mRNA was purified from total RNA using magnetic beads with oligomers (dT). Fragmentation was performed at high temperature using divalent cations in NEBNext first-strand synthesis reaction buffer (5X). First-strand cDNA was synthesized using random hexamer primers and M-MuLV reverse transcriptase (RNase H). Second-strand cDNA was subsequently synthesized using DNA polymerase I and RNase H. Remaining overhangs were converted to blunt ends by exonuclease / polymerase activity. After adding an adenosine nucleotide to the 3' end of the DNA fragment, a NEBNext adapter with a hairpin loop structure was ligated to prepare for hybridization. To select cDNA fragments of optimal length (250–300 bp), library fragments were purified using the AMPure XP system (Beckman Coulter, Beverly, USA). Size-selected, adapter-ligated cDNA was then treated with 3 µl USER enzyme (NEB, USA) at 37°C for 15 minutes, followed by treatment at 95°C for 5 minutes, before PCR was performed. PCR was conducted using Phusion high-fidelity DNA polymerase, universal PCR primers, and Index(X) primers. PCR products were purified using the AMPure XP system, and library quality was assessed using an Agilent Bioanalyzer 2100 system.
[0355] After constructing the paired-end libraries, sequencing was performed using the Illumina NovaSeq 6000 platform by Beijing Novogene Bioinformatics Technology Co., Ltd., generating 150 bp paired-end reads. Reads were aligned and mapped to the mouse genome (mm10). The RNA-seq data presented in this study have been deposited in the Gene Expression Omnibus (GEO) database, accession number GSE299334, belonging to BioProject PRJNA1273918. Differentially expressed genes (DEGs) were defined as log2|fold-change|> 1 and P < 0.01. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed using the R package “Cluster Profiler”. In short, the analysis followed these detailed steps: the count data were converted to per million transcripts (TPM), and all subsequent analyses were performed using the TPM. First, a sample correlation matrix was generated by calculating the Spearman correlation coefficient between each pair of samples. Compared to Pearson, the Spearman correlation coefficient is less sensitive to outliers. Secondly, principal component analysis was performed using the `fviz_pca_ind` function in R. Each component represents a portion of the variance present in the data, with the top six components collectively capturing over 90%. Genes contributing the most to each component were identified by calculating the proportion of variance explained by each gene for each component and selecting those greater than 0.01%. Thirdly, for gene enrichment analysis, normalized expression of each gene was calculated by dividing its expression level across all samples by its observed maximum TPM. Statistically significant GO terms were defined using FDR < 0.01. KEGG pathway analysis was performed using KofamKOALA, and enrichment analysis was performed using the R package `clusterProfiler`.
[0356] Example 2. Design and Construction of Fusion Proteins
[0357] This invention constructs two core CRISPR OFF-EE systems: a system based on dSpCas9 and a system based on Cas-SF01.
[0358] Cas-SF01 is a Cas12i3 variant (derived from BC26312) that achieves nuclease inactivation by mutating E at position 844 to A or D at position 619 to A (dSF01).
[0359] The fusion protein structure was designed as follows: N-terminus - DNMT3A - DNMT3L - dCas protein - KRAB - C-terminus.
[0360] To improve nuclear localization efficiency, this invention introduces optimized nuclear localization sequences (NLS) into the fusion protein, such as the sequences shown in SEQ ID NO: 1-6.
[0361] Preliminary functional validation of the CRISPR-OFF epigenetic editor in mammalian cells via plasmid delivery.
[0362] To rigorously validate the DNA methylation-based epigenetic editing strategy of this invention, two different CRISPR-OFF epigenetic editor (CRISPR-OFF EE) constructs were initially constructed. These comprised either a catalytically inactivated Streptococcus pyogenes Cas9 (dSpCas9) or the more compact Cas12i3 variant Cas-SF01 (dSF01: an E844A mutant Cas-SF01). Each dCas protein had a DNMT3A-DNMT3L DNA methyltransferase effector complex fused to its N-terminus and a KRAB transcriptional repressor domain fused to its C-terminus, yielding the SpCas9-OFF-EE V1 fusion protein (SEQ ID NO: 138) and the SF01-OFF-EE V1 fusion protein (SEQ ID NO: 139). The ability of these plasmid constructs to silence gene expression was evaluated by transfecting them into HEK293T-derived Gapdh-Snrpn-GFP reporter cells. In this established reporting system, GFP expression is directly regulated by the methylation state of the integrated Snrpn promoter element. Figure 1 a). Thirty days post-transfection, flow cytometry analysis of mCherry-positive cells (indicating successful plasmid uptake) showed that approximately 63% and 60.4% of cells, respectively, exhibited GFP inactivation when treated with SpCas9-OFF-EE or SF01-OFF-EE, plus their respective Snrpn-targeted guide RNAs, sgRNA (SEQ ID NO: 140) or crRNA (SEQ ID NO: 141). In contrast, control cells receiving non-targeted guide RNAs NT-sgRNA (SEQ ID NO: 142) or NT-crRNA (SEQ ID NO: 143) showed only 2.5% of GFP-negative cells at baseline. Figure 1 b, Figure 1 c). Consistently, subsequent bisulfite sequencing confirmed that the observed GFP silencing was directly associated with a significant and highly specific increase in DNA methylation targeting the Snrpn promoter region, while no such changes were observed in the adjacent non-targeting Gapdh promoter, highlighting the targeted nature of the methylation induction (Fig. 1d).
[0363] To further compare the efficacy of these plasmid-based editors against endogenous genes, we targeted CD151, which encodes a cell surface protein that is not essential for cell proliferation or survival. Transfection with SpCas9-OFF-EE or SF01-OFF-EE plasmids, and a mixture of the corresponding sgRNA (SEQ ID NO: 13-15) or crRNA (SEQ ID NO: 22-24), induced significant CD151 silencing, achieving at least a 63% (SpCas9-OFF-EE) and 32% (SF01-OFF-EE) reduction in CD151-positive cells, respectively, at 30 days post-transfection. Figure 1 e, Figure 6 Bisulfite sequencing confirmed that this silencing was directly related to CD151 promoter-induced CpG island (CGI) methylation. Figure 1 f). These preliminary findings suggest that, when delivered in plasmid form, both dSpCas9 and dSF01-based CRISPR-OFFEEs can program for efficient and persistent gene silencing through targeted DNA methylation.
[0364] Example 3. Development and optimization of an mRNA-based epigenetic editor platform
[0365] While plasmid delivery has demonstrated persistent epigenetic memory for at least 30 days, mRNA delivery offers a transient editor expression system with improved safety by mitigating the risks associated with prolonged expression or DNA integration. To explore this approach, we first constructed initial in vitro transcription (IVT) vectors for SpCas9-OFF-EE and SF01-OFF-EE mRNAs (designated OFF-EEV1), adding a T7 promoter upstream of the editor sequence. Figure 2 a). Transfection with unoptimized SpCas9-OFF-EE V1 mRNA (SEQ ID NO: 144) induced gene-specific silencing in Snrpn-GFP reporter cells, with over 46.4% of cells silencing GFP; however, at 30 days post-transfection, its silencing effect on the endogenous CD151 gene was less than 15%. Figure 7 , Figure 8 SF01-OFF-EE V1 mRNA (SEQ ID NO: 145) exhibited lower activity, resulting in only about 5.6% silencing of Snrpn-GFP and about 10.3% silencing of CD151 under similar conditions. Figure 7 , Figure 8Crucially, when SpCas9-OFF-EE V1 mRNA and PCSK9-targeting sgRNA (SEQ ID NO: 69), or SF01-OFF-EE V1 mRNA and PCSK9-targeting crRNA (SEQ ID NO: 47) were encapsulated in lipid nanoparticles (LNPs) and intravenously injected into mice, they failed to induce any significant or durable inhibition of circulating PCSK9 protein or any reduction in LDL-cholesterol (LDL-C) levels. Figure 9 ).
[0366] Recognizing the need for significantly improved mRNA performance for in vivo applications, this invention systematically redesigned the OFF-EE mRNA and protein structure to create a "hyperfolded" V2 construct (Figure 2a), and optimized the mRNA sequence encoding the fusion protein using artificial intelligence algorithms. This includes:
[0367] Codon optimization: Improving translation efficiency based on species-specific preferences.
[0368] Structural optimization: The LinearDesign algorithm was used to balance the minimum free energy (MFE) and codon fitness index (CAI) of the mRNA secondary structure to ensure the stability and high expression of the mRNA.
[0369] Chemical modification: During in vitro transcription, UTP was replaced with N1-methylpseuuridine (m1Ψ) to reduce immunogenicity and enhance translation.
[0370] UTR optimization: Includes optimized 5' and 3' untranslated regions and poly(A) tail structure.
[0371] The optimized SpCas9-OFF-EE V2 mRNA sequence is shown in SEQ ID NO: 136, and the SF01-OFF-EE V2 mRNA sequence is shown in SEQ ID NO: 137.
[0372] The SpCas9-OFF-EE V2 fusion protein (SEQ ID NO: 1) and the SF01-OFF-EE V2 fusion protein (SEQ ID NO: 6) were obtained. Optimization strategies included: enhancing the diversity and composition of nuclear localization signals (NLS) (SEQ ID NO: 151-154); optimizing the 5' and 3' UTRs known to affect mRNA stability and translation; employing a split poly(A) tail strategy; and utilizing AI-assisted codon optimization to optimize the EE coding sequence and interdomain connectors. All these designs aimed to produce highly structured and efficiently translated mRNAs. Based on SpCas9-OFF-EE V2, a split-type epigenetic editing system of SpCas9 (consisting of two proteins) was also developed: variants 573-Split-SpCas9-OFF-EE (amino acid sequences SEQ ID NO: 2-3) and 713-Split-SpCas9-OFF-EE (amino acid sequences SEQ ID NO: 4-5).
[0373] Imaging studies of the EE-GFP fusion protein confirmed that the NLS modification included in the V2 construct resulted in the efficient localization of the translation editor protein to the cell nucleus. Figure 10 Temporal analysis of OFF-EE-GFP expression further indicated that, compared to OFF-EEV1, OFF-EE V2 mRNA exhibited higher translation efficiency and / or improved intracellular stability. Figure 11 Importantly, quality control analysis showed that in vitro transcribed SpCas9-OFF-EE V2 mRNA (SEQ ID NO: 136) or SF01-OFF-EE V2 mRNA (SEQ ID NO: 137) contained fewer dsRNA byproducts and exhibited higher purity. Figure 12 Despite these extensive modifications, when prepared using the same lipid composition and microfluidic mixing parameters, the LNP formulation encapsulating OFF-EE V2 mRNA maintained similar particle size and polydispersity index (PDI) as the formulation encapsulating OFF-EE V1 mRNA. Figure 13 ).
[0374] Example 4. Optimized OFF-EE V2 mRNA epigenetic editor mediating highly efficient and persistent [processes] in cultured cells. Gene silencing
[0375] After determining the improved properties of the OFF-EE V2 mRNA constructs, we next evaluated their ability to achieve efficient and persistent gene silencing at different targets in cultured cells. In Gapdh-Snrpn-GFP reporter cells, electrotransfection of OFF-EE V2 mRNA (even under unsaturated restriction mRNA dosage conditions) Figure 14 This resulted in strong GFP silencing; at 3 weeks post-treatment, 87% (SpCas9-OFF-EE V2) and 32% (SF01-OFF-EE V2) of cells were silenced, respectively. Figure 2 b, Figure 2 c). This silencing effect is directly related to the establishment and maintenance of CpG methylation at the Snrpn site, as determined by bisulfite sequencing ( Figure 15 ).
[0376] To further evaluate the OFF-EE V2 platform, we compared the silencing efficiency of OFF-EE V1 and OFF-EE V2 mRNAs targeting three endogenous cell surface markers (CD29, CD81, and CD151) in HEK293T cells. Each target was represented by a pool of three guide RNAs: sgRNA (SEQ ID NO: 7-9) or crRNA (SEQ ID NO: 16-18) targeting CD29; sgRNA (SEQ ID NO: 10-12) or crRNA (SEQ ID NO: 19-21) targeting CD81; and sgRNA (SEQ ID NO: 13-15) or crRNA (SEQ ID NO: 22-24) targeting CD151. The OFF-EE V2 mRNA construct consistently showed significantly improved silencing effects for each gene. Specifically, 21 days after treatment, SpCas9-OFF-EE V2 achieved at least 80% silencing of all three genes, while SF01-OFF-EE V2 mediated silencing efficiency of 10-30%. Figure 2 b, Figure 2 c). Consistent with these functional results, targeted amplicon bisulfite sequencing revealed that, compared to the V1 counterpart, dSpCas9 and dSF01-based OFF-EE-V2 mRNAs induced significantly higher levels of CpG methylation at their respective target gene promoters. Figure 16 , Figure 17 , Figure 18 These data collectively demonstrate that our engineered and optimized OFF-EE V2 mRNA can achieve significantly more efficient and durable gene silencing at multiple endogenous sites in cultured cells.
[0377] Example 5. In vitro screening and optimization of guide RNA for epigenetic silencing of PCSK9
[0378] Given that therapeutic inactivation of PCSK9 is a clinically validated strategy for treating hypercholesterolemia, this invention focuses on identifying the optimal guide RNA for its epigenetic silencing. To this end, we constructed a mouse hepatocellular carcinoma cell line (Hepa 1-6PCSK9). IRES-EGFP EGFP expression serves as a sensitive reporter factor for endogenous PCSK9 transcriptional activity at the single-cell level. Figure 2 d). Delivery of SpCas9-OFF-EE V2 or SF01-OFF-EE V2 mRNA to these reporter cells resulted in a dose-dependent decrease in EGFP signaling at 7 and 21 days post-nuclear transfection. Figure 19 a, Figure 19 b、 Figure 19 c). This reduction in EGFP signaling was closely associated with decreased endogenous PCSK9 mRNA levels and increased PCSK9 promoter CpG methylation, validating the practicality of this reporter system for quantifying targeted epigenetic editing and silencing. Figure 19 d、 Figure 19 e Figure 19 f).
[0379] Using this validated reporting system, we conducted a comprehensive screening to identify the most effective guide RNA sequences. For SF01-OFF-EE, we designed multiple chemically synthesized crRNAs (AS-cr1 to AS-cr18, sequences SEQ ID NO: 25-42, and S-cr1 to S-cr23, sequences SEQ ID NO: 43-65) to be laid flat on the CpG island (CGI) region of the mouse PCSK9 locus (approximately 700 bp relative to the transcription start site (TSS)). Figure 2 e). For SpCas9-OFF-EE, we designed 10 sgRNAs (sg1 to sg10, sequences SEQ ID NO: 66-75), flanking previously known sgRNAs that induce high levels of PCSK9 repression ( Figure 2 e). EGFP silencing was measured by flow cytometry at 7 and 21 days post-electrotransfection. For SpCas9-OFF-EE, the two nearest sgRNAs to the TSS (sgRNA3, sgRNA4, sgRNA5, sgRNA6) were silencing. Figure 2 sg3 in e and 2f; and sgRNA4, Figure 2 The sg4 in e and 2f most effectively inhibits EGFP expression ( Figure 2f, Figure 20 a). Interestingly, for SF01-OFF-EE, crRNAs targeting two different regions (with no obvious strand bias relative to TSS) effectively silenced EGFP expression. Figure 2 e, Figure 2 g, Figure 21 a). It is worth noting that the optimal targeting regions of SF01-OFF-EE and SpCas9-OFF-EE do not overlap ( Figure 2 g). Observed EGFP silencing was closely associated with targeted DNA methylation at the PCSK9 promoter at 7 and 21 days post-treatment. Figure 2 h, Figure 2 i, Figure 20 c, Figure 21 c).
[0380] Consistent with previous studies demonstrating that multiple guides can enhance efficacy, we evaluated the PCSK9 inhibition effects using dual or triple sgRNA / crRNA combinations. Flow cytometry and targeted bisulfite sequencing both showed that, at 7 and 21 days post-nuclear transfection, these combinations induced more potent and durable silencing than single guides. Figure 3 a, Figure 3 b、 Figure 3 c. Figure 3 d, Figure 20 b、 Figure 20 d, Figure 21 b、 Figure 21 d). Based on the criteria of prioritizing maximum silencing on day 21, no perfect off-target match in the mouse genome, and prioritizing the simplicity of dual-guided combinations, we selected the Na-sgRNA (SEQ ID NO: 146) + sgRNA4 (SEQ ID NO: 69) combination for SpCas9-OFF-EE, and the AS-cr7 (SEQ ID NO: 31) + S-cr14 (SEQ ID NO: 56) combination for SF01-OFF-EE for subsequent in vivo studies.
[0381] To rigorously evaluate the specificity of our primary PCSK9-targeting EE configuration, we performed whole-genome bisulfite sequencing (WGBS) and RNA sequencing (RNA-seq) in Hepa 1–6 cells. Cells were treated with SpCas9-OFF-EE, SF01-OFF-EE, or inteptide-mediated split-spCas9 OFF-EE systems (573-Split-SpCas9-OFF-EE and 713-Split-SpCas9-OFF-EE) and compared with simulated-treatment cells (receiving only OFF-EE mRNA). Figure 3 e, Figure 22 c).
[0382] Transient expression of these PCSK9-targeting OFF-EE mRNAs resulted in strong and targeted CpG methylation in the PCSK9 TSS region by day 30. Figure 3 g). Although the targeted methylation features and abundances of the four different EE mRNA constructs were very similar, the global CpG methylation levels remained largely unchanged, with the average delta methylation between editor-treated and simulated-treated cells being less than 1.2% (g). Figure 3 f). Analysis of differentially methylated regions (DMRs; defined as delta methylation ≥0.2 and P≤0.001) confirmed that CpG methylation changes exhibiting the most significant and profound effect sizes were highly specific to the targeted PCSK9 locus (f). Figure 3 h, Figure 3 i, Figure 22 a, Figure 22 b).
[0383] Consistently, RNA-seq data showed that, compared with the simulated control, the PCSK9 mRNA level in OFF-EE treated cells was significantly reduced (approximately 20-30 times). Figure 3 j, Figure 22 c). Intersection analysis of DMRs (FDR ≤ 0.001) and differentially expressed genes (DEGs; |log2 FC| ≥ 1; false discovery rate (FDR) ≤ 0.001) identified 60 and 15 downregulated genes associated with DMRs, respectively, corresponding to SpCas9-OFF-EE and SF01-OFF-EE. Figure 3 k, Figure 3 l). The split-type SpCas9 epigenetic editor also exhibited some off-target transcriptional changes associated with DMR (47 and 74 genes, respectively). Figure 22 d, Figure 22e). Notably, in addition to PCSK9, three other genes (Gfra1, 9030622O22Rik, and Dnajc12) were downregulated and associated with DMR in all four EE-OFF constructs tested. These off-target effects may be attributed to guide-dependent binding of the dCas-editor to unexpected genomic sites with partial sequence homology, or potentially guide-independent interactions of the editor components. Overall, these multi-omics analyses suggest that off-target methylation and gene expression changes can still occur despite high targeting (especially to PCSK9). Importantly, in this cellular setting, SF01-based OFF-EE produced fewer off-target effects than SpCas9-based systems, suggesting that SF01 may have higher specificity.
[0384] Example 6. Highly efficient and specific in vivo silencing of PCSK9 after LNP-mediated epigenome editor delivery.
[0385] Cationic lipids, cholesterol, DSPC, and DMG-PEG2000 were dissolved in ethanol at a molar ratio of 50:10:38.5:1.5. The mRNA and gRNA or crRNA containing the fusion protein (mRNA and gRNA or crRNA weight ratio of 1:1) were dissolved in acetate buffer. During mixing, the aqueous phase to organic solvent ratio was maintained at approximately 3:1, and the flow rate was 12 ml / min. After mixing, the LNPs were diluted with PBS and dialyzed against a 10 kDa filter at 4°C for 12 hours for buffer exchange. The LNPs were concentrated using an Amicon® Ultra centrifuge filter and stored at 4°C for subsequent use. The resulting LNPs exhibited consistent and favorable physicochemical properties, including a particle size of approximately 75 nm, PDI ≈ 0.15, and mRNA encapsulation efficiency exceeding 95%. Figure 23 ).
[0386] The PCSK9 silencing effect in mice was evaluated using LNP delivery of SpCas9-OFF-EE V2 mRNA (SEQ ID NO: 136) + Na-sgRNA (SEQ ID NO: 146) + sgRNA4 (SEQ ID NO: 69) or SF01-OFF-EE V2 mRNA (SEQ ID NO: 137) + AS-crRNA7 (SEQ ID NO: 31) + S-crRNA14 (SEQ ID NO: 56) at three escalating doses (0.75, 1.5, and 3.0 mg / kg total RNA).
[0387] For the splitting epigenetic editing system, LNP delivery of 573-Split-SpCas9-OFF-EE mRNA (SEQ ID NO: 147-148) + Na-sgRNA (SEQ ID NO: 146) + sgRNA4 (SEQ ID NO: 69) or 713-Split-SpCas9-OFF-EE mRNA (SEQ ID NO: 149-150) + Na-sgRNA (SEQ ID NO: 146) + sgRNA4 (SEQ ID NO: 69) was used to evaluate the PCSK9 silencing effect in mice at a high dose (3.0 mg / kg total RNA). Mice treated with LNPs encapsulating Fluc mRNA served as vector controls. One week after injection, we observed a significant dose-dependent decrease in serum PCSK9 protein and LDL-C levels in both editor systems. Figure 4 a, Figure 4 b). Furthermore, at the highest dose (3.0 mg / kg), SpCas9-OFF-EE V2, SF01-OFF-EE V2, 573-Split-OFF EE, and 713-Split-OFF-EE editors (the latter two also formulated with LNP and with their respective wizards) reduced PCSK9 by approximately 92.5%, 78.2%, 54.3%, and 63.2%, respectively, with corresponding reductions in LDL-C of approximately 61.7%, 49.8%, 36.3%, and 41.2%. Figure 4 c, Figure 4 d). RNA sequencing (RNA-seq) of these mouse liver tissues confirmed deep downregulation of PCSK9 mRNA, and gene ontology (GO) bioprocess analysis highlighted significant enrichment of altered cholesterol and sterol metabolic pathways, consistent with effective PCSK9 silencing. Figure 4 e, Figure 4 f). Global CpG methylation levels remained essentially unchanged. Figure 4 g).
[0388] A key benchmark for treatment relevance is the long-term persistence of epigenetic silencing. Therefore, we monitored circulating PCSK9 and LDL-C levels for up to 180 days following a single intravenous injection of LNPs from different EE constructs in C57BL / 6 mice. Figure 5a). Notably, after reaching a trough within a week, cyclic PCSK9 levels maintained stable inhibition of approximately 90% (SpCas9-OFF-EE-V2), 80% (SF01-OFF-EE-V2), 50% (573-Split-OFF-EE), and 65% (713-Split-OFF-EE) over the entire 180-day observation period. Figure 5 b). Correspondingly, plasma LDL-C levels also reached a sustained low, stabilizing at levels approximately 55% (SpCas9), 30% (SF01), 20% (573-Split), and 30% (713-Split) lower than baseline during the study period. Figure 5 c).
[0389] To further establish the mechanistic link between the molecular effects of OFF-EE and the potent inhibition of PCSK9 in vivo, this study used the WGBS method to measure CpG methylation at PCSK9 loci in liver samples obtained at 1 and 4 months post-treatment. Transient expression of OFF-EE resulted in strong methylation of multiple CpG sites in the TSS region at 30 days post-injection, while the methylation rate decreased at 120 days post-injection. Figure 5 d and Figure 5 e). Furthermore, through WGBS ( Figure 5 f, Figure 5 g, Figure 24 Comprehensive in vivo specificity was assessed using RNA-seq of liver tissue and DMRs. The association between DMRs and DEGs further underscores PCSK9 as a major target. Figure 5 f, Figure 5 g, Figure 24 This indicates a strong correlation between in vivo epigenetic modification and efficient gene repression. Furthermore, intersection analysis showed that SF01-OFF-EE generally exhibits lower off-target effects in vivo compared to SpCas9-OFF-EE. Figure 5 h, Figure 5 i, Figure 24 ).
[0390] To assess the specificity of in vivo epigenetic modifications, WGBS analysis of the PCSK9 promoter was performed on genomic DNA extracted from various tissues. The results showed that, compared with the vector control group, EE-treated mice exhibited a liver-specific differential CpG hypermethylation pattern (…). Figure 25 Crucially, the CpG methylation level of the PCSK9 promoter remained unchanged in non-hepatic tissues (including the heart, spleen, lungs, and kidneys), demonstrating the liver-oriented activity and specificity of the LNP delivery system. Figure 25 Furthermore, comprehensive in vivo specificity was assessed using WGBS and RNA-seq of liver tissue. The safety of LNP-EE treatment was evaluated through histopathological analysis and serum biochemistry. No treatment-related tissue abnormalities or signs of inflammation were observed on day 7 in hematoxylin and eosin (H&E) stained sections of major organs (liver, heart, spleen, lung, and kidney). Figure 26 Consistent with these histological findings, serum biochemical markers of liver function (alanine aminotransferase ALT; aspartate aminotransferase AST) and kidney function (creatinine CR; urea) showed no significant differences between mice treated with various epigenetic editors and the vector control group at multiple time points after injection, further validating the good biosafety of the LNP-delivered EE mRNA platform. Figure 27 ).
[0391] Using LNP-encapsulated mRNA encoding Cas-SF01-OFF-EE and crRNA targeting mouse PCSK9 (sequences selected from SEQ ID NO: 31 and 56), a single intravenous injection into C57BL / 6 mice achieved the following results:
[0392] Protein levels: On day 7 post-injection, circulating PCSK9 protein decreased by approximately 80-90%, and LDL-C decreased by approximately 30-55%.
[0393] Duration: This inhibitory effect was observed for at least 180 days after a single dose.
[0394] Mechanism validation: Bisulfite sequencing (WGBS) confirmed that high levels of methylation occurred at the CpG site in the PCSK9 gene promoter region, and this methylation was not observed in non-liver tissues, demonstrating the liver-specific delivery and precision of epigenetic editing of LNPs.
[0395] Specificity: Compared to the SpCas9 system, the Cas-SF01 system exhibits lower off-target effects.
[0396] Example 7. Screening of crRNAs targeting human PCSK9 using the SF01-OFF-EE V2 epigenome editor
[0397] Sixty crRNAs (crRNA1 to crRNA60, with nucleic acid sequences shown in SEQ ID NO: 76-135) were designed to target the 500bp upstream and 500bp downstream regions of the human PCSK9 TSS. The effect of the SF01-OFF-EE V2 epigenome editor on silencing human PCSK9 expression was tested in the human Hep3B cell line. The mRNA encoding SF01-OFF-EE (SEQ ID NO: 137) (3 μg) and crRNA (3 μg) were co-electroplated at 1.0*10-1 cm-1.6 In Hep3B cells, passage was performed when cell confluence reached 90%. PCSK9 mRNA was extracted and quantified by qPCR at 7 and 21 days post-transfection.
[0398] The results showed that on day 21, compared to the control group (only SF01-OFF-EE mRNA was electroporated), the seven crRNAs (crRNA6, crRNA7, crRNA9, crRNA13, crRNA16, crRNA25, and crRNA59) maintained a PCSK9 inhibition efficiency of over 95%. Figure 28 ).
[0399] In this application, we successfully developed and validated an advanced in vivo epigenetic editing platform based on mRNA-LNP. This platform, utilizing the compact Cas-SF01 protein and optimized mRNA structure, overcomes traditional delivery challenges, achieving efficient, durable, and safe silencing of target genes such as PCSK9, demonstrating significant clinical translational potential. Our key innovation lies in the comprehensive optimization of the epigenetic editor architecture (utilizing SpCas9, SF01, and splitting SpCas9 variants) and the mRNA vector itself, ultimately resulting in a significantly enhanced optimized V2 construct. A single intravenous injection of the PCSK9-targeting LNP-formulated SpCas9-OFF-EE-V2 mRNA induced a profound (approximately 90%) and durable reduction in circulating PCSK9 protein in mice, accompanied by a decrease in LDL-C, with effects lasting at least 180 days. This work highlights the therapeutic potential of transiently expressed epigenetic editors in programming persistent phenotypic changes.
[0400] This invention's mRNA-based OFF-EE platform directly addresses key limitations of previous CRISPR-based epigenetic editing strategies. Many early methods relied on plasmid DNA or viral vector delivery, which, while capable of inducing epigenetic changes, carried risks associated with prolonged editor expression (increasing off-target potential) or genome integration (raising safety concerns regarding insertional mutations and oncology). While unmodified mRNA offers a safer, non-integrating alternative, its inherent instability and suboptimal translation have historically limited its application in complex, large editor proteins. This invention's optimized V2 mRNA engineering (including codon optimization, UTR optimization, enhanced NLS design, and poly(A) tail modification) significantly overcomes these problems, improving mRNA stability, translational output, and editor protein function, ultimately enabling powerful in vivo efficacy of transient expression systems. The ability of our V2 OFF-EE mRNA to achieve efficient and persistent silencing in vitro underscores the effectiveness of this optimized mRNA approach. A single dose of LNP-mRNA achieves a PCSK9 silencing duration of >180 days in vivo, which not only exceeds the typical effect duration of other transient approaches (such as siRNA or antisense oligonucleotides targeting PCSK9), but also rivals the persistence sought by more permanent interventions while avoiding their associated risks.
[0401] Although specific embodiments of the invention have been described in detail, those skilled in the art will understand that various modifications and variations can be made to the details based on all the published teachings, and all such changes are within the scope of protection of the invention. The entire scope of the invention is given by the appended claims and any equivalents thereof.
Claims
1. A fusion protein, characterized in that, The fusion protein comprises, from N-terminus to C-terminus, a DNA methyltransferase domain, a Cas protein domain, and a transcriptional repressor domain. The DNA methyltransferase domain includes DNMT3A and DNMT3L; The Cas protein domain is selected from Cas9 protein with nuclease inactivation or Cas12 protein with nuclease inactivation; The transcriptional repression domain is KRAB.
2. The fusion protein according to claim 1, characterized in that, The fusion protein also contains a nuclear localization sequence (NLS). Optionally, the fusion protein comprises an amino acid sequence as shown in any one of SEQ ID NO:1-6, or an amino acid sequence having at least 95% homology with the amino acid sequences shown in SEQ ID NO:1-6.
3. The fusion protein according to claim 1, characterized in that, The Cas9 protein with inactivated nuclease activity is dCas9; the Cas12 protein with inactivated nuclease activity is Cas12i protein.
4. The fusion protein according to claim 3, characterized in that, The Cas12i protein is the Cas-SF01 protein, and preferably, the Cas-SF01 fusion protein contains the amino acid sequence shown in SEQ ID NO:
6.
5. An isolated nucleic acid molecule, characterized in that, The nucleic acid molecule encodes the fusion protein according to any one of claims 1-4.
6. The nucleic acid molecule according to claim 5, characterized in that, The nucleic acid molecule is mRNA; Preferably, the nucleic acid molecule comprises a nucleotide sequence as shown in SEQ ID NO: 136 or SEQ ID NO:
137.
7. A carrier, characterized in that, The carrier comprises the nucleic acid molecule as described in claim 5 or 6.
8. An epigenetic editing system, characterized in that, include: (i) the fusion protein of any one of claims 1-4, or the nucleic acid molecule of any one of claims 5-6, or the vector of claim 7; and (ii) One or more guide RNAs (gRNAs) or their encoded nucleic acids; said gRNAs are capable of targeting the promoter region or the region upstream of the transcription start site of a target gene.
9. The epigenetic editing system according to claim 8, characterized in that, The target gene is the proprotein convertase subtilisin / Kexin type 9 (PCSK9) gene; Preferably, the gRNA contains a targeting sequence complementary to the target sequence, the targeting sequence being selected from any one of SEQ ID NO: 76-135; More preferably, the target sequence of the gRNA is selected from one or more of SEQ ID NO: 81, 82, 84, 88, 91, 100, 134.
10. A pharmaceutical composition, characterized in that, Include: (a) The epigenetic editing system as described in claim 8 or 9; and (b) A pharmaceutically acceptable carrier, preferably a lipid nanoparticle (LNP).
11. The pharmaceutical composition according to claim 10, characterized in that, The epigenetic editing system includes: mRNA encoding the fusion protein of claim 4; and A gRNA targeting PCSK9, wherein the targeting sequence of the gRNA is selected from one or more of SEQ ID NO: 81, 82, 84, 88, 91, 100, 134.
12. The pharmaceutical composition according to claim 10, characterized in that, The lipid nanoparticles (LNPs) are composed of cationic lipids, auxiliary phospholipids, cholesterol, and polyethylene glycol lipids (PEG-lipids). Preferably, the molar percentage ranges of the cationic lipid, auxiliary phospholipid, cholesterol, and polyethylene glycol lipid are as follows: Cationic lipids: 40-60%; Supportive phospholipids: 5-15%; Cholesterol: 30-45%; Polyethylene glycol lipids: 1-3%.
13. Use of the fusion protein of any one of claims 1-4, the nucleic acid molecule of any one of claims 5-6, or the epigenetic editing system of any one of claims 8-9 in the preparation of a medicament for regulating gene expression; Preferably, the drug is used to inhibit the expression of the PCSK9 gene in a subject to treat hypercholesterolemia or cardiovascular disease.