Novel nucleic acid editing protein
A fusion protein combining a nuclease-inactive RNA-guided endonuclease domain with a cytidine deaminase domain addresses the inefficiencies and off-target issues of current genome editing techniques, enabling precise and efficient single-base gene correction in genomic DNA.
Patent Information
- Application Number
- JP2024569521
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-26
- Filing Date
- 2023-05-24
- Publication Date
- 2025-06-05
AI Technical Summary
Current genome editing techniques, such as CRISPR-Cas9, suffer from low efficiency and high off-target activity, making it difficult to achieve precise and efficient single-base gene correction in genomic DNA.
A fusion protein comprising a nuclease-inactive RNA-guided endonuclease domain and a cytidine deaminase domain is used to target specific nucleotide sequences in genomic DNA, enabling efficient and precise base editing without inducing double-strand breaks.
The fusion protein achieves high efficiency in introducing specific base modifications at precise locations in genomic DNA, reducing off-target effects and improving the accuracy of gene editing compared to existing methods.
Smart Images

Figure 2025517515000019 
Figure 2025517515000001 
Figure 2025517515000002
Abstract
Description
[Background technology]
[0001] The present invention relates to a novel nucleic acid editing protein and a base editing system comprising the same.
[0002] Targeted introduction of specific modifications into genomic DNA is a promising approach for studying gene function and has the potential to provide new therapeutic approaches for human genetic diseases. An ideal nucleic acid editing technology would provide high efficiency of introducing the desired modifications, have minimal off-target activity, and have the ability to be guided to precisely edit any site within the genome.
[0003] There are multiple genome engineering tools available, including engineered zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and RNA-guided DNA endonucleases (RGNs). Programmable cleavage can result in mutation of DNA at the cleavage site by non-homologous end joining (NHEJ) or replacement of DNA surrounding the cleavage site by the process of homology-directed repair (HDR).
[0004] One drawback of current techniques is that both NHEJ and HDR typically result in modest gene editing efficiencies, as well as undesired genetic changes that may compete with the desired changes. Because many genetic diseases can in principle be treated by making specific nucleotide changes at specific locations in the genome (e.g., a C to T change in a specific codon of a disease-associated gene), the development of programmable methods to achieve such precise gene editing would not only be a powerful new research tool, but also a potential new approach to human therapeutics based on gene editing.
[0005] The clustered regularly interspaced short palindromic repeats (CRISPR) system is a natural prokaryotic adaptive immune system that has been modified to enable robust and general genome engineering in a variety of organisms and cell lines. The CRISPR-Cas (CRISPR-associated) system is a protein-RNA complex that uses an RNA molecule (gRNA) as a guide to localize the complex to the target DNA sequence via base pairing. In the natural system, the Cas protein acts as an endonuclease to cleave the targeted DNA sequence. The target DNA sequence must be complementary to the sgRNA and also contain a "protospacer adjacent motif" (PAM) sequence at the end of the complementary region for the system to function. Among the known Cas proteins, S. pyogenes Cas9 is the main and widely used tool for genome engineering. This Cas9 protein is a large multi-domain protein that contains two distinct nuclease domains. Point mutations can be introduced into the Cas protein to disable nuclease activity, resulting in a dead Cas (dCas) or nickase (nCas) that still retains the ability to bind DNA guided by a gRNA. When fused to another protein or domain, dCas9 can be targeted to a DNA sequence of interest simply by co-expression with the appropriate gRNA.
[0006] Importantly, 80-90% of human disease-causing protein mutations result from substitutions, deletions or insertions of only a single nucleotide. Current strategies for single-base gene correction include engineered nucleases, which rely on the creation of double-strand breaks (DSBs) followed by stochastic and inefficient homology-directed repair (HDR), and DNA-RNA chimeric oligonucleotides. The latter strategy involves the design of an RNA / DNA sequence that base pairs with a specific sequence in genomic DNA at the position of the nucleotide to be edited. The resulting mismatch is recognized and corrected by the cell's endogenous repair system, resulting in a sequence change. Both strategies suffer from low gene editing efficiency and undesired genetic alterations, as they are subject to both HDR stochasticity and competition between HDR and non-homologous end joining (NHEJ). HDR efficiency varies depending on the location of the target gene in the genome, cell cycle status, and cell / tissue type.
[0007] U.S. Patent No. 10,167,457 discloses several example base editors and provides fusion peptides for targeted base editing.
[0008] The development of a straightforward, programmable method to insert specific types of base modifications at precise locations in genomic DNA with enzyme-like efficiency and without chance represents a powerful new approach to gene-editing-based research tools and human therapeutics. Summary of the Invention
[0009] Some aspects of the present disclosure provide methods, systems, reagents and kits useful for targeted editing of nucleic acids. In some embodiments, a fusion protein of an RNA-guided endonuclease domain and a cytidine deaminase domain is provided. In some embodiments, a method for targeted nucleic acid editing is provided. In some embodiments, a reagent and kit for generating a targeted nucleic acid editing protein, such as a fusion protein of a Cas domain and a deaminase domain, is provided.
[0010] Some aspects of the present disclosure provide a fusion protein comprising (i) a nuclease-inactive RNA-guided endonuclease domain, and (ii) a cytidine deaminase domain. In some embodiments, the nucleic acid editing domain is fused to the N-terminus of the RNA-guided endonuclease domain. In some embodiments, the nucleic acid editing domain is fused to the C-terminus of the RNA-guided endonuclease domain. In some embodiments, the RNA-guided endonuclease domain and the nucleic acid editing domain are fused via a linker.
[0011] Some aspects of the present disclosure provide a method for DNA editing. In some embodiments, the method comprises contacting a DNA molecule with (a) a fusion protein or protein complex comprising a nuclease-inactive RNA-guided endonuclease domain and a cytidine deaminase domain, and (b) a guide RNA that targets the fusion protein to a target nucleotide sequence, wherein the DNA molecule is contacted with the fusion protein or protein complex and the guide RNA in an amount and under conditions suitable for deamination of nucleotide bases.
[0012] In some embodiments, the target DNA sequence comprises a sequence associated with a disease or disorder, and deamination of the nucleotide base results in a sequence that is not associated with the disease or disorder. In some embodiments, the DNA sequence comprises a T>C point mutation. In some embodiments, deamination corrects a point mutation in the sequence associated with the disease or disorder. In some embodiments, the sequence associated with the disease or disorder encodes a protein, and deamination introduces a stop codon or disrupts splicing of the sequence associated with the disease or disorder.
[0013] Some aspects of the disclosure provide kits comprising a nucleic acid construct comprising a sequence encoding a nuclease-inactive RNA-guided endonuclease sequence, a sequence encoding a nucleic acid editing enzyme or enzyme domain, such as a cytidine deaminase, in frame with the sequence encoding the RNA-guided endonuclease, and optionally a sequence encoding a linker located between the Cas coding sequence and the cloning site. Additionally, in some embodiments, the kit comprises suitable reagents, buffers, and / or instructions for use.
[0014] Some aspects of the present disclosure provide kits that include a fusion protein that includes a nuclease-inactive RNA-guided endonuclease domain and a cytidine deaminase domain, and optionally a linker disposed between the RNA-guided endonuclease domain and the deaminase domain. In addition, in some embodiments, the kit includes suitable reagents, buffers, and / or instructions for using the fusion protein, for example, for in vitro or in vivo DNA or RNA editing. In some embodiments, the kit includes instructions for designing and using a suitable gRNA for targeted editing of nucleic acid sequences.
[0015] The invention will now be described with reference to the following drawings. [Brief description of the drawings]
[0016] [Figure 1] A diagram showing an example of the arrangement of fusion protein domains: NLS - nuclear localization signal sequence, NH2 - N-terminus of peptide sequence, COOH - C-terminus of peptide. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0017] Abbreviation [Table 1]
[0018] [Table 2]
[0019] definition The following definitions are used throughout this specification.
[0020] The term "site-specific nuclease" as used herein refers to an enzyme that can specifically recognize and cleave a DNA sequence. Site-specific nucleases may be engineered. Examples of engineered site-specific nucleases include zinc finger nucleases (ZFNs), TAL effector nucleases (TALENs), and CRISPR / Cas9-based systems.
[0021] The term "transcription activator-like effector" or "TALE" as used herein refers to a protein that recognizes and binds to a specific DNA sequence. A "TALE DNA binding domain" refers to a DNA binding domain that contains an array of tandem 33-35 amino acid repeats, each of which specifically recognizes a single base pair of DNA. Such repeats can be arranged in any order to assemble an array that recognizes a specific sequence.
[0022] As used herein, the term "transcription activator-like effector nuclease" or "TALEN" refers to a fusion protein between the catalytic domain of a nuclease and an engineered TALE DNA binding domain that can target custom DNA sequences.
[0023] The term "zinc finger" as used herein refers to a protein that contains a zinc finger domain and recognizes and binds to a DNA sequence. A single zinc finger contains about 30 amino acids, and the domain typically functions by binding three consecutive base pairs of DNA through interactions of a single amino acid side chain per base pair.
[0024] As used herein, the term "zinc finger nuclease" or "ZFN" refers to a chimeric protein molecule that contains at least one zinc finger DNA binding domain operatively linked to at least one nuclease or portion of a nuclease that, when fully transcribed and assembled, is capable of cleaving DNA.
[0025] The term "CRISPR" (clustered regularly interspaced short palindromic repeats) refers to a family of DNA sequences found in the genomes of prokaryotes such as bacteria and archaea. These sequences are derived from DNA fragments of bacteriophages that previously infected the prokaryote. They are used to detect and destroy DNA from similar bacteriophages during subsequent infections.
[0026] The term "CRISPR system" collectively refers to transcripts and other elements involved in expression or directing activity of CRISPR-associated ("Cas") proteins, including sequences encoding Cas proteins, tracr (trans-activating CRISPR) sequences (e.g., tracrRNA or an active partial tracrRNA), tracr mate sequences (including "direct repeats" and, in the context of endogenous CRISPR systems, partial direct repeats processed by tracrRNA), guide sequences (also referred to herein as "spacers" in the context of endogenous CRISPR systems), or other sequences and transcripts from the CRISPR locus.
[0027] The term "type II CRISPR system" refers to an effector system that uses a single effector enzyme, Cas9, to cleave dsDNA and perform targeted DNA double-strand breaks in four sequential steps. Compared to type I and type III effector systems, which require multiple different effectors acting as a complex, type II effector systems can function in alternative contexts, such as eukaryotic cells. Type II effector systems consist of a long pre-crRNA transcribed from a spacer-containing CRISPR locus, Cas9 protein, and a tracrRNA involved in pre-crRNA processing.
[0028] The term "nucleic acid-guided DNA binding protein" refers to any protein that forms a complex with one or more nucleic acids that guides the binding of that protein to a specific region of DNA. An RNA-guided nuclease is an example of a nucleic acid-guided DNA binding protein.
[0029] The terms "RNA-guided endonuclease" or "RGN" are used interchangeably herein and refer to a nuclease that forms a complex with (e.g., binds to or associates with) one or more RNAs that are not targets for cleavage.
[0030] The term "gRNA" is also used interchangeably herein as chimeric single guide RNA ("sgRNA"), to refer to a nucleic acid that is a fusion of two non-coding RNAs: crRNA and tracrRNA. "gRNA" is used interchangeably to refer to a guide RNA that exists as a single molecule or as a complex of two or more molecules. Typically, a gRNA that exists as a single RNA species contains two domains: (1) a domain that shares homology with a target nucleic acid (e.g., directs binding of the Cas9 complex to the target), and (2) a domain that binds to the Cas9 protein.
[0031] The term "Cas9" refers to a type of RGN that cleaves nucleic acids, is encoded by the CRISPR locus, and is part of the type II CRISPR system. The commonly used Cas9 protein is derived from the bacterial species Streptococcus pyogenes. The Cas9 protein can be mutated so that its nuclease activity is partially or completely inactivated.
[0032] The term "dCas9" refers to an inactivated Cas9 protein. An example is dCas9 from Streptococcus pyogenes, which has no nuclease activity. As used herein, "dCas9" refers to a Cas9 protein that has amino acid substitutions that have inactivated its nuclease activity. In the case of S. pyogenes Cas9, these mutations are D10A and H840A.
[0033] The term "nCas9" refers to the Cas9 nickase domain or protein. The term "Cas9 nickase" refers to a modified version of Cas9 that contains a single inactive catalytic domain, either RuvC- or HNH-. With only one active nuclease domain, the Cas9 nickase cuts only one strand of the target DNA, creating a single-strand break or "nick." The Cas9 nickase can still bind to DNA based on gRNA specificity, but the nickase cuts only one of the DNA strands. As an example, nCas9 is derived from Streptococcus pyogenes, and the RuvC domain can be inactivated by amino acid substitution at position D10 (e.g., D10A), and the HNH domain can be inactivated by amino acid substitution at position H840 (e.g., H840A), or at a position corresponding to an amino acid in another protein.
[0034] As used herein, the term "nicking" refers to the reaction of cleaving the phosphodiester bond between two nucleotides in one strand of a double-stranded DNA molecule to generate a 3' hydroxyl group and a 5' phosphate group.
[0035] As used herein, the term "nucleic acid editing domain" refers to a protein or enzyme that can perform one or more modifications (e.g., deamination of a cytidine residue) on a nucleic acid (e.g., DNA or RNA). Exemplary nucleic acid editing domains include, but are not limited to, deaminase, nuclease, nickase, recombinase, methyltransferase, methylase, acetylase, acetyltransferase, transcriptional activator, or transcriptional repressor domains. In some embodiments, the nucleic acid editing domain is a deaminase (e.g., a cytidine deaminase, such as APOBEC or AID deaminase).
[0036] The term "deaminase" refers to an enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is a cytidine deaminase, which catalyzes the hydrolytic deamination of cytidine or deoxycytidine to uracil or deoxyuracil, respectively.
[0037] The term "linker" as used herein refers to a chemical group or molecule that connects two molecules or moieties, such as the binding domain and cleavage domain of a nuclease. In some embodiments, the linker connects the gRNA binding domain of an RNA programmable nuclease and the catalytic domain of a recombinase. In some embodiments, the linker connects a dCas9 and a recombinase. Typically, the linker is located between or adjacent to two groups, molecules, or other moieties and is connected to each via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5-100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.
[0038] As used herein, the terms "nucleic acid", "nucleic acid sequence", "nucleotide sequence", "oligonucleotide" and "polynucleotide" are interchangeable and refer to the polymeric form of nucleotides. The nucleotides may be deoxyribonucleotides (DNA), ribonucleotides (RNA), analogs thereof, or combinations thereof, and may be of any length. Polynucleotides may perform any function and have any secondary and tertiary structure. The term encompasses known analogs of natural nucleotides, as well as nucleotides modified at the base, sugar and / or phosphate moieties. An analog of a particular nucleotide has the same base-pairing specificity (e.g., an analog of A base pairs with T). A polynucleotide may contain one modified nucleotide or multiple modified nucleotides. Examples of modified nucleotides include fluorinated nucleotides, methylated nucleotides and nucleotide analogs. The nucleotide structure may be modified before or after the polymer is assembled. After polymerization, the polynucleotide may be further modified, for example, by conjugation with a labeling moiety or a target binding moiety. The nucleotide sequence may incorporate non-nucleotide components. The term also encompasses nucleic acids containing modified backbone residues or linkages, which are synthetic, naturally occurring, and non-naturally occurring, and which have similar binding properties to the reference polynucleotide (e.g., DNA or RNA). Examples of such analogs include, but are not limited to, phosphorothioates, phosphoramidates, methyl phosphonates, chiral methyl phosphonates, 2-O-methyl ribonucleotides, peptide nucleic acids (PNAs), locked nucleic acids (LNA™) (Exiqon, Inc., Woburn, Mass.) nucleosides, glycol nucleic acids, bridged nucleic acids, and morpholino structures. Polynucleotide sequences are shown herein in the conventional 5' to 3' orientation unless otherwise indicated.
[0039] As used herein, the terms "peptide", "polypeptide" and "protein" are interchangeable and refer to a polymer of amino acids. A polypeptide may be of any length. It may be branched or linear, may be interrupted by non-amino acids, and may contain modified amino acids. The terms may be used to refer to amino acid polymers modified, for example, by acetylation, disulfide bond formation, glycosylation, lipidation, phosphorylation, cross-linking and / or conjugation (e.g., with a labeling moiety or ligand). Polypeptide sequences are presented herein in the conventional N-terminal to C-terminal orientation. Polypeptides and polynucleotides can be made using routine techniques in the field of molecular biology (see, for example, standard texts listed above). Additionally, essentially any polypeptide or polynucleotide can be custom-ordered from commercial sources.
[0040] The terms "target region," "target sequence," or "protospacer," used interchangeably herein, refer to the region of a target gene that is targeted by a CRISPR-based system.
[0041] The term "target site" refers to a sequence within a nucleic acid molecule that is deaminated by a deaminase or a fusion protein that contains a deaminase (e.g., an RGN-cytidine deaminase fusion protein provided herein).
[0042] The term "mutation" as used herein refers to the substitution of a residue in a sequence, e.g., a nucleic acid or amino acid sequence, with another residue, or the deletion or insertion of one or more residues in a sequence. Mutations are typically described herein by identifying the original residue, followed by identifying the position of the residue in the sequence and the identity of the newly substituted residue. Various methods for making the amino acid substitutions (mutations) provided herein are well known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)).
[0043] The term "complement" or "complementary" as used herein means that the nucleic acid can refer to Watson-Crick or Hoogsteen base pairing between nucleotides or nucleotide analogs of a nucleic acid molecule. The term "complementarity" refers to the property shared between two nucleic acid sequences such that when the two sequences are aligned antiparallel to each other, the nucleotide bases at every position are complementary.
[0044] The term "promoter" as used herein refers to a synthetic or naturally occurring molecule that can confer, activate or enhance expression of a nucleic acid in a cell. A promoter may contain one or more specific transcriptional regulatory sequences to further enhance expression and / or alter its spatial and / or temporal expression. A promoter may also contain distal enhancer or repressor elements that may be located as far as several thousand base pairs away from the start site of transcription. Promoters may be derived from sources including viruses, bacteria, fungi, plants, insects and animals.
[0045] The term "enhancer" as used herein refers to a non-coding DNA sequence that contains multiple activator and repressor binding sites. Enhancers range in length from 200 bp to 1 kb and may be either proximal, 5' upstream of the promoter or in the first intron of the regulated gene, or distal, in the intron of an adjacent gene or in an intergenic region far from the locus. Through DNA looping, active enhancers contact the promoter depending on the promoter specificity of the core DNA binding motif. As many as 4-5 enhancers may interact with a promoter.
[0046] The term "operably linked" as used herein means that the expression of a gene is under the control of a promoter with which it is spatially connected. The promoter may be located 5' (upstream) or 3' (downstream) of the gene under its control. The distance between the promoter and the gene may be approximately the same as the distance between the promoter and the gene it controls in the gene from which it is derived. As is known in the art, variations in this distance may be accommodated without loss of promoter function.
[0047] The term "vector" as used herein refers to a nucleic acid sequence that includes an origin of replication. A vector can be a viral vector, a bacteriophage, a bacterial artificial chromosome, or a yeast artificial chromosome. A vector can be a DNA or RNA vector. A vector can be a self-replicating extrachromosomal vector or a DNA plasmid.
[0048] The term "effective amount" as used herein refers to an amount of a biologically active agent sufficient to induce a desired biological response. For example, in some embodiments, an effective amount of a nuclease may refer to an amount of the nuclease sufficient to induce cleavage of a target site specifically bound and cleaved by the nuclease. In some embodiments, an effective amount of a recombinase may refer to an amount of the recombinase sufficient to induce recombination at a target site specifically bound and recombined by the recombinase. As will be appreciated by those skilled in the art, the effective amount of an agent, e.g., a nuclease, a recombinase, a hybrid protein, a fusion protein, a protein dimer, a complex of a protein (or protein dimer) and a polynucleotide, or a polynucleotide, may vary depending on a variety of factors, e.g., the desired biological response, the particular allele, genome, target site, cell, or tissue being targeted, and the agent used.
[0049] The terms "adeno-associated virus" or "AAV," as used interchangeably herein, refer to a small virus belonging to the genus Dependovirus in the family Parvoviridae that infects humans and some other primate species. AAV is not currently known to cause disease, and as a result, the virus provokes a very mild immune response.
[0050] The terms "subject" and "patient", as used interchangeably herein, refer to any vertebrate, including, but not limited to, mammals {e.g., cows, pigs, camels, llamas, horses, goats, rabbits, sheep, hamsters, guinea pigs, cats, dogs, rats and mice, non-human primates (e.g., monkeys such as cynomolgus or rhesus monkeys, chimpanzees), and humans). In some embodiments, the subject can be human or non-human. The subject or patient may be undergoing other forms of treatment.
[0051] The terms "treatment", "treat" and "treating" refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder or one or more symptoms thereof, as described herein. As used herein, the terms "treatment", "treat" and "treating" refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder or one or more symptoms thereof, as described herein. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In other embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay the onset of symptoms or to inhibit the onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, e.g., to prevent or delay recurrence.
[0052] Unless otherwise defined herein, scientific and technical terms used in connection with this disclosure shall have the meanings that are commonly understood by those of ordinary skill in the art.
[0053] Fusion Proteins and Protein Complexes The present disclosure provides a fusion protein comprising (i) a site-specific nuclease domain and (ii) a cytidine deaminase domain.Examples of site-specific nuclease domains are known to those skilled in the art, and include zinc finger nuclease (ZFN), TAL effector nuclease (TALEN), and CRISPR-based systems.Any nucleic acid guided DNA binding domain can be used, as long as the domain is guided to a specific point of interest within the target nucleic acid sequence.CRISPR nuclease domains are suitable for such purposes.
[0054] More specifically, the present disclosure provides a fusion protein comprising (i) a CRISPR nuclease domain and (ii) a cytidine deaminase domain. In some embodiments, the cytidine deaminase comprises one of the above sequences. Suitable CRISPR nuclease domains are described herein, with Cas9 being commonly used. Thus, in certain embodiments, the inactive CRISPR nuclease domain is a dCas9 domain. Other suitable site-specific nuclease domains will be apparent to those skilled in the art based on the present disclosure. Preferably, the CRISPR nuclease domain is a CRISPR nickase domain or an inactive CRISPR nuclease domain.
[0055] The present disclosure provides CRISPR nuclease enzyme / domain fusion proteins with various configurations.In some embodiments, cytidine deaminase enzyme or domain is fused to the N-terminus of CRISPR nuclease domain.In some embodiments, cytidine deaminase enzyme or domain is fused to the C-terminus of CRISPR nuclease domain.
[0056] In some embodiments, the general construct of the Cas fusion proteins provided herein comprises the following structure: ·[NH 2 ]-[cytidine deaminase domain]-[inactive CRISPR nuclease domain]-[COOH], ·[NH 2 ]-[inactive CRISPR nuclease domain]-[cytidine deaminase domain]-[COOH], ·[NH 2 ]-[cytidine deaminase domain]-[CRISPR nickase domain]-[COOH], or ·[NH 2 ]-[CRISPR nickase domain]-[cytidine deaminase domain]-[COOH], Here, NH 2 is the N-terminus of the fusion protein, COOH is the C-terminus of the fusion protein, and "-" is an optional linker.
[0057] In some embodiments, a general construct for a Cas fusion protein comprises the following structure: ·[NH 2 ]-[NLS]-[CRISPR nuclease domain]-[cytidine deaminase domain]-[COOH], ·[NH 2 ]-[NLS]-[cytidine deaminase domain]-[CRISPR nuclease domain]-[COOH], ·[NH 2 ]-[CRISPR nuclease domain]-[cytidine deaminase domain]-[COOH], or ·[NH 2 ]-[cytidine deaminase domain]-[CRISPR nuclease domain]-[COOH], where NLS is the nuclear localization signal and NH 2 is the N-terminus of the fusion protein, COOH is the C-terminus of the fusion protein, and "-" is an optional linker.
[0058] In some embodiments, the NLS is located at the C-terminus of the cytidine deaminase and / or CRISPR nuclease domain. There may be multiple NLSs. In some embodiments, the NLS is located between the cytidine deaminase domain and the Cas domain. In some embodiments, there are multiple NLSs. Preferably, the NLS is located at both the C-terminus and the N-terminus of the cytidine deaminase and / or CRISPR nuclease domain.
[0059] In some embodiments, the CRISPR nuclease domain and the cytidine deaminase domain are fused via a linker. In some embodiments, the linker comprises the sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 31). In some embodiments, the linker comprises the sequence (GGGGS) n (SEQ ID NO: 32), (G) n , (EAAAK) n (SEQ ID NO: 33) or (XP)n In some embodiments, n is independently 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30, or any combination thereof, when more than one linker or more than one linker motif is present. Further suitable linker motifs and linker configurations will be apparent to one of skill in the art. In some embodiments, suitable linker motifs and configurations include those described in Chen et al., Adv Drug Deliv Rev. 2013;65(10):1357-69. In some embodiments, the fusion proteins provided herein comprise the full length amino acid sequence of a nucleic acid editing enzyme, such as one of the sequences provided above. However, in other embodiments, the fusion proteins provided herein do not include the full-length sequence of a nucleic acid editing enzyme, but only a fragment thereof. For example, in some embodiments, the fusion proteins provided herein include a Cas9 domain and a fragment of a nucleic acid editing enzyme, e.g., the fragment includes a nucleic acid editing domain. Exemplary amino acid sequences of nucleic acid editing domains are shown in the sequences above as italicized letters, and further suitable sequences of such domains will be apparent to one of skill in the art.
[0060] In some embodiments, additional features may be present. Such features may be one or more linker sequences between the NLS and the rest of the fusion protein and / or between the cytidine deaminase domain and the CRISPR nuclease domain. Other features may be present, such as, for example, a nuclear localization sequence, a cytoplasmic localization sequence, a transport sequence such as a nuclear export sequence, or other localization sequence. In some embodiments, sequence tags may be present. Such tags are useful for solubilizing, purifying, or detecting the fusion protein. Suitable localization signal sequences and protein tag sequences are provided herein, including, but not limited to, biotin carboxylase carrier protein (BCCP) tag, myc tag, calmodulin tag, FLAG tag, hemagglutinin (HA) tag, polyhistidine tag, also called histidine tag or His tag, maltose binding protein (MBP) tag, nus tag, glutathione-S-transferase (GST) tag, green fluorescent protein (GFP) tag, thioredoxin tag, S tag, Softag (e.g., Softag1, Softag3), strep tag, biotin ligase tag, FlAsH tag, V5 tag, and SBP tag. Further suitable sequences will be apparent to those skilled in the art.
[0061] Further suitable nucleic acid editing enzyme sequences, such as deaminase enzymes and domain sequences, which can be used according to aspects of the present invention, for example, fused to nuclease-inactive CRISPR-associated domains, will be clear to those skilled in the art based on the present disclosure. In some embodiments, such further enzyme sequences comprise a cytidine deaminase domain sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% similar to the sequences provided herein. Further suitable CRISPR nuclease domains, variants and sequences will also be clear to those skilled in the art.
[0062] In some embodiments, the fusion protein provided herein comprises the full-length amino acid sequence of cytidine deaminase, such as one of the sequences provided above. However, in other embodiments, the fusion protein provided herein does not comprise the full-length sequence of cytidine deaminase, but only a fragment thereof. For example, in some embodiments, the fusion protein provided herein comprises a CRISPR nuclease domain (such as, for example, a Cas9 domain) and a fragment of the cytidine deaminase domain, for example, the fragment comprises the cytidine deaminase domain. Further suitable sequences of such domains will be apparent to those skilled in the art.
[0063] Deaminase domain Some aspects of the present disclosure provide fusion proteins and protein complexes comprising a cytidine deaminase domain. The cytidine deaminase domain can catalyze the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the cytidine deaminase domain catalyzes the hydrolytic deamination of cytidine to uracil. In some embodiments, the cytidine deaminase or the cytidine deaminase domain is a naturally occurring cytidine deaminase.
[0064] The present disclosure provides novel cytidine deaminase domains that can be used in fusion proteins or protein complexes comprising any one of SEQ ID NOs: 1-13. Examples of the present disclosure demonstrate cytidine base editing of several cytidine deaminase sequences. CBE07, CBE08, CBE10, CBE11, CBE12, and CBE13 demonstrate active base editing capabilities, with CBE07, CBE11, CBE12, and CBE13 being the most effective.
[0065] [Table 3]
[0066] In some embodiments, the cytidine deaminase or cytidine deaminase domain is a variant of a naturally occurring deaminase from a non-naturally occurring organism. For example, in some embodiments, the cytidine deaminase or cytidine deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the deaminase domains of SEQ ID NOs: 7, 8, 10, 11, 12, or 13.
[0067] In some embodiments, the cytidine deaminase domain is at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the deaminase domain of any one of SEQ ID NOs: 7, 8, 10, 11, 12, or 13. In some embodiments, the cytidine deaminase domain comprises the amino acid sequence of any one of SEQ ID NOs: 7, 8, 10, 11, 12, or 13.
[0068] The cytidine deaminase provided herein can be used for targeted editing of nucleic acid sequence.Such cytidine deaminase is useful for targeted editing of DNA in vitro, for example, generating mutant cells or animals, introducing targeted mutations, for example, correcting genetic defects in ex vivo cells, for example, cells obtained from a subject that are subsequently reintroduced into the same or another subject, and introducing targeted mutations, for example, correcting genetic defects or introducing inactivating mutations in disease-related genes of a subject.
[0069] In some embodiments, the cytidine deaminase domain has a catalytic mutation that reduces but does not eliminate the catalytic activity of the cytidine deaminase domain in the base editing fusion protein. Such a mutation may reduce the likelihood that the cytidine deaminase domain will catalyze the deamination of a residue adjacent to a target residue, thereby narrowing the deamination window. The ability to narrow the deamination window may be useful in preventing undesired deamination of a residue adjacent to a particular target residue, which may be useful in reducing or preventing off-target effects.
[0070] In some embodiments, any of the fusion proteins provided herein comprises a cytidine deaminase domain with reduced catalytic deaminase activity. In some embodiments, any of the fusion proteins provided herein comprises a cytidine deaminase domain with reduced catalytic deaminase activity compared to a suitable control. For example, a suitable control can be the deaminase activity of the cytidine deaminase before introducing one or more mutations into the cytidine deaminase. In other embodiments, a suitable control can be a wild-type deaminase. In some embodiments, a suitable control is a wild-type apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, a suitable control is an APOBEC1 deaminase, an APOBEC2 deaminase, an APOBEC3A deaminase, an APOBEC3B deaminase, an APOBEC3C deaminase, an APOBEC3D deaminase, an APOBEC3F deaminase, an APOBEC3G deaminase, or an APOBEC3H deaminase. In some embodiments, a suitable control is activation induced deaminase (AID). In some embodiments, the deaminase domain can be a deaminase domain that has at least 1%, at least 5%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% less catalytic deaminase activity compared to a suitable control.
[0071] Site-specific nuclease domain Some embodiments of the present disclosure provide fusion proteins and protein complexes that include an RNA-guided endonuclease domain that binds to a guide RNA (gRNA or sgRNA) and then binds to a target nucleic acid sequence via strand hybridization, and a cytidine deaminase domain that can deaminate cytidine.
[0072] Typically, the RNA-guided endonuclease domain of the fusion protein described herein partially lacks nuclease activity or has no nuclease activity. Such a domain can be a fragment of a nuclease-inactive Cas9 protein or a dCas9 protein or domain. Typically, nCas9 (nickase Cas9) is used to cleave (template) the RNA-bound strand of DNA when Cas9 binds. This helps to instruct the fusion protein (or protein complex) to use the non-template (displaced) strand with the deaminase domain acting as a repair template, and helps to fix the mutated cytosine in place with the base-paired adenosine before the uracil is removed.
[0073] In some embodiments, the RNA-guided endonuclease fusion protein or protein complex comprises at least one nuclear localization signal that allows the endonuclease to enter the nucleus of a eukaryotic cell. The RNA-guided endonuclease also comprises at least one nuclease domain and at least one domain that interacts with the guide RNA. The RNA-guided endonuclease is directed to a specific nucleic acid sequence (or target site) by the guide RNA. The guide RNA interacts with the RNA-guided endonuclease and the target site to direct the RNA guide to the nucleic acid sequence of the target site to which the guide RNA is complementary. Because the guide RNA provides specificity to the target site, the endonuclease of the RNA-guided endonuclease is universal and can be used with different guide RNAs for targeted binding to different nucleic acid sequences.
[0074] RNA-guided endonucleases can be derived from clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated (Cas) systems.
[0075] Generally, CRISPR / Cas protein comprises at least one RNA recognition and / or RNA binding domain. The RNA recognition and / or RNA binding domain interacts with guide RNA. CRISPR / Cas protein may also comprise a nuclease domain (i.e., DNase or RNase domain), a DNA binding domain, a helicase domain, an RNAse domain, a protein-protein interaction domain, a dimerization domain, and other domains.
[0076] The CRISPR / Cas system may be a type I, II, III, IV, V or VI system. Non-limiting examples of suitable CRISPR / Cas proteins include Cas1, Cas2, Cas3, Cas4, Cas4, Cas5, Cas7, Cas7, Cas8, Cas9, Cas10, Cas12 (Cpf1), Cas13 (C2c2), Csm and Cmr. For an overview of various examples of CRISPR / Cas proteins, see Makarova, et al. Nat Rev Microbiol 18, 67-83 (2020).
[0077] In one embodiment, the RNA-guided endonuclease is derived from a type II CRISPR / Cas system. In a particular embodiment, the RNA-guided endonuclease is derived from a Cas9 or Cas12 protein.
[0078] The CRISPR / Cas protein may be a wild-type CRISPR / Cas protein, a modified CRISPR / Cas protein, or a fragment of a wild-type or modified CRISPR / Cas protein. The CRISPR / Cas protein may be modified to increase nucleic acid binding affinity and / or specificity, change enzymatic activity, and / or change another property of the protein. For example, the nuclease (i.e., DNase, RNase) domain of the CRISPR / Cas-like protein may be modified, deleted, or inactivated. Alternatively, the CRISPR / Cas protein may be truncated to remove domains that are not essential for the function of the fusion protein or protein complex. The CRISPR / Cas protein may also be truncated or modified to optimize the activity of the effector domain of the fusion protein.
[0079] In some embodiments, the CRISPR / Cas protein may be derived from a wild-type Cas9 protein or a fragment thereof. In other embodiments, the CRISPR / Cas protein may be derived from a modified Cas9 protein. For example, the amino acid sequence of the Cas9 protein may be modified to change one or more properties of the protein (e.g., nuclease activity, affinity, stability, etc.). Alternatively, domains of the Cas9 protein that are not involved in RNA-guided cleavage may be removed from the protein such that the modified Cas9 protein is smaller than the wild-type Cas9 protein.
[0080] Cas9 proteins generally contain at least two nuclease domains. For example, Cas9 proteins may contain a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC and HNH domains cooperate to cleave a single strand and create a double strand break in DNA.
[0081] In some embodiments, the Cas9 protein may be modified to contain only one functional nuclease domain (either RuvC-like or HNH-like nuclease domain). For example, the Cas9-derived protein may be modified such that one of the nuclease domains is deleted or mutated to be non-functional (i.e., no nuclease activity is present). In some embodiments where one of the nuclease domains is inactive, the Cas9-derived protein is capable of introducing nicks into double-stranded nucleic acids (such proteins are called "nickases") but is unable to cleave double-stranded DNA. For example, an aspartic acid to alanine (D10A) conversion in the RuvC-like domain converts the Cas9-derived protein into a nickase. Similarly, an H840A or H839A mutation in the HNH domain converts the Cas9-derived protein into a nickase.
[0082] Each nuclease domain can be modified using well-known methods such as site-directed mutagenesis, PCR-mediated mutagenesis, and total gene synthesis, as well as other methods known in the art.
[0083] Non-limiting exemplary nuclease-inactive Cas9 domains are known to those of skill in the art. One exemplary suitable nuclease-inactivated S. pyogenes Cas9 domain is the D10A / H840A Cas9 domain mutant.
[0084] In a preferred embodiment, the RGN domain is nNme2Cas9 (with a D16A mutation) having the sequence of SEQ ID NO: 34.
[0085] Additional suitable nuclease-inactive CRISPR-associated domains will be apparent to those of skill in the art based on the present disclosure. Additional exemplary suitable nuclease-inactive spCas9 domains include, but are not limited to, D10A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (e.g., Prashant et al. Nature Biotechnology. 2013;31(9):833-838).
[0086] In some embodiments, the Cas9 fusion proteins provided herein comprise the full-length amino acid sequence of the Cas9 protein. However, in other embodiments, the fusion proteins provided herein do not comprise the full-length Cas9 sequence, but only a fragment thereof. For example, in some embodiments, the Cas9 fusion proteins provided herein comprise a Cas9 fragment, which binds to the crRNA and tracrRNA or sgRNA, but does not comprise a functional nuclease domain, e.g., in that it comprises only a truncated form of the nuclease domain or does not comprise the nuclease domain at all. Exemplary amino acid sequences of suitable modified Cas9 domains are described, for example, in Oakes et al., Cell. 2019 Jan 10; 176(1-2):254-267. Further suitable sequences of Cas9 domains and fragments will be apparent to those skilled in the art.
[0087] In some embodiments, Cas9 is used to inhibit Corynebacterium ulcerans (NCBI Reference Nos. NC_015683.1, NC_017317.1), Corynebacterium diphtheria (NCBI Reference Nos. NC_016782.1, NC_016786.1), Spiroplasma syrphidicola (NCBI Reference No. NC_021284.1), Prevotella intermedia (NCBI Reference No. NC_017861.1), Spiroplasma taiwanense (NCBI Reference No. NC_021846.1), Streptococcus iniae (NCBI Reference No. NC_021846.1), iniae (NCBI Reference: NC_021314.1), Belliella baltica (NCBI Reference: NC_018010.1), Psychroflexus torquis I (NCBI Reference: NC_018721.1), Streptococcus thermophilus (NCBI Reference: YP_820832.1), Listeria innocua (NCBI Reference: NP_472073.1), Campylobacter jejuni (NCBI Reference: NP_472073.1), jejuni (NCBI Reference: YP_002344900.1), or Neisseria meningitidis (NCBI Reference: YP_002342100.1).
[0088] Some aspects of the present disclosure provide RNA-guided endonuclease domains with different PAM specificities. Typically, RNA-guided endonuclease proteins, such as the commonly used Cas9 (spCas9) from Streptococcus pyogenes (S. pyogenes), require a canonical NGG PAM sequence to bind to a specific nucleic acid region. Having a nuclease domain that requires a specific PAM sequence may limit the ability to edit a desired base in a genome. In some embodiments, the fusion proteins provided herein may need to be placed at a precise location. For example, in the case of spCas9, the target base is placed within a 4-base region (e.g., the "deamination window") that is about 15 bases upstream of its PAM (Komor, AC, et al., Nature 533, 420-424, 2016). Thus, in some embodiments, any of the fusion proteins or protein complexes provided herein may contain an RGN domain that can bind to nucleotide sequences that do not contain a canonical (e.g., NGG) PAM sequence. RGN domains that bind to non-canonical PAM sequences have been described in the art and will be apparent to one of skill in the art. For example, a Cas9 domain that binds to a non-canonical PAM sequence is described in Kleinstiver, BP, et al., Nature 523, 481-485 (2015).
[0089] The compact orthologous RNA-guided endonuclease from Neisseria meningitidis (NmeCas9) recognizes a simple dinucleotide PAM (nnnnCC) that provides high target site density (Edraki et al. Mol Cell. 2019 Feb 21;73(4):714-726) and is a preferred variant of the Cas9 domain used in the fusion proteins and protein complexes described herein.
[0090] Uracil-protected peptides Some aspects of the present disclosure relate to fusion proteins and protein complexes comprising uracil-protecting peptides (UPPs). Examples of such peptides include uracil glycosylase inhibitor (UGI) (U.S. Pat. No. 10,167,457) and p56. Both UGI and p56 have been shown to inhibit the activity of uracil DNA-glycosylase (UDG). Fusion proteins comprising a cytidine deaminase domain, a dCas (e.g., dCas9) domain, and uracil glycosylase inhibitor (UGI) have been demonstrated to improve the efficiency of deaminating target nucleotides (Komor, et al. Nature 533, 420-424 (2016)). Without wishing to be bound by any particular theory, the cellular DNA repair response to the presence of U:G heteroduplex DNA may be responsible for the reduced nucleobase editing efficiency in cells.
[0091] The present disclosure provides novel UPPs useful in the context of base editing. Preferably, such UPPs comprise the sequence of SEQ ID NO: 43 or 45.
[0092] In some embodiments, any of the fusion proteins provided herein that include an RNA-guided endonuclease domain (e.g., a nuclease-active Cas9 domain, a nuclease-inactive dCas9 domain, or a Cas9 nickase) can be further fused to a UPP, either directly or via a linker.
[0093] Some embodiments of the present disclosure provide cytidine deaminase-dCas9 fusion proteins, cytidine deaminase-nuclease active Cas9 fusion proteins, and cytidine deaminase-Cas9 nickase (nCas9) fusion proteins comprising a UPP. In one embodiment, the present disclosure provides a fusion protein or protein complex comprising (i) an RNA-guided endonuclease domain (such as a nuclease active Cas9 domain, a nuclease inactive dCas9 domain, or a Cas9 nickase), (ii) a cytidine deaminase domain, and (iii) a uracil protecting peptide (UPP) comprising the sequence of SEQ ID NO: 43 or 45.
[0094] Without wishing to be bound by any particular theory, cellular DNA repair response to the presence of U:G heteroduplex DNA may be responsible for the reduced efficiency of nucleic acid base editing in cells. For example, uracil DNA glycosylase (UDG) catalyzes the removal of U from DNA in cells, which can initiate base excision repair, and the most common result is that U:G pair reverts to C:G pair. Thus, the present disclosure contemplates a fusion protein comprising a dCas9 nucleic acid editing domain further fused to a UPP. The present disclosure also contemplates a fusion protein comprising a Cas9 nickase nucleic acid editing domain further fused to a UPP. The use of a UPP may increase the editing efficiency of the cytidine deaminase domain that catalyzes the change from C to U.
[0095] In some embodiments, the fusion protein has the structure: ·[Cytidine deaminase]-[dCas9]-[UPP], · [Cytidine deaminase]-[UPP]-[dCas9], · [UPP]-[cytidine deaminase]-[dCas9], [UPP]-[dCas9]-[cytidine deaminase], [dCas9]-[cytidine deaminase]-[UPP], or [dCas9]-[UPP]-[cytidine deaminase] where "-" is an optional linker sequence.
[0096] In other embodiments, the fusion protein has the structure: ·[Cytidine deaminase]-[nCas9]-[UPP], ·[Cytidine deaminase]-[UPP]-[nCas9], ·[UPP]-[cytidine deaminase]-[nCas9], [UPP]-[nCas9]-[cytidine deaminase], ·[nCas9]-[cytidine deaminase]-[UPP], [nCas9]-[UPP]-[cytidine deaminase] wherein nCas9 is Cas9 nickase, and "-" is an optional linker sequence.
[0097] Exemplary sequences of uracil-protecting peptides are provided in U.S. Pat. No. 10,167,457 and in Hao-Ching Wang et al, Nucleic Acids Research, 42, 2, pp. 1354-1364.
[0098] Guide RNA-fusion protein complex Some aspects of the disclosure provide a complex comprising any of the fusion proteins or protein complexes provided herein and a guide RNA bound to a Cas domain of the fusion protein (e.g., dCas9, nuclease-active Cas9, or Cas9 nickase).
[0099] In some embodiments, the guide RNA is 15-300 nucleotides long and comprises a sequence of at least 10 contiguous nucleotides complementary to the target sequence. In some embodiments, the guide RNA comprises a sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 contiguous nucleotides complementary to the target sequence. In some embodiments, the target sequence is a DNA sequence. In some embodiments, the target sequence is a sequence in a mammalian, plant, or bacterial genome. In some embodiments, the target sequence is a sequence in a human genome. In some embodiments, the 3' end of the target sequence is immediately adjacent to a PAM sequence (e.g., the canonical PAM sequence NGG for SpCas9). In some embodiments, the guide RNA is complementary to a sequence associated with a disease or disorder.
[0100] Uses of Fusion Proteins and Protein Complexes Fusion proteins comprising cytidine deaminase domains can be used for targeted editing of nucleic acid sequences. Such fusion proteins are useful for targeted editing of DNA in vitro, for example, for generating mutant cells or animals, for introducing targeted mutations, for example, for correcting genetic defects in ex vivo cells, for example, cells obtained from a subject that are subsequently reintroduced into the same or another subject, and for introducing targeted mutations, for example, for correcting genetic defects or introducing inactivating mutations in disease-related genes of a subject.
[0101] Some aspects of the present disclosure provide methods of using the cytidine deaminase domain, fusion protein or complex provided herein. In one aspect of the present disclosure, a method is provided that includes contacting a DNA molecule with any of the fusion proteins or protein complexes provided herein and at least one guide RNA, wherein the guide RNA is about 15-300 nucleotides in length and comprises a sequence of at least 10 contiguous nucleotides complementary to a target sequence of interest.
[0102] Alternatively, the method comprises contacting a DNA molecule with a cytidine deaminase domain of a fusion protein provided herein together with at least one gRNA provided herein. In some embodiments, the 3' end of the target sequence is not immediately adjacent to a PAM sequence. In some embodiments, the 3' end of the target sequence is immediately adjacent to a standard PAM sequence (NGG), such as the AGC, GAG, TTT, GTG, or CAA sequence of SpCas9.
[0103] In some embodiments, the target DNA sequence comprises a sequence associated with a disease or disorder. In some embodiments, the target DNA sequence comprises a point mutation associated with a disease or disorder. In some embodiments, the activity of the cytidine deaminase domain, cytidine deaminase fusion protein or complex results in correction of the point mutation. In some embodiments, the target DNA sequence comprises a T→C point mutation associated with a disease or disorder, and deamination of the mutant C base results in a sequence that is not associated with the disease or disorder. In some embodiments, the target DNA sequence encodes a protein, and the point mutation is in a codon, resulting in a change in the amino acid encoded by the mutant codon compared to the wild-type codon. In some embodiments, deamination of the mutant C results in a change in the amino acid encoded by the mutant codon. In some embodiments, deamination of the mutant C results in a codon that encodes the wild-type amino acid. In some embodiments, the contacting is in vivo in a subject. In some embodiments, the subject has a disease or disorder or has been diagnosed with a disease or disorder.
[0104] Some embodiments provide methods for using the cytidine deaminase fusion proteins or protein complexes provided herein. In some embodiments, the fusion proteins are used to introduce point mutations into nucleic acids by deaminating targeted C residues. In some embodiments, deamination of the targeted nucleobase results in the correction of a genetic defect, e.g., a point mutation leading to loss of function in a gene product. In some embodiments, the genetic defect is associated with a disease or disorder. In some embodiments, the methods provided herein are used to introduce inactivating point mutations into genes or alleles that encode gene products associated with a disease or disorder. For example, in some embodiments, methods are provided herein that utilize the DNA editing fusion proteins provided herein to introduce inactivating point mutations into cancer genes (e.g., in the treatment of proliferative diseases). Inactivating mutations, in some embodiments, can generate premature stop codons in the coding sequence, resulting in the expression of truncated gene products, e.g., truncated proteins that lack the function of the full-length protein.
[0105] In some embodiments, the purpose of the methods provided herein is to restore the function of a dysfunctional gene through genome editing. The Cas-cytidine deaminase fusion proteins provided herein can be validated for in vitro gene editing-based human therapeutics, for example, by correcting disease-associated mutations in human cell cultures. It will be understood by those skilled in the art that the fusion proteins provided herein, such as fusion proteins comprising a Cas9 domain and a nucleic acid deaminase domain, can be used to correct any single-point T→C or A→G mutation. In the former case, the mutation is corrected by deamination of the mutant C back to U, and in the latter case, the mutation is corrected by deamination of the C base-paired with the mutant G followed by a series of duplications.
[0106] Base editing system Some aspects of the present disclosure provide base editing systems comprising the fusion proteins and protein complexes of the cytidine deaminase disclosed herein. More specifically, the present disclosure provides base editing systems comprising (1) a fusion protein or protein complex comprising an RGN and a cytidine deaminase domain provided herein, and (2) a guide RNA that binds to the RNA-guided nuclease of the fusion protein or protein complex.
[0107] In some embodiments, the fusion protein further comprises a UPP as disclosed herein.
[0108] Nucleic acids, genetic constructs and vectors for expression of base editing systems The present disclosure further provides nucleic acids encoding components of the base editing system disclosed herein and genetic constructs comprising such nucleic acids. A genetic construct, such as a plasmid or expression vector, can comprise a nucleic acid encoding a fusion protein or protein complex of an RNA-guided nuclease and a cytidine deaminase and / or at least one gRNA targeting a nucleic acid of interest.
[0109] The present disclosure also contemplates compositions of a nucleic acid encoding a modified AAV vector and one or more nucleic acid sequences encoding components of the base editing system disclosed herein. The compositions may include a nucleic acid encoding a modified lentiviral vector.
[0110] The nucleic acid may be present in the cell as a functional extrachromosomal molecule. The genetic construct may be a linear minichromosome containing a centromere, a telomere, or a plasmid or a cosmid. The nucleic acid may also be part of the genome of a recombinant viral vector, including recombinant lentiviruses, recombinant adenoviruses, and recombinant adenovirus-associated viruses. The nucleic acid may be part of the genetic material in an attenuated live microorganism or a recombinant microbial vector.
[0111] The nucleic acid may comprise regulatory elements for gene expression of the coding sequence of the nucleic acid. The regulatory elements may be a promoter, an enhancer, an initiation codon, a stop codon, or a polyadenylation signal.
[0112] The nucleic acid sequence may be in the form of a vector. The vector may be capable of expressing a base editing system provided herein in a mammalian cell. The vector may be recombinant. The vector may comprise a heterologous nucleic acid encoding a fusion protein or protein complex provided herein. The vector may be a plasmid. The vector may be useful for transfecting a cell with a nucleic acid encoding a fusion protein or protein complex.
[0113] The coding sequences of the fusion proteins and protein complexes provided herein may be optimized for stability and high levels of expression. The coding sequences may also be codon-optimized for expression in target cells. In some embodiments, the coding sequences are codon-optimized for expression in a particular cell, such as a eukaryotic cell. The eukaryotic cell may be of or derived from a particular organism, such as a mammal, including but not limited to a human, mouse, rat, rabbit, dog, or non-human primate. In general, codon optimization refers to the process of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon (e.g., more than about or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) of the native sequence with a codon that is more frequently or most frequently used in the genes of the host cell, while maintaining the native amino acid sequence. Different species show a particular bias for certain codons of certain amino acids.
[0114] The vector may comprise a heterologous nucleic acid encoding a base editing system, and may further comprise a start codon, which may be upstream of the base editing system coding sequence, and a stop codon, which may be downstream of the base editing system coding sequence.
[0115] In some embodiments, vectors for expression of the fusion proteins and protein complexes described herein contain one or more nuclear localization sequences (NLSs). In some embodiments, the vectors contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the amino terminus, more than about or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the carboxy terminus, or a combination of one or more NLSs at the amino terminus and one or more NLSs at the carboxy terminus. Typically, an NLS consists of one or more short sequences of positively charged lysines or arginines exposed on the protein surface, although other types of NLSs are known. Non-limiting examples of NLSs include the NLS of the SV40 virus large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO: 29), and the NLS from nucleoplasmin having the amino acid sequence KRPAATKKAGQAKKKK (SEQ ID NO: 30).
[0116] The vector may also include a promoter operably linked to the base editing system coding sequence. The promoter operably linked to the base editing system coding sequence may be a promoter derived from Simian Virus 40 (SV40), a mouse mammary tumor virus (MMTV) promoter, a human immunodeficiency virus (HIV) promoter, such as a bovine immunodeficiency virus (BIV) long terminal repeat (LTR) promoter, a cytomegalovirus (CMV) promoter, such as a CMV immediate early promoter, an Epstein-Barr virus (EBV) promoter, or a Rous sarcoma virus (RSV) promoter. The promoter may also be a promoter derived from a human gene. The promoter may also be a tissue-specific promoter. Examples of such promoters are described in U.S. Patent Application Publication No. 2004 / 017572.
[0117] The vector may also include a polyadenylation signal that may be downstream of the base editing system. The polyadenylation signal may be an SV40 polyadenylation signal, an LTR polyadenylation signal, a bovine growth hormone (bGH) polyadenylation signal, a human growth hormone (hGH) polyadenylation signal, or a human β-globin polyadenylation signal.
[0118] The vector may also include an enhancer upstream of the base editing system component. Examples of enhancers are described in U.S. Pat. No. 5,593,972, U.S. Pat. No. 5,962,428, and WO 94 / 016737. The vector may also include a mammalian origin of replication to maintain the vector extrachromosomally and produce multiple copies of the vector within the cell. The vector may also include regulatory sequences that may be well suited for gene expression in mammalian or human cells to which the vector is administered. The vector may also include a reporter gene and / or a selection marker, such as green fluorescent protein ("GFP").
[0119] In some aspects, the disclosure provides methods that include delivering one or more polynucleotides, e.g., or one or more vectors described herein, one or more transcripts thereof, and / or one or more proteins transcribed therefrom, to a host cell. In some aspects, the invention further provides cells produced by such methods, and organisms (e.g., animals, plants, or fungi) that comprise or are produced from such cells.
[0120] Methods for editing nucleic acids Some aspects of the present disclosure provide a method for editing nucleic acid. In some embodiments, the method is a method for editing the base of a nucleic acid. In some embodiments, the method comprises contacting a target region of a double-stranded nucleic acid (such as DNA) with a fusion protein or protein complex provided herein and a guide RNA that is complementary to the target region, and the target region comprises a targeted nucleic acid base pair to be edited.
[0121] In some embodiments, the method for editing bases of a double stranded nucleic acid comprises: a) contacting a target region of a double-stranded nucleic acid with a complex comprising a fusion protein or protein complex provided herein and a guide RNA, wherein the target region comprises a targeted nucleic acid base pair to be edited; and b) converting a first nucleobase of said target nucleobase pair in one strand of the target region to a second nucleobase. wherein a third nucleobase complementary to the first nucleobase is replaced by a fourth nucleobase complementary to the second nucleobase, and the method results in less than 20% indel formation in the nucleic acid.
[0122] In some embodiments, the first nucleobase is cytidine. In some embodiments, the second nucleobase is deaminated cytidine or uracil. In some embodiments, the third nucleobase is guanine. In some embodiments, the fourth nucleobase is adenine. In some embodiments, the first nucleobase is cytidine, the second nucleobase is deaminated cytidine or uracil, the third nucleobase is guanine, and the fourth nucleobase is adenine. In some embodiments, the method results in less than 19%, 18%, 16%, 14%, 12%, 10%, 8%, 6%, 4%, 2%, 1%, 0.5%, 0.2%, or 0.1% indel formation.
[0123] In some embodiments, the method further comprises replacing the second nucleobase with a fifth nucleobase that is complementary to the fourth nucleobase to generate the intended edited base pair (e.g., C:G→T:A). In some embodiments, the fifth nucleobase is thymine.
[0124] In some embodiments, at least 1% of the intended base pairs are edited, hi some embodiments, at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the intended base pairs are edited.
[0125] In some embodiments, the ratio of intended to unintended products in the target nucleotide is at least 2:1, 5:1, 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or 200:1 or more. In some embodiments, the ratio of intended point mutation to indel formation is greater than 1:1, 10:1, 50:1, 100:1, 500:1, or 1000:1 or more. In some embodiments, the cleaved single strand (the nicked strand) is hybridized to the guide nucleic acid. In some embodiments, the cleaved single strand is opposite the strand that includes the first nucleobase. In some embodiments, the base editor comprises a Cas domain, for example a Cas9 domain.
[0126] In some embodiments, the first base is cytidine. In some embodiments, the second base is not G, C, A, or T. In some embodiments, the second base is uracil. In some embodiments, the fusion proteins or protein complexes provided herein inhibit base excision repair of the edited strand. In some embodiments, the fusion proteins or protein complexes provided herein protect or bind the non-edited strand. In some embodiments, the fusion proteins or protein complexes provided herein protect or bind the edited strand. In some embodiments, the fusion proteins or protein complexes provided herein protect or bind the non-edited strand and the edited strand.
[0127] In some embodiments, the intended editing base pair is upstream of the PAM site. In some embodiments, the intended editing base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides upstream of the PAM site. In some embodiments, the intended editing base pair is downstream of the PAM site. In some embodiments, the intended editing base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides downstream of the PAM site. In some embodiments, the methods do not require a standard spCas9 (e.g., NGG) PAM site.
[0128] In some embodiments, the target sequence comprises a target window, and the target window comprises a target nucleic acid base pair. In some embodiments, the target window comprises 1-10 nucleotides. In some embodiments, the target window is 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, 1-2, or 1 nucleotide long. In some embodiments, the target window is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides long. In some embodiments, the intended editing base pair is within the target window. In some embodiments, the target window comprises the intended editing base pair. In some embodiments, the method is performed using any base editor provided herein. In some embodiments, the target window is a deamination window.
[0129] In some embodiments, the present disclosure provides a method for editing nucleotides. In some embodiments, the present disclosure provides a method for editing nucleic acid base pairs of a double-stranded DNA sequence. In some embodiments, the method comprises: a) contacting a target region of a double-stranded DNA sequence with a complex comprising a fusion protein or protein complex provided herein and a guide nucleic acid (e.g., gRNA), wherein the target region comprises a target nucleic acid base pair; and b) converting a first nucleobase of said target nucleobase pair in one strand of the target region to a second nucleobase. wherein a third nucleobase complementary to the first nucleobase is replaced by a fourth nucleobase complementary to the second nucleobase, and the second nucleobase is replaced by a fifth nucleobase complementary to the fourth nucleobase, thereby generating an intended edited base pair; The efficiency of generating the intended edited base pair is at least 1%.
[0130] In some embodiments, at least 5% of the intended base pairs are edited. In some embodiments, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the intended base pairs are edited. In some embodiments, the method causes less than 19%, 18%, 16%, 14%, 12%, 10%, 8%, 6%, 4%, 2%, 1%, 0.5%, 0.2%, or less than 0.1% indel formation. In some embodiments, the ratio of intended to unintended products at the target nucleotide is at least 2:1, 5:1, 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or 200:1 or more. In some embodiments, the ratio of intended point mutations to indel formation is greater than 1:1, 10:1, 50:1, 100:1, 500:1, or 1000:1 or more. In some embodiments, the cleaved single strand is hybridized to a guide nucleic acid. In some embodiments, the cleaved single strand is opposite the strand that includes the first nucleobase. In some embodiments, the first base is cytidine. In some embodiments, the second nucleobase is not G, C, A, or T. In some embodiments, the second base is uracil. In some embodiments, the base editor inhibits base excision repair of the edited strand. In some embodiments, the base editor protects or binds the non-edited strand. In some embodiments, the nucleobase editor comprises a UPP. In some embodiments, the nucleobase editing comprises a nickase activity. In some embodiments, the intended edited base pair is upstream of the PAM site. In some embodiments, the intended editing base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides upstream of the PAM site. In some embodiments, the intended editing base pair is downstream of the PAM site.In some embodiments, the intended editing base pair is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides downstream of the PAM site. In some embodiments, the method does not require a PAM site. In some embodiments, the nucleobase editor comprises a linker. In some embodiments, the linker is 1-25 amino acids in length. In some embodiments, the linker is 5-20 amino acids in length. In some embodiments, the linker is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids in length. In some embodiments, the target region comprises a target window, and the target window comprises the target nucleobase pair. In some embodiments, the target window comprises 1-10 nucleotides. In some embodiments, the target window is 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, 1-2, or 1 nucleotides in length. In some embodiments, the target window is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length. In some embodiments, the intended editing base pair occurs within the target window. In some embodiments, the target window comprises the intended editing base pair. In some embodiments, the nucleobase editor is any one of the base editors provided herein.
[0131] therapeutic use The present disclosure provides methods for treating diseases or disorders, such as diseases or disorders associated with or caused by point mutations that can be corrected by cytidine deaminase gene editing. Some such diseases are described herein, and further suitable diseases that can be treated by the strategies and fusion proteins provided herein will be apparent to those skilled in the art based on the present disclosure. Exemplary suitable diseases and disorders are listed below. It is understood that the numbering of specific positions or residues in each sequence depends on the specific protein and numbering scheme used. The numbering may be different, for example, in the precursor of the mature protein and the mature protein itself, and sequence differences between species may affect the numbering. Those skilled in the art can identify any homologous proteins and the respective residues in their respective encoding nucleic acids by methods well known in the art, for example, by sequence alignment and determining homologous residues. Exemplary suitable diseases and disorders include, but are not limited to, cystic fibrosis (see, e.g., Schwank et al., Functional repair of CFTR by CRISPR / Cas9 in intestinal stem cell organoids of cystic fibrosis patients. Cell stem cell. 2013;13:653-658; and Wu et.al., Correction of a genetic disease in mouse via use of CRISPR-Cas9. Cell stem cell. 2013;13:659-662;
[0132] Pharmaceutical Compositions The composition of the present invention may be a pharmaceutical composition. The pharmaceutical composition may comprise about 1 ng to about 10 mg of DNA encoding a CRISPR / Cas9-based system or a protein component of a CRISPR / Cas9-based system, i.e., a fusion protein. The pharmaceutical composition may comprise about 1 ng to about 10 mg of DNA of a modified lentiviral vector. The pharmaceutical composition may comprise about 1 ng to about 10 mg of DNA of a modified AAV vector and a nucleotide sequence encoding a site-specific nuclease. The pharmaceutical composition according to the present invention may be formulated according to the mode of administration to be used. When the pharmaceutical compositions are injectable pharmaceutical compositions, they are sterile, pyrogen-free, and particulate-free. Isotonic formulations are preferably used. In general, additives for isotonicity may include sodium chloride, dextrose, mannitol, sorbitol, and lactose. In some cases, isotonic solutions such as phosphate-buffered saline are preferred. Stabilizers include gelatin and albumin. In some embodiments, a vasoconstrictor is added to the formulation.
[0133] The composition may further comprise a pharma- ceutically acceptable excipient. The pharma- ceutically acceptable excipient may be a functional molecule as a vehicle, adjuvant, carrier or diluent. The pharma- ceutically acceptable excipient may be a surface active agent, such as immune stimulating complexes (ISCOMS), Freund's incomplete adjuvant, LPS analogs including monophosphoryl lipid A, muramyl peptides, quinone analogs, vesicles such as squalene and squalene, hyaluronic acid, lipids, liposomes, calcium ions, viral proteins, polyanions, polycations or nanoparticles, or other known transfection facilitating agents.
[0134] The transfection facilitating agent may be a polyanion, a polycation, including poly-L-glutamate (LGS), or a lipid. The transfection facilitating agent is poly-L-glutamate, and more preferably poly-L-glutamate is present in a composition for genome editing in skeletal or cardiac muscle at a concentration of less than 6 mg / ml. The transfection facilitating agent may also include surfactants, such as immune stimulating complexes (ISCOMS), Freund's incomplete adjuvant, LPS analogs, including monophosphoryl lipid A, muramyl peptides, quinone analogs, and vesicles, such as squalene and squalene, and hyaluronic acid may also be used in combination with the gene construct. In some embodiments, the DNA vector encoding the composition may also include transfection facilitating agents, such as lipids, liposomes, including lecithin liposomes or other liposomes known in the art, calcium ions, viral proteins, polyanions, polycations, or nanoparticles, or other known transfection facilitating agents, as a DNA-liposome mixture (see, for example, WO9324640). Preferably, the transfection enhancing agent is a polyanion, a polycation, including poly-L-glutamate (LGS), or a lipid.
[0135] The sequences encompassed by the present invention are shown in Table 4.
[0136] [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5] [Table 4-6] [Table 4-7] EXAMPLES
[0137] Example 1. Identification of novel cytosine base editors The public bacterial genome collection was searched for sequences of less than 200 amino acids that contain the prototypical deaminase fold (Iyer et al Nucleic acids research, 39(22), 9473-9497, 2011) and active site (D / H / C-[X]-E-[X15-45]-PC-[X2]-C), are not part of a larger ORF, and are functionally less than 40% homologous to known cytosine deaminases. Thirteen candidate enzymes were selected for testing along with a known eukaryotic cytosine deaminase (hAPOBEC3A).
[0138] Example 2: Demonstration of base editing activity against endogenous targets in mammalian cells The coding sequence of the identified deaminase is codon-optimized for expression in mammalian cells and introduced into an expression cassette that produces a fusion protein containing a 3xFLAG tag at its N-terminus, operably linked to an NLS at its C-terminus, and operably linked to a putative deaminase sequence at its C-terminus. The putative deaminase is operably linked at its C-terminus to a flexible amino acid linker, which is operably linked at its C-terminus to a known active RNA-guided nuclease (nNme2Cas9_D16A) that is mutated to have an inactive RuvC domain (i.e., mutated to RGN that acts as a nickase). The RNA-guided DNA-binding polypeptide is operably linked at its C-terminus to a second NLS. Each of these expression cassettes is introduced into a vector that can drive the expression of the fusion protein in mammalian cells. Vectors that can express guide RNAs that target the deaminase-RGN fusion protein to the determined genomic location were also generated. These guide RNAs can guide deaminase-RGN fusion proteins to target genomic sequences for base editing.
[0139] Using liposomal transfection, HEK293T cells were transfected with vectors capable of expressing the deaminase-RGN fusion protein and guide RNAs (EGsG0033 and 0034). For liposomal transfection, cells were cultured in 24-well plates at 1.3 × 10 in growth medium (DMEM + 10% fetal bovine serum + 1% penicillin / streptomycin) the day before transfection. 5Distribute cells / well. Transfect 500 ng of deaminase-RGN fusion expression vector and 500 ng of guide RNA expression vector using Lipofectamine® 3000 reagent (Thermo Fisher Scientific) according to the manufacturer's instructions. 48-72 hours after liposomal transfection, harvest genomic DNA from transfected cells and sequence the DNA and analyze for the presence of targeted cytosine base editing mutations using CRISPResso2 (Clement K, Rees H, Canver MC, Gehrke JM, Farouni R, Hsu JY, Cole MA, Liu DR, Joung JK, Bauer DE, Pinello L. CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat Biotechnol. 2019 Mar;37(3):224-226. doi:10.1038 / s41587-019-0032-3. PubMed PMID:30809026).
[0140] Table 5 summarizes the base editing results for both target sequences, while Tables 6, 7, 8, and 9 show the cytidine base editing rates of CBE07, CBE08, CBE10, CBE11, CBE12, and CBE13 deaminases, and the targeted cytosine deamination rate of wild-type Nme2Cas9 targeted to the same region. Active cytosine base editing was defined as INDEL formation greater than the average of the novel CBEs under investigation, compared to the average C>T rate of normal Nme2Cas9 in the targeted region, and greater than 5× C>T SNP base editing at the targeted cytosine. Since the deaminase-RGN fusion protein consists of RGN with an inactive RuvC domain, increased INDEL formation indicates an active cytosine base editor. With only a single active nuclease domain, RGN functions only as a nickase and does not produce detectable INDEL formation by itself. When fused to an active deaminase acting on the opposite strand, cytosines become uracils. Uracil is rapidly removed from DNA, leaving an abasic site, and ultimately a gap, on the strand opposite the one nicked by RGN. This results in a double-strand break and detectable INDEL formation that is repaired through non-homologous end joining (NHEJ). A C>T SNP base edit of greater than 5× at the target cytosine indicates an active cytosine base editor, because if the uracil created by the deaminase is not removed before the nicked strand is repaired, it will be read as a thymine when the nicked strand is repaired and an adenosine will be inserted opposite it. When the uracil is then removed, it will be replaced by a thymine during the excision repair process, fixing the C>T mutation. Each CBE construct has a different editing window, or region of the target sequence over which the cytosine deaminase acts. This is driven by two features: 1) the steric properties of the direct fusion and how accessible the exposed single-stranded DNA is, and 2) what is the cytosine recognition preference of the deaminase itself. Successful cytosine editing occurs only if cytosines with favorable sequence motifs are located in favorable editing windows.TC is a common dinucleotide preference for existing cytosine base editors, and we have observed good editing by APOBEC3A at positions C12 and C14 of EGsG033 (ctggtggccatCtCcactgctagta) (SEQ ID NO: 39), but little editing at positions C8 (GC) or C15 (CC) because these are not preferred motifs. However, a novel cytosine deaminase has been identified in which the preferred editing occurs at C8 (GC).
[0141] [Table 5]
[0142] [Table 6]
[0143] [Table 7]
[0144] [Table 8]
[0145] [Table 9]
[0146] Example 3. Identification of novel uracil-protecting peptides (UPPs) The public genome collection was searched for sequences of less than 200 amino acids with a total of at least 10 aspartic acid and / or glutamic acid residues on the predicted protein surface and negative charges on at least 10% of the predicted surface residues (Wang et al. Nucleic Acids Res. 2014;42(2):1354-1364). Twenty-one candidate uracil-protecting peptides (UPPs) were selected for testing in fusion with a known eukaryotic cytosine deaminase (hAPOBEC3A).
[0147] Example 4: Demonstration of base editing activity of a fusion protein containing the UPP against an endogenous target in mammalian cells The coding sequence of the identified UPP is codon-optimized for expression in mammalian cells and introduced into an expression cassette that produces a fusion protein containing a 3xFLAG tag at its N-terminus, operably linked to an NLS at its C-terminus, and operably linked to a codon-optimized deaminase sequence at its C-terminus. The putative deaminase is operably linked at its C-terminus to a flexible amino acid linker, which is operably linked at its C-terminus to a known active RNA-guided nuclease (nNme2Cas9_D16A) that is mutated to have an inactive RuvC domain (i.e., mutated to RGN that acts as a nickase). The RNA-guided DNA-binding polypeptide is operably linked at its C-terminus to a flexible amino acid linker, which is operably linked to a putative uracil-protecting peptide. The putative uracil-protecting peptide is operably linked at its C-terminus to a flexible amino acid linker, which is operably linked at its C-terminus to a second NLS. Each of these expression cassettes is introduced into a vector capable of driving the expression of the fusion protein in mammalian cells. Vectors capable of expressing guide RNAs that target the deaminase-RGN-UPP fusion protein to the determined genomic location were also generated. These guide RNAs can guide the deaminase-RGN-UPP fusion protein to the target genomic sequence for base editing.
[0148] Using liposomal transfection, HEK293T cells were transfected with vectors capable of expressing the deaminase-RGN-UPP fusion protein and guide RNAs (EGsG0034 and 0041). For liposomal transfection, cells were cultured in 24-well plates at 1.3 × 10 in growth medium (DMEM + 10% fetal bovine serum + 1% penicillin / streptomycin) the day before transfection. 5Distribute cells / well. Transfect 500ng of deaminase-RGN fusion expression vector and 500ng of guide RNA expression vector using Lipofectamine® 3000 reagent (Thermo Fisher Scientific) according to the manufacturer's instructions. 48-72 hours after liposomal transfection, harvest genomic DNA from transfected cells, sequence DNA, and analyze for the presence of targeted cytosine base editing mutations using CRISPResso2 (Clement K, et al Nat Biotechnol. 2019;37(3):224-226).
[0149] Tables 11 and 12 show the cytidine base editing rates of UPP12 (SEQ ID NO: 43) and UPP14 (SEQ ID NO: 45), as well as the targeted cytosine deamination rate of the control deaminase-RGN targeted to the same region. Active cytosine base editing was defined as a reduction in INDEL formation of the novel UPP under investigation, an increase in C>D SNP base editing along the targeting window, and >85% C>T SNP base editing in highly mutated cytosines compared to deaminase-RGN that does not contain a UPP within the target region. Since the deaminase-RGN-UPP fusion protein consists of RGN with an inactive RuvC domain, a reduction in INDEL formation indicates an active UPP. With only a single active nuclease domain, RGN only functions as a nickase and does not produce detectable INDEL formation by itself. When fused to an active deaminase acting on the opposite strand, cytosine becomes uracil. Uracil is rapidly removed from DNA, leaving an abasic site, and ultimately a gap, on the strand opposite the one nicked by RGN. This results in a double-strand break and detectable INDEL formation that is repaired through non-homologous end joining (NHEJ). With the presence of active UPP, the converted uracil is protected from removal, the abasic site is never removed, and NHEJ does not occur. This also results in increased C>D SNP base editing and an increased total amount of C>T conversion, because the uracil created by the deaminase before the nicked strand is repaired is protected and, if not removed, will be read as a thymine when the nicked strand is repaired, and an adenosine will be inserted opposite it. When the uracil is removed, it is then replaced by a thymine during the excision repair process, fixing the C>T mutation.
[0150] [Table 10]
[0151] [Table 11]
[0152]
Table 12
Claims
1. a) a site-specific nuclease domain, and b) a cytidine deaminase domain comprising any one of the sequences set forth in SEQ ID NOs: 7, 8, 10, 11, 12, or 13; A fusion protein or protein complex comprising:
2. The fusion protein or protein complex of claim 1 further comprising a uracil protecting peptide (UPP).
3. The fusion protein or protein complex of claim 2 , wherein the UPP comprises the sequence of SEQ ID NO: 43 or 45.
4. The fusion protein or protein complex according to any one of claims 1 to 4, wherein the site-specific nuclease domain is an RNA-guided endonuclease (RGN) domain.
5. The fusion protein or protein complex of claim 4, wherein the RNA-guided endonuclease domain is a nickase domain or an inactivating endonuclease domain.
6. The fusion protein or protein complex of claim 4, wherein the RNA-guided endonuclease domain is a Cas9 domain.
7. The fusion protein or protein complex of claim 6, wherein the Cas9 domain is an Nme2Cas9 domain.
8. The fusion protein or protein complex of claim 7, wherein the Nme2Cas9 domain comprises the sequence of SEQ ID NO:
34.
9. The fusion protein of claim 1 further comprising a linker sequence.
10. The fusion protein of claim 9, wherein the linker sequence comprises the sequence of SEQ ID NO: 31, 32 or 33.
11. The fusion protein or protein complex of claim 1 further comprising one or more NLS sequences.
12. The fusion protein or protein complex according to claim 11, wherein the NLS sequence is a nucleoplasmin NLS or an SV40 NLS.
13. The fusion protein or protein complex of claim 11, wherein the NLS sequence comprises the sequence of SEQ ID NO: 29 or 30.
14. A polynucleotide sequence encoding a fusion protein or protein complex according to any one of claims 1 to 13.
15. A vector comprising the nucleotide sequence of claim 14.
16. A cell comprising a fusion protein or protein complex according to any one of claims 1 to 13, a polynucleotide according to claim 14, or a vector according to claim 15.
17. A method for editing bases of a target nucleic acid sequence in a target cell, the method comprising introducing into the target cell a fusion protein or protein complex described in any one of claims 1 to 13 or a vector described in claim 15.
18. 1. A method for editing the bases of a nucleic acid, the method comprising: a) contacting a target region of a double-stranded nucleic acid with a complex comprising a fusion protein or protein complex according to any one of claims 1 to 13 and a guide RNA, said target region comprising the targeted nucleic acid base pair to be edited; and b) converting a first nucleobase of said target nucleobase pair in a single strand of said target region to a second nucleobase; Including, a third nucleobase complementary to said first nucleobase is replaced by a fourth nucleobase complementary to said second nucleobase, said method resulting in less than 20% indel formation in the nucleic acid.
19. A pharmaceutical composition comprising a fusion protein or protein complex according to claims 1 to 13, one or more polynucleotides according to claim 14, or one or more vectors according to claim 15.