Programmable DNA base editing with NME2 cas9-deaminase fusion proteins

By fusing the NmeCas9 nuclease with a nucleotide deaminase protein and binding to a specific motif adjacent to the pre-interstitial region sequence, a compact Cas9 fusion construct is formed, which solves the problem of insufficient accuracy and precision of the SpyCas9 platform when targeting single base mutations, and achieves efficient single nucleotide base editing.

CN113166743BActive Publication Date: 2026-04-07UNIV OF MASSACHUSETTS
View PDF 13 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-10-15
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

The existing SpyCas9 base editing platform, due to its limited editing window and high off-target effect, cannot effectively target all single-base mutations, resulting in insufficient accuracy and precision in treating genetic diseases caused by single-base mutations.

Method used

By fusing the NmeCas9 nuclease with a nucleotide deaminase protein and binding to a specific motif adjacent to the pre-interstitial sequence, a compact Cas9 fusion construct is formed, enabling highly accurate and precise single nucleotide base editing of the N4CC nucleotide sequence.

Benefits of technology

It provides a programmable target-specific gene editing platform that can efficiently convert C·G base pairs into T·A base pairs, significantly improving the accuracy and precision of gene editing and reducing off-target effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113166743B_ABST
    Figure CN113166743B_ABST
Patent Text Reader

Abstract

This invention relates to the field of gene editing. Specifically, it targets single nucleotide base editing. For example, such single nucleotide base editing results in the conversion of C·G base pairs to T·A base pairs. The high accuracy and precision of the single nucleotide base gene editor disclosed herein are achieved through an NmeCas9 nuclease fused to a nucleotide deaminase protein. The compact nature of NmeCas9, coupled to motifs adjacent to a large number of compatible pre-interstitial sequences, allows the Cas9 fusion construct envisioned herein to possess a gene editing window that can edit sites that are untargetable by other conventional SpyCas9 base editor platforms.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 745,666, filed October 15, 2018, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This invention relates to the field of gene editing. Specifically, it targets single nucleotide base editing. For example, such single nucleotide base editing results in the conversion of C·G base pairs to T·A base pairs. The high accuracy and precision of the single nucleotide base gene editor disclosed herein are achieved through an NmeCas9 nuclease fused to a nucleotide deaminase protein. The compact nature of NmeCas9, coupled to motifs adjacent to a large number of compatible pre-interstitial sequences, allows the Cas9 fusion construct envisioned herein to possess a gene editing window that can edit sites that are not targeted by other conventional SpyCas9-based editor platforms. Background Technology

[0004] Many human diseases are caused by mutations in a single base. The ability to correct such genetic aberrations is crucial in treating these genetic conditions. Clusters of regularly spaced short palindromic repeats (CRISPR), along with CRISPR-associated (Cas) proteins, constitute the RNA-guided adaptive immune system in archaea and bacteria. These systems provide immunity by targeting and inactivating nucleic acids derived from exogenous genetic elements.

[0005] The SpyCas9 base editing platform cannot be used to target all single-base mutations due to its limited editing window. The editing window is partly limited by the requirements of the NGG PAM and the very precise distance between the edited base and the PAM. SpyCas9 is also inherently associated with high off-target effects in genome editing.

[0006] What is needed in this field is a highly accurate Cas9 single-base editing platform that has programmable target specificity due to its ability to recognize diverse populations of PAM sites. Summary of the Invention

[0007] This invention relates to the field of gene editing. Specifically, it targets single nucleotide base editing. For example, such single nucleotide base editing results in the conversion of C·G base pairs to T·A base pairs. The high accuracy and precision of the single nucleotide base gene editor disclosed herein are achieved through an NmeCas9 nuclease fused to a nucleotide deaminase protein. The compact nature of NmeCas9, coupled to motifs adjacent to a large number of compatible pre-interstitial sequences, enables the Cas9 fusion construct envisioned herein to possess a gene editing window superior to other conventional SpyCas9 base editor platforms.

[0008] In one embodiment, the present invention relates to a mutant NmeCas9 protein comprising a fused nucleotide deaminase and a binding region targeting an N4CC nucleotide sequence. In one embodiment, the protein is Nme2Cas9. In one embodiment, the protein further comprises a nuclear localization signaling protein. In one embodiment, the nucleotide deaminase is a cytidine deaminase. In one embodiment, the nucleotide deaminase is adenosine deaminase. In one embodiment, the protein further comprises a uracil glycosylation inhibitor. In one embodiment, the nuclear localization signaling protein includes, but is not limited to, nucleoplasmic proteins (NLS) and / or SV40 NLS and / or C-myc NLS. In one embodiment, the binding region is a pre-intercalary sequence auxiliary motif interaction domain. In one embodiment, the pre-intercalary sequence auxiliary motif interaction domain contains the mutation. In one embodiment, the mutation is a D16A mutation. In one embodiment, the mutant NmeCas9 protein further comprises CBE4. In one embodiment, the mutant NmeCas9 protein further comprises an adapter. In one embodiment, the adapter is a 73aa adapter. In one embodiment, the adapter is a 3xHA tag.

[0009] In one embodiment, the present invention relates to a construct, wherein the construct is an optimized nNme2Cas9-ABEmax.

[0010] In one embodiment, the present invention relates to a construct, wherein the construct is nNme2Cas9-CBE4.

[0011] In one embodiment, the present invention relates to a construct, wherein the construct is YE1-BE3-nNme2Cas9(D16A)-UGI.

[0012] In one embodiment, the present invention relates to an adeno-associated virus comprising a mutated NmeCas9 protein, said mutated NmeCas9 protein comprising a fused nucleotide deaminase and a binding region targeting an N4CC nucleotide sequence. In one embodiment, the virus is adeno-associated virus 8. In one embodiment, the virus is adeno-associated virus 6. In one embodiment, the protein is Nme2Cas9. In one embodiment, the protein further comprises a nuclear localization signal protein. In one embodiment, the nucleotide deaminase is cytidine deaminase. In one embodiment, the nucleotide deaminase is adenosine deaminase. In one embodiment, the protein further comprises a uracil glycosylation inhibitor. In one embodiment, the nuclear localization signal protein includes, but is not limited to, nucleoplasmic protein (NLS) and / or SV40 NLS and / or C-myc NLS. In one embodiment, the binding region is a pre-intercalary sequence auxiliary motif interaction domain. In one embodiment, the pre-intercalary sequence auxiliary motif interaction domain comprises the mutation. In one embodiment, the mutation is a D16A mutation. In one embodiment, the mutated NmeCas9 protein further comprises CBE4. In one embodiment, the mutated NmeCas9 protein further comprises a linker. In one embodiment, the linker is a 73aa linker. In one embodiment, the linker is a 3xHA tag.

[0013] In one embodiment, the present invention relates to a construct, wherein the construct is an optimized nNme2Cas9-ABEmax.

[0014] In one embodiment, the present invention relates to a construct, wherein the construct is nNme2Cas9-CBE4.

[0015] In one embodiment, the present invention relates to a construct, wherein the construct is YE1-BE3-nNme2Cas9(D16A)-UGI.

[0016] In one embodiment, the present invention relates to a method comprising: a) providing; i) a nucleotide sequence comprising a gene having a mutated single base, wherein the gene is flanked by an N4CC nucleotide sequence; ii) a mutated NmeCas9 protein comprising a fused nucleotide deaminase and a binding region targeting the N4CC nucleotide sequence; b) contacting the nucleotide sequence with the mutated NmeCas9 protein under conditions that cause the binding region to attach to the N4CC nucleotide sequence; and c) replacing the mutated single base with a wild-type base using the mutated NmeCas9 protein. In one embodiment, the protein is Nme2Cas9. In one embodiment, the protein further comprises a nuclear localization signaling protein. In one embodiment, the nucleotide deaminase is a cytidine deaminase. In one embodiment, the nucleotide deaminase is an adenosine deaminase. In one embodiment, the protein further comprises a uracil glycosylation inhibitor. In one embodiment, the nuclear localization signaling protein includes, but is not limited to, nucleoplasmic proteins (NLS) and / or SV40 NLS and / or C-myc NLS. In one embodiment, the binding region is a frontal lobe sequence helper motif interaction domain. In one embodiment, the frontal lobe sequence helper motif interaction domain contains the mutation. In one embodiment, the mutation is a D16A mutation. In one embodiment, the mutated NmeCas9 protein further contains CBE4. In one embodiment, the mutated NmeCas9 protein further contains an adapter. In one embodiment, the adapter is a 73aa adapter. In one embodiment, the adapter is a 3xHA tag. In one embodiment, the gene encodes a tyrosinase. In one embodiment, the gene is Fah. In one embodiment, the gene is c-fos.

[0017] In one embodiment, the present invention relates to a method comprising: a) providing; i) a patient comprising a nucleotide sequence containing a gene having a mutated single base, wherein the gene is flanked by an N4CC nucleotide sequence, wherein the mutated gene causes a genetically based medical condition; ii) an adeno-associated virus comprising a mutated NmeCas9 protein, the mutated NmeCas9 protein comprising a fused nucleotide deaminase and a binding region targeting the N4CC nucleotide sequence; b) treating the patient with the adeno-associated virus under conditions in which the mutated NmeCas9 protein is replaced with a wild-type single base, such that the genetically based medical condition does not develop. In one embodiment, the gene encodes a tyrosinase protein. In one embodiment, the genetically based medical condition is tyrosinemia. In one embodiment, the virus is adeno-associated virus 8. In one embodiment, the virus is adeno-associated virus 6. In one embodiment, the protein is Nme2Cas9. In one embodiment, the protein further comprises a nuclear localization signal protein. In one embodiment, the nucleotide deaminase is a cytidine deaminase. In one embodiment, the nucleotide deaminase is adenosine deaminase. In one embodiment, the protein further comprises a uracil glycosylation inhibitor. In one embodiment, the nuclear localization signaling protein includes, but is not limited to, nucleoplasmic proteins (NLS) and / or SV40 NLS and / or C-myc NLS. In one embodiment, the binding region is a pre-intercalary sequence auxiliary motif interaction domain. In one embodiment, the pre-intercalary sequence auxiliary motif interaction domain contains the mutation. In one embodiment, the mutation is a D16A mutation. In one embodiment, the mutated NmeCas9 protein further comprises CBE4. In one embodiment, the mutated NmeCas9 protein further comprises an adapter. In one embodiment, the adapter is a 73aa adapter. In one embodiment, the adapter is a 3xHA tag. In one embodiment, the gene encodes a tyrosinase. In one embodiment, the gene is Fah. In one embodiment, the gene is c-fos.

[0018] In one embodiment, the present invention relates to a method comprising: a) providing; i) a patient comprising a nucleotide sequence containing a gene having a mutated single base, wherein the gene is flanked by an N4CC nucleotide sequence, wherein the mutated gene causes a genetically based medical condition; ii) an optimized nNme2Cas9-ABEmax comprising a mutated NmeCas9 protein, the mutated NmeCas9 protein comprising a fused nucleotide deaminase and a binding region targeting the N4CC nucleotide sequence; b) treating the patient with the optimized nNme2Cas9-ABEmax under conditions in which the mutated NmeCas9 protein is replaced with a wild-type single base, such that the genetically based medical condition does not develop.

[0019] In one embodiment, the present invention relates to a method comprising: a) providing; i) a patient comprising a nucleotide sequence containing a gene having a mutated single base, wherein the gene is flanked by an N4CC nucleotide sequence, wherein the mutated gene causes a genetically based medical condition; ii) nNme2Cas9-CBE4 comprising a mutated NmeCas9 protein, the mutated NmeCas9 protein comprising a fused nucleotide deaminase and a binding region for the N4CC nucleotide sequence; b) treating the patient with the nNme2Cas9-CBE4 under conditions in which the mutated NmeCas9 protein is replaced with a wild-type single base, such that the genetically based medical condition does not develop.

[0020] In one embodiment, the present invention relates to a method comprising: a) providing; i) a patient comprising a nucleotide sequence containing a gene having a mutated single base, wherein the gene is flanked by an N4CC nucleotide sequence, wherein the mutated gene causes a genetically based medical condition; ii) YE1-BE3-nNme2Cas9(D16A)-UGI comprising a mutated NmeCas9 protein, the mutated NmeCas9 protein comprising a fused nucleotide deaminase and a binding region targeting the N4CC nucleotide sequence; b) treating the patient with the nNme2Cas9-CBE4 under conditions in which the mutated NmeCas9 protein is replaced with a wild-type single base, such that the genetically based medical condition does not develop.

[0021] definition

[0022] To facilitate understanding of the invention, several terms are defined below. The terms defined herein have meanings commonly understood by one of ordinary skill in the art relating to this invention. Terms such as “an,” “a,” and “the” are not intended to refer to only a singular entity, but rather to encompass general categories, specific examples of which may be used for illustrative purposes. The terms used herein are used to describe specific embodiments of the invention, but their use does not limit the invention except as set forth in the claims.

[0023] As used herein, the terms “edited,” “edited,” or “edited” refer to a method of altering the sequence of a polynucleotide nucleic acid (e.g., a wild-type naturally occurring nucleic acid sequence or a mutated naturally occurring sequence) by selectively deleting a specific genomic target. Such specific genomic targets include, but are not limited to, chromosomal regions, genes, promoters, open reading frames, or any nucleic acid sequence.

[0024] As used herein, the term "single base" refers to one and only one nucleotide within a nucleic acid sequence. When used in the context of single base editing, it refers to the substitution of a base at a specific position within a nucleic acid sequence with a different base. Such substitution can occur through a number of mechanisms, including but not limited to substitution or modification.

[0025] As used herein, the term "target" or "target site" refers to a pre-identified nucleic acid sequence of any composition and / or length. Such target sites include, but are not limited to, chromosomal regions, genes, promoters, open reading frames, or any nucleic acid sequence. In some embodiments, the present invention queries these specific genomic target sequences using a complementary sequence of gRNA.

[0026] As used herein, the term “mid-target binding sequence” refers to a subsequence of a specific genomic target that is fully complementary to a programmable DNA binding domain and / or a single guide RNA sequence.

[0027] As used herein, the term “off-target binding sequence” refers to a subsequence of a specific genomic target that can be partially complementary to a programmable DNA binding domain and / or a single guide RNA sequence.

[0028] As used herein, the term "effective amount" refers to a specific amount of a pharmaceutical composition containing a therapeutic agent that achieves a clinically beneficial outcome (i.e., relief of symptoms). The toxicity and therapeutic efficacy of such compositions can be determined using standard pharmaceutical procedures in cell culture or laboratory animals, such as those used to determine the LD50. 50 (The dose that causes 50% mortality in the population) and ED 50 (The dose that is effective in 50% of the population). The dose ratio between toxicity and therapeutic effect is the therapeutic index, and it can be expressed as the ratio LD50. 50 / ED 50Compounds exhibiting a high therapeutic index are preferred. Data obtained from these cell culture assays and other animal studies can be used to formulate a range of dosages for human use. The dosage of such compounds is preferably within the range of ED (extra-exposure protons) containing little or no toxicity. 50 The circulating concentration range. The dosage varies within this range depending on the dosage form used, patient sensitivity, and route of administration.

[0029] As used herein, the term "symptom" refers to any subjective or objective evidence of illness or physical discomfort observed by a patient. For example, subjective evidence is typically based on patient self-report and may include, but is not limited to, pain, headache, visual disturbances, nausea, and / or vomiting. Alternatively, objective evidence is typically the result of medical tests, including but not limited to body temperature, complete blood count, lipid panel, thyroid panel, blood pressure, heart rate, electrocardiogram, tissue and / or body imaging scans.

[0030] As used herein, the term "disease" or "medical condition" refers to any disruption or alteration of the normal state of an organism or plant or part thereof that results in the performance of vital functions. It typically manifests as obvious signs and symptoms, usually in response to: i) environmental factors (such as malnutrition, industrial hazards, or climate); ii) specific infectious agents (such as worms, bacteria, or viruses); iii) inherent defects in the organism (such as genetic abnormalities); and / or iv) combinations of these factors.

[0031] When referring to the expression of any symptom in an untreated subject relative to a treated subject, the terms “relief,” “inhibition,” “reduction,” “blockage,” “lowering,” “prevention,” and grammatically equivalent forms (including “lower,” “smaller,” etc.) mean any amount and / or severity of symptoms in a treated subject that is less than that in an untreated subject by any amount deemed clinically relevant by any medically trained person. In one embodiment, the amount and / or severity of symptoms in a treated subject is at least 10%, at least 25%, at least 50%, at least 75%, and / or at least 90% less than that in an untreated subject.

[0032] As used herein, the term "attachment" refers to any interaction between a medium (or carrier) and a drug. Attachment can be reversible or irreversible. Such attachment includes, but is not limited to, covalent bonding, ionic bonding, van der Waals forces, or frictional forces. A drug attaches to a medium (or carrier) if it is impregnated, incorporated, coated in, in suspension with, in solution with, or mixed with the medium (or carrier).

[0033] As used herein, the terms "drug" or "compound" refer to any pharmacologically active substance that can be administered to achieve a desired effect. Drugs or compounds can be synthetic or naturally occurring non-peptides, proteins or peptides, oligonucleotides or nucleotides, polysaccharides or sugars.

[0034] As used herein, the terms “application” or “administration” refer to any method of providing a composition to a patient so that the composition has the intended effect on the patient. Exemplary methods of application include direct mechanisms such as local tissue application (i.e., extravascular placement), oral administration, transdermal patches, topical application, inhalation, suppositories, etc.

[0035] As used herein, the term "patient" or "subject" refers to a person or animal and does not necessarily require hospitalization. For example, an outpatient or a person in a nursing home is a "patient." Patients can include humans or non-human animals of any age, and therefore include adults and minors (i.e., children). The term "patient" does not imply a need for medical treatment; therefore, patients may participate in experiments, whether voluntarily or involuntarily, whether for clinical or basic scientific research.

[0036] As used herein, the term "affinity" refers to any attractive force between substances or particles that causes them to enter and maintain a chemical combination. For example, inhibitory compounds with high affinity for receptors will provide greater efficacy in preventing the receptor from interacting with its natural ligand compared to inhibitors with low affinity.

[0037] As used herein, the terms “pharmaceuticalally acceptable” or “pharmaceutically acceptable” refer to molecular entities and compositions that do not produce adverse, allergic, or other adverse reactions when administered to animals or humans.

[0038] As used herein, the term "pharmaceutically acceptable carrier" includes any and all solvents or dispersion media, including but not limited to water, ethanol, polyols (e.g., glycerol, propylene glycol, and liquid polyethylene glycol), suitable mixtures thereof, as well as vegetable oils, coatings, isotonic and absorption-delaying agents, liposomes, commercially available detergents, etc. Additional bioactive ingredients may also be incorporated into such carriers.

[0039] The term "viral vector" encompasses any nucleic acid construct derived from a viral genome capable of incorporating a heterologous nucleic acid sequence for expression in a host organism. Such viral vectors can include, but are not limited to, adeno-associated virus vectors, lentiviral vectors, SV40 viral vectors, retroviral vectors, and adenovirus vectors. Although viral vectors are sometimes derived from pathogenic viruses, they can be modified in a manner that minimizes their overall health risk. This typically involves the deletion of a portion of the viral genome involved in viral replication. Such viruses can effectively infect cells, but after infection, the virus may require helper viruses to provide the missing proteins for the production of new virions. Preferably, the viral vector should have minimal impact on the physiology of the cells it infects and exhibit genetically stable properties (e.g., not undergoing spontaneous genomic rearrangements). Most viral vectors are engineered to infect as many cell types as possible. Even so, viral receptors can be modified to target the virus to specific cell species. Viruses modified in this way are called pseudoviruses. Viral vectors are often engineered to incorporate certain genes that help identify cells that have taken up the viral genome. These genes are called marker genes. For example, commonly used marker genes confuse antibiotic resistance to a particular antibiotic.

[0040] As used herein, “ROSA26 gene” or “Rosa26 gene” refers to a human or mouse (respectively) locus widely used for general expression in mice. Targeting of the ROSA26 locus can be achieved by inserting the desired gene into the first intron of the locus at a unique XbaI site approximately 248 bp upstream of the original gene capture line. Constructs can be built using an adenovirus splice acceptor followed by the gene of interest inserted at the unique XbaI site and a polyadenylation site. Neomycin resistance cassettes can also be included in the targeting vector.

[0041] As used herein, the "PCSK9 gene" or "Pcsk9 gene" refers to the human or mouse locus (respectively) that encodes the PCSK9 protein. The PCSK9 gene is located on chromosome 1 at band 1p32.3 and contains 13 exons. This gene can produce at least two isoforms through alternative splicing.

[0042] The terms “proprotein convertase subtilisin / kexin 9” and “PCSK9” refer to proteins encoded by genes that regulate low-density lipoprotein levels. Proprotein convertase subtilisin / kexin 9, also known as PCSK9, is an enzyme encoded by the PCSK9 gene in humans. Seidah et al., “The secretory proprotein convertase neural apoptosis-regulated convertase 1 (NARC-1): liver regeneration and neuronal differentiation” Proc. Natl. Acad. Sci. USA 100(3):928-933 (2003). Similar genes (orthologous genes) have been found in many species. Many enzymes, including PCSK9, are inactive when first synthesized because they have a portion of the peptide chain that blocks their activity; proprotein convertase removes this portion to activate the enzyme. PCSK9 is believed to play a regulatory role in cholesterol homeostasis. For example, PCSK9 can bind to the epidermal growth factor-like repeat A (EGF-A) domain of the low-density lipoprotein receptor (LDL-R), leading to LDL-R internalization and degradation. Clearly, the expected reduction in LDL-R levels results in decreased LDL-C metabolism, which leads to hypercholesterolemia.

[0043] As used herein, the term "hypercholesterolemia" refers to any medical condition in which blood cholesterol levels are elevated above clinically recommended levels. For example, if cholesterol is measured using low-density lipoprotein (LDL), hypercholesterolemia may be present if the measured LDL level is higher than, for example, approximately 70 mg / dL. Alternatively, if cholesterol is measured using free plasma cholesterol, hypercholesterolemia may be present if the measured free cholesterol level is higher than, for example, approximately 200–220 mg / dL.

[0044] As used herein, the term “CRISPR” or “clustered, regularly spaced short palindromic repeats” is an acronym for a DNA locus containing multiple short, direct repeating sequences of bases. Each repeat sequence contains a sequence of bases followed by approximately 30 base pairs, called “spacer DNA.” Spacers are short fragments of viral DNA and can serve as a “memory” of past exposures to facilitate adaptive defenses against future invasions.

[0045] As used in this article, the term "Cas" or "CRISPR-associated (cas)" refers to genes that are typically associated with CRISPR repeat spacer arrays.

[0046] As used herein, the term "Cas9" refers to a nuclease from the type II CRISPR system, specifically designed to create double-strand breaks in DNA. It has two active cleavage sites (HNH and RuvC domains), each targeting one strand of the double helix. Jinek combined tracrRNA and spacer RNA into a "single guide RNA" (sgRNA) molecule, which, when mixed with Cas9, locates and cleaves the DNA target via Watson-Crick pairing between the guide sequence within the sgRNA and the target DNA sequence.

[0047] As used herein, the term "pre-intermediate sequence neighbor motif" (or PAM) refers to the DNA sequence that Cas9 / sgRNA may require to form an R loop to query a specific DNA sequence via its guide RNA to Watson-Crick pairing with the genome. PAM specificity can be a function of the DNA-binding specificity of Cas9 proteins (e.g., the "pre-intermediate sequence neighbor motif recognition domain" at the C-terminus of Cas9).

[0048] As used herein, the term "sgRNA" refers to a single guide RNA used in conjunction with the CRISPR-associated system (Cas). sgRNA is a fusion of crRNA and tracrRNA and contains a nucleotide sequence complementary to the desired target site. (Jinek et al., "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity" Science 337(6096):816-821 (2012).) The Watson-Crick pairing of sgRNA with the target site allows R-loop formation, which, when combined with a functional PAM, allows DNA cleavage, or, in the case of nuclease deficiency, Cas9 allows binding to DNA at that locus.

[0049] As used herein, the term "fluorescent protein" refers to a protein domain comprising at least one organic compound moiety that fluoresces in response to an appropriate wavelength. For example, fluorescent proteins can emit red, blue, and / or green light. Such proteins are readily available commercially and include, but are not limited to: i) mCherry (Clonetech Laboratories): excitation: 556 / 20 nm (wavelength / bandwidth); emission: 630 / 91 nm; ii) sfGFP (Invitrogen): excitation: 470 / 28 nm; emission: 512 / 23 nm; iii) TagBFP (Evrogen): excitation: 387 / 11 nm; emission: 464 / 23 nm.

[0050] As used herein, the term "sgRNA" refers to a single guide RNA used in conjunction with the CRISPR-associated system (Cas). The sgRNA contains a nucleotide sequence complementary to the desired target site. The sgRNA pairs with the Watson-Crick locus at the target site to recruit nuclease-deficient Cas9 to bind DNA at that locus.

[0051] As used herein, the term "orthogonal" refers to targets that do not overlap, are unrelated, or are independent. For example, if two orthogonal nuclease-deficient Cas9 genes fused to different effector domains are implemented, the sgRNAs encoded for each will not interfere with or overlap with each other. Not all nuclease-deficient Cas9 genes function equally, which allows the use of orthogonal nuclease-deficient Cas9 genes fused to different effector domains to provide appropriate orthogonal sgRNAs.

[0052] As used herein, the term “phenotypic change” or “phenotype” refers to a combination of observable characteristics or traits of an organism (e.g., its morphology, development, biochemical or physiological properties, phenology, behavior, and behavioral products). Phenotypes are caused by the expression of an organism's genes, the influence of environmental factors, and the interactions between the two.

[0053] As used herein, “nucleic acid sequence” and “nucleotide sequence” refer to oligonucleotides or polynucleotides and their fragments or portions, as well as DNA or RNA of genomic or synthetic origin that may be single-stranded or double-stranded and represent a sense or antisense strand.

[0054] As used herein, the term "isolated nucleic acid" refers to any nucleic acid molecule that has been removed from its natural state (e.g., removed from a cell and, in a preferred embodiment, does not contain other genomic nucleic acids).

[0055] As used herein, the terms “amino acid sequence” and “polypeptide sequence” are used interchangeably and refer to the amino acid sequence.

[0056] As used herein, the term "part" when referring to a protein (such as in "a part of a given protein") refers to a fragment of the protein. The size of a fragment can range from four amino acid residues to the entire amino acid sequence minus one amino acid.

[0057] When used in relation to nucleotide sequences, the term "fraction" refers to a segment of that nucleotide sequence. The size of a fragment can range from 5 nucleotide residues to the entire nucleotide sequence minus one nucleic acid residue.

[0058] As used herein, the terms “complementary” or “complementarity” refer to “polynucleotides” and “oligonucleotides” (which are interchangeable terms referring to sequences of nucleotides) as described by base pairing rules. For example, the sequence “CAGT” is complementary to the sequence “GTCA”. Complementarity can be “partial” or “complete.” “Partial” complementarity means that one or more nucleic acid bases do not match according to base pairing rules. “Complete” or “full” complementarity between nucleic acids means that each nucleic acid base matches the other base under base pairing rules. The degree of complementarity between nucleic acid chains has a significant impact on the hybridization efficiency and strength between nucleic acid chains. This is particularly important in amplification reactions and detection methods that rely on the binding between nucleic acids.

[0059] As used herein, the terms “homology” and “homogeneous” in relation to nucleotide sequences refer to the degree of complementarity with other nucleotide sequences. Partial or complete homology (i.e., identity) can exist. A nucleotide sequence that is partially complementary to a nucleic acid sequence, i.e., “substantially homologous,” is a nucleotide sequence that at least partially inhibits hybridization between a completely complementary sequence and a target nucleic acid sequence. Inhibition of hybridization between a completely complementary sequence and a target sequence can be examined using hybridization assays (Southern or Northern blotting, solution hybridization, etc.) under low-tightness conditions. Substantially homologous sequences or probes will compete with and inhibit the binding (i.e., hybridization) of completely homologous sequences to the target sequence under low-tightness conditions. This does not mean that low-tightness conditions allow for nonspecific binding; low-tightness conditions require that the binding of two sequences to each other be a specific (i.e., selective) interaction. The presence of nonspecific binding can be tested by using a second target sequence that lacks or even partially complementarizes the sequence (e.g., less than about 30% identity); in the absence of nonspecific binding, the probe will not hybridize with the second non-complementary target.

[0060] As used herein, the terms “homology” and “homogeneous” regarding amino acid sequences refer to the degree of similarity in the primary structure between two amino acid sequences. This degree of similarity can apply to a portion of each amino acid sequence or to the entire length of the amino acid sequence. Two or more amino acid sequences that are “substantially homologous” can have at least 50% similarity, preferably at least 75%, more preferably at least 85%, and most preferably at least 95% or 100%.

[0061] In this paper, an oligonucleotide sequence as a “homology” is defined as an oligonucleotide sequence that exhibits greater than or equal to 50% identity with a sequence having a length of 100 bp or longer.

[0062] Low stringency conditions include conditions equivalent to binding or hybridization at 42°C in a solution of 5x SSPE (43.8 g / L NaCl, 6.9 g / L NaH2PO4·H2O and 1.85 g / L EDTA, pH adjusted to 7.4 with NaOH), 0.1% SDS, 5x Denhardt reagent (each 500 ml of 50x Denhardt contains: 5 g Ficoll (Type 400, Pharmacia), 5 g BSA (Fractional V; Sigma)) and 100 μg / ml denatured salmon sperm DNA, followed by washing at 42°C in a solution containing 5x SSPE and 0.1% SDS. Many equivalent conditions can also be used to construct low-tightness hybridization conditions; factors such as the length and properties of the probe (DNA, RNA, base composition) and the properties of the target (DNA, RNA, base composition, presence in solution or immobilization, etc.), as well as the concentrations of salts and other components (e.g., the presence or absence of formamide, dextran sulfate, polyethylene glycol), and the composition of the hybridization solution can be modified to produce equivalent low-tightness hybridization conditions that differ from those listed above. Additionally, conditions that promote hybridization under high-tightness conditions can be used (e.g., increasing the temperature of the hybridization and / or washing steps, using formamide in the hybridization solution, etc.).

[0063] As used herein, the term "hybridization" refers to the pairing of complementary nucleic acids in any process in which nucleic acid chains bind to complementary chains through base pairing to form a hybridization complex. Hybridization and hybridization strength (i.e., the strength of association between nucleic acids) are influenced by factors such as the degree of complementarity between nucleic acids, the strictness of the conditions involved, and the TT of the resulting heterozygote. m And the G:C ratio within nucleic acids.

[0064] As used herein, the term "hybridization complex" refers to a complex formed between two nucleic acid sequences due to the formation of hydrogen bonds between complementary G and C bases and between complementary A and T bases; these hydrogen bonds can be further stabilized by base stacking interactions. The two complementary nucleic acid sequences are hydrogen-bonded in an antiparallel configuration. Hybridization complexes can form in solution (e.g., COt or R0t analysis) or between one nucleic acid sequence present in solution and another nucleic acid sequence immobilized on a solid support (e.g., Southern and Northern blots, nylon or nitrocellulose membranes used in dot blots, or slides used in in situ hybridization, including FISH (fluorescence in situ hybridization)).

[0065] DNA molecules are described as having both "5' ends" and "3' ends" because mononucleotides react to form oligonucleotides by attaching the 5' phosphate of a mononucleotide's pentose ring to its neighbor's 3' oxygen via a phosphodiester bond. Therefore, if the 5' phosphate of an oligonucleotide is not attached to the 3' oxygen of a mononucleotide's pentose ring, the end of that oligonucleotide is called a "5' end." If the 3' oxygen of an oligonucleotide is not attached to the 5' phosphate of another mononucleotide's pentose ring, the end of that oligonucleotide is called a "3' end." As used herein, even within larger oligonucleotides, nucleic acid sequences can be considered to have both 5' and 3' ends. In linear or circular DNA molecules, discrete elements are referred to as "upstream" or 5' or "downstream" or 3' elements. This terminology reflects the fact that transcription proceeds along the DNA strand in a 5' to 3' manner. Promoter and enhancer elements that direct transcription of linked genes are typically located at the 5' or upstream of the coding region. However, enhancer elements can also function even if located at the 3' of promoter elements and coding regions. Transcription termination and polyadenylation signals are located at 3' or downstream of the coding region.

[0066] The term "transfection" or "transfected" refers to the introduction of foreign DNA into cells.

[0067] As used herein, the terms “nucleic acid molecule coding,” “DNA sequence coding,” and “DNA coding” refer to the sequence or order of deoxyribonucleotides along a chain of deoxyribonucleic acid (DNA). The sequence of these deoxyribonucleotides determines the sequence of amino acids along a polypeptide (protein) chain. Therefore, a DNA sequence codes for an amino acid sequence.

[0068] As used herein, the term "gene" refers to a deoxyribonucleotide sequence containing the coding region of a structural gene and including sequences adjacent to the coding region at approximately 1 kb distance from either the 5' or 3' ends, such that the gene corresponds to the length of full-length mRNA. A sequence located at the 5' end of the coding region and present on the mRNA is called a 5' untranslated sequence. A sequence located at or downstream of the coding region and present on the mRNA is called a 3' untranslated sequence. The term "gene" encompasses both the cDNA and genomic forms of a gene. The genomic form or clone of a gene contains coding regions broken up by non-coding sequences, referred to as "introns," "insertion regions," or "insertion sequences." Introns are segments of genes transcribed into heterogeneous nuclear RNA (hnRNA); introns may contain regulatory elements, such as enhancers. Introns are removed or "spliced ​​out" from nuclear transcripts or primary transcripts; therefore, introns are absent in messenger RNA (mRNA) transcripts. mRNA plays a role in translation to specify the sequence or order of amino acids in the nascent polypeptide.

[0069] In addition to introns, the genomic form of a gene may also contain sequences located at the 5' and 3' ends of sequences present on the RNA transcript. These sequences are called "flanking" sequences or regions (these flanking sequences are located at the 5' or 3' ends of untranslated sequences present on the mRNA transcript). The 5' flanking region may contain regulatory sequences that control or influence gene transcription, such as promoters and enhancers. The 3' flanking region may contain sequences that guide transcription termination, post-transcriptional cleavage, and polyadenylation.

[0070] As used herein, the term "label" or "detectable label" refers to any composition that can be detected by spectroscopic, photochemical, biochemical, immunochemical, electrical, optical, or chemical means. Such labels include biotin used for staining with labeled streptavidin conjugates, magnetic beads (e.g.), and more. Fluorescent dyes (e.g., fluorescein, Texas red, rhodamine, green fluorescent protein, etc.), radioactive labels (e.g.) 3 H, 125 I, 35 S, 14 C or 32 P), enzymes (e.g., horseradish peroxidase, alkaline phosphatase, and other enzymes commonly used in ELISA), and calorimetric labels, such as colloidal gold or colored glass or plastic (e.g., polystyrene, polypropylene, latex, etc.) beads. Patents teaching the use of such labels include, but are not limited to, U.S. Patent Nos. 3,817,837; 3,850,752; 3,939,350; 3,996,345; 4,277,437; 4,275,149; and 4,366,241 (all incorporated herein by reference in their entirety). The labels contemplated in this invention can be detected by a variety of methods. For example, radioactive labels can be detected using film or a scintillation counter, and fluorescent labels can be detected using a photodetector to detect emitted light. Enzyme labels are typically detected by providing a substrate to an enzyme and detecting the reaction product produced by the enzyme's action on the substrate, and calorimetric labels are detected by simply visualizing colored labels.

[0071] Brief description of the attached figures

[0072] The patent or application documents contain at least one color drawing. The Patent Office will, upon request and after payment of the necessary fees, provide a copy of the patent or application publication with the color drawing.

[0073] Figure 1 An exemplary schematic implementation of a single-base editor for an NmeCas9 deaminase fusion protein and an exemplary construction plasmid for the base editor are shown.

[0074] Figure 1 A shows an exemplary YE1-BE3-nNme2Cas9(D16A)-UGI construct.

[0075] Figure 1 B shows an exemplary ABE7.10 nNme2Cas9(D16A) construct.

[0076] Figure 1 C shows an exemplary ABE7.10-nNme2Cas9(D16A) construct containing two SV40 NLS sequences.

[0077] Figure 1 D shows an exemplary nNme2Cas9-CBE4 (also known as BE4-nNme2Cas9(D16A)-UGI-UGI) construct.

[0078] Figure 1 E shows an exemplary optimized nNme2Cas9-ABEmax builder.

[0079] Figure 2 Exemplary data on electroporation of HEK293T cells using a DNA plasmid containing the YE1-BE3-nNme2Cas9(D16A)-UGI fusion protein, which efficiently converts C to T at endogenous target site 25 (TS25) via nuclear transfection in HEK293T cells.

[0080] Figure 2 A shows an exemplary sequence of the TS25 endogenous target site (within the black rectangle). The GN23 sgRNA pairs with the target DNA strand, leaving a substituted DNA strand for editing by cytidine deaminase (e.g., new green nucleotides).

[0081] Figure 2 B shows exemplary sequencing data that shows a bimodal nucleotide peak (the 7th position from the 5' end; arrow), indicating successful single-base editing from cytidine to thymidine (e.g., converting a C·G base pair to a T·A base pair).

[0082] Figure 2 C shows Figure 2 The exemplary quantification of the data shown in B plots the percentage of C→T single-base editing conversion. In samples treated with both the base editor and sgRNA, the percentage of C to T conversion was approximately 40% (p = 6.88 x 10⁻⁶). The “no sgRNA” control shows background noise due to Sanger sequencing. Analysis was performed using EditR (Kluesner et al., 2018).

[0083] Figure 3Exemplary specific UGI target sites are presented, which are integrated into the YE1-BE3-nNme2Cas9 / D16A mutant fusion protein and co-expressed with enhanced green fluorescent protein (EGFP) in a stable K562-derived cell line. Transformed bases are highlighted in orange. Background signal was filtered using a negative control sample (K562 cells transfected with YE1-BE3-nNme2Cas9 without the sgRNA construct). N4CC PAM is boxed. The percentage of total reads showing mutations at the sites targeted by the base editor is shown in the right column.

[0084] Figure 3 A shows an exemplary EGFP-site 1.

[0085] Figure 3 B shows an exemplary EGFP-site 2.

[0086] Figure 3 C shows an exemplary EGFP-site 3.

[0087] Figure 3 D shows an exemplary EGFP-site 4.

[0088] Figure 3 E shows an exemplary deep sequencing analysis demonstrating the location in the endogenous c-fos promoter region where YE1-BE3-nNme2Cas9 converts C residues to T residues. The right column shows the percentage of total reads exhibiting mutations at the sites targeted by the base editor. Converted bases are highlighted in orange or yellow. Background signal was filtered using a negative control sample. The highest editing percentage was 32.50%.

[0089] Figure 3 F shows an exemplary deep sequencing analysis demonstrating that ABE7.10-nNme2Cas9 or ABEmax (Koblan et al., 2018)-nNme2Cas9 converts A residues to G residues in the endogenous c-fos promoter region. The right column shows the percentage of total reads exhibiting mutations at the sites targeted by the base editor. Converted bases are highlighted in orange. Background signal was filtered using a negative control sample. The editing percentage for ABE7.10-nNme2Cas9 was 0.53%, and for ABEmax-nNme2Cas9 (D16A) it was 2.33%.

[0090] Figure 4An exemplary alignment of the wild-type Fah gene with the tyrosinemia Fah mutant gene is presented, showing the AG single-base gene editing target site (position 9). The corresponding SpyCas9 single PAM site and NmeCas9 double PAM site are indicated to demonstrate the suboptimal targeting window relative to the SpyCas9 PAM site.

[0091] Figure 5 Three exemplary compactly related Neisseria meningitidis Cas9 orthologs with different PAMs are shown.

[0092] Figure 5 A shows an exemplary schematic diagram that illustrates the mutated residues (orange balls) between Nme2Cas9 (left) and Nme3Cas9 (right) on the predicted structure mapped to Nme1Cas9, revealing the mutation clusters (black) in the PID.

[0093] Figure 5 B shows an exemplary experimental workflow for in vitro PAM discovery assays using 10-bp random PAM regions. After in vitro digestion, the adaptor is ligated to the cleavage product for library construction and sequencing.

[0094] Figure 5 C shows an exemplary sequence identifier generated by in vitro PAM discovery that reveals enrichment of N4GATTPAM for Nme1Cas9, consistent with its previously established specificity.

[0095] Figure 5 D shows an exemplary sequence identifier indicating that the Nme1Cas9 whose PID is exchanged with the PID of Nme2Cas9 (left) or Nme3Cas9 (right) requires C at PAM position 5. Due to the low cleavage efficiency of the protein chimera with PID exchange, the remaining nucleotides were not determined with high confidence (see [link to documentation]). Figure 6 C).

[0096] Figure 5 E shows an exemplary sequence identifier that demonstrates full-length Nme2Cas9 identification of N4CC PAM, based on efficient substrate cutting of a target library with a fixed C at PAM position 5 and randomized PAM nt 1-4 and 6-8.

[0097] Figure 6 Presented as with Figure 5 Characterization of related Neisseria meningitidis Cas9 orthologs with rapidly developing PID.

[0098] Figure 6A shows an exemplary rootless phylogenetic tree of NmeCas9 orthologs with >80% identity to Nme1Cas9. Three distinct branches emerge, with most mutations clustered in PIDs. Groups 1 (blue), 2 (orange), and 3 (green) have PIDs with >98%, ~52%, and ~86% identity to Nme1Cas9, respectively. Three representative Cas9 orthologs (one per group) are indicated (Nme1Cas9, Nme2Cas9, and Nme3Cas9).

[0099] Figure 6 B shows an exemplary schematic diagram illustrating the CRISPR-Cas loci encoded by strains from (A) that encode three Cas9 orthologs (Nme1Cas9, Nme2Cas9, and Nme3Cas9). The percentage of identity of each CRISPR-Cas component with *Neisseria meningitidis* 8013 (encoding Nme1Cas9) is shown. Blue and red arrows indicate the transcription start sites for pre-crRNA and tracrRNA, respectively.

[0100] Figure 6 C shows exemplary normalized read counts (percentage of total reads) of in vitro assays of cleaved DNA for full-length Nme1Cas9 (grey), for chimeras where the PID of Nme1Cas9 is exchanged with the PIDs of Nme2Cas9 and Nme3Cas9 (mixed colors), and for full-length Nme2Cas9 (orange). Reduced normalized read counts indicate lower cleavage efficiency in chimeras.

[0101] Figure 6 D shows an exemplary sequence identifier for in vitro PAM discovery assays performed on the NNNNNCNNN PAM library by exchanging the PID of Nme1Cas9 with the PID of Nme2Cas9 (left) or Nme3Cas9 (right).

[0102] Figure 7 Exemplary data are presented showing that Nme2Cas9 edited sites adjacent to N4CC PAM using 22-24nt spacers. All experiments were performed in triplicate, and error bars represent the standard error (sem) of the mean.

[0103] Figure 7 A shows an exemplary schematic diagram depicting transient transfection and editing of HEK293T TLR2.0 cells, with mCherry+ cells detected by flow cytometry 72 hours post-transfection.

[0104] Figure 7B shows an exemplary Nme2Cas9 editing of the TLR2.0 reporter factor. Sites with N4CC PAM were targeted with varying efficiencies, while Nme2Cas9 targeting was not observed on N4GATT PAM or in the absence of sgRNA. SpyCas9 (targeting previously validated sites with NGG PAM) and Nme1Cas9 (targeting N4GATT) were used as positive controls.

[0105] Figure 7 C illustrates the exemplary effect of spacer length on Nme2Cas9 editing efficiency. sgRNAs targeting a single TLR2.0 site with spacer lengths of 24–20 nt (including the 5' G required for the U6 promoter) demonstrate that using a 22–24 nt spacer yields the highest editing efficiency.

[0106] Figure 7 D shows an exemplary Nme2Cas9 dual-nicking enzyme that can be used in tandem to generate NHEJ and HDR-based edits in TLR2.0. Plasmids expressing Nme2Cas9 and sgRNA, along with an 800 bp dsDNA donor for homology repair, were electroporated into HEK293T TLR2.0 cells, and the results for NHEJ (mCherry+) and HDR (GFP+) were scored by flow cytometry. (HNH nicking enzyme, Nme2Cas9) D16A RuvC nickase, Nme2Cas9 H588A Use any cleavage enzyme to target and separate the 32bp and 64bp cleavage sites. HNH cleavage enzyme (Nme2Cas9) D16A This produces effective editing, especially for cleavage sites 32 bp apart, while the RuvC cleavage enzyme (Nme2Cas9) produces effective editing. H588A Invalid. Wild-type Nme2Cas9 was used as a control.

[0107] Figure 8 Exemplary data are presented, showing how in mammalian cells, as with Figure 7 The relevant Nme2Cas9-targeted PAM, spacer, and seed requirements were determined. All experiments were performed in triplicate, and error bars represent sample sizes (sem).

[0108] Figure 8 A shows N4C in TLR2.0 D Exemplary Nme2Cas9 targeting at the site, where editing is estimated based on mCherry+ cells. The position of each non-C nucleotide at the test site (N4C) was examined. A N4C T and N4C G Four sites were identified, with the N4CC site used as a positive control.

[0109] Figure 8 B shows N4 in TLR2.0 D An exemplary Nme2Cas9 targeting at site C [similar to (A)].

[0110] Figure 8 C shows the TLR2.0 site with N4CCA PAM (and... Figure 2 The exemplary truncation on (different in C) shows similar length requirements as those observed at other sites.

[0111] Figure 8 D shows the exemplary sensitivity of Nme2Cas9 targeting efficiency to single nucleotide mismatch differences in the seed region of sgRNA. The data demonstrate the effect of single nucleotide sgRNA mismatches traveling along the 23-nt spacer at the TLR2.0 target site.

[0112] Figure 9 Exemplary data are presented demonstrating Nme2Cas9 genome editing at endogenous loci in mammalian cells via various delivery methods. All results represent three independent biological replicates, and error bars represent sem.

[0113] Figure 9 A illustrates exemplary Nme2Cas9 genome editing at endogenous human sites in HEK293T cells following transient transfection with plasmids expressing Nme2Cas9 and sgRNA. Forty sites were initially screened (Table 1); then, 14 sites were reanalyzed in triplicate (selected to include representatives of different editing efficiencies, as measured by TIDE). An Nme1Cas9 target site (with N4GATT PAM) was used as a negative control.

[0114] Figure 9 B shows exemplary data graphs: Left panel: Transient transfection of a single plasmid expressing Nme2Cas9 and sgRNA (targeting the Pcsk9 and Rosa26 loci) enabled editing in Hepa1-6 mouse cells, as detected by TIDE. Right panel: Electroporation of the sgRNA plasmid into K562 cells stably expressing Nme2Cas9 from a lentiviral vector resulted in efficient insertion / deletion formation.

[0115] Figure 9 C illustrates how the exemplary Nme2Cas9 can be electroporated as an RNP complex to induce genome editing. 40 picomols of Cas9 and 50 picomols of in vitro transcribed sgRNA targeting three different loci were electroporated into HEK293T cells. Insertions and deletions were measured using TIDE after 72 hours.

[0116] Figure 10 Example data is presented, showing how it is compared to... Figure 9 The related dose dependence and segmental deletion of Nme2Cas9.

[0117] Figure 10 A shows an exemplary increase in the dosage of the electroporated Nme2Cas9 plasmid (500 ng relative to the target). Figure 3 The 200 ng in A improved editing efficiency at two sites (TS16 and TS6). Data provided in yellow... Figure 9 A can be reused.

[0118] Figure 10 B illustrates an example of Nme2Cas9 used to generate precise segment deletions. Two TLR2.0 targets with cleavage sites 32 bp apart were simultaneously targeted by Nme2Cas9. Most of the resulting lesions were deletions of exactly 32 bp (blue).

[0119] Figure 11 Exemplary data are presented showing that Nme2Cas9 is inhibited in vitro and in cells by a subset of type II-C anti-CRISPR family. All experiments were performed in triplicate, and error bars represent sem.

[0120] Figure 11 Figure A shows an exemplary in vitro cleavage assay of Nme1Cas9 and Nme2Cas9 in the presence of five previously characterized anti-CRISPR proteins (Acr:Cas9 ratio of 10:1). Top panel: Nme1Cas9 efficiently cleaves fragments containing the pre-interstitial region sequence with N4GATT PAM in the absence of Acr or in the presence of the negative control Acr (AcrE2). As expected, all five previously characterized type II-C Acr family proteins inhibit Nme1Cas9. Bottom panel: Nme2Cas9 inhibition is the same as Nme1Cas9, but lacks the inhibition via AcrIIC5. Smu The inhibition.

[0121] Figure 11 B shows an exemplary genome editing in the presence of five previously described anti-CRISPR families. Plasmids expressing Nme2Cas9 (200 ng), sgRNA (100 ng), and each corresponding Acr (200 ng) were co-transfected into HEK293T cells, and genome editing was measured 72 hours post-transfection using traced insertions and deletions by degradation (TIDE). Consistent with our in vitro analyses, except for AcrIIC5… SmuIn addition, all II-C anti-CRISPR agents inhibit genome editing, although with varying efficiencies.

[0122] Figure 11 C shows that the exemplary Acr inhibition of Nme2Cas9 is dose-dependent and has significant apparent power. Nme2Cas9 was inhibited by AcrIIC1 at mass ratios of 2:1 and 1:1 of co-transfected Acr and Nme2Cas9 plasmids. Nme and AcrIIC4 Hpa Complete suppression.

[0123] Figure 12 Example data is presented, showing how it is compared to... Figure 11 The relevant Nme2Cas9 PID swap enables Nme1Cas9 to AcrIIC5 Smu Inhibition is insensitive. In vitro cleavage of the Nme1Cas9-Nme2Cas9PID chimera occurs in the presence of previously characterized Acr protein (10 uM Cas9-sgRNA + 100 uM Acr).

[0124] Figure 13 Example data is presented, showing how it is compared to... Figure 12 The orthogonality and relative accuracy of Nme2Cas9 and SpyCas9 at dual target sites.

[0125] Figure 13 A shows that the exemplary Nme2Cas9 and SpyCas9 guides are orthogonal. TIDE results show the frequency of insertions and deletions generated by the two nucleases targeting DS2 and their homologous sgRNAs or sgRNAs of other orthologs.

[0126] Figure 13 B shows exemplary Nme2Cas9 and SpyCas9, demonstrating comparable mid-target editing efficiencies, as evaluated by GUIDE-seq. The bar graph represents the mid-target readouts from GUIDE-Seq at the three dual sites for each orthologous target. Orange bars represent Nme2Cas9, and black bars represent SpyCas9.

[0127] Figure 13 C shows the target-to-miss reading count for an exemplary SpyCas9 at each site. Orange bars represent target readings, while black bars represent misses.

[0128] Figure 13 D shows the target-to-off-target readings in an exemplary Nme2Cas9 at each site.

[0129] Figure 13The bar chart for E shows exemplary insertion / deletion efficiencies of potential off-target sites predicted by CRISPRSeek (measured by TIDE). Target and off-target site sequences are shown on the left, with PAM regions underlined and sgRNA mismatches and non-shared PAM nucleotides indicated in red.

[0130] Figure 14 Exemplary data are presented, showing that Nme2Cas9 has little or no detectable off-target effects in mammalian cells.

[0131] Figure 14 A shows an exemplary schematic diagram depicting a dual-site (DS) that can be targeted by both SpyCas9 and Nme2Cas9 via their non-overlapping PAMs. The Nme2Cas9 PAM (orange) and SpyCas9 PAM (blue) are highlighted. The 24nt Nme2Cas9 guide sequence is shown in yellow; the corresponding SpyCas9 guide sequence is 4nt shorter at the 5' end.

[0132] Figure 14 B shows exemplary Nme2Cas9 and SpyCas9 that induce insertion deletions at DS. VEGFA (with GN3GN) was selected. 19 Six DS (distribution sites) of the NGGNCC sequence were used for direct comparison of editing by two orthologs. Plasmids expressing each Cas9 (with the same promoter, adapter, tag, and NLS) and its homologous director were transfected into HEK293T cells. Insertion / deletion efficiency was determined by TIDE 72 hours post-transfection. Nme2Cas9 editing was detectable at all six sites and was slightly or significantly more efficient than SpyCas9 at two sites (DS2 and DS6, respectively). SpyCas9 edited four of the six sites (DS1, DS2, DS4, and DS6), with two sites showing significantly higher editing efficiency than Nme2Cas9 (DS1 and DS4). DS2, DS4, and DS6 were selected for GUIDE-Seq analysis because at these sites, Nme2Cas9 showed the same, lower, and higher efficiency compared to SpyCas9, respectively.

[0133] Figure 14 C shows an exemplary Nme2Cas9 genome editing with high precision in human cells. The number of off-target sites detected by GUIDE-Seq at each target site for each nuclease is shown. In addition to the two sites, we also analyzed TS6 in mouse Hepa1-6 cells (due to its high mid-target editing efficiency) as well as Pcsk9 and Rosa26 sites (to measure accuracy in another cell type).

[0134] Figure 14 D shows exemplary targeted deep sequencing used to detect insertions and deletions in edited cells, confirming the high Nme2Cas9 accuracy indicated by GUIDE-seq.

[0135] Figure 14 E shows an exemplary sequence of empirically proven off-target sites for the Rosa26 guide, showing the PAM region (underlined), with three mismatches (red) in the CC PAM dinucleotide (bold) and the distal portion of the spacer region PAM.

[0136] Figure 15 Exemplary data are presented demonstrating in vivo editing of the Nme2Cas9 genome via all-in-one AAV delivery.

[0137] Figure 15 Figure A shows an exemplary workflow for delivering AAV8.sgRNA.Nme2Cas9 to lower cholesterol levels in mice via Pcsk9 targeting. Top: Schematic diagram of a multi-purpose AAV vector expressing Nme2Cas9 and sgRNA (individual genomic elements are not drawn to scale). BGH, bovine growth hormone polymer (A) site; HA, epitope tag; NLS, nuclear localization sequence; h, human codon-optimized. Bottom: Tail vein injection of AAV8.sgRNA.Nme2Cas9 (4 x 10⁻⁶) 11 The timeline of GC was analyzed, cholesterol was measured on day 14, and insertion / deletion, histological and cholesterol analysis was performed on day 28 post-injection.

[0138] Figure 15 B shows an exemplary TIDE analysis to measure insertions and deletions in DNA extracted from the livers of mice injected with AAV8.Nme2Cas9+ sgRNAs targeting the Pcsk9 and Rosa26 (control) loci. The insertion / deletion efficiency of GUIDE-seq at individual off-target sites recognized by these two sgRNAs (Rosa26|OT1) was also evaluated by TIDE.

[0139] Figure 15 C shows an exemplary reduction in serum cholesterol levels in mice injected with a guide targeting Pcsk9 compared to a control targeting Rosa26. P-values ​​were calculated using an unpaired two-tailed t-test.

[0140] Figure 16 Example data is presented, showing the relationship with Figure 15 Related Nme2Cas9AAV delivery and edited PCSK9 knockdown and liver histology.

[0141] Figure 16Figure A shows an exemplary Western blot using an anti-PCSK9 antibody, revealing a significantly reduced level of PCSK9 in the livers of mice treated with sgPcsk9 compared to mice treated with sgRosa26. 2 ng of recombinant PCSK9 was used as a migration standard (leftmost lane), and cross-reactivity bands in the liver samples are indicated by asterisks. GAPDH was used as a loading control (below).

[0142] Figure 16 B shows exemplary H&E staining of livers from mice injected with either AAV8.Nme2Cas9+sgRosa26 (left) or AAV8.Nme2Cas9+sgPcsk9 (right) vectors. Scale bar, 25 μm.

[0143] Figure 17 Exemplary data is presented, and its display is consistent with... Figure 16 Related in vitro editing of Tyr in mouse fertilized eggs.

[0144] Figure 17 A shows two exemplary sites in Tyr, each with N4CC PAM, tested for editing in Hepa1-6 cells. The sgTyr2 guide showed higher editing efficiency and was therefore selected for further testing.

[0145] Figure 17 B shows seven exemplary mice that survived postnatal development and each exhibited a coat color phenotype and mid-target editing, as determined by TIDE.

[0146] Figure 17 C shows exemplary insertion / deletion profiles of tail DNA from each mouse from (B) and from unedited C57BL / 6NJ mice, as shown by TIDE analysis. The efficiency of insertions (positive) and deletions (negative) of various sizes is indicated.

[0147] Figure 18 Exemplary data are presented, demonstrating in vitro Nme2Cas9 genome editing delivered via an all-in-one AAV.

[0148] Figure 18 A illustrates an exemplary workflow for in vitro editing of single AAV Nme2Cas9 to generate albino C57BL / 6NJ mice by targeting the Tyr gene. Fertilized eggs were cultured in KSOM containing AAV6.Nme2Cas9:sgTyr for 5–6 hours, washed in M2, and cultured for one day before being transferred to the oviduct of a pseudopregnancy recipient.

[0149] Figure 18 B shows the result via 3x10 9Exemplary albino (left) and silver-gray (chinchilla) or variegated (middle) mice generated by GC, and 3x10 mice from fertilized eggs with AAV6.Nme2Cas9:sgTyr. 8 Silver-gray or mottled mice generated by GC (right).

[0150] Figure 18 C shows an exemplary summary of Nme2Cas9.sgTyr single-AAV ex vivo Tyr editing experiments at two AAV doses.

[0151] Figure 19 Exemplary mCherry reporter factor assays for the activities of nSpCas9-ABEmax and optimized ABEmax-nNme2Cas9(D16A) are shown.

[0152] Figure 19 A shows an example sequence of the ABE-mCherry reporter factor. There is a TAG stop codon in the mCherry coding region. In stable cell lines with reporter factor integration, there is no mCherry signal. If nSpCas9-ABEmax or the optimized ABEmax-nNme2Cas9(D16A) can convert the TAG to a CAG (which encodes Gln), then the mCherry signal will be displayed.

[0153] Figure 19 B shows exemplary mCherry signal illumination due to the active state of SpCas9-ABE or ABEmax-nNme2Cas9(D16A) in specific regions of the mCherry reporter factor. The top panel is the negative control, the middle panel shows mCherry signal illumination in reporter cells treated with nSpCas9-ABEmax, and the bottom panel shows mCherry signal illumination in reporter cells treated with optimized ABEmax-nNme2Cas9(D16A).

[0154] Figure 19 C shows an exemplary FACS quantification of base editing events in mCherry reporter cells transfected with SpCas9-ABE or ABEmax-nNme2Cas9(D16A). N = 6; error bars represent SD. Results are derived from biological replicates performed in technical replicates.

[0155] Figure 20 Exemplary GFP reporter factor assays for the activity of nSpCas9-CBE4 (Addgene#100802) and CBE4-nNme2Cas9(D16A)-UGI-UGI (CBE4 cloned from Addgene#100802) are shown.

[0156] Figure 20 A shows exemplary sequence information for the CBE-GFP reporter factor. A mutation exists in the core region of the fluorophore in the GFP reporter factor line that converts GYG to GHG. Therefore, no GFP signal is displayed. If nSpCas9-CBE4 or CBE4-nNme2Cas9(D16A)-UGI-UGI can convert CAC to TAC / TAT (histidine to tyrosine), then a GFP signal will be displayed.

[0157] Figure 20 B shows an exemplary GFP signal (green) due to nSpCas9-CBE4 or CBE4-nNme2Cas9(D16A)-UGI-UGI being active in a specific region of the GFP reporter factor. The top panel is the negative control. The middle panel shows the bright mCherry signal in reporter cells treated with CBE4-nNme2Cas9(D16A)-UGI-UGI. The bottom panel shows the bright GFP signal in reporter cells treated with CBE4-nNme2Cas9(D16A)-UGI-UGI.

[0158] Figure 20 C shows an exemplary FACS quantification of base editing events in GFP reporter cells transfected with nSpCas9-CBE4 or CBE4-nNme2Cas9(D16A)-UGI-UGI. N = 6; error bars represent SD. Results are from biological replicates performed in technical replicates.

[0159] Figure 21 This illustrates an exemplary cytosine editing performed via CBE4-nNme2Cas9(D16A)-UGI-UGI. The top figure shows the KANK3 target sequence information for Nme2Cas9 (PAM sequences are shown in red) and base editing in the negative control sample. The bottom figure shows the quantification of the substitution rate for each type of base in the CBE4-nNme2Cas9(D16A)-UGI-UGI editing window for the KANK3 target sequence. The sequence table shows the nucleotide frequency at each position. The expected frequency of C-to-T conversions is highlighted in red.

[0160] Figure 22Exemplary cytosine and adenine editing were performed via CBE4-nNme2Cas9(D16A)-UGI-UGI and optimized ABEmax-nNme2Cas9(D16A), respectively. The top figure shows the PLXNB2 target sequence information for Nme2Cas9 (PAM sequences are shown in red) and base editing in the negative control sample. The middle figure shows the quantification of the substitution rate for each type of base in the optimized ABEmax-nNme2Cas9(D16A) editing window for the PLXNB2 target sequence. The sequence table shows the nucleotide frequency at each position. The expected frequency of A-to-G conversions is highlighted in red. The bottom figure shows the quantification of the substitution rate for each type of base in the CBE4-nNme2Cas9(D16A)-UGI-UGI editing window for the PLXNB2 target sequence. The sequence table shows the nucleotide frequency at each position. The expected frequency of C-to-T conversions is highlighted in red. Invention Details

[0162] This invention relates to the field of gene editing. Specifically, it targets single nucleotide base editing. For example, such single nucleotide base editing results in the conversion of C·G base pairs to T·A base pairs. The high accuracy and precision of the single nucleotide base gene editor disclosed herein are achieved through an NmeCas9 nuclease fused to a nucleotide deaminase protein. The compact nature of NmeCas9, coupled to motifs adjacent to a large number of compatible pre-interstitial sequences, allows the Cas9 fusion construct envisioned herein to edit sites that are inaccessible to conventional SpyCas9 base editor platforms.

[0163] A.NmeCas9 single-base editing

[0164] Cas9 is a programmable nuclease that uses guide RNA to generate double-strand breaks at any desired genomic site. This programmability has been used in biomedicine and therapeutics. However, Cas9-induced breaks often lead to imprecise repair via cellular mechanisms, hindering its therapeutic applications in single-base correction and uniform, precise gene knockout. Furthermore, combining Cas9-induced DNA double-strand breaks with repair templates for homology-directed repair (HDR) to correct genetic mutations in post-mitotic cells, such as neurons, is extremely challenging.

[0165] Single nucleotide base editing is a genome editing method in which an inactivated or damaged Cas9 nuclease (e.g., inactivated Cas9 (dCas9) or nicking enzyme Cas9 (nCas9)) is fused with another enzyme capable of editing nucleotides without causing DNA double-strand breaks. To date, two broad classes of Cas9 base editors have been developed: i) cytosine deaminase (editing C·G base pairs to T·A base pairs) SpyCas9 fusion proteins; and ii) adenosine deaminase (editing A·T base pairs to G·C base pairs) SpyCas9. (Liu et al., “Nucleobase editors and uses thereof” US2017 / 0121693; and Lui et al., “Fusions of Cas9 domains and nucleic acid-editing domains” US 2015 / 0166980, both incorporated herein by reference.)

[0166] However, as mentioned above, the SpyCas9 base editing platform cannot be used to target all single-base mutations due to its limited editing window. The editing window is constrained by the requirements for NGG PAM. SpyCas9 is also inherently associated with high off-target effects in genome editing.

[0167] In one embodiment, the present invention relates to a compact and highly accurate deaminase fusion protein of Nme2Cas9 (several species of Neisseria meningitidis). This Nme2Cas9 has 1,082 amino acids, compared to SpyCas9 which has 1,368 amino acids. This Nme2Cas9 ortholog functions efficiently in mammalian cells, recognizing N4CC PAM with inherently high accuracy. (Edraki et al., Mol Cell. (in preparation))

[0168] While the mechanism of the invention need not be understood, it is believed that the compactness and ultra-accuracy of the NmeCas9 base editor target single-base mutations that were previously unachievable by other Cas9 platforms known in the art. Furthermore, it is considered that the NmeCas9 base editor considered herein targets pathogenic mutations that are infeasible through current base editor platforms and has increased base editing accuracy.

[0169] In one embodiment, the present invention relates to a fusion protein comprising Nme2Cas9 and a deaminase protein, exemplary examples including ABE7.10-nNme2Cas9(D16A); optimized nNme2Cas9-ABEmax; nNme2Cas9-CBE4 (equivalent to BE4-nNme2Cas9(D16A)-UGI-UGI); and ABEmax-nNme2Cas9(D16A). See also, Figure 1 A, Figure 1 B. Figure 1 C Figure 1 D and Figure 1 E.

[0170] Figure 1 An exemplary schematic implementation of a single-base editor for an NmeCas9 deaminase fusion protein and an exemplary construction plasmid for the base editor are shown. Figure 1 A shows an exemplary YE1-BE3-nNme2Cas9(D16A)-UGI construct. Figure 1 B shows an exemplary ABE7.10 nNme2Cas9(D16A) construct. Figure 1 C shows an exemplary ABE7.10-nNme2Cas9(D16A) construct. Figure 1 C shows an exemplary ABE7.10-nNme2Cas9(D16A) construct containing two SV40NLS sequences. Figure 1 D shows an exemplary nNme2Cas9-CBE4 (also known as BE4-nNme2Cas9(D16A)-UGI-UGI) construct. Figure 1 E shows an exemplary optimized nNme2Cas9-ABEmax builder.

[0171] In one embodiment, the deaminase protein is Apobec1 (YE1-BE3). Apobec1 is not intended to be limited to a single organism. In one embodiment, Apobec1 is derived from the rat species. (Kim et al., “Increasing the genome-targeting scope and precision of base editing with engineered Cas9-cytidinedeaminase fusions”. Nature Biotechnology 35 (2017). In one embodiment, Nme2Cas9 comprises the nNme2Cas9 D16A mutant. In one embodiment, the fusion protein also comprises a uracil glycosylase inhibitor protein (UGI). In one embodiment, the fusion protein comprises the YE1-BE3-nNme2Cas9(D16A)-UGI construct.

[0172] In one implementation, the YE1-BE3-nNme2Cas9(D16A)-UGI construct has the following sequence: (SEQ ID NO:1)

[0173]

[0174] YE1-BE3 (underlined); connector (bold), nNme2Cas9 (italic), UGI (bold / underlined), SV40 NLS (no format).

[0175] In one embodiment, the present invention relates to a fusion protein comprising an NmeCas9 / ABE7.10 deaminase protein. In one embodiment, the deaminase protein is TadA. In one embodiment, the deaminase protein is TadA 7.10.

[0176] In one embodiment, the ABE7.10-nNme2Cas9(D16A) construct has the following amino acid sequence: (SEQ ID NO:3)

[0177]

[0178]

[0179] TadA (underlined), TadA 7.10 (underlined / bold), adapter (bold italic), nNme2Cas9 (italic), nucleoplasmic protein NLS (no format).

[0180] In one embodiment, the ABEmax-nNme2Cas9(D16A) construct has the following amino acid sequence: (SEQ ID NO:5)

[0181]

[0182] TadA (underlined), TadA*7.10 (underlined / bold), adapter (bold italic), nNme2Cas9 (italic), nucleoplasmic protein NLS (unformatted) and SV40 NLS (bold).

[0183] In one embodiment, the CBE4-nNme2Cas9(D16A)-UGI-UGI construct has the following amino acid sequence: (SEQ ID NO:6)

[0184]

[0185] rApobec1 (underlined), UGI (underlined / bold), connector (bold italic), nNme2Cas9(D16A) (italic), Cmyc-NLS (unformatted) and SV40 NLS (bold).

[0186] In one implementation, the optimized nNme2Cas9-ABEmax construct refers to an optimized version with improved promoter, NLS sequence, and adapter sequence. In some implementations, the optimized nNme2Cas9-ABEmax construct includes C-myc NLS, 12aa adapter, 15aa adapter, SV40 NLS, TadA, TadA*7.10, 48aa adapter, nNme2Cas9, 73aa adapter (3xHA-tag), 15aa adapter, and C-myc NLS from 5' to 3'. In some implementations, the optimized nNme2Cas9-ABEmax construct further includes at least two alternating C-myc NLS and 12aa adapters at the 3' end. In some implementations, the optimized nNme2Cas9-ABEmax construct further includes at least two alternating 15aa adapters and C-myc NLSs at the 5' end. See, for example, [link to relevant documentation]. Figure 1 E.

[0187] In one embodiment, the optimized nNme2Cas9-ABEmax construct has the following amino acid sequence (SEQ ID NO:7):

[0188]

[0189] hTadA7.10 (underlined), hTadA*7.10 (underlined / bold), connector (bold italic), nNme2Cas9 (italic), Cmyc-NLS (no format), SV40-NLS (bold).

[0190] In some implementations, plasmid nSpCas9-ABEmax (Addgene ID: 112095) is used as an experimental control and for molecular cloning. In some implementations, plasmid nSpCas9-CBE4 (Addgene ID: 100802) is used as an experimental control and for molecular cloning.

[0191] Robust single-base editing from C·G base pairs to T·A base pairs was achieved at the endogenous target site (TS25) by electroporation of HEK293T cells using a DNA plasmid containing the YE1-BE3-Nme2Cas9 nucleotide deaminase fusion protein. See also Figure 2 AC.

[0192] Figure 2 Exemplary data on electroporation of HEK293T cells using a DNA plasmid containing the YE1-BE3-nNme2Cas9(D16A)-UGI fusion protein, which efficiently converts C to T at endogenous target site 25 (TS25) via nuclear transfection in HEK293T cells. Figure 2 A shows an exemplary sequence of the TS25 endogenous target site (within the black rectangle). GN23sgRNA pairs with the target DNA strand, leaving a substituted DNA strand for editing by cytidine deaminase (e.g., new green nucleotides). Figure 2 B shows exemplary sequencing data that shows a bimodal nucleotide peak (the 7th position from the 5' end; arrow), indicating successful single-base editing from cytidine to thymidine (e.g., converting a C·G base pair to a T·A base pair). Figure 2 C shows Figure 2 The exemplary quantification of the data shown in B plots the percentage of C→T single-base editing conversion. In samples treated with both the base editor and sgRNA, the percentage of C to T conversion was approximately 40% (p = 6.88 x 10⁻⁶). The “no sgRNA” control shows background noise due to Sanger sequencing. Analysis was performed using EditR (Kluesner et al., 2018).

[0193] In a stable K562-derived cell line expressing enhanced green fluorescent protein (EGFP), four additional YE1-BE3-nNme2Cas9 / D16A mutant fusion proteins were co-expressed with EGFP. Each YE1-BE3-nNme2Cas9 / D16A mutant fusion protein has a specific UGI target site. See also Figure 3 AD.

[0194] Deep sequencing analysis revealed that YE1-BE3-nNme2Cas9 converts C residues to T residues at each of the four EGFP target sites. The percentage of editing ranged from 0.24% to 2%. The potential base editing window originates from nucleotides 2–8 in the substituted DNA strand, with the 5' (PAM distal) end of the nucleotide counted as nucleotide #1. See also Figure 3 AD.

[0195] Figure 3 Exemplary specific UGI target sites are presented, which are integrated into the YE1-BE3-nNme2Cas9 / D16A mutant fusion protein and co-expressed with enhanced green fluorescent protein (EGFP) in a stable K562-derived cell line. Transformed bases are highlighted in orange. Background signal was filtered using a negative control sample (K562 cells transfected with YE1-BE3-nNme2Cas9 without the sgRNA construct). N4CC PAM is boxed. The percentage of total reads showing mutations at the sites targeted by the base editor is shown in the right column. Figure 3 A shows an example EGFP site 1. Figure 3 B shows an example EGFP site 2. Figure 3 C shows an exemplary EGFP site 3. Figure 3 D shows an exemplary EGFP site 4.

[0196] Electroporation of HEK293T cells with a DNA plasmid containing the YE1-BE3-nNme2Cas9 c-fos promoter enabled robust single-base editing of C·G to T·A base pairs at endogenous target sites within the c-fos promoter. Figure 3 E). Figure 3 E shows an exemplary deep sequencing analysis demonstrating the location in the endogenous c-fos promoter region where YE1-BE3-nNme2Cas9 converts C residues to T residues. The right column shows the percentage of total reads exhibiting mutations at the sites targeted by the base editor. Converted bases are highlighted in orange or yellow. Background signal was filtered using a negative control sample. The highest editing percentage was 32.50%. Figure 3F shows an exemplary deep sequencing analysis demonstrating the location in the endogenous c-fos promoter region where ABE7.10-nNme2Cas9 or ABEmax (Koblan et al., 2018)-nNme2Cas9 converts A residues to G residues. The right column shows the percentage of total reads exhibiting mutations at the sites targeted by the base editor. Converted bases are highlighted in orange. Background signal was filtered using a negative control sample. The editing percentage for ABE7.10-nNme2Cas9 was 0.53%, and for ABEmax-nNme2Cas9 (D16A) it was 2.33%.

[0197] In one embodiment, the present invention relates to the expression of an ABE7.10-nNme2Cas9(D16A) fusion protein for base editing. Although the mechanism of the invention need not be understood, it is believed that Nme2Cas9 base editing can be effectively used for the treatment of tyrosinemia by reversing the G to A point mutation in the Fah gene with the ABE7.10-nNme2Cas9(D16A) fusion protein.

[0198] A G-to-A mutation (red) in the last nucleotide of exon 8 in the Fah gene results in exon skipping. FAH deficiency leads to toxin accumulation and severe liver damage. The location of the SpyCas9 PAM (black rectangle) downstream of the mutation is not optimal for sgRNA design because the A mutation is outside the effective base editing window of ABE7.10 (which is at the 4th-7th nt (underlined) end of the 5' (PAM distal end) (Gaudelli et al., 2017)).

[0199] However, there are two Nme2Cas9 PAMs (red rectangles) in the downstream sequence, which can potentially correct the mutation and reverse the DNA sequence to wild-type via ABE7.10-nNme2Cas9(D16A). See also Figure 4 .

[0200] Figure 4 An exemplary alignment of the wild-type Fah gene with the tyrosinemia Fah mutant gene is presented, showing the AG single-base gene editing target site (position 9). The corresponding SpyCas9 single PAM site and NmeCas9 double PAM site are indicated to illustrate the suboptimal targeting window relative to the SpyCas9 PAM site. This figure serves as a potential example of a site where Nme2Cas9 can overcome the limitations of existing base editors. It is further believed that the NmeCas9 base editor described herein can perform precise base editing that conventional SpyCas9-derived base editors cannot achieve due to their suboptimal base editing window relative to nearby available PAMs.

[0201] Furthermore, we envision extending base editing to a tyrosinemia mouse model for reversing G-to-A point mutations using a viral delivery method of ABEmax-nNme2Cas9 (D16A), where the desired edits cannot be achieved using a SpyCas9-derived base editor due to the suboptimal base editing window relative to nearby available PAMs (e.g., Figure 4 ).

[0202] B.NmeCas9 Construct: Compact and Ultra-Accurate

[0203] Clustered, regularly spaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) proteins together constitute bacterial and archaea adaptive immune pathways against bacteriophages and other mobile genetic elements (MGEs) (Barrangou et al., 2007; Brouns et al., 2008; Marraffini and Sontheimer, 2008). In type II CRISPR systems, CRISPR RNA (crRNA) binds to trans-activating crRNA (tracrRNA) and is loaded onto the Cas9 effector protein, which cleaves MGE nucleic acids complementary to crRNA (Garneau et al., 2010; Deltcheva et al., 2011; Sapranauskas et al., 2011; Gasiusnas et al., 2012; Jinek et al., 2012). crRNA:tracrRNA hybrids can be fused into a single guide RNA (sgRNA) (Jinek et al., 2012). The RNA programmability of the Cas9 endonuclease makes it a powerful genome editing platform in biotechnology and medicine (Cho et al., 2013; Cong et al., 2013; Hwang et al., 2013; Jiang et al., 2013; Jinek et al., 2013; Mali et al., 2013b).

[0204] Besides sgRNA, Cas9 target recognition is typically associated with a 1–5 nucleotide feature downstream of the complementary DNA sequence (called the preinterstitial sequence adjacent motif (PAM)) (Deveau et al., 2008; Mojica et al., 2009). Cas9 orthologs exhibit considerable diversity in PAM length and sequence. Among the characterized Cas9 orthologs, *Streptococcus pyogenes* Cas9 (SpyCas9) is the most widely used, partly because it recognizes short NGG PAMs that provide a high density of targetable sites (Jinek et al., 2012) (where N represents any nucleotide). Nevertheless, the relatively large size of Spy (i.e., 1,368 amino acids) makes this Cas9 difficult to package (along with the sgRNA and promoter) into a single recombinant adeno-associated virus (rAAV). Given the promise of AAV vectors for in vivo gene delivery, this has been shown to be a drawback for therapeutic applications (Keeler et al., 2017). Furthermore, SpyCas9 and its RNA guides require extensive characterization and engineering to minimize the tendency to edit near-cognate, off-target sites (Bolukbasi et al., 2015b; Tsai and Joung, 2016; Tycko et al., 2016; Chen et al., 2017; Casini et al., 2018; Yin et al., 2018). To date, subsequent engineering efforts have not overcome these size limitations.

[0205] Several Cas9 orthologs of less than 1,100 amino acids obtained from different species (including strains of Neisseria meningitidis (NmeCas9, 1,082 aa) (Esvelt et al., 2013; Hou et al., 2013), Staphylococcus aureus (SauCas9, 1,053 aa) (Ran et al., 2015), Campylobacter jejuni (CjeCas9, 984 aa) (Kim et al., 2017), and Geobacillus stearothermophilus (GeoCas9, 1,089 aa) (Harrington et al., 2017b)) have been validated for use in mammalian genome editing. NmeCas9, CjeCas9, and GeoCas9 are representative of type II-C Cas9 (Mir et al., 2018), most of which are <1,100 aa. In addition to GeoCas9, these shorter sequence orthologs have each been successfully used for in vivo editing via all-in-one AAV delivery (where a single vector expresses both the director and the effector) (Ran et al., 2015; Kim et al., 2017; Ibraheim et al., 2018, submitted). Furthermore, NmeCas9 and CjeCas9 have demonstrated natural resistance to off-target editing (Lee et al., 2016; Kim et al., 2017; Amrani et al., 2018, submitted). However, the PAM recognized by compact Cas9 is generally longer than that of SpyCas9, thus significantly reducing the number of targetable sites at or near a given locus; for example, i) N4GAYW / N4GYTT / N4GTCT of NmeCas9 (Esvelt et al., 2013; Hou et al., 2013; Lee et al., 2016; Amrani et al., 2018); ii) N2GRRT of SauCas9 (Ran et al., 2015); iii) N4RYAC of CjeCas9 (Kim et al., 2017); and iv) N4CRAA / N4GMAA of GeoCas9 (Harrington et al., 2017b) (Y = C, T; R = A, G; M = A, C; W = A, T).Smaller subsets of target sites are advantageous for highly accurate and precise gene editing tasks, including but not limited to: i) editing of small targets (e.g., miRNAs); ii) correcting mutations via base editing, which has a very narrow base window relative to PAM changes (Komor et al., 2016; Gaudelli et al., 2017); or iii) precise editing via homology-directed repair (HDR), which is most efficient when the rewritten bases are close to the cleavage site (Gallagher and Haber, 2018). Due to PAM limitations, even with these shorter Cas9 proteins, it is not possible to target many editing sites for in vivo delivery using all-in-one AAV vectors. For example, a SauCas9 mutant with reduced PAM limitation (N3RRT) has been developed (SauCas9). KKH Although such increases in target range typically come at the cost of reduced mid-target editing efficacy, off-target editing is still observed (Kleinstiver et al., 2015).

[0206] The use of Cas9 orthologs and variants that are highly active in human cells, resistant to off-target effects, sufficiently compact for all-in-one AAV delivery, and capable of achieving high-density genomic sites will greatly enhance safe and effective CRISPR-based therapeutic gene editing. In one embodiment, the invention relates to a compact, ultra-precise Cas9 (Nme2Cas9) from different strains of Neisseria meningitidis. In another embodiment, the invention relates to a method for single AAV delivery of Nme2Cas9 and its sgRNA to perform efficient genome editing in vivo and / or in vitro. While the mechanism of the invention need not be understood, it is believed that the ortholog works efficiently in mammalian cells and recognizes N4CC PAM, which provides the same target site density as wild-type SpyCas9 (e.g., an average of 8 bp per DNA strand when considering both strands).

[0207] 1. PAM interacting domains and anti-CRISPR proteins

[0208] PAM recognition by Cas9 orthologs primarily occurs through protein-DNA interactions between the PAM interaction domain (PID) and nucleotides adjacent to the pre-interstitial region (Jiang and Doudna, 2017). PAM mutations often enable phages to evade type II CRISPR immunity (Paez-Espino et al., 2015), placing these systems under selective pressure to acquire new CRISPR spacers and evolve new PAM specificities through PID mutations. Furthermore, some phages and MGE express anti-CRISPR (Acr) proteins that inhibit Cas9 (Pawluk et al., 2016; Hynes et al., 2017; Rauch et al., 2017). PID binding is an effective inhibitory mechanism employed by some Acr strains (Dong et al., 2017; Shin et al., 2017; Yang and Patel, 2017), suggesting that PID mutations may also be driven by selective pressure to evade Acr inhibition. Cas9 PIDs may evolve to allow closely related orthologs to recognize different PAMs, as recently illustrated in two Geobacillus species. Cas9 encoded by Geobacillus thermophilus recognizes the N4CRAA PAM, but when its PID is exchanged with that of Cas9 from strain LC300, its PAM requirement changes to N4GMAA (Harrington et al., 2017b).

[0209] In one embodiment, the present invention relates to multiple meningococcal Cas9 orthologs having distinct PIDs that recognize different PAMs. In one embodiment, the present invention relates to a Cas9 protein having high sequence identity (>80% along its full length) with the Cas9 protein of NmeCas9 strain 8013 (Zhang et al., 2013). As discussed above, Nme1Cas9 also has a smaller size and naturally occurring high accuracy (Lee et al., 2016; Amrani et al., 2018). Comparison revealed clades of three meningococcal Cas9 orthologs, each clade having >98% identity in the N-terminal ~820 amino acid (aa) residues (which include all protein regions except the PID). See also Figure 5 A and Figure 6 A.

[0210] All of these Cas9 orthologs were 1,078–1,082 aa in length. The first clade (Group 1) comprised orthologs with >98% aa sequence identity to Nme1Cas9 extending through the PID. Conversely, the other two groups had PIDs significantly different from Nme1Cas9's PID, with Group 2 and Group 3 orthologs sharing an average of ~52% and ~86% PID sequence identity with Nme1Cas9, respectively. One meningococcal strain from each group was selected for detailed analysis: i) De11444 from Group 2; and ii) 98002 from Group 3, referred to herein as Nme2Cas9 (1,082 aa) and Nme3Cas9 (1,081 aa), respectively. The CRISPR-Cas loci from these two strains had the same repeat sequences and spacer lengths as strain 8013. See also Figure 6 B. This strongly suggests that their mature crRNAs also possess a 24 nt guide sequence and a 24 nt repeat sequence (Zhang et al., 2013). Similarly, the tracrRNA sequences of De11444 and 98002 are 100% identical to the tracrRNA of 8013. See also Figure 6 B. These observations suggest that the same sgRNA sequence scaffold can guide DNA cleavage through all three Cas9s.

[0211] To determine whether these Cas9 orthologs possess different PAMs, the PID of Nme1Cas9 was replaced with the PID of Nme2Cas9 or Nme3Cas9. To determine the corresponding PAM requirements, these protein chimeras were expressed in *E. coli*, purified, and used for in vitro PAM identification (Karvelis et al., 2015; Ran et al., 2015; Kim et al., 2017). In short, a library of DNA fragments containing a pre-interstitial region followed by a 10-nt randomized sequence was cleaved in vitro using recombinant Cas9 and homologous in vitro transcribed sgRNA. See also... Figure 5 B. DNA containing only the Cas9 PAM sequence is expected to be cleaved. The cleavage products are then sequenced to identify the PAM. See also Figure 5 CD.

[0212] The expected N4GATT PAM concordance sequence was verified in the recovered full-length Nme1Cas9. See also Figure 5 C. Chimeric PID-exchanged derivatives exhibit a strong preference for the C residue at position 5 (substituting for the G recognized by Nme1Cas9). See also Figure 5 D.

[0213] In one embodiment, ABE7.10-nNme2Cas9(D16A) is used for single-base editing from A·T base pairs to G·C base pairs. In another embodiment, BEmax-nNme2Cas9(D16A) is used for single-base editing from A·T base pairs to G·C base pairs. (See also...) Figure 3 F).

[0214] Figure 5 Three exemplary compactly related Neisseria meningitidis Cas9 orthologs with different PAMs are shown. Figure 5 A shows an exemplary schematic diagram that illustrates the mutated residues (orange balls) between Nme2Cas9 (left) and Nme3Cas9 (right) on the predicted structure mapped to Nme1Cas9, revealing the mutation clusters (black) in the PID. Figure 5 B shows an exemplary experimental workflow for in vitro PAM discovery assays using 10-bp random PAM regions. After in vitro digestion, the adaptor is ligated to the cleavage product for library construction and sequencing. Figure 5 C shows an exemplary sequence identifier generated by in vitro PAM discovery that reveals enrichment of Nme1Cas9 in N4GATT PAM, consistent with its previously established specificity. Figure 5 D shows an exemplary sequence identifier indicating that the Nme1Cas9 whose PID is exchanged with the PID of Nme2Cas9 (left) or Nme3Cas9 (right) requires C at PAM position 5. Due to the low cleavage efficiency of the protein chimera with PID exchange, the remaining nucleotides were not determined with high confidence (see [link to documentation]). Figure 6 C). Figure 5 E shows an exemplary sequence identifier that demonstrates full-length Nme2Cas9 identification of N4CC PAM, based on efficient substrate cutting of a target library with a fixed C at PAM position 5 and randomized PAM nt 1-4 and 6-8.

[0215] Due to the low cleavage efficiency of the chimeric protein under the conditions used, it was impossible to reliably allocate any remaining PAM nucleotides. See also Figure 6 C. To further resolve the PAM, in vitro assays were performed on a library containing a 7-nt randomized sequence (e.g., 5'-NNNNCNNN-3' on the non-complementary strand of the sgRNA) with an invariant C at the 5th PAM position. This strategy yielded significantly higher cleavage efficiency, and results indicated that Nme2Cas9 and Nme3Cas9 PIDs recognized NNNCC(A) and NNNNCAAA PAMs, respectively. See also Figure 6 CD. The Nme3Cas9 shared sequence is similar to the GeoCas9 shared sequence (Harrington et al., 2017b).

[0216] These tests were repeated using a full-length Nme2Cas9 with the NNNNCNNN DNA library (instead of a chimera with PID exchange), and the NNNNCC(A) concordant sequence was recovered again. See also Figure 5 E. Note that this test has a more effective cut. See also Figure 6 C. These data suggest that one or more of the 15 amino acid variations in Nme2Cas9 (as opposed to Nme1Cas9) outside of the PID support efficient DNA cleavage activity. See also Figure 6 C. Because Nme2Cas9’s unique 2-3nt PAM provides a higher density of potential target sites than the previously described compact Cas9 orthologs, it was chosen for further analysis.

[0217] Figure 6 Presented as with Figure 5 Characterization of related Neisseria meningitidis Cas9 orthologs with rapidly developing PID. Figure 6 A shows an exemplary rootless phylogenetic tree of NmeCas9 orthologs with >80% identity to Nme1Cas9. Three distinct branches emerge, with most mutations clustered in PIDs. Groups 1 (blue), 2 (orange), and 3 (green) have PIDs with >98%, approximately 52%, and approximately 86% identity to Nme1Cas9, respectively. Three representative Cas9 orthologs (one per group) are indicated (Nme1Cas9, Nme2Cas9, and Nme3Cas9). Figure 6 B shows an exemplary schematic diagram illustrating the CRISPR-Cas loci encoded by strains from (A) that encode three Cas9 orthologs (Nme1Cas9, Nme2Cas9, and Nme3Cas9). The percentage of identity of each CRISPR-Cas component with *Neisseria meningitidis* 8013 (encoding Nme1Cas9) is shown. Blue and red arrows indicate the transcription start sites for pre-crRNA and tracrRNA, respectively. Figure 6 C shows exemplary normalized read counts (percentage of total reads) of in vitro assays of cleaved DNA for full-length Nme1Cas9 (grey), for chimeras where the PID of Nme1Cas9 is exchanged with the PIDs of Nme2Cas9 and Nme3Cas9 (mixed colors), and for full-length Nme2Cas9 (orange). Reduced normalized read counts indicate lower cleavage efficiency in chimeras. Figure 6D shows an exemplary sequence identifier for in vitro PAM discovery assays performed on the NNNNNCNNN PAM library by exchanging the PID of Nme1Cas9 with the PID of Nme2Cas9 (left) or Nme3Cas9 (right).

[0218] 2. N4CC PAM-guided gene editing

[0219] To test the efficacy of Nme2Cas9 in human genome editing, a full-length (e.g., un-PID-exchanged) human codon-optimized Nme2Cas9 construct was cloned into a mammalian expression plasmid with an additional nuclear localization signal (NLS) and a previously validated adapter for Nme1Cas9 (Amrani et al., 2018). For initial testing, a modified fluorescence-based traffic light reporter (TLR2.0) (Certo et al., 2011) was used. In short, the disrupted GFP is followed by an out-of-frame T2A peptide and an mCherry box. When a DNA double-strand break (DSB) is introduced into the disrupted GFP box, a subset of non-homologous end joining (NHEJ) repair events leave a +1 frameshift insertion / deletion, placing the mCherry in the reading frame and producing red fluorescence easily quantifiable by flow cytometry. See [link to relevant documentation]. Figure 7 A. Green fluorescence is generated by a DNA donor containing a functional GFP sequence, and homology-mediated repair (HDR) results can also be scored simultaneously (Certo et al., 2011). Since some insertions and deletions do not introduce a +1 frameshift, fluorescence readings often underestimate the true editing efficiency. Nevertheless, the speed, simplicity, and low cost of this assay make it suitable as an initial semi-quantitative measure for genome editing in HEK293T cells carrying a single TLR2.0 locus incorporated via a slow vector.

[0220] For the initial test, the Nme2Cas9 plasmid was transiently co-transfected with one of 15 sgRNA plasmids carrying spacers targeting the TLR2.0 site with N4CC PAM. No HDR donors were included, so only NHEJ-based editing (mCherry) was scored. Most sgRNAs were in the G23 format (i.e., a 5' G to promote transcription followed by a 23nt guide sequence) as is routinely used with Nme1Cas9 (Lee et al., 2016; Pawluk et al., 2016; Amrani et al., 2018; Ibraheim et al., 2018). No sgRNA and an sgRNA targeting N4GATT PAM were used as negative controls, while co-transfection with SpyCas9+sgRNA and Nme1Cas9+sgRNA (targeting NGG and N4GATT pre-interstitial sequences, respectively) was included as a positive control. Editing by SpyCas9 and Nme1Cas9 was readily detectable (~28% and 10% mCherry, respectively). See Figure 7 B.

[0221] For Nme2Cas9, all 15 targets with N4CC PAM are functional, albeit ranging from 4% to 20% mCherry. These 15 sites include instances of each of the four possible nucleotides at the 7th PAM position (e.g., after the CC dinucleotide), indicating a slight preference for A residues observed in vitro. Figure 5 E) Does not reflect the PAM requirements for editing applications in human cells. The N4GATT PAM control produces an mCherry signal similar to that of the sgRNA-free control. See also Figure 7 B.

[0222] To determine whether both C residues in the N4CC PAM are involved in editing, a series of N4DC (D = A, T, G) and N4CD PAM sites were tested in TLR2.0 reporter cells. See also Figure 8 A and 8B. No detectable editing was found at any of these sites, thus providing a preliminary indication that the two C residues of the N4CC PAM concordant sequence are necessary for effective Nme2Cas9 activity.

[0223] The length of the spacer in crRNA varies among Cas9 orthologs and can affect on-target relative off-target activity (Cho et al., 2014; Fu et al., 2014). The optimal spacer length for SpyCas9 is 20 nt, with truncation to 17 nt tolerable (Fu et al., 2014). In contrast, Nme1Cas9 typically has a 24 nt spacer (Hou et al., 2013; Zhang et al., 2013) and tolerates truncation to 18–20 nt (Lee et al., 2016; Amrani et al., 2018). To test the spacer length requirements for Nme2Cas9, guide RNA plasmids were created for a single TLR2.0 site for each target, but with different spacer lengths. See also Figure 7 C and Figure 8 C. Comparable activity was observed using guides G23, G22, and G21, but significantly reduced activity was observed after further truncation to G20 and G19 lengths. See also Figure 7 C. These results validate Nme2Cas9 as a genome editing platform using a 22-24nt guide sequence at the N4CC PAM site in cultured human cells.

[0224] Figure 7 Exemplary data are presented showing that Nme2Cas9 edited sites adjacent to N4CC PAM using 22-24nt spacers. All experiments were performed in triplicate, and error bars represent the standard error (sem) of the mean. Figure 7 A shows an exemplary schematic diagram depicting transient transfection and editing of HEK293T TLR2.0 cells, with mCherry+ cells detected by flow cytometry 72 hours post-transfection. Figure 7 B shows an exemplary Nme2Cas9 editing of the TLR2.0 reporter factor. Sites with N4CC PAM were targeted with varying efficiencies, while Nme2Cas9 targeting was not observed on N4GATT PAM or in the absence of sgRNA. SpyCas9 (targeting previously validated sites with NGG PAM) and Nme1Cas9 (targeting N4GATT) were used as positive controls. Figure 7 C illustrates the exemplary effect of spacer length on Nme2Cas9 editing efficiency. sgRNAs targeting a single TLR2.0 site (spacer lengths of 24–20 nt (including the 5' terminal G required for the U6 promoter)) show that using a 22–24 nt spacer yields the highest editing efficiency. Figure 7D shows that exemplary Nme2Cas9 dual-nickelases can be used in tandem to generate NHEJ and HDR-based edits in TLR2.0. Plasmids expressing Nme2Cas9 and sgRNA, along with an 800 bp dsDNA donor for homology repair, were electroporated into HEK293T TLR2.0 cells, and the results for NHEJ (mCherry+) and HDR (GFP+) were scored by flow cytometry. HNH nickelase, Nme2Cas9D16A; RuvC nickelase, Nme2Cas9H588A. Either nickelase was used to target 32 ​​bp and 64 bp cleavage sites separately. The HNH nickelase (Nme2Cas9D16A) produced efficient edits, especially at 32 bp cleavage sites, while the RuvC nickelase (Nme2Cas9H588A) was ineffective. Wild-type Nme2Cas9 was used as a control.

[0225] 3. Precise editing via HDR and HNH nickases

[0226] The Cas9 enzyme utilizes its HNH and RuvC domains to cleave the guide complementary and non-complementary strands of the target DNA, respectively. SpyCas9 cleavage enzymes (nCas9), in which the HNH or RuvC domains are mutated and inactivated, have been used to induce homology-directed repair (HDR) and to improve genome editing specificity through DSB induction by dual cleavage enzymes (Mali et al., 2013a; Ran et al., 2013).

[0227] To test the efficacy of Nme2Cas9 as a nicking enzyme, Nme2Cas9 was created. D16A (HNH cleavage enzyme) and Nme2Cas9 H588A (RuvC nickases) containing alanine mutations in the catalytic residues of the RuvC and HNH domains, respectively (Esvelt et al., 2013; Hou et al., 2013; Zhang et al., 2013). TLR2.0 cells, along with GFP donor dsDNA, were used to determine whether Nme2Cas9-induced nicks could be precisely edited via HDR induction. Target sites in TLR2.0 were used to test the function of each nickase using guides targeting cleavage sites spaced 32 bp and 64 bp apart. See also Figure 7 D. Targeting a single site with wild-type Nme2Cas9 showed efficient editing, with both NHEJ and HDR being repair outcomes. For the cleavage enzyme, separating the 32bp and 64bp cleavage sites showed editing using Nme2Cas9D16A (HNH cleavage enzyme), but using Nme2Cas9H588A had no effect on either target pair. These results indicate that the Nme2Cas9 HNH cleavage enzyme can be used for efficient genome editing as long as the sites are adjacent.

[0228] Studies of previously characterized Cas9 have identified a specific region near the PAM in which Cas9 activity is highly sensitive to sequence mismatches. This 8- to 12-nt region is termed the seed sequence and has been observed in all characterized Cas9 to date (Gorski et al., 2017). To determine whether Nme2Cas9 also possesses a seed sequence, a series of transient transfections were performed, each targeting the same locus in TLR2.0 but with single nucleotide mismatches at different positions on the guide. See also Figure 8 D. A significant reduction in the number of mCherry-positive cells was observed in the first 10–12 nt near PAM for mismatches, suggesting that Nme2Cas9 has a seed sequence in this region.

[0229] Figure 8 Exemplary data are presented, showing how in mammalian cells, as with Figure 7 The relevant Nme2Cas9-targeted PAM, spacer, and seed requirements were determined. All experiments were performed in triplicate, and error bars represent sample sizes (sem). Figure 8 A shows N4C in TLR2.0 D Exemplary Nme2Cas9 targeting at the site, where editing is estimated based on mCherry+ cells. The position of each non-C nucleotide at the test site (N4C) was examined. A N4C T and N4C G Four sites were identified, with the N4CC site used as a positive control. Figure 8 B shows N4 in TLR2.0 D An exemplary Nme2Cas9 targeting at site C [similar to (A)]. Figure 8 C shows the TLR2.0 site with N4CCAPAM (and... Figure 2 The exemplary truncation on (different in C) shows similar length requirements as those observed at other sites. Figure 8 D demonstrates the exemplary sensitivity of Nme2Cas9 targeting efficiency to single nucleotide mismatches in the seed region of sgRNA. The data show the effect of single nucleotide sgRNA mismatches traveling along the 23-nt spacer at the TLR2.0 target site.

[0230] 4. Delivery methods for mammalian cell types

[0231] The ability of Nme2Cas9 to function in different mammalian cell lines was tested using various delivery methods. As an initial test, forty (40) different sites were used (29 sites were tested using N4CC PAM, and 11 sites were tested using N4CD PAM). Several loci (AAVS1, VEGFA, etc.) were selected, and target sites with N4CC PAM were randomly selected for editing with Nme2Cas9. Editing (%) was determined by transiently transfecting 150 ng of Nme2Cas9 and 150 ng of sgRNA plasmid, followed by TIDE analysis 72 h post-transfection. A subset of sites that showed a range of editing efficiencies in this initial screening were selected for triplet replicate analysis. See also Figure 9 A; and Table 1.

[0232] Figure 9 Exemplary data are presented demonstrating Nme2Cas9 genome editing at endogenous loci in mammalian cells via various delivery methods. All results represent three independent biological replicates, and error bars represent sem. Figure 9 A illustrates exemplary Nme2Cas9 genome editing at endogenous human sites in HEK293T cells following transient transfection with plasmids expressing Nme2Cas9 and sgRNA. Forty sites were initially screened (Table 1); then, 14 sites were reanalyzed in triplicate (selected to include representatives of different editing efficiencies, as measured by TIDE). An Nme1Cas9 target site (with N4GATT PAM) was used as a negative control. Figure 9 B shows exemplary data graphs: Left panel: Transient transfection of a single plasmid expressing Nme2Cas9 and sgRNA (targeting the Pcsk9 and Rosa26 loci) enabled editing in Hepa1-6 mouse cells, as detected by TIDE. Right panel: Electroporation of the sgRNA plasmid into K562 cells stably expressing Nme2Cas9 from a lentiviral vector resulted in efficient insertion / deletion formation. Figure 9 C illustrates that the exemplary Nme2Cas9 can be electroporated as an RNP complex to induce genome editing. 40 picomols of Cas9 and 50 picomols of in vitro transcribed sgRNA targeting three different loci were electroporated into HEK293T cells. Insertions and deletions were measured using TIDE after 72 h.

[0233] Table 1. Exemplary endogenous human genome editing sites targeted by Nme2Cas9.

[0234]

[0235]

[0236]

[0237]

[0238] HEK293T cells were used to support transient transfection, and cells were harvested 72 hours post-transfection for subsequent genomic DNA extraction and selective amplification of target loci. TIDE analysis was used to measure the insertion / deletion efficiency at each locus (Brinkman et al., 2014). Nme2Cas9 editing was detected at most of these sites, although the efficiency varied depending on the target sequence. Table 1. Interestingly, Nme2Cas9 induced insertions / deletions at several genomic sites with N4CD PAM, although the consistency and levels were low. Table 1. Fourteen (14) sites with N4CC PAM were analyzed in triplicate, and consistent editing was observed. See also Figure 9 A. Furthermore, increasing the amount of Nme2Cas9 plasmid delivered can significantly improve editing efficiency, and using two guides can extend this high efficiency to precise fragment deletions. See also Figure 10 A and 10B.

[0239] The ability of Nme2Cas9 to function was tested in mouse Hepa1-6 cells (hepatocellular carcinoma-derived). For Hepa1-6 cells, single plasmids encoding Nme2Cas9 and sgRNA (targeting Rosa26 or Pcsk9) were transiently transfected, and insertions and deletions were measured after 72 hours. Edits were readily observed at both sites. See also Figure 9 B, left. The functionality of Nme2Cas9 was also tested when it was stably expressed in human leukemia K562 cells. For this purpose, a lentiviral construct expressing Nme2Cas9 was created and cells were transduced to stably express Nme2Cas9 under the control of the SFFV promoter. Such stable cell lines showed no visible differences in growth and morphology compared to untransduced cells, indicating that Nme2Cas9 is non-toxic when stably expressed. These cells were transiently electroporated with a plasmid expressing sgRNA and insertion / deletion efficiency was measured by TIDE analysis after 72 hours. Effective (>50%) editing was observed at all three sites tested, validating the ability of Nme2Cas9 to function in K562 cells after lentiviral delivery. See also Figure 9 B.

[0240] Ribonucleoprotein (RNP) delivery of Cas9 and its sgRNA can also be used in some genome editing applications, and the greater transient nature of Cas9 minimizes off-target editing (Kim et al., 2014; Zuris et al., 2015). Furthermore, some cell types (e.g., certain immune cells) are resistant to DNA transfection-based editing (Schumann et al., 2015). To test whether Nme2Cas9 functions via RNP delivery, 6xHis-tagged Nme2Cas9 (fused with three NLS) was cloned into a bacterial expression construct, and the recombinant protein was purified. The recombinant protein was then loaded with sgRNA transcribed by a T7 RNA polymerase targeting three previously validated sites. Electroporation of the Nme2Cas9:sgRNA complex induced successful editing at each of the three target sites in HEK293T cells, as detected by TIDE. See also Figure 9 C. These results collectively demonstrate that Nme2Cas9 can be efficiently delivered to a variety of cell types via plasmids, lentiviruses, or as an RNP complex.

[0241] 5. Anti-CRISPR regulation

[0242] To date, five Acr families from different bacterial species have been shown to inhibit Nme1Cas9 in vitro and in human cells (Pawluk et al., 2016; Lee et al., 2018, submitted). Given the high sequence identity between Nme1Cas9 and Nme2Cas9, at least some of these Acr families should inhibit Nme2Cas9. To test this, recombinant Acrs from all five families were expressed, purified, and the ability of Nme2Cas9 to cleave targets in vitro in the presence of members of each family was tested (Acr:Cas9 molar ratio of 10:1). Inhibitors were used in the IE-type CRISPR system (AcrE2) in *E. coli* as a negative control, while Nme1Cas9 was used as a positive control (Pawluk et al., 2014); (Pawluk et al., 2016). As expected, all five families inhibited Nme1Cas9, while AcrE2 did not. See also Figure 11 A, as shown in the image above. AcrIIC1 Nme AcrIIC2 Nme AcrIIC3 Nme and AcrIIC4 Hpa Complete inhibition of Nme2Cas9. However, surprisingly, AcrIIC5, previously reported as the most potent Nme1Cas9 inhibitor, showed no effect. Smu(Lee et al., 2018) Nme2Cas9 could not be inhibited in vitro even with a 10-fold molar excess. This suggests that it may inhibit Nme1Cas9 through interaction with its PID.

[0243] Figure 10 Example data is presented, showing how it is compared to... Figure 9 The related dose dependence and segmental deletion of Nme2Cas9. Figure 10 A shows an exemplary increase in the dosage of the electroporated Nme2Cas9 plasmid (500 ng relative to the target). Figure 3 The 200 ng in A improved editing efficiency at two sites (TS16 and TS6). Data provided in yellow... Figure 9 A can be reused. Figure 10 B illustrates an example of Nme2Cas9 used to generate precise segment deletions. Two TLR2.0 targets with cleavage sites 32 bp apart were simultaneously targeted by Nme2Cas9. Most of the resulting lesions were deletions of exactly 32 bp (blue).

[0244] Figure 11 Exemplary data are presented showing that Nme2Cas9 is inhibited in vitro and in cells by a subset of type II-C anti-CRISPR family. All experiments were performed in triplicate, and error bars represent sem. Figure 11 Figure A shows an exemplary in vitro cleavage assay of Nme1Cas9 and Nme2Cas9 in the presence of five previously characterized anti-CRISPR proteins (Acr:Cas9 ratio of 10:1). Top panel: Nme1Cas9 efficiently cleaves fragments containing the pre-interstitial region sequence with N4GATT PAM in the absence of Acr or in the presence of the negative control Acr (AcrE2). As expected, all five previously characterized type II-C Acr family proteins inhibit Nme1Cas9. Bottom panel: Nme2Cas9 inhibition is the same as Nme1Cas9, but lacks the inhibition via AcrIIC5. Smu The inhibition. Figure 11 B shows an exemplary genome editing in the presence of five previously described anti-CRISPR families. Plasmids expressing Nme2Cas9 (200 ng), sgRNA (100 ng), and each corresponding Acr (200 ng) were co-transfected into HEK293T cells, and genome editing was measured 72 hours post-transfection using traced insertions and deletions by degradation (TIDE). Consistent with our in vitro analyses, except for AcrIIC5… Smu In addition, all II-C anti-CRISPR agents inhibit genome editing, although with varying efficiencies. Figure 11C shows that the exemplary Acr inhibition of Nme2Cas9 is dose-dependent and has significant apparent power. Nme2Cas9 was completely inhibited by AcrIIC1Nme and AcrIIC4Hpa at co-transfected Acr and Nme2Cas9 plasmid mass ratios of 2:1 and 1:1, respectively.

[0245] To further test this, an Nme1Cas9 / Nme2Cas9 chimera with an Nme2Cas9 PID was tested. See [link / reference] Figure 5 D and Figure 6 D. Due to the reduced activity of this hybrid, a concentration of Cas9 ~30X was used to achieve similar cleavage efficiency while maintaining a Cas9:Acr molar ratio of 10:1. No inhibitory effect of AcrIIC5Smu on the protein chimera was observed. See also Figure 12 This data provides further evidence of a possible PID interaction between AcrIIC5Smu and Nme1Cas9. Regardless of the mechanistic basis of differential inhibition via AcrIIC5Smu, these results suggest that Nme2Cas9 is inhibited by the other four type II-C Acr families.

[0246] Figure 12 Example data is presented, showing how it is compared to... Figure 11 The associated Nme2Cas9 PID exchange renders Nme1Cas9 insensitive to AcrIIC5Smu inhibition. In vitro cleavage of the Nme1Cas9-Nme2Cas9 PID chimera occurs in the presence of previously characterized Acr protein (10 uM Cas9-sgRNA + 100 uM Acr).

[0247] Based on the above in vitro data, it was hypothesized that AcrIIC1Nme, AcrIIC2Nme, AcrIIC3Nme, and AcrIIC4Hpa could be used as an off-switch for Nme2Cas9 genome editing. To test this, HEK293T cells were transfected with Nme2Cas9 / sgRNA plasmids targeting TS16 (150 ng per plasmid) in the presence or absence of Acr expression plasmids, as most Acr cells are reported to inhibit Nme1Cas9 at these plasmid ratios (Pawluk et al., 2016). As expected, AcrIIC1Nme, AcrIIC2Nme, AcrIIC3Nme, and AcrIIC4Hpa inhibited Nme2Cas9 genome editing, while AcrIIC5Smu had no effect. See also Figure 11B. Complete inhibition was observed using AcrIIC3Nme and AcrIIC4Hpa, indicating their high potency against Nme2Cas9 compared to AcrIIC1Nme and AcrIIC2Nme. To further compare the potency of AcrIIC1Nme and AcrIIC4Hpa, we repeated the experiments at various ratios of Acr plasmid to Cas9 plasmid. See [link to article]. Figure 11 C. Data show that the AcrIIC4Hpa plasmid is particularly effective against Nme2Cas9. In summary, these data suggest that several Acr proteins can be used as a switch to turn off Nme2Cas9-based applications.

[0248] 6. Super accuracy

[0249] Nme1Cas9 has demonstrated superior editing fidelity in cell and mouse models (Lee et al., 2016; Amrani et al., 2018; Ibraheim et al., 2018). Furthermore, the similarity of Nme2Cas9 to Nme1Cas9 across most of its length suggests it may be similarly hyperaccurate. However, the higher number of sampling sites in the genome due to dinucleotide PAMs, compared to Nme1Cas9 and its less frequently encountered 4-nucleotide PAMs, may create more opportunities for Nme2Cas9 off-target effects. To assess the off-target profile of Nme2Cas9, GUIDE-seq (genome-wide, unbiased identification of double-strand breaks via sequencing) was used to empirically and unbiasedly identify potential off-target sites (Tsai et al., 2014). Even the best off-target prediction algorithms are prone to false negatives, thus requiring empirical target site analysis methods (Bolukbasi et al., 2015b; Tsai and Joung, 2016; Tycko et al., 2016). GUIDE-seq relies on the integration of double-stranded oligodeoxynucleotides (dsODNs) into DNA double-strand break sites throughout the genome. These insertion sites are then detected by amplification and high-throughput sequencing.

[0250] Because SpyCas9 is a well-characterized ortholog of Cas9, it can be used for multiple applications with other Cas9s and as a benchmark for their editing properties (Jiang and Doudna, 2017; Komor et al., 2017). SpyCas9 and Nme2Cas9 were cloned into the same plasmid backbone with identical UTRs, adapters, NLS, and promoters for parallel transient transfection (along with similarly matched plasmids expressing sgRNA) into HEK293T cells. First, it was confirmed that the RNA directors used for SpyCas9 and Nme2Cas9 are orthogonal, i.e., Nme2Cas9 sgRNA is not edited by SpyCas9-directed editing, and vice versa. See also Figure 13 A. This contradicts earlier reports of results using Nme1Cas9 (Esvelt et al., 2013; Fonfara et al., 2014).

[0251] Next, to identify SpyCas9 as a benchmark for GUIDE-seq, since SpyCas9 and Nme2Cas9 have non-overlapping PAMs, it can potentially edit any dual-site (DS) flanked by a 5'-NGGNCC-3' sequence, satisfying the PAM requirements of both Cas9s. This allows for parallel comparisons of off-target RNA guides that promote editing of the exact same intermediate target site. See also Figure 14 A. Six (6) DS sites in VEGFA were targeted, each with a G at the appropriate 5' position of the PAM, such that both SpyCas9 and Nme2Cas9 directives (driven by the U6 promoter) were 100% complementary to the target sites. TIDE analysis was performed on these sites targeted by each nuclease seventy-two (72) hours post-transfection. Nme2Cas9 induced insertions and deletions at all six sites, although two were less efficient, while SpyCas9 induced insertions and deletions at four of the six sites. See also Figure 14 B. At two of the four effective sites (DS1 and DS4), SpyCas9 induces approximately 7 times more insertions and deletions than Nme2Cas9, while Nme2Cas9 induces approximately 3 times more insertions and deletions at DS6 than SpyCas9. The two Cas9 orthologs edit DS2 with almost equal efficiency.

[0252] For GUIDE-seq, DS2, DS4, and DS6 were selected to sample off-target cuts using the Nme2Cas9 guide, which guides mid-target editing with varying degrees of efficiency compared to the corresponding SpyCas9 guide. In addition to the three dual sites, TS6 was added because it has been observed to be an effective Nme2Cas9 target site for editing, exhibiting approximately 30-50% insertion / deletion efficiency depending on cell type. See also Figure 9 A and 10A. Similar data were seen using mouse Pcsk9 and Rosa26 Nme2Cas9 sites. See also Figure 9 B.

[0253] Each Cas9 cell and its homologous sgRNA and dsODN were transfected with plasmids. Subsequently, GUIDE-seq libraries were prepared as previously described (Amrani et al., 2018). GUIDE-seq analysis showed efficient mid-target editing of the two Cas9 orthologs, with relative efficiency (reflected by GUIDE-seq read counts) similar to that observed via TIDE. Figure 13 B and Table 2. (Tsai et al., 2014; Zhu et al., 2017).

[0254] Figure 13 Example data is presented, showing how it is compared to... Figure 12 The orthogonality and relative accuracy of Nme2Cas9 and SpyCas9 at dual target sites. Figure 13 A shows that the exemplary Nme2Cas9 and SpyCas9 guides are orthogonal. TIDE results show the frequency of insertions and deletions generated by the two nucleases targeting DS2 and their homologous sgRNAs or sgRNAs of other orthologs. Figure 13 B shows exemplary Nme2Cas9 and SpyCas9, demonstrating comparable mid-target editing efficiencies, as evaluated by GUIDE-seq. The bar graph represents the mid-target readouts from GUIDE-Seq at the three dual sites for each orthologous target. Orange bars represent Nme2Cas9, and black bars represent SpyCas9. Figure 13 C shows the target-to-miss reading count for an exemplary SpyCas9 at each site. Orange bars represent target readings, while black bars represent misses. Figure 13 D shows the target-to-off-target readings in an exemplary Nme2Cas9 at each site. Figure 13 The bar chart for E shows exemplary insertion / deletion efficiencies of potential off-target sites predicted by CRISPRSeek (measured by TIDE). Target and off-target site sequences are shown on the left, with PAM regions underlined and sgRNA mismatches and non-shared PAM nucleotides indicated in red.

[0255] Table 2: GUIDE-seq data SpyDS2 (gRNA.name SpyDS2)

[0256]

[0257]

[0258]

[0259]

[0260]

[0261]

[0262]

[0263]

[0264]

[0265]

[0266]

[0267]

[0268]

[0269]

[0270]

[0271]

[0272]

[0273]

[0274]

[0275]

[0276]

[0277]

[0278]

[0279]

[0280]

[0281]

[0282]

[0283]

[0284]

[0285] SpyDS4(gRNA.name SpyDS4)

[0286]

[0287]

[0288]

[0289]

[0290] SpyDS6(gRNA.name SpyDS6)

[0291]

[0292]

[0293]

[0294]

[0295]

[0296]

[0297]

[0298]

[0299]

[0300]

[0301]

[0302]

[0303]

[0304]

[0305]

[0306]

[0307]

[0308]

[0309]

[0310]

[0311]

[0312]

[0313]

[0314]

[0315]

[0316]

[0317]

[0318]

[0319]

[0320]

[0321]

[0322]

[0323]

[0324]

[0325]

[0326]

[0327]

[0328] Nme2DS2

[0329]

[0330]

[0331] Nme2DS4

[0332]

[0333]

[0334] Nme2DS6

[0335]

[0336]

[0337]

[0338] Rosa26

[0339]

[0340]

[0341] PCSK9

[0342]

[0343]

[0344]

[0345] Regarding off-target identification, analysis showed that when analyzing plasmid-based SpyCas9 editing via GUIDE-seq, DS2, DS4, and DS6 SpyCas9 sgRNAs appeared to guide editing with 93, 10, and 118 off-target candidate sites, respectively, within the normal range of off-target events (Fu et al., 2014; Tsai et al., 2014). In stark contrast, DS2, DS4, and DS6Nme2Cas9 sgRNAs appeared to guide editing with 1, 0, and 1 off-target sites, respectively. Figure 14 C and Table 2. The Nme2Cas9 read counts are very low compared to the GUIDE-seq read counts for SpyCas9 off-target effects, further demonstrating the high specificity of Nme2Cas9. See also Figure 13 C, Figure 13 D. Nme2Cas9 GUIDE-seq analysis using TS6, Pcsk9, and Rosa26 yielded similar results (0, 0, and 1 off-target sites, respectively, with smaller read counts for Rosa26-OT1 off-target sites). Figure 13 C, Figure 14 D and Table 2.

[0346] Figure 14 Exemplary data are presented, showing that Nme2Cas9 has little or no detectable off-target effects in mammalian cells. Figure 14 A shows an exemplary schematic diagram depicting a dual-site (DS) that can be targeted by both SpyCas9 and Nme2Cas9 via their non-overlapping PAMs. The Nme2Cas9 PAM (orange) and SpyCas9 PAM (blue) are highlighted. The 24nt Nme2Cas9 guide sequence is shown in yellow; the corresponding SpyCas9 guide sequence is 4nt shorter at the 5' end. Figure 14 B shows exemplary Nme2Cas9 and SpyCas9 that induce insertion deletions at DS. VEGFA (with GN3GN) was selected. 19Six DS (distribution sites) of the NGGNCC sequence were used for direct comparison of editing by two orthologs. Plasmids expressing each Cas9 (with the same promoter, adapter, tag, and NLS) and its homologous director were transfected into HEK293T cells. Insertion / deletion efficiency was determined by TIDE 72 hours post-transfection. Nme2Cas9 editing was detectable at all six sites and was slightly or significantly more efficient than SpyCas9 at two sites (DS2 and DS6, respectively). SpyCas9 edited four of the six sites (DS1, DS2, DS4, and DS6), with two sites showing significantly higher editing efficiency than Nme2Cas9 (DS1 and DS4). DS2, DS4, and DS6 were selected for GUIDE-Seq analysis because at these sites, Nme2Cas9 showed the same, lower, and higher efficiency compared to SpyCas9, respectively. Figure 14 C shows an exemplary Nme2Cas9 genome editing with high precision in human cells. The number of off-target sites detected by GUIDE-Seq at each target site for each nuclease is shown. In addition to the two sites, we also analyzed TS6 in mouse Hepa1-6 cells (due to its high mid-target editing efficiency) as well as Pcsk9 and Rosa26 sites (to measure accuracy in another cell type). Figure 14 D shows exemplary targeted deep sequencing used to detect insertions and deletions in edited cells, confirming the high Nme2Cas9 accuracy indicated by GUIDE-seq. Figure 14 E shows an exemplary sequence of empirically proven off-target sites for the Rosa26 guide, showing the PAM region (underlined), with three mismatches (red) in the CC PAM dinucleotide (bold) and the distal portion of the spacer region PAM.

[0347] To validate off-target sites detected by GUIDE-seq, targeted deep sequencing was performed after GUIDE-seq-independent editing (i.e., without dsODN co-transfection) to measure insertion / deletion formation at the top off-target sites. While SpyCas9 showed considerable editing at most of the tested off-target sites, and in some cases more effectively than at the corresponding intermediate target sites, Nme2Cas9 did not show detectable insertions / deletions at individual DS2 and DS6 candidate off-target sites. See also Figure 14 D. Using Rosa26 sgRNA, Nme2Cas9 induced ~1% editing at the Rosa26-OT1 site in Hepa1-6 cells, compared to ~30% for intermediate-target editing. See also Figure 14 D. Notably, this off-target site shares the common Nme2Cas9 PAM(ACTC) gene. CCT), which has only 3 mismatches at the distal end of the PAM in the guide complement region (i.e., outside the seed). See also Figure 14 E. These data support and reinforce our GUIDE-seq results, demonstrating the high accuracy of Nme2Cas9 genome editing in mammalian cells.

[0348] To further confirm the GUIDE-Seq results above, CRISPRseek was used to computationally predict potential off-target sites for two active Nme2Cas9 sgRNAs targeting TS25 and TS47 (both also located in VEGFA). See [link to documentation]. Figure 9 A; (Zhu et al., 2014). Three (TS25) or four (TS47) of the most closely matched predicted sites had N4CCPAM, and five had N4CA PAM; each had 2–5 mismatches, primarily in the non-seed region distal to their PAMs. See also Figure 13 E. After transfection of the Nme2Cas9+sgRNA plasmid into HEK293T cells, targeted amplification at each locus was performed, followed by TIDE analysis to compare on-target and off-target editing. Consistently, no insertions or deletions were detected by TIDE at the off-target sites of any sgRNA, while effective on-target editing was readily detected in DNA from the same cell population. In summary, our data demonstrate that Nme2Cas9 is a natural, ultra-accurate genome editing platform in mammalian cells.

[0349] 7. Adeno-associated virus delivery

[0350] The compact size, small PAM, and high fidelity of Nme2Cas9 provide major advantages for in vivo genome editing using adeno-associated virus (AAV) delivery. To test whether efficient Nme2Cas9 genome editing could be achieved via single AAV delivery, Nme2Cas9, along with its sgRNA and promoters (U1a and U6, respectively), were cloned into the AAV vector backbone. See also Figure 15 A. A multi-gene AAV was prepared by packaging sgRN-.Nme2Cas9 into a hepatophilic AAV8 capsid to target two genes in mouse liver: i) Rosa26 (a commonly used safe harbor locus for transgene insertion) (Friedrich and Soriano, 1991) as a negative control; and ii) Pcsk9 as a phenotypic target, which is a major regulator of circulating cholesterol homeostasis (Rashid et al., 2005).

[0351] Insertion-deletion of SauCas9 or Nme1Cas9 in Pcsk9 in mouse liver leads to and reduces cholesterol levels, thus providing a useful and easily scored in vivo benchmark for the new editing platform (Ran et al., 2015; Ibraheim et al., 2018). The Nme2Cas9 RNA guide is the same as those used above. See also Figure 9 B. Figure 13 D and Figure 14 Since Rosa26-OT1 is the only validated Nme2Cas9 off-target site in cultured mammalian cells, the Rosa26 guide also provides us with an opportunity to evaluate intermediate-target versus off-target editing in vivo. See also Figure 14 DE. Two groups of mice (n=5) were injected via the tail vein with 4 x 10⁻⁶ Pcsk9 or Rosa26 targeting Pcsk9 or Rosa26. 11 One copy of the AAV8.sgRNA.Nme2Cas9 genome (GC) was collected. Serum was collected at 0, 14, and 28 days post-injection for cholesterol level measurement. Mice were sacrificed 28 days post-injection, and liver tissue was collected. See also Figure 15 A. Targeted deep sequencing of each locus revealed ~38% and ~46% insertion / deletion induction at the Pcsk9 and Rosa26 edit sites in the liver, respectively. See also Figure 15 B. Since hepatocytes account for only 65-70% of the total cells in an adult liver, the hepatocyte editing efficiencies induced by Nme2Cas9 AAV using sgPcsk9 and sgRosa are approximately 54-58% and 66-71%, respectively (Racanelli and Rehermann, 2006).

[0352] Only 2.25% of liver insertions / deletions were detected at the Rosa26-OT1 off-target site (approximately 3-3.5% in hepatocytes), comparable to the 1% edit observed at this site in transfected Hepa1-6 cells. See also Figure 15 B, Figure 14 D. On days 14 and 28 post-injection, Pcsk9 editing was accompanied by a ~44% decrease in serum cholesterol levels, whereas mice treated with AAV expressing sgRosa26 maintained normal cholesterol levels throughout the study. See also Figure 15 The ~44% reduction in serum cholesterol in mice treated with C. Nme2Cas9 / sgPcsk9 AAV was comparable to the ~40% reduction reported when using SauCas9 multi-AAV targeting the same gene (Ran et al., 2015).

[0353] Figure 15 Exemplary data are presented demonstrating in vivo editing of the Nme2Cas9 genome via all-in-one AAV delivery. Figure 15 Figure A shows an exemplary workflow for delivering AAV8.sgRNA.Nme2Cas9 to lower cholesterol levels in mice via Pcsk9 targeting. Top: Schematic diagram of a multi-purpose AAV vector expressing Nme2Cas9 and sgRNA (individual genomic elements are not drawn to scale). BGH, bovine growth hormone polymer (A) site; HA, epitope tag; NLS, nuclear localization sequence; h, human codon-optimized. Bottom: Tail vein injection of AAV8.sgRNA.Nme2Cas9 (4 x 10⁻⁶) 11 The timeline of GC was analyzed, cholesterol was measured on day 14, and insertion / deletion, histological and cholesterol analysis was performed on day 28 post-injection. Figure 15 B shows an exemplary TIDE analysis to measure insertions and deletions in DNA extracted from the livers of mice injected with AAV8.Nme2Cas9+ sgRNAs targeting the Pcsk9 and Rosa26 (control) loci. The insertion / deletion efficiency of GUIDE-seq at individual off-target sites recognized by these two sgRNAs (Rosa26|OT1) was also evaluated by TIDE. Figure 15 C shows an exemplary reduction in serum cholesterol levels in mice injected with a guide targeting Pcsk9 compared to a control targeting Rosa26. P-values ​​were calculated using an unpaired two-tailed t-test. Figure 16 Example data is presented, showing the relationship with Figure 15 Related Nme2Cas9 AAV delivery and edited PCSK9 knockdown and liver histology. Figure 16 Figure A shows an exemplary Western blot using an anti-PCSK9 antibody, revealing a significantly reduced level of PCSK9 in the livers of mice treated with sgPcsk9 compared to mice treated with sgRosa26. 2 ng of recombinant PCSK9 was used as a migration standard (leftmost lane), and cross-reactivity bands in the liver samples are indicated by asterisks. GAPDH was used as a loading control (below). Figure 16 B shows exemplary H&E staining of livers from mice injected with either AAV8.Nme2Cas9+sgRosa26 (left) or AAV8.Nme2Cas9+sgPcsk9 (right) vectors. Scale bar, 25 μm.

[0354] Western blotting was performed using an anti-PCSK9 antibody to assess PCSK9 protein levels in the livers of mice treated with sgPcsk9 and sgRosa26. In mice treated with sgPcsk9, hepatic PCSK9 levels were below the detection limit, while mice treated with sgRosa26 showed normal PCSK9 levels. See also Figure 16A. Hematoxylin and eosin (H&E) staining and histological examination showed no signs of toxicity or tissue damage in either group after Nme2Cas9 expression. See also Figure 16 B. These data validate that Nme2Cas9 is a highly efficient genome editing system in vivo, including when delivered via a single AAV vector.

[0355] AAV vectors have recently been used to generate genome-edited mice without microinjection or electroporation; simply by immersing fertilized eggs in a medium containing AAV vectors before implantation into pseudopregnant females (Yoon et al., 2018). Previous editing was achieved using a dual AAV system, in which SpyCas9 and its sgRNA are delivered in separate vectors (Yoon et al., 2018). To test whether Nme2Cas9 could be accurately and efficiently edited in mouse fertilized eggs using a multi-AAV delivery system, we targeted the tyrosinase (Tyr). Biallelic inactivation of Tyr disrupts melanin production, leading to an albinism phenotype (Yokoyama et al., 1990).

[0356] The efficacy of the Tyr sgRNA was validated, which cleaves the Tyr locus at a position only seventeen (17) bp away from the classic albinism mutation site in Hepa1-6 cells via transient transfection. See also Figure 17 A. Next, the C57BL / 6NJ fertilized eggs were placed in a container containing 3x10 9 Or 3x10 8 The GC-expressing Nme2Cas9 and Tyr sgRNA-infused AAV6 vector were incubated in medium for 5–6 hours. After overnight culture in fresh medium, fertilized eggs that reached the two-cell stage were transferred to the fallopian tubes of pseudopregnancy recipients and allowed to develop to the terminal stage. See also Figure 18 A. Coat color analysis of the pups showed that the mice were albino, silver-gray (indicating a suballelic gene for tyrosinase), or had a mottled coat consisting of albino and silver-gray patches but lacking black pigmentation. See also Figure 18 BC. These results indicate a high frequency of biallelic mutations, since the presence of the wild-type tyrosinase allele should lead to melanin pigmentation. From 3x10 9 A total of 5 offspring (10%) were born in the GC experiment. All of them carried insertions and deletions; phenotypically, two were albino, one was silver-gray, and two had variegated pigmentation, indicating mosaicism.

[0357] According to 3x10 8GC experiments yielded four (4) pups (14%), two of which died at birth, thus preventing coat color or genomic analysis. Coat color analysis of the remaining two pups revealed one silver-grey and one mosaic pup. These results suggest that single AAV delivery of Nme2Cas9 and its director can be used to generate mutations in mouse zygotes without the need for microinjection or electroporation.

[0358] To measure mid-target insertion / deletion formation in the Tyr gene, DNA was isolated from the tail of each mouse, the locus was amplified, and TIDE analysis was performed on it. All mice exhibited high levels of mid-target editing via Nme2Cas9, ranging from 84% to 100%. See also Figure 17 BC. Most lesions in albino mice 9-1 were 1-bp or 4-bp deletions, suggesting mosaicism or trans-heterozygosity, but albino mice 9-2 showed a consistent 2-bp deletion. See also Figure 17 C. Figure 17 Exemplary data is presented, and its display is consistent with... Figure 16 Related in vitro Tyr editing in mouse fertilized eggs. Figure 17 A shows two exemplary sites in Tyr, each with N4CC PAM, tested for editing in Hepa1-6 cells. The sgTyr2 guide showed higher editing efficiency and was therefore selected for further testing. Figure 17 B shows seven exemplary mice that survived postnatal development and each exhibited a coat color phenotype and mid-target editing, as determined by TIDE. Figure 17 C shows exemplary insertion / deletion profiles of tail DNA from each mouse from (B) and from unedited C57BL / 6NJ mice, as shown by TIDE analysis. The efficiency of insertions (positive) and deletions (negative) of various sizes is indicated.

[0359] Figure 18 Exemplary data are presented, demonstrating in vitro Nme2Cas9 genome editing delivered via an all-in-one AAV. Figure 18 A illustrates an exemplary workflow for in vitro editing of single AAV Nme2Cas9 to generate albino C57BL / 6NJ mice by targeting the Tyr gene. Fertilized eggs were cultured in KSOM containing AAV6.Nme2Cas9:sgTyr for 5–6 hours, washed in M2, and cultured for one day before being transferred to the oviduct of a pseudopregnancy recipient. Figure 18 B shows the result via 3x10 9Exemplary albino (left) and silver-gray or mottled (middle) mice generated by GC, and 3x10 mice from fertilized eggs with AAV6.Nme2Cas9:sgTyr. 8 Silver-gray or mottled mice generated by GC (right). Figure 18 C shows an exemplary summary of Nme2Cas9.sgTyr single-AAV ex vivo Tyr editing experiments at two AAV doses.

[0360] The data are inconclusive regarding whether mosaicism was absent in mouse 9-2 or whether other alleles were missing from mouse 9-1, as only tail samples were sequenced, and other tissues may have shown significant lesions. Analysis of tail DNA from silver-gray mice revealed in-read-frame mutations, which may be the cause of the silver-gray coat color. The limited mutational complexity suggests that editing occurred early in the embryonic development of these mice. These results provide a simplified pathway for mammalian mutagenesis by applying a single AAV vector, in which both Nme2Cas9 and its sgRNA are delivered.

[0361] Figure 19 Exemplary mCherry reporter factor assays for the activities of nSpCas9-ABEmax and optimized ABEmax-nNme2Cas9(D16A) are shown. Figure 19 A shows an example sequence of the ABE-mCherry reporter factor. There is a TAG stop codon in the mCherry coding region. In stable cell lines with reporter factor integration, there is no mCherry signaling due to this stop codon. If nSpCas9-ABEmax or an optimized nNme2Cas9-ABEmax can convert the TAG to CAG (which encodes glutamine residues), then mCherry signaling will be activated. Figure 19 B shows exemplary mCherry signal activation due to SpCas9-ABE or Nme2Cas9-ABE activity. Top panel: Negative control (no editing); Middle panel: mCherry activation via nSpCas9-ABEmax; Bottom panel: mCherry activation via optimized nNme2Cas9-ABEmax. Figure 19 C shows an exemplary FACS quantification of base editing events in mCherry reporter cells transfected with SpCas9-ABE or Nme2Cas9-ABE. N = 6; error bars represent SD. Results are from three biological replicates performed in a technical replicate.

[0362] Figure 20An exemplary GFP reporter factor assay is shown for the activity of nSpCas9-CBE4 (Addgene#100802) and nNme2Cas9-CBE4 (the same plasmid backbone as Addgene#100802). Figure 20 A shows exemplary sequence information for the CBE-GFP reporter factor. A mutation exists in the core region of the fluorophore of the GFP reporter factor line that converts GYG to GHG. Due to this mutation, there is no GFP signal. If nSpCas9-CBE4 or nNme2Cas9-CBE4 can convert CAC (encoding histidine) to TAC / TAT (encoding tyrosine), the GFP signal will be activated. Figure 20 B shows exemplary GFP signals activated due to nSpCas9-CBE4 or nNme2Cas9-CBE4 activity. Top panel: Negative control (no editing); Middle panel: GFP activation via nSpCas9-CBE4; Bottom panel: GFP activation via nNme2Cas9-CBE4. Figure 20 C shows an exemplary FACS quantification of base editing events in GFP reporter cells transfected with nSpCas9-CBE4 or nNme2Cas9-CBE4. N = 6; error bars represent SD. Results are from biological replicates performed in technical replicates.

[0363] Figure 21 This section illustrates exemplary cytosine editing performed via nNme2Cas9-CBE4. The top panel shows the KANK3 target sequence information for Nme2Cas9 (PAM sequences are shown in red) and base editing in the negative control sample. The bottom panel shows the quantification of substitution efficiency for each type of base in the nNmeCas9-CBE4 editing window of the KANK3 target sequence. The sequence table shows the nucleotide frequency at each position. The expected frequency of C-to-T conversions is shown in red.

[0364] Figure 22 Exemplary cytosine and adenine editing were performed via nNme2Cas9-CBE4 and nNme2Cas9-ABEmax, respectively. The top figure shows the PLXNB2 target sequence information for Nme2Cas9 (PAM sequences are shown in red) and base editing in the negative control sample. The middle figure shows the quantification of the substitution rate for each type of base in the nNmeCas9-ABEmax editing window of the PLXNB2 target sequence. The sequence table shows the nucleotide frequency at each position. The expected A-to-G conversion frequency is highlighted in red. The bottom figure shows the quantification of the substitution efficiency for each type of base in the nNmeCas9-CBE4 editing window of the PLXNB2 target sequence. The sequence table shows the nucleotide frequency at each position. The expected C-to-T conversion frequency is highlighted in red.

[0365] 8. Sequence

[0366] Comparison of Nme1Cas9 and Nme2Cas9

[0367] Non-PID aa differences (cyan-underlined); PID aa differences (yellow-underlined bold); active site residues (red-bold).

[0368] Nme1Cas9(1-60)(SEQ ID NO:652)

[0369] Nme2Cas9(1-60)(SEQ ID NO:883)

[0370]

[0371] Nme1Cas9(61-120)(SEQ ID NO:653)

[0372] ARRLARSVRRLTRRRAHRLLR T RRLLKREGVLQAA N FDENGLIKSLPNTPWQLRAAALDR

[0373] Nme2Cas9(61-120)(SEQ ID NO:654)

[0374] ARRLARSVRRLTRRRAHRLLR A RRLLKREGVLQAA D FDENGLIKSLPNTPWQLRAAALDR

[0375] Nme1Cas9(121-180)(SEQ ID NO:655)

[0376] KLTPLEWSAVLLHLIKHRGYLSQRKNEGETADKELGALLKGVA G NAHALQTGDFRTPAEL

[0377] Nme2Cas9(121-180)(SEQ ID NO:656)

[0378] KLTPLEWSAVLLHLIKHRGYLSQRKNEGETADKELGALLKGVA N NAHALQTGDFRTPAEL

[0379] Nme1Cas9(181-240)(SEQ ID NO:657)

[0380] ALNKFEKESGHIRNQR S DYSHTFSRKDLQAELILLFEKQKEFGNPHVSGGLKEGIETLLM

[0381] Nme2Cas9(181-240)(SEQ ID NO:658)

[0382] ALNKFEKESGHIRNQR G DYSHTFSRKDLQAELILLFEKQKEFGNPHVSGGLKEGIETLLM

[0383] Nme1Cas9(241-300)(SEQ ID NO:659)

[0384] TQRPALSGDAVQKMLGHCTFEPAEPKAAKNTYTAERFIWLTKLNNLRILEQGSERPLTDT

[0385] Nme2Cas9(241-300)(SEQ ID NO:660)

[0386] TQRPALSGDAVQKMLGHCTFEPAEPKAAKNTYTAERFIWLTKLNNLRILEQGSERPLTDT

[0387] Nme1Cas9(301-360)(SEQ ID NO:661)

[0388] ERATLMDEPYRKSKLTYAQARKLLGLEDTAFFKGLRYGKDNAEASTLMEMKAYHAISRAL

[0389] Nme2Cas9(301-360)(SEQ ID NO:662)

[0390] ERATLMDEPYRKSKLTYAQARKLLGLEDTAFFKGLRYGKDNAEASTLMEMKAYHAISRAL

[0391] Nme1Cas9(361-420)(SEQ ID NO:663)

[0392] EKEGLKDKKSPLNLS P ELQDEIGTAFSLFKTDEDITGRLKDRI QPEILEALLKHISFDKF

[0393] Nme2Cas9(361-420)(SEQ ID NO:664)

[0394] EKEGLKDKKSPLNLS S ELQDEIGTAFSLFKTDEDITGRLKDR V QPEILEALLKHISFDKF

[0395] Nme1Cas9(421-480)(SEQ ID NO:665)

[0396] VQISLKALRRIVPLMEQGKRYDEACAEIYGDHYGKKNTEEKIYLPPIPADEIRNPVVLRA

[0397] Nme2Cas9(421-480)(SEQ ID NO:666)

[0398] VQISLKALRRIVPLMEQGKRYDEACAEIYGDHYGKKNTEEKIYLPPIPADEIRNPVVLRA

[0399] Nme1Cas9(481-540)(SEQ ID NO:667)

[0400] LSQARKVINGVVRRYGSPARIHIETAREVGKSFKDRKEIEKRQEENRKDREKAAAKFREY

[0401] Nme2Cas9(481-540)(SEQ ID NO:668)

[0402] LSQARKVINGVVRRYGSPARIHIETAREVGKSFKDRKEIEKRQEENRKDREKAAAKFREY

[0403] Nme1Cas9(541-600)(SEQ ID NO:669)

[0404] FPNFVGEPKSKDILKLRLYEQQHGKCLYSGKEINL G RLNEKGYVEIDHALPFSRTWDDSF

[0405] Nme2Cas9(541-600)(SEQ ID NO:670)

[0406] FPNFVGEPKSKDILKLRLYEQQHGKCLYSGKEINL V RLNEKGYVEIDHALPFSRTWDDSF

[0407] Nme1Cas9(601-660)(SEQ ID NO:671)

[0408] NNKVLVLGSENQNKGNQTPYEYFNGKDNSREWQEFKARVETSRFPRSKKQRILLQKFDED

[0409] Nme2Cas9(601-660)(SEQ ID NO:672)

[0410] NNKVLVLGSENQNKGNQTPYEYFNGKDNSREWQEFKARVETSRFPRSKKQRILLQKFDED

[0411] Nme1Cas9(661-720)(SEQ ID NO:673)

[0412] GFKE R NLNDTRYVNRFLCQFVAD RMR LTGKGK K RVFASNGQITNLLRGFWGLRKVRAEND

[0413] Nme2Cas9(661-720)(SEQ ID NO:674)

[0414] GFKE C NLNDTRYVNRFLCQFVAD HIL LTGKGK R RVFASNGQITNLLRGFWGLRKVRAEND

[0415] Nme1Cas9(721-780)(SEQ ID NO:675)

[0416] RHHALDAVVVACSTVAMQQKITRFVRYKEMNAFDGKTIDKETG E VLHQKTHFPQPWEFFA

[0417] Nme2Cas9(721-780)(SEQ ID NO:676)

[0418] RHHALDAVVVACSTVAMQQKITRFVRYKEMNAFDGKTIDKETG K VLHQKTHFPQPWEFFA

[0419] Nme1Cas9(781-840)(SEQ ID NO:677)

[0420] QEVMIRVFGKPDGKPEFEEADT L EKLRTLLAEKLSSRPEAVHEYVTPLFVSRAPNRKMSG

[0421] Nme2Cas9(781-840)(SEQ ID NO:678)

[0422] QEVMIRVFGKPDGKPEFEEADT P EKLRTLLAEKLSSRPEAVHEYVTPLFVSRAPNRKMSG

[0423] Nme1Cas9(841-895)(SEQ ID NO:679)

[0424]

[0425] Nme2Cas9(841-899)(SEQ ID NO:680)

[0426]

[0427] Nme1Cas9(896-950)(SEQ ID NO:681)

[0428]

[0429] Nme2Cas9(900–954)(SEQ ID NO:682)

[0430]

[0431] Nme1Cas9(951-1005)(SEQ ID NO:683)

[0432]

[0433] Nme2Cas9(955-1007)(SEQ ID NO:684)

[0434]

[0435] Nme1Cas9(1006-1063)(SEQ ID NO:685)

[0436]

[0437] Nme2Cas9(1008-1063)(SEQ ID NO:686)

[0438]

[0439] Nme1Cas9(1064-1082)(SEQ ID NO:687)

[0440]

[0441] Nme2Cas9(1064-1082)(SEQ ID NO:688)

[0442]

[0443] Comparison of Nme1Cas9 and Nme3Cas9

[0444] Non-PID aa differences (blue-green - underlined); PID aa differences (yellow - bold with underlined); active site residues (red - bold).

[0445] Nme1Cas9 1

[0446]

[0447] Nme3Cas9 1

[0448]

[0449] Nme1Cas9 51 (SEQ ID NO:691)

[0450] VPKTGDSLAMARRLARSVRRLTRRRAHRLLR T RRLLKREGVLQAANFDEN100

[0451] Nme3Cas9 51 (SEQ ID NO:692)

[0452] VPKTGDSLAMARRLARSVRRLTRRRAHRLLR A RRLLKREGVLQAADFDEN100

[0453] Nme1Cas9 101 (SEQ ID NO:693)

[0454] GLIKSLPNTPWQLRAAALDRKLTPLEWSAVLLHLIKHRGYLSQRKNEGET150

[0455] Nme3Cas9 101(SEQ ID NO:694)

[0456] GLIKSLPNTPWQLRAAALDRKLTPLEWSAVLLHLIKHRGYLSQRKNEGET150

[0457] Nme1Cas9 151(SEQ ID NO:695)

[0458] ADKELGALLKGVAGNAHALQTGDFRTPAELALNKFEKE S GHIRNQRSDYS200

[0459] Nme3Cas9 151(SEQ ID NO:696)

[0460] ADKELGALLKGVADNAHALQTGDFRTPAELALNKFEKE C GHIRNQRGDYS200

[0461] Nme1Cas9 201

[0462] HTFSRKDLQAEL I LLFEKQKEFGNPHVSGGLKEGIETLLMTQRPALSGDA250(SEQ ID NO:697)

[0463] Nme3Cas9 201

[0464] HTFSRKDLQAEL N LLFEKQKEFGNPHVSGGLKEGIETLLMTQRPALSGDA250(SEQ ID NO:698)

[0465] Nme1Cas9 251(SEQ ID NO:699)

[0466] VQKMLGHCTFEPAEPKAAKNTYTAERFIWLTKLNNLRILEQGSERPLTDT300

[0467] Nme3Cas9 251(SEQ ID NO:700)

[0468] VQKMLGHCTFEPAEPKAAKNTYTAERFIWLTKLNNLRILEQGSERPLTDT300

[0469] Nme1Cas9 301(SEQ ID NO:701)

[0470] ERATLMDEPYRKSKLTYAQARKLL G LEDTAFFKGLRYGKDNAEASTLMEM 350

[0471] Nme3Cas9 301(SEQ ID NO:702)

[0472] ERATLMDEPYRKSKLTYAQARKLL S LEDTAFFKGLRYGKDNAEASTLMEM 350

[0473] Nme1Cas9 351

[0474] KAYH A ISRALEKEGLKDKKSPLNLSPELQDEIGTAFSLFKTDEDITGRLK 400(SEQ ID NO:703)

[0475] Nme3Cas9 351

[0476] KAYH T ISRALEKEGLKDKKSPLNLSPELQDEIGTAFSLFKTDEDITGRLK 400(SEQ ID NO:704)

[0477] Nme1Cas9 401

[0478] DRIQPEILEALLKHISFDKFVQISLKALRRIVPLMEQGKRYDEACAEIYG 450(SEQ ID NO:705)

[0479] Nme3Cas9 401

[0480] DRIQPEILEALLKHISFDKFVQISLKALRRIVPLMEQGKRYDEACAEIYG 450(SEQ ID NO:706)

[0481] Nme1Cas9 451

[0482] DHYGKKNTEEKIYLPPIPADEIRNPVVLRALSQARKVINGVVRRYGSPAR500(SEQ ID NO:707)

[0483] Nme3Cas9 451

[0484] DHYGKKNTEEKIYLPPIPADEIRNPVVLRALSQARKVINGVVRRYGSPAR500(SEQ ID NO:708)

[0485] Nme1Cas9 501

[0486] IHIETAREVGKSFKDRKEIEKRQEENRKDREKAAAKFREYFPNFVGEPKS550(SEQ ID NO:709)

[0487] Nme3Cas9 501

[0488] IHIETAREVGKSFKDRKEIEKRQEENRKDREKAAAKFREYFPNFVGEPKS550(SEQ ID NO:710)

[0489] Nme1Cas9 551(SEQ ID NO:711)

[0490]

[0491] Nme3Cas9 551(SEQ ID NO:712)

[0492]

[0493] Nme1Cas9 601(SEQ ID NO:713)

[0494] NNKVLVLGSENQNKGNQTPYEYFNGKDNSREWQEFKARVETSRFPRSKKQ 650

[0495] Nme3Cas9 601(SEQ ID NO:714)

[0496] NNKVLVLGSENQNKGNQTPYEYFNGKDNSREWQEFKARVETSRFPRSKKQ 650

[0497] Nme1Cas9 651(SEQ ID NO:715)

[0498] RILLQKFDEDGFKERNLNDTRYVNRFLCQFVADRMRLTGKGKKRVFASNG700

[0499] Nme3Cas9 651(SEQ ID NO:716)

[0500] RILLQKFDEDGFKERNLNDTRYVNRFLCQFVADRMRLTGKGKKRVFASNG700

[0501] Nme1Cas9 701(SEQ ID NO:717)

[0502] QITNLLRGFWGLRKVRAENDRHHALDAVVVACSTVAMQQKITRFVRYKEM 750

[0503] Nme3Cas9 701(SEQ ID NO:718)

[0504] QITNLLRGFWGLRKVRAENDRHHALDAVVVACSTVAMQQKITRFVRYKEM 750

[0505] Nme1Cas9 751(SEQ ID NO:719)

[0506] NAFDGKTIDKETGEVLHQKTHFPQPWEFFAQEVMIRVFGKPDGKPEFEEA800

[0507] Nme3Cas9 751(SEQ ID NO:720)

[0508] NAFDGKTIDKETGEVLHQKTHFPQPWEFFAQEVMIRVFGKPDGKPEFEEA800

[0509] Nme1Cas9 801(SEQ ID NO:721)

[0510] DT L EKLRTLLAEKLSSRPEAVHEYVTPLFVSRAPNRKMSGQGHMETVKSA850

[0511] Nme3Cas9 801(SEQ ID NO:722)

[0512] DT PEKLRTLLAEKLSSRPEAVHEYVTPLFVSRAPNRKMSGQGHMETVKSA850

[0513] Nme1Cas9 851(SEQ ID NO:723)

[0514] KRLDEGVSVLRVPLTQLKLKDLEKMVNREREPKLYEALKARLEAHKDDPA900

[0515] Nme3Cas9 851(SEQ ID NO:724)

[0516] KRLDEGVSVLRVPLTQLKLKDLEKMVNREREPKLYEALKARLEAHKDDPA900

[0517] Nme1Cas9(SEQ ID NO:725)

[0518] 901KAFAEPFYKYDKAGNRTQQVKAVRVEQVQKTGVWVRNHNGIADNATMVRV 950

[0519] Nme3Cas9(SEQ ID NO:726)

[0520] 901KAFAEPFYKYDKAGNRTQQVKAVRVEQVQKTGVWVRNHNGIADNATMVRV 950

[0521] Nme1Cas9 951(SEQ ID NO:727)

[0522]

[0523] Nme3Cas9 951(SEQ ID NO:728)

[0524] Nme1Cas9 1001(SEQ ID NO:729)

[0525]

[0526] Nme3Cas9 1001(SEQ ID NO:884)

[0527]

[0528] Nme2Cas9 expressed by plasmid (SEQ ID NO:732)

[0529] SV40 NLS (yellow - bold); 3X-HA - label (green - (underlined / bold); cMyc-style NLS (turquoise - no format); connector (magenta - bold italic) and Nme2Cas9 (italic).

[0530]

[0531] AAV expression Nme2Cas9 (SEQ ID NO:733)

[0532] SV40 NLS (yellow - bold); 3X-HA - label (green) Underlined / Bold ); Nucleoplasmic protein-like NLS (red- Underlined ); c-myc NLS (blue-green - no format); connector (magenta - bold italic) and Nme2Cas9 (italic).

[0533]

[0534]

[0535] Recombinant Nme2Cas9 (SEQ ID NO:734)

[0536] SV40 NLS (yellow - bold); nucleoplasmic protein-like NLS (red - Underlined ); Connector (magenta-bold italic) and Nme2Cas9 (italic).

[0537]

[0538]

[0539] Recombinant Nme2Cas9 for RNP delivery in mammalian cells: (SEQ ID NO:735)

[0540] SV40 NLS (yellow - bold); nucleoplasmic protein-like NLS (red - Underlined ); Connector (magenta-bold italic) and Nme2Cas9 (italic).

[0541]

[0542]

[0543] 9. Therapeutic applications

[0544] Although compact Cas9 orthologs have previously been validated for genome editing, including via single AAV delivery, their longer PAMs have limited therapeutic development due to lower target site frequencies compared to the more widely adopted SpyCas9. Furthermore, SauCas9 and its KKH variant, with its more lenient PAM requirements (Kleinstiver et al., 2015), tend to employ off-target editing of some sgRNAs (Friedland et al., 2015; Kleinstiver et al., 2015). These limitations are exacerbated when using target loci requiring editing within narrow sequence windows or precise fragment deletions. We have identified Nme2Cas9 as a compact and highly accurate Cas9 for in vivo genome editing via AAV delivery, with a less restrictive dinucleotide PAM. The development of Nme2Cas9 has significantly expanded the range of genomes that can be edited in vivo, particularly via viral vector delivery. The Nme2Cas9 all-in-one AAV delivery platform established in this study can, in principle, target a wide range of sites as broadly as SpyCas9 (due to the optimal density of N4CC and NGG PAM), without requiring the delivery of two separate vectors to the same target cell. The availability of the catalytically inactivated form of Nme2Cas9 (dNme2Cas9) also promises to expand its applications, such as CRISPRi, CRISPRa, base editing, and related methods (Dominguez et al., 2016; Komor et al., 2017). Furthermore, the ultra-accuracy of Nme2Cas9 enables precise editing of target genes, potentially mitigating safety concerns arising from off-target activities. Contrary to common sense, the higher target site density of Nme2Cas9 (compared to Nme1Cas9) does not necessarily lead to a relative increase in off-target editing. Similar results have been recently reported, with SpyCas9 variants evolving to have shorter PAMs (Hu et al., 2018). Type II-C Cas9 orthologs are typically slower nucleases than SpyCas9 in vitro (Ma et al., 2015; Mir et al., 2018); interestingly, enzymatic principles suggest that reduced epigenetic k cat (Within limits) can improve the on-target relative off-target specificity of RNA-guided nucleases (Bisaria et al., 2017).

[0545] The discovery of Nme2Cas9 and Nme3Cas9 relied on unexplored Cas9 sequences highly correlated (outside of PID) with previously validated orthologs used for human genome editing (Esvelt et al., 2013; Hou et al., 2013; Lee et al., 2016; Amrani et al., 2018). The association of Nme2Cas9 and Nme3Cas9 with Nme1Cas9 offers the added benefit of using the exact same sgRNA scaffold, thus avoiding the need for separate identification and validation of functional tracrRNA sequences. The accelerated evolution of novel PAM-specific sequences in the context of natural CRISPR immunity may reflect selective pressures to recover phages and MGE targeting that have evaded interference through PAM mutations (Deveau et al., 2008; Paez-Espino et al., 2015). Our analysis of AcrIIC5... Smu The observation of inhibition of Nme1Cas9 but not Nme2Cas9 suggests a second, non-mutually exclusive basis for accelerated PID mutations: escape from CRISPR inhibition. We further speculate that accelerated variability may not be limited to PID, but could be due to escape from CRISPR-resistant selective pressures binding to other Cas9 domains. Cas9 inhibitors binding to more conserved regions of Cas9 (e.g., AcrIIC1) may exhibit fewer pathways for mutational escape and thus a broader repressive spectrum (Harrington et al., 2017a). Regardless of the source of the selective pressures driving the co-evolution of Acr and Cas9, the availability of validated Nme2Cas9 inhibitors (e.g., AcrIIC1-4) offers opportunities for additional levels of control over their activity.

[0546] The method used in this study (i.e., searching for rapidly evolving domains in Cas9) can be implemented elsewhere, especially in bacterial species with well-sampled genome sequences. This method can also be applied to other CRISPR-Cas effector proteins, such as Cas12 and Cas13, which have also been developed for genome or transcriptome engineering and other applications. As with Nme1Cas9, this strategy is particularly compelling when using Cas proteins closely associated with orthologs that have demonstrated efficacy in heterologous contexts (e.g., in eukaryotic cells). Applying this method to the meningococcal Cas9 ortholog resulted in a novel genome editing platform, Nme2Cas9, with a unique combination of features promising to accelerate the development of genome editing tools for general and therapeutic applications (compact size, dinucleotide PAM, ultra-accuracy, single AAV delivery, and Acr sensitivity).

[0547] Table 3. Exemplary sequences of plasmids and oligonucleotides disclosed herein are given below.

[0548]

[0549]

[0550]

[0551]

[0552]

[0553]

[0554]

[0555]

[0556]

[0557]

[0558]

[0559]

[0560]

[0561]

[0562]

[0563]

[0564]

[0565]

[0566] RNP delivery for mammalian genome editing

[0567] For RNP experiments, the Neon electroporation system was used exactly as described (Amrani et al., 2018). Briefly, 40 picomol of 3xNLS-Nme2Cas9 and 50 picomol of T7-transcribed sgRNA were combined in buffer R and electroporated using 10 μL of Neon tip. After electroporation, cells were seeded into preheated 24-well plates containing appropriate antibiotic-free medium. Electroporation parameters (voltage, width, pulse number) were 1150 V, 20 ms, 2 pulses for HEK293T cells; and 1000 V, 50 ms, 1 pulse for K562 cells.

[0568] In vivo AAV8.Nme2Cas9+sgRNA delivery and liver tissue processing

[0569] For AAV8 vector injection, 4 x 10⁴ AAV8 vectors were injected into 8-week-old female C57BL / 6NJ mice via the tail vein. 11 One genome copy / mouse, where the sgRNA targets a validated site in Pcsk9 or Rosa26. Mice were sacrificed 28 days after vector administration, and liver tissue was collected for analysis. Liver tissue was fixed overnight in 4% formalin, embedded in paraffin, sectioned, and stained with hematoxylin and eosin (H&E). Blood was drawn from the facial vein at 0, 14, and 28 days post-injection, and serum was separated using a serum separator (BD, catalog number 365967) and stored at -80°C until assay. Infinity was used according to the manufacturer's protocol and as previously described (Ibraheim et al., 2018). TM Serum cholesterol levels were measured using a colorimetric endpoint assay (Thermo-Scientific). For anti-PCSK9 Western blotting, 40 μg of protein from tissue or 2 ng of recombinant mouse PCSK9 protein (R&DSystems, 9258-SE-020) was loaded into... TGX TM On a pre-prepared gel (Bio-Rad). The separated bands were transferred to a PVDF membrane and blocked for 2 hours at room temperature with 5% Blocking-Grade Blocker solution (Bio-Rad). Next, the membrane was incubated overnight with rabbit anti-GAPDH (Abcam ab9485, 1:2,000) or goat anti-PCSK9 (R&D Systems AF3985, 1:400) antibody. The membrane was washed in TBST and incubated for 2 hours at room temperature with horseradish peroxidase (HRP) conjugated goat anti-rabbit (Bio-Rad 1706515, 1:4,000) and donkey anti-goat (R&D Systems HAF109, 1:2,000) secondary antibodies. The membrane was washed again in TBST and incubated using an M35A XOMAT processor (Kodak) with Clarity.TM Western ECL substrate (Bio-Rad) visualization.

[0570] Delivering AAV6.Nme2Cas9 in mouse fertilized eggs

[0571] The fertilized egg was placed in a container containing 3x10 9 Or 3x10 8 GC-mediated incubation of 15 μl drops (4 fertilized eggs per drop) of KSOM (Potassium-Supplemented Simplex Optimized Medium, Millipore, catalog number MR-106-D) containing the AAV6.Nme2Cas9.sgTyr vector was incubated for 5–6 hours. After incubation, the fertilized eggs were washed in M2 and transferred to fresh KSOM for overnight culture. The next day, embryos that had reached the 2-cell stage were transferred to the fallopian tubes of pseudopregnant recipients and allowed to develop to the late stage.

[0572] experiment

[0573] Example I

[0574] Discovery of Cas9 orthologs with different PIDs

[0575] The Nme1Cas9 peptide sequence was used as the query in a BLAST search to find all Cas9 orthologs in the Neisseria meningitidis species. Orthologs with >80% identity to Nme1Cas9 were selected for the remainder of this study. The PIDs were then compared with the Nme1Cas9 PIDs (residues 820-1082) using ClustalW2, and those with mutation clusters in the PIDs were selected for further analysis. A rootless phylogenetic tree of NmeCas9 orthologs was constructed using FigTree (http: / / tree.bio.ed.ac.uk / software / figtree / ).

[0576] Example II

[0577] Cloning, expression, and purification of Cas9 and Acr orthologs

[0578] Table 3 lists examples of plasmids and oligonucleotides used in this study. The PIDs for Nme2Cas9 and Nme3Cas9 were ordered as gBlocks (IDT) to replace the PID for Nme1Cas9 in the bacterial expression plasmid pMSCG7 (Zhang et al., 2015), which encodes Nme1Cas9 with a 6xHis tag, using Gibson Assembly (NEB). As previously described (Pawluk et al., 2016), the constructs were transformed into *E. coli* for expression and purification. Briefly, Rosetta (DE3) cells containing the respective Cas9 plasmids were grown at 37°C to an OD of 0.6. 600 Protein expression was induced at 18°C ​​for 16 hours using 1 mM IPTG. Cells were harvested and lysed by sonication in a lysis buffer supplemented with a mixture of 1 mg / mL lysozyme and protease inhibitor (Sigma) [50 mM Tris-HCl (pH 7.5), 500 mM NaCl, 5 mM imidazole, 1 mM DTT]. The lysates were then run through Ni... 2+ -NTA agarose column (Qiagen), and the bound protein was eluted with 300 mM imidazole and dialyzed into storage buffer [20 mM HEPES-NaOH (pH 7.5), 250 mM NaCl, 1 mM DTT]. For Acr protein, 6xHis-labeled protein was expressed in E. coli strain BL21 Rosetta (DE3). Cells were grown in a shaking incubator at 37°C until an optical density (OD) of 0.6 was reached. 600 Bacterial cultures were cooled to 18°C ​​and protein expression was induced overnight by adding 1 mM IPTG. The next day, cells were harvested and resuspended in lysis buffer supplemented with a mixture of 1 mg / mL lysozyme and protease inhibitors (Sigma), and the protein was purified using the same protocol as for Cas9. The 6xHis tag was removed to isolate unlabeled Acr by incubating resin-bound protein with tobacco etch virus (TEV) protease overnight at 4°C.

[0579] Example III

[0580] In vitro PAM detection assay

[0581] A dsDNA target library with random PAM sequences was generated by overlap PCR, with the forward primer containing a 10-nt random PAM region. The library was purified by gel electrophoresis and in vitro cleavage using purified Cas9 and T7-transcribed sgRNA. A 300 nM Cas9:sgRNA complex was used to cleave the 300 nM target fragment for 1 h at 37 °C in 1X NEBuffer 3.1 (NEB). The reaction mixture was then treated with proteinase K at 50 °C for 10 min and electrophoresed on a 4% agarose / 1xTAE gel. The cleavage products were excised, eluted, and cloned using a modified previously described protocol (Zhang et al., 2012). Briefly, DNA ends were repaired, a non-templated 2'-deoxyadenosine tail was added, and a Y-adaptor was ligated. Following PCR, the products were quantified using the KAPA Library Quantification Kit and sequenced using NextSeq 500 (Illumina) to obtain 75 nt paired-end reads. Sequences were analyzed using a custom script and R.

[0582] Example IV

[0583] Transfection and mammalian genome editing

[0584] Human codon-optimized Nme2Cas9 was cloned into the pCDest2 plasmid backbone previously used for Nme1Cas9 and SpyCas9 expression via Gibson Assembly (Pawluk et al., 2016; Amrani et al., 2018). HEK293T and HEK293T-TLR2.0 cells were transfected as previously described (Amrani et al., 2018). For Hepa1-6 transfection, cells cultured for 24 hours prior to transfection were transfected with 500 ng of the all-in-one AAV.sgRNA.Nme2Cas9 plasmid in 24-well plates (~10⁵ cells / well) using Lipofectamine LTX. For K562 cells stably expressing Nme2Cas9 delivered via a lentiviral vector (see below), 50,000–150,000 cells were electroporated with 500 ng of sgRNA plasmid using 10 μL Neon tip. To measure insertions and deletions in all cells 72 hours after transfection, cells were harvested and genomic DNA was extracted using the DNaesy Blood and Tissue Kit (Qiagen). Targeted loci were amplified by PCR, followed by Sanger sequencing (Genewiz), and analysis was performed using the Desktop Genetics web-based interface (http: / / tide.deskgen.com) via TIDE (Brinkman et al., 2014).

[0585] Example V

[0586] Lentiviral transduction of K562 cells to stably express Nme2Cas9

[0587] As previously described for Nme1Cas9 (Amrani et al., 2018), K562 cells stably expressing Nme2Cas9 were generated. For lentivirus production, lentiviral vectors, along with packaging plasmids (Addgene 12260 and 12259), were co-transfected into HEK293T cells in 6-well plates using the TransIT-LT1 transfection reagent (Mirus Bio). 24 hours later, the culture medium was aspirated from the transfected cells and replaced with 1 mL of fresh DMEM. The next day, the virus-containing supernatant was collected and filtered through a 0.45 μm filter. 10 μL of the undiluted supernatant was mixed with 2.5 μg of polybrene for transduction of ~10 cells in 6-well plates. 6 K562 cells were selected. Transduced cells were selected using culture medium supplemented with 2.5 μg / mL puromycin.

[0588] Example VI

[0589] RNP delivery for mammalian genome editing

[0590] For RNP experiments, the Neon electroporation system was used exactly as described (Amrani et al., 2018). Briefly, 40 picomol of 3xNLS-Nme2Cas9 and 50 picomol of T7-transcribed sgRNA were combined in buffer R and electroporated using 10 μL of Neon tip. After electroporation, cells were seeded into preheated 24-well plates containing appropriate antibiotic-free medium. Electroporation parameters (voltage, width, pulse number) were 1150 V, 20 ms, 2 pulses for HEK293T cells; and 1000 V, 50 ms, 1 pulse for K562 cells.

[0591] Example VII

[0592] GUIDE-seq

[0593] GUIDE-seq experiments were performed as previously described (Tsai et al., 2014), with minor modifications (Bolukbasi et al., 2015a). Briefly, HEK293T cells were transfected with 200 ng of Cas9 plasmid, 200 ng of sgRNA plasmid, and 7.5 pmol of annealed GUIDE-seq oligonucleotides using Polyfect (Qiagen). Alternatively, Hepa1-6 cells were transfected as described above. 72 hours post-transfection, genomic DNA was extracted using the DNeasy Blood and Tissue Kit (Qiagen) according to the manufacturer's protocol. Library preparation and sequencing were performed exactly as previously described (Bolukbasi et al., 2015a). For analysis, all sequences with up to ten mismatches with the target site and those with a C (N4CN) at the fifth PAM position were considered potential off-target sites. Data were analyzed using the Bioconductor software package GUIDEseq version 1.1.17 (Zhu et al., 2017).

[0594] Example VIII

[0595] Targeted deep sequencing and analysis

[0596] We used targeted deep sequencing to confirm the GUIDE-seq results and measure insertion / deletion rates with maximum accuracy. We used two-step PCR amplification to generate DNA fragments for each target and off-target site. For SpyCas9 editing on DS2 and DS6, we selected top-ranked off-target sites based on GUIDE-seq read counts. For SpyCas9 editing on DS4, fewer candidate off-target sites were identified by GUIDE-seq, and only those with NGG (DS4|OT1, DS4|OT3, DS4|OT6) or NGC (DS4|OT2) PAMs were examined by sequencing. In the first step, we used locus-specific primers carrying universal overhangs with ends complementary to the adaptor. Fragments with overhangs were generated using 2x PCR premix (NEB) in the first step. In the second step, the purified PCR products were amplified using universal forward primers and indexed reverse primers. The full-size products (~250 bp) were gel purified and sequenced in paired-end mode on an Illumina MiSeq. As previously described (Pinello et al., 2016; Ibraheim et al., 2018), MiSeq data analysis was performed.

[0597] Example IX

[0598] Off-target analysis using CRISPRseek

[0599] Global off-target prediction for TS25 and TS47 was performed using the Bioconductor software package CRISPRseek. Minor modifications were made to accommodate the characteristics of Nme2Cas9, which is not shared with SpyCas9. Specifically, we used the following changes: gRNA.size = 24, PAM = "NNNNCC", PAM.size = 6, RNA.PAM.pattern = "NNNNCN", and collected candidate off-target sites with fewer than 6 mismatches. The most probable off-target sites were selected based on the number and location of mismatches. Genomic DNA from cells targeted by each respective sgRNA was used to amplify each candidate off-target locus and then analyzed using TIDE.

[0600] Example X

[0601] Mouse strains and embryo collection

[0602] All animal experiments were conducted under the guidance of the Institutional Animal Care and Use Committee (IACUC) of the University of Massachusetts Medical School. Mice were C57BL / 6NJ (Stock No. 005304), obtained from The Jackson Laboratory. All animals were kept in a 12-hour light cycle. The midpoint of the light cycle on the day the mating plug was observed was considered Embryo Day 0.5 (E0.5) of gestation. Fertilized eggs were collected at E0.5 by tearing open the ampulla with forceps and incubating in M2 medium containing hyaluronidase to remove cumulus cells.

[0603] Example XI

[0604] In vivo AAV8.Nme2Cas9+sgRNA delivery and liver tissue processing

[0605] For AAV8 vector injection, 4 x 10⁴ AAV8 vectors were injected into 8-week-old female C57BL / 6NJ mice via the tail vein. 11 One genome copy / mouse, where the sgRNA targets a validated site in Pcsk9 or Rosa26. Mice were sacrificed 28 days after vector administration, and liver tissue was collected for analysis. Liver tissue was fixed overnight in 4% formalin, embedded in paraffin, sectioned, and stained with hematoxylin and eosin (H&E). Blood was drawn from the facial vein at 0, 14, and 28 days post-injection, and serum was separated using a serum separator (BD, catalog number 365967) and stored at -80°C until assay. Infinity was used according to the manufacturer's protocol and as previously described (Ibraheim et al., 2018). TMSerum cholesterol levels were measured using a colorimetric endpoint assay (Thermo-Scientific). For anti-PCSK9 Western blotting, 40 μg of protein from tissue or 2 ng of recombinant mouse PCSK9 protein (R&DSystems, 9258-SE-020) was loaded into... TGX TM On a pre-prepared gel (Bio-Rad). The separated bands were transferred to a PVDF membrane and blocked for 2 hours at room temperature with 5% Blocking-Grade Blocker solution (Bio-Rad). Next, the membrane was incubated overnight with rabbit anti-GAPDH (Abcam ab9485, 1:2,000) or goat anti-PCSK9 (R&D Systems AF3985, 1:400) antibody. The membrane was washed in TBST and incubated for 2 hours at room temperature with horseradish peroxidase (HRP) conjugated goat anti-rabbit (Bio-Rad 1706515, 1:4,000) and donkey anti-goat (R&D Systems HAF109, 1:2,000) secondary antibodies. The membrane was washed again in TBST and incubated using an M35A XOMAT processor (Kodak) with Clarity. TM Western ECL substrate (Bio-Rad) visualization.

[0606] Example XII

[0607] Delivering AAV6.Nme2Cas9 in mouse fertilized eggs

[0608] The fertilized egg was placed in a container containing 3x10 9 Or 3x10 8 GC-mediated incubation of 15 μl drops (4 fertilized eggs per drop) of KSOM (Potassium-Supplemented Simplex Optimized Medium, Millipore, catalog number MR-106-D) containing the AAV6.Nme2Cas9.sgTyr vector was incubated for 5–6 hours. After incubation, the fertilized eggs were washed in M2 and transferred to fresh KSOM for overnight culture. The next day, embryos that had reached the 2-cell stage were transferred to the fallopian tubes of pseudopregnant recipients and allowed to develop to the late stage.

[0609] The references are all incorporated into this paper through citation:

[0610] Amrani,N.,Gao,XD,Liu,P,Edraki,A,Mir,A,Ibraheim,R,Gupta,A,Sasaki,KE,Wu,T,Donohoue,PD,People(2018).NmeCas9is an intrinsically high-fidelity genome editing platform.BioRxiv,https: / / doi.org / 10.1101 / 172650.

[0611] Barrangou, R., Fremaux, C., Deveau, H., Richards, M., Boyaval, P., Monkey, S., Romero, DA, and Horvath, P. (2007).

[0612] http: / / dx.doi.org / 10.1037 / 0021-843X.101.2.213 Bisaria, N., Jarmoskaite, I., & Herschlag, D. (2017).

[0613] Bolukbasi, MF, Gupta, A., Oikemus, S., Derr, AG, Garber, M., Brodsky, MH, Zhu, LJ, and Wolfe, SA(2015a).

[0614] Bolukbasi, M.F., Gupta, A., and Wolfe, S.A. (2015b). Creating and evaluating accurate CRISPR-Cas9 scalpels for genomic surgery. Nat. Methods 13, 41-50.

[0615] Brinkman, E.K., Chen, T., Amendola, M., and van Steensel, B. (2014). Easy quantitative assessment of genome editing by sequence trace decomposition. Nucleic Acids Res. 42, e168.

[0616] Brouns, S.J., Jore, M.M., Lundgren, M., Westra, E.R., Slijkhuis, R.J., Snijders, A.P., Dickman, M.J., Makarova, K.S., Koonin, E.V., and van der Oost, J. (2008). Small CRISPR RNAs guide antiviral defense in prokaryotes. Science 321, 960-964.

[0617] Casini, A., Olivieri, M., Petris, G., Montagna, C., Reginato, G., Maule, G., Lorenzin, F., Prandi, D., Romanel, A., Demichelis, F., et al. (2018). A highly specific SpCas9 variant is identified by in vivo screening in yeast. Nat. Biotechnol. 36, 265-271.

[0618] Certo, M.T., Ryu, B.Y., Annis, J.E., Garibov, M., Jarjour, J., Rawlings, D.J., and Scharenberg, A.M. (2011). Tracking genome engineering outcome at individual DNA breakpoints. Nat. Methods 8, 671 - 676.

[0619] Chen, J.S., Dagdas, Y.S., Kleinstiver, B.P., Welch, M.M., Sousa, A.A., Harrington, L.B., Sternberg, S.H., Joung, J.K., Yildiz, A., and Doudna, J.A. (2017). Enhanced proofreading governs CRISPR - Cas9 targeting accuracy. Nature 550, 407 - 410.

[0620] Cho, S.W., Kim, S., Kim, J.M., and Kim, J.S. (2013). Targeted genome engineering in human cells with the Cas9 RNA - guided endonuclease. Nat. Biotechnol. 31, 230 - 232.

[0621] Cho, S.W., Kim, S., Kim, Y., Kweon, J., Kim, H.S., Bae, S., and Kim, J.S. (2014). Analysis of off - target effects of CRISPR / Cas - derived RNA - guided endonucleases and nickases. Genome Res. 24, 132 - 141.

[0622] Cong, L., Ran, F.A., Cox, D., Lin, S., Barretto, R., Habib, N., Hsu, P.D., Wu, X., Jiang, W., Marraffini, L.A., et al. (2013). Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819 - 823.

[0623] Deltcheva, E., Chylinski, K., Sharma, C. M., Gonzales, K., Chao, Y., Pirzada, Z. A., Eckert, M. R., Vogel, J., and Charpentier, E. (2011). CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III. Nature 471, 602 - 607.

[0624] Deveau, H., Barrangou, R., Garneau, J. E., Labonte, J., Fremaux, C., Boyaval, P., Romero, D. A., Horvath, P., and Moineau, S. (2008). Phage response to CRISPR-encoded resistance in Streptococcus thermophilus. J. Bacteriol. 190, 1390 - 1400.

[0625] Dominguez, A. A., Lim, W. A., and Qi, L. S. (2016). Beyond editing: repurposing CRISPR-Cas9 for precision genome regulation and interrogation. Nat. Rev. Mol. Cell Biol. 17, 5 - 15.

[0626] Dong, Guo, M., Wang, S., Zhu, Y., Wang, S., Xiong, Z., Yang, J., Xu, Z., and Huang, Z. (2017). Structural basis of CRISPR-SpyCas9 inhibition by an anti-CRISPR protein. Nature 546, 436 - 439.

[0627] Esvelt, K.M., Mali, P., Braff, J.L., Moosburner, M., Yaung, S.J., and Church, G.M. (2013). Orthogonal Cas9 proteins for RNA-guided gene regulation and editing. Nat. Methods 10, 1116-1121.

[0628] Fonfara, I., Le Rhun, A., Chylinski, K., Makarova, K.S., Lecrivain, A.L., Bzdrenga, J., Koonin, E.V., and Charpentier, E. (2014). Phylogeny of Cas9 determines functional exchangeability of dual-RNA and Cas9 among orthologous type II C RISPR-Cas systems. Nucleic Acids Res. 42, 2577-2590.

[0629] Friedland, A.E., Baral, R., Singhal, P., Loveluck, K., Shen, S., Sanchez, M., Marco, E., Gotta, G.M., Maeder, M.L., Kennedy, E.M., et al. (2015). Characterization of Staphylococcus aureus Cas9: a smaller Cas9 for all-in-one adeno-associated virus delivery and paired nickase applications. Genome Biol. 16, 257.

[0630] Friedrich, G., and Soriano, P. (1991). Promoter traps in embryonic stem cells: a genetic screen to identify and mutate developmental genes in mice. Genes Dev. 5, 1513-1523.

[0631] Fu, Y., Sander, J. D., Reyon, D., Cascio, V. M., and Joung, J. K. (2014). Improving CRISPR-Cas nuclease specificity using truncated guide RNAs. Nat. Biotechnol. 32, 279-284.

[0632] Gallagher, D. N., and Haber, J. E. (2018). Repair of a Site-Specific DNA Cleavage: Old-School Lessons for Cas9-Mediated Gene Editing. ACS Chem. Biol. 13, 397-405.

[0633] Garneau, J. E., Dupuis, M. E., Villion, M., Romero, D. A., Barrangou, R., Boyaval, P., Fremaux, C., Horvath, P., Magadan, A. H., and Moineau, S. (2010). The CRISPR / Cas bacterial immune system cleaves bacteriophage and plasmid DNA. Nature 468, 67-71.

[0634] Gasiunas, G., Barrangou, R., Horvath, P., and Siksnys, V. (2012). Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria. Proc. Natl. Acad. Sci. USA 109, E2579-2586.

[0635] Gaudelli, N. M., Komor, A. C., Rees, H. A., Packer, M. S., Badran, A. H., Bryson, D. I., and Liu, D. R. (2017). Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage. Nature 551, 464-471.

[0636] Ghanta, K., Dokshin, G., Mir, A., Krishnamurthy, P., Gneid, H., Edraki, A., Watts, J., Sontheimer, E., and Mello, C. (2018). 5′ Modifications Improve Potency and Efficacy of DNA Donors for Precision Genome Editing. Biorxiv 354480.

[0637] Gorski, S. A., Vogel, J., and Doudna, J. A. (2017). RNA-based recognition and targeting: sowing the seeds of specificity. Nat. Rev. Mol. Cell Biol. 18, 215 - 228.

[0638] Harrington, L. B., Doxzen, K. W., Ma, E., Liu, J. J., Knott, G. J., Edraki, A., Garcia, B., Amrani, N., Chen, J. S., Cofsky, J. C., et al. (2017a). A Broad-Spectrum Inhibitor of CRISPR-Cas9. Cell 170, 1224 - 1233.

[0639] Harrington, L. B., Paez-Espino, D., Staahl, B. T., Chen, J. S., Ma, E., Kyrpides, N. C., and Doudna, J. A. (2017b). A thermostable Cas9 with increased lifetime in human plasma. Nat. Commun. 8, 1424.

[0640] Hou, Z., Zhang, Y., Propson, NE, Howden, SE, Chu, LF, Sontheimer, EJ, and Thomson, JA (2013).Efficient genome engineering in human pluripotent stem cells using Cas9 from Neisseria meningitidis.

[0641] Hu, JH, Miller, SM, Geurts, MH, Tang, W., Chen, L., Sun, N., Zeina, CM, Gao, X., Rees, HA, Lin, Z., Account(2018).Evolved Cas9 variants with broad PAMcompatibility and high DNA specificity.Nature 556,57-63.

[0642] Hwang, WY, Fu, Y., Reyon, D., Maeder, ML, Tsai, SQ, Sander, JD, Peterson, RT, Yeh, JR, and Joung, JK(2013).Efficient genome editing in zebrafish using the CRISPR-Cas system.Nat.Biotechnol.31,227-229.

[0643] Hynes,AP,Rousseau,GM,Lemay,M.-L.,Horvath,P.,Romero,DA,Fremaux,C.,andMoineau,S.(2017).An anti-CRISPR from a virulent streptococcal phageinhibits Streptococcus pyogenes Cas9.Nat.Microbiol.2,1374-1380.

[0644] Ibraheim, R., Song, C.-Q., Mir, A., Amrani, N., Xue, W., and Sontheimer, E. J. (2018). All-in-One Adeno-associated Virus Delivery and Genome Editing by Neisseria meningitidis Cas9 in vivo. BioRxiv, https: / / doi.org / 10.1101 / 295055.

[0645] Jiang, F., and Doudna, J. A. (2017). CRISPR-Cas9 Structures and Mechanisms. Annu. Rev. Biophys. 46, 505-529.

[0646] Jiang, W., Bikard, D., Cox, D., Zhang, F., and Marraffini, L. A. (2013). RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nat. Biotechnol. 31, 233-239.

[0647] Jinek, M., Chylinski, K., Fonfara, I., Hauer, M., Doudna, J. A., and Charpentier, E. (2012). A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science 337, 816-821.

[0648] Jinek, M., East, A., Cheng, A., Lin, S., Ma, E., and Doudna, J. (2013). RNA-programmed genome editing in human cells. eLife 2, e00471.

[0649] Karvelis, T., Gasiunas, G., Young, J., Bigelyte, G., Silanskas, A., Cigan, M., and Siksnys, V. (2015). Rapid characterization of CRISPR-Cas9 protospacer adjacent motif sequence elements. Genome Biol. 16, 253.

[0650] Keeler, A.M., ElMallah, M.K., and Flotte, T.R. (2017). Gene Therapy 2017: Progress and Future Directions. Clin. Transl. Sci. 10, 242 - 248.

[0651] Kim, E., Koo, T., Park, S.W., Kim, D., Kim, K.-E., Kim, K., Cho, H.-Y., Song, D.W., Lee, K.J., Jung, M.H., et al. (2017). In vivo genome editing with a small Cas9 ortholog derived from Campylobacter jejuni. Nat. Commun. 8, 14500.

[0652] Kim, S., Kim, D., Cho, S.W., Kim, J., and Kim, J.S. (2014). Highly efficient RNA-guided genome editing in human cells via delivery of purified Cas9 ribonucleoproteins. Genome Res. 24, 1012 - 1019.

[0653] Kim, B., Komor, A., Levy, J., Packer, M., Zhao, K., and Liu, D. (2017). Increasing the genome-targeting scope and precision of base editing with engineered Cas9-cytidine deaminase fusions. Nature Biotechnology 35.

[0654] Kleinstiver, B.P., Prew, M.S., Tsai, S.Q., Nguyen, N.T., Topkar, V.V., Zheng, Z., and Joung, J.K. (2015). Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition. Nat. Biotechnol. 33, 1293-1298.

[0655] Kluesner, M., Nedveck, D., Lahr, W., Garbe, J., Abrahante, J., Webber, B., and Moriarity, B. (2018). EditR: A Method to Quantify Base Editing from Sanger Sequencing. The CRISPR Journal 1, 239-250.

[0656] Koblan, L., Doman, J., Wilson, C., Levy, J., Tay, T., Newby, G., Maianti, J., Raguram, A., and Liu, D. (2018). Improving cytidine and adenine base editors by expression optimization and ancestral reconstruction. Nat Biotechnol 36, 843.

[0657] Komor, A.C., Badran, A.H., and Liu, D.R. (2017). CRISPR-Based Technologies for the Manipulation of Eukaryotic Genomes. Cell 168, 20-36.

[0658] Komor, A.C., Kim, Y.B., Packer, M.S., Zuris, J.A., and Liu, D.R. (2016). Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420-424.

[0659] Lee, C.M., Cradick, T.J., and Bao, G. (2016). The Neisseria meningitidis CRISPR-Cas9 system enables specific genome editing in mammalian cells. Mol. Ther. 24, 645-654.

[0660] Lee, J., Mir, A., Edraki, A., Garcia, B., Amrani, N., Lou, H.E., Gainetdinov, I., Pawluk, A., Ibraheim, R., Gao, X.D., et al. (2018). Potent Cas9 inhibition in bacterial and human cells by new anti-CRISPR protein families. BioRxiv, https: / / www.biorxiv.org / content / early / 2018 / 2006 / 2020 / 350504.

[0661] Ma, E., Harrington, L.B., O'Connell, M.R., Zhou, K., and Doudna, J.A. (2015). Single-Stranded DNA Cleavage by Divergent CRISPR-Cas9 Enzymes. Mol. Cell 60, 398-407.

[0662] Mali, P., Aach, J., Stranges, P.B., Esvelt, K.M., Moosburner, M., Kosuri, S., Yang, L., and Church, G.M. (2013a). CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nat. Biotechnol. 31, 833-838.

[0663] Mali, P., Yang, L., Esvelt, K. M., Aach, J., Guell, M., DiCarlo, J. E., Norville, J. E., and Church, G. M. (2013b). RNA-guided human genome engineering via Cas9. Science 339, 823 - 826.

[0664] Marraffini, L. A., and Sontheimer, E. J. (2008). CRISPR interference limits horizontal gene transfer in staphylococci by targeting DNA. Science 322, 1843 - 1845.

[0665] Mir, A., Edraki, A., Lee, J., and Sontheimer, E. J. (2018). Type II-C CRISPR-Cas9 biology, mechanism and application. ACS Chem. Biol. 13, 357 - 365.

[0666] Mojica, F. J., Diez-Villasenor, C., Garcia-Martinez, J., and Almendros, C. (2009). Short motif sequences determine the targets of the prokaryotic CRISPR defence system. Microbiology 155, 733 - 740.

[0667] Paez-Espino, D., Sharon, I., Morovic, W., Stahl, B., Thomas, B. C., Barrangou, R., and Banfield, J. F. (2015). CRISPR immunity drives rapid phage genome evolution in Streptococcus thermophilus. mBio 6.

[0668] Pawluk,A.,Amrani,N.,Zhang,Y.,Garcia,B.,Hidalgo-Reyes,Y.,Lee,J.,Edraki,A.,Shah,M.,Sontheimer,EJ,Maxwell,KL,People(2016).Naturally occurringoff-switches for CRISPR-Cas9.Cell 167,1829-1838and1829.

[0669] Pawluk, A., Bondy-Denomy, J., Cheung, VH, Maxwell, KL, and Davidson, AR(2014).mBio 5,e00896.

[0670] Pinello, L., Canver, MC, Hoban, MD, Orkin, SH, Kohn, DB, Bauer, DE, and Yuan, GC(2016).Analyzing CRISPR Genome-Editing Experiments withCRISPResso.Nat.Biotechnol.34,695-697.

[0671] Racanelli, V., and Rehermann, B. (2006).The liver as an immunological organ.Hepatology 43,S54-62.

[0672] Ran,FA,Cong,L.,Yan,WX,Scott,DA,Gootenberg,JS,Kriz,AJ,Zetsche,B.,Shalem,O.,Wu,X.,Makarova,KS,People(2015).In vivo genome editing of Staphylococcus aureus Cas9.Nature 520,186-191.

[0673] Ran, F.A., Hsu, P.D., Lin, C.Y., Gootenberg, J.S., Konermann, S., Trevino, A.E., Scott, D.A., Inoue, A., Matoba, S., Zhang, Y., et al. (2013). Double nicking by RNA-guided CRISPR Cas9 for enhanced genome editing specificity. Cell 154, 1380-1389.

[0674] Rashid, S., Curtis, D.E., Garuti, R., Anderson, N.N., Bashmakov, Y., Ho, Y.K., Hammer, R.E., Moon, Y.A., and Horton, J.D. (2005). Decreased plasma cholesterol and hypersensitivity to statins in mice lacking Pcsk9. Proc. Natl. Acad. Sci. USA 102, 5374-5379.

[0675] Rauch, B.J., Silvis, M.R., Hultquist, J.F., Waters, C.S., McGregor, M.J., Krogan, N.J., and Bondy-Denomy, J. (2017). Inhibition of CRISPR-Cas9 with Bacteriophage Proteins. Cell 168, 150-158e110.

[0676] Sapranauskas, R., Gasiunas, G., Fremaux, C., Barrangou, R., Horvath, P., and Siksnys, V. (2011). The Streptococcus thermophilus CRISPR / Cas system provides immunity in Escherichia coli. Nucleic Acids Res. 39, 9275-9282.

[0677] Schumann, K., Lin, S., Boyer, E., Simeonov, DR, Subramaniam, M., Gate, RE, Haliburton, GE, Ye, CJ, Bluestone, JA, Doudna, JA, Specialist(2015).Generation ofknock-in primary human T cells using Cas9 ribonucleoproteins.Proc.Natl.Acad.Sci.USA 112,10437–10442.

[0678] Shin,J.,Jiang,F.,Liu,JJ,Bray,NL,Rauch,BJ,Baik,SH,Walnut,E,Bondy-Denomy,J,Corn,JE,&Doudna,JA(2017).Disabling Cas9 by an anti-CRISPR DNA mimic.Sci.Adv.3,e1701620.

[0679] http: / / dx.doi.org / 10.1037 / 0021-843X.111.2.113 Tsai, SQ, & Joung, JK(2016).

[0680] Tsai,SQ,Zheng,Z.,Nguyen,NT,Liebers,M.,Topkar,VV,Thapar,V.,Wyvekens,N.,Khayter,C.,Iafrate,AJ,Le,LP,Expert(2014).GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Casnucleases.Nat.Biotechnol.33,187-197.

[0681] Tycko, J., Myer, VE, and Hsu, PD(2016).Methods for optimizing CRISPR-Cas9 genome editing specificity.Mol.Cell 63,355-370.

[0682] Yang, H., and Patel, D. J. (2017). Inhibition Mechanism of an Anti-CRISPR Suppressor AcrIIA4 Targeting SpyCas9. Mol Cell 67, 117 - 127e115.

[0683] Yin, H., Song, C. Q., Suresh, S., Kwan, S. Y., Wu, Q., Walsh, S., Ding, J., Bogorad, R. L., Zhu, L. J., Wolfe, S. A., et al. (2018). Partial DNA-guided Cas9 enables genome editing with reduced off-target activity. Nat. Chem. Biol. 14, 311 - 316.

[0684] Yokoyama, T., Silversides, D. W., Waymire, K. G., Kwon, B. S., Takeuchi, T., and Overbeek, P. A. (1990). Conserved cysteine to serine mutation in tyrosinase is responsible for the classical albino mutation in laboratory mice. Nucleic Acids Res. 18, 7293 - 7298.

[0685] Yoon, Y., Wang, D., Tai, P. W. L., Riley, J., Gao, G., and Rivera-Perez, J. A. (2018). Streamlined ex vivo and in vivo genome editing in mouse embryos using recombinant adeno-associated viruses. Nat. Commun. 9, 412.

[0686] Zhang, Y., Heidrich, N., Ampattu, B. J., Gunderson, C. W., Seifert, H. S., Schoen, C., Vogel, J., and Sontheimer, E. J. (2013). Processing-independent CRISPR RNAs limit natural transformation in Neisseria meningitidis. Mol. Cell 50, 488 - 503.

[0687] Zhang, Y., Rajan, R., Seifert, H. S., Mondragón, A., and Sontheimer, E. J. (2015). DNase H activity of Neisseria meningitidis Cas9. Mol. Cell 60, 242 - 255.

[0688] Zhang, Z., Theurkauf, W. E., Weng, Z., and Zamore, P. D. (2012). Strand-specific libraries for high throughput RNA sequencing (RNA-Seq) prepared without poly(A) selection. Silence 3, 9.

[0689] Zhu, L. J., Holmes, B. R., Aronin, N., and Brodsky, M. H. (2014). CRISPRseek: a bioconductor package to identify target-specific guide RNAs for CRISPR-Cas9 genome-editing systems. PLoS One 9, e108424.

[0690] Zhu, L. J., Lawrence, M., Gupta, A., Pagés, H., Kucukural, A., Garber, M., and Wolfe, S. A. (2017). GUIDEseq: a bioconductor package to analyze GUIDE-Seq datasets for CRISPR-Cas nucleases. BMC Genomics 18, 379.

[0691] Zuris, JA, Thompson, DB, Shu, Y., Guilinger, JP, Bessen, JL, Hu, JH, Maeder, ML, Joung, JK, Chen, Z.-Y., and Liu, DR (2015). Cationic lipid-mediated delivery of proteins enables efficient protein-based genome editing in vitro and in vivo. Nat. Biotechnol. 33, 73-80.

[0692] All publications and patents mentioned in the foregoing specification are incorporated herein by reference. Various modifications and variations to the methods and systems described herein will be apparent to those skilled in the art without departing from the scope and spirit of the invention. Although the invention has been described in conjunction with specific preferred embodiments, it should be understood that the claimed invention should not be unduly limited to such specific embodiments. Indeed, various modifications to the described methods for carrying out the invention that will be apparent to those skilled in the art of biological control, biochemistry, molecular biology, entomology, plankton, fisheries systems, and freshwater ecology, or related fields, are within the scope of the following claims.

Claims

1. A fusion protein comprising an Nme2Cas9 cleavage enzyme protein and a nucleotide deaminase, wherein the Nme2Cas9 cleavage enzyme protein comprises a pre-intercalation sequence neighbor motif interaction domain, wherein the pre-intercalation sequence neighbor motif interaction domain binds to an N4CC pre-intercalation sequence neighbor motif, and wherein the Nme2Cas9 cleavage enzyme protein is Nme2Cas9 D16A Cleavage enzyme protein or Nme2Cas9 H588A Cutting enzyme protein.

2. The fusion protein of claim 1, further comprising a nuclear localization signaling protein.

3. The fusion protein of claim 1, wherein the nucleotide deaminase is cytidine deaminase.

4. The fusion protein of claim 1, wherein the nucleotide deaminase is adenosine deaminase.

5. The fusion protein of claim 1, further comprising a uracil glycosylation inhibitor.

6. The fusion protein of claim 2, wherein the nuclear localization signaling protein is selected from nucleoplasmic proteins and SV40.

7. The fusion protein of claim 1, wherein the preinterstitial sequence adjacent motif interaction domain contains a mutation.

8. An adeno-associated virus encoding a fusion protein comprising an Nme2Cas9 cleavage enzyme protein and a nucleotide deaminase, wherein the Nme2Cas9 cleavage enzyme protein comprises a pre-interstitial sequence adjacent motif interaction domain, wherein the pre-interstitial sequence adjacent motif interaction domain binds to an N4CC pre-interstitial sequence adjacent motif, and wherein the Nme2Cas9 cleavage enzyme protein is Nme2Cas9 D16A Cleavage enzyme protein or Nme2Cas9 H588A Cutting enzyme protein.

9. The virus of claim 8, wherein the virus is adeno-associated virus 8.

10. The virus of claim 8, wherein the virus is adeno-associated virus 6.

11. The virus of claim 8, wherein the Nme2Cas9 nickase protein further comprises a nuclear localization signal protein.

12. The virus of claim 8, wherein the nucleotide deaminase is cytidine deaminase.

13. The virus of claim 8, wherein the nucleotide deaminase is adenosine deaminase.

14. The virus of claim 8, wherein the fusion protein further comprises a uracil glycosylation inhibitor.

15. The virus of claim 11, wherein the nuclear localization signal protein is selected from nucleoplasmic proteins and SV40.

16. A nucleic acid molecule encoding the fusion protein of claim 1.

Citation Information

Patent Citations

  • Transistor circuit arrangement for relays

    GB1000451A

  • Improvements in the manufacture of phosphoric acid

    GB1000453A

  • Fusions of CAS9 domains and nucleic acid-editing domains

    US20150166980A1

  • Nucleobase editors and uses thereof

    US20170121693A1

  • Enzyme amplification assay

    US3817837A