A high-fidelity cas protein and uses thereof
The SuperFi-Cas9 system addresses the trade-off in CRISPR systems by using a 5′ extended gRNA and mutated Cas9 protein to enhance editing efficiency and fidelity, enabling precise nucleic acid editing in mammalian cells.
Patent Information
- Application Number
- PCT/CN2024/085903
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-03
- Publication Date
- 2025-10-09
AI Technical Summary
Existing CRISPR systems, particularly high-fidelity Cas nucleases, face a trade-off between high fidelity and editing efficiency, limiting their potential in gene editing applications.
Development of a SuperFi-Cas9 system with a 5′ extended guide RNA (gRNA) that enhances editing efficiency while maintaining high fidelity, utilizing a SuperFi-Cas9 protein with specific mutations in the RuvC loop to stabilize target recognition and reduce off-target editing.
The SuperFi-Cas9 system achieves effective nucleic acid editing in mammalian cells with improved editing efficiency and reduced off-target effects, making it suitable for precise gene editing and therapeutic applications.
Smart Images

Figure CN2024085903_09102025_PF_FP_ABST
Abstract
Description
A High-Fidelity Cas Protein and Uses Thereof
[0001] FIELD OF DISCLOSURE
[0002] The present disclosure relates to a high-fidelity Cas protein and a gene editing system comprising a high-fidelity Cas protein and guide RNAs and uses thereof.
[0003] SEQUENCE LISTING
[0004] This application contains a Sequence Listing electronically submitted as an XML file entitled “Seq. xml” having a size of 94KB and created on April 3, 2024. The information contained in the Sequence Listing is incorporated by reference herein.BACKGROUND
[0005] Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) are a family of DNA sequences found in prokaryotic organisms such as bacteria and archaea. They consist of a series of short, highly conserved direct repeat (DR) sequences interspersed with similarly sized spacer sequences. These CRISPR sequences are transcribed and processed by associated proteins to produce CRISPR RNA (crRNA) . In the vicinity of CRISPR sequences, there are a series of conserved CRISPR-associated genes (Cas genes) . The proteins encoded by Cas genes (Cas proteins) contain nucleic acid-related functional domains. Cas proteins and crRNA collaborate to participate in the prokaryotic CRISPR immune defense process.
[0006] Based on the core functional elements of Cas genes, the CRISPR system can be classified into three major types. In the first and third classes of CRISPR systems, the formation of a complex of multiple Cas proteins is required to degrade foreign nucleic acids. While in the second class of CRISPR systems, a single effector protein is sufficient to achieve this goal. Therefore, the second class of CRISPR system, as a simpler nucleic acid targeting tool, has been engineered to become an important tool for gene editing, playing a crucial role in the fields of gene therapy and disease research.
[0007] The second class of CRISPR systems can be further categorized into three types: Type II, Type V, and Type VI. The most well-characterized endonuclease is the Cas9 endonuclease derived from Streptococcus pyogenes (SpCas9) , that uses a single guide RNA to target nucleic acids. While Type II and other CRISPR systems have enormous potential to allow for targeting of previously undruggable targets and for the editing of abnormal nucleic acids in vivo or ex vivo, safety and efficacy remain a concern.
[0008] In its original function in bacterial adaptive immunity, CRISPR benefits from lower specificity than would be desired in a targeted application of CRISPR, allowing for a certain level of flexibility during pathogen recognition. Thus, many efforts have been made to amend CRISPR systems to improve specificity by developing high-fidelity Cas nucleases, optimizing the design and selection of gRNA, and fine-tuning the time window of editing.
[0009] Multiple research groups have focused on developing high-fidelity Cas nucleases to reduce off-target editing events, and a variety of nucleases have been reported. These high-fidelity Cas9 enzymes were developed through methods such as structure-guided engineering1-3 and high throughput screening4-6. These methods aim to modify the structure and function of Cas9 nuclease to introduce mutations to weaken the interactions between the protein and the targeted or untargeted DNA strand or weaken the intermolecular interactions between the domains of the protein7. Some examples of high-fidelity Cas9 enzymes include eSpCas92, HypaCas93, Sniper-Cas95, HiFi Cas96, LZ3 Cas98, and SpCas9-NG9.
[0010] Although the use of high-fidelity Cas nuclease can reduce off-target editing, it usually comes with a significant decrease in editing efficiency. In a paper published in 2022, kinetics-guided cryo-electron microscopy was used to investigate a special intermediate conformation formed by Cas9, gRNA, and mismatched targets, which prevented Cas9 activation10. This discovery pinpointed residues within the RuvC loop that contact and stabilize a distorted duplex between gRNA and a mismatched target containing 18-20 mismatches, which are absent in on-target structures. This understanding of the structural basis of Cas9 editing of mismatched targets guided the authors to design a proof-of-concept high-fidelity enzyme, SuperFi-Cas9, in which they mutated seven residues to aspartic acid (7-D mutant) within the RuvC loop that are responsible for mismatch stabilization. The SuperFi-Cas9 endonuclease was later reported to show low editing activity in mammalian cells, across 26 endogenous sites in cell line, despite being considered a high-fidelity Cas9 protein7.
[0011] Thus, there is a need for developing a CRISPR system that exhibits both high fidelity and high editing efficiency if the potential of CRISPR is to be fully realized. The present disclosure addresses this need.SUMMARY
[0012] In one aspect, the present disclosure provides a new SuperFi-Cas9 system. The present disclosure shows that the SuperFi-Cas9 system described herein is capable of cleaving nucleic acids effectively in mammalian cells. The present disclosure provides new options for nucleic acid editing tools.
[0013] In some embodiments, the present disclosure provides a system comprising: (a) a SuperFi-Cas9 protein or variant thereof; and (b) a guide RNA (gRNA) comprising a polynucleotide sequence encoding a spacer that is complementary to a target nucleic acid sequence and 20 nucleotides in length, wherein the spacer further comprises a 5′ extension. In some embodiments, the SuperFi-Cas9 protein comprises the amino acid sequence of SEQ ID NO: 1, or wherein the variant thereof comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%identity to SEQ ID NO: 1. In some embodiments, the 5′ extension is 1 nucleotide in length. In some embodiments, the 5′ extension is 2 nucleotides in length. In some embodiments, the 5′ extension is 3 nucleotides in length. In some embodiments, the 5′ extension is 4 nucleotides in length. In some embodiments, the 5′ extension is complementary to the target nucleic acid sequence. In some embodiments, the gRNA is capable of binding to SuperFi-Cas9 or a variant thereof. In some embodiments, the 5′ extension does not decrease SuperFi-Cas9 binding to the gRNA relative to SuperFi-Cas9 binding to a gRNA that does not comprise a 5′ extension.
[0014] In some embodiments, the present disclosure provides a vector comprising: (a) a polynucleotide encoding a SuperFi-Cas9 protein or variant thereof as described herein; and (b) a polynucleotide encoding a gRNA as described herein. In some embodiments, the vector further comprises at least one regulatory element operably linked to (a) , (b) , or (a) and (b) . In some embodiments, the vector is a plasmid or a viral vector.
[0015] In some embodiments, the present disclosure provides a ribonucleoprotein comprising: (a) a SuperFi-Cas9 protein or variant thereof as described herein; and (b) a polynucleotide encoding a gRNA as described herein.
[0016] In some embodiments, the present disclosure provides a system comprising: (a) a SuperFi-Cas9 protein or variant thereof as described herein; (b) a polynucleotide encoding a gRNA as described herein or a prime editing guide RNA (pegRNA) ; and (c) one or more effector molecules. In some embodiments, the pegRNA comprises a spacer that is complementary to a target nucleic acid sequence and wherein the spacer is 20, 21, 22, 23, or 24 nucleotides in length. In some embodiments, the pegRNA is capable of binding to SuperFi-Cas9 or a variant thereof. In some embodiments, the one or more effector molecules are selected from a ribonuclease, a nickase, a base editor, an epigenetic modifier, a transposase, a recombinase, and a reverse transcriptase. In some embodiments, the one or more effector molecules comprise a base editor. In some embodiments, the base editor is a cytosine deaminase or an adenine deaminase. In some embodiments, the system comprises a SuperFi-Adenine base editor (SuperFi-ABE) comprising an amino acid sequence of SEQ ID NO: 2, or having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%identity to SEQ ID NO: 2. In some embodiments, the system comprises a SuperFi-Cytosine base editor (SuperFi-CBE) comprising an amino acid sequence of SEQ ID NO: 3, or having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identity to SEQ ID NO: 3. In some embodiments, the system comprises a pegRNA and the one or more effector molecules comprise a reverse transcriptase. In some embodiments, the system comprises a SuperFi-Prime editor (SuperFi-PE) having an amino acid sequence of SEQ ID NO: 4, or having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%identity to SEQ ID NO: 4.
[0017] In some embodiments, the present disclosure provides a vector comprising: (a) a polynucleotide encoding a SuperFi-Cas9 protein or variant thereof as described herein; and (b) a polynucleotide encoding a gRNA as described herein or a pegRNA as described herein; and (c) a polynucleotide encoding one or more effector molecules as described herein. In some embodiments, the vector further comprises at least one regulatory element operably linked to one or more of (a) - (c) . In some embodiments, the vector is a plasmid or a viral vector.
[0018] In some embodiments, the present disclosure provides a ribonucleoprotein comprising: (a) a SuperFi-Cas9 protein or variant thereof as described herein; (b) a polynucleotide encoding a gRNA as described herein or a pegRNA as described herein; and (c) one or more effector molecules as described herein.
[0019] In some embodiments, the present disclosure provides a method for targeting a nucleic acid sequence, comprising contacting the nucleic acid sequence with a system as described herein, a vector as described herein, and / or a ribonucleoprotein as described herein, wherein the spacer hybridizes with a nucleic acid. In some embodiments, targeting the nucleic acid sequence comprises one or more of cutting and / or cleaving the nucleic acid sequence, nicking the nucleic acid sequence, enhancing expression of the nucleic acid sequence, decreasing expression of the nucleic acid sequence, visualizing or detecting the nucleic acid sequence, labeling the nucleic acid sequence, binding the nucleic acid sequence, enriching the nucleic acid sequence, depleting the nucleic acid sequence, editing the nucleic acid sequence, trafficking the nucleic acid sequence, splicing the nucleic acid sequence, and / or masking the nucleic acid sequence. In some embodiments, the nucleic acid sequence is a DNA sequence. In some embodiments, the DNA sequence is double stranded. In some embodiments, the DNA sequence is chromosomal DNA. In some embodiments, targeting the DNA sequence induces a double-stranded break. In some embodiments, the DNA sequence is edited through non-homologous end joining or homologous recombination. In some embodiments, the nucleic acid sequence is an RNA sequence. In some embodiments, the method increases indel frequency for SuperFi-Cas9 relative to when a gRNA is 20 nucleotides in length and does not comprise a 5′ extension.
[0020] In some embodiments, the present disclosure provides a cell comprising a system as described herein, a vector as described herein, and / or a ribonucleoprotein as described herein. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a stem cell. In some embodiments, the cell is a somatic cell.
[0021] BRIEF DESCRIPTION OF FIGURES
[0022] Fig. 1 shows a violin plot of the efficiency of CRISPR editing from the same set of 91, 603 gRNAs. The editing efficiency of six Cas nucleases were plotted for evaluation: SpCas9 (wild-type) , SpCas9-NG, Sniper-Cas9, LZ3-Cas9, HiFi-Cas9, and SuperFi-Cas9.
[0023] Figs. 2A-2B show the nucleotide preferences of SuperFi-Cas9 / gRNA (Fig. 2A) and SpCas9 / gRNA (Fig. 2B) , cleavage activity in the target sequences, and context sequences in the flanking genomic regions. At each position along the x-axis, gRNAs are ranked in descending order according to indel frequencies. The percentage of a particular nucleotide in the first quartile at each position was divided by its percentage in the fourth quartile. The log2-transformed relative ratio is denoted as the log-odds score (y-axis) , representing relative abundance in the 1st and 4th quartiles.
[0024] Figs. 3A-3C show the illustration of three types of gRNA expression cassettes. Fig. 3A shows gRNA transcribed under a U6 promoter, which requires guanine (G) as the first nucleotide of its transcript and results in a gRNA with a 5′ leading G; Fig. 3B shows a tRNA-gRNA transcript transcribed under a U6 promoter, in which the tRNA is released from the transcript that is processed by RNase P and RNase Z, respectively, which also releases the gRNA; and Fig. 3C shows a hammerhead-gRNA transcript transcribed under a U6 promoter, where a hammerhead ribozyme cuts at the 3′ of hammerhead, which releases the gRNA.
[0025] Figs. 4A-4D show the nucleotide preferences of Sniper-Cas9 / gRNA (Fig. 4A) , HiFi-Cas9 / gRNA (Fig. 4B) , SpCas9-NG (Fig. 4C) , and LZ3-Cas9 (Fig. 4D) cleavage activity in the target sequences and context sequences in the flanking genomic regions. At each position along the x-axis, gRNAs were ranked in descending order according to their indel frequencies. The percentage of a particular nucleotide in the first quartile at each position was divided by its percentage in the fourth quartile. The log2-transformed relative ratio is denoted as the log-odds score (y-axis) , representing relative abundance in the 1st and 4th quartiles.
[0026] Figs. 5A-5D show the mutation map of four representative gRNAs, where each gRNA comprises one of the four nucleotides in the -1 position (19 positions in the plot) : guanine (Fig. 5A) , adenine (Fig. 5B) , thymine (Fig. 5C) , and cystine (Fig. 5D) . Each block in the heatmap shows the ratio of SuperFi-Cas9 editing efficiencies of a nucleotide matched to the target sequence and the three other nucleotides mismatched to the target sequence. The ratios of editing efficiency between the maximumly changed mismatched nucleotide and the matched nucleotide were plotted below the heatmap. The 20-nucleotide target sequences and 40 nucleotides surrounding context sequences (20 nucleotides upstream and 20 nucleotides downstream of the target sequences in the human genome) were shown.
[0027] Fig. 6 shows a schematic drawing of gRNA and its targeting DNA. Nucleotides are colored: target strand (dark grey) , non-target strand (light grey) , gRNA (green) , PAM sequence (red) , 5′ mismatched nucleotide of gRNA (red) .
[0028] Fig. 7 shows a bar plot summarizing the indel frequency of SuperFi-Cas9 in HEK293T cells at twenty endogenous genome sites (x-axis: Site 1 to Site 20) with gRNAs that are 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, and 24 nucleotides in length. The genome sites were ordered by the editing efficiency of the SuperFi-Cas9 / 20-nucleotide gRNA in a descending order.
[0029] Fig. 8 shows a violin plot summarizing the rank of gRNAs by their editing efficiency at different lengths, determined from twenty endogenous targets in the human genome. At each of the twenty sites, gRNAs of 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, and 24 nucleotides were used with SuperFi-Cas9, and gRNAs were ranked by their editing efficiencies in a descending order. Overall ranks were plotted in the violin plot, against the length of gRNAs.
[0030] Fig. 9 shows the editing efficiency of the 21-or 22-nucleotide gRNA normalized to the 20-nucleotide gRNA, for the same twenty endogenous genome sites shown in Fig. 7.
[0031] Fig. 10 shows a bar plot summarizing the indel frequency of SuperFi-Cas9 at two endogenous genome sites (x-axis: Site 11 and Site 12) with gRNAs that were 20 nucleotides, 21 nucleotides, 22 nucleotides, and 23 nucleotides in length. “tRNA” refers to gRNAs that were released by RNase P and RNase Z from a tRNA-gRNA transcript (Fig. 3B) , and “HH” refers to gRNAs that were released by the hammerhead ribozyme from a hammerhead-gRNA transcript (Fig. 3C) .
[0032] Fig. 11 shows a bar plot summarizing the indel frequency of SuperFi-Cas9 or wildtype Cas9 (WT-Cas9) complexed with gRNA as RNP (ribonucleotide protein) at three endogenous genome sites. SuperFi-Cas9 protein was complexed with a 20-nucleotide gRNA, a 21-nucleotide gRNA, a 22-nucleotide gRNA, a 23-nucleotide gRNA, and a 24-nucleotide gRNA. WT-Cas9 protein was complexed with a 20-nucleotide gRNA and a 22-nucleotide gRNA. “Site 5” , “Site 9” , and “Site 18” refer to the three endogenous sites.
[0033] Figs. 12A-12C show three bar plots summarizing the A-to-G base substitution frequency of SuperFi-ABE (Adenine Base Editor) at three endogenous sites, Site 4, Site 5, and Site 9, respectively. The sequence of the target site was shown below each bar plot, in which black arrow indicated the most 5’ nucleotide of a 20-nucleotide gRNA, the letter A with underline indicated the potential targets of ABE, and the number underlying the letter A marked the position of adenosine with detectable substitution efficiency. Five gRNAs with different length, including a 20-nucleotide gRNA, a 21-nucleotide gRNA, a 22-nucleotide gRNA, a 23-nucleotide gRNA, and a 24-nucleotide gRNA, were tested in this experiment.
[0034] Figs. 13A-13C show three bar plots summarizing the C-to-T base substitution frequency of SuperFi-CBE (Cytosine Based Editor) at three endogenous sites, Site 4, Site 5, and Site 9, respectively. The sequence of the target site was shown below each bar plot, in which black arrow indicated the most 5’ nucleotide of a 20-nucleotide gRNA, the letter C with underline indicated the potential targets of CBE, and the number underlying the letter C marked the position of cytosine with detectable substitution efficiency. Five gRNAs with different length, including a 20-nucleotide gRNA, a 21-nucleotide gRNA, a 22-nucleotide gRNA, a 23-nucleotide gRNA, and a 24-nucleotide gRNA, were tested in this experiment.
[0035] Fig. 14 shows a bar plot summarizing the desired edits of SuperFi-PE (Prime Editor) at three endogenous sites, Site 5, HEK3, and EMX1. SuperFi-PE was complexed with pegRNA with a 20-nucleotide spacer, a 21-nucleotide spacer, a 22-nucleotide spacer, a 23-nucleotide spacer, and a 24-nucleotide spacer, respectively.DETAILED DESCRIPTION
[0036] All publications cited in this specification are herein incorporated by reference as though fully set forth. If certain content of a reference cited herein contradicts or is inconsistent with the present disclosure, the present disclosure controls.
[0037] Definitions
[0038] In the present disclosure, unless otherwise specified, the scientific and technical terms used herein have the meanings generally understood by a person skilled in the art. Although any methods and materials similar or equivalent to those described herein find use in the practice of the present disclosure, the preferred methods and materials are described herein. Accordingly, the terms defined herein are more fully described by reference to the Specification as a whole.
[0039] As used herein, the singular terms “a, ” “an, ” and “the” include the plural reference unless the context clearly indicates otherwise.
[0040] As used herein, “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative ( “or” ) . Moreover, the present invention also contemplates that in some embodiments of the invention, any feature or combination of features set forth herein can be excluded or omitted.
[0041] Unless the context requires otherwise, the terms “comprise, ” “comprises, ” and “comprising, ” or similar terms are intended to mean a non-exclusive inclusion, such that a recited list of elements or features does not include those stated or listed elements solely, but may include other elements or features that are not listed or stated.
[0042] Unless otherwise indicated, nucleic acids are written left to right in the 5′ to 3′ orientation, and amino acid sequences are written left to right in amino to carboxy orientation, respectively.
[0043] It is to be understood that this disclosure is not limited to the particular methodology, protocols, and reagents described, as these may vary, depending upon the context in which they are used by those skilled in the art.
[0044] As used herein, the terms “percent identity” and “%identity, ” as applied to nucleic acid or polynucleotide sequences, refer to the percentage of residue matches between at least two nucleic acid or polynucleotide sequences aligned using a standardized algorithm. Such an algorithm may insert, in a standardized and reproducible way, gaps in the sequences being compared in order to optimize alignment between two sequences, and therefore achieve a more meaningful comparison of the two sequences.
[0045] Percent identity between nucleic acid or polynucleotide sequences may be determined using a suite of commonly used and freely available sequence comparison algorithms provided by the National Center for Biotechnology Information (NCBI) Basic Local Alignment Search Tool (BLAST) (Altschul, S.F. et al. (1990) J. Mol. Biol. 215: 403-410) , which is available from several sources, including the NCBI, Bethesda, Md., and on the Internet at http: / / www. ncbi. nlm. nih. gov / BLAST / .
[0046] Nucleic acid or polynucleotide sequences that do not show a high degree of identity may nevertheless encode similar amino acid sequences due to the degeneracy of the genetic code. It is understood that changes in a nucleic acid sequence can be made using this degeneracy to produce multiple nucleic acid sequences that all encode substantially the same protein. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al. (1991) Nucleic Acid Res 19: 5081; Ohtsuka et al. (1985) J Biol Chem 260: 2605-2608; Cassol et al. (1992) ; Rossolini et al. (1994) Mol Cell Probes 8: 91-98) . The term “nucleic acid” refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single-or double-stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogues of natural nucleotides which have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. The term nucleic acid is used interchangeably with polynucleotide, and (in appropriate contexts) gene, cDNA, and mRNA encoded by a gene.
[0047] As used herein, “percent (%) amino acid sequence identity” with respect to a peptide, polypeptide or protein sequence is defined as the percentage of amino acid residues in a candidate sequence that are identical with the amino acid residues in another peptide or polypeptide sequence, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity, and not considering any conservative substitutions as part of the sequence identity. Percent amino acid sequence identity in the current disclosure is measured using BLAST software. Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared.
[0048] An amino acid substitution refers to the replacement of one amino acid in a polypeptide with another amino acid. Amino acid substitutions can be conservative or non-conservative substitutions. Exemplary substitutions are shown in Table 1. Amino acid substitutions may be introduced into a protein of interest and the products screened for a desired activity, for example, retained / improved biological activity.
[0049] Table 1.
[0050] Amino acids may be grouped according to common side-chain properties:
[0051] (1) hydrophobic: Norleucine, Met, Ala, Val, Leu, Ile;
[0052] (2) neutral hydrophilic: Cys, Ser, Thr, Asn, Gln;
[0053] (3) acidic: Asp, Glu;
[0054] (4) basic: His, Lys, Arg;
[0055] (5) residues that influence chain orientation: Gly, Pro;
[0056] (6) aromatic: Trp, Tyr, Phe.
[0057] As used herein, the term “polypeptide” is intended to encompass a singular “polypeptide” as well as plural “polypeptides, ” and refers to a molecule composed of monomers (amino acids) linearly linked by amide bonds (also known as peptide bonds) . The term “polypeptide” refers to any chain or chains of two or more amino acids, and does not refer to a specific length of the product. Thus, “peptides, ” “protein” , or any other term used to refer to a chain or chains of two or more amino acids, are included within the definition of “polypeptide, ” and the term “polypeptide” may be used instead of, or interchangeably with any of these terms. The term “polypeptide” is also intended to refer to the products of post-expression modifications of the polypeptide, including without limitation glycosylation, acetylation, phosphorylation, amidation, derivatization by known protecting / blocking groups, proteolytic cleavage, or modification by non-naturally occurring amino acids. A polypeptide may be derived from a natural biological source or produced by recombinant technology, but is not necessarily translated from a designated nucleic acid sequence. It may be generated in any manner, including by chemical synthesis.
[0058] As used herein, the term “encode” or “encoding” as it is applied to polynucleotides refers to a polynucleotide which is said to “encode” a polypeptide if, in its native state or when manipulated by methods well known to those skilled in the art, it can be transcribed and / or translated to produce the mRNA for the polypeptide and / or a fragment thereof. The antisense strand is the complement of such a nucleic acid, and the encoding sequence can be deduced therefrom.
[0059] A “guide RNA” (gRNA) refers to a synthetic or expressed RNA sequence that is capable of hybridizing to a target nucleic acid. In some embodiments, the guide RNA is capable of binding to a Cas protein to form a Cas complex and direct the Cas complex to the target nucleic acid.
[0060] As used herein, the term “variant” refers to varied form of a subject, which includes wild-type forms, naturally occurring forms, or artificially mutant forms. In some embodiments, the variant has the same or similar function of the original subject.
[0061] As used herein, the term “fidelity” refers to the accuracy of CRISPR activity (e.g., genome editing) . In some embodiments, the systems and methods disclosed herein have high and / or increased fidelity, meaning that they are highly accurate (e.g., containing few or no errors) .
[0062] SuperFi-Cas9 Endonuclease
[0063] Prokaryotic adaptive immune systems use Clustered Regularly Interspaced Short Palindromic Repeats (CRISPRs) and CRISPR-associated (Cas) proteins for RNA-guided cleavage of foreign genetic elements. Type II CRISPR–Cas systems contain a single protein Cas9 that is assembled with a CRISPR RNA (crRNA) and trans-activating crRNA (tracrRNA) to form a crRNA-guided nucleic acid-targeting effector complex.
[0064] SpCas9, derived from Streptococcus pyogenes, has been widely used and studied for genomic editing, including target gene disruption, transcriptional repression and activation, epigenetic modulation, and single base-pair conversion. Notably, these functions in CRISPR-Cas9-mediated genomic editing are disrupted by off-target DNA cleavage. It was recently reported that a specific conformation of the gRNA-DNA duplex formed in the presence of a mismatch prevent Cas9 activation. Specifically, substrates containing mismatches that were distal to the protospacer adjacent motif (PAM) were stabilized by a reorganization of a loop in the RuvC domain in Cas9. It was found that mutagenesis of 7 residues associated with mismatch stabilization reduced off-target DNA cleavage while maintaining on-target DNA cleavage10. The mutagenesis of these residues formed the basis for the development of the SuperFi-Cas9 plasmid, a high-fidelity variant of Cas9. However, as previously noted, SuperFi-Cas9 was later reported to show low editing activity in mammalian cells, across 26 endogenous sites in cell line7.
[0065] The present disclosure provides SuperFi-Cas9 (encoded by SEQ ID NO: 1) that is capable of cleaving nucleic acid effectively in cells, e.g., mammalian cells.
[0066] The present disclosure also provides variants of the SuperFi-Cas9 protein disclosed herein. As used herein, “variant” includes derivative, functional fragment, homolog, ortholog, and paralog. In some embodiments, the present disclosure provides a SuperFi-Cas9 protein variant having an amino acid sequence of at least 80%identity to SEQ ID NO: 1. In some embodiments, the present disclosure provides a SuperFi-Cas9 protein variant having an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%identity to SEQ ID NO: 1. In some embodiments, the SuperFi-Cas9 protein variant disclosed herein comprises conserved amino acid residue substitutions. In some embodiments, the SuperFi-Cas9 protein variant disclosed herein comprises only conserved amino acid residue substitutions (i.e., all amino acid substitutions in the derivative are conserved substitutions, and there is no substitution that is not conserved) .
[0067] In some embodiments, the SuperFi-Cas9 protein variant disclosed herein retains at least one of the functions of SuperFi-Cas9. In some embodiments, the SuperFi-Cas9 protein variant disclosed herein is capable of forming a complex with a gRNA. In some embodiments, the SuperFi-Cas9 protein variant disclosed herein has at least one nucleic acid catalytic activity. In some embodiments, the SuperFi-Cas9 protein variant disclosed herein is capable of cleaving target nucleic acid. In some embodiments, the SuperFi-Cas9 protein variant disclosed herein is capable of catalyzing nucleic acid degradation. In some embodiments, the SuperFi-Cas9 protein variant disclosed herein is capable of catalyze degradation of target nucleic acids and non-target nucleic acids.
[0068] Guide RNA (gRNA)
[0069] In an aspect, the present disclosure provides a guide RNA (gRNA) capable of hybridizing to a target nucleic acid region. A gRNA comprises two main components: a crRNA that is complementary to a target nucleic acid region (also referred to as a spacer) and a tracrRNA that serves as the binding scaffold for the Cas endonuclease (also referred to as a scaffold) . The gRNA binds to the complementary target nucleic acid sequence when a protospacer adjacent motif (PAM) is present in the non-targeted nucleic acid strand.
[0070] In some embodiments, the gRNA comprises a spacer that is 20 nucleotides long, and the spacer comprises an additional 5′ extension comprising one or more additional nucleotides, thus resulting in a spacer that comprises at least 20 nucleotides. In some embodiments, the 5′ extension of the spacer is 1, 2, 3, or 4 nucleotides in length, and the spacer is 21, 22, 23, or 24 nucleotides total in length. In some embodiments, the nucleotides in the 5′ nucleotide extension are complementary to the target nucleic acid region. In some embodiments, the nucleotides in the 5′ nucleotide extension are mismatched to the target nucleic acid region.
[0071] In some embodiments, the gRNA is capable of binding to SuperFi-Cas9 (SEQ ID NO: 1) or a variant thereof, wherein the variant thereof has an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 1.
[0072] In some embodiments, the 5′ extension of the spacer does not decrease SuperFi-Cas9 binding to the gRNA, relative to SuperFi-Cas9 binding to a gRNA that does not comprise a 5′ extension. In some embodiments, the 5′ extension of the spacer does not decrease SuperFi-Cas9 genome editing fidelity, relative to SuperFi-Cas9 genome editing fidelity when a gRNA does not comprise a 5′ extension. In some embodiments, the 5′ extension of the spacer improves SuperFi-Cas9 genome editing fidelity, relative to SuperFi-Cas9 genome editing fidelity when a gRNA does not comprise a 5′ extension. In some embodiments, the 5′ extension of the spacer does not decrease SuperFi-Cas9 indel frequency, relative to SuperFi-Cas9 indel frequency when a gRNA does not comprise a 5′ extension. In some embodiments, the 5′ extension of the spacer increases SuperFi-Cas9 indel frequency, relative to SuperFi-Cas9 indel frequency when a gRNA does not comprise a 5′ extension. In some embodiments, the 5′ extended spacer improves indel frequency by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, or at least 80%relative to a 20-nucleotide gRNA that does not comprise a 5′ extension.
[0073] In some embodiments, the gRNA described herein is a prime editing guide RNA (pegRNA) , comprising a spacer that is complementary to a target nucleic acid and is at least 20 nucleotides in length. In some embodiments, the spacer is 20, 21, 22, 23, or 24 nucleotides total in length.
[0074] In some embodiments, the target nucleic acid is DNA or RNA. In some embodiments, the target nucleic acid is DNA. The target DNA may be any suitable form of DNA, including naturally occurring DNA and engineered DNA.
[0075] For example, in some embodiments, the SuperFi-Cas9 protein recognizes and cleaves DNA targets located on the coding strand of open reading frames (ORFs) . In some embodiments, the target DNA is associated with a disease or condition of disease. In some embodiments, the CRISPR-Cas9 systems described herein can be used to treat a condition or disease by targeting a relevant DNA. For instance, the target DNA associated with a condition or disease may be a toxic DNA and / or a mutated DNA (e.g., a DNA molecule having a splicing defect or a mutation) . The target nucleic acid may also be a DNA that is specific for a particular microorganism (e.g., a pathogenic bacteria) .
[0076] Recently, it has been shown that Cas9 endonucleases from subtypes II-A and II-C can recognize and cleave single stranded RNA independent of a PAM11. In some embodiments, the target nucleic acid is RNA. The target RNA may be any suitable form of RNA, including naturally occurring RNA and engineered RNA. For example, non-limiting examples of target RNA include mRNA, tRNA, ribosomal RNA (rRNA) , non-coding RNA, lncRNA (long non-coding RNA) , microRNA (miRNA) , interfering RNA (siRNA) , viral RNA, circular RNA, and nuclear RNA. For example, in some embodiments, the SuperFi-Cas9 protein recognizes and cleaves RNA targets located on the coding strand of open reading frames (ORFs) . In some embodiments, the target RNA is associated with a disease or condition of disease. In some embodiments, the CRISPR-Cas9 systems described herein can be used to treat a condition or disease by targeting a relevant RNA. For instance, the target RNA associated with a condition or disease may be an RNA molecule that is overexpressed in a diseased cell (e.g., a cancer or tumor cell) . The target nucleic acid may also be a toxic RNA and / or a mutated RNA (e.g., an mRNA molecule having a splicing defect or a mutation) . The target nucleic acid may also be an RNA that is specific for a particular microorganism (e.g., a pathogenic bacteria) .
[0077] CRISPR-Cas9 Complex
[0078] In an aspect, the present disclosure provides a CRISPR-Cas complex, comprising: a gRNA comprising a spacer capable of hybridizing to a target nucleic acid and further comprising a 5′ extension, and a SuperFi-Cas9 protein (SEQ ID NO: 1) or a variant thereof, wherein the variant thereof has an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 1, and wherein the SuperFi-Cas9 protein or the variant thereof is capable of binding to the gRNA. In some embodiments, the spacer is 21-24 nucleotides in length. In some embodiments, the spacer is 21 nucleotides long. In some embodiments, the spacer is 22 nucleotides long. In some embodiments, the spacer is 23 nucleotides long. In some embodiments, the spacer is 24 nucleotides long.
[0079] In an aspect, the present disclosure provides a system comprising (1) a gRNA or a polynucleotide encoding thereof, wherein the guide RNA comprises a spacer, wherein the spacer is capable of hybridizing to a target nucleic acid and further comprises a 5′ extension; (2) and a SuperFi-Cas9 protein or a variant thereof or a polynucleotide encoding the SuperFi-Cas9 (SEQ ID NO: 1) or the variant thereof, wherein the variant thereof has an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 1. In some embodiments, the spacer is 21-24 nucleotides in length. In some embodiments, the spacer is 21 nucleotides long. In some embodiments, the spacer is 22 nucleotides long. In some embodiments, the spacer is 23 nucleotides long. In some embodiments, the spacer is 24 nucleotides long.
[0080] In an aspect, the present disclosure provides a ribonucleoprotein (RNP) comprising (1) a gRNA or a polynucleotide encoding the gRNA, wherein the guide RNA comprises a spacer, wherein the spacer is capable of hybridizing to a target nucleic acid and further comprises a 5′ extension; (2) and a SuperFi-Cas9 protein (SEQ ID NO: 1) or a variant thereof, wherein the SuperFi-Cas9 has an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%sequence identity to SEQ ID NO: 1. In some embodiments, the spacer is 21-24 nucleotides in length. In some embodiments, the spacer is 21 nucleotides long. In some embodiments, the spacer is 22 nucleotides long. In some embodiments, the spacer is 23 nucleotides long. In some embodiments, the spacer is 24 nucleotides long.
[0081] Polynucleotide
[0082] In some embodiments, the present disclosure provides an engineered polynucleotide comprising a sequence encoding a SuperFi-Cas9 protein (SEQ ID NO: 1) or a variant thereof as described herein, wherein the variant thereof has an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 1. In some embodiments, the present disclosure provides an engineered polynucleotide encoding a gRNA as described herein, wherein the gRNA comprises a spacer that is capable of hybridizing to a target nucleic acid and further comprises a 5′ extension, wherein the spacer is at least 20 nucleotides in length.
[0083] In some embodiments, the polynucleotide is codon-optimized. In some embodiments, the polynucleotide is codon-optimized to increase GC content. In some embodiments, the polynucleotide is codon-optimized by replacing rare codons with frequent codons. In some embodiments, the polynucleotide is humanized by replacing rare codons in human genome with frequent codons in human genome. Codon optimization methods are known in the art and may be useful in efforts to achieve one or more of several goals. These goals include to match codon frequencies in target and host organisms to ensure proper folding, bias GC content to increase mRNA stability or reduce secondary structures, minimize tandem repeat codons or base runs that may impair gene construction or expression, customize transcriptional and translational control regions, insert or remove protein trafficking sequences, remove / add post translation modification sites in encoded protein (e.g., glycosylation sites) , add, remove or shuffle protein domains, insert or delete restriction sites, modify ribosome binding sites and mRNA degradation sites, to adjust translational rates to allow the various domains of the protein to fold properly, or to reduce or eliminate problem secondary structures within the mRNA. Codon optimization tools, algorithms and services are known in the art, non-limiting examples include services from GeneArt (Life Technologies) and / or DNA2.0 (Menlo Park Calif. ) .
[0084] In some embodiments, the polynucleotide is an RNA. In some embodiments, the RNA comprises a 5′-untranslated region (UTR) and a 3′-UTR. In some embodiments, the mRNA further comprises a polyA sequence at 3′ end and / or a 5′-cap. Untranslated region (UTR) refers to the region flanking the protein coding sequence on either the 3′ side or the 5′ side on an RNA. It is called 5′-UTR if it is upstream of the protein coding sequence (i.e., on the 5′ side) . It is called 3′-UTR if it is downstream of the protein coding sequence (i.e., on the 3′ side) . UTRs are important regulatory elements with a strong impact on the post-transcriptional regulation of gene expression. A poly (A) tail is a long chain of adenine nucleotides that is added to the 3′ end of a mRNA molecule. In some embodiments, the length of the poly (A) tail is at least 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides. 5′ cap refers to a specially altered nucleotide on the 5′ end of the RNA sequence.
[0085] In some embodiments, the RNA is an in vitro transcribed (IVT) mRNA. In vitro transcription is a procedure that allows for template-directed synthesis of RNA molecules of any sequence from short oligonucleotides to those of several kilobases in μg to mg quantities. In some embodiments, it is based on the engineering of a template that includes a bacteriophage promoter sequence (e.g., from the T7 coliphage) upstream of the sequence of interest followed by transcription using the corresponding RNA polymerase. Techniques for in vitro transcription is well known in the art12.
[0086] In some embodiments, the RNA comprises modified nucleotide. For example, modifications can comprise one or more nucleotides modified at the 2′ position of the sugar, for example, a 2′-O-alkyl, 2′-O-alkyl-O-alkyl, or 2′-fluoro-modified nucleotides. In some examples, RNA modifications can comprise 2′-fluoro, 2′-amino or 2′ O-methyl modifications on the ribose of pyrimidines, abasic residues, or an inverted base at the 3′ end of the RNA.
[0087] The polynucleotides disclosed herein can be obtained by methods known in the art. For example, the polynucleotide can be obtained from cloned DNA (e.g., from a DNA library) , by chemical synthesis, by cDNA cloning, or by the cloning of genomic DNA or fragments thereof, purified from the desired cell. When the polynucleotides are produced by recombinant means, any method known to those skilled in the art for identification of nucleic acids that encode desired genes can be used. Any method available in the art can be used to obtain a full length (i.e., encompassing the entire coding region) cDNA or genomic DNA encoding a desired protein, such as from a cell or tissue source. Modified or variant polynucleotides can be engineered from a wildtype polynucleotide using standard recombinant DNA methods. Polynucleotides can be cloned or isolated using any available methods known in the art for cloning and isolating nucleic acid molecules. Such methods include PCR amplification of nucleic acids and screening of libraries, including nucleic acid hybridization screening, antibody-based screening, and activity-based screening.
[0088] Methods for amplification of polynucleotides can be used to isolate polynucleotides encoding a desired protein, including for example, polymerase chain reaction (PCR) methods. PCR can be carried out using any known methods or procedures in the art. Exemplary methods include use of a Perkin-Elmer Cetus thermal cycler and Taq polymerase (Gene Amp) . A nucleic acid containing a gene of interest can be used as a source material from which a desired polypeptide-encoding nucleic acid molecule can be amplified. For example, DNA and mRNA preparations, cell extracts, tissue extracts from an appropriate source (e.g., testis, prostate, breast) , fluid samples (e.g., blood, serum, saliva) , samples from healthy and / or diseased subjects can be used in amplification methods. The source can be from any eukaryotic species including, but not limited to, vertebrate, mammalian, human, porcine, bovine, feline, avian, equine, canine, and other primate sources. Nucleic acid libraries also can be used as a source material. Primers can be designed to amplify a desired polynucleotide. For example, primers can be designed based on expressed sequences from which a desired polynucleotide is generated. Primers can be designed based on back-translation of a polypeptide amino acid sequence. If desired, degenerate primers can be used for amplification. Oligonucleotide primers that hybridize to sequences at the 3′ and 5′ termini of the desired sequence can be used as primers to amplify by PCR from a nucleic acid sample. Primers can be used to amplify the entire full-length polynucleotide, or a truncated sequence thereof. Nucleic acid molecules generated by amplification can be sequenced and confirmed to encode a desired polypeptide.
[0089] Vector
[0090] In an aspect, the present disclosure provides a vector comprising a polynucleotide described herein.
[0091] In an aspect, the present disclosure provides a vector comprising a polynucleotide encoding a gRNA, wherein the gRNA comprises a spacer that is capable of hybridizing to a target nucleic acid and further comprises a 5′ extension, wherein the spacer is at least 20 nucleotides in length; and a polynucleotide encoding a SuperFi-Cas9 protein (SEQ ID NO: 1) or a variant thereof, wherein the variant thereof has an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 1.
[0092] In some embodiments of the vector described herein, the vector is a plasmid or a viral vector.
[0093] In some embodiments, the vector is an AAV vector. AAV is a non-enveloped virus that can be engineered to deliver DNA to target cells. AAV comprises a protein shell surrounding and protecting a small, single-stranded DNA genome of approximately 4.8 kilobases (kb) . Recombinant AAV (rAAV) , which lacks viral DNA, is essentially a protein-based nanoparticle engineered to traverse the cell membrane, where it can ultimately traffic and deliver its DNA cargo into the nucleus of a cell. In the absence of Rep proteins, ITR-flanked transgenes encoded within rAAV can form circular concatemers that persist as episomes in the nucleus of transduced cells. Because recombinant episomal DNA does not integrate into host genomes, it will eventually be diluted over time as the cell undergoes repeated rounds of replication. This will eventually result in the loss of the transgene and transgene expression, with the rate of transgene loss dependent on the turnover rate of the transduced cell. These characteristics make rAAV ideal for certain gene therapy applications.
[0094] Any methods known in the art for the insertion of DNA fragments into a vector can be used to construct expression vectors comprising a polynucleotide disclosed herein. These methods can include in vitro recombinant DNA and synthetic techniques and in vivo (genetic) recombination. The polynucleotide disclosed herein can be operably linked to control sequences in the expression vector (s) to ensure protein expression. Such control sequences may include, but are not limited to, leader or signal sequences, promoters (e.g., naturally associated or heterologous promoters) , ribosomal binding sites, enhancer or activator elements, translational start and termination sequences, and transcription start and termination sequences, and are chosen to be compatible with the host cell chosen to express the proteins. Constitutive or inducible promoters as known in the art are also contemplated. The promoters may be either naturally occurring promoters, hybrid promoters that combine elements of more than one promoter, or synthetic promoters. An expression construct may be present in a cell on an episome, such as a plasmid, or the expression construct may be inserted in a chromosome such as in a gene locus. In some embodiment, the expression vector includes a selectable marker gene to allow the selection of transformed host cells. In some embodiments, the vector is an expression vector comprising a nucleotide sequence encoding a variant polypeptide operably linked to at least one regulatory control sequence. Regulatory control sequences for use herein include promoters, enhancers, and other expression control elements. In some embodiments, the expression vector is designed for the choice of the host cell to be transformed, the particular variant polypeptide desired to be expressed, the vector's copy number, the ability to control that copy number, and / or the expression of any other protein encoded by the vector, such as antibiotic markers.
[0095] The vector can include, but is not limited to, viral vectors and plasmid DNA. Viral vectors can include, but are not limited to, adenoviral vectors, lentiviral vectors, retroviral vectors, and adeno-associated viral vectors. Commonly, expression vectors contain selection markers such as ampicillin-resistance, hygromycin-resistance, tetracycline resistance, kanamycin resistance, or neomycin resistance to permit detection of those cells transformed with the desired DNA sequences. Suitable vectors, promoter, and enhancer elements are known in the art; many are commercially available for generating subject recombinant constructs. In some embodiments, the vector is a polycistronic vector. In some embodiments, the vector is a bicistronic vector or a tricistronic vector. Bicistronic or polycistronic expression vectors may include (1) multiple promoters fused to each of the open reading frames; (2) insertion of splicing signals between genes; (3) fusion of genes whose expressions are driven by a single promoter; and (4) insertion of proteolytic cleavage sites between genes (self-cleavage peptide) or insertion of internal ribosomal entry sites (IRESs) between genes. A polycistronic vector is used to co-express multiple genes in the same cell. In a strategy commonly used to prepare a polycistronic vector, an Internal Ribosome Entry Site (IRES) , acting as another ribosome recruitment site, allows initiation of translation from an internal region of the mRNA. Thus, two proteins are translated from one mRNA. IRES elements are quite large (usually 500-600 bp) 13, 14. Cell
[0096] In an aspect, the present disclosure provides a cell comprising the CRISPR-Cas complex described herein.
[0097] In an aspect, the present disclosure provides a cell comprising the polynucleotide described herein.
[0098] In an aspect, the present disclosure provides a cell comprising the system described herein.
[0099] In an aspect, the present disclosure provides a cell comprising the vector described herein.
[0100] In some embodiments of the cell described herein, the cell is a eukaryotic cell. In some embodiments of the cell described herein, the cell is a prokaryotic cell.
[0101] In some embodiments of the cell described herein, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a plant cell.
[0102] In some embodiments of the cell described herein, the cell is a stem cell.
[0103] In some embodiments of the cell described herein, the cell is a somatic cell.
[0104] In some embodiments, the SuperFi-Cas9 protein is introduced directly into the cell. In some embodiments, the SuperFi-Cas9 protein is expressed in a recombinant cell, such as E. coli, and purified. The resulting purified SuperFi-Cas9 protein, along with at least one appropriate gRNA specific for one or more target nucleic acids, is then introduced into a cell or organism where one or more nucleic acids can be targeted. In some embodiments, the SuperFi-Cas9 protein and gRNA are introduced as separate components into the target cell / organism. In some embodiments, the purified SuperFi-Cas9 protein is complexed with the guide RNA, and this RNP complex is introduced into the target cell (e.g., using transfection or injection) . In some embodiments, the SuperFi-Cas9 protein and guide molecule are injected into an embryo (such as a human, mouse, zebrafish, or Xenopus embryo) .
[0105] In some embodiments, the SuperFi-Cas9 protein is expressed from a polynucleotide in the cell. In some embodiments, the SuperFi-Cas9 protein is expressed from a vector, such as a viral vector or plasmid introduced into a cell. This results in the production of the SuperFi-Cas9 protein in the cell. In some embodiments, the polynucleotide encoding the SuperFi-Cas9 is co-expressed in the cell with the guide RNA. In some embodiments, multiple plasmids or vectors are used to deliver the SuperFi-Cas9 protein and the guide RNA into the cell. For example, the polynucleotide encoding the SuperFi-Cas9 can be provided on one vector or plasmid, and the guide RNA on another plasmid or vector. Multiple plasmids or viral vectors can be mixed and introduced into cells (or a cell free system) at the same time, or separately. In some examples, multiple polynucleotides are expressed from a single vector or plasmid. For example, a single vector can include the polynucleotide encoding the SuperFi-Cas9 and the guide RNA. In some embodiments, the vector (s) are introduced into the cell by transduction. In some embodiments, the vector (s) are introduced into the cell by microinjection or electroporation. In some embodiments, the vector (s) are introduced into the cell by chemical-based transfection.
[0106] Method
[0107] The SuperFi-Cas9 protein or variant thereof described herein can be used in applications as other Cas9 proteins and Cas proteins in Type II CRISPR systems. In some embodiments, the system comprising a SuperFi-Cas9 protein or variant thereof and guide RNAs described herein can be used in applications in Type II CRISPR systems. In some embodiments, the SuperFi-Cas9 protein or variant thereof is further modified, for example by mutation in a catalytic domain (e.g., RuvC or HNH-like domain) , or by fusion to or conjugation to another effector (e.g., a label, an affinity tag, a deaminase) . In an aspect, the present disclosure provides a method for targeting a target nucleic acid, comprising contacting the nucleic acid with the CRISPR-Cas complex described herein, the system of described herein, and / or the vector described herein. In some embodiments, the target nucleic acid is a DNA.
[0108] In some embodiments, targeting the target DNA comprises one or more of cutting the target DNA, nicking the target DNA, visualizing or detecting the target DNA, labeling the target DNA, binding target DNA, enriching the target DNA, depleting the target DNA, editing the target DNA, splicing the target DNA, and masking the target DNA.
[0109] Cas9 has found widespread application in areas such as DNA editing. In some embodiments, DNA editing mediated by SuperFi-Cas9 is performed through non-homologous end joining or by homologous recombination.
[0110] The CRISPR systems described herein have a wide variety of utilities including modifying (e.g., deleting, inserting, translocating, inactivating, and activating) a target polynucleotide or nucleic acid in a multiplicity of cell types. The CRISPR systems have a broad spectrum of applications in, e.g., DNA / RNA detection (e.g., specific high sensitivity enzymatic reporter unlocking (SHERLOCK) ) , tracking and labeling of nucleic acids, enrichment assays (extracting desired sequence from background) , controlling interfering RNA or miRNA, detecting circulating tumor DNA, preparing next generation library, drug screening, disease diagnosis and prognosis, and treating various genetic disorders.
[0111] In some embodiments, the method of targeting the target nucleic acid allows for one or more nucleotide base substitutions, nucleotide base edits, nucleotide base deletions, nucleotide base insertions, or combinations thereof, in the target nucleic acid. In some embodiments, the method of targeting the target nucleic acid allows for knockdown of a gene.
[0112] In some embodiments, the SuperFi-Cas9 protein or variant thereof described herein is used in place of a Cas9 nickase or other Cas protein in a base editing system (e.g., Adenine base editor (SuperFi-ABE) or Cytosine base editor (SuperFi-CBE) ) 15, 16 to allow for base editing of specific nucleotides in a target nucleic acid. In some embodiments, the present disclosure provides a SuperFi-ABE or variant thereof having an amino acid sequence of SEQ ID NO: 2, or having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%identity to SEQ ID NO: 2. In some embodiments, the present disclosure provides a SuperFi-CBE or variant thereof having an amino acid sequence of SEQ ID NO: 3, or having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%identity to SEQ ID NO: 3. In some embodiments, the SuperFi-ABE protein, SuperFi-CBE protein, or variant disclosed herein comprises conserved amino acid residue substitutions. In some embodiments, the SuperFi-ABE protein, SuperFi-CBE protein, or variant protein disclosed herein comprises only conserved amino acid residue substitutions (i.e., all amino acid substitutions in the derivative are conserved substitutions, and there is no substitution that is not conserved) .
[0113] In some embodiments, the SuperFi-Cas9 protein or variant thereof described herein is used in a complex with a pegRNA as described herein to perform prime editing (SuperFi-PE) 17. Prime editing allows for editing of a target nucleic acid without double strand breaks or donor DNA templates, and therefore could provide for greater flexibility in editing. In some embodiments, the present disclosure provides a SuperFi-PE having an amino acid sequence of SEQ ID NO: 4, or having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%identity to SEQ ID NO: 4. In some embodiments, the SuperFi-PE protein or variant disclosed herein comprises conserved amino acid residue substitutions. In some embodiments, the SuperFi-PE protein or variant protein disclosed herein comprises only conserved amino acid residue substitutions (i.e., all amino acid substitutions in the derivative are conserved substitutions, and there is no substitution that is not conserved) .
[0114] In some embodiments, the method of targeting the target nucleic acid results in detecting, visualizing, or labeling the target nucleic acid. For example, by using a SuperFi-Cas9 variant described herein that does not have endonuclease activity and a crRNA with a spacer specific for the target nucleic acid, and an effector module, the target nucleic acid will be recognized by the SuperFi-Cas9 variant described herein but will not be cut or nicked while the effector module becomes activated. In some embodiments, the effector module is fused to the SuperFi-Cas9 variant described herein. In some embodiments, the effector module is linked to the SuperFi-Cas9 variant described herein, optionally with a linker. In some embodiments, the effector module is a fluorescent protein or other detectable label. Binding of the SuperFi-Cas9 variant to the target nucleic acid can be visualized by microscopy or other methods of imaging. Such a method can be used in a cell or cell free system to determine if a target nucleic acid is present, such as in a tumor cell.
[0115] In some embodiments, the method of targeting the target nucleic acid results in editing the sequence of a target nucleic acid. For example, by using the SuperFi-Cas9 protein or variant thereof disclosed herein and a crRNA with a spacer specific for the target nucleic acid as described herein, the target nucleic acid can be cut or nicked at a precise location. In some examples, such a method is used to edit a target nucleic acid, which can correct a mutation associated with the corresponding protein. Such a method can be used in a cell where a mutated nucleic acid is associated with a disease and correction of the gene is required.
[0116] In some embodiments, the methods described herein are in vivo methods. In some embodiments, the methods described herein are ex vivo methods. In some embodiments, the methods described herein are in vitro methods.
[0117] Table 2.
[0118] REFERENCES
[0119] 1. Kleinstiver, B. P. et al., High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects. Nature 529, 490-495 (2016) .
[0120] 2. Slaymaker, I.M. et al., Rationally engineered Cas9 nucleases with improved specificity. Science 351, 84-88 (2016) .
[0121] 3. Chen, J.S. et al., Enhanced proofreading governs CRISPR-Cas9 targeting accuracy. Nature 550, 407-410 (2017) .
[0122] 4. Casini, A. et al., A highly specific SpCas9 variant is identified by in vivo screening in yeast. Nat Biotechnol 36, 265-271 (2018) .
[0123] 5. Lee, J.K. et al., Directed evolution of CRISPR-Cas9 to increase its specificity. Nat Commun 9, 3048 (2018) .
[0124] 6. Vakulskas, C.A. et al., A high-fidelity Cas9 mutant delivered as a ribonucleoprotein complex enables efficient gene editing in human hematopoietic stem and progenitor cells. Nat Med 24, 1216-1224 (2018) .
[0125] 7. Kulcsar, P.I., Talas, A., Ligeti, Z., Krausz, S.L. &Welker, E., SuperFi-Cas9 exhibits remarkable fidelity but severely reduced activity yet works effectively with ABE8e. Nat Commun 13, 6858 (2022) .
[0126] 8. Schmid-Burgk, J.L. et al., Highly Parallel Profiling of Cas9 Variant Specificity. Mol Cell 78, 794-800 e798 (2020) .
[0127] 9. Nishimasu, H. et al., Engineered CRISPR-Cas9 nuclease with expanded targeting space. Science 361, 1259-1262 (2018) .
[0128] 10. Bravo, J.P.K. et al., Structural basis for mismatch surveillance by CRISPR-Cas9. Nature 603, 343-347 (2022) .
[0129] 11. Strutt, S.C. et al., RNA-dependent RNA targeting by CRISPR-Cas9. eLife, &: e32724 (2018) .
[0130] 12. Beckert, B. and Masquida, B., Synthesis of RNA by in vitro transcription. RNA. Methods in Molecular Biology, 703: 29-41 (2011) .
[0131] 13. Pelletier, J. &Sonenberg, N., Internal initiation of translation of eukaryotic mRNA directed by a sequence derived from poliovirus RNA. Nature 334, 320-325 (1988) .
[0132] 14. Jang, S.K. et al., A segment of the 5′ nontranslated region of encephalomyocaditis virus RNA directs internal entry of ribosomes during in vitro translation. J Virol 62, 2636-2646 (1988) .
[0133] 15. Komor, A.C. et al., Programmable editing of a target base in genomic DNA without double-strnaded DNA cleavage. Nature 533, 420-424 (2016) .
[0134] 16. Gaudelli, N.M. et al., Programmable base editing of A-T to G-C in genome DNA without DNA cleavage. Nature 551, 464-471 (2017) .
[0135] 17. Anzalone, A.V. et al., Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019) .
[0136] 18. Zhang, H. et al., Deep sampling of gRNA in the human genome and deep-learning-informed prediction of gRNA activities. Cell Discov 9, 48 (2023) .
[0137] 19. Batard, P., Jordan, M. &Wurm, F., Transfer of high copy number plasmid into mammalian cells by calcium phosphate transfection. Gene 270, 61-68 (2001) .
[0138] 20. Kingston, R.E., Chen, C.A. &Okayama, H., Calcium phosphate transfection. Curr Protoc Immunol Chapter 10, Unit 10 13 (2001) .
[0139] 21. Clement, K. et al., CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat Biotechnol 37, 224-226 (2019) .
[0140] EXAMPLES
[0141] Example 1. High-throughput screening to identify SuperFi-Cas9 preference in the 5′ extension.
[0142] To profile the cleavage activity of SuperFi-Cas9 / gRNA in mammalian genome, a synthetic gRNA-target paired library was applied to K562 cells. This library included 91, 603 gRNA-target pairs, covering 19, 074 protein-coding genes and 5, 579 non-coding genes in the human genome. These gRNA-target pairs were transduced as a lentivirus library into K562 cells with stable SuperFi-Cas9 expression. The cleavage activity of individual gRNA was quantified from its paired synthetic target by high throughput sequencing (HTS) . In total, cleavage activity data was generated for 62, 083 gRNA in vivo after quality control. The Pearson correlation between the two biological replicates demonstrated good data quality, but the overall indel frequency of SuperFi-Cas9 / gRNA was relatively low compared to wild-type SpCas9 and other SpCas9 variants (Fig. 1) , which is consistent with a previous argument to the impaired editing efficiency of SuperFi-Cas9 compromised by the high fidelity property7.
[0143] To visualize the nucleotide preference of SuperFi-Cas9 / gRNA, a log-odds score was defined to identify the favored nucleotide along the target sequence. When compared to the nucleotide preference of SpCas9 / gRNA18, SuperFi-Cas9 / gRNA exhibited a strong preference at the 21st nucleotide at the PAM-distal region (Fig. 2) . This observation was unexpected and interesting, as the 21st nucleotide of the target sequence goes beyond the range of canonical base pairing of the gRNA-target duplex. Given the design of the synthetic cassette, the expressed gRNAs had a leading guanine at the -1 position (5′) of the 20-nucleotide gRNA (Fig. 3A) . The -1 guanine of gRNA could perfectly pair with a cytosine or sub-optimally pair with a thymine, which were reflected as a guanine or an adenine at the 21st position of the target sequence. It was thus reasoned that the observed guanine and adenine preference was an indication of a positive contribution to the editing efficiency of a 21 base-pair gRNA-target duplex, in which the most 5′ nucleotide of a 21-nucleotide gRNA paired with the target sequence.
[0144] To determine whether the beneficial -1 base pairing is unique to SuperFi-Cas9 / gRNA, the nucleotide preference of four SpCas9 variants, including SpCas9-NG, Sniper-Cas9, HiFi-Cas9, and LZ3-Cas9 was further studied (Fig. 4) . Interestingly, none of these SpCas9 variants showed nucleotide preference at the 21st position of the PAM-distal region (Fig. 4) . It was noted that SuperFi-Cas9 / gRNA showed a guanine and adenine preference at the 21st position of the PAM-distal region of target sequence, suggesting it may utilize a gRNA longer than 20 nucleotides. And extending gRNA-target duplex at the 5′ of gRNA may increase the cleavage activity of this high-fidelity Cas nuclease.
[0145] Example 2. Determination of SuperFi-Cas9’s nucleotide preference in a 5′ extension
[0146] Based on the indel frequency and nucleotide preference of gRNAs from high throughput profiling, a deep-learning (DL) model was generated using the high throughput profiling data and mutation maps were generated to predict the indel frequency when the -1 nucleotide of gRNA was changed. The model was trained using a pre-train and fine-tune frame, and the resultant model performance of predicting the indel frequency is 0.826 (Pearson correlation) and 0.768 (Spearman correlation) on the testing dataset.
[0147] The DL model was used to establish mutation maps for four gRNAs, each representing one type of -1 nucleotide. One example gRNA targeted a genomic sequence where the 21st nucleotide at the PAM-distal region was guanine. The mutation map generated from the prediction model indicated that the indel frequency decreased when the -1 nucleotide of that gRNA changed from guanine to the other three nucleotides (Fig. 5A) . Among them, G-to-T and G-to-C were more deleterious compared to G-to-A, which was consistent with the observed nucleotide preference from the high throughput profiling data. When the 21st nucleotide at the PAM-distal region was adenine, A-to-T and A-to-C change of the -1 nucleotide of gRNA were more deleterious to the indel frequency than A-to-G, according to the predicted scores (Fig. 5B) . Interestingly, the mutation map demonstrated an increased indel frequency of a gRNA, in which the 21st nucleotide at the PAM-distal region was thymine. And a T-to-G or T-to-A nucleotide swap increased the indel frequency of that gRNA from ~10%to ~30% (Fig. 5C) . The mutation map of a gRNA with a cytosine as the 21st target nucleotide echoed the pattern of thymine (Fig. 5D) .
[0148] Example 3. Determination of the impact of 5′ extensions on indel frequencies
[0149] Both the high throughput profiling and DL-predicted mutation map supported the superior editing efficiency of the 5′-extended gRNA. To evaluate the impact of a 5′ extension in a cellular environment, gRNAs were designed to target twenty endogenous genomic locations, as set forth in Table 3 (Fig. 6) . For each gRNA, a canonical 20-nucleotide gRNA (targeting the bolded sequences in Table 3) , whose editing efficiency was used as baseline, and 21-24-nucleotide gRNAs with 5′-extended nucleotides matched with the target sequences were designed. In some cases, a 21-nucleotide gRNA was designed, in which the most 5′ nucleotide mismatched with the target sequences. To exclude the influence from the leading guanine (Fig. 3A) , a U6-tRNA-gRNA expression cassette was used to express the gRNA with the desired length (Fig. 3B) .
[0150] In general, the 5′-extended nucleotides increased the indel frequency compared to the 20-nucleotide gRNA at all the twenty endogenous sites (Fig. 7) . Among them, the 21-nucleotide gRNA presented superior indel frequency at nineteen endogenous sites compared to the canonical 20-nucleotide gRNA; the 22-nucleotide gRNA outperformed the 20-nucleotide gRNA at seventeen sites; the 23-nucleotide gRNA outperformed the 20-nucleotide gRNA at eighteen sites; and the 24-nucleotide gRNA outperformed the 20-nucleotide gRNA at ten sites.
[0151] The 21-nucleotide gRNA outperformed all other lengths of gRNA at ten of the twenty sites, the 22-nucleotide gRNA outperformed all other gRNAs at nine of the twenty sites, and the 23-nucleotide gRNA outperformed all other gRNAs at one of the twenty sites (Fig. 7-8) . Thus, it was concluded that gRNAs with 5′-extended and matched nucleotides increase the indel frequency of SuperFi-Cas9, and the SuperFi-Cas9 / 21-nucleotide gRNA and SuperFi-Cas9 / 22-nucleotide gRNA showed superior editing performances among the 20-24-nucleotide gRNA that were examined.
[0152] The increase in indel frequency was quantified by comparing the 5′-extended gRNA to the canonical 20-nucleotide gRNA. At each endogenous targeting site and for each of the 5′-extended gRNA, the relative editing efficiency was calculated by normalizing to the indel frequency of the corresponding 20-nucleotide gRNA. Interestingly, with respect to the 20-nucleotide gRNA, the increase in indel frequency was greater when the editing efficiency is low (Fig. 9) . Table 3 provides sequences for the target region in this analysis, with the gRNAs designed to be complementary to the bolded 20-nucleotide sequence and / or 5′ extension nucleotides.
[0153] Table 3.
[0154] The 5′-extended gRNAs expressed from the U6-hammerhead-gRNA expression cassette also recapitulated the improved indel frequency compared to the 20-nucleotide gRNA (Fig. 10) .
[0155] Example 4. Use of SuperFi-Cas9 in a ribonucleoprotein complex, base editor, and prime editor
[0156] SuperFi-Cas9 or wildtype Cas9 (WT-Cas9) were complexed with gRNA as an RNP, and indel frequency was assessed at three endogenous genome sites (i.e., Site 5, Site 9, and Site 18) . SuperFi-Cas9 protein was complexed with a 20-nucleotide gRNA, a 21-nucleotide gRNA, a 22-nucleotide gRNA, a 23-nucleotide gRNA, and a 24-nucleotide gRNA, whereas WT-Cas9 was complexed with a 20-nucleotide gRNA and 22-nucleotide gRNA. While SuperFi-Cas9 editing efficiency, as shown by indel frequency, was low when a 20-nucleotide gRNA was used, when a 5′ extended gRNA was used, SuperFi-Cas9 editing efficiency was markedly increased (Fig. 11) . For example, SuperFi-Cas9 editing efficiency at Site 5 was approximately twice as effective when a 5′ extended gRNA (21-nucleotide, 22-nucleotide, 23-nucleotide, or 24-nucleotide gRNA) was used relative to a canonical 20-nucleotide gRNA. Indel frequency was also comparable to WT-Cas9 at all three sites for SuperFi-Cas9 when using a 5′ extended gRNA.
[0157] SuperFi-Cas9 was also modified into an Adenine base editor (ABE) , a Cytosine base editor (CBE) , and a Prime Editor. In ABE, CBE, and Prime Editors, a Cas9 nickase (Cas9n) or other Cas9 protein, along with other effector proteins, is used to edit specific nucleic acid targets. For example, the ABE targets adenine residues for substitution. In each of these embodiments, i.e., ABE, CBE, and Primer Editor, the SuperFi-Cas9 protein replaced the Cas9n or Cas9 protein. Figs. 12A-C show three bar plots summarizing the A-to-G base substitution frequency of SuperFi-ABE (Adenine Base Editor; SEQ ID NO: 2) at three endogenous sites, i.e., Site 4, Site 5, and Site 9, using five gRNAs of different lengths (a 20-nucleotide gRNA, a 21-nucleotide gRNA, a 22-nucleotide gRNA, a 23-nucleotide gRNA, and a 24-nucleotide gRNA) . While SuperFi-ABE was able to edit adenines at the -1, -2, and -5 positions, editing efficiency was the highest for adenines at the 3, 5, and 11 positions.
[0158] Figs. 13A-C show three bar plots summarizing the C-to-T base substitution frequency of SuperFi-CBE (Cytosine Base Editor; SEQ ID NO: 3) at three endogenous sites, i.e., Site 4, Site 5, and Site 9. Similar to the testing of SuperFi-ABE, SuperFi-CBE was evaluated with five gRNAs having different lengths (a 20-nucleotide gRNA, a 21-nucleotide gRNA, a 22-nucleotide gRNA, a 23-nucleotide gRNA, and a 24-nucleotide gRNA) . SuperFi-CBE showed greater editing efficiency of cytosine residues outside of the 20-nucleotide gRNA (i.e., residues at -7, -6, -4, and -2) than SuperFi-ABE, with relatively consistent editing efficiency across different length gRNA.
[0159] Fig. 14 shows a bar plot summarizing the desired edits of SuperFi-PE (Prime Editor; SEQ ID NO: 4) at three endogenous sites, Site 5, HEK3, and EMX1. SuperFi-PE was complexed with a prime editing guide RNA (pegRNA) with a 20-nucleotide spacer, a 21-nucleotide spacer, a 22-nucleotide spacer, a 23-nucleotide spacer, and a 24-nucleotide spacer. While SuperFi-PE editing was most efficient at Site 5, SuperFi-PE was able to introduce desired edits at all tested sites when complexed with pegRNA comprising a spacer that was between 20-24 nucleotides in length.
[0160] Methods and Materials
[0161] Screening Library. 91, 603 gRNA-target pairs were designed and synthesized to produce an oligo pool (GenScript) to represent these sequences. Each oligo was 149 nucleotides in length, including a gRNA, its putative genomic target, and appendix sequences required by the downstream molecular cloning. Among them, the putative genomic target included the 20-nucleotide target sequence, 3-nucleotide PAM sequence, and the ±20-nucleotide flanking sequence.
[0162] The plasmid library preparation was the same as described in Zhang et al18. In brief, a barcode oligo with a random 20-nucleotide sequence was amplified by PCR (primers: Barcode-PCR-F and Barcode-PCR-R) , and the purified PCR products were cloned into the backbone vector (ON-OFF-Backbone_lentiGuide-Puro-Mkate2) by Golden Gate. The product vector library was named Lig1-ON-OFF-Barcode, and more than 1.2×108 colonies were collected. The primers used for the barcode oligo amplification and the sequence of the barcode oligo are listed in Table 4.
[0163] The oligo pool was synthesized by amplification in 24 PCR reactions, with 40 ng oligo template used in each reaction. The PCR products were purified and cloned into the Lig1-ON-OFF-Barcode vector by Golden Gate. The product vector library was named Lig2-ON-OFF-Barcode-Oligo, and ~1.4×108 colonies were collected. ) .
[0164] An ON-OFF-Insert-Scaffold vector, which includes the gRNA scaffold, was ligated with the Lig2-ON-OFF-Barcode-Oligo vector library by Golden Gate. The product vector library was the final plasmid library, ready for lentivirus packing. The final plasmid library was composed of ~3.2×107 colonies, which covered 300x of the number of oligos. Two independent plasmid libraries were prepared, and each library was used in one replicate of the two biological replicates of high throughput screening.
[0165] A NGS library of the gRNA-target pairs and barcode cassette in the final plasmid library was established and sequenced as a starting reference. The PCR reaction included 50 μL reaction included 40 ng plasmid, 250 nM forward primer (Fwd-libseq) , 250 nM reverse primer (Rev-libseq) , 25 μl of NEBNext Ultra II Q5 Master Mix, and nuclease-free H2O. The PCR program was set as following: 30 s at 98 ℃; 6 cycles of 10 s at 98 ℃, 30 s at 64 ℃, and 10 s at 72 ℃; 8 cycles of 10 s at 98 ℃, 30 s at 71 ℃, and 10 s at 72 ℃; 2 mins at 72 ℃; 4 ℃ hold. The PCR products were purified using AMPure XP beads and subjected for NGS sequencing. The primers used for the NGS library preparation are listed in Table 4.
[0166] To prepare the Lig2-ON-OFF-Barcode-Oligo lentivirus library, the transfer plasmid, psPAX2, and pMD2. G were mixed at a weight ratio of 5: 3: 2, and a total of 192 ug plasmid mixture were transfected to HEK293T cells at 90%confluence in a T175 flask using the calcium phosphate method according to the published protocols19, 20. The culture media was changed eight hours after transfection, and the cells were washed with PBS to remove residual plasmids and transfection reagents. After 48 hours and 72 hours of transfection, the supernatants containing the lentivirus were collected. The supernatants were filtered through a 0.45 μm polyvinylidene fluoride filter and spun in an ultracentrifuge at 20,000 rpm for 2 hours. The precipitation was resuspended in 1 mL PBS and stored at -80 ℃ in the proper size of aliquots.
[0167] Table 4. Sequences and primers used in library generation.
[0168] High Throughput Screening in K562 Cells. To conduct high throughput screening, SuperFi-Cas9 stable expression K562 cells was established by transducing lenti-UCOE-SuperFI-Cas9-BSD lentivirus into the wildtype K562 cells. The K562-SuperFi-Cas9 cells were grown from a monoclonal cell, and the successful integration of SuperFi-Cas9 was examined by Sanger sequencing at DNA level, and Western Blot at protein level. 15 million K562-SuperFi-Cas9 cells were resuspended in 6 mL RPMI in the presence of 8 μg / mL polybrene and then split into 6-well plates with 2.5 million cells per well. The cells were infected by Lig2-ON-OFF-Barcode-Oligo lentivirus at MOI=5 by spinfection, which was conducted by spinning cells at 600g for 2 hours at 32 ℃. At the end of spinfection, 1 mL of pre-warmed RPMI with 15%FBS and 8 μg / mL polybrene was immediately added into each well. 36 hours post-transduction, the transduced K562-SuperFi-Cas9 cells were selected under 3 μg / mL puromycin for 2 days. 84 hours post-transduction, the K562-SuperFi-Cas9 cells that were successfully transduced by the Lig2-ON-OFF-Barcode-Oligo lentivirus were harvested and ready for NGS library preparation.
[0169] In total, 36 million cells were harvested from the 1st replicate and 40 million cells were harvested from the 2nd replicate. To prepare the NGS library for the gRNA-target pair and barcode cassette, genomic DNA (gDNA) was extracted from the collected cells using DNeasy Blood &Tissue Kits (Qiagen) . The gRNA-target pairs and barcode cassette were amplified by multiple PCR reactions to use up the extracted gDNA. Each 50 μL reaction included 2.5 μg gDNA, 250 nM forward primer (Fwd-libseq) , 250 nM reverse primer (Rev-libseq) , 25 μL NEBNext Ultra II Q5 Master Mix (New England Biolabs) , and nuclease-free H2O. The PCR program was set as following: 30 s at 98 ℃; 10 cycles of 10 s at 98 ℃, 30 s at 64 ℃, and 10 s at 72 ℃; 12 cycles of 10 s at 98 ℃, 30 s at 71 ℃, and 10 s at 72 ℃; 2 mins at 72 ℃; 4 ℃ hold. The PCR products were combined and concentrated to 100-200 μL by Amicon Ultra 0.5-mL Centrifugal Filters (Sigma) . SPRI beads (Beckman Coulter) were used to select the products with expected size at 420 bp. The size-selected products were subjected for NGS sequencing. The primers used for the NGS library preparation are listed in Table 4.
[0170] Bioinformatic analyses. The data analysis was conducted following the same pipeline as described in Zhang et al18. In brief, sequencing reads were subjected for adaptor removal, alignment and quality control using in-house Python scripts. The gRNA-target and barcode cassette in the plasmid library was used as starting reference to determine the sequence species that pre-existed before the high throughput screening in cells, and thus non-existing gRNA-target and barcode cassettes were considered as errors introduced by plasmid or PCR amplification and removed from the screened (edited) library. The gRNA-target and barcode cassette in the plasmid library was also used to determine the barcode sequences that associated with a gRNA. A cutoff of > 100 sequencing reads and > 5 BC reads was used to filter high-quality reads. In total of 62, 083 gRNA passed quality filtration and were retained in the following analysis and modeling.
[0171] To determine the indel frequency of each gRNA on their synthetic target, the number of sequencing reads aligned to the 20 base-pair synthetic target sequence were counted in the edited library. The indel frequency of each gRNA was calculated as:
[0172] Deep-learning modeling. Pre-train followed by fine-tune was used to train the prediction model based on the indel frequency data of SuperFi-Cas9 / gRNA based on the 6-mer pre-training model from DNABERT.
[0173] Determining indel efficiency at endogenous genomic locations. Twenty endogenous targets were chosen to verify the editing efficiency of gRNAs. For each target, gRNAs were designed with 21-24-nucleotide matched sequences and a 21-nucleotide gRNA with an unmatched 5′ nucleotide (Table 4) . The gRNAs were cloned into a V6-opt-test1 backbone vector. A tRNA was placed between the U6 promoter and gRNA, allowing the incorporation of the designed leading nucleotide at the 5′ of the expressed gRNA.
[0174] K562-SuperFi-Cas9 cells were cultured in 10%FBS and 5 μg / mL blasticidin in RPMI1640 (Sigma R8758) . For electroporation, 1 million cells were resuspended in 98 μL LONZA SF Nucleofector Solution (with the provided supplement supplied in the SF Cell Line 4D-NucleofectorTM X Kit) together with 1.5-5 μg plasmids. Cells were electroporated using program FF-120 of the LONZA 4D-NucleofectorTM X Unit following manufacturer’s instructions. Pre-warmed (37 ℃) culture media was added into cells immediately after electroporation, and cells were transferred and cultured in a 24-well plate.
[0175] For RNP electroporation, gRNAs in different length were ordered from GenScript or Azenta. gRNAs were resuspended to 100 μM using nuclease-free H2O (Thermo 10977015) . Cas9 protein was ordered from Takara (632641) . 1 μL 100 μM gRNA and 6 μg Cas9 protein were vortexed to mix then incubated at 25 ℃ for 20 min. The RNP complex was used to transfect 0.2 million K562 cells using LONZA SF Nucleofector X kit (LONZA V4XC-2032) under program FF-120 following instruction.
[0176] Cells were harvested 48 hours post electroporation, and gDNA was extracted using the Blood / Cells gDNA extraction kit from TIANGEN (TIANGEN DP304-03) . Amplicon-seq was then used to quantify the editing efficiency of validating gRNA. In brief, the target sequences were enriched from the gDNA by two rounds of PCR amplification. The 1st round of PCR was conducted in a 50 μL reaction, including 150 ng gDNA (~1 μL) , 250 nM forward primer, 250 nM reverse primer, 25 μL of Equinox Amplification Master Mix (2X) (WATCHMAKER GENOMICS 7K0014) , and nuclease-free H2O. The PCR program was set as following: 45 s at 98 ℃; 32 cycles of 15 s at 98 ℃, 30 s at 60 ℃, and 15 s at 72 ℃; 2 mins at 72 ℃; 4 ℃ hold. The 2nd round of PCR was conducted in a 50 μL reaction, including 2.5 μL products of the 1st round of PCR, 250 nM forward primer, 250 nM reverse primer, 25 μL of Equinox Amplification Master Mix (2X) (WATCHMAKER GENOMICS 7K0014) , and nuclease-free H2O. The PCR program was set as following: 45 s at 98 ℃; 12 cycles of 15 s at 98 ℃, 30 s at 62 ℃, and 15 s at 72 ℃; 2 mins at 72 ℃; 4 ℃ hold. AMPure XP beads (Beckman Coulter) was used to purify the PCR products, which were subjected for NGS sequencing. The primers used for the NGS library preparation was listed in Table 4.
[0177] The Amplicon-seq data was analyzed using CRISPResso 2.13 (--max_paired_end_reads_overlap 140 --min_paired_end_reads_overlap 10 --exclude_bp_from_left 0 --exclude_bp_from_right 0 --plot_window_size 40 --min_frequency_alleles_around_cut_to_plot 0.1)21.
Claims
1.A system comprising:a. a SuperFi-Cas9 protein or variant thereof; andb. a guide RNA (gRNA) comprising a polynucleotide sequence encoding a spacer that is complementary to a target nucleic acid sequence and 20 nucleotides in length, wherein the spacer further comprises a 5′ extension.2.The system of claim 1, wherein the SuperFi-Cas9 protein comprises the amino acid sequence of SEQ ID NO: 1, or wherein the variant thereof comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%identity to SEQ ID NO: 1.3.The system of claim 1 or 2, wherein the 5′ extension is 1 nucleotide in length.4.The system of claim 1 or 2, wherein the 5′ extension is 2 nucleotides in length.5.The system of claim 1 or 2, wherein the 5′ extension is 3 nucleotides in length.6.The system of claim 1 or 2, wherein the 5′ extension is 4 nucleotides in length.7.The system of any one of claims 3-6, wherein the 5′ extension is complementary to the target nucleic acid sequence.8.The system of any one of claims 1-7 wherein the gRNA is capable of binding to SuperFi-Cas9 or a variant thereof.9.The system of claim 8, wherein the 5′ extension does not decrease SuperFi-Cas9 binding to the gRNA relative to SuperFi-Cas9 binding to a gRNA that does not comprise a 5′ extension.10.A vector comprising:a. a polynucleotide encoding the SuperFi-Cas9 protein or variant thereof of claim 1 or 2; andb. a polynucleotide encoding the gRNA of any one of claims 1-9.11.The vector of claim 10, further comprising at least one regulatory element operably linked to (a) , (b) , or (a) and (b) .12.The vector of claim 10 or 11, wherein the vector is a plasmid or a viral vector.13.A ribonucleoprotein comprising:a. the SuperFi-Cas9 protein or variant thereof of claim 1 or 2; andb. a polynucleotide encoding the gRNA of any one of claims 1-9.14.A system comprising:a. the SuperFi-Cas9 protein or variant thereof of claim 1 or 2;b. a polynucleotide encoding the gRNA of any one of claims 1-9 or a prime editing guide RNA (pegRNA) ; andc. one or more effector molecules.15.The system of claim 14, wherein the pegRNA comprises a spacer that is complementary to a target nucleic acid sequence and wherein the spacer is 20, 21, 22, 23, or 24 nucleotides in length.16.The system of claim 14 or 15, wherein the pegRNA is capable of binding to SuperFi-Cas9 or a variant thereof.17.The system of any one of claims 14-16, wherein the one or more effector molecules are selected from a ribonuclease, a nickase, a base editor, an epigenetic modifier, a transposase, a recombinase, and a reverse transcriptase.18.The system of claim 14 or 17, wherein the one or more effector molecules comprise a base editor.19.The system of claim 18, wherein the base editor is a cytosine deaminase or an adenine deaminase.20.The system of claim 19, wherein the system comprises a SuperFi-Adenine base editor (SuperFi-ABE) comprising an amino acid sequence of SEQ ID NO: 2, or having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%identity to SEQ ID NO: 2.21.The system of claim 19, wherein the system comprises a SuperFi-Cytosine base editor (SuperFi-CBE) comprising an amino acid sequence of SEQ ID NO: 3, or having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%identity to SEQ ID NO: 3.22.The system of any one of claims 14-17, wherein the system comprises a pegRNA and the one or more effector molecules comprise a reverse transcriptase.23.The system of claim 22, wherein the system comprises a SuperFi-Prime editor (SuperFi-PE) having an amino acid sequence of SEQ ID NO: 4, or having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%identity to SEQ ID NO: 4.24.A vector comprising:a. a polynucleotide encoding the SuperFi-Cas9 protein or variant thereof of claim 1 or 2; andb. a polynucleotide encoding the gRNA of any one of claims 1-9 or the pegRNA of any one of claims 14-15; andc. a polynucleotide encoding the one or more effector molecules of any one of claims 14-23.25.The vector of claim 24, further comprising at least one regulatory element operably linked to one or more of (a) - (c) .26.The vector of claim 24 or 25, wherein the vector is a plasmid or a viral vector.27.A ribonucleoprotein comprising:a. the SuperFi-Cas9 protein or variant thereof of claim 1 or 2;b. a polynucleotide encoding the gRNA of any one of claims 1-9 or the pegRNA of any one of claims 14-16; andc. the one or more effector molecules of any one of claims 14-23.28.A method for targeting a nucleic acid sequence, comprising contacting the nucleic acid sequence with the system of any one of claims 1-9 or 14-23, the vector of any one of claims 10-12 or 24-26, and / or the ribonucleoprotein of claim 13 or 27, wherein the spacer hybridizes with the nucleic acid.29.The method of claim 28, wherein targeting the nucleic acid sequence comprises one or more of cutting and / or cleaving the nucleic acid sequence, nicking the nucleic acid sequence, enhancing expression of the nucleic acid sequence, decreasing expression of the nucleic acid sequence, visualizing or detecting the nucleic acid sequence, labeling the nucleic acid sequence, binding the nucleic acid sequence, enriching the nucleic acid sequence, depleting the nucleic acid sequence, editing the nucleic acid sequence, trafficking the nucleic acid sequence, splicing the nucleic acid sequence, and / or masking the nucleic acid sequence.30.The method of claim 28 or 29, wherein the nucleic acid sequence is a DNA sequence.31.The method of claim 30, wherein the DNA sequence is double stranded.32.The method of claim 30 or 31, wherein the DNA sequence is chromosomal DNA.33.The method of any one of claims 30-32, wherein targeting the DNA sequence induces a double-stranded break.34.The method of claim 33, wherein the DNA sequence is edited through non-homologous end joining or homologous recombination.35.The method of claim 28 or 29, wherein the nucleic acid sequence is an RNA sequence.36.The method of any one of claims claim 28-35, wherein the method increases indel frequency for SuperFi-Cas9 relative to when a gRNA is 20 nucleotides in length and does not comprise a 5′ extension.37.A cell comprising the system of any one of claims 1-9 or 14-23, the vector of any one of claims 10-12 or 24-26, and / or the ribonucleoprotein of claim 13 or 27.38.The cell of claim 37, wherein the cell is a eukaryotic cell.39.The cell of claim 38, wherein the cell is a mammalian cell.40.The cell of claim 38, wherein the cell is a stem cell.41.The cell of claim 38, wherein the cell is a somatic cell.
Citation Information
Patent Citations
Extended single guide RNA and use thereof
EP3744844A1
RNA Modification to Engineer Cas9 Activity
US20150376587A1