Fusion protein with expanded pam recognition range and uses thereof
By fusing the non-PI domain of xCas9 with the PI domain of a SpCas9 variant, a novel fusion protein was constructed, which solved the problem of the limited PAM recognition range of the SpCas9 system and enabled efficient editing of multiple PAM targets.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SICHUAN UNIV
- Filing Date
- 2022-12-29
- Publication Date
- 2026-05-01
AI Technical Summary
The existing SpCas9 system exhibits bias in recognizing PAM sequences, which limits its target recognition range and makes it difficult to effectively edit non-NGG type PAM targets.
By fusing the non-PI domain of xCas9 with the PI domain of SpCas9 variants, novel fusion proteins, including Cas9-NG, non-G SpCas9, and SpRY, were constructed, expanding the PAM recognition range and improving editing efficiency.
It enables efficient editing of PAMs such as NRN, NRCH, NRTH, and NRTH, expands the PAM recognition range, and improves gene editing efficiency.
Smart Images

Figure CN116286736B_ABST
Abstract
Description
A fusion protein with expanded PAM recognition range and its applications
[0001] This application claims priority to Chinese invention patent application CN202210047419.7, filed on January 17, 2022, entitled "Application of a fusion protein in base editing", which is incorporated herein by reference in its entirety. Technical Field
[0002] This invention relates to the field of protein engineering, and more specifically to a fusion protein with an expanded PAM recognition range and its applications. Background Technology
[0003] Clustered regularly spaced short palindromic repeats (CRISPR), together with CRISPR-associated proteins (Cas), constitute the adaptive immune system of microorganisms, protecting them from invading bacteriophages and plasmids. PAMs (protospacer adjacent motifs) play a crucial role in immune acquisition and target recognition. Utilizing the target recognition mechanism of the CRISPR system to edit desired targets enables gene editing in a variety of organisms. Among the discovered CRISPR systems, the Cas9 system (SpCas9), derived from Streptococcus pyogenes, has become the most widely used gene editing system due to its simplicity and high efficiency in mammalian cells.
[0004] Typically, wild-type SpCas9 recognizes purine-rich PAMs, favoring NGG PAMs (where N = A, T, G, or C), thus limiting its target recognition range. Structural studies have revealed that the PI domain (particularly the β5-β7 plate) is responsible for sequence-specific binding of PAMs. To alter SpCas9's PAM preference, a set of SpCas9 variants with certain amino acid substitutions within the PI domain have been designed based on the structural characteristics of SpCas9. Through protein engineering, several variants have achieved novel PAM specificities. For example, the VQR, EQR, and VRER SpCas9 variants change the PAM specificity from NGG to NGA, NGAG, or NGCG, respectively. Non-GSpCas9s include SpCas9-NRRH, SpCas9-NRTH, and SpCas9-NRCH, which collectively recognize NRNH PAMs (where R = A or G, H = A, C, or T). Other variants, including Cas9-NG and SpRY, relax the PAM restriction from NGG to NGN and NNN (NRN > NYN PAMs), respectively. Most of the mutations in these variants are concentrated in the PI domain. Summary of the Invention
[0005] In a first aspect, the present invention provides a fusion protein comprising:
[0006] a) The amino acid sequence from the non-PI domain of xCas9;
[0007] b) The amino acid sequence of the PI domain from the SpCas9 variant.
[0008] Furthermore, the SpCas9 variants include Cas9-NG, non-G SpCas9, or SpRY.
[0009] Furthermore, the non-G SpCas9 includes one or more of SpCas9-NRRH, SpCas9-NRCH, and SpCas9-NRTH.
[0010] Furthermore, the amino acid sequence from the non-PI domain of xCas9 contains the following mutations: A262T / R324L / S409I / E480K / E543D / M694I.
[0011] Furthermore, the amino acid sequence of the PI domain from Cas9-NG contains the following mutations: R1335V / L1111R / D1135V / G1218R / E1219F / A1322R / T1337R mutations.
[0012] Furthermore, the amino acid sequence of the fusion protein includes one or more of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4 and SEQ ID NO:114.
[0013] Furthermore, the fusion protein further includes a heterologous functional domain.
[0014] Furthermore, the heterologous functional domain is a deaminase domain.
[0015] Furthermore, the deaminase domain includes adenine deaminase or cytosine deaminase.
[0016] In a second aspect, the present invention provides a method for improving base editing efficiency and / or expanding the PAM recognition range, characterized in that the method includes coupling a non-PI domain from xCas9 and a PI domain from a SpCas9 variant.
[0017] Furthermore, the amino acid sequence from the non-PI domain of xCas9 contains the following mutations: A262T / R324L / S409I / E480K / E543D / M694I.
[0018] Furthermore, the SpCas9 variants include Cas9-NG, non-G SpCas9, or SpRY.
[0019] Furthermore, the non-G SpCas9 includes one or more of SpCas9-NRRH, SpCas9-NRCH, and SpCas9-NRTH.
[0020] The method provided by this invention can improve the editing efficiency of Cas9-NG in NRN PAM, and can improve the editing efficiency of non-GSpCas9 and SpRY in NRN PAM (where R = A or G). The method provided by this invention can expand the PAM recognition range of Cas9-NG to TTAH, TTCV, TTTA, CTTC, YCAM, TCAT, NCCC, MCCA, and TCTC PAM (where H = A, C, or T; V = A, C, or G; Y = T or C; M = A or C), expand the PAM recognition range of SpCas9-NRRH to TAHD, AAYT, AAAG, SACG, and GACA PAM (where H = A, C, or T; Y = T or C; S = G or C), and expand the PAM recognition range of SpCas9-NRCH to TTTN, TTMR, MTCA, CTAG, ATCG, TCHA, HCYG, BCAG, and CCCA PAM (where N = A, T, C, or G; M = A or C; H = A, C, or T; Y = T or C; B = G, T, or C).
[0021] Thirdly, the present invention also provides the use of the above-mentioned fusion protein in gene editing.
[0022] Furthermore, the uses include altering the gene sequence of target cells by delivering the fusion protein or the nucleotide sequence encoding the fusion protein in vivo or in vitro, to develop gene therapy drugs and / or cell therapy drugs.
[0023] Beneficial effects
[0024] This invention fuses the mutated portion of the non-PI domain of xCas9 with the mutated portion of the PI domain of SpCas9 variants (SpCas9-NG, SpCas9-NRRH, SpCas9-NRCH, SpCas9-NRTH, and SpRY) to construct new mutants (xCas9-NG, xCas9-NRRH, xCas9-NRCH, xCas9-NRTH, and xCas9-RY), i.e., fusion proteins. These fusion proteins exhibit more relaxed PAM restrictions than xCas9, Cas9-NG, non-G SpCas9, and SpRY, and significantly improve base editing efficiency. For example, the PAM recognition range of Cas9-NG is expanded from NG to NR, and xCas9-NG significantly improves the editing efficiency containing NAN PAM; non-G SpCas9 and SpRY also improve the editing efficiency of NRN PAM, and the PAM recognition range of non-G SpCas9 is further expanded. By delivering the fusion proteins provided by this invention in vivo or in vitro, the gene sequences of target cells can be altered, thereby enabling the development of gene therapy drugs and / or cell therapy drugs.
[0025] As used herein, the term "PAM (pre-spacer adjacent motif)" refers to a short DNA sequence of approximately 2–8 base pairs that is an important target region for nucleases such as Cas9. Typically, the PAM sequence is located on the non-target strand and downstream of the Cas9 cleavage site in the 5' to 3' direction.
[0026] As used herein, the terms "PAM interaction domain" or "PI domain" refer to the portion of the Cas9 protein that recognizes and binds to PAM sequences located on the non-target strand. "Non-PI domain" refers to the portion of the Cas9 protein excluding the PI domain, including, for example, the HNH domain (which cleaves nucleic acid strands complementary to the guide RNA), the RuvC domain (which cleaves nucleic acid strands not complementary to the guide RNA), and the REC domain (which recognizes the target).
[0027] As used herein, the term "non-G SpCas9" refers to a SpCas9 variant capable of identifying a non-G PAM sequence. In some embodiments, the non-G PAM may be NRRH, NRTH, or NRCH (where R = A or G; H = A, C, or T).
[0028] As used herein, “SpCas9 variant” refers to a protein that has any variation relative to the wild-type SpCas9 protein sequence. It should be noted that the term “variant” refers to a change in the wild-type protein sequence, regardless of whether the change alters the function of the protein (e.g., adds, reduces, or confers a new function) or whether the change has no effect on the protein’s function (e.g., the mutation or variant is silent). Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0030] Figure 1 shows an experimental plot comparing the editing efficiency of xCas9, Cas9-NG, and xCas9-NG on endogenous targets containing NGN PAMs.
[0031] Figure 2 shows the experimental results of characterizing the PAM preferences of xCas9, Cas9-NG, and xCas9-NG by editing three independent PAM libraries using a cytosine base editor;
[0032] Figure 3 shows an experimental diagram comparing the editing efficiency of non-G SpCas9s and SpRY and their xCas9 modified variants at endogenous gene sites;
[0033] Figure 4 shows the experimental results of characterizing the PAM preference of xCas9-modified non-G SpCas9s variants by editing the PAM library using a cytosine base editor;
[0034] Figure 5 is an experimental graph comparing the editing efficiency of non-G SpCas9s and their x-Cas9 modified variants (taking SpCas9-NRCH as an example). Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0036] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0037] As used in this specification, the term "about" typically means + / - 5% of the value, more typically + / - 4%, more typically + / - 3%, more typically + / - 2%, even more typically + / - 1%, even more typically + / - 0.5% of the value.
[0038] In this specification, some embodiments may be disclosed in a range-bound format. It should be understood that this "range-bound" description is merely for convenience and brevity and should not be construed as a rigid limitation on the disclosed range. Therefore, the description of a range should be considered as specifically disclosing all possible subranges and the individual numerical values within that range. For example, a description of the range 1–6 should be considered as specifically disclosing subranges such as 1–3, 1–4, 1–5, 2–4, 2–6, 3–6, etc., and the individual numbers within that range, such as 1, 2, 3, 4, 5, and 6. This rule applies regardless of the breadth of the range.
[0039] Detailed description of the attached figures
[0040] Figure 1: (a) Schematic diagram of the SpCas9 / RNA / DNA complex (PDB ID: 4UN3). Mutant amino acids in the non-PI domain are marked with red spheres (A262T, R324L, S409I, E480K, E543D, M694I), and E1219V in the PI domain is marked with green spheres. The PI domain is shown in gray, and the non-PI domain is shown in yellow. (b) Location of mutated amino acids in SpCas9 variants. The D10A mutation, which disrupts the activity of the target strand nuclease, is marked with a blue background. Amino acids in the non-PI domain are marked in gray, and amino acids in the PI domain are marked in orange. ce shows the improved editing efficiency of xCas9-NG on PAMs of NGN, NAN, and NYN (compared to Cas9-NG and xCas9). (c) Adenine base editing efficiency of xCas9, Cas9-NG, and xCas9-NG at endogenous gene sites containing NGN PAMs. The x-axis markings represent the location of target A (NGN PAM counted as 21-23). Values and error bars reflect the mean and SD of three independent biological samples. Differences between groups were tested using multiple t-tests. ns P>0.05, *P<0.05, **P<0.01, ***P<0.001. (d) Adenine base editing efficiency of xCas9, Cas9-NG, and xCas9-NG at endogenous sites containing NAN PAMs. Values and error bars reflect the mean and SD of three biologically independent samples. Differences between groups were tested using multiple t-tests. ns P>0.05, *P<0.05, **P<0.01, ***P<0.001. The location of target A is shown (NGN PAMs are counted as 21-23). (e) Adenine base editing at endogenous sites containing NTN and NCN PAMs by xCas9, Cas9-NG, and xCas9-NG. Values and error bars reflect the mean and SD of three biologically independent samples. Differences between groups were tested using multiple t-tests. ns P>0.05, *P<0.05, **P<0.01, ***P<0.001. The x-axis markings represent the position of target A (NGN PAM counted as 21-23).
[0041] Figure 2: (a) Schematic diagram of the PAM screening assay using cytosine base editing. Three independent PAM libraries containing eight random bases on the sides of the spacer region were used for screening. (b) Violin plot showing the cytosine base editing activity of xCas9, Cas9-NG, and xCas9-NG on the PAM libraries. PAMs were classified according to the type of the second and / or third base. Solid lines represent the median, and dashed lines represent the first and third quartiles. On NNNN PAMs, xCas9-NG showed higher editing efficiency than Cas9-NG and xCas9. (c) Heatmap showing the preference of Cas9-NG, xCas9, and xCas9-NG for the first–4 bases of 8N PAMs. The values for each NNNN PAM are the average editing efficiency of the three libraries. In summary, xCas9-NG expands the recognition range of PAMs.
[0042] Figure 3: (a) Amino acids mutated in SpCas9 variants (xCas9, xCas9-NRTH, xCas9-NRRH, xCas9-NRCH, and xCas9-RY(R61A)). The D10A mutation, which disrupts the activity of the target strand nuclease, is marked with a blue background; amino acids located in the non-PI domain are marked in gray, and amino acids located in the PI domain are marked in orange. (b) Adenine base editing efficiency of xCas9, SpCas9-NRTH, and xCas9-NRTH at endogenous gene sites containing NRTN PAMs. The location of target A is shown (NGN PAMs are counted as 21–23). Values and error bars reflect the mean and SD of three biologically independent samples. Differences between groups were tested using multiple t-tests. ns P>0.05, *P<0.05, **P<0.01, ***P<0.001. (c) Adenine base editing efficiency of xCas9, SpCas9-NRRH, and xCas9-NRRH at endogenous sites containing NRRN PAMs. (d) Adenine base editing efficiency of xCas9, SpCas9-NRCH, and xCas9-NRCH at endogenous sites containing NRCN PAMs. (e) Adenine base editing efficiency of xCas9, SpRY, and xCas9-RY(R61A) at endogenous sites containing NAN, NGN, or NCN PAMs. In summary, the editing efficiency of xCas9-modified variants is improved at most endogenous sites.
[0043] Figure 4: Violin plots show the base editing activities of SpCas9-NRRH(a), SpCas9-NRCH(c), SpCas9-NRTH(e), and their xCas9-modified counterparts on three PAM libraries. PAMs are classified according to the type of the second and / or third base. Solid lines represent the median, and dashed lines represent the first and third quartiles. Heatmaps show the preference of SpCas9-NRRH(b), SpCas9-NRCH(d), SpCas9-NRTH(f), and their xCas9-modified counterparts for the first to fourth bases of 8N PAMs. The value for each NNN PAM is the average of the editing efficiencies of the three libraries. In summary, the range of PAMs identified by the xCas9-modified non-G SpCas9 variants is expanded compared to the original Cas9.
[0044] Example 1: Materials and Methods
[0045] 1.1. Plasmid Construction: The steps for constructing the ABE / CBE3-xcas9-NG vector include: (1) designing primers; (2) using ABE7.10 / CBE3-NG and ABE7.10 / CBE3-xCas9 as templates, and... (3) Amplification using the Max Super-Fidelity DNA Polymerase system to obtain the target fragment; (4) Gel recovery of the amplified target fragment; (5) Seamless cloning and ligation; (6) Transformation into E. coli (DH5α), plating, and selection of single clones for sequencing verification.
[0046] The steps for constructing sgRNA include: (1) designing primers based on the spacer sequence; (2) phosphorylation at 37℃ for half an hour, followed by natural cooling to room temperature at 95℃; (3) digestion of SpCas9 sgRNA-scaffold with Bbs1; (4) ligation with T4 enzyme; (5) transformation with E. coli (DH5α), plating, and selecting single clones for sequencing verification.
[0047] Plasmids encoding pCMV-CBE3 (#73021), xCas9 (3.7)-ABE (7.10) (#108382), and NG-ABEmax (#124163) were obtained from addgene. Plasmids encoding pUC19-Cas9-NRCH, pUC19-Cas9-NRTH, and pUC19-Cas9-NRRH were synthesized by General Biosystems Ltd. The corresponding spacer sequences were inserted into SpCas9sgRNA-scaffolds digested with BBSI. The remaining plasmids (including ABE-xCas9-NG, ABE-xCas9-NRRH, ABE-xCas9-NRCH, ABE-xCas9-NRTH, CBE3-NG, CBE3-xCas9, CBE3-xCas9-NRRH, CBE3-xCas9-NRTH, CBE3-xCas9-NRCH, etc.) were constructed using seamless cloning (ClonExpress II One Step Cloning Kit, Vazyme Biotech Co., Ltd.). All plasmids were validated using Sanger sequencing.
[0048] 1.2. Cell Culture: HEK 293T cells were cultured in Dulbecco's modified Eagles's medium (ThermoFisher Scientific) supplemented with 10% (v / v) fetal bovine serum (Life Technologies) and 1% penicillin / streptomycin (Boster Biological Technology Co., Ltd). The temperature was maintained at 37°C, 100% humidity, and 5% CO2 in a CO2 incubator. Cells were passaged every 48-72 hours until they reached 95% confluence.
[0049] 1.3. Plasmid Transfection: HEK293T cells were seeded into 96-well plates (BIOFIL). 16-24 hours after seeding, when cell confluence was approximately 70-80%, cells were transfected using Transeasy plasmid transfection technology. TMCells were transfected using Transeasy (Forgene). For ABE (Adenine Base Editors) and CBE (Cytosine Base Editors) experiments, 300 ng of base editor expression plasmid and 100 ng of sgRNA expression plasmid were transfected with 0.75 μl Transeasy (Forgene). For PAM library editing experiments, 300 ng of base editor expression plasmid, 100 ng of sgRNA expression plasmid, and 45 ng of PAM library plasmid were transfected with 0.75 μl Transeasy (Forgene). 72 h post-transfection, the culture medium was removed, and cells were washed three times with 1×PBS (Meilunbio). Genomic DNA was then extracted using 30 μl of freshly prepared lysis buffer (10 mM Tris-HCl, pH 7.5; 0.05% SDS; 20 μg / mL Proteinase K (Beyotime)). The mixture was incubated at 55 °C for 10 min, followed by heat inactivation at 95 °C for 10 min. The resulting DNA lysis buffer was stored at -20°C for PCR amplification and subsequent analysis.
[0050] 1.4. Base editing analysis using Sanger sequencing and EditR software: The target genomic region was amplified by PCR and analyzed using Sanger sequencing. The sequencing map was then further quantified using EditR software (https: / / moriaritylab.shinyapps.io / editr_v10 / ). Primers used to amplify each target site are listed in Table 2. In this invention, "target site" and "target" have the same meaning unless otherwise stated.
[0051] Table 2: Primers used to amplify each target site.
[0052]
[0053]
[0054]
[0055] 1.5. PAM Library
[0056] A PAM library containing an 8-nt random PAM sequence at the 3' end of the target site was generated by PCR (Table 3). The obtained PCR product and pUC19 vector were digested with EcoRI and HindIII, and then ligated using T4 ligase (TAKARA). The ligated plasmids were transformed into DH5α cells by electroporation. The transformed cells were grown overnight in 100 ml LB medium at 37°C. The PAM library was extracted using the midi-prep plasmid extraction kit (Omega).
[0057] Table 3: Primers used for PAM library construction
[0058]
[0059]
[0060] 1.6. PAM Screening Method for High-Throughput Sequencing (HTS)
[0061] Cells were harvested 72 hours post-transfection, and genomic DNA was extracted using freshly prepared lysis buffer. The target genomic region was amplified by PCR, with different barcodes on the primer sides (Table 4). Products were purified using a DNA gel extraction kit (Thermo Fisher Scientific) and quantified using a Nanodrop spectrophotometer (Thermo-fisher). Samples were commercially sequenced using an Ilumina Novaseq 6000 platform (Personalbio, Shanghai, China). Sequencing reads were analyzed using Microsoft Excel. First, an 8-bp fragment upstream of the target sequence was used to identify reads containing spacer sequences. Then, the identified reads were grouped into 256 groups of individual 4-bp PAMs. The editing efficiency of each PAM was calculated based on the 6th base of the spacer sequence in each group (C = non-edited, T = edited). The resulting efficiencies were then analyzed and plotted using GraphPadPrism 9 software.
[0062] Table 4: HTS primers used in PAM screening
[0063]
[0064]
[0065]
[0066] Example 2: Non-PI fragments of xCas9 improve the editing efficiency of PI domains of Cas9-NG for NRN PAM.
[0067] To test whether the non-PI fragment of xCas9 expands the PAM restriction of Cas9 variants, this embodiment constructs an xCas9 and Cas9-NG mosaic protein in which a non-PI fragment from xCas9 (amino acids 1-1098, containing A262T / R324L / S409I / E480K / E543D / M694I substitutions) is fused with a PI domain from Cas9-NG (amino acids 1099-1368, containing R1335V / L1111R / D1135V / G1218R / E1219F / A1322R / T1337R substitutions) (Figures 1a-1b). The resulting variant (referred to as xCas9-NG) is generated on the ABE7.10 backbone (ABE-xCas9-NG) to facilitate the quantification of editing efficiency. Table 1 lists the SpCas9 variants involved in this invention.
[0068] Table 1: SpCas9 variants involved in this invention.
[0069]
[0070]
[0071]
[0072]
[0073]
[0074] Note: Underlined: Amino acid resulting from mutation.
[0075] To compare the editing activity of ABE-xCas9-NG with ABE-Cas9-NG (hereinafter referred to as "ABE-NG") and ABE-xCas9, these constructs were co-transfected with a set of 30 sgRNAs (13 NGN PAMs and 16 NAN PAMs) targeting endogenous genomic sites. Since both Cas9-NG and xCas9 prefer NGN PAMs, this example begins with comparisons using targets containing NGN PAMs. As shown in Figure 1c, ABE-xCas9-NG outperformed ABE-Cas9-NG in editing efficiency across almost all tested NGN targets, but not ABE-xCas9. The average editing efficiency of ABE-xCas9-NG reached 48.23±19.8% for the NGG target, 23.95±11.45% for the NGA target, 15.95±6.9% for the NGC target, and 25.73±9.3% for the NGT target, which were 1.47, 1.24, 1.14, and 1.16 times that of ABE-Cas9-NG, respectively. In contrast, ABE-xCas9-NG showed comparable editing activity to ABE-xCas9 for the NGG (xCas9-NG:xCas9 = 46.33%: 40.00%) and NGC (18.07%: 15.07%) targets, while showing higher activity for the NGA (xCas9-NG:xCas9 = 24.28%: 13.56%) and NGT (26.10%: 12.67%) targets.
[0076] Then, their performance was tested on 16 endogenous targets containing NAN PAMs, which could be identified by Cas9-NG and xCas9 but were not highly efficient at editing (i.e., not favored). Notably, ABE-xCas9-NG showed significantly higher activity than ABE-Cas9-NG or ABE-xCas9 in most of the validated NAN targets (Fig. 1d). In the NAG target, ABE-xCas9-NG had an average editing efficiency of 32.53 ± 15.5%, which was approximately 3.32 and 4.74 times higher than ABE-Cas9-NG (9.8 ± 3.9%) and ABE-xCas9 (6.87 ± 4.8%), respectively. All three ABEs showed only mild activity in the remaining NAN (NAA, NAC, and NAT) targets (Fig. 1d). Among them, ABE-xCas9-NG was effective at 3 out of 5 NAT targets (which can also be understood as "statistically significant"), with an average editing efficiency of 15.78±8.41%, while the average editing efficiencies of ABE-Cas9-NG and ABE-xCas9 were 8.33±2.5% and 6.56±2.4%, respectively. That is, ABE-xCas9-NG increased the efficiency by 1.9 and 2.41 times compared with ABE-NG and ABE-xCas9, respectively. ABE-xCas9-NG was effective at only 1 out of 4 NAA targets and only 1 out of 4 NAC targets, but its editing efficiency was higher than that of ABE-Cas9-NG and ABE-xCas9 in both cases. The results show that the xCas9-NG, after introducing a mutation in the non-PI domain of xCas9, not only improves the editing efficiency in NGN PAMs but also expands the PAM recognition range from NGN to NAN. Furthermore, its improvement effect on NAN is superior to that of NGN, especially compared to xCas9 (xCas9's editing efficiency is almost negligible at all 16 NAN PAM targets). The overall trend of the three editors' editing effects on NAN PAMs is ABE-xCas9-NG > ABE-Cas9-NG > ABE-xCas9. In addition, this embodiment also evaluated the editing efficiency of ABE-xCas9-NG on 5 NYN PAMs, finding that ABE-xCas9-NG is effective at one site containing GCTG PAMs. In contrast, ABE-Cas9-NG only performed slight editing at the same site, while ABE-xCas9 showed poor editing performance at this site (Figure 1e). The editing efficiency of CBE3-xCas9-NG on 5 NYN PAMs was also evaluated, with results showing effectiveness only in one GTGC PAM.
[0077] Furthermore, studies on the editing window show that ABE-xCas9-NG has a similar editing window to ABE-xCas9 and ABE-Cas9-NG (bases 4 to 8, with PAM counts of 21-23).
[0078] This embodiment also constructed corresponding cytosine base editors, including CBE3-xCas9, CBE3-Cas9-NG, and CBE3-xCas9-NG. These base editors were directly compared using 10 endogenous target sites. The results showed that CBE3-xCas9-NG had higher editing efficiency at most target sites than CBE3-xCas9 and CBE3-Cas9-NG. The average base editing efficiency of CBE3-xCas9-NG was 1.91 times and 1.19 times that of CBE3-NG and CBE3-xCas9 at NGN PAM sites, respectively. In summary, these results indicate that the non-PI fragments of xCas9 generally improve the editing efficiency of the PI domain of Cas9-NG on NRN PAM (R=G or A).
[0079] Example 3: Non-PI fragments of xCas9 expand the PAM recognition range of Cas9-NG.
[0080] Next, this embodiment attempts to test whether xCas9-NG expands the PAM recognition range beyond NRN. It starts with a set of 4 endogenous sites that have previously been shown to contain non-NRN PAMs and can be edited by the SpRY variant (xCas9-NG can only edit the target sites of GCTG PAMs (Figure 1e and data not shown)).
[0081] To gain a deeper understanding of the overall PAM recognition after the introduction of the non-PI domain of xCas9 into SpCas9-NG (xCas9-NG), PAM screening experiments were conducted using three different PAM libraries (8N random sequences located downstream of three independent spacer regions, with the sixth base of each spacer sequence being cytosine). Three target sites where the sixth base in the editing window is C were selected, and three PAM libraries containing NNNNNNNN were constructed using these target sites. The three 8N PAM libraries, along with the corresponding sgRNAs and cytosine base editors from Cas9-NG, xCas9, and xCas9-NG (CBE3-Cas9-NG (hereinafter referred to as "CBE3-NG"), CBE3-xCas9, and CBE3-xCas9-NG), were transfected into HEK293T cells. After the genome was collected, it was evaluated by high-throughput sequencing (Figure 2a).
[0082] Consistent with observations of endogenous targets, HTS analysis of the edited libraries showed that CBE3-xCas9-NG exhibited better editing performance (i.e., average editing efficiency) than both CBE3-xCas9 and CBE3-Cas9-NG in almost every PAM, whether on NRN or NYN PAMs (R includes bases A and G, Y includes bases C and T). To facilitate comparison of the activity of each Cas9 on different PAMs, the PAMs were categorized according to the type of the second and / or third base (Figure 2b). In terms of average editing efficiency, CBE3-xCas9-NG achieved 14.9% editing efficiency on NGN PAMs and 12.6% on NAN PAMs, which were comparable; this is significantly higher than that of CBE3-Cas9 NG (8.3% on NGN, 6.8% on NAN) or CBE3-xCas9 (9.9% on NGN, 5.6% on NAN). CBE3-xCas9 hardly recognizes NYN PAMs, but NYG appears to be more easily recognized than NYH (H includes bases A, T, and C) (Fig. 2b). On NGN PAMs recognized by both CBE3-Cas9-NG and CBE3-xCas9-NG, CBE3-xCas9-NG shows a significant improvement over both CBE3-Cas9-NG and CBE3-xCas9 (efficiency ranking from highest to lowest: CBE3-xCas9-NG > CBE3-Cas9-NG > CBE3-xCas9) (Fig. 2c). In NAN PAMs, CBE3-xCas9-NG is more likely to recognize NAG. CBE3-Cas9-NG shows no preference on NAN PAMs and its activity is lower than that of CBE3-xCas9-NG. For NTN PAMs, CBE3-xCas9-NG also exhibits much higher editing efficiency than both CBE3-xCas9 and CBE3-Cas9-NG. Both CBE3-Cas9-NG and CBE3-xCas9-NG show a preference for NTG PAM, with the latter being more efficient in editing.
[0083] To determine whether CBE3-xCas9-NG broadens the recognition of PAMs, a detailed analysis of the PAM preferences of CBE3-Cas9-NG, CBE3-xCas9, and CBE3-xCas9-NG was conducted (Figure 2c). For NRNN PAMs, both CBE3-Cas9-NG and CBE3-xCas9-NG can recognize them; compared to CBE3-Cas9-NG, CBE3-xCas9-NG shows improved editing efficiency (Figure 2c). For NTNN PAMs, CBE3-Cas9-NG does not recognize TTAH, TTCV, TTTA, and CTTC (H = A, C, or T; V = A, C, or G, with editing efficiency below 2%), but CBE3-xCas9-NG can recognize these PAMs, indicating that the CBE3-xCas9 modification broadens the recognition of these PAMs by CBE3-Cas9-NG. For NCNN PAMs, CBE3-xCas9-NG also demonstrated recognition of PAMs including YCAM, TCAT, NCCC, MCCA, TCCG, CCCT, CCTM, and TCTY (Y = C or T; M = A or C). Conversely, CBE3-xCas9 clearly failed to recognize NTN or NCN PAMs. Furthermore, while the C at positions 2 and 3 of the PAM is exclusive to all variants, CBE3-xCas9-NG can edit PAMs containing G at positions 1 and / or 4 (Figure 2c). In PAMs where the second digit is T, CBE3-Cas9-NG and CBE3-xCas9-NG show similar preferences for NTGN, with CBE3-xCas9-NG exhibiting improved efficiency compared to CBE3-Cas9-NG. For NTHN PAMs, in addition to improved efficiency for PAMs that CBE3-Cas9-NG could already recognize, CBE3-xCas9-NG also shows activity for some PAMs that Cas9-NG cannot recognize, such as VTAA, GTGG, GTTY, and MTTG. CBE3-Cas9-NG rarely recognizes NCNN PAMs. However, CBE3-xCas9-NG can recognize GCNG PAMs where the first and fourth digits are both G, and also shows activity for ACCA, ACGG, and ACTG PAMs. Among the 256 PAMs in NNNN, CBE3-xCas9-NG shows improved efficiency compared to CBE3-Cas9-NG for the vast majority of PAMs.
[0084] Example 4: The non-PI fragment of xCas9 improves the recognition ability of SpRY and non-G SpCas9s PI domains for NRN PAM.
[0085] Inspired by the xCas9-NG results, this embodiment next attempts to extend the above findings to other recently developed SpCas9 variants, including non-G SpCas9s and SpRY. Coupled with the non-PI domain fragments of xCas9 to the PI domains of non-G SpCas9s and SpRY, four variants were generated: xCas9-NRRH, xCas9-NRCH, xCas9-NRTH, and xCas9-RY(R61A) (Figure 3a). These xCas9-modified variants were then compared with the unmodified variants when editing different target sites.
[0086] The performance of x-non-G SpCas9s (including xCas9-NRRH, xCas9-NRCH, and xCas9-NRTH) was evaluated using adenine base editing. ABE-xCas9-NRTH was compared with SpCas9-NRTH or xCas9 in adenine base editing at eight NRTN PAM target sites, and ABE-xCas9-NRTH outperformed both ABE-NRTH and ABE-xCas9. Firstly, ABE-xCas9-NRTH exhibited higher editing activity than ABE-NRTH and ABE-xCas9 at most of these target sites (7 out of 8). In particular, ABE-xCas9-NRTH demonstrated potent activity in editing the IDH2-3 site (PAM=AGTT), IDH2-4 site (PAM=GGTG), PIK3CA-5 site (PAM=CATT), and IDH2-1 site (PAM=CATC), all of which could not be edited by ABE-NRTH or ABE-xCas9. At the target sites that can be edited by ABE-NRTH (the PAM for PIK3CA-3 is TGTT, the PAM for AKT1-6 is TGTC, and the PAM for AKT1-5 is GATC), the average editing efficiency of ABE-NRTH is 11.83±7.3%, and the average editing efficiency of ABE-xCas9-NRTH is 24.90±11.3% (Figure 3b).
[0087] The activities of ABE-NRRH, ABE-xCas9, and ABE-xCas9-NRRH at a set of five target sites containing NRRH PAM were then compared (Fig. 3c). However, this embodiment found that ABE-xCas9-NRRH was more efficient than ABE-NRRH and ABE-xCas9 at most sites (4 out of 5 sites) (Fig. 3c). Similar phenomena were observed when comparing ABE-xCas9-NRCH, ABE-xCas9-RY (R61A), and their corresponding unmodified variants, showing that xCas9 modification significantly improved the editing efficiency at sites identifiable by the unmodified variants (Fig. 3d and Fig. 3e). On average, xCas9 modification improved the editing efficiency of ABE-NRCH and ABE-RY by approximately 90% and 50%, respectively (Fig. 3d and Fig. 3e).
[0088] Example 5: Non-PI fragments of xCas9 enhance PAM recognition of non-G SpCas9s
[0089] Next, cytosine base editors from non-G SpCas9 and its xCas9-modified variants were constructed, and PAM screening experiments were performed using the same PAM library. Similar to the observations on endogenous targets, xCas9 modification not only generally improved the editing efficiency of each original version for identifiable PAMs, but also improved the editing efficiency for unidentifiable PAMs.
[0090] A direct comparison of CBE-NRRH and CBE-xCas9-NRRH revealed that, except for NGGN PAM, CBE-xCas9-NRRH exhibited higher editing activity than CBE-NRRH on all NRNN PAMs, with a more significant improvement on NANN PAM than on NGNN PAM (Figure 4a). For NCNN and NTNN PAMs, although neither CBE-xCas9-NRRH nor CBE-NRRH performed extensive editing, CBE-xCas9-NRRH showed slightly higher editing activity than CBE-NRRH (Figure 4a). Detailed analysis at each location showed that CBE-xCas9-NRRH had an expanded PAM recognition range compared to CBE-NRRH (Figure 4b). For example, for NANN PAMs, CBE-xCas9-NRRH could recognize TAHD, AAYT, AAAG, SACG, and GACA PAMs, while SpCas9-NRRH could not.
[0091] A comparison between CBE-NRCH and CBE-xCas9-NRCH shows that the latter performs significantly better than the former in almost all PAMs (Figure 5). CBE-NRCH primarily edits NRNN PAMs, with little editing activity on NYNN PAMs (Figure 4c). In terms of editing efficiency, CBE3-xCas9-NRCH outperforms CBE3-NRCH on NRNN PAMs (12.2% vs. 8.4%). Furthermore, CBE-xCas9-NRCH exhibits significantly higher editing activity on NYNNs than CBE-NRCH (9.7% vs. 3.9%), and its editing efficiency on NYNN PAMs is comparable to its editing efficiency on NRNN PAMs (Figures 4c and 4d). CBE-xCas9-NRCH expands its recognition range for NYNN PAMs to include TTTN, TTMR, MTCA, CTAG, ATCG, TCHA, HCYG, BCAG, and CCCA (Figure 5). A comparison between CBE-NRTH and CBE-xCas9-NRTH shows that they both favor NRNN to NYNN PAM (Figure 4f). However, CBE-xCas9-NRTH is more efficient than CBE-NRTH (9.8% vs. 6.2%) (Figure 4e).
[0092] This invention constructs a series of SpCas9 variants, in which non-PI domains from xCas9 3.7 are fused with PI domains from Cas9-NG, SpRY, and non-G, respectively. Compared to the original version, the resulting xCas9-modified variants exhibit higher editing activity and more PAM recognition. These results demonstrate that non-PI domain fragments contribute to PAM constraint in Cas9 systems and provide novel editing tools with enhanced activity and expanded target range.
[0093] Summarize:
[0094] While the PAM restriction of the Cas9 system allows bacteria to distinguish between the host genome and the invading genome, it also limits the target range of the Cas9 protein for genome editing. The PI domain, which interacts directly with the PAM sequence, is generally considered the sole reason for PAM specificity.
[0095] This invention begins with a widely used variant, Cas9-NG, and constructs a mosaic SpCas9 protein (xCas9-NG) by fusing the PI domain of Cas9-NG with the non-PI portion of xCas9. Unexpectedly, this invention reveals that xCas9-NG exhibits more relaxed PAM restrictions than both xCas9 and Cas9-NG in terms of base editing. Specifically, the non-PI portion of the Cas9 protein restricts the PAM recognition specificity of the PI domain, while the non-PI fragment of xCas9 expands the PAM recognition of the Cas9-NG PI domain. Furthermore, xCas9-NG significantly improves the editing efficiency of targets containing NAN PAM.
[0096] Furthermore, this invention also finds that replacing the non-PI portion of xCas9 can relax the PAM specificity of the PI domain in other SpCas9 variants. That is, the above findings also apply to other xCas9-modified SpCas9 variants, including SpRY and non-GSpCas9 (e.g., SpCas9-NRRH, SpCas9-NRCH, and SpCas9-NRTH). In other words, a series of experiments in the embodiments of this invention demonstrate that the non-PI domain fragment plays a role in the PAM constraint of the Cas9 system.
[0097] In summary, this invention couples non-PI fragments of xCas9 with PI domains from different SpCas9s, finding that the resulting xCas9-modified SpCas9 variants significantly relax the PAM restrictions for each individual PI domain. In addition to broadening the PAM restrictions, xCas9 modification also generally enhances editing activity on the original PAM recognized by each PI domain. Based on this, this invention also provides a novel editing tool with expanded targeting range.
[0098] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
[0099] References
[0100] 1. Jinek, M., et al., A programmable dual-RNA-guided DNA endonuclease inadaptive bacterial immunity. Science, 2012.337(6096): p.816-21.
[0101] 2.Doudna,J.A.and E.Charpentier,Genome editing.The new frontier ofgenome engineering with CRISPR-Cas9.Science,2014.346(6213):p.1258096.
[0102] 3.Hille,F.,et al.,The Biology of CRISPR-Cas:Backward andForward.Cell,2018.172(6):p.1239-1259.
[0103] 4.Hwang,W.Y.,et al.,Efficient genome editing in zebrafish using aCRISPR-Cas system.Nat Biotechnol,2013.31(3):p.227-9.
[0104] 5.Sander,J.D.and J.K.Joung,CRISPR-Cas systems for editing,regulatingand targeting genomes.Nat Biotechnol,2014.32(4):p.347-55.
[0105] 6.Cho,S.W.,et al.,Targeted genome engineering in human cells with theCas9 RNA-guided endonuclease.Nat Biotechnol,2013.31(3):p.230-2.
[0106] 7.Lim,K.R.Q.,C.Yoon,and T.Yokota,Applications of CRISPR / Cas9 for theTreatment of Duchenne Muscular Dystrophy.J Pers Med,2018.8(4).
[0107] 8.Jiang,W.,et al.,RNA-guided editing of bacterial genomes usingCRISPR-Cas systems.Nat Biotechnol,2013.31(3):p.233-9.
[0108] 9.Komor,A.C.,A.H.Badran,and D.R.Liu,CRISPR-Based Technologies for theManipulation of Eukaryotic Genomes.Cell,2017.168(1-2):p.20-36.
[0109] 10.Jiang,F.and J.A.Doudna,The structural biology of CRISPR-Cassystems.Curr Opin Struct Biol,2015.30:p.100-111.
[0110] 11.Kleinstiver,B.P.,et al.,Engineered CRISPR-Cas9 nucleases withaltered PAM specificities.Nature,2015.523(7561):p.481-5.
[0111] 12.Miller,S.M.,et al.,Continuous evolution of SpCas9 variantscompatible with non-G PAMs.Nat Biotechnol,2020.38(4):p.471-481.
[0112] 13.Nishimasu,H.,et al.,Engineered CRISPR-Cas9 nuclease with expandedtargeting space.Science,2018.361(6408):p.1259-1262.
[0113] 14.Walton,R.T.,et al.,Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9variants.Science,2020.368(6488):p.290-296.
[0114] 15.Hu,J.H.,et al.,Evolved Cas9 variants with broad PAM compatibilityand high DNA specificity.Nature,2018.556(7699):p.57-63.
Claims
1. A fusion protein, characterized in that, The fusion protein comprises: a) a non-PI domain from xCas9; b) a PI domain from a SpCas9 variant; the amino acid sequence of the fusion protein comprises one or more of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:
4.
2. The fusion protein as described in claim 1, characterized in that, The fusion protein further includes a heterologous functional domain, which may optionally be a deaminase domain.
3. The fusion protein as described in claim 2, characterized in that, The deaminase domain includes adenine deaminase or cytosine deaminase.
4. The use of the fusion protein according to any one of claims 1-3 in gene editing, characterized in that, The uses include delivering the fusion protein or the nucleotide sequence encoding the fusion protein in vivo or in vitro to alter the gene sequence of target cells, and the uses are for non-disease therapeutic purposes.
Citation Information
Patent Citations
Recombinant crispr-cas9 nucleases with altered PAM specificity
US20210301269A1