Crispr-frcas9 protein mutant and use thereof
By analyzing the FrCas9 protein structure and performing specific amino acid mutations, a CRISPR-FrCas9 protein mutant was developed, which solved the problems of CRISPR/Cas9 gene editing tools in targeting range and off-target effects, and achieved efficient and safe gene editing effects.
Patent Information
- Application Number
- PCT/CN2024/087116
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-07
- Filing Date
- 2024-04-11
- Publication Date
- 2025-09-11
AI Technical Summary
Existing CRISPR/Cas9 gene editing tools have problems with limited targeting range and off-target effects, making them difficult to use effectively in clinical treatment. Existing modification methods often sacrifice editing efficiency while improving targeting specificity.
By analyzing the structure of the FrCas9 protein and using cryo-electron microscopy to analyze the three-dimensional structure of the FrCas9-sgRNA-DNA complex, key amino acid sites were identified and targeted mutations were performed to develop CRISPR-FrCas9 protein mutants, especially the V1103K and N732A mutations, which formed a more compact sgRNA-DNA heteroduplex structure and improved editing efficiency and specificity.
The targeting specificity of the CRISPR/Cas9 system was significantly improved without reducing the editing efficiency, and the Cas9 nuclease with the highest editing efficiency and the highest fidelity was obtained, which is suitable for clinical treatment.
Smart Images

Figure CN2024087116_12092025_PF_FP_ABST
Abstract
Description
A CRISPR-FrCas9 protein mutant and its application Technical Field
[0001] The present invention relates to the field of genetic engineering technology, and in particular to a CRISPR-FrCas9 protein mutant and its application. Background Art
[0002] As a new programmable gene editing method, CRISPR / Cas9 technology combines the advantages of high editing efficiency and simple operation. It can perform genome editing in an efficient and precise manner, showing great advantages in clinical disease research. However, CRISPR / Cas9 gene editing tools still have some problems. First, they are limited by the PAM recognition sequence, which restricts the targetable range on the genome and prevents precise editing of some disease sites. Second, existing CRISPR / Cas9 gene editing systems have significant off-target effects, which pose risks in the gene therapy process and also limit their application in clinical treatment.
[0003] Currently, the editing efficiency and targeting specificity of the CRISPR / Cas9 system can be improved through rational protein structure-guided directed evolution. Several high-fidelity versions of SpCas9 have been obtained through this approach. For example, eSpCas9 enhances the specificity of Cas protein targeting DNA by neutralizing positively charged amino acid residues; SpCas9-HF1 minimizes the off-target effects of SpCas9 by reducing nonspecific interactions between the SpCas9 protein and the target DNA site; HypaCas9 reduces off-target effects by mutating a conserved residue cluster in the REC3 domain that senses RNA-DNA interactions to alanine to trigger the conformation of the HNH nuclease domain; and SuperFi-Cas9 reduces the rate of Cas9 protein cleavage of mismatched DNA by 500-fold by mutating a loop in RuvC. However, although these high-fidelity SpCas9 variants can improve the accuracy of genomic targeting, the enhanced specificity of these variants is achieved at the expense of target cleavage efficiency, indicating that current modifications require a trade-off between DNA targeting specificity and editing efficiency. Therefore, while continuously modifying SpCas9, there is an urgent need to develop a new Cas9 gene editing tool with both high specificity and high editing activity.
[0004] Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a CRISPR-FrCas9 protein mutant and its application.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] In a first aspect, the present invention provides a CRISPR-FrCas9 protein mutant, wherein the CRISPR-FrCas9 protein mutant has at least one mutation of V1103K, V1103R, D660A, N532A, Y763A, W692A, T615A, K545A, Y550A, N694A, N732A, K986A, N725A, and Q728A based on the FrCas9 protein.
[0008] The present invention analyzes the mechanism of specific recognition of natural FrCas9 and target DNA and identifies the key amino acid sites in the FrCas9 protein involved in heteroduplex binding. Starting from the perspective of protein structure, the present invention uses cryo-electron microscopy technology to analyze the three-dimensional structure of the FrCas9-sgRNA-DNA complex to analyze why natural FrCas9 has high cutting efficiency and low off-target effect: FrCas9 is able to accommodate a longer spacer in almost the same length as SpCas9, that is, FrCas9 forms a more compact sgRNA-DNA heteroduplex structure, requiring more stringent heteroduplex pairing, thereby reducing mismatches, which may be the main reason for the strong specificity of natural FrCas9. The above-selected amino acid sites with interactions were mutated and efficiency tested in turn to obtain a new type 2 FrCas9 variant with significantly increased editing efficiency and high specificity.
[0009] As a preferred embodiment of the CRISPR-FrCas9 protein mutant of the present invention, the CRISPR-FrCas9 protein mutant undergoes V1103K and / or N732A mutations based on the FrCas9 protein.
[0010] The present invention uses structure-guided protein information to mutate amino acid residue sites that interact with the sgRNA-DNA heteroduplex, ultimately identifying two amino acid sites, V1103 located in PLL and N732 located at the distal end of the PAM, and performing targeted mutations. V1103 is mutated to a K (lysine) with a longer side chain to increase the binding force of the phosphate lock ring, and N732 is mutated to an A (alanine) with a larger and flatter aromatic ring to increase the interaction with the TS chain. Both sites can now increase the editing efficiency by about 2 times while maintaining high specificity, thus obtaining the eFrCas9 variant: FrCas9-V1103K-N732A. Whether compared with all currently reported SpCas9 variants or natural FrCas9, the FrCas9 variant obtained by the present invention is the Cas9 nuclease with the highest editing efficiency, while also retaining the superior high fidelity of natural FrCas9.
[0011] As a preferred embodiment of the CRISPR-FrCas9 protein mutant of the present invention, the CRISPR-FrCas9 protein mutant undergoes V1103K and N732A mutations on the basis of the FrCas9 protein, and its amino acid sequence is shown in SEQ ID NO.6.
[0012] The present invention mutated amino acid sites in PLL and discovered that V1103 is a key amino acid site involved in FrCas9-specific editing. Purposeful FrCas9-V1103K modification can improve the protein's targeting specificity. Similarly, the present invention also analyzed the key amino acid sites in FrCas9 involved in sgRNA-DNA heteroduplex recognition and cleavage. By combining V1103K with combined mutations, and through functional experiments and screening, an eFrCas9 variant was obtained that increased editing efficiency by approximately 2-fold while maintaining high fidelity.
[0013] In a second aspect, the present invention provides a nucleic acid molecule comprising a nucleic acid sequence encoding the CRISPR-FrCas9 protein mutant.
[0014] As a preferred embodiment of the nucleic acid molecule of the present invention, its nucleotide sequence is shown as SEQ ID NO.5.
[0015] In a third aspect, the present invention provides an expression vector comprising the nucleic acid molecule.
[0016] In a fourth aspect, the present invention provides a cell comprising the expression vector.
[0017] In a fifth aspect, the present invention provides a base editor comprising the CRISPR-FrCas9 protein mutant.
[0018] As a preferred embodiment of the base editor described in the present invention, it also includes at least one gRNA or sgRNA that can bind to the CRISPR-FrCas9 protein mutant.
[0019] In a sixth aspect, the present invention provides a base editing method, which performs base editing using the base editor.
[0020] Preferably, the CRISPR-FrCas9 protein mutant and guide RNA are used to form a CRISPR-Cas9 system for gene editing, or,
[0021] The CRISPR-Cas9 system is composed of the CRISPR-FrCas9 protein mutant and the guide RNA formed by processing the tandem guide RNA to perform simultaneous gene editing of multiple gene sites.
[0022] In a seventh aspect, the present invention provides a CRISPR complex comprising the CRISPR-FrCas9 protein mutant, gRNA or sgRNA, and a target sequence bound to the gRNA.
[0023] As a preferred embodiment of the CRISPR complex of the present invention, the gRNA or sgRNA includes a direct repeat sequence capable of binding to the CRISPR-FrCas9 protein mutant and a guide sequence capable of targeting the target sequence.
[0024] As a preferred embodiment of the CRISPR complex of the present invention, the nucleotide sequence of the sgRNA is shown in SEQ ID NO.2.
[0025] In an eighth aspect, the present invention provides a kit for gene editing, gene targeting or gene cutting, comprising the CRISPR-FrCas9 protein mutant, or the nucleic acid molecule, or the expression vector, or the cell, or the base editor, or the CRISPR complex.
[0026] In a ninth aspect, the present invention uses the CRISPR-FrCas9 protein mutant, the nucleic acid molecule, the expression vector, the cell, the base editor, the CRISPR complex, and the kit in gene editing.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] (1) The present invention does not modify the Cas9 nuclease by the result-oriented random mutation method or phage selection evolution method used in the prior art. Instead, it independently analyzes the FrCas9 protein structure and uses a clear and definite structure as a guide to perform protein modification. It analyzes the FrCas9 recognition specific PAM and unique PLL recognition mechanism, and clarifies the difference between FrCas9 and SpCas9 in R-loop extension and domain conformation change. FrCas9 can accommodate a longer spacer in almost the same length as SpCas9, forming a more compact sgRNA-DNA heteroduplex structure, explaining the main reason why FrCas9-WT has high specificity. The protein modification theory combining this structure and functional mechanism adopted by the present invention further clarifies the recognition and cutting mechanism of FrCas9. The editing efficiency of natural FrCas9 is comparable to that of SpCas9, and its specificity is significantly higher than SpCas9. Previous modifications of SpCas9 have improved its fidelity at the expense of editing efficiency. The present invention obtains a FrCas9 variant with better performance through protein modification. eFrCas9 is the Cas9 nuclease with the highest cutting efficiency and strongest fidelity known to date.
[0029] (2) The FrCas9 described in the present invention differs significantly from SpCas9 in PAM recognition, PLL recognition, and R-loop extension. Furthermore, by combining mutations in PLL and PAM distal amino acid sites, highly efficient and specific eFrCas9 variants were screened, indicating that PLL and PAM distal amino acid sites have a synergistic effect. This invention utilizes the conformational mechanism of FrCas9 to design novel eFrCas9 variant modification methods, which may provide a reference for modifying other Cas9 proteins. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is the structure of the FrCas9-sgRNA-DNA ternary complex analyzed by cryo-electron microscopy in the present invention.
[0031] Figure 2 is a structural analysis of the FrCas9-specific PAM recognition and a diagram showing the editing efficiency after amino acid site mutation according to the present invention.
[0032] Figure 3 is a structural analysis of FrCas9 PLL recognition and a diagram showing editing efficiency detection after amino acid site mutation according to the present invention.
[0033] Figure 4 is a diagram showing the editing efficiency after mutation of the amino acid site involved in sgRNA-DNA heteroduplex recognition by FrCas9 according to the present invention.
[0034] Figure 5 is a diagram showing the editing efficiency detection after multi-site combined mutation of FrCas9 according to the present invention. DETAILED DESCRIPTION
[0035] To better illustrate the purpose, technical solutions and advantages of the present invention, the present invention will be further described below in conjunction with specific embodiments. Those skilled in the art should understand that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0036] Unless otherwise specified, the experimental methods used in the examples are conventional methods; the materials, reagents, etc. used are all available from commercial sources unless otherwise specified.
[0037] FrCas9 from Faecalibaculum rodentium (its amino acid sequence can be found in Chinese invention patent document 202110331485.2) recognizes a unique PAM with the sequence 5'-NRTA (R=A / G), which increases the applicable scenarios of existing gene editing tools. The cutting activity of FrCas9 is comparable to or even slightly higher than that of SpCas9. More importantly, the off-target effect of FrCas9 is significantly lower than that of SpCas9. This is also the only new type 2 CRISPR nuclease with a cutting efficiency comparable to SpCas9 but a significantly lower off-target effect than SpCas9. Comparing the editing efficiency and specificity of FrCas9 and SpCas9 at 11 endogenous human sites, it was found that FrCas9 has higher specificity than SpCas9 without reducing the editing efficiency. Although FrCas9 has a lower off-target effect than SpCas9, its mechanism of action is still unclear.
[0038] This paper analyzes the three-dimensional crystal structure of FrCas9 based on high-resolution cryo-electron microscopy (Cryo-EM), analyzes the structural mechanism of FrCas9 protein binding to sgRNA, PAM recognition and target DNA-dependent conformational rearrangement, and optimizes natural FrCas9 through rational protein modification. The purpose of this invention is to analyze the structure of the FrCas9-sgRNA-dsDNA ternary complex, use the structure as a guide, and through rational design and modification, screen for FrCas9 mutants with higher editing efficiency and stronger fidelity, providing a new, efficient and safe option for the application of CRISPR / Cas9 in clinical treatment.
[0039] Example 1: Preparation of FrCas9-sgRNA-dsDNA ternary complex
[0040] The precise cutting of target DNA by the FrCas9-sgRNA complex is a highly collaborative process, including PAM recognition, R-loop extension, target DNA cutting and final product release.
[0041] In order to understand the recognition mechanism of FrCas9 in detail, the catalytically inactive FrCas9 (H877A) was used to reconstruct the three-dimensional structure of two different target DNA complexes by cryo-electron microscopy: Cryo-electron microscopy analysis of 43 bp paired dsDNA substrates bound to sgRNA was performed at a resolution of The 26nt target single-stranded DNA (TS chain) sequence and the 4nt PAM single-stranded DNA (NTS chain) were analyzed by cryo-electron microscopy at a resolution of 100 nm.
[0042] (1) Expression and purification of FrCas9 protein
[0043] The prokaryotic optimized nucleotide sequence corresponding to the catalytically inactive FrCas9-H877A protein (sequence shown in SEQ ID NO.1) was cloned into the psumo protein expression vector by homologous recombination. The correctly sequenced recombinant plasmid was transformed into E. coli Rosseta 2 (full gold, CD811-02) competent cells. After activation, Kana resistance (50 μg / mL) culture dishes were coated. The next day, single colonies were picked and cultured in culture medium until the OD600 of the bacterial solution was between 0.8 and 1.0. IPTG was added to the bacterial solution to a final concentration of 0.1 mM, and the induction culture was continued for 14-16 hours. The bacterial solution was collected by centrifugation. Resuspension buffer (40 mM Tris-Cl, 500 mM NaCl, 20 mM imidazole, 5% glycerol, pH 8.0) was added to resuspend the precipitate, followed by high-pressure sterilization. After Ni column affinity chromatography and molecular sieve purification, the target protein was eluted with 250 mM imidazole and identified by SDS-PAGE electrophoresis. The final FrCas9 protein was concentrated to 30.5 mg / mL, quickly frozen in liquid nitrogen, and stored at -80 degrees.
[0044] (2) sgRNA preparation
[0045] The sgRNA (sequence shown in SEQ ID NO. 2) was transcribed from a dsDNA template using T7 RNA polymerase. The transcription reaction was incubated at 37°C for 4 h, purified by 8% denaturing TBE-urea PAGE, and dissolved in enzyme-free water. The 26nt TS ssDNA (sequence shown in SEQ ID NO. 3) and 4nt NTS ssDNA (TGTA) were synthesized by Genwi Biotech.
[0046] (3) Synthesis of 43 bp dsDNA complex
[0047] The purified sgRNA was pre-annealed at 94°C for 5 minutes and then slowly cooled to room temperature. FrCas9 protein and sgRNA were incubated at 25°C for 20 minutes at a ratio of 1.75:1. The dsDNA (sequence shown in SEQ ID NO.4) was then heated to 95°C for 5 minutes and cooled at room temperature. The ternary complex was composed of FrCas9, sgRNA, and dsDNA mixed in a ratio of 1.75:1:1.2 and incubated at 25°C for 20 minutes. Finally, the complex was purified using a Superdex200 10 / 300 GL column (GE Healthcare) in buffer (20mM Tris HCl, 150mM NaCl, 5mM DTT, 1% glycerol, pH 8.0).
[0048] (4) Combination of 26nt TS-4nt NTS complex
[0049] The purified sgRNA was added to EDTA at a final concentration of 1 mM and pre-annealed at 94°C for 5 minutes, then slowly cooled to room temperature. FrCas9 protein and sgRNA were incubated at 25°C for 20 minutes at a ratio of 1.75:1. 26nt TS ssDNA and 4nt NTS ssDNA were mixed at a ratio of 1:1 and hydrated in buffer (20mM Tris-Cl, 150mM NaCl, 5mM DTT, pH 8.0). The dsDNA was then heated to 95°C for 5 minutes and cooled at room temperature. The ternary complex consisted of Cas9, sgRNA, and ssDNA complex mixed at a ratio of 1:1.5:2.3 and incubated at 25°C for 20 minutes. Finally, the complex was purified using a Superdex200 10 / 300 GL column (GE Healthcare). Peak fractions were collected every 50 μL. MgCl2 was then added to the fractions at a final concentration of 10 mM and incubated on ice for 3 hours. Prior to preparing frozen samples, the ternary complex was thawed at room temperature for 20 min.
[0050] Example 2: Analyzing the structure of the FrCas9-sgRNA-DNA ternary complex
[0051] The FrCas9-sgRNA-DNA complex with uniform and stable properties obtained in Example 1 was taken, and the structure of the protein complex with high resolution was obtained using single-particle cryo-electron microscopy technology.
[0052] (1) Preparation of cryo-EM samples
[0053] a) Sample preparation was performed using a Vitrobot Mark IV. Liquid nitrogen was first used to pre-cool the metal reactor. Ethane gas was then slowly introduced to cool the ethane to a solid-liquid mixture. The Vitrobot was then turned on and parameters were set.
[0054] B) Before preparing negative staining and cryo-electron microscopy samples, the copper mesh carrying the sample is first hydrophilized, and a mixed gas of different components is used as a medium for brilliant discharge so that the surface of the mesh carries an electric charge. The mesh after hydrophilization is now ready for use. The mesh is made to have a carbon film side facing up, and an appropriate amount of 4 μL FrCas9-sgRNA-DNA complex sample prepared in Example 1 is drawn and added dropwise to the mesh so that it naturally diffuses and covers the entire front of the mesh. After absorbing it for 15 seconds, 5 μL heavy metal salts (uranyl acetate) are drawn and added dropwise to the mesh so that it also naturally diffuses. The copper mesh is transferred to a sample loader and fixed, and added to liquid nitrogen-cooled ethane, that is, the steps of adding samples, adsorption, and quick freezing are completed in sequence.
[0055] (2) Cryo-electron microscopy data collection and processing
[0056] The data of the FrCas9-sgRNA-DNA ternary complex were collected using a FEI Titan Krios cryo-electron microscope equipped with a Gatan K3 Summit direct electron detector at 300kV. After optimizing the protein complex and cryo-EM samples, suitable and reproducible conditions were screened to prepare cryo-EM samples with stable and uniform particles. The sample loader was pushed into the sample stage, and an area with a better ice layer was found to take a picture to check whether the sample was in good condition. The appropriate area was screened and the data of Cas9-sgRNA-dsDNA was collected using a FEI Titan Krios cryo-electron microscope equipped with a Gatan K3 Summit direct electron detector at 300kV and ×105000 magnification (0.83 pixel size). Finally, in super-resolution mode, the images were automatically recorded using EPU, and the pixel size of the 26nt TS and 4nt NTS data sets (2956 micrographs) was 43bp dsDNA dataset (8670 micrographs).
[0057] (3) Model establishment and refinement
[0058] The collected raw images were imported into RELION software for offset correction, and particles from the better-quality images were selected using the automatic particle selection function. Particle extraction was then performed, and the particles were classified using the 2D classification function. The selected particles underwent multiple rounds of 3D reconstruction and optimization using the 3D initial model function. The initial model was derived from the high-resolution structures of PDB models 4OO8 and 6O0X. Manual FrCas9 domain and nucleic acid placement and adjustment were completed using Chimera 59, COOT 60, and phenix 61. The final model was validated using statistics of MolProbity scores, EMRinger scores, and Romachandran plots. The local resolution of the ternary complex was assessed using ResMap 62.
[0059] Structural analysis: The distribution of each domain of FrCas9 is shown in Figure 1a, respectively. Cryo-electron microscopy analysis of catalytically inactive FrCas9 (H877A) bound to sgRNA with a 43 bp dsDNA substrate was performed (a, b, d in Figure 1); Cryo-electron microscopy analysis of 22nt ssDNA and 4ntPAM ssDNA complexes bound to sgRNA was performed at a resolution of 100 nm (c, e in Figure 1). For structures with 26nt TS and 4nt NTS, most of the RNP complex can be clearly resolved except for the HNH domain and PAM. The electron density (EM density) of the HNH domain is poor, and the PAM is not visible in the electron density map. In addition, the HNH domain is not visible in the electron density map of the 43bp dsDNA structure, which indicates the flexibility of the HNH domain. Due to molecular motion, the gRNA-TS distal to the PAM is missing in the electron density, and because the electron density of the adjacent REC2 and REC3 domains is also poor, the structure of the FrCas9-sgRNA-DNA ternary complex can be resolved by combining the two forms of structure.
[0060] Example 3: Analyzing the mechanism of FrCas9-specific cleavage of target DNA
[0061] Based on the high-resolution cryo-electron microscopy structure of the FrCas9-sgRNA-DNA ternary complex, the mechanisms of FrCas9 recognition of PAM, binding of sgRNA to target DNA, and DNA-dependent conformational rearrangement were elucidated. To clarify the importance of some key amino acids specifically recognized in the complex, important amino acid sites were mutated using the three-dimensional structural information as a guide. The unmutated natural FrCas9 was designated FrCas9-WT. The editing efficiency of the mutants was tested by amplicon high-throughput sequencing, or on-target and off-target effects in eukaryotic cells were detected using GUIDE-seq by inserting dsODN. The specific steps are as follows:
[0062] The FrCas9 protein was codon-optimized for human origin (its amino acid sequence is shown in SEQ ID NO. 7), and the corresponding mutant FrCas9 nucleotide sequence was cloned into the PX330 eukaryotic expression vector (Addgene, 59909) to construct the PX330-FrCas9 protein eukaryotic expression plasmid.
[0063] Construct a site plasmid. Taking mammalian HEK293T cells as an example, select the endogenous gene site and construct a sequence format of 5'-spacer sequence complementary to the target site-scaffold-3' to construct the PXZ-target target plasmid.
[0064] Table 1 Target names and sequences
[0065] For amplicon detection: PX330-FrCas9 protein plasmid and PXZ-target plasmid were lipofected into 24-well plates of well-grown HEK293T cells. After 72 hours, the cells were harvested and genomic DNA was extracted. A pair of amplicon primers were designed upstream and downstream of the selected target site for high-throughput library construction. 150bp paired-end sequencing data was generated using the Illumina HiSeq 2500 platform. The indels (insertions and deletions) in each read of the considered region were calculated by extracting sequences of the ±10bp flanking region of the cleavage site using the CRISPRMatch software package. The editing efficiency of the mutant was compared with that of the FrCas9-WT.
[0066] For GUIDE-seq detection: PX330-FrCas9 protein plasmid, PXZ-target plasmid, and 1.2uL dsODN were electroporated into a 24-well plate of HEK293T cells in good growth condition. After 72 hours, the cells were harvested to extract genomic DNA, and GUIDE-seq library was constructed. Next-generation sequencing was performed on the machine, and the open source GUIDE-seq software was used to analyze the data with mismatch ≤ 7 to detect on-target and off-target situations.
[0067] (1) Mechanism of FrCas9 recognition of specific PAM 5'-NRTA-3'
[0068] The applicant's previous work showed that FrCas9 uses the TA-rich palindromic sequence PAM region to edit the target gene. In order to clarify the recognition mode, the PAM duplex with PAM 5'-TGTA-3' was used to analyze the mechanism of FrCas9 recognition of specific PAM. The structure of the FrCas9-sgRNA-DNA complex shows that the PAM duplex is nested in the positive charge groove of the PI domain, and the first dT1* of the PAM forms a hydrogen bond with the side chain of Y1198. At the same time, the constraints of Y1198 and R1213 on the phosphate backbone jointly stabilize the recognition of the PAM by the PI domain (a, b in Figure 2). dG2* of 5'-NRTA-3' PAM is recognized by the side chain of N1137, dT3* is recognized by the amino group of K1321 through a water-mediated hydrogen bond, N6 of dA4* bonds to the hydroxyl group of T1214, and dT4* forms a non-base-specific π-π stacking interaction with Q1334 (a, b in Figure 2).
[0069] To clarify the importance of amino acid residues involved in PAM-specific recognition in FrCas9, the amino acid residue sites involved in PAM recognition were mutated. Amplicon sequencing and GUIDE-seq high-throughput sequencing showed that all mutations (T1198R, N1137A, T1214W, T1214F, T1214Y, Q1334R, Q1334T, K1321A and R1213A) significantly reduced the targeted cutting efficiency of FrCas9, demonstrating the importance of these amino acid sites in PAM recognition (ce in Figure 2).
[0070] (2) Analyzing the unique recognition mechanism of the FrCas9 phosphate lock ring
[0071] From the three-dimensional structure of SpCas9 reported in many previous literatures and the structural analysis of FrCas9 in this example, it can be seen that: although the amino acid sequence of FrCas9 is only 41.39% similar to that of SpCas9, their 3D structures are surprisingly similar. FrCas9 adopts a typical bilobed structure with an α-helical recognition (REC) lobe and a nuclease (NUC) lobe. The sgRNA-target DNA heteroduplex is bound in the central channel, just like other Cas9 homologs (a in Figure 3). The phosphate locking loop (PLL) in the Cas9 protein is responsible for PAM-dependent target DNA binding and the formation of RNA-DNA hybrid chains. It is an important element for regulating the off-target sensitivity and editing efficiency of the Cas9 protein.
[0072] The structure of the resolved FrCas9-sgRNA-DNA ternary complex shows that in the contact between the PLL of FrCas9 and the +1 phosphate of the TS chain, the amino acid residues D1101 and S1102 in the PLL play a major role in the interaction, and the amino acid residue V1103 has less contact with the +1 phosphate, which is different from the contact mechanism of the amino acid residue 1107-KES-1109PLL element in SpCas9. It is worth noting that there is a certain electron density cavity at the V1103 amino acid residue of FrCas9 (b in Figure 3), which provides more possibilities for modification of the V1103 residue site on the PLL element. In order to determine the importance of the V1103 amino acid site, the present invention generated 19 other amino acid substituted mutants, and used amplicons and detection mutant editing efficiencies, and detected on-target and off-target situations in eukaryotic cells using GUIDE-seq by insertion of dsODN.
[0073] The amplicon detection results are shown in c in Figure 3: mutations to proline and aromatic amino acids significantly reduced the editing efficiency of FrCas9, while other linear side chain amino acids, except aspartic acid and glutamic acid, had almost no effect on the editing ability of FrCas9. Next, to explore whether the length change of the amino acid side chain at site 1103 would affect the specificity of FrCas9, the present invention selected two positively charged amino acid mutants from the long side chain amino acid mutants: lysine and arginine, namely FrCas9-V1103K and FrCas9-V1103R, and performed on-target and off-target detection of the specificity of these two mutants by GUIDE-seq. The results showed that both FrCas9-V1103K and FrCas9-V1103R increased the on-target readings of the three target sites, except for HEK293SITE Except for the x2 site, no off-target sites were generated. Given that FrCas9-V1103K significantly improved the editing efficiency of FrCas9 compared to FrCas9-V1103R (d in Figure 3), this example continued to test the specificity of FrCas9-V1103K at six other genomic sites and found that the overall fidelity of FrCas9-V1103K at nine target sites increased by 32.25% (eh in Figure 3). These results indicate that structure-guided PLL modification can improve the specificity of FrCas9.
[0074] (3) Analysis of the recognition mechanism of sgRNA-DNA heteroduplexes
[0075] Structural analysis shows that FrCas9 and SpCas9 have different R-loop extension mechanisms. FrCas9 can accommodate a longer spacer in a length almost the same as SpCas9 (SpCas9 spacer length is 20nt, FrCas9 spacer length is 22nt), that is, FrCas9 forms a more compact sgRNA-DNA heteroduplex structure, requiring more stringent heteroduplex pairing, thereby reducing mismatches. This may be the main reason for the strong specificity of natural FrCas9 (a in Figure 4). The three-dimensional structure analysis of the FrCas9-sgRNA-DNA complex provides structural guidance for rationally improving the specificity of FrCas9 and sgRNA-DNA heteroduplexes. In this example, the structures involved in sgRNA-DNA heteroduplex recognition at the PAM proximal and PAM distal ends in the FrCas9-sgRNA-DNA complex were analyzed. The sgRNA-DNA heteroduplex in the PAM-proximal seed region is primarily recognized by the conserved bridge helix and REC1 domains. The amino acid sites involved in heteroduplex interactions are D660, W692, N532, and Y763. Additionally, R61, R64, R68, R69, and R72 participate in recognition of the phosphate backbone of the sgRNA seed region. The PAM-distal region of the sgRNA-DNA heteroduplex is primarily recognized by the REC flap and RuvC domains. In the protein complex described in this example, the main amino acid sites distal to the PAM involved in sgRNA-DNA heteroduplex interaction are T615, K545, Y550, N694, N732, K986, N725, and Q728.
[0076] The amino acid sites involved in the sgRNA-DNA heteroduplex interaction were mutated to alanine (Ala, A), and the editing efficiency and fidelity were tested by amplicon sequencing and GUIDE-seq to evaluate the importance of these amino acid sites. The results showed:
[0077] Mutants D660A, N532A, and Y763A increased FrCas9 cleavage efficiency to varying degrees, while W692A significantly reduced cleavage of target substrates in vivo, indicating that W692 plays a key role in fixing the seed-terminal region. Compared with the arginines (R61A / R69A) that recognize the seed region, the arginines at the end of the seed region (R64A / R68A / R72A) are more important for FrCas9 cleavage activity (Figure 4b). Alanine substitutions at some amino acid sites distal to the PAM can improve target cleavage efficiency without significant off-target effects, such as T615A, K545A, and Y550A / N694A. However, N732A and K986A of FrCas9 (equivalent to the high-fidelity variants of SpCas9, HypCas9 H698A and Q695A) both reduced the number of on-target sites and increased the number of off-target reads, which is very different from the specificity mechanism of SpCas9. Consistent with SpCas9, the N725A (equivalent to HypaCas9 N692A) and Q728A (equivalent to HypCaCas9 Q695A) mutants reduced off-target effects (cf. in Figure 4).
[0078] In summary, the structural information of FrCas9 and the in vivo functional experimental results of mutants both indicate that although their structures are very similar, FrCas9 and SpCas9 have different off-target recognition mechanisms.
[0079] Example 4: eFrCas9 variants
[0080] Guided by the structure, FrCas9 was rationally designed to obtain eFrCas9 variants with improved editing efficiency and fidelity.
[0081] The structural analysis of Example 1 and Example 2 provides a framework for the formation of the FrCas9-sgRNA-DNA complex, and the structural data reveal the structural rearrangement during the extension of the R-loop in FrCas9. In Example 3, V1103K in the PLL element enhances the interaction between FrCas9 and the +1bp heteroduplex to improve the fidelity of FrCas9. In addition, Example 3 found that alanine substitution of partially recognized FrCas9 heteroduplex residues can improve target cutting efficiency (a in Figure 5). Therefore, this embodiment assumes that a rationally designed FrCas9 variant with a heteroduplex interaction cluster site may show improved efficiency and fidelity in human cells. In order to detect whether V1103K and PAM distal amino acid site mutations have a synergistic effect, this embodiment uses a double-site or triple-site joint mutation to analyze the editing efficiency and fidelity of multiple FrCas9 combined variants.
[0082] The specific operations are as follows:
[0083] Using the PX330-FrCas9 protein eukaryotic expression plasmid from Example 3 as a template, primers were designed at the target mutation amino acid sites, and FrCas9 protein plasmids with dual or triple mutations were constructed using the Gibson method. The PXZ-target plasmid and amplicon assays for editing efficiency, as well as GUIDE-seq fidelity testing using dsODN insertion, were performed as described in Example 3.
[0084] The results showed that both amplicon and GUIDE-seq showed that compared with FrCas9-WT, the dual-site variant FrCas9-V1103K-N732A (eFrCas9 variant, nucleotide sequence as shown in SEQ ID NO. 5, protein sequence as shown in SEQ ID NO. 6) increased the editing efficiency by approximately 2-fold while reducing off-target effects, showing high specificity and high editing activity (bd in Figure 5).
[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A CRISPR-FrCas9 protein mutant, characterized in that The CRISPR-FrCas9 protein mutant has at least one mutation of V1103K, V1103R, D660A, N532A, Y763A, W692A, T615A, K545A, Y550A, N694A, N732A, K986A, N725A, and Q728A based on the FrCas9 protein.
2. The CRISPR-FrCas9 protein mutant according to claim 1, characterized in that The CRISPR-FrCas9 protein mutant undergoes V1103K and / or N732A mutations based on the FrCas9 protein.
3. The CRISPR-FrCas9 protein mutant according to claim 2, characterized in that The CRISPR-FrCas9 protein mutant undergoes V1103K and N732A mutations based on the FrCas9 protein, and its amino acid sequence is shown in SEQ ID NO.
6.
4. A nucleic acid molecule, characterized in that The nucleic acid molecule contains a nucleic acid sequence encoding the CRISPR-FrCas9 protein mutant according to claim 1.
5. The nucleic acid molecule according to claim 4, characterized in that Its nucleotide sequence is shown in SEQ ID NO.
5.
6. An expression vector, characterized in that Comprising the nucleic acid molecule according to claim 4 or 5.
7. A cell, characterized in that Comprising the expression vector according to claim 6.
8. A base editor, characterized in that Comprising the CRISPR-FrCas9 protein mutant according to any one of claims 1 to 3.
9. The base editor according to claim 8, characterized in that It also includes at least one gRNA or sgRNA that can bind to the CRISPR-FrCas9 protein mutant.
10. A base editing method, characterized in that: Base editing is performed using the base editor described in claim 8 or 9.
11. A CRISPR complex, characterized in that It includes the CRISPR-FrCas9 protein mutant described in any one of 1-3, gRNA or sgRNA, and a target sequence bound to the gRNA or sgRNA.
12. The CRISPR complex according to claim 11, wherein The gRNA or sgRNA includes a direct repeat sequence capable of binding to the CRISPR-FrCas9 protein mutant and a guide sequence capable of targeting the target sequence.
13. The CRISPR complex according to claim 12, wherein The nucleotide sequence of the sgRNA is shown in SEQ ID NO.
2.
14. A kit for gene editing, gene targeting or gene cleavage, characterized in that: Comprising the CRISPR-FrCas9 protein mutant of any one of claims 1 to 3, or the nucleic acid molecule of claim 4 or 5, or the expression vector of claim 6, or the cell of claim 7, or the base editor of claim 8 or 9, or the CRISPR complex of any one of claims 11 to 13.
15. Use of the CRISPR-FrCas9 protein mutant according to any one of claims 1 to 3, the nucleic acid molecule according to claim 4 or 5, the expression vector according to claim 6, the cell according to claim 7, the base editor according to claim 8 or 9, the CRISPR complex according to any one of claims 11 to 13, and the kit according to claim 14 in gene editing.
Citation Information
Patent Citations
S. pyogenes cas9 mutant genes and polypeptides encoded by same
CN110462034A
Construction of chimeric SaCas9 based on evolutionary information for enhancing and extending recognition of PAM loci
CN110684755A
Construction method of homologous type 2 CRISPR / Cas gene editing system
CN112331264A
Temperature-sensitive RNA-guided endonuclease
CN113939588A
Type 2 CRISPR / Cas9 gene editing system and application thereof
CN114075559A