Modularized progenitor reconstructed wide-spectrum PAM Cas9 variant and preparation method thereof
Through modular ancestral sequence reconstruction technology, the PAM recognition domain of Cas9 was reconstructed, and the Cas9 variant RHCas9 with wide spectrum PAM was obtained, which solved the problem of limited recognition range of the existing Cas9 protein PAM, and achieved broadening the substrate recognition range and improving editing efficiency while maintaining high activity and low off-target rate.
Patent Information
- Application Number
- CN202510360283.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-05-06
AI Technical Summary
The existing Cas9 protein has a limited range of PAM recognition, which is difficult to adapt to the editing of different genomic loci, and its macromolecular size limits the application of gene therapy in vivo.
Through modular ancestral sequence reconstruction technology, only the PAM recognition domain of Cas9 was reconstructed, and the native sequences of other domains were retained, and the Cas9 variant RHCas9 of broad spectrum PAM was obtained.
While maintaining the high activity of nucleases, RHCas9 broadens the scope of substrate recognition, increases the coverage of targeted sites, and achieves optimization of specific functions, breaking through the targeting limitations of traditional CRISPR tools.
Smart Images

Figure CN119931989A_ABST
Abstract
Description
Technical Field
[0001] The present invention provides a wide-spectrum PAM Cas9 variant for modular ancestral reconstruction, and performs modular ancestral sequence reconstruction on the Cas9 protein. The present invention further discloses a method for preparing the variant, belonging to the technical field of gene editing. Technical Background
[0002] CRISPR / Cas9 technology introduces DNA double-strand breaks (DSBs) at target sites to achieve gene editing at specific sites, which has played a revolutionary role in the field of biomedicine. The Cas9 system (SpCas9) derived from Streptococcus pyogenes is the most widely used gene editing system. Its natural form only recognizes NGG-type PAM and can only edit about 6.5% of human genome sites. In addition, the large molecular size of SpCas9 (about 4.2 kb) makes it difficult to adapt to the delivery capacity of adeno-associated virus (AAV) vectors (≤4.7kb), which seriously restricts its application in in vivo gene therapy. The small-sized Cas9 variants discovered later (such as SpaCas9) are easier to deliver due to their compact size, but their complex PAM recognition requirements (such as SpaCas9 requires NNGTGA-type PAM) and the limitations of traditional modification methods have further aggravated the bottleneck of the application of gene editing tools.
[0003] Existing technologies mainly use strategies such as directed evolution screening (such as bacteria or phage-assisted continuous evolution), homologous PI domain replacement and rational structural design mutation for optimization. However, it still faces multiple technical bottlenecks: First, the PAM extension range of most variants is limited; the design of variants is highly dependent on the predicted protein-DNA complex structure information and empirical trial and error, the development cycle is long and difficult to extend to other Cas proteins; the obtained variants often have a significant reduction in cutting activity due to global perturbations, which further amplifies the defects of traditional methods.
[0004] Ancestral sequence reconstruction (ASR) technology calculates the sequences of internal nodes of the phylogenetic tree by inferring the phylogenetic relationships between modern homologous sequences and applying statistical models of amino acids. The reconstructed ancestral proteins show enhanced thermal or mechanical stability, more extreme pH ranges, increased activity, and substrate promiscuity. Summary of the invention
[0005] The purpose of the present invention is to provide a modular ancestral reconstruction of a wide-spectrum PAM Cas9 variant and a preparation method, which only reconstructs the ancestral sequence of the PAM recognition domain of Cas9 and retains the natural sequences of other domains to obtain a wide-spectrum PAM Cas9 variant, which has the characteristics of broadening the substrate recognition range while maintaining high nuclease activity, breaking through the targeting limitations of traditional CRISPR tools.
[0006] The modular ancestral reconstruction of the broad-spectrum PAM Cas9 variant described in the present invention is named RHCas9, and its amino acid sequence is shown in SEQ ID NO.1.
[0007] The method for precise gene editing of a modular ancestral reconstruction broad-spectrum PAM Cas9 variant described in the present invention comprises the following steps: 1) Homology search: The SpaCas9 protein sequence of Streptococcus pasteurianus was submitted to NCBI, and PSI-BLAST (site-specific iterative alignment) was used for homology search; 2) Homologous sequence comparison and construction of developmental tree: Multiple sequence alignment: Generate PID region multiple sequence alignment for all 35 homologous sequences using MAFFT (L-INS-i algorithm); manually delete gaps in the multiple sequence alignment to ensure alignment quality; Evolutionary tree construction: IQ-TREE was used to automatically select the best amino acid substitution model, and the branch support rate was evaluated by 1000 bootstraps to construct the maximum likelihood tree; 3) Reconstruction of ancestral sequence: The evolutionary path of PID was traced through PAML software, and the amino acid residues of each node were determined based on the maximum likelihood probability. The branch node sequence that eventually differentiated into the target Cas9 was selected as the candidate ancestral variant.
[0008] 4) Obtaining RHCas9: Recombining the reconstructed PID with the remaining domain of wild-type SpaCas9 to obtain a modular ancestral reconstructed broad-spectrum PAM Cas9 variant, named: RHCas9 (SEQ ID: NO.1); 5) Validate its PAM recognition preference through a targeted PAM library system combined with deep sequencing, and evaluate the editing efficiency in mammalian cells.
[0009] Experimental data show that the RHCas9 variant provided by the present invention expands the PAM recognition range of the wild-type SpaCas9 from NNGT to NNRH (R is purine, H is non-G base), such as Figure 2 C, the coverage of targeted sites increased by 6 times; in human cells, RHCas9 showed comparable or higher editing efficiency for wild-type non-editable PAM sites (such as NNGA, NNGC, NNAT, etc.) than wild-type preferred PAM-NNGT sites, up to 74%. Figure 3 ; RHCas9 expands the PAM while maintaining editing efficiency, achieving optimization of specific functions.
[0010] The PAM recognition (PI domain), DNA unwinding (REC domain) and cleavage activity (HNH / RuvC domain) of Cas9 described in the invention have clear region independence. Based on this, phylogenetic tracing and ancestral sequence reconstruction are performed for a single functional domain (such as the PI domain recognized by PAM), while retaining the native sequence of other functional domains (such as HNH / RuvC), without sacrificing the cleavage activity of the HNH domain or the DNA binding specificity of the REC domain.
[0011] The positive effects of the present invention are: Through a modular ancestral reconstruction, a Cas9 variant with a wide spectrum of PAM is obtained: RHCas9, which avoids the loss of catalytic activity or non-targeted functional confusion (such as ssDNA cutting) caused by global transformation; the substrate recognition range is broadened under the premise of maintaining high nuclease activity. Compared with the wild-type enzyme, RHCas9 has a 6-fold increase in substrate recognition types, and does not affect the editing efficiency of the original preferred sites. It has equivalent or higher efficiency at the newly added editing sites than the original preferred sites. It can not only avoid the functional confusion caused by global ASR, but also accurately release the evolutionary potential of the target domain, thereby breaking through the PAM limitation while maintaining high editing efficiency and low off-target rate. In addition, modular ASR does not require the prior knowledge of high-resolution structural information, and is generally suitable for the directional optimization of various Cas proteins, providing a new paradigm for the development of efficient, safe and delivery-compatible gene editing tools. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 . Flowchart of modular ancestral reconstruction of a broad spectrum of PAM Cas9 variants; Figure 2 . Flow chart and results of RHCas9 detection of PAM in HEK293T cells; Figure 3 . Result diagram of RHCas9 at 23 endogenous sites. DETAILED DESCRIPTION
[0013] The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention. The present invention is further described in detail below in conjunction with specific examples and with reference to data. It should be understood that these examples are only for illustrating the present invention, and are not intended to limit the scope of the present invention in any way. If not specifically indicated, the technical means used in the examples are conventional means well known to those skilled in the art. The materials, reagents, etc. used in the following examples, if not specifically stated, can all be obtained from commercial sources.
[0014] Example 1: Method for preparing broad-spectrum PAM Cas9 variants by modular ancestral reconstruction 1. Obtaining homologous sequences The SpaCas9 protein sequence of Streptococcus pasteurianus was submitted to NCBI, and PSI-BLAST (site-specific iterative alignment) was used for homology search to obtain the Top 500 homologous protein sequences. 35 sequences belonging to the genus Streptococcus were retained, and proteins of other genera or with large sequence differences were excluded.
[0015] Table 1. NCBI numbers of homologous sequences
[0016] 2. Homologous sequence comparison and construction of developmental tree Multiple sequence alignment: Generate high-precision PID region multiple sequence alignment for all 35 sequences using MAFFT (L-INS-i algorithm); manually delete gaps in the multiple sequence alignment to ensure alignment quality; Evolutionary tree construction: IQ-TREE was used to automatically select the best amino acid substitution model (the WAG+G4 model was finally selected), and the branch support rate was evaluated by 1000 bootstraps to construct the maximum likelihood tree; Figure 1 A; 3. SpaCas9 PID ancestral sequence reconstruction Based on the above multiple sequence comparison and evolutionary tree, PAML (CodeML module) was used to reconstruct the ancestral sequence. The amino acid residues of each node were determined based on the maximum likelihood probability, and the key node sequences were selected along the SpaCas9 differentiation path; 4. Obtaining RHCas9 The reconstructed PID was recombined with the remaining domains of wild-type SpaCas9 to obtain a modular ancestral reconstructed broad-spectrum PAM Cas9 variant, named: RHCas9 (SEQ ID: NO.1); Structural prediction: Use AlphaFold3 to predict the structure of the ternary complex of RHCas9, sgRNA, and target DNA, focus on the PID region (red mark), and find the amino acid that interacts with the PAM region in the target DNA, which is R1086. Figure 1 C; whether the PAM recognized by the new variant has changed is predicted based on whether this amino acid and its surrounding amino acids have mutated, such as Figure 1 As shown in D; The results are as follows Figure 1As shown, A is the phylogenetic tree of the sequence used to generate RHCas9; B is the domain structure of SpaCas9, showing the sequence comparison in the PAM recognition domain. C is the structure prediction of SpaCas9 using AlphaFold3, the red mark is PID, and the enlarged part is the amino acid residues in PID that interact with PAM. D is the key amino acid residues in PID that interact with PAM and their nearby sequence comparison.
[0017] Example 2: PAM determination of RHCas9 in mammalian cells
[0018] PAM detection system GenScript synthesized human codon-optimized RHCas9 (SEQ ID: NO.1) and used the previously developed PAM detection system PAM-DOSE forward screening system to identify functional PAMs in human cells. pmTmG contains two fluorescent protein genes (tdTomoto and EGFP) and two target sequences (sloxP and Spacer N8) (Figure 2A). Spacer N8 contains a spacer and 8 random nucleotides located downstream of the tdTomato gene cassette, which are used to empirically determine the PAM recognition ability of Cas proteins in human cells. After the cleavage, the tdTomato gene cassette will be excised and the CAG promoter will drive the expression of the EGFP gene. The random PAM sequence after excision can be analyzed by Sanger sequencing or high-throughput sequencing (NGS), and the frequency of each PAM type is counted based on the sequencing data.
[0019] 1.2 sgRNA vector construction sgRNA was designed according to the upstream sequence of the PAM library, and a CACC sequence linker was added to the 5' end of the upstream sequence of each sgRNA, and an AGAC sequence linker was added to the 5' end of the downstream sequence.
[0020] The synthesized sgRNA oligos are shown in Table 2: Table 2. HEK293T cell crRNA oligos
[0021] After synthesis, the upstream and downstream were annealed by a procedure (95 °C, 5 min; room temperature cooling for 30 min), and the annealed products were connected to the upstream BnS The Spa-V5 backbone vector (SEQ ID: NO.2) was digested and then transformed, single clones were picked and sequenced to obtain the correct sgRNA, and high-purity plasmids were obtained by large-scale extraction.
[0022] 1.3 Cell culture and transfection HEK293T cells were inoculated in DMEM medium with 10% FBS, containing 1% double antibody, and in an atmosphere of 5% CO 2 The cells were cultured in a 37 ℃ cell culture incubator. The cells used for transfection were inoculated in a 12-well plate the day before and cultured. The cells were observed the next day. When the cell density was about 75%, the medium was replaced with a medium containing 10% FBS and no double antibody. HieffTransTM Liposomal Transfection Reagent was used for transfection. SpaCas9, sgRNA and pmTmG reporter plasmid (1:1:1:1) were co-transfected. The total amount of each well of the 12-well plate during transfection was 1 μg. After mixing the plasmids, they were diluted and mixed with 100 μl Opti-MEM as reagent A. At the same time, 3 μl of Hieff transfection reagent and 100 μl of Opti-MEM were diluted and mixed as reagent B, and left to stand for 5 min. The above reagents A and B were mixed, pipetted and mixed, and then added dropwise to the 12-well plate cells to be transfected after standing for 20 min. They were returned to the 37 ℃ incubator for culture, and fresh culture medium was replaced after 24 h and 48 h. Fluorescence images were taken with a microscope (TS100; Nikon, Tokyo, Japan) at 72 h. Cells were harvested and lysed with cell lysis buffer (Novozyme OneStep Mouse Genotyping Kit).
[0023] 1.4 Editing efficiency detection The above cleavage products were used as templates to perform PCR amplification on the sequences near the target site, and the fragments flanking the target sequence were amplified by two rounds of PCR with barcodes for next-generation sequencing. PCR primers are shown in Table 3: Table 3. Primers for PCR identification of sites in HEK293T cells
[0024] Extract PAM regions without indels within three bases near the PAM. Count the PAMs and use them to generate sequence identifiers.
[0025] The results are as follows Figure 2 As shown, A is a schematic diagram of PAM library PAM-DOSE. B is a representative fluorescence microscopy image of SpaCas9 PAM analysis. C is a sequence mark for identifying PAM based on next-generation sequencing (NGS) results. RHCas9 produces more green fluorescence in HEK293T cells than the wild-type SpaCas9, indicating its stronger cutting ability. Next-generation sequencing shows that PAM is expanded from the original NNGT to NNRH; the coverage of targeted sites is increased by 6 times.
[0026] Example 3: RHCas9 endogenous site editing
[0027] 1.1 sgRNA vector construction 24 endogenous gene sites of HEK293T cells were selected to cover the NNRH site according to the PAM assay results. sgRNA was designed, and a CACC sequence linker was added to the 5' end of the upstream sequence of each crRNA, and an AGAC sequence linker was added to the 5' end of the downstream sequence. The synthesized sgRNA oligos are shown in Table 4: Table 4. HEK293T cell crRNA oligos
[0028] After synthesis, the upstream and downstream were annealed by a program (95°C for 5 min and room temperature for 30 min), and the annealed products were connected to the upstream BnS The Spa-V5 backbone vector was digested and then transformed, single clones were picked and sequenced to obtain the correct sgRNA, and high-purity plasmids were obtained by large-scale extraction.
[0029] 1.2 Cell culture and transfection HEK293T cells were inoculated in DMEM medium containing 10% FBS, which contained 1% double antibody, and cultured in a 37°C cell culture incubator containing 5% CO2. The cells used for transfection were inoculated in 12-well plates for culture the day before. The cells were observed the next day. When the cell density was about 75%, the medium was replaced with a medium containing 10% FBS and no double antibody. HieffTransTM Liposomal Transfection Reagent was used for transfection, and Cas9 and sgRNA (1:1) were co-transfected. The amount of each well of the 12-well plate during transfection was 1 μg. After mixing the plasmids, dilute and mix with 100 μl Opti-MEM as reagent A. At the same time, 3 μl of Hieff transfection reagent and 100 μl of Opti-MEM were diluted and mixed as reagent B, and left to stand for 5 min. Mix the above reagents A and B, pipette and mix well, let it stand for 20 minutes, add dropwise to the 12-well plate cells to be transfected, return to the 37°C incubator for culture, and replace with fresh culture medium after 24 hours and 48 hours. Harvest the cells after 72 hours and lyse them with cell lysis buffer (Novozyme One Step Mouse Genotyping Kit).
[0030] 1.3 RHCas9 editing efficiency detection The above cleavage products were used as templates to perform PCR amplification on the sequences near the target site. The PCR primers are shown in Table 5: Table 5. Primers for PCR identification of sites in HEK293T cells
[0031] After the amplified product was subjected to gel electrophoresis to detect the target band, it was sent to Shanghai Bioengineering Company for first-generation sequencing. The PCR amplification system for the target site was as follows: 2×taq (Tiangen) 12.5μl, F (10 pmol / μl) 1μl, R (10 pmol / μl) 1μl, template 1μl, ddH2O filled to 25μl. After sequencing, the sequencing results were analyzed for efficiency using the EditR analysis website.
[0032] The results are as follows Figure 3 As shown: The results in the endogenous sites are consistent with the PAM detected by the reporter system, and 23 endogenous sites can be effectively edited, while expanding the recognition range while maintaining efficient editing activity. At the GC1 site, 74% editing activity can be achieved.
Claims
1. A modular ancestral reconstruction of a broad-spectrum PAM Cas9 variant, named RHCas9, whose amino acid sequence is shown in SEQ ID NO.
1.
2. The method for precise gene editing of a modular ancestral reconstruction broad spectrum PAM Cas9 variant as claimed in claim 1, comprising the following steps: 1) Homology search: The SpaCas9 protein sequence of Streptococcus pasteurianus was submitted to NCBI, and a homology search was performed using PSI-BLAST site-specific iterative alignment; 2) Homologous sequence comparison and construction of developmental tree: Multiple sequence alignment: Generate multiple sequence alignment of PID region for all 35 homologous sequences through MAFFT; Delete gaps in multiple sequence alignment to ensure alignment quality; Evolutionary tree construction: Use IQ-TREE to automatically select the best amino acid substitution model, evaluate the branch support rate through 1000 Bootstraps, and construct the maximum likelihood tree ML tree; structure prediction and function evaluation; 3) Reconstruct the ancestral sequence: trace the evolutionary path of PID through PAML software, determine the amino acid residues at each node based on the maximum likelihood probability, and select the branch node sequence that eventually differentiates into the target Cas9 as the candidate ancestral variant; 4) Obtaining RHCas9: The reconstructed PID was recombined with the remaining domains of the wild-type SpaCas9 to obtain a modular ancestral reconstructed broad-spectrum PAM Cas9 variant, named: RHCas9, whose amino acid sequence is shown in: SEQ ID NO.1.