Structure-oriented PAM recognition key domain ancestor sequence reconstruction-based Cas9 protein and preparation method thereof

The ancestral sequence of the PAM recognition key domain of Cas9 protein is reconstructed through a structure-oriented method, solving the problems of limitations in the PAM recognition range and functional confusion in the prior art, and achieving the broad-spectrum PAM recognition ability and efficient gene editing effect of Cas9.

CN120098968APending Publication Date: 2025-06-06JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510360282.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art has efficiency limitations and functional confusion in broadening the range of PAM recognition and improving the accuracy of gene editing tools, especially in terms of activity defects and non-targeted functional risks of small-sized Cas9 variants.

Method used

Using a structure-oriented approach, the PAM of Cas9 protein is reconstructed to identify key domain ancestral sequences by integrating structural homologous search and multi-dimensional sequence alignment, breaking through sequence similarity limitations, identifying functional homologous proteins, and obtaining variants compatible with more PAMs through modular ancestral reconstruction.

Benefits of technology

The PAM recognition capability of Cas9 was expanded to NNRR, the target site coverage species was increased by 4 times, and the editing efficiency reached 44.3% in mammalian cells, while maintaining the activity and specificity of the enzyme.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120098968A_ABST
    Figure CN120098968A_ABST
Patent Text Reader

Abstract

The invention provides a Cas9 protein based on structure-oriented PAM recognition key domain ancestor sequence reconstruction and a preparation method thereof, through the technology of integrating structure homologous search and multi-dimensional sequence alignment to carry out ancestor sequence reconstruction, optimizing a PAM recognition key domain of a gene editing tool SpaCas9 and expanding the recognition range of the PAM recognition key domain, based on structure information of PID, structure homologous screening is carried out, and the structure-oriented PAM recognition key domain ancestor sequence reconstruction is obtained. The limitation of sequence similarity is broken through, and functional homologous protein is identified. A PAM recognition key domain of proteins with function homology but different PAM recognition preferences is subjected to ancestral reconstruction to obtain a variant compatible with more PAMs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a Cas9 protein for structure-guided PAM recognition key domain ancestral sequence reconstruction, and an engineering method for preparing the same, which integrates structural homology search and multi-dimensional sequence alignment to reconstruct the ancestral sequence, optimizes the PAM recognition key domain of the gene editing tool SpaCas9 and expands its recognition range, and belongs to the field of gene editing technology. Technical Background

[0002] CRISPR / Cas9 technology introduces DNA double-strand breaks (DSBs) at target sites to achieve site-specific gene editing, which has played a revolutionary role in the biomedical field. The protospacer adjacent motif (PAM) is a short DNA sequence that must be recognized by CRISPR-Cas proteins before binding and cutting DNA targets. PAM binding is essential for initiating DNA unwinding, R-loop formation, and efficient positioning of genomic targets. In nature, the co-evolution of bacteria and bacteriophages has promoted the diversification of Cas proteins, enabling them to recognize a wide range of PAMs. In genome editing, PAM is the key to specificity, but it also limits the range of editable genomic sites, which poses a challenge to technologies such as base editing and homology-directed repair that require precise positioning of Cas proteins.

[0003] At present, there are many studies related to broadening PAM, but they are mainly focused on SpCas9 (1368aa). The main strategies used for optimization are directed evolution screening (such as bacteria or phage-assisted continuous evolution), homologous PAM recognition domain (PID) replacement, and rational structural design mutation. However, these methods require a large number of repetitive experiments, a large number of empirical trial and error, a long development cycle, and are difficult to extend to other Cas proteins. And the recent deep learning CRISPR-CAS PAM recognition engineered Protein2PAM, Protein2PAM can quickly and accurately predict PAM specificity directly from the Cas proteins of type I, type II, and type V CRISPR-Cas systems, but when deep mutation scanning is used to compute evolution of PAM, iterative mutations will introduce unnatural mutations, resulting in low variant efficiency, emphasizing the importance of natural mutations.

[0004] Ancestral sequence reconstruction (ASR) technology infers and calculates the sequence of internal nodes of the phylogenetic tree. The reconstructed ancestral proteins show enhanced stability, increased activity and substrate promiscuity, which can break through the functional limitations of modern proteins. Studies have shown that when ASR is performed on Cas proteins, the reconstructed ancestral sequence of Cas9 (FCA anCas) and the reconstructed ancestral sequence of Cas12a (ReChb) break through the PAM limitations of modern proteins, but due to the perturbation of the whole sequence, non-specific functional confusion (such as non-targeted nuclease activation or sgRNA scaffold-dependent changes) is caused, resulting in a significant increase in the risk of off-target, which cannot meet the needs of precise gene editing.

[0005] In recent years, the intersection of structural biology and computational biology has provided new ideas for protein engineering. The structure of the Cas9-DNA complex analyzed by cryo-electron microscopy shows that the PAM recognition ability is determined by the topological conformation of PID, and its evolution may achieve PAM diversity through the coordinated variation of key residues. At the same time, the protein function mining strategy based on structural homology has made important progress. Gao Caixia's team systematically analyzed the functional sites of deaminases through AI-driven protein tertiary structure clustering; Jennifer Doudna's team used AlphaFold2 structure comparison to discover new nucleases such as Cas13an and reveal their evolutionary path; Zhang Feng's research group combined structure and sequence evolution tracking and proposed that the CRISPR-Cas13 system originated from the ancient RNA toxin-antitoxin system. These studies have proved that homology searches based on three-dimensional structures can break through sequence restrictions and effectively mine distant proteins with significant sequence differences but conservative functions. However, structural information has not yet been combined with ASR, and existing ASR lacks structure-guided modular design, making it difficult to accurately release the evolutionary potential of target domains (such as PID) while maintaining the natural activity of other functional domains (such as HNH / RuvC).

[0006] In summary, a new method that integrates structural constraints and modular ancestral reconstruction is urgently needed to address the functional promiscuity of traditional ASR, the efficiency limitations of existing PAM extension technology, and the activity defects of small-sized Cas9 variants, thereby achieving the development of efficient and precise broad-spectrum gene editing tools. Summary of the invention

[0007] The present invention discloses a Cas9 protein based on structure-guided PAM recognition key domain ancestral sequence reconstruction and a preparation method. Based on the structural information of the PAM recognition key domain, structural homology screening is performed to break through the sequence similarity limitation and identify functional homologous proteins; proteins with functional homology but different PAM recognition preferences are ancestrally reconstructed to obtain variants that are compatible with more PAMs.

[0008] The Cas9 protein (named as FAnPID-SpaCas9) based on structure-guided PAM recognition key domain ancestral sequence reconstruction of the present invention is characterized in that its gene sequence is: SEQ ID: NO.1.

[0009] The method for preparing a Cas9 protein based on structure-guided PAM recognition key domain ancestral sequence reconstruction of the present invention comprises the following steps: 1. Using the SpaCas9 protein of Streptococcus pasteurianus as a prototype, the SpaCas9, target DNA, and sgRNA ternary complex was predicted by Alphafold3. The key regions where SpaCas9 interacted directly or indirectly with the PAM of the target DNA were observed. The key PAM recognition domain was 90 amino acids from G1040 to K1130 of SpaCas9, and the key residue was R1086. 2. Use PyMOL software to open the ternary complex structure of SpaCas9, and then extract the PID part (P967 to K1130) separately. The DALI program is a tool for comparing protein structures, which can help find other proteins with similar structures to the target protein. The DALI program uses the PDB50 database to search for structural similarity of the isolated SpaCas9 PID. From a large number of search results, Cas9 proteins that recognize different PAMs compared to SpaCas9 are found; a total of 6 qualified proteins are obtained, and the PDB numbers (and protein names) are as follows: 6RJ9 (St1Cas9), 5X2G (CjCas9), 6JDQ (Nme1Cas9), 8UZA (GeoCas9), 8D2K (AceCas9), 8WMM (CbCas9); 3. Use Pymol to extract the PAM recognition key domain structure of each protein separately, use DALI to compare the PAM recognition key domains of all 7 proteins ALL aginst ALL, obtain the structural similarity (Z-score) score, generate a similarity matrix, and iteratively merge adjacent nodes through a hierarchical clustering algorithm to construct a structure tree; DALI is used to align the 6 Cas9s with SpaCas9 to generate a structural sequence comparison of the PAM recognition key domain; 4. The structural multiple sequence alignment of the PAM recognition key domains of the seven Cas9 proteins and the structure tree of the seven Cas9 proteins were used as the input for PAML ancestral reconstruction, and CodeML was used to infer the ancestral sequence and select the node representing the ancestor of the sequence; 5. The reconstructed ancestral PAM recognition key domain was recombined with the remaining domains of the wild-type SpaCas9 to obtain a Cas9 protein reconstructed based on the structure-guided PAM recognition key domain ancestral sequence, named: FAnPID-SpaCas9 (SEQ ID: NO.1); 6. Validate its PAM recognition preference through a targeted PAM library system combined with deep sequencing, and evaluate the editing efficiency in endogenous sites in mammalian cells.

[0010] The present invention provides a method of effectively multi-dimensionally comparing homologous proteins with large sequence differences through structural search, and then reconstructing the PAM recognition key domain of Cas9 through modular ancestral sequence reconstruction to obtain a variant with mixed PAM recognition, that is, a variant with broad-spectrum PAM recognition ability. It gets rid of the limitation of sequence homology, further expands the difference, and obtains variants with more mixed functions and wider recognition. It does not need to sacrifice the cleavage activity of the HNH domain or the DNA binding specificity of the REC domain, and maximizes the functional breadth of a single domain, so that the PAM recognition ability of Cas9 has greater compatibility.

[0011] The positive effects of the present invention are: based on the structural information of PID, structural homology screening is performed to break through the sequence similarity limitation and identify functional homologous proteins. The PAM recognition key domains of proteins with functional homology but different PAM recognition preferences are reconstructed to obtain variants that are compatible with more PAMs, which are specifically manifested in the following characteristics: 1. Breaking through the limitation of sequence homology: The lowest sequence homology of homologous proteins based on structure search is 26.67%. And through the alignment of structure-guided PAM recognition domains, the key residues of PAM recognition can be better aligned, while the deviation based on sequence alignment is large; 2. Broad-spectrum PAM recognition capability: The reconstructed ancestral Cas9 variant can recognize non-classical PAM 5'-NNRR -3', and the coverage of genomic targets is increased by 4 times, such as Figure 2 As shown in C; 3. Efficient editing: The editing efficiency in mammalian cells can reach 44.3%, such as Figure 3 shown.

[0012] Through a structure-guided modular ancestral reconstruction, a broad spectrum of PAM Cas9 variants was obtained. Structure guidance breaks away from the limitations of sequence homology, further expands the differences, and obtains variants with more mixed functions and wider recognition; modularity avoids the loss of catalytic activity or non-targeted functional confusion (such as ssDNA cutting) caused by global transformation. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 , oneA flowchart of structure-guided reconstruction of PAM recognition domain ancestors; Figure 2 , Flow chart and result diagram of FAnPID-SpaCas9 detecting PAM in HEK293T cells; Figure 3 , the results of FAnPID-SpaCas9 at 13 endogenous sites. DETAILED DESCRIPTION

[0014] The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention. The present invention is further described in detail below in conjunction with specific examples and with reference to data. It should be understood that these examples are only for illustrating the present invention, and are not intended to limit the scope of the present invention in any way. If not specifically indicated, the technical means used in the examples are conventional means well known to those skilled in the art. The materials, reagents, etc. used in the following examples, if not specifically stated, can all be obtained from commercial sources.

[0015] Example 1

[0016] Step 1: Determine the key domain and key amino acids for PAM recognition Using the SpaCas9 protein of Streptococcus pasteurianus as a prototype, the SpaCas9, target DNA, and sgRNA ternary complex was predicted by Alphafold3, and the key regions where SpaCas9 interacted directly or indirectly with the PAM of the target DNA were observed. The key PAM recognition domain is the 90 amino acids from G1040 to K1130 of SpaCas9. The 1086th R in SpaCas9 will specifically interact with the third G and the fourth T in the PAM region of the target DNA, and R1086 is the key amino acid for PAM recognition.

[0017] Step 2: DALI program performs structural similarity search on SpaCas9 PID The ternary complex structure of SpaCas9 was opened with PyMOL software, and the PID part was extracted separately. The isolated SpaCas9 PID structure was searched for structural similarity using the PDB50 database using the DALI program. From a large number of search results, Cas9 proteins that recognize different PAMs compared to SpaCas9 were found; a total of 6 proteins that meet the criteria were obtained, as shown in Table 1: Table 1. PDB and PAM preferences of proteins with similar structures to SpaCas9 PID

[0018] Step 3: DALI generates structural sequence comparison and structural evolution tree The PID structure of each protein was extracted using Pymol, and the PAM recognition key domains of all seven proteins were compared using DALI to obtain the structural similarity (Z-score) score, generate a similarity matrix, and construct a structure tree by iteratively merging adjacent nodes through a hierarchical clustering algorithm, such as Figure 1 As shown in B; DALI was used to align the six Cas9s with SpaCas9 as the first, generating a structural sequence comparison of the PAM recognition key domain; Figure 1 As shown in C, the key amino acids for PAM recognition of each Cas9 protein can be well aligned ( Figure 1 C), while the use of MAFFT alignment based solely on sequence results in a large shift in the position of key amino acids ( Figure 1 C); Step 4: PAML ancestral sequence reconstruction The structural multiple sequence alignment of the PAM recognition key domains of 7 Cas9 proteins and the structure tree of 7 Cas9 proteins were used as the input for PAML ancestral reconstruction. CodeML was used to infer the ancestral sequence and select the nodes representing the ancestral sequence. The structure-guided PAM recognition domain ancestral reconstruction sequence was obtained and named FAnPID. Figure 1 D and 1E.

[0019] Step 5: Generate FAnPID-SpaCas9 The reconstructed ancestral PAM recognition key domain FAnPID was recombined with the remaining domains of the wild-type SpaCas9 to obtain the Cas9 protein reconstructed based on the structure-guided PAM recognition domain ancestral, FAnPID-SpaCas9 (SEQ ID: NO.1).

[0020] The structure-guided PAM recognition domain ancestor reconstruction flow chart of the present invention is as follows: Figure 1 As shown: A, Flowchart of structure-guided modular ancestral reconstruction of the key domain of SpaCas9 PAM recognition; B, Structural similarity heat map and structure tree of the all-to-all comparison of the seven Cas9 PAM recognition key domains; C. Structure-guided alignment results and sequence-based alignment results; D, partial magnified image of PID of FAnPID-SpaCas9; E. Comparison of the key amino acid residues that interact with PAM in the PAM recognition key domain and their surrounding sequences.

[0021] Example 2: PAM assay of FAnPID-SpaCas9 in mammalian cells PAM detection system

[0022] GenScript synthesized FAnPID-SpaCas9 with humanized codon optimization; used the previously developed PAM detection system PAM-DOSE positive screening system to identify functional PAMs in human cells. pmTmG contains two fluorescent protein genes (tdTomoto and EGFP) and two target sequences (sloxP and Spacer N4), see Figure 2: A, schematic diagram of PAM library PAM-DOSE; B, representative fluorescence microscopy image of SpaCas9 PAM analysis; C, sequence marks for identifying PAMs based on next-generation sequencing (NGS) results. The random PAM sequences after excision can be recovered and analyzed by Sanger sequencing or high-throughput sequencing (NGS), and the frequency of each PAM type is counted based on the sequencing data.

[0023] 1.2 sgRNA vector construction sgRNA was designed according to the upstream sequence of the PAM library, and a CACC sequence linker was added to the 5' end of the upstream sequence of each crRNA, and an AGAC sequence linker was added to the 5' end of the downstream sequence.

[0024] The synthesized sgRNA oligos are shown in Table 2: Table 2. HEK293T cell crRNA oligos Primers Sequence (5'-3') PAM-sgF CACCggctcgtgcgaacagttcag PAM-sgR AGACctgaactgttcgcacgagcc After synthesis, the upstream and downstream were annealed by a procedure (95 °C, 5 min; room temperature cooling for 30 min), and the annealed products were connected to the upstream BnS The Spa-V5 backbone vector (SEQ ID: NO.2) was digested and then transformed, single clones were picked and sequenced to obtain the correct sgRNA, and high-purity plasmids were obtained by large-scale extraction.

[0025] 1.3 Cell culture and transfection HEK293T cells were inoculated in DMEM medium with 10% FBS, containing 1% double antibody, and in an atmosphere of 5% CO 2The cells were cultured in a 37 ℃ cell culture incubator. The cells used for transfection were inoculated in a 12-well plate the day before and cultured. The cells were observed the next day. When the cell density was about 75%, the medium was replaced with a medium containing 10% FBS and no double antibody. HieffTransTM Liposomal Transfection Reagent was used for transfection. SpaCas9, sgRNA and pmTmG reporter plasmid (1:1:1:1) were co-transfected. The total amount of each well of the 12-well plate during transfection was 1 μg. After mixing the plasmids, they were diluted and mixed with 100 μl Opti-MEM as reagent A. At the same time, 3 μl of Hieff transfection reagent and 100 μl of Opti-MEM were diluted and mixed as reagent B, and left to stand for 5 min. The above reagents A and B were mixed, pipetted and mixed, and then added dropwise to the 12-well plate cells to be transfected after standing for 20 min. They were returned to the 37 ℃ incubator for culture, and fresh culture medium was replaced after 24 h and 48 h. Fluorescence images were taken with a microscope (TS100; Nikon, Tokyo, Japan) at 72 h. Cells were harvested and lysed with cell lysis buffer (Novozyme OneStep Mouse Genotyping Kit).

[0026] 1.4 PAM Preference Detection The above cleavage products were used as templates to perform PCR amplification on the sequences near the target site, and the fragments flanking the target sequence were amplified by two rounds of PCR with barcodes for next-generation sequencing. PCR primers are shown in Table 3: Table 3. Primers for PCR identification of sites in HEK293T cells

[0027] Extract PAM regions without indels within three bases near the PAM. Count the PAMs and use them to generate sequence identifiers.

[0028] The results are as follows Figure 2 As shown in the figure, FAnPID-SpaCas9 produces more green fluorescence in HEK293T cells than the wild-type SpaCas9, indicating its stronger cutting ability. Next-generation sequencing showed that the PAM of FAnPID-SpaCas9 was expanded from the original NNGT to NNRR, and the coverage of targeted sites increased by 4 times. This proves that the FAnPID-SpaCas9 version retains the enzyme activity to the greatest extent under the same PAM expansion range.

[0029] Example 3: FAnPID-SpaCas9 endogenous site editing efficiency 1.1 sgRNA vector construction 13 endogenous gene sites of HEK293T cells were selected, and the NNGN and NNRG sites were covered according to the PAM assay results. sgRNA was designed, and a CACC sequence linker was added to the 5' end of the upstream sequence of each crRNA, and an AGAC sequence linker was added to the 5' end of the downstream sequence. The synthesized sgRNA oligos are shown in Table 4: Table 4. HEK293T cell crRNA oligos

[0030] After synthesis, the upstream and downstream were annealed by a program (95°C for 5 min and room temperature for 30 min), and the annealed products were connected to the upstream BnS The Spa-V5 backbone vector was digested and then transformed, single clones were picked and sequenced to obtain the correct sgRNA, and high-purity plasmids were obtained by large-scale extraction.

[0031] 1.3 Cell culture and transfection HEK293T cells were inoculated in DMEM medium containing 10% FBS, which contained 1% double antibody, and cultured in a 37°C cell culture incubator containing 5% CO2. The cells used for transfection were inoculated in 12-well plates for culture the day before. The cells were observed the next day. When the cell density was about 75%, the medium was replaced with a medium containing 10% FBS and no double antibody. HieffTransTM Liposomal Transfection Reagent was used for transfection, and Cas9 and sgRNA (1:1) were co-transfected. The amount of each well of the 12-well plate during transfection was 1 μg. After mixing the plasmids, dilute and mix with 100 μl Opti-MEM as reagent A. At the same time, 3 μl of Hieff transfection reagent and 100 μl of Opti-MEM were diluted and mixed as reagent B, and left to stand for 5 min. Mix the above reagents A and B, pipette and mix well, let it stand for 20 minutes, add dropwise to the 12-well plate cells to be transfected, return to the 37°C incubator for culture, and replace with fresh culture medium after 24 hours and 48 hours. Harvest the cells after 72 hours and lyse them with cell lysis buffer (Novozyme One Step Mouse Genotyping Kit).

[0032] 1.4 FAnPID-SpaCas9 editing efficiency detection The above cleavage products were used as templates to perform PCR amplification on the sequences near the target site. The PCR primers are shown in Table 5: Table 5. Primers for PCR identification of sites in HEK293T cells

[0033] After the target band was detected by gel electrophoresis, the amplified product was sent to Shanghai Bioengineering Company for first-generation sequencing. The PCR amplification system for the target site was as follows: 2× Taq (Tiangen) 12.5 μl, F (10 pmol / μl) 1 μl, R (10 pmol / μl) 1 μl, template 1 μl, ddH 2 O was added to 25 μl. After sequencing, the sequencing results were analyzed for efficiency using the EditR analysis website. The test results were as follows: Figure 3 As shown: The results in the endogenous site are consistent with the PAM detected by the reporter system, and the PAM site of NNRR can be effectively edited, while expanding the recognition range and maintaining the editing activity. The activity can reach 44.3% (GC2 site).

[0034] Experimental data show that the alignment of the structure-guided PAM recognition domain in the present invention can better align the key residues for PAM recognition, while the deviation based on sequence alignment is large. The FAnPID-SpaCas9 variant developed by the present invention expands the PAM recognition range of the wild-type SpaCas9 from NNGT to NNRR (R is purine). It can effectively edit the G at the fourth position of the variant obtained based on traditional sequence reconstruction that cannot be edited. In human cells, FAnPID-SpaCas9 showed effective editing of traditional non-editable PAM sites, up to 44.3%. FAnPID-SpaCas9 widens the PAM while maintaining editing efficiency, achieving optimization of specific functions.

[0035] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. A Cas9 protein based on structure-guided PAM recognition key domain ancestral sequence reconstruction, characterized in that: The gene sequence is as follows: SEQ ID: NO.

1.

2. The method for preparing the Cas9 protein based on structure-guided PAM recognition key domain ancestral sequence reconstruction as claimed in claim 1, characterized in that: 1) Using the SpaCas9 protein of Streptococcus pasteurianus as a prototype, the SpaCas9, target DNA, and sgRNA ternary complex was predicted by Alphafold3, and the key regions where SpaCas9 directly or indirectly interacted with the PAM of the target DNA were observed. The key PAM recognition domain was 90 amino acids from G1040 to K1130 of SpaCas9, and the key residue was R1086; 2) Use PyMOL software to open the ternary complex structure of SpaCas9 and extract the PID part separately; use the DALI program to search the structural similarity of the isolated SpaCas9 PID structure using the PDB50 database to obtain a Cas9 protein that recognizes a different PAM compared to SpaCas9: A total of 6 qualified proteins were obtained, and their PDB numbers (and protein names) are as follows: 6RJ9 (St1Cas9), 5X2G (CjCas9), 6JDQ (Nme1Cas9), 8UZA (GeoCas9), 8D2K (AceCas9), and 8WMM (CbCas9); 3) Pymol was used to extract the PAM recognition key domain structure of each protein separately, and DALI was used to compare the PAM recognition key domains of all 7 proteins ALL aginst ALL to obtain the structural similarity (Z-score) score, generate a similarity matrix, and iteratively merge adjacent nodes through a hierarchical clustering algorithm to construct a structure tree; DALI was used to align the 6 Cas9s with SpaCas9 to generate a structural sequence comparison of the PAM recognition key domain; 4) The structural multiple sequence alignment of the PAM recognition key domains of the seven Cas9 proteins and the structure tree of the seven Cas9 proteins were used as input for PAML ancestral reconstruction, and CodeML was used to infer the ancestral sequence and select the nodes representing the sequence ancestors; 5) The reconstructed ancestral PAM recognition key domain is recombined with the remaining domains of the wild-type SpaCas9 to obtain a Cas9 protein reconstructed based on the structure-guided PAM recognition key domain ancestral sequence; 6) Validate its PAM recognition preference through a targeted PAM library system combined with deep sequencing and evaluate the editing efficiency in endogenous sites in mammalian cells.