Cas protein pam analysis method based on local geometry and machine learning

CN121483365BActive Publication Date: 2026-04-10ZHEJIANG LAB
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies for identifying the PAM sequence of Cas12a protein involve complex and costly experimental methods, and high-throughput screening is insufficient for identifying non-classical PAM sequences. Traditional methods for analyzing protein structures are time-consuming and laborious, making it difficult to objectively and accurately process and analyze data.

Method used

By combining local geometry and machine learning, the structure of the ternary complex was constructed by obtaining the sequences of Cas protease, crRNA, and target DNA using structure prediction tools. The PAM sequence was analyzed using geometric vector perception neural network and unsupervised clustering algorithm, and identified by combining prior biological knowledge.

Benefits of technology

It enables more objective and reliable PAM sequence identification, reduces experimental bias, provides a comprehensive and comparable analytical method, lowers experimental costs, and is suitable for PAM prediction of novel Cas proteases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121483365B_ABST
    Figure CN121483365B_ABST
Patent Text Reader

Abstract

The application discloses a Cas protein PAM analysis method based on local geometric structure and machine learning, and belongs to the fields of biotechnology and artificial intelligence. The method comprises the following steps: constructing an initial structure of a Cas protein-DNA-RNA ternary complex by using a structure prediction tool, and generating a mutant structure data set containing different PAM sequences by calculation and optimization; extracting and embedding the local structure of a PAM site to obtain a structure feature vector, performing dimension reduction and cluster analysis by using an unsupervised machine learning method, realizing natural classification of the PAM sequence, and identifying the cluster result in combination with P prior biological knowledge. The application creatively combines the three-dimensional local geometric structure of a protein and machine learning, can predict and classify the PAM preference of a Cas protein from a structure source, has the advantages of objectivity, comprehensiveness, and the ability to reveal hidden patterns, and can be verified by experimental data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of biotechnology and artificial intelligence, and designs a Cas protein PAM analysis method based on local geometric structure and machine learning. BACKGROUND

[0002] CRISPR-Cas system (Clustered Regularly Interspaced Short Palindromic Repeats-CRISPR associated proteins) is an adaptive immune defense mechanism evolved by bacteria and archaea, and is also an important tool in the field of gene editing. CRISPR sequence is a special DNA sequence composed of a series of short palindromic repeat sequences (repeat) and spacer sequences (spacer). Cas protein is a class of proteins related to CRISPR sequence, which plays a key role in CRISPR system. Cas protein is diverse, among which commonly used are Cas9, Cas12 and other proteins. Cas protein is a nuclease that can recognize and bind to specific DNA sequences, and then cut DNA through its nuclease activity, thereby achieving editing of target DNA.

[0003] Cas12a (originally called Cpf1) is a type II V nuclease in the CRISPR-Cas family, which has unique properties compared to the commonly used Cas9, providing more tool options for gene editing. Cas12a protease can specifically target nucleic acids and activate transcleavage activity, non-specifically and continuously cutting single-stranded nucleic acid reporter genes. The characteristics of this "all-purpose" protease make it have wide application potential in gene editing and molecular diagnosis. Cas12a achieves target recognition by recognizing the protospacer adjacent motif (PAM) sequence, forming an R-loop structure, activating nuclease activity and cutting target DNA. Among them, the recognition of PAM sequence is the initial step of Cas12a target recognition, which determines the DNA region that Cas12a can bind to. Cas12a recognizes PAM sequences rich in thymine (T), such as TTTV (AsCas12a or LbCas12a), while Cas9 recognizes PAM sequences rich in guanine (G) (such as NGG).

[0004] Currently, the identification of PAM (Protospacer Adjacent Motif) sequences of Cas12a protein mainly relies on the strategy combining experimental methods and bioinformatics prediction. Experimental methods include intracellular library screening and in vitro library screening. Intracellular screening refers to expressing Cas12a protein and crRNA in cells, targeting specific gene sites, and analyzing the indel (indel) frequency of the target site by deep sequencing to infer the activity of the PAM sequence. The advantage of this method is that it can directly evaluate the activity of the PAM sequence in the cell environment, and the result is closer to the actual application situation. However, it is complex to construct a PAM library in cells, the experimental period is long, and complex cell culture and sequencing analysis are required. In vitro high-throughput PAM screening experiments, such as PAM-SCANR / PAM-DEPENDENT, the basic principle is to construct a DNA library containing random PAM sequences, incubate with Cas12a-crRNA complex, and then analyze the cut DNA fragments by sequencing to determine the preferred PAM sequence of Cas12a. However, the cost of synthesis and sequencing is high, and high-throughput screening is sensitive to strong PAM (such as TTTV), but may not be sufficient for the recognition of non-classical PAM, resulting in an incomplete preference map.

[0005] The structure of a protein determines its function. Currently, the identification of PAM of Cas protein is often based on experimental testing, but protein structure is the basis of its activity, so structural feature data also has important reference and reference significance in PAM prediction. Traditional analysis of a complex structure is time-consuming and laborious, but the development of computational biology makes it possible to complete structure prediction in a very short time, and predicting a large number of protein structure data is no longer a distant goal. It is of great significance to make full use of structural data to predict PAM. Machine learning is currently changing the methods and progress of scientific research with its powerful data processing and analysis capabilities. We creatively introduced geometric vector perception neural network into the processing of Cas protein complex structure data, and used machine learning to naturally group PAM in the data. This method can find hidden patterns, reduce experimental bias, and more objectively process and analyze data, which are not possessed by traditional methods. SUMMARY

[0006] To overcome the shortcomings of the prior art, the present application provides a Cas proteinase PAM analysis method based on local geometric structure and machine learning. The present application combines Cas proteinase structure characteristics and activity data to predict PAM, which can find natural grouping of PAM in data, find hidden patterns, reduce experimental bias, more objectively process and analyze data, and the conclusion is more reliable.

[0007] The technical scheme of the present application is as follows:

[0008] In a first aspect, the present application discloses a Cas protein PAM analysis method based on local geometric structure and machine learning, which comprises:

[0009] S1: Obtain target Cas protein, crRNA and target DNA sequence, and use a structure prediction tool to construct an initial structure of a Cas protein-double-stranded DNA-RNA ternary complex;

[0010] S2: Use the initial structure as a template, select different PAM region lengths according to the Cas protein species, and perform site-directed mutagenesis and structure optimization on the PAM site of the target DNA through a calculation tool to generate a mutant ternary complex structure dataset containing different PAM sequences;

[0011] S3: For each structure in the mutant ternary complex structure dataset, extract the local structure centered on the PAM mutation site; use a geometric vector perception graph neural network to extract and encode the features of the local structure to generate a structure feature embedding vector corresponding to each PAM mutation site;

[0012] S4: Input the structure feature embedding vector obtained in step S3 into a machine learning model; first perform dimension reduction processing, and then use an unsupervised clustering algorithm to perform clustering analysis on the dimension-reduced data to obtain natural grouping of different PAM sequences, identify the clustering results in combination with prior biological knowledge of PAM sequences, and obtain PAM sequence groups that can be recognized by the target Cas protein.

[0013] In the present application, the prior biological knowledge of PAM sequences refers to biological knowledge that is well known and recognized in the industry; for example, PAM sequences that have been verified by in vitro experiments to be preferred by Cas proteinases or PAM sequences that cannot be recognized; for example, Cas12a prefers T-rich sequences and Cas9 prefers G-rich sequences, and it can also be a known classic PAM sequence that can be recognized by different Cas proteinases (such as the classic PAM sequence TTTV of Cas12a proteinase). According to these prior biological knowledge, the clustering results are interpreted and identified, and the PAM cluster that best fits the prior biological knowledge is selected as the PAM sequence group that can be recognized by the target Cas protein.

[0014] Typically but not limitatively, a number of prior biological knowledge can be determined first, and then the prior biological knowledge is used to score each clustering result, and each clustering structure is sorted according to the total score of each clustering result.

[0015] In addition, the method of the present application can also calculate and predict the results (PAM classification) for completely new Cas subtypes, thereby providing clear guidance for subsequent experimental design and verification.

[0016] In a second aspect, the present application further discloses a Cas protein PAM analysis device based on local geometric structure and machine learning, comprising:

[0017] one or more processors;

[0018] a memory for storing one or more programs;

[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned Cas protein PAM analysis method based on local geometric structure and machine learning.

[0020] Compared with the prior art, the present application has the following beneficial effects:

[0021] The present application uses structural prediction tools and computational tools to obtain the structure of Cas12a-DNA-RNA mutant complex, uses geometric vector perception neural network to embed the structure data, and then uses unsupervised machine learning to reduce dimension and cluster the embedded data, thereby discovering the PAM classification of Cas enzyme. The present application makes full use of the structural features of Cas protein-DNA-RNA, and analyzes and predicts the recognition PAM classification of protease from the structural source. The present application creatively introduces geometric vector perception neural network and PCA, t-SNE, GMM and K-means unsupervised machine learning into Cas protein structure feature data processing and PAM identification. This method can discover the natural grouping of PAM in data, more objectively processes and analyzes data, and the conclusion is more reliable. These are not possessed by traditional methods, and provide a comprehensive and comparable analysis method for PAM identification. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is the structure of Cas12a-DNA-crRNA complex;

[0023] Figure 2 is the structure obtained by AlphaFold3 and pyrosetta optimization;

[0024] Figure 3 is the 4-cluster and 6-cluster result based on the combination of PCA and GMM;

[0025] Figure 4 is the 4-cluster and 6-cluster result based on the combination of t-SNE and GMM;

[0026] Figure 5 is the 4-cluster and 6-cluster result based on the combination of t-SNE and K-means;

[0027] Figure 6 is the clustering result based on activity data;

[0028] Figure 7 is the Cas12a protein cleavage reaction result under different PAMs in vitro;

[0029] Figure 8 is the PAM cluster result of Cas9;

[0030] Figure 9 is the overall flowchart of the scheme of the present application. DETAILED DESCRIPTION

[0031] The present application will be further described and illustrated with reference to the specific embodiments. The embodiments are only exemplary and do not circumscribe the scope of the present disclosure. The technical features of each embodiment of the present application can be combined accordingly without conflict.

[0032] As shown in Figure 9 , the Cas proteinase PAM analysis method based on local geometric structure and machine learning of the present application mainly includes the following steps:

[0033] S1: Obtain target Cas proteinase, crRNA and target DNA sequence, and use a structure prediction tool to construct the initial structure of the Cas proteinase-double-stranded DNA-RNA ternary complex;

[0034] The ternary complex of Cas proteinase-double-stranded DNA-RNA is the core executor of the CRISPR-Cas system to exert its gene editing function. For Cas12a proteinase and Cas13 proteinase, the RNA part is called crRNA; for Cas9 proteinase, it is called gRNA. The target DNA sequence contains at least a PAM sequence and a Protospacer sequence part, wherein the Protospacer sequence part corresponds to the spacer sequence in the RNA sequence.

[0035] The PAM region length is selected as n, and the value of n is determined according to the type of Cas proteinase, n = 2-8; the number of ternary complex structures in the data set is 4 n . If 4 bases are selected as PAM, 4 4 complex structure data sets are generated for each initial structure; the preferred range of the number of PAM sequence positions is 3-8. If it is a new Cas enzyme and the number of PAM bases is unknown, 6 or 8 can be selected.

[0036] Typically but not limitedly, the present application uses AlphaFold3 online prediction to obtain the initial structure of the ternary complex. When the Cas proteinase is Cas12a protein, the PAM sequence in the initial structure of the ternary complex can be selected as TTTG; when the Cas proteinase is Cas9 protein, the PAM sequence in the initial structure of the ternary complex can be selected as GGGC, and the PAM sequence is selected by one base, which verifies whether the method is sensitive to the fourth base.

[0037] S2: using the initial structure as a template, performing site-directed mutation and structure optimization on the PAM site of the target DNA by a computing tool to generate a mutant ternary complex structure dataset containing different PAM sequences;

[0038] According to a preferred scheme of the present application, PyRosetta tool is used for site-directed mutation of PAM site, and side chain rearrangement and energy optimization in the adjacent range of the mutation region to generate the mutant ternary complex structure dataset.

[0039] Specifically, the PyRosetta tool is used to mutate the PAM site on the DNA double strand of the initial structure from TTTG to other PAM sequences and perform side chain rearrangement to generate an initial mutant structure; then the side chain rearrangement is performed in the 4 Å range adjacent to the mutation region to eliminate the side chain collision introduced by the mutation; the energy minimization is performed in the 6 Å range adjacent to the mutation region; the Relax optimization is performed in the 8 Å range adjacent to the mutation region to obtain the final low-energy conformation.

[0040] S3: for each structure in the mutant ternary complex structure dataset, extract the local structure centered on the PAM mutation site; use a geometric vector perception graph neural network to extract and encode the features of the local structure to generate a structure feature embedding vector corresponding to each PAM mutation site;

[0041] The extracted local structure is centered on the PAM mutation site, and the local structure within an 8 Å radius range is extracted, and the graph adjacency relationship between atoms is constructed based on a 6 Å radius neighborhood. The node features in the adjacency graph include: scalar features composed of one-hot encoded residue types and coordinates of atoms relative to the mutation center, and vector features composed of normalized direction vectors.

[0042] The scalar features and the vector features are processed by two independent branch encoders respectively: the scalar branch is mapped to a 512-dimensional vector by a multilayer perceptron, and the vector branch is mapped to a 128-dimensional vector; after the outputs of the two branches are spliced at the node level, they are further integrated by a fusion layer, and finally a 512-dimensional structure feature embedding vector is generated by global average pooling.

[0043] S4: The structural feature embedding vector obtained in step S3 is input into the machine learning model; dimensionality reduction is performed first, and then unsupervised clustering algorithm is used to perform cluster analysis on the dimensionality-reduced data to obtain natural groupings of different PAM sequences. The clustering results are then identified by combining prior biological knowledge of PAM sequences (e.g., Cas12a prefers T-rich sequences, Cas9 prefers G-rich sequences); thus, the PAM sequence groupings that the target Cas protease can recognize are obtained. This invention can provide the necessary prerequisite information for subsequent use of Cas enzymes in nucleic acid detection and gene editing. For completely novel Cas subtypes, the computational prediction results of this invention can provide clear guidance for subsequent experimental design and validation, and reduce or avoid the cost of in vitro experiments.

[0044] The dimensionality reduction process is principal component analysis or t-distributed random neighborhood embedding; the unsupervised clustering algorithm is a Gaussian mixture model or a K-means clustering algorithm.

[0045] The steps of the present invention will be described in detail below with reference to specific embodiments.

[0046] 1. Construction of Structural Feature Dataset

[0047] First, the Cas12a protease sequence for which PAM needs to be predicted is obtained from the database. In this embodiment, the AsCas12a protease sequence is selected. The crRNA sequence of the AsCas12a protease is obtained from the database, and the corresponding target DNA sequences are selected. In this embodiment, six target DNA sequences are selected: EMX1, DNMT1, MerS, eGFP-1, eGFP-3, and FANCF. AlphaFold3 is used to predict the structure of the ternary complex composed of Cas12a protease, double-stranded DNA, and crRNA. Because six targets are used, the PAM sequence TTTG is used as the wild type. The complex structures corresponding to the six target DNA sequences Mers_TTTG, DNMT1_TTTG, EMX1_TTTG, eGFP-1_TTTG, eGFP-3_TTTG, and FANCF_TTTG are predicted and used as the initial wild-type structures, such as... Figure 1 As shown.

[0048] After obtaining 6 initial wild-type structures, use PyRosetta tools to mutate the PAM site on the DNA double strand from TTTG to other PAM sequences and perform side chain rearrangement using the wild-type protein-DNA-RNA ternary complex as a template to generate initial mutant structures; then perform side chain rearrangement in the 4 Å range adjacent to the mutant region to quickly eliminate side chain collisions introduced by mutation. Perform energy minimization in the 6 Å range adjacent to the mutant region to fine-tune the backbone and side chain together to further reduce energy. Perform Relax optimization in the 8 Å range adjacent to the mutant region to simulate annealing optimization in a larger range to obtain the final low-energy conformation. Through mutation and structure optimization, 1536 ternary complex structures corresponding to PAM site mutations are obtained from 6 initial wild-type structures. As shown in Figure 2 Figure 6, if the AlphaFold3 predicted structure is used as the main structure, the complex structure obtained by structure optimization has no large conformational change from the structure obtained by AlphaFold. We take DNMT1_AACA as an example, the overall structure RMSD obtained by the two methods is 0.593, which can be considered that the structure optimization by Pyrosetta after base mutation does not cause drastic changes in the structure of the proteinase. And the RMSD in the 8 Å range of the local mutation region is 1, which is within an acceptable range.

[0049] 2. Geometric vector transfer graph neural network embedding processing of structure data

[0050] Geometric vector transfer graph neural network (GVP-GNN) is a graph neural network architecture specifically designed for three-dimensional geometric data. The core innovation is the introduction of geometric vector perceptron (GVP), which can handle both scalar features and vector features, and has properties such as rotation and translation. GVP-GNN performs exceptionally well in processing geometric structure data such as protein 3D structures and molecules. In the representation of protein local mutations, a hybrid geometric perception graph neural network architecture is adopted, which uses a decoupled dual-branch processing strategy: the scalar branch integrates residue type encoding and relative spatial coordinates to capture the local chemical microenvironment; the vector branch focuses on modeling normalized direction information to reflect the spatial arrangement pattern of atoms around the mutation center. By independently processing scalar (residue type + relative position) and vector (normalized direction) features and then fusing them, the embedding representation of protein local mutations is generated. The features of the two branches are integrated through a post-fusion layer. This simplified architecture balances theoretical rigor and engineering practicality by explicitly separating chemical information and geometric information, effectively extracting key features of the local mutation region while ensuring model stability and computational efficiency, avoiding the complexity and instability of strict GVP-GNN implementation.

[0051] After obtaining 1536 protein-DNA-RNA ternary complex structures, we extracted the local structure centered at the mutation site within a 8 Å radius and constructed the graph adjacency based on the 6 Å radius neighborhood. We set the edge length of the adjacency graph to 6 Å, aiming to balance the biophysical reality and computational efficiency. This distance can cover the key short-range interactions (such as hydrogen bonds and strong van der Waals forces) that determine the local conformation, while avoiding the noise introduced by long-range weak interactions, ensuring that the graph neural network can focus on the direct microenvironment of the mutation site and control the computational complexity. The node features consist of two parts: the scalar features include the concatenation of 28-dimensional residue type one-hot encoding (covering 20 standard amino acids and 8 nucleotide acids (A / T / C / G for DNA and A / U / C / G for RNA)) and 3-dimensional relative coordinates; the vector features are the normalized 3D direction vectors. These two features are processed by two independent branch encoders: the scalar branch is mapped to a 512-dimensional vector by a multi-layer perceptron (MLP), and the vector branch is mapped to a 128-dimensional vector. After concatenation at the node level, the outputs of the two branches are further integrated by a fusion layer, and finally a 512-dimensional local mutation embedding is generated by global average pooling.

[0052] This feature design realizes the decoupled encoding of chemical properties and spatial information: the 28-dimensional one-hot encoding preserves the chemical identity independence of each type of residue in the form of a sparse vector, avoiding the introduction of false sequence relationships; the 3-dimensional relative coordinates describe the Euclidean position of the atom relative to the mutation center, providing the model with local spatial conformation perception. After concatenation, both the inherent physical properties of the residues (such as hydrophobicity, charge, and volume) and their relative distance and orientation in space can be quantified, providing the graph neural network with a clear and complete representation of the local structure.

[0053] In terms of dimension setting, the scalar branch outputs 512 dimensions, providing sufficient expression ability for complex chemical-spatial patterns while avoiding over-parameterization on limited data sets due to high dimensions (such as 1024 dimensions). The vector branch outputs 128 dimensions, which not only encodes the direction information but also reflects the dominance of chemical information through a dimension ratio of about 4:1. The fusion layer compresses the 512+128-dimensional features to 512 dimensions, and this "bottleneck" structure encourages the model to focus on key features and improve information density. The final 512-dimensional embedding has good storage, retrieval, and downstream task compatibility, and its embedding dimension is consistent with that of mainstream pre-trained models, achieving a reasonable balance between model expression ability and computational efficiency.

[0054] In addition, the local structure of the mutation position is selected instead of the whole structure because only the PAM position is different in the corresponding mutant structure of each target. Embedding all parts of the structure will cause the structure signal of the PAM position to be submerged and unable to effectively distinguish the characteristics of the structure around the PAM.

[0055] 3. Cluster analysis of PAM using structure feature embedding data

[0056] The structure of a protein determines its function, and when the structure of a complex is formed, the active feature is also determined. Using the structure feature embedding data obtained by GVP-GNN processing, a machine learning model is used to analyze the PAM of Cas proteinase. After obtaining 512-dimensional local mutation embedding, in order to perform visual analysis and cluster mining, we first use principal component analysis to reduce the dimensionality. PCA is a classic unsupervised linear dimensionality reduction method, which converts the original high-dimensional features into a set of linearly uncorrelated principal components through orthogonal transformation, and determines the importance of each component according to the variance contribution rate. Here, we select the first two principal components to reduce the embedding to 2 dimensions, aiming to maximize the retention of the original data variation information while achieving intuitive two-dimensional visualization, so as to preliminarily observe the distribution and clustering trend of the mutation samples in the structure space.

[0057] Based on the two-dimensional features after dimensionality reduction, we further use Gaussian mixture model for cluster analysis. GMM is a clustering method based on probability model, which assumes that the data is generated by a mixture of multiple Gaussian distributions, and learns the parameters (mean, covariance, mixing weight) of each distribution through expectation maximization algorithm, and then provides the probability of each sample belonging to each cluster. We set 4 and 6 mixed components (i.e. cluster number) respectively to explore the natural grouping of the local structure of the mutation at different granularities: 4 categories focus on identifying macro functional or structural categories; 6 categories may further reveal more detailed subcategories. The results of 4 categories and 6 categories based on PCA and GMM are shown in Figure 3 No matter 4 clusters or 6 clusters, there is a class (cluster 0) with 64 PAMs, including TTTN with higher activity and NTTN PAM. In fact, the 64 PAMs in this cluster can be represented as NYYN (N represents A, T, G, C, and Y represents C or T), and the classic PAM of AsCas12a is TTTV, which shows that the current clustering has all the classic PAMs together, and the Cas12a protein also has recognition effect on non-classical PAMs containing C, which shows the effectiveness of the clustering.

[0058] To further explore the nonlinear structure relationship and probabilistic clustering pattern in local mutation embedding, we also tried the process combining t-distributed Stochastic Neighbor Embedding (t-SNE) and Gaussian Mixture Model (GMM). t-SNE is a nonlinear dimensionality reduction algorithm that focuses on preserving the local neighborhood structure by optimizing the probability distribution of data point similarities between high-dimensional and low-dimensional spaces, and is particularly good at revealing complex manifold and cluster distribution in high-dimensional data. Here, we reduced the 512-dimensional mutation embedding to 2-dimensional to visualize its potential nonlinear clustering pattern on a two-dimensional plane. The results are shown in Figure 4 From the results, the clustering results are consistent compared with PCA and GMM processing, but there are slight differences in distribution, while the PCA and GMM are better in distribution.

[0059] To obtain clear and separable mutation categories and perform rapid clustering division, we also tried t-distributed Stochastic Neighbor Embedding and K-means clustering method. First, we used t-SNE to nonlinearly map the 512-dimensional mutation embedding to 2-dimensional space, which expanded the originally complex cluster structure in high-dimensional space and made it visualizable by preserving the local similarity between high-dimensional samples. Then, K-means clustering was performed on the two-dimensional representation after t-SNE dimensionality reduction. K-means is a distance-based partitioning hard clustering algorithm that optimizes cluster centers and sample assignments iteratively, finally explicitly divides each mutation sample into one of the specified K clusters, forming classes that are separated from each other and compact within. This process takes into account the intuitiveness of nonlinear visualization and the high interpretability of clustering results, and is suitable for mutation pattern analysis that requires clear classification boundaries and high distinction between categories. The results are shown in Figure 5 From the clustering results, this method is consistent with t-SNE and GMM processing after dimensionality reduction for each PAM position, but due to the different clustering methods of K-means and GMM, the clustering results are significantly different, with only 48 PAMs in cluster 3 of TTTV, which is more focused and contains GTTN and TTTN PAMs. In the 6-classification, cluster 5 groups NTTN PAMs together, and the more active ones are focused together.

[0060] Using structure feature embedding data to perform PAM analysis, it can be seen that the combination of PCA and GMM works best, and the classic and non-classical PAM reported by Cas12a is gathered together. But t-SNE+GMM, and t-SNE+K-means combination will also give acceptable clustering results. Compared with the PAM obtained by wet experiment method before, the calculation method is more time-saving and labor-saving, and also reduces various biases in the experiment, and gives more objective results. Using this method, PAM analysis can be performed on completely new Cas proteinases, only the protein sequence, RNA and DNA sequence need to be given, among which the DNA target can choose the commonly used fragment, and the other part of the RNA sequence is determined by the Cas enzyme species, such as Cas12a only needs repeats sequence, and Cas9 needs tracrRNA and repeat sequence. Using this method combined with the existing prior knowledge of Cas enzyme, such as Cas9 is more likely to recognize G-rich PAM, and Cas12a recognizes T-rich PAM, the clustering results can be judged.

[0061] 4. Using experimental data to verify the clustering analysis of PAM

[0062] CN202310173577.1 discloses a method for identifying a protospacer adjacent motif, which has obtained 1536 fluorescence detection signals corresponding to 1536 structure features in this example through experiments, that is, experimental data is obtained. We normalize the experimental data and select PCA (principal component analysis) for dimension reduction. In this example, the data is reduced to 2 features, K-means is used for clustering, and finally 4 clusters are obtained, as shown in Figure 6 The second class of results, we found the classic PAM sequence TTCV of Cas12a, and the structure analysis obtained as TTCV and GTTV are also in class 2, which also shows that the PAM sequence obtained by structure prediction has certain consistency with the experimental data.

[0063] 5. Experimental identification of PAM sequence

[0064] To verify the PAMs screened, we designed a CRISPR / Cas12a-based in vitro gene cleavage activity detection experiment. In the experiment, crRNA is composed of a 21 nt region that interacts with Cas12a and a 23 nt programmable guide region for target DNA recognition, and the crRNA sequence is: UAAUUUCUACUAAGUGUAGAUCUGAUGGUCCAUGUCUGUUACUC (SEQ ID No. 3). The target DNA is a DNMT1 gene fragment, and at the PAM position in front of the target sequence, we mutate the point mutation experiment to design detection substrates containing different PAMs. In the gene editing system, once the crRNA mediates the Cas proteinase to target the target DNA molecule, the cleavage activity of the proteinase will be activated, and then the target DNA fragment is cut.

[0065] The detection system uses 400 ng Cas12a-68, 100 ng target DNA fragment, and 5 pmol of crRNA, and the final volume of the reaction is 10 μL. The reaction is incubated at 37°C for one hour, heated at 95°C for 5-10 minutes to inactivate the Cas12a proteinase, and then subjected to agarose gel electrophoresis to observe the cut target DNA molecules. The linearized gene fragment is cut by Cas to form two DNA fragments with lengths of about 4500 bp and 1500 bp. If Cas12a cannot recognize PAM, the fragment remains unchanged, with a single band with a size of about 6000 bp.

[0066] The specific experimental results are shown in Figure 7 From Figure 7 It can be seen that in the in vitro cleavage activity experiment, Cas12a can cut the double-stranded part of the target DNA fragment, and its corresponding PAM is TTCC, TCTA, GTTA, TTCA and TTTC. The above PAMs are all in the same category as the classic TTTV, and the in vitro cleavage function of Cas12a is also observed. TGGA and TGGG are rarely grouped together with TTTV, that is, the recognition ability of Cas12a for them is weak, as shown in the results, Cas12a cannot correctly recognize these PAMs in vitro. The above results also demonstrate the effectiveness of our PAM identification data based on machine learning processing.

[0067] 6. Cas9 PAM clustering based on the above method

[0068] To test the applicability of the identification method, we also tested SpCas9, whose PAM is located at the 3' end of the target sequence, which is different from Cas12a, but does not affect the analysis by this method. The classic PAM of Cas9 is NGG, but it can also recognize non-classic PAMs containing A, as shown in the literature (Genome Biol. 2015 Nov 19; 16: 253. doi: 10.1186 / s13059-015-0818-7), and its ability to recognize non-classic PAMs containing A increases when the concentration of Cas9 is increased in the reaction. According to our previous analysis of Cas12a, we simply clustered the PAM of Cas9, using a target sequence (EMX1_GGGC: GAAGGGCCTGAGTCCGAGCAGAAGAAGAAGGGCTCCCATCACATCAACCGGTGGC; SEQ ID No. 1), sgRNA (GAGUCCGAGCAGAAGAAGAAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC; SEQ ID No. 2), and Cas9 PAM is 3 bases such as NGG, the first base can be any base, and in order to consider more possibilities, we performed mutations at the PAM position of 4 bases, i.e. NNNN base combination (256), and simply used t-SNE dimensionality reduction and K-means clustering analysis, and the results are shown in Figure 8 In the 4-cluster, we see that in class 2, there are 64 PAMs whose PAM is NRRN (R represents A or G). In the 6-cluster, in class 0, there are 48 PAMs whose PAM can be summarized as: NAG (N), NGA (N) and NGG (N), which excludes NAA (N), indicating that Cas9 has weaker recognition ability for NAA relative to G-rich PAM. Moreover, selecting a PAM sequence length of 4 does not affect the results, and Cas9 has no selectivity for the fourth base. The classification of Cas9 also demonstrates the effectiveness of our analysis of Cas enzyme PAM based on local geometric structure and artificial intelligence.

[0069] In an embodiment of the present application, a PAM identification device for Cas protein based on local geometric structure and machine learning is also provided, comprising a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the PAM identification method for Cas protein based on machine learning.

[0070] For the system embodiment, since it basically corresponds to the method embodiment, the relevant part is described in the method embodiment, and the implementation method of the remaining modules is not described here. The system embodiment described above is only illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present application according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0071] The system embodiment of the present application can be applied to any device with data processing capability, which can be a device or apparatus such as a computer. The system embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logical device, it is formed by reading the corresponding computer program instructions in the non-volatile memory into the memory and running by the processor of the device with data processing capability.

[0072] The embodiment of the present application also provides a computer readable storage medium, which stores a program, and the program is executed by a processor to realize the Cas protein PAM identification method and device based on local geometry structure and machine learning.

[0073] The computer readable storage medium can be an internal storage unit of any device with data processing capability, such as a hard disk or a memory. The computer readable storage medium can also be an external storage device of any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. Further, the computer readable storage medium can include both the internal storage unit and the external storage device of any device with data processing capability. The computer readable storage medium is used to store the computer program and other programs and data required by the device with data processing capability, and can also be used to temporarily store data that has been output or will be output.

[0074] Obviously, the above-mentioned embodiments and drawings are only some examples of the present application, and for those skilled in the art, the present application can also be applied to other similar situations according to the drawings without the need for creative labor. In addition, it can be understood that although the work done in the development process here may be complex and long, but for those skilled in the art, some design, manufacture or production changes according to the technical content disclosed in the present application are only routine technical means and should not be regarded as insufficient disclosure of the present application. Without departing from the concept of the present application, a number of modifications and improvements can also be made, which are within the scope of protection of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. A method for analyzing PAM of Cas protein based on local geometry and machine learning, characterized in that, The method comprises the following steps: S1: Obtain a target Cas protein, a crRNA, and a target DNA sequence, and use a structure prediction tool to construct an initial structure of a Cas protein-dual-stranded DNA-RNA ternary complex; S2: Use the initial structure as a template, select different PAM region lengths according to the type of the Cas protein, and perform site-directed mutagenesis and structure optimization on the PAM site of the target DNA by using a calculation tool to generate a mutant ternary complex structure dataset containing different PAM sequences; S3: For each structure in the mutant ternary complex structure dataset, extract a local structure centered on the PAM mutation site; use a geometric vector perception graph neural network to extract and encode features of the local structure, and generate a structure feature embedding vector corresponding to each PAM mutation site; The local structure extracted in S3 is centered on the PAM mutation site, and the local structure within an 8 Å radius range is extracted, and a graph adjacency relationship between atoms is constructed based on a 6 Å radius neighborhood. The node features in the adjacency graph include scalar features composed of one-hot encoding of residue types and coordinates of atoms relative to the mutation center, and vector features composed of normalized direction vectors; The scalar features and the vector features are processed by two independent branch encoders respectively: the scalar branch is mapped to a 512-dimensional vector by a multilayer perceptron, and the vector branch is mapped to a 128-dimensional vector; after the two branches are spliced at the node level, the fusion layer is further integrated, and finally the global average pooling is used to generate a 512-dimensional structure feature embedding vector; S4: Input the structure feature embedding vector obtained in step S3 into a machine learning model; first perform dimension reduction processing, and then use an unsupervised clustering algorithm to perform clustering analysis on the reduced data to obtain natural grouping of different PAM sequences, and identify the PAM sequence clustering results combined with prior biological knowledge of the PAM sequence; obtain the PAM sequence grouping that can be recognized by the target Cas protein.

2. The method of claim 1, wherein, The PAM region length is selected as n in the step S2, the n value is determined according to the type of Cas proteinase, n = 2-8; the number of ternary complex structures in the data set is 4. n ​ 3. The method of claim 1, wherein, In step S2, the PyRosetta tool is used for site-directed mutagenesis of the PAM site, and side chain rearrangement and energy optimization in the adjacent range of the mutation region are performed to generate the mutant ternary complex structure dataset.

4. The method of claim 1, wherein, In step S2, the PyRosetta tool is used for site-directed mutagenesis of the PAM site, and side chain rearrangement and energy optimization in the adjacent range of the mutation region are performed to generate the mutant ternary complex structure dataset.

5. The method of claim 1, wherein, The scalar features are obtained by splicing 28-dimensional residue type one-hot encoding and 3-dimensional relative coordinates, and the 28-dimensional residue type one-hot encoding covers 20 standard amino acids and 8 nucleic acid nucleotides, and the 8 nucleic acid nucleotides are A / T / C / G of DNA and A / U / C / G of RNA.

6. The method of claim 1, wherein, The dimension reduction processing in the step S3 is principal component analysis or t-distribution stochastic neighbor embedding; and the unsupervised clustering algorithm is Gaussian mixture model or K-means clustering algorithm.

7. The method of claim 1, wherein, The Cas proteinase is Cas9, Cas12, Cas13 proteinase and Cascade complex of type I in CRISPR-Cas category 1; and the target DNA sequence contains at least PAM sequence and Protospacer sequence part.

8. A device for identifying PAM of Cas protein based on local geometry and machine learning, characterized in that, The method comprises the following steps: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the Cas proteinase PAM analysis method based on local geometry structure and machine learning as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • A method for identifying a protospacer adjacent motif

    CN116083536B

  • Optimum design method of structure-based CRISPR protein

    CN111893104A

  • Protein-DNA binding site prediction method based on multi-view graph embedding fusion

    CN118824353A

  • Method for identifying effectiveness of Cas target spot by combining gene large model and ML

    CN120636549A