Cas protein recognition method and system based on protein language model and multi-scale convolutional neural network
Through the method based on the protein language model and multi-scale convolutional neural network, Cas proteins are recognized, which solves the problem of time-consuming and low accuracy of traditional methods, and achieves fast and accurate Cas protein recognition, reducing the recognition cost.
Patent Information
- Application Number
- CN202510140713.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-08
AI Technical Summary
Traditional Cas protein recognition methods take time and are limited in recognition accuracy, making it difficult to achieve high-throughput applications, and end-to-end automated feature extraction and classification prediction, which increases the complexity of practical applications.
The Cas protein recognition method based on protein language model and multi-scale convolutional neural network is adopted to encode protein sequences through AATP feature channel and ESM-2 feature channel to extract features, and feature compression and deep extraction are used for neural network full-connection layer and multi-scale convolutional neural network, and finally feature fusion and model training are carried out.
Fast and accurate Cas protein recognition is achieved, reducing recognition costs, simplifying the application process, and efficiently screening out possible Cas proteins.
Smart Images

Figure CN120072053A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics technology, and particularly to a method and system for identifying Cas proteins based on a protein language model and a multi-scale convolutional neural network. Background Art
[0002] Cas proteins are key components in the CRISPR / Cas system. The Cas system is a natural adaptive immune system that widely exists in prokaryotes. The CRISPR / Cas system consists of CRISPR sequences and multiple Cas proteins, and the Cas proteins are responsible for performing specific biological functions. The identification of Cas proteins mainly relies on traditional methods based on sequence homology, such as BLAST and HMMER. These methods identify Cas proteins by detecting conserved regions similar to known sequences and are suitable for analyzing proteins closely related to sequences in the database. In addition to relying on traditional methods based on sequence homology, the remaining traditional machine learning methods still rely on manually designed features, such as CASPredict, CRISPRCasStack, and CASPredict-SVM, and it is difficult to fully mine the deep features in the sequences.
[0003] However, for Cas proteins or novel Cas proteins found in bacteria and archaea with insufficient research, these methods show obvious limitations. They are time-consuming, have limited identification accuracy, and are difficult to achieve high-throughput applications. Moreover, methods such as CASPredict, CRISPRCasStack, and CASPredict-SVM for identifying Cas proteins cannot achieve end-to-end automated feature extraction and classification prediction, increasing the complexity of practical applications.
[0004] Therefore, traditional Cas protein identification methods often have problems of low identification efficiency and high identification cost. Summary of the Invention
[0005] Based on this, in order to solve the above technical problems, a method and system for identifying Cas proteins based on a protein language model and a multi-scale convolutional neural network are provided, which can simply and quickly screen out possible Cas proteins with high efficiency and low cost.
[0006] A method for identifying Cas proteins based on a protein language model and a multi-scale convolutional neural network, the method comprising:
[0007] Obtain different collected Cas protein datasets, and divide the Cas protein datasets into two target datasets, and the target datasets contain protein sequences;
[0008] Determine the AATP feature channel and the ESM-2 feature channel. Encode the input protein sequence through the AATP feature channel to characterize the protein sequence as AATP features; encode the input protein sequence through the ESM-2 feature channel to extract the protein embedding features in the protein sequence;
[0009] Use the fully connected layer of the neural network to compress the AATP features to obtain the compressed AATP features; and use the multi-scale convolutional neural network to deeply extract the protein embedding features to obtain ESM-2 features;
[0010] Fuse the compressed AATP features and ESM-2 features to obtain the fused features; train the Cas protein recognition model based on the fused features, and perform Cas protein recognition through the trained Cas protein recognition model.
[0011] In one embodiment, obtain different Cas protein datasets collected, and divide the Cas protein datasets into two target datasets, including:
[0012] Obtain different Cas protein datasets collected, and divide the Cas protein datasets into the CRISPRCasStack dataset and the CASPredict dataset;
[0013] The protein sequences in the CRISPRCasStack dataset are clustered through a first similarity threshold;
[0014] The protein sequences in the CASPredict dataset are clustered through a second similarity threshold; wherein, the CRISPRCasStack dataset is used for hyperparameter optimization of the model, and the CASPredict dataset and the CRISPRCasStack dataset are jointly used to evaluate the performance of the model under different training conditions.
[0015] In one embodiment, encoding the input protein sequence through the AATP feature channel to characterize the protein sequence as AATP features includes:
[0016] Capture the conservation and evolutionary information at different positions in the protein sequence through the position-specific scoring matrix in the AATP feature channel;
[0017] Based on the values in the position-specific scoring matrix, extract the preference information of different types of amino acids in the protein sequence;
[0018] Extract the AATP features from the position-specific scoring matrix according to the preference information, conservativeness, and evolutionary information and output them.
[0019] In one embodiment, the input protein sequence is encoded through the ESM-2 feature channel to extract the protein embedding features in the protein sequence, including:
[0020] Input the protein sequence into the ESM-2 feature channel, where a protein language model ESM-2 is set in the ESM-2 feature channel;
[0021] Extract the protein embedding features from the protein sequence through the protein language model.
[0022] In one embodiment, use a neural network fully connected layer to compress the AATP features to obtain the compressed AATP features, including:
[0023] Use a two-layer neural network fully connected layer to extract the AATP features;
[0024] Perform dimensionality reduction on the AATP features through the neural network fully connected layer to obtain the compressed AATP features.
[0025] In one embodiment, use a multi-scale convolutional neural network to deeply extract the protein embedding features to obtain ESM-2 features, including:
[0026] Input the protein embedding features into the multi-scale convolutional neural network, and deeply extract the protein embedding features through different convolutional scales in the multi-scale convolutional neural network to obtain deep features;
[0027] Perform compression and dimensionality reduction on the deep features to obtain ESM-2 features.
[0028] In one embodiment, fuse the compressed AATP features and ESM-2 features to obtain fused features, including:
[0029] Convert the feature dimensions of the compressed AATP features and the ESM-2 features to the same feature dimension;
[0030] Perform normalization on the compressed AATP features and ESM-2 features after feature dimension conversion to obtain target AATP features and target ESM-2 features;
[0031] Fuse the target AATP features and target ESM-2 features according to the feature dimension to obtain fused features.
[0032] In one of the embodiments, training the Cas protein recognition model based on the fusion feature includes:
[0033] Input the fusion feature into the Cas protein recognition model, and calculate the loss based on the target data set using the cross-entropy loss function and the backpropagation strategy to determine the loss function;
[0034] Use the five-fold cross-validation method to verify the Cas protein recognition model to obtain the verification result; determine the hyperparameters of the Cas protein recognition model according to the verification result;
[0035] Based on the loss function, update the parameters of the Cas protein recognition model, and complete the training of the Cas protein recognition model in combination with the hyperparameters.
[0036] A Cas protein recognition system based on a protein language model and a multi-scale convolutional neural network, the system includes:
[0037] A data collection module for obtaining different Cas protein data sets collected and dividing the Cas protein data set into two target data sets, and the target data sets contain protein sequences;
[0038] A protein sequence characterization module for determining the AATP feature channel and the ESM-2 feature channel, encoding the input protein sequence through the AATP feature channel, and characterizing the protein sequence as AATP features; encoding the input protein sequence through the ESM-2 feature channel, and extracting the protein embedding features in the protein sequence;
[0039] A feature extraction module for using a neural network fully connected layer to compress the AATP features to obtain compressed AATP features; and using a multi-scale convolutional neural network to deeply extract the protein embedding features to obtain ESM-2 features;
[0040] A feature fusion and model training module for fusing the compressed AATP features and the ESM-2 features to obtain fusion features; training the Cas protein recognition model based on the fusion features, and performing Cas protein recognition through the trained Cas protein recognition model.
[0041] The above-mentioned Cas protein recognition method and system based on a protein language model and a multi-scale convolutional neural network encode the input protein sequence through two feature channels respectively, extract two features and then perform compression, dimensionality reduction and fusion processing, so that the trained Cas protein recognition model can quickly and simply screen out possible Cas proteins, with high efficiency and low cost. Description of the Drawings
[0042] Figure 1 It is an application environment diagram of the Cas protein recognition method based on a protein language model and a multi-scale convolutional neural network in an embodiment;
[0043] Figure 2 It is a schematic flowchart of the Cas protein recognition method based on a protein language model and a multi-scale convolutional neural network in an embodiment;
[0044] Figure 3 It is a structural block diagram of the Cas protein recognition system based on a protein language model and a multi-scale convolutional neural network in an embodiment;
[0045] Figure 4 It is a schematic diagram of the structural framework of the Cas protein recognition based on a protein language model and a multi-scale convolutional neural network in an embodiment;
[0046] Figure 5 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0047] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0048] It can be understood that the terms "first", "second", etc. used in the present application can be used in this article to describe similarity thresholds, but these similarity thresholds are not limited by these terms. These terms are only used to distinguish the first similarity threshold from another similarity threshold. For example, without departing from the scope of the present application, the first similarity threshold can be called the second similarity threshold, and similarly, the second similarity threshold can be called the first similarity threshold. Both the first similarity threshold and the second similarity threshold are similarity thresholds, but they are not the same similarity threshold.
[0049] The Cas protein recognition method based on a protein language model and a multi-scale convolutional neural network provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. As Figure 1As shown, the application environment includes a computer device 110. The computer device 110 can obtain different collected Cas protein data sets, divide the Cas protein data sets into two target data sets, and the target data sets contain protein sequences; the computer device 110 can determine the AATP feature channel and the ESM-2 feature channel, encode the input protein sequence through the AATP feature channel, and represent the protein sequence as AATP features; encode the input protein sequence through the ESM-2 feature channel to extract the protein embedding features in the protein sequence; the computer device 110 can use the fully connected layer of the neural network to compress the AATP features to obtain the compressed AATP features; and use the multi-scale convolutional neural network to deeply extract the protein embedding features to obtain the ESM-2 features; the computer device 110 can fuse the compressed AATP features and the ESM-2 features to obtain the fused features; train the Cas protein recognition model based on the fused features, and perform Cas protein recognition through the trained Cas protein recognition model. Among them, the computer device 110 can be but is not limited to various personal computers, laptop computers, smart phones, robots, tablet computers and other devices.
[0050] In one embodiment, as Figure 2 shown, a Cas protein recognition method based on a protein language model and a multi-scale convolutional neural network is provided, including the following steps:
[0051] Step 202, obtain different collected Cas protein data sets, divide the Cas protein data sets into two target data sets, and the target data sets contain protein sequences.
[0052] Among them, the different Cas protein data sets can be collected from previous studies. Specifically, the different collected Cas protein data sets are two different Cas protein data sets collected from two previous studies, namely the CRISPRCasStack data set and the CASPredict data set.
[0053] In one embodiment, a method for identifying Cas proteins based on a protein language model and a multi-scale convolutional neural network may further include the process of collecting a data set. The specific process includes: obtaining different Cas protein data sets collected, and dividing the Cas protein data sets into a CRISPRCasStack data set and a CASPredict data set; clustering the protein sequences in the CRISPRCasStack data set through a first similarity threshold; clustering the protein sequences in the CASPredict data set through a second similarity threshold; wherein, the CRISPRCasStack data set is used for hyperparameter optimization of the model, and the CASPredict data set and the CRISPRCasStack data set are jointly used to evaluate the performance of the model under different training conditions.
[0054] Among them, the CRISPRCasStack data set contains 209 positive samples and manually screened non-Cas protein negative samples. The protein sequence clustering process is carried out through the CD-HIT tool, and the first similarity threshold is set to 70% to ensure low similarity of the sequences and exclude non-standard amino acid sequences. Finally, 150 positive samples and 150 negative samples are obtained for the training set, and 59 positive samples and 59 negative samples are obtained for the test set.
[0055] The CASPredict data set is based on the UniProt database. Initially, 293 Cas protein sequences are extracted. After manually checking and removing the sequences containing ambiguous residues, further screening is carried out through the CD-HIT tool with a second similarity threshold of 30%. Finally, 155 high-quality positive samples are retained. At the same time, non-Cas proteins are manually screened from the UniProt database as negative samples, and CD-HIT is used to remove redundancy with a second similarity threshold of 40% to ensure the diversity and length distribution matching of the negative samples. Finally, 155 positive samples and 155 negative samples are obtained as the training set, and 64 positive samples and 64 negative samples are obtained as the test set.
[0056] In this embodiment, the CRISPRCasStack data set is mainly used for hyperparameter optimization of the model, and the CASPredict data set and the CRISPRCasStack data set are jointly used to comprehensively evaluate the performance of the model under different training conditions. These data sets have been strictly processed and screened, providing high-quality and well-balanced training and test data.
[0057] Step 204, determine the AATP feature channel and the ESM-2 feature channel, encode the input protein sequence through the AATP feature channel, and characterize the protein sequence as AATP features; encode the input protein sequence through the ESM-2 feature channel, and extract the protein embedding features in the protein sequence.
[0058] To comprehensively capture the functional and evolutionary information of protein sequences, a dual-channel feature representation method is adopted in this embodiment, and the input protein sequences are encoded through the AATP feature channel and the ESM-2 feature channel respectively.
[0059] In one embodiment, a Cas protein recognition method based on a protein language model and a multi-scale convolutional neural network may further include the process of extracting AATP features. The specific process includes: capturing the conservation and evolutionary information at different positions in the protein sequence through the position-specific scoring matrix in the AATP feature channel; extracting the preference information of different types of amino acids in the protein sequence based on the values in the position-specific scoring matrix; and extracting and outputting AATP features from the position-specific scoring matrix according to the preference information, conservation and evolutionary information.
[0060] Among them, the position-specific scoring matrix (PSSM) is a matrix representing the evolutionary information of protein sequences. PSSM captures the conservation and evolutionary information at different positions within the sequence, and AATP features can be obtained from the PSSM through different calculation methods. The core concept of AATP is to analyze and summarize the values in the PSSM matrix to extract the preference information of different types of amino acids in the sequence. By statistically capturing the trend of specific amino acid occurrences, AATP reduces data redundancy while retaining the key sequence information related to amino acid preferences. In the AATP feature channel, the protein sequence to be recognized is characterized as 420-dimensional AATP features.
[0061] In one embodiment, a Cas protein recognition method based on a protein language model and a multi-scale convolutional neural network may further include the process of extracting protein embedding features. The specific process includes: inputting the protein sequence into the ESM-2 feature channel, where the protein language model ESM-2 is set; and extracting protein embedding features from the protein sequence through the protein language model.
[0062] In the ESM-2 feature channel, the protein sequence generates 320-dimensional ESM-2 features through the ESM-2 (Evolutionary Scale Modeling2) protein large language model. ESM-2 is a large-scale pre-trained model based on the Transformer architecture, which is used to extract deep embedding features from protein sequences. In bioinformatics, protein sequences and text data have structural similarities, so the representation method based on the language model can be directly applied to protein sequences, and the pre-trained model can learn semantic information from large-scale protein sequences and generate rich protein embeddings.
[0063] Step 206: Use the fully connected layer of the neural network to compress the AATP features to obtain the compressed AATP features; and use the multi-scale convolutional neural network to deeply extract the protein embedding features to obtain the ESM-2 features.
[0064] Among them, different feature extraction strategies are adopted for the two feature channels, and different feature processing methods are used for feature processing.
[0065] In one embodiment, a method for identifying Cas proteins based on a protein language model and a multi-scale convolutional neural network may further include a process of processing the AATP features. The specific process includes: using a two-layer fully connected layer of the neural network to extract the AATP features; performing dimensionality reduction processing on the AATP features through the fully connected layer of the neural network to obtain the compressed AATP features.
[0066] For the generated 420-dimensional AATP features, the AATP features use a two-layer fully connected layer of the neural network to further extract and reduce the dimensions of the AATP features, and compress the AATP features to 64 dimensions to obtain the compressed AATP features.
[0067] In one embodiment, a method for identifying Cas proteins based on a protein language model and a multi-scale convolutional neural network may further include a process of extracting protein embedding features. The specific process includes: inputting the protein embedding features into the multi-scale convolutional neural network, and deeply extracting the protein embedding features through different convolution scales in the multi-scale convolutional neural network to obtain deep features; performing compression and dimensionality reduction processing on the deep features to obtain the ESM-2 features.
[0068] For the 320-dimensional protein embedding features, the multi-scale convolutional neural network can be used to deeply extract the features of the protein embedding features through different convolution scales (9, 11, 13), and finally convert them into 256-dimensional features, that is, the ESM-2 features.
[0069] Step 208: Fuse the compressed AATP features and ESM-2 features to obtain the fused features; train the Cas protein recognition model based on the fused features, and perform Cas protein recognition through the trained Cas protein recognition model.
[0070] For the features of the two channels after feature extraction, feature fusion is performed, merged into 300-dimensional fused features, and then the features are input into a two-layer fully connected layer for prediction. Finally, the prediction results of different protein sequences can be obtained.
[0071] Specifically, in one embodiment, a method for identifying Cas proteins based on a protein language model and a multi-scale convolutional neural network may further include a feature fusion process. The specific process includes: converting the feature dimensions of the compressed AATP features and the ESM-2 features into the same feature dimension; performing normalization processing on the compressed AATP features and the ESM-2 features after the feature dimension conversion to obtain target AATP features and target ESM-2 features; and fusing the target AATP features and the target ESM-2 features according to the feature dimension to obtain fused features.
[0072] In one embodiment, a method for identifying Cas proteins based on a protein language model and a multi-scale convolutional neural network may further include a model training process. The specific process includes: inputting the fused features into the Cas protein identification model, calculating the loss based on the target dataset using the cross-entropy loss function and the backpropagation strategy, and determining the loss function; validating the Cas protein identification model using the five-fold cross-validation method to obtain a validation result; determining the hyperparameters of the Cas protein identification model according to the validation result; updating the parameters of the Cas protein identification model based on the loss function, and completing the training of the Cas protein identification model in combination with the hyperparameters.
[0073] In this embodiment, based on the target dataset, the five-fold cross-validation method is used, combined with the cross-entropy loss function and the backpropagation strategy to update the function to calculate the loss. The hyperparameters of the Cas protein identification model are determined by observing the performance of the validation set divided in the five-fold cross-validation. Then, the complete model training is performed through the training set of the target dataset, the model parameters are updated, and finally the Cas protein identification model is determined. After that, the Cas protein identification model is set to the validation mode and no longer participates in the backpropagation. Different protein sequences are input to complete the prediction of Cas proteins.
[0074] In one embodiment, the same test set is used on two target datasets respectively, and the effects obtained on different models are compared as follows:
[0075] On the CRISPRCasStack dataset, using the same test set, comparing the HMMCAS, CASPredict, and CRISPRCasStack models, the obtained effect data is shown in the following table:
[0076]
[0077] On the CASPredict dataset, using the same test set, comparing the CASPredict-SVM model and our Cas protein identification model, the comparison results are shown as follows:
[0078]
[0079] Benefiting from the pre-training characteristics of the protein language model ESM-2, the multi-scale feature learning ability of the multi-scale convolutional neural network MSCNN, and the performance ability of the generated amino acid type features AATP in terms of Cas proteins, the Cas protein recognition model in this application shows significant recognition advantages when recognizing Cas proteins.
[0080] It should be understood that although the steps in the above flowcharts are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0081] In one embodiment, as Figure 3 shown, a Cas protein recognition system based on a protein language model and a multi-scale convolutional neural network is provided, including: a data collection module 310, a protein sequence characterization module 320, a feature extraction module 330, and a feature fusion and model training module 340, where:
[0082] The data collection module 310 is used to obtain different Cas protein data sets collected and divide the Cas protein data set into two target data sets, and the target data sets contain protein sequences;
[0083] The protein sequence characterization module 320 is used to determine the AATP feature channel and the ESM-2 feature channel, encode the input protein sequence through the AATP feature channel, and characterize the protein sequence as AATP features; encode the input protein sequence through the ESM-2 feature channel, and extract the protein embedding features in the protein sequence;
[0084] The feature extraction module 330 is used to compress the AATP features using the fully connected layer of the neural network to obtain the compressed AATP features; and deeply extract the protein embedding features using the multi-scale convolutional neural network to obtain the ESM-2 features;
[0085] The feature fusion and model training module 340 is used to fuse the compressed AATP features and ESM-2 features to obtain fused features; train the Cas protein recognition model based on the fused features, and recognize Cas proteins through the trained Cas protein recognition model.
[0086] In one embodiment, a Cas protein recognition system based on a protein language model and a multi-scale convolutional neural network can be applied in a structural framework as Figure 4 shown. As Figure 4 shown, this structural framework may include A, the Cas protein recognition model structure (Overview of CasPro-ESM2 Model); B, protein language model and multi-scale convolutional neural network feature extraction (ESM-2MSCNN Feature Extraction); C, model interpretability (Model Interpretability); D, web server architecture (Web Server Architecture). The protein sequence is encoded through the AATP feature channel to represent the protein sequence as AATP features; encoded through the ESM-2 feature channel to extract the protein embedding features in the protein sequence; the AATP features are compressed using the fully connected layer of the neural network to obtain the compressed AATP features; and the multi-scale convolutional neural network is used to deeply extract the protein embedding features to obtain ESM-2 features; the compressed AATP features and ESM-2 features are fused to obtain fused features; the Cas protein recognition model is trained based on the fused features to obtain the trained Cas protein recognition model.
[0087] In one embodiment, the data collection module 310 is further configured to obtain different Cas protein data sets collected, and divide the Cas protein data sets into CRISPRCasStack data sets and CASPredict data sets; the protein sequences in the CRISPRCasStack data sets are clustered through a first similarity threshold; the protein sequences in the CASPredict data sets are clustered through a second similarity threshold; wherein, the CRISPRCasStack data set is used for hyperparameter optimization of the model, and the CASPredict data set and the CRISPRCasStack data set are jointly used to evaluate the performance of the model under different training conditions.
[0088] In one embodiment, the protein sequence characterization module 320 is further configured to capture the conservation and evolutionary information at different positions in the protein sequence through the position-specific scoring matrix in the AATP feature channel; extract the preference information of different types of amino acids in the protein sequence based on the values in the position-specific scoring matrix; and extract the AATP features from the position-specific scoring matrix according to the preference information, conservation and evolutionary information and output them.
[0089] In one embodiment, the protein sequence characterization module 320 is further configured to input the protein sequence into the ESM-2 feature channel, where the protein language model ESM-2 is set in the ESM-2 feature channel; and extract the protein embedding features from the protein sequence through the protein language model.
[0090] In one embodiment, the feature extraction module 330 is further configured to extract the AATP features using a two-layer neural network fully connected layer; and perform dimensionality reduction processing on the AATP features through the neural network fully connected layer to obtain the compressed AATP features.
[0091] In one embodiment, the feature extraction module 330 is further configured to input the protein embedding features into a multi-scale convolutional neural network, and deeply extract the protein embedding features through different convolutional scales in the multi-scale convolutional neural network to obtain deep features; and perform compression and dimensionality reduction processing on the deep features to obtain the ESM-2 features.
[0092] In one embodiment, the feature fusion and model training module 340 is further configured to convert the feature dimensions of the compressed AATP features and the ESM-2 features into the same feature dimension; perform normalization processing on the compressed AATP features and the ESM-2 features after the feature dimension conversion to obtain the target AATP features and the target ESM-2 features; and fuse the target AATP features and the target ESM-2 features according to the feature dimension to obtain the fused features.
[0093] In one embodiment, the feature fusion and model training module 340 is further configured to input the fused features into the Cas protein recognition model, calculate the loss based on the target dataset using the cross-entropy loss function and the backpropagation strategy to determine the loss function; verify the Cas protein recognition model using the five-fold cross-validation method to obtain the verification result; determine the hyperparameters of the Cas protein recognition model according to the verification result; update the parameters of the Cas protein recognition model based on the loss function, and complete the training of the Cas protein recognition model in combination with the hyperparameters.
[0094] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 5As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a Cas protein recognition method based on a protein language model and a multi-scale convolutional neural network. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0095] Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0096] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it realizes the steps of a Cas protein recognition method based on a protein language model and a multi-scale convolutional neural network.
[0097] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it realizes the steps of a Cas protein recognition method based on a protein language model and a multi-scale convolutional neural network.
[0098] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0099] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0100] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A Cas protein recognition method based on a protein language model and a multi-scale convolutional neural network, characterized in that: The method comprises: Acquire different collected Cas protein data sets, and divide the Cas protein data sets into two target data sets, wherein the target data sets contain protein sequences; Determine an AATP feature channel and an ESM-2 feature channel, encode the input protein sequence through the AATP feature channel, and characterize the protein sequence as an AATP feature; encode the input protein sequence through the ESM-2 feature channel, and extract protein embedding features in the protein sequence; Using a neural network fully connected layer to compress the AATP feature to obtain a compressed AATP feature; and using a multi-scale convolutional neural network to deeply extract the protein embedding feature to obtain an ESM-2 feature; The compressed AATP feature and ESM-2 feature are fused to obtain a fusion feature; a Cas protein recognition model is trained based on the fusion feature, and Cas protein recognition is performed using the trained Cas protein recognition model.
2. The Cas protein recognition method based on protein language model and multi-scale convolutional neural network according to claim 1, characterized in that: Obtain the collected different Cas protein data sets, and divide the Cas protein data sets into two target data sets, including: Obtaining different collected Cas protein data sets, and dividing the Cas protein data sets into a CRISPRCasStack data set and a CASPredict data set; The protein sequences in the CRISPRCasStack data set are clustered according to a first similarity threshold; The protein sequences in the CASPredict data set are clustered according to a second similarity threshold; Among them, the CRISPRCasStack dataset is used for hyperparameter optimization of the model, and the CASPredict dataset and the CRISPRCasStack dataset are used together to evaluate the performance of the model under different training conditions.
3. The Cas protein recognition method based on protein language model and multi-scale convolutional neural network according to claim 1, characterized in that: The input protein sequence is encoded by the AATP feature channel, and the protein sequence is characterized as an AATP feature, including: Capturing conservation and evolutionary information at different positions in the protein sequence through a position-specific scoring matrix in the AATP feature channel; extracting preference information of different types of amino acids in the protein sequence based on the values in the position-specific scoring matrix; AATP features are extracted from the position-specific scoring matrix according to the preference information, conservation and evolution information and output.
4. The Cas protein recognition method based on protein language model and multi-scale convolutional neural network according to claim 1, characterized in that: The input protein sequence is encoded through the ESM-2 feature channel to extract the protein embedding features in the protein sequence, including: Inputting the protein sequence into the ESM-2 feature channel, wherein the ESM-2 feature channel is provided with a protein language model ESM-2; Protein embedding features are extracted from the protein sequence by using the protein language model.
5. The Cas protein recognition method based on protein language model and multi-scale convolutional neural network according to claim 1, characterized in that: The AATP features are compressed using a neural network fully connected layer to obtain compressed AATP features, including: The AATP features are extracted using a two-layer neural network fully connected layer; The AATP feature is subjected to dimensionality reduction processing through the neural network fully connected layer to obtain a compressed AATP feature.
6. The Cas protein recognition method based on protein language model and multi-scale convolutional neural network according to claim 1, characterized in that: The protein embedding features are deeply extracted using a multi-scale convolutional neural network to obtain ESM-2 features, including: Inputting the protein embedding features into a multi-scale convolutional neural network, performing deep extraction on the protein embedding features through different convolution scales in the multi-scale convolutional neural network to obtain deep features; The deep features are compressed and dimensionally reduced to obtain ESM-2 features.
7. The Cas protein recognition method based on protein language model and multi-scale convolutional neural network according to claim 1, characterized in that: The compressed AATP features and ESM-2 features are fused to obtain fused features, including: Converting the feature dimension of the compressed AATP feature and the feature dimension of the ESM-2 feature to the same feature dimension; The compressed AATP features and ESM-2 features after feature dimension conversion are standardized to obtain target AATP features and target ESM-2 features; The target AATP feature and the target ESM-2 feature are fused according to feature dimensions to obtain fused features.
8. The Cas protein recognition method based on protein language model and multi-scale convolutional neural network according to claim 1, characterized in that: Training the Cas protein recognition model based on the fusion feature includes: Inputting the fusion feature into the Cas protein recognition model, calculating the loss based on the target data set using the cross-over loss function and the directional propagation strategy, and determining the loss function; The Cas protein recognition model is verified using a five-fold cross-validation method to obtain a verification result; and the hyperparameters of the Cas protein recognition model are determined according to the verification result; Based on the loss function, the parameters of the Cas protein recognition model are updated, and the Cas protein recognition model is completed in combination with the hyperparameter training.
9. A Cas protein recognition system based on protein language model and multi-scale convolutional neural network, characterized in that: The system comprises: A data collection module, used to obtain different collected Cas protein data sets, and divide the Cas protein data sets into two target data sets, and the target data sets contain protein sequences; A protein sequence characterization module is used to determine an AATP feature channel and an ESM-2 feature channel, encode the input protein sequence through the AATP feature channel, and characterize the protein sequence as an AATP feature; encode the input protein sequence through the ESM-2 feature channel, and extract protein embedded features in the protein sequence; A feature extraction module is used to compress the AATP feature using a neural network fully connected layer to obtain a compressed AATP feature; and to deeply extract the protein embedding feature using a multi-scale convolutional neural network to obtain an ESM-2 feature; The feature fusion and model training module is used to fuse the compressed AATP features and ESM-2 features to obtain fusion features; train the Cas protein recognition model based on the fusion features, and perform Cas protein recognition through the trained Cas protein recognition model.
Citation Information
Patent Citations
MHC identification method and system based on protein language model
CN118280444A
DNA binding protein and RNA binding protein classification method based on ESM-2 and dual-path neural network
CN119229982A
Machine learning systems and methods for deep learning of genomic contexts
US20240312558A1
Cited By
Method for identifying effectiveness of Cas target spot by combining gene large model and ML
CN120636549A
Method for identifying effectiveness of cas target points by combining gene large model and ml
CN120636549B