Protein Sequence Analysis Method, Apparatus, Device, Medium and Product

By grouping and feature extraction of ion channel protein sequences based on artificial intelligence, the problem of the inability to systematically analyze ion channel protein responses to physical field stimulation in the prior art is solved, and efficient sequence feature functional fragment analysis is achieved, which improves the efficiency of protein sequence analysis.

CN118571322BActive Publication Date: 2025-07-29TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410636883.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-21
Publication Date
2025-07-29
Estimated Expiration
2044-05-21

AI Technical Summary

Technical Problem

The prior art cannot systematically and efficiently analyze the response sequence characteristics and gating mechanism of ion channel proteins to physical field stimuli, resulting in limitations in the extraction of protein evolution information characteristics.

Method used

Using an artificial intelligence-based method, protein sequences are grouped and featured extracted through a pre-trained physical field stimulus response functional prediction model, and a protein language model and a multi-layer perceptron network are used to combine K-space amino acid pair analysis algorithms and evolutionary trees to identify functional fragments that respond to physical stimuli.

Benefits of technology

It has achieved rapid and efficient analysis of potential sequence characteristic functional fragments in protein sequences, improved the efficiency of protein sequence analysis, promoted the research on ion sensory physics mechanism, and provided a data basis for the transformation and design of ion channel proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118571322B_ABST
    Figure CN118571322B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a protein sequence analysis method, device, equipment, medium and product. Among them, the method includes: inputting the protein sequence vectors corresponding to each protein sequence in the protein sequence set to be analyzed into a pre-trained physical field stimulation response function prediction model to obtain a prediction result on whether each protein sequence has the function of responding to a preset physical stimulus; grouping the protein sequences in the protein sequence set to be analyzed according to the prediction result to obtain a first protein sequence group having the function of responding to the preset physical stimulus and a second protein sequence group not having the function of responding to the preset physical stimulus; respectively performing sequence feature extraction and sequence feature analysis on the first protein sequence group and the second protein sequence group to obtain target characteristic protein sequence fragments. The technical solution of the embodiment of the present invention can realize the rapid and efficient mining analysis and prediction of protein sequence fragments with specific functions in protein sequences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of bioengineering technology, and in particular to a protein sequence analysis method, device, equipment, medium and product. Background Art

[0002] Currently, research on ion channel proteins is limited to the study of a single channel protein or a single family. For example, machine learning algorithms are used to align the sequences of a single family, analyze the sequence conservation of a family, and conduct mutation experiments based on the conservation.

[0003] However, this approach cannot systematically explore the response sequence characteristics and gating mechanisms of ion channels to physical field stimuli, and cannot fully extract the characteristics of protein evolutionary information, which has certain limitations in the study of ion channel proteins. Summary of the invention

[0004] The embodiments of the present invention provide a protein sequence analysis method, apparatus, device, medium and product, which can analyze and mine potential sequence feature functional fragments in protein sequences in an artificial intelligence-based manner, thereby improving the efficiency of protein sequence analysis.

[0005] In a first aspect, an embodiment of the present invention provides a protein sequence analysis method, the method comprising:

[0006] Inputting the protein sequence vector corresponding to each protein sequence in the set of protein sequences to be analyzed into a pre-trained physical field stimulus response function prediction model to obtain a prediction result of whether each protein sequence has the function of responding to a preset physical stimulus;

[0007] Grouping the protein sequences in the set of protein sequences to be analyzed according to the prediction results to obtain a first protein sequence group having a function of responding to the preset physical stimulus and a second protein sequence group having no function of responding to the preset physical stimulus;

[0008] Sequence feature extraction and sequence feature analysis are performed on the first protein sequence group and the second protein sequence group, respectively, to obtain target characteristic protein sequence fragments.

[0009] In a second aspect, an embodiment of the present invention provides a protein sequence analysis device, the device comprising:

[0010] A protein sequence function prediction module is used to input the protein sequence vector corresponding to each protein sequence in the protein sequence set to be analyzed into a pre-trained physical field stimulus response function prediction model to obtain a prediction result of whether each protein sequence has the function of responding to a preset physical stimulus;

[0011] A protein sequence grouping module, configured to group the protein sequences in the protein sequence set to be analyzed according to the prediction result, so as to obtain a first protein sequence group with the function of responding to the preset physical stimulus and a second protein sequence group without the function of responding to the preset physical stimulus;

[0012] A protein sequence feature extraction module, configured to perform sequence feature extraction and sequence feature analysis on the first protein sequence group and the second protein sequence group respectively, so as to obtain target feature protein sequence fragments.

[0013] In a third aspect, an embodiment of the present invention further provides a computer device, where the computer device includes:

[0014] One or more processors;

[0015] A memory, configured to store one or more programs;

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the protein sequence analysis method provided in any embodiment of the present invention.

[0017] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the protein sequence analysis method provided in any embodiment of the present invention.

[0018] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the protein sequence analysis method provided in any embodiment of the present invention.

[0019] The embodiments in the above-mentioned invention have the following advantages or beneficial effects:

[0020] In an embodiment of the present invention, by inputting the protein sequence vectors corresponding to each protein sequence in the protein sequence set to be analyzed into a pre-trained physical field stimulation response function prediction model, a prediction result of whether each protein sequence has the function of responding to a preset physical stimulus can be obtained, and the protein sequences can be preliminarily predicted and analyzed quickly and efficiently; according to the prediction results, the protein sequences in the protein sequence set to be analyzed are grouped to obtain a first protein sequence group having the function of responding to the preset physical stimulus and a second protein sequence group not having the function of responding to the preset physical stimulus; for the first protein sequence group and the second protein sequence group, sequence feature extraction and sequence feature analysis are respectively performed to obtain target characteristic protein sequence fragments, and potential sequence feature function fragment analysis is performed on the protein sequences having the function of responding to the preset physical stimulus and not having the function of responding to the preset physical stimulus based on the grouping results, and the physical field stimulation response function prediction model is decoded. The technical solution of the embodiment of the present invention solves the problem that the current potential sequence feature function fragment analysis cannot be systematically and efficiently performed on some protein sequences, and can realize the analysis and excavation of potential sequence feature function fragments in protein sequences based on an artificial intelligence method, improve the efficiency of protein sequence analysis, contribute to the research on the ion sensing physical field mechanism, promote the development of physical field stimulation related technologies, and provide a data basis for the modification and design of ion channel proteins. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a flowchart of a protein sequence analysis method provided by an embodiment of the present invention;

[0022] Figure 2 is a flowchart of a protein sequence analysis method provided by an embodiment of the present invention;

[0023] Figure 3 is a schematic diagram of a motif prediction process for an ion channel to respond to physical stimuli provided by an embodiment of the present invention;

[0024] Figure 4 is a model framework for functional annotation of an ion channel protein sequence based on a model algorithm provided by an embodiment of the present invention;

[0025] Figure 5 is a radar chart for comparing the performances of different machine learning algorithms provided by an embodiment of the present invention;

[0026] Figure 6 is a radar chart for comparing different deep learning algorithms provided by an embodiment of the present invention;

[0027] Figure 7 is a visualization schematic diagram of a motif prediction model screening process for an ion channel to respond to physical stimuli provided by an embodiment of the present invention;

[0028] Figure 8 It is a comparison chart of the data volume before and after the expansion of the protein sequence data to be analyzed provided by the embodiment of the present invention;

[0029] Figure 9 It is a comparison chart of the F1 score of feature mining provided by the embodiment of the present invention;

[0030] Figure 10 It is a schematic diagram of the evolutionary tree of voltage-gated ion channels provided by the embodiment of the present invention;

[0031] Figure 11 It is a schematic diagram of the structural visualization of sequence features provided by the embodiment of the present invention;

[0032] Figure 12 It is a schematic diagram of the structure of a protein sequence analysis device provided by the embodiment of the present invention;

[0033] Figure 13 It is a schematic diagram of the structure of a computer device provided by the embodiment of the present invention. Detailed implementation manners

[0034] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the convenience of description, only parts related to the present invention are shown in the drawings, rather than all the structures.

[0035] Figure 1 It is a flowchart of a protein sequence analysis method provided by the embodiment of the present invention. This embodiment is applicable to the scenario of protein sequence analysis, especially for the prediction of specific functional fragments in protein sequences. This method can be executed by a protein sequence analysis device, which can be implemented in software and / or hardware and integrated in a computer device with application development functions.

[0036] As Figure 1 shown, the protein sequence analysis method of this embodiment includes the following steps:

[0037] S110. Input the protein sequence vectors corresponding to each protein sequence in the protein sequence set to be analyzed into a pre-trained physical field stimulation response function prediction model, and obtain a prediction result on whether each of the protein sequences has the function of responding to a preset physical stimulus.

[0038] Among them, the protein sequence set to be analyzed can be obtained in advance from a protein database, and it is necessary to mine and analyze potential sequence feature functional fragments in the protein sequences. For example, in some scientific research scenarios, it is necessary to explore the special functional sequence fragments of ion channel proteins that sense physical field stimuli. Then, the protein sequence set to be analyzed can include any protein sequences whose ability to sense physical field stimuli is not clear. Of course, protein sequences that are clearly capable of sensing physical field stimuli, that is, known ion channel proteins, can also be added, which can be used as a reference for positive data during the analysis process to improve the accuracy of the analysis results.

[0039] Ion channel proteins are a class of functional proteins embedded in the cell membrane that control the transmembrane transport of ions. These proteins have the function of sensing physical field stimuli such as electromagnetic fields, mechanical forces, temperature, and light. Ion channels respond quickly and can sense a variety of stimuli, so they are important target proteins for techniques such as transcranial magnetic stimulation, transcranial electrical stimulation, and optogenetics. The function of ion channels to sense physical fields is determined by the sequence. In the protein sequence, proteins with specific functions contain relatively fixed amino acid combinations, called motifs. For example, the CRAC motif is the binding site for cholesterol molecules.

[0040] The protein sequence vector corresponding to each protein sequence in the protein sequence set to be analyzed can be determined by encoding with a protein language model. A protein sequence vector is a way to describe a protein sequence using an effective mathematical vector or discrete model. Among them, the protein language model is obtained by migrating a natural language processing model to the protein field because there is a high similarity between protein sequences and natural languages.

[0041] In an alternative embodiment, the ProtT5-XL-U50 protein language model can be used to express each protein sequence in the protein sequence set to be analyzed as a corresponding protein sequence vector. This model is based on the transformer architecture and is pre-trained using the protein database Uniref50. The Uniref50 database contains 45 million protein sequences consisting of 15 billion amino acids, which fully ensures that ProtT5 can capture the structural and functional connections between different types or species of proteins. The attention stack of ProtT5 contains 24 layers, each layer contains 32 attention heads, and the size of the hidden layer is 1024 dimensions. This stacking pattern allows each layer to operate on the output of the previous layer. Through such repeated combinations of word embeddings, ProtT5 can achieve a very rich representation at the bottom layer of the model.

[0042] Among them, the pre-trained physical field stimulation response function prediction model can predict whether each protein sequence in the protein sequence set to be analyzed has the function of responding to a preset physical stimulus, and obtain corresponding prediction results of yes or no and corresponding probability values. Among them, the physical field stimulation response function prediction model can be implemented by any machine learning model or deep learning model with a judgment function.

[0043] The preset physical stimulus can be any one of physical field stimuli such as electromagnetic, mechanical force, temperature, or light. Of course, different stimulus parameter values can also be set in any type of physical field stimulus. Different stimulus parameter values can be regarded as different stimuli under the same type of physical field stimulus.

[0044] S120. Group the protein sequences in the protein sequence set to be analyzed according to the prediction results, and obtain a first protein sequence group with the function of responding to the preset physical stimulus and a second protein sequence group without the function of responding to the preset physical stimulus.

[0045] The prediction result indicates whether the protein sequence in the protein sequence set to be analyzed has the function of responding to the preset physical stimulus. Grouping according to the prediction result can facilitate subsequent mining and distinguishing of the sequence difference features that can sense physical field stimuli and those that cannot.

[0046] S130. Perform sequence feature extraction and sequence feature analysis on the first protein sequence group and the second protein sequence group respectively to obtain target characteristic protein sequence fragments.

[0047] Among them, the characteristic protein sequence refers to a short sequence pattern with a specific function or structure in the protein sequence. In this embodiment, the target characteristic protein sequence fragment can be a protein sequence fragment with the function of responding to the preset physical field stimulus. Studying the characteristic protein sequence can help us better understand the function and structure of proteins and their roles in biological processes.

[0048] The process of performing sequence feature extraction and sequence feature analysis within each group respectively can use methods such as sequence alignment algorithms, pattern recognition algorithms, machine learning algorithms, or deep learning algorithms to compare and analyze the similarity of each protein sequence in the same group, so as to discover the common features and differences among them. For example, according to the content of different features in the protein sequences of different groups, select similar fragments containing sequence features that distinguish physical field stimuli for functional analysis to determine the target characteristic protein sequence fragment.

[0049] The technical solution of this embodiment is to input the protein sequence vectors corresponding to each protein sequence in the protein sequence set to be analyzed into a pre-trained physical field stimulation response function prediction model, and obtain the prediction results of whether each protein sequence has the function of responding to a preset physical stimulus, so as to quickly and efficiently conduct a preliminary prediction analysis on the protein sequences; group the protein sequences in the protein sequence set to be analyzed according to the prediction results, and obtain a first protein sequence group with the function of responding to the preset physical stimulus and a second protein sequence group without the function of responding to the preset physical stimulus; for the first protein sequence group and the second protein sequence group, perform sequence feature extraction and sequence feature analysis respectively to obtain target characteristic protein sequence fragments, and perform potential sequence feature function fragment analysis on the protein sequences with the function of responding to the preset physical stimulus and those without the function of responding to the preset physical stimulus based on the grouping results, and decode the physical field stimulation response function prediction model. The technical solution of the embodiment of the present invention solves the problem that the current method cannot systematically and efficiently perform potential sequence feature function fragment analysis on some protein sequences, and can realize the analysis and excavation of potential sequence feature function fragments in protein sequences based on artificial intelligence, improve the efficiency of protein sequence analysis, contribute to the research on the ion-sensing physical field mechanism, promote the development of physical field stimulation-related technologies, and provide a data basis for the modification and design of ion channel proteins.

[0050] Figure 2 FIG. is a flowchart of a protein sequence analysis method provided by an embodiment of the present invention. This embodiment and the protein sequence analysis method in the above embodiment belong to the same inventive concept, and further describes the training process of the physical field stimulation response function prediction model. This method can be executed by a protein sequence analysis device, which can be implemented in a software and / or hardware manner and integrated in a computer device with application development functions.

[0051] As Figure 2 shown, the protein sequence analysis method of this embodiment includes the following steps:

[0052] S210. Obtain a preset protein sequence sample set, and perform sequence encoding on each protein sequence sample in the target protein sequence sample set through a pre-trained protein language model to obtain protein sequence sample vectors.

[0053] The preset protein sequence sample set is a sample set used to train the physical field stimulation response function prediction model. The protein sequence samples included in this set can be sourced from a manually annotated protein database to ensure high-quality and accurate protein annotation information. The protein sequence samples can be protein sequences with the function of responding to a certain preset physical field stimulation, as well as protein sequences that clearly do not have the function of responding to a certain preset physical field stimulation.

[0054] Each protein sequence sample in the target protein sequence sample set is encoded by a pre-trained protein language model to obtain a protein sequence sample vector. The protein sequence sample vector enables the physical field stimulus response function prediction model to better identify and analyze the information in the protein sequence.

[0055] S220. Input the protein sequence sample vector into the physical field stimulus response function prediction model to be trained, and obtain the model learning output result.

[0056] In this embodiment, the physical field stimulus response function prediction model to be trained includes a multi-layer perceptron network structure for further mining the context information of the protein sequence.

[0057] In an alternative embodiment, the multi-layer perceptron may include five hidden layers and one classification layer. The activation function used between the hidden layers is the rectified linear unit (ReLU).

[0058] After inputting the protein sequence sample vectors corresponding to the protein sequences in the preset protein sequence sample set into the physical field stimulus response function prediction model to be trained, a model learning output result will be obtained in response.

[0059] S230. Calculate the learning loss based on the model learning output result and the sample label corresponding to the protein sequence sample vector, and update the parameters of the physical field stimulus response function prediction model to be trained according to the learning loss calculation result to complete the model training process and obtain a pre-trained physical field stimulus response function prediction model.

[0060] Among them, the weight parameter in the loss function for calculating the learning loss is determined according to the positive and negative sample ratio in the target protein sequence sample set. Since there is an imbalance in the number of positive and negative samples in the preset protein sequence sample set, a class weight parameter (class_weight) is added to the multi-layer perceptron model in this embodiment. The class weight parameter can be a parameter used to adjust the weight of each class (with or without the physical field stimulus response function) in the loss function.

[0061] When the physical field stimulus response function prediction model after parameter update can meet the preset learning loss convergence threshold or other model training process end conditions, the model training process can be ended to obtain a pre-trained physical field stimulus response function prediction model.

[0062] S240. Input the protein sequence vectors corresponding to each protein sequence in the protein sequence set to be analyzed into the pre-trained physical field stimulation response function prediction model, and obtain the prediction results of whether each of the protein sequences has the function of responding to the preset physical stimulation.

[0063] After the physical field stimulation response function prediction model is trained, it can be used in the process of predicting and analyzing the characteristic sequences that respond to the preset physical field stimulation.

[0064] Among them, the protein sequence set to be analyzed can be protein sequences from an unannotated protein database, which are protein sequence data that have not been experimentally verified and whose function of responding to the preset physical stimulation is not clear.

[0065] In an optional implementation manner, the protein sequence set to be analyzed can be further expanded by adding the preset protein sequence sample set used to train the physical field stimulation response function prediction model to the protein sequence set to be analyzed. The protein sequence set to be analyzed includes the protein sequences in the preset protein sequence sample set. Further, each protein sequence in the protein sequence set to be analyzed can be screened to filter out non-homologous protein sequences with a similarity greater than the preset similarity threshold. For example, preprocessing for eliminating redundancy is performed on the expanded protein sequence set, and non-homologous proteins with a similarity greater than 60% are removed to improve the data quality of the analysis data.

[0066] S250. Group the protein sequences in the protein sequence set to be analyzed according to the prediction results, and obtain a first protein sequence group with the function of responding to the preset physical stimulation and a second protein sequence group without the function of responding to the preset physical stimulation.

[0067] S260. Respectively extract the feature vectors of each protein sequence in the first protein sequence group and the second protein sequence group based on the composition analysis algorithm of K-space amino acid pairs.

[0068] The composition analysis algorithm of K-space amino acid pairs (CKSAAP) is a sequence-based feature extraction method. This algorithm can use the composition ratio of residue pairs with a k-interval distance in a protein sequence fragment in the sequence to establish a mathematical model and extract feature vectors, so as to achieve the purpose of predicting ubiquitin.

[0069] Feature extraction is performed on each protein sequence in the first protein sequence group and the second protein sequence group, and corresponding feature vectors can be obtained. Specifically, it can be to use the composition ratio of k residue pairs with a certain interval distance in each protein sequence fragment in the sequence to establish a mathematical model, extract the feature vector, and then perform a series of downstream tasks.

[0070] S270. Use a preset feature screening algorithm to screen features from the feature vectors to obtain a target feature combination.

[0071] Among them, the target feature combination is relative to the protein sequence set to be analyzed. This protein sequence set to be analyzed is composed of different protein families, so the features in the characteristic vector are also the sum of the features of different families, called the target feature combination. This target feature combination can be considered to be able to distinguish whether the protein sequences in the protein sequence set to be analyzed have the function of responding to a preset physical field stimulus or do not have the function of responding to a preset physical field stimulus.

[0072] The preset feature screening algorithm can be the MRMD (Max-Relevance-Max-Distance based Dimensionality Reduction) algorithm, which is used to sort each feature in the feature vector, and then screen out multiple features with higher rankings as the target feature combination.

[0073] S280. Classify each protein sequence in the first protein sequence group and the second protein sequence group, and extract the conserved segments in the protein sequences of each category.

[0074] The conserved region in a protein sequence usually refers to the region where the sequence is relatively stable and changes little in different species or different individuals of the same species.

[0075] Evolutionary trees corresponding to each group can be constructed based on the amino acid sequences of each protein sequence in the first protein sequence group and the second protein sequence group. Then, group the protein sequences within the group based on the corresponding evolutionary tree to obtain the protein sequence classification results of different protein sequence families. Furthermore, the sequence conserved segments in the protein sequences of different protein sequence families can be obtained through the expectation maximization algorithm.

[0076] S290. Use the conserved segments containing the target feature combination in the conserved segments as target feature protein sequence segments.

[0077] In this step, screening can be performed in the conserved fragments, and the conserved fragments containing the target feature combination in the conserved fragments are used as the target feature protein sequence fragments. It can be understood that in this process, the target feature protein sequence fragments containing the target feature combination of the first protein sequence group are found in the conserved fragments corresponding to the first protein sequence group; the target feature protein sequence fragments containing the target feature combination of the second protein sequence group are found in the conserved fragments corresponding to the second protein sequence group.

[0078] The technical solution of this embodiment is as follows: By obtaining a preset protein sequence sample set, and performing sequence encoding on each protein sequence sample in the target protein sequence sample set through a pre-trained protein language model to obtain protein sequence sample vectors; inputting the protein sequence sample vectors into a physical field stimulation response function prediction model to be trained to obtain a model learning output result; calculating a learning loss based on the model learning output result and the sample label corresponding to the protein sequence sample vector, and updating the parameters of the physical field stimulation response function prediction model to be trained according to the learning loss calculation result to complete the model training process and obtain a pre-trained physical field stimulation response function prediction model; inputting the protein sequence vectors corresponding to each protein sequence in the protein sequence set to be analyzed into the pre-trained physical field stimulation response function prediction model to obtain a prediction result on whether each protein sequence has the function of responding to a preset physical stimulus; grouping the protein sequences in the protein sequence set to be analyzed according to the prediction result to obtain a first protein sequence group having the function of responding to the preset physical stimulus and a second protein sequence group not having the function of responding to the preset physical stimulus; respectively extracting the feature vectors of each protein sequence in the first protein sequence group and the second protein sequence group based on the composition analysis algorithm of K-space amino acid pairs; performing feature screening from the feature vectors by using a preset feature screening algorithm to obtain a target feature combination; classifying the protein sequences in the first protein sequence group and the second protein sequence group by using an evolutionary tree, and extracting the conserved fragments in the protein sequences of each category; using the conserved fragments containing the target feature combination in the conserved fragments as the target feature protein sequence fragments. The technical solution of the embodiment of the present invention solves the problem that the current analysis of potential sequence feature functional fragments in some protein sequences cannot be systematic and efficient. It can train a neural network model to analyze and mine the potential sequence feature functional fragments in protein sequences in an artificial intelligence-based manner, improve the efficiency of protein sequence analysis, and is a systematic analysis of protein sequences in multiple protein families during the analysis process, and the obtained analysis results have high reliability, which helps to improve the research on the physical field mechanism of ion sensing, promote the development of physical field stimulation-related technologies, and provide a data basis for the modification and design of ion channel proteins.

[0079] In a specific embodiment, the prediction of ion channel motifs capable of responding to voltage stimulation is taken as an example for illustration. Specifically, the embodiment of the method for predicting motifs of ion channel proteins that sense physical field stimuli includes the following content:

[0080] The overall scheme of the method for predicting motifs of ion channel proteins that sense physical field stimuli is as Figure 3 shown. The overall scheme includes three stages: data preparation, data augmentation, and data mining.

[0081] Among them, in the data preparation stage, the corresponding protein sequences are obtained from a preset protein database. In this embodiment, two protein sequence sets need to be prepared in advance. One is the data set for training the physical field stimulus response function prediction model, and the other is the data to be analyzed that has not been experimentally verified and has unclear functions.

[0082] Commonly used protein databases are: UniProt database and Gene Ontology (GO) database. UniProt (Universal Protein Resource) is a comprehensive protein database, including UniProtKB (UniProt Knowledgebase), UniParc (UniProt Archive), and UniRef (UniProt Reference Clusters). Among them, UniProtKB (UniProt Knowledgebase) is a main component of the UniProt database family. UniProtKB includes two main parts: Swiss-Prot and TrEMBL. Swiss-Prot is a manually annotated protein database to ensure high-quality and accurate protein annotation information. TrEMBL is a computationally annotated protein database containing a large number of protein sequences that have not been manually verified. The goal of TrEMBL is to cover all known proteins as much as possible and provide researchers with a wider range of protein data resources. In this embodiment, two data sets are compiled in total. The data set for model training is from Swiss-Prot, which is a protein sequence verified by experiments. The data set that needs to be analyzed by the model is from TrEMBL, which is data that has not been experimentally verified and has unclear functions.

[0083] In the data preparation stage, keywords (protein marker information): Ion channel (KW-0407), Voltage-gated channel (KW-0851), voltage-gated channel activity (GO:0022832), and voltage-gated sodium channel activity (GO:0005248) can be used to screen protein sequences. In addition, a series of basic protein screening criteria are used to collect a high-quality dataset for model training, including 514 voltage-gated channel proteins, 872 non-voltage-gated channel proteins, and 5551 ion channel sequence data that have not been experimentally functionally tested.

[0084] Next, using the prepared dataset, a function prediction model with high classification accuracy is built. The function prediction model constructed in the present invention is as Figure 4 shown. The first step of the prediction model is to solve how to convert protein sequence information into vector samples that can be used for function prediction. Therefore, when constructing a sequence-based function prediction algorithm, a key step is to use effective mathematical vectors or discrete models to describe protein sequences. Since there is a high similarity between protein sequences and natural language, natural language processing models can be well migrated to the protein field, giving rise to protein language models. The ProtT5-XL-U50 protein language model used in this embodiment is based on the transformer architecture and pre-trained using the protein database Uniref50. The Uniref50 database contains 45 million protein sequences composed of 15 billion amino acids, fully ensuring that ProtT5 can capture the structural and functional connections between different types or races of proteins. The attention stack of ProtT5 contains 24 layers, each layer contains 32 attention heads, and the size of the hidden layer is 1024 dimensions. This stacking pattern enables each layer to operate on the output of the previous layer. Through such repeated word embedding combinations, ProtT5 can achieve a very rich representation at the bottom layer of the model. Therefore, the embedding of the last layer of the attention stack can be extracted as the feature representation. A multi-layer perceptron model is used to continue to mine the context information of protein sequences. The multi-layer perceptron mainly includes five hidden layers and a classification layer. Specifically:

[0085] The activation function used between the hidden layers is the Rectified Linear Unit (ReLU).

[0086]

[0087] The Adam (Adaptive Moment Estimation) algorithm can be used for optimization. The Adam optimization algorithm automatically adjusts the learning rate based on the first and second moment estimates of the gradient of each parameter. The loss function used is the Binary Cross Entropy Loss with Logits, which is a commonly used loss function in classification algorithms. It combines the Sigmoid activation function and the binary cross entropy loss. The Sigmoid function maps the original output of the network to the range (0, 1), representing the probability value.

[0088] loss(z,y)=meanl0,……,l N-1 ;

[0089] l n =-(y n *log(z n )+(1 - y n )*log(1 - z n ))。

[0090] The backpropagation algorithm is used to further optimize the parameters. The key of the backpropagation algorithm is the use of the chain rule. By propagating the error backward layer by layer, the contribution of each parameter to the loss is calculated, so as to realize the adjustment of the parameters. For the function y = f(g(x)), the derivative of y with respect to x is

[0091]

[0092] Due to the sample imbalance in the dataset, in this embodiment, a class weight parameter (class_weight) is added to the multi-layer perceptron model. The class weight is a parameter used to adjust the weight of each class in the loss function. Experimental verification is performed on this parameter of the functional prediction model for the response voltage stimulus, and finally a ratio value of 4:1 for positive samples and negative samples is given.

[0093] Next, the classification performance of the functional prediction model is evaluated. In this embodiment, the ten-fold cross-validation evaluation method is used, and six commonly used metrics for evaluating model performance are used, namely the True Positive Rate (TPR), the True Negative Rate (TNR), the Accuracy (ACC), the Matthews Correlation Coefficient (MCC), and the Area Under Curve (AUC) of the ROC curve. The specific comparison results are as Figure 5 and Figure 6 shown. Among them, Figure 5 Radar chart of performance comparison of different machine learning algorithms,Figure 6 It is a radar chart for comparing different deep learning algorithms.

[0094] Among them, VGIs-Pred (voltage gated ion channels) refers to Figure 4 the pre-trained functional prediction model of voltage-gated channel proteins. After comparative analysis, the performance of VGIs-Pred is relatively better. Further, as Figure 7 shown, through the non-linear dimensionality reduction technique of t-SNE (t-Distributed Stochastic Neighbor Embedding), the dimensionality of the protein sequence information is reduced, and the classification process of the protein sequences in the dataset is visualized by VGIs-Pred. This model can well classify the data that responds to voltage stimulation and the data that does not respond.

[0095] In the data augmentation stage, using the built functional prediction model, the functions of 5551 ion channel protein sequences in the TrEMBL database are predicted, and functional information is annotated for the ion channel sequences. The functional prediction model that responds to voltage stimulation successfully predicts 2121 voltage-gated channel proteins and 3430 non-voltage-gated channel proteins. These two newly separated types of proteins will be merged with the protein dataset used for model training, with a total of 2635 voltage-gated channel proteins and 4302 non-voltage-gated channel proteins, for the next step of predicting the structural functions of voltage-gated sequences. The data expansion results are shown as Figure 8 shown, the comparison chart of the data volume before and after data expansion. Figure 8 In it, Training Dataset represents the protein sequence data used for model training, Classified Dataset represents the protein sequence data used for classification predicted by the model, and Merged Dataset represents the protein sequence data after the merger of the first two. VGICs represents voltage-gated channel proteins, and Non-VGICs represents non-voltage-gated channel proteins.

[0096] Then in the data mining stage, first use the CD-HIT redundancy removal algorithm, select a similarity threshold of 60%, and after screening, the dataset includes 695 voltage-controlled channel proteins and 1352 non-voltage-gated channel proteins, which are used for the subsequent visualization feature analysis of the CKSAAP feature extraction algorithm.

[0097] Specifically: By using the composition ratio of k-spaced residue pairs in a protein sequence fragment in the sequence, a mathematical model is established to extract feature vectors for a series of downstream tasks. When k = 0, that is, each amino acid and the next adjacent amino acid form a pair for extraction, which means the spacing between these two amino acids is k = 0 amino acids. And so on, when k = 1, the residue pairs to be extracted are spaced k = 1 amino acid apart. In this method, considering the helical domain of the protein, k is set to 5. When the sequence length is N and the spacing is k, the total number of residue pairs that can be extracted is N - k - 1, denoted as

[0098] N Total = N - k - 1.

[0099] There are 20 basic amino acids, so the number of residue pairs that can be formed is 20×20 = 400. The CKSAAP algorithm counts the probabilities of these residue pairs appearing in this protein sequence, thus generating a 400-dimensional feature vector, that is

[0100]

[0101] Therefore, when k = 5, that is, when taking values of 0, 1, 2, 3, 4, or 5 respectively, the length of the vector is 400×6 = 2400, and the features range from AA.gap0 to YY.gap5. When the spacing is 5, a total of 7 amino acids are involved. Generally, a helix has a cycle of 7 amino acids, and the composition of 7 amino acids can reflect the characteristics in the periodic fragment. Therefore, a maximum gap of 5 can better accommodate sequence and structural information.

[0102] Next, the MRMD3.0 algorithm is used to rank the importance of the 2400 features and screen important feature combinations. The screening process is as Figure 9 shown, Figure 9 showing the F1 score comparison chart of each feature during the feature mining process. Based on Figure 9 the important feature combinations in response to voltage stimulation, the final results are shown in Table 1, and a total of 36 important features are screened out.

[0103] Table 1

[0104]

[0105]

[0106] Next, an evolutionary tree was constructed for the amino acid sequences involved in the feature combination analysis, and the maximum expectation algorithm was used to obtain the sequence conserved fragments. The conserved fragments containing the feature combination were retained to become the predicted motif results. In the example of the response to voltage stimulation, the evolutionary tree of voltage-gated ion channels is as Figure 10 shown. The obtained positive families mainly include voltage-gated potassium channels (Kv, Kir, Slo, K2p), voltage-gated sodium channels (Cav), voltage-gated sodium channels (Nav), some chloride ion channels (Clv), cyclic nucleotide-gated channels (CNG), and glutamate receptor channels (iGluR). The results can be seen in Table 2.

[0107] Table 2

[0108]

[0109]

[0110]

[0111] The sequence features were projected onto the structure using a visualization tool to further localize the features. This will greatly facilitate biologists' interpretation of the mechanism of ion channel response to physical field stimulation. The sequence features are expressed on the structure as Figure 11 shown. In Figure 11 , the left figure a exemplifies the schematic structure of the voltage-gated potassium channel protein Kv, and the right figure b shows the 5 feature sequences contained in Kv.

[0112] In this example, theoretical demonstration is first carried out. The collected data is analyzed and sorted to obtain a reliable data set with biological and statistical significance. Then, using this data set, a functional prediction model with high classification accuracy is built to identify and classify new data, so as to expand the data set and provide more data support for subsequent functional analysis. In order to further predict the motif of ion channels in response to physical field stimulation, the redundancy removal algorithm, CKSAAP algorithm for feature extraction, MRMD3.0 algorithm for feature ranking, expectation maximization algorithm for constructing family classification, and expectation maximization algorithm for screening conserved sequences are jointly used to predict the motif in response to physical field stimulation. First of all, this example can be used to identify the channels in ion channel proteins that sense physical field stimulation, realize the functional annotation of all sequences of ion channel proteins, provide more data for subsequent motif prediction, and also be able to predict the functions of new ion channel protein sequences. Secondly, this example proposes a new research method process for ion channel proteins, integrating various existing machine learning methods to analyze the mapping relationship of the sequence functions of ion channel proteins and predict the motif of the function of sensing physical stimulation. This example can not only use artificial intelligence algorithms with high accuracy to perform functional annotation on all ion channel protein data sets based on high-quality small samples, but also use machine learning algorithms to decode the features of the artificial intelligence algorithm model, alleviating the contradiction that artificial intelligence algorithms are not interpretable. The present invention predicts the motif of ion channels in response to physical field stimulation, and these features can help biologists provide directions in the research of ion channel proteins sensing physical field stimulation and assist them in analyzing and modifying the ability of ion channels to respond to physical field stimulation.

[0113] Figure 12 FIG. is a schematic structural diagram of a protein sequence analysis device provided by an embodiment of the present invention. This embodiment is applicable to the scenario of protein sequence analysis, especially for the case of predicting specific functional fragments in protein sequences. The protein sequence analysis device can be implemented in the form of software and / or hardware and integrated in a computer terminal device with application development functions.

[0114] As Figure 12 shown, the protein sequence analysis device includes: a protein sequence function prediction module 310, a protein sequence grouping module 320, and a protein sequence feature extraction module 330.

[0115] Among them, the protein sequence function prediction module 310 is used to input the protein sequence vectors corresponding to each protein sequence in the protein sequence set to be analyzed into a pre-trained physical field stimulation response function prediction model, and obtain the prediction result of whether each of the protein sequences has the function of responding to a preset physical stimulus; the protein sequence grouping module 320 is used to group the protein sequences in the protein sequence set to be analyzed according to the prediction result, and obtain a first protein sequence group with the function of responding to the preset physical stimulus and a second protein sequence group without the function of responding to the preset physical stimulus; the protein sequence feature extraction module 330 is used to perform sequence feature extraction and sequence feature analysis on the first protein sequence group and the second protein sequence group respectively, and obtain target feature protein sequence fragments.

[0116] The technical solution of this embodiment can quickly and efficiently perform preliminary prediction analysis on protein sequences by inputting the protein sequence vectors corresponding to each protein sequence in the protein sequence set to be analyzed into a pre-trained physical field stimulation response function prediction model, and obtaining the prediction result of whether each protein sequence has the function of responding to a preset physical stimulus; group the protein sequences in the protein sequence set to be analyzed according to the prediction result, and obtain a first protein sequence group with the function of responding to the preset physical stimulus and a second protein sequence group without the function of responding to the preset physical stimulus; perform in-group sequence feature extraction and sequence feature analysis on the first protein sequence group and the second protein sequence group respectively, and obtain target feature protein sequence fragments, and perform potential sequence feature function fragment analysis on the protein sequences with the function of responding to the preset physical stimulus and the protein sequences without the function of responding to the preset physical stimulus based on the grouping result, and decode the physical field stimulation response function prediction model. The technical solution of the embodiment of the present invention solves the problem that the current cannot systematically and efficiently perform potential sequence feature function fragment analysis on some protein sequences, and can realize the analysis and excavation of potential sequence feature function fragments in protein sequences based on artificial intelligence, improve the efficiency of protein sequence analysis, help improve the research on the physical field mechanism of ion sensing, promote the development of physical field stimulation-related technologies, and provide a data basis for the transformation and design of ion channel proteins.

[0117] In an optional implementation manner, the protein sequence feature extraction module 330 is specifically used for:

[0118] Extract the feature vectors of each protein sequence in the first protein sequence group and the second protein sequence group respectively based on the composition analysis algorithm of amino acid pairs in the K space;

[0119] Use a preset feature screening algorithm to screen features from the feature vectors to obtain a target feature combination;

[0120] Classify each protein sequence in the first protein sequence group and the second protein sequence group, and extract the conserved fragments in the protein sequences of each category;

[0121] Use the conserved fragments containing the target feature combination in the conserved fragments as the target feature protein sequence fragments.

[0122] In an alternative embodiment, the protein sequence feature extraction module 330 is further specifically configured to:

[0123] Construct a phylogenetic tree based on the amino acid sequences of each protein sequence in the first protein sequence group and the second protein sequence group;

[0124] Perform intra-group protein sequence grouping based on the phylogenetic tree to obtain the protein sequence classification results of different protein sequence families;

[0125] Obtain the sequence conserved fragments in the protein sequences of different protein sequence families through the expectation maximization algorithm.

[0126] In an alternative embodiment, the protein sequence analysis device further includes a model training module for training a physical field stimulation response function prediction model. The training process of the physical field stimulation response function prediction model includes:

[0127] Obtain a preset protein sequence sample set, and perform sequence encoding on each protein sequence sample in the target protein sequence sample set through a pre-trained protein language model to obtain protein sequence sample vectors;

[0128] Input the protein sequence sample vectors into the physical field stimulation response function prediction model to be trained to obtain a model learning output result;

[0129] Calculate the learning loss based on the model learning output result and the sample label corresponding to the protein sequence sample vector, and update the parameters of the physical field stimulation response function prediction model to be trained according to the learning loss calculation result to complete the model training process;

[0130] Among them, the weight parameter in the loss function for calculating the learning loss is determined according to the positive and negative sample ratio in the target protein sequence sample set.

[0131] In an alternative embodiment, the protein sequence set to be analyzed includes the protein sequences in the target protein sequence sample set. Before grouping the protein sequences in the protein sequence set to be analyzed according to the prediction result, the protein sequence feature extraction module 330 may also be specifically configured to:

[0132] For each protein sequence in the protein sequence set to be analyzed, non-homologous protein sequences with a similarity greater than a preset similarity threshold are filtered out.

[0133] In an alternative embodiment, the physical field stimulation response function prediction model includes a multi-layer perceptron network structure.

[0134] In an alternative embodiment, the preset physical stimulation is at least one of electrical stimulation, light stimulation, temperature stimulation, and mechanical force stimulation.

[0135] The protein sequence analysis device provided by the embodiments of the present invention can execute the protein sequence analysis method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0136] Figure 13 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Figure 13 It shows a block diagram of an exemplary computer device 12 suitable for implementing the embodiments of the present invention. Figure 13 The shown computer device 12 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention. The computer device 12 can be any terminal device with computing capabilities, such as intelligent controllers, servers, mobile phones, and other terminal devices.

[0137] As Figure 13 shown, the computer device 12 is presented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0138] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0139] The computer device 12 typically includes a variety of computer system-readable media. These media can be any available media accessible by the computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0140] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computing device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 13 not shown, typically referred to as a "hard disk drive"). Although Figure 13 not shown in, a disk drive for reading and writing on removable non-volatile disks (such as a "floppy disk"), and an optical disk drive for reading and writing on removable non-volatile optical disks (such as CD-ROM, DVD-ROM or other optical media) can be provided. In these cases, each drive can be connected to the bus 18 through one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present invention.

[0141] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 generally execute the functions and / or methods in the embodiments described in the present invention.

[0142] The computing device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the computing device 12, and / or communicate with any device that enables the computing device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Also, the computing device 12 can communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the computing device 12 through the bus 18. It should be understood that although Figure 13 not shown in, other hardware and / or software modules can be used in conjunction with the computing device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0143] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing the protein sequence analysis method provided in the embodiments of the present invention. The method includes:

[0144] Input the protein sequence vectors corresponding to each protein sequence in the protein sequence set to be analyzed into a pre-trained physical field stimulation response function prediction model, and obtain the prediction results of whether each of the protein sequences has the function of responding to a preset physical stimulus;

[0145] Group the protein sequences in the protein sequence set to be analyzed according to the prediction results, and obtain a first protein sequence group having the function of responding to the preset physical stimulus and a second protein sequence group not having the function of responding to the preset physical stimulus;

[0146] For the first protein sequence group and the second protein sequence group, perform sequence feature extraction and sequence feature analysis respectively to obtain target feature protein sequence fragments.

[0147] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the protein sequence analysis method provided in any embodiment of the present invention. The method includes:

[0148] Input the protein sequence vectors corresponding to each protein sequence in the protein sequence set to be analyzed into a pre-trained physical field stimulation response function prediction model, and obtain the prediction results of whether each of the protein sequences has the function of responding to a preset physical stimulus;

[0149] Group the protein sequences in the protein sequence set to be analyzed according to the prediction results, and obtain a first protein sequence group having the function of responding to the preset physical stimulus and a second protein sequence group not having the function of responding to the preset physical stimulus;

[0150] For the first protein sequence group and the second protein sequence group, perform sequence feature extraction and sequence feature analysis respectively to obtain target feature protein sequence fragments.

[0151] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable media may be computer-readable signal media or computer-readable storage media. The computer-readable storage media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage media include: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this document, the computer-readable storage media may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component.

[0152] The computer-readable signal media may include data signals propagated in a baseband or as part of a carrier wave, which carry computer-readable program codes. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal media may also be any computer-readable medium other than the computer-readable storage media, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, device, or component.

[0153] The program codes contained on the computer-readable media can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0154] The computer program codes for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program codes can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0155] Embodiments of the present disclosure also provide a computer program product, including a computer program which, when executed by a processor, implements the protein sequence analysis method provided in any one of the embodiments of the present disclosure.

[0156] In the process of implementing the computer program product, computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0157] Those of ordinary skill in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. Optionally, they can be implemented with program code executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. Thus, the present invention is not limited to any specific combination of hardware and software.

[0158] Note that the above is only the preferred embodiment of the present invention and the applied technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A method for protein sequence analysis, characterized in that, Including: Inputting the protein sequence vectors corresponding to each protein sequence in the protein sequence set to be analyzed into a pre-trained physical field stimulation response function prediction model to obtain a prediction result on whether each of the protein sequences has the function of responding to a preset physical stimulation; each protein sequence in the protein sequence set to be analyzed is an ion channel protein; wherein, the protein sequence vector is obtained based on a pre-trained protein language model, the protein language model is trained based on the transformer architecture, the attention stack of the protein language model contains 24 layers, each layer contains 32 attention heads, and the size of the hidden layer is 1024 dimensions; Grouping the protein sequences in the protein sequence set to be analyzed according to the prediction results to obtain a first protein sequence group having the function of responding to the preset physical stimulation and a second protein sequence group not having the function of responding to the preset physical stimulation, and merging the first protein sequence group and the second protein sequence group with the protein data having the function of responding to the preset physical stimulation and the protein data not having the function of responding to the preset physical stimulation in the protein dataset used to train the physical field stimulation response function prediction model respectively to obtain an updated first protein sequence group and an updated second protein sequence group; Respectively extracting the feature vectors of each protein sequence in the updated first protein sequence group and the updated second protein sequence group based on the composition analysis algorithm of K-space amino acid pairs; Performing feature screening from the feature vectors by using a preset feature screening algorithm to obtain a target feature combination; Classifying each protein sequence in the updated first protein sequence group and the updated second protein sequence group, and extracting the conserved segments in the protein sequences of each category; Taking the conserved segments containing the target feature combination in the conserved segments as target feature protein sequence segments.

2. The method according to claim 1, wherein The classifying each protein sequence in the updated first protein sequence group and the updated second protein sequence group, and extracting the conserved segments in the protein sequences of each category includes: Constructing an evolutionary tree based on the amino acid sequences of each protein sequence in the updated first protein sequence group and the updated second protein sequence group; Performing intra-group protein sequence grouping based on the evolutionary tree to obtain the protein sequence classification results of different protein sequence families; Obtaining the sequence conserved segments in the protein sequences of different protein sequence families through the expectation maximization algorithm.

3. The method according to claim 1, wherein The training process of the physical field stimulation response function prediction model includes: Obtaining a preset protein sequence sample set, and performing sequence encoding on each protein sequence sample in the target protein sequence sample set through a pre-trained protein language model to obtain a protein sequence sample vector; Inputting the protein sequence sample vector into the physical field stimulation response function prediction model to be trained to obtain a model learning output result; Calculate the learning loss according to the output result of model learning and the sample label corresponding to the protein sequence sample vector, and update the parameters of the physical field stimulation response function prediction model to be trained according to the calculation result of the learning loss, so as to complete the model training process; Among them, the weight parameter in the loss function used for calculating the learning loss is determined according to the positive and negative sample ratio in the target protein sequence sample set.

4. The method according to claim 3, characterized in that, The protein sequence set to be analyzed includes the protein sequences in the preset protein sequence sample set. Before grouping the protein sequences in the protein sequence set to be analyzed according to the prediction result, the method further includes: Screen each protein sequence in the protein sequence set to be analyzed, and filter out non-homologous protein sequences with similarity greater than the preset similarity threshold.

5. The method according to claim 3, characterized in that, The physical field stimulation response function prediction model includes a multi-layer perceptron network structure.

6. The method according to any one of claims 1-5, characterized in that, The preset physical stimulation is at least one of electrical stimulation, light stimulation, temperature stimulation and mechanical force stimulation.

7. A protein sequence analysis device, characterized in that, It includes: A protein sequence function prediction module, configured to input the protein sequence vector corresponding to each protein sequence in the protein sequence set to be analyzed into a pre-trained physical field stimulation response function prediction model, and obtain a prediction result on whether each protein sequence has the function of responding to the preset physical stimulation; each protein sequence in the protein sequence set to be analyzed is an ion channel protein; wherein, the protein sequence vector is obtained based on a pre-trained protein language model, the protein language model is trained based on the transformer architecture, the attention stack of the protein language model contains 24 layers, each layer contains 32 attention heads, and the size of the hidden layer is 1024 dimensions; A protein sequence grouping module, configured to group the protein sequences in the protein sequence set to be analyzed according to the prediction result, obtain a first protein sequence group with the function of responding to the preset physical stimulation and a second protein sequence group without the function of responding to the preset physical stimulation, and combine the first protein sequence group and the second protein sequence group with the protein data with the function of responding to the preset physical stimulation and the protein data without the function of responding to the preset physical stimulation in the protein dataset used to train the physical field stimulation response function prediction model respectively, to obtain an updated first protein sequence group and an updated second protein sequence group; A protein sequence feature extraction module, configured to respectively extract the feature vectors of each protein sequence in the updated first protein sequence group and the updated second protein sequence group based on the composition analysis algorithm of K-space amino acid pairs; perform feature screening from the feature vectors by using a preset feature screening algorithm to obtain a target feature combination; classify each protein sequence in the updated first protein sequence group and the updated second protein sequence group, and extract the conserved fragments in each category of protein sequences; use the conserved fragments containing the target feature combination in the conserved fragments as the target feature protein sequence fragments.

8. A computer device, characterized in that, The computer device includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the protein sequence analysis method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the protein sequence analysis method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the protein sequence analysis method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method for identifying thermophilic protein based on machine learning

    CN110517730A

  • Thermophilic protein identification method based on ensemble learning, storage medium and equipment

    CN113971985A