Gene prediction and functional peptide recognition method based on multi-feature fusion
By using a multi-feature fusion method combined with a deep learning model to extract local and global information of genome and peptide sequences, the problem of low efficiency in gene prediction and functional peptide identification in existing technologies is solved, and accurate prediction of multiple functional peptides and discovery of new peptides are achieved.
Patent Information
- Application Number
- CN202510978554.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-26
AI Technical Summary
The existing gene prediction and functional peptide identification models are inefficient, costly, and rely on local sequence features while ignoring contextual semantics and structural information. In addition, the functional peptide identification models lack a unified multi-classification framework, making it difficult to identify multiple functional peptides.
A multi-feature fusion method is adopted, combining convolutional neural networks and Transformer encoders to extract the local structure and global context information of the genomic sequence, and the BERT language model is used to perform contextual semantic modeling of the peptide sequence. The global sequence dependency is captured through the Bi-LSTM module, and finally a multi-layer perceptron is used for multi-classification prediction.
It achieves efficient automatic feature extraction, combines local and global information, accurately predicts the activity of genes and functional peptides, can identify multiple functional peptides, reduce dependence on database comparison, and discover new genes and new anticancer and antibacterial peptides.
Smart Images

Figure CN120708695A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a gene prediction and functional peptide identification method, belonging to the technical field of bioinformatics. Background Art
[0002] Developing therapeutics with novel mechanisms of action, high bioactivity, and low risk of drug resistance has become an effective approach to combat cancer and drug-resistant infections. Since the discovery of functional peptides (such as ACPs and AMPs) is premised on the identification of their encoding genes, a complete bioinformatics processing chain is usually required: first, the coding region is identified from genomic data using a gene prediction model, the predicted gene sequence is translated into a protein sequence, and finally, a functional peptide recognition model is used to screen candidate peptides with anticancer or antibacterial activity from the protein sequence. However, existing gene prediction and functional peptide recognition models still have some problems: traditional prediction methods are unable to automatically extract features, resulting in low efficiency and high cost; most models rely on local sequence features, ignoring contextual semantics and structural information, and are unable to fully capture the potential bioactivity signals of the sequence, facing the problem of insufficient information consideration; existing functional peptide recognition models are mostly single functional peptide predictions, lacking a unified, generalizable multi-classification framework, making it difficult to identify multiple functional peptides. Summary of the Invention
[0003] The present invention aims to solve the problems that traditional prediction methods cannot automatically extract features, resulting in low efficiency and high cost; most models rely on local sequence features and cannot fully capture the potential biological activity signals of the sequence, resulting in insufficient information consideration, and thus propose a gene prediction and functional peptide identification method based on multi-feature fusion.
[0004] The technical solution adopted by the present invention to solve the above problems is: the present invention includes a gene prediction part and a functional peptide recognition part, wherein the gene prediction part includes: The feature extraction module is used to extract the open reading frame information and coding sequence from the input genome sequence and extract the corresponding artificial features; Multi-feature fusion module, used to expand the sequence encoding results of ORF and artificial feature vectors and concatenate them into a combined vector; Convolutional neural network module, used to perform one-dimensional convolution operation on the combined vector to extract local structural features in the sequence; The Transformer encoder module is used to further model the CNN output results and extract long-range dependencies and global contextual semantic information; The classification prediction module, including the fully connected layer and the Softmax layer, is used to predict the probability of whether each ORF is a coding region; The protein translation and storage module is used to translate the ORFs predicted as coding regions into corresponding protein sequences and store them for subsequent biological function mining or analysis; The functional peptide recognition part includes: A peptide sequence representation module is used to receive and perform tokenization or tokenization on the original peptide sequence; The Bert module uses the BERT language model, which has been pre-trained in the field of biological sequences, to perform contextual semantic modeling on peptide sequences and output a semantic vector representation of each amino acid residue; CNN module, used to extract local structural features from BERT output; Bi-LSTM modules, used to process BERT outputs in parallel and capture global sequence dependencies; The feature fusion and classification module is used to fuse the output results of CNN and Bi-LSTM and input them into the multi-layer perceptron classifier to complete the multi-classification prediction task of anticancer peptides, antibacterial peptides and non-functional peptides.
[0005] Furthermore, the working process of the feature extraction module is: Step 1: Single codon usage, use a vector to represent a fragment: , , Where, Indicates the The frequency of occurrence of a single codon in the ORF, Indicates the The number of codons in the ORF, represents the total number of all codons in the ORF; Step 2: Double codon usage, using vector Represents a fragment: , , Where, Indicates double codon The frequency of occurrence in ORF, Indicates double codon number in ORF; Step 3: Feature vector Expressed as: , , Where, represents a sequence window centered around a potential start codon, represents the one-hot encoding function, represents the normalization function, Indicates the prediction of whether there is a TIS translation start site in the sequence window; Step 4: ORF length, using feature vector Represents metagenomic fragments: , Where, Indicates the ORF length, Indicates the length of the fragment; Step 5: GC content, using vector Represents a sequence fragment: , , Where, Indicates the The GC content percentage of the sequence, represents the total number of sequences, Indicates the G content in the sequence, Indicates the content of C in the sequence, Indicates the total number of GCs in the ORF; Step 6: The basic base content, the content ratio of adenine, thymine, guanine and cytosine are respectively calculated using 、 、 、 express: , ,
[0006] , , , Where, Indicates the total number of bases in the sequence, represents the number of adenines in the sequence, represents the number of thymines in the sequence, Indicates the amount of guanine, represents the number of cytosines.
[0007] Furthermore, the co-location process of the multi-feature fusion module is as follows: The encoded ORF, single codon usage, double codon usage, ORF length, TIS, GC content, base content, and all the features of the tag are concatenated into a set of one-dimensional feature vectors to represent the input sequence fragment, which is expressed as: , Where, represents the ORF feature extraction vector, represents the single codon usage feature extraction vector, represents the dicodon usage feature extraction vector, represents the TIS feature extraction vector, Indicates the length feature extraction vector of ORF, represents the GC content feature extraction vector, represents the basic base content feature extraction vector, Indicates a label.
[0008] Furthermore, the CNN module performs a one-dimensional convolution operation on the combined vector to extract local structural features in the sequence, specifically including: using two convolutional layers and two maximum pooling layers.
[0009] Furthermore, the Transformer encoder module further models the CNN output results to extract long-distance dependencies and global contextual semantic information, specifically including: using position encoding, multi-head attention mechanism, residual connection + layer normalization, feedforward fully connected network and residual connection + normalization layer.
[0010] Furthermore, the peptide sequence representation module works as follows: each amino acid is converted into its corresponding digital ID, using 26 tokens in the language vocabulary, and the digital ID is used as the input of the peptide language model.
[0011] Furthermore, the Bert module uses the BERT language model, which has been pre-trained in the field of biological sequences, to perform contextual semantic modeling on peptide sequences and output a semantic vector representation of each amino acid residue. Specifically, it selects OntoProtein as the base model and further fine-tunes it based on the working dataset.
[0012] Furthermore, four datasets were used to construct the gene prediction dataset, including: the first dataset consisted of 131 fully sequenced bacterial and archaeal genomes and 33 Gram bacterial genomes randomly selected from the NCBI RefSeq database, which contained a total of 164 complete genomes in a 7:3 ratio for training and validating the model; The second dataset consists of 10 complete genomes and is used to debug the model; The third dataset is an independent dataset containing two complete genomes of archaea and seven bacteria, with multiple complete genome samples for each species; The fourth dataset consisted of 100 new genomes of Gram-forming bacteria and Staphylococci, which were randomly divided into 5 subsets.
[0013] Furthermore, the functional peptide recognition dataset was constructed using three datasets, including two anticancer peptide datasets Dataset1 and Dataset2 provided by the AntiCP 2.0 method, and an antimicrobial peptide dataset Dataset3 provided by the C_AMP method. The model was trained and tested with an 8:2 ratio. The integrated datasets were derived from multiple recognized peptide databases, including: DADP, CAMP, APD, APD2, UniProt and Swiss-Prot; In the anticancer peptide prediction task, positive samples consist of experimentally verified anticancer peptides, and the data is compiled by combining the AMP database and the CancerPPD database; negative samples are selected from antimicrobial peptide sequences that lack anticancer activity in the AMP database to ensure that they do not have known anticancer properties; In the antimicrobial peptide prediction task, to eliminate bias, peptides with only anticancer activity were not included in the AMP positive class to ensure that all samples had antimicrobial activity; non-AMP data were obtained from the UniProt database, which screened cytoplasmic localized proteins and excluded entries containing keywords such as "antibacterial, antiviral, and antifungal." After removing duplicates and excluding sequences identical to AMP, non-AMP sequences were finally retained.
[0014] Furthermore, open reading frames were extracted from the genomic sequence and truncated or padded to a fixed length of 700 base pairs; these fixed-length ORFs were one-hot encoded and artificial features were extracted simultaneously; Then, the encoded open reading frame sequence and the extracted artificial feature vector are expanded into one-dimensional arrays respectively, and these arrays are concatenated into a longer vector and input into the CNN model to further extract local features; Flatten the feature map output by CNN and then input it into the Transformer encoder to extract global features; The candidate open reading frames are predicted through the fully connected layer and the softmax layer. The predicted coding regions are translated into protein sequences and tokenized. Each amino acid is represented by its corresponding token ID, using special tags [CLS] and [SEP]. The tokenized sequence is then input into a pre-trained language model, which maps each amino acid into a vector representation containing rich biological semantics and contextual information. The BERT output is then fed into two feature extraction paths, the CNN module and the Bi-LSTM module, for local feature extraction and global dependency feature extraction. Finally, the outputs of CNN and Bi-LSTM are concatenated through concat, and the concatenated vector is input into the multi-layer perceptron for functional peptide prediction.
[0015] The present invention has the following beneficial effects: It utilizes a deep learning-based model architecture, using genome and peptide sequences from databases such as NCBI as training and testing data, to develop an artificial intelligence model capable of accurately predicting genes and functional peptides across the genomes of most species. Compared to existing technologies, this invention better handles automatic feature extraction, combining local feature extraction with global information considerations to accurately predict the activities of genes and two functional peptides. It also eliminates the reliance of traditional prediction methods on database comparisons and facilitates the discovery of new genes and novel anticancer and antimicrobial peptides. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flow chart of the present invention; Figure 2 It is a structural schematic diagram of the present invention; Figure 3 is a schematic diagram of the evaluation of the gene prediction model; Figure 4 Schematic diagram of the evaluation of the functional peptide recognition model. DETAILED DESCRIPTION
[0017] Specific embodiment 1: A gene prediction and functional peptide identification method based on multi-feature fusion, the method includes gene prediction, functional peptide identification, data set construction and model evaluation, such as Figure 1 As described above, gene prediction includes: Step A1: The feature extraction module is used to extract the open reading frame information and coding sequence from the input genome sequence, and extract the corresponding artificial features; Step A2: The multi-feature fusion module expands the sequence encoding result of the open reading frame and the artificial feature vector and concatenates them into a combined vector; Step A3: The CNN module performs a one-dimensional convolution operation on the combined vector to extract local structural features in the sequence; Step A4: The Transformer encoder module further models the CNN output and extracts long-range dependencies and global contextual semantic information. Step A5: The classification prediction module includes a fully connected layer and a softmax layer, which is used to predict the probability of whether each open reading frame is a coding region; Step A6: The protein translation and storage module translates the ORF predicted as the coding region into the corresponding protein sequence and stores it for subsequent biological function mining or analysis; Functional peptide identification includes: Step B1: The peptide sequence representation module is used to receive and perform word segmentation or tokenization on the original peptide sequence; Step B2: The BERT module uses the BERT language model that has been pre-trained in the field of biological sequences to perform contextual semantic modeling on the peptide sequence and output a semantic vector representation of each amino acid residue; Step B3: The CNN module extracts local structural features from the BERT output; Step B4: The Bi-LSTM module processes the BERT output in parallel to capture global sequence dependencies. Step B5: The feature fusion and classification module fuses the output results of CNN and Bi-LSTM and inputs them into the multi-layer perceptron classifier to complete the multi-classification prediction task of anticancer peptides (ACP), antimicrobial peptides (AMP) and non-functional peptides.
[0018] The constructed datasets include gene prediction datasets and functional peptide identification datasets. The gene prediction dataset uses four datasets, specifically: The first dataset consists of 131 fully sequenced bacterial and archaeal genomes and 33 Gram-negative bacterial genomes randomly selected from the NCBI RefSeq database, which contains a total of 164 complete genomes in a 7:3 ratio for training and validation models; The second dataset consists of 10 complete genomes and is used to debug the model; The third dataset is an independent dataset containing the complete genomes of two archaea and seven bacteria, with each species containing thousands to tens of thousands of complete genome samples; The fourth dataset consisted of unannotated real genes from CAMI and Sharon, which were compared with the NCBI RefSeq database using blast. This yielded 100 highly similar new genomes, including Gram-like bacteria and Staphylococci. For ease of testing, dataset four was randomly divided into five subsets.
[0019] The functional peptide identification dataset utilizes three datasets: Dataset 1 and Dataset 2, two anticancer peptide datasets provided by the AntiCP 2.0 method, and Dataset 3, an antimicrobial peptide dataset provided by the C_AMP method. The model was trained and tested with an 8:2 ratio. These integrated datasets come from a wide range of sources, covering multiple recognized peptide databases, including DADP, CAMP, APD, APD2, UniProt, and Swiss-Prot. For the anticancer peptide prediction task, positive samples consist of experimentally validated anticancer peptides, compiled from a combination of the AMP and CancerPPD databases. Negative samples were selected from antimicrobial peptide sequences lacking anticancer activity in the AMP database to ensure they lack known anticancer properties.
[0020] In the antimicrobial peptide prediction task, to eliminate bias, peptides with only anticancer activity were not included in the AMP positive class to ensure that all samples had antimicrobial activity; non-AMP data were obtained from the UniProt database, which screened cytoplasmic localized proteins and excluded entries containing keywords such as "antibacterial, antiviral, and antifungal." After removing duplicates and excluding sequences identical to AMP, non-AMP sequences were finally retained.
[0021] Model evaluation includes gene prediction and functional peptide identification: The accuracy of gene prediction and functional peptide identification were compared with the benchmark or latest prediction methods. Figure 3 、 Figure 4 As shown, Figure 3 Each number on the horizontal axis represents a bacterial species, namely 1 for A. fulgidus, 2 for B. pseudomallei, 3 for B. subtilis, 4 for C. tepidum, 5 for E. coli, 6 for H. pylori, 7 for N. pharaonis, 8 for P. aeruginosa, and 9 for W. endosymbiont. Figure 4 Dataset1 and Dataset2 represent two data sets of anticancer peptides, and Dataset3 represents the antibacterial peptide dataset.
[0022] It should be noted that the prediction process is that the model reads the data file containing the data to be predicted, and finally outputs and saves the prediction results.
[0023] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement of the above embodiments made according to the technical essence of the present invention, within the spirit and principles of the present invention, without departing from the content of the technical solution of the present invention, shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A gene prediction and functional peptide identification method based on multi-feature fusion, characterized in that: It includes a gene prediction part and a functional peptide identification part, wherein the gene prediction part includes: The feature extraction module is used to extract the open reading frame information and coding sequence from the input genome sequence and extract the corresponding artificial features; Multi-feature fusion module, used to expand the sequence encoding results of ORF and artificial feature vectors and concatenate them into a combined vector; Convolutional neural network module, used to perform one-dimensional convolution operation on the combined vector to extract local structural features in the sequence; The Transformer encoder module is used to further model the CNN output results and extract long-range dependencies and global contextual semantic information; The classification prediction module, including the fully connected layer and the Softmax layer, is used to predict the probability of whether each ORF is a coding region; The protein translation and storage module is used to translate the ORFs predicted as coding regions into corresponding protein sequences and store them for subsequent biological function mining or analysis; The functional peptide recognition part includes: A peptide sequence representation module is used to receive and perform tokenization or tokenization on the original peptide sequence; The Bert module uses the BERT language model, which has been pre-trained in the field of biological sequences, to perform contextual semantic modeling on peptide sequences and output a semantic vector representation of each amino acid residue; CNN module, used to extract local structural features from BERT output; Bi-LSTM modules, used to process BERT outputs in parallel and capture global sequence dependencies; The feature fusion and classification module is used to fuse the output results of CNN and Bi-LSTM and input them into the multi-layer perceptron classifier to complete the multi-classification prediction task of anticancer peptides, antibacterial peptides and non-functional peptides.
2. A gene prediction and functional peptide identification method based on multi-feature fusion according to claim 1, characterized in that: The working process of the feature extraction module is: Step 1: Single codon usage, using vector Represents a fragment: , , Where, Indicates the The frequency of occurrence of a single codon in the ORF, Indicates the The number of codons in the ORF, represents the total number of all codons in the ORF; Step 2: Double codon usage, using vector Represents a fragment: , , Where, Indicates double codon The frequency of occurrence in ORF, Indicates double codon number in ORF; Step 3: Feature vector Expressed as: , , Where, represents a sequence window centered around a potential start codon, represents the one-hot encoding function, represents the normalization function, Indicates the prediction of whether there is a TIS translation start site in the sequence window; Step 4: ORF length, using feature vector Represents metagenomic fragments: , Where, Indicates the ORF length, Indicates the length of the fragment; Step 5: GC content, using vector Represents a sequence fragment: , , Where, Indicates the The GC content percentage of the sequence, represents the total number of sequences, Indicates the G content in the sequence, Indicates the content of C in the sequence, Indicates the total number of GCs in the ORF; Step 6: The basic base content, the content ratio of adenine, thymine, guanine and cytosine are respectively calculated using 、 、 、 express: , , , , , Where, Indicates the total number of bases in the sequence, represents the number of adenines in the sequence, represents the number of thymines in the sequence, Indicates the amount of guanine, represents the number of cytosines.
3. The method for gene prediction and functional peptide identification based on multi-feature fusion according to claim 1, characterized in that: The co-location process of the multi-feature fusion module is: The encoded ORF, single codon usage, double codon usage, ORF length, TIS, GC content, base content, and all the features of the tag are concatenated into a set of one-dimensional feature vectors to represent the input sequence fragment, which is expressed as: , Where, represents the ORF feature extraction vector, represents the single codon usage feature extraction vector, represents the dicodon usage feature extraction vector, represents the TIS feature extraction vector, Indicates the length feature extraction vector of ORF, represents the GC content feature extraction vector, represents the basic base content feature extraction vector, Indicates a label.
4. The method for gene prediction and functional peptide identification based on multi-feature fusion according to claim 1, characterized in that: The CNN module performs a one-dimensional convolution operation on the combined vector to extract local structural features in the sequence, specifically using two convolutional layers and two maximum pooling layers.
5. The method for gene prediction and functional peptide identification based on multi-feature fusion according to claim 1, characterized in that: The Transformer encoder module further models the CNN output results to extract long-distance dependencies and global contextual semantic information, including: using position encoding, multi-head attention mechanism, residual connection + layer normalization, feedforward fully connected network and residual connection + normalization layer.
6. The method for gene prediction and functional peptide identification based on multi-feature fusion according to claim 1, characterized in that: The peptide sequence representation module works as follows: each amino acid is converted into its corresponding numerical ID, using 26 tokens in the language vocabulary, and the numerical ID is used as the input of the peptide language model.
7. The method for gene prediction and functional peptide identification based on multi-feature fusion according to claim 1, characterized in that: The BERT module uses the BERT language model, which has been pre-trained in the field of biological sequences, to perform contextual semantic modeling on peptide sequences and output a semantic vector representation of each amino acid residue. Specifically, it selects OntoProtein as the base model and further fine-tunes it based on the working dataset.
8. The method for gene prediction and functional peptide identification based on multi-feature fusion according to claim 1, characterized in that: Four datasets were used to construct the gene prediction datasets, including: the first dataset consisted of 131 fully sequenced bacterial and archaeal genomes and 33 Gram bacterial genomes randomly selected from the NCBI RefSeq database, which contained a total of 164 complete genomes in a ratio of 7:3 for training and validation models; The second dataset consists of 10 complete genomes and is used to debug the model; The third dataset is an independent dataset containing two complete genomes of archaea and seven bacteria, with multiple complete genome samples for each species; The fourth dataset consisted of 100 new genomes of Gram-forming bacteria and Staphylococci, which were randomly divided into 5 subsets.
9. The method for gene prediction and functional peptide identification based on multi-feature fusion according to claim 1, characterized in that: The functional peptide identification dataset was constructed using three datasets, including two anticancer peptide datasets Dataset1 and Dataset2 provided by the AntiCP 2.0 method, and an antimicrobial peptide dataset Dataset3 provided by the C_AMP method. The model was trained and tested with an 8:2 ratio. The integrated datasets were derived from multiple recognized peptide databases, including: DADP, CAMP, APD, APD2, UniProt, and Swiss-Prot. In the anticancer peptide prediction task, positive samples consist of experimentally verified anticancer peptides, and the data is compiled by combining the AMP database and the CancerPPD database; negative samples are selected from antimicrobial peptide sequences that lack anticancer activity in the AMP database to ensure that they do not have known anticancer properties; In the antimicrobial peptide prediction task, to eliminate bias, peptides with only anticancer activity were excluded from the AMP positive class to ensure that all samples had antimicrobial activity. Non-AMP data were obtained from the UniProt database. Cytoplasmic-localized proteins were screened and entries containing the keywords "antibacterial, antiviral, and antifungal" were excluded. After removing duplicates and excluding sequences identical to AMPs, non-AMP sequences were finally retained.
10. The method for gene prediction and functional peptide identification based on multi-feature fusion according to claim 1, characterized in that: Open reading frames were extracted from the genomic sequence and truncated or padded to a fixed length of 700 base pairs; these fixed-length ORFs were one-hot encoded and artificial features were extracted simultaneously; Then, the encoded open reading frame sequence and the extracted artificial feature vector are expanded into one-dimensional arrays respectively, and these arrays are concatenated into a longer vector and input into the CNN model to further extract local features; Flatten the feature map output by CNN and then input it into the Transformer encoder to extract global features; The candidate open reading frames are predicted through the fully connected layer and the softmax layer. The predicted coding regions are translated into protein sequences and tokenized. Each amino acid is represented by its corresponding token ID, using special tags [CLS] and [SEP]. The tokenized sequence is then input into a pre-trained language model, which maps each amino acid into a vector representation containing rich biological semantics and contextual information. The BERT output is then fed into two feature extraction paths, the CNN module and the Bi-LSTM module, for local feature extraction and global dependency feature extraction. Finally, the outputs of CNN and Bi-LSTM are concatenated through concat, and the concatenated vector is input into the multi-layer perceptron for functional peptide prediction.
Citation Information
Patent Citations
Multi-modal fusion deep learning model and multifunctional bioactive peptide prediction method
CN116013404A
Antibacterial peptide prediction method based on BERT feature coding technology and deep learning combination model
CN117292749A
Method and system for predicting capacity of small open reading window coding polypeptide in non-coding RNA (Ribonucleic Acid)
CN118038995A
Virus identification method based on multi-head attention mechanism and graph isomorphic neural network
CN118506872A
Neuropeptide prediction method and system based on multi-modal features and twin network
CN118629516A