Nitrogen cycle function gene prediction method and system based on multi-feature fusion

By adopting a multi-character fusion method in the prediction of nitrogen circulation functional genes, combining the bidirectional long and short-term memory network and self-attention mechanism, the problem of difficult extraction of key information of protein sequences in the existing technology is solved, and more accurate prediction of nitrogen circulation functional genes is achieved.

CN120164528APending Publication Date: 2025-06-17UNIV OF JINAN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510234436.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing method for predicting nitrogen cycle functional genes has the problem that the bidirectional long and short-term memory network cannot highlight key information in protein sequences, and the lack of protein properties leads to structural incompleteness, which in turn leads to functional changes.

Method used

A multi-feature fusion method is used to extract context and global features in protein sequences through bidirectional long and short-term memory networks and self-attention mechanisms, and to fuse these features through deep neural networks to predict the nitrogen cycle functional genes.

Benefits of technology

It effectively solves the problem that the two-way long and short-term memory network cannot highlight key information in protein sequences, and provides information on the physicochemical properties of amino acids through manual characteristics, improving the accuracy of prediction and structural integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164528A_ABST
    Figure CN120164528A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information prediction, and provides a nitrogen cycle function gene prediction method and system based on multi-feature fusion, and the method comprises the steps: obtaining protein sequence data of a nitrogen metabolism gene; analyzing the protein sequence data, screening out protein sequence data with class imbalance, performing protein sequence data enhancement based on a pseudo mutation strategy, and forming a balanced protein sequence data set by a minority of class protein sequences after data enhancement and protein sequences without class imbalance; balanced protein sequence data sets with different lengths are processed based on a bidirectional long-short-term memory network, context features of amino acid residues are extracted, and a relationship between any two amino acid residues is constructed through a self-attention mechanism to extract global features of the amino acid residues in the protein sequence. The context features and the global features are fused through a deep neural network, the prediction probability of the gene family is calculated based on the fused features, and the nitrogen cycle function genes are predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of prediction information technology, and particularly relates to a nitrogen cycle functional gene prediction method and system based on multi-feature fusion. Background Art

[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] To systematically and standardly describe protein functions, Gene Ontology (GO) and Function Categories (FunCat) have been proposed. Currently, GO is one of the most widely accepted and commonly used protein function classification systems. The GO protein classification system covers various functions and locations of proteins in cells and organisms. GO can classify proteins into three main aspects: Molecular Function (MF), Biological Process (BP), and Cellular Component (CC); Molecular Function describes the specific functions exhibited by protein molecules, such as catalytic reactions and binding to other molecules; Biological Process describes events related to protein functions, such as metabolic pathways and cell signal transduction; Cellular Component describes the specific locations and structures where proteins are located in cells, such as the nucleus and cell membrane. Each GO is a functional label, and the process of predicting protein functions is the process of determining the labels possessed by proteins. Proteins in databases such as Uniprot, Ensembl, and InterPro are annotated with GO functional labels, which can conveniently provide functional annotations for protein sequences. The essence of protein function prediction lies in accurately determining the similarity in sequence, function, etc. between proteins with unknown functions and proteins with known functions. Therefore, the methods for protein function prediction can be divided into three main directions: methods for predicting functions based on protein sequences, methods for predicting functions based on protein structures, and methods for predicting functions based on protein interaction networks.

[0004] Function prediction based on protein sequences can use traditional BLAST alignment tools, which may be able to predict conserved nitrogen cycle functional genes, but may not be able to detect new genes or genes with low sequence identity to known nitrogen cycle functional genes. Tools based on machine learning algorithms such as hidden Markov models also have deficiencies in learning the high-level semantic and structural representation similarities of gene sequences, limiting the accuracy of prediction results.

[0005] A structure - function prediction method based on a convolutional neural network can be adopted to predict protein functions from the tertiary structure of the active sites of heme proteins, so as to study the relationship between structure and function. By converting the tertiary structure of the heme - binding site into the xy - plane and dividing the space into small cubic regions (voxels); however, it is only applicable to the prediction of a large number of proteins with unknown functions. The deep - learning algorithm for predicting the protein tertiary structure from the amino - acid sequence can accurately predict the structure of the heme - binding site in heme proteins. If the challenge of predicting the heme - binding site from the amino - acid sequence can be overcome, the function of proteins can be directly predicted by deep - learning methods using the amino - acid sequence of heme proteins.

[0006] Convolutional neural networks can be used to learn, extract, and integrate features such as protein sequences, protein domains, and PPI networks, and are used for protein function prediction; by using a sub - model based on convolutional neural networks to extract protein sequence information, PPI network information, and the features of protein domains; the general features of protein domains, such as type, quantity, and location information, etc., can be integrated; however, some longer - range sequence - related features will be ignored.

[0007] In summary, the existing nitrogen - cycle functional gene prediction has the following defects:

[0008] 1) Regarding the problem that the bidirectional long short - term memory network BiLSTM cannot highlight the key information in the protein sequence;

[0009] 2) The lack of protein properties leads to incomplete structures, which in turn causes functional changes. Summary of the Invention

[0010] To solve the above problems, the present invention proposes a nitrogen - cycle functional gene prediction method and system based on multi - feature fusion, which extracts the key information of protein sequences based on the multi - feature fusion method, complements the protein sequence structure, and predicts the nitrogen - cycle functional genes.

[0011] According to some embodiments, the first solution of the present invention provides a nitrogen - cycle functional gene prediction method based on multi - feature fusion, adopting the following technical solutions:

[0012] A nitrogen - cycle functional gene prediction method based on multi - feature fusion includes:

[0013] Obtain protein sequence data of nitrogen - containing metabolic genes;

[0014] Analyze the obtained protein sequence data and screen out the protein sequence data with class imbalance;

[0015] Perform protein sequence data augmentation based on the pseudo-mutation strategy on the selected protein sequence data with class imbalance. The protein sequences of the minority class after data augmentation and the protein sequence data without class imbalance constitute a balanced protein sequence dataset;

[0016] Based on the balanced protein sequence dataset of different lengths obtained by processing with a bidirectional long short-term memory network, extract the context features of amino acid residues in the protein sequence. Construct the relationship between any two amino acid residues extracted through the self-attention mechanism to extract the global features of amino acid residues in the protein sequence. Fuse the extracted context features and global features through a deep neural network to obtain fused features, and calculate the prediction probability of the gene family based on the obtained fused features to complete the prediction of nitrogen cycle functional genes based on multi-feature fusion.

[0017] As a further technical limitation, perform preprocessing of data merging and redundancy on the obtained protein sequence data of nitrogen metabolism genes, analyze the preprocessed protein sequence data, screen out the protein sequence data with class imbalance, and obtain the protein sequences of the minority class.

[0018] Furthermore, perform data augmentation on the protein sequence data with class imbalance to expand the dataset; the methods of data augmentation at least include random substitution, random deletion, random exchange, and random addition.

[0019] As a further technical limitation, convert each amino acid in the protein sequence into an embedding vector of a preset length through an embedding layer, input the obtained embedding vector into the bidirectional long short-term memory network, and obtain the context features of amino acid residues in the protein sequence through the bidirectional long short-term memory network.

[0020] As a further technical limitation, use the self-attention mechanism to calculate the similarity between each amino acid residue and other amino acid residues, obtain the weighted distribution of each amino acid residue, and capture the global information of amino acid residues in the protein sequence; process the sequence output from the previous layer of the bidirectional long short-term memory network through the next layer of the bidirectional long short-term memory network, and then strengthen the global features through the self-attention mechanism; perform global average pooling on the output values of each layer of the bidirectional long short-term memory network through the self-attention mechanism to extract the global features of amino acid residues in the protein sequence.

[0021] As a further technical limitation, the deep neural network adopts a multi-layer unsupervised neural network, uses the output features of the previous layer as the input of the next layer for feature learning, and through layer-by-layer feature mapping, maps the features of the existing space samples to another feature space to complete the feature transformation of non-linear mapping.

[0022] According to some embodiments, the second solution of the present invention provides a nitrogen cycle functional gene prediction system based on multi-feature fusion, adopting the following technical solution:

[0023] A nitrogen cycle functional gene prediction system based on multi-feature fusion, comprising:

[0024] An acquisition module, which is configured to acquire protein sequence data of nitrogen metabolism genes;

[0025] A screening module, which is configured to analyze the acquired protein sequence data and screen out the protein sequence data with class imbalance;

[0026] A dataset construction module, which is configured to perform protein sequence data augmentation based on a pseudo-mutation strategy on the screened protein sequence data with class imbalance, and the protein sequences of the minority classes after data augmentation and the protein sequences without class imbalance constitute a balanced protein sequence dataset;

[0027] A prediction module, which is configured to process the obtained balanced protein sequence dataset of different lengths based on a bidirectional long short-term memory network, extract the context features of amino acid residues in the protein sequence, construct the relationship between any two amino acid residues extracted through a self-attention mechanism to extract the global features of amino acid residues in the protein sequence, fuse the extracted context features and global features through a deep neural network to obtain fusion features, and calculate the prediction probability of the gene family based on the obtained fusion features to complete the prediction of nitrogen cycle functional genes based on multi-feature fusion.

[0028] According to some embodiments, the third solution of the present invention provides a computer-readable storage medium, adopting the following technical solution:

[0029] A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in a method for predicting nitrogen cycle functional genes based on multi-feature fusion as described in the first solution of the present invention.

[0030] According to some embodiments, the fourth solution of the present invention provides an electronic device, adopting the following technical solution:

[0031] An electronic device, comprising a memory, a processor, and a program stored on the memory and running on the processor, and when the processor executes the program, it implements the steps in a method for predicting nitrogen cycle functional genes based on multi-feature fusion as described in the first solution of the present invention.

[0032] According to some embodiments, the fifth solution of the present invention provides a computer program product, adopting the following technical solution:

[0033] A computer program product includes software code, and the program in the software code executes the steps in a nitrogen cycle functional gene prediction method based on multi-feature fusion as described in the first solution of the present invention.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] The present invention fuses a bidirectional long short-term memory network with a self-attention mechanism, automatically assigns weights according to the global context information in the protein sequence, so that the model can focus on the input sequence with discrimination and emphasis; further solves the problem that the bidirectional long short-term memory network cannot highlight the key information in the protein sequence;

[0036] The present invention provides information on the physicochemical properties of amino acids through manual features, and uses a deep neural network to extract and fuse the manually made features; further solves the problem that the structure is incomplete due to the lack of protein properties, which in turn leads to changes in function. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings forming a part of this embodiment are used to provide a further understanding of this embodiment. The schematic embodiments and descriptions thereof of this embodiment are used to explain this embodiment and do not constitute an improper limitation to this embodiment.

[0038] Figure 1 It is a flowchart of a nitrogen cycle functional gene prediction method based on multi-feature fusion in Embodiment 1 of the present invention;

[0039] Figure 2 It is a step diagram of a nitrogen cycle functional gene prediction method based on multi-feature fusion in Embodiment 1 of the present invention;

[0040] Figure 3 It is a schematic diagram of gene fragment screening in Embodiment 1 of the present invention;

[0041] Figure 4 It is a schematic diagram of the distribution of nitrogen cycle functional gene families in Embodiment 1 of the present invention;

[0042] Figure 5 It is a schematic diagram of a protein sequence generation method based on a pseudo-mutation strategy in Embodiment 1 of the present invention;

[0043] Figure 6 It is a schematic diagram of the distribution of nitrogen cycle functional gene families after data augmentation in Embodiment 1 of the present invention;

[0044] Figure 7 It is an overall architecture diagram of a nitrogen cycle functional gene prediction method based on multi-feature fusion in Embodiment 1 of the present invention;

[0045] Figure 8Schematic diagram of the bidirectional long short-term memory network framework structure in Embodiment 1 of the present invention;

[0046] Figure 9 Schematic diagram of the self-attention mechanism framework structure in Embodiment 1 of the present invention;

[0047] Figure 10 Schematic diagram of manual feature extraction in Embodiment 1 of the present invention;

[0048] Figure 11 Block diagram of a nitrogen cycle functional gene prediction system based on multi-feature fusion in Embodiment 2 of the present invention. Detailed implementation manners

[0049] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0050] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations for the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0051] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0052] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relationship terms determined for the convenience of describing the structural relationship of each component or element of the present invention and do not specifically refer to any component or element of the present invention. It should not be construed as a limitation to the present invention.

[0053] In the present invention, terms such as "fixed connection", "connected", "connection", etc. should be understood in a broad sense, indicating that it can be a fixed connection, an integral connection or a detachable connection; it can be directly connected or indirectly connected through an intermediate medium. For those related scientific research or technical personnel in the field, the specific meaning of the above terms in the present invention can be determined according to specific circumstances and should not be construed as a limitation to the present invention.

[0054] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0055] Term explanation

[0056] Nitrogen cycle functional genes refer to genes that play a key role in regulating the conversion and utilization of nitrogen elements in organisms. Nitrogen is one of the essential elements in organisms, but the forms and states of nitrogen elements often change and need to be converted and utilized through the nitrogen cycle process. Nitrogen cycle functional genes include a series of genes such as nitrogen fixation, amino acid synthesis, nitrification, and denitrification. They jointly participate in the conversion, utilization, and regeneration of nitrogen elements in organisms, maintaining the normal operation of the life system. Nitrogen cycle functional genes not only play an important role in the natural nitrogen cycle process but also have important application values in fields such as agricultural production and environmental protection.

[0057] Example 1

[0058] Example 1 of the present invention introduces a prediction method for nitrogen cycle functional genes based on multi-feature fusion.

[0059] As Figure 1 shown, a prediction method for nitrogen cycle functional genes based on multi-feature fusion includes:

[0060] Obtain protein sequence data of nitrogen-containing metabolic genes;

[0061] Analyze the obtained protein sequence data and screen out protein sequence data with class imbalance;

[0062] Perform protein sequence data enhancement based on the pseudo-mutation strategy on the screened protein sequence data with class imbalance. The enhanced minority-class protein sequences and the protein sequences without class imbalance form a balanced protein sequence dataset;

[0063] Based on the bidirectional long short-term memory network to process the obtained balanced protein sequence dataset of different lengths, extract the context features of amino acid residues in the protein sequence, construct the relationship between any two amino acid residues extracted through the self-attention mechanism to extract the global features of amino acid residues in the protein sequence, fuse the extracted context features and global features through a deep neural network to obtain fusion features, and calculate the prediction probability of the gene family based on the obtained fusion features to complete the prediction of nitrogen cycle functional genes based on multi-feature fusion.

[0064] As Figure 2As shown in the figure, in this embodiment, the blastx function in bioinformatics software is used to compare the genome containing nitrogen metabolism genes with the nitrogen cycle database. The sequence identity needs to be customized by the user (set to 85% in this embodiment). When the comparison result may be on the negative strand of DNA, reverse complementation is required (if it is on the positive strand, it is not required). Then, combined with the ORF finder bioinformatics tool, the potential protein-coding regions are found (that is, the DNA sequence in the previous step is converted into a protein sequence, that is, the amino acid sequence that can formally express the gene), and the obtained sequence is input into the model for prediction (automatically calculate the length of this sequence. If the amino acid residue length is within 0 - 100, DAbilstm_100 is called; if the amino acid residue length is within 100 - 300, DAbilstm_300 is called; if the amino acid residue length is within 300 - 500, DAbilstm_500 is called; if the amino acid residue length is within 500 - 1000, DAbilstm_1000 is called).

[0065] As Figure 3 shown, for some protein sequences in the obtained protein sequence data, there may be a situation where the starting site of the alignment region on the query sequence (Query ID) is greater than the ending site (as Figure 3 shown, the 7th column represents the starting site of the alignment region on the query sequence, and the 8th column represents the ending site of the alignment region on the query sequence); this indicates that this alignment region may be on the negative strand of DNA. Therefore, to ensure the accuracy of the direction, these sequences need to be processed by complementing the negative strand and uniformly converted into positive strand sequences.

[0066] In this embodiment, the nitrogen cycle-related gene database as Figure 4 shown is obtained. After merging and removing redundancy of the data, 118,086 sequences are obtained, with a total of 72 gene families; it is found that there is a phenomenon of class imbalance in the data analysis process.

[0067] In this embodiment, to address the problem of data class imbalance, a data augmentation method is used to expand the dataset. The protein sequence data augmentation method based on the pseudo-mutation strategy in this embodiment mainly includes the following operations:

[0068] (1) Random substitution: Randomly select n positions in the amino acid sequence and randomly replace the amino acids at the corresponding positions, as Figure 5 shown in (a) of the figure;

[0069] (2) Random deletion: Randomly select n positions in the amino acid sequence and delete the amino acids at the selected positions, as Figure 5 shown in (b) of the figure;

[0070] (3) Random exchange: Randomly select n pairs of amino acids in the amino acid sequence and exchange the positions of each pair of amino acids, as shown in (c) of Figure 5 ;

[0071] (4) Random addition: Randomly select n positions in the amino acid sequence and randomly add amino acids at the corresponding positions, as shown in (d) of Figure 5 ;

[0072] In this embodiment, the minimum sample number threshold is set (50% of the average number of all data). The number of categories below this threshold is enhanced through four methods: 25% random replacement, 25% random deletion, 25% random exchange, and 25% random addition to achieve the effect of balanced data. Finally, 141,216 sequences are divided into a training set of 98,850 sequences (70% of each category) and a test set of 42,366 sequences (30% of each category). The class distribution of the enhanced dataset is as shown in Figure 6 ;

[0073] As shown in Figure 7 , in this embodiment, a nitrogen cycle functional gene prediction model based on multi-feature fusion is constructed. The constructed model consists of a bidirectional long short-term memory network, a self-attention mechanism, and a deep neural network. Specifically:

[0074] (1) Bidirectional long short-term memory network

[0075] To process sequences of different lengths, in this embodiment, the sequence length is limited to 100. For sequences longer than 100, they are truncated to 100. For sequences shorter than 100, the existing amino acids are repeated until the sequence length reaches 100. Each amino acid is converted into a 100-dimensional embedding vector using an embedding layer. The vocabulary size of the embedding layer is set to 21, representing all possible amino acid types, and any unknown amino acid will be mapped to 0. The 100-dimensional embedding vector is input into the bidirectional long short-term memory network (BiLSTM) as shown in Figure 8 . The hidden layer dimension is set to 128 to extract the context features of amino acid residues.

[0076] The bidirectional long short-term memory network in this embodiment is an improved recurrent neural network (RNN), which is specifically designed to process sequential data. BiLSTM can capture bidirectional dependencies in the sequence by combining the outputs of two LSTM networks, one forward and one backward. Bi-LSTM is an extension of LSTM and involves two LSTMs running in parallel. The LSTM neural network generally consists of an input layer, a hidden layer, and an output layer. Bilstm includes an input gate, an output gate, a forget gate, and a memory cell block; specifically:

[0077] 1) Forward LSTM calculation:

[0078] (a) Forget gate calculation: f t = σ(W if x t + b if + W hf h t-1 + b hf );

[0079] After weighted summation of the current input and the previous hidden state and activation by Sigmoid, it determines the information to be forgotten in the previous cell state. Among them, f t represents the value of the forget gate at time t. The forget gate determines which information in the previous cell state will be forgotten; σ is the Sigmoid function, which is an activation function that can map the input value to between 0 and 1 and is often used in the gating mechanism to output a probability value representing the opening degree of the gate; W if is a weight matrix used to weight the input x t at time t; x t is the input vector at time t, which contains the information received by the network at the current moment; b if is the bias vector related to the input x t used to adjust the calculation result; W hf is another weight matrix used to weight the hidden state h t-1 at the previous time t - 1; h t-1 is the hidden state vector at the previous time t - 1, which contains the internal state information of the network at the previous moment; b hf is the bias vector related to the previous hidden state h t-1 used to adjust the calculation result.

[0080] (b) Input gate calculation: i t = σ(W ii x t + b ii + W hi ht-1 +b hi );

[0081] Determine the proportion of new information input into the cell state by weighted summation of the current input and the hidden state at the previous moment and activation by Sigmoid. Here, i t represents the value of the input gate at time t; σ is the Sigmoid function, which is an activation function that can map the input value to between 0 and 1 and is often used in the gating mechanism to output a probability value representing the opening degree of the gate; W ii is a weight matrix used to weight the input x t at time t; x t is the input vector at time t, which contains the information received by the network at the current moment; b ii is the bias vector related to the input x t for adjusting the calculation result; W hi is another weight matrix used to weight the hidden state h t-1 at the previous moment t - 1; h t-1 is the hidden state vector at the previous moment t - 1, which contains the internal state information of the network at the previous moment; b hi is the bias vector related to the hidden state h t-1 at the previous moment for adjusting the calculation result.

[0082] (c) Update unit state calculation: g t = tanh(W ig x t +b ig +W hg h t-1 +b hg );

[0083] Generate new information that may be added to the cell state by weighted summation of the current input and the hidden state at the previous moment and activation by hyperbolic tangent. Here, g t Generally in related architectures such as the Long Short-Term Memory network (LSTM), it represents the candidate cell state at time t (Candidate Cell State), which is the new information that may be added to the cell state; tanh is the hyperbolic tangent function, which is an activation function that maps the input value to between -1 and 1 and performs a non-linear transformation on the data; W ig is a weight matrix used to weight the input x t at time t; x t is the input vector at time t, which contains the information received by the network at the current moment; b ig is the bias vector related to the input x t for adjusting the calculation result; Whg is another weight matrix used to weight the hidden state h at the previous time step t - 1 t-1 ; h t-1 is the hidden state vector at the previous time step t - 1, which contains the internal state information of the network at the previous time step; b hg is the bias vector related to the hidden state h at the previous time step t-1 used to adjust the calculation result.

[0084] (d) Cell state calculation: c t = f t ⊙ c t-1 + i t ⊙ g t ;

[0085] Update the cell state by performing element-wise multiplication and summation on the cell state at the previous time step, the result of the forget gate, the input gate, and the candidate cell state result. Here, C t represents the cell state at time step t, which is the key carrier of information transmission in the long short-term memory network (LSTM) and is continuously updated at different time steps; C t-1 refers to the cell state at the previous time step t - 1, which carries the information of the previous time steps; ⊙ is element-wise multiplication, also known as the Hadamard product, which multiplies the corresponding elements of two vectors of the same dimension. f t and g t are as shown above.

[0086] (e) Output gate calculation: o t = σ(W io x t + b io + W ho h t-1 + b ho );

[0087] Determine the part of the cell state output by weighted summation of the current input and the hidden state at the previous time step and activation by the Sigmoid function. Here, O t represents the value of the output gate at time step t. The output gate is used to determine which part of the cell state will be output as the hidden state at the current time step; σ is the Sigmoid function, which is an activation function that can map the input value to the range between 0 and 1 and is often used in the gating mechanism to output a probability value representing the opening degree of the gate; W iO is a weight matrix used to weight the input x at time step t t ; x tis the input vector at time t, which contains the information received by the network at the current time; b iO is the bias vector related to the input x t , used to adjust the calculation result; W hO is another weight matrix, used to weight the hidden state h t-1 at the previous time t - 1; h t-1 is the hidden state vector at the previous time t - 1, which contains the internal state information of the network at the previous time; b hO is the bias vector related to the hidden state h t-1 at the previous time, used to adjust the calculation result.

[0088] (f) Hidden state update: h t = o t ⊙ tanh(c t );

[0089] The cell state is multiplied by the output gate result element after being activated by the hyperbolic tangent to obtain the hidden state at the current time. Among them, h t represents the hidden state at time t (Hidden State), which is one of the outputs of the network at the current time and will participate in the calculations at the next time and the final prediction and other tasks.

[0090] 2) Reverse LSTM calculation:

[0091] For the reverse LSTM, starting from the last element of the sequence, a calculation similar to the forward LSTM is performed in reverse time order to obtain the output hidden state h t ' at each time step.

[0092] (2) Self-attention mechanism

[0093] As Figure 9 shown, in this embodiment, the relationship between any two amino acid residues in the input sequence is modeled through the self-attention mechanism. Specifically, by calculating the similarity between each residue and other residues, the weighted distribution of each residue is obtained, so as to capture the global information in the sequence. Immediately afterwards, the second layer of BiLSTM processes the sequence output from the first layer of BiLSTM and further strengthens the global features by applying the self-attention mechanism again. Subsequently, the model applies global average pooling to the output of each layer of BiLSTM to extract global features.

[0094] This embodiment adopts a special attention mechanism, which allows the model to consider the relationship between each element in the sequence and all other elements when processing a sequence. This mechanism can help the model better understand the context information in the sequence, so as to process the sequence data more accurately; specifically:

[0095] (a) Calculate query, key, and value: For each element in the input sequence, calculate a query vector, a key vector, and a value vector; these vectors are obtained through a linear transformation of the learned weight matrices and the input elements; assume the input sequence is X = [x1, x2,..., x n and the weight matrices are W Q , W K , W V . For each element, Q i = W Q · x i , K i = W K · x i , V i = W V · x i , where Q i , K i , V i represent the query, key, and value vectors of the i-th element respectively.

[0096] (b) Calculate attention scores: For each pair of elements x i and x j , calculate an attention score, indicating the degree of attention of x i to x j . The attention score is obtained by taking the dot product of the query vector and the key vector, and then dividing by a scaling factor (usually the square root of the dimension of the key vector): score(Q i , K i ) = (Q i · K i ) / √d k , where d k is the dimension of the key vector.

[0097] (c) Calculate attention weights: Through the softmax function, convert the attention scores into values between 0 and 1 and sum to 1, thereby obtaining the attention weights W ij = softmax(score(Q i , K j ))

[0098] (d) Calculate the output: Multiply each element's value vector by its corresponding attention weight, and then sum to obtain the final output z i = ∑ n j=1 W ij V jThese weights are used to combine the input word vectors to generate a new context-related word vector. The resulting word vector contains not only the information of the current word, but also the information of its context, so as to better understand the meaning of each word in a specific context. This embodiment uses self-attention to extract 256-dimensional features.

[0099] (3) Deep Neural Networks

[0100] like Figure 10 As shown, this embodiment inputs the manually designed 631-dimensional features into the deep neural network (DNN), obtains 256-dimensional features after processing, and fuses them with the features extracted by BiLSTM. To avoid overfitting, this embodiment uses Dropout layers at multiple key positions; the fused features are classified through the fully connected layer, and the predicted probabilities of 72 gene families are output; through joint training, the local context information of amino acid residues, global dependency information and manual features can be effectively combined, thereby improving the accuracy and generalization ability of protein sequence classification.

[0101] The deep neural network (DNN) in this embodiment adopts a multi-layer unsupervised neural network, and uses the output features of the previous layer as the input of the next layer for feature learning. After layer-by-layer feature mapping, the features of the existing spatial samples are mapped to another feature space, so as to learn better feature expression for the existing input.

[0102] Deep neural networks have multiple nonlinear mapping feature transformations and can fit highly complex functions. In a typical three-layer fully connected neural network (input layer, hidden layer, and output layer), assume that: X = [x1, x2, ..., x n ] T : is the input vector; W ij [R] : The weight between the jth neuron in the Rth layer and the ith neuron in the R+1th layer; b i [R] : The bias term of the i-th neuron in the l-th layer; then the conversion from the first layer to the second layer of the network can be completed by the following formula: j [1] =∑ n i=1 (W ji [1] *x i )+b j [1] , then through the activation function f(z), the output of the second layer is obtained: h j [1] =f(z j [1])。Based on the same method applied to the relationship between the second layer and the third layer: z k [2] = ∑ m j=1 (W kj [2] *h j [1] ) + b k [2] , and apply the activation function again to obtain the final output: y k = f(z k [2] ).

[0103] This embodiment can use an activation function for de - linearization. Specifically:

[0104] (1) Sigmoid function: The value range is (0, 1), and it is a monotonically increasing function.

[0105] (2) Tanh function: The value range is (-1, 1), and it is a monotonically increasing function.

[0106] (3) ReLU function: F(x) = max(0, x). When X > 0, F(x) = x, and the function slope is 1; when X <= 0, F(x) = 0. It is not a monotonically increasing function and is not differentiable at x = 0.

[0107] This embodiment uses two fully - connected layers, and applies a ReLU (Rectified Linear Unit) activation function after the first fully - connected layer. The first layer converts 631 - dimensional hand - crafted features into 512 - dimensional feature vectors. Then, the activation function is applied to perform a non - linear transformation on each element, enabling the network to learn more complex non - linear transformations. The 512 - dimensional vector after ReLU activation is converted into 256 output features, and the torch.cat method is used to fuse the three features.

[0108] To verify the prediction results, this embodiment uses the following evaluation metrics to evaluate the prediction results, that is:

[0109] (1) Precision

[0110] Among all the samples determined as positive samples by the model, the proportion of samples that are indeed positive samples, that is

[0111] (2) Accuracy

[0112] Among all samples, the proportion of correctly classified samples, that is, the ratio of the number of correctly classified samples to the total number of samples, that is

[0113] (3) Recall Rate

[0114] It reflects the proportion of samples that are actually positive samples and can be correctly determined as positive samples by the model, that is

[0115] (4) F1 Score

[0116] To comprehensively consider the performance of precision and recall rate, this embodiment introduces the F1 score evaluation index, which is obtained by calculating the harmonic mean of precision and recall rate, aiming to balance the importance of the two; that is

[0117] Case Analysis

[0118] In this embodiment, the neural network is trained and evaluated through the PyTorch framework. All experiments are carried out on a platform equipped with an Nvidia GeForce RTX 3060 graphics card. To accelerate the convergence speed of the model and improve stability, the Adam optimizer is used to train the network, the label smoothing loss is used as the loss function (the smoothing factor is taken as 0.2), the learning rate is set to 0.001, each training batch contains 128 samples, and a total of 50 iterations of training are carried out.

[0119] The method in this embodiment is compared with PhageScanner-CNN, PhageScanner-RNN, and ProtICNN-BiLSTM. The comparison results are shown in Tables 1, 2, 3, and 4 respectively. The results show that the method in this embodiment shows the best performance on the same dataset, significantly superior to other advanced methods.

[0120] Table 1 Comparison Results of Different RNNs on the Same Test Set

[0121]

[0122] Table 2 Comparison Results of Adding Self-Attention Mechanism on the Same Test Set

[0123]

[0124]

[0125] Table 3 Comparison Results of Adding Manual Features on the Same Test Set

[0126]

[0127] Table 4 Comparison Results of Different Methods on the Same Test Set

[0128]

[0129] In this embodiment, a bidirectional long short-term memory network is fused with a self-attention mechanism to automatically assign weights according to the global context information in the protein sequence, enabling the model to focus on the input sequence in a discriminative and focused manner; further solving the problem that the bidirectional long short-term memory network cannot highlight the key information in the protein sequence.

[0130] In this embodiment, the physicochemical property information of amino acids is provided through manual features, and the manually crafted features are extracted and fused using a deep neural network; further solving the problem that the lack of protein properties leads to incomplete structures and thus functional changes.

[0131] Embodiment 2

[0132] Embodiment 2 of the present invention introduces a nitrogen cycle functional gene prediction system based on multi-feature fusion.

[0133] As Figure 11 shown, a nitrogen cycle functional gene prediction system based on multi-feature fusion includes:

[0134] An acquisition module configured to acquire protein sequence data of nitrogen metabolism genes;

[0135] A screening module configured to analyze the acquired protein sequence data and screen out the protein sequence data with class imbalance;

[0136] A dataset construction module configured to perform protein sequence data augmentation based on a pseudo-mutation strategy on the screened protein sequence data with class imbalance, and the augmented minority-class protein sequences and the protein sequences without class imbalance form a balanced protein sequence dataset;

[0137] A prediction module configured to process the obtained balanced protein sequence dataset of different lengths based on a bidirectional long short-term memory network, extract the context features of amino acid residues in the protein sequence, construct the relationship between any two amino acid residues extracted through a self-attention mechanism to extract the global features of amino acid residues in the protein sequence, fuse the extracted context features and global features through a deep neural network to obtain fusion features, and calculate the prediction probability of the gene family based on the obtained fusion features to complete the nitrogen cycle functional gene prediction based on multi-feature fusion.

[0138] The detailed steps are the same as those of a nitrogen cycle functional gene prediction method provided in Embodiment 1 and will not be elaborated here.

[0139] Embodiment 3

[0140] Embodiment 3 of the present invention provides a computer-readable storage medium.

[0141] A computer-readable storage medium stores a program thereon, and when the program is executed by a processor, it implements the steps in a method for predicting nitrogen cycle functional genes based on multi-feature fusion as described in Embodiment 1 of the present invention.

[0142] The detailed steps are the same as those of a method for predicting nitrogen cycle functional genes based on multi-feature fusion provided in Embodiment 1, and will not be elaborated herein.

[0143] Embodiment 4

[0144] Embodiment 4 of the present invention provides an electronic device.

[0145] An electronic device includes a memory, a processor, and a program stored on the memory and running on the processor. When the processor executes the program, it implements the steps in a method for predicting nitrogen cycle functional genes based on multi-feature fusion as described in Embodiment 1 of the present invention.

[0146] The detailed steps are the same as those of a method for predicting nitrogen cycle functional genes based on multi-feature fusion provided in Embodiment 1, and will not be elaborated herein.

[0147] Embodiment 5

[0148] Embodiment 5 of the present invention provides a computer program product.

[0149] A computer program product includes software code, and the program in the software code implements the steps in a method for predicting nitrogen cycle functional genes based on multi-feature fusion as described in Embodiment 1 of the present invention.

[0150] The detailed steps are the same as those of a method for predicting nitrogen cycle functional genes based on multi-feature fusion provided in Embodiment 1, and will not be elaborated herein.

[0151] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented in various computer languages, for example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.

[0152] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0153] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0155] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0156] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

[0157] The above description is only the preferred embodiments of this example and is not used to limit this example. For those skilled in the art, this example can have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this example shall be included within the protection scope of this example.

Claims

1. A method for predicting nitrogen cycle functional genes based on multi-feature fusion, characterized in that: include: Obtain protein sequence data for nitrogen metabolism genes; Analyze the obtained protein sequence data and screen out protein sequence data with class imbalance; The protein sequence data with class imbalance that has been screened out are enhanced by using a pseudo mutation strategy. The minority class protein sequences after data enhancement and the protein sequences without class imbalance constitute a balanced protein sequence data set. Based on the balanced protein sequence data sets of different lengths obtained by bidirectional long short-term memory network processing, the contextual features of amino acid residues in protein sequences are extracted, and the relationship between any two extracted amino acid residues is constructed through the self-attention mechanism to extract the global features of amino acid residues in protein sequences. The extracted contextual features and global features are fused through deep neural networks to obtain fused features. The prediction probability of gene families is calculated based on the obtained fused features, and the prediction of nitrogen cycle functional genes based on multi-feature fusion is completed.

2. A method for predicting nitrogen cycle functional genes based on multi-feature fusion as described in claim 1, characterized in that: The obtained protein sequence data of nitrogen-containing metabolic genes are subjected to data merging and redundant preprocessing, the preprocessed protein sequence data are analyzed, the protein sequence data with class imbalance are screened out, and the minority class protein sequences are obtained.

3. A method for predicting nitrogen cycle functional genes based on multi-feature fusion as described in claim 2, characterized in that: Data enhancement is performed on protein sequence data with class imbalance to expand the data set; the data enhancement method at least includes random replacement, random deletion, random exchange and random addition.

4. A method for predicting nitrogen cycle functional genes based on multi-feature fusion as described in claim 1, characterized in that: Each amino acid in the protein sequence is converted into an embedding vector of a preset length through an embedding layer, and the obtained embedding vector is input into the bidirectional long short-term memory network, and the contextual features of the amino acid residues in the protein sequence are obtained through the bidirectional long short-term memory network.

5. A method for predicting nitrogen cycle functional genes based on multi-feature fusion as claimed in claim 1, characterized in that: The self-attention mechanism is used to calculate the similarity between each amino acid residue and other amino acid residues, and the weighted distribution of each amino acid residue is obtained to capture the global information of amino acid residues in the protein sequence. The sequence output from the previous bidirectional long short-term memory network is processed by the next layer of bidirectional long short-term memory network, and the global features are strengthened by the self-attention mechanism. The output values ​​of each layer of the bidirectional long short-term memory network are globally averaged pooled through the self-attention mechanism to extract the global features of amino acid residues in the protein sequence.

6. A method for predicting nitrogen cycle functional genes based on multi-feature fusion as claimed in claim 1, characterized in that: The deep neural network adopts a multi-layer unsupervised neural network, takes the output features of the previous layer as the input of the next layer for feature learning, and maps the features of the existing spatial samples to another feature space through layer-by-layer feature mapping, completing the feature transformation of nonlinear mapping.

7. A nitrogen cycle functional gene prediction system based on multi-feature fusion, characterized in that: include: an acquisition module configured to acquire protein sequence data of nitrogen-containing metabolic genes; A screening module is configured to analyze the acquired protein sequence data and screen out protein sequence data with class imbalance; Constructing a dataset module, which is configured to perform protein sequence data enhancement based on a pseudo mutation strategy on the screened protein sequence data with class imbalance, wherein the minority class protein sequences and the protein sequences without class imbalance after data enhancement constitute a balanced protein sequence dataset; The prediction module is configured to extract the contextual features of amino acid residues in the protein sequence based on balanced protein sequence data sets of different lengths obtained by processing the bidirectional long short-term memory network, construct the relationship between any two extracted amino acid residues through the self-attention mechanism to extract the global features of the amino acid residues in the protein sequence, fuse the extracted contextual features and global features through a deep neural network to obtain a fused feature, calculate the prediction probability of the gene family based on the obtained fused feature, and complete the prediction of nitrogen cycle functional genes based on multi-feature fusion.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of a method for predicting nitrogen cycle functional genes based on multi-feature fusion as described in any one of claims 1 to 6 are implemented.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the steps of a nitrogen cycle functional gene prediction method based on multi-feature fusion as described in any one of claims 1-6 are implemented.

10. A computer program product comprising software code, characterized in that The program in the software code executes the steps of a method for predicting nitrogen cycle functional genes based on multi-feature fusion as described in any one of claims 1 to 6.