A key miRNA identification method based on multi-head self-attention mechanism and sequence
By combining the multi-head self-attention mechanism and sequence features with BiLSTM to extract temporal features and construct deep learning features, the problem of insufficient accuracy in key miRNA identification in existing technologies is solved, efficient identification is achieved, and a foundation is provided for disease research.
Patent Information
- Application Number
- CN202210838647.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-18
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-07-18
AI Technical Summary
Existing technologies lack predictive accuracy in identifying key miRNAs, especially due to insufficient utilization of the temporal characteristics of miRNA nucleotide sequences and deep learning technology, resulting in high human and financial costs and low research efficiency.
The multi-head self-attention mechanism is combined with sequence features. The temporal features are extracted through BiLSTM. The multi-head self-attention and weighted attention mechanisms are combined to construct miRNA deep learning features, which are input into the MLP model for key miRNA identification.
It improves the recognition accuracy of key miRNAs, saves manpower and financial costs, improves research efficiency, and provides a basic basis for subsequent research on miRNA-related diseases.
Smart Images

Figure CN115206432B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and in particular to a key miRNA identification method based on a multi-head self-attention mechanism and sequences. Background Art
[0002] MicroRNA (miRNA) is a class of endogenous small RNAs, approximately 20-24 nucleotides in length, that play a variety of important regulatory roles within cells. Each miRNA can have multiple target genes, and several miRNAs can regulate the same gene. Combinations of miRNAs can finely control gene expression. It is estimated that miRNAs regulate one-third of human genes. As a typical non-coding RNA, miRNAs (microRNAs) are approximately 22 nucleotides in length. Recent studies have demonstrated that miRNAs play a crucial role in key biological processes such as cell proliferation and differentiation. Furthermore, miRNAs can regulate target mRNAs by destabilizing them and inhibiting their translation. Recent research has shown that miRNAs are closely linked to numerous human diseases, particularly complex diseases such as cancer. Given the current importance of miRNAs to human disease and the time and cost constraints of traditional biomedical experiments, computational approaches to studying miRNAs have made significant progress. Current research on miRNAs primarily focuses on the relationship between miRNAs and diseases and the identification of key miRNAs. However, due to limited basic data, the identification of key miRNAs still leaves much room for improvement. MicroRNAs exist in various forms. The most primitive is pri-miRNA, which is approximately 300-1000 bases long. Pri-miRNAs undergo a single process to become pre-miRNAs, or microRNA precursors, approximately 70-90 bases long. Pre-miRNAs are then cleaved by the Dicer enzyme to become mature miRNAs, approximately 20-24 nucleotides long.
[0003] Currently, the identification of key miRNAs is mainly considered from the perspectives of biomedical experiments and computational methods.
[0004] (1) Biomedical experimental measurement methods
[0005] Traditional biomedical experimental methods primarily rely on gene editing and other techniques in animals to identify key miRNAs. These methods generally offer high accuracy, but they also have significant drawbacks. As we all know, there are numerous candidate key miRNAs, and validating them through such biomedical experiments requires significant human and material resources. However, the progress currently being made by researchers in this area provides an important foundation for building a basic database for identifying key miRNAs.
[0006] (2) Recognition methods based on computing technology
[0007] The rapid development of computational technology, especially the significant achievements in machine learning and deep learning in recent years, has provided a significant impetus for the study of miRNAs using computational methods. These methods transform the identification of key miRNAs into a binary classification problem. Based on the characteristics of the mature miRNA production process, that is, from pre-miRNA to splicing, the pre-miRNA and miRNA sequences are used to extract the key miRNA feature set and input it into the computational model to identify key miRNAs.
[0008] In the typical key miRNA identification method miES, the sequence and structural features of pre-miRNA and miRNA are integrated, and key miRNAs are ultimately identified through a logistic regression model. In addition, PESM is also a typical machine learning method for key miRNA identification, which extracts more sequence features than the miES method. In addition, in terms of prediction model selection, it uses XGBoost in the Gradient Boosting Machine (GBM) model to predict key miRNAs. PESM also further improves its prediction performance.
[0009] However, while these methods have made significant progress in identifying key miRNAs and achieved impressive predictive performance, providing fundamental guidance for basic biomedical experiments and laying the foundation for subsequent computational research on key miRNA identification, there is still room for improvement in both prediction accuracy and methodologies. For example, in miRNA feature extraction, current methods primarily focus on static statistical and structural features, while insufficiently considering the temporal characteristics of miRNA nucleotide sequences. Furthermore, previous methods primarily consider traditional machine learning methods in selecting prediction models, while insufficiently leveraging the unprecedented success of deep learning technologies. Therefore, given the importance of current key miRNA identification research and the shortcomings of computational methods, new methods are urgently needed to address these issues and improve their prediction accuracy. Summary of the Invention
[0010] The technical problem to be solved by the present invention is: to address the shortcomings of the existing technology, to provide a key miRNA identification method based on a multi-head self-attention mechanism and sequence, which can effectively identify key miRNAs and provide a basic basis for subsequent biomedical experiments on key miRNAs, while saving human and financial costs and improving research efficiency. In order to solve the above problems, the technical solution is as follows:
[0011] The present invention provides a key miRNA identification method based on a multi-head self-attention mechanism and sequence, the key miRNA identification method comprising the following steps:
[0012] Step S1, based on the characteristics that mature miRNAs are derived from pre-miRNA splicing and pre-miRNA has a stem-loop hairpin structure, construct the nucleotide sequence statistical structural feature F of pre-miRNA based on pre-miRNA and miRNA sequence information s ;
[0013] Step S2: Based on the temporal characteristics of the nucleotide sequence of the pre-miRNA sequence, construct the miRNA deep learning initial feature F based on BiLSTM do ;
[0014] Step S3, miRNA deep learning initial feature F based on multi-head self-attention mechanism and BiLSTM do Obtain miRNA multi-head self-attention feature F dm ;
[0015] Step S4, statistically analyzing the structural features F of the nucleotide sequence according to the nucleotide subsequences in the pre-miRNA sequence s The importance of weighted attention mechanism is used to calculate the final feature representation F of miRNA deep learning. d ;
[0016] Step S5: Calculate the structure features F of the acquired miRNA nucleotide sequence s and miRNA deep learning final feature representation F d The final miRNA signature F was obtained by splicing f And input into the MLP model to identify key miRNAs based on the existing key miRNA sample basic data set.
[0017] Furthermore, in step S1, the nucleotide sequence statistical structural feature F s The calculation steps are:
[0018] Step S101, based on the mature miRNA generation process, a pre-miRNA that can generate miRNA is divided into two parts: (1) mature miRNA, and (2) non-mature miRNA remaining after removing the mature miRNA sequence from the pre-miRNA sequence; a total of 38 dimensions of miRNA statistical structural features are obtained;
[0019] Step S102, calculating the statistics of three types of basic nucleotides, U, C, and G, in the pre-miRNA sequence, miRNA sequence, and non-mature miRNA sequence to obtain three feature representations with three dimensions; in addition, obtaining two feature representations with one dimension on the miRNA and non-mature miRNA sequences respectively;
[0020] Step S103: Calculate a feature representation with a dimension of 1 for the splicing sites in the miRNA production process. The specific rules are: 1: all cleavage sites generated by the miRNA are U, 0: not all cleavage sites are U, -1: all cleavage sites are non-U;
[0021] Step S104: For nucleotide types U, C, and G, the frequency statistics of their nine combinations are calculated on the pre-miRNA and miRNA sequences, respectively, to obtain two features with nine dimensions. Based on the stem-loop hairpin structure of the pre-miRNA, a feature with two dimensions is calculated using the minimum free energy of the pre-miRNA, which includes the minimum free energy and the average minimum free energy based on the sequence length.
[0022] Step S105, based on the base pairing attributes, three base pairing attributes and three base pairing attribute features based on the nucleotide sequence length are calculated, and the feature dimension obtained is 6;
[0023] An overview of the statistical structural features of miRNA sequences is shown in Table 1:
[0024]
[0025]
[0026] Furthermore, in step S2, the initial features of miRNA deep learning based on BiLSTM are Where L is the length of the initial embedding of the sequence after BiLSTM processing, d b To obtain the feature dimension, the specific calculation process is:
[0027] Step S201, considering the temporal characteristics of the nucleotide sequence and the successful application of the k-mer model, the nucleotide sequence of the pre-miRNA is initialized using the 3-mer method, and the pre-miRNA sequence S = s1, s2, s3, s4..., s |S| , the features expressed by 3-mer are initially represented as:
[0028] X 1,2,3 ,X 2,3,4 ,...,X |S|-2,|S|-1,|S|
[0029] Among them, is the initialization feature representation of the i-th 3-gram; taking the sequence "AUGGUCC" as an example, its 3-mer sequence representation is "AUG,UGG,GGU,GUC,UCC".
[0030] In step S202, embedding features are initialized based on the pre-miRNA nucleotide sequence and input into the BiLSTM. The BiLSTM combines the forward and backward hidden layers to access the historical and future sequence information of the pre-miRNA. LSTM uses short-term memory as a hidden unit, effectively solving the problems of vanishing and exploding gradients, and plays a key role in extracting complex text information in the field of natural language.
[0031] Step S203: After t iterations of the BiLSTM network, the initial features of miRNA deep learning are obtained. The initialization characteristic process is where d b =128.
[0032] Furthermore, in step S3, miRNA multi-head self-attention feature F dm The acquisition process is as follows:
[0033] Step S301: Map three matrices into one query (Qurey, Q), and express the initial feature F of deep learning in the form of key (Key, K)-value (Value, V) pairs. do , these three mapping matrices are learned through back propagation;
[0034] In step S302, the multi-head attention mechanism works as follows: First, h sets of different linear projections are independently learned to transform the query, key, and value. These h sets of transformed queries, keys, and values are then pooled in parallel. Finally, the outputs of these h attention pools are concatenated and transformed using another learned linear projection to produce the final output. Figure 3 Its basic process is described in
[15] , where the scaling dot product attention mechanism (single attention mechanism) process is defined as follows:
[0035] The feature representations of Q, K, and V of each nucleotide sequence are obtained through BiLSTM, where the dimensions of the Q and K matrices are d k , the matrix dimension of V is d v , calculate the dot product of each Q and K, and then divide by Finally, the weight value is obtained through the softmax function, and its output matrix is:
[0036]
[0037] The calculation process of multi-head attention is as follows:
[0038] MultiHead(Q,K,V)=Concat(head1,...,head h )W o
[0039] where head1=Attention(QW i Q ,KW i k ,VW i v )
[0040] in and They are mapping parameter matrices, and the specific parameters are set to d input =d b =128,d k =d v =64,d output =38, where h=4, and the multi-head self-attention feature F of miRNA is obtained through the multi-head attention mechanism process. dm ={F dm1 ,F dm2 ,....,F dm|L|}.
[0041] Furthermore, in step S4, the miRNA deep learning final feature representation F based on the weighted attention mechanism is dThe process of obtaining is as follows:
[0042] Step S401: Statistical structural features F of nucleotide sequences obtained based on pre-miRNA and miRNA sequences s Based on the s Compared with the miRNA deep learning initial feature F based on multi-head self-attention mechanism and BiLSTM do Obtain miRNA multi-head self-attention feature F dm The weighted attention mechanism is used to obtain the final feature representation F of miRNA deep learning d ;
[0043] Step S402: Considering the representation vector F of the statistical structural characteristics of the miRNA sequence s And the implicit vector set of the pre-miRNA sequence based on the multi-head self-attention mechanism, the subsequences in the sequence are considered based on the weight to obtain the final feature representation F of the deep learning of miRNA d , that is, calculating the importance of subsequences in the pre-miRNA sequence to the representation of statistical structural features, giving greater weight to important subsequences. The calculation process is as follows:
[0044] h m =f(W inter F s +b inter ),
[0045] h i =f(W inter F dmi +b inter ),
[0046]
[0047] Among them, W inter and b inter are weight matrix and bias vector respectively, α i is the attention weight value, which represents the degree of association between a pre-miRNA subsequence and its statistical feature representation;
[0048] Step S403: Based on the above weights, the final feature F of miRNA deep learning is obtained. d The method of obtaining it is as follows:
[0049]
[0050] Furthermore, in step S5, the key miRNA identification process is as follows:
[0051] Step S501: Obtain the 38-dimensional statistical structure features F sAnd the final feature F of deep learning d Splicing is performed to obtain the final miRNA signature F f ∈R d , d = 76;
[0052] In step S502, the final miRNA feature vector is input into the MLP classification model to identify key miRNAs. In the iterative process of MLP:
[0053] h t =f h (W h h t-1 +b h ),
[0054] Among them, W h and b h are the weight matrix and bias vector respectively;
[0055] In step S503, the key miRNA identification problem is a binary classification problem. The specific definition of the miRNA feature vector obtained by MLP is as follows:
[0056] z=w output y con +b output ,
[0057] Among them, w output ∈R m×2d and b output ∈R m are the weight matrix and bias vector, respectively. When predicting key miRNAs, m is set to m = 2;
[0058] Step S504, based on the output feature vector z = [o0, o1], the specific process definition of calculating the key miRNA possibility score through the softmax function is:
[0059]
[0060] Among them, l∈{0,1} is the label, p l is the probability value of label l; therefore, the key miRNA is predicted by the softmax function;
[0061] In step S505, cross-entropy loss is used as the loss function, which is defined as follows:
[0062]
[0063] Among them, y i,l and p i,lare the one-hot vector representations of a single sample on the label. If the i-th sample belongs to l label types, then y i,l =1, otherwise y i,l =0; N is the number of key miRNA samples in the training set, and the training objective is to minimize the following loss function l:
[0064]
[0065] Among them, θ is the set of all weight matrices and bias vectors, and the parameter λ is the L2 regularization hyperparameter.
[0066] The beneficial effects of the prediction method provided by the present invention are:
[0067] The present invention proposes a key miRNA identification method based on a multi-head attention mechanism and sequence to identify key miRNAs. First, based on the fact that mature miRNAs are produced by pre-miRNA splicing and the stem-loop structure characteristics of pre-miRNAs, the statistical structural characteristics of the miRNA sequence are calculated. Next, further considering the temporal characteristics of miRNAs, Bi-directional Long Short-Term Memory (BiLSTM) is used to extract the initial miRNA deep learning features based on the pre-miRNA sequence. Then, based on the miRNA deep learning initial features, the miRNA multi-head self-attention features are obtained based on the multi-head self-attention mechanism. Based on the importance of miRNA subsequences to the miRNA statistical structural features, the deep learning features of miRNAs are obtained based on the statistical structural features and multi-head attention features using a weighted attention mechanism. Finally, the final deep learning features and statistical structural features are spliced and input into the MLP (Multilayer Perceptron) model to identify key miRNAs.
[0068] However, the present invention differs from previous prediction methods in that, based on the existing statistical structural features and the temporal characteristics of the sequence, it calculates sequence features using BiLSTM. Furthermore, on this basis, it further extracts features of the multi-head self-attention mechanism. Furthermore, based on the varying importance of mi subsequences to the statistical structural features, a weighted attention mechanism is employed to obtain the final deep learning features of the miRNA. Finally, the miRNA statistical structural features and the final deep learning features are concatenated and input into the MLP classification model. The likelihood scores of key miRNAs are calculated using a binary classification problem to obtain the final key miRNA prediction results.
[0069] The present invention provides a new method for predicting key miRNAs, which is helpful for the subsequent understanding of the pathogenic mechanism of miRNA-related diseases, diagnosis and treatment of diseases, and drug development. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0071] Figure 1 This is the overall flow chart of the key miRNA identification method based on multi-head self-attention mechanism and sequence of the present invention.
[0072] Figure 2 Schematic diagram of the BiLSTM model of the present invention.
[0073] Figure 3 Schematic diagram of the multi-head self-attention mechanism of the present invention.
[0074] Figure 4 It is the AUC graph of the present invention in five-fold cross validation. DETAILED DESCRIPTION
[0075] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention and to make the above-mentioned objects, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are further described below with reference to the accompanying drawings.
[0076] It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0077] Please refer to Figures 1 to 4 In this embodiment, a key miRNA identification method based on a multi-head self-attention mechanism and sequence is provided. The key miRNA identification method includes the following steps:
[0078] Step S1, based on the characteristics that mature miRNAs are derived from pre-miRNA splicing and pre-miRNA has a stem-loop hairpin structure, construct the nucleotide sequence statistical structural feature F of pre-miRNA based on pre-miRNA and miRNA sequence information s ;
[0079] Step S2: Based on the temporal characteristics of the nucleotide sequence of the pre-miRNA sequence, construct the miRNA deep learning initial feature F based on BiLSTM do ;
[0080] Step S3, miRNA deep learning initial feature F based on multi-head self-attention mechanism and BiLSTM do Obtain miRNA multi-head self-attention feature F dm ;
[0081] Step S4, statistically analyzing the structural features F of the nucleotide sequence according to the nucleotide subsequences in the pre-miRNA sequence s The importance of weighted attention mechanism is used to calculate the final feature representation F of miRNA deep learning. d ;
[0082] Step S5: Calculate the structure features F of the acquired miRNA nucleotide sequence s and miRNA deep learning final feature representation F d The final miRNA signature F was obtained by splicing f And input into the MLP model to identify key miRNAs based on the existing key miRNA sample basic data set.
[0083] Specifically, in step S1, the nucleotide sequence statistical structural feature F s The calculation steps are:
[0084] Step S101, based on the mature miRNA generation process, a pre-miRNA that can generate miRNA is divided into two parts: (1) mature miRNA, and (2) non-mature miRNA remaining after removing the mature miRNA sequence from the pre-miRNA sequence; a total of 38 dimensions of miRNA statistical structural features are obtained;
[0085] Step S102, calculating the statistics of three types of basic nucleotides, U, C, and G, in the pre-miRNA sequence, miRNA sequence, and non-mature miRNA sequence to obtain three feature representations with three dimensions; in addition, obtaining two feature representations with one dimension on the miRNA and non-mature miRNA sequences respectively;
[0086] Step S103: Calculate a feature representation with a dimension of 1 for the splicing sites in the miRNA production process. The specific rules are: 1: all cleavage sites generated by the miRNA are U, 0: not all cleavage sites are U, -1: all cleavage sites are non-U;
[0087] Step S104: For nucleotide types U, C, and G, the frequency statistics of their nine combinations are calculated on the pre-miRNA and miRNA sequences, respectively, to obtain two features with nine dimensions. Based on the stem-loop hairpin structure of the pre-miRNA, a feature with two dimensions is calculated using the minimum free energy of the pre-miRNA, which includes the minimum free energy and the average minimum free energy based on the sequence length.
[0088] Step S105, based on the base pairing attributes, three base pairing attributes and three base pairing attribute features based on the nucleotide sequence length are calculated, and the feature dimension obtained is 6;
[0089] An overview of the statistical structural features of miRNA sequences is shown in Table 1:
[0090]
[0091]
[0092] In step S2, the initial features of miRNA deep learning based on BiLSTM Where L is the length of the initial embedding of the sequence after BiLSTM processing, d b To obtain the feature dimension, the specific calculation process is:
[0093] Step S201: Considering the temporal characteristics of nucleotide sequences and the successful application of the k-mer model, the 3-mer method is used to initialize the features of the pre-miRNA nucleotide sequence. Taking the sequence "AUGGUCC" as an example, its 3-mer sequence is represented as "AUG, UGG, GGU, GUC, UCC". For a pre-miRNA sequence of length |S|, S = s1, s2, s3, s4..., s |S| , the features expressed by 3-mer are initially represented as:
[0094] X 1,2,3 ,X 2,3,4 ,...,X |S|-2,|S|-1,|S|
[0095] Among them, is the initialization feature representation of the i-th 3-gram;
[0096] Step S202, based on the pre-miRNA nucleotide sequence, the embedding features are initialized and input into BiLSTM. LSTM uses short-term memory as a hidden unit, effectively solving the problem of gradient vanishing and gradient exploding. It plays a key role in extracting complex text information in the field of natural language. BiLSTM is developed on LSTM, which combines forward and backward hidden layers. It not only accesses the historical sequence of pre-miRNA, but also accesses future sequence information. The processing process of BiLSTM is as follows Figure 2 As shown, the final output depends on two hidden layers S and S' that are forward and backward propagated respectively;
[0097] Step S203: After t iterations of the BiLSTM network, the initial features of miRNA deep learning are obtained. The initialization characteristic process is where d b =128.
[0098] In step S3, miRNA multi-head self-attention feature F dm The acquisition process is as follows:
[0099] Step S301: Map three matrices into one query (Qurey, Q), and express the initial feature F of deep learning in the form of key (Key, K)-value (Value, V) pairs. do , these three mapping matrices are learned through back propagation;
[0100] In step S302, the multi-head attention mechanism works as follows: First, h sets of different linear projections are independently learned to transform the query, key, and value. These h sets of transformed queries, keys, and values are then pooled in parallel. Finally, the outputs of these h attention pools are concatenated and transformed using another learned linear projection to produce the final output. Figure 3 Its basic process is described in
[15] , where the scaling dot product attention mechanism (single attention mechanism) process is defined as follows:
[0101] The feature representations of Q, K, and V of each nucleotide sequence are obtained through BiLSTM, where the dimensions of the Q and K matrices are d k , the matrix dimension of V is d v , calculate the dot product of each Q and K, and then divide by Finally, the weight value is obtained through the softmax function, and its output matrix is:
[0102]
[0103] The calculation process of multi-head attention is as follows:
[0104] MultiHead(Q,K,V)=Concat(head1,...,head h )W o
[0105] where head1=Attention(QW i Q ,KW i k ,VW i v )
[0106] in and They are mapping parameter matrices, and the specific parameters are set to d input =d b =128,d k =d v =64,d output =38, where h=4, and the multi-head self-attention feature F of miRNA is obtained through the multi-head attention mechanism process. dm ={F dm1 ,F dm2 ,....,F dm|L|}.
[0107] In step S4, the final feature representation F of miRNA deep learning based on the weighted attention mechanism is d The process of obtaining is as follows:
[0108] Step S401: Statistical structural features F of nucleotide sequences obtained based on pre-miRNA and miRNA sequences s Based on the s Compared with the miRNA deep learning initial feature F based on multi-head self-attention mechanism and BiLSTM do Obtain miRNA multi-head self-attention feature F dm The weighted attention mechanism is used to obtain the final feature representation F of miRNA deep learning d ;
[0109] Step S402: Considering the representation vector F of the statistical structural characteristics of the miRNA sequence s And the implicit vector set of the pre-miRNA sequence based on the multi-head self-attention mechanism, the subsequences in the sequence are considered based on the weight to obtain the final feature representation F of the deep learning of miRNA d , that is, calculating the importance of subsequences in the pre-miRNA sequence to the representation of statistical structural features, giving greater weight to important subsequences. The calculation process is as follows:
[0110] h m =f(W inter F s +b inter ),
[0111] h i =f(W inter F dmi +b inter ),
[0112]
[0113] Among them, W inter and b inter are weight matrix and bias vector respectively, α i is the attention weight value, which represents the degree of association between a pre-miRNA subsequence and its statistical feature representation;
[0114] Step S403: Based on the above weights, the final feature F of miRNA deep learning is obtained. d The method of obtaining it is as follows:
[0115]
[0116] In step S5, the key miRNA identification process is as follows:
[0117] Step S501: Obtain the 38-dimensional statistical structure features F s And the final feature F of deep learning d Splicing is performed to obtain the final miRNA signature F f ∈R d , d = 76;
[0118] In step S502, the final miRNA feature vector is input into the MLP classification model to identify key miRNAs. In the iterative process of MLP:
[0119] h t =f h (W h h t-1 +b h ),
[0120] Among them, W h and b h are the weight matrix and bias vector respectively;
[0121] In step S503, the key miRNA identification problem is a binary classification problem. The specific definition of the miRNA feature vector obtained by MLP is as follows:
[0122] z=woutput y con +b output ,
[0123] Among them, w output ∈R m×2d and b output ∈R m are the weight matrix and bias vector, respectively. When predicting key miRNAs, m is set to m = 2;
[0124] Step S504, based on the output feature vector z = [o0, o1], the specific process definition of calculating the key miRNA possibility score through the softmax function is:
[0125]
[0126] Among them, l∈{0,1} is the label, p l is the probability value of label l; therefore, the key miRNA is predicted by the softmax function;
[0127] In step S505, cross-entropy loss is used as the loss function, which is defined as follows:
[0128]
[0129] Among them, y i,l and p i,l are the one-hot vector representations of a single sample on the label. If the i-th sample belongs to l label types, then y i,l =1, otherwise y i,l =0; N is the number of key miRNA samples in the training set, and the training objective is to minimize the following loss function l:
[0130]
[0131] Among them, θ is the set of all weight matrices and bias vectors, and the parameter λ is the L2 regularization hyperparameter.
[0132] In order to evaluate the predictive performance of the present invention, the prediction performance evaluation indicators were calculated based on AUC (areas under ROC curves), F1-score and ACC (accuracy) through a five-fold cross-validation method. During the five-fold cross-validation experiment, the 77 key miRNA positive samples that have been experimentally verified in the benchmark data set and the same number of randomly selected secondary samples were randomly divided into 5 parts. One of the groups was selected as the test set in turn, and the other 4 groups were used as the test set. The final average value was obtained as the prediction result. Similarly, we also adopted the same verification method and the same evaluation indicators for the comparative calculation method.
[0133] Table 2 Prediction performance of the present invention and other methods
[0134]
[0135] Table 2 compares the proposed method with other methods in a five-fold cross-validation test. Table 2 shows that the proposed method outperforms other methods in terms of AUC, ACC, and F1-score, reaching 0.9557, 0.8966, and 0.8927, respectively. These performances are superior to the best comparison method, PESM (0.9117, 0.0.8516, and 0.8572).
[0136] Figure 4 The AUC plot for the present invention and the comparative method in a five-fold cross-validation analysis is shown. The plot shows that the present invention achieved the highest AUC value (0.9557). The comparative methods achieved AUC values of 0.9147, 0.8837, 0.8720, and 0.8571, respectively. Therefore, based on the AUC values, the present invention achieved the best prediction results.
[0137] The cross-validation test experiments and the comparison with the other four calculation methods have proved that the present invention can more effectively identify key miRNAs, and can also provide important help for the subsequent understanding of the pathogenic mechanism of miRNA-related diseases and the development of drugs related to their diagnosis and treatment.
[0138] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0139] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. It will be apparent to those skilled in the art that various changes, modifications, substitutions, and variations made to these embodiments without departing from the principles and spirit of the present invention are still within the scope of protection of the present invention.
Claims
1. A key miRNA identification method based on multi-head self-attention mechanism and sequence, characterized by: The key miRNA identification method comprises the following steps: Step S1, based on the characteristics that mature miRNAs are derived from pre-miRNA splicing and pre-miRNA has a stem-loop hairpin structure, the nucleotide sequence statistical structure feature F of pre-miRNA is constructed based on the pre-miRNA and miRNA sequence information. s ; Step S2: Based on the temporal characteristics of the nucleotide sequence of the pre-miRNA sequence, construct the miRNA deep learning initial feature F based on BiLSTM do ; Step S3, miRNA deep learning initial feature F based on multi-head self-attention mechanism and BiLSTM do Obtain miRNA multi-head self-attention feature F dm ; Step S4, statistically analyzing the structural features F of the nucleotide sequence according to the nucleotide subsequences in the pre-miRNA sequence s The importance of weighted attention mechanism is used to calculate the final feature representation F of miRNA deep learning. d ; Based on the statistical structural features of nucleotide sequences obtained from pre-miRNA and miRNA sequences F s Based on the s Compared with the miRNA deep learning initial feature F based on multi-head self-attention mechanism and BiLSTM do Obtain miRNA multi-head self-attention feature F dm The weighted attention mechanism is used to obtain the final feature representation F of miRNA deep learning d ; Step S5: Calculate the structure features F of the acquired miRNA nucleotide sequence s and miRNA deep learning final feature representation F d The final miRNA signature F was obtained by splicing f And input into the MLP model to identify key miRNAs based on the existing key miRNA sample basic data set.
2. The key miRNA identification method according to claim 1, characterized in that In step S1, the nucleotide sequence statistical structural feature F s The calculation steps are: Step S101, based on the mature miRNA generation process, a pre-miRNA that can generate miRNA is divided into two parts: (1) mature miRNA, and (2) non-mature miRNA remaining after removing the mature miRNA sequence from the pre-miRNA sequence; a total of 38 dimensions of miRNA statistical structural features are obtained; Step S102: Calculate the statistics of three types of basic nucleotides, U, C, and G, in the pre-miRNA sequence, miRNA sequence, and non-mature miRNA sequence to obtain three feature representations with three dimensions; in addition, obtain two feature representations with one dimension on the miRNA and non-mature miRNA sequences respectively; Step S103: Calculate a feature representation with a dimension of 1 for the splicing sites in the miRNA production process. The specific rules are: 1: all cleavage sites generated by the miRNA are U, 0: not all cleavage sites are U, -1: all cleavage sites are non-U; Step S104: For nucleotide types U, C, and G, the frequency statistics of their nine combinations are calculated on the pre-miRNA and miRNA sequences, respectively, to obtain two features with nine dimensions. Based on the stem-loop hairpin structure of the pre-miRNA, a feature with two dimensions is calculated using the minimum free energy of the pre-miRNA, which includes the minimum free energy and the average minimum free energy based on the sequence length. Step S105, based on the base pairing attributes, three base pairing attributes and three base pairing attribute features based on the nucleotide sequence length are calculated, and the feature dimension obtained is 6; An overview of the statistical structural features of miRNA sequences is shown in Table 1:
3. The key miRNA identification method according to claim 2, characterized in that In step S2, the initial features of miRNA deep learning based on BiLSTM Where L is the length of the initial embedding of the sequence after BiLSTM processing, d b To obtain the feature dimension, the specific calculation process is: Step S201, considering the temporal characteristics of the nucleotide sequence and the successful application of the k-mer model, the nucleotide sequence of the pre-miRNA is initialized using the 3-mer method, and the pre-miRNA sequence S = s1, s2, s3, s4..., s |S| , the features expressed by 3-mer are initially represented as: X 1,2,3 ,X 2,3,4 ,...,X |S|-2,|S|-1,|S| Among them, is the initialization feature representation of the i-th 3-gram; Step S202: Initialize embedding features based on the pre-miRNA nucleotide sequence, input the initialized embedding features into the BiLSTM, and the BiLSTM combines the forward and backward hidden layers to access the historical and future sequence information of the pre-miRNA; Step S203: After t iterations of the BiLSTM network, the initial features of miRNA deep learning are obtained. The initialization characteristic process is where d b =128.
4. The key miRNA identification method according to claim 3, characterized in that In step S3, miRNA multi-head self-attention feature F dm The acquisition process is as follows: Step S301: Map three matrices into one query (Qurey, Q), and express the initial feature F of deep learning in the form of key (Key, K)-value (Value, V) pairs. do , these three mapping matrices are learned through back propagation; In step S302, the multi-head attention mechanism process is as follows: first, based on independent learning, h sets of different linear projections are obtained to transform the query, key, and value; then, these h sets of transformed queries, keys, and values are parallelly attention-pooled; finally, the outputs of these h attention pools are concatenated and transformed by another learnable linear projection to produce the final output; wherein the scaled dot product attention mechanism process is defined as follows: The feature representations of Q, K, and V of each nucleotide sequence are obtained through BiLSTM, where the dimensions of the Q and K matrices are d k , the matrix dimension of V is d v , calculate the dot product of each Q and K, and then divide by Finally, the weight value is obtained through the softmax function, and its output matrix is: The calculation process of multi-head attention is as follows: MultiHead(Q,K,V)=Concat(head1,...,head h )W o where head1=Attention(QW i Q ,KW i k ,VW i v ) in and They are mapping parameter matrices, and the specific parameters are set to d input =d b =128,d k =d v =64,d output =38, where h=4, and the multi-head self-attention feature F of miRNA is obtained through the multi-head attention mechanism process. dm ={F dm1 ,F dm2 ,....,F dm|L| }.
5. The key miRNA identification method according to claim 1, characterized in that In step S4, the final feature representation F of miRNA deep learning based on the weighted attention mechanism is d The process of obtaining is as follows: Step S401: Considering the representation vector F of the statistical structural characteristics of the miRNA sequence s And the implicit vector set of the pre-miRNA sequence based on the multi-head self-attention mechanism, the subsequences in the sequence are considered based on the weight to obtain the final feature representation F of the deep learning of miRNA d , that is, calculating the importance of subsequences in the pre-miRNA sequence to the representation of statistical structural features, giving greater weight to important subsequences. The calculation process is as follows: h m =f(W inter F s +b inter ), h i =f(W inter F dmi +b inter ), Among them, W inter and b inter are weight matrix and bias vector respectively, α i is the attention weight value, which represents the degree of association between a pre-miRNA subsequence and its statistical feature representation; Step S402: Based on the above weights, the final feature F of miRNA deep learning is obtained. d The method of obtaining it is as follows:
6. The key miRNA identification method according to claim 1, characterized in that: In step S5, the key miRNA identification process is as follows: Step S501: Obtain the 38-dimensional statistical structure features F s And the final feature F of deep learning d Splicing is performed to obtain the final miRNA signature F f ∈R d , d = 76; In step S502, the final miRNA feature vector is input into the MLP classification model to identify key miRNAs. In the iterative process of MLP: h t =f h (W h h t-1 +b h ), Among them, W h and b h are the weight matrix and bias vector respectively; In step S503, the key miRNA identification problem is a binary classification problem. The specific definition of the miRNA feature vector obtained by MLP is as follows: z=w output y con +b output , Among them, w output ∈R m×2d and b output ∈R m are the weight matrix and bias vector, respectively. When predicting key miRNAs, m is set to m = 2; Step S504, based on the output feature vector z = [o0, o1], the specific process definition of calculating the key miRNA possibility score through the softmax function is: Among them, l∈{0,1} is the label, p l is the probability value of label l; therefore, the key miRNA is predicted by the softmax function; In step S505, cross entropy loss is used as the loss function, which is defined as follows: Among them, y i,l and p i,l are the one-hot vector representations of a single sample on the label. If the i-th sample belongs to l label types, then y i,l =1, otherwise y i,l =0; N is the number of key miRNA samples in the training set, and the training objective is to minimize the following loss function l: Among them, θ is the set of all weight matrices and bias vectors, and the parameter λ is the L2 regularization hyperparameter.
Citation Information
Patent Citations
Multi-scale CNN-BiLSTM non-coding RNA interaction relationship prediction method with introducing attention
CN111341386A
RNA-protein binding site prediction method and system based on self-attention mechanism
CN114023376A