Human important microRNA recognition method based on deep field adversarial learning framework
Through the DeepHEM model, the deep domain adversarial learning framework and multimodal feature extractor are used to solve the problem of lack of tags and model limitations in human miRNA importance prediction, and more efficient and reliable miRNA importance prediction is achieved.
Patent Information
- Application Number
- CN202510264230.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art lacks effective computational models to predict the importance of human miRNAs, mainly because human miRNAs do not have real importance tags, and the structural and algorithmic limitations of traditional domain adaptive models are difficult to capture complex data features.
Develop an important human miRNA recognition method based on the deep domain adversarial learning framework, DeepHEM, through the construction of a multimodal feature extractor and a deep domain adaptive fusion model, break through the limitations of traditional models, deeply explore hidden feature information in human and mouse miRNA data, and accurately align the distribution of the two domains.
A more comprehensive miRNA feature representation and learning is achieved, the accuracy and reliability of human miRNA importance prediction is improved, and the shortcomings of traditional models in feature representation and learning ability are overcome.
Smart Images

Figure CN120220816A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and specifically to a method for identifying important human microRNAs based on a deep domain adversarial learning framework. Background Art
[0002] MicroRNA (miRNA) is a class of single-stranded non-coding RNA molecules with a length of about 22 nucleotides. It can bind to messenger RNA, inhibit its translation or promote its degradation, thereby regulating gene expression at the post-transcriptional level. A large number of studies have shown that miRNAs are widely involved in a series of key biological processes such as the development, proliferation, metabolism, differentiation and death of animal and plant cells, and are closely related to the occurrence and development of complex diseases such as cancer. In addition, the knockout of some miRNA families will lead to various abnormal phenotypes, indicating that these miRNAs play a crucial role in the growth and development of multicellular organisms. For example, when miR-199a-2 is knocked out, the activity of mammalian target of rapamycin (mTOR) in the mouse brain will decrease, and various Rett syndrome-related phenotypes will appear; after the miR-181 family is deleted, the thymus and peripheral natural killer T cells in mice will be absent; while knocking out miR-451 will interfere with the generation of mouse red blood cells, ultimately leading to anemia symptoms. These examples fully demonstrate that specific miRNAs play an irreplaceable role in the development and homeostasis maintenance of multicellular organisms. Accurately identifying key molecules in human miRNAs is of great significance for in-depth exploration of their functions in complex biological processes and diseases. At present, a large number of miRNAs have been discovered. Precise identification of important molecules from these miRNAs helps to narrow the research scope and improve research efficiency.
[0003] However, traditional methods for identifying miRNA importance mainly rely on biological knockout experiments. Such experiments are not only costly and time-consuming, but also difficult to directly apply to human research due to ethical and technical limitations. Therefore, developing efficient computational methods has become a necessary strategy for identifying important miRNAs. Using computational models to identify the importance of human miRNAs can not only break through the ethical and technical limitations of traditional experiments, but also greatly improve the efficiency of identifying important miRNAs. This enables researchers to focus on the identified important miRNAs, thereby accelerating the research process of disease mechanism analysis and target discovery.
[0004] In the early stage of identifying important miRNAs, researchers constructed some biological indicators to quantify the importance of miRNAs. For example, Cui et al. used the experimentally validated human miRNA-disease association data in the HMDD v3.0 database to calculate the Disease Spectrum Width (DSW) score. The definition of this score is the ratio between the number of diseases associated with a miRNA and the total number of diseases associated with all human miRNAs. The higher the DSW score, the more important the miRNA. Based on the theory that gene importance is positively correlated with evolutionary conservation, Wang et al. calculated the conservation score of human miRNAs, which is proportional to the number of miRNA family members. The higher the conservation score, the more important the miRNA. In recent years, many studies have focused on predicting the importance of miRNAs by inputting miRNA sequences. Song et al. developed the microRNA importance score (miES) model. First, they established a gold standard dataset for identifying the importance of mouse miRNAs, and then used a logistic regression model to predict the importance score by extracting sequence features. In addition, Yan et al. used a gradient boosting machine to predict the importance score of mouse miRNAs, and extracted 6 structural features and 18 dinucleotide frequency features in addition to sequence features. The model accuracy was better than that of miES. The RFEM calculation framework developed by Wang et al. introduced a miRNA and target gene interaction dataset to calculate miRNA functional features, and used a rotation forest model to evaluate the importance of miRNAs after constructing multiple features. There are also studies that used a voting method to integrate 60 different classification models (including 5 classification algorithms and 12 feature extraction methods) for prediction. Min et al. used the XGBoost algorithm to aggregate the prediction results of 5 base classifiers and designed the XGEM calculation framework for prediction. Yan et al. developed a deep learning framework for prediction based on miRNA sequences by integrating bidirectional long short-term memory, multi-head self-attention mechanism, and weighted attention mechanism. Generally speaking, the process of predicting the importance of mouse miRNAs in the existing technology is as follows: using mouse miRNA sequences as samples, designing various feature extraction methods to obtain feature representations, and then using different classifiers for binary classification prediction to identify important miRNAs.
[0005] At present, there is still a lack of effective computational models to predict the importance of human miRNAs, mainly because human miRNAs do not have a true importance label. Existing studies have mostly focused on developing models for identifying important mouse miRNAs, and computational models for identifying important human miRNAs remain to be explored. Although indicators such as DSW scores and conservation scores have been used to characterize the importance of human miRNAs, due to the lack of importance labels for human miRNAs, there are few studies that use computational models to predict the importance of human miRNAs. The DSW score relies on known miRNA-disease association information, and human miRNAs without association information cannot be calculated for DSW scores. Recently, Cui et al. attempted to use the DSW scores of known human miRNAs as the importance labels of human miRNAs and established a MIC algorithm based on random forest regression to predict the importance of human miRNAs with unknown DSW scores. However, directly using DSW as a true label for modeling lacks rigor and reliability, and human miRNA importance identification modeling faces huge challenges.
[0006] Human and mouse genes are highly homologous and closely related. Mouse miRNAs have importance labels, but human ones do not. Using transfer learning to transfer the annotation knowledge in the mouse field to the human field and fully combining the existing achievements in mouse miRNA importance identification is expected to achieve cross-species knowledge transfer. This research idea is theoretically feasible. However, traditional domain adaptive models are not competent; due to their simple structure, they cannot capture complex data features; the computational complexity is high when processing large-scale or high-dimensional input data; for the initial manually extracted features, only shallow or linear transformations are used, and the feature representation and learning capabilities are weak; the adaptive strategy is single, mainly using metrics such as the maximum mean difference to align the distribution of the two domains.
[0007] These problems have limited the development of modeling in the field of human miRNA importance identification. Therefore, it is necessary to develop a new method to solve this problem. Summary of the invention
[0008] This paper focuses on the deficiencies of the existing technology and aims to develop a new fusion model based on multimodal feature extractor and deep domain adaptation. This model can not only break through the structural and algorithmic limitations of traditional domain adaptation models, deeply mine the hidden feature information in human and mouse miRNA data, but also accurately align the distribution of the two domains, effectively solving the key problem of the lack of human miRNA importance labels.
[0009] The first aspect of the present invention provides a method and framework for identifying important human miRNAs based on a deep domain adversarial learning framework, DeepHEM, which provides a new idea and strategy for predicting the importance of human miRNAs using computational methods.
[0010] A method for identifying important human miRNAs based on a deep domain adversarial learning framework, comprising the steps of:
[0011] S1. Dataset construction: Construct a benchmark dataset of miRNA importance and a miRNA-target gene interaction dataset (MTI dataset) for mice and humans;
[0012] S2. The training phase is divided into two key steps. One is to construct a multi-modal feature extractor for feature extraction, and the other is to construct a loss function to achieve feature alignment between the two domains and model training;
[0013] S3. Prediction phase: Input the target domain data into the multi-modal feature extractor to obtain the miRNA feature representation of the target domain, and then obtain the importance score of the predicted target domain samples through the label predictor;
[0014] Furthermore, step S2 includes step S21 and step S22;
[0015] S21. Feature extraction is performed by constructing a multi-modal feature extractor to obtain the feature representations of the source domain and the target domain, where the data of mice and humans are regarded as the source domain data and the target domain data respectively;
[0016] S22. Based on the feature representations of the source domain and the target domain obtained in step S21, construct a loss function to achieve feature alignment between the two domains and model training;
[0017] Furthermore, in step S21, the feature representation is obtained by splicing the sequence feature representation, the intrinsic feature representation, and the MTI feature representation of the miRNA;
[0018] Even further, in step S22, the loss function includes a classification loss, a Coral loss, and an adversarial loss.
[0019] Furthermore, in step S1, the benchmark dataset of importance is obtained from the miRBase database and includes the precursor and mature miRNA sequences of mice and human miRNAs; and the important miRNAs in the mouse miRNA samples are marked as positive samples, and the same number of miRNA samples as the important miRNAs are selected from the remaining unknown miRNA samples as negative samples;
[0020] In step S1, the MTI dataset includes the interactions between miRNAs and target genes in mice, and the interactions between human miRNAs and target genes; the MTI dataset is obtained from the miRTarBase database, and the interactions with abnormal miRNA names and duplicate interactions confirmed by different experimental literatures are removed to construct the MTI dataset.
[0021] In one embodiment, 1226 mouse miRNAs and 1913 human miRNAs were collected from the miRBase database. Referring to the professional review by Bartel, 91 important miRNAs were marked out from the mouse miRNA samples, and the remaining 1135 were regarded as unknown miRNAs, from which 91 non-important miRNAs were selected. At the same time, the precursor and mature miRNA sequences of all miRNAs in mice and humans were obtained from the miRbase database.
[0022] In one embodiment, in order to further enrich the data dimension, the miRTarBase database updated in 2025 was utilized. This database contains more than 380,000 directly downloadable MTI datasets, and each piece of data has been strictly verified by experimental literature, with extremely high credibility. In the present invention, by eliminating the interactions with abnormal miRNA names and removing the duplicate interactions confirmed by different experimental literatures, finally, 398,858 interactions between 1172 miRNAs and 14,117 target genes in mice, and 1,731,969 interactions between 2989 miRNAs and 16,979 target genes in humans were screened from the miRTarBase database to construct the MTI dataset.
[0023] Furthermore, step S21 includes: based on the dataset constructed in step S1, the multimodal feature extractor takes the sequence data and MTI data of miRNAs as inputs, and extracts three parts of features: the sequence feature representation, the intrinsic feature representation, and the MTI feature representation of miRNAs;
[0024] Furthermore, the sequence feature representation is f seq , which is obtained by encoding the 3-mer frequency vector of the sequence through a transformer encoder;
[0025] For a sequence S=(s1, s2,..., s L ), s i ∈{A, U, G, C}, the 3-mer combinations of the sequence are Comb(S, 3)={s i , s i+1 , s i+2 |i = 1, 2,..., L - 2}, from which the 3-mer frequency vector of sequence S can be calculated:
[0026] f mer =[f1, f2,..., f 64 (1)
[0027] where the dimension of f mer is 64, and f iis the occurrence frequency of the i-th 3-mer sequence in the combination; subsequently, the 3-mer frequency vector is encoded by a transformer encoder to obtain f seq , whose dimension is 128.
[0028] Furthermore, the intrinsic feature is denoted as f inherent , and an 18-dimensional intrinsic feature f 18_vector can be obtained according to the input miRNA sequence. It is converted into an intrinsic feature representation with a dimension of 128 through a multi-layer perceptron module:
[0029] f inherent = ReLU(W inherent × f 18_vector + b inherent ) (2)
[0030] where W inherent and b inherent are the weight matrix and bias vector for the transformation of the intrinsic feature, respectively.
[0031] Furthermore, the MTI feature is denoted as f mti , and based on the collected MTI data corresponding to mouse and human miRNAs, an MTI feature f mti_vector with a dimension of 128 is obtained through one-hot encoding and principal component analysis methods;
[0032] Preferably, in the MTI feature representation part, first organize the target genes corresponding to each miRNA, perform one-hot encoding to obtain an interaction binary vector (if there is an interaction between target gene i and miRNA, the i-th element of the vector is 1, otherwise it is 0), and obtain 14117-dimensional MTI features for mice and 16979-dimensional MTI features for humans; considering the data sparsity and high dimension, both are reduced to 128 dimensions by principal component analysis methods. Then, an MTI feature representation with a dimension of 128 is obtained through a multi-layer perceptron module:
[0033] f mti = ReLU(W mti × f mti_vector + b mti ) (3)
[0034] where W mti and b mti are the weight matrix and bias vector for the transformation of the MTI feature, respectively;
[0035] Furthermore, the sequence feature representation, intrinsic feature representation, and MTI feature representation are concatenated to obtain the final miRNA feature representation with a dimension of 384:
[0036] f final = Concat(fseq , f inherent , f mti ) (4).
[0037] Further, in step S21, the input of the Transformer encoder module is the 3-mer frequency vector f mer , which first passes through the embedding layer and the position encoding layer to obtain the encoded representation of the frequency vector;
[0038] The calculation process of the embedding layer is expressed as follows:
[0039] f emb = W emb × f mer + b emb (5)
[0040] Among them, W emb and b emb are the weight matrix and bias vector of the embedding layer respectively;
[0041] The calculation formula of the position encoding layer is as follows:
[0042]
[0043] where pos is the position index of the word (if the sentence length is L, then pos = 0, 1,..., L - 1), i is a certain dimension of the word vector, and d model is the dimension of the word vector; in some embodiments, i is a certain dimension of the word vector (in the present invention, i ∈ [0, 63)), and d model is the dimension of the word vector (in the present invention, it is taken as 128). The position encoding layer represents the position information through sine and cosine functions of different frequencies, assigns a unique representation to each token in the sequence, enables the model to understand the structure and order of the sequence, and improves the prediction accuracy.
[0044] Further, the vectors obtained by the embedding layer and the position encoding layer are added element-wise and input into a 4-layer Transformer encoder; the encoder consists of a multi-head attention module, an addition and regularization module, and a feed-forward network layer.
[0045] Furthermore, the input of the multi-head attention module, namely the query (Query), key (Key), and value (Value) vectors Q, K, V, then passes through a linear layer transformation to obtain QW q , KW k and VW v , where W q , W k and V i W v are the weight matrices of the query, key, and value respectively;
[0046] The output of the multi-head attention module is:
[0047] MultiHead(Q, K, V) = Concat(head1,..., head h )W 0 (7)
[0048] where head i = Attention(Q i W q , K i W k , V i W v ) is the representation of the i-th head, and W 0 is the linear transformation matrix. In some embodiments of the present invention, the number of heads h is taken as 8.
[0049] The attention mechanism of the multi-head attention module adopts scaled dot-product attention. Taking Q, K, and V as inputs, after matrix multiplication, scaling operation, and Softmax, the final output can be obtained. The specific formula is as follows:
[0050]
[0051] where Q i , K i , and V i are the query, key, and value matrices of the i-th head, and d k is the matrix dimension; the purpose is to make the Softmax function more stable.
[0052] After that, residual connection is adopted, that is, the result of adding the embedding layer and the position encoding layer is added element-wise to the result of the multi-head attention module processing. Layer normalization is used for regularization, that is, normalization is performed in the sample direction to prevent gradient disappearance or gradient explosion; subsequently, through the forward propagation network layer, which mainly includes a linear layer and a ReLU activation function; finally, the input and output of the forward propagation network layer are subjected to residual connection and normalization processing to obtain the output of the final transformer encoding, that is, the sequence feature representation f seq .
[0053] After obtaining the feature representations of the source domain and the target domain, the present invention designs three loss functions to align the distributions of the two-domain features and ensure classification accuracy, namely classification loss, Coral (correlation alignment) loss, and adversarial loss.
[0054] Further, the source domain features are input into the label predictor to obtain the importance labels of the predicted source domain samples, and its formula is:
[0055]
[0056] where is the final feature representation of the source domain miRNA, respectively represent the sequence feature representation, intrinsic feature, and MTI feature representation of miRNA in the source domain; are the weight matrix and bias vector of the label predictor respectively;
[0057] Then, the classification loss is calculated by comparing the predicted importance scores of the source domain samples with the true importance labels of the mouse miRNA samples:
[0058]
[0059] where is the true importance label of the i-th miRNA sample in the source domain (1 and 0 represent positive and negative samples respectively), is the predicted probability of the i samples predicted by the model, and n S is the number of source domain samples;
[0060] In one implementation, the label predictor is a fully connected layer network.
[0061] The Coral (correlation alignment) loss: First, calculate the covariance matrix of the miRNA feature representations of the two domains, and then calculate the Coral loss:
[0062]
[0063] where d is the dimension of the source domain features, and ||·|| F represents the Frobenius norm; C S and C T are the covariance matrices of the source domain and target domain features respectively, and the specific formulas are:
[0064]
[0065] where X S and X T are the source domain and target domain feature matrices, I is an n S dimensional vector with all elements being 1, and n T is the number of target domain samples; The Coral loss reduces the feature distribution difference between the two domains by minimizing the F norm between the covariance matrices of the two domains, enabling the classifier trained on the source domain to better generalize to the target domain.
[0066] The adversarial loss is calculated from the domain cross-entropy loss between the source domain and the target domain;
[0067]
[0068] where is the cross-entropy loss of the source domain, is the cross-entropy loss of the target domain; for the i-th source domain sample feature and the j-th target domain sample feature, their respective domain cross-entropy losses can be expressed as:
[0069]
[0070] where G d represents the domain prediction label output by the domain discriminator using the sigmoid activation function; by inputting the feature representations of miRNAs in the source domain and the target domain into the domain discriminator, the domain prediction labels of the two domains are obtained respectively; based on the binary classification task, if the feature representation comes from the source domain, its domain label is 0, otherwise it is 1;
[0071] represent the final feature representations of the i-th sample in the source domain and the j-th sample in the target domain respectively, represents the true domain labels of the i-th sample feature in the source domain and the j-th sample feature in the target domain, which are 1 and 0 respectively;
[0072] During the model training process, the classification loss, Coral loss, and adversarial loss are summed up for backpropagation to update the model parameters; the gradient reversal layer performs a gradient inversion operation on the gradient during the backpropagation of the adversarial loss, so that the parameters of the domain discriminator are optimized in the direction of decreasing gradient, while the parameters of the feature extractor are optimized in the direction of increasing gradient, thus forming an adversarial relationship, prompting the feature extractor to learn feature representations that are useful for the task and insensitive to domain changes; through multiple iterative trainings, the parameters of the feature extractor and the label predictor are continuously adjusted until the preset number of training epochs is reached and the training stops.
[0073] Furthermore, in step S3,
[0074] the importance score of the target domain samples:
[0075]
[0076] where is the final feature representation of the miRNA in the target domain, represent the sequence feature representation, intrinsic feature representation, and MTI feature representation of miRNAs in the target domain respectively, are the weight matrix and bias vector of the label predictor respectively.
[0077] The second aspect of the present invention lies in providing an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the human important microRNA recognition method described above is implemented.
[0078] The third aspect of the present invention lies in providing a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the human important microRNA recognition method is implemented.
[0079] The beneficial effects of the present invention are at least as follows:
[0080] (1) The present invention innovatively designs a brand-new multi-modal feature extractor to comprehensively capture the features of miRNA from multiple perspectives. This extractor integrates the sequence features of miRNA based on k-mer and transformer encoder, the MTI features based on MTI data, and the inherent features based on the sequence, and can extract more modal miRNA features, thus representing miRNA features more comprehensively.
[0081] (2) The adversarial training method provided by the present invention enables the domain classifier to distinguish the domain source of the input features and optimize in the direction of decreasing gradient. At the same time, it enables the feature extractor to generate features that confuse the domain classifier and optimize in the direction of increasing gradient. In this way, the two-domain feature distributions are gradually aligned. On this basis, the Coral loss is further combined to minimize the difference between the two-domain feature distributions, and finally the model learns domain-invariant features. Innovatively combining the adversarial training strategy with the Coral loss and embedding it into the human important miRNA recognition framework based on deep domain adaptation is the first in this field and the key improvement of the present invention.
[0082] (3) Based on the multi-modal feature extractor and domain adversarial learning method of the present invention, the DeepHEM model developed by the present invention is an end-to-end deep domain adaptation framework, which combines the deep network with domain adaptation, can more effectively learn the miRNA feature representation, align the two-domain feature distributions, and further improve the accuracy and reliability of human miRNA importance prediction. Description of the Drawings
[0083] The drawings are used to provide a further understanding of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings:
[0084] Figure 1 : The human important miRNA recognition framework diagram based on the deep domain adversarial learning framework;
[0085] Figure 2 : Framework diagram of the multi-modal feature extractor. Specific implementation manner
[0086] To make the above objects, features, and advantages of the present invention more obvious and understandable, the specific implementation manner of the present invention will be described in detail below in conjunction with the embodiments of the specification. Many specific details are set forth in the following description in order to fully understand the present invention, but the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar promotions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0087] Evaluation method: In view of the fact that there are no real labels in the research target domain of the present invention, this study uses this model to predict the importance of unknown human miRNAs. Subsequently, the predicted importance scores are correlated with biological indicators (including DSW scores and conservation scores) to obtain corresponding correlation indicators (including Spearman rank correlation coefficient, Kendall correlation coefficient, Pearson correlation coefficient) and their corresponding p-values, and these indicators are used as the evaluation basis for the model performance.
[0088] Method for identifying important microRNAs in humans based on a deep domain adversarial learning framework
[0089] S1. First, construct a benchmark dataset of miRNA importance for mice and humans, as well as the MTI dataset;
[0090] Specifically, the miRNA importance benchmark dataset is obtained from the miRBase database. The precursor and mature miRNA sequences of all miRNAs in mice and humans are obtained from the miRbase database. A total of 1226 mouse miRNAs and 1913 human miRNAs are collected; and referring to the professional review of Bartel, 91 important miRNAs are marked in the mouse miRNA samples, and 91 non-important miRNAs are selected from the remaining 1135 unknown miRNAs.
[0091] Using the miRRTarBase database updated in 2025. Remove the interactions with abnormal miRNA names and remove the repeated interactions confirmed by different experimental literatures. Finally, 398,858 interactions between 1172 miRNAs and 14,117 target genes in mice, and 1,731,969 interactions between 2989 miRNAs and 16,979 target genes in humans are screened from the miRRTarBase database to construct the MTI dataset.
[0092] Such as Figure 1As shown, the method for identifying important human miRNAs based on the deep domain adversarial learning framework includes a training stage and a prediction stage: Step S2. Training stage: It includes constructing a multi-modal feature extractor for feature extraction (Step S21), and constructing a loss function to achieve feature alignment between the two domains and model training (Step S22).
[0093] First, based on the miRNA sequence data and MTI data of mice and humans collected in Step S1, the data of mice is regarded as source domain data, and the data of humans is regarded as target domain data; a multi-modal feature extractor is constructed to extract features of miRNAs, obtaining feature representations of the source domain and the target domain.
[0094] As Figure 2 (a) shows, the multi-modal feature extractor takes the sequence data and MTI data of miRNAs as input, and extracts three parts of features: the sequence feature representation, the intrinsic feature representation, and the MTI feature representation of miRNAs.
[0095] The multi-modal feature extractor of the present invention has the advantage of more comprehensive feature extraction compared with traditional methods; traditional methods usually only focus on single-modal features (such as sequence features) of miRNAs, while the present invention can simultaneously extract and fuse multiple modal features of miRNAs by designing a multi-modal feature extractor, so as to more comprehensively capture its rich biological information. This way of multi-modal feature fusion helps to more accurately understand the function and importance of miRNAs, and further improves the accuracy of prediction.
[0096] The multi-modal feature extractor of the present invention is further described in detail below:
[0097] In the sequence feature representation part, for a sequence S=(s1, s2,..., s L ), s i ∈{A, U, G, C}, its 3-mer combination of the sequence is Comb(S, 3)={s i , s i+1 , s i+2 |i = 1, 2,..., L - 2}, from which the 3-mer frequency vector of the sequence S can be calculated:
[0098] f mer =[f1, f2,..., f 64 (1)
[0099] where the dimension of f mer is 64, and f i is the occurrence frequency of the i-th 3-mer sequence in the combination. Subsequently, after the 3-mer frequency vector is encoded by the transformer encoder, the sequence feature representation of the miRNA is obtained as f seq, with a dimension of 128.
[0100] Figure 2 (b) shows the transformer encoder module of the multimodal feature extractor, whose input is the 3-mer frequency vector f mer , which first passes through the embedding layer and the position encoding layer to obtain the encoded representation of the frequency vector.
[0101] Specifically, the calculation process of the embedding layer is expressed as:
[0102] f emb = W emb × f mer + b emb (5)
[0103] where W emb and b emb are the weight matrix and bias vector of the embedding layer, respectively.
[0104] The calculation formula of the position encoding layer is as follows:
[0105]
[0106] where pos is the position index of the word (if the sentence length is L, then pos = 0, 1,..., L - 1), i is a certain dimension of the word vector (in the present invention, i ∈ [0, 63)), and d model is the dimension of the word vector (in the present invention, it is taken as 128).
[0107] Then, the vectors obtained from the embedding layer and the position encoding layer are added element-wise and input into the 4-layer transformer encoder. Each encoder consists of a multi-head attention module, an addition and regularization module, and a feed-forward network layer.
[0108] The multi-head attention module is as shown in Figure 2 (c): First, the inputs of the multi-head attention module are obtained, namely the query (Query), key (Key), and value (Value) vectors Q, K, V, and then they are transformed through three different linear layers to obtain QW q , KW k and VW v , where W q , W k and V i W v are the query, key, and value weight matrices, respectively. The output of the multi-head attention is:
[0109] MultiHead(Q, K, V) = Concat(head1,..., head h )W 0 (7)
[0110] Among them, head i = Attention(Q i W q , K i W k , V i W v ) is the representation of the i-th head, and W 0 is the linear transformation matrix. In the embodiments of the present invention, the number of heads h is taken as 8.
[0111] As Figure 2 (d) shows: The attention mechanism adopted is scaled dot-product attention. Scaled dot-product attention takes Q, K, and V as inputs, and after matrix multiplication, scaling operation, and Softmax, the final output can be obtained. The specific formula is as follows:
[0112]
[0113] Among them, Q i , K i and V i are the query, key, and value matrices of the i-th head, and d k is the matrix dimension.
[0114] After that, residual connection is adopted, that is, the sum of the embedding layer and the position encoding layer is added element-wise to the result of the multi-head attention module. Layer normalization is used for regularization, that is, normalization is performed in the sample direction to prevent gradient disappearance or gradient explosion. Subsequently, through the feed-forward network layer, which mainly includes a linear layer and a ReLU activation function. Finally, the input and output of the feed-forward network layer are subjected to residual connection and normalization processing to obtain the output of the final transformer encoding, that is, the sequence feature representation f seq .
[0115] As Figure 2 (a)-(ii) shows: In the inherent feature representation part, according to the input miRNA sequence, an 18-dimensional inherent feature f 18_vector can be obtained, and it is converted into an inherent feature representation with a dimension of 128 through a multi-layer perceptron module:
[0116] f inherent = ReLU(W inherent × f 18_vector + b inherent ) (2)
[0117] Among them, W inherent and b inherent are the weight matrix and bias vector of the inherent feature transformation respectively.
[0118] As shown in Figure 2 (a)-(iii): In the MTI feature representation part, according to the collected MTI data corresponding to mouse and human miRNAs, the MTI feature f with a dimension of 128 is obtained through one-hot encoding and principal component analysis methods mti_vector .
[0119] Specifically, first organize the target genes corresponding to each miRNA, perform one-hot encoding to obtain interaction binary vectors (if there is an interaction between target gene i and miRNA, the i-th element of the vector is 1, otherwise it is 0), and obtain MTI features of 14,117 dimensions for mice and 16,979 dimensions for humans. Considering the sparse and high-dimensional data, the principal component analysis method is used to reduce the dimension to 128. Then, through the multi-layer perceptron module, the MTI feature representation with a dimension of 128 is obtained: f mti = ReLU(W mti × f mti_vector + b mti )
[0120] where W mti and b mti are the weight matrix and bias vector for MTI feature transformation respectively. Concatenate the above three feature representations to obtain the final miRNA feature representation with a dimension of 384:
[0121] f final = Concat(f seq , f inherent , f mti ) (4)
[0122] Step S22: Construct a loss function based on the feature representations of the source domain and the target domain obtained in step S21 to achieve feature alignment between the two domains and model training
[0123] In cross-domain or cross-dataset applications, the difference in feature distributions between the source domain and the target domain may lead to a decline in model performance. The present invention improves feature alignment by combining adversarial training and Coral loss. The DeepHEM model designed by the present invention adopts an adversarial training method, and through domain adversarial learning, the feature distributions of the two domains are gradually aligned, and domain-invariant features are learned. This mechanism helps to reduce the differences between domains and improve the generalization ability of the model on different datasets. At the same time, the present invention also introduces Coral loss for joint optimization to further narrow the difference between the feature distributions of the two domains. This combination not only enhances the model's ability to learn domain-invariant features, but also significantly improves the generalization performance and prediction accuracy of the model on different datasets
[0124] As shown in Figure 1As shown in the figure, the present invention designs three loss functions to align the distributions of the two-domain features and ensure classification accuracy, namely classification loss, Coral (correlation alignment) loss, and adversarial loss.
[0125] The source domain features are input into the label predictor to obtain the importance labels of the predicted source domain samples, and its formula is:
[0126]
[0127] Where, is the final feature representation of the source domain miRNA, respectively represent the sequence feature representation, intrinsic feature, and MTI feature representation of miRNA in the source domain; are the weight matrix and bias vector of the label predictor respectively; then, the classification loss is calculated by comparing the predicted importance scores of the source domain samples with the true importance labels of the mouse miRNA samples.
[0128] Classification loss:
[0129] Where, is the true importance label of the i-th miRNA sample in the source domain (1 and 0 represent positive and negative samples respectively), is the predicted probability of the i-th sample predicted by the model, and n S is the number of source domain samples.
[0130] Coral loss The calculation formula of is:
[0131]
[0132] First, calculate the covariance matrices of the miRNA feature representations in the two domains, and then calculate the Coral loss:
[0133]
[0134] C S and C T are the covariance matrices of the source domain and target domain features respectively; where d is the dimension of the source domain features, and ||g|| F represents the Frobenius norm. X S and X T are the source domain and target domain feature matrices, I is an n S dimensional vector with all elements equal to 1, and n T is the number of target domain samples; the Coral loss reduces the difference in feature distributions between the two domains by minimizing the F norm between the covariance matrices of the two domains, enabling the classifier trained on the source domain to better generalize to the target domain.
[0135] Adversarial loss : Calculated from the domain cross-entropy loss between the source domain and the target domain;
[0136]
[0137] Among them, is the cross-entropy loss of the source domain, is the cross-entropy loss of the target domain; for the i-th source domain sample feature and the j-th target domain sample feature, their respective domain cross-entropy losses can be expressed as:
[0138]
[0139]
[0140] Among them, G d represents the domain prediction label output by the domain discriminator using the sigmoid activation function; by inputting the feature representations of miRNAs in the source domain and the target domain into the domain discriminator, the domain prediction labels of the two domains are obtained respectively; based on the binary classification task, if the feature representation comes from the source domain, its domain label is 0, otherwise it is 1;
[0141] respectively represent the final feature representations of the i-th sample in the source domain and the j-th sample in the target domain, represents the true domain labels of the i-th sample feature in the source domain and the j-th sample feature in the target domain, which are 1 and 0 respectively.
[0142] Such as Figure 1 (iii) shows that during the model training process, the classification loss, Coral loss, and adversarial loss are summed up for backpropagation to update the model parameters; the gradient reversal layer performs an inverse operation on the gradient during the backpropagation of the adversarial loss, so that the parameters of the domain discriminator are optimized in the direction of decreasing gradient, while the parameters of the feature extractor are optimized in the direction of increasing gradient, thus forming an adversarial relationship, prompting the feature extractor to learn feature representations that are useful for the task and insensitive to domain changes; through multiple iterative trainings, the parameters of the feature extractor and the label predictor are continuously adjusted until the preset number of training rounds is reached and the training stops.
[0143] Such as Figure 1 (2) shows that in the prediction stage of step S3: based on the model trained in step S1, the target domain data is input into the multi-modal feature extractor to obtain the miRNA feature representation of the target domain, and then the importance score of the predicted target domain sample is obtained through the label predictor:
[0144]
[0145] Among them, is the final feature representation of the miRNA in the target domain, which respectively represent the sequence feature representation, the intrinsic feature, and the MTI feature representation of the miRNA in the target domain. are respectively the weight matrix and the bias vector of the label predictor.
[0146] The end-to-end deep domain adaptation framework provided by the present invention: Compared with some traditional, phased, or handcrafted feature extraction-based methods, the DeepHEM framework of the present invention tightly combines a deep network with domain adaptation technology. This design simplifies the model training process, reduces the amount of human intervention and feature engineering work. At the same time, the deep network can automatically learn high-level, multi-modal feature representations, improving the model's learning ability for miRNA features. The domain adaptation technology further enhances the model's adaptability on different datasets, making DeepHEM show higher accuracy and reliability in predicting the importance of human miRNAs.
[0147] To further verify the model performance, a comparative experiment was carried out, comparing the performance of the model developed by the present invention with Transfer Component Analysis (TCA), Deep domain confusion (DDC), and Deep Adaptation Networks (DAN). The experimental results of the comparative experiment are shown in Table 1.
[0148] Table 1: Results of the comparative experiment
[0149]
[0150] It can be clearly seen from the data in Table 1 that in the correlation analysis experiments based on the DSW score and the conservation score, the prediction effect of this model is basically better than that of other comparative models. This result fully verifies the effectiveness of the deep domain adversarial learning framework proposed by the present invention.
[0151] In addition, ablation experiments were also carried out by ablating multiple modules (Coral loss, loss function, feature fusion method, and feature representation module) in the DeepHEM model developed by the present invention. Correlation analysis based on the DSW score and correlation analysis based on the conservation score were respectively performed, and Spearman rank correlation coefficient, Kendall correlation coefficient, Pearson correlation coefficient, and the corresponding p-value were used as evaluation criteria. The results of the ablation experiments are shown in Table 2 (Table 2-1 and Table 2-2).
[0152] Table 2-1: Results of the ablation experiment
[0153]
[0154] Table 2-2: Results of Ablation Experiments
[0155]
[0156] As can be seen from the results in Table 2 (Table 2-1 and Table 2-2), the DeepHEM model developed in the present invention outperforms other variant models in the three related analysis experiments based on DSW scores and conservation scores.
[0157] In addition, to further verify the prediction performance of the DeepHEM model, the present invention also conducted a case study. Based on miRNA data of mice and humans, DeepHEM was used to score the importance of all human miRNAs, and the predicted scores were ranked in descending order. The top 10 miRNAs were selected for literature verification to see if the miRNA has a certain important function, indirectly proving their importance. The results of the case study are shown in Table 3.
[0158] Table 3: Results of Case Study
[0159] Rank Name Score PMID Rank in DSW score list Rank in conservativeness score list 1 hsa-mir-548d-1 0.9948 31937753 1332 175 2 hsa-mir-548d-2 0.9926 31937753 1331 174 3 hsa-mir-548c 0.9895 Unconfirmed 306 169 4 hsa-mir-515-1 0.9883 33243974 1292 69 4 hsa-mir-515-2 0.9883 33243974 1293 70 6 hsa-mir-3613 0.9864 29689704 1744 700 7 hsa-mir-3688-1 0.9863 Unconfirmed 1456 882 8 hsa-mir-3688-2 0.9860 Unconfirmed 1455 881 9 hsa-mir-548x 0.9845 32196088 809 181 10 hsa-mir-548a-3 0.9841 Unconfirmed 1281 170
[0160] As can be obtained from Table 3, among the top 10 human important miRNAs predicted by the DeepHEM model of the present invention, 6 are confirmed by relevant literature to have important biological functions. In addition, the confirmed miRNAs are ranked outside the 809th place and outside the 69th place in the DSW score list and the conservation score list respectively, proving that the present model can identify miRNAs with important biological functions that could not be identified by previous metrics, further verifying the prediction performance of the DeepHEM model.
[0161] Device for a method of identifying important human miRNAs based on a deep domain adversarial learning framework:
[0162] On the other hand, the present invention provides a device for a method of identifying important human miRNAs based on a deep domain adversarial learning framework, including a memory and a processor connected in communication. The memory is used to store a computer program, and the processor is used to read the computer program and execute the method of identifying important human miRNAs based on the deep domain adversarial learning framework described above.
[0163] In some embodiments, the drug and pathway association prediction device based on supervised learning includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in any of the above methods of identifying important human miRNAs based on the deep domain adversarial learning framework.
[0164] Those skilled in the art can understand that the schematic diagram is only an example of the terminal device, which does not constitute a limitation on the terminal device. It may include more or fewer components than those shown in the figure, or combine some components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.
[0165] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the present invention.
Claims
1. A method for identifying important human miRNAs based on a deep domain adversarial learning framework, characterized in that: Includes steps: S1. Dataset construction: Construct mouse and human miRNA importance benchmark datasets and MTI datasets; S2. Training phase: including step S21-constructing a multimodal feature extractor for feature extraction, step S22-constructing a loss function to achieve two-domain feature alignment and model training; S21. Extract features by constructing a multimodal feature extractor to obtain feature representations of the source domain and the target domain, wherein the mouse and human data are regarded as the source domain data and the target domain data, respectively; S22. Based on the feature representations of the source domain and the target domain obtained in step S21, a loss function is constructed to achieve feature alignment and model training of the two domains; S3. Prediction stage: Input the target domain data into the multimodal feature extractor to obtain the target domain miRNA feature representation, and then pass it through the label predictor to obtain the predicted importance score of the target domain samples; In step S21, the feature representation is obtained by splicing the sequence feature representation, the intrinsic feature representation and the MTI feature representation of the miRNA; In step S22, the distribution of features of the two domains is aligned and the classification accuracy is ensured through a loss function; the loss function includes classification loss, Coral loss, and adversarial loss.
2. The method for identifying important human miRNAs based on a deep domain adversarial learning framework according to claim 1, characterized in that: In step S1, the importance benchmark data set is obtained from the miRBase database, including the precursor and mature miRNA sequences of mouse and human miRNAs; and the important miRNAs are marked as positive samples for the mouse miRNA samples, and the same number of miRNA samples as the important miRNAs are selected from the remaining unknown miRNA samples as negative samples; In step S1, the MTI dataset is a miRNA-target gene interaction dataset obtained from the miRTarBase database: including the interaction between mouse miRNA and target genes, and the interaction between human miRNA and target genes; and the interactions with abnormal miRNA names and repeated interactions confirmed by different experimental literature are eliminated to construct the MTI dataset.
3. The method for identifying important human miRNAs based on a deep domain adversarial learning framework according to claim 1, characterized in that: Step S21 comprises: based on the data set constructed in step S1, the multimodal feature extractor takes the sequence data and MTI data of miRNA as input, and extracts three features of the sequence feature representation, the inherent feature representation and the MTI feature representation of miRNA; The sequence feature is denoted as f seq , obtained by encoding the 3-mer frequency vector of the sequence through the transformer encoder; For a sequence S=(s1,s2,...,s L ),s i ∈{A,U,G,C}, the 3-mer combination of its sequence is Comb(S,3)={s i ,s i+1 ,s i+2 |i=1,2,…,L-2}, from which the 3-mer frequency vector of sequence S can be calculated: f mer =[f1,f2,...,f 64 ] (1) where f mer The dimension is 64, f i is the frequency of occurrence of the i-th 3-mer sequence in the combination; then, the 3-mer frequency vector is encoded by the transformer encoder to obtain f seq , whose dimension is 128; The inherent characteristic is denoted as f inherent , according to the input miRNA sequence, we obtain the 18-dimensional intrinsic feature f 18_vector , which is converted into an intrinsic feature representation with a dimension of 128 through a multi-layer perceptron module: f inherent =ReLU(W inherent ×f 18_vector +b inherent ) (2) Among them, W inherent and b inherent are the weight matrix and bias vector of the intrinsic feature transformation respectively; The MTI feature is represented by f mti According to the collected MTI data corresponding to mouse and human miRNA, the MTI feature f with a dimension of 128 was obtained by one-hot encoding and principal component analysis. mti_vector ; f mti =ReLU(W mti ×f mti_vector +b mti ) (3) Among them, W mti and b mti They are the weight matrix and bias vector of MTI feature transformation respectively; The sequence feature representation, the intrinsic feature representation and the MTI feature representation are concatenated to obtain the final miRNA feature representation with a dimension of 384: f final =Concat(f seq ,f inherent ,f mti ) (4)。 4. The method for identifying important human miRNAs based on a deep domain adversarial learning framework according to claim 3, characterized in that: Extract the MTI feature representation step. First, sort out the target genes corresponding to each miRNA and perform one-hot encoding to obtain the interaction binary vector. If there is an interaction between target gene i and miRNA, the i-th element of the vector is 1, otherwise it is 0. The 14117-dimensional MTI features of mice and 16979-dimensional MTI features of humans were obtained; the principal component analysis method was used to reduce their dimensions to 128 dimensions; and the MTI feature representation with a dimension of 128 was obtained through the multi-layer perceptron module.
5. The method for identifying important human miRNAs based on a deep domain adversarial learning framework according to claim 3, characterized in that: In step S21, the transformer encoder module takes as input the 3-mer frequency vector f mer , first pass through the embedding layer and the position encoding layer to obtain the encoded representation of the frequency vector; The calculation process of the embedding layer is expressed as follows: f emb =W emb ×f mer +b emb (5) Among them, W emb and b emb are the weight matrix and bias vector of the embedding layer respectively; The calculation formula of the position encoding layer is as follows: Where pos is the position index of the word, i is a dimension of the word vector, and d model is the dimension of the word vector; The vectors obtained by the embedding layer and the position encoding layer are added element by element and input into a transformer encoder; the encoder is composed of a multi-head attention module, an addition and regularization module, and a forward propagation network layer; The input of the multi-head attention module includes query, key and value vectors Q, K, V, which are then transformed by a linear layer to obtain QW q ,KW k and VW v , where W q ,W k and V i W v are the weight matrices for query, key, and value respectively; The output of the multi-head attention module is: MultiHead(Q,K,V)=Concat(head1,...,head h )W 0 (7) Among them, head i =Attention(Q i W q ,K i W k ,V i W v ) is the representation of the i-th head, W 0 is the linear transformation matrix.
6. The method for identifying important human miRNAs based on a deep domain adversarial learning framework according to claim 5, characterized in that: The attention mechanism of the multi-head attention module adopts scaled dot product attention, taking Q, K and V as input, and obtaining the final output after matrix multiplication, scaling operation and Softmax. The specific formula is as follows: Among them, Q i ,K i and V i is the query, key, and value matrix of the i-th head, d k is the matrix dimension; After that, residual connection is used, that is, the sum of the embedding layer and the position encoding layer is added to the result of the multi-head attention module processing element by element, and the regularization part adopts layer normalization; then, it passes through the forward propagation network layer, which mainly includes two parts: linear layer and ReLU activation function; finally, the input and output of the forward propagation network layer are residually connected and normalized to obtain the final transformer encoded output, that is, the sequence feature representation f of miRNA seq .
7. The method for identifying important human miRNAs based on a deep domain adversarial learning framework according to claim 1, characterized in that: In step S22, the source domain features are input into the label predictor to obtain the predicted importance label of the source domain sample, and the formula is: in, is the final feature representation of the source domain miRNA, and They represent the sequence feature representation, intrinsic feature representation, and MTI feature representation of miRNA in the source domain, respectively; and are the weight matrix and bias vector of the label predictor respectively; Then, the classification loss is calculated by combining the importance scores of the predicted source domain samples with the true mouse miRNA sample importance labels: in, is the true importance label of the i-th miRNA sample in the source domain, is the predicted probability of the i samples predicted by the model, n S is the number of source domain samples; The Coral loss: first calculate the covariance matrix of the two-domain miRNA feature representation, and then calculate the Coral loss: Where d is the dimension of the source domain features, ||·|| F represents the Frobenius norm; C S and C T are the covariance matrices of the source domain and target domain features respectively, and the specific formula is: Where X S and X T is the feature matrix of the source domain and the target domain, I is an n- S dimensional vector, n T is the number of samples in the target domain; Coral loss is to reduce the difference in feature distribution between the two domains by minimizing the F norm between the covariance matrices of the two domains, so that the classifier trained in the source domain can be better generalized to the target domain; The adversarial loss It is calculated by the domain cross entropy loss between the source domain and the target domain; in, is the cross entropy loss of the source domain, is the cross entropy loss of the target domain; for the i-th source domain sample feature and the j-th target domain sample feature, their respective domain cross entropy losses can be expressed as: Among them, G d represents the domain prediction label output by the domain discriminator using the sigmoid activation function; by inputting the feature representation of miRNA in the source domain and the target domain into the domain discriminator, the domain prediction labels of the two domains are obtained respectively; based on the binary classification task, if the feature representation comes from the source domain, its domain label is 0, otherwise it is 1; and They represent the final feature representations of the i-th sample in the source domain and the j-th sample in the target domain, respectively. and Represents the true domain labels of the i-th sample feature in the source domain and the j-th sample feature in the target domain, which are 1 and 0 respectively; During the model training process, the loss function is back-propagated to update the model parameters.
8. The method for identifying important human miRNAs based on a deep domain adversarial learning framework according to claim 1, characterized in that: In step S3, Importance score of target domain samples: in, is the final feature representation of the target domain miRNA, and They represent the sequence feature representation, inherent feature representation and MTI feature representation of miRNA in the target domain respectively. and are the weight matrix and bias vector of the label predictor, respectively.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the important miRNA identification method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the important miRNA identification method as described in any one of claims 1 to 8 is implemented.