MiRNA mutation harmfulness prediction system based on improved Transformer

By improving the Transformer model and GMM clustering model, combining multiple miRNA characteristics, the problem of insufficient ability to extract and generalize miRNA mutation prediction feature in the prior art is solved, and more accurate prediction of harmful miRNA mutations is achieved.

CN120089211AActive Publication Date: 2025-06-03SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510250465.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-03
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The prior art lacks the ability to extract and generalize features and models in miRNA mutation prediction, resulting in inaccurate prediction of harmful miRNA mutations.

Method used

The miRNA mutation harmful prediction system based on improved Transformer was adopted, and harmful scores of miRNA mutation data sets were integrated, miRNA energy characteristics, miRNA interaction characteristics and existing prediction methods were extracted, and the Transformer model was improved using local attention mechanisms and dynamic sparse attention mechanisms, and harmful prediction was carried out in combination with GMM clustering model.

Benefits of technology

It improves the accuracy and generalization ability of harmful prediction of miRNA mutations, and can more accurately classify miRNA mutations into harmful, uncertain and harmless categories, supporting accurate diagnosis and personalized treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089211A_ABST
    Figure CN120089211A_ABST
Patent Text Reader

Abstract

The invention discloses a miRNA mutation harmfulness prediction system based on an improved Transform, and the system comprises a miRNA data set construction module which integrates miRNA mutation information and classification tags from different sources to construct a miRNA mutation data set; the miRNA feature extraction module is used for extracting three types of features, namely an energy feature and an interaction feature of the miRNA mutation data set and a miRNA mutation harmfulness score of an existing prediction method; the training module is used for capturing mutation information and classification labels from the miRNA mutation data set on the basis of the improved Transform model, learning the mutual relation among the three types of features and finally obtaining the trained improved Transform model; and the harmfulness prediction module is used for realizing the harmfulness prediction of the miRNA by introducing a GMM clustering model into a feature matrix obtained by the improved Transform model. The miRNA mutation harmfulness prediction can be effectively realized, a precise method is provided for clinical diagnosis and precise medical treatment, and the miRNA mutation harmfulness prediction method has the advantages of high accuracy and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and biomedicine, and in particular to a miRNA mutation harmfulness prediction system based on an improved Transformer. Background Art

[0002] At present, how to screen out clinically valuable variant information from a vast amount of gene variant data has become a challenging task. In addition, the diversity and complexity of these vast amounts of mutant gene data itself bring many difficulties to the research on gene mutations, especially the harmfulness of miRNA mutations, becoming the pain points and difficulties that need to be solved urgently in this field. Therefore, the prediction of the harmfulness of miRNA mutations has become a research topic that needs to be solved urgently.

[0003] MicroRNA, as an endogenous non-coding RNA, is crucial for gene expression regulation. In disease research, miRNA mutations can change regulatory functions and affect cell physiology. miRNAs have sequence conservation, tissue specificity, and temporal expression, and can undergo various mutations, which may interfere with their functions. Among them, mutations in the seed region and precursor structure are often harmful. However, miRNA mutation data is diverse and complex, scattered in different databases and literatures, with huge and fragmented information, making it difficult to obtain and requiring integration and cleaning. In addition, the dataset often has the problem of type imbalance, which affects model training and performance evaluation. Not all miRNA mutations are pathogenic, and many mutations have little functional impact or are harmless, resulting in inaccurate mutation classification in clinical detection and affecting precision medicine. Accurately mining harmful miRNA mutations can provide clinical markers for precision diagnosis, guide the understanding of disease mechanisms, drug design, and personalized medication, and has important medical and social significance. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art, and propose a miRNA mutation harmfulness prediction system based on an improved Transformer, to solve the problems of insufficient feature extraction and model generalization ability in miRNA mutation prediction by traditional machine learning. By deeply studying the advantages and disadvantages of various algorithm models, a new model building idea that can effectively solve the shortcomings of traditional machine learning algorithms is finally explored. At the same time, to assist researchers in judging the harmfulness of miRNA mutations, realizing artificial intelligence-assisted screening and diagnosis, deepening people's understanding of miRNA functions and disease occurrence mechanisms, providing a new perspective for research in related fields, and implementing a miRNA mutation harmfulness prediction system based on an improved Transformer model.

[0005] To achieve the above purpose, the technical solution provided by the present invention is: A miRNA mutation harmfulness prediction system based on an improved Transformer, comprising:

[0006] The miRNA dataset construction module is used to integrate miRNA mutation information from different sources and its classification labels of "harmful", "uncertain", and "harmless". After data cleaning, standardization, and deduplication, a miRNA mutation dataset is obtained;

[0007] The miRNA feature extraction module extracts three types of features through various miRNA prediction tools, namely, the miRNA energy feature, miRNA interaction feature of the miRNA mutation dataset, and the miRNA mutation harmfulness score of existing prediction methods;

[0008] The training module, based on the improved Transformer model, captures its mutation information and classification labels from the miRNA mutation dataset, learns the mutual relationship among the above three types of features, and finally obtains a trained improved Transformer model; this improved Transformer model improves the feature processor and attention mechanism of the Transformer model; the improvement of the feature processor is that the above three types of features are introduced into the input of the feature processor; the improvement of the attention mechanism is that the original single attention mechanism is changed to a local attention mechanism and a dynamic sparse attention mechanism. The local attention mechanism captures features at adjacent positions by restricting the attention calculation within a local window, and the dynamic sparse attention mechanism calculates by dynamically selecting key positions;

[0009] The harmfulness prediction module realizes the harmfulness prediction of miRNA by introducing a GMM clustering model into the feature matrix obtained by the improved Transformer model, and classifies the mutations in the miRNA mutation dataset into three key categories, namely, harmful, uncertain, and harmless.

[0010] Furthermore, the miRNA dataset construction module integrates four databases related to miRNA mutations, namely miRbase, ClinVar, dbSNP, and COSMIC, and then adopts the union processing method to summarize miRNA mutation data, miRNA sequences before and after mutation, and classification labels with clinical significance to obtain a miRNA mutation dataset; then, the miRNA mutation dataset after union is cleaned: first, miRNA mutation data beyond the normal range and missing information are removed according to the statistical method Z-score, and then the classification labels are standardized and unified into "harmful", "uncertain", and "harmless", and then deduplication is performed to ensure that each miRNA mutation data is unique in the miRNA mutation dataset; finally, a miRNA mutation dataset with miRNA mutation data, miRNA sequences before and after mutation, and classification labels is obtained.

[0011] Further, the miRNA feature extraction module performs the following operations:

[0012] Obtain miRNA energy features: Through the RNAFold prediction tool, input the miRNA mutation data in the miRNA mutation dataset to obtain the energy values and energy change values before and after miRNA mutation.

[0013] Obtain miRNA interaction features: Obtain the FASTA files of all sequences of three types of target genes, namely lncRNA, 3'UTR, and circRNA, from the LncBase database, TargetScan database, and CircBase database respectively. Then, input the above files and the miRNA sequences before and after mutation into the miRanda prediction tool to obtain a list of three types of target genes, lncRNA, 3'UTR, and circRNA, that interact with the miRNA mutation dataset. Statistically obtain the quantity values and change values of the interactions between miRNA before and after mutation and three types of target genes, lncRNA, 3'UTR, and circRNA, respectively, and generate one-hot encoding.

[0014] Obtain the miRNA mutation harmfulness scores of existing prediction methods: Input the miRNA mutation dataset into existing prediction tools related to miRNA mutation harmfulness prediction, and extract the harmfulness scores generated by each method.

[0015] Further, the training module performs the following operations:

[0016] a. Input the miRNA mutation dataset into the improved Transformer model, add positional encoding to retain the positional information of the miRNA sequences before and after mutation in the miRNA mutation dataset. The miRNA mutation dataset is converted from discrete data to continuous vector H through the embedding layer of the improved Transformer model embed , and the mapping of each data is realized through the embedding matrix of the embedding layer. Then, input the continuous vector H embed output by the embedding layer and the features obtained by the miRNA feature extraction module into the feature processor to generate token-level embedding H token ; Then, after layer normalization and linear transformation of the model, it is output to the encoder of the improved Transformer model. Among them, the improvement of the feature processor is: introducing 3 types of features related to miRNA: miRNA energy feature F energy , miRNA interaction feature F interaction and the miRNA mutation harmfulness score F score of existing prediction methods, increasing the diversity of input features, which is beneficial to generating richer token-level embedding H token , improving the accuracy of the improved Transformer model, as shown below:

[0017] H token = Concat(H embed , F energy , F interaction , F score )

[0018] b. The encoder receives the token-level embeddings H token output by the embedding layer and processes them with an improved attention mechanism; the encoder transforms and combines each token-level embedding multiple times, extracts deeper features, and generates new feature representations through a feed-forward neural network; finally, the encoder completes the processing of all token-level embeddings and outputs a feature matrix, where each row of the matrix corresponds to the feature representation of each miRNA mutation in the miRNA mutation dataset; among them, the improvement to the attention mechanism is: changing the original single attention mechanism to a multi-head attention mechanism, which is beneficial to taking advantage of the advantages of different attention mechanisms, improving the efficiency of the model in processing long sequence data, reducing the computational complexity, and at the same time retaining important global information.

[0019] Further, the multi-head attention mechanism consists of a local attention mechanism and a dynamic sparse attention mechanism;

[0020] The local attention mechanism defines a window w of a fixed size, calculates the attention weights within this window, and after inputting the token-level embeddings, only focuses on the current token i and its surrounding tokens to extract local context information; the query vector Q i of the current token i will be calculated with the key vectors K [i-w,i+w] within the window to generate the attention weights for the surrounding tokens; then, the attention weights are used to weight the value vectors V [i-w,i+w] within the window, thereby obtaining the representation of the local context; the process of the local attention mechanism is represented as follows:

[0021]

[0022] In the formula, LocalAttention represents the local attention mechanism, and d k is the dimension of the key vector;

[0023] The dynamic sparse attention mechanism uses the query matrix Q of all tokens of the token-level embeddings to calculate the attention weights related to the selected set of key tokens S; then, based on the attention weights, the corresponding key matrix K S and value matrix V S are selected to reduce the computational amount while retaining the key information; the process of the dynamic sparse attention mechanism is represented as follows:

[0024]

[0025] Wherein, SparseAttention represents a dynamic sparse attention mechanism.

[0026] Furthermore, the harmfulness prediction module performs the following operations:

[0027] a. Input the feature matrix output by the encoder of the improved Transformer model into the GMM clustering model for clustering operations, where the parameters of the GMM clustering model are initialized, the covariance in the parameters is set as the identity matrix, and the initial values of the mixing coefficients are set as equal values;

[0028] b. Use the EM algorithm for iterative optimization: In the E-step, calculate the posterior probability that the feature vector passes through the i'-th Gaussian mixture component in the GMM clustering model to obtain whether it belongs to harmful or harmless; in the M-step, calculate the GMM coefficients of X classification clusters; after one iteration of the EM algorithm, update the model parameters and enter the next iteration until the change in the model parameters is less than 10 -4 ;

[0029] c. Use the optimized GMM clustering model to calculate the probability P that each miRNA mutation data in the miRNA mutation dataset belongs to the harmful class. If P > 0.5, then classify this miRNA mutation data as harmful, otherwise classify it as harmless; traverse each miRNA mutation data in the harmful and harmless classes obtained by the GMM clustering model. If the posterior probability is greater than the preset threshold, it is considered that this miRNA mutation data is acceptable at the confidence level of this category and is retained in the original category; otherwise, it is considered that the confidence is insufficient and it is assigned to the uncertain category;

[0030] d. On the basis of completing GMM clustering and obtaining the probability P of each miRNA mutation data, calculate the miRNA mutation score S, S ∈ [0, 1], where 0 is considered harmless and 1 is considered harmful. If the clustering category of a certain miRNA mutation data is harmful, then S = P; if the clustering category is harmless, then S = 1 - P.

[0031] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0032] 1. The present invention innovatively combines deep learning and machine learning. The Transformer model in the deep learning framework can mine the connections between miRNA mutation information. The introduction of the self-attention mechanism enables the model to deeply explore the connections between data layer by layer, and at the same time improves the generalization ability of the model.

[0033] 2. The introduction of the local attention mechanism and the dynamic sparse attention mechanism improves the efficiency and performance of the Transformer model in processing long sequence data.

[0034] 3. The machine learning framework utilizes an unsupervised clustering model (GMM clustering model) to classify miRNA mutations by receiving the matrix information obtained from the encoder of the Transformer model. The unsupervised classification solves the problem of a small number of sample labels.

[0035] 4. The combination of deep learning and machine learning is beneficial for exploring the internal relationships between mutation information, alleviates the limitations brought by a small number of sample labels, and provides a new research scheme for predicting the harmfulness of miRNA mutations. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic diagram of the relationships between the various modules of the system of the present invention.

[0037] Figure 2 It is a flowchart of the training and prediction of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The present invention will be further described below in conjunction with specific embodiments.

[0039] This embodiment discloses a system for predicting the harmfulness of miRNA mutations based on an improved Transformer, which uses a large machine learning model as the basis for predicting the harmfulness of miRNA mutations. The relationships between the various modules of the system are as Figure 1 shown, and the flowcharts of the training and prediction of the system are as Figure 2 shown. It includes:

[0040] An miRNA dataset construction module, which is used to integrate miRNA mutation information from different sources and its classification labels of "harmful", "uncertain", and "harmless", and obtain an miRNA mutation dataset after data cleaning, standardization, and deduplication;

[0041] An miRNA feature extraction module, which extracts three types of features through various miRNA prediction tools, namely, the miRNA energy feature, the miRNA interaction feature of the miRNA mutation dataset, and the miRNA mutation harmfulness score of the existing prediction methods;

[0042] The training module, based on the improved Transformer model, captures its mutation information and classification labels from the miRNA mutation dataset, learns the interrelationships among the above three types of features, and finally obtains a trained improved Transformer model; this improved Transformer model improves the feature processor and attention mechanism of the Transformer model; the improvement to the feature processor is that in the input of the feature processor, the above three types of features are introduced; the improvement to the attention mechanism is that the original single attention mechanism is changed to a local attention mechanism and a dynamic sparse attention mechanism, where the local attention mechanism captures features at adjacent positions by restricting the attention calculation within a local window, and the dynamic sparse attention mechanism calculates by dynamically selecting key positions;

[0043] The harmfulness prediction module realizes the harmfulness prediction of miRNA by introducing a GMM clustering model into the feature matrix obtained by the improved Transformer model, and classifies the mutations in the miRNA mutation dataset into three key categories: harmful, uncertain, and harmless.

[0044] Specifically, the miRNA dataset construction module integrates four databases related to miRNA mutations, namely miRbase, ClinVar, dbSNP, and COSMIC, and then adopts the union processing method to summarize miRNA mutation data, miRNA sequences before and after mutation, and classification labels with clinical significance to obtain a miRNA mutation dataset; then, the miRNA mutation dataset after union is cleaned: first, according to the statistical method Z-score, miRNA mutation data beyond the normal range and with missing information are removed, and then the classification labels are standardized and unified into "harmful", "uncertain", and "harmless", and then the deduplication operation is performed to ensure that each miRNA mutation data is unique in the miRNA mutation dataset; finally, a miRNA mutation dataset with miRNA mutation data, miRNA sequences before and after mutation, and classification labels is obtained.

[0045] Specifically, the miRNA feature extraction module performs the following operations:

[0046] Obtain miRNA energy features: Through the RNAFold prediction tool, input the miRNA mutation data in the miRNA mutation dataset to obtain the energy values and energy change values before and after miRNA mutation;

[0047] Obtain miRNA interaction features: Obtain the FASTA files of all sequences of three types of target genes, namely lncRNA, 3'UTR, and circRNA, from the LncBase database, TargetScan database, and CircBase database respectively. Then, input the above files and the miRNA sequences before and after mutation into the miRanda prediction tool to obtain a list of three types of target genes, namely lncRNA, 3'UTR, and circRNA, that interact with the miRNA mutation dataset. Count the numerical values and change values of the interactions between miRNA before and after mutation and the three types of target genes, namely lncRNA, 3'UTR, and circRNA, and generate one-hot encoding.

[0048] Obtain the miRNA mutation harmful scores of existing prediction methods: Input the miRNA mutation dataset into existing prediction tools related to miRNA mutation harm prediction, and extract the harmful scores generated by each method.

[0049] Specifically, the training module performs the following operations:

[0050] a. Input the miRNA mutation dataset into the improved Transformer model, add positional encoding to retain the positional information of the miRNA sequences before and after mutation in the miRNA mutation dataset. The miRNA mutation dataset is converted from discrete data to continuous vector H through the embedding layer of the improved Transformer model embed , and the mapping of each data is realized through the embedding matrix of the embedding layer. Then, input the continuous vector H embed output by the embedding layer and the features obtained by the miRNA feature extraction module into the feature processor to generate token-level embedding H token ; Then, after layer normalization and linear transformation of the model, it is output to the encoder of the improved Transformer model. Among them, the improvement of the feature processor is: introduce 3 types of features related to miRNA: miRNA energy feature F energy , miRNA interaction feature F interaction and the miRNA mutation harmful score F score of existing prediction methods, which increases the diversity of input features, is conducive to generating richer token-level embedding H token , improves the accuracy of the improved Transformer model, and is expressed as follows:

[0051] H token = Concat(H embed , F energy , F interaction , F score )

[0052] b. The encoder receives the token-level embedding H output by the embedding layertoken and processed by an improved attention mechanism; the encoder transforms and combines each token-level embedding multiple times, extracts deeper features, and generates a new feature representation through a feed-forward neural network; finally, the encoder completes the processing output of all token-level embeddings to obtain a feature matrix, where each row of the matrix corresponds to the feature representation of each miRNA mutation in the miRNA mutation dataset; among them, the improvement of the attention mechanism is: changing the original single attention mechanism to a multi-head attention mechanism, which is beneficial to taking advantage of the advantages of different attention mechanisms, improving the efficiency of the model in processing long sequence data, reducing computational complexity, and at the same time retaining important global information.

[0053] Specifically, the multi-head attention mechanism consists of a local attention mechanism and a dynamic sparse attention mechanism;

[0054] The local attention mechanism defines a window w of a fixed size, calculates the attention weights within this window, and after inputting the token-level embeddings, only focuses on the current token i and its surrounding tokens to extract local context information; the query vector Q of the current token i i will be calculated with the key vectors K [i-w,i+w] within the window to generate the attention weights for the surrounding tokens; then, the attention weights are used to weight the value vectors V [i-w,i+w] within the window, thereby obtaining the representation of the local context; the process of the local attention mechanism is represented as follows:

[0055]

[0056] In the formula, LocalAttention represents the local attention mechanism, d k is the dimension of the key vector;

[0057] The dynamic sparse attention mechanism uses the query matrix Q of all tokens of the token-level embeddings to calculate the attention weights related to the selected set of key tokens S; then, based on the attention weights, the corresponding key matrix K S and value matrix V S are selected to reduce the computational amount while retaining the key information; the process of the dynamic sparse attention mechanism is represented as follows:

[0058]

[0059] In the formula, SparseAttention represents the dynamic sparse attention mechanism.

[0060] Specifically, the harmfulness prediction module performs the following operations:

[0061] a. Input the feature matrix obtained by improving the encoder output of the Transformer model into the GMM clustering model for clustering operations. Initialize the parameters of the GMM clustering model, set the covariance in the parameters as the identity matrix, and set the initial values of the mixing coefficients as equal values.

[0062] b. Use the EM algorithm for iterative optimization: In the E-step, calculate the posterior probability that the feature vector belongs to harmful or harmless after passing through the i'-th Gaussian mixture component in the GMM clustering model; in the M-step, calculate the GMM coefficients of X classification clusters. After one iteration of the EM algorithm, update the model parameters and enter the next iteration until the change in the model parameters is less than 10 -4 ;

[0063] c. Use the optimized GMM clustering model to calculate the probability P that each miRNA mutation data in the miRNA mutation dataset belongs to the harmful class. If P > 0.5, then classify this miRNA mutation data as the harmful class; otherwise, classify it as the harmless class. Traverse each miRNA mutation data in the harmful and harmless classes obtained by the GMM clustering model. If the posterior probability is greater than the preset threshold, it is considered that this miRNA mutation data is acceptable at the confidence level of this category and is retained in the original category; otherwise, it is considered that the confidence is insufficient and it is assigned to the uncertain class.

[0064] d. On the basis of completing GMM clustering and obtaining the probability P of each miRNA mutation data, calculate the miRNA mutation score S, where S ∈ [0, 1]. Here, 0 is considered harmless and 1 is considered harmful. If the clustering category of a certain miRNA mutation data is harmful, then S = P; if the clustering category is harmless, then S = 1 - P.

[0065] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. The miRNA mutation harmfulness prediction system based on improved Transformer is characterized by: include: The miRNA dataset construction module is used to integrate miRNA mutation information from different sources and their classification labels of "harmful", "uncertain" and "harmless". After data cleaning, standardization and deduplication, the miRNA mutation dataset is obtained; The miRNA feature extraction module uses various miRNA prediction tools to extract three types of features, namely, miRNA energy features, miRNA interaction features, and miRNA mutation harmfulness scores of existing prediction methods. The training module, based on the improved Transformer model, captures the mutation information and classification labels from the miRNA mutation dataset, learns the relationship between the above three types of features, and finally obtains a trained improved Transformer model; the improved Transformer model is an improvement on the feature processor and attention mechanism of the Transformer model; the improvement on the feature processor is: the above three types of features are introduced into the input of the feature processor; The improvement of the attention mechanism is: the original single attention mechanism is changed into a local attention mechanism and a dynamic sparse attention mechanism. The local attention mechanism captures the features of adjacent positions within a local window by limiting the attention calculation, and the dynamic sparse attention mechanism performs calculation by dynamically selecting key positions. The harmfulness prediction module predicts the harmfulness of miRNA by introducing the GMM clustering model into the feature matrix obtained by the improved Transformer model, and divides the mutations in the miRNA mutation dataset into three key categories: harmful, uncertain, and harmless.

2. The miRNA mutation harmfulness prediction system based on improved Transformer according to claim 1, characterized in that: The miRNA dataset construction module integrates four databases related to miRNA mutations, namely miRbase, ClinVar, dbSNP, and COSMIC, and then uses a union processing method to summarize miRNA mutation data, sequences before and after miRNA mutations, and classification labels with clinical significance to obtain a miRNA mutation dataset; then the miRNA mutation dataset after the union is cleaned: first, miRNA mutation data that exceeds the normal range and has missing information are eliminated according to the statistical method Z-score, and then the classification labels are standardized and unified into "harmful", "uncertain" and "harmless", and then a deduplication operation is performed to ensure that each miRNA mutation data is unique in the miRNA mutation dataset; finally, a miRNA mutation dataset with miRNA mutation data, sequences before and after miRNA mutations, and classification labels is obtained.

3. The miRNA mutation harmfulness prediction system based on improved Transformer according to claim 2, characterized in that: The miRNA feature extraction module performs the following operations: Obtain miRNA energy characteristics: Use the RNAFold prediction tool to input the miRNA mutation data in the miRNA mutation dataset to obtain the energy value before and after the miRNA mutation and the energy change value; Obtain miRNA interaction features: Obtain the FASTA files of the complete sequences of three target genes, lncRNA, 3'UTR and circRNA, from the LncBase database, TargetScan database and CircBase database respectively, then input the above files and the sequences before and after miRNA mutation into the miRanda prediction tool to obtain the list of three target genes, lncRNA, 3'UTR and circRNA, that interact with the miRNA mutation dataset; Count the number and change values ​​of the three target genes, lncRNA, 3'UTR and circRNA, that interact with miRNA before and after mutation, and generate a one-hot encoding; Obtain the harmfulness scores of miRNA mutations from existing prediction methods: Input the miRNA mutation dataset into existing prediction tools related to the harmfulness prediction of miRNA mutations, and extract the harmfulness scores generated by each method.

4. The miRNA mutation harmfulness prediction system based on improved Transformer according to claim 3, characterized in that: The training module performs the following operations: a. Input the miRNA mutation dataset into the improved Transformer model, add position encoding to retain the position information of the sequences before and after the miRNA mutation in the miRNA mutation dataset; the miRNA mutation dataset is converted from discrete data to a continuous vector H through the embedding layer of the improved Transformer model. embed , and the mapping of each data is realized through the embedding matrix of the embedding layer; then, the continuous vector H output by the embedding layer is embed The features obtained by the miRNA feature extraction module are input into the feature processor to generate the token-level embedding H token ; Then, after the model layer normalization and linear transformation, it is output to the encoder of the improved Transformer model; Among them, the improvement of the feature processor is: introducing three types of features related to miRNA: miRNA energy feature F energy , miRNA interaction characteristics F interaction The harmfulness score F of miRNA mutations compared with existing prediction methods score , which increases the diversity of input features and is conducive to generating richer token-level embeddings H token , improve the accuracy of the improved Transformer model, expressed as follows: H token =Concat(H embed ,F energy ,F interaction ,F score ) b. The encoder receives the token-level embedding H output by the embedding layer token , and processed by an improved attention mechanism; the encoder transforms and combines each token-level embedding multiple times to extract deeper features, and generates new feature representations through a feedforward neural network; finally, the encoder completes the processing and output of all token-level embeddings to obtain a feature matrix, in which each row of the matrix corresponds to the feature representation of each miRNA mutation in the miRNA mutation dataset; among them, the improvement of the attention mechanism is: changing the original single attention mechanism to a multi-head attention mechanism, which is conducive to utilizing the advantages of different attention mechanisms, improving the efficiency of the model in processing long sequence data, reducing computational complexity, and retaining important global information.

5. The miRNA mutation harmfulness prediction system based on improved Transformer according to claim 4, characterized in that: The multi-head attention mechanism consists of a local attention mechanism and a dynamic sparse attention mechanism; The local attention mechanism defines a fixed-size window w, calculates the attention weight within the window, and after inputting the token-level embedding, only pays attention to the current token i and its surrounding tokens to extract local context information; the query vector Q of the current token i i will be combined with the key vector K in the window [i-w,i+w] Perform calculations to generate attention weights for surrounding tokens; The attention weights are then used to weight the value vector V within the window [i-w,i+w] , thereby obtaining the representation of the local context; the process of the local attention mechanism is expressed as follows: Where LocalAttention represents the local attention mechanism, d k is the dimension of the key vector; The dynamic sparse attention mechanism uses the query matrix Q of all tokens embedded at the token level to calculate the attention weights associated with the selected key token set S; then, the corresponding key matrix K is selected based on the attention weights S Sum value matrix V S , to reduce the amount of computation while retaining key information; the process of the dynamic sparse attention mechanism is as follows: Where SparseAttention represents the dynamic sparse attention mechanism.

6. The miRNA mutation harmfulness prediction system based on improved Transformer according to claim 5, characterized in that: The harmfulness prediction module performs the following operations: a. Input the feature matrix output by the encoder of the improved Transformer model into the GMM clustering model for clustering operation, wherein the parameters of the GMM clustering model are initialized, the covariance in the parameters is set to the unit matrix, and the initial value of the mixing coefficient is set to an equal value; b. Use the EM algorithm for iterative optimization: Step E calculates the eigenvector through the i′th Gaussian mixture component in the GMM clustering model to obtain the posterior probability of it being harmful or harmless; Step M calculates the GMM coefficients of X classification clusters; After completing one EM algorithm iteration, update the model parameters and enter the next iteration until the change in model parameters is less than 10 -4 ; c. Use the optimized GMM clustering model to calculate the probability P of each miRNA mutation data in the miRNA mutation data set belonging to the harmful class. If P>0.5, the miRNA mutation data is classified into the harmful class, otherwise it is classified into the harmless class. Traverse each miRNA mutation data in the harmful class and harmless class obtained by the GMM clustering model. If the posterior probability is greater than the preset threshold, it is considered that the miRNA mutation data is in this category at an acceptable confidence level and is retained in the original category; otherwise, it is considered that the confidence level is insufficient and is assigned to the uncertain class. d. After completing GMM clustering and obtaining the probability P of each miRNA mutation data, calculate the miRNA mutation score S, S∈[0,1], where 0 is considered harmless and 1 is considered harmful. If the clustering category of a miRNA mutation data is harmful, then S=P; if the clustering category is harmless, then S=1-P.

Citation Information

Patent Citations

  • Rotation forest algorithm based miRNA-disease correlation predicting method

    CN110400600A

  • Transform attention mechanism and 2D-3D feature cross fusion-based drug response prediction model

    CN117912590A

  • Multi-channel attention mechanism lncRNA-miRNA association prediction method

    CN119207579A

  • Sparse code multiple access encoding and decoding system based on generative adversarial network

    WO2024016424A1