MiRNA subcellular localization prediction method and system based on hypergraph and contrast learning

By constructing a hypergraph network and performing comparative learning and fusion features, the problem of insufficient information utilization in miRNA subcellular localization prediction is solved, and higher prediction accuracy is achieved.

CN120673854APending Publication Date: 2025-09-19HUNAN UNIV OF CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510554284.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing miRNA subcellular localization prediction methods fail to fully utilize the association information between miRNA, mRNA and disease, resulting in low prediction accuracy.

Method used

A hypergraph network was constructed, combining miRNA sequence similarity, functional similarity and disease association, and a miRNA subcellular localization prediction model was established through comparative learning and fusion features.

Benefits of technology

The accuracy of miRNA subcellular localization prediction is improved, and the prediction effect is enhanced by comprehensively utilizing multiple similarity data and embedded features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673854A_ABST
    Figure CN120673854A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, and discloses an miRNA subcellular localization prediction method and system based on hypergraph and contrast learning. The method comprises the following steps: extracting miRNA-disease associated characteristics and miRNA-mRNA associated network characteristics of a miRNA function similarity network, a miRNA-mRNA associated network and a miRNA-disease associated network by adopting a constructed hypergraph; carrying out comparative learning fusion processing on the miRNA sequence features, the miRNA-disease associated features and the miRNA-mRNA associated network features to obtain comparative learning fusion features; and establishing a miRNA subcellular localization prediction model according to the contrast learning fusion features, and carrying out localization based on the miRNA subcellular localization prediction model. The correlation between miRNA and mRNA and diseases is fully utilized, and positive and negative samples are compared by contrast learning, so that the accuracy of miRNA subcellular localization prediction is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning technology, and in particular to a method and system for predicting miRNA subcellular localization based on hypergraph and comparative learning. Background Art

[0002] MicroRNAs (miRNAs) are a class of small, non-coding RNA molecules that play a key role in regulating gene expression. MiRNAs have shown potential new functions in other subcellular compartments. For example, miRNAs mature in the cytoplasm and can translocate back to the nucleus. The most important function of these nuclear miRNAs is their ability to bind to complementary sequences within cis-regulatory elements (CREs), such as promoters and enhancers, thereby regulating the transcriptional activation or repression of target genes. MiRNAs are also present in the nucleolus, where they can influence ribosomal RNA (rRNA) in the cell to exert their biological activities. When miRNAs enter the mitochondria, they bind to mtDNA-encoded mRNAs, thereby regulating mitochondrial-related functions, particularly in cancer cells. Membrane-derived extracellular vesicles (EVs), such as microvesicles, exosomes, and extracellular vesicles, are key mediators of intercellular communication. MiRNAs have specific physiological roles in different cellular sites, and their subcellular localization is crucial for a deeper understanding of their physiological functions.

[0003] Currently, some researchers extract miRNA sequence features and conservation information and combine them with support vector machines to predict subcellular localization. Others, based on deep learning and graph neural networks, utilize miRNA sequence features and functional network graphs to capture topological information for prediction. Still others employ statistical features and traditional machine learning methods, such as random forests and support vector machines, to extract key features from nucleotide composition and sequence properties for prediction. Others have constructed a network-based method, MiRLoc, for detecting miRNA subcellular localization, leveraging miRNA-associated mRNAs and their subcellular localization. Others have proposed the DAmiRLocGNet method, which extracts miRNA features from sequence and miRNA-disease associations. It outperforms MiRLoc in detecting miRNA subcellular localization. While these methods demonstrate good performance in determining miRNA subcellular localization, they also have significant limitations. Most methods utilize only a single miRNA feature, resulting in incomplete representation of the miRNA. This incomplete representation can compromise prediction performance.

[0004] As can be seen, existing research on miRNA subcellular localization prediction is effective, but it still has many shortcomings. On the one hand, some methods do not fully utilize the useful information of miRNA-related small molecules and do not comprehensively consider various similarity data. On the other hand, some methods fail to make good use of the learned miRNA embedding features, resulting in unsatisfactory prediction results. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for predicting miRNA subcellular localization based on hypergraph and comparative learning, so as to solve the technical problem that the existing miRNA subcellular localization prediction technology does not make sufficient use of information, resulting in low prediction accuracy.

[0006] To achieve the above objectives, the present invention provides a first aspect of a method for predicting miRNA subcellular localization based on a hypergraph and contrastive learning, comprising the following steps:

[0007] S1: Determine miRNA sequence similarity network, miRNA-mRNA association network and miRNA-disease association network; and extract miRNA sequence features from the miRNA sequence similarity network;

[0008] S2: Construction of miRNA functional similarity network;

[0009] S3: Construction of hypergraph based on miRNA functional similarity network, miRNA-mRNA association network and miRNA-disease association network;

[0010] S4: extracting miRNA-disease association features and miRNA-mRNA association network features of the miRNA function similarity network, miRNA-mRNA association network, and miRNA-disease association network based on the hypergraph; and performing comparative learning and fusion processing on the miRNA sequence features, miRNA-disease association features, and miRNA-mRNA association network features to obtain comparative learning and fusion features;

[0011] S5: establishing a miRNA subcellular localization prediction model according to the comparative learning fusion features, and performing localization based on the miRNA subcellular localization prediction model.

[0012] In a second aspect, the present application provides a miRNA subcellular localization prediction system based on hypergraph and contrastive learning, comprising:

[0013] An extraction module is used to determine the miRNA sequence similarity network, the miRNA-mRNA association network, and the miRNA-disease association network; and extract miRNA sequence features from the miRNA sequence similarity network;

[0014] The first building block is used to construct the miRNA functional similarity network;

[0015] The second building module is used to construct a hypergraph based on the miRNA functional similarity network, miRNA-mRNA association network and miRNA-disease association network;

[0016] A learning module is used to extract miRNA-disease association features and miRNA-mRNA association network features of the miRNA functional similarity network, miRNA-mRNA association network, and miRNA-disease association network based on the hypergraph; and perform comparative learning and fusion processing on the miRNA sequence features, miRNA-disease association features, and miRNA-mRNA association network features to obtain comparative learning fusion features;

[0017] A positioning module is used to establish a miRNA subcellular localization prediction model according to the comparative learning fusion feature, and perform positioning based on the miRNA subcellular localization prediction model.

[0018] The present invention has the following beneficial effects:

[0019] The present invention's hypergraph-based and contrastive learning-based miRNA subcellular localization prediction method uses a constructed hypergraph to extract miRNA-disease association features and miRNA-mRNA association network features from miRNA functional similarity networks, miRNA-mRNA association networks, and miRNA-disease association networks; then performs contrastive learning fusion processing on miRNA sequence features, miRNA-disease association features, and miRNA-mRNA association network features to obtain contrastive learning fusion features; establishes a miRNA subcellular localization prediction model based on the contrastive learning fusion features, and performs localization based on the miRNA subcellular localization prediction model. In this way, the association between miRNAs, mRNAs, and diseases is fully utilized, and contrastive learning is used to compare positive and negative samples, effectively improving the accuracy of miRNA subcellular localization prediction.

[0020] In addition to the above-described objects, features and advantages, the present invention has other objects, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0022] Figure 1 This is one of the flow charts of the miRNA subcellular localization prediction method based on hypergraph and comparative learning in a preferred embodiment of the present invention;

[0023] Figure 2 This is the second flow chart of the miRNA subcellular localization prediction method based on hypergraph and comparative learning in a preferred embodiment of the present invention;

[0024] Figure 3 This is a structural block diagram of a miRNA subcellular localization prediction system based on hypergraph and comparative learning in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0025] The embodiments of the present invention are described in detail below with reference to the accompanying drawings. However, the present invention can be implemented in many different ways as defined and covered by the claims.

[0026] Currently, some methods do not fully utilize useful information about miRNA-related small molecules and do not comprehensively consider various similarity data. On the other hand, some methods fail to make good use of the learned miRNA embedding features, resulting in suboptimal prediction results. Based on this, this application provides a method for predicting miRNA subcellular localization based on hypergraphs and comparative learning.

[0027] See Figure 1-Figure 2 , this application provides a miRNA subcellular localization prediction method based on hypergraph and contrastive learning, including:

[0028] S1: Determine miRNA sequence similarity network, miRNA-mRNA association network and miRNA-disease association network; and extract miRNA sequence features from the miRNA sequence similarity network;

[0029] S2: Construction of miRNA functional similarity network;

[0030] S3: Construction of hypergraph based on miRNA functional similarity network, miRNA-mRNA association network and miRNA-disease association network;

[0031] S4: extracting miRNA-disease association features and miRNA-mRNA association network features of the miRNA function similarity network, miRNA-mRNA association network, and miRNA-disease association network based on the hypergraph; and performing comparative learning and fusion processing on the miRNA sequence features, miRNA-disease association features, and miRNA-mRNA association network features to obtain comparative learning and fusion features;

[0032] S5: establishing a miRNA subcellular localization prediction model according to the comparative learning fusion features, and performing localization based on the miRNA subcellular localization prediction model.

[0033] It should be understood that the miRNA-disease association data in this application are obtained from the Human miRNA Disease Database (HMDD, version 3.2), and the miRNA-mRNA association data are extracted from the database miRTarBase 2020. The miRNA sequence information is all derived from the miRBase database.

[0034] Extracting miRNA sequence features from the miRNA sequence similarity network can be done by using Node2Vec to extract key features from the miRNA sequence similarity network, the miRNA-mRNA association network, and the miRNA-disease association network, respectively, to construct an initial miRNA sequence feature representation.

[0035] The aforementioned hypergraph-based and contrastive learning miRNA subcellular localization prediction method uses a constructed hypergraph to extract miRNA-disease association features and miRNA-mRNA association network features from the miRNA functional similarity network, miRNA-mRNA association network, and miRNA-disease association network. The miRNA sequence features, miRNA-disease association features, and miRNA-mRNA association network features are then subjected to contrastive learning and fusion processing to obtain contrastive learning fusion features. A miRNA subcellular localization prediction model is established based on the contrastive learning fusion features, and localization is performed based on the miRNA subcellular localization prediction model. This method fully utilizes the associations between miRNAs, mRNAs, and diseases, and uses contrastive learning to compare positive and negative samples, effectively improving the accuracy of miRNA subcellular localization prediction.

[0036] As a further improvement of the method of the present invention: the miRNA sequence similarity network uses the Smith-Waterman algorithm to measure the similarity of two miRNA sequences, satisfying the following relationship; the miRNA-mRNA association network is extracted from the database miRTarBase 2020; the miRNA-disease association network is extracted from the miRNA and disease associations reported in the Human miRNA Disease Database (HMDD, version 3.2).

[0037]

[0038] Among them, sp(m i , m j ) represents m i and m j The Smith-Waterman alignment score between sequences, and sp(m i , m i ) and sp(m j , m j ) are the comparison scores of miRNA itself.

[0039] Among them, the miRNA functional similarity network is constructed based on miRNA diseases, including the following steps:

[0040] In the Medical Subject Headings (MeSH), all diseases are organized into a directed acyclic graph (DAG). Each disease d i It can be represented as a subgraph G(d i ), including diseases i and all ancestor nodes. Define a disease d t For disease i Contribution value:

[0041]

[0042] Where λ is the attenuation factor, which is used to measure the contribution of the child node to the parent node. In one example, the recommended value is 0.5. i The semantic value SV(d i ):

[0043]

[0044] Among them, A(d i ) is included d i and the set of all its ancestor nodes. Based on the semantic value of the disease, define the disease d i and d j Semantic similarity:

[0045]

[0046] In the formula, SV(d j ) indicates disease d j The semantic value of A(d j ) is included d j and the set of all its ancestor nodes, Indicates disease t For disease j contribution value.

[0047] In this application, we use disease similarity to construct a miRNA functional similarity network, and use the association information between miRNA and disease to map disease similarity to miRNA. i and m j Represents miRNAi and miRNAj, defining the m of miRNA i and m j The set of diseases associated with each is D(m i ) and D(m j ). The functional similarity of two miRNAs is defined as:

[0048]

[0049] In order to reduce the proportion of zero values ​​in the similarity matrix, GIP core similarity is introduced into the fusion: if the functional similarity of two miRNAs Sim m (m i ,m j )>0, it is used directly; otherwise, the GIP kernel similarity GIP(m i ,m j ):

[0050]

[0051] The similarity matrix Sim between miRNAs is formed based on similarity m fused , based on the above similarity matrix Sim m fused , construct a miRNA functional similarity network: nodes represent miRNAs, and edges represent functional similarities. Set a threshold TTT to binarize the continuous values:

[0052]

[0053] Furthermore, hyperedges are generated. In this embodiment, hyperedge generation is based on a similarity threshold. First, the node similarity matrix E[i, j] is input and hyperedges are constructed with the miRNA-mRNA association network and the miRNA-disease association network, respectively. A threshold value is preset. The similarity matrix is ​​then thresholded, with elements in E[i, j] greater than or equal to the threshold value marked as 1, and all other elements marked as 0.

[0054]

[0055] Next, we traverse each node i and find the set of neighboring nodes N(i) with a similarity of 1. We then combine node i and the set of neighboring nodes into a hyperedge, and further obtain a hyperedge list based on the hyperedge. We need to ensure that each hyperedge contains at least two nodes to preserve valid high-order relationships. The formula can be expressed as:

[0056] E k ={i|i∈N(i), A[i,j]=1}.

[0057] In order to further construct the hypergraph, it is necessary to generate the adjacency matrix A of the hypergraph from the hyperedge list HyperGraph The matrix describes the connection relationship between nodes in the hypergraph. The specific steps are as follows:

[0058] First, initialize an all-zero matrix A of size N×N HyperGraph, where N is the total number of nodes. Traverse the hyperedge list, for each hyperedge E k , connect all nodes in it. If the node pair (i, j) in the hyperedge satisfies i≠j, then set A in the adjacency matrix HyperGraph [i,j]=1 and A HyperGraph [j,i]=1 ensures that the adjacency matrix is ​​symmetric, thus accurately representing the undirected nature of the hypergraph.

[0059] After constructing the hypergraph adjacency matrix, a hypergraph convolutional network (HGCN) is constructed based on the hypergraph adjacency matrix, and then the hypergraph convolutional network (HGCN) is used to extract the embedded features of the nodes. The hypergraph convolutional network gradually extracts the high-order features of the nodes through multi-layer feature propagation. The feature propagation formula for each layer is:

[0060]

[0061] Where, represents the normalized adjacency matrix, D represents the degree matrix, and H (l+1) represents the feature representation of the l+1th layer, σ represents the activation function, and H (l) represents the feature representation of the lth layer, W (l) represents the weight matrix.

[0062] Furthermore, the feature fusion is performed through contrast learning, which includes the following steps:

[0063] In the process of implementing contrastive learning, the input features need to be normalized first in order to calculate the cosine similarity. The purpose is to map each feature vector onto the unit sphere, thereby ensuring that the similarity calculation is not affected by the size of the vector. Next, positive and negative sample pairs are generated through data augmentation and random sampling. For positive sample pairs, the augmentation function Augment() is usually used to extract the original feature z. i Construct the enhanced sample z j The negative sample pairs are randomly selected from the feature set, which are different from z i Sample z k After obtaining the sample pairs, the similarity between the positive and negative sample pairs is calculated to measure the degree of similarity between the samples. The similarity between the positive and negative sample pairs satisfies the following relationship:

[0064]

[0065] Then, based on the generated similarity sin(z i , z j ), and optimize it using the NT-Xent loss function. The core of NT-Xent is to maximize the similarity of positive sample pairs while minimizing the similarity difference between positive and negative samples. Its formula is:

[0066]

[0067] Among them, τ is the temperature parameter used to control the smoothness of the similarity distribution. N represents the number of samples, sin(z i , z k ) represents the similarity of negative sample pairs, sin(z i , z j ) represents the similarity of the positive sample pairs. In the implementation, the similarity of the positive sample pairs is extracted by constructing the identity matrix, and the similarity of the negative sample pairs is extracted by excluding the positive samples. Finally, the loss value is calculated. To improve the nonlinear representation capability of features, a projection head is added to the model. The projection head is implemented using a multi-layer perceptron (MLP) and maps the input feature z into a new embedding space. Its formula is:

[0068] h=W2·ReLU(W1·z+b1)+b2;

[0069] Where z is the input feature vector, W1, W2 are weight matrices, and b1, b2 are bias terms used to adjust the output.

[0070] The miRNA subcellular localization prediction comprises the following steps:

[0071] After the contrastive learning pre-training is completed, the generated high-quality embedding z is input into the downstream task model for training. For example, in a multi-label classification task, the probability of each label is predicted using the fully connected layer, as follows:

[0072]

[0073] in, is the output of the model, representing the predicted probability of the i-th category. For multi-label classification problems, the model outputs the probability value of each label (obtained through the sigmoid function). Its value is between (0,1), indicating the probability that the sample belongs to that category. The loss function uses binary cross entropy:

[0074]

[0075] Where L represents the value of the loss function, which measures the accuracy of the model prediction. N represents the total number of samples. In multi-label classification, it usually refers to the number of samples in a batch. i This is the true label of the i-th label. For multi-label classification problems, y i The value of is usually 0 or 1, indicating whether the sample belongs to the label. If the sample belongs to the label, y i =1, otherwise y i =0.

[0076] As a general technical concept, Figure 3 As shown, the present application also provides a miRNA subcellular localization prediction system based on hypergraph and comparative learning, comprising:

[0077] An extraction module is used to determine the miRNA sequence similarity network, the miRNA-mRNA association network, and the miRNA-disease association network; and extract miRNA sequence features from the miRNA sequence similarity network;

[0078] The first building block is used to construct the miRNA functional similarity network;

[0079] The second building module is used to construct a hypergraph based on the miRNA functional similarity network, miRNA-mRNA association network and miRNA-disease association network;

[0080] A learning module is used to extract miRNA-disease association features and miRNA-mRNA association network features of the miRNA functional similarity network, miRNA-mRNA association network, and miRNA-disease association network based on the hypergraph; and perform comparative learning and fusion processing on the miRNA sequence features, miRNA-disease association features, and miRNA-mRNA association network features to obtain comparative learning fusion features;

[0081] A positioning module is used to establish a miRNA subcellular localization prediction model according to the comparative learning fusion feature, and perform positioning based on the miRNA subcellular localization prediction model.

[0082] Next, an experiment is conducted to verify the effectiveness of the above-mentioned miRNA subcellular localization prediction based on hypergraph and contrastive learning:

[0083] In order to evaluate the performance of the model, this example uses two comprehensive indicators, AUC (area under the curve) and AUPR (average precision-recall rate). The prediction accuracy of the miRNA subcellular localization prediction method is evaluated as follows:

[0084] Model training employed 10-fold cross-validation to ensure robustness and generalization. Area Under the Circumstances (AUC) and Average Percentile (AUPR) were used to measure prediction performance. The model training process was iterated 100 times to fully optimize parameters and achieve convergence. The average of all evaluation metrics was calculated after each training run. The experimental results are shown in Tables 1 and 2.

[0085] Table 1: AUC experimental results of HCLMSL of the present invention and other methods

[0086]

[0087]

[0088] Table 2: AUPR experimental results of HCLMSL of the present invention and other methods

[0089]

[0090] As can be seen from Tables 1 and 2, the miRNA subcellular localization prediction method of the present invention achieved an AUC of 0.9467 and an AUPR of 0.8664. All indicators of the HCLMSL model of the present invention are superior to the PMiSLocMF model.

[0091] The above describes in detail the preferred embodiments of the present application. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present application without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art based on the concepts of the present application through logical analysis, reasoning, or limited experimentation on the basis of the prior art should be within the scope of protection defined by the claims.

Claims

1. A miRNA subcellular localization prediction method based on hypergraph and contrastive learning, characterized in that: include: S1: Determine miRNA sequence similarity network, miRNA-mRNA association network and miRNA-disease association network; and extract miRNA sequence features from the miRNA sequence similarity network; S2: Construction of miRNA functional similarity network; S3: Construction of hypergraph based on miRNA functional similarity network, miRNA-mRNA association network and miRNA-disease association network; S4: extracting miRNA-disease association features and miRNA-mRNA association network features of the miRNA function similarity network, miRNA-mRNA association network, and miRNA-disease association network based on the hypergraph; and performing comparative learning and fusion processing on the miRNA sequence features, miRNA-disease association features, and miRNA-mRNA association network features to obtain comparative learning and fusion features; S5: establishing a miRNA subcellular localization prediction model according to the comparative learning fusion features, and performing localization based on the miRNA subcellular localization prediction model.

2. The miRNA subcellular localization prediction method based on hypergraph and contrastive learning according to claim 1, characterized in that The S2 includes: Organize all diseases in the medical subject headings into a directed acyclic graph DAG, and divide each disease d i Represented as a subgraph G(d i ), each subgraph includes disease d i and disease i For all corresponding ancestor nodes, define a disease d t For disease i The contribution value satisfies the following relationship: Among them, λ is the attenuation factor, which is used to measure the contribution of the child node to the parent node. Represents the child node d′ t The contribution value of disease d i The semantic value SV(d i ), satisfying the following relationship: Among them, A(d i ) is included d i and the set of all its ancestor nodes; Based on the semantic value of the disease, define the disease d i and d j The semantic similarity of satisfies the following relationship: In the formula, SV(d j ) indicates disease d j The semantic value of A(d j ) is included d j and the set of all its ancestor nodes, Indicates disease t For disease j Contribution value; Using the association information between miRNA and disease, we can map disease similarity to miRNA and define miRNAm i and m j The set of diseases associated with each is D(m i ) and D(m j ), the functional similarity of two miRNAs is defined as satisfying the following relationship: If the functional similarity of two miRNAs is m (m i ,m j )>0, it is used directly; otherwise, the GIP kernel similarity GIP(m i ,m j ), satisfying the following relationship: The similarity matrix Sim between miRNAs is formed based on similarity m fused , based on the similarity matrix Sim m fused A miRNA functional similarity network was constructed, where nodes represented miRNAs and edges represented functional similarities.

3. The miRNA subcellular localization prediction method based on hypergraph and contrastive learning according to claim 2, characterized in that After S2, the method further includes: Set a threshold T to binarize the similarity matrix to satisfy the following relationship:

4. The miRNA subcellular localization prediction method based on hypergraph and contrastive learning according to claim 1, characterized in that The S3 includes: Hyperedges were constructed based on miRNA functional similarity network, miRNA-mRNA association network and miRNA-disease association network; Generating an adjacency matrix of a hypergraph according to the hyperedges, wherein the adjacency matrix is ​​used to describe the connection relationship between nodes in the hypergraph; Generate a hypergraph convolutional network as described.

5. The method for predicting miRNA subcellular localization based on hypergraph and contrastive learning according to claim 4, wherein: The hyperedges are constructed based on the miRNA functional similarity network, the miRNA-mRNA association network and the miRNA-disease association network, including: Input the binarized similarity matrix E[i,j] and construct hyperedges with the miRNA-mRNA association network and miRNA-disease association network respectively. A threshold is preset, and then the similarity matrix is ​​thresholded. Elements in E[i,j] that are greater than or equal to the threshold are marked as 1, and other elements are marked as 0. Traverse each node i, find the set of neighboring nodes N(i) with a similarity of 1, and form a hyperedge with node i and the set of neighboring nodes. It is necessary to ensure that each hyperedge contains at least two nodes and satisfies the following relationship: E k ={i|i∈N(i),A[i,j]=1}。 6. The method for predicting miRNA subcellular localization based on hypergraph and contrastive learning according to claim 4, wherein: Generating an adjacency matrix of a hypergraph according to the hyperedges includes: Initialize an N×N all-zero matrix A HyperGraph , where N is the total number of nodes; Traverse the list of hyperedges, for each hyperedge E k , connect all nodes in pairs. If the node pair (i, j) in the hyperedge satisfies i≠j, then set A in the adjacency matrix HyperGraph [i,j]=1 and A HyperGraph [j,i]=1 ensures that the adjacency matrix is ​​symmetric; Generate the final adjacency matrix based on the traversal results.

7. The method for predicting miRNA subcellular localization based on hypergraph and contrastive learning according to claim 1, wherein The S4 includes: The embedded features of nodes are extracted using a hypergraph convolutional network. The hypergraph convolutional network gradually extracts high-order features of nodes through multi-layer feature propagation. The feature propagation formula of each layer satisfies the following relationship: Where, represents the normalized adjacency matrix, D represents the degree matrix, and H (l+1) represents the feature representation of the l+1th layer, σ represents the activation function, and H (l) represents the feature representation of the lth layer, W (l) represents the weight matrix; In the process of implementing contrastive learning, the embedded features of the extracted nodes are first taken as input and normalized; Then, positive and negative sample pairs are generated through data enhancement and random sampling. For the positive sample pairs, the enhancement function Augment() is used to extract the original feature z i Construct the enhanced sample z j ; For negative sample pairs, randomly select different i Sample z k structure; After obtaining the sample pairs, the similarity between the positive and negative sample pairs is calculated to measure the degree of similarity between the samples. The similarity between the positive and negative sample pairs satisfies the following relationship: Then, based on the generated similarity sin(z i , z j ), and optimize using the NT-Xent loss function. During optimization, the goal is to maximize the similarity of positive sample pairs while minimizing the similarity difference between positive and negative samples, satisfying the following relationship: Among them, τ is the temperature parameter used to control the smoothness of the similarity distribution, N is the number of samples, sin(z i , z k ) represents the similarity of negative sample pairs, sin(z i , z j ) represents the similarity of the positive sample pair. During the calculation, a projection head is added to the model to map the input feature z to a new embedding space, satisfying the following relationship: h=W2·ReLU(W1·z+b1)+b2; Among them, z is the input feature vector, W1 and W2 are weight matrices, b1 and b2 are bias terms used to adjust the output.

8. The method for predicting miRNA subcellular localization based on hypergraph and contrastive learning according to claim 1, wherein The S5 includes: After the contrastive learning pre-training is completed, the generated high-quality embedding h is input into the downstream task model for training to obtain the final miRNA subcellular localization prediction model, which satisfies the following relationship: in, Is the output of the model, which represents the predicted probability of the i-th category. For multi-label classification problems, the model outputs the probability value of each label, which is between (0,1), indicating the probability that the sample belongs to this category. The loss function satisfied by the model is calculated using binary cross entropy, which satisfies the following relationship: Where L represents the value of the loss function, which measures the accuracy of the model prediction, N represents the total number of samples, and y i is the true label of the i-th label. If the sample belongs to this label, y i =1, otherwise y i =0.

9. A miRNA subcellular localization prediction system based on hypergraph and contrastive learning, characterized in that: include: An extraction module is used to determine the miRNA sequence similarity network, the miRNA-mRNA association network, and the miRNA-disease association network; and extract miRNA sequence features from the miRNA sequence similarity network; The first building block is used to construct the miRNA functional similarity network; The second building module is used to construct a hypergraph based on the miRNA functional similarity network, miRNA-mRNA association network and miRNA-disease association network; A learning module is used to extract miRNA-disease association features and miRNA-mRNA association network features of the miRNA functional similarity network, miRNA-mRNA association network and miRNA-disease association network based on the hypergraph; and Performing comparative learning and fusion processing on the miRNA sequence features, miRNA-disease association features and miRNA-mRNA association network features to obtain comparative learning fusion features; A positioning module is used to establish a miRNA subcellular localization prediction model according to the comparative learning fusion feature, and perform positioning based on the miRNA subcellular localization prediction model.