Optimization method of non-coding RNA (Ribonucleic Acid) and disease association prediction
By constructing heterogeneous graphs through a self-supervised graph contrastive learning method and extracting node embedding features using a graph convolutional network, the time-consuming and labor-intensive problem of predicting the association between non-coding RNA and diseases was solved, achieving high-precision and robust prediction results.
Patent Information
- Application Number
- CN202510714170.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies are time-consuming and labor-intensive in predicting the association between non-coding RNA and diseases, with a low success rate, making it difficult to achieve accurate predictions.
A self-supervised graph contrastive learning method is adopted to construct heterogeneous graphs and extract node embedding features through graph convolutional networks. Positive and negative sample views are generated by using data augmentation and data destruction. The fused GCN encoder is trained and combined with the XGBoost model for correlation information prediction.
It achieves high-precision prediction of non-coding RNA and disease association, improves the robustness and prediction accuracy of the model, and enables accurate prediction on different data sets.
Smart Images

Figure CN120636841A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to an optimization method for predicting the association between non-coding RNA and diseases. Background Art
[0002] Although long non-coding RNAs (lncRNAs) do not encode proteins in the human body, they possess a variety of biological functions. Experiments have confirmed that lncRNAs significantly influence many human diseases and cancers. Therefore, predicting the association between lncRNAs and diseases is of great significance. However, existing approaches to exploring the association between lncRNAs and diseases are time-consuming and labor-intensive, with low success rates. Although researchers have made significant efforts in the field of lncRNA-disease association prediction and have proposed numerous lncRNA-disease association prediction methods, they have failed to capture accurate lncRNA-disease information, making precise prediction difficult. Summary of the Invention
[0003] In order to solve the problems existing in the above-mentioned prior art, the purpose of the present invention is to propose an optimization method for predicting the association between non-coding RNA and disease, which fully captures node features based on self-supervised graph comparative learning and achieves high-precision lncRNA-disease association prediction.
[0004] To achieve the above object, the present invention provides the following solutions:
[0005] An optimization method for predicting the association between non-coding RNA and disease, comprising:
[0006] Obtaining a similarity matrix of diseases and non-coding RNAs, and constructing a heterogeneous graph based on the similarity matrix;
[0007] Processing the heterogeneous graph to generate different views, respectively inputting the different views into an encoder fused with a GCN to obtain embedded feature vectors; the encoder fused with the GCN is trained using a first training set; the first training set includes: original views;
[0008] Based on the embedded feature vector, association information between the disease and the non-coding RNA is obtained.
[0009] Optionally, obtaining the similarity matrix of the disease includes:
[0010]
[0011] Among them, DS(i,j) is the similarity matrix of the disease, DSS(i,j) is the semantic similarity of different diseases, and DGS(i,j) is the Gaussian kernel similarity of the disease.
[0012] Optionally, obtaining the similarity matrix of the non-coding RNA includes:
[0013]
[0014] Among them, LS(i,j) is the similarity matrix of non-coding RNA, LFS(i,j) is the functional similarity of non-coding RNA, and LGS(i,j) is the Gaussian kernel similarity of non-coding RNA.
[0015] Optionally, constructing the heterogeneous graph includes:
[0016] G H =(V H ,E H )
[0017] V H =V lncRNA ∪V disease
[0018] E H =E IncRNA-disease ∪E IncRNA-IncRNA ∪E disease-disease
[0019] Among them, G H is a heterogeneous graph, V H is a node set, E H is the edge set, V lncRNA is the non-coding RNA node set, V disease is the disease node set, E IncRNA-disease is the edge set of non-coding RNA nodes and diseases, E IncRNA-IncRNA is the edge set of different non-coding RNA nodes, E disease-disease is the edge set of different diseases.
[0020] Optionally, generating the different views includes:
[0021] Perform data augmentation on the heterogeneous graph to obtain a positive sample view:
[0022] G H ′=Dropout(G H ,p)
[0023] Among them, G H ' is the positive sample view, p is the dropout probability;
[0024] Perform data destruction on the heterogeneous graph to obtain a negative sample view:
[0025]
[0026] in, is the negative sample view, G H It is a heterogeneous graph.
[0027] Optionally, obtaining the embedded feature vector includes:
[0028] Z=f GCNEncoder (HLD,X)
[0029] Z A =f AugmentEncoder (HLD-A,X A )
[0030] Z C =f CorruptEncoder (HLD-C,X C )
[0031] Among them, HLD is a heterogeneous graph, HLD-A is a positive sample view, HLD-C is a negative sample view, Z is the node embedding vector of the fused heterogeneous graph, and f GCNEncoder The extracted node embedding vector that integrates the original heterogeneous graph information, Z A f AugmentEncoder The node embedding vector of the extracted fused enhanced graph information, Z C f CorruptEncoder The node embedding vector that integrates the destruction graph information, f GCNEncoder For the GCN encoder that processes the original heterogeneous graph, f AugmentEncoder To process the GCN encoder for enhanced heterogeneous graphs, f CorruptEncoder To process the GCN encoder that destroys the heterogeneous graph, X is the original heterogeneous graph, X A To enhance heterogeneous graphs, X C To destroy heterogeneous graphs.
[0032] Optionally, obtaining the encoder of the fused GCN includes:
[0033] Using the first training set to train the encoder of the original fused GCN to obtain the encoder of the fused GCN;
[0034] During the training process, the encoder of the original fusion GCN is trained with the objective function of maximizing the similarity between the heterogeneous graph embedding and the positive sample view embedding while minimizing the similarity between the heterogeneous graph embedding and the negative sample view embedding.
[0035] Optionally, the objective function includes:
[0036] L contrastive =(1-cos(Z,Z A )+max(0,cos(Z,Z C )-m))
[0037] Among them, L contrastive is the contrastive learning loss function, cos(Z,Z A ) is the similarity between the original node embedding vector and the enhanced node embedding vector, cos(Z,Z C ) is the similarity between the original node embedding vector and the corrupted node embedding vector, and m is the minimum similarity difference used to control the positive and negative samples.
[0038] Optionally, obtaining the association information between the disease and the non-coding RNA includes:
[0039] The embedded feature vector is reconstructed to obtain reconstructed features, and the reconstructed features are input into an XGBoost model to obtain association information between the disease and the non-coding RNA; the XGBoost model is trained using a second training set, and the second training set includes: original reconstructed features.
[0040] The beneficial effects of the present invention are:
[0041] The present invention fully captures node features based on self-supervised graph comparative learning and achieves high-precision lncRNA-disease association prediction.
[0042] This paper mainly uses a graph comparative learning framework, which can effectively fuse lncRNA and disease-related information to form a heterogeneous graph. It uses a graph comparative learning method to effectively extract the global information of lncRNAs and diseases, and achieve more accurate predictions on different data sets.
[0043] The present invention constructs a data augmentation graph encoder and a data erosion graph encoder to generate data augmented embedding vectors and data eroded embedding vectors respectively, so that the model can perform good learning without relying on certain features, which can greatly improve the robustness of the model and prevent the model from overfitting.
[0044] The present invention constructs positive and negative sample pairs based on the embedded vector information extracted by the graph convolutional encoder, forming a graph contrastive learning self-supervised architecture. This allows the model to be trained independently of the original label information, bringing similar sample pairs closer together and distinguishing dissimilar sample pairs, thereby effectively extracting global information about lncRNA-disease. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0046] Figure 1 This is a flow chart of an optimization method for predicting the association between non-coding RNA and disease according to an embodiment of the present invention. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0048] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0049] like Figure 1 As shown, this embodiment discloses an optimization method for predicting the association between non-coding RNA and disease, including: obtaining a similarity matrix between the disease and the non-coding RNA, and constructing a heterogeneous graph based on the similarity matrix and the association information between the disease and the non-coding RNA; performing data processing on the heterogeneous graph to generate different views, and inputting the different views into an encoder fused with GCN respectively to obtain embedded feature vectors; the encoder fused with GCN is trained using a first training set; the first training set includes: original views; and obtaining the association information between the disease and the non-coding RNA based on the embedded feature vectors.
[0050] Furthermore, obtaining the similarity matrix of diseases and non-coding RNAs includes:
[0051] The present invention constructs the semantic similarity of diseases. Based on the grid database, a directed acyclic graph (DAGs) related to the disease is constructed, where the vertices of the DAG represent the diseases and the edges of the DAG represent the relationships between the vertices. Taking disease A as an example, the DAG A =(T A ,E A ), where T represents a node set consisting of the current node and its ancestor nodes, and E represents a set of edges connecting parent nodes and child nodes. i The contribution to the semantic value of disease A is defined as:
[0052]
[0053] Where Δ is the semantic attenuation factor, usually set to 0.5, D A (d i ) represents node d i The score to node A, D A (d i ′) represents node d′ iThe score to node A.
[0054] The semantic value of disease A is defined as:
[0055]
[0056] The similarity between disease A and disease B is calculated:
[0057]
[0058] Among them, D B (d i ) is the disease d i The contribution to the semantic value of disease B, DV(B) is the semantic value of disease B.
[0059] Based on the fact that diseases may be associated with lncRNAs with similar functions, the functional similarity of lncRNAs was calculated:
[0060]
[0061] Among them, S(d,D) is the similarity between node d and set D, S(d,d i ) are nodes d and d i The similarity value of , k is the number of elements in the set, LFS(l1,l2) is the functional similarity value of nodes l1 and l2, S(d i ,D2) is node d i The similarity value with D2, S(d i ,D1) is node d i The similarity value with D1, m and n are the number of elements between the two structures.
[0062] Constructing Gaussian similarity LGS of lncRNAs:
[0063] LGS=exp(-γ l ||LD(i,:)-LD(j,:)|| 2 ) (6)
[0064] Among them, γ l It is used to control the kernel bandwidth, LD(i,:) is the feature of the i-th lncRNA, and LD(j,:) is the feature of the j-th lncRNA.
[0065] Controlling kernel bandwidth γ l are the parameters of the Gaussian kernel:
[0066]
[0067] Here, n represents the number of lncRNAs.
[0068] Construct DGS based on the Gaussian kernel similarity corresponding to the disease:
[0069] DGS=exp(-γ d ||LD(:,i)-LD(:,j)|| 2 ) (8)
[0070]
[0071] Among them, γ d is the parameter of the Gaussian kernel, and m is the total number of diseases.
[0072] The final disease similarity matrix DS is obtained by the semantic similarity of the disease and the GIPK similarity of the disease:
[0073]
[0074] The similarity matrix of lncRNAs is obtained by the functional similarity of lncRNAs and the GIPK similarity of lncRNAs:
[0075]
[0076] Furthermore, the construction of the heterogeneous graph includes:
[0077] When studying the complex relationship between lncRNAs and diseases, the construction of a heterogeneous graph can effectively represent these different types of biological entities and their interactions. The present invention uses the lncRNA-disease association information and the disease similarity information DS and lncRNA similarity information LS constructed according to the above method to construct a heterogeneous network G H .
[0078] The nodes of the heterogeneous graph are divided into two types, namely lncRNA node set and disease node set. The lncRNA node set is set as V lncRNA ={l1,l2,...,l n}, the disease node set is set to V disease ={d1,d2,...,d n There are three types of edges including lncRNA-disease edge, let its edge set be E IncRNA-disease lncRNA-lncRNA edge, let its edge set be E IncRNA-IncRNA . disease-disease edge, let its edge set be E disease-disease .
[0079] According to the above definitions of nodes and edges, the present invention constructs a heterogeneous graph HLD containing different types of nodes and edges:
[0080] G H =(V H ,E H ) (12)
[0081] The node set:
[0082] V H =V lncRNA ∪V disease (13)
[0083] Edge collection:
[0084] E H =E IncRNA-disease ∪E IncRNA-IncRNA ∪E disease-disease (14)
[0085] Finally, the heterogeneous graph G H =(V H ,E H )’s adjacency matrix HLD:
[0086]
[0087] Among them, LD is the lncRNA-disease association matrix, LS is the lncRNA-lncRNA similarity matrix, DS is the disease similarity matrix, LD T is the disease-lncRNA association matrix, and T is the transposition symbol.
[0088] Furthermore, generating different views includes: performing data enhancement processing on the heterogeneous graph to obtain positive sample views; and performing data destruction processing on the heterogeneous graph to obtain negative sample views.
[0089] Specifically, to enhance the model's ability to learn data, enable the model to learn different aspects of the data, and better extract embedded feature vectors, since graph contrastive learning requires positive and negative sample views, the present invention performs data augmentation and data destruction on the heterogeneous graphs to construct different views for graph contrastive learning.
[0090] The present invention uses data enhancement to generate positive sample views. The core idea of generating positive sample views is to transform or enhance the original graph to a certain extent, but without changing the core structure and information of the graph. The advantage of doing so is that it can help the model learn the feature representations of different aspects of the graph, enable the model to better understand the intrinsic structure and characteristics of the data, and can capture features well without relying on certain specific information, thereby enhancing the robustness of the model. The present invention uses dropout to randomly discard some edges and values to perform data enhancement. Formula 16 represents the process of data enhancement, where p is the probability of dropout, which determines the probability of each edge being discarded.
[0091] G H ′=Dropout(G H ,p) (16)
[0092] Unlike the generation of positive sample views, the purpose of negative sample views is to generate a graph structure that is semantically dissimilar to the original image or positive sample view by destroying the original image. The negative sample view generation method of the present invention is data destruction, such as completely randomly shuffling the nodes or edges of the graph. Specifically, this method will disrupt the order of the nodes in the graph or randomly reconnect the edges in the graph, making the structure of the graph completely different, but keeping the number of nodes and edges unchanged. The graph structure generated in this way will lose its original semantic relationship and structural pattern, ensuring that it is dissimilar to the original image or positive sample view, so the present invention uses the data-destroyed sample as the negative sample view:
[0093]
[0094] Furthermore, obtaining the fused GCN encoder includes: using the first training set to train the encoder of the original fused GCN to obtain the fused GCN encoder; during the training process, maximizing the similarity between the heterogeneous graph embedding and the positive sample view embedding while minimizing the similarity between the heterogeneous graph embedding and the negative sample view embedding as the objective function, training the encoder of the original fused GCN.
[0095] Specifically, this paper proposes a framework based on graph contrastive learning, in which the GCN encoder and contrastive learning process are key components. This framework uses a graph convolutional network encoder to extract node embedding information from the graph structure of different views. Using the contrastive learning framework to train the model, the model not only learns the inherent connections between similar samples, but also effectively distinguishes the differences between dissimilar samples.
[0096] First, the original graph (HLD) generates two variants through data augmentation and data corruption: an enhanced graph (HLD-A) and a corrupted graph (HLD-C). These variant graphs are processed by their respective GCN encoders. The design of the data augmentation encoder and the data corruption encoder ensures that the features of the graph are captured from different perspectives and enhances the robustness of the model to changes in the graph structure. The present invention uses an encoder fused with GCN to extract node embedding information from different views. The present invention passes the graph structure information to each node through the graph convolution layer and learns the embedded representation of each node. The graph convolution network formula of the lth layer is as follows:
[0097]
[0098] Among them, σ is the activation function, D is the degree matrix of the heterogeneous graph, A is the adjacency matrix, H (l) is the feature matrix of the lth layer, W (l) are the neural network parameters.
[0099] Since different views contain different semantic information, this paper uses three groups of GCN encoders to extract node embedding information of different views respectively, specifically:
[0100] The constructed heterogeneous graph structure and original features are input into the constructed GCN encoder. The original features are updated, and the encoder is trained by multi-image comparison. The encoder parameters are updated iteratively with loss. Each set of encoders uses two layers of GCN for feature extraction, and three different sets of embedding vectors are obtained:
[0101] Z=f GCNEncoder (HLD,X) (19)
[0102] Z A =f AugmentEncoder (HLD-A,X A ) (20)
[0103] Z C =f CorruptEncoder (HLD-A,X C ) (twenty one)
[0104] After extracting the embedding vectors, the contrastive learning method aims to maximize the similarity between the original graph embedding and the enhanced graph embedding, while minimizing the similarity with the corrupted graph embedding. This approach strengthens the model's ability to capture useful information and effectively distinguish subtle differences in graph structure, thereby improving the model's ability to distinguish dissimilar samples.
[0105] L contrastive =(1-cos(Z,Z A )+max(0,cos(Z,Z C)-m)) (22)
[0106] Through this structured training approach, the model not only learns the intrinsic connections between similarities between nodes in a graph, but also effectively discerns and resists the influence of erroneous or biased information when faced with structural perturbations. This contrastive learning framework demonstrates superior performance in various graph data analysis tasks, particularly those requiring precise node and graph classification.
[0107] Furthermore, obtaining the association information between the disease and the non-coding RNA includes: reconstructing the embedded feature vector to obtain the reconstructed features, inputting the reconstructed features into the XGBoost model, and obtaining the association information between the disease and the non-coding RNA; the XGBoost model is trained using the second training set, and the second training set includes: the original reconstructed features.
[0108] Specifically, the present invention uses the aforementioned graph contrastive learning self-supervised learning framework to extract embedded feature vectors for lncRNAs and diseases. This method not only captures complex network structures but also explores deep underlying connections between different nodes. These feature vectors capture high-level information about each node.
[0109] By reconstructing features based on the node numbers of lncRNAs and diseases, the present invention achieves a more refined and efficient feature representation. These reconstructed features are then used as input to the XGBoost model to predict the association between lncRNAs and diseases. This process improves prediction accuracy and ensures that the model extracts more potential biological information from the reconstructed features.
[0110] This paper used four different datasets to verify the robustness and generalization of the model, achieving AUC values of 0.9816, 0.9940, 0.9412, and 0.9789, respectively. These results significantly outperform other advanced models, demonstrating the excellent predictive capabilities of HGCLDA. The paper conducted extensive experiments on various datasets, evaluating the model using a variety of metrics. The experimental results demonstrate that HGCLDA achieves superior results compared to other advanced models.
[0111] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. An optimization method for predicting the association between non-coding RNA and disease, characterized in that: include: Obtaining a similarity matrix of diseases and non-coding RNAs, and constructing a heterogeneous graph based on the similarity matrix; Processing the heterogeneous graph to generate different views, inputting the different views into the encoder of the GCN fusion to obtain embedded feature vectors; The GCN-fused encoder is trained using the first training set; The first training set includes: original views; Based on the embedded feature vector, association information between the disease and the non-coding RNA is obtained.
2. The optimization method for predicting the association between non-coding RNA and disease according to claim 1, characterized in that Obtaining the similarity matrix of the disease includes: Among them, DS(i, j) is the similarity matrix of the disease, DSS(i, j) is the semantic similarity of different diseases, and DGS(i, j) is the Gaussian kernel similarity of the disease.
3. The optimization method for predicting the association between non-coding RNA and disease according to claim 1, characterized in that Obtaining the similarity matrix of the non-coding RNA includes: Among them, LS(i, j) is the similarity matrix of non-coding RNA, LFS(i, j) is the functional similarity of non-coding RNA, and LGS(i, j) is the Gaussian kernel similarity of non-coding RNA.
4. The optimization method for predicting the association between non-coding RNA and disease according to claim 1, characterized in that Constructing the heterogeneous graph includes: G H =(V H ,E H ) V H =V lncRNA ∪V disease AND H =And IncRNA-disease ∪E IncRNA-IncRNA ∪E disease-disease Among them, G H is a heterogeneous graph, V H is a node set, E H is the edge set, V lncRNA is the non-coding RNA node set, V disease is the disease node set, E IncRNA-disease is the edge set of non-coding RNA nodes and diseases, E IncRNA-IncRNA is the edge set of different non-coding RNA nodes, E disease-disease is the edge set of different diseases.
5. The optimization method for predicting the association between non-coding RNA and disease according to claim 1, characterized in that Generating the different views includes: Perform data augmentation on the heterogeneous graph to obtain a positive sample view: G H ′=Dropout(G H ,p) Among them, G H ' is the positive sample view, p is the dropout probability; Perform data destruction on the heterogeneous graph to obtain a negative sample view: in, is the negative sample view, G H It is a heterogeneous graph.
6. The optimization method for predicting the association between non-coding RNA and disease according to claim 1, characterized in that Obtaining the embedded feature vector includes: Z=f GCNEncoder (HLD,X) From A =f AugmentEncoder (HLD-A,X A ) Z C =f CorruptEncoder (HLD-C,X C ) Among them, HLD is a heterogeneous graph, HLD-A is a positive sample view, HLD-C is a negative sample view, Z is the node embedding vector of the fused heterogeneous graph, and f GCNEncoder The extracted node embedding vector that integrates the original heterogeneous graph information, Z A f AugmentEncoder The node embedding vector of the extracted fused enhanced graph information, Z C f CorruptEncoder The node embedding vector that integrates the destruction graph information, f GCNEncoder For the GCN encoder that processes the original heterogeneous graph, f AugmentEncoder To process the GCN encoder for enhanced heterogeneous graphs, f CorruptEncoder To process the GCN encoder that destroys the heterogeneous graph, X is the original heterogeneous graph, X A To enhance heterogeneous graphs, X C To destroy heterogeneous graphs.
7. The optimization method for predicting the association between non-coding RNA and disease according to claim 1, characterized in that Obtaining the encoder of the fused GCN includes: Using the first training set to train the encoder of the original fused GCN to obtain the encoder of the fused GCN; During the training process, the encoder of the original fusion GCN is trained with the objective function of maximizing the similarity between the heterogeneous graph embedding and the positive sample view embedding while minimizing the similarity between the heterogeneous graph embedding and the negative sample view embedding.
8. The optimization method for predicting the association between non-coding RNA and disease according to claim 7, characterized in that: The objective function includes: L contrastive =(1-cos(Z,Z A )+max(0,cos(Z,Z C )-m)) Among them, L contrastive is the contrastive learning loss function, cos(Z,Z A ) is the similarity between the original node embedding vector and the enhanced node embedding vector, cos(Z,Z C ) is the similarity between the original node embedding vector and the corrupted node embedding vector, and m is the minimum similarity difference used to control the positive and negative samples.
9. The optimization method for predicting the association between non-coding RNA and disease according to claim 1, characterized in that: Obtaining the association information between the disease and the non-coding RNA includes: The embedded feature vector is reconstructed to obtain reconstructed features, and the reconstructed features are input into an XGBoost model to obtain association information between the disease and the non-coding RNA; the XGBoost model is trained using a second training set, and the second training set includes: original reconstructed features.