Method for predicting correlation between lncRNA and disease

By constructing heterogeneous networks and introducing graph attention networks and transformer technology, the accuracy and reliability problems of traditional methods in predicting the association between lncRNA and diseases are solved, and more accurate disease diagnosis and treatment basis is achieved.

CN120199321APending Publication Date: 2025-06-24SOUTHWEST MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510208037.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When traditional methods predict the association of lncRNA with diseases, they ignore complex interdependence and interaction relationships, and it is difficult to integrate multi-dimensional and heterogeneous data, resulting in limited accuracy and reliability of prediction results.

Method used

By constructing heterogeneous networks, integrating multi-dimensional information of lncRNA and disease, deep learning technologies such as graph attention network (GAT) and transformer (Transformer) are used to deeply process heterogeneous networks to comprehensively capture the complex relationship between lncRNA and disease.

Benefits of technology

It significantly improves the accuracy and reliability of lncRNA and disease association prediction, providing a more accurate basis for the diagnosis and treatment of diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199321A_ABST
    Figure CN120199321A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting relevance between lncRNA and diseases, and relates to the technical field of biological information, and the method comprises the following specific steps: collecting multi-dimensional lncRNA related data from a biological database and literature, completing preprocessing through cleaning, missing value processing and standardization, similarity matrix calculation and the like, constructing a heterogeneous network and optimizing, and obtaining the relevance between the lncRNA and the diseases. Feature representation is generated through GAT and transformer processing, MLP classification prediction is carried out, and the optimal effect is achieved through parameter adjustment; according to the method, the advantages of a heterogeneous network, a graph attention network (GAT) and a transformer are organically combined, the complex relation between long-chain non-coding RNA (lncRNA) and diseases can be comprehensively and deeply captured, the local neighborhood relation between the lncRNA and the diseases is considered, the global long-range dependency relation is captured, and the prediction accuracy and reliability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics technology, and specifically to a method for predicting the association between lncRNA and diseases. Background Art

[0002] Long non-coding RNA (lncRNA), as an important class of non-coding RNA molecules, plays a key regulatory role in various biological processes of organisms. In recent years, with the rapid development of high-throughput sequencing technology and bioinformatics, more and more studies have revealed the complex associations between lncRNA and various diseases. These associations not only involve the abnormal expression of lncRNA and the occurrence and development of diseases, but also delve into how lncRNA affects the disease process through mechanisms such as regulating gene expression and participating in signal transduction pathways. Therefore, accurately predicting the association between lncRNA and diseases is of great significance for the early diagnosis of diseases, the formulation of personalized treatment plans, and the in-depth understanding of the disease occurrence mechanism.

[0003] Although significant progress has been made in the research on predicting the association between lncRNA and diseases, traditional prediction methods still have many limitations. On the one hand, some traditional methods rely too much on simple similarity metrics, such as methods based on sequence similarity or expression profile similarity. These methods ignore the complex interdependent and interacting relationships between lncRNA and diseases, resulting in limited accuracy and reliability of the prediction results. On the other hand, some single-model-based methods often cannot fully exploit the potential information in the data when dealing with large-scale and high-dimensional data, restricting the improvement of prediction performance. In addition, traditional methods lack flexibility in dealing with heterogeneous data and are difficult to integrate lncRNA and disease data from different sources and different types, further limiting their generality and accuracy in practical applications.

[0004] In view of the above problems, it is necessary to optimize the existing methods for predicting the association between lncRNA and diseases. By constructing a heterogeneous network, integrating multi-dimensional information of lncRNA and diseases to form a complex network structure, providing rich topological information for subsequent analysis, and introducing graph attention network and transformer technology to deeply process the heterogeneous network and comprehensively capture the complex relationships between lncRNA and diseases. Therefore, it is of great significance to develop a method for predicting the association between lncRNA and diseases that can comprehensively achieve the above characteristics. Summary of the Invention

[0005] The object of the present invention is to make up for the deficiencies of the prior art and provide a method for predicting the association between lncRNA and diseases. It can construct a heterogeneous network, integrate various characteristic information of lncRNA and diseases, form a complex network structure, provide rich topological information for subsequent analysis, and at the same time introduce advanced deep learning technologies such as graph attention network (GAT) and Transformer to extract and transform the node features in the heterogeneous network, comprehensively capture the complex relationship between lncRNA and diseases. The graph attention network can assign attention weights to neighbor nodes according to the node features and network structure to extract richer local feature information, while the Transformer uses the self-attention mechanism to perform global analysis on the feature sequence to capture the long-range dependence relationship and interaction between features. Finally, a multi-layer perceptron (MLP) classifier is used for classification prediction, realizing the accurate prediction of the association between lncRNA and diseases, significantly improving the accuracy and reliability of the prediction, and providing a more accurate basis for the diagnosis and treatment of diseases.

[0006] To solve the above technical problems, the present invention provides the following technical solution: A method for predicting the association between lncRNA and diseases, the method comprising the following specific steps:

[0007] Data collection and preprocessing: Collect multi-dimensional data covering lncRNA sequences, expression profiles, and disease clinical and pathological information from biological databases and scientific literature, remove noise through cleaning operations, handle missing values according to the data characteristics, and unify the data range through standardization means. Use algorithms to calculate the lncRNA-disease similarity matrix, and at the same time sort out the known lncRNA-disease association information to lay a foundation for subsequent network construction and analysis;

[0008] Heterogeneous network construction: Using lncRNA and diseases as nodes, establish edges between the corresponding nodes based on the known lncRNA-disease association information to construct a heterogeneous network. Fully consider the different attributes of the nodes and the weights of the edges, and assign different weights to the nodes according to the functional importance of lncRNA, its participation in key biological processes, as well as the severity and incidence factors of diseases. And adopt graph theory algorithms to remove redundant edges and nodes, and at the same time optimize the network structure to improve the information transmission efficiency;

[0009] Graph Attention Network Processing: The Graph Attention Network (GAT) is used to process the heterogeneous network. Based on the node's own features and network structure information, the coefficients of each adjacent node are determined, and the multi-head attention mechanism is introduced to perceive the network structure from multiple perspectives. The attention coefficients are combined with the corresponding adjacent node features to determine the new embedding representation of the target node in the multi-head attention mechanism. During the training process, according to the actual network scale and data complexity, through experiments and tuning, the number of GAT layers and the number of neurons are determined, and the validation set is used to evaluate the performance of GAT in real time, so as to generate a highly representative node feature representation;

[0010] Transformer Processing: The features output by GAT are input into the transformer and associated with the query Q, key K, and value V vectors to calculate its attention output. The multi-head attention mechanism is used, and multiple attention heads work in parallel to analyze the feature relationships from different subspaces, calculate the attention scores, and concatenate the outputs of multiple heads in sequence to generate a comprehensive feature representation. At the same time, for different tasks and data features, the number of transformer heads and the hidden layer dimension parameters are adjusted, and the learning rate decay strategy is adopted to ensure stable training, so as to strengthen the understanding and characterization of the complex relationship between lncRNA and diseases;

[0011] Classification Prediction: Using the feature representation processed by the transformer, classification prediction is performed through the MLP classifier. The MLP judges whether there is an association between lncRNA and diseases according to different feature combinations. When training the MLP, the cross-entropy loss function is used to calculate the difference between the prediction result and the true result. According to the difference, the parameters of the MLP are continuously adjusted, and the Adam optimizer is used to optimize the learning rate parameter. When the performance of the model on the validation set no longer improves, the training is stopped, the model is verified and evaluated through the test set, and the hyperparameters are fine-tuned accordingly to achieve the best prediction effect.

[0012] Furthermore, in the data collection and preprocessing step, an algorithm is used to calculate the lncRNA-disease similarity matrix, and its algorithm formula is: where FS represents the functional similarity of two lncRNAs, DS refers to the disease semantic similarity of two given diseases, d(ix) and d(jy) are the disease elements in the disease sets D i and D j associated with lncRNA i and j respectively, and m and n refer to the number of diseases in the D i and D j groups respectively.

[0013] Even further, in the heterogeneous network construction step, lncRNAs and diseases are used as nodes, and based on the known lncRNA-disease association information, edges are established between the corresponding nodes to construct a heterogeneous network, and its heterogeneous network matrix is: Among them, G(V, E) represents the constructed heterogeneous network, where V represents the set of nodes in the network, E represents the set of edges, LS is the lncRNA functional similarity matrix, recording the functional similarity degree between different lncRNAs, DS is the disease similarity matrix, reflecting the similarity degree between different diseases, and A is the adjacency matrix of known lncRNA-disease associations. When there is an association between lncRNAi and disease j, A ij = 1, otherwise it is 0, and A T is the transpose of the adjacency matrix A.

[0014] Furthermore, in the graph attention network processing step, GAT is used to process the heterogeneous network. According to the node's own characteristics and network structure information, the attention coefficient of each adjacent node is determined. Specifically, the attention score of the adjacent node is calculated, and its formula is: e ij = LeakyReLU(a T [Wh i ||Wh j ), where e ij represents the attention score from adjacent node j to node i, a T represents the transpose of the learnable weight vector, W is the weight matrix, h i and h j are the feature embeddings of node i and adjacent node j respectively, ∥ represents the concatenation operation of feature vectors, and LeakyReLU is the activation function. The attention coefficient is obtained by normalizing the attention score, and its normalization formula is: where α ij is the normalized attention coefficient, N is the set of adjacent nodes of node j, and σ is the non-linear activation function.

[0015] Furthermore, in the graph attention network processing step, a multi-head attention mechanism is introduced to perceive the network structure from multiple perspectives. The attention coefficient is combined with the corresponding adjacent node features to determine the new embedding representation of the target node in the multi-head attention mechanism. Its new embedding representation formula is: where h i i is the new embedding of node i under the multi-head attention mechanism, M is the number of attention heads, is the attention coefficient from adjacent node j to node i in the m-th attention head, W m is the weight matrix of the m-th attention head, h j is the feature embedding of adjacent node j, N i is the set of adjacent nodes of node i, and σ is the non-linear activation function.

[0016] Furthermore, in the transformer processing step, the features output by GAT are input into the transformer and associated with the query Q, key K, and value V vectors to calculate its attention output. The attention calculation formula is as follows: where Attention(Q, K, V) is the calculated attention output, Q is the query vector, K is the key vector, which is matched with the query vector Q to calculate the attention score, and its dimension is d k , V is the value vector, and softmax is the normalization function.

[0017] Furthermore, in the transformer processing step, the multi-head attention mechanism is used. Multiple attention heads work in parallel to analyze the feature relationships from different subspaces, calculate the attention scores, and connect the outputs of multiple heads in sequence. Specifically, through the formula the output of each attention head is calculated, and through the formula MultiHead(QKV) = Concat(head1,..., head h )W O the outputs of multiple heads are connected in sequence, where MultiHead(QKV) is the final output of the multi-head attention mechanism, head i is the output of the i-th attention head, are the weight matrices for linearly transforming Q, K, and V respectively when the i-th attention head calculates the attention score, and W O is the weight matrix for linearly transforming the outputs of multiple attention heads in the multi-head attention mechanism, and Concat is the concatenation operation.

[0018] Furthermore, in the classification and prediction step, when training the MLP, the cross-entropy loss function is used to calculate the difference between the prediction result and the true result. The function formula is as follows: where Loss is the value of the loss function, N is the total number of samples, y i is the actual label of sample i, is the predicted probability of the model for sample i.

[0019] Compared with the prior art, the method for predicting the association between lncRNA and diseases has the following beneficial effects:

[0020] 1. By organically combining the advantages of heterogeneous networks, graph attention networks (GAT), and transformers, the present invention can comprehensively and deeply capture the complex relationships between long non-coding RNAs (lncRNAs) and diseases. It not only considers the local neighborhood relationships between lncRNAs and diseases but also captures the global long-range dependencies, significantly improving the accuracy and reliability of predictions. This is of crucial significance for the early diagnosis of diseases, the formulation of treatment plans, and the in-depth understanding of disease mechanisms, and can provide more accurate information and decision-making basis for biomedical research and clinical practice.

[0021] 2. By adjusting the model parameters and architecture, the present invention can process lncRNA and disease data from different sources and types. With the continuous emergence of new bioinformatics data, this method can be continuously optimized and updated to adapt to new data characteristics and research needs, thus being applicable to a wider range of biomedical research and clinical application scenarios.

[0022] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0024] Figure 1 It is a flowchart operation diagram of a method for predicting the association between lncRNA and disease;

[0025] Figure 2 It is a flowchart of a method for predicting the association between lncRNA and disease. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention objective, the following, in combination with the accompanying drawings and preferred embodiments, details the specific embodiments, structures, features, and their effects of the present invention as follows.

[0027] Example 1

[0028] Collect data from authoritative databases in cancer research, including but not limited to gene expression data, clinical data, methylation data, etc. Specifically, collect lncRNA expression profile data of various cancer types (such as breast cancer, lung cancer, colorectal cancer, gastric cancer, liver cancer, etc.), as well as the corresponding patient clinical information. The clinical information includes the patient's age, gender, cancer stage (stage I, stage II, stage III, stage IV), cancer tissue type (such as adenocarcinoma, squamous cell carcinoma, etc.), tumor size, whether there is metastasis, survival time, treatment effect, etc. At the same time, collect cancer-related lncRNA function information and disease characteristic information from relevant literature and other bioinformatics databases. Use data cleaning rules to remove incorrect data, such as outliers in the gene expression profile (which may be caused by experimental errors or data recording errors), and incomplete records in the clinical data (such as patient information lacking key clinical indicators). For missing values in the gene expression data, according to the expression distribution of lncRNA, use mean filling or similarity-based filling methods, and perform standardization processing on it. According to the collected data, use algorithms to calculate the functional similarity matrix between lncRNAs and the similarity matrix between cancer diseases.

[0029] For example, for two lncRNAs l i and l j , which are respectively associated with disease sets D i and D j , calculate the functional similarity according to the formula

[0030] Take different lncRNAs and cancer types as nodes. For each cancer patient sample, if it is known that a certain lncRNA is abnormally expressed in this sample and this sample has a certain cancer, establish an edge between the corresponding lncRNA node and cancer node, thereby constructing a heterogeneous network G(V, E), and its heterogeneous network matrix is: Among them, V represents the set of nodes in the network, E represents the set of edges, LS is the functional similarity matrix of lncRNAs, recording the functional similarity degree between different lncRNAs, DS is the disease similarity matrix, reflecting the similarity degree between different diseases, A is the adjacency matrix of known lncRNA-disease associations. When there is an association between lncRNAi and disease j, A ij = 1, otherwise it is 0, A T ​is the transpose of the adjacency matrix A. According to the importance and functional characteristics of lncRNAs in cancer, weights are assigned to the nodes. For example, for lncRNAs that play a key role in cancer occurrence and development (such as participating in key signaling pathways or acting as regulators of oncogenes), higher node weights are assigned. For cancer types with a higher tendency to metastasize, higher disease node weights are assigned. For the weights of the edges, they are set according to the strength of the associated experimental evidence and the consistency of clinical observations. For example, for lncRNA-cancer associations found in multiple independent studies, the weights of their edges are higher. Graph theory algorithms are used to optimize the network. For example, the minimum spanning tree algorithm is adopted to remove redundant edges, making the network structure more concise. The connected component algorithm is used to ensure the connectivity of the network and avoid the appearance of isolated nodes. For example, in the network, there may be some lncRNA nodes that are less connected to other nodes due to experimental errors or data biases. The connected component algorithm can connect them to the main part of the network to better reflect the true biological relationships.

[0031] Input the heterogeneous network into the Graph Attention Network (GAT). For each cancer-related node i and its adjacent node j, calculate the attention score e from the adjacent node j to the node i ij , the formula is i j = LeakyReLU(a T [Wh i ||Wh j ), where a T represents the transpose of the learnable weight vector, W is the weight matrix, h i and h j are the feature embeddings of nodes i and j, and ∥ represents the concatenation operation of the feature vectors. Taking breast cancer as an example, for lncRNA nodes related to breast cancer, calculate the attention scores between them according to their own characteristics (such as expression levels, functional characteristics) and the characteristics of neighbor nodes (such as other related lncRNAs and breast cancer nodes at different stages). Use the SoftMax function to normalize the attention scores to obtain the normalized attention coefficient α ij , the formula is where N i is the set of adjacent nodes of node i, and the sum of the attention coefficients of its adjacent nodes is 1, so it can be used as a weight. Adopt the multi-head attention mechanism (assuming M = 4 heads) to calculate the new node embedding under multi-head attention In the scenarios of different cancers, different attention heads can focus on different aspects. For example, one head can focus on the associations of lncRNAs in the early stage of cancer, and another head can focus on the associations during the cancer metastasis process, etc. During the training process, according to the actual network scale and data complexity, through experiments and tuning, determine the number of GAT layers and neurons, and use the validation set to evaluate the performance of GAT in real time, so as to generate a highly representative node feature representation.

[0032] Use the node embeddings processed by GAT as the input of the Transformer. First, normalize these node embeddings to ensure the consistency and stability of the input data, and use the formula to calculate the attention scores of the query Q, key K, and value V vectors. Among them, is the dimension of K. For cancer data, Q, K, and V can represent the feature vectors of different cancer stages or different lncRNAs. For multi-head attention calculation, first calculate the output of each attention head through , and then connect the outputs of multiple attention heads in sequence, that is, MultiHead(QKV) = Concat(head1,…,head h )W O . For example, when analyzing different cancer types and stages, the multi-head attention mechanism can capture global information from multiple perspectives. One head may focus on the similarities of lncRNA expression in different cancers, and another head may focus on the common features and differences of different cancer types. At the same time, for different tasks and data characteristics, adjust the number of Transformer heads and the hidden layer dimension parameters, and adopt a learning rate decay strategy to ensure stable training, so as to strengthen the understanding and representation of the complex relationship between lncRNAs and diseases.

[0033] Use the feature representation processed by the Transformer for classification prediction through the MLP classifier. The MLP judges whether there is an association between lncRNAs and diseases according to different feature combinations. When training the MLP, use the cross-entropy loss function to calculate the difference between the prediction result and the true result, where y i is the actual cancer label of sample i (such as whether having a certain cancer, represented by 0 or 1), is the predicted probability, N is the total number of samples. Using the Adam optimizer, it is trained through multiple training cycles (such as 2000 cycles) and learning rate (such as 0.001). The model parameters are adjusted according to the performance on the validation set. Metrics such as the area under the receiver operating characteristic curve (AUC) and the area under the precision-recall curve (AUPRC) are used to evaluate the performance of the model on cancer data. By comparing the prediction results of different cancer types, the prediction accuracy and reliability of the model for different cancers are evaluated. For example, for breast cancer, the AUC value predicted by the model is calculated. If the AUC value is high, it indicates that the model has a strong prediction ability for lncRNAs related to breast cancer.

[0034] Example Two

[0035] Data is widely collected from professional cardiovascular disease databases, such as the UK Biobank and the database of the American Heart Association, as well as numerous clinical cardiovascular disease research projects. This data contains rich information, such as information on various cardiovascular diseases like coronary heart disease, myocardial infarction, heart failure, arrhythmia, etc. For patient information, it comprehensively covers basic information including age, gender, family medical history, clinical indicators such as blood pressure, blood lipids, blood sugar, heart rate, electrocardiogram indicators, etc., and imaging examination results like coronary angiography results. At the same time, the expression profile data of lncRNAs related to the cardiovascular system is also collected. In addition, from various bioinformatics databases and relevant academic literatures, information on the functions of lncRNAs in the cardiovascular system and the pathogenesis of cardiovascular diseases and other aspects of information are obtained to form a comprehensive data resource. The collected data is carefully checked, and incomplete and inaccurate data is removed. For example, for blood pressure values that are significantly outside the normal range or seriously inconsistent with other patient indicators, which may be measurement errors or recording errors, they will be removed from the dataset. For duplicate patient information, according to the time of information update and the reliability of the data source, the latest or most reliable data is selected for retention. For handling missing values, for clinical indicators, they are reasonably filled based on the clinical knowledge of cardiovascular diseases and data distribution and standardized. When calculating the similarity matrix, according to the characteristics of cardiovascular diseases and the functional information of lncRNAs, algorithms are used to calculate the functional similarity matrix between lncRNAs and the similarity matrix between cardiovascular diseases. For example, for lncRNAs, considering their functions in aspects such as angiogenesis, regulation of cardiomyocyte function, proliferation and differentiation of vascular smooth muscle cells, etc., the functional similarity matrix is calculated. For cardiovascular diseases, according to factors such as the etiology of the disease (such as the severity of atherosclerosis, the extent of myocardial infarction, the type of arrhythmia), symptoms (such as the frequency and severity of chest pain, the severity of dyspnea), etc., the similarity matrix between cardiovascular diseases is calculated.

[0036] When constructing a heterogeneous network, cardiovascular diseases and related lncRNAs are used as nodes, and the edges of the network are constructed based on existing research results and clinical evidence. For example, it is known that certain lncRNAs play key roles in the occurrence and development of coronary heart disease, such as participating in processes like vascular endothelial function regulation and inflammatory response. In the network, the associated edges between them and coronary heart disease will be assigned higher weights. For some rare cardiovascular diseases, corresponding associated edges will also be established based on limited research data, but the weights may be relatively low. As more research data accumulates, the weights of these edges will be adjusted accordingly.

[0037] When using the Graph Attention Network (GAT) for processing, for each node i related to cardiovascular diseases and its adjacent node j, calculate the attention score e from adjacent node j to node i ij , and its calculation formula is e ij = LeakyReLU(a T [Wh i ||Wh j ), where a T represents the transpose of the learnable weight vector, W is the weight matrix, h i and h j are the feature embeddings of nodes i and j respectively, and ∥ represents the concatenation operation of feature vectors. Taking heart failure as an example, for nodes related to heart failure, calculate the attention scores between them according to their own characteristics (such as the expression characteristics of lncRNAs related to heart failure and the clinical index characteristics of heart failure) and the characteristics of adjacent nodes (such as nodes of other cardiovascular diseases concurrent with heart failure and related lncRNA nodes). This score reflects the importance and influence degree of adjacent nodes on this node. Then, use the SoftMax function to normalize the attention scores to obtain the normalized attention coefficient α ij , and the formula is where N i is the set of adjacent nodes of node i. This can ensure that for each node, the sum of the attention coefficients of its adjacent nodes is 1, enabling these coefficients to be used as weights for subsequent processing. Adopt the multi-head attention mechanism (assuming 4 heads), and calculate the new node embedding through the formula , where and W m are the attention coefficient and weight matrix of the m-th attention head respectively. Different attention heads can focus on local relationships in cardiovascular diseases from different perspectives. For example, one attention head can focus on the association between lncRNAs and cardiovascular diseases in the early stage of the disease, and another head can focus on the association of lncRNAs under different complication conditions during the disease progression, thereby mining local information from multiple perspectives.

[0038] In the processing stage of the Transformer, by using its multi-head attention mechanism, the complex relationship between cardiovascular diseases and lncRNAs is comprehensively analyzed from a global perspective. For some lncRNAs related to multiple cardiovascular diseases, they may have different global association patterns in different diseases. For example, a certain lncRNA may be mainly related to the inflammatory response of blood vessel walls in coronary heart disease and related to the energy metabolism of cardiomyocytes in heart failure. The Transformer can capture these complex relationships from a global perspective. Different attention heads can focus on different information dimensions respectively. Some attention heads can focus on the expression patterns of lncRNAs in different tissues in different cardiovascular diseases, and some can focus on their overall impact on signal pathways in the cardiovascular system. Through the comprehensive processing of different attention heads, information from different dimensions is integrated, presenting us with more comprehensive association information.

[0039] Finally, prediction is carried out through the MLP classifier, providing clinicians and researchers with information in multiple aspects. For a newly discovered lncRNA, through this model, it can be predicted whether it may become a new diagnostic marker for a certain cardiovascular disease. For example, if the prediction result shows that a certain lncRNA presents a unique expression pattern in coronary heart disease patients and is closely related to the occurrence and development of the disease, then it may be used as a potential diagnostic marker to assist clinical diagnosis. At the same time, this model can also provide potential target information for the development of therapeutic drugs for cardiovascular diseases. For example, if it is found that a certain lncRNA is related to the abnormal proliferation of cardiomyocytes, then it may become a target for developing drugs to inhibit myocardial hypertrophy. In addition, this model can also help researchers better understand the pathogenesis of cardiovascular diseases, such as discovering the potential roles of certain lncRNAs in the genetic susceptibility and environmental factor induction mechanisms of cardiovascular diseases, and providing a theoretical basis for the formulation of personalized treatment plans. For example, for patients with specific lncRNA expression characteristics, more targeted treatment plans are formulated, including drug selection, adjustment of treatment time and dosage, etc.

[0040] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed as above with a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to it as equivalent embodiments within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A method for predicting the association between lncRNA and disease, characterized in that: The method comprises the following specific steps: Data collection and preprocessing: Collect multi-dimensional data covering lncRNA sequences, expression profiles, and clinical and pathological aspects of diseases from biological databases and scientific literature, remove noise through cleaning operations, process missing values ​​according to data characteristics, unify data ranges through standardization, use algorithms to calculate lncRNA and disease similarity matrices, and organize known lncRNA-disease association information to lay the foundation for subsequent network construction and analysis; Heterogeneous network construction: lncRNA and disease are used as nodes. Based on the known lncRNA-disease association information, edges are established between corresponding nodes to construct heterogeneous networks. Different attributes of nodes and weights of edges are fully considered. Different weights are assigned to nodes based on the functional importance of lncRNA, the degree of involvement in key biological processes, and the severity and incidence of the disease. Graph theory algorithms are used to remove redundant edges and nodes, while optimizing the structure of the network and improving the efficiency of information transmission. Graph attention network processing: GAT is used to process heterogeneous networks. The coefficient of each adjacent node is determined based on the node's own characteristics and network structure information. A multi-head attention mechanism is introduced to perceive the network structure from a multi-dimensional perspective. The attention coefficient is combined with the corresponding adjacent node characteristics to determine the new embedding representation of the target node in the multi-head attention mechanism. During the training process, the number of GAT layers and neurons is determined through experiments and tuning according to the actual network scale and data complexity, and the performance of GAT is evaluated in real time using a validation set to generate a highly representative node feature representation. Transformer processing: The features output by GAT are input into the transformer and associated with the query Q, key K and value V vectors, and its attention output is calculated. The multi-head attention mechanism is used, and multiple attention heads work in parallel to analyze feature relationships from different subspaces, calculate attention scores and connect the outputs of multiple heads in sequence to generate a comprehensive feature representation. At the same time, the number of transformer heads and hidden layer dimension parameters are adjusted for different tasks and data characteristics, and the learning rate decay strategy is used to ensure training stability, so as to enhance the understanding and characterization of the complex relationship between lncRNA and disease; Classification prediction: The feature representation processed by the transformer is used to perform classification prediction through the MLP classifier. MLP judges whether there is an association between lncRNA and disease based on different feature combinations. When training MLP, the cross entropy loss function is used to calculate the difference between the predicted result and the actual result. The parameters of MLP are continuously adjusted according to the difference, and the Adam optimizer is used to optimize the learning rate parameters. When the performance of the model on the validation set is no longer improved, the training is stopped, and the model is verified and evaluated through the test set, and the hyperparameters are fine-tuned accordingly to achieve the best prediction effect.

2. A method for predicting the association between lncRNA and disease according to claim 1, characterized in that: In the data collection and preprocessing steps, an algorithm is used to calculate the lncRNA and disease similarity matrix, and the algorithm formula is: Among them, FS represents the functional similarity of two lncRNAs, DS refers to the disease semantic similarity of two given diseases, and d(ix) and d(jy) are the disease sets D associated with lncRNAi and j, respectively. i , D j The disease elements in the equation, m and n refer to D i and D j The number of diseases in the group.

3. The method for predicting the association between lncRNA and disease according to claim 1, characterized in that: In the heterogeneous network construction step, lncRNA and disease are used as nodes, and edges are established between corresponding nodes based on known lncRNA-disease association information to construct a heterogeneous network. The heterogeneous network matrix is: Among them, G(V,E) represents the constructed heterogeneous network, V represents the set of nodes in the network, E represents the set of edges, LS is the lncRNA functional similarity matrix, which records the functional similarity between different lncRNAs, DS is the disease similarity matrix, which reflects the similarity between different diseases, and A is the adjacency matrix of known lncRNA-disease associations. When there is an association between lncRNAi and disease j, A ij =1, otherwise 0, A T is the transpose of the adjacency matrix A.

4. The method for predicting the association between lncRNA and disease according to claim 1, characterized in that: In the graph attention network processing step, GAT is used to process the heterogeneous network, and the attention coefficient of each adjacent node is determined according to the node's own characteristics and network structure information. Specifically, the attention score of the adjacent node is calculated, and the formula is: ij =LeakyReLU(a T [Wh i ||Wh j ]), where e ij represents the attention score of the adjacent node j to node i, a T represents the transpose of the learnable weight vector, W is the weight matrix, and h i and h j are the feature embeddings of node i and adjacent node j respectively, ∥ represents the connection operation of feature vectors, LeakyReLU is the activation function, and the attention coefficient is obtained by normalizing the attention score. The normalization formula is: Among them, α ij is the normalized attention coefficient, N i is the set of adjacent nodes of node i, and softmax is the normalization function.

5. The method for predicting the association between lncRNA and disease according to claim 1, characterized in that: In the graph attention network processing step, a multi-head attention mechanism is introduced to perceive the network structure from multiple perspectives, and the attention coefficient is combined with the corresponding adjacent node features to determine the new embedding representation of the target node in the multi-head attention mechanism. The new embedding representation formula is: Among them, h i ′ is the new embedding of node i under the multi-head attention mechanism, M is the number of attention heads, is the attention coefficient from adjacent node j to node i in the mth attention head, W m is the weight matrix of the mth attention head, h j is the feature embedding of neighboring node j, N is the set of neighboring nodes of node j, and σ is a nonlinear activation function.

6. The method for predicting the association between lncRNA and disease according to claim 1, characterized in that: In the transformer processing step, the features output by GAT are input into the transformer and associated with the query Q, key K and value V vectors to calculate its attention output. The attention calculation formula is: Among them, Attention(Q,K,V) is the calculated attention output, Q is the query vector, K is the key vector, and the attention score is calculated by matching the query vector Q, and its dimension is d k , V is the value vector, and softmax is the normalization function.

7. The method for predicting the association between lncRNA and disease according to claim 1, characterized in that: In the transformer processing step, a multi-head attention mechanism is used. Multiple attention heads work in parallel to analyze feature relationships from different subspaces, calculate attention scores, and connect the outputs of multiple heads in sequence. Specifically, through the formula head i =Attention(QW i Q ,KW i K ,VW i v ) calculates the output of each attention head and uses the formula MultiHead(QKV)=Concat(head1,…,head h )W O Connect the outputs of multiple heads in sequence, where MultiHead(QKV) is the final output of the multi-head attention mechanism, head i is the output of the ith attention head, W i Q , W i K , W i v are the weight matrices of the linear transformation of Q, K, and V by the i-th attention head when calculating the attention score, and W O It is a weight matrix used to linearly transform the outputs of multiple attention heads in the multi-head attention mechanism, and Concat is a concatenation operation.

8. The method for predicting the association between lncRNA and disease according to claim 1, characterized in that: In the classification prediction step, when training the MLP, the cross entropy loss function is used to calculate the difference between the predicted result and the actual result, and its function formula is: Among them, Los is the loss function value, N is the total number of samples, and y i is the actual label of sample i, is the model's predicted probability for sample i.