A prediction method for lncRNA-disease association based on the KAN network

Through the KAN network combined with Transformer and GraphSAGE, the global characteristics and topological information of lncRNA and disease were extracted, which solved the problem of insufficient feature extraction in the existing methods and achieved more accurate lncRNA-disease association prediction.

CN119905141BActive Publication Date: 2025-07-22ANHUI AGRICULTURAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510406936.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-22
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The existing lncRNA-disease association prediction methods fail to fully explore multi-source feature information and insufficient feature extraction methods, resulting in limited ability to portray complex biological relationships and unable to effectively capture long-distance dependence information in heterogeneous networks.

Method used

Using a KAN network-based method, combining Transformer to extract global sequence features and GraphSAGE to mine heterogeneous network topology information, model nonlinear relationships through KAN network, and fuse multi-level information to improve prediction accuracy.

Benefits of technology

It improves the accuracy and generalization ability of lncRNA-disease association prediction, can more accurately characterize complex biological relationships, and improves the prediction ability and generalization performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119905141B_ABST
    Figure CN119905141B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the field of lncRNA-disease association prediction, and specifically provides a lncRNA-disease association prediction method based on the KAN network, including the following steps: fusing lncRNA-disease association network information, various similarity information of lncRNA and diseases to construct a comprehensive heterogeneous network; constructing a Transformer feature extractor based on the multi-head attention mechanism for capturing global serialized features from lncRNA and disease nodes; constructing a GraphSAGE feature extractor based on the heterogeneous graph network for capturing topological structure features of lncRNA and disease nodes; fusing the global serialized features and topological structure features to obtain comprehensive lncRNA-disease association pair feature information; constructing a KAN network for predicting the association score between lncRNA and diseases. This method combines Transformer to extract global sequence features, GraphSAGE to mine heterogeneous network topological information, and models the non-linear relationship through the KAN network, fully mining the multi-level information of lncRNA and diseases, and improving the accuracy of lncRNA-disease association prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of lncRNA-disease association prediction, and particularly relates to a method for predicting lncRNA-disease association based on a KAN network. Background Art

[0002] LncRNA is a class of non-coding RNA molecules that do not encode proteins but have important regulatory functions. They are widely involved in gene transcription, epigenetic regulation, and RNA interaction, and play a key role in cancer occurrence, development, and drug resistance; lncRNA can also affect cancer-related signaling pathways and be used as biomarkers for cancer diagnosis, prognosis evaluation, and targeted therapy, showing broad application prospects.

[0003] In related technologies, for example, LINC01608 has been experimentally determined to be a reliable prognostic biomarker for hepatocellular carcinoma; the overexpression of lncRNACTA-929C8 in the brain tissue may contribute to the development of Alzheimer's disease. Therefore, accurately identifying disease-related lncRNAs is of great significance for disease diagnosis and revealing the pathogenesis. In addition, for example, WMFLDA designed a weighted matrix factorization method to infer disease-related lncRNAs and thus predict the association information between them; MCA-Net proposed a multi-feature encoding method, combining six similarity features to construct the association features between lncRNAs and diseases, and using a convolutional neural network to infer the potential association between lncRNAs and diseases. Therefore, the method based on graph neural network has become an effective way to solve the lncRNA-disease association prediction problem. However, there is a problem that the performance of the model based on a single graph neural network is limited due to over-smoothing.

[0004] Currently, traditional lncRNA-disease association prediction methods mainly include experimental-based and computational biology-based methods. The experimental-based methods require expensive and time-consuming experimental processes to measure the association between lncRNAs and diseases, while the computational biology-based methods use known lncRNA and disease data for prediction. However, these traditional methods have problems such as incomplete data, limited feature representation, and limitations of prediction models; moreover, the availability and quality of data are limited, the feature representation lacks comprehensiveness, and the prediction model cannot fully consider the complex association between lncRNAs and diseases. Therefore, the existing prediction methods have the following problems:

[0005] First, the existing methods mainly extract node features based on lncRNA sequence similarity and disease semantic similarity, but do not fully mine the multi-source feature information of lncRNAs and diseases, restricting the model's ability to depict complex biological relationships;

[0006] Second, the feature extraction methods of existing methods cannot effectively extract the information of lncRNA and diseases. For example, convolutional neural networks are mainly good at local feature extraction and cannot effectively capture the long-range dependence information in heterogeneous networks. Summary of the Invention

[0007] The purpose of the embodiments of the present invention is to provide an lncRNA-disease association prediction method based on the KAN network. This method combines Transformer to extract global sequence features, GraphSAGE to mine heterogeneous network topology information, and models the non-linear relationship through the KAN network, fully mining the multi-level information of lncRNA and diseases and improving the accuracy of association prediction.

[0008] To achieve the above purpose, the present invention provides the following technical solutions:

[0009] An lncRNA-disease association prediction method based on the KAN network, comprising the following steps:

[0010] S1. Integrate the lncRNA-disease association network information, various similarity information of lncRNA and diseases to construct a comprehensive heterogeneous network;

[0011] S2. Construct a Transformer feature extractor based on the multi-head attention mechanism for capturing the global serialized features from lncRNA and disease nodes;

[0012] S3. Construct a GraphSAGE feature extractor based on the heterogeneous graph network for capturing the topological structure features of lncRNA and disease nodes;

[0013] S4. Integrate the global serialized features and topological structure features to obtain the comprehensive lncRNA-disease association pair feature information; construct a KAN network for predicting the association score between lncRNA and diseases. The KAN network includes a B-Spline transformation layer and a deep KAN structure. Use the learnable B-spline interpolation of the transformation layer to learn the non-linear mapping relationship of the lncRNA-disease association pair feature information, and use the deep KAN structure to predict the association score between lncRNA and diseases.

[0014] Further, in step S1, use an m×n association network to represent the association information of lncRNA-disease, where m represents the number of lncRNAs and n represents the number of diseases. When lncRNA l i and disease d j are associated, A(l i , d jis 1, otherwise 0;

[0015] The associated network definition is as follows:

[0016] ;

[0017] Calculate the preliminary feature information of lncRNAs and diseases. The preliminary feature information includes the lncRNA sequence similarity lS seq , the lncRNA functional similarity lS func and the lncRNA Gaussian interaction profile kernel similarity lS GIPK , the disease semantic similarity DS sem and the disease Gaussian interaction profile kernel similarity DS GIPK ;

[0018] Fuse the preliminary feature information of lncRNAs and diseases respectively to obtain the lncRNA comprehensive similarity feature lS and the disease comprehensive similarity feature DS. Among them, , ;

[0019] Finally, fuse the lncRNA-disease associated network information, the comprehensive similarity feature lS and the disease comprehensive similarity feature DS to construct an lncRNA-disease heterogeneous network: , where lS and DS represent the lncRNA comprehensive similarity feature matrix and the disease comprehensive similarity feature matrix respectively, and A represents their adjacency relationship matrix. A T represents the transpose of the adjacency matrix.

[0020] Furthermore, in the calculation process of the lncRNA sequence similarity lS seq , the Smith-Waterman sequence alignment algorithm is used to evaluate the sequence similarity between lncRNAs. The formula for calculating the sequence similarity is as follows:

[0021] ;

[0022] where l i and l j represent the i-th lncRNA and the j-th lncRNA respectively, and SW(l i , l j ) represents the sequence alignment score calculated based on the Smith-Waterman alignment algorithm.

[0023] Furthermore, in the calculation process of the lncRNA functional similarity lS func , it is expressed by the formula:

[0024] ;

[0025] Among them, p and q respectively represent the number of diseases associated with lncRNA l1 and l2. de represents a single disease, and DE represents a set of diseases, DE = {de1, de2,...}. S(de, DE) is used to calculate the relationship between diseases, expressed as: .

[0026] Furthermore, during the calculation of the Gaussian interaction profile kernel similarity lS of lncRNAs GIPK , the calculation formula for the GIPK kernel similarity between lncRNAs is as follows:

[0027] ;

[0028] Among them, A(l i ,) and A(l j ,) respectively represent the i-th row vector and the j-th row vector of A; λ n is the kernel width coefficient, defined as: , where N represents the total number of lncRNAs, and A(l k ,) represents the k-th row vector of A.

[0029] Furthermore, during the calculation of the disease semantic similarity DS sem , the disease ontology DO is used to measure the semantic similarity between related diseases, expressed as a directed acyclic graph. By utilizing the hierarchical structure of diseases in DO, the semantic similarity is calculated based on the directed acyclic graph, and the calculation formula is expressed as:

[0030] ;

[0031] Among them, d i and d j respectively represent the i-th disease and the j-th disease, T i and T j respectively represent a set containing the i-th and j-th diseases, and S di (t) represents the semantic impact of disease t on the i-th disease.

[0032] Furthermore, during the calculation of the disease Gaussian interaction profile kernel similarity DS GIPK , the calculation formula is expressed as:

[0033] ;

[0034] Among them, A(,d i ) and A(,d j ) respectively represent the i-th column vector and the j-th column vector of A, λ dExpressed as: , where D represents the total number of diseases, and A(,n l ) represents the l-th column vector of A.

[0035] Furthermore, in step S2, the Transformer feature extractor includes three encoding layers. In the h-th attention head of one of the encoding layers, there are three linear transformation matrices: the query matrix , the key matrix , and the value matrix . The specific linear transformation is:

[0036] ;

[0037] Among them, represents the weight matrix of the query matrix; represents the weight matrix of the key matrix; represents the weight matrix of the value matrix; H l-1 represents the output of the (l - 1)-th layer;

[0038] Calculate the attention score matrix . The output of each layer is obtained by normalizing the attention score matrix and the value matrix. The specific formula is expressed as:

[0039] ;

[0040] Among them, d represents and 's matrix dimension; represents the output of the l-th layer in the h-th attention head;

[0041] The final result of each layer is obtained by concatenating the results of each attention head through multiple attention heads; after the output of the three-layer Transformer encoding layer, the feature representation F1 is extracted.

[0042] Furthermore, in step S3, the GraphSAGE feature extractor consists of three GraphSAGE calculation layers, and feature extraction is achieved through the neighbor sampling and aggregation mechanism; among them, for each lncRNA and disease node, its neighbor node set is randomly sampled, and the mean of its features is calculated, expressed as:

[0043] ;

[0044] Among them, W k is the weight matrix, u is the set of adjacent nodes of node v represented by N(v), MEAN represents the averaging operation, and σ is the non-linear activation function, using RELU; The node representation after calculation for the k-th layer;

[0045] By adopting multi-layer calculation, the output of each layer is used as the input of the next layer, and finally the representation feature matrix F2 of the heterogeneous network is obtained.

[0046] Furthermore, in the step of fusing the global sequence feature and the topological structure feature to obtain the comprehensive lncRNA-disease association pair feature information, the fusion method is expressed as:

[0047] ;

[0048] where the coefficients α1 and α2 represent the importance degree of the features;

[0049] The calculation method of the B-Spline transformation layer is expressed as: y * = W·f(x) + b; where f(x) represents the feature after B-Spline interpolation, and W and b are trainable weights and bias terms;

[0050] The loss function is expressed as follows:

[0051] ;

[0052] where y i represents the true label of the lncRNA-disease association pair, represents the predicted label of the lncRNA-disease association pair, and N represents the number of lncRNA-disease association pairs.

[0053] Compared with the prior art, the technical advantages of the lncRNA-disease association prediction method based on the KAN network of the present invention are shown in the following aspects:

[0054] First, the lncRNA-disease association prediction method proposed by the present invention is used to identify potential lncRNA-disease associations. This method combines Transformer to extract global sequence features, GraphSAGE to mine heterogeneous network topological information, and models the non-linear relationship through the KAN network, fully mining the multi-level information of lncRNA and diseases, and improving the accuracy of association prediction;

[0055] Second, the present invention combines two feature learning extractors, Transformer and GraphSAGE, which are respectively used to capture the global feature information of lncRNA and diseases and the network topological structure information; the two modules cooperate with each other, enabling TransSAGE-KAN-LDA to more accurately depict the complex relationship between lncRNA and diseases, and improving the accuracy and generalization ability of prediction.

[0056] In summary, the present invention combines two feature learning modules, Transformer and GraphSAGE, which are respectively used to extract the global feature information of lncRNAs and diseases and capture the topological structure information of the heterogeneous network; uses the KAN network for feature association prediction, fully integrating multi-source information to more accurately identify potential lncRNA-disease associations; through the complementary learning of Transformer and GraphSAGE, this method can effectively characterize the complex relationship between lncRNAs and diseases, improving the prediction ability and generalization performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention.

[0058] Figure 1 It is a schematic diagram of the implementation process of the lncRNA-disease association prediction method based on the KAN network of the present invention;

[0059] Figure 2 It is a schematic diagram of the ROC curve on the dataset under five-fold cross-validation of the present invention;

[0060] Figure 3 It is a schematic diagram of the PR curve on the dataset under five-fold cross-validation of the present invention;

[0061] Figure 4 It is a system architecture diagram of the lncRNA-disease association prediction method based on the KAN network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further details the present invention with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0063] The following describes the specific implementation of the present invention in detail with reference to specific embodiments.

[0064] Please refer to Figure 1 and Figure 4 , in an embodiment provided by the present invention, a lncRNA-disease association prediction method based on the KAN network is provided, and the method includes the following steps:

[0065] S1. Construct a comprehensive heterogeneous network by integrating lncRNA-disease association network information, various similarity information of lncRNAs and diseases. In this step, it is necessary to first collect data, including lncRNA sequence data, disease DOID data, and known lncRNA-disease association information in public databases.

[0066] S2. Construct a Transformer feature extractor based on the multi-head attention mechanism to capture the global serialized features from lncRNA and disease nodes.

[0067] S3. Construct a GraphSAGE feature extractor based on the heterogeneous graph network to capture the topological structure features of lncRNA and disease nodes.

[0068] S4. Fuse the global serialized features and topological structure features to obtain comprehensive lncRNA-disease association pair feature information. Construct a KAN network for predicting the association score between lncRNAs and diseases. The KAN network includes a B-Spline transformation layer and a deep KAN structure. Use the learnable B-spline interpolation of the transformation layer to learn the non-linear mapping relationship of lncRNA-disease association pair feature information, and use the deep KAN structure to predict the association score between lncRNAs and diseases.

[0069] Therefore, the association prediction method provided by the present invention combines two feature learning modules, Transformer and GraphSAGE, which are respectively used to extract the global feature information of lncRNAs and diseases and capture the topological structure information of the heterogeneous network. Further, use the KAN network for feature association prediction, fully integrating multi-source information to more accurately identify potential lncRNA-disease associations.

[0070] Through the complementary learning of Transformer and GraphSAGE, the present method can effectively characterize the complex relationship between lncRNAs and diseases, improving the prediction ability and generalization performance of the model.

[0071] In addition, the association prediction method provided by the present invention also optimizes the model parameters through backpropagation and gradient descent algorithms.

[0072] In step S1 of the present invention, collect lncRNA sequence data, disease DOID data, and known lncRNA-disease association information in public databases.

[0073] Among them, during the data collection process, RNADisease v4.0 contains 3,428,058 lncRNA-disease associations, covering 17,820 lncRNAs and 4,090 diseases; Lnc2Cancer 3.041 collects a total of 1,057 associations, involving 531 lncRNAs and 86 human cancers; LncRNADisease v3.0 collects 25,440 experimentally supported lncRNA-disease associations, covering 6,066 lncRNAs and 566 diseases.

[0074] After merging the above three datasets, simple screening was performed on the lncRNA data, and duplicates were removed. Finally, 886 lncRNA-disease associations were obtained, involving 406 lncRNAs and 255 diseases.

[0075] Subsequently, in step S1 of the present invention, an m×n association network A∈R m×n is used to represent the lncRNA-disease association information, where m represents the number of lncRNAs, n represents the number of diseases. When lncRNA l i and disease d j are associated, A(l i , d j ) at the corresponding position is 1, otherwise it is 0; among them, the association network is defined as follows:

[0076] ;

[0077] Calculate the preliminary feature information of lncRNAs and diseases. The preliminary feature information includes lncRNA sequence similarity lS seq , lncRNA functional similarity lS func and lncRNA Gaussian interaction spectrum kernel similarity lS GIPK , disease semantic similarity DS sem and disease Gaussian interaction spectrum kernel similarity DS GIPK ;

[0078] Fuse the preliminary feature information of lncRNAs and diseases respectively to obtain the lncRNA comprehensive similarity feature lS and the disease comprehensive similarity feature DS. Among them, , ;

[0079] Finally, fuse the lncRNA-disease association network information, the comprehensive similarity feature lS and the disease comprehensive similarity feature DS to construct an lncRNA-disease heterogeneous network: , where \(l_S\) and \(D_S\) respectively represent the comprehensive similarity feature matrix of lncRNAs and the comprehensive similarity feature matrix of diseases, and \(A\) represents their adjacency relationship matrix, \(A\) T represents the transpose of the adjacency matrix;

[0080] In addition, the present invention also calculates five types of similarity information related to lncRNAs and diseases; the five types of similarity information include: lncRNA sequence similarity, lncRNA functional similarity, and lncRNA Gaussian interaction spectrum kernel similarity, disease semantic similarity, and disease Gaussian interaction spectrum kernel similarity;

[0081] Among them, in one implementation, the lncRNA sequence similarity \(l_S\) seq In the calculation process, for lncRNA sequence similarity, according to the principle that lncRNAs with similar sequences are more likely to have similar functions, the present invention uses the Smith-Waterman sequence alignment algorithm to evaluate the sequence similarity between lncRNAs; the formula for calculating the sequence similarity is as follows:

[0082] ;

[0083] where \(l\) i and \(l\) j respectively represent the \(i\)-th lncRNA and the \(j\)-th lncRNA; \(SW(l\) i , \(l\) j ) represents the sequence alignment score between lncRNA \(l\) i and lncRNA \(l\) j calculated based on the Smith-Waterman alignment algorithm.

[0084] Among them, in one implementation, the lncRNA functional similarity \(l_S\) func In the calculation process, for lncRNA functional similarity, this method adopts a direct method to reflect their functional similarity, aiming to reveal the potential association between them and diseases with similar pathological phenomena or symptoms; the functional similarity of lncRNAs can be calculated and measured by evaluating the similarity between the diseases associated with them, and is expressed by the formula:

[0085] ;

[0086] where \(p\) and \(q\) respectively represent the number of diseases associated with lncRNA \(l_1\) and lncRNA \(l_2\), \(de\) represents a disease, \(DE\) is defined as a set of diseases, \(DE = \{de_1, de_2,...\}\), and \(S(de, DE)\) is used to calculate the relationship between diseases, expressed as: .

[0087] Among them, in one implementation, the Gaussian interaction spectrum kernel similarity lS of lncRNA GIPK In the calculation process, for the Gaussian interaction spectrum kernel similarity of lncRNA, the Gaussian interaction feature kernel similarity is usually used to evaluate the similarity between two similar nodes in the non-coding RNA-disease association prediction task, indicating that in a specific disease, lncRNAs with similar functions have similar patterns. The calculation formula for the GIPK kernel similarity between lncRNAs is as follows:

[0088] ;

[0089] Among them, A(l i ,) and A(l j ,) respectively represent the i-th row vector and the j-th row vector of A; λ n is the kernel width coefficient, defined as: , where N represents the total number of lncRNAs, and A(l k ,) represents the k-th row vector of A.

[0090] Furthermore, in one implementation, in the calculation process of the disease semantic similarity DS sem , for the disease semantic similarity, the present invention uses the Disease Ontology (DO) to measure the semantic similarity between related diseases, which is usually represented as a directed acyclic graph. By utilizing the hierarchical structure of diseases in DO and calculating their semantic similarity based on the directed acyclic graph, the more parent diseases shared between diseases, the greater their similarity. The calculation formula is expressed as:

[0091] ;

[0092] Among them, d i and d j respectively represent the i-th disease and the j-th disease, T i and T j respectively represent a set containing the i-th disease and the j-th disease, and S di (t) represents the semantic impact of disease t on the i-th disease.

[0093] Among them, in one implementation, in the calculation process of the disease Gaussian interaction spectrum kernel similarity DS GIPK , for the disease Gaussian interaction spectrum kernel similarity, similar to the Gaussian interaction spectrum kernel similarity of lncRNA, only the formula is given, and the calculation formula is expressed as:

[0094] ;

[0095] where A(,d i ) and A(,d j ) represent the i-th column vector and the j-th column vector of A respectively, and the definition of λ d is: , D represents the total number of diseases, and A(,n l ) represents the l-th column vector of A.

[0096] After that, the present invention obtains a comprehensive lncRNA similarity network and a comprehensive disease similarity network by fusing the similarity networks of lncRNA and diseases respectively.

[0097] Furthermore, in step S2, the present invention designs a Transformer feature extractor based on the multi-head attention mechanism, aiming to capture the global sequential features from lncRNA and disease nodes. By introducing the multi-head attention mechanism, the model is allowed to learn multiple different representations in parallel, reducing the dependence on a single representation and making the learning process more robust;

[0098] Specifically, the Transformer feature extractor provided by the present invention includes three encoding layers. In the h-th attention head of one of the encoding layers, it contains three linear transformation matrices: the query matrix , the key matrix and the value matrix . The specific linear transformation is:

[0099] ;

[0100] where, represents the weight matrix of the query matrix; represents the weight matrix of the key matrix; represents the weight matrix of the value matrix; H l-1 represents the output of the (l - 1)-th layer;

[0101] Calculate the attention score matrix , and obtain the output of each layer by normalizing the attention score matrix and the value matrix. The specific formula is:

[0102] ;

[0103] where, d represents and matrix dimensions; represents the output of the l-th layer in the h-th attention head;

[0104] The results of each attention head are concatenated through multiple attention heads to obtain the final result of each layer; after the output of the three-layer Transformer encoding layer, the feature representation F1 is extracted.

[0105] Further, in step S3, the present invention designs a GraphSAGE feature extractor based on a heterogeneous graph network, aiming to capture the topological structure information of the heterogeneous network. The GraphSAGE feature extractor provided by the present invention consists of three layers of GraphSAGE calculation layers, and feature extraction is realized through a neighbor sampling and aggregation mechanism; wherein, for each lncRNA and disease node, its neighbor node set is randomly sampled, and the mean value of its features is calculated, which is expressed as:

[0106] ;

[0107] Among them, where, W k is the weight matrix, u is the set of adjacent nodes of node v represented by N(v), MEAN represents the averaging operation, σ is the non-linear activation function. Preferably, the activation function σ adopts the RELU function; is the node representation after the k-th layer calculation; this process can effectively aggregate the local information of lncRNA and disease nodes, so that the final node representation can synthesize the features of neighbors; by adopting multi-layer calculation, the output of each layer is used as the input of the next layer, and finally the representation feature matrix F2 of the heterogeneous network is obtained.

[0108] Further, in the step of fusing the global serialized feature and the topological structure feature to obtain the comprehensive lncRNA-disease association pair feature information, the fusion method is expressed as:

[0109] ;

[0110] Among them, the coefficients α1 and α2 represent the importance degree of the features;

[0111] In an implementation manner of the present invention, a new network KAN (Kolmogorov-Arnold Network) is constructed, and its core components include a B-Spline transformation layer and a deep KAN structure. By using learnable B-splines interpolation to learn the non-linear mapping relationship of data, the limitation of the fixed activation function is avoided, and the parameter efficiency is improved at the same time;

[0112] Specifically, the calculation method of the B-Spline transformation layer is expressed as: y * =W·f(x)+b;

[0113] Among them, f(x) represents the feature after B-Spline interpolation, and W and b are trainable weights and bias terms;

[0114] Furthermore, the loss function is expressed as follows:

[0115] ;

[0116] where y i represents the true label of the lncRNA-disease association pair, represents the predicted label of the lncRNA-disease association pair, and N represents the number of lncRNA-disease association pairs.

[0117] Furthermore, please refer to Figure 3 and Figure 4 , in the validation of the effectiveness of the embodiments of the present invention, the present invention uses five-fold cross-validation to evaluate the performance of the model, which helps to comprehensively evaluate the generalization ability of the model and can effectively prevent the model from overfitting.

[0118] Specifically, the present invention regards all known lncRNA-disease associations as positive samples. Unknown lncRNA-disease associations mean that there is no relevant biological experiment to verify their correlation with each other. Therefore, all unknown lncRNA-disease associations are regarded as negative samples. To ensure the balance of positive and negative samples, the present invention randomly selects the same number of negative samples as the positive samples from all negative samples.

[0119] Since AUPR and AUC can represent the performance of the model at different thresholds, the present invention selects AUPR and AUC as the main evaluation indicators.

[0120] In addition, the present invention also uses other threshold-based evaluation indicators, such as F1 score F1_score, accuracy Accuracy, recall Recall, and precision Precision. The relevant mathematical formulas are expressed as follows:

[0121] ;

[0122] ;

[0123] ;

[0124] ;

[0125] where TP, FP, TN, and FN represent true positive, false positive, true negative, and false negative, respectively.

[0126] Among them, AUC refers to the area under the curve of the receiver operating characteristic (ROC) curve, which quantitatively reflects the performance of the model measured based on the ROC curve; AUPR represents the area under the Precision-Recall (PR) curve. The larger the area under the PR curve, the better the performance of the model. The results of the model performance are shown in Table 1.

[0127] Table 1 Five-fold cross-validation result table

[0128]

[0129] To evaluate the prediction performance of the model, the method TransSAGE-KAN-LDA of the present invention is compared with several existing methods. The experimental results are shown in Table 2.

[0130] Table 2 Performance comparison of different methods under five-fold cross-validation

[0131]

[0132] Analyzing the above validation results of the effectiveness of the embodiments of the present invention, it can be seen that the lncRNA-disease association prediction method TransSAGE-KAN-LDA of the present invention based on multi-layer Transformer, GraphSAGE and KAN network is used to identify potential lncRNA-disease associations. This method fully exploits five types of similarity information from lncRNAs and diseases. The present invention introduces two feature learning extractors, Transformer and GraphSAGE, from different perspectives, which are used to capture the global feature information of lncRNAs and diseases and the network topology information respectively. Through the interaction and complementarity between each module, TransSAGE-KAN-LDA can capture the complex relationship between lncRNAs and diseases more accurately. Under five-fold cross-validation on the dataset, TransSAGE-KAN-LDA of the present invention reaches AUC values of 0.9868 and 0.9803 respectively, exceeding other models.

[0133] Although the embodiments of the present invention have been disclosed as above, they are not limited to only the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the examples shown and described herein.

Claims

1. A method for predicting lncRNA-disease associations based on the KAN network, characterized in that, It includes the following steps: S1. Construct a comprehensive heterogeneous network by integrating lncRNA-disease association network information and various similarity information of lncRNAs and diseases; S2. Construct a Transformer feature extractor based on the multi-head attention mechanism to capture the global serialized features from lncRNA and disease nodes; S3. Construct a GraphSAGE feature extractor based on the heterogeneous graph network to capture the topological structure features of lncRNA and disease nodes; S4. Fuse the global serialized features and topological structure features to obtain the comprehensive lncRNA-disease association pair feature information; construct a KAN network for predicting the association score between lncRNAs and diseases. The KAN network includes a B-Spline transformation layer and a deep KAN structure. Use the learnable B-spline interpolation of the transformation layer to learn the non-linear mapping relationship of the lncRNA-disease association pair feature information, and use the deep KAN structure to predict the association score between lncRNAs and diseases; In step S2, the Transformer feature extractor includes three encoding layers. In the h-th attention head of one of the encoding layers, there are three linear transformation matrices: the query matrix , the key matrix , and the value matrix . The specific linear transformation is as follows: ; Among them, represents the weight matrix of the query matrix; represents the weight matrix of the key matrix; represents the weight matrix of the value matrix; H l-1 represents the output of the (l-1)-th layer; Calculate the attention score matrix , and obtain the output of each layer by normalizing the attention score matrix and the value matrix. The specific formula is as follows: ; where d represents and matrix dimensions; represents the output of the l-th layer in the h-th attention head; Obtain the final result of each layer by concatenating the results of each attention head through multiple attention heads; after the output of the three-layer Transformer encoding layer, the feature representation F1 is extracted; In step S3, the GraphSAGE feature extractor consists of three layers of GraphSAGE calculation layers, and feature extraction is achieved through the neighbor sampling and aggregation mechanism; among them, for each lncRNA and disease node, randomly sample its neighbor node set and calculate the mean of its features, expressed as: ; Among them, W k is the weight matrix, u is the set of adjacent nodes of node v represented by N(v), MEAN represents the averaging operation, σ is the non-linear activation function, and RELU is adopted; is the node representation after the calculation of the k-th layer; Through multi-layer calculation, the output of each layer is used as the input of the next layer, and finally the representation feature matrix F2 of the heterogeneous network is obtained; In the step of fusing the global serialized features and topological structure features to obtain the comprehensive lncRNA-disease association pair feature information, the fusion method is expressed as: ; Among them, the coefficients α1 and α2 represent the importance degree of the features; The calculation method of the B-Spline transformation layer is expressed as: y * = W·f(x) + b; Among them, f(x) represents the feature after B-Spline interpolation, and W and b are trainable weights and bias terms; The loss function is expressed as follows: ; where y i represents the true label of the lncRNA-disease association pair, represents the predicted label of the lncRNA-disease association pair, and N represents the number of lncRNA-disease association pairs.

2. The lncRNA-disease association prediction method based on the KAN network according to claim 1, wherein In step S1, use an m×n association network to represent the lncRNA-disease association information, where m represents the number of lncRNAs and n represents the number of diseases; When lncRNA l i is associated with disease d j , A(l i , d j ) at the corresponding position is 1, otherwise it is 0; the association network is defined as follows: ; Calculate the preliminary feature information of lncRNA and diseases. The preliminary feature information includes the lncRNA sequence similarity lS seq , the lncRNA functional similarity lS func and the lncRNA Gaussian interaction profile kernel similarity lS GIPK , the disease semantic similarity DS sem and the disease Gaussian interaction profile kernel similarity DS GIPK ; Fuse the preliminary feature information of lncRNA and diseases respectively to obtain the lncRNA comprehensive similarity feature lS and the disease comprehensive similarity feature DS; among them, , ; Integrate lncRNA-disease association network information, comprehensive similarity feature lS, and disease comprehensive similarity feature DS to construct an lncRNA-disease heterogeneous network: , where lS and DS represent the comprehensive similarity feature matrix of lncRNAs and the comprehensive similarity feature matrix of diseases, respectively, and A represents the adjacency relationship matrix between lS and DS, and A T represents the transpose of the adjacency matrix.

3. A method for predicting lncRNA-disease associations based on the KAN network according to claim 2, wherein lncRNA sequence similarity lS seq During the calculation process, the Smith-Waterman sequence alignment algorithm is used to evaluate the sequence similarity between lncRNAs; The formula for calculating sequence similarity is as follows: ; where l i and l j represent the i-th lncRNA and the j-th lncRNA respectively, and SW(l i , l j ) represents the sequence alignment score between lncRNA l i and l j .

4. The lncRNA-disease association prediction method based on the KAN network according to claim 3, wherein lncRNA functional similarity lS func In the calculation process, it is expressed by the formula as follows: ; Among them, p and q respectively represent the number of diseases associated with lncRNAs l1 and l2, de represents a disease, DE represents a set of diseases, DE = {de1, de2,...}, and S(de, DE) is used to calculate the relationship between diseases, expressed as: 。 5. A method for predicting lncRNA-disease associations based on the KAN network according to claim 4, wherein lncRNA Gaussian Interaction Profile Kernel Similarity lS GIPK During the calculation of, the GIPK kernel similarity calculation formula between lncRNAs is as follows: ; Among them, A(l i ,) and A(l j ,) represent the i-th row vector and the j-th row vector of A respectively; λ n is the kernel width coefficient, defined as: , where N represents the total number of lncRNAs, and A(l k ,) represents the k-th row vector of A.

6. The lncRNA-disease association prediction method based on the KAN network according to claim 5, wherein Disease Semantic Similarity DS sem In the calculation process of, the Disease Ontology DO is used to measure the semantic similarity between related diseases, which is represented as a directed acyclic graph. By utilizing the hierarchical structure of diseases in DO, the semantic similarity is calculated based on the directed acyclic graph, and the calculation formula is expressed as: ; where, d i and d j represent the i-th disease and the j-th disease respectively, T i and T j represent a set containing the i-th and j-th diseases respectively, and S di (t) represents the semantic impact of disease t on the i-th disease.

7. A method for predicting lncRNA-disease associations based on the KAN network according to claim 6, characterized in that Disease Gaussian Interaction Spectrum Kernel Similarity DS GIPK In the calculation process, the calculation formula is expressed as: ; Among them, A(, d i ) and A(, d j ) represent the i-th column vector and the j-th column vector of A respectively, and λ d is expressed as: , D represents the total number of diseases, A(,n l ) represents the l-th column vector of A.

Citation Information

Patent Citations

  • MiRNA-disease association prediction based on graph neural network

    CN114242237A

  • IncRNA and disease association prediction method based on multi-view data and hypergraph learning

    CN116913374A

  • Multi-channel attention mechanism lncRNA-miRNA association prediction method

    CN119207579A