Long non-coding RNA-disease association prediction method and system based on self-attention mechanism
By introducing a self-attention mechanism in the prediction of long-chain non-coding RNA-disease associations, extracting the multi-hop topological pathway features and building a prediction model, the problem of existing methods ignoring topological relationships is solved, and the prediction accuracy is significantly improved.
Patent Information
- Application Number
- CN202210818621.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-07-12
AI Technical Summary
The existing long-chain non-coding RNA-disease association prediction method based on deep learning ignores the potential topological relationships in the network during the feature extraction stage and fails to represent the multi-hop path information between the prediction targets, resulting in insufficient prediction accuracy.
The method based on the self-attention mechanism is used to extract the multi-hop topological path features in the heterogeneous network, and a prediction model is constructed through the self-attention network to focus on the interdependence between topological paths from a global perspective.
The prediction accuracy of the model was improved and higher AUC and AUPR indicators were obtained. Compared with the best algorithm GAERF, the prediction accuracy was improved by 1.4 percentage points and 21.8 percentage points.
Smart Images

Figure CN115171780B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics technology, and particularly relates to a method and system for predicting long non-coding RNA-disease associations based on self-attention mechanism. Background Art
[0002] Long non-coding RNA is a type of non-coding RNA with a length greater than 200 nucleotides. As an important regulatory factor in organisms, long non-coding RNA participates in many key gene regulation processes and is closely related to the occurrence and development of various human diseases. For example, the apoptosis-related transcript in bladder cancer (AATBC) promotes the metastasis of nasopharyngeal carcinoma by regulating the desmosome-related protein (Pinin).
[0003] In recent years, some long non-coding RNA-disease association relationships have been verified by wet biological experiments. However, the high time and resource costs have greatly restricted the further development of wet biological experimental methods. Based on the known experimental data, researchers are increasingly using computational methods to predict long non-coding RNA-disease association relationships, so as to provide reliable guidance for wet biological experimental verification.
[0004] According to the categories of the technologies used by the prediction models, the existing computational methods in this field can be divided into two categories: prediction methods based on traditional machine learning and prediction methods based on deep learning. The prediction methods based on traditional machine learning were earlier applied to the analysis of long non-coding RNA-disease associations, and through matrix operations, network propagation, and classifier algorithms, feature extraction and binary classification prediction were realized. Deep learning has powerful non-linear fitting capabilities and is the primary choice for constructing models in current computational methods. Compared with machine learning, the prediction methods based on deep learning are more flexible in structure and can end-to-end realize low-dimensional feature embedding representation and classification prediction, thereby improving the prediction accuracy of the model.
[0005] However, there are still some deficiencies in the existing long non-coding RNA-disease association prediction methods based on deep learning. The processing flow of the prediction methods based on deep learning can be divided into a feature extraction stage and a model stage. In the feature extraction stage, the current encoding method integrates the known long non-coding RNA and disease association situations into a heterogeneous network and calculates the intra-class similarity for nodes of the same type. However, only two-hop information between the prediction targets is extracted, ignoring the potential topological relationships in the network and failing to represent the multi-hop pathway information between the prediction targets. In the model stage, existing models mostly adopt a combination of multi-layer perceptron (MLP), convolutional neural network (CNN), and graph neural network (GNN), but there are problems such as complex model structure and confusing feature semantics, so the prediction accuracy still needs to be further improved. Summary of the Invention
[0006] To solve the above problems, the purpose of the present invention is to provide a method and system for predicting long non-coding RNA-disease associations based on the self-attention mechanism, which extracts multi-hop topological pathway features in a heterogeneous network and focuses on the interdependence between topological pathways from a global perspective, thereby improving the prediction accuracy of the model. The technical solutions are as follows:
[0007] A method for predicting long non-coding RNA-disease associations based on the self-attention mechanism, comprising the following steps:
[0008] S1: Data acquisition and preprocessing: Obtain disease semantic information and known long non-coding RNA-disease association situations, construct a heterogeneous network, and calculate the similarity between homogeneous nodes;
[0009] S2: Topological feature extraction: Represent the similarity and association situation as a weighted adjacency matrix of the heterogeneous network, and according to the definition of the power of the adjacency matrix, obtain the multi-hop topological pathway features between target long non-coding RNA-disease node pairs;
[0010] S3: Predicting association relationships based on the self-attention network: Construct a neural network based on the self-attention mechanism, obtain feature embeddings and perform classification through a single-layer fully connected network, convert the output matrix of the neural network based on the self-attention mechanism into a probability distribution, and predict whether there is an association relationship;
[0011] S4: Model training: Perform model training on the training set based on the backpropagation algorithm to obtain a long non-coding RNA-disease association prediction model;
[0012] S5: Predicting association relationships: Process target long non-coding RNA-disease node pairs through the prediction model to determine whether there is an association relationship.
[0013] Further, the step S1 specifically includes:
[0014] S11: Integrating long non-coding RNA-disease association relationships:
[0015] For l long non-coding RNAs and d diseases, let A LD ∈R l×d represent their association relationships; if a long non-coding RNA-disease node pair is known to be associated, the corresponding position value in A LD is 1, otherwise the value is 0;
[0016] S12: Calculating disease similarity:
[0017] Obtain the parent-child relationship between diseases and represent it as a directed acyclic graph; for disease t1, denote T1 as the set containing t1 and its ancestral diseases, then the semantic contribution value of any disease t in T1 relative to t1 is:
[0018]
[0019] Among them, children of t refers to the set of sub-diseases of disease t;
[0020] Let Then the calculation process of the similarity between diseases t1 and t2 is as follows:
[0021]
[0022] Among them, SV(t1) and SV(t2) respectively represent the sum of the disease semantic contribution values of sets T1 and T2;
[0023] For the d diseases in the dataset, calculate the similarity between each pair of them to obtain the disease similarity matrix S D ∈R d×d ;
[0024] S13: Calculate the similarity of long non-coding RNAs:
[0025] It is known that long non-coding RNAs r1 and r2 are respectively related to n1 and n2 diseases, and these diseases are denoted as t 1k and t 2m , 1 ≤ k ≤ n1, 1 ≤ m ≤ n2; then the calculation process of the similarity between long non-coding RNAs r1 and r2 is as follows:
[0026]
[0027] Among them, Sim(t 1k , t 2m ) and Sim(t 2m , t 1k ) represent the disease similarity;
[0028] For the l long non-coding RNAs in the dataset, calculate the similarity between each pair of them to obtain the long non-coding RNA similarity matrix S L ∈R l×l .
[0029] Furthermore, the specific steps of step S2 are as follows:
[0030] S21: Integrate the weighted adjacency matrix:
[0031] Integrate the long non-coding RNA-disease association matrix A LD , the long non-coding RNA similarity matrix S L and the disease similarity matrix S D into the weighted adjacency matrix A of the heterogeneous network:
[0032]
[0033] Among them, is the transpose of A LD ;
[0034] S22: Process the multi-hop adjacency matrix:
[0035] Based on the weighted adjacency matrix A, obtain the normalized multi-hop adjacency matrix according to the following calculation process:
[0036]
[0037]
[0038] where n h represents the number of hops, A diag0 means setting the values on the main diagonal of A to 0; max(A diag0 ) represents the maximum value in the A diag0 matrix;
[0039] S23: Topological feature extraction:
[0040] For the i-th node of the target long non-coding RNA - the j-th node of the disease, extract topological features through the following calculation process:
[0041]
[0042] where and represent the i-th column and the j-th column of the n h hop adjacency matrix , is the linear transformation matrix, is the matrix input for the subsequent model, d model is the dimension parameter of the matrix input.
[0043] Furthermore, the specific steps of step S3 are as follows:
[0044] S31: Construct a neural network based on the self-attention mechanism to extract topological path features:
[0045] The neural network based on the self-attention mechanism consists of N identical layers, each layer containing two sub-layers: the multi-head self-attention sub-layer and the position feed-forward network sub-layer. The output of each sub-layer is processed through residual connection and layer normalization, and each sub-layer generates an output with a dimension of d model ; The formula for normalization processing is as follows:
[0046] H(X) = LayerNorm(X + Sublayer(X))
[0047] Among them, Sublayer(X) represents the function implemented by the sublayer, and X represents the matrix input of each sublayer;
[0048] The dot-product-based attention includes three matrix inputs Q, K, and V, and its calculation process is as follows:
[0049]
[0050] Among them, represents the channel dimension of K. The multi-head self-attention sublayer splits the calculation process of the dot-product-based attention into multiple heads; the multi-head attention calculation formula is as follows:
[0051]
[0052] head i = Attention(Q i , K i , V i )
[0053] Among them, n head represents the number of heads, 1 ≤ i ≤ n head ; is the linear transformation matrix;
[0054] The three matrices of each head are obtained through the following linear transformation:
[0055] Q i = XW i Q
[0056] K i = XW i K
[0057] V i = XW i V
[0058] Among them, W i Q , W i K , are three linear transformation matrices; Q i , K i and V i are the three parameter matrices after linear transformation respectively;
[0059] The position feed-forward network sublayer includes two linear transformations, which are activated by the ReLU activation function in the middle:
[0060] FFN(X) = max(0, XW1 + b1)W2 + b2
[0061] Among them, and are conversion parameters; here, d ff = d model / 2, representing the size of the middle layer; b1 and b2 are offset constants;
[0062] S33: Predict the association situation:
[0063] Flatten the output matrix of the self-attention network into a one-dimensional array, and then calculate the predicted association probability p between long non-coding RNA and diseases through a linear transformation and a sigmoid activation function. The calculation formula is as follows:
[0064] p = sigmoid(flatten(X encoded )W out + b out )
[0065] sigmoid(x) = 1 / (1 + e -x )
[0066] Among them, is the parameter of the linear transformation matrix, b out is the offset constant; flatten(·) means connecting the matrix row by row from beginning to end and flattening it into a one-dimensional array.
[0067] Furthermore, the specific steps of step S4 are as follows:
[0068] S41: Use dropout to reduce overfitting of the model:
[0069] S411: After calculating the matrix X, add a dropout function to process the matrix X:
[0070] dropout(X, p drop )
[0071] Among them, p drop is the dropout probability;
[0072] S412: After completing the sub-layer calculation, add a dropout function to process the output:
[0073] dropout(Sublayer(X), p drop )
[0074] Among them, Sublayer(X) means processing the input matrix through a multi-head self-attention sub-layer or a position feed-forward network sub-layer;
[0075] S413: When After flattening into a one-dimensional array, add a dropout function to process the output:
[0076] dropout(flatten(X encoded ),p drop )
[0077] S42: Accelerate model convergence: Use the adam optimizer to accelerate model convergence and train the model with the cross-entropy loss function;
[0078] Loss=-∑[ylog(p)+(1-y)log(1-p)]
[0079] where p represents the predicted association probability and y represents the training label.
[0080] A system for constructing a prediction model for long non-coding RNA-disease associations includes a processor, a memory, and a computer program stored on the memory. The computer program is executed on the processor to implement the above method.
[0081] The beneficial effects of the present invention are:
[0082] 1) In the feature extraction stage, the present invention designs a simple and novel topological feature extraction process to obtain multi-hop topological paths between target long non-coding RNA-disease node pairs, providing more effective features for the model;
[0083] 2) In the model stage, the present invention introduces a self-attention mechanism to construct a prediction model, focuses on topological features from a global perspective, and assigns higher weights to key association paths to enable the network to fully learn key features, thereby improving the model prediction accuracy;
[0084] 3) Compared with existing models on the same dataset, the optimal prediction metrics (AUC = 0.994, AUPR = 0.709) are obtained. Compared with the average AUC value of 0.980 and the average AUPR value of 0.491 of the previous best algorithm GAERF (the model with the highest prediction accuracy on the same dataset in the currently disclosed technical solutions), it is improved by 1.4 percentage points and 21.8 percentage points respectively;
[0085] 4) The present invention trains a neural network through existing experimental data to predict whether there is an association between unvalidated long non-coding RNA-diseases to guide wet biological experiments, effectively reducing experimental time and financial losses. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 is a flowchart of the method and system for predicting long non-coding RNA-disease associations based on the self-attention mechanism of the present invention.
[0087] Figure 2 This is a schematic diagram of the model structure of the method and system for predicting long non-coding RNA-disease associations based on the self-attention mechanism of the present invention. Detailed implementation manners
[0088] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0089] The present invention proposes a method and system for predicting long non-coding RNA-disease associations based on the self-attention mechanism. This method extracts multi-hop topological paths between target long non-coding RNA-disease node pairs, and introduces the self-attention mechanism to construct a prediction model, assigning higher weights to key topological paths so that the network can fully learn key features, thereby improving the prediction accuracy of the model.
[0090] This embodiment provides a method for predicting long non-coding RNA-disease associations based on the self-attention mechanism, referring to Figure 1 and Figure 2 , which is implemented based on python3.8.8-pytorch1.10.0. This method includes:
[0091] I. Data acquisition and preprocessing
[0092] Obtain disease semantic information and long non-coding RNA-disease association data and perform data preprocessing, construct a heterogeneous network, and calculate the similarity between homogeneous nodes.
[0093] Obtain disease semantic information and long non-coding RNA-disease association data verified by biological wet experiments from public databases, and calculate disease similarity and long non-coding RNA similarity.
[0094] 1. Obtain data. The disease semantic information comes from the disease subordination relationship data summarized by the Disease Ontology international project, which represents the ancestor-descendant information between diseases in the structure of a directed acyclic graph. The long non-coding RNA-disease association data comes from a public dataset, which records "long non-coding RNA, disease, experimentally verified association" in tabular format.
[0095] 2. Integrate known association data. For l long non-coding RNAs and d diseases, let A LD ∈R l×d represent their association relationship. If a long non-coding RNA-disease node pair is known to be associated, the corresponding position value in A LD is 1, otherwise the value is 0.
[0096] 3. Calculate disease similarity. Based on the directed acyclic graph of disease terms, calculate the semantic similarity between diseases. For disease t1, let T1 be the set of terms that includes t1 and its ancestral diseases. Then, the semantic contribution value of any disease term t in T1 relative to t1 is:
[0097]
[0098] Let Similarly for SV(t2). Then, the calculation process for the similarity between diseases t1 and t2 is:
[0099]
[0100] For the d diseases in the dataset, calculate the similarity between each pair of them to obtain the disease similarity matrix S D ∈R d×d .
[0101] 4. Calculate long non-coding RNA similarity. Based on the known association relationships and the disease similarity calculated above, calculate the long non-coding RNA similarity. Given that long non-coding RNAs r1 and r2 are respectively associated with n1 and n2 diseases, denote these diseases as t 1k , (1 ≤ k ≤ n1) and t 2m , (1 ≤ m ≤ n2). Then, the calculation process for the similarity between long non-coding RNAs r1 and r2 is:
[0102]
[0103] For the l long non-coding RNAs in the dataset, calculate the similarity between each pair of them to obtain the long non-coding RNA similarity matrix S L ∈R l×l .
[0104] II. Topological Feature Extraction
[0105] Represent the similarity and association situation as the weighted adjacency matrix of a heterogeneous network, and according to the definition of the power of the adjacency matrix, obtain the multi-hop topological path features between target long non-coding RNA-disease node pairs.
[0106] 1. Construct a weighted adjacency matrix. Integrate the long non-coding RNA-disease association matrix A LD , the long non-coding RNA similarity matrix S L and the disease similarity matrix S D into the weighted adjacency matrix A:
[0107]
[0108] Among them, is ALD Transpose of
[0109] 2. Process the multi-hop adjacency matrix. Based on the constructed weighted adjacency matrix A, obtain the normalized multi-hop adjacency matrix:
[0110]
[0111]
[0112] where the power n h represents the number of hops, and A diag0 means setting the values on the main diagonal of A to 0. Here, n h is set to 3.
[0113] 3. Feature extraction. For the target long non-coding RNA (the i-th node) - disease (the j-th node), extract topological features as follows to obtain the matrix input of the model
[0114]
[0115] where and represent the i-th and j-th columns of the n h hop adjacency matrix , and is the linear transformation matrix.
[0116] III. Predict the association relationship based on the self-attention network
[0117] Construct a neural network based on the self-attention mechanism, obtain feature embeddings and classify through a single-layer fully connected network. Convert the output matrix of the neural network based on the self-attention mechanism (self-attention mechanism) into a probability distribution to predict whether there is an association relationship.
[0118] 1. Feature embedding of topological pathways. Complete the feature embedding of topological pathways through the self-attention network. This network consists of N identical layers, and each layer contains two sub-layers: the multi-head self-attention (Multi-Head Attention) sub-layer and the position-wise feed-forward networks (Position-wise Feed-Forward Networks) sub-layer. The output of each sub-layer is processed through a residual connection and layer normalization, and the formula is as follows:
[0119] H(X) = LayerNorm(X + Sublayer(X))
[0120] where Sublayer(X) represents the function implemented by this sublayer. To simplify the calculation of the residual connection, each sublayer generates an output with a dimension of d model .
[0121] 1) Multi-Head Self-Attention Sub-layer
[0122] The multi-head self-attention sub-layer contains n head heads. First, each head performs three parallel linear transformations on the matrix input X:
[0123] Q i = XW i Q
[0124] K i = XW i K
[0125] V i = XW i V
[0126] where 1 ≤ i ≤ n head , W i Q , W i K , are linear transformation matrices, whose parameters are randomly initialized and optimized later through the backpropagation algorithm of the model.
[0127] Multi-head attention allows the self-attention network to focus on information from different representation subspaces, which helps the model learn more key features. The calculation formula for multi-head attention is as follows:
[0128]
[0129] head i = Attention(Q i , K i , V i )
[0130] where is the linear propagation parameter. Here, n head = 2 is set.
[0131] 2) Position-wise Feed-Forward Network Sub-layer
[0132] The position-wise feed-forward network sub-layer consists of two linear transformations, activated by the ReLU activation function in the middle:
[0133] FFN(x) = max(0, xW1 + b1)W2 + b2
[0134] where and are transformation parameters, where d ff = d model / 2.
[0135] 2. Predict the association relationship. Classify the encoded features through a single-layer fully connected network, convert the output matrix of the self-attention network into a probability distribution, and predict whether there is a binding site. Flatten the output matrix into a one-dimensional array, and then calculate the probability p through a linear transformation with sigmoid as the activation function. The calculation formula is as follows:
[0136] p = sigmoid(flatten(X encoded )W out + b out )
[0137] sigmoid(x) = 1 / (1 + e -x )
[0138] where the parameters of the linear transformation matrix are randomly initialized and optimized later through the backpropagation algorithm of the model.
[0139] IV. Model Training
[0140] Perform model training on the training set based on the backpropagation algorithm, use dropout to reduce overfitting, and accelerate model convergence through the adam optimizer to obtain a long non-coding RNA-disease association relationship prediction model.
[0141] 1. Reduce overfitting. Use dropout to reduce overfitting of the model, and add the dropout function at three positions of the model. In this embodiment, the dropout probability is set to p drop = 0.05.
[0142] 1) After calculating the matrix X: After calculating the matrix X, add the dropout function to process the matrix X.
[0143] dropout(X, p drop )
[0144] 2) After calculating Sublayer(X): After completing the sublayer calculation, add the dropout function to process the output.
[0145] dropout(Sublayer(X), p drop )
[0146] 3) After flattening the output matrix X encoded : After flattening the output matrix After flattening into a one-dimensional array, add a dropout function to process the output.
[0147] dropout(flatten(X encoded ), p drop )
[0148] 2. Accelerate model convergence. Use the Adam optimizer to accelerate model convergence and train the model using the cross-entropy loss function;
[0149] Loss = -∑[ylog(p) + (1 - y)log(1 - p)]
[0150] V. Predicting the association relationship
[0151] Process the target long non-coding RNA-disease node pair through the prediction model to obtain the predicted association probability p. If p > 0.5, then predict that the target node pair has an association relationship; otherwise, it does not have an association relationship.
[0152] In the specific implementation of the prediction method provided by the present invention, the process can be automatically run in a software manner. The device for running the process should also be within the protection scope of the present invention.
[0153] The beneficial effects of the present invention are verified through the following comparative experiments.
[0154] The data used in this experiment is extracted from a public dataset, including 240 long non-coding RNAs, 412 diseases, and 2697 long non-coding RNA-disease association relationships between them.
[0155] We compared the performance of the method proposed by the present invention with six typical methods in this field, namely DMFLDA (Method 1), GAMCLDA (Method 2), CNNDLP (Method 3), VADLP (Method 4), GCRFLDA (Method 5), and GAERF (Method 6). We evaluated the performance of the model through the area under the ROC curve (AUC) and the area under the PR curve (AUPR). The AUC and AUPR performance of the model are shown in Table 1.
[0156] Table 1 Results of comparative experiments
[0157]
[0158]
[0159] As can be seen from Table 1, the method of the present invention is superior to the above six methods. Specifically, the average AUC of our method is 0.994, exceeding that of DMFLDA, GAMCLDA, CNNDLP, VADLP, GCRFLDA, and GAERF by 6.4%, 7.3%, 2.5%, 3.8%, 3.5%, and 1.4% respectively; the average AUPR is 0.709, exceeding that of DMFLDA, GAMCLDA, CNNDLP, VADLP, GCRFLDA, and GAERF by 43.8%, 29.2%, 42.3%, 26.0%, 30.4%, and 21.8% respectively, indicating that the method of the present invention has stronger long non-coding RNA-disease association prediction ability.
[0160] It can be concluded that compared with the existing long non-coding RNA-disease association prediction methods, the method of the present invention has higher prediction accuracy.
[0161] In summary, the present invention designs a long non-coding RNA-disease association prediction method and system based on the self-attention mechanism, which can effectively improve the performance of predicting long non-coding RNA-disease association relationships. The research results of the present invention can be applied to the field of biomedicine. With the help of this method, researchers can guide the verification of long non-coding RNA-disease association relationships, and then conduct more in-depth research on the occurrence and development of related diseases. In addition, since the self-attention mechanism is good at learning the interdependent features between topological pathways, the research results of the present invention can not only be applied to the prediction of long non-coding RNA-disease association relationships, but also to the prediction of other regulatory factor-disease association relationships.
Claims
1. A method for predicting long non-coding RNA-disease associations based on the self-attention mechanism, characterized in that, It includes the following steps: S1: Data acquisition and preprocessing: Acquire disease semantic information and known long non-coding RNA-disease association situations, construct a heterogeneous network, and calculate the similarity between homogeneous nodes. S2: Topological feature extraction: Represent the similarity and association situation as a weighted adjacency matrix of the heterogeneous network, and according to the definition of the power of the adjacency matrix, obtain the multi-hop topological path features between target long non-coding RNA-disease node pairs. S3: Predicting the association relationship based on the self-attention network: Construct a neural network based on the self-attention mechanism, obtain feature embeddings and perform classification through a single-layer fully connected network, convert the output matrix of the neural network based on the self-attention mechanism into a probability distribution, and predict whether there is an association relationship. S4: Model training: Perform model training on the training set based on the backpropagation algorithm to obtain a long non-coding RNA-disease association prediction model. S5: Predicting the association relationship: Process the target long non-coding RNA-disease node pairs through the prediction model to determine whether there is an association relationship.
2. The long non-coding RNA-disease association prediction method based on the self-attention mechanism according to claim 1, wherein The specific content of step S1 is as follows: S11: Integrate long non-coding RNA-disease association relationships. For l long non-coding RNAs and d diseases, let A LD ∈R l×d represent their association relationship; if it is known that a long non-coding RNA-disease node pair is associated, then the corresponding position value in A LD is 1, otherwise the value is 0; S12: Calculate disease similarity. Obtain the parent-child relationship between diseases and represent it as a directed acyclic graph; for disease t1, denote T1 as the set containing t1 and its ancestor diseases, then the semantic contribution value of any disease t in T1 relative to t1 is: where children of t refers to the set of sub-diseases of disease t. Suppose Then the calculation process of the similarity between diseases t1 and t2 is as follows: where SV(t1) and SV(t2) respectively represent the sum of the disease semantic contribution values of sets T1 and T2. For the d diseases in the dataset, calculate the similarity between each pair of them to obtain the disease similarity matrix S D ∈R d×d ; S13: Calculate long non-coding RNA similarity. It is known that long non-coding RNAs r1 and r2 are respectively associated with n1 and n2 diseases, and these diseases are denoted as t 1k and t 2m , 1 ≤ k ≤ n1, 1 ≤ m ≤ n2; then the similarity calculation process between long non-coding RNAs r1 and r2 is as follows: Among them, Sim(t 1k , t 2m ) and Sim(t 2m , t 1k ) represent disease similarity; For the l long non-coding RNAs in the dataset, calculate the similarity between each pair of them to obtain the long non-coding RNA similarity matrix S L ∈R l×l 。 3. The method for predicting the association between long non-coding RNA and disease based on the self-attention mechanism according to claim 2, wherein The specific content of step S2 is as follows: S21: Integrate the weighted adjacency matrix. Integrate the long non-coding RNA-disease association matrix A LD , the long non-coding RNA similarity matrix S L and the disease similarity matrix S D into the weighted adjacency matrix A of the heterogeneous network: Among them, is the transpose of A LD ; S22: Process the multi-hop adjacency matrix. Based on the weighted adjacency matrix A, according to the following calculation process, obtain the normalized multi-hop adjacency matrix. where n h represents the hop count, and A diag0 means setting the values on the main diagonal of A to 0; max(A diag0 ) represents the maximum value in the A diag0 matrix; S23: Topological feature extraction. For the i-th node of the target long non-coding RNA - the j-th node of the disease, extract topological features through the following calculation process. Among them, and represent the i-th column and the j-th column of the n h hop adjacency matrix , is a linear transformation matrix, is the matrix input for the subsequent model, and d model is the dimension parameter of the matrix input.
4. The method for predicting the association between long non-coding RNA and disease based on the self-attention mechanism according to claim 3, wherein The specific content of step S3 is as follows: S31: Construct a neural network based on the self-attention mechanism to extract topological path features. The neural network based on the self-attention mechanism consists of N identical layers, each layer containing two sub-layers: the multi-head self-attention sub-layer and the position-wise feed-forward network sub-layer. The output of each sub-layer is processed by residual connection and layer normalization, and each sub-layer generates an output with a dimension of d model The formula for normalization is as follows: H(X) = LayerNorm(X + Sublayer(X)) where Sublayer(X) represents the function implemented by this sublayer, and X represents the matrix input of each sublayer. The dot-product-based attention includes three matrix inputs Q, K, and V, and its calculation process is: Among them, represents the channel dimension of K; and the multi-head self-attention sub-layer splits the calculation process of dot-product-based attention into multiple heads; the multi-head attention calculation formula is as follows: head i = Attention(Q i , K i , V i ) where n head represents the number of heads, 1 ≤ i ≤ n head ; is a linear transformation matrix; The three matrices of each head are obtained through the following linear transformation: Among them, are three linear transformation matrices; Q i , K i and V i are respectively three parameter matrices after linear transformation; The position feed-forward network sublayer includes two linear transformations, activated by the ReLU activation function in the middle: FFN(X) = max(0, XW1 + b1)W2 + b2 Among them, and are conversion parameters; here, d ff = d model / 2, representing the size of the intermediate layer; b1 and b2 are offset constants; S33: Predict the association situation. Flatten the output matrix of the self-attention network into a one-dimensional array, and then calculate the predicted association probability p between long non-coding RNA and disease through a linear transformation and the sigmoid activation function. The calculation formula is as follows: p = sigmoid(flatten(X encoded )W out + b out ) sigmoid(x) = 1 / (1 + e -x ) Among them, is the parameter of the linear transformation matrix, and b out is the offset constant; flatten(·) means to connect the matrix row by row from beginning to end and flatten it into a one-dimensional array.
5. The method for predicting the association between long non-coding RNA and disease based on the self-attention mechanism according to claim 4, wherein The specific content of step S4 is as follows: S41: Use dropout to reduce model overfitting. S411: After calculating matrix X, add the dropout function to process matrix X. dropout(X, p drop ) where p drop is the dropout probability; S412: After completing the sublayer calculation, add the dropout function to process the output. dropout(Sublayer(X), p drop ) Among them, Sublayer(X) represents processing the input matrix through a multi-head self-attention sublayer or a position feed-forward network sublayer; S413: After flattening the output matrix into a one-dimensional array, add a dropout function to process the output: dropout(flatten(X encoded ),p drop ) S42: Accelerate model convergence: Use the adam optimizer to accelerate model convergence and train the model using the cross-entropy loss function; Loss = -∑[ylog(p)+(1 - y)log(1 - p)] Among them, p represents the predicted correlation probability, and y represents the training label.
6. A long non-coding RNA-disease association prediction system based on the self-attention mechanism, characterized in that, It includes a processor, a memory, and a computer program stored on the memory. The computer program is executed on the processor to implement the method according to any one of claims 1 to 5.