Method and system for predicting interaction between lncRNA gene sequence and protein
By constructing a heterogeneous information network and extracting sub-graphs using meta-paths for graph convolution, feature representations are generated, and the problem of insufficient prediction accuracy and adaptability of lncRNA and protein interactions in the prior art is solved, and a more efficient prediction effect is achieved.
Patent Information
- Application Number
- CN202510117225.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-23
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
AI Technical Summary
The existing lncRNA-protein interaction prediction methods have insufficient accuracy and adaptability, especially when dealing with sparse LPI networks, the model has low biological significance and generalization ability.
A prediction method based on heterogeneous information network is proposed. By obtaining sequence data of lncRNA, miRNA and protein, calculating their similarity, constructing a heterogeneous information network, and extracting sub-graphs from the network using multiple metapaths for graph convolution operations to generate feature representations. Then the feature representation is fused and the trained prediction model is input for prediction.
It significantly improves the predictive ability of lncRNA and protein interactions, enhances the quality of node representations, and can more accurately predict the interaction between lncRNA and proteins, improving the accuracy and adaptability of predictions.
Smart Images

Figure CN120072038A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and more specifically, to a method and system for predicting the interaction between lncRNA gene sequences and proteins. Background Art
[0002] Long non-coding RNAs (lncRNAs) are a class of RNA molecules with a length exceeding 200 nucleotides. Although they do not encode proteins, they play important roles in biological processes such as gene expression regulation, cell differentiation, tumor growth, and immune response. In recent years, with the progress of high-throughput sequencing technology, researchers have been able to obtain a large amount of transcriptome and proteome data, revealing potential lncRNA-protein interactions (LPIs). However, due to traditional experimental methods such as RNA immunoprecipitation (RIP), etc., which face high costs and complex experimental procedures, it is difficult to meet the needs of large-scale LPI research. Therefore, computational methods have gradually become the main means for studying the interaction between lncRNAs and proteins.
[0003] Existing computational methods are mainly divided into three categories: similarity network-based methods, machine learning-based methods, and graph neural network (GNN)-based methods. Similarity network-based prediction methods can identify potential LPIs to a certain extent by assuming that lncRNAs with similar sequences or functions are more likely to interact with the same proteins. However, these methods usually fail to effectively utilize the topological information of the LPI heterogeneous network and have limitations in dealing with non-linear relationships. Machine learning-based prediction methods, especially deep learning methods, although able to automatically extract high-level features and minimize data noise, their performance usually depends on the encoding quality of sequence data and ignores the influence of topological information. Graph neural network (GNN) methods can utilize node attributes and the topological structure of the graph, but still face challenges in dealing with sparse LPI networks, and there are deficiencies in the biological significance and generalization ability of the model, generally having problems of low prediction accuracy and adaptability, and it is difficult to fully exert the potential of predicting the interaction between lncRNAs and proteins. Summary of the Invention
[0004] In order to improve the accuracy and adaptability of the technology for predicting the interaction between lncRNAs and proteins, the present invention proposes the following technical solutions: In the first aspect, the present invention proposes a method for predicting the interaction between lncRNA gene sequences and proteins, including: Obtaining the sequence data of lncRNAs and miRNAs, as well as protein data; Calculating the similarities of lncRNAs, miRNAs, and proteins respectively; Construct a heterogeneous information network with lncRNAs, miRNAs, and proteins as nodes and the similarity relationships among lncRNAs, miRNAs, and proteins as edges according to the similarities of lncRNAs, miRNAs, and proteins. Extract subgraphs from the heterogeneous information network according to the preset meta-paths, and perform graph convolutional operations on the subgraphs to generate feature representations of lncRNAs, miRNAs, and proteins. Perform fusion processing on the feature representations of lncRNAs, miRNAs, and proteins to obtain fused feature representations. Input the fused feature representations into a trained prediction model for prediction to obtain the prediction results of the interactions between lncRNAs and proteins.
[0005] As a preferred technical solution, before constructing the heterogeneous information network, the method further includes: Use the known interaction pairs of lncRNAs and proteins as the positive sample set , where represents an lncRNA, represents a protein; Use miRNAs as mediators to count the number of common miRNAs in the potential interaction pairs of lncRNAs and proteins to obtain the set of all potential interaction pairs , where n is a set threshold; Exclude the interaction pairs with the number of common miRNAs exceeding the threshold n from the set , and then randomly sample a certain proportion of negative samples from the remaining set to form a negative sample set N .
[0006] As a preferred technical solution, based on the Gaussian kernel function, calculate the similarities between lncRNAs and miRNAs respectively:
[0007]
[0008] where is the similarity between nodes and in the lncRNA sequence, is the similarity between nodes and in the miRNA sequence, is the parameter of the Gaussian kernel function, is the embedding vector of node in the lncRNA sequence, is the embedding vector of node The embedding vector of is the embedding vector of the node in the miRNA sequence, is the embedding vector of the node in the miRNA sequence, represents the Euclidean norm.
[0009] As a preferred technical solution, calculating the similarity of proteins includes: Obtaining the annotation information of each protein in the Gene Ontology database; For any two protein nodes, calculate the Jaccard similarity of their annotation information according to the following formula:
[0010] where and respectively represent the sets of annotation information of protein and protein in the Gene Ontology database.
[0011] As a preferred technical solution, the preset meta - path includes: The first meta - path, indicating that an lncRNA node establishes an indirect interaction with another lncRNA node through a protein node; The second meta - path, indicating that an lncRNA node establishes an indirect interaction with another lncRNA node through an miRNA node; The third meta - path, indicating a direct interaction between two lncRNA nodes based on sequence similarity; The fourth meta - path, indicating that a protein node establishes an indirect interaction with another protein node through an miRNA node; The fifth meta - path, indicating a direct interaction between miRNA nodes based on sequence similarity.
[0012] As a preferred technical solution, according to the preset meta - path, extracting a sub - graph from the heterogeneous information network and performing graph convolution operations on the sub - graph to generate the feature representations of lncRNA, miRNA, and protein, includes: According to the preset meta - path, extracting a sub - graph composed of the nodes and edges involved in the meta - path from the heterogeneous information network; Using a graph neural network to update the nodes of the sub - graph, and its expression is as follows:
[0013] In the formula, respectively represent the nodes in the meta - path and node In the feature representation of the -th layer of the graph neural network, is the set of neighbor nodes of node , is the adjacency matrix of the subgraph, representing the connection weight between node and node , indicating the connection between nodes; and are the number of edges connected to node and respectively, is the weight matrix of the -th layer; Fuse the node feature representations guided by different meta-paths through the semantic attention mechanism to obtain the fusion weight , and its expression is as follows:
[0014] where, is the weight matrix to be trained, represents the node feature representation under the meta-path , is the activation function; Normalize the fusion weight , and its expression is as follows:
[0015] Perform weighted summation on the node feature representations of different meta-paths to obtain the final embedded feature representation of different nodes: .
[0016] As a preferred technical solution, fuse the feature representations of lncRNA, miRNA, and protein to obtain the fusion embedded feature representation, including: Obtain the interaction matrix between lncRNA and miRNA, and the interaction matrix between protein and miRNA; Generate the query matrix, key matrix, and value matrix of the multi-head attention mechanism using the feature representations of lncRNA, miRNA, and protein; Calculate the interaction weight between lncRNA and protein according to the query matrix, key matrix, and value matrix, and its expression is as follows:
[0017] where, represents the softmax function, is the query matrix of lncRNA, is the key matrix of miRNA, is the key matrix of the protein, is the scaling factor, is the bias term; Calculate the context feature representations of lncRNAs and the context feature representations of proteins respectively:
[0018]
[0019] In the formula, is the query matrix of the protein, is the value matrix of miRNA; According to the interaction weights between lncRNAs and proteins, the context feature representations of lncRNAs and the context feature representations of proteins are fused to obtain the fused context feature representations and the fused context feature representations , and their expressions are as follows:
[0020]
[0021] The fused context feature representations and the fused context feature representations are respectively concatenated with the original lncRNA and protein representations to obtain the final fused embedding feature representations L and the fused embedding feature representations P , and their expressions are as follows:
[0022] .
[0023] As a preferred technical solution, the fused feature representations are input into a trained prediction model for prediction to obtain the prediction results of the interaction between lncRNAs and proteins, including: The fused embedding feature representations L and the fused embedding feature representations P are concatenated into a feature representation vector :
[0024] where, represents the concatenation operation; Using ELU as the activation function, the feature representation vector is linearly transformed to obtain the output feature :
[0025] Among them, is the weight matrix, is the bias vector, is the activation function; The output feature is passed through a linear layer and a LogSoftmax operation to obtain the predicted probability distribution , and its expression is as follows:
[0026] Among them, is the weight matrix, is the bias vector, is used to convert the output feature into a probability distribution, where each value represents the probability of whether the lncRNA and the protein interact.
[0027] As a preferred technical solution, before inputting the fused feature representation into the trained prediction model for prediction, the method further includes training and optimizing the prediction model using a comprehensive loss function; The expression of the comprehensive loss function is as follows:
[0028] Among them, is the feature representation learning loss of lncRNA, is the feature representation learning loss of miRNA, is the feature representation learning loss of protein, is the sample weighted loss of the gradient harmonization mechanism, is the training sample set, is the sample i and the sample j negative log-likelihood loss, is the sample i and the sample j weights; , , and are the weight coefficients of each loss respectively.
[0029] In a second aspect, the present invention also proposes a prediction system for the interaction between lncRNA gene sequences and proteins, which is applied to the prediction method for the interaction between lncRNA gene sequences and proteins according to any one of the schemes in the first aspect, and includes: An acquisition module, configured to acquire the sequence data of lncRNA and miRNA, as well as protein data; A calculation module for calculating the similarities of lncRNA, miRNA, and protein respectively; A construction module for constructing a heterogeneous information network with lncRNA, miRNA, and protein as nodes and the similarity relationships between lncRNA, miRNA, and protein as edges according to the similarities of lncRNA, miRNA, and protein; A generation module for extracting subgraphs from the heterogeneous information network according to preset meta-paths and performing graph convolution operations on the subgraphs to generate feature representations of lncRNA, miRNA, and protein; A fusion module for performing fusion processing on the feature representations of lncRNA, miRNA, and protein to obtain a fused feature representation; A prediction module for inputting the fused feature representation into a trained prediction model for prediction to obtain a prediction result of the interaction between lncRNA and protein.
[0030] In a third aspect, the present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. Wherein, when the processor executes the program, it implements the operations performed by the building complex power load prediction method based on the multi-modal deep forest algorithm according to any one of the solutions in the first aspect.
[0031] The beneficial effects of the present invention at least include: The present invention calculates the similarities between lncRNA, miRNA, and protein, and uses multiple meta-paths to extract subgraphs from the heterogeneous network for graph convolution operations to generate feature representations, enhancing the ability to capture the relationships between different biomolecules. Through miRNA as the intermediary information, it can effectively fuse the connections between lncRNA and protein with miRNA, thereby improving the quality of node representations. By optimizing the utilization of the heterogeneous information network, the present invention can significantly improve the prediction ability when dealing with sparse data. At the same time, through contrastive learning between meta-paths, it reduces the information differences between different subgraphs and alleviates the negative impact of information imbalance on the prediction results. The adaptability of the present invention on diverse datasets is enhanced, and it can more accurately predict the interaction between lncRNA and protein, effectively improving the accuracy of predicting the interaction between lncRNA and protein, thus providing effective support for related biological research and applications, such as molecular breeding, disease research, drug development, etc. Description of the Drawings
[0032] Figure 1 It is a schematic flowchart of the method for predicting the interaction between lncRNA gene sequence and protein provided by the embodiment of the present invention.
[0033] Figure 2This is the architecture diagram of the lncRNA gene sequence and protein interaction prediction system provided by the embodiments of the present invention.
[0034] Figure 3 This is the structural schematic diagram of the electronic device provided by the embodiments of the present invention. Detailed implementation manners
[0035] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred technical solutions. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred technical solutions are only for illustrating the present invention rather than for limiting the protection scope of the present invention.
[0036] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, number, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0037] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0038] Embodiment 1 This embodiment proposes a method for predicting the interaction between lncRNA gene sequences and proteins. As Figure 1 shown, Figure 1 This is the flowchart of a method for predicting the interaction between lncRNA gene sequences and proteins provided by this embodiment. The method includes the following steps: S1: Obtain the sequence data of lncRNA and miRNA, as well as protein data; S2: Calculate the similarities of lncRNA, miRNA, and proteins respectively; S3: According to the similarities of lncRNA, miRNA, and proteins, construct a heterogeneous information network with lncRNA, miRNA, and proteins as nodes and the similarity relationships between lncRNA, miRNA, and proteins as edges. S4: Extract subgraphs from the heterogeneous information network according to the preset meta-paths, and perform graph convolution operations on the subgraphs to generate feature representations of lncRNAs, miRNAs, and proteins; S5: Perform fusion processing on the feature representations of lncRNAs, miRNAs, and proteins to obtain fused feature representations; S6: Input the fused feature representations into the trained prediction model for prediction to obtain the prediction results of the interactions between lncRNAs and proteins.
[0039] It can be understood that by calculating the similarities between lncRNAs, miRNAs, and proteins, and using multiple meta-paths to extract subgraphs from the heterogeneous network for graph convolution operations to generate feature representations, the ability to capture the relationships between different biomolecules is enhanced. Through miRNAs as intermediate information, the connections between lncRNAs and proteins with miRNAs can be effectively fused, thereby improving the quality of node representations. By optimizing the utilization of the heterogeneous information network, the present invention can significantly enhance the prediction ability when dealing with sparse data. At the same time, through contrastive learning between meta-paths, the information differences between different subgraphs are reduced, and the negative impact of information imbalance on the prediction results is alleviated. The adaptability of the present invention on diverse datasets is enhanced, and it can more accurately predict the interactions between lncRNAs and proteins, effectively improving the accuracy of predicting lncRNA-protein interactions, thereby providing effective support for related biological research and applications, such as molecular breeding, disease research, drug development, etc.
[0040] Example 2 This example makes improvements on the basis of the method for predicting the interaction between lncRNA gene sequences and proteins proposed in Example 1.
[0041] In this example, after collecting lncRNA, miRNA sequence data, and protein data from the database, the k-mer slicing method is used to segment the lncRNA sequences and miRNA sequences. Among them, 3-mer is used for lncRNA sequences, and 2-mer is used for miRNA sequences. At the same time, the Word2Vec model is used to train the sliced sequences, with some parameters being vector_size = 300, min_count = 3, epoch = 100, and the rest using default parameters.
[0042] In this embodiment, since most of the existing databases only record positive samples, i.e., interaction data, for other unrecorded pairs, they are potential interaction samples, that is, it is unknown whether they interact. Compared with the number of all pairs, the positive and negative samples are unbalanced. This embodiment adopts a negative sample sampling strategy. First, all known lncRNA-protein interactions are regarded as positive samples. Then, using miRNA as an intermediary, select those lncRNA-protein pairs with common miRNAs as potential positive samples, and exclude such samples first in the screening of negative samples. By excluding lncRNA-protein pairs whose common miRNA number exceeds a certain threshold for sampling. Finally, randomly select a certain proportion of negative samples to ensure the balance of the number of positive and negative samples in the training set. Specifically: Take the interaction pairs of known lncRNAs and proteins as the positive sample set , where represents lncRNA, represents protein; Use miRNA as an intermediary to count the number of common miRNAs in potential lncRNA and protein interaction pairs to obtain the set of all potential interaction pairs , where n is the set threshold; Exclude interaction pairs with a common miRNA number exceeding the threshold n from the set , and then randomly sample a certain proportion of negative samples from the remaining set to form the negative sample set N .
[0043] Before model training, use 5-fold cross-validation to divide the data into a training set and a test set. In each iteration, select a part of it as the test set, and the rest as the training set to ensure the reliability and generalization ability of the results. This embodiment adopts an 8:2 ratio to divide the training set and the test set.
[0044] In this embodiment, in order to capture the similarities between lncRNA, miRNA, and protein, similarity matrices based on the Gaussian kernel function are constructed for lncRNA and miRNA sequences respectively. Based on the Gaussian kernel function, the similarities between lncRNA and miRNA are calculated respectively, and their expressions are as follows:
[0045]
[0046] Among them, represents the node in the lncRNA sequence and node Similarity of indicating nodes in the miRNA sequence and nodes Similarity of is the parameter of the Gaussian kernel function is the embedding vector of node in the lncRNA sequence is the embedding vector of node in the lncRNA sequence is the embedding vector of node in the miRNA sequence is the embedding vector of node in the miRNA sequence represents the Euclidean norm. Based on these similarities, in this embodiment, a K-Nearest Neighbor (KNN) graph is constructed to represent the similarity relationship between lncRNA and miRNA. For lncRNA nodes, the 10 most similar nodes are saved, and for miRNA nodes, the 20 most similar nodes are saved.
[0047] In this embodiment, calculating the similarity of proteins includes: Obtaining the annotation information of each protein in the Gene Ontology database; For any two protein nodes, according to the following formula, calculate the Jaccard similarity of their annotation information:
[0048] where and respectively represent the sets of annotation information of protein and protein in the Gene Ontology database.
[0049] In this embodiment, the design of the meta-path is to reflect the multiple interaction relationships between nodes. Considering both node types and edge types, in the heterogeneous information network of lncRNA, miRNA, and proteins, the preset meta-paths include: The first meta-path: , indicating that an lncRNA node establishes an indirect interaction with another lncRNA node through a protein node; The second meta-path: , indicating that an lncRNA node establishes an indirect interaction with another lncRNA node through a miRNA node; The third meta-path: , indicating a direct interaction between two lncRNA nodes based on sequence similarity; The fourth meta-path: indicates that a protein node establishes an indirect interaction with another protein node through an miRNA node; The fifth meta-path: indicates a direct interaction between miRNA nodes based on sequence similarity.
[0050] In this embodiment, meta-path convolution is learned through a Graph Convolutional Network (GCN) on the subgraphs induced by the meta-paths. For each meta-path, subgraphs in the heterogeneous graph are first extracted, and these subgraphs are composed of the nodes and edges involved in the path. Then, graph convolution is used to perform convolution operations on these subgraphs in order to learn node representations from the local structure.
[0051] Specifically, for a set of meta-paths corresponding adjacency matrix sets can be extracted from the heterogeneous information network where each adjacency matrix represents the topological structure of the subgraph corresponding to the meta-path According to the preset meta-paths, subgraphs are extracted from the heterogeneous information network, and graph convolution operations are performed on the subgraphs to generate feature representations of lncRNAs, miRNAs, and proteins, including: According to the preset meta-paths, subgraphs composed of the nodes and edges involved in the meta-paths are extracted from the heterogeneous information network; Use a graph neural network to update the nodes of the subgraph, and its expression is as follows:
[0052] In the formula, respectively represent the feature representations of nodes and node in the -th layer of the graph neural network, is the set of neighbor nodes of node for node , is the adjacency matrix of the subgraph, representing the connection weight between node and node , indicating the connection between nodes; and are respectively the number of edges connected to node and , is the weight matrix of the -th layer; After different subgraphs guided by meta-paths pass through the Graph Convolutional Network (GCN), a series of node representations will be generated. These representations capture the features of nodes in different meta-path views. To integrate this information and enable the model to learn the complex relationships between lncRNAs, miRNAs, and proteins more comprehensively, it is necessary to fuse the features of these different views to obtain a unified node representation. Feature fusion is achieved through a semantic attention mechanism. Taking proteins as an example, we get 。
[0053] Fuse the node feature representations guided by different meta-paths through a semantic attention mechanism to obtain the fusion weights ,whose expression is as follows:
[0054] where is the weight matrix to be trained, represents the node feature representation under the meta-path , is the activation function; Normalize the fusion weights , and its expression is as follows:
[0055] Perform weighted summation on the node feature representations of different meta-paths to obtain the final embedded feature representations of different nodes : 。
[0056] In this embodiment, contrastive learning is performed between meta-paths. In the contrastive learning between meta-paths, the goal is to maximize the mutual information between different meta-path views, specifically as follows: For each node, its corresponding representation in other views is regarded as a positive sample. Construct a K-Nearest Neighbor (KNN graph). First, use the KNN algorithm to construct a graph based on node similarity to identify the most similar nodes. For each node, find its most similar neighbors in different views and use these neighbors to construct positive and negative sample sets. All nodes not in the same KNN subgraph are regarded as negative samples.
[0057] Use cosine similarity to calculate the similarity between nodes. For the representations of node in view and view , denoted as and , the similarity is calculated as follows:
[0058] where and respectively represent the node in the view and embedding representations, represents the norm of the vector.
[0059] To maximize the similarity of the same nodes in different views while minimizing the similarity between different nodes, a noise contrastive estimation (NCE) loss is used. The specific formula is as follows:
[0060] where: is the loss function for cross-view contrastive learning, , represents different meta-path views, , represents the nodes in the graph, represents the cosine similarity, is the temperature parameter used to adjust the sensitivity of the contrastive loss.
[0061] In this embodiment, the lncRNA-miRNA (LMI) and protein-miRNA (PMI) interaction matrices are used as attention masks. In the multi-head attention mechanism, the attention mask is used to ensure that only known miRNA interaction pairs contribute to the attention scores, without affecting the noise information during model training. Specifically, a mask is used to filter out irrelevant miRNAs, thus ensuring that the attention mechanism focuses on the miRNAs important for the interaction between lncRNA and proteins. The feature representations of lncRNA, miRNA, and protein are fused to obtain the fused embedding feature representation, including: Obtain the interaction matrix between lncRNA and miRNA, and the interaction matrix between protein and miRNA; Using the feature representations of lncRNA, miRNA, and protein, generate the query matrix, key matrix, and value matrix of the multi-head attention mechanism:
[0062] where, , and represent the query, key, and value matrices respectively, and these matrices are generated from the representations of lncRNA, miRNA, and protein. is the dimension of the key vector. It is a matrix used to represent the attention mask, ensuring that attention is only applied to relevant miRNAs. In the mask matrix, all non-interactive positions are set to a large negative number to ensure that these positions are effectively ignored in the softmax operation, thus ensuring that the model only focuses on the actual miRNA interaction information.
[0063] Based on the query matrix, key matrix, and value matrix, calculate the interaction weights between lncRNAs and proteins, and its expression is as follows:
[0064] where represents the softmax function, is the query matrix of lncRNA, is the key matrix of miRNA, is the key matrix of protein, is the scaling factor, is the bias term; Calculate the context feature representations of lncRNA and protein context feature representation respectively:
[0065]
[0066] In the formula, is the query matrix of protein, is the value matrix of miRNA; According to the interaction weights between lncRNA and protein, fuse the context feature representation of lncRNA and protein context feature representation to obtain the fused context feature representation and fused context feature representation , and its expression is as follows:
[0067]
[0068] Concatenate the fused context feature representation and fused context feature representation with the original lncRNA and protein representations respectively to obtain the final fused embedding feature representations L and fused embedding feature representation P , and its expression is as follows:
[0069] 。
[0070] In this embodiment, the prediction model is a multi-layer perceptron (MLP). The structure of the MLP includes two linear layers. The first layer is used for feature transformation and dimensionality reduction, and the second layer is used for classification. A non-linear activation function is added between layers to enhance the expressive power of the model. The fused feature representation is input into the trained prediction model for prediction to obtain the prediction results of the interaction between lncRNA and protein, including: The fused embedding feature representation L and the fused embedding feature representation P are concatenated into a feature representation vector :
[0071] where, represents the concatenation operation; The ELU is used as the activation function to perform a linear transformation on the feature representation vector to obtain the output feature :
[0072] where, is the weight matrix, is the bias vector, is the activation function; The output feature passes through a linear layer and a LogSoftmax operation to obtain the predicted probability distribution , and its expression is as follows:
[0073] where, is the weight matrix, is the bias vector, is used to convert the output feature into a probability distribution, where each value represents the probability of whether lncRNA and protein interact.
[0074] Through the output of the MLP, the predicted probability of whether there is an interaction between the lncRNA and protein pair can be obtained. It is a one-dimensional vector of length 2. The first value represents the probability of the existence of an interaction, and the second value represents the probability of the absence of an interaction.
[0075] In this embodiment, since there is usually an imbalance in the number of positive and negative samples in the interaction data between lncRNA and protein, this may affect the performance of the model. To address this imbalance, a gradient harmonizing mechanism (GHMC) is introduced into the loss function to dynamically adjust the weights of positive and negative samples, thereby ensuring that the model can effectively handle the imbalance problem during training.
[0076] GHMC (Gradient Harmonizing Mechanism-C) first calculates the gradient value of each sample. The magnitude of the gradient is calculated using the formula , where is the predicted probability of the model for the pairing of lncRNA and protein, is the true label, and the gradient value represents the deviation of the model's prediction for a certain pairing. Next, all samples are binned according to their gradient values. The boundaries of these bins are defined by the formula , where is the number of partitions, which is used to control the fineness of the gradient binning. Then, weights are assigned to each sample according to the number of samples in each bin. The weights are calculated as, , where represents the total number of samples, represents the number of samples in bin , is a very small constant used to avoid division by zero, which ensures that rare gradients (i.e., difficult-to-learn samples) are given greater weights. is updated according to the formula , where is the momentum term, which is used to smooth the update of the current and previous gradient distributions. The final GHMC loss function is
[0077] where,
[0078] Here is the negative log-likelihood loss. After combining with the weight , the model pays more attention to difficult-to-learn samples during training, especially in the case of imbalance between positive and negative samples, improving the learning effect and prediction ability of the model.
[0079] In summary, the prediction model is trained and optimized using the comprehensive loss function. The expression of the comprehensive loss function is as follows:
[0080] where, is the feature representation learning loss of lncRNA, is the feature representation learning loss of miRNA, is the feature representation learning loss of the protein, is the sample weighted loss of the gradient harmonization mechanism, is the training sample set, is the sample i and the sample j negative log-likelihood loss, is the sample i and the sample j weights; , , and are the weight coefficients of each loss respectively.
[0081] In this embodiment, in order to enable the model to dynamically adjust the weights of each part of the loss according to the contribution of different loss terms to the training effect during the training process, the dynamic weight averaging (DWA) method is adopted in this embodiment. The DWA method adaptively adjusts the weights by comparing the changes of different task losses at two consecutive time steps.
[0082] The formula for calculating the weights by DWA is as follows:
[0083] In the formula, and respectively represent the values of the loss term in the and the iteration. is the weight ratio of the th task at the time step , calculated by the loss changes at two consecutive time steps. is the temperature parameter, used to control the sensitivity of weight adjustment, is a constant, used to normalize the weights. is the final weight of the th task at the time step , normalized using the softmax function to ensure that the sum of all weights is 1.
[0084] It can be understood that many existing methods often only rely on a single network relationship to learn the representations of lncRNAs and proteins, but seriously neglect the crucial role played by miRNAs as mediators in complex biological interaction processes. The present invention ingeniously integrates the information of common miRNAs and applies it as key context information to the embedded representations of lncRNA and protein pairing. This unique information integration method can automatically focus on the important miRNAs closely related to the interaction between lncRNAs and proteins, thereby strongly enhancing the prediction ability of the model. In this way, the present invention successfully achieves a more accurate modeling of the potential biological mechanism of interaction, greatly improving the model's ability to capture real interaction relationships and making its prediction effect in this field even better.
[0085] The present invention also employs a cross-view contrastive learning strategy. This strategy maximizes the representational similarity of the same node under different meta-path views, effectively ensuring that the model can always maintain the consistency and complementarity of node representations in different meta-paths. Traditional prediction methods usually only use a single view to construct the model, completely ignoring the rich diversity of information contained in different views and their mutual complementarity. The present invention uses contrastive learning between meta-paths, enabling the model to organically integrate diverse information from different meta-paths, thereby generating richer and more accurate node representations. In addition, the noise contrastive estimation (NCE) loss adopted in the contrastive learning can ensure that the model efficiently learns information from sparse network structures, and thus can still exhibit excellent performance even in the case of severe imbalance between positive and negative samples.
[0086] The present invention also makes full use of deep learning techniques. Compared with the original model in the reference solution, although the original model also introduced miRNA as an intermediate substance in an attempt to improve the prediction ability, it did not use advanced technologies such as deep learning to predict interactions. In contrast, the present invention uses a heterogeneous information network. Specifically, by carefully constructing a heterogeneous information network, the similarity network and the interaction network are cleverly combined, and graph convolutional training is carried out with the help of meta-path-guided subgraphs. These subgraphs cover different types of nodes and the intricate relationships between them. By performing representation learning on these meta-path-guided subgraphs through GCN (Graph Convolutional Network), the complex interaction information between nodes can be accurately captured. This method can make better use of various biological information compared with the traditional calculation method based on simple network relationships, significantly improving the prediction accuracy of the interaction between lncRNA and protein. At the same time, deep learning methods such as the multi-head attention mechanism and cross-view contrast learning are combined to further enhance the feature expression ability, enabling the model to still maintain high prediction performance and strong robustness when facing complex and sparse biological networks, providing a more solid and reliable technical support and theoretical basis for the research and application in this field.
[0087] As an exemplary illustration, the dataset of this embodiment is sourced from the RAIDv2.0 database. As a comprehensive resource, this database covers RNA-related interaction information across multiple organisms. When processing it, through filtering operations, lncRNAs and proteins not associated with miRNA are excluded, and finally 1,097 lncRNAs, 144 proteins, and 2,132 miRNAs are obtained. The interaction information contained in the dataset is specifically 33,130 lncRNA-miRNA interaction instances, 2,464 protein-miRNA interaction instances, and 2,763 lncRNA-protein interaction instances, and all these interactions are reliable results verified through experiments.
[0088] The model conducts training work in a supervised learning manner, and the samples involved include both positive and negative samples. Given that the database only provides positive samples and lacks corresponding negative samples, great care must be taken during the sampling process of negative instances. The specific steps for sampling negative samples are as follows: First step, clearly identify all the recorded interactions in the document as positive samples; Second step, exclude lncRNA and protein pairs sharing the same miRNA and the already determined positive samples; Third step, randomly sample negative samples in different proportions to ensure that the dataset used for model training has good balance and representativeness in the distribution of positive and negative samples, providing a solid data foundation for the effective training of the model.
[0089] In terms of sequence information extraction, the k-mer segmentation technique is adopted to extract the sequence information of lncRNA and miRNA. For the relatively long lncRNA sequences, a 3-mer segmentation method is selected; while for the shorter miRNA sequences, a 2-mer segmentation method is used. Subsequently, the Word2Vec technique is employed to embed and transform the data segmented by k-mer into a 180-dimensional vector representation form, providing standardized and informative basic data for subsequent vector-based data analysis and processing.
[0090] In the similarity network construction stage, for each lncRNA and miRNA, the Gaussian kernel function is used to calculate the similarity values between them, and a K-Nearest Neighbor (KNN) graph is constructed based on these values to visually represent the similarity associations between lncRNA and miRNA. For proteins, with the help of GO annotation information from UniProt, the GO similarity between proteins is calculated through the Jaccard coefficient, and then a protein similarity graph is constructed to clearly present the similarity characteristics between proteins.
[0091] Integrate the similarity networks constructed above to form a heterogeneous network containing three different node types: lncRNA, miRNA, and protein. This heterogeneous network completely covers the known lncRNA-miRNA, lncRNA-protein, and protein-miRNA interaction relationships. To effectively extract valuable subgraphs from this heterogeneous network for in-depth learning, a series of different types of meta-paths are specifically designed, and through these meta-paths, the complex relationships between nodes in the heterogeneous network are systematically explored. These meta-paths specifically include the interaction meta-path between lncRNA and protein, the interaction meta-path between lncRNA and miRNA, and the self-loop meta-path based on similarity, etc. By reasonably using these meta-paths, multiple subgraphs with specific meanings are accurately defined in the heterogeneous network, providing a clear structural framework and path guidance for subsequent node representation learning.
[0092] On the meta-path-guided subgraphs, the Graph Convolutional Network (GCN) is used to deeply learn the representation of nodes. For each meta-path, a corresponding set of adjacency matrices can be obtained, and a single-layer GCN is used to process these sets of adjacency matrices to effectively capture the topological information in different views. After obtaining the node features under different meta-paths, the semantic attention mechanism is used to perform weighted aggregation operations on the features of different meta-paths, organically integrating the information contained in each meta-path, and finally generating the final node representation that can comprehensively and accurately reflect the node characteristics.
[0093] To further enhance the representation learning effect of lncRNA and protein interactions, a miRNA information fusion module based on the multi-head attention mechanism was specifically designed. By carefully calculating the attention weights of each miRNA, this module can efficiently integrate miRNA information related to lncRNA and protein. Through this mechanism, the model can automatically identify and focus on important miRNAs that have a key impact on lncRNA and protein interactions, thereby significantly improving the modeling ability of interaction relationships and providing strong technical support for revealing the internal mechanisms of biomolecular interactions.
[0094] In addition, a cross-view contrast learning strategy was introduced. By maximizing the mutual information of the same nodes under different meta-path views, it is ensured that the model can maintain the consistency and complementarity of node representations when facing different meta-path views. In this way, the model can comprehensively and deeply learn the information in the heterogeneous network from multiple perspectives, thus more accurately grasping the complex relationships between nodes. After completing the node representation learning, a multi-layer perceptron (MLP) is used to predict the interaction probability between lncRNA and protein. Considering the problem of imbalance between positive and negative samples in the training data, the GHMC loss function is adopted to optimize the training process of the model, effectively improving the prediction accuracy of the model in the case of imbalance between positive and negative samples.
[0095] Table 1 Performance of the present invention under different ratios of positive and negative samples and different common miRNA thresholds
[0096] Table 2 Performance metrics of the present invention after removing miRNA information fusion and cross-view contrast learning
[0097] Table 3 Performance comparison of the present invention with other models
[0098] To comprehensively verify the effectiveness of the model, a five-fold cross-validation experiment was strictly carried out on the above-mentioned carefully prepared dataset. Throughout the experiment, the performance of the model was meticulously and comprehensively evaluated and compared using a variety of performance metrics, including key metrics such as the area under the receiver operating characteristic curve (AUROC), the area under the precision-recall curve (AUPRC), and F1-Score. The relevant experimental results are shown in Table 1.
[0099] It can be clearly seen from Table 1 that under different positive-negative sample ratios (such as 1:1, 1:3, 1:5, 1:7, etc.) and different common miRNA thresholds (3, 5, 7, 9, etc.), the performance indicators of the model show a certain trend of change. For example, when the positive-negative sample ratio is 1:1 and the common miRNA threshold is 3, the AUROC reaches 0.9919, the AUPR is 0.9902, and the score is 0.9713, indicating that the model has high performance under this condition. With the change of the positive-negative sample ratio and the common miRNA threshold, although the performance indicators of the model fluctuate, they still remain at a relatively high level on the whole. This shows that the model has a certain stability and adaptability under different parameter settings and can provide relatively reliable prediction results under various conditions.
[0100] Table 2 focuses on showing the changes in the performance indicators of the present invention under the removal of miRNA information fusion and cross-view contrast learning. Among them, "Nofusionmodule" represents the removal of miRNA information fusion, at this time the AUROC is 0.9885, the AUPR is 0.9578, and the F1-Score is 0.9267; "Nocontrastivelearning" represents the removal of cross-view contrast learning, and the corresponding AUROC is 0.9916, the AUPR is 0.9709, and the F1-Score is 0.9340; "withoutboth" is to remove both of these means at the same time, and its performance indicators are AUROC 0.9876, AUPR 0.9539, and F1-Score 0.9251; while "Fullmodel" is the complete model that uses miRNA fusion and contrast learning, and its AUROC is 0.9926, the AUPR is 0.9756, and the F1-Score is 0.9383. Through comparison, it can be clearly found that the complete model is significantly better than the versions that remove these means in all the set performance indicators. This result fully and strongly proves that the means of miRNA information fusion and cross-view contrast learning play a crucial role in improving the model performance. Their existence significantly enhances the model's prediction ability and accuracy for RNA-related interactions, and effectively demonstrates the effectiveness and superiority of the designed innovative means and strategies.
[0101] Meanwhile, to further deeply verify the performance advantages of this model compared with other models in this context, three of the most advanced models in the current field were carefully selected, and a rigorous five-fold cross-validation experiment was also carried out under the same dataset. After rigorous and meticulous experimental comparison and in-depth data analysis, the specific experimental results are shown in Table 3. In Table 3, the comparison of this invention (i.e., the proposed model) with other models (LPICGAE, LPIGAE, CCGNN, LPIDF, BiHo-GNN) in various performance indicators can be seen. This invention reached 0.9944, 0.9823, 0.9498, 0.9783, 0.9121 respectively in indicators such as AUC, AUPR, F1-score, Recall, Precision, while the values of other models in these indicators were all lower than this invention. For example, the AUC of the LPICGAE model was 0.9763, the AUPR was 0.9496, the F1-score was only 0.7862, the Recall was 0.8437, and the Precision was 0.6601, showing an obvious gap compared with this invention. This comparison result clearly and strongly indicates that this invention has significant performance advantages in the field of RNA-related interaction prediction. Whether it is the ability to distinguish positive and negative samples (AUC, AUPR), or the comprehensive prediction accuracy (F1-score), as well as the recall ability for positive samples (Recall) and the precision of prediction (Precision), etc., it is far superior to other existing advanced models. Through this series of comprehensive and in-depth experimental verifications, the advancement, practicality, and strong competitiveness of this invention in the field of RNA-related interaction prediction are fully demonstrated, providing a more efficient and accurate model and method for the research and application in this field.
[0102] Example 3 As Figure 2 shown, this embodiment proposes a prediction system for the interaction between lncRNA gene sequences and proteins, which is applied to the method for predicting the interaction between lncRNA gene sequences and proteins as described in the above embodiment, and includes: an acquisition module 100, a calculation module 200, a construction module 300, a generation module 400, a fusion module 500, and a prediction module 600.
[0103] Among them, the acquisition module 100 is used to acquire the sequence data of lncRNA and miRNA, as well as protein data; the calculation module 200 is used to calculate the similarities of lncRNA, miRNA, and protein respectively; the construction module 300 is used to construct a heterogeneous information network with lncRNA, miRNA, and protein as nodes and the similarity relationships among lncRNA, miRNA, and protein as edges according to the similarities of lncRNA, miRNA, and protein; the generation module 400 is used to extract subgraphs from the heterogeneous information network according to preset meta-paths and perform graph convolution operations on the subgraphs to generate the feature representations of lncRNA, miRNA, and protein; the fusion module 500 is used to perform fusion processing on the feature representations of lncRNA, miRNA, and protein to obtain a fused feature representation; the prediction module 600 is used to input the fused feature representation into a trained prediction model for prediction to obtain the prediction result of the interaction between lncRNA and protein.
[0104] It should be noted that the foregoing explanation of the embodiment of the method for predicting the interaction between lncRNA gene sequence and protein also applies to the system for predicting the interaction between lncRNA gene sequence and protein in this embodiment, and will not be elaborated here.
[0105] Embodiment 4 Figure 3 The following is a schematic structural diagram of the electronic device 700 provided in this embodiment. The electronic device 700 includes: a memory 701, a processor 702, and a computer program stored on the memory 701 and executable on the processor 702.
[0106] When the processor 702 executes the program, it implements the method for predicting the power load of a building complex based on the multi-modal deep forest algorithm provided in the foregoing embodiment.
[0107] Furthermore, the electronic device 700 further includes: a communication interface 703, which is used for communication between the memory 701 and the processor 702.
[0108] The memory 701 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0109] If the memory 701, the processor 702, and the communication interface 703 are implemented independently, the communication interface 703, the memory 701, and the processor 702 can be interconnected via a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 only a thick line is used in Figure 3 , but it does not mean that there is only one bus or one type of bus.
[0110] Optionally, in a specific implementation, if the memory 701, the processor 702, and the communication interface 703 are integrated on a single chip, the memory 701, the processor 702, and the communication interface 703 can communicate with each other via an internal interface.
[0111] The processor 702 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0112] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or N embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0113] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0114] Any process or method description depicted in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code that includes one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present invention includes additional implementations where functions may be performed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0115] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, and the like.
[0116] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried out in the method of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0117] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not intended to limit the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the embodiments here. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for predicting the interaction between lncRNA gene sequence and protein, characterized in that: include: Obtain lncRNA and miRNA sequence data, as well as protein data; The similarities of lncRNA, miRNA, and protein were calculated separately; According to the similarity of lncRNA, miRNA and protein, a heterogeneous information network is constructed with lncRNA, miRNA and protein as nodes and the similarity relationship between lncRNA, miRNA and protein as edges; According to the preset meta-path, subgraphs are extracted from the heterogeneous information network and graph convolution operations are performed on the subgraphs to generate feature representations of lncRNA, miRNA and protein; The feature representations of lncRNA, miRNA and protein are fused to obtain a fused feature representation; The fusion feature representation is input into the trained prediction model for prediction to obtain the prediction results of the interaction between lncRNA and protein.
2. The method for predicting the interaction between lncRNA gene sequence and protein according to claim 1, characterized in that: Before constructing the heterogeneous information network, the method further includes: The known lncRNA and protein interaction pairs are used as positive sample sets ,in Indicates lncRNA, Indicates protein; Using miRNA as an intermediary, we counted the number of common miRNAs in potential lncRNA-protein interaction pairs. , get the set of all potential interaction pairs , where n is the set threshold; From the collection Exclude the interacting pairs whose number of common miRNAs exceeds the threshold n, and then select the remaining pairs from the set Randomly sample a certain proportion of negative samples to form a negative sample set N .
3. The method for predicting the interaction between lncRNA gene sequence and protein according to claim 1, characterized in that: Based on the Gaussian kernel function, the similarity of lncRNA and miRNA was calculated separately: in, To represent the nodes in the lncRNA sequence and nodes The similarity of Represents a node in the miRNA sequence and nodes The similarity of is the parameter of the Gaussian kernel function, Node in the lncRNA sequence The embedding vector of Node in the lncRNA sequence The embedding vector of Node in miRNA sequence The embedding vector of Node in miRNA sequence The embedding vector of represents the Euclidean norm.
4. The method for predicting the interaction between lncRNA gene sequence and protein according to claim 1, characterized in that: Calculate protein similarity, including: Obtain annotation information of each protein in the Gene Ontology database; For any two protein nodes, the Jaccard similarity of their annotation information is calculated according to the following formula: in, and Represents protein and protein A collection of annotation information in the Gene Ontology database.
5. The method for predicting the interaction between lncRNA gene sequence and protein according to claim 1, characterized in that: The preset meta-path includes: The first-element path indicates that a lncRNA node establishes an indirect interaction with another lncRNA node through a protein node; The second meta-path indicates that the lncRNA node establishes an indirect interaction with another lncRNA node through the miRNA node; The third-element path represents the direct interaction between two lncRNA nodes based on sequence similarity; The fourth-element path indicates that a protein node establishes an indirect interaction with another protein node through a miRNA node; The fifth meta-path represents the direct interaction between miRNA nodes based on sequence similarity.
6. The method for predicting the interaction between lncRNA gene sequence and protein according to claim 1, characterized in that: According to the preset meta-path, subgraphs are extracted from the heterogeneous information network, and graph convolution operations are performed on the subgraphs to generate feature representations of lncRNA, miRNA, and protein, including: According to the preset meta-path, a sub-graph consisting of nodes and edges involved in the meta-path is extracted from the heterogeneous information network; The graph neural network is used to update the nodes of the subgraph. The expression is as follows: In the formula, Represents the meta path Midpoint and nodes In the graph neural network The feature representation of the layer, For Node The set of neighbor nodes of is the adjacency matrix of the subgraph, representing the nodes and nodes The connection weight represents the connection between nodes; and Respectively with the node and The number of connected edges, For the The weight matrix of the layer; The node feature representations guided by different meta-paths are fused through the semantic attention mechanism to obtain the fusion weight , whose expression is as follows: in, is the weight matrix for training, Represents a meta path The node feature representation below is: is the activation function; Fusion weight After normalization, the expression is as follows: The node feature representations of different meta-paths are weighted summed to obtain the final embedded feature representation of different nodes : 。 7. The method for predicting the interaction between lncRNA gene sequence and protein according to claim 6, characterized in that: The feature representations of lncRNA, miRNA and protein are fused to obtain fused embedded feature representations, including: Obtain the interaction matrix between lncRNA and miRNA, and the interaction matrix between protein and miRNA; Using the feature representations of lncRNA, miRNA, and protein, we generate the query matrix, key matrix, and value matrix of the multi-head attention mechanism. According to the query matrix, key matrix and value matrix, the interaction weight between lncRNA and protein is calculated, and its expression is as follows: in, represents the softmax function, is the query matrix of lncRNA, Key matrix for miRNA, is the bond matrix of the protein, is the scaling factor, is the bias term; Calculate the context feature representation of lncRNA separately and protein context feature representation : In the formula, is the query matrix of proteins, is the value matrix of miRNA; According to the interaction weight between lncRNA and protein, the context feature representation of lncRNA and protein context feature representation Perform fusion processing to obtain fusion context feature representation and fusion context feature representation , whose expression is as follows: The fusion context feature representation and fusion context feature representation The original lncRNA and protein representations are spliced together to obtain the final fusion embedded feature representation. L and fusion embedding feature representation P , whose expression is as follows: 。 8. The method for predicting the interaction between lncRNA gene sequence and protein according to claim 7, characterized in that: The fusion feature representation is input into the trained prediction model for prediction, and the prediction results of the interaction between lncRNA and protein are obtained, including: Embedding the fusion into feature representation L and fusion embedding feature representation P Concatenate into feature representation vector : in, Represents a splicing operation; Use ELU as the activation function to represent the feature vector Perform linear transformation to obtain output features : in, is the weight matrix, is the bias vector, is the activation function; The output features After a linear layer and LogSoftmax operation, the predicted probability distribution is obtained , whose expression is as follows: in, is the weight matrix, is the bias vector, It is used to convert the output features into a probability distribution, where each value represents the probability of whether the lncRNA and protein interact.
9. The method for predicting the interaction between lncRNA gene sequence and protein according to claim 8, characterized in that: Before inputting the fused feature representation into the trained prediction model for prediction, the method further includes using a comprehensive loss function to train and optimize the prediction model; The expression of the comprehensive loss function is as follows: in, is the feature representation learning loss of lncRNA, is the feature representation learning loss of miRNA, is the feature representation learning loss of proteins, is the sample weighted loss of the gradient reconciliation mechanism, is the training sample set, For sample i and samples j Negative log-likelihood loss, For sample i and samples j The weight of , , and are the weight coefficients of each loss respectively.
10. A lncRNA gene sequence and protein interaction prediction system, characterized in that: include: The acquisition module is used to obtain the sequence data of lncRNA and miRNA, as well as protein data; The calculation module is used to calculate the similarity of lncRNA, miRNA and protein respectively; A construction module is used to construct a heterogeneous information network based on the similarities of lncRNA, miRNA and protein, with lncRNA, miRNA and protein as nodes and similarity relationships between lncRNA, miRNA and protein as edges; The generation module is used to extract subgraphs from heterogeneous information networks according to preset meta-paths, and perform graph convolution operations on the subgraphs to generate feature representations of lncRNA, miRNA, and protein; A fusion module is used to fuse the feature representations of lncRNA, miRNA and protein to obtain a fusion feature representation; The prediction module is used to input the fusion feature representation into the trained prediction model for prediction, and obtain the prediction results of the interaction between lncRNA and protein.
Citation Information
Cited By
IncRNA-protein interaction prediction method based on bidirectional intention
CN121148467A
Bi-directional intention based lncrna-protein interaction prediction method
CN121148467B