PiRNA and disease association prediction method, device and equipment based on comparative learning

By constructing a piRNA-disease heterogeneous graph network, combining hierarchical attention and Transformer model, the model is optimized to solve the problems of data sparseness and category imbalance, and high-accurate piRNA-disease association prediction is achieved, breaking through the local limitations of traditional models and improving the prediction effect.

CN120544693APending Publication Date: 2025-08-26XIAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510647501.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing piRNA-disease association prediction technology faces the problems of data sparsity, category imbalance and nonlinear correlation characteristics, resulting in weak model training foundation and low prediction accuracy, making it difficult to give full play to its advantages in clinical prediction and mechanism research.

Method used

A method based on contrast learning is adopted, combined with the hierarchical attention mechanism and the Transformer model, a piRNA-disease heterogeneous graph network is constructed, and local and global features are captured through the hierarchical attention mechanism and the multi-head self-attention mechanism, and the model is optimized through comparison loss, class balance loss and regularization techniques to improve prediction accuracy.

Benefits of technology

It significantly improves the accuracy and robustness of piRNA-disease association prediction, enhances the ability of multi-level association characterization, can effectively deal with sparse data and category imbalance problems, improves AUC and AUPR indicators, and enhances the generalization performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544693A_ABST
    Figure CN120544693A_ABST
Patent Text Reader

Abstract

The invention discloses a pi RNA and disease association prediction method, device and equipment based on comparative learning. The method comprises the following steps: step S1, constructing a p RNA-disease heterogeneous graph network; s2, fusing the multi-level semantic information; step S3, according to the node representation, introducing a Transform model, and enhancing the global association modeling capability of the piRNA and the disease node; s4, constructing a topological graph and a semantic graph; step S5, obtaining an accurate p iRNA-disease association prediction score; the prediction device comprises a heterogeneous graph construction unit, a node embedding learning unit and a node embedding learning unit. An input and output unit, a storage unit, a communication unit, an RAM unit, an ROM unit and a GPU of the electronic equipment are connected with one another through a bus, and the requirements for complex calculation and data interaction of a p-RNA and disease associated prediction task are met. The method has the characteristic of high prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of piRNA function identification, and in particular relates to a method, device and equipment for predicting the association between piRNA and disease based on comparative learning. Background Art

[0002] Piwi-interacting RNA (piRNA), as an important class of small RNA molecules, plays a key role in biological processes such as gene regulation, stem cell development and disease occurrence. In recent years, breakthroughs in high-throughput sequencing technology have revealed that piRNA has significant regulatory functions in multiple medical fields such as tumorigenesis, neurodegenerative diseases, cardiovascular diseases and reproductive system diseases. According to statistics from the MNDR v3.0 database, there are only 11,981 piRNA-disease association entries that have been rigorously experimentally verified, while there are as many as 180,850 unverified predicted association records. This order of magnitude difference (verified data is less than 6.6% of predicted data) clearly shows that there is still a huge gap in the current systematic understanding of piRNA-disease associations.

[0003] Existing piRNA-disease association prediction technologies mainly include the following typical methods: 1) The iPiDA-GCN model based on graph convolutional networks (Hou et al.) analyzes the complex relationship between piRNAs and diseases by constructing heterogeneous networks; 2) The iPiDA-LTR tool (Zhang et al.) adopts a learning ranking strategy, which has the ability to predict the loss of known associations and process emerging piRNA associations; 3) The SPRDA model based on structural perturbation (Zheng et al.) uses a two-layer network architecture combined with positive sample unlabeled learning technology; 4) The deep learning framework piRDA (Ali et al.) adopts a two-step forward unlabeled learning and guidance technology to avoid false negative predictions.

[0004] However, existing technologies face three core flaws: first, the extreme sparsity of experimental validation data (the validation ratio is less than about 5%) leads to a weak foundation for model training; second, severe class imbalance problems (the amount of data associated with different diseases can differ by two orders of magnitude) reduce the accuracy of the prediction system in identifying low-frequency disease associations; third, the inherent nonlinear association characteristics of piRNA-disease make the prediction error rate of traditional linear models generally high. It is particularly worth noting that when faced with complex scenarios that simultaneously contain sparse data, class imbalance and nonlinear relationships, the existing optimal algorithms find it difficult to exert their advantages, which significantly limits their application value in clinical prediction and mechanism research. Summary of the Invention

[0005] To overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a method, device and equipment for predicting the association between piRNA and disease based on contrastive learning. This method uses contrastive learning technology, combines a hierarchical attention mechanism and a Transformer model, and introduces a joint optimization strategy to achieve accurate prediction of the association between piRNA and disease, with the characteristic of high prediction accuracy.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] A method for predicting the association between piRNA and disease based on contrastive learning, characterized by comprising the following steps:

[0008] Step S1, constructing a piRNA-disease heterogeneous graph network based on the known piRNA-disease interaction network, piRNA sequence alignment similarity network, and disease semantic similarity network;

[0009] Step S2: Based on the piRNA-disease heterogeneous graph network constructed in step S1, a hierarchical attention mechanism is used to dynamically learn the complex correlation features in the piRNA-disease heterogeneous graph; a node-level attention mechanism is used based on the piRNA-disease heterogeneous graph network to learn the relative importance of nodes and their meta-path neighbors to capture local structure; a semantic-level attention mechanism is used based on the piRNA-disease heterogeneous graph network to learn the importance of different meta-paths and integrate multi-level semantic information; and an embedded representation of piRNA and disease nodes is obtained;

[0010] In step S3, based on the piRNA and disease node embedding representations obtained in step S2, the Transformer model is introduced. This model can effectively capture the long-distance dependencies across meta-paths through a multi-head self-attention mechanism, breaking through the local limitations of traditional graph neural networks and realizing the global correlation characteristics across meta-paths.

[0011] Step S4: Based on the embedded information of piRNAs and disease nodes captured in step 3, a topological map and a semantic map are constructed. The topological map is obtained by determining whether two piRNA or disease pairs share the same piRNA or disease node. The semantic map is obtained by combining the piRNA-disease association matrix and the feature matrix using the KNN algorithm, effectively alleviating the data sparsity problem.

[0012] In step S5, a contrastive learning framework is constructed based on the topological graph and semantic graph information. By jointly optimizing contrastive loss, class balance loss, and regularization techniques, the discriminative ability and generalization performance of the contrastive learning framework are improved to obtain accurate piRNA-disease association prediction scores.

[0013] The step S2 specifically includes the following steps:

[0014] Step S2-1, extracting a set of meta-paths of multiple relationship patterns between nodes based on the combination relationship between node types and edge types in the heterogeneous graph;

[0015] Initialize the meta-path set and define the meta-path set C = (p1, p2, p3…p k ), these meta-paths are combined according to node types and edge types to represent various relationship patterns between nodes in the graph. Meta-paths are used to guide the flow of semantic information between nodes in a heterogeneous graph and help capture associated information in the heterogeneous graph. Meta-paths represent various relationships between nodes in a heterogeneous graph through different types of nodes and edges;

[0016] Step S2-2: For each meta-path, weights are assigned to neighboring nodes based on node features through similarity analysis and node-level attention mechanism.

[0017] For each node v, calculate its similarity S with its neighbor node u vu :

[0018]

[0019] Among them, h v and h u are the feature representations of nodes v and u respectively, ||·|| represents the norm of the vector, and T represents the transpose;

[0020] For each meta-path p i ∈C, generates specific meta-path instances based on node similarity. Specifically, the meta-path p i =(o1,o2,...,o n ) in each o n Represents the node type, starting from node v, according to the type sequence of the meta-path o1, o2, ..., o n , traverse its neighbor nodes u, and select the first three neighbor nodes u with higher similarity to the current node as the next node u i+1 , and no more than three-order neighbors, gradually constructing a complete meta-path instance P v , for each node v, generate a set of dynamic meta-path sets P v ={p v1 ,p v2 ,…,p vn}, in this step, for each dynamic meta-path p vi ∈P v , calculate the feature representation h of the meta-path by aggregating the features of all nodes on the path p , the aggregation operation adopts the attention pooling strategy, by assigning an attention weight α to each node v u , and then perform weighted summation on the node features to obtain the feature representation of the meta-path:

[0021]

[0022] Among them, node v is the neighbor set p v A node in , the eigenvector h of node v u According to the attention weight α u Perform weighted summation to obtain the updated feature representation h' of node v v ;

[0023] Step S2-3: According to step S2-2, the updated feature representation h' of node v is v Projecting to a shared semantic space, different weights are assigned to different meta-paths through a semantic-level attention mechanism to achieve weighted aggregation of node features and further enrich the feature representation of node v;

[0024] To calculate the importance of each meta-path, the node feature representation needs to be projected into a specific semantic space through a multi-layer perceptron: is the updated feature representation of node v; for each meta-path through the node-level attention mechanism, P i , calculate the semantic importance score of each meta-path: Where W a is the attention weight matrix, capturing the interaction relationship in the semantic space, is the semantic attention vector, b a is the bias term, tanh(·) represents the hyperbolic tangent function, and T represents the transpose;

[0025] The similarity scores are normalized using the softmax function to obtain the attention weight for each meta-path:

[0026]

[0027] in, is the metapath P i Semantic importance score, Represents the meta-path P i The contribution weight of node v, F represents the number of meta-paths, and f represents the sum index variable;

[0028] Finally, the node feature vector under each meta-path is multiplied by its corresponding attention weight, and the results of all meta-paths are weighted summed to obtain the final node representation: Among them, K represents K element paths, Represents the high-order semantic features after MLP nonlinear transformation. Similarly, the final piRNA node represents h p .

[0029] The step S3 specifically includes the following steps:

[0030] Step S3-1, according to step S2-3, obtain the final piRNA node representation h p and disease nodes represent h d As the input of the Transformer model, for each meta-path r constructed by step S2, a learnable relation embedding vector e is generated r , the vector e r Used to encode the semantic information of the meta-path r; for the meta-path r, the query, key, and value vectors are calculated, where in are three independent trainable projection matrices of the meta-path r to ensure meta-path specificity; || represents the splicing operation, h i and h j Represents the characteristic matrix of node i and node j;

[0031] After getting the query vector Q r and key vector K r After that, the attention score is calculated as follows:

[0032]

[0033] Among them, Q r Representation query vector, K r represents the key vector, d k is the dimension of the key vector, represents the attention score of node i and node j under meta-path r, and T represents transposition;

[0034] Step S3-2, in the Transformer model, the gate coefficient Dynamically adjust the weights of different node information to more flexibly capture the dependencies between nodes:

[0035]

[0036] Among them, q i is the query vector of node i, e r is to generate a learnable relation embedding vector, W r is the weight matrix, σ is the activation function;

[0037] Finally, the attention score is:

[0038]

[0039] in, The score between the i-th query and the j-th key under meta-path r, is the gating coefficient; Qr and Denote the query and key matrices, W r represents the weight matrix, d k is the dimension of the key vector, used to scale the dot product result;

[0040] In step S3-3, to further enhance the Transformer model's ability to distinguish between different types of relations, a different attention head is assigned to each meta-path type. The weight matrix corresponding to each attention head is learned independently to adapt to the feature representation of different types of meta-paths. Specifically, it is divided into local meta-paths and global meta-paths. The output of each attention head is calculated as follows: Among them, V r is a value vector, through Calculated, Z h is the output of the h-th attention head, and the outputs of each attention head are aggregated through the relationship-aware weights: Among them, α h is the weight of the h-th attention head, Z h is the output of the h-th attention head, H is the total number of attention heads, α h is the weight of the h-th attention head, Z h is the output of the h-th attention head, H is the total number of attention heads, and l represents the sum index variable;

[0041] In step S3-4, relative position encoding is introduced to capture the spatial relationship between nodes, further improving the Transformer model's ability to capture the spatial relationship between nodes and obtaining the final embedded representation of piRNAs and diseases. The attention score is calculated in two parts: position-based attention score and relationship-based attention score. The formula is as follows:

[0042]

[0043] Among them, y ij is the relative position code of node i and node j, d k is the dimension of the key vector,

[0044] The relation-based attention score is embedded in the vector e via the meta-path r And node characteristics are calculated, the specific formula is as follows:

[0045]

[0046] Among them, W r is a weight matrix used to capture the relationship information of the meta-path.

[0047] In order to dynamically adjust the strength of position encoding, the strength of position information is controlled by g. ij =σ(Wp [h i ||h j ]), the final attention score is:

[0048]

[0049] in, is the position-based attention score, is the relation-based attention score, g ij It controls the strength of position information and is used to weight the two attention scores.

[0050] The step S4 specifically includes the following steps:

[0051] Step S4-1, constructing a topology map based on whether two piRNA-disease pairs share the same piRNA or the same disease;

[0052] In the process of constructing the heterogeneous network topology, we first operate based on the piRNA-disease association matrix and the piRNA-disease feature matrix, and obtain the final embedding representation of piRNA and disease according to steps S2-4; the piRNA node embedding is represented by P n , the disease node embedding is represented as D m , to construct the piRNA and disease feature matrix F, which is defined as: A topological graph G2 is constructed using the piRNA-disease association matrix A and the piRNA-disease feature matrix F. In this topological graph G2, nodes include piRNAs and diseases, and edges represent the associations between piRNAs and diseases. If two piRNA-disease pairs contain the same piRNA or the same disease, there is an edge between them, indicating that the two piRNA-disease pairs share certain common characteristics; otherwise, there is no edge. The existence of such an edge reflects the potential association between the nodes and provides a basis for further analysis.

[0053] Step S4-2, constructing a semantic graph using the piRNA-disease association matrix and the piRNA-disease feature matrix according to the KNN algorithm;

[0054] A prediction association matrix A' was constructed using the KNN algorithm, and a topological graph G2 was constructed based on the piRNA-disease feature matrix F. The prediction association matrix A' represents the association between piRNAs and diseases constructed using the KNN algorithm, while the piRNA-disease feature matrix F contains the characteristic information of piRNAs and diseases. The specific implementation process is as follows: the cosine similarity between each piRNA and all diseases is calculated using the KNN algorithm. Based on the calculated cosine similarity, the top three most similar diseases are selected as adjacent nodes of the piRNA. The corresponding positions in the adjacency matrix are defined as 1 to indicate an association between the piRNA and these diseases, and 0 otherwise.

[0055] The step S5 specifically includes the following steps:

[0056] In step S5-1, based on the topological graph of step S4-1 and the semantic graph information of step S4-2, a contrastive loss function is constructed to capture the intrinsic structure and semantic information of the data by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs. The contrastive loss function of the topological graph is:

[0057]

[0058] Among them, k represents the index of traversing all nodes, U i Positive sample set, including piRNA node i itself and its first-order neighbor nodes; N i represents the negative sample set, which contains nodes that are not related to piRNA node i; Representation of node i in the topology graph; The representation of node j in the semantic graph; τ represents the temperature parameter; N represents the total number of piRNA-diseases; according to step S5-1, the topological graph information is replaced with the semantic graph information to obtain the contrast loss function L of the semantic graph s ;

[0059] Step S5-2: Based on the data class imbalance, a class-balanced loss function is designed to adjust the weights of positive and negative samples, so that the minority class contributes more to the loss. This gives the minority data class a higher weight during training, effectively alleviating the data class imbalance problem. A modulation factor is introduced based on class balance to reduce the loss contribution of easy-to-classify samples and further focus on minority data class samples:

[0060]

[0061] Where: γ is a modulation factor used to adjust the weight of difficult samples; the balanced weight when considering the number of samples is: Where β∈[0,1) is the smoothing parameter, Q classis the number of piRNA and disease pairs; 1-α is the weight of negative samples, and α is the weight of positive samples; Indicates the number of positive samples; p(x i ) Model for sample x i The predicted probability of x is the probability that the sample belongs to the positive class. i represents the feature vector of the i-th sample; θ i represents the true label of sample i; N represents the total number of piRNA-diseases;

[0062] In step S5-3, a regularization term is designed to combine the class balance loss and contrastive learning loss to effectively suppress the overfitting phenomenon of the model. Ultimately, the model optimization objective consists of three parts: class balance loss, contrastive learning loss, and regularization term R(Θ):

[0063]

[0064] Among them, R(Θ) is the sum of the squares of the weight values, Θ represents all trainable model parameters, w is the weight parameter of the regularization term, which is used to control the strength of the regularization, and L cb is the class balance loss, L t is the topology contrast loss, L s Semantic graph contrast loss, R(Θ) is the regularization term; λ is the weight coefficient used to balance L t and L s contribution;

[0065] In step 5-4, by optimizing the loss function and using the gradient descent method to update the parameters of the contrastive learning model, the association score between piRNA and disease is finally obtained, providing prediction results for subsequent research.

[0066] A prediction device for the method includes a heterogeneous graph construction unit, a node embedding learning unit, a topology map and semantic map construction unit, and a joint optimization prediction unit; the heterogeneous graph construction unit, the node embedding learning unit, the topology map and semantic map construction unit, and the joint optimization prediction unit can be implemented through hardware modules or software modules.

[0067] The prediction device and the heterogeneous graph construction unit are responsible for integrating the association data and related feature data of piRNA and disease, and constructing the piRNA-disease heterogeneous graph network through the known piRNA-disease interaction network, piRNA sequence alignment similarity network, and disease semantic similarity network;

[0068] The node embedding learning unit leverages heterogeneous graph information and employs a heterogeneous neural network to learn node-level and semantic-level attention mechanisms guided by meta-paths, effectively capturing local features of the heterogeneous graph. The Transformer model is introduced to address long-range dependencies across meta-paths, capturing global features and ultimately obtaining embeddings for piRNAs and disease nodes.

[0069] The topological map and semantic map construction unit constructs a topological map and a semantic map by capturing the embedded information of piRNAs and disease nodes through the Transformer model;

[0070] The joint optimization prediction unit fuses topological graph and semantic graph information through comparative learning to enhance the discriminability of node representation; introduces class balance loss to solve the problem of category imbalance, and combines regularization to improve model stability, and finally obtains the association prediction score through joint optimization.

[0071] An electronic device used in the method includes a bus, an input / output unit, a storage unit, a communication unit, a RAM unit, a ROM unit, and a GPU. These components are interconnected via the bus to form an efficient and collaborative computing system that supports the complex computing and data interaction requirements of the piRNA-disease association prediction task.

[0072] The bus serves as the core communication channel, connecting all functional units to ensure efficient transmission of data and control signals. The input and output units are connected to the storage unit, communication unit, RAM unit, ROM unit, and GPU through the bus. The input and output units are responsible for data input and output. The output device is used to display the prediction results.

[0073] The storage unit shares data with the communication unit, RAM unit, ROM unit, and GPU through the bus. The storage unit also supports data reading, writing, and backup to ensure data security and accessibility.

[0074] The communication unit is used to realize data exchange between the electronic device and external devices or networks;

[0075] The RAM unit serves as the control core, ensuring that all units operate in coordination according to predetermined program instructions, including scheduling computing tasks, managing memory allocation, and monitoring system status;

[0076] The ROM unit is used to store firmware and startup programs to ensure that the electronic device can start normally and load the operating system when it is turned on;

[0077] As the computing core of electronic equipment, the GPU is responsible for executing program instructions. The GPU interacts with storage units, ARM units, and other units through the bus to exchange data, including reading data from storage units, receiving scheduling instructions from ARM units, and transmitting calculation results to output devices or communication units.

[0078] The GPU is replaced by a CPU.

[0079] The beneficial effects of the present invention are:

[0080] Based on the known piRNA-disease interaction network, piRNA sequence alignment similarity network, and disease semantic similarity network, this paper constructs a piRNA-disease heterogeneous graph network. By combining the heterogeneous graph neural network layered attention mechanism and the Transformer model for feature extraction, it effectively captures local and global features, thereby better understanding the complex relationship between piRNAs and diseases. To address the problem of data sparsity, the model's ability to handle sparse data is enhanced by constructing topological and semantic graphs. Through a joint optimization strategy, contrast loss, class balance loss, and regularization techniques are introduced to improve the model's discriminative power and generalization performance, thereby obtaining accurate piRNA-disease association prediction scores. The advantages are specifically reflected in the following aspects:

[0081] 1) The ability to represent multi-level associations has been significantly improved

[0082] Since the present invention adopts a hierarchical graph attention framework, it dynamically learns the relative importance of meta-path neighbors in the piRNA-disease heterogeneous graph through the node-level attention mechanism (which can quantify the improvement of the key node feature weight allocation accuracy by 2.28%). At the same time, it adaptively fuses the semantic information of different meta-paths through the semantic-level attention mechanism (experimental data show that semantic feature fusion improves the accuracy by 2.35%). It comprehensively captures the multi-level association patterns between piRNAs and diseases, and solves the problem of insufficient representation caused by the single attention mechanism in the existing technology.

[0083] 2) Breakthrough in global interaction modeling capabilities

[0084] Through the innovatively designed Transformer model:

[0085] The multi-head self-attention mechanism is integrated to achieve global dependency modeling across meta-paths (breaking through the k-hop locality limitation of traditional GNNs, improving AUC by 0.92% and AUPR by 1.45%).

[0086] This technical solution effectively enhances the model's depth of exploration of potential piRNA-disease associations in complex biological networks.

[0087] 3) Robustness optimization and generalization performance enhancement

[0088] The proposed joint optimization strategy has three technical effects:

[0089] The contrastive loss function optimizes the representation distance between positive and negative samples, which improves the model's performance by 1.04% and AUPR by 4.78% under sparse data, significantly enhancing the discrimination ability.

[0090] To address the problem of imbalanced biological data categories, the loss weight is dynamically adjusted, which improves the AUC of the minority class by 8.23% and the AUPR by 13.54%.

[0091] By adopting an adaptive regularization strategy, the model's AUC stability on the test set is improved by 1.21% (overfitting rate is reduced), and AUPR is increased by 1.38%.

[0092] 4) Technical compatibility and scalability

[0093] This framework:

[0094] Supports graph structure input of any heterogeneous biological network (compatible with mainstream database formats such as STRING and KEGG)

[0095] Modular design enables key components (such as attention mechanism, optimization strategy) to be independently replaced and upgraded

[0096] It provides a standardized technical solution for the prediction of diagnostic targets for piRNA-related diseases (the test set covers 12 major diseases). BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Figure 1 It is a flow chart of the method of the present invention.

[0098] Figure 2 This is a schematic diagram of the self-attention mechanism based on relative position encoding and dynamic position gating mechanism of the present invention.

[0099] Figure 3 It is a principle block diagram of the method of the present invention.

[0100] Figure 4 It is a structural schematic diagram of a piRNA and disease association prediction device based on comparative learning of the present invention.

[0101] Figure 5 Schematic diagram of an electronic device of the present invention. DETAILED DESCRIPTION

[0102] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0103] See also Figure 1 A method for predicting the association between piRNA and disease based on contrastive learning, comprising the following steps:

[0104] Step S1, constructing a piRNA-disease heterogeneous graph network based on the known piRNA-disease interaction network, piRNA sequence alignment similarity network, and disease semantic similarity network;

[0105] The step S1 specifically includes the following steps:

[0106] Step S1-1: Download the currently available human piRNA-disease association data from the MNDR4.0 (https: / / doi.org / 10.1093 / nar / gkac814), piRDiseaseV1.0 (https: / / doi.org / 10.1093 / database / baz052), and MNDRv3.0 (http: / / www.rnadisease.org / ) databases, and construct a piRNA-disease association matrix based on the piRNA-disease association data;

[0107] Among them, piRNA-disease association data were downloaded from the MNDR4.0 dataset, and after screening, 9502 clinically and experimentally verified piRNA-disease associations between 8178 piRNAs and 15 diseases were obtained; piRNA-disease association data were downloaded from the MNDR v3.0 dataset, and after screening, 11981 experimentally verified piRNA-disease associations between 10149 piRNAs and 19 diseases were obtained. The purpose is to demonstrate the generalization ability on multiple datasets.

[0108] Based on the piRNA-disease association data downloaded above, a piRNA-disease association matrix was constructed. Let the piRNA-disease association matrix be A, where the rows of the matrix represent piRNAs and the columns of the matrix represent diseases. The formula is as follows:

[0109]

[0110] Among them, a m,n Indicates whether there is an association between the mth piRNA and the nth disease. If there is an association, then a m,n =1, otherwise a m,n =0;

[0111] Step S1-2: Calculate piRNA similarity using the Smith-Waterman alignment algorithm. By aligning each pair of piRNAs, calculate the local similarity between them, and construct a piRNA sequence alignment similarity network:

[0112]

[0113] Among them, SW(p n ,pm ) represents the sequence ratio between the nth piRNA and the mth piRNA, p n represents the nth piRNA, p m represents the mth piRNA, SW(p n ,p n ) represents the sequence ratio SW of the nth piRNA to itself (p m ,p m ) represents the sequence ratio of the mth piRNA to itself.

[0114] Step S1-3: construct a structured representation of the disease using a directed acyclic graph (DAG), quantify the similarity between diseases by calculating the semantic contribution value of disease nodes in the DAG, and construct a disease semantic similarity network;

[0115] Two diseases i and d j The semantic similarity of is calculated by the semantic contribution of their common ancestor nodes in the DAG. The specific formula is:

[0116]

[0117] in, Indicates disease i The disease semantic value of node t, Indicates disease i and the set of all its ancestor nodes, t is the common ancestor node, Dc(d i ) indicates disease d i With total semantic value, Dc(d j ) indicates disease d j With total semantic value, the formula is as follows:

[0118]

[0119] Among them, Δ is the semantic contribution attenuation parameter, which means that as the distance between the disease and the ancestor increases, the semantic contribution gradually decreases; Indicates disease i The disease semantic value of node t'; t' represents the child node of t.

[0120] Step S1-4, constructing a piRNA-disease heterogeneous graph network containing multiple relationships based on the piRNA-disease interaction network, piRNA sequence alignment similarity network, and disease semantic similarity network constructed in steps S1-1, S1-2, and S1-3;

[0121] Specifically, the piRNA-disease heterogeneity map The node set ψ and the edge set The node set ψ is divided into two categories: piRNA node set ψ p , disease node set ψ d ; The edges in the graph include three types: the association edge between piRNA and disease Used to characterize the interaction between piRNA and disease; similarity edges between piRNA nodes Quantify the similarity of piRNA sequences; similarity edges between disease nodes Reflects the degree of association at the semantic level of diseases; by calculating the association between piRNAs and diseases, the similarity of piRNAs, and the semantic similarity between diseases, association is performed for each type of edge, and this information is integrated into a multi-level, structured heterogeneous graph to reflect the complex relationship between piRNAs and diseases.

[0122] In step S2, based on the piRNA-disease heterogeneous graph network constructed in step S1, a hierarchical attention mechanism is used to dynamically learn the complex correlation features in the piRNA-disease heterogeneous graph; based on the piRNA-disease heterogeneous graph network, a node-level attention mechanism is used to learn the relative importance of nodes and their meta-path neighbors to capture local structure; based on the piRNA-disease heterogeneous graph network, a semantic-level attention mechanism is used to learn the importance of different meta-paths to integrate multi-level semantic information; and the embedded representation of piRNA and disease nodes is obtained.

[0123] The step S2 specifically includes the following steps:

[0124] Step S2-1, extracting a set of meta-paths of multiple relationship patterns between nodes based on the combination relationship between node types and edge types in the heterogeneous graph;

[0125] Initialize the meta-path set and define the meta-path set C = (p1, p2, p3…p k ), these meta-paths are combined according to node and edge types to represent various relationship patterns between nodes in the graph. These meta-paths will be used to guide the flow of semantic information between nodes in the heterogeneous graph and help capture the associated information in the heterogeneous graph. For example: p1 = piRNA → disease → piRNA, p2 = piRNA → piRNA → disease, p3 = disease → piRNA → disease. These meta-paths represent various relationships between nodes in the heterogeneous graph through different types of nodes and edges.

[0126] Step S2-2: For each meta-path, weights are assigned to neighboring nodes based on node features through similarity analysis and node-level attention mechanism.

[0127] For each node v, calculate its similarity S with its neighbor node u vu :

[0128]

[0129] Among them, h v and h u are the feature representations of nodes v and u respectively, ||·|| represents the norm of the vector, and T represents the transpose.

[0130] For each meta-path p i ∈C, generates specific meta-path instances based on node similarity. Specifically, the meta-path p i =(o1,o2,...,o n ) in each o n Represents the node type, starting from node v, according to the type sequence of the meta-path o1, o2, ..., o n , traverse its neighbor nodes u, and select the first three neighbor nodes with higher similarity to the current node as the next node u i+1 , and no more than three-order neighbors, gradually constructing a complete meta-path instance P v , for each node v, generate a set of dynamic meta-path sets P v ={p v1 ,p v2 ,…,p vn}, in this step, for each dynamic meta-path p vi ∈P v , calculate the feature representation h of the meta-path by aggregating the features of all nodes on the path p , the aggregation operation adopts the attention pooling strategy, by assigning an attention weight α to each node u , and then perform weighted summation on the node features to obtain the feature representation of the meta-path:

[0131]

[0132] Among them, node v is the neighbor set p v A node in , the eigenvector h of node v u According to the attention weight α u Perform weighted summation to obtain the updated feature representation h' of node v v ;

[0133] Step S2-3, according to step S2-2, update the feature table h' of node v v Projecting to a shared semantic space, different weights are assigned to different meta-paths through a semantic-level attention mechanism to achieve weighted aggregation of node features and further enrich the feature representation of the node;

[0134] To calculate the importance of each meta-path, the node feature representation needs to be projected into a specific semantic space through a multi-layer perceptron: is the updated feature representation of node v; for each meta-path through the node-level attention mechanism, P i , calculate the semantic importance score of each meta-path: Where W a is the attention weight matrix, capturing the interaction relationship in the semantic space, is the semantic attention vector, b a is the bias term, tanh(·) represents the hyperbolic tangent function, and T represents the transpose;

[0135] The similarity scores are normalized using the softmax function to obtain the attention weight for each meta-path:

[0136]

[0137] in, is the metapath P i Semantic importance score, Represents the meta-path P i The contribution weight of node v, F represents the number of meta-paths, and f represents the sum index variable;

[0138] Finally, the node feature vector under each meta-path is multiplied by its corresponding attention weight, and the results of all meta-paths are weighted summed to obtain the final node representation: Among them, K represents K element paths, Represents the high-order semantic features after MLP nonlinear transformation. Similarly, the final piRNA node represents h p .

[0139] In step S3, based on the piRNA and disease node embedding representations obtained in step S2, the Transformer model is introduced. This model can effectively capture the long-distance dependencies across meta-paths through a multi-head self-attention mechanism, breaking through the local limitations of traditional graph neural networks and realizing the global correlation characteristics across meta-paths.

[0140] The step S3 specifically includes the following steps:

[0141] Step S3-1, according to step S2-3, obtain the final piRNA node representation h p and disease nodes represent h d As the input of Transformer, for each meta-path r constructed by step S2, a learnable relation embedding vector e is generated r , the relationship embedding vector e rUsed to encode the semantic information of the meta-path r; for the meta-path r, the query, key, and value vectors are calculated, where in are three independent trainable projection matrices of the meta-path r to ensure meta-path specificity; || represents the splicing operation, h i and h j Represents the characteristic matrix of node i and node j;

[0142] After getting the query vector Q r and key vector K r After that, the attention score is calculated as follows:

[0143]

[0144] Among them, Q r Representation query vector, K r represents the key vector, d k is the dimension of the key vector, represents the attention score of node i and node j under meta-path r, and T represents transposition;

[0145] Step S3-2: In the Transformer model, the gate coefficient Ability to dynamically adjust the weights of different node information to better capture the relationship between nodes:

[0146]

[0147] Among them, q i is the query vector of node i, e r is to generate a learnable relation embedding vector, W r is the weight matrix, σ is the activation function;

[0148] Finally, the attention score is:

[0149]

[0150] in, The score between the i-th query and the j-th key under meta-path r, is the gating coefficient; Q r and Denote the query and key matrices, W r represents the weight matrix, d k is the dimension of the key vector, used to scale the dot product result;

[0151] In step S3-3, to further enhance the Transformer model's ability to distinguish between different types of relations, a different attention head is assigned to each meta-path type. The weight matrix corresponding to each attention head is learned independently to adapt to the feature representation of different types of meta-paths. Specifically, it can be divided into local meta-paths and global meta-paths. The output of each attention head is calculated as follows: Among them, V r is a value vector, through Calculated, Z h is the output of the h-th attention head, and the outputs of each attention head are aggregated through the relationship-aware weights: Among them, α h is the weight of the h-th attention head, Z h is the output of the h-th attention head, H is the total number of attention heads, and l represents the sum index variable;

[0152] In step S3-4, relative position encoding is introduced to capture the spatial relationship between nodes, further improving the Transformer model's ability to capture the spatial relationship between nodes and obtaining the final embedded representation of piRNAs and diseases. The attention score is calculated in two parts: position-based attention score and relationship-based attention score. The formula is as follows:

[0153]

[0154] Among them, y ij is the relative position code of node i and node j, d k is the dimension of the key vector.

[0155] The relation-based attention score is embedded in the vector e via the meta-path r And node characteristics are calculated, the specific formula is as follows:

[0156]

[0157] Among them, W r is a weight matrix used to capture the relational information of meta-paths.

[0158] In order to dynamically adjust the strength of position encoding, the strength of position information is controlled by g. ij =σ(W p [h i ||h j ]), the final attention score is:

[0159]

[0160] in, is the position-based attention score, is the relation-based attention score, g ij It controls the strength of position information and is used to weight the two attention scores;

[0161] In step S4, a topological map and a semantic map are constructed based on the embedded information of piRNAs and disease nodes captured by the Transformer model in step 3. The topological map is obtained by determining whether two piRNA or disease pairs share the same piRNA or disease node. The semantic map is obtained by combining the piRNA-disease association matrix and the feature matrix according to the KNN algorithm, effectively alleviating the data sparsity problem.

[0162] The step S4 specifically includes the following steps:

[0163] Step S4-1, constructing a heterogeneous network topology map based on whether two piRNA-disease pairs share the same piRNA or the same disease;

[0164] In the process of constructing the topological map, we first operate based on the piRNA-disease association matrix and the piRNA-disease feature matrix, and obtain the final embedding representation of piRNA and disease according to steps S2-4; the piRNA node embedding is represented by P n , the disease node embedding is represented as D m , in order to construct the piRNA and disease feature matrix F, it is defined as: A topological graph G2 is constructed using the piRNA-disease association matrix A and the piRNA-disease feature matrix F. In this topological graph G2, nodes include piRNAs and diseases, and edges represent the associations between piRNAs and diseases. If two piRNA-disease pairs contain the same piRNA or the same disease, there is an edge between them, indicating that the two piRNA-disease pairs share certain common characteristics; otherwise, there is no edge. The existence of such an edge reflects the potential association between the nodes and provides a basis for further analysis.

[0165] Step S4-2, constructing a semantic graph using the piRNA-disease association matrix and the piRNA-disease feature matrix according to the KNN algorithm;

[0166] The KNN algorithm is used to construct a prediction association matrix A', and a topological graph G2 is constructed based on the piRNA and disease feature matrix F. The prediction association matrix A' represents the association relationship between piRNAs and diseases constructed by the KNN algorithm, and the prediction association matrix F contains the feature information of piRNAs and diseases. The specific implementation process is as follows: the cosine similarity between each piRNA and all diseases is calculated using the KNN algorithm. Based on the calculated cosine similarity, the top three most similar diseases are selected as adjacent nodes of the piRNA, and the corresponding positions in the adjacency matrix are defined as 1, indicating that there is an association between the piRNA and these diseases, otherwise it is 0.

[0167] In step S5, a contrastive learning framework is constructed based on the topological graph and semantic graph information. By jointly optimizing contrastive loss, class balance loss, and regularization techniques, the discriminative ability and generalization performance of the contrastive learning framework are improved to obtain accurate piRNA-disease association prediction scores.

[0168] The step S5 specifically includes the following steps:

[0169] Step S5-1: Based on the topological graph of step S4-1 and the semantic graph information of step S4-2, as shown in FIG. Figure 3 As shown, a contrast loss function is constructed to better capture the intrinsic structure and semantic information of the data by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs. The contrast loss function of the topological map is as follows:

[0170]

[0171] Among them, k represents the index of traversing all nodes, U i Positive sample set, including piRNA node i itself and its first-order neighbor nodes; N i represents the negative sample set, which contains nodes that are not related to piRNA node i; Representation of node i in the topology graph; The representation of node j in the semantic graph; τ represents the temperature parameter; N is the total number of all piRNA-diseases; similarly, the contrast loss function L of the semantic graph can be obtained s ;

[0172] Step S5-2: Based on the data class imbalance, a class-balanced loss function is designed to adjust the weights of positive and negative samples, so that the minority class contributes more to the loss. This gives the minority data class a higher weight during training, effectively alleviating the data class imbalance problem. A modulation factor is introduced on the basis of class balance to reduce the loss contribution of easy-to-classify samples and further focus on minority data class samples:

[0173]

[0174] Where: γ is a modulation factor used to adjust the weight of difficult samples; the balanced weight when considering the number of samples is: Where β∈[0,1) is the smoothing parameter, Q class is the number of piRNA and disease pairs; 1-α is the weight of negative samples, and α is the weight of positive samples; Indicates the number of positive samples; p(x i ) Model for sample x i The predicted probability of x is the probability that the sample belongs to the positive class. i represents the feature vector of the i-th sample; θ i represents the true label of sample i; N is the total number of all piRNA-diseases;

[0175] In step S5-3, a regularization term is designed to combine the class balance loss and contrastive learning loss to effectively suppress the overfitting phenomenon of the model. Ultimately, the model optimization objective consists of three parts: class balance loss, contrastive learning loss, and regularization term R(Θ):

[0176]

[0177] Among them, R(Θ) is the sum of the squares of the weight values, Θ represents all trainable model parameters, w is the weight parameter of the regularization term, which is used to control the strength of the regularization, and L cb is the class balance loss, L t is the topology contrast loss, L s Semantic graph contrast loss, R(Θ) is the regularization term; λ is the weight coefficient used to balance L t and L s contribution;

[0178] In step 5-4, by optimizing the loss function and using the gradient descent method to update the parameters of the contrastive learning model, the association score between piRNA and disease is finally obtained, providing prediction results for subsequent research.

[0179] To further illustrate the benefits of the method provided by the present invention, we conducted a performance test on the MNDR4.0, piRDisease v1.0, and MNDRv3.0 datasets, using five-fold cross-validation to evaluate the effectiveness and superiority of this method in predicting piRNA-disease associations. The specific steps are as follows:

[0180] Each dataset is randomly divided into five non-overlapping subsets (i.e., quintiles), each containing roughly the same number of samples to ensure uniform data distribution. In each round of cross-validation, one of the subsets is selected as the test set, and the remaining four subsets are used as training sets, ensuring that each subset has an opportunity to serve as the test set. In each round of cross-validation, the contrastive learning model is trained using the training set, and the model's performance is evaluated using the test set. This process is repeated five times, ensuring that each subset is used as the test set at least once, resulting in five sets of performance evaluation results.

[0181] To comprehensively evaluate the performance of the learning model, the average AUC and average AUPR from five-fold cross-validation were used as the primary evaluation metrics. AUC plots the true positive rate against the false positive rate to evaluate the model's performance at different classification thresholds. AUPR plots the precision against the recall rate to evaluate the model's performance on a class-imbalanced dataset. By calculating the average of the AUC and AUPR from the five-fold cross-validation, we can verify the model's robustness on sparse and class-imbalanced datasets. The experimental results are shown in the following table:

[0182] Table 1: Five-fold cross-validation experimental results of SCLPDA and other methods under MNDR4.0

[0183] method AUC AUPR iPiDi-PUL 0.7842 0.1652 piRDA 0.9122 0.3385 iPiDA-GCN 0.8852 0.4275 iPiDA-SWGCN 0.9203 0.4615 PUTransGCN 0.9301 0.5982 SCLPDA 0.9115 0.8779

[0184] Table 2: Five-fold cross-validation experimental results of SCLPDA and other methods under MNDRv3.0

[0185]

[0186]

[0187] As shown in Tables 1 and 2, on the MNDR4.0 dataset, our method, SCLPDA, achieved AUC and AUPR of 0.9115 and 0.8779, respectively. While the AUC was slightly lower, the AUPR was 0.2797 higher than existing methods, demonstrating that the model maintains a high recall rate while maintaining high precision. On the MNDRv3.0 dataset, both the AUC and AUPR were higher than those of existing methods, demonstrating the robustness of the model across multiple datasets. Existing methods generally achieved low AUPR values ​​on the MNDR4.0 dataset, indicating significant shortcomings in their performance with sparse data. Specifically, existing methods often struggle to effectively capture key features of piRNAs and disease positive samples when processing sparse data, resulting in poor prediction performance, especially difficulty maintaining a high recall rate while maintaining high precision. In contrast, our method, through a contrastive learning mechanism, integrates multi-source information from topological and semantic graphs, significantly alleviating the data sparsity issue. This fusion strategy not only enhances the model's ability to model sparse data, but also enables it to maintain a high recall rate under high precision, thus showing significant advantages in the task of predicting the association between piRNA and disease.

[0188] At least one implementation of this embodiment provides a piRNA-disease association prediction device based on comparative learning. Figure 4 The schematic diagram of the prediction device is shown in Figure 2. Figure 4 As shown, the piRNA-disease association prediction device 400 includes a heterogeneous graph construction unit 410, a node embedding learning unit 420, a topology graph and semantic graph construction unit 430, and a joint optimization prediction unit 440. These units can be implemented as hardware modules or software modules, for example, using a CPU, GPU, or other hardware device with data processing capabilities and corresponding computer instructions. This embodiment does not specifically limit this.

[0189] Among them, the heterogeneous graph construction unit 410 is responsible for integrating the association data and related feature data of piRNA and disease, and constructing the piRNA-disease heterogeneous graph network through the known piRNA-disease interaction network, piRNA sequence alignment similarity network, and disease semantic similarity network; its specific implementation method can refer to the relevant description of S110 and will not be repeated here.

[0190] The node embedding learning unit 420 utilizes heterogeneous graph information and adopts a heterogeneous neural network to perform meta-path-guided node-level and semantic-level attention mechanism learning, effectively capturing the local features of the heterogeneous graph. The Transformer model is introduced to resolve long-distance dependencies across meta-paths, capture global features, and obtain the final piRNA and disease node embeddings. The specific implementation method can be referred to the relevant description of S120 and will not be repeated here.

[0191] The topology map and semantic map construction unit 430 constructs a topology map and a semantic map based on the embedded information of piRNAs and disease nodes captured by the Transformer model. The specific implementation method can refer to the relevant description of S130 and will not be repeated here.

[0192] Joint optimization prediction unit 440 enhances the discriminability of node representations by fusing topological and semantic graph information through comparative learning. It also introduces a class balance loss to address class imbalance and combines it with regularization to improve model stability. Ultimately, it achieves an association prediction score through joint optimization. The specific implementation method can be found in the description of S140 and will not be further elaborated here.

[0193] The present invention also provides an electronic device for predicting piRNA-disease associations based on contrastive learning. The electronic device for predicting piRNA-disease associations of the present invention can be implemented in various forms, such as desktop computers, laptop computers, embedded computers, dedicated hardware accelerators (such as GPUs or TPUs), and high-performance computing clusters. It should be noted that the components, their connections, and their functions described herein are provided for illustrative purposes only and do not constitute limitations on the scope of the present invention. In practical applications, the device configuration can be adjusted and optimized based on specific needs.

[0194] like Figure 5 As shown, the electronic device 500 may include a bus 510, an input / output unit 520, a storage unit 530, a communication unit 540, a RAM unit 550, a ROM unit 560, and a GPU / CPU 570. These components are interconnected via the bus 510 to form an efficient collaborative computing system that can support the complex computing and data interaction requirements of the piRNA-disease association prediction task.

[0195] The bus 510 serves as the core communication channel, connecting all functional units and ensuring efficient transmission of data and control signals. The input and output unit 520 connects to other units through the bus 510 and is responsible for data input and output. Input devices can be keyboards, mice, touch screens, etc. Output devices can be displays, projectors, or data interfaces for displaying prediction results.

[0196] The storage unit 530 can be a hard disk, a magnetic disk, an optical disk, or other storage media, and can share data with other units via the bus 510. In addition, the storage unit 530 also supports data reading, writing, and backup to ensure data security and accessibility.

[0197] The communication unit 540 can be Ethernet, optical fiber, wireless transceiver, etc., for realizing data exchange between the electronic device 500 and external devices or networks;

[0198] The RAM unit 550 acts as the control core, ensuring that the various units operate in coordination according to predetermined program instructions, such as scheduling computing tasks, managing memory allocation, and monitoring system status;

[0199] ROM unit 560 is used to store firmware and startup programs, ensuring that electronic device 500 can boot normally and load the operating system when it is turned on. The GPU / CPU, as the computing core of electronic device 500, is responsible for executing program instructions. The GPU / CPU exchanges data with storage unit 530, ARM unit 550, and other units through bus 510, for example, reading data from storage unit 530, receiving scheduling instructions from ARM unit 550, and transmitting calculation results to output device 520 or communication unit 540. Through the coordinated work of the above-mentioned units, electronic device 500 can efficiently and accurately complete the task of predicting the association between piRNA and disease, providing reliable technical support for biomedical research.

[0200] The GPU 570 may be replaced by a CPU.

Claims

1. A piRNA-disease association prediction method based on contrastive learning, characterized in that: The following steps are involved: Step S1, constructing a piRNA-disease heterogeneous graph network based on the known piRNA-disease interaction network, piRNA sequence alignment similarity network, and disease semantic similarity network; Step S2, based on the piRNA-disease heterogeneous graph network constructed in step S1, a hierarchical attention mechanism is used to dynamically learn the complex correlation features in the piRNA-disease heterogeneous graph; Based on the piRNA-disease heterogeneous graph network, a node-level attention mechanism is used to learn the relative importance of nodes and their meta-path neighbors to capture local structure. Based on the piRNA-disease heterogeneous graph network, a semantic-level attention mechanism is used to learn the importance of different meta-paths and integrate multi-level semantic information. The embedded representation of piRNA and disease nodes is obtained. In step S3, based on the piRNA and disease node embedding representations obtained in step S2, the Transformer model is introduced. This model effectively captures long-range dependencies across meta-paths through a multi-head self-attention mechanism, breaking through the local limitations of traditional graph neural networks and realizing global correlation features across meta-paths. Step S4: Based on the embedded information of piRNAs and disease nodes captured in step 3, a topological map and a semantic map are constructed. The topological map is obtained by determining whether two piRNA or disease pairs share the same piRNA or disease node. The semantic map is obtained by combining the piRNA-disease association matrix and the feature matrix using the KNN algorithm, effectively alleviating the data sparsity problem. In step S5, a contrastive learning framework is constructed based on the topological graph and semantic graph information. By jointly optimizing contrastive loss, class balance loss, and regularization techniques, the discriminative ability and generalization performance of the contrastive learning framework are improved to obtain accurate piRNA-disease association prediction scores.

2. According to claim 1, it is characterized in that The method for predicting the association between piRNA and disease based on contrastive learning, step S2, specifically includes the following steps: Step S2-1, extracting a set of meta-paths of multiple relationship patterns between nodes based on the combination relationship between node types and edge types in the heterogeneous graph; Initialize the meta-path set and define the meta-path set C = (p1, p2, p3…p k ), these meta-paths are combined according to node types and edge types to represent various relationship patterns between nodes in the graph. Meta-paths are used to guide the flow of semantic information between nodes in a heterogeneous graph and help capture associated information in the heterogeneous graph. Meta-paths represent various relationships between nodes in a heterogeneous graph through different types of nodes and edges; Step S2-2: For each meta-path, weights are assigned to neighboring nodes based on node features through similarity analysis and node-level attention mechanism. For each node v, calculate its similarity S with its neighbor node u vu : Among them, h v and h u are the feature representations of nodes v and u respectively, ||·|| represents the norm of the vector, and T represents the transpose; For each meta-path p i ∈C, generates specific meta-path instances based on node similarity. Specifically, the meta-path p i =(o1,o2,...,o n ) in each o n Represents the node type, starting from node v, according to the type sequence of the meta-path o1, o2, ..., o n , traverse its neighbor nodes u, and select the first three neighbor nodes u with higher similarity to the current node as the next node u i+1 , and no more than three-order neighbors, gradually constructing a complete meta-path instance P v , for each node v, generate a set of dynamic meta-path sets P v ={p v1 ,p v2 ,…,p vn }, in this step, for each dynamic meta-path p vi ∈P v , calculate the feature representation h of the meta-path by aggregating the features of all nodes on the path p , the aggregation operation adopts the attention pooling strategy, by assigning an attention weight α to each node v u , and then perform weighted summation on the node features to obtain the feature representation of the meta-path: Among them, node v is the neighbor set p v A node in , the eigenvector h of node v u According to the attention weight α u Perform weighted summation to obtain the updated feature representation h' of node v v ; Step S2-3: According to step S2-2, the updated feature representation h' of node v is v Projecting to a shared semantic space, different weights are assigned to different meta-paths through a semantic-level attention mechanism to achieve weighted aggregation of node features and further enrich the feature representation of node v; To calculate the importance of each meta-path, the node feature representation needs to be projected into a specific semantic space through a multi-layer perceptron: is the updated feature representation of node v; for each meta-path through the node-level attention mechanism, P i , calculate the semantic importance score of each meta-path: Where W a is the attention weight matrix, capturing the interaction relationship in the semantic space, is the semantic attention vector, b a is the bias term, tanh(·) represents the hyperbolic tangent function, and T represents the transpose; The similarity scores are normalized using the softmax function to obtain the attention weight for each meta-path: in, is the metapath P i Semantic importance score, Represents the meta-path P i The contribution weight of node v, F represents the number of meta-paths, and f represents the sum index variable; Finally, the node feature vector under each meta-path is multiplied by its corresponding attention weight, and the results of all meta-paths are weighted summed to obtain the final node representation: Among them, K represents K element paths, Represents the high-order semantic features after MLP nonlinear transformation. Similarly, the final piRNA node represents h p .

3. The method for predicting the association between piRNA and disease based on contrastive learning according to claim 1, characterized in that: The step S3 specifically includes the following steps: Step S3-1, according to step S2-3, obtain the final piRNA node representation h p and disease nodes represent h d As the input of the Transformer model, for each meta-path r constructed by step S2, a learnable relation embedding vector e is generated r , the relationship embedding vector e r Used to encode the semantic information of the meta-path r; for the meta-path r, the query, key, and value vectors are calculated, where in are three independent trainable projection matrices of the meta-path r to ensure meta-path specificity; || represents the splicing operation, h i and h j Represents the characteristic matrix of node i and node j; After getting the query vector Q r and key vector K r After that, the attention score is calculated as follows: Among them, Q r Representation query vector, K r represents the key vector, d k is the dimension of the key vector, represents the attention score of node i and node j under meta-path r, and T represents transposition; Step S3-2, in the Transformer model, the gate coefficient Dynamically adjust the weights of different node information to more flexibly capture the dependencies between nodes: Among them, q i is the query vector of node i, e r is to generate a learnable relation embedding vector, W r is the weight matrix, σ is the activation function; Finally, the attention score is: in, The score between the i-th query and the j-th key under meta-path r, is the gating coefficient; Q r and Denote the query and key matrices, W r represents the weight matrix, d k is the dimension of the key vector, used to scale the dot product result; In step S3-3, to further enhance the Transformer model's ability to distinguish between different types of relations, a different attention head is assigned to each meta-path type. The weight matrix corresponding to each attention head is learned independently to adapt to the feature representation of different types of meta-paths. Specifically, it is divided into local meta-paths and global meta-paths. The output of each attention head is calculated as follows: Among them, V r is a value vector, through Calculated, Z h is the output of the h-th attention head, and the outputs of each attention head are aggregated through the relationship-aware weights: Among them, α h is the weight of the h-th attention head, Z h is the output of the h-th attention head, H is the total number of attention heads, and l represents the sum index variable; In step S3-4, relative position encoding is introduced to capture the spatial relationship between nodes, further improving the Transformer model's ability to capture the spatial relationship between nodes and obtaining the final embedded representation of piRNAs and diseases. The attention score is calculated in two parts: position-based attention score and relationship-based attention score. The formula is as follows: Among them, y ij is the relative position code of node i and node j, d k is the dimension of the key vector, The relation-based attention score is embedded in the vector e via the meta-path r And node characteristics are calculated, the specific formula is as follows: Among them, W r is a weight matrix used to capture the relational information of meta-paths, In order to dynamically adjust the strength of position encoding, the strength of position information is controlled by g. ij =σ(W p [h i ||h j ]), the final attention score is: in, is the position-based attention score, is the relation-based attention score, g ij It controls the strength of position information and is used to weight the two attention scores.

4. The method for predicting the association between piRNA and disease based on contrastive learning according to claim 1, characterized in that: The step S4 specifically includes the following steps: Step S4-1, constructing a topology map based on whether two piRNA-disease pairs share the same piRNA or the same disease; In the process of constructing the topological map, we first operate based on the piRNA-disease association matrix and the piRNA-disease feature matrix, and obtain the final embedding representation of piRNA and disease according to steps S2-4; the piRNA node embedding is represented by P n , the disease node embedding is represented as D m , to construct the piRNA and disease feature matrix F, which is defined as: A topological graph G2 is constructed using the piRNA-disease association matrix A and the piRNA-disease feature matrix F. In this topological graph G2, nodes include piRNAs and diseases, and edges represent the associations between piRNAs and diseases. If two piRNA-disease pairs contain the same piRNA or the same disease, there is an edge between them, indicating that the two piRNA-disease pairs share certain common characteristics; otherwise, there is no edge. The existence of such an edge reflects the potential association between the nodes and provides a basis for further analysis. Step S4-2, constructing a semantic graph using the piRNA-disease association matrix and the piRNA-disease feature matrix according to the KNN algorithm; The KNN algorithm was used to construct a prediction association matrix A', and a topological graph G2 was constructed based on the piRNA and disease feature matrix F. The prediction association matrix A' represents the association relationship between piRNAs and diseases constructed by the KNN algorithm; the piRNA and disease feature matrix F contains the characteristic information of piRNAs and diseases. The specific implementation process is as follows: the cosine similarity between each piRNA and all diseases is calculated using the KNN algorithm. Based on the calculated cosine similarity, the top three most similar diseases are selected as adjacent nodes of the piRNA, and the corresponding positions in the adjacency matrix are defined as 1, indicating that there is an association between the piRNA and these diseases, otherwise it is 0.

5. The method for predicting the association between piRNA and disease based on contrastive learning according to claim 1, characterized in that: The step S5 specifically includes the following steps: In step S5-1, based on the topological graph of step S4-1 and the semantic graph information of step S4-2, a contrastive loss function is constructed to capture the intrinsic structure and semantic information of the data by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs. The contrastive loss function of the topological graph is: Among them, k represents the index of traversing all nodes, U i Positive sample set, including piRNA node i itself and its first-order neighbor nodes; N i represents the negative sample set, which contains nodes that are not related to piRNA node i; Representation of node i in the topological graph; The representation of node j in the semantic graph; τ represents the temperature parameter; N represents the total number of piRNA-diseases; according to step S5-1, the topological graph information is replaced with the semantic graph information to obtain the contrast loss function L of the semantic graph s ; Step S5-2: Based on the data class imbalance, a class-balanced loss function is designed to adjust the weights of positive and negative samples, so that the minority class contributes more to the loss. This gives the minority data class a higher weight during training, effectively alleviating the data class imbalance problem. A modulation factor is introduced based on class balance to reduce the loss contribution of easy-to-classify samples and further focus on minority data class samples: Where: γ is a modulation factor used to adjust the weight of difficult samples; the balanced weight when considering the number of samples is: Where β∈[0,1) is the smoothing parameter, Q class is the number of piRNA and disease pairs; 1-α is the weight of negative samples, and α is the weight of positive samples; Indicates the number of positive samples; p(x i ) Model for sample x i The predicted probability of x is the probability that the sample belongs to the positive class. i represents the feature vector of the i-th sample; θ i represents the true label of sample i; N represents the total number of piRNA-diseases; In step S5-3, a regularization term is designed to combine the class balance loss and contrastive learning loss to effectively suppress the overfitting phenomenon of the model. Ultimately, the model optimization objective consists of three parts: class balance loss, contrastive learning loss, and regularization term R(Θ): Among them, R(Θ) is the sum of the squares of the weight values, Θ represents all trainable model parameters, w is the weight parameter of the regularization term, which is used to control the strength of the regularization, and L cb is the class balance loss, L t is the topology contrast loss, L s Semantic graph contrast loss, R(Θ) is the regularization term; λ is the weight coefficient used to balance L t and L s contribution; In step 5-4, by optimizing the loss function and using the gradient descent method to update the parameters of the contrastive learning model, the association score between piRNA and disease is finally obtained, providing prediction results for subsequent research.

6. A prediction device for the method according to claim 1, characterized in that: The system includes a heterogeneous graph construction unit (410), a node embedding learning unit (420), a topology map and semantic map construction unit (430), and a joint optimization prediction unit (440); the heterogeneous graph construction unit (410), the node embedding learning unit (420), the topology map and semantic map construction unit (430), and the joint optimization prediction unit (440) can be implemented through hardware modules or software modules.

7. The prediction device according to claim 6, characterized in that The heterogeneous graph construction unit (410) is responsible for integrating the association data between piRNA and disease and related feature data, and constructing the piRNA-disease heterogeneous graph network through the known piRNA-disease interaction network, piRNA sequence alignment similarity network, and disease semantic similarity network; The node embedding learning unit (420) utilizes heterogeneous graph information and adopts a heterogeneous neural network to perform meta-path-guided node-level and semantic-level attention mechanism learning, effectively capturing the local features of the heterogeneous graph, introducing a Transformer model to resolve long-distance dependencies across meta-paths, capturing global features, and obtaining the final embedding of piRNAs and disease nodes; The topological map and semantic map construction unit (430) constructs a topological map and a semantic map using the embedded information of piRNAs and disease nodes captured by the Transformer model; The joint optimization prediction unit (440) enhances the discriminability of node representation by fusing topological graph and semantic graph information through comparative learning; introduces class balance loss to solve the class imbalance problem, and combines regularization to improve model stability, and finally obtains the association prediction score through joint optimization.

8. An electronic device used in the method according to claim 1, characterized in that: The system comprises a bus (510), an input / output unit (520), a storage unit (530), a communication unit (540), a RAM unit (550), a ROM unit (560) and a GPU (570). These components are interconnected via the bus (510) to form an efficient collaborative computing system that supports the complex computing and data interaction requirements of the piRNA-disease association prediction task.

9. The electronic device according to claim 8, wherein: The bus (510) serves as the core communication channel, connecting all functional units to ensure efficient transmission of data and control signals. The input / output unit (520) is connected to the storage unit (530), the communication unit (540), the RAM unit (550), the ROM unit (560), and the GPU (570) via the bus (510). The input / output unit (520) is responsible for data input and output. Output devices are used to display prediction results; The storage unit (530) shares data with the communication unit (540), the RAM unit (550), the ROM unit (560) and the GPU (570) via the bus (510). The storage unit (530) also supports reading, writing and backing up data to ensure data security and accessibility. The communication unit (540) is used to implement data interaction between the electronic device (500) and an external device or a network; The RAM unit (550) serves as the control core, ensuring that each unit operates in coordination according to predetermined program instructions, including scheduling computing tasks, managing memory allocation, and monitoring system status; The ROM unit (560) is used to store firmware and startup programs to ensure that the electronic device (500) can normally start and load the operating system when it is turned on; The GPU (570) serves as the computing core of the electronic device (500) and is responsible for executing program instructions. The GPU exchanges data with the storage unit (530), the ARM unit (550) and other units through the bus (510), including reading data from the storage unit (530), receiving scheduling instructions from the ARM unit (550), and transmitting calculation results to the output device (520) or the communication unit (540).

10. The electronic device according to claim 8 or 9, characterized in that: The GPU (570) is replaced by a CPU.

Citation Information

Cited By

  • Catalyst prediction method and system based on molecular characterization contrast learning

    CN121034442A