System for carrying out snoRNA and disease association prediction based on multi-graph SAGE network
Through the multi-graph SAGE network, the heterogeneous network and GraphSAGE architecture are integrated, combined with graph attention mechanism and small batch loader technology, the existing model has been solved with high cost and limited applicability, and efficient and flexible snoRNA and disease association prediction is achieved.
Patent Information
- Application Number
- CN202510247688.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-08-26
AI Technical Summary
The existing snoRNA-disease association prediction models are costly and time-consuming, making it difficult to fully cover all combinations, and the model needs to be reconstructed when facing new snoRNAs or diseases, resulting in limited applicability and generalization capabilities.
Multi-graph SAGE network is used to integrate multiple heterogeneous networks, combine the feature learning ability of GraphSAGE architecture, improve computing efficiency through small batch loader technology, enhance model adaptability and scalability, and adjust the aggregation method using graph attention mechanism and bidirectional message delivery to capture the complex relationship between snoRNA and disease.
It improves the prediction accuracy of snoRNA-disease association, reduces training time, enhances the adaptability and generalization performance of the model, and can calculate the embedding of newly introduced snoRNA or disease nodes, providing more efficient and flexible prediction tools.
Smart Images

Figure CN120544677A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to a system for predicting the association between snoRNA and diseases based on a multi-graph SAGE network. Background Art
[0002] In the field of bioinformatics, small nucleolar RNA (snoRNA), as an important structural component of the eukaryotic nucleolus, plays a vital role in the maturation and post-transcriptional modification of ribosomal RNA. snoRNA is not only involved in the complex process of ribosome synthesis, but is also closely related to the development of various human diseases. In recent years, with the development of high-throughput sequencing technology, more and more snoRNAs have been identified, but the specific association mechanism between snoRNA and disease is still not fully understood. Therefore, accurately predicting the association between snoRNA and disease is of great significance for understanding the occurrence and development of diseases and promoting disease diagnosis, treatment and drug development.
[0003] Currently, the method for identifying the association between snoRNA and disease mainly relies on classical biological experiments. However, this method is not only costly and time-consuming, but also difficult to fully cover all snoRNA and disease combinations, which seriously hinders the in-depth research. In addition, although some computational models have been used to predict the association between snoRNA and disease, these models have defects to varying degrees. Although the PSnoD method has achieved certain predictive performance through 5-fold stratified shuffling, there is still room for improvement. Although the iSnoDi-LSGT model uses network embedding technology, the effect in practical applications is not ideal. Although the GCNSDA model has achieved high prediction accuracy based on graph neural networks, when faced with new snoRNAs or diseases, it is necessary to rebuild the graph and training model, resulting in limited applicability and generalization ability of the model.
[0004] To address the above issues, it is necessary to optimize the existing prediction model by integrating multiple heterogeneous networks to comprehensively capture the characteristic information of snoRNA and disease from different perspectives, and combining the feature learning capabilities of the GraphSAGE architecture to effectively capture the complex relationship between the two. Therefore, it is of great significance to develop a system based on multi-graph SAGE networks to predict the association between snoRNA and disease that can comprehensively realize the above characteristics. Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a system for predicting the association between snoRNA and diseases based on a multi-graph SAGE network. It can comprehensively capture the characteristic information of snoRNA and diseases from different angles by integrating multiple heterogeneous networks, and effectively capture the complex relationship between the two by combining the powerful feature learning ability of the GraphSAGE architecture. By adopting small batch loader technology, the computing efficiency is improved, the training time is reduced, and the generalization performance of the model is enhanced. In addition, the system can calculate the embedding of newly introduced snoRNA or disease nodes without retraining the model, overcoming the limitations of existing models and having stronger adaptability and scalability. It not only improves the prediction accuracy of snoRNA-disease association, but also provides strong support for disease diagnosis, treatment and drug development, and has broad application prospects.
[0006] To solve the above technical problems, the present invention provides the following technical solution: a system for predicting the association between snoRNA and diseases based on a multi-graph SAGE network, the system comprising the following components:
[0007] Data acquisition and preprocessing module: This module obtains data containing snoRNA and disease association information from MNDRv3.1, performs feature extraction on snoRNA sequences, calculates snoRNA pairwise similarity using the Tanimoto coefficient, and uses a DAG-based semantic similarity method to derive disease pairwise similarity.
[0008] Heterogeneous network construction module: Using snoRNA and disease as nodes, similarity data and association networks are integrated to construct a heterogeneous graph. PyTorch Geometric functions are used to divide the graph data into training, validation, and test sets. The types of network nodes and edges are defined, and mapping functions are trained to capture structural and semantic links.
[0009] Model building and training module: Builds the SAGESDA model based on the GraphSAGE architecture, initializes the weight matrix, updates node embeddings through sampling and aggregation operations, introduces GAT's attention mechanism and bidirectional message passing to adjust the aggregation method, trains the model on multiple heterogeneous networks, introduces penalty factors to prevent overfitting, uses a point integral classifier for prediction, and optimizes the training process with the help of a mini-batch loader;
[0010] Prediction and evaluation module: The data to be predicted is input into the trained model to obtain the possibility of snoRNA-disease association. The model performance is evaluated based on the test data, mainly based on AUC, combined with AUPRC and F1 score;
[0011] Result analysis and application module: Combine biological knowledge and authoritative databases to verify the accuracy of prediction results and conduct experimental research on potential associations.
[0012] Furthermore, the data acquisition and preprocessing module obtains data containing snoRNA and disease association information from MNDRv3.1, including 27 diseases, 220 snoRNAs, and 459 known snoRNA-disease association information.
[0013] Furthermore, the data acquisition and preprocessing module uses the Tanimoto coefficient to calculate the pairwise similarity of snoRNA, and the calculation formula is: Among them, s i 、s j Refers to the vector representation of different snoRNAs, i , s j > is s i and s j The vector dot product of is calculated as where s i,k and s j,k They are s i and s j The kth element in the vector, n is the vector dimension, ||s i || is s i Euclidean distance between vectors, The similarity between snoRNAs was calculated based on the Tanimoto coefficient, which comprehensively considered the vector dot product and vector length to evaluate the pairwise similarity of snoRNAs.
[0014] Furthermore, the heterogeneous network construction module uses snoRNA and disease as nodes, integrates similarity data and association networks to construct a heterogeneous graph. Specifically, snoRNA and disease are used as two types of nodes, and the pairwise similarity data of snoRNA and disease and the association network between them are integrated to construct a heterogeneous graph G = (V, E). The association information is represented by the adjacency matrix. If there is a known association between snoRNA and disease, the corresponding position in the matrix is marked as 1, otherwise it is 0, and the Random Link of PyTorchGeometric is used. The Split function divides the graph data into 80% training set, 10% validation set and 10% test set. Among the training edges, 70% are used for message passing and 30% are used for supervision. At the same time, in order to evaluate the performance of the model, fixed negative edges are generated at a ratio of 0.1. During training, small batches of negative edges are dynamically generated at a negative sampling rate of 0.2. The heterogeneous network is further represented as G = (V, E, D, R), where V is a node set containing snoRNA and disease nodes, E is an edge set, D is a node type set, and R is an edge type set. Each node belongs to a specific node type and each edge belongs to a specific relationship type. In addition, the mapping function is trained to generate a new vector representation of the node to capture the structure and semantic links between different nodes.
[0015] Furthermore, the model construction and training module constructs a SAGESDA model based on the GraphSAGE architecture. Specifically, the SAGESDA model is constructed using the GraphSAGE architecture with an attention mechanism. GraphSAGE collects node information through an aggregator function. Each iteration includes two operations: sampling and aggregation. During sampling, a fixed number of neighbors are selected for each node based on a random walk strategy. During aggregation, the embeddings of the sampled neighbors are combined to create a new representation for the central node. The multi-head attention mechanism of the graph attention network is introduced, so that the model can selectively pay attention to different neighbors according to the importance of the input feature elements and the situations of different heads, and learn multiple representations of the graph. At the same time, the aggregation operation of GraphSAGE is adjusted to realize two-way message passing, and the neighbor node representation of the node sending the message to the node and the incoming message are used as the aggregation function. The input is processed to obtain the updated representation of the nodes in the next layer of the network. Combined with the similarity and direct interaction information of snoRNA and disease nodes, multiple heterogeneous networks are trained. The pairwise similarity data of snoRNA and disease and their known associations are mapped into three heterogeneous networks. A three-layer GraphSAGE network is used for training to generate snoRNA and disease embeddings. The three heterogeneous networks are constructed based on snoRNA sequence similarity, disease semantic similarity and known association networks respectively. The model expressiveness is enhanced by integrating different feature types. A penalty factor is introduced through the linear layer applied to the output embedding. The results are normalized to obtain the final node embeddings of snoRNAs and diseases. A point integral classifier is used for link-level prediction to identify possible snoRNA-disease associations. A mini-batch loader is used to generate a subgraph of the model input to assist model training.
[0016] Furthermore, the model construction and training module introduces the multi-head attention mechanism of the graph attention network, and the embedding representation is: in, represents the embedding representation of node u in the kth layer under the multi-head attention mechanism. σ(.) is a nonlinear activation function used to increase the nonlinear expression ability of the model. W(k) is the weight matrix of the kth layer, which is learned during the model training process and is used to adjust the weight of the features. is the embedding representation of node v in the k-1 layer and the set of neighbor nodes of node u. MEAN is the average aggregation operation, which averages the embedding of node u in the previous layer. and the embedding of its neighbor node v in the previous layer Perform average calculation, α vu is the weight factor that determines the importance of the message from node v to node u, and the calculation formula is
[0017] Furthermore, the model building and training module adjusts the aggregation operation of GraphSAGE to achieve two-way message passing. The adjustment formula is: in, represents the embedded representation of node j at the kth layer, N(i) is the set of neighbor nodes of node i, that is, all nodes directly connected to node i, is the incoming message from node j to node i in layer k, is a normalization factor, where |N(i)| represents the number of elements in set N(i), |N′(i)| represents the number of elements in set N′(i), AGGREGATE represents the aggregation operation, and N′(i) represents the set of incoming message neighbors of node i, that is, the set of nodes that send messages to node i.
[0018] Furthermore, the model building and training module introduces a penalty factor via a linear layer applied to the output embedding, which is formulated as: Where M is the output embedding, which can be normalized to obtain the final node embedding of snoRNAs and diseases respectively. λ is the penalty factor used to adjust the contribution ratio of SS and DS in M. SS is the snoRNA pairwise similarity matrix, which contains the pairwise similarity information between all snoRNAs. A is the snoRNA-disease association adjacency matrix. T is the transposed matrix of A, and DS is the disease pairwise similarity matrix, which contains the pairwise similarity information between all diseases.
[0019] Compared with existing technologies, this system for predicting snoRNA-disease associations based on multi-graph SAGE networks has the following benefits:
[0020] First, this paper integrates pairwise similarity data of snoRNAs and diseases and their association networks to construct a multi-graph heterogeneous network. It also uses the GraphSAGE architecture with an attention mechanism to comprehensively capture the characteristic information of snoRNAs and diseases, effectively capturing the complex relationship between the two, and significantly improving the prediction accuracy of snoRNA-disease associations. It achieves higher AUC, AUPRC, and F1 scores on the test set, demonstrating stronger classification ability and prediction reliability, providing biomedical researchers with more accurate prediction results.
[0021] Second, the present invention uses small batch loader technology to divide large graphs into small subgraphs for processing, thereby improving the computational efficiency of the model, reducing training time, helping the model better adapt to complex biological network data, reducing the risk of over-smoothing, and improving generalization performance. In addition, the SAGESDA model is based on the GraphSAGE architecture and can calculate the embedding of newly introduced snoRNA or disease nodes without retraining the model, enhancing the adaptability and scalability of the model. This allows the present invention to maintain stable performance when facing large-scale data and the continuous discovery of new snoRNAs and diseases, providing a more efficient and flexible prediction tool for biomedical research.
[0022] Other advantages, objects and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art based on an examination of the following or may be learned from the practice of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0024] Figure 1 Schematic diagram of the system for predicting the association between snoRNA and diseases based on the multi-graph SAGE network;
[0025] Figure 2 Flowchart of the system for predicting snoRNA-disease association based on multi-graph SAGE network. DETAILED DESCRIPTION
[0026] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0027] Example 1
[0028] Comprehensive cardiovascular disease-related snoRNA and disease data were collected from the MNDRv3.1 database, covering not only common cardiovascular diseases such as coronary heart disease and cardiomyopathy, but also rare cardiovascular disease data to enrich data diversity. 4-mer features were extracted from snoRNA sequences and converted into fixed-length numerical vectors. After that, the snoRNA pairwise similarity was calculated. The calculation formula is: Among them, si 、s j Refers to the vector representation of different snoRNAs, i , s j > is s i and s j The vector dot product of is calculated as where s i,k and s j,k They are s i and s j The kth element in the vector, n is the vector dimension, ||s i || is s i Euclidean distance between vectors, By calculating the similarity between snoRNAs, based on the Tanimoto coefficient, and taking into account the vector dot product and vector length, the pairwise similarity of snoRNAs was evaluated. The results and data of each step of the calculation were recorded in detail for subsequent analysis and verification. A semantic similarity method based on DAG was used, combined with the disease MeSHID, to determine the pairwise similarity between cardiovascular diseases. At the same time, referring to the latest cardiovascular disease research literature, the calculation results were manually reviewed and adjusted to ensure accuracy. The similarity data was integrated with known association information, represented by an adjacency matrix, and a heterogeneous graph was constructed. During the construction process, not only direct associations were considered, but also potential indirect connections established through common biological pathways or gene regulatory networks were explored.
[0029] A SAGESDA model based on the GraphSAGE architecture was built and the weight matrix was initialized. During training, neighbor nodes were sampled through random walks, and the node embedding was updated using an aggregation function with an attention mechanism. With the help of the backpropagation algorithm and the Adam optimizer, the weights were adjusted based on the error between the model prediction and the true label. A small learning rate was used in the early stage of training to avoid model oscillation. As training progressed, the learning rate was gradually adjusted to accelerate convergence. When calculating the attention weights, a variety of node features and edge attributes were comprehensively considered to enable the model to more accurately capture the complex relationships in the graph. The GraphSAGE aggregation operation was adjusted to achieve two-way message passing. The adjustment formula is: in, represents the embedded representation of node j at the kth layer, N(i) is the set of neighbor nodes of node i, that is, all nodes directly connected to node i, is the incoming message from node j to node i in layer k, is the normalization factor, where |N(i)| represents the number of elements in set N(i), |N′(i)| represents the number of elements in set N′(i), and AGGREGATE represents the aggregation operation. The neighbor node embedding and incoming message update node representation are integrated. During the bidirectional message transmission process, the message transmission path and information integration method are recorded in detail. The efficiency of the model in utilizing different information is analyzed. The data is mapped to three heterogeneous networks. The model is trained using a three-layer GraphSAGE network, and a penalty factor is introduced (the optimal value is determined through cross-validation) to prevent overfitting. The formula is: Where M is the output embedding, which can be normalized to obtain the final node embedding of snoRNAs and diseases respectively. λ is the penalty factor used to adjust the contribution ratio of SS and DS in M. SS is the snoRNA pairwise similarity matrix, which contains the pairwise similarity information between all snoRNAs. A is the snoRNA-disease association adjacency matrix. T is the transposed matrix of A, DS is the disease pairwise similarity matrix, which contains the pairwise similarity information between all diseases. Cross-validation uses multiple random partitioning of the data set to ensure the reliability of the verification results. The small batch loader is used to improve training efficiency and model generalization performance. The batch size is dynamically adjusted during training, and the optimal value is selected according to the model training situation and data characteristics to balance training speed and model performance.
[0030] Biological samples are collected from patients suspected of cardiovascular disease, and relevant snoRNA data are extracted using a variety of advanced experimental technologies such as high-throughput sequencing and quantitative PCR to ensure the accuracy and completeness of the data. This data is then input into the trained SAGESDA model to predict the possibility of an association between snoRNA and cardiovascular disease. The prediction results are verified by combining traditional diagnostic methods such as clinical symptoms, electrocardiograms, and blood tests, as well as authoritative cardiovascular disease databases. For example, the model predicts that a patient's specific snoRNA is strongly associated with coronary heart disease. Further examination reveals that the patient's related cardiovascular indicators are abnormal, such as dyslipidemia and elevated myocardial enzymes. The electrocardiogram also shows signs of myocardial ischemia. These results, taken together, support the model's predictions. In addition, the predictions can be tracked and verified over a long period of time to observe the patient's disease progression over a period of time and further evaluate the model's predictive accuracy.
[0031] Based on the model's predictions, doctors can more accurately determine a patient's risk of developing a specific cardiovascular disease. For patients predicted to be at higher risk, closer monitoring and further examinations, such as cardiac ultrasound and coronary angiography, can be arranged to detect the disease early and take intervention measures. For example, for patients predicted to be at high risk of coronary heart disease, in addition to regular cardiac ultrasound examinations, coronary CT angiography (CTA) is also arranged to more accurately assess the degree of coronary artery stenosis. The prediction results can also provide a reference for personalized treatment plans, exploring potential targeted treatments or adjusting existing therapeutic drugs for snoRNA associated with the disease. If the model predicts that a certain snoRNA is related to arrhythmia, doctors may consider using antisense oligonucleotide drugs targeting the snoRNA, or adjusting the dosage and type of existing antiarrhythmic drugs.
[0032] Example 2
[0033] We obtained data on snoRNAs and diseases related to diabetes and its complications, including type 1 diabetes, type 2 diabetes, and complications such as diabetic nephropathy and retinopathy, from the MNDRv3.1 database. We also collected relevant information from other relevant databases, such as the diabetes-related gene database and the protein-protein interaction database, to expand the data sources. We extracted 4-mer features from snoRNAs, calculated their pairwise similarities, and used a DAG-based semantic similarity method to calculate the similarity between diseases. During the calculation process, we used a variety of computational tools and algorithms for comparative verification to ensure the accuracy of the similarity calculation. We integrated the data to construct a heterogeneous graph, clarified the node (snoRNA and disease) and edge (association and similarity) information, and divided the data into training, validation, and test sets. We used a stratified sampling method to divide the data set to ensure that each subset contained data on different types of diabetes and its complications, as well as snoRNA-disease pairs with different degrees of association.
[0034] The SAGESDA model was constructed and the weights were initialized. During the training process, random walk sampling, average aggregation function, attention mechanism, and bidirectional message passing were used to optimize the model performance. The model training loss, verification indicators (such as AUC, AUPRC, and F1 scores) and other data were recorded regularly, and training curves were drawn to promptly detect training anomalies. The contribution of different similarity matrices was balanced by adjusting the penalty factor to prevent overfitting. When adjusting the penalty factor, a comprehensive judgment was made based on the model training process and generalization performance. A small batch loader was used to accelerate training and improve the generalization ability of the model. Data augmentation techniques were used in the loading process, such as random perturbation of snoRNA data and reorganization of disease data features, to increase data diversity and improve model robustness. After training, the model performance was evaluated using indicators such as AUC, AUPRC, and F1 scores, and compared with other diabetes-related prediction models to evaluate the advantages and disadvantages of the SAGESDA model.
[0035] Diabetes-related snoRNA data are input into the trained model to predict potential snoRNA-diabetes and its complications associations. The prediction results are then analyzed in depth, and combined with the pathophysiological mechanisms of diabetes and relevant research literature, snoRNAs closely related to the key pathological processes of diabetes are screened out. For example, the model predicts that certain snoRNAs are associated with core pathological links of diabetes, such as insulin resistance and abnormal blood sugar metabolism. The functions of these snoRNAs are further verified through experiments in diabetic cell models or animal models, such as knocking down or overexpressing these snoRNAs through gene editing technology, and observing changes in blood sugar metabolism indicators, insulin sensitivity, etc. in cells or animals. Bioinformatics tools are used to analyze the relationship between these snoRNAs and known diabetes drug targets to explore whether they are in the same signaling pathway or regulatory network.
[0036] Using the predicted key snoRNAs as potential drug targets, drug developers use computer-aided drug design techniques, such as molecular docking and virtual screening, to rapidly screen for potentially active compounds. They then conduct in vitro and animal experiments to evaluate the regulatory effects of these compounds on the target snoRNAs and their improvement of diabetes-related pathological indicators. By modulating the functions of these snoRNAs, they intervene in the progression of diabetes. At the same time, they re-evaluate existing drugs to explore whether regulating these snoRNAs can enhance their efficacy or reduce side effects. For example, for existing insulin sensitizers, they study whether they work by regulating the predicted snoRNAs. If so, they further optimize the drug dosage or combine it with other drugs targeting the same snoRNAs. Furthermore, the prediction results can provide a basis for patient stratification in drug clinical trials, improving the efficiency and success rate of clinical trials. Based on the patient's snoRNA expression profile and model prediction results, patients can be divided into different subgroups, and personalized clinical trial plans can be designed for different subgroups to improve the targetedness and effectiveness of drug development.
[0037] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A system for predicting the association between snoRNA and diseases based on a multi-graph SAGE network, characterized by: The system includes the following components: Data acquisition and preprocessing module: This module obtains data containing snoRNA and disease association information from MNDRv3.1, performs feature extraction on snoRNA sequences, calculates snoRNA pairwise similarity using the Tanimoto coefficient, and uses a DAG-based semantic similarity method to derive disease pairwise similarity. Heterogeneous network construction module: Using snoRNA and disease as nodes, similarity data and association networks are integrated to construct a heterogeneous graph. PyTorch Geometric functions are used to divide the graph data into training, validation, and test sets. The types of network nodes and edges are defined, and mapping functions are trained to capture structural and semantic links. Model building and training module: Builds the SAGESDA model based on the GraphSAGE architecture, initializes the weight matrix, updates node embeddings through sampling and aggregation operations, introduces GAT's attention mechanism and bidirectional message passing to adjust the aggregation method, trains the model on multiple heterogeneous networks, introduces penalty factors to prevent overfitting, uses a point integral classifier for prediction, and optimizes the training process with the help of a mini-batch loader; Prediction and evaluation module: The data to be predicted is input into the trained model to obtain the possibility of snoRNA-disease association. The model performance is evaluated based on the test data, mainly based on AUC, combined with AUPRC and F1 score; Result analysis and application module: Combine biological knowledge and authoritative databases to verify the accuracy of prediction results and conduct experimental research on potential associations.
2. The system for predicting the association between snoRNA and disease based on a multi-graph SAGE network according to claim 1, characterized in that: The data acquisition and preprocessing module obtains data containing snoRNA and disease association information from MNDRv3.1, including 27 diseases, 220 snoRNAs, and 459 known snoRNA-disease association information.
3. The system for predicting the association between snoRNA and disease based on a multi-graph SAGE network according to claim 1, characterized in that: The data acquisition and preprocessing module uses the Tanimoto coefficient to calculate the pairwise similarity of snoRNAs, and the calculation formula is: Among them, s i 、s j Refers to the vector representation of different snoRNAs, (s i , s j > is s i and s j The vector dot product of is calculated as where s i,k and s j,k They are s i and s j The kth element in the vector, n is the vector dimension, ||s i || is s i Euclidean distance between vectors, The similarity between snoRNAs was calculated based on the Tanimoto coefficient, which comprehensively considered the vector dot product and vector length to evaluate the pairwise similarity of snoRNAs.
4. The system for predicting the association between snoRNA and disease based on a multi-graph SAGE network according to claim 1, characterized in that: The heterogeneous network construction module uses snoRNA and disease as nodes, integrates similarity data and association networks to construct a heterogeneous graph. Specifically, snoRNA and disease are used as two types of nodes, and the pairwise similarity data of snoRNA and disease and the association network between them are integrated to construct a heterogeneous graph G = (V, E). The association information is represented by the adjacency matrix. If there is a known association between snoRNA and disease, the corresponding position in the matrix is marked as 1, otherwise it is 0, and the Random The LinkSplit function divides the graph data into 80% training set, 10% validation set and 10% test set. Among the training edges, 70% are used for message passing and 30% are used for supervision. At the same time, to evaluate the performance of the model, fixed negative edges are generated at a ratio of 0.
1. During training, small batches of negative edges are dynamically generated at a negative sampling rate of 0.
2. The heterogeneous network is further represented as G = (V, E, D, R), where V is a node set containing snoRNA and disease nodes, E is an edge set, D is a node type set, and R is an edge type set. Each node belongs to a specific node type and each edge belongs to a specific relationship type. In addition, the mapping function is trained to generate a new vector representation of the node to capture the structure and semantic links between different nodes.
5. The system for predicting the association between snoRNA and disease based on a multi-graph SAGE network according to claim 1, characterized in that: The model construction and training module constructs a SAGESDA model based on the GraphSAGE architecture. Specifically, the SAGESDA model is constructed using the GraphSAGE architecture with an attention mechanism. GraphSAGE collects node information through an aggregator function. Each iteration includes two operations: sampling and aggregation. During sampling, a fixed number of neighbors are selected for each node based on a random walk strategy. During aggregation, the embeddings of the sampled neighbors are combined to create a new representation for the central node. The multi-head attention mechanism of the graph attention network is introduced, so that the model can selectively focus on different neighbors according to the importance of the input feature elements and the conditions of different heads, and learn multiple representations of the graph. At the same time, the aggregation operation of GraphSAGE is adjusted to realize two-way message passing. The neighbor node representation of the node sending a message to the node and the incoming message are used as the input of the aggregation function. After processing, an updated representation of the nodes in the next layer of the network is obtained. Combined with the similarity and direct interaction information of snoRNA and disease nodes, multiple heterogeneous networks are trained. The pairwise similarity data of snoRNA and disease and their known associations are mapped into three heterogeneous networks. A three-layer GraphSAGE network is used for training to generate snoRNA and disease embeddings. The three heterogeneous networks are constructed based on snoRNA sequence similarity, disease semantic similarity and known association networks, respectively. The model expressiveness is enhanced by integrating different feature types. A penalty factor is introduced through the linear layer applied to the output embedding, and the results are normalized to obtain the final node embeddings of snoRNAs and diseases. A point integral classifier is used for link-level prediction to identify possible snoRNA-disease associations, and a mini-batch loader is used to generate a subgraph of the model input to assist model training.
6. The system for predicting the association between snoRNA and disease based on a multi-graph SAGE network according to claim 5, characterized in that: The model construction and training module introduces the multi-head attention mechanism of the graph attention network, and the embedding representation is: in, represents the embedding representation of node u in the kth layer under the multi-head attention mechanism, σ(.) is a nonlinear activation function used to increase the nonlinear expression ability of the model, W( k ) is the weight matrix of the kth layer, which is learned during the model training process and is used to adjust the weight of the features. is the embedding representation of node v in the k-1 layer and the set of neighbor nodes of node u. MEAN is the average aggregation operation, which averages the embedding of node u in the previous layer. and the embedding of its neighbor node v in the previous layer Perform average calculation, α vu is the weight factor that determines the importance of the message from node v to node u, and the calculation formula is 7. The system for predicting the association between snoRNA and diseases based on a multi-graph SAGE network according to claim 5, characterized in that: The model building and training module adjusts the aggregation operation of GraphSAGE to achieve two-way message transmission. The adjustment formula is: in, represents the embedded representation of node j at the kth layer, N(i) is the set of neighbor nodes of node i, that is, all nodes directly connected to node i, is the incoming message from node j to node i in layer k, is a normalization factor, where |N(i)| represents the number of elements in set N(i), |N′(i)| represents the number of elements in set N′(i), AGGREGATE represents the aggregation operation, and N′(i) represents the set of incoming message neighbors of node i, that is, the set of nodes that send messages to node i.
8. The system for predicting the association between snoRNA and diseases based on a multi-graph SAGE network according to claim 5, characterized in that: The model building and training module introduces a penalty factor via a linear layer applied to the output embedding, which is formulated as: Where M is the output embedding, which can be normalized to obtain the final node embedding of snoRNAs and diseases respectively. λ is the penalty factor used to adjust the contribution ratio of SS and DS in M. SS is the snoRNA pairwise similarity matrix, which contains the pairwise similarity information between all snoRNAs. A is the snoRNA-disease association adjacency matrix. T is the transposed matrix of A, and DS is the disease pairwise similarity matrix, which contains the pairwise similarity information between all diseases.