Construction method, prediction system and prediction method of circular RNA and disease association prediction model based on sharing unit

By constructing a circular RNA and disease association prediction model based on shared units and combining multi-view interaction, attention mechanism and contrastive learning, the problem of insufficient utilization of view information in existing technologies is solved, and higher performance circular RNA-disease association prediction is achieved.

CN119400252BActive Publication Date: 2025-10-10YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA

Patent Information

Application Number
CN202510012421.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-10-10
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

Existing technologies fail to effectively integrate multiple types of biological data when predicting the association between circular RNA and diseases, do not fully utilize the potential information between views, and do not consider the importance differences between views, resulting in limited prediction performance.

Method used

A shared unit-based approach is used to construct similarity networks and meta-path networks between circular RNA and diseases. Through multi-view interaction and fusion, combined with a multi-channel attention mechanism and contrastive learning strategy, feature representation and prediction performance are optimized.

Benefits of technology

It improves the accuracy and generalization ability of circular RNA-disease association prediction, significantly outperforms existing methods, and can more accurately predict complex associations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119400252B_ABST
    Figure CN119400252B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on shared unit's circular RNA and the construction method of disease association prediction model, prediction system and prediction method, and the association prediction is realized by the circular RNA and the disease association prediction model constructed, comprising: based on known association dataset constructs circular RNA meta-path network, circular RNA similarity network, disease meta-path network and disease similarity network, the feature extraction of the aforementioned network respectively obtains circular RNA similarity feature, circular RNA meta-path feature, disease similarity feature and disease meta-path feature;Similarity feature and meta-path feature are simultaneously input into shared unit;The similarity feature output by shared unit is input into multilayer perception machine to carry out the association prediction of circular RNA and disease and update model parameter.Share unit is designed for model, and the similarity network and meta-path network of circular RNA and disease are constructed, and potential cross-view information is captured in the process of multi-view feature fusion, so as to improve the prediction performance of model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer bioinformatics, and in particular relates to a method for constructing a circular RNA and disease association prediction model based on shared units, a prediction system and a prediction method. Background Art

[0002] Circular RNAs, a special class of non-coding RNA molecules, play a key role in diverse biological processes within cells. They interact with molecules such as miRNAs and proteins to regulate gene expression and signaling, thereby influencing the onset and progression of diseases. Predicting the associations between circRNAs and specific diseases can help us uncover the molecular mechanisms underlying diseases and understand their complex biological networks. This understanding is fundamental to the development of new therapies, early diagnostic tools, and personalized treatment plans. Identifying disease-associated circRNAs can lead to the development of new biomarkers that can be used for early disease detection, disease progression monitoring, and therapeutic efficacy assessment, thereby improving the precision of treatment. Furthermore, this association information may reveal new therapeutic targets and provide guidance for the development of targeted drugs, particularly for diseases where traditional treatments have limited efficacy. Furthermore, this predictive ability can help predict patient responses to specific drugs and optimize clinical treatment regimens. More importantly, understanding the role of circRNAs in disease can provide new strategies for disease prevention, such as reducing disease risk by modulating the expression of specific circRNAs. In summary, the prediction of circRNA-disease associations is not only of great significance to basic scientific research but also has profound implications for clinical practice, public health policymaking, and economic benefits. This prediction can facilitate the transition from basic research to clinical application, driving advances in medical science and ultimately improving human health.

[0003] Currently, there are mainly the following methods for predicting the association between circular RNA (circRNA) and diseases:

[0004] Network-based methods predict associations by building heterogeneous networks based on the full utilization of different types of biological data. However, how to effectively integrate multiple types of data is a problem that requires in-depth research, and there is currently no good solution.

[0005] Based on traditional machine learning methods, this method predicts circRNA-disease associations by integrating different features and combining traditional machine learning methods, such as gradient decision trees and matrix decomposition. However, this method relies on different similarity strategies when extracting features, and not all strategies are effective for the model.

[0006] A deep learning-based approach can learn and predict potential disease-associated circular RNAs. It can automatically extract high-level features contained in descriptors and accurately predict new circular RNA-disease associations. However, existing deep learning-based methods fail to fully utilize the potential information between views and do not consider the differences in view importance, resulting in limited prediction performance. Summary of the Invention

[0007] The purpose of the present invention is to address the problems existing in the prior art and to propose a method for constructing a circular RNA and disease association prediction model based on shared units, a prediction system and a prediction method.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] A method for constructing a circular RNA and disease association prediction model based on shared units, comprising:

[0010] Construct circular RNA metapathway networks, circular RNA similarity networks, disease metapathway networks, and disease similarity networks based on known association datasets;

[0011] Feature extraction was performed on the circular RNA similarity network, circular RNA metapath network, disease similarity network and disease metapath network to obtain circular RNA similarity features, circular RNA metapath features, disease similarity features and disease metapath features respectively;

[0012] CircRNA similarity features and circRNA metapath features are simultaneously input into the shared unit for multi-view interaction and fusion;

[0013] Disease similarity features and disease meta-path features are simultaneously input into the shared unit for multi-view interaction and fusion;

[0014] The circRNA similarity features and disease similarity features output by the shared unit are input into the multi-layer perceptron (MLP) to predict the association between circRNA and disease.

[0015] Based on the model loss function The model parameters are updated according to the association prediction results to enable the model to have association prediction capabilities.

[0016] In the above-mentioned method for constructing a circular RNA and disease association prediction model based on a shared unit, at least two circular RNA similarity networks are included, and the circular RNA similarity features of at least one circular RNA similarity network and the circular RNA metapath features are simultaneously input into the shared unit;

[0017] The circular RNA similarity features after the shared unit and the other circular RNA similarity features are input into the multi-channel attention mechanism, and then input into the multi-layer perceptron after being processed by the multi-channel attention mechanism;

[0018] At least two disease similarity networks are included, and the disease similarity features of at least one disease similarity network and the disease meta-path features are simultaneously input into the sharing unit;

[0019] The disease similarity view features and other disease similarity features that have passed through the shared unit are input into the multi-channel attention mechanism, and then processed by the multi-channel attention mechanism and input into the multi-layer perceptron.

[0020] In the above-mentioned method for constructing a shared unit-based circular RNA and disease association prediction model, the method also includes a comparative learning process:

[0021] Construct positive and negative sample pairs in contrastive learning, select the features of the same circular RNA / the same disease in different views as positive sample pairs, and select the features of different entities in different views as negative sample pairs. In the process of updating model parameters, the contrastive learning loss function of circular RNA is used to calculate the positive and negative sample pairs. and disease contrast learning loss function Increase the similarity between positive sample pairs and reduce the similarity between negative sample pairs;

[0022] Based on the loss function during model training Update model parameters.

[0023] In the above-mentioned method for constructing a shared unit-based circular RNA and disease association prediction model, the multi-channel attention mechanism's output of multiple disease similarity features and the disease meta-pathway features processed by the shared unit participate in the comparative learning process of the disease;

[0024] The multi-channel attention mechanism participates in the comparative learning process of circular RNA by outputting similarity features of multiple circular RNAs and the circular RNA meta-path features processed by shared units.

[0025] In the above-mentioned method for constructing a circular RNA and disease association prediction model based on shared units, two shared units are included, corresponding to the disease and the circular RNA respectively;

[0026] CircRNA similarity features and circRNA metapathway features are simultaneously input into a shared unit;

[0027] Disease similarity features and disease meta-path features are simultaneously input into another shared unit;

[0028] Includes two multi-channel attention mechanisms, corresponding to diseases and circular RNAs respectively;

[0029] The circular RNA similarity features after the shared unit and the other circular RNA similarity features are input into a multi-channel attention mechanism for processing;

[0030] The disease similarity view features and other disease similarity features that pass through the shared unit are input into another multi-channel attention mechanism for processing.

[0031] In the above-mentioned method for constructing a circular RNA and disease association prediction model based on shared units, the circular RNA similarity network includes a circular RNA function similarity network and a circular RNA GIP similarity network;

[0032] The circular RNA GIP similarity features and circular RNA metapath features extracted from the circular RNA GIP similarity network are simultaneously input into the shared unit;

[0033] The disease similarity network includes a disease semantic similarity network and a disease GIP similarity network;

[0034] The disease GIP similarity features and disease meta-path features extracted from the disease GIP similarity network are simultaneously input into the shared unit.

[0035] In the above-mentioned method for constructing a shared unit-based circular RNA and disease association prediction model, the disease semantic similarity network is constructed in the following manner:

[0036]

[0037] in: and Represents diseases and diseases ancestral disease collection;

[0038] Indicates ancestral disease right The semantic contribution of Indicates ancestral disease right semantic contribution;

[0039] The disease GIP similarity network is constructed as follows:

[0040] in: and Diseases and Column vector in the circRNA-disease association matrix;

[0041] is the normalized bandwidth parameter;

[0042] The circular RNA functional similarity network was constructed as follows:

[0043]

[0044] in: and Circular RNA and a collection of related diseases;

[0045] Indicates disease With collection Similarity scores of all diseases in ; Indicates disease With collection Similarity scores of all diseases in ;

[0046] The circular RNA GIP similarity network was constructed as follows:

[0047]

[0048] and circular RNA and row vectors in the incidence matrix; is the normalized bandwidth parameter.

[0049] In the above-mentioned method for constructing a circular RNA and disease association prediction model based on shared units, constructing the circular RNA metapathway network based on a known association dataset includes:

[0050]

[0051] The disease meta-pathway network is constructed based on the known association data set, including:

[0052]

[0053] in: is the correlation matrix of the known correlation data set; The association matrix between circular RNA and disease The transpose of .

[0054] The above association data set is in the form of an association matrix, or the association data set is converted into an association matrix.

[0055] A shared unit-based circular RNA and disease association prediction system includes a processor and a storage medium, wherein the storage medium stores a prediction model constructed by the above method, and the processor is used to execute the prediction model, that is, to execute the algorithms and calculations in the prediction model.

[0056] A method for predicting the association between circular RNA and diseases based on shared units uses a model constructed by the method described above to perform association prediction. The input of the model includes the target circular RNA and target disease whose association relationship is to be predicted, as well as the known association between the target circular RNA and the disease, and the known association between the target disease and the circular RNA. Based on the input, the model constructs a meta-path network and a similarity network of the target circular RNA, as well as a meta-path network and a similarity network of the target disease, and performs feature extraction. Finally, a multi-layer perceptron performs association prediction based on the extracted features and outputs a prediction result.

[0057] The advantages of the present invention are:

[0058] 1) This approach designs shared units for the model. By constructing a similarity network and a meta-path network between circRNAs and diseases, this approach enables information interaction and fusion between similarity views and meta-path views. By capturing potential cross-view information during multi-view feature fusion, the model's performance in predicting circRNA-disease associations is improved.

[0059] 2) This solution uses a multi-channel attention mechanism while designing shared units. It dynamically assigns weights based on the importance of features in different similarity networks, effectively integrating the features of multiple similarity networks and improving the model's ability to model complex relationships.

[0060] 3) In addition, this solution introduces a contrastive learning strategy that fully utilizes the complementary information between multiple views. By maximizing the similarity between positive samples and minimizing the similarity between negative samples, it further optimizes feature representation and improves the model's prediction accuracy and generalization ability.

[0061] 4) This solution solves the problem that the model fails to fully utilize the potential information between different views. It proposes an optimization algorithm that combines multi-view interaction and multi-channel attention mechanism. Compared with existing methods, MSMCDA shows significant advantages in the accuracy and generalization ability of circular RNA-disease association prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 Shown is a flow chart of a method for predicting the association between circular RNA and diseases based on shared units provided by an embodiment of the present invention;

[0063] Figure 2 FIG2 is a schematic diagram of a method for predicting the association between circular RNA and diseases based on shared units according to an embodiment of the present invention;

[0064] Figure 3 The comparative results of AUC, AUPR, ACC and F1 of the present application on 5 independent test data sets are shown;

[0065] Figure 4 The comparative results of AUC of the present application and the prior art on 5 training data sets are shown

[0066] Figure 5 The comparative results of the ablation experiment of the present application on the CircR2disease data set are shown;

[0067] Figure 6 The comparative results of the ablation experiment of the present application on the Circ2disease data set are shown;

[0068] Figure 7 The comparative results of the ablation experiment of the present application on the CircRNAdisease data set are shown;

[0069] Figure 8 The comparative results of the ablation experiment of the present application on the CircRDS data set are shown;

[0070] Figure 9 The comparative results of the ablation experiment of the present application on the CircR2diseasev2.0 data set are shown. DETAILED DESCRIPTION

[0071] The present scheme provides a circular RNA and disease association prediction method based on shared units, which can realize higher performance association prediction of circular RNA and disease. First, a construction method of a circular RNA and disease association prediction model based on shared units for realizing the method is proposed. Secondly, a circular RNA and disease association prediction system based on shared units is proposed, which includes a storage medium and a processor for storing and executing the aforementioned constructed model.

[0072] For convenience of description, the constructed model is referred to as MSMCDA here. The method constructs multi-view features of circular RNA and disease, combines shared units and attention mechanism, and optimizes feature representation using a contrastive learning strategy, so as to more accurately predict the potential association between circular RNA and disease. As shown in Figure 1 The method provided by the present scheme includes:

[0073] Collect and organize known association data sets of circular RNA and disease in multiple public databases;

[0074] Based on the obtained association data set, a similarity network and a meta-path network of circular RNA and a similarity network and a meta-path network of disease are constructed.

[0075] The constructed similarity network and meta-path network are input into a sharing unit, and the sharing unit is used to promote information interaction between multi-view features, so as to capture potential information between different views and enhance the information expression ability across views.

[0076] The multi-channel attention mechanism is used to adaptively adjust the importance weight of the multi-view features of circular RNA and disease, and the expression quality of each feature is optimized.

[0077] The contrast learning strategy is introduced to maximize the feature similarity of positive samples from the same view and minimize the similarity between negative sample pairs across views, further enhancing the feature representation ability.

[0078] The features of circular RNA and disease are extracted through the above processes, and the multi-layer perception MLP performs association prediction based on these features. Finally, the model parameters are updated based on the loss function According to the association prediction result, the model parameters are updated to make the model have the association prediction ability.

[0079] The trained model can be used for circular RNA-disease association prediction. Finally, the prediction results of the model are verified combined with actual biological data, and compared with other methods. The results show that the scheme is superior to other methods in each index, and has the highest overall performance.

[0080] Specifically, as shown in Figure 2 , the construction method of the similarity network and the meta-path network, and the construction process thereof are as follows:

[0081] (1) Similarity network construction

[0082] In this embodiment, the similarity network of the disease includes a disease semantic similarity network and a disease Gaussian interaction kernel (GIP) similarity network.

[0083] The disease semantic similarity network is constructed, the DOID (Disease Ontology Identifier) is used to retrieve the identifier (DOID) of each disease from the disease ontology, and the semantic similarity between diseases and is calculated :

[0084] (1)

[0085] wherein, and represent diseases and ancestral disease collection;

[0086] Indicates ancestral disease right The semantic contribution of is calculated as follows:

[0087] (2)

[0088] In the disease semantic similarity network, the weight of the edge is the semantic similarity value.

[0089] Construct a disease GIP similarity network and use Gaussian interaction kernel (GIP) similarity to calculate the similarity between diseases:

[0090] (3)

[0091] in: and Diseases and Column vector in the circRNA-disease association matrix;

[0092] To normalize the bandwidth parameter, the calculation formula is:

[0093] (4)

[0094] in, is the number of columns of the incidence matrix, is the norm of each column vector.

[0095] In the disease GIP similarity network, the edge weight is the GIP similarity value.

[0096] In this embodiment, the circular RNA similarity network includes a circular RNA function similarity network and a circular RNA GIP similarity network.

[0097] CircRNAs were calculated by and Functional similarity:

[0098] (5)

[0099] in: and Circular RNA and a collection of related diseases;

[0100] Indicates disease With collection Similarity scores for all diseases in .

[0101] In the circular RNA functional similarity network, the weight of the edge is the functional similarity value.

[0102] The GIP similarity of circular RNA was calculated as follows:

[0103] (6)

[0105] and circular RNA and row vectors in the incidence matrix;

[0106] To normalize the bandwidth parameter, the calculation formula is:

[0107] (7)

[0108] in, is the number of rows of the incidence matrix, is the norm of each row vector.

[0109] In the circular RNA GIP similarity network, the edge weight is the GIP similarity value.

[0110] (2) Meta-path construction:

[0111] Constructing a metapathway network of circular RNAs using an association matrix :

[0112] (8)

[0113] in: is the correlation matrix; is the transpose of the association matrix between circular RNA and disease.

[0114] Meta-path network In the example, the weight of the edge is the functional similarity value.

[0115] The meta-path network of the disease is also constructed using the association matrix :

[0116] (9)

[0117] Meta-path network In the example, the weight of the edge is the functional similarity value.

[0118] Furthermore, the multi-view feature interaction and fusion method of the shared unit includes the following process:

[0119] A convolution operation is performed on the circular RNA functional similarity network and the GIP similarity network respectively. The updated feature calculation process of the two networks is as follows:

[0120] (10)

[0121] in, , is the adjacency matrix of the circular RNA functional similarity network or the GIP similarity network, is the added identity matrix;

[0122] is the corresponding degree matrix; Represents learnable parameters is the activation function; Represents the circular RNA node features in the lth layer.

[0123] After a layer of GCN convolution for the disease semantic similarity network and the disease GIP similarity network, the updated features are calculated as follows:

[0124] (11)

[0125] in, , is the adjacency matrix of the semantic similarity network of the disease or the GIP similarity network, and I is the added identity matrix;

[0126] is the corresponding degree matrix; represents a learnable parameter; is the activation function; Represents the disease node features in the lth layer.

[0127] Similarly, the two meta-path networks are also convolved, and the updated features of the meta-path network are calculated as follows:

[0128] (12)

[0129] in, , is the adjacency matrix of the meta-path network, and I is the added identity matrix;

[0130] is the corresponding degree matrix; represents a learnable parameter; is the activation function; Indicates the Disease node features or circular RNA node features in the layer.

[0131] For circular RNA and disease, the obtained GIP similarity features and their meta-path features are input into the shared unit together. Four independent linear modules are used in the shared unit to perform element-by-element operations on the multi-view features, adjust the feature weights, and output the updated features. The feature update process is as follows:

[0132] (13)

[0133] (14)

[0134] in, and Represents the updated similarity view features and metapath view features of circular RNA or disease respectively; operator represents element-wise multiplication; , , and is a trainable parameter.

[0135] Furthermore, the specific method of optimizing feature weights using the multi-channel attention mechanism includes the following steps:

[0136] The multiple similarity view features of circular RNA are input into the multi-channel attention mechanism. In the multi-channel attention mechanism, the input similarity view features are first globally averaged and pooled, and then the weight coefficient is calculated. The calculation formula is as follows:

[0137] (15)

[0138] (16)

[0139] in: Indicates functional similarity features; Indicates GIP similarity features; Represents a global average pooling operation.

[0140] and is the weight matrix of the fully connected layer; Represents the ReLU activation function; Represents the Sigmoid activation function.

[0141] Using the calculated weight coefficient , weighted processing of the similarity features of circular RNA and integrated view features:

[0142] (17)

[0143] in: Represents a two-dimensional convolutional neural network used to fuse multi-view features.

[0144] The multiple similarity view features of the disease are input into the multi-channel attention mechanism. In the multi-channel attention mechanism, the input similarity view features are first globally averaged and pooled, and then the weight coefficient of the disease similarity view is calculated:

[0145] (18)

[0146] (19)

[0147] in: Represents semantic similarity features; Indicates GIP similarity features; Represents a global average pooling operation.

[0148] and is the weight matrix of the fully connected layer; Represents the ReLU activation function; Represents the Sigmoid activation function.

[0149] Using the calculated weight coefficient , weight the similarity features of the disease and integrate the features:

[0150] (20)

[0151] in: Represents a two-dimensional convolutional neural network used to fuse multi-view features.

[0152] The processing of the multilayer perceptron is as follows:

[0153] (twenty one)

[0154] in: and circular RNA and diseases The embedded features of , i.e., the corresponding circular RNA features and corresponding disease features processed by shared units, multi-channel attention mechanism and contrastive learning;

[0155] represents element-by-element multiplication; FNN is a fully connected layer used to extract features from the combined features; sigmoid is an activation function used to normalize the prediction results to the range of [0, 1], indicating circular RNA and disease The probability score of the association.

[0156] Furthermore, the specific method of contrastive learning to optimize feature representation includes the following process:

[0157] Construct positive and negative pairs for contrastive learning. Positive pairs are selected based on features of the same entity (circRNA or disease) in different views, while negative pairs are selected based on features of different entities in different views. During model updates, the similarity between positive pairs is increased, while the similarity between negative pairs is decreased.

[0158] The contrastive learning loss function for circular RNA is calculated by the following formula:

[0159] (twenty two)

[0160] in: and circular RNA features in similarity networks and metapath networks; represents the set of negative samples i≠j, sim(a,b) is the cosine similarity function, and the calculation formula is:

[0161] (twenty three)

[0162] The contrastive learning loss function for diseases has the same form as that for circular RNAs and is defined as:

[0163] (twenty four)

[0164] in: and Diseases features in similarity networks and metapath networks;

[0165] represents the set of negative samples i≠j, and sim(a,b) is the cosine similarity function.

[0166] Specifically, this model adopts the binary cross entropy loss function To optimize:

[0167] (25)

[0168] Where: Y + and Y - Represent the positive samples and negative samples in the training data respectively;

[0169] when , the actual label ;when , the actual label .

[0170] The final loss function Combining binary cross entropy loss and contrastive learning loss:

[0171]

[0172] Furthermore, this scheme uses 5-fold cross-validation to optimize the classification performance of the model. First, the training data is divided into five subsets, one subset is selected as the test set, and the remaining four subsets are used as training sets. The process is repeated five times, and the average index is calculated to evaluate the model performance.

[0173] After training the circular RNA-disease association prediction model, the process of using this model for prediction is as follows:

[0174] The model is fed with the target circular RNA and target disease whose association is to be predicted, as well as the known association between the target circular RNA and the disease, and the known association between the target disease and the circular RNA. For example, taking circular RNA1 and disease 1 as an example, assuming that the association between circular RNA1 and disease 1 is to be predicted, and disease 1 is known to be associated with circular RNA2 and circular RNA3, and circular RNA1 is known to be associated with disease 2 and disease 3, then the following association matrix can be input:

[0175] Disease 1 Disease 2 Disease 3 circular RNA1 To be tested 1 1 circular RNA2 1 0 1 circular RNA3 1 0 1

[0176] In the table, 1 indicates that the corresponding circular RNA is associated with the corresponding disease, and 0 indicates an unknown association.

[0177] The model constructs the meta-path network and similarity network of the target circular RNA based on the input, as well as the meta-path network and similarity network of the target disease and performs feature extraction. Finally, the multi-layer perceptron is used to Figure 2 The features extracted in step B are used for association prediction and the prediction results are output.

[0178] The proposed method, MSMCDA, differs from traditional circular RNA and disease association prediction networks by introducing shared units to promote feature interaction between different views, enabling complementary integration during the fusion process while capturing potential cross-view information. Furthermore, this approach utilizes an attention mechanism to assign different importance weights to features from different similar networks, thereby enhancing the model's adaptability to multi-view information. Finally, MSMCDA incorporates a contrastive learning approach to fully utilize cross-view complementary information and further improve feature representation capabilities.

[0179] In order to verify the effectiveness and prediction accuracy of the proposed method, this scheme was used for verification on different test sets, and the performance was compared with six currently advanced methods, including AMHMDA, MDGF-MCEC, Bi-SGTAR, GMNN2CD, GraphCDA, and DMFCDA, on the same dataset. Ablation experiments were also conducted:

[0180] In each figure, ACC represents the prediction accuracy, that is, the ratio of the number of samples correctly classified by the classifier to the total number of samples. F1-score comprehensively evaluates the accuracy and recall of the classifier. AUC is the area under the ROC curve, which objectively expresses the classification ability of the model. AUPR is the area under the PR curve. In the task of predicting the association between circular RNA and disease, associated circular RNA and disease pairs are positive samples, and unassociated circular RNA and disease pairs are negative samples. Precision represents the proportion of all samples predicted as positive that are actually positive, measuring how many of the samples predicted as positive by the model are correct. Recall represents the proportion of all samples that are actually positive that are correctly predicted by the model.

[0181] Figure 3 This is the result of validating this method on 5 independent test set data of different sizes. From the results, MSMCDA has high indicators on all 5 data sets, proving that the method proposed in this scheme has high generalization ability for data sets of different sizes.

[0182] Figure 4 The performance of six advanced methods, including AMHMDA, MDGF-MCEC, Bi-SGTAR, GMNN2CD, GraphCDA, and DMFCDA, was compared on the same dataset. As can be seen from the figure, MSMCDA achieved the highest AUC value across all datasets, outperforming the other compared methods, demonstrating the effectiveness of the proposed method.

[0183] Figure 5-Figure 9 The following are schematic diagrams comparing the results of ablation experiments on the datasets CircR2disease, Circ2disease, CircRNAdisease, CircRDS, and CircR2diseasev2.0. The models include:

[0184] MSMCDA, a circular RNA-disease association prediction model using multi-view shared units and contrastive learning strategy, also known as the method of this scheme;

[0185] MSMCDA-noatten: The model without the attention mechanism of MSMCDA, where all views are considered equally important.

[0186] MSMCDA-noct: MSMCDA model without contrastive learning, directly using the aggregated similarity features for prediction.

[0187] MSMCDA-no share: MSMCDA model without sharing unit, directly applying two layers of GCN on the original input features without any feature sharing process.

[0188] By Figure 5-Figure 9 It can be seen that MSMCDA is superior to other methods in various indicators on the above five data sets. It can be seen that the present scheme still has outstanding advantages and higher prediction performance compared with the current mainstream and advanced models.

[0189] The specific embodiments described herein are merely illustrative of the spirit of the present application. Those skilled in the art of the present application can make various modifications or supplements to the described specific embodiments or replace them with similar ways, but will not deviate from the spirit of the present application or exceed the scope defined by the appended claims.

Claims

1. A method for constructing a circular RNA and disease association prediction model based on shared units, characterized in that: include: Construct circular RNA metapathway networks, circular RNA similarity networks, disease metapathway networks, and disease similarity networks based on known association datasets: The circular RNA metapathway network is constructed based on a known association dataset, including: The disease meta-pathway network is constructed based on the known association data set, including: in: is the correlation matrix of the known correlation data set; The association matrix between circular RNA and disease The transpose of Feature extraction was performed on the circular RNA similarity network, circular RNA metapath network, disease similarity network and disease metapath network to obtain circular RNA similarity features, circular RNA metapath features, disease similarity features and disease metapath features respectively; The circular RNA similarity network includes a circular RNA functional similarity network and a circular RNA GIP similarity network; the circular RNA GIP similarity features and the circular RNA metapath features extracted from the circular RNA GIP similarity network are simultaneously input into a shared unit for multi-view interaction; The disease similarity network includes a disease semantic similarity network and a disease GIP similarity network; The disease GIP similarity features and disease meta-path features extracted from the disease GIP similarity network are simultaneously input into a shared unit for multi-view interaction; The circular RNA similarity features processed by the shared unit and the other circular RNA similarity features are input into the multi-channel attention mechanism, and then input into the multi-layer perceptron after fusion processing by the multi-channel attention mechanism; The disease similarity view features and other disease similarity features that have passed through the shared unit are input into the multi-channel attention mechanism, and after fusion processing by the multi-channel attention mechanism, are input into the multi-layer perceptron. Multilayer perceptron predicts the association between circular RNA and disease based on input; Based on the model loss function The model parameters are updated according to the association prediction results to enable the model to have association prediction capabilities.

2. The method for constructing a shared unit-based circular RNA and disease association prediction model according to claim 1, wherein: For circular RNA and disease, the obtained GIP similarity features and their meta-path features are input into the shared unit together. Four independent linear modules are used in the shared unit to perform element-by-element operations on the multi-view features, adjust the feature weights, and output the updated features. The feature update process is as follows: (13) (14) in, and Represents the updated similarity view features and metapath view features of circular RNA or disease respectively; operator represents element-wise multiplication; , , and is a trainable parameter.

3. The method for constructing a shared unit-based circular RNA and disease association prediction model according to claim 2, wherein: This method also includes a contrastive learning process: Construct positive and negative sample pairs in contrastive learning, select the features of the same circular RNA / the same disease in different views as positive sample pairs, and select the features of different entities in different views as negative sample pairs. In the process of updating model parameters, the contrastive learning loss function of circular RNA is used to calculate the positive and negative sample pairs. and disease contrast learning loss function Increase the similarity between positive sample pairs and reduce the similarity between negative sample pairs; Based on the loss function during model training Update model parameters.

4. The method for constructing a circular RNA and disease association prediction model based on shared units according to claim 3, characterized in that The multi-channel attention mechanism outputs similar features of multiple diseases and the disease meta-path features processed by shared units participate in the comparative learning process of diseases; The multi-channel attention mechanism participates in the comparative learning process of circular RNA by outputting similarity features of multiple circular RNAs and RNA meta-path features processed by shared units.

5. The method for constructing a circular RNA and disease association prediction model based on shared units according to claim 4, characterized in that: The disease semantic similarity network is constructed as follows: in: and Represents diseases and diseases ancestral disease collection; Indicates ancestral disease right The semantic contribution of Indicates ancestral disease right semantic contribution; The disease GIP similarity network is constructed as follows: in: and Diseases and Column vector in the circRNA-disease association matrix; is the normalized bandwidth parameter; The circular RNA functional similarity network was constructed as follows: in: and Circular RNA and a collection of related diseases; Indicates disease With collection Similarity scores of all diseases in ; Indicates disease With collection Similarity scores of all diseases in ; The circular RNA GIP similarity network was constructed as follows: and circular RNA and row vectors in the incidence matrix; is the normalized bandwidth parameter.

6. A circular RNA and disease association prediction system based on shared units, characterized in that: It includes a processor and a storage medium, wherein the storage medium stores a prediction model constructed by the method described in any one of claims 1 to 5, and the processor is used to execute the prediction model.

7. A method for predicting the association between circular RNA and disease based on shared units, characterized in that: Association prediction is performed using a model constructed using the method described in any one of claims 1 to 5. The input of the model includes the target circular RNA and target disease whose association relationship is to be predicted, as well as the known association between the target circular RNA and the disease, and the known association between the target disease and the circular RNA. Based on the input, the model constructs a meta-path network and a similarity network of the target circular RNA, as well as a meta-path network and a similarity network of the target disease, and performs feature extraction. Finally, a multi-layer perceptron performs association prediction based on the extracted features and outputs a prediction result.

Citation Information

Patent Citations

  • Disease-related circular RNA (Ribonucleic Acid) identification method based on graph attention

    CN114944192A

  • Ring RNA-disease association prediction method and device based on weighted graph attention and heterogeneous graph neural network, and medium

    CN115798730A

  • Drug recommendation method based on Siamese contrast learning network

    CN118398240A

Cited By

  • Annular RNA and drug association prediction method and system

    CN121237229A

  • Circular RNA and drug association prediction method and system

    CN121237229B