CircRNA-disease association prediction model based on MAML and GCN

By combining the circRNA-disease association prediction model of MAML and GCN, a comprehensive similarity matrix was constructed and multi-view feature fusion was carried out, which solved the problem of insufficient feature extraction in circRNA-disease association prediction, and achieved higher prediction accuracy and generalization ability, which was suitable for early diagnosis and precision medicine.

CN120412754APending Publication Date: 2025-08-01XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510605565.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing circRNA-disease association prediction models have the problem of insufficient feature extraction when dealing with nonlinear relationships and large-scale data, and are insufficiently adaptable in the case of scarce or imbalance of data, making it difficult to fully capture the key features between circRNA and disease.

Method used

The circRNA-disease association prediction model based on MAML and GCN is adopted. By constructing a comprehensive similarity matrix, combining the GCN layer with multi-view feature fusion and the MAML structure of the double-layer internal circulation, the positive and negative sample ratio is dynamically adjusted, the data balance is used using the RandomOverSample method, and the model parameters are updated through internal and external circulation to achieve correlation prediction.

Benefits of technology

The feature expression diversity and accuracy of association prediction of circRNA and disease relationships are improved, and good generalization ability and adaptability across data sets are shown, and high accuracy rates are achieved on multiple benchmark data sets, especially in miRNA disease association predictions on HMDD_v2.0 and HMDD_v3.2 data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412754A_ABST
    Figure CN120412754A_ABST
Patent Text Reader

Abstract

The invention relates to a circRNA (Ribonucleic Acid)-disease association prediction model based on MAML (Maximum Amplified Markup Language) and GCN ( The invention discloses a circRNA-disease association prediction model based on MAML and GCN. The circRNA-disease association prediction model comprises a similarity information module, a prediction module and a prediction module, wherein the similarity information module is used for constructing the comprehensive similarity of circRNA and diseases; the data preprocessing module comprises data standardization, feature interaction and dynamic balance data; and the association prediction module comprises a GCN layer with multi-view feature fusion, an MAML formed by using double-layer internal circulation and single-layer external circulation, the GCN layer as a basic model and the MAML formed by the double-layer internal circulation and the single-layer external circulation as a meta-learning model, and the GCN layer and the MAML are fused to realize association prediction. According to the circRNA-disease association prediction model based on the MAML and the GCN, the GCN layer with multi-view feature fusion is used, a double-layer internal circulation structure is used on the basis of the MAML, the accuracy and generalization ability are improved, and it is indicated that the circRNA-disease association prediction model has good application prospects in circRNA-related diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biomedical data processing, and particularly relates to a circRNA-disease association prediction model based on MAML and GCN. Background Art

[0002] Circular RNAs (circRNAs) are a class of non-coding RNAs with special circular structures. They are formed by back-splicing, are not easily degraded by RNA exonucleases, and have high stability. Although the specific functions of circRNAs in cells have not been fully elucidated, existing studies have shown that they can act as "sponges" for miRNAs to participate in processes such as gene transcription and splicing. In addition, circRNAs are closely related to the occurrence and development of various diseases, providing new targets for disease diagnosis and treatment. Currently, there are various circRNA-related databases, such as circad, circ2Disease, and CircR2Cancer. Many researchers have gradually established and improved these databases by manually collecting and biologically verifying the relationships between circRNAs and diseases. However, manual experiments require a large amount of time and effort, involving complex designs, meticulous operations, and long-term data analysis, all of which can affect the research results. This method not only increases the time and effort costs but also reduces the research efficiency. Therefore, scientific researchers are developing new prediction algorithms and experimental verification methods by integrating various data resources to reveal the association mechanism between circRNAs and diseases. In recent years, breakthroughs have been made in circRNA-disease association prediction models developed in matrix factorization, machine learning, and deep learning methods.

[0003] In the research on the prediction of the association between circRNAs and diseases, matrix factorization techniques are used to integrate data from different sources, such as gene-disease association data, gene-gene interaction data, and circRNA and gene expression data. For example, the DWNMF model, the RNMFLP model, etc. However, with the increase in data volume and complexity, traditional matrix factorization methods still face challenges in dealing with non-linear relationships and large-scale data. Therefore, more and more research has begun to combine machine learning techniques to more effectively capture complex patterns and potential relationships in the data. For example, the MLCDA model combines machine learning techniques, the XGBCDA model, a circRNA-disease association prediction method based on a multiple heterogeneous network, the GBDTCDA model based on gradient boosting decision trees to predict the association between circRNAs and diseases, the RWRKNN model that combines restart random walk and the K-nearest neighbor method, the DCDA model, the MAMLCDA model, etc. However, these methods may be limited by scarce or unbalanced data in practical applications, or there are problems such as insufficient adaptability and failure to fully capture the key features between circRNAs and diseases.

[0004] To solve these problems, the present invention proposes a circRNA-disease association prediction model based on MAML and GCN, which is a method that combines MAML and GCN, and this circRNA-disease association prediction model is the BIMLGCDA model. Summary of the Invention

[0005] The object of the present invention is to provide a circRNA-disease association prediction model based on MAML and GCN, which solves the problem of insufficient extraction of key problem features in circRNA-disease association prediction and can better capture the relationship between circRNAs and diseases.

[0006] To achieve the above object, the technical solution adopted is:

[0007] A circRNA-disease association prediction model based on MAML and GCN, comprising:

[0008] Similarity information module: constructing the comprehensive similarity of circRNAs and diseases;

[0009] Data preprocessing module, including: data standardization, feature interaction, and dynamic data balancing;

[0010] Association prediction module, including: a GCN layer for multi-perspective feature fusion, a MAML composed of a double-layer inner loop and a single-layer outer loop, using the GCN layer as the base model and the MAML composed of the double-layer inner loop and the single-layer outer loop as the meta-learning model, and fusing them to achieve association prediction.

[0011] Furthermore, the comprehensive similarity is obtained by integrating multiple similarity information between circRNA and diseases, including: GIP core similarity, sequence similarity, functional similarity, disease semantic similarity and comprehensive similarity of GIP core similarity.

[0012] Furthermore, the formula for the comprehensive similarity of the GIP kernel similarity is as follows:

[0013]

[0014] Where, MixDS(d i ,d j ) indicates disease d i and disease j The comprehensive similarity between them is of dimension MixDS∈R n ×n , n is the number of diseases. When the disease DO semantic similarity is equal to 0, the value of the disease's GIP kernel similarity is taken. Otherwise, the disease DO semantic similarity is taken.

[0015]

[0016] Where, MixCS(c i ,c j ) represents the comprehensive similarity of circRNA, MixCS∈R m×m ,m is the number of circRNAs. When the CFS value of circRNA is equal to 0, the similarity of the GIP core of circRNA is taken. Otherwise, the CFS value is taken.

[0017] Furthermore, in the dynamic balance data, the RandomOverSample method is used to dynamically adjust the ratio of positive and negative samples. The formula is as follows:

[0018]

[0019] N co pi es =N minority,target -N minority ;

[0020] Where N majority is the number of majority class samples, N minority is the number of minority class samples, and sampling_strategy is the oversampling ratio.

[0021] Furthermore, the GCN layer for multi-view feature fusion includes an input layer, a feature fusion layer, a fully connected layer, and an output layer;

[0022] Among them, the feature fusion layer respectively convolves the feature matrices MixCS, MixDSE, and CSDSE to obtain H c , H d and H cd , and the formula is as follows:

[0023]

[0024] In the formula, is the normalized adjacency matrix after adding self-loops, and W c , W d and W cd respectively represent the weight matrices of MixCS, MixDSE, and CSDSE, and σ is the activation function;

[0025] The feature fusion layer performs feature fusion operations, and the formula is as follows:

[0026] F = σ(α·H c + β·H d + (1 - α - β)·H cd );

[0027] In the formula, α and β are learnable parameters used to control the proportions of H c , H d and H cd , and σ is the activation function;

[0028] In the fully connected layer, Dropout is added to the fused features, and the formula is: F = Dropout(F).

[0029] Furthermore, in the MAML with a double-layer inner loop structure, the first layer optimizes the disease feature embedding, and the second layer fine-tunes the circRNA feature embedding to gradually complete feature adaptation. The formula is as follows:

[0030]

[0031]

[0032] In the formula, θ' d is the parameter related to the disease, θ is the initial parameter of the model, α is the learning rate, and L d is the loss function, is the gradient of the loss function L d with respect to the parameter θ; similarly, is the gradient of the loss function L c with respect to the parameter θ' d ;

[0033] The formula for the outer loop update method is as follows:

[0034]

[0035] In the formula, β is the learning rate, is the meta-loss function L that sums over all τ summations meta Gradient of the parameter θ.

[0036] Furthermore, in the association prediction module, the inner loop is independently performed for each task, and is quickly updated by creating a model copy and using task data; after the inner loop terminates, it enters the outer loop stage;

[0037] Among them, the inner loop update calculates the loss using BCELoss and optimizes the parameters of the model copy through gradient descent;

[0038] The formula of BCELoss is as follows:

[0039]

[0040] In the formula, N is the total number of circRNA-disease samples, and y i represents the true label of the i-th circRNA-disease sample, and p i represents the probability that the BIMLGCDA model predicts the i-th sample as a positive class.

[0041] The second object of the present invention is to provide a circRNA-disease association prediction device based on MAML and GCN, which adopts the above-mentioned circRNA-disease association prediction model.

[0042] The third object of the present invention is to provide a computer device that can implement the operation of the above-mentioned drug-disease association prediction model.

[0043] In order to achieve the above object, the technical solution adopted is:

[0044] A computer device includes a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the operation of the above-mentioned circRNA-disease association prediction model.

[0045] The fourth object of the present invention is to provide a computer-readable storage medium that can implement the operation of the above-mentioned drug-disease association prediction model.

[0046] In order to achieve the above object, the technical solution adopted is:

[0047] A computer-readable storage medium stores computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the operation of the above-mentioned drug-disease association prediction model.

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0049] Circular RNA (circRNA) is a non-coding RNA with a closed structure, which is known to participate in many biological processes and is related to the occurrence and development of various diseases. Due to the important role of circRNA in diseases, predicting its relationship with diseases is very helpful for early diagnosis and treatment. The present invention proposes a new method - a circRNA-disease association prediction model based on MAML and GCN, which is BIMLGCDA based on model-agnostic meta-learning (MAML) and graph convolutional network (GCN). The advantages of the technical solution of the present invention are as follows:

[0050] 1. The present invention considers various similarities of circRNA, including Gaussian Interaction Profile (GIP) kernel similarity, sequence similarity, and functional similarity, and constructs a comprehensive similarity for diseases based on Disease Ontology (DO) semantic similarity and GIP kernel similarity, which is beneficial to improving the diversity of feature expression and the accuracy of association prediction.

[0051] 2. In order to better capture the relationship between circRNA and diseases, the present invention adopts a GCN layer with multi-perspective feature fusion, and uses a double-layer inner loop structure based on MAML to achieve task-level optimization in the outer loop.

[0052] 3. The technical solution of the present invention is subjected to five-fold cross-validation on four benchmark datasets, and the method shows good results. In order to verify the generalization ability of the improved meta-learning, the present invention also predicts miRNA-disease associations on the HMDD_v2.0 and HMDD_v3.2 datasets, and achieves accuracies of 0.9866 and 0.9837 respectively. The high accuracy and generalization ability of the method indicate its good application prospect in the identification of circRNA-related diseases, and can provide help for disease early screening devices and precision medicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a flowchart of BIMLGCDA;

[0054] Figure 2 It is a flowchart of BIML;

[0055] Figure 3 The effects of different parameters on the performance of the BIML model, where (a) the influence of the inner-loop learning rate on the overall performance of the model, (b) the influence of the outer-loop learning rate on the overall performance of the model, (c) the change of the learning rate in the model training and evaluation stages, (d) the influence of the GCN hidden layer dimension on the model performance;

[0056] Figure 4 Figures (a) and (b) in the middle show the performance on the dataset HMDD_v2.0, and Figures (c) and (d) show the performance on the dataset HMDD_v3.2;

[0057] Figure 5 The comparative experiments of the BIMLGCDA model on the datasets circ2Disease (a), circR2Disease (b), and circRNADisease (c). Detailed implementation manners

[0058] In order to further elaborate a circRNA-disease association prediction model based on MAML and GCN according to the present invention and achieve the expected invention purpose, the following combines preferred embodiments to detail the specific implementation manners, structures, features and effects of a circRNA-disease association prediction model based on MAML and GCN proposed according to the present invention. In the following description, different "one embodiment" or "embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0059] Before elaborating in detail a circRNA-disease association prediction model based on MAML and GCN according to the present invention, it is necessary to further explain the relevant prior art mentioned in the present invention to achieve better results.

[0060] In recent years, there have been breakthroughs in the circRNA disease association prediction models developed in terms of matrix factorization, machine learning and deep learning methods.

[0061] In the research on predicting the association between circRNAs and diseases, matrix factorization techniques are used to integrate data from different sources, such as gene-disease association data, gene-gene interaction data, and circRNA and gene expression data. For example, the DWNMF model combines deep walk and non-negative matrix factorization. It calculates the similarity between circRNAs and diseases in the network through deep walk, while incorporating functional and semantic similarities, and uses the weighted k-nearest neighbor algorithm to preprocess the network, thereby adjusting the non-negative association relationship. In addition, the researchers also introduced the 1-norm, bi-graph regularization, and Frobenius norm, etc., to improve the prediction accuracy. Based on the DWNMF model, the RNMFLP model further improved the prediction method. It combines non-negative matrix factorization and label propagation, and reduces the false negative impact by optimizing the adjacency matrix, thereby identifying more accurate circRNA-disease association pairs. These two methods have different focuses, but both have made improvements in data processing and error reduction, complementing each other. The NMFCDA model combines randomized neural network pseudo-inverse learning and non-negative matrix factorization, integrates circRNA sequence, disease semantics, and GIP kernel similarity information, and uses randomized neural network pseudo-inverse learning to optimize the solution process. This method adds a more complex learning mechanism on the basis of the previous two models, making the prediction more accurate. The NMFMSN model is optimized in capturing similarity, processing sparse noise data, and utilizing multi-angle biological information. This model calculates the similarity between circRNA sequence data and disease semantic information by optimizing similarity capture, sparse noise data processing, and multi-angle biological information fusion. By predicting the missing connections using the interactions between neighboring circRNAs and diseases, and reconstructing the association network, and finally integrating these similarity networks into the non-negative matrix factorization framework, the model's ability to process complex data is enhanced, and potential circRNA-disease associations are revealed. These methods are gradually improved on their respective bases, improving data processing efficiency and prediction accuracy. Each method has its unique advantages and makes up for the deficiencies of other methods. However, with the increase in data volume and complexity, traditional matrix factorization methods still face challenges in dealing with non-linear relationships and large-scale data. Therefore, more and more research has begun to combine machine learning techniques to more effectively capture complex patterns and potential relationships in the data.

[0062] The MLCDA model proposed by Wang et al. combines machine learning techniques and can effectively predict potential unknown associations by using circRNA sequences and disease ontology information. Shen et al. proposed a circRNA-disease association prediction method XGBCDA based on a multiple heterogeneous network. This method integrates a circRNA similarity matrix, a disease similarity matrix, and a circRNA-disease association matrix, extracts statistical features and graph theory features, inputs them into an XGBoost classifier for training, and represents potential features through the tree structure and leaf node index learned by the XGBoost model. Finally, combining the potential features learned by XGBoost with the original features, a final model is trained to predict the association between circRNAs and diseases. In addition, Lei et al. used the GBDTCDA model based on gradient boosting decision trees to predict the association between circRNAs and diseases. This model was evaluated by leave-one-out cross-validation and showed good prediction performance. They also proposed an RWRKNN model that combines restart random walk and the K-nearest neighbor method. By weighting the global network features through the random walk algorithm and then using K-nearest neighbor classification, a prediction score is generated for each pair of circRNA-disease. With the improvement of computing power, deep learning-based methods have also begun to be widely used in circRNA-disease association prediction. These methods further improve the prediction accuracy by automatically learning more complex features and patterns.

[0063] The DCDA model combines a feed-forward neural network and a denoising autoencoder to effectively identify disease-related circRNAs. The iGRLCDA model uses a graph convolutional network (GCN) and a graph decomposition deep learning model to extract features and predicts the association between circRNAs and diseases through a random forest. In addition, the LGCDA model alleviates the data sparsity problem by generating k-hop closed subgraphs and combines a graph neural network (GNN) and cosine similarity to accurately capture the global features of circRNAs and diseases. The Bi-SGTAR model splits the adjacency matrix into two perspectives, introduces a sparse gating encoder to evaluate the credibility of known circRNA-disease associations, and evaluates the probability of true associations through an encode-reconstruct-regress framework. Similarly, the KGRACDA model combines the explicit structure and implicit embedding information of the knowledge graph, optimizes the attention mechanism to extract local depth features, and thus achieves accurate prediction of circRNA-disease associations. In terms of feature extraction, the NSECDA model regards the circRNA sequence as a biological language, uses natural language processing techniques to analyze its deep meaning, fuses disease characteristics and GIP kernel characteristics, captures important features with a graph attention network (GAT), and finally uses a rotation forest algorithm for prediction. The GGCDA model uses a multi-head attention mechanism to assign weights to different features, combines GCN to aggregate adjacent node and self-features, generates feature representations from multiple perspectives, and finally realizes prediction through a multi-layer fully connected neural network. The MNMDCDA model predicts the association between circRNAs and diseases through a high-order GCN and a deep neural network, using a multi-source similarity network and high-order neighborhood information. The KGETCDA model constructs a biological knowledge graph, combines Transformer representation learning and an attention mechanism to extract high-order features, and makes predictions through a multi-layer perceptron. Although existing models have achieved certain results on specific datasets, they usually require a large amount of labeled data, which may be limited by data scarcity or imbalance in practical applications. To address these challenges, meta-learning has received extensive attention in recent years and has shown potential in multiple fields. The MAMLCDA model proposed by Tian et al. combines model-agnostic meta-learning (MAML) with a convolutional neural network (CNN). By integrating the similarity features of circRNAs and diseases, it classifies samples using k-means clustering, transforms the features into images through probabilistic principal component analysis dimensions, and then processes them through MAML-CNN. This method effectively reduces the dependence on a large amount of labeled data and improves performance in the case of data scarcity. However, the engineering complexity of the MAMLCDA model is relatively high, and its adaptability is insufficient when dealing with large differences in features within a single layer, failing to fully capture the key features between circRNAs and diseases. To solve these problems, the present invention proposes a method combining MAML and GCN, called BIMLGCDA.At present, the method based on MAML combined with GCN is the first attempt in the field of circRNA-disease association prediction.

[0064] After understanding the relevant prior arts mentioned in the present invention, the following will further introduce in detail a circRNA-disease association prediction model based on MAML and GCN of the present invention in combination with specific embodiments:

[0065] Having understood the importance of the stability and extensive biological functions of circRNAs in the occurrence and development of diseases, especially in cancers, cardiovascular diseases and neurological diseases, the abnormal expression of circRNAs is closely related to disease progression. Based on this background, the present invention proposes a circRNA-disease association prediction model and method - BIMLGCDA based on the model MAML and the graph convolutional network GCN, for exploring the association between circRNAs and diseases, so as to provide a new theoretical basis and practical guidance for the formulation of early diagnosis and treatment strategies. This technical solution constructs a comprehensive similarity matrix by integrating various similarity information of circRNAs and diseases, and combines a GCN layer with multi-perspective feature fusion and a MAML structure with a double-layer inner loop, effectively improving the prediction performance. In the five-fold cross-validation experiment on four benchmark datasets, BIMLGCDA shows relatively excellent performance in multiple evaluation metrics such as F1, AUC, ACC, AUPR and Recall, demonstrating strong prediction ability. In addition, the verification experiments of BIMLGCDA on the miRNA datasets HMDDv2.0 and HMDDv3.2 further prove its good generalization ability and cross-dataset adaptability. The comparison results with existing methods show that BIMLGCDA has relatively balanced and superior performance on multiple datasets, is better than or close to other advanced models, verifying its potential in circRNA-disease association prediction. The specific embodiments are as follows:

[0066] Example 1.

[0067] The specific operation steps are as follows:

[0068] Method A

[0069] Overview of the BIMLGCD model: The core of this model lies in the combination of bilayer inner-loop meta-learning (BIML) and multi-view GCN, aiming to address the key issue of insufficient feature extraction in circRNA-disease association prediction. First, regarding the data imbalance problem, the model adopts the RandomOverSample method to dynamically adjust the ratio of positive and negative samples, ensuring that the model can fully learn the features of minority-class samples during training. Second, to capture the complex relationships between circRNAs and diseases, the model introduces multi-view GCN. By fusing various features such as the sequence similarity, functional similarity of circRNAs, and semantic similarity of diseases, a comprehensive similarity matrix is constructed, thereby enhancing the comprehensiveness of feature expression.

[0070] The BIMLGCDA model mainly includes the following modules:

[0071] (1) Similarity information module: Calculation of the comprehensive similarity between circRNAs and diseases.

[0072] (2) Data preprocessing module, including: data standardization, feature interaction, and dynamic data balancing, which are used to optimize the quality of input data, address the data imbalance problem, and enhance the model's prediction ability for circRNA-disease associations.

[0073] (3) Association prediction module, including BIML composed of a bilayer inner-loop and a single-layer outer-loop and a multi-layer GCN module. Taking GCN as the basic model and MAML composed of a bilayer inner-loop and a single-layer outer-loop as the meta-learning model, they are fused to achieve association prediction.

[0074] The flowchart of the BIMLGCDA is as Figure 1 shown. Through the collaborative action of these modules, the BIMLGCDA model can effectively predict the associations between circRNAs and diseases, providing a theoretical basis for the early diagnosis and treatment of diseases. Specifically:

[0075] I. Similarity information module

[0076] Similarity calculation, including:

[0077] (1) Human CDAs

[0078] According to the characteristics of circRNA and disease data, considering that a single similarity is not sufficient to fully characterize the features of each sample, in this embodiment, the association data between circRNA and diseases is collected from multiple public databases and multiple similarities are integrated to fully characterize circRNA and diseases. In this embodiment, the known human circRNA-disease associations are extracted from four known databases: LncRNADisease, circRNADisease, circR2Disease, and circ2disease. After removing redundancy, the dataset shown in Table 1 below is finally obtained.

[0079] Table 1 Dataset

[0080]

[0081] In this embodiment, the association matrix of circRNA-disease is defined as CDM ∈ R m×n , where m represents the number of circRNAs and n represents the number of diseases. When circRNA c i is associated with disease d j then otherwise

[0082] (2) GIP kernel similarity

[0083] Based on the known association information between diseases and circRNAs, in this embodiment, an association matrix is constructed and the Gaussian kernel method is used to calculate the similarity between circRNAs. The core idea of this method is that if two circRNAs are associated with the same disease, they may be functionally similar and thus have a higher similarity. The specific calculation formula is as follows:

[0084]

[0085]

[0086] In the formula, CG(c i , c j ) is the GIP kernel similarity between circRNA c i and circRNA c j , L(c i ) and L(c j ) are the row vectors of the circRNA-disease association matrix respectively, and λ c is the bandwidth parameter, which is mainly used to normalize the scale of similarity calculation.

[0087] Similarly, the GIP kernel similarity of diseases is as shown in the formula:

[0088]

[0089]

[0090] In the formula, DG(d i , d j ) represents the GIP kernel similarity between disease d i and disease d j . H(d i ) and H(d j ) are respectively the column vectors of the circRNA-disease association matrix, and λ d is the bandwidth parameter.

[0091] (3) Disease semantic similarity

[0092] To measure the semantic similarity between diseases, in this embodiment, the disease ontology identifier (DOID) numbers of diseases are obtained from the Disease Ontology (DO, https: / / disease-ontology.org) database, and the doSim in DOSE (R package) based on DO is used to calculate the semantic similarity between diseases according to Wang's method. The similarity calculation method is as follows:

[0093]

[0094] In the formula, S W (d1, d2) represents the semantic similarity between diseases d1 and d2, d p is the common ancestor node of d1 and d2, and are respectively the ancestor nodes of d1 and d2. Among them, is defined as follows:

[0095]

[0096]

[0097] Here, is the weight of the edge connecting d1 and d i , d i is the child node of d p , and CA(d1, d2) represents the set of the most recent common ancestors of diseases d1 and d2, and A(d1) represents the set of all ancestor nodes related to disease d1.

[0098] (4) circRNA functional similarity

[0099] The functional similarity of circRNAs generally refers to the degree of similarity in biological functions among different circRNA molecules, which may be based on aspects such as the biological processes they participate in, their interactions with proteins or miRNAs, and their associations in diseases. In this example, the functional similarity CF(c i ,c j ) of circRNAs is calculated by analyzing the disease association matrix CDM of circRNAs and the disease semantic similarity, as shown in the formula:

[0100]

[0101] In the formula, CF(c i ,c j ) represents the functional similarity between circRNA c i and circRNA c j . Gr(c i ) and Gr(c j ) respectively represent the sizes of the disease sets related to circRNA c i and circRNA c j . and respectively represent the diseases associated with circRNA c i and circRNA c j . The molecular part calculates the sum of the maximum semantic similarities among all the diseases associated with circRNA c i and circRNA c j . The denominator is the sum of the sizes of the two disease groups, which is used to normalize the calculated similarity to prevent the similarity calculated for circRNAs with larger disease sets from being too high. The basic idea is that if there is a high semantic similarity among the diseases associated with two circRNAs, then these two circRNAs may also have similarity in biological functions.

[0102] (5) circRNA sequence similarity

[0103] With the development of sequencing technology, the sequence information of circRNAs has been gradually improved. In this example, the sequences of the used circRNAs are obtained from the circBank database. In this example, the Levenshtein Distance method is used to calculate the sequence similarity between circRNAs, which is defined as the minimum number of single-base edits (insertions, deletions, or substitutions) required to convert one sequence into another. In this example, the editing costs for insertions and deletions are 1, and the editing cost for substitutions is 2. The Levenshtein Distance method can be expressed as the formula:

[0104]

[0105] In the formula, CS(i, j) represents the similarity between the first i bases of sequence S1 and the first j bases of sequence S2.

[0106] As can be seen from the formula, a smaller distance indicates a higher similarity. To convert the measurement of similarity into a unified and comparable range, a normalization method can be used. Then the sequence similarity between two circRNAs can be defined as:

[0107]

[0108] In the formula, CS represents the Levenshtein Distance calculated according to a specific cost, and len(S1) and len(S2) represent the sequence lengths of the two circRNAs respectively.

[0109] (6) Comprehensive similarity

[0110] In this embodiment, some diseases are encountered, and their corresponding DOID numbers cannot be retrieved from the database. This results in the semantic similarity of these diseases and the functional similarity between circRNAs being recorded as 0 values. To solve this data missing problem, this embodiment uses the GIP kernel similarity metric to fill in these missing similarity values. As shown in the formula:

[0111]

[0112] In the formula, MixDS(d i ,d j ) represents the comprehensive similarity between disease d i and disease d j , and its dimension is MixDS ∈ R n ×n , n is the number of diseases. When the disease DO semantic similarity is equal to 0, the value of the GIP kernel similarity of the disease is taken. Otherwise, the disease DO semantic similarity is taken.

[0113]

[0114] In the formula, MixCS(c i ,c j ) represents the circRNA comprehensive similarity, MixCS ∈ R m×m , m is the number of circRNAs. When the CFS value of the circRNA is equal to 0, the value of the GIP kernel similarity of the circRNA is taken. Otherwise, the CFS value is taken. The CFS is defined as follows:

[0115]

[0116] II. Data Preprocessing Module

[0117] 1. Data Standardization

[0118] In this embodiment, the circRNA comprehensive similarity matrix and the disease comprehensive similarity matrix are first standardized. The Z-score standardization method (StandardScaler) is used to transform each feature into a distribution with a mean of 0 and a variance of 1, so as to eliminate the dimensional differences of data from different sources and ensure the stability and fairness of model training. Specifically, features such as the GIP kernel similarity, sequence similarity of circRNA, and semantic similarity of diseases are standardized respectively, avoiding features with a large numerical range from dominating the model learning process, while retaining the distribution characteristics of the original data, providing inputs with a unified scale for subsequent feature fusion.

[0119] 2. Feature Interaction

[0120] In this embodiment, the feature interaction technology is used to explicitly model the association pattern between circRNA and diseases. The specific implementation process is as follows:

[0121] In this embodiment, the row vectors of the disease comprehensive similarity matrix are extended to be the same size as the row vectors of the circRNA comprehensive similarity matrix , denoted as Next, each row of MixCS is concatenated with each row vector of MixDSE row by row to obtain the matrix In addition, the Hadamard product of each row of the matrix MixCS and each row of MixDSE is calculated to obtain the matrix Finally, the matrices CSDSE and DCSDSE are concatenated by rows to obtain the fusion features of each circRNA-disease pair As shown in the formula:

[0122] MixDSE = [MixDS::::] (14)

[0123] CSDSE ij = [MixCS i , MixDSE j (15)

[0124] MCD i = [CSDSE i , DCSDSE i (16)

[0125] In the formula, MixCS i represents the i-th row of the circRNA comprehensive similarity matrix MixCS, MixDSEj represents the j-th row of the comprehensive disease similarity matrix MixDSE after row expansion, CSDSE i represents the i-th row of the matrix CSDSE, DCSDSE i represents the i-th row of the matrix DCSDSE, :::: represents the row expansion of MixDSE.

[0126] This interaction method can not only effectively maintain the discriminative information of the original features, but also accurately capture the complex association rules between circRNA molecular characteristics and disease phenotype characteristics, providing a more expressive input representation for subsequent deep feature learning.

[0127] 3. Dynamic balance data

[0128] Aiming at the problem of positive and negative sample imbalance commonly existing in circRNA-disease association prediction, this embodiment proposes a data balance strategy based on dynamic oversampling. Since the number of verified circRNA-disease association samples (positive samples) is significantly less than that of un-verified association samples (negative samples), this embodiment introduces the dynamic balance data RandomOverSample method, as shown in the formula:

[0129]

[0130] N copies = N minority,target - N minority (18)

[0131] In the formula, N majority is the number of majority class samples, N minority is the number of minority class samples, and sampling_strategy is the oversampling ratio. By randomly selecting N copies minority class samples and copying them into the dataset to increase the number of minority class samples until the target number N minority,target is reached. The sampling_strategy value of this model is 0.5. After replication, the ratio of known associated circRNA-disease to unknown associated circRNA-disease sample sizes reaches 1:2. This adaptive oversampling mechanism not only effectively alleviates the model bias problem caused by class imbalance, but also avoids the information loss caused by simple undersampling, thus ensuring that the model can fully learn the key feature patterns of positive samples and improving the prediction accuracy of potential circRNA-disease associations.

[0132] III. Association prediction module

[0133] 1. Multi-perspective feature fusion GCN

[0134] GNN has been a popular direction in bioinformatics in recent years and has shown great potential in processing graph-structured data. Based on GNN, this embodiment improves a GCN layer with multi-view feature fusion, introduces disease similarity features and circRNA similarity features based on nodes, etc., and integrates features from different views into a unified graph embedding space to make feature expression more comprehensive. This model mainly includes an input layer, a feature fusion layer, a fully connected layer, and an output layer.

[0135] The feature fusion layer respectively convolves the feature matrices MixCS, MixDSE, and CSDSE to obtain H c 、H d and H cd , as shown in the formula:

[0136]

[0137]

[0138]

[0139] In the formula, is the normalized adjacency matrix after adding self-loops, and W c 、W d and W cd respectively represent the weight matrices of MixCS, MixDSE, and CSDSE. σ is the activation function.

[0140] Then, the feature fusion layer performs a feature fusion operation, as shown in the formula:

[0141] F = σ(α·H c +β·H d +(1-α-β)·H cd ) (22)

[0142] In the formula, α and β are learnable parameters used to control the proportions of H c 、H d and H cd , and σ is the activation function.

[0143] Add Dropout to the fused features as shown in the formula:

[0144] F = Dropout(F) (23)

[0145] Here, to prevent overfitting, this embodiment sets the random dropout rate to 0.2.

[0146] The fully connected layer further processes the fused feature F as shown in the formula:

[0147] H (2) = σ(F·W(2) ) (24)

[0148] Y = sig(H (2) ·W (3) ) (25)

[0149] where W (2) and W (3) are the weight matrices of the fully connected layer, sig is the sigmoid activation function, and Y is the final predicted value.

[0150] 2. Double - layer inner - loop meta - learning

[0151] The core idea of meta - learning is to improve the learning effect when facing new tasks by learning from the experiences obtained from multiple tasks. The core goal of meta - learning is to optimize the learning process of tasks, which can generally be divided into two levels. First is the external learning process, which is mainly responsible for extracting general knowledge from multiple tasks so that the model can quickly adapt when facing new tasks. Second is the internal learning process. In a new task, the model will use the knowledge obtained from the external learning process to perform rapid learning. Taking MAML as an example, it optimizes the initial parameters of the model so that the model can quickly adjust and effectively learn when encountering new tasks. In recent years, experts in various fields have begun to combine meta - learning with previous models to adapt to different training tasks.

[0152] The advantages of meta - learning are that it can help the model quickly adapt to new tasks, especially in the case of scarce data, and can obtain good learning effects through a small number of samples. It improves the generalization ability of the model, enables the model to share useful knowledge when facing different tasks, and enhances the overall performance. In addition, meta - learning can also automatically optimize hyperparameters, avoid cumbersome manual adjustment, and can improve the synergy between tasks in multi - task learning. Previously, MAML was mainly applied in the field of image processing and adopted a single inner - loop and outer - loop update method, achieving good prediction results. However, in the prediction of circRNA - disease associations in graph data, the MAML method has hardly been used. Therefore, in this embodiment, a meta - learning model is constructed according to the characteristics of these data such as circRNA, disease, and association matrix. In this embodiment, a double - layer inner - loop is used. The first layer optimizes the disease feature embedding, and the second layer fine - tunes the circRNA feature embedding to gradually complete feature adaptation, as shown in the formula:

[0153]

[0154] where θ' d is the parameter related to the disease, θ is the initial parameter of the model, α is the learning rate, L d is the loss function, is the gradient of the loss function L d with respect to the parameter θ. Similarly, is the loss function L c with respect to the parameter θ' d The gradient of the circRNA-related parameter θ' c is updated based on the update of the disease parameter θ' d and further adjusted to minimize the circRNA-related loss function L c .

[0155] Then the outer loop update method is similar to the previous outer loop update, as shown in the formula:

[0156]

[0157] In the formula, β is another learning rate, is the meta-loss function L that sums over all τ The gradient of the parameter θ meta with respect to the parameter θ.

[0158] 3. GCN + MAML

[0159] In this embodiment, the above-mentioned GCN is used as the basic model, and the double-layer inner loop meta-learning algorithm is used as the meta-learning model. The two are combined to achieve prediction. The specific details are as Figure 2 shown.

[0160] As can be seen from Figure 2 , the inner loop is performed independently for each task. By creating a model copy and using the task data for rapid update, it can adapt to specific tasks. The inner loop update uses BCELoss (Binary Cross-Entropy Loss) to calculate the loss and optimizes the parameters of the model copy through gradient descent. The number of update steps is determined by inner_steps, which is set to 3 in this implementation to achieve rapid adaptation. After the inner loop terminates, it enters the outer loop stage. For each task, the outer loop loss is calculated using BCELoss and the gradients are accumulated, and then the gradients of all tasks are averaged to update the global model parameters. This process is called meta-update. And the termination condition of the outer loop is to complete the processing of all tasks. Through this mechanism, the BIMLGCDA model can learn more general features on multiple tasks, thereby improving the ability to quickly adapt to new tasks and the generalization performance.

[0161] The BIMLGCDA model also uses BCELoss to calculate the loss between the model prediction value and the true label during the training phase, guiding the optimizer to adjust the model parameters. BCELoss is as shown in the formula:

[0162]

[0163] In the formula, N is the total number of circRNA-disease samples, y idenotes the true label of the i-th circRNA-disease sample, p i denotes the probability that the BIMLGCDA model predicts the i-th sample as the positive class.

[0164] B Results and Discussion

[0165] 1. Model Evaluation

[0166] In this embodiment, the evaluation metrics of common classification models: F1, AUC, Accuracy, AUPR, Recall are used to measure the BIMLGCDA model, as shown in the formula:

[0167]

[0168]

[0169]

[0170]

[0171]

[0172] In the formula, TP represents the number of correctly predicted positive samples, FN represents the number of positive samples predicted as negative samples, TN represents the number of correctly predicted negative samples, FP represents the number of negative samples predicted as positive samples, Precision represents the precision rate, FPR i and TPR i respectively represent the false positive rate and true positive rate of the i-th point, FPR i-1 and TPR i-1 respectively represent the false positive rate and true positive rate of the i-1-th point. In this embodiment, five-fold cross-validation is used to evaluate the performance of the BIMLGCDA model in predicting the association between circRNA and diseases.

[0173] 2. Model Classification Performance

[0174] To prove the influence of the key modules, in this embodiment, BIML is removed from the BIMLGCDA model respectively, and the GCN basic learner is replaced with an MLP with two fully connected layers, and ablation experiments are carried out through five-fold cross-validation on the dataset. The classification performance of the BIMLGCDA model on four benchmark datasets is comprehensively evaluated. The results are shown in Table 2.

[0175] The experimental results are shown in Table 2, indicating that the BIMLGCDA model achieved excellent performance on the circ2disease dataset with an F1 score of 0.9881, an AUC value of 0.9989, an accuracy of 0.992, an AUPR value of 0.9965, and a Recall value of 1.0. On the circRNADisease dataset, the F1 score of this model reached 0.9884, the AUC value was 0.9992, the accuracy was 0.9922, the AUPR value was 0.9978, and the Recall value was 0.9949. For the circR2Disease dataset, the BIMLGCDA model also performed excellently, with an F1 score of 0.9657, an AUC value of 0.9962, an accuracy of 0.9765, an AUPR value of 0.9893, and a Recall value of 0.9897. In addition, on the lncRNAdisease dataset, the model also demonstrated stable performance, with an F1 score of 0.9564, an AUC value of 0.9935, an accuracy of 0.97, an AUPR value of 0.9837, and a Recall value of 0.9873. These experimental results fully verified the effectiveness of the BIMLGCDA model in the circRNA-disease association prediction task.

[0176] Table 2 Results of ablation experiments

[0177]

[0178] It can be seen from the results that on the circR2Disease dataset, the BIMLGCDA model outperformed the GCN model in all metrics: the F1 score increased by 6.8%, the AUC increased by 1.5%, the ACC increased by 4.7%, the AUPR increased by 3.7%, and the Recall increased by 4.2%. In addition, on other datasets, the BIMLGCDA model also had varying degrees of improvement in metrics such as F1, AUC, ACC, AUPR, and Recall. The BIML module and the GCN base learner are the core factors contributing to the performance improvement of the BIMLGCDA model. Their synergistic effect enhanced the model's prediction ability for circRNA-disease associations, especially when dealing with complex graph-structured data and mining potential associations. This finding provides an important basis for the further optimization and improvement of the model.

[0179] 3. Comparison with advanced methods

[0180] To further evaluate the prediction performance of the BIMLGCDA model, this embodiment will conduct comparative experiments with some excellent models in recent years. In this embodiment, on the dataset circ2Disease, it is compared with the MDGF-MCEC, DCDA, MGRCDA, GSLCDA, and MAGCDA models; on the dataset circR2Disease, it is compared with the MPCLCDA, SGFCCDA, MDGF-MCEC, MLNGCF, RDGAN, MSPCD, iGRLCDA, GSLCDA, and NSL2CD models; on the dataset circRNADisease, it is compared with the DETHACDA, GCNMFCDA, MNMDCDA, MAGCDA, and GSLCDA. All experiments use 5-fold cross-validation and compare the performance of each model in terms of three evaluation metrics: F1, AUC, and ACC. The specific results are as Figure 5 shown.

[0181] As can be seen Figure 5 from it, the BIMLGCDA model performs excellently in terms of F1 score, AUC, and accuracy on the dataset circ2Disease, reaching 0.9881, 0.9989, and 0.992 respectively, which is better than other models. The BIMLGCDA and MPCLCDA models perform the most prominently on the dataset circR2Disease, with F1 scores of 0.9657 and 0.9878 respectively, AUC values of 0.9962 and 0.9877 respectively, and accuracy rates of 0.9765 and 0.9876 respectively. While other models have their own advantages and disadvantages in different metrics. The F1 score of the BIMLGCDA model on the dataset circRNADisease is 0.9884, the AUC value reaches 0.9992, and the accuracy rate is 0.9922. The F1 score of the DETHACDA model is 0.9508, the AUC value is 0.9882, and the accuracy rate is 0.9508. The effect of other models in predicting circRNA-disease associations is significantly lower. Although most models are superior to other models in a certain metric, they perform significantly worse in other metrics. Therefore, considering the performance of each metric comprehensively, the BIMLGCDA model demonstrates a relatively balanced prediction ability and overall excellent performance.

[0182] 4. Model Parameter Analysis

[0183] In this embodiment, the inner loop learning rate (inner_lr), the outer loop learning rate (outer_lr), the learning rate (lr) of the AdamW optimizer in the training stage, and the size of the hidden layer (hidden_dim) of the GCN layer are of observational significance to the model performance. Therefore, in this embodiment, systematic experiments were conducted on these key parameters on the dataset circRNADisease, and the optimal parameter configuration was selected to ensure that the model can achieve the best performance in the circRNA-disease classification task. This embodiment adopts the single-variable method to analyze the influence of each hyperparameter on the model one by one, and the results are as Figure 3 shown.

[0184] As Figure 3 can be seen, choosing different values for different hyperparameters affects the model performance to a certain extent. The value of inner_lr has a weak impact on the model performance, while outer_lr, lr, and hidden_dim have a greater impact on the model performance. Considering the changes in the model performance and the requirement of small computing power, this embodiment finally selects inner_lr = 1e-5, outer_lr = 1e-4, lr = 1e-4, and hidden_dim = 128.

[0185] 5. Generalization experiment

[0186] To comprehensively evaluate the adaptability and generalization ability of the BIMLGCDA model in a multi-task environment, this embodiment conducted five-fold cross-validation experiments on the miRNA datasets HMDD_v2.0 and HMDD_v3.2. The results are shown in Table 3, Figure 4 as shown. These datasets contain rich miRNA-disease association information and are from the same datasets used by Yulian Ding in his VGAMF model. Among them, Yulian Ding adopted a different strategy from the method of the present invention in data preprocessing. Therefore, conducting generalization experiments on these datasets can not only test the cross-domain adaptability of the model but also provide more powerful verification.

[0187] Table 3 Performance on miRNA datasets

[0188]

[0189] As can be seen from Table 3, on the HMDD_v2.0 dataset, the F1 value of the BIMLGCDA model is 0.9801, the AUC is 0.9946, the ACC is 0.9866, the AUPR is 0.9848, and the Recall is 0.9930; while on the HMDD_v3.2 dataset, the F1 value is 0.9760, the AUC is 0.9942, the ACC is 0.9837, the AUPR is 0.9815, and the Recall is 0.9950. These metrics indicate that the BIMLGCDA model has high accuracy and robustness on both datasets. Especially in differentiating positive and negative samples, the AUC values both reach above 0.9940, showing the strong ability of the model in identifying the differences between positive and negative samples. In addition, despite the differences in dataset versions and the data preprocessing methods being different from previous studies, the model can still stably output high-precision results. This shows that the BIMLGCDA model not only performs well on specific datasets but also has good cross-domain adaptability and generalization ability, and can effectively handle the complex features and correlation information in different versions of datasets.

[0190] Example 2: Case Analysis

[0191] Breast cancer and gastric cancer are the main causes of cancer-related deaths globally, posing a serious threat to public health. Breast cancer is one of the most common cancers in women, with a complex pathogenesis involving genetic factors, hormonal level changes, and environmental factors. Breast cancer usually starts in the breast ducts or glands, and after the tumor cells spread, they can invade the surrounding tissues and metastasize to other organs through the lymphatic system or blood. The harm of breast cancer is not only reflected in its high incidence and fatality rate but also because its early symptoms are not obvious, and many patients are diagnosed at a late stage.

[0192] The incidence and mortality of gastric cancer are also relatively high globally, especially in East Asia where the incidence of gastric cancer is significant. Its pathogenesis is closely related to Helicobacter pylori infection, chronic inflammation of the gastric mucosa, genetic susceptibility, and environmental factors. The early symptoms of gastric cancer are usually not obvious, and after the disease progresses, symptoms such as stomach pain and loss of appetite may occur, but they are mostly not taken seriously enough, resulting in patients missing the best treatment opportunity.

[0193] Early detection and accurate prediction of breast cancer and gastric cancer are crucial for prevention and treatment. In the context of challenges in the early diagnosis and prognosis assessment of breast cancer and gastric cancer, it is particularly important to explore new biomarkers, especially circRNAs as potential disease-related factors. CircRNAs play a key role in various diseases, with strong stability and specificity, and can support early diagnosis, personalized treatment, and disease prognosis assessment. Therefore, in this example, a case study of breast cancer and gastric cancer was carried out on the circR2Disease dataset, aiming to verify the application potential of the BIMLGCDA model in these two cancer types. In this example, the top 10 circRNAs with the highest correlation with breast cancer and gastric cancer were selected, and their association with specific diseases was confirmed through literature retrieval. The results are shown in Table 4:

[0194] Table 4 The top 10 circRNAs with the highest predicted values for breast cancer and gastric cancer

[0195]

[0196]

[0197] As can be seen from Table 4, the association between the circRNAs predicted by the BIMLGCDA model and specific diseases has been verified to varying degrees. For breast cancer, hsa_circ_104689, circVMA21, circPABPN1, hsa_circ_0001283, hsa_circRNA_002271, circUBR5 / hsa_circ_0001819, and circ-FBXW7 were predicted as potentially relevant circRNAs and have been confirmed by the literature. For gastric cancer, hsa_circ_0002343, hsa_circRNA_100269, circETFA, circBRAF, Titin circRNAs, hsa_circ_0004458, and circRNA_Atp9b were predicted as potentially relevant circRNAs and supported by the literature. This indicates that the BIMLGCDA model has a certain degree of accuracy and reliability in predicting the association between circRNAs and diseases, and can provide valuable references for the discovery of biomarkers for related diseases.

[0198] The above description is only a preferred embodiment of the embodiments of the present invention, and does not impose any form of limitation on the embodiments of the present invention. Any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the embodiments of the present invention still fall within the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A circRNA-disease association prediction model based on MAML and GCN, characterized in that, It includes: Similarity information module: constructing the comprehensive similarity between circRNA and diseases; Data preprocessing module, including: data standardization, feature interaction, and dynamic data balancing; Association prediction module, including: a GCN layer for multi-perspective feature fusion, and a MAML composed of a double-layer inner loop and a single-layer outer loop. Using the GCN layer as the base model and the MAML composed of the double-layer inner loop and the single-layer outer loop as the meta-learning model, they are fused to achieve association prediction.

2. The circRNA-disease association prediction model according to claim 1, wherein the comprehensive similarity is obtained by integrating various similarity information of circRNA and diseases, including: GIP kernel similarity, sequence similarity, functional similarity, disease semantic similarity, and the comprehensive similarity of GIP kernel similarity.

3. The circRNA-disease association prediction model according to claim 2, wherein the formula for the comprehensive similarity of the GIP kernel similarity is as follows: where MixDS(d i , d j ) represents the comprehensive similarity between disease d i and disease d j , and its dimension is MixDS ∈ R n×n , n is the number of diseases. When the semantic similarity of disease DO is equal to 0, the GIP kernel similarity value of the disease is taken; otherwise, the semantic similarity of disease DO is taken. where MixCS(c i , c j ) represents the comprehensive similarity of circRNA, MixCS ∈ R m×m , m is the number of circRNAs. When the CFS value of a circRNA is equal to 0, the similarity of the GIP nucleus of the circRNA is taken; otherwise, the CFS value is taken.

4. The circRNA-disease association prediction model according to claim 1, wherein in the dynamic data balancing, the RandomOverSample method is adopted to dynamically adjust the positive and negative sample ratios, and the formula is as follows: N minority = "N majority × sampling_strategy];" N copies = N minority,target - N minority ; where N majority is the number of majority class samples, N minority is the number of minority class samples, and sampling_strategy is the oversampling ratio.

5. The circRNA-disease association prediction model according to claim 1, wherein the GCN layer for multi-perspective feature fusion includes an input layer, a feature fusion layer, a fully connected layer, and an output layer; Among them, the feature fusion layer respectively convolves the feature matrices MixDS, MixDSE, and CSDSE to obtain H c , H d and H cd , and the formula is as follows: In the formula, is the normalized adjacency matrix after adding self-loops, and W c , W d and W cd represent the weight matrices of MixCS, MixDSE, and CSDSE respectively, and σ is the activation function; the feature fusion layer performs feature fusion operations, and the formula is as follows: F = σ(α·H c + β·H d +(1 - α - β)·H cd ); Where α and β are learnable parameters used to control the ratios of H c , H d and H cd , and σ is the activation function; in the fully connected layer, Dropout is added to the fused features, and the formula is: F = Dropout(F).

6. The circRNA-disease association prediction model according to claim 1, wherein in the MAML with a double-layer inner loop structure, the first layer optimizes the disease feature embedding, and the second layer fine-tunes the circRNA feature embedding to gradually complete feature adaptation, and the formula is as follows: where θ' d is a parameter related to the disease, θ is the initial parameter of the model, α is the learning rate, and L d is the loss function, and d is the gradient of the loss function L with respect to the parameter θ; similarly, c is the gradient of the loss function L d with respect to the parameter θ'. The formula for the outer loop update method is as follows: where β is the learning rate, is the meta-loss function L τ summed over all meta and ∇θL is the gradient of L with respect to the parameter θ.

7. The circRNA-disease association prediction model according to claim 1, wherein in the association prediction module, the inner loop is independently performed for each task, and the model copy is created and quickly updated using the task data; after the inner loop terminates, it enters the outer loop stage; wherein, the inner loop update calculates the loss using BCELoss and optimizes the parameters of the model copy through gradient descent; The formula for BCELoss is as follows: where N is the total number of circRNA-disease samples, and y i represents the true label of the i-th circRNA-disease sample, and p i represents the probability that the BIMLGCDA model predicts the i-th sample as the positive class.

8. A circRNA-disease association prediction device based on MAML and GCN, characterized in that The circRNA-disease association prediction device adopts the circRNA-disease association prediction model according to any one of claims 1-7.

9. A computer device, characterized in that, It includes a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the operation of the circRNA-disease association prediction model according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when called and executed by a processor, cause the processor to implement the operation of the drug-disease association prediction model according to any one of claims 1-7.