Multi-level information fusion method based on high-order weighted disturbance

Through the method of high-order weighted perturbation and deep non-negative matrix decomposition combined with hypergraph convolution and multi-layer perceptron, the problem of insufficient information mining and single feature learning in circRNA-disease association prediction is solved, and higher prediction accuracy and generalization ability are achieved.

CN120492837APending Publication Date: 2025-08-15CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510556836.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art has problems of insufficient advanced information mining and single feature learning methods in circRNA-disease association prediction.

Method used

The advanced-order weighted perturbation method is used to combine deep non-negative matrix decomposition and dual-path feature learning, and the higher-order association and nonlinear features of circRNA and disease are captured through hypergraph convolution and multi-layer perceptrons, and multi-layer perceptrons are trained to improve prediction accuracy.

Benefits of technology

The accuracy and generalization ability of circRNA-disease association prediction have been significantly improved, and the verification results show significant advantages on multiple public data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005383625350000011
    Figure HDA0005383625350000011
Patent Text Reader

Abstract

The invention provides a multi-level information fusion method based on high-order weighted perturbation. The multi-level information fusion method is used for circRNA-disease association prediction. According to the high-order weighted perturbation method, the key proximity relation and long-distance association are balanced by calculating association information of different orders and giving decreasing weights, and therefore the high-order information mining capacity is enhanced. A depth non-negative matrix factorization method is combined with high-order correlation information, a global topology mode is captured through multilayer mapping, the representation capability of a correlation matrix is improved, and more accurate input is provided for subsequent analysis. The two-stage matrix decomposition method and the two-path feature learning method are combined to realize diversification of feature learning, linear and nonlinear structures of data are fully mined, and the modeling capability of the method for multi-level information is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of circRNA-disease association prediction, and in particular to a multi-level information fusion method based on high-order weighted perturbation. Background Art

[0002] Circular RNAs (circRNAs) are a class of non-coding RNAs with unique structures and functions, holding significant value in revealing disease mechanisms and developing early prevention and treatment strategies. In recent years, computational methods have made significant progress in predicting potential circRNA-disease associations, but they still face challenges such as insufficient high-order information mining and a single feature learning approach. To address this, this paper proposes a multi-level information fusion method based on high-order weighted perturbations. First, in a two-stage matrix factorization, high-order correlation information is extracted using a high-order weighted perturbation method. Multi-level structural features are then obtained using deep non-negative matrix factorization to reconstruct the correlation matrix. Linear features are then extracted using a single-layer non-negative matrix factorization. Next, in a dual-path feature learning approach, the circRNA and disease similarity networks are fed into hypergraph convolution and multilayer perceptrons, respectively, to capture complex nonlinear features. Finally, for classification prediction, the concatenated features are fed into the multilayer perceptron for training. Summary of the Invention

[0003] The purpose of the present invention is to address the difficulties in the field of circRNA-disease association prediction mentioned above and provide a multi-level information fusion method based on high-order weighted perturbations to help researchers make accurate predictions. The technical solution of the present invention is as follows:

[0004] A. Obtain circRNA-disease related data from multiple public and reliable databases and construct a similarity matrix between four circRNAs and diseases.

[0005] B. A high-order weighted perturbation method is used to extract high-order correlation information, combined with deep non-negative matrix decomposition to obtain multi-level structural features, thereby reconstructing the correlation matrix, and extracting linear features through single-layer non-negative matrix decomposition.

[0006] C. The circRNA and disease similarity networks are fed into hypergraph convolution and multilayer perceptron, respectively, to capture complex nonlinear features.

[0007] D. Input the spliced features into the multi-layer perceptron for training to complete the prediction task.

[0008] E. Performance evaluation is conducted on multiple public datasets. Experimental results show that the proposed method has significant advantages in prediction accuracy, and case studies further verify its reliability.

[0009] According to the data acquisition and integration described in Claim A, circRNA-disease related data were obtained from multiple public and reliable databases to construct four similarity matrices: a disease semantic similarity matrix, a circRNA functional similarity matrix, a circRNA (disease) GIP core similarity matrix, a circRNA (disease) cosine similarity matrix, and a circRNA (disease) information entropy and mutual information similarity matrix. By integrating these matrices, a comprehensive similarity network between circRNAs and diseases was obtained.

[0010] 1. Based on the two-stage matrix factorization described in claim B, first, the original correlation matrix and the high-order correlation matrix are input into the deep non-negative matrix factorization model in combination with the high-order weighted perturbation method to reconstruct the correlation matrix. This process uses a multi-layer mapping structure to map from the high-dimensional network to the low-dimensional latent space, deeply exploring data features and improving the integrity of the correlation information. In the second stage, non-negative matrix factorization is used to extract the low-dimensional feature matrix of circRNAs and diseases from the reconstructed correlation matrix.

[0011] 2. Based on the dual-path feature learning described in Claim C, hypergraph convolution effectively captures high-order relationships between multiple nodes, revealing the complex associations between circRNAs and diseases. Multilayer perceptrons enhance feature representation through nonlinear mapping, further enriching feature extraction. Hypergraph convolution and multilayer perceptrons extract information from the perspectives of high-order relationships and nonlinear feature mapping, respectively, supplementing and strengthening feature representation.

[0012] 3. Based on the classification prediction described in Claim D, the circRNAs extracted through two-stage matrix factorization and dual-path feature learning are concatenated with disease features and then input into a multilayer perceptron with two hidden layers and one output layer. This integrates linear and nonlinear features to enhance the multi-level representation of the data, thereby improving prediction accuracy and generalization.

[0013] 4. Based on the performance evaluation and validation described in Claim E, we conducted parameter analysis, ablation experiments, and comparative experiments on multiple public and reliable datasets. The experimental results demonstrate that our method exhibits significant advantages in accuracy, robustness, and interpretability. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a multi-level information fusion method process based on high-order weighted perturbations. DETAILED DESCRIPTION

[0015] To make the objectives, technical solutions and advantages of the present invention more clear, the association prediction method of the present invention is further described in detail below with reference to the accompanying drawings.

[0016] First, in a two-stage matrix factorization, high-order weighted perturbation is used to extract high-order correlation information, and deep non-negative matrix factorization is used to obtain multi-level structural features to reconstruct the correlation matrix. Subsequently, linear features are extracted through a single-layer non-negative matrix factorization. Then, in a dual-pathway feature learning, the circRNA and disease similarity networks are input into hypergraph convolution and multi-layer perceptron, respectively, to capture complex nonlinear features.

[0017] Finally, in the classification prediction, the splicing features are input into the multi-layer perceptron for training to complete the prediction task.

Claims

1. A multi-level information fusion method based on high-order weighted perturbations for circRNA-disease association prediction, involving the fields of bioinformatics, machine learning, and deep learning. The method mainly includes the following steps: A. Data acquisition and integration: CircRNA-disease related data were obtained from multiple public and reliable databases, and similarity matrices between four circRNAs and diseases were constructed. B. Two-stage matrix factorization: A high-order weighted perturbation method is used to extract high-order correlation information, combined with deep non-negative matrix factorization to obtain multi-level structural features, thereby reconstructing the correlation matrix, and extracting linear features through single-layer non-negative matrix factorization. C. Dual-path feature learning: The circRNA and disease similarity networks are input into hypergraph convolution and multi-layer perceptron respectively to capture complex nonlinear features. D. Classification prediction: Input the spliced features into the multi-layer perceptron for training to complete the prediction task. E. Performance Evaluation and Validation: Performance evaluation is conducted on multiple public datasets. Experimental results show that the proposed method has significant advantages in prediction accuracy, and case studies further verify its reliability.

2. Data acquisition and integration according to claim 1, by obtaining circRNA and disease-related data from multiple public and reliable databases, constructing four similarity matrices, namely: disease semantic similarity matrix, circRNA functional similarity matrix, circRNA (disease) GIP core similarity matrix, circRNA (disease) cosine similarity matrix, and circRNA (disease) information entropy and mutual information similarity matrix. By integrating these matrices, a comprehensive similarity network of circRNAs and diseases is obtained.

3. The two-stage matrix factorization according to claim 1 first combines a high-order weighted perturbation method to input the original correlation matrix and the high-order correlation matrix into a deep non-negative matrix factorization model to reconstruct the correlation matrix. This process uses a multi-layer mapping structure to map from a high-dimensional network to a low-dimensional latent space, deeply exploring data features and improving the integrity of correlation information. In the second stage, non-negative matrix factorization is used to extract a low-dimensional feature matrix of circRNAs and diseases from the reconstructed correlation matrix.

4. In the dual-path feature learning method described in claim 1, the hypergraph convolution method effectively captures high-order relationships between multiple nodes, revealing the complex associations between circRNAs and diseases. The multilayer perceptron enhances feature expression through nonlinear mapping, further enriching feature extraction. Hypergraph convolution and multilayer perceptron extract information from the perspectives of high-order relationships and nonlinear feature mapping, respectively, supplementing and strengthening feature expression.

5. The classification prediction method according to claim 1 comprises concatenating the circRNAs extracted through two-stage matrix factorization and dual-path feature learning with disease features and then inputting the concatenated data into a multilayer perceptron with two hidden layers and one output layer. This method integrates linear and nonlinear features, enhances the multi-level representation of the data, and thus improves prediction accuracy and generalization.

6. To evaluate and validate the performance of claim 1, we conducted parameter analysis, ablation experiments, and comparative experiments on multiple publicly available and reliable datasets. These results demonstrate that this method exhibits significant advantages in accuracy, robustness, and interpretability, and has broad application prospects in bioinformatics, potentially improving the efficiency and accuracy of circRNA-disease association prediction.