Microorganism-disease associated prediction system based on sparse variational auto-encoder
By extracting the nonlinear and linear features of microorganisms and diseases through a sparse variational autoencoder model, and combining graph convolution and sparse penalty, the problem of low accuracy in microorganism-disease association prediction is solved, efficient multi-source heterogeneous data integration and feature extraction are achieved, and prediction accuracy and interpretability are improved.
Patent Information
- Application Number
- CN202510755140.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-09
AI Technical Summary
Existing microbiome-disease association prediction methods have problems such as low prediction accuracy, high cost, lack of model interpretability, and inability to effectively integrate multi-source heterogeneous data.
A sparse variational autoencoder (PGVAEMDA) model was used to construct a microbe-disease association matrix to obtain the Gaussian interaction spectrum kernel similarity, symptom similarity, and sequence similarity matrices between microbes and diseases. Graph convolution and sparse penalty were combined to extract nonlinear and linear features, and the CatBoost model was used for prediction.
It improves the accuracy of microbe-disease association prediction, enhances feature extraction capabilities and model interpretability, solves the problem of integrating multi-source heterogeneous data, and reduces the impact of noise.
Smart Images

Figure CN120613017A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of bioinformatics and electronic data processing, and in particular to a microbe-disease association prediction system based on a sparse variational autoencoder. Background Art
[0002] Microorganisms play a vital role in the human body, often called its "forgotten organ." The human microbiome plays a key role in maintaining physiological homeostasis, participating in metabolic processes, and regulating immune function. Imbalances in the microbiome are closely linked to a variety of conditions, including diabetes, cardiovascular disease, intestinal disorders, and neurodegenerative diseases. Therefore, studying the relationship between microorganisms and disease is crucial.
[0003] Current methods for studying microbial-disease associations primarily rely on traditional biological experiments. However, due to the vast diversity of microbial species and the complex relationships between these microbes and diseases, these methods rely on long-term experimental validation, often requiring years to draw conclusions. This is time-consuming, labor-intensive, and costly, making traditional biological experiments inefficient and costly for predicting microbial-disease associations. With the advancement of technology, machine learning, network topology, and deep learning methods have become the mainstream models for predicting microbial-disease associations. These models typically construct similarity networks of microbes or diseases based on known association data and predict potential associations through network propagation or low-dimensional embedding (e.g., matrix factorization). However, microbial-disease associations are influenced by multiple factors, and current prediction models are unable to effectively integrate heterogeneous data from diverse sources, such as genetics, environment, and lifestyle, resulting in low accuracy in predicting microbial-disease associations. Furthermore, microbiome data are typically high-dimensional and sparse. Existing prediction models often miss key information or introduce significant errors when processing such data, further contributing to low accuracy in predicting microbial-disease associations. With the development of deep learning, deep learning-based microbe-disease association prediction methods have achieved breakthroughs in prediction accuracy. However, their "black box" characteristics make the models lack of interpretability, making it impossible for researchers to obtain key microbial characteristics or potential biological mechanisms. As a result, they still face challenges in effectively integrating multi-source heterogeneous data, limited feature extraction, and processing high-dimensional sparse data, resulting in low overall model prediction accuracy and thus low microbe-disease association prediction accuracy. Summary of the Invention
[0004] The purpose of the present invention is to solve the problem of low accuracy of disease-microorganism association prediction in existing microorganism-disease association prediction methods, and propose a microorganism-disease association prediction system based on sparse variational autoencoders.
[0005] A microbe-disease association prediction system based on sparse variational autoencoders, comprising: an association matrix construction module, a module for obtaining final embedded features to be tested, and a module for predicting microbe-disease association relationships; The association matrix building module is used to obtain the known microorganism-disease association matrix , and the matrix Input to the final embedding feature acquisition module to be tested; Microbe-Disease Association Matrix Middle i Rank j The elements of the column are: in, It is i disease, It's a sign of disease. It is microorganisms, is the microbe-disease association matrix Middle i Rank j Elements of the column; The final embedded feature acquisition module to be tested uses the microorganism-disease association matrix Obtaining a final embedding feature matrix, and using the final embedding feature matrix to obtain a final embedded feature to be tested, and inputting the final embedded feature to be tested into a microorganism-disease association prediction module; The method of using the final embedded feature matrix to obtain the final embedded features to be tested is specifically as follows: obtaining the rows corresponding to the disease to be tested and the microorganism to be tested in the final embedded feature matrix , and As the final embedded feature to be tested; The microorganism-disease association prediction module inputs the final embedded features to be tested into the trained CatBoost model to obtain the predicted value of the association between the disease to be tested and the microorganism.
[0006] Furthermore, the final embedding feature matrix and the trained CatBoost unit are obtained by: A1. Using the Microbe-Disease Association Matrix Obtaining the microbial Gaussian interaction spectrum kernel similarity matrix and disease Gaussian interaction spectrum kernel similarity matrix : in, It is j microorganisms and The Gaussian interaction spectrum kernel similarity of microorganisms, It is diseases and The Gaussian interaction spectrum kernel similarity of each disease, is the normalized microbial nuclear bandwidth, is the matrix A column vector, is the matrix A column vector, It is a microbial nuclear broadband, is the total number of columns in matrix A, is the normalized disease kernel bandwidth, is the matrix A row vector, is the matrix A's row vector, It is the disease core broadband, is the total number of rows in matrix A; A2. Obtaining disease symptom similarity matrix and sequence similarity matrices of microorganisms ; Disease Symptom Similarity Matrix The elements in are: in, Indicates disease and diseases Based on the similarity of symptoms, and Represents diseases and diseases association vectors with respective symptoms; The association vector stores the association value between the current disease and the corresponding symptom, and the association value includes: 0 and 1; 0 indicates that the current disease is not related to the current symptom, and 1 indicates that the current disease is related to the current symptom; The elements in the sequence similarity matrix of the microorganism are obtained by searching the STRINGv11 database to obtain the similarity between the protein sequences of the microorganisms. ; in, It is a microorganism and microorganisms The similarity between protein sequences; A3. Using the Matrix and matrix Obtaining microbial fusion similarity matrix , using the matrix and matrix Obtaining the disease fusion similarity matrix : in, It is a disease and disease The fusion similarity, It is a microorganism and microorganisms The fusion similarity; A4. Fusion of microbe-disease similarity matrix , disease fusion similarity matrix Combined with the microorganism-disease association matrix A to obtain the graph convolution adjacency matrix H ; A5. Constructing the initial embedding matrix ,use , graph convolution adjacency matrix H and matrix A to form a training set, use the training set to train the PGVAEMDA model, obtain the trained PGVAEMDA model and output the final embedding feature matrix; Initial embedding matrix , specifically: The PGVAEMDA model includes: a nonlinear feature extraction unit, a linear feature extraction unit, an embedding matrix output unit, and a CatBoost unit; The nonlinear feature extraction unit is used to obtain the microbial nonlinear feature matrix and the disease nonlinear characteristic matrix ; The linear feature extraction unit uses the microorganism-disease association matrix Obtaining microbial linear feature matrix and disease linear feature matrix ; The embedding matrix output unit converts the microorganism-disease association matrix , disease nonlinear characteristic matrix , microbial nonlinear characteristic matrix , disease linear feature matrix and microbial linear feature matrix Concatenate to obtain the final embedding feature matrix and output it; The CatBoost unit outputs a disease-microbe association value based on the final embedded feature matrix.
[0007] Furthermore, the microorganism-disease fusion similarity matrix in A4 , disease fusion similarity matrix Combined with the microorganism-disease association matrix A to obtain the graph convolution adjacency matrix H , specifically: 。
[0008] Furthermore, the nonlinear feature extraction unit includes: an encoder layer and a decoder layer; The encoder layer is a graph convolutional network GCN, using and H Get the intermediate layer hidden vector Z , and send the intermediate layer hidden vector to the decoder layer; The utilization and H Get the intermediate layer hidden vector Z , specifically: First, use the following formula to obtain : Then, use the Hidden vector of the middle layer Z : in, l is the number of graph convolution layers, It is the potential feature data The mean of is an intermediate variable, yes The variance of is the forward propagation function of the mean in the encoding layer, is the forward propagation function of the variance in the coding layer, e Obey standard Gaussian distribution N (0,1), L is the total number of graph convolutional layers; The decoder layer is used to utilize the intermediate layer latent vector Z Pair Matrix H Reconstruct and obtain the reconstructed matrix : in, is the forward propagation function from the encoder layer to the decoder layer.
[0009] Furthermore, the intermediate layer hidden vectorZ The spatial distribution is as follows: Among them, refers to the conditional probability distribution of the latent vector Z, is the distribution of data in the matrix H, is the spatial distribution of the intermediate layer hidden vector Z.
[0010] Furthermore, the nonlinear feature extraction unit is trained using the following loss function: N = N d + N m in, Loss T is the overall loss function of the nonlinear feature extraction unit, is the KL divergence loss; is the sparse penalty loss; ||·||1 is the L1 norm, SpTh is the sparse threshold, λ is the sparse penalty coefficient, , N is the total number of microorganism and disease samples, is the mean of the feature distribution, is the standard deviation of the characteristic distribution, is the loss function of the decoder layer, is a matrix H The n Row features, is a matrix H' Middle n Reconstruction features.
[0011] Furthermore, the linear feature extraction unit uses the microorganism-disease association matrix Obtaining microbial linear feature matrix and disease linear feature matrix , specifically: Correlation Matrix A Perform negative matrix decomposition to obtain the microbial linear feature matrix and disease linear feature matrix : in, r is the characteristic dimension.
[0012] Furthermore, the objective function of the linear feature extraction unit is: where ||·||F is the Frobenius norm.
[0013] Furthermore, the embedded feature matrix output unit converts the microorganism-disease association matrix , disease nonlinear characteristic matrix , microbial nonlinear characteristic matrix , disease linear feature matrix and microbial linear feature matrix The final embedding feature matrix is obtained by concatenation, specifically: First, the concatenation matrix 、 and Obtain the final disease embedding feature matrix , the splicing matrix 、 and Obtain the final microbial embedding feature matrix : in, is the final embedding feature matrix of the disease, It is the final embedding feature matrix of microorganisms; Then, obtain the microorganism-disease association matrix The microorganism-disease association pair corresponding to the value 1 is the positive example. At the same time, the number of positive examples is obtained. Then, the negative examples with the same number as the positive examples are randomly selected. The selected positive and negative examples are combined into the first set. The row labels corresponding to the microorganism-disease pairs in the first set in the matrix A are obtained. - Column number , then splice Middle j Line and Middle i Rows are used as the final embedding feature matrix t A line in: in, yes No. OK, yes No. OK, is the first s OK, The microorganism-disease pair corresponding to the value 0 is a negative example.
[0014] Furthermore, the CatBoost unit is specifically: in, It is m A decision tree model, M is the total number of microorganism-disease samples finally embedded in the feature matrix, is the target value, It is before m -1 prediction value of the decision tree model, It is the s The predicted value of samples; The following loss function is used to train the CatBoost unit: Among them, among them, is the positive sample label in the final embedded feature matrix, express The probability of being predicted as a positive sample, is the negative sample label in the final embedded feature matrix, express The probability of being predicted as a negative sample; Positive examples indicate that there is an association relationship, and negative examples indicate that there is no association relationship.
[0015] The beneficial effects of the present invention are: The present invention uses a method that fuses microbial Gaussian interaction spectrum kernel similarity with microbial sequence similarity to obtain a microbial fusion similarity matrix; and a method that fuses disease symptom similarity with disease Gaussian interaction spectrum kernel similarity to obtain a disease fusion similarity matrix. The fusion of the two similarities fills in the potential correlation information. The encoder layer in the nonlinear feature extraction unit of the PGVAEMDA model of the present invention implements a transformation mapping by constraining the latent variable Z to obey the standard normal distribution (KL divergence) and reconstructing the data, thereby extracting the latent features of the data through the latent distribution of the data. The PGVAEMDA model also introduces graph convolution and sparsity penalty. Through graph convolution, the adjacent or skipped neighbor information of the node is learned, and the node features are embedded in a richer representation space. The sparsity penalty is used to reduce the impact of data noise on model performance, thereby extracting more expressive features. The present invention uses a sparse variational autoencoder to extract nonlinear features, uses non-negative matrix decomposition technology to obtain linear features, and splices the linear features, nonlinear features and correlation matrix into CatBoost. Without losing the original information, the potential linear and nonlinear features are connected in series, revealing the unknown correlation relationships hidden in the data. The present invention solves the problems of existing microorganism-disease association prediction methods, such as the inability to effectively integrate multi-source heterogeneous data, limited feature extraction capabilities, and low prediction accuracy under high-dimensional sparse data, thereby improving the accuracy of prediction of microorganism-disease association relationships. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A general flow chart constructed for the microbe-disease association; Figure 2 Construct detailed process maps for microbe-disease associations; Figure 3 A detailed diagram of the construction process of the graph convolutional network (GCN); Figure 4 This is a detailed process diagram for constructing a sparse variational autoencoder; Figure 5 Schematic diagram of non-negative matrix factorization (NMF); Figure 6 This is a diagram of data splicing before PGVAEMDA model training; Figure 7 This is the ROC graph obtained by the PGVAEMDA model under 2 fold, 5 fold, and 10 fold. DETAILED DESCRIPTION
[0017] Specific implementation method 1: Figure 1-Figure 2 As shown, in this embodiment, a microorganism-disease association prediction system based on sparse variational autoencoder includes: an association matrix construction module, a module for obtaining the final embedded features to be tested, and a microorganism-disease association relationship prediction module; The association matrix building module is used to obtain the known microorganism-disease association matrix , and the matrix Input into the final embedding feature acquisition module to be tested; Microbe-Disease Association Matrix Middle i Rank j The elements of the column are: in, It is i disease, It's a sign of disease. It is microorganisms, is the microbe-disease association matrix Middle i Rank j Elements of the column; The final embedded feature acquisition module to be tested uses the microorganism-disease association matrix Obtaining a final embedding feature matrix, and using the final embedding feature matrix to obtain a final embedded feature to be tested, and inputting the final embedded feature to be tested into a microorganism-disease association prediction module; In the final embedding feature matrix t Get the rows corresponding to the disease to be tested and the microorganism to be tested , and As the final embedded feature to be tested; The microorganism-disease association prediction module inputs the final embedded features to be tested into the trained CatBoost unit to obtain the predicted value of the association between the disease to be predicted and the microorganism.
[0018] Specific implementation method 2: The final embedded feature matrix and the trained CatBoost unit are obtained by the following method: A1. Using the Microbe-Disease Association Matrix Obtaining the microbial Gaussian interaction spectrum kernel similarity matrix and disease Gaussian interaction spectrum kernel similarity matrix : in, It is j microorganisms and The Gaussian interaction spectrum kernel similarity of microorganisms, It is diseases and The Gaussian interaction spectrum kernel similarity of each disease, is the normalized microbial nuclear bandwidth, is the matrix A Column vector (microorganisms Gaussian interaction spectrum), is the matrix A column vector, is the original microbial nuclear broadband (before normalization), is the total number of columns in matrix A, is the normalized disease kernel bandwidth, is the matrix A Row vector (disease Gaussian interaction spectrum), is the matrix A row vector, It is the original disease kernel broadband, is the total number of rows in matrix A; In this step, the Gaussian kernel function is a commonly used similarity measurement method, which can determine the similarity between two vectors by calculating the Euclidean distance between them. Controls the smoothness of similarity calculation. The larger the value, the larger the local influence range of the Gaussian kernel function. If the value is too small, it is easy to overfit in the classification task. A2. Use the Medical Subject Headings (MeSH) database to obtain the relationship between different diseases and thus obtain the disease symptom similarity matrix , using the STRINGv11 database to retrieve microbial protein sequences and obtain the sequence similarity matrix of microorganisms : A2-1. The relationship between diseases was obtained from the Medical Subject Headings (MeSH) database (https: / / www.ncbi.nlm.nih.gov / ). The symptom-based similarity between two diseases is closely related to the number of genes they share and the degree of protein interaction. In addition, the diversity of clinical manifestations of diseases may be related to the connection pattern of the underlying protein interaction network. Disease Symptom Similarity Matrix The elements in are: in, Indicates disease and diseases Based on the similarity of symptoms, and Represents diseases and diseases The correlation vector with each symptom. If a disease is associated with a symptom, the correlation value between them is 1, and if not, the correlation value is 0.
[0019] A2-2. Similarity between protein sequences of microorganisms retrieved from the STRINGv11 database (protein-protein functional interaction network https: / / string-db.org) , thereby obtaining the sequence similarity matrix of microorganisms ; in, It is a microorganism and microorganisms The similarity between protein sequences; A3. Using the Matrix and matrix Obtaining microbial fusion similarity matrix , using the matrix and matrix Obtaining the disease fusion similarity matrix , specifically: in, It is a disease and disease The fusion similarity, It is a microorganism and microorganisms fusion similarity.
[0020] A4. Obtained microorganism-disease fusion similarity matrix 、 Combined with the associated feature matrix A to obtain the graph convolution adjacency matrix H , specifically: A5. Constructing the initial embedding matrix ,use , graph convolution adjacency matrix H The training set is composed of matrix A, and the PGVAEMDA model is trained using the training set to obtain the trained PGVAEMDA model and output the final embedding feature matrix; The initial embedding matrix as follows: The PGVAEMDA model includes: a nonlinear feature extraction unit, a linear feature extraction unit, an embedding matrix output unit, and a CatBoost unit; The nonlinear feature extraction unit is a variational autoencoder VAE (Variational Autoencoder, VAE), which is used to obtain the microbial nonlinear feature matrix and the disease nonlinear characteristic matrix ; The linear feature extraction unit uses the microorganism-disease association matrix Obtaining microbial linear feature matrix and disease linear feature matrix : The embedding matrix output unit converts the microorganism-disease association matrix , disease nonlinear characteristic matrix , microbial nonlinear characteristic matrix Disease linear feature matrix and microbial linear feature matrix Concatenate to obtain the final embedding feature matrix t , and output the final embedding feature matrix; The CatBoost unit is based on the final embedding feature matrix t Output disease-microorganism association values.
[0021] Specific embodiment three: the nonlinear feature extraction unit is an autoencoder VAE, including: an encoder layer and a decoder layer; The encoder layer is a graph convolutional network, using and H Get the intermediate layer hidden vector Z , and send the intermediate layer hidden vector to the decoder layer; Graph Convolutional Network (GCN) is a neural network that can learn low-dimensional representations and is mainly used in graph structures. By learning the adjacent or skip-neighbor information of nodes through graph convolution, the features of nodes are embedded into a richer representation space. The specific details are as follows Figure 3 Given the adjacency matrix of the heterogeneous network defined above H , the GCN of this heterogeneous network can be defined as: in, It is a node l Layer embedding, where l =1,…; for H The degree matrix of is a diagonal matrix, For the l The weight matrix of the layer nodes, kl is the embedding dimension, set to kl =128; Represents the graph convolution forward propagation function, and ReLU represents the activation function; Hidden vector of the middle layer Z Specifically: First, use the following formula to obtain : Then, use the Hidden vector of the middle layer Z : in, L= 3 is the number of graph convolution layers, It is the potential feature data The mean of is an intermediate variable, yes The variance of is the forward propagation function of the mean in the encoding layer, is the forward propagation function of the variance in the coding layer, e Obey standard Gaussian distribution N (0,1), Z is a vector with a potential dimension of 32, l is the number of graph convolutional layers.
[0022] Hidden vector of the middle layer Z The spatial distribution of is: Among them, refers to the conditional probability distribution of the latent vector Z, is the distribution of data in the matrix H, is the spatial distribution of the intermediate layer hidden vector Z; The decoder layer is used to utilize the intermediate layer latent vector Z Pair Matrix H Reconstruct and obtain the reconstructed matrix , specifically: in, is the complete forward propagation function from the encoder layer to the decoder layer; The nonlinear feature extraction unit is trained using the following loss function: The loss function of the encoder layer is divided into KL Divergence and sparsity penalty losses. KL Divergence provides an indicator to measure the degree of match between two probability distributions. The smaller the difference between the two distributions, the better. KLThe smaller the divergence, the sparse penalty loss can encourage the model to select and retain the most important features, resist the interference of noise, and improve the performance of the model. σ Sparse penalties are performed to obtain more expressive features: N = N d + N m in, is the KL divergence loss; is the sparse penalty loss; ||·||1 is the L1 norm, SpTh is the sparse threshold, λ is the sparse penalty coefficient, , N is the total number of microorganism and disease samples, is the mean of the feature distribution, is the standard deviation of the characteristic distribution; The present invention sets SpTh =10, λ =0.1; The loss function for training the decoder layer is: in, is the loss function of the decoder layer; is a matrix H The n Row features, is a matrix H' Middle n Reconstruction features.
[0023] Therefore, the overall loss of the nonlinear feature extraction unit is specifically: Where, Loss T is the overall loss function of the nonlinear feature extraction unit.
[0024] In this step, graph convolution is a powerful neural network architecture specifically designed to process graph-structured data. Its role is to effectively capture the dependencies between nodes (microorganisms or diseases) in the graph and reveal complex associations by learning the embedded representation of the nodes; the sparse penalty constraint can force many parameters or feature weights in the model to become zero, thereby automatically performing feature selection. This means that the model only retains the features that contribute most to the prediction target, while ignoring those unimportant features, reducing noise and outliers in the data, and thus improving the performance of the model. The introduction of graph convolution and sparse penalty terms in the variational autoencoder can significantly improve the performance of the model in extracting microbial and disease features, enhance the accuracy and generalization ability of feature representation, and at the same time improve the efficiency and interpretability of the model through automatic feature selection and reducing the risk of overfitting, providing strong support for disease discovery and precision medicine. We use a variational autoencoder model that integrates graph convolution and sparse penalty terms to extract the potential nonlinear features of the data. The general structure is as follows: Figure 4 As shown in the figure, the variational autoencoder (VAE) of the present invention realizes the conversion mapping by constraining the latent variable Z to obey the standard normal distribution (KL divergence) and reconstructing the data, thereby extracting the potential features of the data through the potential distribution of the data; Specific embodiment 4: the linear feature extraction unit uses the microorganism-disease association matrix Obtaining microbial linear feature matrix and disease linear feature matrix , specifically: Negative Matrix Factorization (NMF) is a commonly used feature extraction method that can decompose a non-negative matrix into the product of two or more non-negative matrices. Figure 5 As shown, this step is to A Decompose it as shown in the formula: in, r Represents the feature dimension, setting r =32; and are the non-negative feature matrices of diseases and microorganisms, respectively.
[0025] The goal of NMF is to minimize the difference between the original matrix A and the approximation matrix The difference between , we use Euclidean distance to measure the difference, specifically: Where ||·||F represents the Frobenius norm; This step is as follows Figure 5 As shown, we only focus on the front of A r singular values, and continuously looking for two non-negative r-dimensional matrices to approximate the original matrix A. During the optimization process, and The update of is achieved through an iterative multiplicative algorithm, which ensures the non-negativity of the two matrices during the decomposition process. and To approximately replace the original matrix A, the low-dimensional linear feature extraction function of the data is realized.
[0026] Specific implementation method five: Figure 6 As shown, the embedding matrix output unit uses the microorganism-disease association matrix , disease nonlinear characteristic matrix and microbial nonlinear characteristic matrix , the obtained disease linear feature matrix and microbial linear feature matrix The final embedding feature matrix is obtained by concatenation and output, specifically: First, the concatenation matrix 、 and Obtain the final disease embedding feature matrix , the splicing matrix 、 and Obtain the final microbial embedding feature matrix : in, is the final embedding feature matrix of the disease, It is the final embedding feature matrix of microorganisms; Then, obtain the microorganism-disease association matrix The microorganism-disease association pair corresponding to the value 1 is the positive example. At the same time, the number of positive examples is obtained, and then the negative examples with the same number of positive examples are randomly selected ( The microorganism-disease pairs corresponding to the value 0 in the matrix are selected to form the first set of positive and negative examples, and the row labels corresponding to the microorganism-disease pairs in the first set in the matrix A are obtained. - Column number , then splice Middle j Line and Middle i Rows are used as the final embedding feature matrix t The s OK : in, is the first s OK; In the present invention, the final embedding feature matrix input to the CatBoost unit in each round of training is the matrix corresponding to the current training set t The embedding feature matrix output by the embedding feature matrix output unit is the embedding feature of all microorganisms and disease pairs learned by the model after multiple rounds of training.
[0027] Specific implementation method 4: The CatBoost unit is specifically: in, It is m decision tree model; for each sample , calculate the corresponding predicted value With the previous m -1 prediction value of the decision tree model the difference between the calculated weighted values; M is the total number of microorganism-disease samples finally embedded in the feature matrix, is the target value.
[0028] The following loss function is used to train the CatBoost unit: in, is the positive sample label in the final embedded feature matrix, express The probability of being predicted as a positive sample, is the negative sample label in the second training set, express The probability of being predicted as a negative example.
[0029] Positive examples indicate that there is an association relationship, and negative examples indicate that there is no association relationship. Example: In order to verify the beneficial effects of the present invention, the present invention conducted the following experiments: This example uses 2-, 5-, and 10-fold cross validation to evaluate the performance of the evaluation model PGVAEMDA proposed in this invention. The ROC images obtained based on 2-CV, 5-CV, and 10-CV are as follows: Figure 7 The comparison of AUC values based on 2-CV, 5-CV, 10-CV and other models is shown in Table 1.
[0030] Table 1 Comparison of AUC values of PGVAEMDA model and other models under 2-fold, 5-fold, and 10-fold cross validation on the same dataset In the prediction results, the present invention demonstrated the top 15 microorganisms associated with Crohn's disease and the top 15 microorganisms associated with psoriatic arthritis. The verification results are shown in Tables 2 and 3.
[0031] Table 2 Top 15 microorganisms associated with Crohn's disease
[0032] Table 3 Top 15 microorganisms associated with psoriatic arthritis The present invention first obtains their correlation matrix based on the known microorganism-disease correlation relationship , and then the Gaussian interaction spectrum kernel matrix of the disease and microorganisms was obtained by calculation ; Use the Medical Subject Headings (MeSH) database and STRINGv11 to obtain disease symptom similarity matrices and microbial sequence similarity matrix . The known association information is enriched by integrating the Gaussian similarity matrix and the symptom similarity (sequence similarity) matrix.
[0033] This paper innovatively improves the application of traditional variational autoencoders (VAEs). Unlike traditional VAEs, which are primarily used for data generation, this paper uses them to learn the underlying distribution of data, thereby obtaining a deep nonlinear feature representation of the data. To further enhance feature extraction capabilities, we innovatively combine graph convolutional networks (GCNs) with a sparse penalty mechanism to construct a more powerful nonlinear feature extraction module. Furthermore, by introducing non-negative matrix factorization (NMF) technology, we simultaneously obtain explicit linear feature representations of microorganisms and diseases. Ultimately, by performing multimodal feature fusion of the linear and nonlinear features of microorganisms and diseases with the known association matrix A, we construct a comprehensive feature space for training the CatBoost model.
[0034] Through performance evaluation, the PGVAEMDA model has certain advantages over previous methods of constructing association relationships, and the prediction results show that the association prediction system has certain reliability.
Claims
1. A microbe-disease association prediction system based on sparse variational autoencoders, characterized by The system includes: an association matrix construction module, a module for obtaining the final embedded features to be tested, and a module for predicting the microorganism-disease association relationship; The association matrix building module is used to obtain the known microorganism-disease association matrix , and the matrix Input to the final embedding feature acquisition module to be tested; Microbe-Disease Association Matrix Middle i Rank j The elements of the column are: in, It is i disease, It's a sign of disease. It is microorganisms, is the microbe-disease association matrix Middle i Rank j Elements of the column; The final embedded feature acquisition module to be tested uses the microorganism-disease association matrix Obtaining a final embedding feature matrix, and using the final embedding feature matrix to obtain a final embedded feature to be tested, and inputting the final embedded feature to be tested into a microorganism-disease association prediction module; The method of using the final embedded feature matrix to obtain the final embedded features to be tested is specifically as follows: obtaining the rows corresponding to the disease to be tested and the microorganism to be tested in the final embedded feature matrix , and As the final embedded feature to be tested; The microorganism-disease association prediction module inputs the final embedded features to be tested into the trained CatBoost unit to obtain the predicted value of the association between the disease to be tested and the microorganism.
2. The microbe-disease association prediction system based on sparse variational autoencoders according to claim 1, characterized in that: The final embedding feature matrix and the trained CatBoost unit are obtained by: A1. Using the Microbe-Disease Association Matrix Obtaining the microbial Gaussian interaction spectrum kernel similarity matrix and disease Gaussian interaction spectrum kernel similarity matrix : in, It is j microorganisms and The Gaussian interaction spectrum kernel similarity of microorganisms, It is diseases and The Gaussian interaction spectrum kernel similarity of each disease, is the normalized microbial nuclear bandwidth, is the matrix A column vector, is the matrix A column vector, It is a microbial nuclear broadband, is the total number of columns in matrix A, is the normalized disease kernel bandwidth, is the matrix A row vector, is the matrix A's row vector, It is the disease core broadband, is a matrix A The total number of rows; A2. Obtaining disease symptom similarity matrix and sequence similarity matrices of microorganisms ; Disease Symptom Similarity Matrix The elements in are: in, Indicates disease and diseases Based on the similarity of symptoms, and Represents diseases and diseases association vectors with respective symptoms; The association vector stores the association value between the current disease and the corresponding symptom, and the association value includes: 0 and 1; 0 indicates that the current disease is not related to the current symptom, and 1 indicates that the current disease is related to the current symptom; The elements in the sequence similarity matrix of the microorganism are obtained by searching the STRINGv11 database to obtain the similarity between the protein sequences of the microorganisms. ; in, It is a microorganism and microorganisms The similarity between protein sequences; A3. Using the Matrix and matrix Obtaining microbial fusion similarity matrix , using the matrix and matrix Obtaining the disease fusion similarity matrix : in, It is a disease and disease The fusion similarity, It is a microorganism and microorganisms The fusion similarity; A4. Fusion of microbe-disease similarity matrix , disease fusion similarity matrix Combined with the microorganism-disease association matrix A to obtain the graph convolution adjacency matrix H ; A5. Constructing the initial embedding matrix ,use , graph convolution adjacency matrix H and matrix A to form a training set, use the training set to train the PGVAEMDA model, obtain the trained PGVAEMDA model and output the final embedding feature matrix; Initial embedding matrix , specifically: The PGVAEMDA model includes: a nonlinear feature extraction unit, a linear feature extraction unit, an embedded feature matrix output unit, and a CatBoost unit; The nonlinear feature extraction unit is used to obtain the microbial nonlinear feature matrix and the disease nonlinear characteristic matrix ; The linear feature extraction unit uses the microorganism-disease association matrix Obtaining microbial linear feature matrix and disease linear feature matrix ; The embedded feature matrix output unit converts the microorganism-disease association matrix , disease nonlinear characteristic matrix , microbial nonlinear characteristic matrix , disease linear feature matrix and microbial linear feature matrix Concatenate to obtain the final embedding feature matrix and output it; The CatBoost unit outputs a disease-microbe association value based on the final embedded feature matrix.
3. The microbe-disease association prediction system based on sparse variational autoencoders according to claim 2, characterized in that: The microbe-disease fusion similarity matrix in A4 , disease fusion similarity matrix Combined with the microorganism-disease association matrix A to obtain the graph convolution adjacency matrix H , specifically: 。 4. The microbe-disease association prediction system based on sparse variational autoencoders according to claim 3, characterized in that: The nonlinear feature extraction unit includes: an encoder layer and a decoder layer; The encoder layer is a graph convolutional network GCN, using and H Get the intermediate layer hidden vector Z , and send the intermediate layer hidden vector to the decoder layer; The utilization and H Get the intermediate layer hidden vector Z , specifically: First, use the following formula to obtain : Then, use the Hidden vector of the middle layer Z : in, l is the number of graph convolution layers, It is the potential feature data The mean of is an intermediate variable, yes The variance of is the forward propagation function of the mean in the encoding layer, is the forward propagation function of the variance in the coding layer, e Obey standard Gaussian distribution N (0,1), L is the total number of graph convolutional layers; The decoder layer uses the intermediate layer latent vector Z Pair Matrix H Reconstruct and obtain the reconstructed matrix : in, is the forward propagation function from the encoder layer to the decoder layer.
5. The microbe-disease association prediction system based on sparse variational autoencoders according to claim 4, characterized in that: The intermediate layer hidden vector Z The spatial distribution is as follows: in, Refers to the conditional probability distribution of the latent vector Z, is a matrix H The distribution of data in is the spatial distribution of the intermediate layer hidden vector Z.
6. The microbe-disease association prediction system based on sparse variational autoencoders according to claim 5, characterized in that: The nonlinear feature extraction unit is trained using the following loss function: N = N d + N m in, Loss T is the overall loss function of the nonlinear feature extraction unit, is the KL divergence loss; is the sparse penalty loss; ||·||1 is the L1 norm, SpTh is the sparse threshold, λ is the sparse penalty coefficient, , N is the total number of microorganism and disease samples, is the mean of the feature distribution, is the standard deviation of the characteristic distribution, is the loss function of the decoder layer, is a matrix H The n Row features, is a matrix H' Middle n Reconstruction features.
7. The microbe-disease association prediction system based on sparse variational autoencoders according to claim 6, characterized in that: The linear feature extraction unit uses the microorganism-disease association matrix Obtaining microbial linear feature matrix and disease linear feature matrix , specifically: Correlation Matrix A Perform negative matrix decomposition to obtain the microbial linear feature matrix and disease linear feature matrix : in, r is the characteristic dimension.
8. The microbe-disease association prediction system based on sparse variational autoencoders according to claim 7, characterized in that: The objective function of the linear feature extraction unit is: where ||·||F is the Frobenius norm.
9. The microbe-disease association prediction system based on sparse variational autoencoders according to claim 8, characterized in that: The embedded feature matrix output unit converts the microorganism-disease association matrix , disease nonlinear characteristic matrix , microbial nonlinear characteristic matrix , disease linear feature matrix and microbial linear feature matrix The final embedding feature matrix is obtained by concatenation, specifically: First, the concatenation matrix 、 and Obtain the final disease embedding feature matrix , the splicing matrix 、 and Obtain the final microbial embedding feature matrix : in, is the final embedding feature matrix of the disease, It is the final embedding feature matrix of microorganisms; Then, obtain the microorganism-disease association matrix The microorganism-disease association pair corresponding to the value 1 is the positive example. At the same time, the number of positive examples is obtained. Then, the negative examples with the same number as the positive examples are randomly selected. The selected positive and negative examples are combined into the first set. The row labels corresponding to the microorganism-disease pairs in the first set in the matrix A are obtained. - Column number , then splice Middle j Line and Middle i Rows are used as the final embedding feature matrix t A line in: in, yes No. OK, yes No. OK, is the first s OK, The microorganism-disease pair corresponding to the value 0 is a negative example.
10. The microbe-disease association prediction system based on sparse variational autoencoder according to claim 9, characterized in that: The CatBoost unit is specifically: in, It is m A decision tree model, M is the total number of microorganism-disease samples finally embedded in the feature matrix, is the target value, It is before m -1 prediction value of the decision tree model, It is the s The predicted value of samples; The following loss function is used to train the CatBoost unit: Among them, among them, is the positive sample label in the final embedded feature matrix, express The probability of being predicted as a positive sample, is the negative sample label in the final embedded feature matrix, express The probability of being predicted as a negative sample; Positive examples indicate that there is an association relationship, and negative examples indicate that there is no association relationship.