A method and system for screening bladder cancer metabolic markers based on deep learning
In the analysis of metabolomics data of bladder cancer, using a missing value interpolation method combining deep learning with clustering and regression models, and using a variational autoencoder for feature screening and visualization, the complex problems of missing value processing and feature screening in the existing technology are solved, and efficient metabolic marker screening is achieved.
Patent Information
- Application Number
- CN202210600879.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-05-30
AI Technical Summary
When prior art screening metabolic markers of bladder cancer from metabolomics data, there are problems such as complex processing of missing values and difficult to visualize and interpret feature screening.
A deep learning-based method is adopted, combined with fuzzy C-mean clustering, K nearest neighbor method and Bayesian ridge regression model for missing value interpolation, and a variational autoencoder model is used for feature visualization and screening to achieve importance analysis and dimensionality reduction of m/z features.
It effectively solves the complexity of missing value interpolation and feature screening, reduces feature dimensions, and improves the screening efficiency and accuracy of metabolic markers of bladder cancer.
Smart Images

Figure CN114997303B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of metabolomics, and relates to a method and system for screening bladder cancer metabolic markers based on deep learning. Background Art
[0002] Bladder cancer is one of the top ten common cancers in the world. Although its etiology and fatal risk factors have been widely studied, due to its high recurrence rate and great clinical complexity, diagnosis and treatment are more difficult. Metabolomics has a unique prospect in the early diagnosis of bladder cancer using metabolic markers. It uses advanced mass spectrometry instruments to obtain mass spectrometry data, and uses statistical learning and other methods to find the mass-to-charge ratio (m / z) of a group of metabolite ions between bladder cancer patients and healthy control groups from the mass spectrometry data. Their mass spectrometry intensities are significantly different between different groups, and this group of differential ions can effectively separate bladder cancer patient samples from healthy control samples. By analyzing and identifying the names of the metabolites represented by this group of m / z characteristics through biotechnology, and determining their metabolic changes in cancer patients, this group of metabolites can be used as metabolic markers for the early diagnosis of bladder cancer.
[0003] Mass spectrometry technology (MS) has the advantages of high sensitivity and high resolution, and can simultaneously detect thousands of metabolite ions in a large dynamic range. At present, liquid chromatography-mass spectrometry (LC-MS) technology has been applied in the field of bladder cancer metabolomics. Metabolomics data has complex characteristics such as high-dimensional features, small samples, and high noise. There is a possibility of obtaining false positive results in feature screening by current conventional methods. How to correctly analyze metabolomics data is the key, and extracting valuable information from metabolomics data and screening potential metabolic markers are the main problems to be solved in metabolomics data processing.
[0004] The XCMS platform is the most widely used open-source program for processing LC-MS metabolomics data. After the original data is processed by the program through peak selection, alignment, and grouping, it is converted into clean data that can be used for statistical analysis, but there is still a problem of missing values, which may affect the performance of the data mining model. Based on the iterative imputation method of the Bayesian ridge regression model, each missing feature is predicted as a function of other features, and this process of estimating the feature value is repeated multiple times. Repeating allows the use of optimized estimated values of other features as input in subsequent iterations of predicting missing values, and this process increases the computational complexity due to the overly large mass spectrometry dataset.
[0005] Performing feature selection after preprocessing the clean data is the key to determining bladder cancer disease metabolic markers. However, when screening features from the clean dataset through a deep learning model, using intensity values for calculation on neurons loses the key mass-to-charge ratio information. The selected features are difficult to visualize and interpret, which is not conducive to the discovery of metabolic markers.
[0006] Therefore, this field urgently hopes to propose a feature visualization method for neural networks to identify m / z features that have a significant impact on classification results, achieve feature dimension reduction, and facilitate the search for feature subsets that achieve the best classification results. In response to the problem of missing values in mass spectrometry data, a new interpolation method needs to be proposed to mine useful information in the data and facilitate the subsequent discovery of metabolic markers. Summary of the invention
[0007] The first purpose of the present invention is to solve the problems existing in the above-mentioned background technology, and propose a method for screening bladder cancer metabolic markers based on deep learning, which includes missing value interpolation and feature visualization based on a variational autoencoder model. The missing value interpolation integrates fuzzy C-means clustering, K nearest neighbor method and iterative interpolation algorithm based on Bayesian Ridge regression model to metabolomics data set missing values. In the first stage, fuzzy C-means clustering is used to perform unsupervised clustering, and samples in the data set are divided into different categories according to similarity; in the second stage, the best nearest neighbor set of each missing record is found according to the similarity of the sample by KNN algorithm calculation; in the third stage, each feature is modeled using a data set consisting of the selected sample and the nearest neighbor set, and the Bayesian Ridge regression model is used to iteratively interpolate the missing values in the sample set. Based on the feature visualization of the variational autoencoder model, unsupervised training is performed on bladder cancer samples and healthy control samples, and a weight threshold analysis based on back propagation is proposed for the weight value of the encoder. The importance of the input layer m / z feature is determined according to the weight size of each neuron, and the m / z feature that meets the threshold condition is screened out to achieve feature dimensionality reduction, which is convenient for diagnosing the discovery of bladder cancer metabolic markers.
[0008] In order to achieve the above purpose, the technical method adopted by the present invention is as follows:
[0009] Based on non-targeted bladder cancer metabolomics data, the present invention pre-processes the raw data of mass spectrometry to obtain clean data, and improves the performance of data mining through pre-processing methods of missing value interpolation and data normalization. The pre-processed data is used to train a variational autoencoder model, and back propagation is used to determine the neuron weights of the previous layer from the encoding layer weight matrix of the model, and a weight threshold is set to achieve m / z feature screening, so as to further determine the m / z features that have the ability to distinguish between cancer samples and control samples.
[0010] Furthermore, the specific steps are as follows:
[0011] Step (1): Raw data pre-processing
[0012] 1-1 Obtain the original mass spectrometry data, and then use the peak selection and alignment functions in the IPO (Isotopologue Parameter Optimization) toolkit to find the best parameters for the original mass spectrometry data, obtaining the best mass spectrometry parameters;
[0013] 1-2 Input the best mass spectrometry parameters into the XCMS program package to obtain clean data from the original mass spectrometry data of all samples; this clean data can be represented as , where b in represents the intensity value corresponding to the nth mass-to-charge ratio feature of the ith sample.
[0014] Step (2): Perform missing value imputation on the clean data obtained in step (1)
[0015] 2-1 Perform fuzzy C-means clustering (FCM) on the clean data to obtain multiple subsets of samples;
[0016] The above operations make the subsamples in the same cluster more similar to each other. Imputing the multiple subsets of samples in turn also reduces the computational complexity of the Bayesian ridge regression model imputation and improves the computational efficiency.
[0017] 2-2 Traverse all subsets of samples, and screen out the subsamples R i ;
[0018] 2-3 Randomly find a non-empty element L in the subsample R i , and update the initial intensity value in this element L to a null value.
[0019] 2-4 Use the KNN-Impute algorithm to perform missing value imputation on the subset of samples where the updated subsample R i is located. Set the value of K to satisfy 2 to Np - 1, where Np represents the number of samples in the subset of samples where the subsample R i is located, and obtain the intensity imputation value of the element L in the subsample R i .
[0020] 2-5 Calculate the error between the initial intensity and the imputed value intensity of the element L under different values of K, and find the best value of K corresponding to the minimum error.
[0021] 2-6 According to the Euclidean distance formula, find the K nearest subsamples to the subsample R i in the subset of samples where it is located, and combine them into an imputation subset S. i
[0022] 2-7 Perform iterative imputation on the imputation subset S using the Bayesian ridge regression model to achieve imputation of all missing values in all subsamples in the imputation subset S, and save the imputed subsample Ri data
[0023] Repeat steps 2-3 to 2-7 from 2 to 8 until all the subsamples with missing values in the subsample set where subsample R i is located are all imputed.
[0024] 2-9 Return to step 2-2, select the next subsample set until all subsample sets are imputed.
[0025] Step (3), Data normalization
[0026] Perform min-max normalization on the data imputed in step (2) to measure different mass-to-charge ratio intensity information for easy calculation and comparison.
[0027] Step (4), Build a variational autoencoder model
[0028] 4-1 Variational autoencoder network construction
[0029] The variational autoencoder network (VAE) uses a fully connected layer neural network architecture, which consists of 5 layers: an input layer (L 1 ), three hidden layers (h 1 , h 2 and h 3 ) and an output layer (L 2 ).
[0030] 4-2 Loss function construction
[0031] The unsupervised learning of this model mainly consists of the Kullback-Leible divergence (KL-divergence) between the Gaussian distribution and the actual distribution of the encoding and the marginal likelihood modeled as categorical cross-entropy to minimize the reconstruction loss between the original data and the reconstructed data.
[0032] Step (5), m / z feature screening based on backpropagation-based weight threshold analysis
[0033] A threshold analysis method based on backpropagation is proposed for the weight parameter W 1 of the L 1 layer neural network. This method mainly starts from the weight matrix of each neuron in the h 2 layer, sorts according to the weight value size, determines the neuron with the largest weight in the h 1 layer, performs threshold analysis on the weight matrix of this neuron, and realizes the m / z feature screening of the input layer L 1 .
[0034] The specific method steps are as follows:
[0035] 5-1 Obtain the standard deviation and mean of each mass-to-charge ratio intensity in the dataset after normalization in step (3)
[0036] 5-2 Traverse the hidden layer h 2 The weight matrix W of the i-th neuron in the layer 2 , first sort by the magnitude of the weight values, take the maximum weight value to identify the previous hidden layer h 1 of the j-th neuron, and use the weight matrix W of the j-th neuron 1 j Multiply by the standard deviation of the intensities of all mass-to-charge ratios to obtain a new weight matrix variable W 1 ′ j . Among them, the weight matrix W 1 j is calculated from the intensity values of the input layer, and each neuron represents an m / z feature.
[0037] 5-3 Use the weight matrix variable W 1 ′ j to calculate the threshold T, where T = mean(W 1 ′ j ) + gama * std(W 1 ′ j ), gama = [1, 3]
[0038] 5-4 Screen out the neurons corresponding to the weight values greater than the threshold T in the weight matrix variable W 1 ′ j until all neurons in the hidden layer h 2 are traversed, and finally all m / z features that meet the threshold conditions are obtained to get the set M.
[0039] 5-5 Take the mean of the intensities of each m / z feature in all samples, perform peak query according to the mean of all mass-to-charge ratio intensities, determine the m / z set N corresponding to its local maximum value, and take the intersection of the sets M and N to obtain the reduced m / z feature set P.
[0040] Step (6), Metabolite biomarker screening
[0041] Perform secondary screening on the screened feature set P in step (5), adopt the recursive feature elimination method, and combine with the random forest model to calculate the importance of the m / z features. When the classification accuracy of the random forest model reaches the highest, screen and obtain the optimal m / z features, and these features can be used as potential metabolite biomarkers for serum bladder cancer metabolomics.
[0042] The second object of the present invention is to provide a bladder cancer metabolite biomarker screening system based on deep learning, including:
[0043] A data preprocessing module for obtaining clean data from the original mass spectrometry data;
[0044] The missing value filling module is used to fill in the missing values of the data processed by the data preprocessing module;
[0045] The data normalization module normalizes the data processed by the missing value filling module;
[0046] The m / z feature screening module uses the trained variational autoencoder network VAE to implement weight threshold analysis based on backpropagation to screen m / z features;
[0047] The potential metabolite biomarker screening module uses the m / z features screened by the m / z feature screening module and adopts the recursive feature elimination method combined with the random forest model to screen and obtain the optimal m / z features.
[0048] The third object of the present invention is to provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method.
[0049] The fourth object of the present invention is to provide a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method is implemented.
[0050] The beneficial effects of the present invention are as follows:
[0051] 1. The present invention proposes to use backpropagation threshold analysis for visual interpretation of mass spectrometry features of the variational autoencoder network, realizes finding differential metabolite ion features related to bladder cancer patients from a large number of mass spectrometry features, and greatly reduces the feature dimension; subsequently, the recursive feature elimination method is used to shorten the time consumed to find a set of feature subsets with the highest classification accuracy. This set of feature subsets has the ability to well distinguish cancer patients from healthy people as potential metabolite biomarkers.
[0052] 2. The present invention proposes a method for imputing missing values for mass spectrometry data, improves the ability to mine useful information from data, and avoids that data containing null values will make the modeling process fall into chaos and lead to unreliable outputs; reliable results are output for subsequent visual interpretation of features of the variational autoencoder network. Description of the Drawings
[0053] Figure 1 is the flow chart of the selection of metabolic biomarkers based on bladder cancer metabolomics of the present invention;
[0054] Figure 2 is the flow chart of missing value imputation based on fuzzy C-means clustering, KNN imputation, and iterative imputation based on Bayesian ridge regression model in the present invention;
[0055] Figure 3 is the neuron architecture diagram of the variational autoencoder in the present invention;
[0056] Figure 4 It is the structural diagram of the encoding method of the variational autoencoder in the present invention;
[0057] Figure 5 It is the loss function diagram of the variational autoencoder training in the example of the present invention;
[0058] Figure 6 It is the visual flowchart of feature screening based on backpropagation threshold analysis of the present invention. Detailed implementation manners
[0059] By referring to exemplary embodiments, the objects and functions of the present invention and the methods for achieving these objects and functions are clarified. However, the present invention is not limited to the disclosed exemplary embodiments and can be implemented in different forms. The essence of the specification is only to help those skilled in the relevant art comprehensively understand the specific details of the present invention.
[0060] The following further describes the invention with reference to the accompanying drawings.
[0061] This embodiment is to be implemented using the method of the present invention. A metabolomics dataset of bladder cancer based on LC-MS is collected. This dataset includes sample data of 64 bladder cancer patients, sample data of 60 healthy control group patients, and 7 equal-volume mixed solutions divided according to the sample injection order as QC sample data.
[0062] A method for screening metabolic markers of bladder cancer based on deep learning, the overall process is as Figure 1 shown, and specifically includes the following steps.
[0063] Step (1), preprocessing of raw data
[0064] 1-1 For the raw mass spectrometry data in the.raw format obtained from the mass spectrometer, use the MSConvert software to convert it into the.mzXML format; in order to improve the accuracy of peak selection and alignment and save the time consumed for finding program parameters, use the QC samples to find the optimal parameters of the peak selection function and the alignment function.
[0065] Input the mass spectrometry data of 7 QC samples, import it into the IPO package, and use the built-in peak selection function CentWave to find parameters. The best parameter results are: peakwidth=(6.64,27.3), mzdiff=0.00835, prefilter=(3,383.56), ppm=73, signal-to-noise=10, noise=231.4; the best parameter results of the alignment function obiWarp are: profstep=1, respon=29.88, gapInit=0.548, gapExtend=2.62; the feature grouping parameters are: bw=0.88, minFrac=0.2, minSample=1, mzwid=0.0344, maxFeatures=50.
[0066] 1-2 Use the best parameters in step 1-1, and use the CentWave and Obiwrap methods in the XCMS package to perform peak selection and alignment on the 124 original sample data. Then, obtain the clean data through the peak grouping method. The clean data can be represented by where b in represents the intensity value corresponding to the nth mass-to-charge ratio feature of the ith sample. The mass-to-charge ratio range is 50-1000 Da.
[0067] Step (2): Perform missing value imputation on the clean data obtained in step (1)
[0068] As Figure 2 shown, the input of this algorithm is the clean data obtained in step (1). Combining fuzzy C-means clustering, KNN imputation, and an iterative imputation method based on the Bayesian ridge regression model, it outputs complete data without missing values. The specific steps are as follows:
[0069] 2-1 First, through fuzzy C-means clustering, perform fuzzy C-means clustering on the clean data L of the samples containing 64 bladder cancers and 60 healthy controls. The clustering category is 4, and it is divided into 4 sub-sample sets D j .
[0070] 2-2 Traverse all the sub-sample sets D j , take the ith missing value-containing sub-sample R j in each sub-sample set D i , and remove the ith sub-sample data from the sub-sample set D j to form a new data set P.
[0071] 2-3 Randomly find a non-missing value element in R i , save its true value AV, then update the value of this element to be empty, and update the updated R iCombine with P to form a new dataset temp.
[0072] 2-4 Use the KNN-Impute algorithm to impute the new dataset temp, set different K values, where the K value satisfies 2 to Np-1, and Np represents the number of samples in the new dataset temp, to obtain the updated R under different K values i of the imputed value IV
[0073] 2-5 According to the error RMSE calculation formula, determine the minimum RMSE from the true value and the imputed value, and use the corresponding K value as the optimal value, where the RMSE calculation formula:
[0074]
[0075] 2-6 According to the Euclidean distance formula, find the K sub-samples in dataset P that are closest to the unupdated sub-sample R i to form dataset S.
[0076] 2-7 Use the Bayesian Ridge Regression model (BayesianRidge) to perform iterative imputation on dataset θ, and the missing values of sub-sample R i are imputed, and the imputed sub-sample R is saved i .
[0077] 2-8 Take the next sub-sample with missing values in sub-sample set D j , update the sub-sample R in temp to the imputed data, and repeat steps 2-3 to 2-7 until all samples in sub-sample set D i are imputed. j All samples in
[0078] 2-9 Return to step 2-2, select the next sub-sample set until all missing values in all sub-sample sets are imputed. At this time, the clean data L has no missing values, and proceed to the next step.
[0079] Step (3), Data normalization
[0080] Perform min-max normalization on each value in the imputed clean data L, and its calculation method is as follows:
[0081]
[0082] where min and max represent the minimum and maximum intensity values in the intensity vector, and b n represents the vector formed by the intensities of the nth mass-to-charge ratio in all samples, represents the normalized data.
[0083] Step (4), Build a variational autoencoder model
[0084] 4-1 Construction of Variational Autoencoder Network
[0085] As Figure 3 、 Figure 4 shown, the proposed neural network architecture consists of 5 layers, including an input layer (L 1 ), three hidden layers (h 1 , h 2 and h 3 ) and an output layer (L 2 ). The dimension of the input layer L 1 is 965. h 1 is a fully connected layer with 256 neurons and the activation function selu. Instead of directly generating the encoding given the input, the h 2 layer of the encoder generates the mean encoding μ and the standard deviation ρ, and then the actual encoding is randomly sampled from the Gaussian distribution with mean μ and standard deviation ρ, with the number of neurons being 8. The structure of the decoder h 3 layer is the same as that of the h 1 layer, with the number of neurons being 256. The number of neurons in the output layer L 2 is 965, and the activation function is sigmoid.
[0086] 4-2 Construction of Loss Function
[0087] The loss function consists of two parts. The first is the usual reconstruction loss, which forces the autoencoder to reproduce its input by calculating the cross-entropy loss between the original input and the reconstructed output. The second is the latent loss, which makes the encoding of the autoencoder look like it is sampled from a simple Gaussian distribution: it is the KL divergence between the target distribution and the actual distribution of the encoding.
[0088] The reconstruction loss is calculated as follows:
[0089]
[0090] where y i represents the class of sample i, with the positive class being 1 and the negative class being 0, p i represents the probability that sample i is predicted as the positive class, N represents the number of samples, and K represents the number of features in the input layer.
[0091] The latent loss is calculated as follows:
[0092]
[0093] where σ i and μ i are the mean and standard deviation of the weights of the i-th neuron in the encoder
[0094] Final loss function: Loss = Reco_Loss + KL_Loss
[0095] The preprocessed data set is trained using the above variational autoencoder network framework and loss function. Each batch has 16 samples and is trained 100 times in total. Finally, the loss function converges. The loss function effect is as follows: Figure 5 Shown
[0096] Step (5): m / z feature screening based on weight threshold analysis of back propagation
[0097] like Figure 6 As shown, the present invention proposes a method based on neural network weight matrix analysis to associate the encoded features with the input m / z features. Through threshold analysis, the screened features support the clustering and classification tasks required for metabolic markers. The specific steps are as follows:
[0098] 5-1 Calculate the intensity standard deviation set std_spec and mean set mean_spec for each m / z in the data set trained by the variational autoencoder, extract the trained model parameters, and obtain the encoder h 1 Layer weight matrix W 1 and h 2 Layer weight matrix W 2 ,
[0099] 5-2 Traversing the hidden layer h 2 The weight matrix W 2 For each column, sort the weight values of the column by size, select the neuron j corresponding to the maximum weight value, and this neuron belongs to the hidden layer h 1 , the weight matrix of the neural cloud The input layer L 1 Calculate and convert the weight matrix W 1 j The new weight matrix W is formed by multiplying the mass-to-charge ratio intensity standard deviation std_spec 1 ' j .
[0100] 5-3 Set the threshold T, which is calculated as follows: T = mean (W 1 ' j )+gama*std(W 1 ' j ), gama takes a constant of 2.5.
[0101] 5-4 When W 1 ' j When the value in is greater than T, the value obtained in W 1 ' j The position j in the input layer L 1For the neurons, these neurons are screened out, and each neuron corresponds to an m / z feature. Return to step 5-2, return all neurons that meet the threshold conditions, and screen out the corresponding set M of m / z features.
[0102] 5-5 Use the local maximum method (argrelextrema) to query the mean intensity of each m / z. A total of 317 maximum peaks are found, and the corresponding set N of m / z features is determined.
[0103] 5-6 Take the intersection P of the m / z features in steps 5-4 and 5-5 to achieve the process of reducing the dataset from high dimension to low dimension.
[0104] Step (6), Screening of metabolic markers
[0105] 6-1 Re-screen the screened feature set P in step (5). Use the recursive elimination algorithm, and use the Gini index of the random forest model to evaluate the importance of m / z features. Remove the least important features in turn, and calculate the classification accuracy of the model. When the classification accuracy is the largest, obtain 18 m / z features trained by the model at this time. This set of m / z features can be used as potential metabolic markers for bladder cancer metabolomics (see Table 1).
[0106] 6-2 Use the 18 potential metabolic marker m / z features obtained by screening in step 6-1, and use classification models such as random forest (RF), support vector machine (SVM), decision tree (DT), and gradient boosting decision tree (GBDT) for 5-fold cross-validation. The evaluation indicators are the area under the ROC curve (AUC), sensitivity, and specificity. The average values of the evaluation indicators are shown in Table 2.
[0107] Table 1 Mass-to-charge ratios of metabolite ions of potential metabolic markers screened in this example
[0108]
[0109] In the Olatomiwa article, it was found that there are 8 metabolic markers that can be used as bladder cancer metabolomics. Four of their corresponding mass-to-charge ratios were found in Table 1, namely 207.10, 475.24, 257.15, and 330.22. Therefore, the method proposed in this study for discovering metabolic markers for bladder cancer diagnosis has a certain effect.
[0110] Table 2 Average values of evaluation indicators for 5-fold cross-validation under different classification models
[0111]
[0112]
[0113] Overall evaluation conclusion:
[0114] (1) The present invention is directed to the analysis of bladder cancer metabolomics data, and provides a data imputation method based on the similarity between samples. First, the fuzzy C-means clustering is used to screen out the first-level similarity sample set, the KNN is used to find the nearest neighbor set in the first-level similarity sample set as the second-level similarity sample set, and the Bayesian ridge regression model is used to perform Bayesian ridge regression model imputation on the second-level similarity sample set. This not only helps to improve the imputation accuracy, but also simplifies the computational complexity of the iterative imputation with the reduction of the sample quantity and shortens the imputation time.
[0115] (2) The present invention proposes a method for extracting features based on the variational autoencoder model for bladder cancer mass spectrometry data. By using the feature weight value matrix in the middle of the encoder, the weights of the neurons in the previous layer are inversely inferred, and the threshold cut-off condition is set, and the m / z features of the input layer are mapped according to the magnitude of the weight values. This method reduces the m / z feature dimension, facilitates the implementation of the m / z feature subset with the best classification effect, achieves a good classification effect in bladder cancer and healthy control samples, and the identified m / z features have been found in the Olatomiwa study.
Claims
1. A method for screening bladder cancer metabolic markers based on deep learning, characterized in that it includes the following steps: Step (1), preprocessing of original data 1-1 Obtain the original mass spectrometry data, and then use the peak selection and alignment functions in the IPO toolkit to find the best parameters for the original mass spectrometry data to obtain the best mass spectrometry parameters; 1-2 Input the optimal mass spectrometry parameters into the XCMS R package program to obtain clean data from all the original mass spectrometry data; this clean data is represented by , where b in represents the intensity value of the nth mass-to-charge ratio feature of the ith sample; Step (2), perform missing value imputation on the clean data obtained in step (1); 2-1 Perform fuzzy C-means clustering on the clean data to obtain multiple subsample sets; 2-2 Traverse all subsets of samples, and filter out the subset of samples R with missing values from each subset of samples i ; 2-3 in the sub-sample R i Randomly find an element L with a non-empty value in it, and update the initial strength value in this element L to an empty value; 2-4 Use the KNN-Impute algorithm to perform missing value imputation on the updated subsample set where R is located. Set the value of K to satisfy 2 to Np-1, where Np represents the number of samples in the subsample set where R is located, and obtain the intensity imputation value of element L in subsample i. i Perform missing value imputation on the subsample set where i is located. Set the value of K to satisfy 2 to Np-1, where Np represents the number of samples in the subsample set where R is located. i Obtain the intensity imputation value of element L in subsample i. 2-5 Calculate the error between the initial intensity and the imputed value intensity of element L under different K values, and find the best K value corresponding to the minimum error; 2-6 According to the Euclidean distance formula, from the sub-sample R i find the K nearest sub-samples in the sub-sample set where it is located and combine them into the imputation subset S; i Iteratively impute the imputation subset S using the Bayesian ridge regression model to achieve imputation of all missing values in all subsamples of the imputation subset S, and save the subsample R i after imputation; Repeat steps 2-3 to 2-7 from 2-8 until all the subsamples with missing values in the subsample set where i subsample R is located are all imputed; 2-9 Return to step 2-2, select the next subsample set until all subsample sets are completely imputed; Step (3), normalize the data imputed in step (2); Step (4), build a variational autoencoder model; The variational autoencoder network VAE adopts a fully connected layer neural network architecture, including an input layer L 1 , three hidden layers h 1 , h 2 , h 3 and an output layer L 2 ; Step (5), screening of mass-to-charge ratio m / z features based on weight threshold analysis of backpropagation, specifically as follows: 5-1 Obtain the standard deviation and mean of the intensity of each mass-to-charge ratio m / z in the dataset after normalization in step (3); 5-2 Traverse the hidden layer h 2 The weight matrix W of the i-th neuron in the layer 2 , first sort by the magnitude of the weight values, and take the largest weight value to identify the j-th neuron in the previous hidden layer h 1 , multiply the weight matrix of the j-th neuron by the standard deviation of the intensities of all mass-to-charge ratios to obtain a new weight matrix variable W 1 ′ j ; 5-3 Use the weight matrix variable W 1 ′ j Calculate the threshold T, where T = mean(W 1 ′ j ) + gama * std(W 1 ′ j ), where gama represents a constant; 5-4 Screen out the weight matrix variable W 1 ′ j the neurons corresponding to the weight values greater than the threshold T until the hidden layer h is traversed 2 All neurons finally form a set M consisting of all m / z features that meet the threshold conditions; 5-5 Take the mean of the intensities of each m / z feature in all samples, and then perform peak query according to the mean of all mass-to-charge ratio intensities to determine the mass-to-charge ratio corresponding to the local maximum in the mean of all mass-to-charge ratio intensities, thereby forming the m / z set N; finally, find the intersection of sets M and N to obtain the reduced m / z feature set P; Step (6), screening of metabolic markers Perform secondary screening on the feature set P screened in step (5), adopt the recursive feature elimination method, and combine the random forest model to calculate the importance of m / z features. When the classification accuracy of the random forest model reaches the highest, screen and obtain the optimal m / z feature, and this feature is used as the potential metabolic marker of serum bladder cancer metabolomics.
2. A method for screening bladder cancer metabolic markers based on deep learning according to claim 1, characterized in that the variational autoencoder network minimizes the reconstruction loss between the original data and the reconstructed data through the VAE cost function composed of the KL-divergence and the marginal likelihood modeled as categorical cross-entropy.
3. A deep learning-based bladder cancer metabolic marker screening system for implementing the method described in any one of claims 1-2, including: A data preprocessing module for obtaining clean data from the original mass spectrometry data; A missing value filling module for filling in missing values in the data processed by the data preprocessing module; A data normalization module for performing normalization processing on the data processed by the missing value filling module; An m / z feature screening module for implementing the screening of m / z features based on weight threshold analysis of backpropagation using the trained variational autoencoder network VAE; A potential metabolic marker screening module for screening and obtaining the optimal m / z feature by using the recursive feature elimination method in combination with the random forest model for the m / z features screened by the m / z feature screening module.
4. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed on a computer, the computer is made to execute the method described in any one of claims 1-2.
5. A computing device, comprising a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 1-2 is implemented.
Citation Information
Patent Citations
Defect high-risk module identification method based on software network
CN110147321A
Biomarker database generation and use
IN201817041006A