Cancer subtype classification and prognosis prediction method based on generic cancer multi-omics data
By building a CA-CAE model to integrate multi-omics data, the problem of insufficient data integration in cancer subtype classification and prognosis prediction in existing technologies has been solved, achieving more accurate cancer subtype identification and survival prediction.
Patent Information
- Application Number
- CN202510717016.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Existing technologies have difficulty in effectively integrating multi-omics data in cancer subtype classification and prognosis prediction, resulting in limited predictive capabilities and an inability to reveal the interactions and combined effects between different omics.
A deep learning-based method was used to construct a convolutional autoencoder (CA-CAE) model with a channel attention mechanism. This model integrated multi-omics data, including mRNA, DNA methylation, and miRNA data, and performed cancer subtype classification and prognosis prediction through deep neural networks.
It improves the accuracy of cancer subtype classification and the precision of survival prediction, can stably identify cancer subtypes in high-dimensional, heterogeneous data environments, and improve the performance of prediction models.
Smart Images

Figure CN120670897A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a cancer profiling classification and prognosis prediction method, and belongs to the technical field of bioinformatics. Background Art
[0002] Identification of disease subtypes is a core part of modern precision medicine practice, and is particularly valuable in the clinical diagnosis and treatment of tumors. Current studies have confirmed that malignant tumors of the same anatomical classification (such as breast cancer, lung cancer, etc.) can be further subdivided into biological subtypes with significant heterogeneity through molecular characteristics. In the field of prognostic assessment, the prognostic risk stratification model constructed by integrating genomic variations, transcriptional regulatory networks and microenvironmental characteristics can achieve dynamic prediction of patient survival. The key driver molecules identified by such models not only provide a basis for the screening of therapeutic targets, but also reveal abnormalities in the core signaling pathways in the evolution of tumor clones. The analysis of the association between this molecular mechanism and clinical phenotype will effectively promote the optimization of personalized monitoring programs and lay a theoretical foundation for the development of combination therapies targeting the tumor ecosystem.
[0003] In recent years, in order to classify cancer subtypes and predict prognosis from genomic data, researchers have developed a variety of computational strategies based on molecular feature recognition and machine learning, such as iClusterPlus, PINSPlus, SNF and PCA. With the rapid development of cancer multi-omics data and deep learning technology, research on cancer subtype classification and prognosis prediction has become more refined and accurate. These advanced technologies allow researchers to use large-scale gene expression, DNA methylation and miRNA expression data from databases such as UCSC Xena and TCGA to deeply explore the heterogeneity and potential subtypes among cancer patients. Genomic analysis can accurately identify the molecular characteristics of patient tumors and reveal the key molecular differences between different cancer subtypes, while feature selection methods based on machine learning (such as LASSO, random forest, XGBoost) can screen key genes related to prognosis and improve the accuracy of the model in predicting patient survival.
[0004] However, multi-omics data often contain omissions and errors. For example, due to variations in sample processing and sequencing, data may be inaccurate and missing. Moreover, most previous studies have mainly processed single-omics data, which can only provide information about a certain omics level. Due to the multi-level regulatory mechanisms of complex diseases such as cancer, this single perspective may miss a lot of key information and fail to reveal the interactions and combined effects between these different omics, limiting the predictive ability of the prognostic model. By combining deep learning technology, more accurate subtype classification and prognosis prediction models can be constructed, revealing the key molecular mechanisms between cancer subtypes and prognosis, thereby providing solid theoretical support for personalized treatment and precision medicine for cancer patients. Therefore, there is a need for a cancer subtype classification and prognosis prediction method that can integrate multi-omics data and deep learning algorithms. Summary of the Invention
[0005] The present invention aims to solve the problem of cancer subtype classification and prognosis prediction in pan-cancer multi-omics, and further proposes a method for cancer subtype classification and prognosis prediction based on pan-cancer multi-omics data.
[0006] The technical solution adopted by the present invention to solve the above problems is: the steps of the present invention include:
[0007] Step 1: Obtain a multi-omics pan-cancer dataset, wherein the multi-omics cancer dataset includes a dataset consisting of mRNA, DNA methylation, miRNA, and clinical prognosis data;
[0008] Step 2: Preprocess the multi-omics cancer dataset in step 1;
[0009] Step 3: Build a deep neural network model;
[0010] Step 4: Train the deep neural network model constructed in step 3 based on the multi-omics pan-cancer dataset preprocessed in step 2;
[0011] Step 5: Based on the deep neural network model trained in step 4, the cancer subtype classification and prognosis prediction are performed on the test data.
[0012] Furthermore, the steps of deep neural network model training in step 4 include:
[0013] Step 401: normalization of multi-omics data;
[0014] Step 402: Constructing a cancer subtype classification model by combining the channel attention mechanism with CAE to construct a CAE-based cancer subtype classification and prognosis prediction model (CA-CAE).
[0015] Step 403: Construct a prognosis prediction model, combine the features extracted from the model with clinical data, and evaluate the correlation between the features extracted from the model and the patient's survival time and prognosis.
[0016] Furthermore, the normalization method in step 401 is an alignment method of an algorithm combining standard deviation screening and MinMax normalization, and the steps include:
[0017] Step 40101: extract feature vectors from each omics data to form a feature-sample matrix;
[0018] Step 40102: Calculate the standard deviation of each feature in all samples and sort them from largest to smallest based on the standard deviation.
[0019] Step 40103: After filtering the first N features with the highest standard deviations, perform MinMax normalization on the selected features and map the feature values to the [0, 1] interval.
[0020] Furthermore, the MinMax normalization calculation steps in step 40103 are:
[0021] Step A: Determine the minimum value x for each feature min and the maximum value x max , used for scaling calculations:
[0022] x min =min(x1,x2,...,x n ),x max =max(x1,x2,...,x n ),
[0023] Among them, x1,x2,...,x n Indicates the value of this feature in all samples;
[0024] Step B: Normalize each eigenvalue x using the MinMax normalization formula:
[0025]
[0026] Where x′ is the normalized value;
[0027] Step C: Repeat the above calculation for all features and normalize the entire data matrix so that all its columns are scaled to the same numerical range.
[0028] Furthermore, the steps for constructing the cancer type classification model in step 402 are as follows:
[0029] Step a: Combine the channel attention mechanism with CAE to build an end-to-end model. From the output of the encoder, the model learns to generate the distribution of the latent variable z, which is expressed as follows:
[0030] H1=σ(Conv2D(X,W1)+b1),
[0031] H2=σ(CA(Conv2D(H1,W2)+b2)),
[0032] H3=σ(CA(Conv2D(H2,W3)+b3)),
[0033] Z=Flatten(H3),
[0034] Where X is the input data, H1, H2, H3 are the feature representations after convolution, W1, W2, W3 are convolution kernels, b1, b2, b3 are bias terms, CA is the channel attention mechanism used to enhance key features, σ() is the ReLU activation function to ensure nonlinear transformation, and Z is the final latent variable representation for subsequent decoding or cancer subtype classification;
[0035] CA is the channel attention mechanism, which is expressed as follows:
[0036] CA(H)=σ(W2·ReLU(W1·GAP(H))+W2·ReLU(W1·GMP(H))),
[0037] Among them, GAP(H) is average pooling, GMP(H) is maximum pooling, W1 and W2 are the weights of the fully connected layer for feature transformation, and the ReLU+Sigmoid combination is used for nonlinear mapping;
[0038] The loss function of CA-CAE uses mean square error loss as the main optimization objective to measure the difference between the input data X and the reconstructed data Xi:
[0039]
[0040] Among them, X i is the original data, is the reconstructed data, N is the total number of samples;
[0041] Step b: inputting multiple omics data and performing feature learning on different omics data respectively;
[0042] Step c: The decoder uses a symmetrical deconvolution + channel attention mechanism to recover features. The expression is as follows:
[0043]
[0044]
[0045] in, is the deconvolution weight, which is used to restore the data structure. CA() continues to optimize the feature expression in the decoding stage. Restructure the data.
[0046] Furthermore, CA is the channel attention mechanism, which is expressed as follows
[0047] CA(H)=σ(W2·ReLU(W1·GAP(H))+W2·ReLU(W1·GMP(H))),
[0048] Among them, GAP(H) is average pooling, GMP(H) is maximum pooling, W1 and W2 are the weights of the fully connected layer, which are used for feature transformation.
[0049] Furthermore, the loss function of CA-CAE uses the mean square error loss as the main optimization objective to measure the difference between the input data X and the reconstructed data Xi:
[0050]
[0051] Among them, X i is the original data, is the reconstructed data, and N is the total number of samples.
[0052] Furthermore, the steps for constructing the prognosis prediction model in step 403 are as follows:
[0053] Step 40301: Use the Cox proportional hazards model and Lasso regression to screen features and extract features related to survival. First, use Lasso regression:
[0054]
[0055] in, is the estimated value of the regression coefficient. The first term is the residual sum of squares of the ordinary least squares method, which represents the difference between the predicted value and the true value. The second term is the regularization term, and λ controls the strength of the regularization.
[0056] Then use Cox-PH screening:
[0057] h(t|x)=h0(t)exp(x Τ β),
[0058] Among them, h(t|x) is the conditional risk function, which represents the risk of an individual under characteristic x, h0(t) is the baseline risk function, and x T β is a linear combination of features x;
[0059] Step 40302: Use K-means clustering to identify different cancer subtypes and classify patients into different risk groups. The expression is as follows:
[0060]
[0061] Where J is the sum of the squares of the distances from all samples to their respective cluster centers, C i is the i-th cluster, x j is the jth sample, μ i is the center point of the i-th cluster, ||x j -μ i || 2 Represents sample x j Its corresponding cluster center μ i The square of the Euclidean distance between them;
[0062] Step 40303: First, perform survival analysis using the Cox proportional hazards regression model, calculate the statistical significance P value, and screen out important molecular features that affect survival. Then, use 5-fold cross-validation to calculate the C-index to measure the consistency of the Cox model in ranking patient survival. The expression is as follows:
[0063]
[0064] in, represents the risk score of the i-th individual, I(T i <T j ) is an indicator function, when T i Less than T j Returns 1 if yes, otherwise returns 0.
[0065] Furthermore, in step 1, for the multi-omics data, data cleaning is first performed to remove genes with all zero values and missing genes as a preliminary screening step. Subsequently, the screened gene expression matrix is standardized to reduce data bias and improve analysis accuracy. For clinical prognosis data, samples with missing survival data are removed.
[0066] Furthermore, the deep neural network model described in step 3 includes an encoder, a channel attention module, and a decoder. The encoder maps the input data to a low-dimensional potential representation through convolution operations; the channel attention module can weight key features and enhance the model's attention to important channels; the decoder reconstructs information from the potential representation into an output consistent with the shape of the original data through deconvolution operations.
[0067] The beneficial effects of the present invention are as follows: The present invention proposes a convolutional autoencoder (CA-CAE) model that integrates channel attention mechanisms to integrate multi-omics data and accurately predict the prognosis of cancer patients. By jointly learning the potential nonlinear features of mRNA-seq, miRNA-seq, and DNA methylation data, CA-CAE can stably identify cancer subtypes in high-dimensional, heterogeneous data environments and improve the accuracy of survival prediction. By comparing with other state-of-the-art cancer subtype classification and prognosis prediction models, the CA-CAE model has demonstrated superior performance across multiple evaluation indicators. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 It is a flowchart of the present invention; Figure 2 It is a structural model diagram of the CA-CAE model of the present invention; Figure 3 is a schematic diagram of the results of the present invention on a multi-omics pan-cancer dataset; Figure 4 Schematic diagram of the comparison results of CA-CAE and other methods on a multi-omics pan-cancer dataset; Figure 5 It is a schematic diagram of the CA-CAE of the present invention in constructing a classifier. DETAILED DESCRIPTION
[0069] Specific implementation method 1: Figures 1 to 5 As shown in FIG, a method for cancer profiling and prognosis prediction based on pan-cancer multi-omics data is used, and the specific steps include: S1: Obtain a pan-cancer multi-omics dataset, comprising mRNA-seq, miRNA-seq, DNA methylation data, and clinical prognosis data. Specifically, this implementation compares CA-CAE with three cancer subtype classification methods: ProgCAE, DeepProg, and PCA. P-value and C-index metrics are used to evaluate the effectiveness of these models in cancer subtype classification and prognosis prediction on these datasets. We obtained 15 real-world multi-omics datasets from existing technologies. They are adrenocortical carcinoma (ACC), bladder urothelial carcinoma (BLCA), cervical squamous cell carcinoma and adenocarcinoma (CESC), intrahepatic bile duct carcinoma (CHOL), colorectal adenocarcinoma (COAD), chromophobe renal cell carcinoma (KICH), acute myeloid leukemia (LAML), lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), mesothelioma (MESO), sarcoma (SARC), gastric adenocarcinoma (STAD), thyroid carcinoma (THCA), endometrial carcinoma (UCEC), and uveal melanoma (UVM).
[0070] S2: Preprocess the multi-omics cancer dataset described in S1. For multi-omics data, first perform data cleaning to remove genes with all zero values or missing values as a preliminary screening step. Then, standardize the screened gene expression matrix to reduce data bias and improve analysis accuracy. For clinical prognosis data, remove samples with missing survival data.
[0071] The pan-cancer multi-omics dataset obtained in this embodiment is as follows:
[0072] The ACC dataset data were processed according to the procedures described in the original study, and the final dataset contained 79 samples.
[0073] The BLCA dataset was processed according to the procedures described in the original study, and the final dataset contained 419 samples.
[0074] The CESC dataset was processed according to the procedures described in the original study, and the final dataset contained 296 samples.
[0075] The CHOL dataset was processed according to the procedures described in the original study, and the final dataset contained 45 samples.
[0076] The COAD dataset was processed according to the procedures described in the original study, and the final dataset contained 298 samples.
[0077] The KICH dataset was processed according to the procedures described in the original study, and the final dataset contained 64 samples.
[0078] The LAML dataset was processed according to the procedures described in the original study, and the final dataset contained 85 samples.
[0079] The LUAD dataset was processed according to the procedures described in the original study, and the final dataset contained 453 samples.
[0080] The LUSC dataset was processed according to the procedures described in the original study, and the final dataset contained 360 samples.
[0081] The MESO dataset was processed according to the procedures described in the original study, and the final dataset contained 84 samples.
[0082] The SARC dataset was processed according to the procedures described in the original study, and the final dataset contained 258 samples.
[0083] The STAD dataset was processed according to the procedures described in the original study, and the final dataset contained 324 samples.
[0084] The THCA dataset was processed following the procedures described in the original study, and the final dataset contained 557 samples.
[0085] The UVM dataset was processed according to the procedures described in the original study, and the final dataset contained 80 samples.
[0086] The CA-CAE framework uses datasets consisting of mRNA-seq, miRNA-seq, DNA methylation data, and clinical prognosis data as its primary input. For the gene expression matrix, we first perform data cleaning to remove genes with all zero values or missing values as an initial screening step. Subsequently, we normalize the filtered gene expression matrix to reduce data bias and improve analysis accuracy.
[0087] S3: Build a deep neural network model, which includes an encoder, a channel attention module, and a decoder. The encoder maps input data to a low-dimensional latent representation through convolution operations; the channel attention module can weight key features and enhance the model's attention to important channels; the decoder reconstructs information from the latent representation into an output consistent with the original data shape through deconvolution operations;
[0088] Specifically, such as Figure 2 As shown in the figure, the CA-CAE model framework is used to classify cancer subtypes and predict prognosis for pan-cancer multi-omics data. The encoder is a 3-layer graph convolutional network, and the decoder is a symmetric deconvolutional decoder. The first stage of CA-CAE consists of dimensionality reduction and feature transformation, using Pearson correlation normalization and a convolutional autoencoder with an attention mechanism for dimensionality reduction. Each of the three omics data is modeled by a convolutional autoencoder with an attention mechanism to achieve flexibility and scalability for heterogeneous data types. After dimensionality reduction, the transformed features are first subjected to LASSO regression analysis to filter out features with Lasso coefficients not equal to 0, and then the features are fitted with univariate Cox-PH to further select features related to survival. Finally, the survival features of each omics are combined for subsequent survival analysis and cancer subtype selection.
[0089] In S4, the deep neural network model constructed in S3 is trained based on the multi-omics pan-cancer dataset preprocessed by S2. The steps of training the deep neural network model include:
[0090] S401: Normalization of Multi-omics Data: To maintain data consistency and equality, the most important step is to normalize multi-omics data to facilitate subsequent analysis and processing. Here, we use an algorithm that combines standard deviation screening with MinMax normalization to normalize multi-omics data. Compared with other algorithms, this algorithm can better handle noise and variability in the data, thereby increasing data usability and being more suitable for complex multi-omics data.
[0091] First, use the Pearson correlation coefficient to process the data. The Pearson correlation coefficient is used to evaluate the linear correlation between two variables. Based on the correlation between columns, find the columns with stronger correlation. The calculation formula of the Pearson correlation coefficient is as follows:
[0092]
[0093] Where r is the Pearson correlation coefficient between two variables x and y; x i and y i is the i-th value in the dataset; and are the means of x and y respectively.
[0094] We calculated the geometric mean of each row of the correlation matrix and reordered the features according to their numerical values.
[0095]
[0096] Where ρ is the geometric mean of the cumulative correlation coefficients. ρ1, ρ2, ...ρ n Represents the absolute value of all correlation coefficients corresponding to a row in the correlation matrix.
[0097] Then, the Min-Max normalization is used to linearly transform the data value to a fixed range. By scaling, the minimum value of the data becomes 0 and the maximum value becomes 1, while other values are linearly transformed according to this range. The Min-Max normalization formula is as follows:
[0098]
[0099] Where x′ is the normalized value, x is the original data, and x min is the minimum value, x max is the maximum value.
[0100] Through the above steps, multi-omics data can be successfully normalized.
[0101] S402: Construction of a cancer subtype classification model, combining the channel attention mechanism with CAE to construct a CAE-based cancer subtype classification and prognosis prediction model CA-CAE.
[0102] Here, we constructed a CAE-based cancer subtype classification and prognosis prediction model (CA-CAE). By combining the channel attention mechanism with a convolutional autoencoder (CAE), we constructed an end-to-end model that simultaneously learns important features and latent representations from multi-omics data. This combination enhances the model's feature extraction capabilities for high-dimensional, heterogeneous cancer data and effectively identifies key biomarkers.
[0103] Cancer subtype data are usually characterized by high dimensionality, high noise, and missing data. CAE uses an encoder to reduce dimensionality, extract the potential expression features of the data, and reconstruct features through a decoder, thereby enhancing data integrity and improving the robustness and generalization ability of the model.
[0104] a. Combine the channel attention mechanism with CAE to build an end-to-end model. From the output of the encoder, the model learns to generate the distribution of the latent variable z, which is expressed as follows:
[0105] H1=σ(Conv2D(X,W1)+b1)
[0106] H2=σ(CA(Conv2D(H1, W2)+b2))
[0107] H3=σ(CA(Conv2D(H2, W3)+b3))
[0108] Z=Flatten(H3)
[0109] Where X is the input data (multi-omics data matrix), H1, H2, and H3 are the feature representations after convolution. W1, W2, and W3 are the convolution kernels, and b1, b2, and b3 are the bias terms. CA is the channel-wise attention mechanism, which is used to enhance key features. σ() is the ReLU activation function, which ensures nonlinear transformations. Z is the final latent variable representation, which is used for subsequent decoding or cancer subtype classification.
[0110] Further CA is the channel attention mechanism, which is expressed as follows
[0111] CA(H)=σ(W2·ReLU(W1·GAP(H))+W2·ReLU(W1·GMP(H)))
[0112] Where GAP(H) is average pooling, GMP(H) is maximum pooling, W1 and W2 are the weights of the fully connected layer for feature transformation, and the ReLU+Sigmoid combination is used for nonlinear mapping.
[0113] The loss function of CA-CAE further adopts the mean square error (MSE) loss as the main optimization objective to measure the difference between the input data X and the reconstructed data Xi:
[0114]
[0115] where X i is the original data, is the reconstructed data, and N is the total number of samples.
[0116] b. Input multiple omics data and perform feature learning on different omics data separately;
[0117] c. The decoder uses symmetric deconvolution (Conv2DTranspose) + channel attention mechanism (CA) for feature recovery, which is expressed as follows:
[0118]
[0119] in is the deconvolution weight, which is used to restore the data structure. CA() continues to optimize the feature expression in the decoding stage. The reconstructed data should be as similar as possible to the input X.
[0120] S403: Construct a prognostic prediction model, combine the features extracted from the model with clinical data, and evaluate the association between the features extracted from the model and the patient's survival time and prognosis.
[0121] The Cox proportional hazard model (Cox-PH) and Lasso regression were used to screen features and extract features related to survival. First, Lasso regression was used:
[0122]
[0123] in, is the estimated value of the regression coefficient. The first term is the sum of squared residuals of the ordinary least squares method, which represents the difference between the predicted value and the true value. The second term is the regularization term, and λ controls the strength of the regularization.
[0124] Next we use Cox-PH screening:
[0125] h(t|x)=h0(t)exp(x T β)
[0126] Where h(t|x) is the conditional risk function, which represents the risk of an individual under characteristic x, and h0(t) is the baseline risk function. T β is a linear combination of the features x.
[0127] We use K-means clustering to identify different cancer subtypes and divide patients into different risk groups. The expression is as follows:
[0128]
[0129] Where J is the sum of the squares of the distances from all samples to their respective cluster centers, C i is the i-th cluster, x j is the jth sample, μ i is the center point of the i-th cluster, ||x j -μ i || 2 Represents sample x j Its corresponding cluster center μ i The square of the Euclidean distance between them;
[0130] We used the Cox proportional hazards regression model for survival analysis, calculated the statistical significance P value, screened out important molecular features that affect survival, and then used 5-fold cross-validation (KFold) to calculate the C-index to measure the consistency of the Cox model in ranking patient survival. The expression is as follows:
[0131]
[0132] in represents the risk score of the i-th individual, I(T i <T j ) is an indicator function, when T i Less than T j Returns 1 if yes, otherwise returns 0.
[0133] S5: Based on the deep neural network model trained in S4, we perform cancer subtype classification and prognosis prediction on the test data. After obtaining pan-cancer multi-omics data and successfully building the model, we use CA-CAE to perform cancer subtype classification and prognosis prediction.
[0134] This implementation compares CA-CAE with ProgCAE, DeepProg, and PCA for cancer subtype classification. P-value and C-index metrics are used to evaluate the effectiveness of these models in cancer subtype classification and prognosis prediction on a dataset consisting of mRNA-seq, miRNA-seq, DNA methylation data, and clinical prognosis. The datasets for adrenocortical carcinoma (ACC), bladder urothelial carcinoma (BLCA), cervical squamous cell carcinoma and adenocarcinoma (CESC), intrahepatic bile duct carcinoma (CHOL), colorectal adenocarcinoma (COAD), chromophobe renal cell carcinoma (KICH), acute myeloid leukemia (LAML), lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), mesothelioma (MESO), sarcoma (SARC), gastric adenocarcinoma (STAD), thyroid carcinoma (THCA), endometrial carcinoma (UCEC), and uveal melanoma (UVM) were downloaded from the data portal: https: / / xenabrowser.net / datapages / .
[0135] Example 1
[0136] like Figure 3 As shown in Figure 2, CA-CAE was evaluated for its ability to classify cancer subtypes and predict prognosis using multi-omics analysis of 15 cancer types. It can be clearly observed that CA-CAE has powerful capabilities in cancer subtype classification and prognosis prediction.
[0137] Example 2
[0138] We compared CA-CAE with other methods and used P-value to evaluate performance. In most cancers, CA-CAE outperformed other methods. The results are as follows:
[0139] Table 1 Comparison of P-values of CA-CAE and other three methods on 15 real cancer datasets
[0140]
[0141]
[0142] Example 3
[0143] We compared CA-CAE with other methods and used C-index for performance evaluation. CA-CAE outperformed other methods in most cancers. The results are as follows:
[0144] Table 2 Comparison of C-index of CA-CAE with other three methods on 15 real cancer datasets
[0145]
[0146] The results of Examples 2 and 3 demonstrate that CA-CAE can more efficiently identify cancer subtypes in cancer datasets and exhibits high accuracy. These results suggest that CA-CAE may be more accurate than existing methods in identifying cancer subtypes and has the potential to provide more effective diagnoses for cancer prognosis.
[0147] Example 4
[0148] like Figure 4 To demonstrate the effectiveness of CA-CAE in validating cluster analysis results in supervised learning, we applied CA-CAE to construct classifiers based on multi-omics data. We used the labels obtained by clustering to construct classifiers for ACC, BLCA, and LUAD data. First, the raw data were preprocessed according to the previously described method. Next, the top 50 features with the largest standard deviations were selected from each omics dataset and concatenated to form a new feature matrix. This new matrix was partitioned into 70% for the training set and 30% for the test set. A support vector machine (SVM) model was then trained on the training set, hyperparameters were optimized using grid search, and predictions were performed on the test set. Based on the prediction results, Kaplan-Meier curves were plotted for each test dataset, and P values were calculated. All P values for the test sets were less than 0.05. These results demonstrate that the features selected by our method have robust survival classification performance. Our supervised predictions were significant, validating the effectiveness of the clustering results in supervised learning.
[0149] Example 5
[0150] like Figure 5 As shown in the figure, we evaluated the effectiveness of CA-CAE in predicting cancer patient survival using single-omic and multi-omic signatures. Specifically, we extracted mRNA omics signatures from LUAD for survival prediction, compared the differences in survival prediction between single-omic and multi-omic genes, and calculated the significance of genes through multivariate Cox regression analysis.
[0151] The forest plot shows the significance of single-omics (transcriptomics) signature genes obtained through multivariate Cox regression analysis, demonstrating the effect of different genes in predicting survival. Genes in the figure, such as CRLF1 and IGHG4, showed significant P values in LUAD patients and had strong predictive ability. The two Kaplan-Meier survival curves on the right show the predictive effects of using multi-omics gene signatures and single-omics gene signatures in different survival groups. Judging from the distribution of the survival curves, the multi-omics model has a better effect in distinguishing different survival subgroups, with a smaller P value and a relatively high C-index, indicating that the model has better predictive accuracy. Although the single-omics model's discriminatory effect is also statistically significant, its P value is relatively large and its C-index is also low, indicating that the predictive effect is inferior to that of the multi-omics model.
[0152] This result suggests that integrated multi-omics analysis can better capture complementary information between different omics, thereby improving the accuracy and robustness of survival predictions. The fusion of multi-omics features can help reveal the biological heterogeneity of cancer patients and provide a more reliable predictive basis for precision medicine. While single-omics analysis can have significant effects on specific genes, the lack of supplementary data from other omics limits the predictive effectiveness of the model.
[0153] The results of Examples 1 to 5 show that CA-CAE demonstrates the accuracy of identifying cancer subtypes in pan-cancer multi-omics data.
[0154] Example 6
[0155] To further understand the genetic differences caused by cancer subtypes, we extracted and studied the top 20 genes in the three omics of LUAD cancer. The results are as follows:
[0156] Table 5 Summary of the top 20 genes of CA-CAE in LUAD subtypes
[0157]
[0158]
[0159] In mRNA omics, SFTPA1 and SFTPA2 were shown to increase the risk of cancer in mutation carriers, GPX2 expression was closely associated with the prognosis of lung adenocarcinoma patients
[27] , and CEACAM5 can be used as a reliable marker for pre-screening of lung adenocarcinoma tumor cells based on RT-PCR
[28] . MSLN expression is a molecular marker of tumor aggressiveness and a potential therapeutic target
[29] . S100P was shown to be useful for cancer prognosis analysis by combining single-cell features of lung adenocarcinoma heterogeneity and microenvironment. AKR1C1-STAT3 levels were highly correlated with poor prognosis in NSCLC patients. Bidirectional MR and colocalization analysis revealed a negative correlation between plasma proteins SFTPB and KDELC2 in lung adenocarcinoma (LUAD). PIGR is an independent prognostic biomarker associated with acquired EGFR-TKI resistance and tumor immune cell infiltration in lung adenocarcinoma. TFF3 is a potent promoter of lung adenocarcinoma progression. Targeting TFF3 with novel small molecule inhibitors alone or in combination with traditional MEK1 / 2 inhibitors is a potential strategy to improve the prognosis of lung adenocarcinoma. MUC5B-AS1 promotes cell migration and invasion by forming RNA-RNA duplexes with MUC5B, thereby increasing MUC5B expression in lung adenocarcinoma. High MUC5B expression is significantly associated with poor outcome in lung adenocarcinoma. CPS1 is a promising therapeutic target in combination with other chemotherapeutic agents and a prognostic biomarker, enabling personalized treatment of lung adenocarcinoma. CTSE is overexpressed and hypomethylated in LUAD, which corresponds to normal lung tissue. Methylation of the TFPI-2 gene is an independent factor for poor prognosis in patients with NSCLC. CXCL14 is a potential diagnostic biomarker for lung adenocarcinoma with a microscopic pattern. Two common risk genes (SFTPD and HLA-DRA) at different stages can serve as prognostic markers in LUAD. The number of TILs within tumor islands is an independent prognostic marker for patients with brain metastases from lung adenocarcinoma. Comparative transcriptomics revealed that the expression of three surfactant metabolism-related genes (SFTPA1, SFTPB, and NAPSA) is closely correlated with the number of TILs. Upregulation of genes AKR1C1 / C2, TM4SF1, and NR0B1 in lung adenocarcinoma A549 cells SP may be targets for poor prognosis in anticancer therapy.
[0160] In summary, we conducted an in-depth study of the top 20 genes across the three omics categories for LUAD cancer and discovered several potential biomarkers and therapeutic targets with important clinical implications. These genes not only reveal the molecular mechanisms of LUAD but also provide new insights into molecular biomarkers for cancer prognosis.
[0161] Example 7
[0162] To investigate the association between clinical cancer traits and subtype clusters, and to assess the potential value of model-assigned cluster labels in reflecting cancer course characteristics and clinical stage, we performed chi-square tests on four sets of clinical data (Stage, T, N, and M) for BLCA, COAD, and LUAD cancers and analyzed their associations with the model-assigned cluster labels. The results are as follows:
[0163]
[0164] The P value between Stage and cluster in BLCA was 0.0005, significantly less than 0.05, indicating a significant association between disease stage and cancer subtype clustering. Different clusters may represent different stages of disease progression, providing a basis for personalized treatment. The P value for T classification in BLCA was 0.0004, also significantly less than 0.05, indicating a significant association between the extent of primary tumor extension and clustering, suggesting that different clusters may have specific characteristics regarding tumor spread. In COAD cancers, except for T, the other three clinical P values were significantly less than 0.05, indicating a significant association between disease stage, lymph node metastasis, and distant metastasis, respectively. For LUAD, the P value for Stage was 0.0263, and the P value for N classification was 0.0128, both significantly less than 0.05, indicating a significant association between disease stage and lymph node metastasis, respectively, and clustering, providing potential directions for personalized treatment and clinical intervention.
[0165] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement of the above embodiments made according to the technical essence of the present invention, within the spirit and principles of the present invention, without departing from the content of the technical solution of the present invention, shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for cancer profiling and prognosis prediction based on pan-cancer multi-omics data, characterized in that: The specific steps include: Step 1: Obtain a multi-omics pan-cancer dataset, wherein the multi-omics cancer dataset includes a dataset consisting of mRNA, DNA methylation, miRNA, and clinical prognosis data; Step 2: Preprocess the multi-omics cancer dataset in step 1; Step 3: Build a deep neural network model; Step 4: Train the deep neural network model constructed in step 3 based on the multi-omics pan-cancer dataset preprocessed in step 2; Step 5: Based on the deep neural network model trained in step 4, the cancer subtype classification and prognosis prediction are performed on the test data.
2. The method for cancer profiling and prognosis prediction based on pan-cancer multi-omics data according to claim 1, characterized in that: The steps of deep neural network model training in step 4 include: Step 401: normalization of multi-omics data; Step 402: Constructing a cancer subtype classification model by combining the channel attention mechanism with CAE to construct a CAE-based cancer subtype classification and prognosis prediction model (CA-CAE). Step 403: Construct a prognosis prediction model, combine the features extracted from the model with clinical data, and evaluate the correlation between the features extracted from the model and the patient's survival time and prognosis.
3. The method for cancer profiling and prognosis prediction based on pan-cancer multi-omics data according to claim 2, characterized in that: In step 401, the normalization method is an alignment method that combines standard deviation screening with MinMax normalization, and the steps include: Step 40101: extract feature vectors from each omics data to form a feature-sample matrix; Step 40102: Calculate the standard deviation of each feature in all samples and sort them from largest to smallest based on the standard deviation. Step 40103: After filtering the first N features with the highest standard deviations, perform MinMax normalization on the selected features and map the feature values to the [0, 1] interval.
4. The method for cancer profiling and prognosis prediction based on pan-cancer multi-omics data according to claim 3, characterized in that: The MinMax normalization calculation steps in step 40103 are: Step A: Determine the minimum value x for each feature min and the maximum value x max , used for scaling calculations: x min =min(x1,x2,...,x n ),x max =max(x1,x2,...,x n ), Among them, x1,x2,...,x n Indicates the value of this feature in all samples; Step B: Normalize each eigenvalue x using the MinMax normalization formula: Where x′ is the normalized value; Step C: Repeat the above calculation for all features and normalize the entire data matrix so that all its columns are scaled to the same numerical range.
5. The method for cancer profiling and prognosis prediction based on pan-cancer multi-omics data according to claim 2, characterized in that: The steps for constructing the cancer type classification model in step 402 are as follows: Step a: Combine the channel attention mechanism with CAE to build an end-to-end model. From the output of the encoder, the model learns to generate the distribution of the latent variable z, which is expressed as follows: H1=σ(Conv2D(X,W1)+b1), H2=σ(CA(Conv2D(H1,W2)+b2)), H3=σ(CA(Conv2D(H2,W3)+b3)), Z=Flatten(H3), Where X is the input data, H1, H2, H3 are the feature representations after convolution, W1, W2, W3 are convolution kernels, b1, b2, b3 are bias terms, CA is the channel attention mechanism used to enhance key features, σ() is the ReLU activation function to ensure nonlinear transformation, and Z is the final latent variable representation for subsequent decoding or cancer subtype classification; CA is the channel attention mechanism, which is expressed as follows: CA(H)=σ(W2·ReLU(W1·GAP(H))+W2·ReLU(W1·GMP(H))), Among them, GAP(H) is average pooling, GMP(H) is maximum pooling, W1 and W2 are the weights of the fully connected layer for feature transformation, and the ReLU+Sigmoid combination is used for nonlinear mapping; The loss function of CA-CAE uses mean square error loss as the main optimization objective to measure the difference between the input data X and the reconstructed data Xi: Among them, X i is the original data, is the reconstructed data, N is the total number of samples; Step b: inputting multiple omics data and performing feature learning on different omics data respectively; Step c: The decoder uses a symmetrical deconvolution + channel attention mechanism to recover features. The expression is as follows: Among them, W′1, W′2, and W′3 are deconvolution weights used to restore the data structure. CA() continues to optimize the feature expression in the decoding stage. Restructure the data.
6. The method for cancer profiling and prognosis prediction based on pan-cancer multi-omics data according to claim 5, characterized in that: CA is the channel attention mechanism, the expression is as follows CA(H)=σ(W2·ReLU(W1·GAP(H))+W2·ReLU(W1·GMP(H))), Among them, GAP(H) is average pooling, GMP(H) is maximum pooling, W1 and W2 are the weights of the fully connected layer, which are used for feature transformation.
7. The method for cancer profiling and prognosis prediction based on pan-cancer multi-omics data according to claim 5, characterized in that: The loss function of CA-CAE uses mean square error loss as the main optimization objective to measure the difference between the input data X and the reconstructed data Xi: Among them, X i is the original data, is the reconstructed data, and N is the total number of samples.
8. The method for cancer profiling and prognosis prediction based on pan-cancer multi-omics data according to claim 2, wherein: The steps for constructing the prognosis prediction model in step 403 are as follows: Step 40301: Use the Cox proportional hazards model and Lasso regression to screen features and extract features related to survival. First, use Lasso regression: in, is the estimated value of the regression coefficient. The first term is the residual sum of squares of the ordinary least squares method, which represents the difference between the predicted value and the true value. The second term is the regularization term, and λ controls the strength of the regularization. Then use Cox-PH screening: h(t|x)=h0(t)exp(x Τ β), Among them, h(t|x) is the conditional risk function, which represents the risk of an individual under characteristic x, h0(t) is the baseline risk function, and x T β is a linear combination of features x; Step 40302: Use K-means clustering to identify different cancer subtypes and classify patients into different risk groups. The expression is as follows: Where J is the sum of the squares of the distances from all samples to their respective cluster centers, C i is the i-th cluster, x j is the jth sample, μ i is the center point of the i-th cluster, ||x j -μ i || 2 Represents sample x j Its corresponding cluster center μ i The square of the Euclidean distance between them; Step 40303: First, perform survival analysis using the Cox proportional hazards regression model, calculate the statistical significance P value, and screen out important molecular features that affect survival. Then, use 5-fold cross-validation to calculate the C-index to measure the consistency of the Cox model in ranking patient survival. The expression is as follows: in, represents the risk score of the i-th individual, I(T i <T j ) is an indicator function, when T i Less than T j Returns 1 if yes, otherwise returns 0.
9. The method for cancer profiling and prognosis prediction based on pan-cancer multi-omics data according to claim 1, wherein: In step 1, for multi-omics data, data cleaning is first performed to remove genes with all zero values or missing genes as a preliminary screening step. Subsequently, the screened gene expression matrix is standardized to reduce data bias and improve analysis accuracy. For clinical prognosis data, samples with missing survival data are removed.
10. The method for cancer profiling and prognosis prediction based on pan-cancer multi-omics data according to claim 1, characterized in that: The deep neural network model described in step 3 includes an encoder, a channel attention module, and a decoder. The encoder maps the input data to a low-dimensional potential representation through convolution operations; the channel attention module can weight key features and enhance the model's attention to important channels; the decoder reconstructs information from the potential representation into an output consistent with the shape of the original data through deconvolution operations.
Citation Information
Patent Citations
Hash sample balance cancer labeling method for histopathologic image
CN112906804A
Cancer subtype classification method and system based on multi-omics data
CN118366551A
Intelligent lung cancer metastasis prediction system based on GCAVE-GAN and multi-mode fusion
CN118628462A
Construction method of risk prediction model for prognosis of gastric cancer
US20240318254A1
Cited By
Incomplete multi-omics cancer subtype data clustering method
CN121034428A