Lung cancer prognosis prediction method based on multi-omics data fusion

Through the fusion of multiomic data and multi-level feature extraction methods, the problem that single anomic data is difficult to fully reveal the pathogenic mechanism of lung cancer is solved, achieving higher prognosis prediction accuracy and personalized treatment support.

CN120015323AActive Publication Date: 2025-05-16CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510168161.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-16
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The prior art lung cancer prognosis prediction methods based on single antonyms data are difficult to fully reveal the complex pathogenic mechanism of lung cancer, and the prediction accuracy is limited.

Method used

The lung cancer prognosis prediction method based on multiomics data fusion was adopted, and the multiomics data was mapped to low-dimensional space through an autoencoder, and the VSOEnetbag algorithm combined with elastic network and bagging ideas was used for feature selection. The samples were clustered risk subtypes using K-means clustering, and a multivariable Cox prognosis model was constructed.

Benefits of technology

It improves the accuracy and reliability of lung cancer prognosis prediction, reveals the potential biological heterogeneity of lung cancer through multi-level feature extraction and clustering analysis, and supports the formulation of personalized treatment plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015323A_ABST
    Figure CN120015323A_ABST
Patent Text Reader

Abstract

The invention relates to a lung cancer prognosis prediction method based on multi-omics data fusion, which comprises the following steps: acquiring multi-omics data and survival data of lung cancer patients, and preprocessing the multi-omics data of each lung cancer patient; splicing the multi-omics data of each lung cancer patient to obtain a multi-omics integrated feature matrix of each training sample; mapping the multi-omics integrated feature matrix of the training sample to a low-dimensional space through an auto-encoder to obtain a multi-omics low-dimensional feature matrix of the training sample; performing feature selection on the multi-omics low-dimensional feature matrix of the training sample by using a VSOEnetbag algorithm combining an elastic network and a bagging idea according to survival data of the lung cancer patient to obtain a significant expression feature matrix of the training sample; performing risk subtype clustering on the training sample by using K-means clustering according to the significant feature expression matrix of the training sample to obtain a subtype clustering result; taking a subtype clustering result of the training sample as a label, and taking RNA-Seq expression profile data of the lung cancer patient as an independent variable to construct a multi-variable-based cox prognosis model; rNA-Seq expression profile data of a to-be-detected lung cancer patient is input into the multi-variable-based cox prognosis model, and a prognosis result of the lung cancer patient is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of bioinformatics and computer science, and in particular relates to a lung cancer prognosis prediction method based on multi-omics data fusion. Background Art

[0002] Lung cancer is one of the most common malignant tumors in the world and a major threat to human health. Its prognosis and survival rate are low, posing a serious threat to the patient's quality of life and life. Accurately predicting the prognosis risk of lung cancer patients is crucial to guiding clinical diagnosis and treatment decisions and improving survival rates. Traditional lung cancer prognosis prediction methods mainly rely on clinical indicators such as age, tumor grade, and surgical resection range, but the predictive ability of these indicators is limited. With the development of high-throughput sequencing technology, genomic data has been widely used in lung cancer prognosis research, and some prediction models based on molecular features such as gene expression and DNA methylation have been proposed one after another. These models use machine learning to screen out key features related to prognosis from high-dimensional molecular features, significantly improving the prediction accuracy.

[0003] Early lung cancer prediction models mainly used single-omics data such as genomics, proteomics and metabolomics data to analyze tumors, and used some biochemical indicators for prediction. For example, HER2 gene mutations and EGFR gene mutations in genomics have been widely used in clinical testing, benefiting a large number of patients. Xie et al. developed a lung cancer auxiliary diagnosis and prognosis evaluation system based on exosome protein markers (patent document publication number: CN202411139237.8), which includes information acquisition, calculation and diagnosis modules to improve the sensitivity and specificity of lung cancer diagnosis. SU et al. developed a non-small cell lung cancer recurrence prognosis risk scoring model (patent document publication number: CN202410747151.7), which constructed a risk scoring model based on 6 genes such as HLA-DOB, SPP1, TWIST2 and 2 clinical characteristics, which can accurately distinguish high-risk and low-risk recurrence patients. Luo et al. proposed a method for detecting platelet markers in non-small cell lung cancer (patent document publication number: CN202410299001.4), and screened out ten characteristic genes closely related to the prognosis of non-small cell lung cancer based on differential gene analysis, providing new clues for prognosis evaluation. Wang et al. created a gene signature for radiotherapy prognosis of non-small cell lung cancer (patent document publication number: CN202410846192.1), and used LASSO logistic regression to construct a gene signature based on five genes such as MDM2 and THBS1 to predict radiotherapy prognosis and support the formulation of personalized radiotherapy plans. In addition, Wang et al. also developed a lung cancer malignant risk stratification model (patent document publication number: CN202410793134.7), and the Densenet-Self-Attention-DeepSurv neural network was applied, combining medical imaging and clinical data to accurately predict the malignant risk and prognosis of lung cancer. These systems and models provide important technical support for the diagnosis, prognosis evaluation and personalized treatment of lung cancer.

[0004] However, research based on a single omics data cannot fully reveal the entire mechanism of lung cancer, a complex disease. With the advancement of biological science and the improvement of measurement technology, researchers can obtain and analyze molecular data at many different levels, including different "omics" data such as genome, transcriptome, proteome and metabolome. These data reflect biological information at different levels such as genes, RNA, proteins and metabolites, respectively, providing the possibility of a comprehensive analysis of the pathogenic mechanism of lung cancer. On the other hand, the occurrence and development of lung cancer is a complex process driven by the interaction of changes at multiple molecular levels. There are highly correlated regulatory relationships between molecules at different omics levels. For example, gene mutations may cause changes in RNA expression, which in turn affects the synthesis of specific proteins and ultimately leads to disorders in cellular metabolism. Therefore, a single omics data can only provide a one-sided analysis of the pathogenic mechanism of lung cancer, and it is difficult to fully reveal the full picture and complexity of the disease. Summary of the invention

[0005] In order to solve the problems existing in the background technology and improve the accuracy of lung cancer prognosis prediction, one aspect of the present invention provides a lung cancer prognosis prediction method based on multi-omics data fusion, comprising:

[0006] S1: Obtain multi-omics data and survival data of lung cancer patients, and pre-process the multi-omics data of each lung cancer patient;

[0007] S2: The multi-omics data of each lung cancer patient are spliced ​​to obtain the multi-omics integrated feature matrix of each training sample;

[0008] S3: The multi-omics integrated feature matrix of the training samples is mapped to a low-dimensional space through an autoencoder to obtain a multi-omics low-dimensional feature matrix of the training samples;

[0009] S4: Based on the survival data of lung cancer patients, the VSOEnetbag algorithm combining elastic network and bagging idea is used to select the multi-omics low-dimensional feature matrix of training samples to obtain the significant expression feature matrix of training samples;

[0010] S5: Use K-means clustering based on the significant feature expression matrix of the training samples to cluster the training samples into risk subtypes and obtain the subtype clustering results;

[0011] S6: The subtype clustering results of the training samples were used as labels, and the RNA-Seq expression profile data of lung cancer patients were used as independent variables to construct a multivariate-based Cox prognostic model;

[0012] S7: The RNA-Seq expression profile data of the lung cancer patient to be tested is input into the multivariate-based Cox prognostic model to obtain the prognostic results of the lung cancer patient.

[0013] Another aspect of the present invention provides a lung cancer prognosis prediction device based on multi-omics data fusion, comprising a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory, and is used to execute the computer program stored in the memory, so that the lung cancer prognosis prediction device based on multi-omics data fusion performs the lung cancer prognosis prediction method based on multi-omics data fusion.

[0014] Another aspect of the present invention provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the method for predicting lung cancer prognosis based on multi-omics data fusion.

[0015] The present invention has at least the following beneficial effects

[0016] The present invention develops an integrated framework of multi-omics data for performing cancer subtype clustering, and realizes step-by-step dimensionality reduction and efficient integration of high-dimensional multi-omics data through multi-level feature extraction. Compared with single-omics analysis, multi-omics integration can improve the accuracy and reliability of framework prediction. At the same time, the present invention proposes an innovative multi-level extraction strategy, which effectively addresses the computational difficulties of large-scale multi-omics data and reduces the computational complexity of the model. This strategy not only improves the efficiency of data processing by hierarchical processing and optimized data representation, but also helps to reduce storage requirements and processing time, making large-scale multi-omics data analysis more feasible and efficient. The present invention proposes a feature selection method VSOEnetbagging based on elastic network (Elastic Net) regression and bagging (Bagging), which accurately screens out features significantly related to lung cancer prognosis on the basis of multiple iterations and cross-validation. By combining LASSO regression with multiple subspace iterations, overfitting can be effectively avoided, and the most representative and predictive feature variables can be screened out by statistical frequency. Compared with the traditional single feature selection method, this method has obvious advantages in improving the stability and accuracy of feature selection. Based on the feature matrix, the present invention uses the K-means clustering method to subtype lung cancer samples. This clustering process can divide samples into multiple subtypes, thereby revealing the potential biological heterogeneity of lung adenocarcinoma. This classification not only provides a new perspective for the biological understanding of cancer subtypes, but also provides an important basis for the formulation of individualized treatment plans. Compared with traditional simple classification methods, this method can more accurately capture the complex characteristics of cancer through cluster analysis of multi-omics data, and improve the ability to identify and stratify cancer subtypes. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the method flow of the present invention;

[0018] Figure 2To evaluate the optimal number of clusters based on Silhouette Score;

[0019] Figure 3 Visualize the K-means clustering results;

[0020] Figure 4 Figure 3 is the Kaplan-Meier survival curve (high risk vs. low risk) of the multi-omics lung cancer prognosis model based on autoencoder and VSOenetbagging; predicted_survival=0 indicates high risk, and predicted_survival=1 indicates low risk. DETAILED DESCRIPTION

[0021] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0022] See also Figure 1 One aspect of the present invention provides a method for predicting lung cancer prognosis based on multi-omics data fusion, comprising:

[0023] S1: Obtain multi-omics data and survival data of lung cancer patients, and pre-process the multi-omics data of each lung cancer patient;

[0024] In this embodiment, multiple omics data of cancer and corresponding survival data of patients are obtained from reliable data sources (e.g., the public database Cancer Genome Atlas TCGA), missing values ​​are processed for each omics data, and then key variables closely related to the occurrence and development of cancer are screened out using statistical methods and computational tools. Next, standardization is performed to integrate these data sets into a comprehensive multi-omics data set as the input of the cross-omics integrated autoencoder; in a specific embodiment, a lung adenocarcinoma (LUAD) data set from TCGA is used. Considering the complexity and heterogeneity of cancer itself, four types of omics data are selected, including mRNA expression data of lung adenocarcinoma patients, that is, gene expression data (Geneexpression, GE), miRNA expression data (miRNAexpression, miRNA), DNA methylation data (DNAmethylation, ME) and protein expression data (Protein Expression, PE). The generation platforms of these four multi-omics data are different, among which GE and miRNA data are generated by IlluminaHiSeq, ME data are generated by IlluminaInfiniumHumanMethylation450, and PE data are generated by Reverse Phase ProteinArray (RPPA). These data are all TCGA level 3 data.

[0025] Preferably, the multi-omics data of the lung cancer patients include: RNA-Seq expression profile data, DNA methylation data, miRNA expression data and protein expression data of the lung cancer patients; the survival data of the lung cancer patients include: survival time and survival outcome of the lung cancer patients.

[0026] S2: The multi-omics data of each lung cancer patient are spliced ​​to obtain the multi-omics integrated feature matrix of each training sample;

[0027] In this embodiment, based on the statistical data set in step S1, the omics information of lung cancer patients is obtained. These matrices are concatenated to obtain a multi-omics integrated feature matrix as the input of the feature fusion module (autoencoder);

[0028] S3: The multi-omics integrated feature matrix of the training samples is mapped to a low-dimensional space through an autoencoder to obtain a multi-omics low-dimensional feature matrix of the training samples;

[0029] Based on the multi-omics integrated feature matrix obtained in step S2, an autoencoder is used to perform data fusion for the high-dimensional feature data set. The autoencoder used includes an input layer, a first convolutional layer, a bottleneck layer, a second convolutional layer, and an output layer; for the nodes of each layer, the number of input layer nodes is equal to the dimension of the feature data set obtained in S2 (44581), the number of first convolutional layer nodes is 1000, the number of bottleneck layer nodes is 500, the number of second convolutional layer nodes is 1000, and the number of output layer nodes is equal to the number of input layer nodes, which is 44581. The autoencoder learns valid data by minimizing the error between the reconstructed output data and the original input data. In this example, the bottleneck layer of the autoencoder is extracted to represent the new features of the original omics data.

[0030] In this embodiment, except for the output layer, the activation function used between each layer is the Relu activation function, specifically including the activation function between the input layer and the first convolutional layer, the activation function between the first convolutional layer and the bottleneck layer, the activation function between the bottleneck layer and the second convolutional layer, and the activation function between the second convolutional layer and the output layer;

[0031] The details are shown in the following formula:

[0032] f(x)=max(0,x)

[0033] Where x is the input real value or vector.

[0034] For real-valued x, the ReLU function turns negative values ​​to zero, while for non-negative values ​​x, the ReLU function remains unchanged. Therefore, if x is greater than or equal to zero, the output is x; if x is less than zero, the output is zero. For real-valued x, the output of the ReLU function can be expressed as:

[0035] For vector x, the ReLU function applies the ReLU operation to each element in the vector to obtain the corresponding output vector. Therefore, x can represent a single real value or a vector composed of multiple real values ​​in the ReLU function.

[0036] In this embodiment, the activation function used in the output layer of the autoencoder is Sigmoid, as shown in the following formula:

[0037]

[0038] Among them, y i is the output of the i-th neuron in the output layer, is the weight matrix of the output layer, is the input of the output layer, and is the bias term of the output layer.

[0039] Specifically, the data is fused in S3, and the number of hidden layer nodes and the number of bottleneck layer nodes are adjusted. In this example, the hidden layer nodes are configured between 1000 nodes and 1500 nodes, and the bottleneck layer nodes are configured between 500 nodes and 700 nodes. The specific process includes randomly dividing the input data set into a training set and a validation set, of which the validation set accounts for 70%. For each combination of the number of hidden layer nodes and the number of bottleneck layer nodes, an autoencoder model is constructed respectively. For each model configuration, an early stopping strategy is used to prevent overfitting. When the validation loss no longer decreases in 5 consecutive epochs, the training stops automatically, thereby avoiding the risk of long training time and overfitting. For each model configuration, after the training is completed, its best loss value on the validation set is recorded. The validation losses of all model configurations are compared, and the configuration with the smallest loss is selected as the best model. In this example, 1000 hidden layer nodes and 500 bottleneck layer nodes are selected as the best configuration.

[0040] S4: Based on the survival data of lung cancer patients, the VSOEnetbag algorithm combining elastic network and bagging idea is used to select the multi-omics low-dimensional feature matrix of training samples to obtain the significant expression feature matrix of training samples;

[0041] Preferably, the step S4 comprises:

[0042] S41: construct T sub-sample subspaces by sampling from the training set through the bagging algorithm;

[0043] Preferably, constructing the M sample subspaces includes: performing sampling with replacement from the training set, sampling with the same number of samples each time, and repeating the sampling T times to obtain T sample subspaces.

[0044] In this embodiment, several sample subspaces are generated, and each sample subspace is screened by multiple filters including the univariate Cox model, and LASSO regression is performed on each subset to perform feature screening on the variables through cross validation. After repeated use of LASSO regression on multiple subsets, the selection frequency of each feature will be recorded; the important features are screened out according to the cutoff threshold determined by the selection frequency of the feature. A certain proportion (80%) of samples are randomly selected from 314 samples to form a subspace. Repeat this operation 100 times to generate 100 subspaces. A certain proportion of samples (80%) are randomly selected from the training set as a subspace, and this subset usually contains repeated samples. Repeat the process 100 times, and each time a different sample subset is extracted, each subspace contains about 80% of the samples, that is, about 251 samples. In this way, 100 sample subspaces are generated through Bagging, each of which contains different samples from the training set.

[0045] S42: Traverse each sample subspace, take each feature element in the multi-omics low-dimensional feature matrix of the training sample as an independent variable, take the survival data of lung cancer patients as the dependent variable, and construct a univariate Cox regression model for each feature element;

[0046] In this embodiment, the first subspace is taken as an example, and the subspace has 251 samples and 500 features. For each of the 500 features, it is used as an independent variable, and the survival data of lung cancer patients corresponding to the 251 samples in the subspace is used as a dependent variable to construct a univariate Cox regression model.

[0047] For example, for feature 1, a Cox regression model is constructed with the value of feature 1 of all samples in the subspace as the independent variable and the corresponding survival data as the dependent variable; then the same operation is performed on feature 2 to construct another univariate Cox regression model until the corresponding models are constructed for all 500 features.

[0048] S43: Calculate the likelihood ratio test P value of each characteristic element based on the Cox regression model of each characteristic element;

[0049] Based on the Cox regression model constructed for each characteristic element, the likelihood ratio test P value of each characteristic element was calculated.

[0050] For example, for the 500 univariate Cox regression models constructed above, each model can calculate the likelihood ratio test P value for the corresponding feature, so that 500 P values ​​are obtained, corresponding to 500 features. In the context of univariate Cox regression models, the likelihood ratio test is used to compare the goodness of fit of two nested models.

[0051] S44: taking the feature elements whose likelihood ratio test P value is less than the set threshold as the initial features, and obtaining the initial feature set of each sample subspace;

[0052] Assume that the threshold is set to 0.05, and the feature elements with a likelihood ratio test P value less than 0.05 are used as initial features. For example, among the 500 calculated P values, 30 features have P values ​​less than , then these 30 features constitute the initial feature set of the subspace. This operation is performed on 100 subspaces, and each subspace obtains its own initial feature set.

[0053] S45: using the cross-validation elastic network to further perform feature selection on the initial feature set of each sample subspace, and obtain the candidate feature set after each bagging of the cross-validation elastic network;

[0054] Preferably, step S45 includes: in each bagging process, dividing the samples of each sample subspace into two parts according to a preset ratio, wherein the part with a small number of samples is used as a sub-validation set, and the part with a large number of samples is used as a sub-training set; using the survival data of the training samples in the sub-training set as labels, using the screened initial features as input, using the sub-training set to train the elastic network regression model, and after the training is completed, calculating the regression coefficient of the elastic network regression model for each initial feature through the sub-validation set, using the initial features whose regression coefficients are not zero as candidate features, and constructing a candidate feature set.

[0055] Preferably, the loss function during training of the elastic network regression model includes:

[0056]

[0057] Among them, L represents the loss function, y i represents the survival data of the i-th sample, x ij represents the jth initial feature of the i-th sample, β j represents the regression coefficient of the jth initial feature, λ1 and λ2 represent the coefficients of the regularization terms of L1 loss and L2 loss respectively, J represents the number of initial features, and n represents the number of samples.

[0058] In this embodiment, it is assumed that the preset ratio is to divide the samples into 70% and 30%. In the first subspace, 251 samples are divided according to this ratio, about 176 samples are used as sub-training sets, and 75 samples are used as sub-validation sets.

[0059] The survival data of the samples in the sub-training set are used as labels, and the selected initial features (such as the 30 initial features in the previous example) are used as input to train the elastic network regression model using the sub-training set. The loss function used during training is:

[0060]

[0061] Assume that λ1=0.5, λ2=0.3. After the training is completed, the regression coefficient of the elastic network regression model for each initial feature is calculated through the sub-validation set. The initial features with non-zero regression coefficients are taken as candidate features to construct a candidate feature set. For example, after the 30 initial features are calculated, the regression coefficients of 20 features are not zero. These 20 features constitute the candidate feature set after the subspace is bagged this time. This operation is performed on 100 subspaces in each bagging process. After the 100 subspaces have undergone multiple bagging processes, the number of times each feature element is selected as a candidate feature is counted. For example, feature A was selected into the candidate feature set 30 times during the bagging process of 100 subspaces. Then the frequency f of its selection as a candidate feature is A=30÷100÷P, where P represents the number of times each subspace is bagged.

[0062] S46: Calculate the frequency of each feature element being selected as a candidate feature according to the candidate feature sets of all sample subspaces in all bagging processes;

[0063] S47: Determine the significant expression feature matrix of the training sample using the PST method according to the frequency of each feature element being selected as a candidate feature;

[0064] Preferably, the step S47 includes:

[0065] S471: For each feature element j, the frequency f of its selection as a candidate feature j Calculate the corresponding number of selections k = T × f j , T represents the number of subspaces;

[0066] S472: Calculate the significance of feature element j:

[0067]

[0068] Among them, d j represents the significance of feature element j, X j represents the number of selections of feature element j in T subspaces; P(X j ≥k) represents the probability that the number of selections of feature element j in T subspaces is greater than k; p = 0.5 represents the probability that each feature element is randomly selected in a subspace; P i represents p to the power of i;

[0069] S473: Sort the significance of all feature elements in ascending order, and calculate the significance of feature element j after sorting. Among them, m represents the number of characteristic elements, l j The position of feature element j after sorting;

[0070] S474: If the significance of feature element j If it is less than the set threshold, the feature element j is taken as the significant expression feature, so as to construct the significant expression feature matrix of the training sample according to all the significant expression features of the training sample;

[0071] In this embodiment, it is assumed that the threshold is set to If it is less than the set threshold value of 0.01, feature A is taken as a significant expression feature. In this way, all significant expression features are determined, and thus a significant expression feature matrix of the training sample is constructed based on all significant expression features of the training sample.

[0072] S5: Use K-means clustering based on the significant feature expression matrix of the training samples to cluster the training samples into risk subtypes and obtain the subtype clustering results;

[0073] Assume that we have obtained the significant feature expression matrix of the training samples through the previous feature selection step. The matrix contains 200 lung cancer patient samples (training samples), each of which is described by 10 significant features. We hope to classify these patient samples into different risk subtypes through the K-means clustering algorithm.

[0074] S6: The subtype clustering results of the training samples were used as labels, and the RNA-Seq expression profile data of lung cancer patients were used as independent variables to construct a multivariate-based Cox prognostic model;

[0075] S7: The RNA-Seq expression profile data of the lung cancer patient to be tested is input into the multivariate-based Cox prognostic model to obtain the prognostic results of the lung cancer patient.

[0076] Another aspect of the present invention provides a lung cancer prognosis prediction device based on multi-omics data fusion, comprising a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory, and is used to execute the computer program stored in the memory, so that the lung cancer prognosis prediction device based on multi-omics data fusion performs the lung cancer prognosis prediction method based on multi-omics data fusion.

[0077] Another aspect of the present invention provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the method for predicting lung cancer prognosis based on multi-omics data fusion.

[0078] In this example, we used a combination of K-means clustering, LASSO feature selection and Cox survival analysis to build a model for lung cancer patient survival prediction. First, the top 30 key genes were extracted from the multi-omics data and standardized to suit subsequent analysis. Then, the suitability of different cluster numbers was evaluated based on the Silhouette Score, and the optimal cluster number was finally determined to be k = 2. Please refer to Figure 2 ,Through K-means clustering, we divided the patients into high-risk and low-risk groups, and verified the rationality of the clustering through t-SNE dimensionality reduction visualization, as shown in Figure 4 At the same time, we calculated the Jaccard similarity index (0.78), proving that this method can better capture the distribution characteristics of survival status.

[0079] In order to further optimize the model, we used LASSO regression for feature screening and finally selected 12 vectors with the greatest predictive value. After re-clustering with K-means based on the screened features, we found that the Jaccard index rose to 0.9382716 and the Silhouette Score increased to 0.1338907. In addition, we used Kaplan-Meier survival analysis to verify the survival significance of the clustering results in one step and found that the survival rate of the high-risk group was significantly lower than that of the low-risk group, with a p value of 0.022, indicating that the clustering method can effectively distinguish patients with different survival outcomes, such as Figure 3 shown.

[0080] In order to further quantify the ability to predict survival, we used the Cox proportional hazard regression model to evaluate the survival prediction effect of the screened genes. Cox analysis showed that the C-index (consistency index) of the model was as high as 0.78, indicating that the model had a strong ability to predict patient survival time. At the same time, in order to further evaluate the classification performance of the model, we used the ROC curve to calculate the AUC (area under the curve) of survival prediction, and finally obtained AUC = 0.81, Figure 4 It further shows that the model can accurately predict the survival outcomes of patients. In summary, the K-means+LASSO+Cox survival analysis model constructed in this study can not only effectively stratify patients by risk, but also provide reliable survival prediction information, which is helpful for the formulation of precision medicine and personalized treatment plans.

[0081] In summary, the present invention develops an integrated framework of multi-omics data for performing cancer subtype clustering, and realizes step-by-step dimensionality reduction and efficient integration of high-dimensional multi-omics data through multi-level feature extraction. Compared with single-omics analysis, multi-omics integration can improve the accuracy and reliability of framework prediction. At the same time, the present invention proposes an innovative multi-level extraction strategy, which effectively copes with the computational difficulties of large-scale multi-omics data and reduces the computational complexity of the model. This strategy not only improves the efficiency of data processing by hierarchical processing and optimized data representation, but also helps to reduce storage requirements and processing time, making large-scale multi-omics data analysis more feasible and efficient. The present invention proposes a feature selection method VSOEnetbagging based on elastic network (ElasticNet) regression and bagging (Bagging), which accurately screens out features significantly related to lung cancer prognosis on the basis of multiple iterations and cross-validation. By combining LASSO regression with multiple subspace iterations, overfitting can be effectively avoided, and the most representative and predictive feature variables can be screened out by statistical frequency. Compared with the traditional single feature selection method, this method has obvious advantages in improving feature selection stability and accuracy. Based on the feature matrix, the present invention uses the K-means clustering method to subtype lung cancer samples. This clustering process can divide samples into multiple subtypes, thereby revealing the potential biological heterogeneity of lung adenocarcinoma. This classification not only provides a new perspective for the biological understanding of cancer subtypes, but also provides an important basis for the formulation of individualized treatment plans. Compared with traditional simple classification methods, this method can more accurately capture the complex characteristics of cancer through cluster analysis of multi-omics data, and improve the ability to identify and stratify cancer subtypes.

[0082] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A lung cancer prognosis prediction method based on multi-omics data fusion, characterized in that: include: S1: Obtain multi-omics data and survival data of lung cancer patients, and pre-process the multi-omics data of each lung cancer patient; S2: The multi-omics data of each lung cancer patient are spliced ​​to obtain the multi-omics integrated feature matrix of each training sample; S3: The multi-omics integrated feature matrix of the training samples is mapped to a low-dimensional space through an autoencoder to obtain a multi-omics low-dimensional feature matrix of the training samples; S4: Based on the survival data of lung cancer patients, the VSOEnetbag algorithm combining elastic network and bagging idea is used to select the multi-omics low-dimensional feature matrix of training samples to obtain the significant expression feature matrix of training samples; S5: Use K-means clustering based on the significant feature expression matrix of the training samples to cluster the training samples into risk subtypes and obtain the subtype clustering results; S6: The subtype clustering results of the training samples were used as labels, and the RNA-Seq expression profile data of lung cancer patients were used as independent variables to construct a multivariate-based Cox prognostic model; S7: The RNA-Seq expression profile data of the lung cancer patient to be tested is input into the multivariate-based Cox prognostic model to obtain the prognostic results of the lung cancer patient.

2. The method for predicting lung cancer prognosis based on multi-omics data fusion according to claim 1, characterized in that: The multi-omics data of lung cancer patients include: RNA-Seq expression profile data, DNA methylation data, miRNA expression data and protein expression data of lung cancer patients; the survival data of lung cancer patients include: survival time and survival outcome of lung cancer patients.

3. The method for predicting lung cancer prognosis based on multi-omics data fusion according to claim 1, characterized in that: The autoencoder comprises: an encoder and a decoder, wherein the encoder is composed of an input layer, a first hidden layer and a bottleneck layer, and is used to map the input high-dimensional features to a low-dimensional space; the decoder is composed of a second hidden layer and an output layer, and is used to restore the features mapped to the low-dimensional space to the original high-dimensional features.

4. The method for predicting lung cancer prognosis based on multi-omics data fusion according to claim 1, characterized in that: The step S4 comprises: S41: construct T sub-sample subspaces by sampling from the training set through the bagging algorithm; S42: Traverse each sample subspace, take each feature element in the multi-omics low-dimensional feature matrix of the training sample as an independent variable, take the survival data of lung cancer patients as the dependent variable, and construct a univariate Cox regression model for each feature element; S43: Calculate the likelihood ratio test P value of each characteristic element based on the Cox regression model of each characteristic element; S44: taking the feature elements whose likelihood ratio test P value is less than the set threshold as the initial features, and obtaining the initial feature set of each sample subspace; S45: using the cross-validation elastic network to further perform feature selection on the initial feature set of each sample subspace, and obtain the candidate feature set after each bagging of the cross-validation elastic network; S46: Calculate the frequency of each feature element being selected as a candidate feature according to the candidate feature sets of all sample subspaces in all bagging processes; S47: Determine the significant expression feature matrix of the training sample using the PST method according to the frequency of each feature element being selected as a candidate feature.

5. The method for predicting lung cancer prognosis based on multi-omics data fusion according to claim 4, characterized in that: The constructing of the M sample subspaces includes: performing sampling with replacement from the training set, the number of samples in each sampling is the same, and repeating the sampling T times to obtain T sample subspaces.

6. The method for predicting lung cancer prognosis based on multi-omics data fusion according to claim 4, characterized in that: The step S45 includes: in each bagging process, dividing the samples of each sample subspace into two parts according to a preset ratio, wherein the part with a small number of samples is used as a sub-validation set, and the part with a large number of samples is used as a sub-training set; using the survival data of the training samples in the sub-training set as labels, using the screened initial features as input, using the sub-training set to train the elastic network regression model, and after the training is completed, calculating the regression coefficient of the elastic network regression model for each initial feature through the sub-validation set, using the initial features with non-zero regression coefficients as candidate features, and constructing a candidate feature set.

7. The method for predicting lung cancer prognosis based on multi-omics data fusion according to claim 6, characterized in that: The loss function during the training of the elastic network regression model includes: Among them, L represents the loss function, y i represents the survival data of the i-th sample, x ij represents the jth initial feature of the i-th sample, β j represents the regression coefficient of the jth initial feature, λ1 and λ2 represent the coefficients of the L1 and L2 regularization terms respectively, J represents the number of initial features, and n represents the number of samples.

8. The method for predicting lung cancer prognosis based on multi-omics data fusion according to claim 4, characterized in that: The step S47 comprises: S471: For each feature element j, the frequency f of its selection as a candidate feature j Calculate the corresponding number of selections k = T × f j , T represents the number of subspaces; S472: Calculate the significance of feature element j: Among them, d j represents the significance of feature element h, X j represents the number of selections of feature element j in T subspaces; P(X j ≥k) represents the probability that the number of selections of feature element j in T subspaces is greater than k; p = 0.5 represents the probability that each feature element is randomly selected in a subspace; P i represents p to the power of i; S473: Sort the significance of all feature elements in ascending order, and calculate the significance of feature element j after sorting. Among them, m represents the number of characteristic elements, l j The position of feature element j after sorting; S474: If the significance of feature element j If it is less than the set threshold, the feature element j is taken as the significant expression feature, so as to construct the significant expression feature matrix of the training sample according to all the significant expression features of the training sample.

9. A lung cancer prognosis prediction device based on multi-omics data fusion, characterized in that: It comprises a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the lung cancer prognosis prediction device based on multi-omics data fusion executes the lung cancer prognosis prediction method based on multi-omics data fusion according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the method for predicting lung cancer prognosis based on multi-omics data fusion according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • A method for detecting platelet markers in non-small cell lung cancer

    CN118155715B

  • Breast cancer prediction method based on penalty COX regression

    CN114141360A

  • Cancer lifetime prediction method and system, terminal and storage medium

    CN115579133A

  • Method for detecting platelet marker of non-small cell lung cancer

    CN118155715A

  • Construction method of lung cancer malignant risk hierarchical model, model and system

    CN118609818A

Cited By

  • Feature dynamic screening-based secondary blood index data set generation method and system

    CN121051466A

  • Secondary blood index dataset generation method and system based on feature dynamic screening

    CN121051466B

  • A Method and System for Auxiliary Diagnosis of Lung Nodules Based on Multi-View Fuzzy Clustering

    CN122575687A