A method and system for identifying cancer subtypes based on multi-omics data

By combining multiomics data, unsupervised multitasking clustering model and XGBOOST model, identifying the subtype of liver cancer is solved, and the problems of insufficient processing capabilities and poor feature interpretability in the existing technology are solved, and a more accurate and personalized liver cancer treatment strategy is achieved.

CN118762756BActive Publication Date: 2025-05-13FOSHAN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410729630.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-06
Publication Date
2025-05-13
Estimated Expiration
2044-06-06

AI Technical Summary

Technical Problem

When identifying liver cancer subtypes, the high-dimensional nonlinear data processing capabilities are limited, the computational complexity is high, and the characteristics of deep learning models are poor, making it difficult to find key gene characteristics that affect liver cancer subtypes.

Method used

Using the identification method based on multiomics data, the multiomics gene data is obtained through the unsupervised multi-task clustering model and the XGBOOST model, the multiomics gene data is obtained for pre-processing, clustering and fitting the cancer subtype tag, providing characteristic contribution, and improving the accuracy and interpretability of the recognition.

Benefits of technology

It improves the accuracy and interpretability of liver cancer subtype identification, optimizes treatment strategies, improves patient survival, and provides strong support for targeted drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118762756B_ABST
    Figure CN118762756B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for identifying cancer subtypes based on multi-omics data, the method comprising: obtaining multi-omics gene data, and preprocessing the multi-omics gene data to obtain initial data; clustering the initial data based on an unsupervised multi-task clustering model to obtain cancer subtype labels; and fitting the cancer subtype labels using an XGBOOST model based on mRNA features to obtain feature contribution. It can be seen that the present invention effectively identifies cancer subtypes and provides interpretability, which is beneficial to personalized treatment of patients and development of targeted drugs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics, and in particular to a method and device for identifying cancer subtypes based on multi-omics data. Background Art

[0002] One of the main reasons why liver cancer patients are difficult to treat is the heterogeneity of treatment for different tumor subtypes. The same treatment method leads to different survival outcomes for patients of different subtypes, and targeted treatment is needed for different subtypes. However, traditional liver cancer subtype identification mainly relies on clinicians to observe the patient's cancer histological characteristics and clinical manifestations, which often leads to low consistency of results. At the same time, with the rapid development of multi-omics technologies such as genomics, transcriptomics, proteomics and metabolomics, multi-omics data provides rich biological information. At the same time, with the rapid development of artificial intelligence technology, it has become possible to identify tumor subtypes by integrating multi-omics data and using machine learning methods to describe the molecular characteristics of liver cancer patients from multiple angles. Accurately identifying liver cancer subtypes by combining machine learning methods with multi-omics data can help optimize treatment strategies, improve treatment effects and patient survival rates.

[0003] However, the performance of the mainstream machine learning methods currently used to identify liver cancer subtypes is limited by high-dimensional nonlinear data, and there is much room for improvement. In particular, the cumulative number of multi-omics data features for cancer patients often reaches tens of thousands, which increases the computational complexity. Traditional machine learning methods are not good at processing such high-dimensional complex features, and although deep neural networks have certain advantages in processing such high-dimensional nonlinear features, such methods cannot well identify hidden distribution characteristics from high-noise patient multi-omics data, which reduces the performance of the model. On the other hand, deep learning models are usually presented in the form of black boxes, which means that it is difficult for researchers to understand the calculation process inside the neural network and to explain how the model comes up with specific prediction results, that is, the feature interpretability problem of deep learning algorithms. Black box models make it difficult to discover key gene features that affect liver cancer subtypes, which greatly increases the difficulty of medical decision-making and targeted drug development. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a method and device for identifying cancer subtypes based on multi-omics data, which can effectively identify cancer subtypes and provide explainability, which is beneficial to personalized treatment of patients and development of targeted drugs.

[0005] In order to solve the above technical problems, the first aspect of the present invention discloses a method for identifying cancer subtypes based on multi-omics data, the method comprising:

[0006] A method for identifying cancer subtypes based on multi-omics data, characterized in that the method comprises:

[0007] Acquiring multi-omics gene data, and preprocessing the multi-omics gene data to obtain initial data;

[0008] Performing clustering processing on the initial data based on an unsupervised multi-task clustering model to obtain cancer subtype labels;

[0009] Based on the mRNA features, the XGBOOST model was used to fit the cancer subtype labels to obtain feature contributions.

[0010] As an optional embodiment, in the first aspect of the present invention, the multi-omics gene data includes one or more of genomics data, transcriptomics data, proteomics data and metabolomics data.

[0011] As an optional embodiment, in the first aspect of the present invention, the preprocessing of the multi-omics gene data to obtain initial data includes:

[0012] Deleting samples and features with more than 20% missing in the multi-omics gene data;

[0013] The remaining missing items were imputed using the median;

[0014] Gene alignment of imputed multi-omics gene data and data for independent validation;

[0015] Standardize multi-omics gene data to obtain initial data.

[0016] As an optional implementation, in the first aspect of the present invention, clustering the initial data based on an unsupervised multi-task clustering model to obtain cancer subtype labels includes:

[0017] Performing fusion processing on the multi-omics data to obtain multi-omics fusion data;

[0018] Processing the multi-omics fusion data through a multi-head self-attention mechanism to obtain self-attention output features;

[0019] Processing the self-attention output features through a variational self-encoder to obtain multi-omics low-dimensional features;

[0020] The multi-omics low-dimensional features are clustered to obtain cancer subtype labels.

[0021] As an optional embodiment, in the first aspect of the present invention, clustering the multi-omics low-dimensional features to obtain cancer subtype labels includes:

[0022] Based on the clustering algorithm and the initial data, determining the number of clusters;

[0023] The multi-omics low-dimensional features are clustered according to the number of clusters and the feature similarity loss function to obtain cancer subtype labels.

[0024] As an optional implementation, in the first aspect of the present invention, the determining the number of clusters based on the clustering algorithm and the initial data includes:

[0025] Using the k-means clustering algorithm and the initial data, the WSS values ​​of different numbers of clusters are calculated;

[0026] An elbow diagram is drawn according to the obtained WSS values, and the number of clusters is determined according to the elbow diagram.

[0027] As an optional implementation, in the first aspect of the present invention, the fusing the multi-omics data to obtain the multi-omics fusion data includes:

[0028] Performing linear transformation processing on the multi-omics data to obtain transformation features;

[0029] The transformation features are weighted to obtain multi-omics fusion data.

[0030] The second aspect of the present invention discloses a cancer subtype identification system based on multi-omics data, the system comprising:

[0031] An acquisition module, the acquisition module is used to acquire multi-omics gene data and pre-process the multi-omics gene data to obtain initial data;

[0032] A cluster identification module, wherein the cluster identification module is used to perform clustering processing on the initial data based on an unsupervised multi-task clustering model to obtain cancer subtype labels;

[0033] A fitting module is used to fit the cancer subtype label based on the mRNA feature using the XGBOOST model to obtain the feature contribution.

[0034] As an optional implementation, in the second aspect of the present invention, the acquisition module preprocesses the multi-omics gene data to obtain initial data, including:

[0035] Deleting samples and features with more than 20% missing in the multi-omics gene data;

[0036] The remaining missing items were imputed using the median;

[0037] Gene alignment of imputed multi-omics gene data and data for independent validation;

[0038] Standardize multi-omics gene data to obtain initial data.

[0039] As an optional implementation, in the second aspect of the present invention, the cluster identification module performs clustering processing on the initial data based on an unsupervised multi-task clustering model to obtain a cancer subtype label, including:

[0040] Performing fusion processing on the multi-omics data to obtain multi-omics fusion data;

[0041] Processing the multi-omics fusion data through a multi-head self-attention mechanism to obtain self-attention output features;

[0042] Processing the self-attention output features through a variational self-encoder to obtain multi-omics low-dimensional features;

[0043] The multi-omics low-dimensional features are clustered to obtain cancer subtype labels.

[0044] As an optional implementation, in the second aspect of the present invention, the cluster identification module performs clustering processing on the multi-omics low-dimensional features to obtain cancer subtype labels, including:

[0045] Based on the clustering algorithm and the initial data, determining the number of clusters;

[0046] The multi-omics low-dimensional features are clustered according to the number of clusters and the feature similarity loss function to obtain cancer subtype labels.

[0047] As an optional implementation, in the second aspect of the present invention, the cluster identification module determines the number of clusters based on the clustering algorithm and the initial data, including:

[0048] Using the k-means clustering algorithm and the initial data, the WSS values ​​of different numbers of clusters are calculated;

[0049] An elbow diagram is drawn according to the obtained WSS values, and the number of clusters is determined according to the elbow diagram.

[0050] As an optional implementation, in the second aspect of the present invention, the cluster identification module performs fusion processing on the multi-omics data to obtain multi-omics fusion data, including:

[0051] Performing linear transformation processing on the multi-omics data to obtain transformation features;

[0052] The transformation features are weighted to obtain multi-omics fusion data.

[0053] The third aspect of the present invention discloses another cancer subtype identification device based on multi-omics data, the device comprising:

[0054] A memory storing executable program code;

[0055] a processor coupled to the memory;

[0056] The processor calls the executable program code stored in the memory to execute part or all of the steps in a method for identifying cancer subtypes based on multi-omics data disclosed in the first aspect of an embodiment of the present invention.

[0057] A fourth aspect of an embodiment of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute some or all of the steps in a cancer subtype identification method based on multi-omics data disclosed in the first aspect of an embodiment of the present invention.

[0058] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: by integrating multi-omics data and using deep learning technology to accurately identify liver cancer subtypes, the personalized level of treatment strategies and patient survival rates are improved. It can not only extract key information from high-dimensional nonlinear data, but also provide interpretability through the XGBoost model, which helps to discover key genes that affect liver cancer subtypes and provides strong support for clinical medical decision-making and targeted drug development. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0060] Figure 1 is a flow chart of a method for identifying cancer subtypes based on multi-omics data disclosed in an embodiment of the present invention;

[0061] Figure 2 is a schematic diagram of the structure of a cancer subtype identification system based on multi-omics data disclosed in an embodiment of the present invention;

[0062] Figure 3 is a schematic diagram of the structure of a cancer subtype identification device based on multi-omics data disclosed in an embodiment of the present invention;

[0063] Figure 4 It is an elbow diagram for determining the number of clusters by WSS (within-cluster sum of squares) in a cancer subtype identification method based on multi-omics data disclosed in an embodiment of the present invention;

[0064] Figure 5It is a survival curve (Kaplan-Meier curve) of cell gene expression data and clustering model patient subtype labels in a cancer subtype identification method based on multi-omics data disclosed in an embodiment of the present invention;

[0065] Figure 6 It is a survival curve (Kaplan-Meier curve) of an independent validation experiment in a cancer subtype identification method based on multi-omics data disclosed in an embodiment of the present invention;

[0066] Figure 7 It is a gene importance graph of the top ten interpretable feature importances in a cancer subtype identification method based on multi-omics data disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0067] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0068] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, device, product or end including a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or ends.

[0069] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0070] The present invention discloses a method and system for identifying cancer subtypes based on multi-omics data, which can effectively identify cancer subtypes and provide explainability, which is beneficial to personalized treatment of patients and development of targeted drugs. The following are detailed descriptions.

[0071] Embodiment 1

[0072] See also Figure 1 , Figure 1 : is a flow chart of a method for identifying cancer subtypes based on multi-omics data disclosed in an embodiment of the present invention. Figure 1 The described method is applied to a cancer subtype identification device based on multi-omics data. The identification device can be a corresponding identification terminal, identification equipment or server, and the server can be a local server or a cloud server, which is not limited in the embodiment of the present invention.

[0073] like Figure 1 As shown, the method for identifying cancer subtypes based on multi-omics data may include the following operations:

[0074] 101. Obtain multi-omics gene data, and pre-process the multi-omics gene data to obtain initial data.

[0075] In the embodiment of the present invention, the multi-omics gene data includes genomics data, transcriptomics data, proteomics data and metabolomics data.

[0076] 102. The initial data is clustered based on the unsupervised multi-task clustering model to obtain cancer subtype labels.

[0077] 103. Based on mRNA features, the XGBOOST model is used to fit cancer subtype labels and obtain feature contributions.

[0078] It can be seen that the method described in the embodiment of the present invention can effectively identify cancer subtypes through unsupervised multi-task clustering models and multi-omics data, and provide interpretability through mRNA features and XGBOOST models, so that researchers can understand which features are most critical for the identification of liver cancer subtypes, which is beneficial to personalized treatment of patients and targeted drug development.

[0079] In an optional embodiment, in step 101, preprocessing the multi-omics gene data to obtain initial data includes:

[0080] Delete samples and features with more than 20% missing in multi-omics gene data;

[0081] The remaining missing items were imputed using the median;

[0082] Gene alignment of imputed multi-omics gene data and data for independent validation;

[0083] Standardize multi-omics gene data to obtain initial data.

[0084] In an optional embodiment, in the above step 102, clustering the initial data based on the unsupervised multi-task clustering model to obtain cancer subtype labels includes:

[0085] Perform fusion processing on multi-omics data to obtain multi-omics fusion data;

[0086] The multi-omics fusion data is processed through the multi-head self-attention mechanism to obtain the self-attention output features;

[0087] The self-attention output features are processed by the variational self-encoder to obtain multi-omics low-dimensional features;

[0088] Multi-omics low-dimensional features are clustered to obtain cancer subtype labels.

[0089] In an optional embodiment, in the above step, fusing the multi-omics data to obtain the multi-omics fusion data includes:

[0090] Perform linear transformation on multi-omics data to obtain transformation features;

[0091] The transformed features are weighted to obtain multi-omics fusion data.

[0092] In an embodiment of the present invention, a Squeeze-and-Excitation structure is used to linearly transform the multi-omics data through a fully connected layer, and the weighted results of the multi-omics data are output through a Sigmoid function. It is worth noting that you can choose whether to use cross-modal flow and whether to share weights. In addition, it also supports recording the average value of the weighting coefficient during operation to dynamically adjust the influence of different omics and adapt to the input of different omics by learning to adjust the weights.

[0093] In an optional embodiment, in the above steps, the multi-omics fusion data is processed by a multi-head self-attention mechanism to obtain self-attention output features.

[0094] In an embodiment of the present invention, the multi-omics fusion data is input into a multi-head self-attention mechanism and the representation of the query, key, and value is obtained through linear transformation. Then, by calculating the attention score, the attention weight is generated using the Softmax function and applied to the value to obtain the self-attention output. Finally, the original input multi-omics fusion data is connected with the self-attention output for output to obtain the self-attention output feature.

[0095] In an optional embodiment, in the above steps, the self-attention output features are processed by a variational self-encoder to obtain multi-omics low-dimensional features.

[0096] In the embodiment of the present invention, after multi-omics fusion input is performed on multi-omics data, its multi-omics low-dimensional features are obtained through an encoder structure including a linear layer, a multi-head self-attention layer, a latent variable layer, and a heavy parameter layer. The output of the encoder is represented as a shared feature, which is optimized and back-propagated by multiple loss functions (i.e., multi-task modules). The neural network uses LeakyReLU as a nonlinear activation function, h l+1 =LeakyRelu(W×h l +b), W is the weight matrix and b is the bias.

[0097] In an optional embodiment, in the above step, clustering the multi-omics low-dimensional features to obtain cancer subtype labels includes:

[0098] Based on the clustering algorithm and initial data, determine the number of clusters;

[0099] Clustering multi-omics low-dimensional features according to the number of clusters and feature similarity loss function to obtain cancer subtype labels;

[0100] The number of clusters is determined based on the clustering algorithm and the initial data, including:

[0101] Using the k-means clustering algorithm and initial data, the WSS values ​​of different numbers of clusters are calculated;

[0102] An elbow diagram is drawn based on the obtained WSS values, and the number of clusters is determined based on the elbow diagram.

[0103] In the embodiment of the present invention, the obtained multi-omics low-dimensional features are optimized by a data decoding loss function (i.e., a data encoding and decoding task), and the process of the encoder constructing the multi-omics low-dimensional features is defined as: Z=E(X), where X is the result of the input initial multi-omics data after multi-omics fusion processing, i.e., multi-omics fusion data, and the decoded features are expressed as X'=D(Z); then the data decoding loss function can be expressed as:

[0104] The multi-omics low-dimensional features after dimensionality reduction by the encoder are clustered in the neural network through the feature similarity loss function (i.e., subtype unsupervised clustering task), and the samples are divided into K clusters, so that the sample similarity within each cluster is high, while the feature differences between different clusters are large. The number of clusters K is determined by the WSS (intra-cluster sum of squares) elbow diagram obtained by calculating different values ​​of K using the K-MEANS (k-means clustering algorithm). In each iteration, the sample is assigned to the closest cluster and the center point of the cluster is updated, and this is repeated until the cluster center no longer changes significantly. The similarity loss function in the network layer can be expressed as: Where t is used to measure the cluster center μ j and the embedding point z jThe similarity between them is recorded as: s ij The target distribution can be expressed as: In order to understand the degree of difference between the actual clustering effect and the target distribution, the accuracy of the clustering task can be evaluated.

[0105] Considering that this method is a multi-task optimized deep learning framework, and the multi-omics low-dimensional features after dimensionality reduction are sampled from the latent space of the variational autoencoder, in order to sample Z more efficiently and with latent space regularity, the KL divergence loss function (latent space regularization task) is used to quantify the difference between the latent variable distribution learned by the variational autoencoder and the standard normal distribution, which is recorded as: where σ 2 is the mean sampled from the latent space, u 2 is the variance sampled from the latent space;

[0106] The overall loss function of the entire neural network for the above multi-task optimization strategy can be expressed as: L = L D +β×L p +γ×L c Among them, β and γ are hyperparameters that can be controlled and selected to balance the importance of these tasks.

[0107] It can be seen that the implementation of this optional embodiment can accurately identify liver cancer subtypes by integrating multi-omics data and using deep learning technology, thereby improving the personalization level of treatment strategies and patient survival rates. It can not only extract key information from high-dimensional nonlinear data, but also provide interpretability through the XGBoost model, which helps to discover key genes that affect liver cancer subtypes and provides strong support for clinical medical decision-making and targeted drug development.

[0108] After building the XGboost lightweight model, in order to verify the accuracy of the model, we downloaded liver cancer patient tumor data including gse10141, gse14520, and gse54236 from the GEO database as independent validation data. Figures 4 to 7 , the independent verification results show that the p-value of the survival curve between different subtypes is 0.002086673, and the p-value of different clustering labels for survival time is less than the pre-set significance level of 0.05, indicating that there are significant survival differences between different subtypes, proving the superiority of the XGBoost prediction model and further confirming the effectiveness of the clustering model. For the feature importance results obtained by the XGBoost model, we analyzed its feature explanatory results and extracted the top ten gene features, which are VTN, SPRR3, SLC7A5, GPLD1, SLC12A2, TIMP4, PIK3R1, GGCX, TCEA2, UAP1

[0109] A survey of the top ten gene features ranked by importance on CNKI, PubMed, and Google Scholar found that five genes, VTN, SLC7A5, SLC12A2, PIK3R1, and GGCX, have been confirmed to be related to liver cancer. Studies have shown that they play an important role in the occurrence and development of cancer, as follows:

[0110] Regarding the VTN gene, "In vitro and in vivo experimental studies on the inhibition of hepatocellular carcinoma by low expression of VTN" mentioned that VTN may play an important role in the malignant growth of hepatocellular carcinoma, and this role has obvious gender differences.

[0111] Regarding the SLC7A5 gene, "Yes-associated protein 1and transcriptional coactivator with PDZ-binding motif activates the mammalian target of rapamycin complex 1pathway by regulating amino acid transporters in hepatocellular carcinoma" mentioned that high expression of SLC7A5 is significantly correlated with shorter survival of patients with hepatocellular carcinoma.

[0112] Regarding the SLC12A2 gene, "Quantitative proteome analysis of HCC cell lines with different metastatic potentials by SILAC" mentioned that in their study, SLC12A2 expression was significantly upregulated in highly metastatic tumors, which was further confirmed in liver cancer cells and culture supernatants, and in patient serum.

[0113] Regarding the PIK3R1 gene, "The expression of PIK3R1 in human primary hepatocellular carcinoma and its relationship with tumor epithelial-mesenchymal transition" mentioned that the expression level of PIK3R1 is increased in HCC tissues. High levels of PIK3R1 may promote the occurrence of tumor epithelial-mesenchymal transition and are related to the poor prognosis of HCC.

[0114] Regarding the GGCX gene, the article "GAMMA-GLUTAMYL CARBOXYLASE INHIBITS CELL MIGRATION INHEPATOCELLULAR CARCINOMACELL LINES" mentioned that GGCX expression has an impact on cell growth and invasion, especially in hepatocellular carcinoma (HCC). Understanding the mechanism of these molecular interactions may provide useful insights into the treatment of hepatocellular carcinoma and may reveal some potential therapeutic targets.

[0115] In addition to the above five genes that have been confirmed to be related to the occurrence and development of liver cancer, there are clear studies showing that there is a relationship between the TIMP4 gene and breast cancer, the SPRR3 gene and pancreatic cancer, and the UAP1 gene and bladder cancer. Esophageal cancer, pancreatic cancer and liver cancer are all digestive tract tumors. This result shows that they are very likely to be related to liver cancer.

[0116] Embodiment 2

[0117] See also Figure 2 , Figure 2 is a schematic diagram of a cancer subtype identification device based on multi-omics data disclosed in an embodiment of the present invention. Figure 2 The described device can be applied to a corresponding identification terminal, X identification device or server, and the server can be a local server or a cloud server, which is not limited in the embodiment of the present invention. Figure 2 As shown, the device may include:

[0118] An acquisition module 100 is used to acquire multi-omics gene data and pre-process the multi-omics gene data to obtain initial data;

[0119] In the embodiment of the present invention, the multi-omics gene data includes genomics data, transcriptomics data, proteomics data and metabolomics data.

[0120] A cluster identification module 200, which is used to perform clustering processing on the initial data based on an unsupervised multi-task clustering model to obtain cancer subtype labels;

[0121] The fitting module 300 is used to fit the cancer subtype label based on the mRNA feature using the XGBOOST model to obtain the feature contribution.

[0122] It can be seen that the method described in the embodiment of the present invention can effectively identify cancer subtypes through unsupervised multi-task clustering models and multi-omics data, and provide interpretability through mRNA features and XGBOOST models, so that researchers can understand which features are most critical for the identification of liver cancer subtypes, which is beneficial to personalized treatment of patients and targeted drug development.

[0123] In an optional embodiment, the acquisition module 100 preprocesses the multi-omics gene data to obtain initial data, including:

[0124] Delete samples and features with more than 20% missing in multi-omics gene data;

[0125] The remaining missing items were imputed using the median;

[0126] Gene alignment of imputed multi-omics gene data and data for independent validation;

[0127] Standardize multi-omics gene data to obtain initial data.

[0128] In an optional embodiment, the cluster identification module 200 performs clustering processing on the initial data based on an unsupervised multi-task clustering model to obtain cancer subtype labels, including:

[0129] Perform fusion processing on multi-omics data to obtain multi-omics fusion data;

[0130] The multi-omics fusion data is processed through the multi-head self-attention mechanism to obtain the self-attention output features;

[0131] The self-attention output features are processed by the variational self-encoder to obtain multi-omics low-dimensional features;

[0132] Multi-omics low-dimensional features are clustered to obtain cancer subtype labels.

[0133] In an optional embodiment, the cluster identification module 200 performs fusion processing on the multi-omics data to obtain multi-omics fusion data, including:

[0134] Perform linear transformation on multi-omics data to obtain transformation features;

[0135] The transformed features are weighted to obtain multi-omics fusion data.

[0136] In an embodiment of the present invention, a Squeeze-and-Excitation structure is used to linearly transform the multi-omics data through a fully connected layer, and the weighted results of the multi-omics data are output through a Sigmoid function. It is worth noting that you can choose whether to use cross-modal flow and whether to share weights. In addition, it also supports recording the average value of the weighting coefficient during operation to dynamically adjust the influence of different omics and adapt to the input of different omics by learning to adjust the weights.

[0137] In an optional embodiment, the cluster identification module 200 processes the multi-omics fusion data through a multi-head self-attention mechanism to obtain self-attention output features.

[0138] In an embodiment of the present invention, the multi-omics fusion data is input into a multi-head self-attention mechanism and the representation of the query, key, and value is obtained through linear transformation. Then, by calculating the attention score, the attention weight is generated using the Softmax function and applied to the value to obtain the self-attention output. Finally, the original input multi-omics fusion data is connected with the self-attention output for output to obtain the self-attention output feature.

[0139] In an optional embodiment, the cluster identification module 200 processes the self-attention output features through a variational self-encoder to obtain multi-omics low-dimensional features.

[0140] In the embodiment of the present invention, after multi-omics fusion input is performed on multi-omics data, its multi-omics low-dimensional features are obtained through an encoder structure including a linear layer, a multi-head self-attention layer, a latent variable layer, and a heavy parameter layer. The output of the encoder is represented as a shared feature, which is optimized and back-propagated by multiple loss functions (i.e., multi-task modules). The neural network uses LeakyReLU as a nonlinear activation function, h l+1 =Leaky Relu(W×h l +b), W is the weight matrix and b is the bias.

[0141] In an optional embodiment, the cluster identification module 200 performs clustering processing on multi-omics low-dimensional features to obtain cancer subtype labels, including:

[0142] Based on the clustering algorithm and initial data, determine the number of clusters;

[0143] Clustering multi-omics low-dimensional features according to the number of clusters and feature similarity loss function to obtain cancer subtype labels;

[0144] The number of clusters is determined based on the clustering algorithm and the initial data, including:

[0145] Using the k-means clustering algorithm and initial data, the WSS values ​​of different numbers of clusters are calculated;

[0146] An elbow diagram is drawn based on the obtained WSS values, and the number of clusters is determined based on the elbow diagram.

[0147] In the embodiment of the present invention, the obtained multi-omics low-dimensional features are optimized by a data decoding loss function (i.e., a data encoding and decoding task), and the process of the encoder constructing the multi-omics low-dimensional features is defined as: Z=E(X), where X is the result of the input initial multi-omics data after multi-omics fusion processing, i.e., multi-omics fusion data, and the decoded features are expressed as X'=D(Z); then the data decoding loss function can be expressed as:

[0148] The multi-omics low-dimensional features after dimensionality reduction by the encoder are clustered in the neural network through the feature similarity loss function (i.e., subtype unsupervised clustering task), and the samples are divided into K clusters, so that the sample similarity within each cluster is high, while the feature differences between different clusters are large. The number of clusters K is determined by the WSS (intra-cluster sum of squares) elbow diagram obtained by calculating different values ​​of K using the K-MEANS (k-means clustering algorithm). In each iteration, the sample is assigned to the closest cluster and the center point of the cluster is updated, and this is repeated until the cluster center no longer changes significantly. The similarity loss function in the network layer can be expressed as: Where t is used to measure the cluster center μ j and the embedding point z j The similarity between them is recorded as: s ij The target distribution can be expressed as: In order to understand the degree of difference between the actual clustering effect and the target distribution, the accuracy of the clustering task can be evaluated.

[0149] Considering that this method is a multi-task optimized deep learning framework, and the multi-omics low-dimensional features after dimensionality reduction are sampled from the latent space of the variational autoencoder, in order to sample Z more efficiently and with latent space regularity, the KL divergence loss function (latent space regularization task) is used to quantify the difference between the latent variable distribution learned by the variational autoencoder and the standard normal distribution, which is recorded as: where σ 2 is the mean sampled from the latent space, u 2 is the variance sampled from the latent space;

[0150] The overall loss function of the entire neural network for the above multi-task optimization strategy can be expressed as: L = L D +β×L p +γ×L c Among them, β and γ are hyperparameters that can be controlled and selected to balance the importance of these tasks.

[0151] It can be seen that the implementation of this optional embodiment can accurately identify liver cancer subtypes by integrating multi-omics data and using deep learning technology, thereby improving the personalization level of treatment strategies and patient survival rates. It can not only extract key information from high-dimensional nonlinear data, but also provide interpretability through the XGBoost model, which helps to discover key genes that affect liver cancer subtypes and provides strong support for clinical medical decision-making and targeted drug development.

[0152] Embodiment 3

[0153] See also Figure 3 , Figure 3 FIG. 1 is a schematic diagram of another cancer subtype identification device based on multi-omics data disclosed in an embodiment of the present invention. Figure 3 As shown, the device may include:

[0154] A memory 301 storing executable program codes;

[0155] a processor 302 coupled to the memory 301;

[0156] The processor 302 calls the executable program code stored in the memory 301 to execute part or all of the steps in the method for identifying cancer subtypes based on multi-omics data disclosed in the first embodiment of the present invention.

[0157] Embodiment 4

[0158] An embodiment of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute some or all of the steps in the cancer subtype identification method based on multi-omics data disclosed in Example 1 of the present invention.

[0159] The device embodiments described above are only illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, i.e., they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative work.

[0160] Through the specific description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution can be essentially or partly contributed to the prior art in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0161] Finally, it should be noted that the cancer subtype identification method and device based on multi-omics data disclosed in the embodiments of the present invention disclose only the preferred embodiments of the present invention, which are only used to illustrate the technical scheme of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical schemes described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical schemes from the spirit and scope of the technical schemes of the embodiments of the present invention.

Claims

1. A method for identifying cancer subtypes based on multi-omics data, characterized in that: The method comprises: Acquiring multi-omics gene data, and preprocessing the multi-omics gene data to obtain initial data; Performing clustering processing on the initial data based on an unsupervised multi-task clustering model to obtain cancer subtype labels; Based on the mRNA features, the XGBOOST model is used to fit the cancer subtype labels to obtain feature contributions; Wherein, the preprocessing of the multi-omics gene data to obtain initial data includes: Deleting samples and features with more than 20% missing in the multi-omics gene data; The remaining missing items were imputed using the median; Gene alignment of imputed multi-omics gene data and data for independent validation; Standardize multi-omics gene data to obtain initial data; The clustering process of the initial data based on the unsupervised multi-task clustering model to obtain cancer subtype labels includes: Performing fusion processing on the multi-omics data to obtain multi-omics fusion data; Processing the multi-omics fusion data through a multi-head self-attention mechanism to obtain self-attention output features; Processing the self-attention output features through a variational self-encoder to obtain multi-omics low-dimensional features; performing clustering processing on the multi-omics low-dimensional features to obtain cancer subtype labels; The clustering process of the multi-omics low-dimensional features to obtain cancer subtype labels includes: Based on the clustering algorithm and the initial data, determining the number of clusters; Clustering the multi-omics low-dimensional features according to the number of clusters and the feature similarity loss function to obtain cancer subtype labels; The determining the number of clusters based on the clustering algorithm and the initial data includes: Using the k-means clustering algorithm and the initial data, the WSS values ​​of different numbers of clusters are calculated; An elbow diagram is drawn according to the obtained WSS values, and the number of clusters is determined according to the elbow diagram.

2. The method for identifying cancer subtypes based on multi-omics data according to claim 1, characterized in that: The multi-omics gene data includes one or more of genomics data, transcriptomics data, proteomics data and metabolomics data.

3. The method for identifying cancer subtypes based on multi-omics data according to claim 1, characterized in that: The fusing the multi-omics data to obtain multi-omics fusion data includes: Performing linear transformation processing on the multi-omics data to obtain transformation features; The transformation features are weighted to obtain multi-omics fusion data.

4. A cancer subtype identification system based on multi-omics data, characterized in that: The system comprises: An acquisition module, the acquisition module is used to acquire multi-omics gene data and pre-process the multi-omics gene data to obtain initial data; A cluster identification module, wherein the cluster identification module is used to perform clustering processing on the initial data based on an unsupervised multi-task clustering model to obtain cancer subtype labels; A fitting module, wherein the fitting module is used to fit the cancer subtype label based on the mRNA feature using the XGBOOST model to obtain a feature contribution; Wherein, the preprocessing of the multi-omics gene data to obtain initial data includes: Deleting samples and features with more than 20% missing in the multi-omics gene data; The remaining missing items were imputed using the median; Gene alignment of imputed multi-omics gene data and data for independent validation; Standardize multi-omics gene data to obtain initial data; The clustering process of the initial data based on the unsupervised multi-task clustering model to obtain cancer subtype labels includes: Performing fusion processing on the multi-omics data to obtain multi-omics fusion data; Processing the multi-omics fusion data through a multi-head self-attention mechanism to obtain self-attention output features; Processing the self-attention output features through a variational self-encoder to obtain multi-omics low-dimensional features; performing clustering processing on the multi-omics low-dimensional features to obtain cancer subtype labels; The clustering process of the multi-omics low-dimensional features to obtain cancer subtype labels includes: Based on the clustering algorithm and the initial data, determining the number of clusters; Clustering the multi-omics low-dimensional features according to the number of clusters and the feature similarity loss function to obtain cancer subtype labels; The determining the number of clusters based on the clustering algorithm and the initial data includes: Using the k-means clustering algorithm and the initial data, the WSS values ​​of different numbers of clusters are calculated; An elbow diagram is drawn according to the obtained WSS values, and the number of clusters is determined according to the elbow diagram.

5. A cancer subtype identification device based on multi-omics data, characterized in that: The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the cancer subtype identification method based on multi-omics data according to any one of claims 1 to 3.

6. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the cancer subtype identification method based on multi-omics data according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Multi-omics cancer subtype classification method, system, device, medium and program product

    CN118296442A