A data processing method for multi-omics cancer subtype classification

By generating an adversarial network fusion gene expression, protein expression and methylation signal matrix, and generating mRNA expression profiles, the problem of insufficient data integration in multiomic cancer research was solved, and the accuracy of cancer subtype was improved.

CN116779040BActive Publication Date: 2025-07-25CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310702675.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-14
Publication Date
2025-07-25
Estimated Expiration
2043-06-14

AI Technical Summary

Technical Problem

Existing multiomic cancer studies have failed to effectively integrate different omic data, resulting in low accuracy of analysis results.

Method used

Generative adversarial network (GAN) is used to fuse the gene expression matrix, protein expression matrix and methylation signal matrix to generate mRNA expression profiles with protein information and apparent group information, and perform subtype clustering.

Benefits of technology

The classification accuracy of cancer subtype classification is improved, and the mRNA expression profile generated by the adversarial network is generated to achieve efficient fusion of multiomic data, which improves the accuracy of the analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116779040B_ABST
    Figure CN116779040B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of deep learning, and specifically relates to a data processing method for multi-omics cancer subtype classification, including: obtaining an initial gene expression matrix, a methylation signal value matrix, and a protein expression matrix of a target object; constructing the action relationship between genes and proteins; constructing the action relationship between methylation sites and genes according to the initial gene expression matrix and the methylation signal value matrix; inputting the initial gene expression matrix, the protein expression matrix, and the action relationship between genes and proteins into a generative adversarial network to obtain a first gene expression matrix; inputting the first gene expression matrix, the methylation signal value matrix, and the action relationship between methylation sites and genes into a generative adversarial network to obtain a second gene expression matrix; performing subtype clustering on the second gene expression matrix; The present invention fuses multi-omics data, uses a generative adversarial network to generate an mRNA expression profile with protein information and epigenomic information, and improves the accuracy of classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and particularly to a data processing method for multi-omics cancer subtype classification. Background Art

[0002] With the development of computer technology, the use of computer technology for auxiliary diagnosis and treatment in the medical field is gradually becoming popular. As a difficult-to-cure disease, more and more research has been carried out on cancer. Multi-omics data refers to data considering multiple levels (such as genes, proteins, metabolites, etc.) simultaneously. Cancer multi-omics research is a cancer research method based on multi-omics data, aiming to understand the molecular mechanisms of cancer occurrence and development. Cancer multi-omics data usually includes data in aspects such as genomics, transcriptomics, proteomics, and metabolomics; by integrating these multi-omics data, biomarkers such as genes, proteins, and metabolites related to cancer can be identified, so as to more deeply understand the molecular mechanisms of cancer; at the same time, the integration of multi-omics data can also help predict the treatment response and survival rate of patients, providing a reference basis for personalized treatment.

[0003] However, most current multi-omics cancer research analyzes different omics data separately and finally integrates the analysis results, without fusing different omics data at the data level, resulting in low accuracy of the analysis results. Summary of the Invention

[0004] To solve the above problems existing in the prior art, the present invention proposes a data processing method for multi-omics cancer subtype classification, which includes:

[0005] S1: Obtain multi-omics data of a target object, where the multi-omics data includes an initial gene expression matrix, a methylation signal value matrix, and a protein expression matrix of the target object;

[0006] S2: Construct the relationship between genes and proteins according to the initial gene expression matrix and the protein expression matrix;

[0007] S3: Construct the relationship between methylation sites and genes according to the initial gene expression matrix and the methylation signal value matrix;

[0008] S4: Input the initial gene expression matrix, the protein expression matrix, and the relationship between genes and proteins into a generative adversarial network to obtain a first gene expression matrix;

[0009] S5: Input the first gene expression matrix, the methylation signal value matrix, and the relationship between methylation sites and genes into a generative adversarial network to obtain a second gene expression matrix;

[0010] S6: Perform subtype clustering on the second gene expression matrix;

[0011] S7: Perform data analysis based on the subtype clustering results.

[0012] Advantages of the present invention:

[0013] The present invention fuses multi-omics data and uses a generative adversarial network (GAN) to generate an mRNA expression profile with protein information and epigenomic information; subtype classification of cancer is performed based on the mRNA expression profile, improving the accuracy of classification. Description of the Drawings

[0014] Figure 1 is the overall flowchart of the present invention;

[0015] Figure 2 is the PCA clustering diagram of five LUAD molecular subtypes of the present invention;

[0016] Figure 3 is the waterfall diagram of mutation information of five LUAD molecular subtypes of the present invention;

[0017] Figure 4 is the survival curve of five LUAD molecular subtypes of the present invention. Detailed Embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] A data processing method for multi-omics cancer subtype classification, as Figure 1 shown, the method includes:

[0020] S1: Obtain multi-omics data of a target object, where the multi-omics data includes an initial gene expression matrix, a methylation signal value matrix, and a protein expression matrix of the target object;

[0021] S2: Construct the relationship between genes and proteins according to the initial gene expression matrix and the protein expression matrix;

[0022] S3: Construct the relationship between methylation sites and genes according to the initial gene expression matrix and the methylation signal value matrix;

[0023] S4: Input the initial gene expression matrix, the protein expression matrix, and the relationship between genes and proteins into a generative adversarial network to obtain a first gene expression matrix;

[0024] S5: Input the first gene expression matrix, methylation signal value matrix, and the relationship between methylation sites and genes into a generative adversarial network to obtain a second gene expression matrix;

[0025] S6: Perform subtype clustering on the second gene expression matrix;

[0026] S7: Conduct data analysis based on the subtype clustering results.

[0027] In this embodiment, obtaining the multi-omics data of the target object includes: downloading the gene expression matrix (n = 576) and methylation signal value matrix (n = 492) of lung adenocarcinoma (LUAD) from the TCGA database. Downloading the protein expression matrix (n = 362) from the TCPA database. Taking the intersection of the matrix samples of the three omics, including 305 samples.

[0028] Obtain the protein-protein interaction network (PPI) from the STRING database. Obtain the gene annotation file (GTF) from the GENECOED database, and use the gene-protein correspondence therein to combine with the PPI to obtain the gene-protein interaction relationship. Construct an adjacency matrix: record the gene-protein relationship with an effect as 1, and the relationship without an effect as 0.

[0029] Obtain the 450k methylation chip annotation file from the GEO database, which includes the names of methylation sites, the coordinates of methylation sites, and the genes affected by methylation sites, etc. Since the intensity of the effect of methylation sites in different regions on genes is different, it is generally considered that the methylation sites near the transcription start site have a higher intensity of effect on genes; at the same time, there are also differences in the intensity of the effect on genes when they occur in the gene region and upstream of the gene; and in order to compare the relative scores of different methylation-gene pairs, the gene length is also taken into account. Therefore, for the methylation sites occurring upstream of the gene, the following formula is used to calculate the methylation-gene effect score:

[0030] score = 1 + |methsite - start| / width

[0031] The methylation-gene score formula for the methylation sites in the gene region is:

[0032] score = 1 - |methsite - start| / width

[0033] Among them, methsite refers to the methylation site coordinate, start refers to the gene transcription start site coordinate, and width refers to the gene length.

[0034] Input the protein expression matrix, gene expression matrix, and the relationship between genes and proteins into a generative adversarial network to obtain a gene expression matrix that includes proteomic and transcriptomic information. Specifically, the protein expression matrix is derived from reverse-phase protein array (RPPA) technology and reflects the abundance of proteins in samples. The gene expression matrix is derived from sequencing data in The Cancer Genome Atlas (TCGA) database and reflects the level of gene expression. The relationship between genes and proteins is derived from the Search Tool for the Retrieval of Interacting Genes / Proteins (STRING) database, where a relationship is defined as 1 and no relationship is defined as 0. Specifically, the generative adversarial network includes a generator and a discriminator. The generator includes 4 fully connected layers, and the learning rate is 5×10 -6 , and the discriminator includes 3 fully connected layers, and the learning rate is 5×10 -5 . In addition, for the gene expression matrix that integrates proteomics and transcriptomics, tumor clinical information also needs to be combined to construct a classifier to evaluate the quality of the generated data. If the area under the ROC curve (AUC) is higher than that of the original gene expression data when the generated gene expression data is used for tumor clinical information classification, the quality of the generated data is qualified.

[0035] Input the gene expression matrix that integrates proteomics and transcriptomics, methylation matrix, and the relationship between methylation sites and genes into a generative adversarial network. Through training and learning, the generative adversarial network obtains a gene expression matrix that includes proteomic, transcriptomic, and epigenomic information. Specifically, the gene expression matrix that integrates proteomics and transcriptomics is derived from the output of the first generative adversarial network. Specifically, the methylation signal value matrix is derived from microarray data in The Cancer Genome Atlas (TCGA) database and reflects the effect intensity of methylation sites in samples. The relationship between methylation sites and genes is derived from the annotation file of the 450k methylation chip in the Gene Expression Omnibus (GEO) database. Specifically, the generative adversarial network includes a generator and a discriminator. The generator includes 4 fully connected layers, and the learning rate is 5×10 -6 , and the discriminator includes 3 fully connected layers, and the learning rate is 5×10 -5 . In addition, for the gene expression matrix that integrates proteomic, transcriptomic, and epigenomic information, tumor clinical information also needs to be combined to construct a classifier to evaluate the quality of the generated data. If the area under the ROC curve (AUC) is higher than that of the gene expression data generated in step 1 when the generated gene expression data is used for tumor clinical information classification, the quality of the generated data is qualified.

[0036] The generative adversarial network includes a generator and a discriminator. An attention mechanism is introduced into the generator to make the generator pay more attention to important feature regions and generate a more accurate gene expression matrix. The generator includes 3 deconvolution blocks and an attention module. The first deconvolution block consists of a deconvolution layer, a batch normalization layer, and a Leaky Relu activation layer, and the size of its convolution kernel is 1*1; the structure of the second deconvolution block is the same as that of the first deconvolution block; the third deconvolution block consists of a deconvolution layer and a Tanh activation layer, and the size of the convolution kernel is 1*1; the attention module includes a convolution layer and a Sigmoid activation layer; the discriminator includes 4 convolution blocks. The first convolution block, the second convolution block, and the third convolution block are all composed of a convolution layer and a Leaky Relu activation layer, and the size of their convolution kernels is 3*3; the fourth convolution block consists of a convolution layer, and the size of the convolution kernel is 1*1.

[0037] The process of using the generative adversarial network to process the initial gene expression matrix, protein expression matrix, and the relationship between genes and proteins includes: inputting the initial gene expression matrix into the discriminator; inputting the protein expression matrix and the relationship between genes and proteins into the generator to generate a gene expression matrix with fusion information; inputting the generated gene expression matrix into the discriminator, and the discriminator calculates the loss function of the discriminator based on the initial gene expression matrix and the generated gene expression matrix and passes the loss function of the discriminator to the generator; the generator constructs the loss function of the generator according to the loss function of the discriminator and performs iterative training on the generator. When the generator and the discriminator converge, a gene expression matrix with protein information is output.

[0038] The loss function of the generator is:

[0039] y g =-C(G(Y)) + α||G(Y)-X||2

[0040] Y = S*P

[0041] The loss function of the discriminator is:

[0042] C(G(Y)) - C(X)

[0043] Among them, y g is the loss function of the generator, X represents the initial gene expression matrix, α represents the weight parameter. G(Y) represents the gene expression matrix generated by the generator, Y represents the fusion information of the protein matrix and the relationship between genes and proteins, S represents the protein expression matrix, and P represents the relationship between genes and proteins.

[0044] In this embodiment, the α value is optimized; specifically, the Bayesian optimization method is used to optimize the α value. The specific process includes: setting the loss function of the generator as the objective function, and at the same time using the tree-structured model as the surrogate model, setting the range of α from 0 to 0.5; inputting the parameter α range, the surrogate model, and the objective function into the Bayesian optimizer for iterative calculation. The specific process includes: first, randomly select some α values from 0 to 0.5 as the initial sample points and input them into the objective function for evaluation, and the obtained function values are used to construct the initial surrogate model; then, through the acquisition function, determine the next α to be evaluated, input it into the objective function for evaluation, and obtain the function value under this α configuration according to the output of the objective function; then add the new parameter configuration α and the function value to the existing evaluation data and update the surrogate model. By continuously updating the surrogate model, selecting new α values, and evaluating the objective function, it gradually converges to the global optimal solution. In this process, the entropy acquisition function is used to measure the uncertainty of the parameter configuration, and the parameter configuration with the greatest uncertainty is selected. When the objective function converges, the optimal α value is selected. At the same time, an accelerator is introduced to optimize the entropy acquisition function, and multi-threading and multi-process computing are used to parallelly compute these subtasks to accelerate the computing speed of the acquisition function.

[0045] For the gene expression matrix that integrates proteomics, transcriptomics, and epigenomics, unsupervised clustering methods such as K-means, SNF, and PAM are used to cluster the samples to explore the molecular subtypes of tumors. There are two methods for selecting the clustering number k value. One is the elbow method: for a set of data, as the K value increases, the clustering effect will gradually improve, but as the K value further increases, the improvement of the clustering effect will gradually slow down. Therefore, plot the relationship curve between the K value and the clustering error, and select the K value corresponding to the elbow bend. The other is the silhouette coefficient method: the silhouette coefficient is an index used to evaluate the clustering effect, which reflects the similarity of each sample to its own cluster and the dissimilarity to other clusters. For a set of data, different K values can be tried, calculate the silhouette coefficient of each sample, and calculate the average value of the silhouette coefficients of all samples. Select the K value with the largest average silhouette coefficient as the optimal K value. The clusters obtained by clustering are the different molecular subtypes of the tumor.

[0046] The K-means method is used to cluster the generated expression matrix including three omics information. By comparing the silhouette coefficient and the clustering error, the optimal clustering number is determined to be 5, and finally 5 LUAD molecular subtypes are obtained. Among them, subtype 1 includes 70 samples, subtype 2 includes 75 samples, subtype 3 includes 79 samples, subtype 4 includes 43 samples, and subtype 5 includes 38 samples. The PCA diagram of the clustering results is shown in the appendix Figure 2 .

[0047] The mutation data of LUAD were downloaded from the TCGA database. The data format obtained was the mutation annotation format (MAF). The mutation data included 194,729 entries and 141 variables. The mutation analysis was based on the R package maftools. For the above five LUAD subtypes, the top 20 gene mutations were calculated. The results are shown in the attached figure. Figure 3 .

[0048] The clinical information of each sample was downloaded from the TCGA database, including the survival time and survival status of the sample, including 641 LUAD patient samples. The survival time and survival status were used as dependent variables, and different LUAD subtypes were used as independent variables to construct the Kaplan-Meier survival curve. The results are shown in the attached figure. Figure 4 .

[0049] The single-sample gene set enrichment (ssGSEA) method was used to combine the immune gene set to evaluate the differences in immune cells and immune functions among LUAD molecular subtypes.

[0050] For different subtypes, mutation analysis, survival analysis, and immune infiltration analysis were performed to explore the biological differences between different tumor subtypes.

[0051] The specific steps of mutation analysis are: download the mutation data of samples from the TCGA database, and obtain the data in the mutation annotation format (MAF). Mutation analysis is based on the R package maftools. For different cancer subtypes, the differences in mutation status are calculated and the results are displayed in Figure 3 middle.

[0052] The specific steps of survival analysis are as follows: Download the clinical information of each sample from the TCGA database, including the survival time and survival status of the sample. Use the survival time and survival status as dependent variables, and the different subtypes as independent variables to construct the Kaplan-Meier survival curve, and use the Log-Rank test to calculate the P value to obtain significance.

[0053] The specific steps of immune infiltration analysis are as follows: for the generated gene expression matrix including multi-omics information, the single sample gene set enrichment (ssGSEA) method is used, combined with the immune gene set to evaluate the differences in immune cells and immune functions between tumor molecular subtypes.

[0054] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A data processing method for multi-omics cancer subtype classification, characterized in that, Including: S1: Obtain multi-omics data of the target object, where the multi-omics data includes the initial gene expression matrix, methylation signal value matrix, and protein expression matrix of the target object; S2: Construct the relationship between genes and proteins based on the initial gene expression matrix and protein expression matrix; S3: Construct the relationship between methylation sites and genes based on the initial gene expression matrix and methylation signal value matrix; S4: Input the initial gene expression matrix, protein expression matrix, and the relationship between genes and proteins into a generative adversarial network to obtain the first gene expression matrix; Specifically including: Input the initial gene expression matrix into the discriminator; Input the protein expression matrix and the relationship between genes and proteins into the generator to generate a gene expression matrix with fusion information; Input the generated gene expression matrix into the discriminator, and the discriminator calculates the loss function of the discriminator based on the initial gene expression matrix and the generated gene expression matrix, and passes the loss function of the discriminator to the generator; The generator constructs the loss function of the generator based on the loss function of the discriminator and performs iterative training on the generator. When the generator and the discriminator converge, output the gene expression matrix with protein information; The loss function of the generator is: y g = -C(G(Y)) + α‖G(Y) - X‖² Y = S * P The loss function of the discriminator is: C(G(Y)) - C(X) Among them, y g is the loss function of the generator, X represents the initial gene expression matrix, α represents the weight parameter, G(Y) represents the gene expression matrix generated by the generator, Y represents the protein expression matrix and the fusion information of the relationship between genes and proteins, S represents the protein expression matrix, and P represents the relationship between genes and proteins; S5: Input the first gene expression matrix, methylation signal value matrix, and the relationship between methylation sites and genes into a generative adversarial network to obtain the second gene expression matrix; S6: Perform subtype clustering on the second gene expression matrix; S7: Perform data analysis based on the subtype clustering results.

2. The data processing method for multi-omics cancer subtype classification according to claim 1, characterized in that The process of constructing the relationship between genes and proteins includes: Obtain a protein-protein interaction network and a gene annotation file, where the protein-protein interaction network contains the network relationship between proteins, and the gene annotation file contains the protein information translated from genes; Deduce the relationship between genes and proteins based on the protein information translated from genes and the relationship between proteins; Construct an adjacency matrix based on the relationship between genes and proteins, where the gene-protein relationship with an effect is recorded as 1, and the relationship without an effect is recorded as 0; The deduced relationship between genes and proteins is: If gene A translates to produce protein B, and protein B has an interaction relationship with protein C, then gene A has an effect on protein C.

3. A data processing method for multi-omics cancer subtype classification according to claim 1, characterized in that Constructing the relationship between methylation sites and genes includes: Obtain a 450k methylation chip annotation file, which includes the name of the methylation site, the coordinates of the methylation site, and the genes affected by the methylation site; Calculate the methylation-gene effect score of the methylation site upstream of the gene and the methylation-gene score of the methylation site in the gene region based on the name of the methylation site, the coordinates of the methylation site, and the genes affected by the methylation site; Obtain the relationship between methylation sites and genes based on the calculated methylation-gene scores.

4. The data processing method for multi-omics cancer subtype classification according to claim 3, characterized in that The formula for the methylation-gene effect score of the methylation site upstream of the gene is: score = 1 + |methsite - start| / width The methylation-gene score formula for the methylation sites in the gene region is as follows: score = 1 - |methsite - start| / width where methsite refers to the methylation site coordinate, start refers to the gene transcription start site coordinate, and width refers to the gene length.

5. A data processing method for multi-omics cancer subtype classification according to claim 1, characterized in that, The generative adversarial network includes a generator and a discriminator; the generator includes 3 transposed convolution blocks and an attention module. The first transposed convolution block consists of a transposed convolution layer, a batch normalization layer, and a Leaky Relu activation layer, and the size of its convolution kernel is 1*1; The structure of the second transposed convolution block is the same as that of the first transposed convolution block; the third transposed convolution block consists of a transposed convolution layer and a Tanh activation layer, and the size of the convolution kernel is 1*1; the attention module includes a convolution layer and a Sigmoid activation layer; the discriminator includes 4 convolution blocks. The first convolution block, the second convolution block, and the third convolution block are all composed of a convolution layer and a Leaky Relu activation layer in front, and the size of their convolution kernels is 3*3; the fourth convolution block consists of a convolution layer, and the size of the convolution kernel is 1*1.

6. The data processing method for multi-omics cancer subtype classification according to claim 1, characterized in that The process of the generator processing the input data includes: inputting the protein expression matrix and the relationship between genes and proteins into 3 transposed convolution blocks to extract feature relationships and obtain feature maps; inputting the feature maps into the attention module, converting the feature maps into attention maps through the convolution layer, and normalizing the attention maps generated by the convolution layer to 0-1 using the sigmoid activation function; multiplying the feature maps by the normalized weights to weight the intermediate feature maps, thereby improving the accuracy of the generated matrix.

7. A data processing method for multi-omics cancer subtype classification according to claim 1, characterized in that Subtype clustering of the second gene expression matrix includes: using the silhouette coefficient method and the elbow method to determine the optimal number of clusters K; performing Kmeans clustering on the second gene expression matrix according to the determined K value to obtain different subtypes of cancer; the specific process of clustering includes: according to the selected K value, randomly selecting K samples as the cluster centers; then calculating the distance from each sample to the center point and dividing the sample into the category where the center point is located; then calculating the centroid of each category and using it as the new center point; continuing to calculate the distance from each sample to the new center point and dividing it into the new category until convergence.