A tumor heterogeneity identification method, device, electronic equipment and storage medium

By locating tumor risk genes using tumor copy number variation data and transcription profile data, unsupervised deconvolution analysis and biological function analysis are performed to construct a tumor prognostic model, solving the trauma problem caused by puncture biopsy and realizing non-invasive tumor heterogeneity analysis and survival prediction.

CN115375640BActive Publication Date: 2026-01-02HARBIN MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210964997.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-11
Publication Date
2026-01-02
Estimated Expiration
2042-08-11

AI Technical Summary

Technical Problem

Puncture biopsy is traumatic to patients and may increase the risk of tumor cells entering the bloodstream. Current technology makes it difficult to detect tumor heterogeneity non-invasively.

Method used

By locating tumor risk genes with consistent alterations using tumor copy number variation data and tumor transcription profile data, unsupervised deconvolution analysis is performed to identify subclone-specific genes. Combined with biological function and survival analysis, an optimal tumor prognosis model is constructed and analyzed using tumor MRI images.

Benefits of technology

This technology enables quantitative analysis of tumor heterogeneity using tumor MRI images without causing trauma to the patient, and predicts the survival time of cancer patients, providing a theoretical basis for precision oncology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115375640B_ABST
    Figure CN115375640B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of tumor identification, and provides a tumor heterogeneity identification method, device, electronic equipment and storage medium. The application firstly locates tumor risk genes with consistency change, then identifies subclone specific genes related to tumor risk gene expression, determines subclone specific genes related to survival of a patient to a specified correlation degree, performs consistency clustering analysis on sample patients to obtain a classification label, constructs an optimal tumor prognosis model and screens an optimal image genome feature according to the tumor MRI image and the classification label of the sample patient, and finally analyzes an external tumor MRI image through the optimal tumor prognosis model and the optimal image genome feature, so that tumor heterogeneity quantitative analysis is realized on the premise that the patient is not caused trauma, the survival time of a tumor patient is predicted only through the tumor MRI image, and important theoretical basis and application value are provided for tumor precision medicine.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of tumor recognition, and particularly relates to a tumor heterogeneity recognition method and device, electronic equipment and storage medium. BACKGROUND

[0002] Tumor is one of the main causes of death worldwide, causing serious social burden. Tumor intra-heterogeneity refers to the existence of morphological and functional differences between tumor cell subpopulations in different regions of the same primary tumor or between the primary focus and the metastatic focus. Due to the heterogeneity of tumors in time and space, it brings great difficulty to select ideal tumor markers and achieve precise treatment in clinic.

[0003] At present, detection of tumor heterogeneity usually requires puncture biopsy of directly obtained disease tissues, which is the gold standard for clinical diagnosis. This method first needs to obtain tumor tissue samples by puncture, separate single cells, and perform transcriptomic analysis by single-cell RNA sequencing (scRNA-seq) to understand the diversity of cell states and the heterogeneity of cell populations.

[0004] Due to the consideration of existing data resources, some studies choose to analyze heterogeneity using the transcriptional profile of tissue samples. In most current studies, the premise of deconvolution of patient gene expression profiles is to first deconvolute the expression profiles of mixed different tumor samples to obtain the reference expression profile of the tumor, and then perform supervised deconvolution according to the obtained reference expression profile to analyze the intra-tumor heterogeneity of the patient.

[0005] However, the applicant of the present application found at least the following defects in implementing the above technical solutions:

[0006] Puncture biopsy can cause trauma to the patient, may lead to complications, and even increase the risk of tumor cell entering the blood. SUMMARY

[0007] The purpose of the embodiments of the present application is to provide a tumor heterogeneity recognition method, which aims to solve the problems mentioned in the background.

[0008] The embodiments of the present application are implemented in the following way: a tumor heterogeneity recognition method, comprising the following steps:

[0009] locating tumor risk genes with consistent changes according to tumor copy number variation data and tumor transcriptional profile data;

[0010] performing unsupervised deconvolution analysis on the expression profile of the tumor risk gene to identify subclone-specific genes related to the expression of the tumor risk gene;

[0011] biological function analysis and survival analysis on the subclone-specific genes to determine subclone-specific genes that are associated with patient survival to a specified degree of association;

[0012] performing consistency clustering analysis on the sample patients based on the subclone-specific genes that are associated with survival function to a specified degree of association to obtain a classification label, and constructing an optimal tumor prognosis model and screening an optimal image genome feature according to the tumor MRI image of the sample patient and the classification label;

[0013] analyzing the external tumor MRI image through the optimal tumor prognosis model and the optimal image genome feature.

[0014] Preferably, the step of locating the consistently altered tumor risk genes according to the tumor copy number variation data and the tumor transcriptomic data comprises:

[0015] identifying copy number significantly variant regions to a specified degree according to the tumor copy number variation data;

[0016] locating the consistently altered tumor risk genes on the copy number significantly variant regions in combination with the tumor transcriptomic data.

[0017] Preferably, the step of performing biological function analysis on the subclone-specific genes comprises:

[0018] performing gene enrichment analysis on the subclone-specific genes through a biological pathway database, and marking the corresponding subclone-specific genes as biological functions through the most significantly enriched biological pathways.

[0019] Preferably, the step of constructing an optimal tumor prognosis model and screening an optimal image genome feature according to the tumor MRI image of the sample patient and the classification label comprises:

[0020] constructing a tumor prognosis model and screening an image genome feature through a supervised deconvolution algorithm, a machine learning algorithm and a convolutional neural network algorithm according to the tumor MRI image of the sample patient and the classification label, and selecting an optimal tumor prognosis model and an optimal image genome feature from the tumor prognosis model and the image genome feature according to an evaluation index; the machine learning algorithm comprises a plurality of data screening modes and a plurality of model structures.

[0021] Another object of the embodiment of the present application is to provide a tumor heterogeneity identification device, which comprises:

[0022] a tumor risk gene locating module configured to locate consistently altered tumor risk genes according to tumor copy number variation data and tumor transcriptomic data;

[0023] subclone-specific gene recognition module, configured to perform unsupervised deconvolution analysis on the expression profile of the tumor risk gene to identify a subclone-specific gene associated with expression of the tumor risk gene;

[0024] biological function analysis and survival analysis module, configured to perform biological function analysis and survival analysis on the subclone-specific gene to determine a subclone-specific gene associated with survival of a patient to a specified correlation degree;

[0025] tumor prognosis model construction module, configured to perform consistency clustering analysis on a sample patient based on the subclone-specific gene associated with survival function to a specified correlation degree, to obtain a classification label, and to construct an optimal tumor prognosis model and screen an optimal image genomic feature according to a tumor MRI image of the sample patient and the classification label;

[0026] tumor image analysis module, configured to analyze an external tumor MRI image through the optimal tumor prognosis model and the optimal image genomic feature.

[0027] Preferably, the tumor risk gene positioning module comprises:

[0028] copy number significant variation region identification subunit, configured to identify a copy number significant variation region to a specified degree according to tumor copy number variation data;

[0029] tumor risk gene positioning subunit, configured to position a tumor risk gene with consistent changes in the copy number significant variation region in combination with tumor transcription profile data.

[0030] Another object of the embodiments of the present application is to provide an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the tumor heterogeneity identification method according to any one of the above.

[0031] Another object of the embodiments of the present application is to provide a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the tumor heterogeneity identification method according to any one of the above.

[0032] The tumor heterogeneity identification method provided by the embodiment of the present application firstly locates tumor risk genes with consistent changes according to tumor copy number variation data and tumor transcriptome data, then performs unsupervised deconvolution analysis on the expression profile of the tumor risk genes to identify subclone-specific genes related to the expression of the tumor risk genes, then performs biological function analysis and survival analysis on the subclone-specific genes to determine subclone-specific genes with a specified correlation degree with patient survival, then performs consistent clustering analysis on sample patients based on the subclone-specific genes with a specified correlation degree with survival function to obtain classification tags, and constructs an optimal tumor prognosis model and screens optimal image genomic features according to the tumor MRI images of the sample patients and the classification tags, and finally analyzes external tumor MRI images through the optimal tumor prognosis model and the optimal image genomic features, so as to realize tumor heterogeneity quantitative analysis without causing trauma to patients, predict the survival time of tumor patients only through tumor MRI images, and provide important theoretical basis and application value for tumor precision medicine. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 A step flowchart of the tumor heterogeneity identification method provided by the embodiment of the present application is provided.

[0034] Figure 2 A step flowchart of locating tumor risk genes with consistent changes provided by the embodiment of the present application is provided.

[0035] Figure 3 A structure block diagram of the tumor heterogeneity identification device provided by the embodiment of the present application is provided.

[0036] Figure 4 A structure block diagram of the tumor risk gene positioning module provided by the embodiment of the present application is provided.

[0037] Figure 5 A flowchart of tumor heterogeneity identification provided by the embodiment of the present application is provided.

[0038] Figure 6 A work flowchart of the tumor heterogeneity identification platform provided by the embodiment of the present application is provided.

[0039] Figure 7 A structure schematic diagram of an electronic device provided by the embodiment of the present application is provided. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0041] The specific implementation of the present application is described in detail below in combination with specific embodiments.

[0042] As shown in FIG. 1, a tumor heterogeneity identification method provided by an embodiment of the present application includes the following steps: Figure 1 and 5 As shown in FIG. 1, a tumor heterogeneity identification method provided by an embodiment of the present application includes the following steps:

[0043] S100, locating a consistently changed tumor risk gene according to tumor copy number variation data and tumor transcriptome data.

[0044] Copy number refers to the number of a certain gene in a certain biological genome. Copy number variation is caused by rearrangement of a genome, generally refers to an increase or decrease in the copy number of a large fragment of a genome with a length of more than 1000 bases, and is an important component of genomic structural variation.

[0045] Tumor transcriptome data is obtained by transcriptome sequencing technology (RNA-seq), is a method of extracting total RNA and performing sequencing analysis by high-throughput sequencing technology, thereby reflecting the expression level of each gene and revealing the biological function of the gene.

[0046] In the present embodiment, first, a consistently changed tumor risk gene is located, and the specific locating method is locating according to tumor copy number variation data and tumor transcriptome data.

[0047] S200, performing unsupervised deconvolution analysis on the expression profile of the tumor risk gene to identify subclone-specific genes related to the expression of the tumor risk gene.

[0048] Subclone refers to the difference in growth rate, invasion ability, etc. of cell offspring due to partial changes in genetic material under the continuous proliferation of a single tumor cell. At this time, the tumor cell population is composed of multiple cell populations with heterogeneity, i.e., composed of multiple subclones.

[0049] Deconvolution is the inverse operation of convolution, which aims to solve another input under the condition that the output and a certain input are known. This algorithm is also called supervised deconvolution. Unsupervised deconvolution refers to the operation of solving two inputs under the condition that only the output is known. In the present embodiment, by performing unsupervised deconvolution analysis on the expression profile of the tumor risk gene (output), the subclone-specific genes and their proportion in each patient (two inputs) are solved.

[0050] In the present embodiment, the unsupervised deconvolution can specifically take the form of unsupervised convex analysis of mixture (CAM). CAM assumes that the sequencing-derived gene expression information can be regarded as a mixed expression of multiple potential gene subclones, and is a weighted expression profile of each potential subclone according to its proportion. By plotting a simplex scatter plot of the RNA expression profile, the apex of each simplex is taken as a potential subclone, differential expression analysis is performed on each gene of the subclone, key subclone-specific genes that are significantly correlated with tumor risk gene expression are identified, and the composition proportion matrix of different subclones in tumor samples and the specific expression matrix of each subclone are calculated by using the standardized mean and least squares method.

[0051] S300, biological function analysis and survival analysis are performed on the subclone-specific genes to determine the subclone-specific genes that are significantly correlated with patient survival.

[0052] In the present embodiment, the correlation between the proportion of different subclone-specific genes and survival function can be tested by log-rank test. The log-rank test is a statistical method used to evaluate the effect of a variable on patient survival. For each subclone-specific gene of all patients, the patients are grouped according to the different proportions of the subclone-specific gene in all patients and subjected to log-rank test. In order to avoid systematic errors caused by multiple random sampling during the test process, the Benjamini-Hochberg method is used to correct the p-value representing the significance of the correlation between the variable and survival obtained by the log-rank test, and the threshold value of the group with the smallest corrected p-value is taken as the final grouping threshold value of the subclone-specific gene. Kaplan-Meier survival curve is plotted, hazard ratio (HR) and its 95% confidence interval are calculated, and the significance of the composition proportion of the subclone-specific gene and patient survival is determined. Gene differential expression analysis is performed on the two groups of patients under the classification standard to explain the survival difference caused by different proportions of the subclone-specific gene.

[0053] In the present embodiment, the biological function of different subclone-specific genes is identified to interpret the subclone heterogeneity and its important role in tumor progression.

[0054] S400, consistency clustering analysis is performed on the sample patients based on the subclone-specific genes that are significantly correlated with survival function, a classification label is obtained, and an optimal tumor prognosis model is constructed and an optimal image genomic feature is screened according to the tumor MRI image of the sample patients and the classification label.

[0055] In the embodiment, the sample patients are subjected to consistency clustering analysis based on the subclone-specific genes reaching a specified correlation degree with survival function, the clustering results are determined based on a cumulative distribution function (CDF) curve, the patients are grouped, and a grouping label is obtained. Then, the optimal tumor prognosis model and the optimal imaging genomic feature are constructed based on the tumor MRI images of the sample patients and the classification label, so as to depict the relationship between the subclone-specific genes and survival, and to determine the productivity of the patients subsequently.

[0056] S500, analyzing the external tumor MRI images based on the optimal tumor prognosis model and the optimal imaging genomic feature.

[0057] In the embodiment, the optimal tumor prognosis model and the optimal imaging genomic feature determined in step S400 are used to analyze the external tumor MRI images, so as to evaluate the prognostic efficiency of the imaging genomic feature, to realize the quantitative analysis of tumor heterogeneity based on the tumor MRI images without causing trauma to the patients, to predict the survival time of the tumor patients, and to provide important theoretical basis and application value for tumor precision medicine.

[0058] In the embodiment, first, the tumor risk genes with consistent changes are located based on the tumor copy number variation data and the tumor transcriptomic data, then, the expression profile of the tumor risk genes is subjected to unsupervised deconvolution analysis, the subclone-specific genes related to the expression of the tumor risk genes are identified, then, the subclone-specific genes are subjected to biological function analysis and survival analysis, the subclone-specific genes reaching a specified correlation degree with patient survival are determined, then, the sample patients are subjected to consistency clustering analysis based on the subclone-specific genes reaching a specified correlation degree with survival function, a classification label is obtained, and the optimal tumor prognosis model and the optimal imaging genomic feature are constructed based on the tumor MRI images of the sample patients and the classification label, finally, the external tumor MRI images are analyzed based on the optimal tumor prognosis model and the optimal imaging genomic feature, so as to realize the quantitative analysis of tumor heterogeneity based on the tumor MRI images without causing trauma to the patients, to predict the survival time of the tumor patients, and to provide important theoretical basis and application value for tumor precision medicine.

[0059] Furthermore, in this embodiment, a pan-cancer imaging genomics database can be pre-constructed. Specifically, data related to 11 types of cancer, including RNA expression profiles, copy number variation profiles, patient survival information, and MRI imaging data, are obtained from TCGA (The Cancer Genome Atlas - Cancer Genome). This data is further integrated with resources such as the GEO (Gene Expression Omnibus) database, the TCIA (The Cancer Imaging Archive) database, the CGGA (Chinese Glioma Genome Atlas), and biological pathway databases (KEGG, Reactome, SMPDB, etc.) to improve the pan-cancer imaging genomics database.

[0060] In this embodiment, 11 types of cancer omics data from TCGA (MRI image data, RNA-seq gene expression profiles, copy number variation profiles, and clinical survival information) can be selected, as shown in the table below.

[0061] TCGA cancer types MRI RNA-seq Copy Number Variation Clinical Bladder urothelial carcinoma (BLCA) 20 406 412 412 Breast invasive carcinoma (BRCA) 137 1095 1098 1097 Cervical squamous cell carcinoma and endocervical adenocarcinoma (CESC) 54 304 302 307 Glioblastoma multiforme (GBM) 262 166 599 599 Head and neck squamous cell carcinoma (HNSC) 227 503 526 528 Kidney renal clear cell carcinoma (KIRC) 62 532 534 537 Kidney renal papillary cell carcinoma (KIRP) 17 290 290 291 Brain lower grade glioma (LGG) 199 514 515 515 Liver hepatocellular carcinoma (LIHC) 40 371 376 337 Prostate adenocarcinoma (PRAD) 10 497 498 500 Uterine corpus endometrial carcinoma (UCEC) 8 557 558 548

[0062] like Figure 2 As shown, in one embodiment, the step of locating tumor risk genes with consistent alterations based on tumor copy number variation data and tumor transcription profile data includes:

[0063] S101, based on tumor copy number variation data, identify regions with significant copy number variations that reach a specified level;

[0064] S102, combined with tumor transcriptional profiling data, locates tumor risk genes with consistent alterations in the regions of significant copy number variation.

[0065] In this embodiment, first, the tumor copy number variation profile of TCGA sample patients (take breast cancer as an example) is obtained from the GDC platform (Genomic Data Commons Data Portal). The data is generated by identifying the repeated genomic regions using relevant hardware devices (such as Affymetrix SNP 6.0 chip) and calculating the copy number of these regions through subsequent analysis. The GDC platform further converts the copy number into segment mean format, equal to log2[(copy number) / 2]. Humans are diploid, the segment mean is 0, the segment mean of the amplified region is positive, and the segment mean of the deleted region is negative. After obtaining the copy number variation data in the segment mean format, the GISTIC (Genome Identification of Significant Targets in Cancer) 2.0 module of the GenePattern platform is selected to analyze the tumor copy number variation data. The module considers the frequency and intensity of copy number variation at the same time, determines the frequently occurring copy number variation region by scoring the significance of the variation, and identifies the genes with significant copy number variation in combination with the GRCh38 human reference genome information. The RNA-seq data of healthy samples and breast cancer samples is obtained, and differential expression analysis is performed to identify genes with significantly changed expression after breast cancer. The data of significant copy number variation and significant expression change are integrated, and genes that are both increased or decreased are screened to locate the consistently changed breast cancer risk genes.

[0066] In one case of the present embodiment, the step of performing biological function analysis on the subclone-specific genes comprises:

[0067] The gene set enrichment analysis is performed on the subclone-specific genes by the biological pathway database, and the most significant enriched biological pathway is marked as the biological function of the corresponding subclone-specific gene.

[0068] In the present embodiment, pathway analysis is performed on each gene subclone-specific gene, and the relevance of each subclone-specific gene to a biological pathway database (which can be selected from KEGG, Reactome, SMPDB, etc.) is measured using hypergeometric testing to analyze its biological mechanism. Using the biological pathway database, gene set enrichment analysis (GSEA) is performed on the subclone-specific genes. This method first sorts the identified subclone-specific genes according to the degree of specificity, then observes the ranking of the genes in the subclone-specific genes in the gene sets of multiple pathways included in the biological pathway database, calculates the enrichment score of each gene set according to the ranking, and finally selects the pathway gene set according to the enrichment score as the biological pathway enriched by the subclone. The most significant pathway is selected as the name of the subclone, and is used as the biological function of the subclone.

[0069] In one case of the present embodiment, the step of constructing an optimal tumor prognosis model and screening optimal image genomic features according to the tumor MRI images of the sample patients and the classification labels comprises:

[0070] According to the tumor MRI images of the sample patients and the classification labels, a tumor prognosis model and image genomic features are constructed by a supervised deconvolution algorithm, a machine learning algorithm and a convolutional neural network algorithm, and the optimal tumor prognosis model and the optimal image genomic features are selected from the tumor prognosis model and the image genomic features according to the evaluation index; the machine learning algorithm includes multiple data screening methods and multiple model structures.

[0071] In the present embodiment, the evaluation index can be AUC, decision curve, etc. Specifically, after constructing multiple tumor prognosis models and screening image genomic features, an ROC curve is drawn, and the classification threshold is determined according to the Youden index (sensitivity + specificity - 1). The classification effect of the model is evaluated according to AUC, decision curve, etc., and the image genomic feature and the corresponding model with the best classification effect are selected as the final selection. After correcting the clinical variables such as age, ER, PR, HER2 and tumor size, a multivariate Cox regression model is used to analyze whether the image genomic feature is an independent prognostic factor for overall survival or recurrence-free survival. The likelihood ratio test is performed on the survival model using the image genomic feature as a variable and the survival model using only clinical variables to evaluate the improvement of the method on the prognosis efficiency of the model.

[0072] In the present embodiment, multiple data screening methods such as GBDT, LASSO, RF, XGBoost, etc. can be used, and multiple model structures such as RF, GBDT, Adaboost, LR, NB, SVM, DT, KNM, etc. can be used.

[0073] In machine learning tasks, the selection method of image genomic features and the structure of the tumor prognosis model directly affect the predictive performance of the model. To minimize the influence of training data and model structure and improve the model's classification effect on the proportion of patient gene subclonal composition, four feature selection methods (GBDT, LASSO, RF, and XGBoost) and eight different model structures (RF, GBDT, Adaboost, LR, NB, SVM, DT, and KNM) were selected for training. During feature selection, image features were selected based on the importance ranking of variables given by GBDT, RF, and XGBoost, while LASSO can reduce the regression coefficients of irrelevant variables to zero, retaining variables with non-zero coefficients. Based on this, image features were selected for classifying patients with different subclonal composition proportions. In terms of model construction, PyRadiomics was first used to quantify tumor information based on tumor size, shape, and boundary clarity, outputting gray-scale matrices, first-order statistics, and second-order statistics, which reflect numerical variables of tumor image information. Then, combined with classification labels, a machine learning model was constructed to classify patients.

[0074] With the advancement of computer performance and the digitization of medical images, numerous studies have emerged utilizing neural networks for model building. Compared to traditional machine learning tasks, Convolutional Neural Networks (CNNs) can directly understand and select from image data, avoiding errors caused by insufficient image information extraction. This method does not require quantifying image data; it can directly learn from the input image information. By setting a loss function, the model adjusts the weights of each node during multiple iterations, reducing classification errors. To fully extract image genomic features that reflect tumor heterogeneity, a t-test can be used to assess the significant differences in optimal image genomic features across different patient groups, characterizing the intrinsic relationship between optimal image genomic features and tumor heterogeneity genomic features. After training, heatmaps are generated to visualize the CNN's regions of interest, enhancing model interpretability.

[0075] As attached Figure 3 and 5 As shown, a tumor heterogeneity identification device according to an embodiment of the present invention includes:

[0076] The tumor risk gene localization module 100 is used to locate tumor risk genes with consistent alterations based on tumor copy number variation data and tumor transcription profile data.

[0077] In this embodiment, the first step is to locate tumor risk genes with consistent alterations. The specific location method is based on tumor copy number variation data and tumor transcription profile data.

[0078] sub-clonal specific genes associated with the expression of the tumor risk genes.

[0079] In this embodiment, the unsupervised deconvolution can be specifically implemented as unsupervised convex analysis of mixture (CAM). CAM assumes that the sequencing-derived gene expression information can be regarded as the mixed expression of multiple potential gene sub-clones, and is the weighted expression of the specific expression profile of each potential sub-clone according to its proportion. By plotting a simplex scatter plot of the RNA expression profile, the apex of each simplex is taken as a potential sub-clone, differential expression analysis is performed on each gene of the sub-clone, key sub-clonal specific genes significantly associated with the expression of the tumor risk genes are identified, and the composition proportion matrix of different sub-clones in the tumor sample and the specific expression matrix of each sub-clone are calculated by using the standardized mean and least square method.

[0080] The biological function analysis and survival analysis module 300 is used for biological function analysis and survival analysis of the sub-clonal specific genes, and determines the sub-clonal specific genes associated with survival of patients to a specified degree of correlation.

[0081] In this embodiment, the log-rank test can be used to analyze the correlation between the proportion of different sub-clonal specific genes and survival function. The log-rank test is a statistical method used to evaluate the effect of a variable on the survival of patients. For each sub-clonal specific gene of all patients, the patients are grouped according to the different proportions of the sub-clonal specific gene in all patients and subjected to log-rank test. In order to avoid systematic errors caused by multiple random sampling in the test process, the Benjamini-Hochberg method is used to correct the p-value representing the significance of the correlation between the variable and survival obtained by the log-rank test, and the threshold value of the group with the smallest corrected p-value is taken as the final grouping threshold value of the sub-clonal specific gene. The Kaplan-Meier survival curve is plotted, the hazard ratio (HR) and its 95% confidence interval are calculated, and the significance of the composition proportion of the sub-clonal specific gene and the survival of patients is determined. Gene differential expression analysis is performed on the two groups of patients under the classification standard, and the survival difference caused by different proportions of the sub-clonal specific gene is explained.

[0082] In this embodiment, the biological function of different sub-clonal specific genes is identified, thereby explaining the sub-clonal heterogeneity and its important role in tumor progression.

[0083] The tumor prognosis model construction module 400 is configured to perform consistent clustering analysis on the sample patients based on the subclone-specific genes that have a specified correlation degree with survival function, obtain a classification label, and construct an optimal tumor prognosis model and screen an optimal image genome feature according to the tumor MRI image of the sample patients and the classification label.

[0084] In this embodiment, the sample patients are grouped and a grouping label is obtained by performing consistent clustering analysis on the sample patients based on the subclone-specific genes that have a specified correlation degree with survival function, and determining the clustering result according to a cumulative distribution function (CDF) curve. Furthermore, an optimal tumor prognosis model is constructed and an optimal image genome feature is screened according to the tumor MRI image of the sample patients and the classification label, so as to depict the relationship between the subclone-specific genes and survival, and determine the productivity of the patients subsequently.

[0085] The tumor image analysis module 500 is configured to analyze the external tumor MRI image by using the optimal tumor prognosis model and the optimal image genome feature.

[0086] In this embodiment, the optimal tumor prognosis model and the optimal image genome feature determined by the tumor prognosis model construction module 400 are used to analyze the external tumor MRI image, so as to evaluate the prognostic efficiency of the image genome feature, and realize the quantitative analysis of tumor heterogeneity by using only the tumor MRI image without causing trauma to the patients, thereby predicting the survival time of the tumor patients and providing important theoretical basis and application value for tumor precision medicine.

[0087] In this embodiment, first, the tumor risk genes with consistent changes are located according to the tumor copy number variation data and the tumor transcription profile data, then the expression profile of the tumor risk genes is subjected to unsupervised deconvolution analysis to identify the subclone-specific genes related to the expression of the tumor risk genes, then the subclone-specific genes are subjected to biological function analysis and survival analysis to determine the subclone-specific genes that have a specified correlation degree with survival function, then the sample patients are subjected to consistent clustering analysis based on the subclone-specific genes that have a specified correlation degree with survival function, a classification label is obtained, and an optimal tumor prognosis model is constructed and an optimal image genome feature is screened according to the tumor MRI image of the sample patients and the classification label, and finally the external tumor MRI image is analyzed by using the optimal tumor prognosis model and the optimal image genome feature, so as to realize the quantitative analysis of tumor heterogeneity by using only the tumor MRI image without causing trauma to the patients, predict the survival time of the tumor patients, and provide important theoretical basis and application value for tumor precision medicine.

[0088] As shown in FIG. 1, the tumor prognosis model construction module 400 is configured to perform consistent clustering analysis on the sample patients based on the subclone-specific genes that have a specified correlation degree with survival function, obtain a classification label, and construct an optimal tumor prognosis model and screen an optimal image genome feature according to the tumor MRI image of the sample patients and the classification label. Figure 4As shown, in one case of the embodiment, the tumor risk gene positioning module 100 includes:

[0089] The copy number significant variation region identification subunit 101 is configured to identify a copy number significant variation region reaching a specified degree according to tumor copy number variation data.

[0090] The tumor risk gene positioning subunit 102 is configured to position a consistently changed tumor risk gene on the copy number significant variation region in combination with tumor transcription spectrum data.

[0091] In the embodiment, first, tumor copy number variation expression spectrum of TCGA sample patients (taking breast cancer as an example) is obtained from the GDC platform (Genomic Data Commons Data Portal). The data is generated by identifying repeated genomic regions using relevant hardware devices (for example, Affymetrix SNP 6.0 chip) and calculating the copy number of the regions through subsequent analysis. The GDC platform further converts the copy number into segment mean format, equal to log2[(copy number) / 2]. Humans are diploid, the segment mean is 0, the segment mean of the amplified region is positive, and the segment mean of the deleted region is negative. After obtaining the copy number variation expression spectrum in the segment mean format, the GISTIC (Genome Identification of Significant Targets in Cancer) 2.0 module of the GenePattern platform is selected to analyze the tumor copy number variation data. The module simultaneously considers the frequency and intensity of the copy number variation, determines the frequently occurring copy number variation region by the variation significance score, and identifies the copy number significant variation gene in combination with the GRCh38 human reference genome information. The RNA-seq data of the healthy sample and the breast cancer sample is obtained, and the differential expression analysis is performed to identify the genes whose expression is significantly changed after suffering from breast cancer. The copy number significant variation and expression significant change data are integrated, and the genes that are both increased or decreased are screened to position the consistently changed breast cancer risk gene.

[0092] In one case of the embodiment, the biological function analysis and survival analysis module 300 includes:

[0093] The gene enrichment analysis unit is configured to perform gene enrichment analysis on the subclone-specific genes by using a biological pathway database, and mark the corresponding subclone-specific genes with the most significant enriched biological pathway as the biological function. In this embodiment, the pathway analysis is performed on the specific genes of each gene subclone, and the hypergeometric test is used to measure the correlation between the specific genes of each subclone and the biological pathway database (which can be KEGG, Reactome, SMPDB, etc.) to analyze the biological mechanism. The gene set enrichment analysis (GSEA) is performed on the subclone-specific genes by using the biological pathway database. This method first sorts the identified subclone-specific genes according to the specificity, then observes the genes in the multiple pathway gene sets included in the biological pathway database in the ranking of the subclone-specific genes, calculates the enrichment score of each gene set according to the ranking, and finally selects the pathway gene set according to the enrichment score, which is used as the biological pathway enriched by the subclone. The most significant enriched pathway is used to name the subclone, and is used as the biological function of the subclone.

[0094] The survival analysis unit is configured to perform survival analysis on the subclone-specific genes by using the patient survival information, and identify the subclone-specific genes that are correlated with the survival of the patients to a specified degree.

[0095] In this embodiment, the survival analysis is performed on the specific genes of each gene subclone, and the log-rank test is used to measure the correlation between the proportion of different subclone-specific genes and the survival function. The log-rank test is a statistical method used to evaluate the effect of a variable on the survival of patients. For all patients of each subclone-specific gene, the patients are grouped according to the different proportions of the subclone-specific genes in all patients and subjected to log-rank test. To avoid systematic errors caused by multiple random sampling during the test, the Benjamini-Hochberg method is used to correct the p-value representing the significance of the correlation between the variable and the survival obtained by the log-rank test, and the threshold value of the group with the smallest corrected p-value is taken as the final grouping threshold value of the subclone-specific gene. The Kaplan-Meier survival curve is drawn, the hazard ratio (HR) and its 95% confidence interval are calculated, and the significance of the proportion of the subclone-specific genes and the survival of the patients is determined.

[0096] In one case of this embodiment, the tumor prognosis model construction module 400 includes:

[0097] The module construction and feature selection unit is configured to construct a tumor prognosis model and screen an image genomic feature according to the tumor MRI image of the sample patient and the classification label by using a supervised deconvolution algorithm, a machine learning algorithm, and a convolutional neural network algorithm, and select an optimal tumor prognosis model and an optimal image genomic feature from the tumor prognosis model and the image genomic feature according to an evaluation index. The machine learning algorithm includes multiple data screening methods and multiple model structures.

[0098] In this embodiment, the evaluation index can be AUC, decision curve, etc. Specifically, after constructing multiple tumor prognosis models and screening image genomic features, an ROC curve is drawn, and a classification threshold is determined according to the Youden index (sensitivity + specificity - 1). The image genomic feature and the corresponding model with the best classification effect are selected as the final selection according to the AUC, decision curve, etc. After correcting the clinical variables such as age, ER, PR, HER2, and tumor size, a multivariate Cox regression model is used to analyze whether the image genomic feature is an independent prognostic factor for overall survival or recurrence-free survival. The likelihood ratio test is performed on the survival model using the image genomic feature as a variable and the survival model using only the clinical variables to evaluate the improvement of the model prognosis performance of the method.

[0099] In this embodiment, multiple data screening methods such as GBDT, LASSO, RF, XGBoost, etc. can be used, and multiple model structures such as RF, GBDT, Adaboost, LR, NB, SVM, DT, KNM, etc. can be used.

[0100] In the machine learning task, the screening method of the image genomic feature and the structure of the tumor prognosis model directly affect the prediction performance of the model. In order to reduce the influence of the training data and the model structure as much as possible and improve the classification effect of the model on the proportion of patient gene subclone composition, four feature screening methods (GBDT, LASSO, RF, XGBoost) and eight models with different structures (RF, GBDT, Adaboost, LR, NB, SVM, DT, KNM) are selected for training. In the feature screening process, the image features are selected according to the variable importance ranking given by GBDT, RF, and XGBoost, while LASSO can reduce the regression coefficients of irrelevant variables to zero, and the variables with non-zero coefficients are retained, which are used to classify patients with different subclone composition proportions. In terms of model construction, first, the tumor information is numerized according to the size, shape, and boundary definition of the tumor by using the PyRadiomics tool, and the gray matrix, first-order statistics, and second-order statistics are outputted, which reflect the numerical variables of the tumor image information. Then, the machine learning model is constructed in combination with the classification label to classify the patients.

[0101] With the development of computer performance and medical image digitization, many studies have emerged that use neural networks to build models. Compared with traditional machine learning tasks, convolutional neural networks (CNN) can directly understand and select image data, avoiding errors caused by insufficient extraction of image information. This method does not require numerical image data, can directly learn the input image information, and through the setting of the loss function, the model will adjust the weights of each node in multiple iterations to reduce classification errors. To fully extract image genomic features that can reflect tumor heterogeneity, t-test can be used to evaluate the significance of differences in optimal image genomic features in different patient groups, and to depict the internal relationship between optimal image genomic features and tumor heterogeneity genomic features. After training, a heat map is drawn to visualize the attention area of CNN and enhance the model's interpretability.

[0102] In another embodiment, a tumor heterogeneity identification platform is provided, which is constructed based on the optimal tumor prognosis model and the optimal image genomic features obtained from the tumor heterogeneity identification method provided in the above embodiments.

[0103] For different cancers, image genomic features with prognostic efficacy are identified, and then data management is performed based on a MySQL database. Docker is used as a running software carrier to encapsulate all program running and development environments into Docker containers and package them into images. Java language is used to write the front-end website page to build a user-oriented and interface-friendly pan-cancer image genomic feature prognosis analysis platform. Figure 6 The workflow of the platform is shown. The platform mainly includes the following analysis modules: 1) pan-cancer MRI data resource library and expandable other disease database interface; 2) copy number variation driven pan-cancer subclone specific gene query analysis module; 3) copy number variation driven pan-cancer subclone function analysis module; 4) image genomic feature prognosis online real-time analysis module.

[0104] Combined with the drawings Figure 7 In another embodiment, an electronic device is provided, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the tumor heterogeneity identification method provided in the above embodiments is implemented.

[0105] In the embodiment, the data transmission between the memory and the processor is implemented through the communication interface. The memory can include a high-speed RAM memory, and can also include a non-volatile memory such as at least one disk memory. The processor is used to execute the computer program to implement the tumor heterogeneity identification method provided in the above embodiment. If the memory, the processor and the communication interface are independently implemented, the communication interface, the memory and the processor can be connected with each other through a bus and complete the communication therebetween. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 7 In the figure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.

[0106] Optionally, in a specific implementation, if the memory, the processor and the communication interface are integrated on a chip, the memory, the processor and the communication interface can complete the communication therebetween through an internal interface.

[0107] The processor can be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiment.

[0108] In another embodiment, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program is executed by the processor to implement the tumor heterogeneity identification method provided in the above embodiment.

[0109] It should be understood that, although the steps in the flowcharts of the embodiments of the present application are shown in a certain order according to the arrows, the steps are not necessarily executed in the order of the arrows. Unless otherwise specified herein, the execution of the steps is not strictly limited in order, and the steps can be executed in other orders. Moreover, at least some of the steps in the embodiments can include a plurality of sub-steps or a plurality of stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the sub-steps or stages is not necessarily sequential, but can be round-robin or alternately executed with at least some of the other steps or sub-steps or stages of the other steps.

[0110] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the program can be stored in a non-volatile computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of the methods. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0111] The technical features of the above-mentioned embodiments can be combined in any way. In order to make the description concise, all possible combinations of the technical features in the above-mentioned embodiments are not described, but as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.

[0112] The above embodiments only express several implementation manners of the present application, which are described in a more specific and detailed manner, but should not be understood as a limitation on the patent scope of the present application. It should be noted that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

[0113] The above merely describes the preferred embodiments of the present application and should not be used to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A method for identifying tumor heterogeneity, characterized in that, The method comprises the following steps: locating tumor risk genes with consistent changes according to tumor copy number variation data and tumor transcriptome data, wherein the step of locating tumor risk genes with consistent changes according to tumor copy number variation data and tumor transcriptome data comprises the following steps: identifying copy number significantly variable regions reaching a specified degree according to tumor copy number variation data; locating tumor risk genes with consistent changes on the copy number significantly variable regions in combination with tumor transcriptome data; performing unsupervised deconvolution analysis on the expression profile of the tumor risk genes to identify subclonal specific genes related to the expression of the tumor risk genes; performing biological function analysis on the subclonal specific genes to determine subclonal specific genes with a specified correlation degree to survival function, wherein the step of performing biological function analysis on the subclonal specific genes comprises the following steps: performing gene enrichment analysis on the subclonal specific genes through a biological pathway database, and marking the corresponding subclonal specific genes as biological functions with the most significant biological pathways enriched; performing consistent clustering analysis on sample patients based on the subclonal specific genes with the specified correlation degree to survival function to obtain classification labels, and constructing an optimal tumor prognosis model and screening an optimal image genome feature according to tumor MRI images of the sample patients and the classification labels; analyzing external tumor MRI images through the optimal tumor prognosis model and the optimal image genome feature.

2. The tumor heterogeneity identification method of claim 1, wherein, The step of constructing an optimal tumor prognosis model and screening an optimal image genome feature according to tumor MRI images of sample patients and the classification labels comprises the following steps: constructing a tumor prognosis model and screening an image genome feature through a supervised deconvolution algorithm, a machine learning algorithm and a convolutional neural network algorithm according to tumor MRI images of sample patients and the classification labels, and selecting an optimal tumor prognosis model and an optimal image genome feature from tumor prognosis models and image genome features according to evaluation indexes; the machine learning algorithm comprises multiple data screening methods and multiple model structures.

3. A tumor heterogeneity identification apparatus, comprising: The method comprises the following steps: a tumor risk gene locating module for locating tumor risk genes with consistent changes according to tumor copy number variation data and tumor transcriptome data, wherein the step of locating tumor risk genes with consistent changes according to tumor copy number variation data and tumor transcriptome data comprises the following steps: identifying copy number significantly variable regions reaching a specified degree according to tumor copy number variation data; locating tumor risk genes with consistent changes on the copy number significantly variable regions in combination with tumor transcriptome data; a subclonal specific gene identification module for performing unsupervised deconvolution analysis on the expression profile of the tumor risk genes to identify subclonal specific genes related to the expression of the tumor risk genes; a biological function analysis module for performing biological function analysis on the subclonal specific genes to determine subclonal specific genes with a specified correlation degree to survival function, wherein the step of performing biological function analysis on the subclonal specific genes comprises the following steps: performing gene enrichment analysis on the subclonal specific genes through a biological pathway database, and marking the corresponding subclonal specific genes as biological functions with the most significant biological pathways enriched; a biological function analysis module configured to perform biological function analysis on the subclone-specific genes, determine subclone-specific genes that are associated with survival function to a specified degree, and label the subclone-specific genes that are associated with survival function to a specified degree as having biological functions, wherein the step of performing biological function analysis on the subclone-specific genes comprises: performing gene enrichment analysis on the subclone-specific genes by using a biological pathway database, and labeling the subclone-specific genes that are associated with the most significant biological pathways as having biological functions; a tumor prognosis model construction module configured to perform consistency clustering analysis on sample patients based on the subclone-specific genes that are associated with survival function to a specified degree, obtain a classification label, and construct an optimal tumor prognosis model and screen an optimal image genomic feature according to tumor MRI images of the sample patients and the classification label; a tumor image analysis module configured to analyze foreign tumor MRI images by using the optimal tumor prognosis model and the optimal image genomic feature.

4. An electronic device, comprising: The application also provides a computer readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement the tumor heterogeneity identification method according to any one of claims 1 to 2. The computer program is executed by a processor to implement the tumor heterogeneity identification method according to any one of claims 1 to 2.

5. A computer-readable storage medium having stored thereon a computer program, characterized in that, ​

Citation Information

Patent Citations

  • Methods and systems for determining somatic mutation clonality

    CN110770838A

  • Tumor molecule typing method and device, terminal equipment and readable storage medium

    CN112766428A