A consensus molecular subtype classification system for esophageal squamous cell carcinoma based on histopathological images.
The imECMS system addresses limitations in ESCC classification by using deep learning to extract spatial features from histological images, improving reliability and versatility in ESCC subtype identification for clinical applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SHANXI MEDICAL UNIV
- Filing Date
- 2025-07-14
- Publication Date
- 2026-05-01
Smart Images

Figure 2026073928000001_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of medical image processing and computer vision, and specifically relates to an esophageal squamous cell carcinoma consensus molecular subtype classification system based on histopathological images.
Background Art
[0002] Esophageal squamous cell carcinoma (ESCC) is a common but highly heterogeneous malignant tumor, and the treatment response and prognosis vary depending on differences among patients. Currently, clinically, multiple ESCC classification systems have already been proposed, aiming to analyze its heterogeneity and guide the formulation of treatment strategies. Conventionally, the classification of ESCC mainly depends on histological morphological features such as cell morphology and tissue structure, and such classification methods sometimes involve subjectivity and limitations.
[0003] In recent years, with the development of computer vision and deep learning technologies, computer-aided diagnosis (CAD) systems based on histological images have gradually become the focus of research. These systems can automatically extract features from tissue section images and perform classification and analysis of tumors using machine learning algorithms, thereby improving the accuracy and efficiency of diagnosis. However, currently, ESCC subtype classification systems based on histological images still face problems. First, conventional systems sometimes perform classification using only a few features and cannot fully represent the heterogeneity of tumors. Second, these systems lack unified standards and reliable verification methods, causing instability in classification results and low reliability. Therefore, there is a need for a more comprehensive and reliable ESCC classification system based on histological images, which is required to better guide clinical treatment and prognosis evaluation.
Summary of the Invention
[0004] This invention provides a consensus molecular subtype classification system for esophageal squamous cell carcinoma based on histopathological images. The classification system, referred to as imECMS, first uses deep learning to identify eight tissue types from H&E-stained histological images, including background (BACK), connective tissue (CON), squamous epithelium (EPI), glandular tissue (GLA), lymphocytes (LYM), smooth muscle (MUS), cancer-associated stroma (STR), and tumor (TUM). Subsequently, based on the distribution of the eight tissue types in the histopathological images, multiple spatial tissue features (SOFs) are extracted. These features are then used to train a random forest classifier used for ESCC subtype classification, and its accuracy is validated with a large number of individual cues.
[0005] This invention utilizes deep learning to quantify spatial tissue features from automatically labeled H&E-stained section images. It demonstrates a powerful potential for ESCC subtype classification using imECMS with multiple individual cues. This classification, based on histopathological images, is a common data type in clinical trials and can be easily obtained from living tissue specimens. In contrast, multi-omics data requires additional sample processing and experimental procedures and can be affected by factors such as technical variability and experimental conditions. The image classification method is directly applicable to conventional histopathological diagnostic processes and offers greater versatility and reliability.
[0006] The present invention is realized by the following technical solution: a consensus molecular subtype classification system for esophageal squamous cell carcinoma based on histopathological images, wherein the system divides molecular subtypes into four types, denoted as ECMS1, ECMS2, ECMS3, and ECMS4, and the characteristics of the four types are as follows: ECMS1 is characterized by metabolic pathway abnormalities, NFE2L2 activation, and selection of drugs used as NFE2L2 inhibitors; ECMS2 is characterized by upregulation of the classical signaling pathway of the tumor and low methylation; ECMS3 is characterized by fewer copy number change events, low tumor mutation burden, high PD-1 expression, and beneficial treatment with immunosuppressants; and ECMS4 may be characterized by activation of the epithelial-mesenchymal transition pathway.
[0007] Furthermore, the method for obtaining the classification system is as follows: One method for spatial tissue feature extraction utilizes deep learning to automatically extract contours from histological images of esophageal squamous cell carcinoma. Based on these contours, it identifies eight tissue types from H&E-stained histological images, including background (BACK), connective tissue (CON), squamous epithelium (EPI), glandular LA (GLA), lymphocytes (LYM), smooth muscle (MUS), cancer-associated stroma (STR), and tumor (TUM). Subsequently, based on the distribution of these eight tissue types in the histopathological images, it extracts multiple multiscale spatial tissue features (SOFs) including primary, secondary, and higher-order features. Using a random forest algorithm, these SOFs are combined with the extracted spatial tissue features to construct a random forest classifier used for ESCC subtype classification, and its accuracy is validated using a large number of individual cues.
[0008] The specific method for obtaining the classification system is as follows: (1) Spatial tissue feature extraction, which involves the collection and automatic contour extraction of histopathological images of the samples to be classified, utilizes deep learning to identify eight tissue types from H&E-stained histological images and selects two cues, SXM-I and SXM-II. Here, each sample in the SXM-I cue includes the results of automatic contour extraction of tissue types from multiple labeled histopathological sections and the molecular subtype typing results ECMS1-4 of multi-omics data, and is used for building and training a molecular subtype classification model. Each sample in the SXM-II cue includes the results of automatic contour extraction of tissue types from multiple labeled histopathological sections and corresponding clinical survival information, and is used for model validation. In addition, multiple spatial tissue features, including primary, secondary, and higher-order features, are extracted, specifically as follows: A. By quantifying the proportion of each of the eight tissue types within the total tissue, the primary characteristics of the distribution of the eight tissue types are shown. B. Overall SOFs are calculated for each tissue type by calculating the relative proportion of four parts in the overall section image: the intratumor region, the proximal tumor region, the distal tumor region, and the overall section image. First, based on a Gaussian mixture model, the distance from the centroid of the tumor region for various tissue distributions is fitted, and the overall histopathological image is divided into three types of distributions: intratumor region, proximal tumor region, and distal tumor region. Then, the SOFs of the four sub-regions are quantified by the ratio of the proportion of tissue blocks classified into the corresponding category to the total number of tissue blocks in the overall histopathological image. C. The Organizational Variance Index (WCDI) is a measure of regional variance within a certain range. Defined by TIFF2026073928000002.tif17170, where A represents a set of tissue blocks of a particular type, and d(i,j) represents the Euclidean distance between tissue blocks i and j, where i is located within the adjacent region. D is a secondary feature that shows interactions between different types of tissues, such as tumor-matrix and tumor-lymphocytes. For each tumor tissue block, a one-to-one tissue interaction between tumor and non-tumor tissue is calculated, and these interactions are defined as the relative proportion of non-tumor tissue blocks distributed to different tissue types within each ROI. We calculate the intergroup variance index BCDI to evaluate the overall distance between two different types of organizations. Defined by TIFF2026073928000003.tif16170, where A and B represent two different types of tissue, nA and nB are the number of tissue blocks of these two types of tissue, respectively, and d(i,j) represents the Euclidean distance calculated based on the coordinates of the i-th tissue block of tissue A and the j-th tissue block of tissue B. E, a higher-order feature that indicates domain / global diversity. The Spatial Diversity Index (SDI) is an ecological statistic that quantifies the degree of spatial heterogeneity among various cell / tissue types. Calculated using TIFF2026073928000004.tif10170, where m is the number of tissue types and pi is the percentage of the image block of the i-th tissue type. Kullback-Leibler variance (KLD) is used to quantify nuclear heterogeneity determined by the difference from the overall single-cell phenotypic distribution of the patient. Nuclear heterogeneity within an ROI is approximated by calculating the KLD of the average tissue type distribution of all ROIs within the WSI from the distribution of tissue types from the ROI, i.e., the proportion of each tissue type. The Morisita-Horn Similarity Index (MH index) is an ecological measure used to quantify the degree of coexistence between tumors and other types of tissue in WSIs. Spatial correlations are calculated using the Pearson correlation and the Morisita-Horn similarity index, where the number of tumors and other types of tissue in each ROI is... TIFF2026073928000005.tif6170 and TIFF2026073928000006.tif6170 shows the proportion of other types of tissue and tumors in polygon i, respectively. This refers to TIFF2026073928000007.tif12170, Texture features include the gray-level co-occurrence matrix GLCM, the gray-level run-length matrix GLRLM, and the gray-level size region matrix GLSZM, GLCM is used to analyze the gray level spatial relationships of image textures based on statistical methods. The calculation steps are as follows: select one gray level (L), determine the distance (d) and direction (θ) between two pixel points, statistically calculate the difference (Δg) in gray levels and the covariance (C) of gray levels for pixel pairs (i,j) that satisfy all conditions in the image for a given d and θ, construct a GLCM matrix, where the element P(i,j;d,θ) represents the frequency of occurrence of pixel pairs of gray levels i and j at a given distance and direction, and P(i,j;d,θ)=(1 / N)*Σ[Σ[g(i)-g(j)]^2=Δg^2]*[g(i)and g(j)are at distance d and angleθ], where N is the total number of pixel pairs that satisfy the conditions. GLRLM is the length of consecutive gray levels in an image, and its calculation steps are as follows: select a gray level threshold k, statistically calculate the length (l) of consecutive pixel sequences in the image where all gray levels are k or greater, construct a GLRLM matrix, where element P(l;k) represents the frequency of consecutive pixel sequences with a gray level of k or greater length l, and P(l;k) = (1 / N) * Σ [number of runs of length l with gray level >= k], where N is the total number of consecutive sequences satisfying the condition, GLSZM analyzes the gray level distribution of regions of different sizes in an image. The calculation steps are as follows: select a gray level threshold k, statistically calculate the area A of regions in the image where all gray levels are k or greater, construct a GLSZM matrix, where element P(A;k) represents the frequency at which the area A of regions with a gray level of k or greater is equal to the value of k, and P(A;k) = (1 / N) * Σ [number of regions with area A and gray level >= k], where N is the total number of regions that satisfy the condition. Ultimately, the tumor region is scanned using windows of different sizes to obtain region-of-interest (ROIs). ROIs are identified by local tissue areas, containing at least 25% tumor (TUM) blocks and less than 50% background (BACK) blocks. The total SOFs, WCDI, 1:1 tissue interaction, SDI, KLD, Morisita-Horn index, gray-level co-occurrence matrix (GLCM), gray-level run-length matrix (GLRLM), and gray-level size region matrix (GLSZM) are calculated for all identified ROIs, and the average values are taken to obtain overall estimates for the entire tissue section and the patient. (2) A subtype classification based on spatial tissue features SOF, in which the accuracy and reliability of the image classifier are ensured by training and internal validation of the image classifier using H&E sections of esophageal squamous cell carcinoma and multi-omics molecular typing, In an external individual cue containing only H&E sections and clinical information, the relationship between observed clinicopathological features and multi-omics subtypes is examined in the discovery cue. This involves z-score standardization for SOFs, removing features with low variability between samples, and removing scale differences between different features. For SOFs, unsupervised clustering is performed using similarity network fusion SNF, and the median of the sample contour is selected as the cutoff value to identify samples with high representativeness. The performance of the predictive model is evaluated by using an extra-tree classifier to randomly sample 90% of the samples, predict the remaining 10% of the samples, removing intercepts with more than 90% misclassification as outliers, and performing 10x cross-validation on the remaining samples. This involves training a subtype classifier imECMS based on selected core H&E intercepts and evaluating the generalizability of the classifier to other individual cues, and For individuals with multiple intersections, an overall imECMS prediction is determined. Specifically, the overall imECMS prediction for individuals with multiple intersections is determined as a category corresponding to the ECMS subtype that most frequently appeared in the multiple intersections.
[0009] If ECMS4 is present in all of the multiple sections, it is classified as an ECMS4 subtype.
[0010] We trained and validated an ESCC molecular subtype classification model based on random forests, and established an image-based classifier that predicts four types of ECMS subtypes using an extra tree based on SOFs. The extra tree had the following parameter settings: ntree = 100, mtry = 1 feature considered at each node, numrandomcuts = 1 randomly selected feature, nodesize = 2 minimum samples per leaf node. By repeating 10x cross-validation 100 times, the performance of the prediction model reached an AUC of 0.80 or higher for each subtype.
[0011] A computer-readable storage medium is provided, the storage medium stores a computer program, and when the computer program is executed by a computer, the computer performs the esophageal squamous cell carcinoma consensus molecular subtype classification system based on histopathological images as described above.
[0012] In this invention, the spatial tissue feature extraction module includes advanced statistical methods used to quantify the spatial relationships between tissue blocks and feature engineering strategies used to identify interactions and spatial relationships between different tissue types. The random forest algorithm used in the subtype classification module combines 310 feature-selected spatial tissue features, and the accuracy and generalization ability of the classification model are verified through cross-validation and testing on independent datasets. The sample selection process implemented during the modeling process of the random forest algorithm used in the subtype classification module is for identifying core samples in pathological tissue sections from multiple sites.
[0013] Finally, the present invention scans the tumor region using windows of different sizes to obtain regions of interest (ROIs). The ROIs are identified by local tissue regions and include at least 25% tumor (TUM) blocks and 50% or less background (BACK) blocks. The overall SOFs, WCDI, one-to-one tissue interactions, SDI, KLD, Morisita-Horn index, gray-level co-occurrence matrix (GLCM), gray-level run-length matrix (GLRLM), and gray-level size zone matrix (GLSZM) of all identified ROIs are calculated, and the average values are taken as the overall tissue section and the overall estimated value of the patient. Specifically, the present invention uses an ROI-based quantification method at different magnifications (10x, 20x, 40x) to identify a total of 317 SOFs.
Brief Description of the Drawings
[0014] [Figure 1] It is a flowchart of the acquisition step of the esophageal squamous cell carcinoma consensus molecular subtype classification system based on the histopathological images described in the present invention. [Figure 2] It is a flowchart for obtaining a core sample set used for final modeling through the SNF clustering and iterative modeling method. [Figure 3] It is a survival curve diagram for verifying the overall survival period and disease-free survival period of the predicted imECMS subtypes in the SXM-I cohort. [Figure 4] It is a sankey diagram for verifying the correspondence between the predicted imECMS subtypes and the actual ECMS subtypes in the SXM-I cohort. [Figure 5] It is a survival curve diagram for verifying the overall survival period of the predicted imECMS subtypes in the SXM-II cohort.
Modes for Carrying Out the Invention
[0015] To further clarify the objectives, technical solutions, and advantages of the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and comprehensively described below. Of course, the described embodiments are some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative labor are included in the scope of the claims of the present invention.
[0016] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. All the cited references and the materials they cite disclosed herein are incorporated by reference in the form of citation.
[0017] Equivalent technologies of the specific embodiments described that are understandable by those skilled in the art or can be grasped through ordinary experiments are all included in this application.
[0018] Unless otherwise specified, the experimental methods in the following embodiments are all ordinary methods. Unless otherwise specified, the equipment used in the following embodiments are all ordinary laboratory equipment, and the experimental materials used in the following embodiments are all purchased from ordinary stores.
[0019] (Example 1) As shown in FIG. 1, it is a flowchart of the acquisition steps of an esophageal squamous cell carcinoma consensus molecular subtype classification system based on histopathological images according to the present invention, including the following steps.
[0020] In step 101, two cues, SXM-I and SXM-II, are selected from the hospital to collect and auto-contour pathology samples from patients to be classified and to train the ESCC molecular subtype classification model. Here, each patient sample in the SXM-I cue includes the results of auto-contour extraction of tissue types from multiple labeled histopathology sections and the molecular subtype typing results ECMS1-4 from multi-omics data, and is used to build and train the molecular subtype classification model. Each patient sample in the SXM-II cue includes the results of auto-contour extraction of tissue types from multiple labeled histopathology sections and corresponding clinical survival information, and is used to validate the model.
[0021] In step 102, to automatically extract spatial features from the automatic contour extraction results and construct the features of each image, this embodiment extracts multiple spatial organizational features, including primary features, secondary features, and higher-order features, specifically as follows.
[0022] By quantifying the proportion of each of the eight tissue types within the total tissue, we can reveal the primary characteristics of the distribution of these eight tissue types.
[0023] Overall SOFs are calculated for each tissue type by determining the relative proportion of four regions in the overall section image: internal tumor, proximal tumor, distal tumor, and the entire tumor region. First, based on a Gaussian mixture model, the distance from the centroid of the tumor region is fitted for various tissue distributions, dividing the overall histopathological image into three types of distributions: internal tumor, proximal tumor, and distal tumor. Then, the SOFs for the four subregions are quantified by the ratio of the proportion of tissue blocks classified into the corresponding category to the total number of tissue blocks in the overall histopathological image.
[0024] The Intra-Organizational Variance Index (WCDI) is a measure of regional variance within a certain range. Defined by TIFF2026073928000008.tif15170, where A represents a set of tissue blocks of a particular type, and d(i,j) represents the Euclidean distance between tissue blocks i and j, which are located within the adjacent region of i.
[0025] These are secondary features that show interactions between different types of tissues (e.g., tumor-matrix and tumor-lymphocyte interactions).
[0026] For each tumor tissue block, a one-to-one tissue interaction between tumor and non-tumor tissue is calculated, and these interactions are defined as the relative proportion of non-tumor tissue blocks distributed to different tissue types within each ROI.
[0027] We calculate the intergroup variance index (BCDI) to assess the overall distance between two different types of organizations. Defined by TIFF2026073928000009.tif12170, where A and B represent two different types of tissue, nA and nB are the number of tissue blocks in these two types of tissue, and d(i,j) is the Euclidean distance calculated based on the coordinates of the i-th tissue block of tissue A and the j-th tissue block of tissue B.
[0028] This is a higher-order feature that indicates regional / global diversity.
[0029] The Spatial Diversity Index (SDI) is an ecological statistic that quantifies the degree of spatial heterogeneity among various cell / tissue types. Calculated using TIFF2026073928000010.tif9170, where m is the number of tissue types and pi is the percentage of the image block of the i-th tissue type.
[0030] Kullback-Leibler variance (KLD) is used to quantify nuclear heterogeneity determined by the difference from the overall single-cell phenotypic distribution of the patient. Nuclear heterogeneity within an ROI is approximated by calculating the KLD of the average tissue type distribution of all ROIs within the WSI from the distribution of tissue types (proportion of each tissue type) from the ROI.
[0031] The Morisita-Horn Similarity Index (MH index) is an ecological measure used to quantify the degree of coexistence between tumors and other types of tissue in WSIs. Spatial correlations are calculated using the Pearson correlation and the Morisita-Horn similarity index, where the number of tumors and other types of tissue in each ROI is... TIFF2026073928000011.tif6170 and TIFF2026073928000012.tif6170 shows the percentage of other types of tissue and tumors in polygon i, respectively. This shows TIFF2026073928000013.tif12170.
[0032] Texture features include the gray-level co-occurrence matrix (GLCM), the gray-level run-length matrix (GLRLM), and the gray-level size-region matrix (GLSZM).
[0033] GLCM is used to analyze the gray level spatial relationships of image textures based on statistical methods, and its calculation steps are as follows: (1) Select one gray level (L) and determine the distance (d) and direction (θ) between two pixel points; (2) For given d and θ, statistically calculate the difference (Δg) in gray levels and the covariance (C) of gray levels for pixel pairs (i,j) in the image that satisfy all conditions; (3) Construct a GLCM matrix in which the element P(i,j;d,θ) indicates the frequency of occurrence of pixel pairs of gray levels i and j at a given distance and direction, where P(i,j;d,θ)=(1 / N)*Σ[Σ[g(i)-g(j)]^2=Δg^2]*[g(i)and g(j)are at distance d and angleθ], where N is the total number of pixel pairs that satisfy the conditions.
[0034] GLRLM is the length of continuous occurrences of gray levels in an image of interest. The calculation steps are as follows: (1) select a threshold (k) for one gray level; (2) statistically calculate the length (l) of continuous pixel sequences in the image where all gray levels are greater than or equal to k; and (3) construct a GLRLM matrix where element P(l;k) represents the frequency of continuous pixel sequences of length l with gray levels greater than or equal to k, and P(l;k) = (1 / N) * Σ [number of runs of length l with gray level >= k], where N is the total number of continuous sequences that satisfy the condition.
[0035] GLSZM analyzes the gray level distribution of regions of different sizes in an image. The calculation steps are as follows: select a gray level threshold (k), statistically calculate the area (A) of regions in the image where all gray levels are k or greater, construct a GLSZM matrix where element P(A;k) represents the frequency at which the area of a region with a gray level of k or greater is A, and P(A;k) = (1 / N) * Σ [number of regions with area A and gray level >= k], where N is the total number of regions that satisfy the condition.
[0036] This embodiment ultimately scans the tumor region using windows of different sizes to acquire regions of interest (ROIs), which are identified by local tissue regions and include at least 25% tumor (TUM) blocks and less than 50% background (BACK) blocks. The total SOFs, WCDI, one-to-one tissue interaction, SDI, KLD, Morisita-Horn index, gray-level co-occurrence matrix (GLCM), gray-level run-length matrix (GLRLM), and gray-level size region matrix (GLSZM) are calculated for all identified ROIs, and the average values are taken to obtain the overall estimates for the entire tissue section and the patient. Specifically, this embodiment uses ROI-based quantification methods at different magnifications (10x, 20x, 40x) to identify a total of 317 SOFs.
[0037] In step 103, a core sample set to be used for the final model is obtained through SNF clustering and iterative modeling methods. Specifically, to resolve intratumoral heterogeneity, this embodiment designs a process as shown in Figure 2.
[0038] Specifically, the study included 152 patients with multi-omics subtypes. To eliminate the impact of data size, this embodiment performed z-score standardization and filtered out features with seven low inter-sample variations (median absolute deviation, MAD=0). Since the subtype classification of the SXM-I cues was obtained by unsupervised clustering using similarity network fusion (SNF), the mean of the contour coefficients of all patients was used as an indicator of the degree of close grouping of all data in the clustering. To determine the samples with high representativeness of the unsupervised multi-omics subtypes in this embodiment, this embodiment selected the median contour coefficient of the SXM-I samples as the cutoff value. 161 H&E sections from 37 patients were removed, and 465 H&E sections from the remaining 115 patients were used for subsequent analysis. Considering the heterogeneity of multiple histological sections within each individual's tumor, a molecular subtype label is assigned to each section based on the overall multi-omics subtype classification results of the corresponding tumor sample, and then outliers are identified by a random sampling method. Specifically, in this embodiment, an extra-tree classifier is constructed for 80% of randomly selected sections, and 100 iterative predictions are performed on the other 20% of sections. Sections that are misclassified in more than 90% of iterations are removed as outliers. After removing 125 sections, this embodiment evaluates the performance of the predictive model by performing 100 tenfold cross-validation with the retained samples (106 patients, 340 H&E sections). Subsequently, this embodiment trains an extra-tree-based subtype classifier (imECMS) using all 340 H&E sections and evaluates the generalizability of the classifier to other individual cues (SXM-II). For patients with multiple sections, the overall imECMS prediction is determined by the ECMS subtype that most frequently appeared in the sections. Specifically, in this embodiment, since a significant impact of ECMS4 on prognosis was observed, any patient in which ECMS4 appeared in multiple sections is classified as having the ECMS4 subtype.
[0039] In step 104, we train and validate an ESCC molecular subtype classification model based on random forests. Specifically, the performance of the predictive model is as described above. This embodiment establishes an image-based classifier that predicts four types of ECMS subtypes using an extra tree based on SOFs.
[0040] The extra tree in this embodiment has the following parameter settings: ntree (number of trees) is 100, mtry (number of features to consider at each node) is 1, numrandomcuts (number of random cuts for each randomly selected feature) is 1, and nodesize (minimum number of samples per leaf node) is 2.
[0041] These parameters are used to construct an extra-tree model, where each tree is trained using a randomly selected subset of features, and each node randomly selects splitting features and splitting points. The minimum number of samples for a leaf node is 2, meaning that if a node has fewer than 2 samples, further splitting is stopped.
[0042] In this embodiment, by repeating 10x cross-validation 100 times, the performance of the prediction model reaches an AUC of 0.80 or higher for various subtypes (Figure 3).
[0043] In step 105, the automated contour extraction results are input into the target system, and the molecular subtype classification results are automatically obtained.
[0044] In this example, the model's performance was evaluated for all samples in both the SXM-I queue (152 patients, 626 pathological sections) and the SXM-II queue (358 patients, 1262 pathological sections). This example excluded individuals with an even number of pathological sections and no subtype that showed significant advantage in the classification results. Stratification of patients using SOFs-based classification labels was found to be significantly associated with OS and DFS in all SXM-I samples (all P<0.05, Figure 3). The imECMS classification results were in high agreement with ECMS (Figure 4) and were significantly associated with OS in the independent SXM-II queue (P=0.006, Figure 5).
[0045] (Example 2) A storage medium containing a stored computer program, wherein when the computer program is running, the device on which the storage medium is installed causes the sorting method described in any one of the embodiments described above to be executed.
[0046] In this embodiment, the storage medium is a computer-readable storage medium, and the esophageal squamous cell carcinoma pathology image classification system for an endoscope examination quality evaluation device is implemented in the form of a software function unit and may be stored on the computer-readable storage medium when sold or used as an independent product. Based on this understanding, the present invention can implement all or some of the processes in the methods of the embodiments described above, and can also complete these by instructing the relevant hardware with a computer program, the computer program can be stored on the computer-readable storage medium, and when the computer program is executed by a processor, it can implement the steps of the embodiments of each method described above. Here, the computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device, recording medium, USB flash memory, removable hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, electrical signals, and software distribution media that can carry the computer program code.
[0047] Finally, it should be noted that the above embodiments are merely for illustrating, and not limiting, the technical solutions of the present invention. While the present invention will be described in detail with reference to the above embodiments, it is still possible to modify the technical solutions described in the above embodiments, or to make equivalent changes to some or all of their technical features, as will be understandable to those skilled in the art. Such modifications or transformations will ensure that the essence of the corresponding technical solutions falls within the scope of the technical solutions of each embodiment of the present invention.
Claims
1. A consensus molecular subtype classification system for esophageal squamous cell carcinoma based on histopathological images, wherein the system divides molecular subtypes into four types, denoted as ECMS1, ECMS2, ECMS3, and ECMS4, and the characteristics of each of the four types are as follows: In ECMS1, metabolic pathway abnormalities, NFE2L2 activation, and the selection of drugs to be used as NFE2L2 inhibitors are considered. In ECMS2, the classical signaling pathways of the tumor are upregulated, and the degree of methylation is low. ECMS3 has fewer copy number change events, lower tumor mutation burden, high PD-1 expression, and is beneficial for immunosuppressant therapy. A consensus molecular subtype classification system for esophageal squamous cell carcinoma based on histopathological images, characterized by the fact that the epithelial-mesenchymal transition pathway may be activated in ECMS4.
2. The method for obtaining the classification system is as follows: A consensus molecular subtype classification system for esophageal squamous cell carcinoma based on histopathological images, characterized in that it includes a spatial tissue feature extraction method that utilizes deep learning to automatically extract contours from histological images of esophageal squamous cell carcinoma, identifying eight types of tissue from H&E-stained histological images, including background (BACK), connective tissue (CON), squamous epithelium (EPI), glandular LA, lymphocytes (LYM), smooth muscle (MUS), cancer-associated stroma (STR), and tumor (TUM); extracting multiple multiscale spatial tissue features (SOFs) including primary, secondary, and higher-order features based on the distribution of the eight types of tissue in the histopathological image; constructing a random forest classifier used for ESCC subtype classification by combining the extracted spatial tissue features using a random forest algorithm; and verifying its accuracy with a large number of individual cues.
3. The specific method for obtaining the classification system is as follows: (1) Spatial tissue feature extraction, in which deep learning is used for the collection and automatic contour extraction of histopathological images of the samples to be classified, eight types of tissue are identified from H&E stained histological images, and two cues, SXM-I and SXM-II, are selected. Here, each sample in the SXM-I cue includes the results of automatic contour extraction of tissue types from multiple labeled histopathological sections and the molecular subtype typing results ECMS1-4 of multi-omics data, and is used for the construction and training of a molecular subtype classification model. Each sample in the SXM-II cue includes the results of automatic contour extraction of tissue types from multiple labeled histopathological sections and corresponding clinical survival information, and is used for model validation. There are also methods for extracting multiple spatial tissue features, including primary, secondary, and higher-order features, specifically as follows: A. By quantifying the proportion of each of the eight tissue types in the total tissue, the primary characteristics of the distribution of the eight tissue types are shown. B. Overall SOFs are calculated for each tissue type by calculating the relative proportion of four parts in the overall section image: the tumor interior, proximal tumor, distal tumor, and the entire tumor region. First, based on a Gaussian mixture model, the distance from the centroid of the tumor region for various tissue distributions is fitted, and the overall histopathological image is divided into three types of distributions: the tumor interior, proximal tumor, and distal tumor region. Then, the SOFs for the four sub-regions are quantified by the ratio of the proportion of tissue blocks classified into the corresponding category to the total number of tissue blocks in the overall histopathological image. C. The intra-organizational variance index (WCDI) is a measure of regional variance within a certain range. Defined by, where A represents a set of tissue blocks of a particular type, and d(i,j) represents the Euclidean distance between tissue blocks i and j, where i is located within the adjacent region. D, a secondary feature that shows interactions between different types of tissues, such as tumor-matrix and tumor-lymphocytes. For each tumor tissue block, a one-to-one tissue interaction between tumor and non-tumor tissue is calculated, and these interactions are defined as the relative proportion of non-tumor tissue blocks distributed to different tissue types within each ROI. The intergroup variance index (BCDI) is calculated to assess the overall distance between two different types of organizations. Defined by, where A and B represent two different types of tissue, nA and nB are the number of tissue blocks of these two types of tissue, respectively, and d(i,j) represents the Euclidean distance calculated based on the coordinates of the i-th tissue block of tissue A and the j-th tissue block of tissue B. E, higher-order features that indicate domain / global diversity, The Spatial Diversity Index (SDI) is an ecological statistic that quantifies the degree of spatial heterogeneity among various cell / tissue types. It is calculated by, where m is the number of tissue types and pi is the proportion of the image block of the i-th tissue type, Kullback-Leibler variance (KLD) is used to quantify nuclear heterogeneity determined by the difference from the phenotypic distribution of single cells across the entire patient. Nuclear heterogeneity within an ROI is approximated by calculating the KLD of the average tissue type distribution of all ROIs within the WSI from the distribution of tissue types from the ROI, i.e., the proportion of each tissue type. The Morisita-Horn similarity index (M-H index) is an ecological measure used to quantify the degree of coexistence between tumors and other types of tissue in WSIs. The spatial correlation is calculated using the Pearson correlation and the Morisita-Horn similarity index, where the number of tumors and other types of tissue in each ROI is... and These represent the proportion of other types of tissue and tumors in polygon i, respectively. Things that show, Texture features include the gray-level co-occurrence matrix GLCM, the gray-level run-length matrix GLRLM, and the gray-level size region matrix GLSZM, GLCM is used to analyze the gray level spatial relationships of image textures based on statistical methods. The calculation steps are as follows: select one gray level (L), determine the distance (d) and direction (θ) between two pixel points, statistically calculate the difference (Δg) in gray levels and the covariance (C) of gray levels for pixel pairs (i,j) that satisfy all conditions in the image for a given d and θ, construct a GLCM matrix, where element P(i,j;d,θ) indicates the frequency of occurrence of pixel pairs of gray levels i and j at a given distance and direction, and P(i,j;d,θ) = (1 / N) * Σ[Σ[g(i) - g(j)]^2 = Δg^2] * [g(i) and g(j) are at distance d and angle θ], where N is the total number of pixel pairs that satisfy the conditions. GLRLM is the length of continuous gray levels in an image, and its calculation steps are as follows: select a gray level threshold k, statistically calculate the length (l) of continuous pixel sequences in the image where all gray levels are k or greater, construct a GLRLM matrix, where element P(l;k) represents the frequency at which the length of continuous pixel sequences with gray levels k or greater is l, and P(l;k) = (1 / N) * Σ[number of runs of length l with gray level >= k], where N is the total number of continuous sequences satisfying the condition, GLSZM analyzes the gray level distribution of regions of different sizes in an image, and its calculation steps are as follows: select a gray level threshold k, statistically calculate the area A of regions in the image where all gray levels are k or greater, construct a GLSZM matrix, where element P(A;k) represents the frequency at which the area of regions with gray levels k or greater is A, and P(A;k) = (1 / N) * Σ [number of regions with area A and gray level >= k], where N is the total number of regions that satisfy the condition, and Finally, the tumor region is scanned using windows of different sizes to obtain region-of-interest ROIs, which are identified by local tissue regions and include at least 25% tumor (TUM) blocks and less than 50% background (BACK) blocks. The total SOFs, WCDI, 1:1 tissue interaction, SDI, KLD, Morisita-Horn index, gray-level co-occurrence matrix GLCM, gray-level run-length matrix GLRLM, and gray-level size region matrix GLSZM are calculated for all identified ROIs, and the average values are taken to obtain the overall estimates for the entire tissue section and the patient. (2) A subtype classification based on spatial tissue features (SOF), in which the accuracy and reliability of the image classifier are ensured by training and internal validation of the image classifier using H&E sections of esophageal squamous cell carcinoma and multi-omics molecular typing, In an external individual cue containing only H&E sections and clinical information, the relationship between observed clinicopathological features and multi-omics subtypes is examined in the discovery cue. This involves z-score standardization of SOFs to remove features with low inter-sample variability and to remove scale differences between different features. For SOFs, unsupervised clustering is performed using similarity network fusion SNF, and the median of the sample contour is selected as the cutoff value to identify samples with high representativeness. This method evaluates the performance of the predictive model by using an extra-tree classifier to randomly sample 90% of the samples, predict the remaining 10% of the samples, removing intercepts with more than 90% misclassifications as outliers, and performing 10x cross-validation on the remaining samples. This involves training the subtype classifier imECMS based on selected core H&E intercepts and evaluating the generalizability of the classifier to other individual cues, The esophageal squamous cell carcinoma consensus molecular subtype classification system for histopathological images according to claim 2, characterized in that for each individual having multiple sections, an overall imECMS prediction is determined, specifically, the overall imECMS prediction for each individual with multiple sections is determined as a category corresponding to the ECMS subtype that most frequently appeared in the multiple sections, and in cases where the most frequently appearing ECMS subtype could not be determined, the overall imECMS prediction is "undeterminable".
4. The consensus molecular subtype classification system for histopathological images of esophageal squamous cell carcinoma according to claim 3, characterized in that if ECMS4 is present in all of the multiple sections, it is classified as an ECMS4 subtype.
5. The esophageal squamous cell carcinoma consensus molecular subtype classification system for histopathological images according to claim 2, characterized in that the performance of the prediction model reaches an AUC of 0.80 or higher for each subtype by training and validating an ESCC molecular subtype classification model based on random forests and using an extra tree based on SOFs, the extra tree has the following parameter settings: number of trees ntree = 100, number of features considered at each node mtry = 1, number of random cuts of each randomly selected feature numrandomcuts = 1, minimum number of samples per leaf node node = 2, and by repeating 10x cross-validation 100 times, the performance of the prediction model reaches an AUC of 0.80 or higher for each subtype.
6. A computer-readable storage medium, wherein a computer program is stored in the storage medium, and when the computer program is executed by a computer, the computer performs the esophageal squamous cell carcinoma consensus molecular subtype classification system based on histopathological images as described in claim 1.