Intelligent decision-making method and system for T stage of gastric cancer CT image based on deep learning

By employing deep learning methods and utilizing improved residual networks and Grad-CAM technology, gastric cancer CT image features are automatically extracted. Combined with radiomics scores and clinical parameters, dynamic nomograms are generated, which solves the problems of subjectivity and reliance on manual annotation in traditional gastric cancer T staging and achieves rapid and accurate automated decision-making for gastric cancer T staging.

CN121767293APending Publication Date: 2026-03-31CHIMEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional T-staging for gastric cancer relies on subjective interpretation by doctors, which has significant subjectivity and limitations. Manually labeled data acquisition is costly and inefficient, making it difficult to achieve accurate and rapid fully automated T-staging.

Method used

A deep learning-based approach is used to extract features through an improved residual network and Grad-CAM technology. Combined with radiomics scores and clinical parameters, a dynamic nomogram is generated, which outputs gastric cancer T-staging prediction results and a visualized heatmap, enabling automatic localization of tumor regions without manual annotation.

Benefits of technology

It achieves automated, rapid (12 seconds per case) and accurate (95% accuracy in identifying T4 stage) T staging of gastric cancer, and enhances clinical interpretability and physician trust through visualization tools, solving the problems of subjectivity and reliance on manual annotation in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767293A_ABST
    Figure CN121767293A_ABST
Patent Text Reader

Abstract

The invention provides a gastric cancer CT image T stage intelligent decision-making method and system based on deep learning, and the method comprises the steps: obtaining CT image data, carrying out the preprocessing of the CT image data, and obtaining a standardized venous phase CT image; the standardized vein phase CT image is input into an improved residual network for feature extraction, feature maps of different levels are obtained, and a gastric cancer T-stage prediction result is obtained through the feature maps; obtaining a visual thermodynamic diagram; constructing a radiomics score through the regression model, and generating a dynamic column line graph according to the radiomics score in combination with the clinical parameters; and outputting a gastric cancer T stage prediction result, a visual thermodynamic diagram and a dynamic column diagram. The method gets rid of the dependence of manual marking, automatically locates the tumor area and extracts the features directly based on the original CT image, and avoids the deviation caused by the difference of operators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of medical image processing and artificial intelligence, specifically to a deep learning-based intelligent decision-making method and system for T-staging of gastric cancer CT images. Background Technology

[0002] Gastric cancer is one of the most common malignant tumors worldwide, and accurate T staging is crucial for developing treatment plans and assessing prognosis. Traditional gastric cancer T staging relies primarily on the subjective interpretation of CT images by physicians. However, this method is highly subjective and limited, with significant differences in diagnostic accuracy between different physicians. For example, one study showed that the average accuracy rate of physicians in T staging gastric cancer was only 42.4%. Furthermore, manually annotating CT images is time-consuming and labor-intensive, and the accuracy of the annotation is also affected by human factors.

[0003] As deep learning technology is increasingly applied in the medical field, some deep learning-based methods have been attempted for gastric cancer T-staging. However, most of these methods rely on large amounts of manually labeled data, and obtaining high-quality manually labeled data faces problems such as high cost and low efficiency. Currently, there is an urgent need for a technical solution that can eliminate manual labeling and achieve accurate, fast, and fully automated gastric cancer T-staging. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a deep learning-based intelligent decision-making method and system for T-staging of gastric cancer CT images, in order to solve the above-mentioned problems.

[0005] To achieve the above objectives, in a first aspect, the present invention provides a deep learning-based intelligent decision-making method for T-staging of gastric cancer CT images, comprising the following steps: Acquire CT image data and preprocess the CT image data to obtain standardized venous phase CT images; The standardized venous phase CT image is input into an improved residual network. The improved residual network is used to extract features from the standardized venous phase CT image to obtain feature maps at different levels. The gastric cancer T-stage prediction result is obtained through the feature maps. Calculate the gradient of each feature map pair for the predicted class in the last convolutional layer of the residual network to obtain a visual heatmap; Key features are selected from the N-dimensional deep features extracted from the fully connected layer of the residual network using a regression model. Radiomics scores are obtained based on the key features. Dynamic nodal plots are generated based on the radiomics scores and clinical parameters. Output the gastric cancer T-stage prediction results, the visualized heatmap, and the dynamic nomogram.

[0006] Secondly, embodiments of the present invention provide a deep learning-based intelligent decision-making system for T-staging of gastric cancer CT images, comprising the following modules: The data processing module is used to acquire CT image data and preprocess the CT image data to obtain standardized venous phase CT images. An improved network architecture module is used to input the standardized venous phase CT image into an improved residual network, extract features from the standardized venous phase CT image through the improved residual network, obtain feature maps at different levels, and obtain gastric cancer T-staging prediction results through the feature maps; A visualization heatmap generation module is used to calculate the gradient of each feature map pair for the predicted class in the last convolutional layer of the residual network and obtain a visualization heatmap. The dynamic nomogram generation module is used to select key features from the N-dimensional deep features extracted from the fully connected layer of the residual network through a regression model, obtain radiomics scores based on the key features, and generate dynamic nomograms based on the radiomics scores and clinical parameters. The output module is used to output the gastric cancer T-stage prediction results, the visualized heatmap, and the dynamic nomogram.

[0007] Thirdly, embodiments of the present invention provide an electronic device, including: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to implement a semantic intent recognition method that integrates contextual information as described in any of the first aspects.

[0008] Fourthly, embodiments of the present invention provide a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements a semantic intent recognition method that integrates contextual information as described in any of the first aspects.

[0009] The above technical solution has the following beneficial effects: The embodiments of the present invention eliminate the reliance on manual annotation, directly and automatically locate the tumor region and extract features based on the original CT image, avoid the deviation caused by operator differences, and the analysis time for a single case is only 12 seconds, which is dozens of times more efficient than manual diagnosis. This invention overcomes the limitations of single imaging modalities by integrating radiomics features and clinical indicators (such as carcinoembryonic antigen CEA and tumor differentiation degree) through an ordered logistic regression model to construct a dynamic nomogram tool that quantifies the contribution of multi-dimensional parameters to staging and provides more comprehensive predictive information. The embodiments of this invention solve the "black box" problem by using a visual heatmap to intuitively display the tumor area of ​​interest of the model, and combining it with a nomogram to dynamically analyze the decision-making logic, thereby enhancing clinical interpretability and physician trust. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart of an intelligent decision-making method for T-staging of gastric cancer CT images based on deep learning, according to an embodiment of the present invention. Figure 2 yes Figure 1 Flowchart of step S12; Figure 3 yes Figure 1 Flowchart of step S13; Figure 4 This is a schematic diagram of Grad-CAM visualization of tumor features in enhanced CT images of T1-T4 stage gastric cancer according to an embodiment of the present invention. Figure 5 is a schematic diagram of a dynamic nodal plot generated by selecting the top 30 key radiomics features according to an embodiment of the present invention.

[0012] Figure 6 This is a functional block diagram of a deep learning-based intelligent decision-making system for T-staging of gastric cancer CT images, according to an embodiment of the present invention. Figure 7 This is a functional block diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0013] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0014] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0015] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0019] The purpose of this invention is to provide an end-to-end AI vision model to achieve unmanned analysis of T-staging of gastric cancer using CT images. This aims to change the current situation where manual observation of gastric cancer CT images is difficult to accurately assess T-staging, providing a strong basis for preoperative decision-making in gastric cancer. This invention uses CT images of 1792 patients in the venous phase before treatment, selecting the slice containing the largest tumor cross-section for image preprocessing. A multi-classification model is constructed using a modified ResNet152 algorithm and trained using the Adam optimizer. Combining radiomics features and clinical indicators such as Lauren classification, an ordered logistic regression model (Proportiona Odds Model) is constructed, generating a joint prediction nomogram. In addition, a visualization heatmap and a multimodal nomogram tool are generated using Grad-CAM.

[0020] Example 1 like Figure 1 As shown in the figure, an intelligent decision-making method for T-staging of gastric cancer CT images based on deep learning is provided in this embodiment of the invention, which includes the following steps: Step S11: Acquire CT image data and preprocess the CT image data to obtain standardized venous phase CT images.

[0021] This embodiment expands the diversity of the dataset and improves the robustness of the model through standardized preprocessing (linear normalization based on spleen parenchyma CT values ​​and isotropic resampling) and data augmentation operations such as random translation, rotation, and mirror flipping in the data processing flow. This enables the model to adapt to CT images under different poses, angles, and imaging conditions, accurately extract tumor features, and thus ensure that the subsequent Grad-CAM technology can reliably locate the tumor region on different image data without the need for manual annotation and adjustment for different images.

[0022] Step S12: Input the standardized venous phase CT image into the improved residual network, extract features from the standardized venous phase CT image through the improved residual network, obtain feature maps at different levels, and obtain the gastric cancer T-staging prediction results through the feature maps.

[0023] Specifically, this embodiment employs an improved ResNet152 network architecture to receive standardized 224×224 pixel venous phase CT images. Initial dimensionality reduction is achieved using 7×7 convolutional kernels and 3×3 max-pooling layers. Further feature extraction is then performed using 16 residual modules with bottleneck structures. Simultaneously, an adaptive multi-scale feature fusion mechanism integrates shallow texture and deep semantic features to enhance sensitivity to tumor invasion depth. Finally, global average pooling layers and fully connected layers output T1-T4 four-class classification probabilities to predict the T stage of gastric cancer. In this process, the network architecture extracts features at various levels from the image, which serve as the basis for the model's staging judgment.

[0024] In this embodiment, the ResNet152 network architecture processes the input CT image layer by layer. Convolutional and pooling layers extract basic features such as edges and textures from the image, while the residual module learns deeper semantic features while preserving shallow features. These features cover information such as the shape, size, and boundaries of gastric tissue and tumors, forming the basis for the model's subsequent analysis and judgment.

[0025] This embodiment also introduces an adaptive multi-scale feature fusion mechanism into the network, which integrates shallow texture and deep semantic features extracted from different levels. Shallow features capture detailed information of the image, while deep features focus on more abstract semantic concepts. The fusion of the two enhances the model's sensitivity to tumor invasion depth, obtaining a more comprehensive, accurate image feature representation that is highly correlated with the T stage of gastric cancer. After processing by global average pooling layers and fully connected layers, the network maps the extracted and integrated features to T1-T4 four-class probabilities. That is, through the analysis and processing of the input CT image, the probability of the gastric cancer corresponding to the image being in the T1-T4 stage is finally obtained, thereby realizing fully automated intelligent decision-making for the T stage of gastric cancer and providing quantitative basis for the formulation of clinical diagnosis and treatment plans.

[0026] Step S13: Calculate the gradient of each feature map pair in the last convolutional layer of the ResNet152 residual network to predict the class, and obtain a visual heatmap.

[0027] Specifically, this embodiment employs Grad-CAM technology to calculate the gradient of each feature map in the last convolutional layer of the ResNet152 residual network for the predicted class. Grad-CAM (Gradient-weighted Class Activation Mapping) is a technique for visualizing the decision-making process of convolutional neural networks. By backpropagating the gradient information of the last convolutional layer, it calculates the importance of the feature map of the last convolutional layer to the prediction result. In the gastric cancer T-staging model, after the model analyzes CT images and outputs the T1-T4 four-class probabilities, Grad-CAM starts from the category corresponding to the prediction result and backpropagates the gradient of each feature map in the last convolutional layer. By weighted summing these gradients with the feature maps, the contribution of each pixel to the prediction result can be obtained (the contribution represents the importance of each pixel to the model's final T-staging prediction result. If the contribution of a pixel is high, it means that the image region corresponding to that pixel plays a key role in the model's determination of which stage of gastric cancer it is in (T1-T4), and a heatmap is generated. On a heatmap, areas with darker colors and higher heat intensity are the regions that the model focuses on when making staging decisions—that is, the areas where the tumor is located—thus achieving automatic tumor localization. This localization method does not rely on manual marking of tumor locations on the image; it is entirely based on the model's own calculations and analysis. Doctors can quickly understand the key image areas that the model focuses on when predicting T-staging of gastric cancer by observing the heatmap, helping to determine the location and potential extent of tumor invasion.

[0028] In this embodiment, after the ResNet152 network outputs its prediction results, the Grad-CAM technology starts from the last convolutional layer of the network and backpropagates the gradient information corresponding to the predicted category. By calculating the weighted sum of the gradient and the feature map, the importance score of each pixel to the prediction result is obtained, thereby generating a Grad-CAM heatmap. Since the ResNet152 network learns tumor-related feature patterns during training, the regions it focuses on when making staging judgments (i.e., tumor regions) will show high heat on the heatmap. Therefore, Grad-CAM can automatically locate tumor regions without manual annotation based on the features extracted by the ResNet152 network, intuitively displaying the tumor regions of interest to the network.

[0029] The generated Grad-CAM heatmap, together with the staging results output by the ResNet152 network architecture, provides information for doctors. The staging results are presented as probability values ​​for stages T1-T4, providing a quantitative diagnostic reference; the heatmap shows the image regions on which the model made staging predictions, helping doctors understand the model's decision-making process, solving the "black box" problem of deep learning models, enabling doctors to have more confidence in the model's predictions, and also helping to further optimize the model.

[0030] In this embodiment, the ResNet152 network uses weighted fusion of shallow texture and deep semantic features. Grad-CAM can be applied to fusion features at different scales to generate multi-scale heatmaps. The final tumor localization result can integrate heatmap information at multiple scales to improve localization accuracy.

[0031] Step S14: Select key features from the N-dimensional deep features extracted from the fully connected layer of the residual network using a regression model, obtain radiomics scores based on the key features, and generate a dynamic nomogram based on the radiomics scores and clinical parameters.

[0032] In this embodiment, by fusing radiomics features with clinical parameters (such as CEA and Lauren classification), 2048-dimensional deep features extracted from the ResNet152 fully connected layer are used to screen key features via LASSO (Least Absolute Shrinkage and Selection Operator) regression to construct a radiomics score (Radscore). An ordered logistic regression model is then used to quantify the contributions of multi-dimensional parameters, generating a dynamic nomogram tool to intuitively analyze the weight of each variable on the staging. The radiomics score, through LASSO regression, compresses the 2048-dimensional features into a few key features, solving the feature redundancy problem and avoiding model overfitting. The ordered logistic regression then builds a predictive model based on these condensed features, improving the model's stability and interpretability.

[0033] Specifically, a 2048-dimensional feature vector X=[x1,x2,……x] is extracted from the output of the global average pooling layer before the fully connected layer of ResNet152. 2048 These features include information such as tumor morphology, texture, and spatial relationships. Clinical indicators such as CEA levels and Lauren classification are used as supplementary features (Y=y1,y2,...y...). n This forms a multimodal feature set F=[X;Y]. LASSO (Minimum Absolute Shrinkage and Selection Operator) regression is applied for feature dimensionality reduction, with the optimization objective being: ; in, For the true T-stage of sample i, is the j-th eigenvalue of sample i, is the coefficient of feature j, is the penalty coefficient. By adjusting , key features strongly correlated with the T stage of gastric cancer are screened out, redundant and irrelevant features are removed, and generally about 30 - 50 key features are finally screened out.

[0034] The screened features are standardized, the LASSO coefficient weights of each feature are calculated, and a radiomics score (Radscore) is constructed: ; where is the standardized key feature value. By multiplying each key feature by its corresponding weight and summing them up, the final radiomics score is obtained, which quantifies the comprehensive contribution of imaging features and clinical parameters to the T stage.

[0035] Taking the T stage of gastric cancer (ordinal categorical variable, T1 < T2 < T3 < T4) as the dependent variable and the radiomics score (Radscore) and clinical parameters as independent variables, an ordinal logistic regression model is constructed: ; where, is the threshold parameter, is the regression coefficient, and the regression coefficient represents the impact of a one-unit change in feature on the T stage. Through model training, the regression coefficient values of each parameter are calculated and converted into odds ratios (OddsRatio) to visually show the contribution intensity of each parameter to staging. A dynamic nomogram is generated to visualize the weights of each parameter and assist doctors in understanding the model's decision-making logic.

[0036] Step S15, output the prediction result of the T stage of gastric cancer, the visualized heat map, and the dynamic nomogram.

[0037] In this embodiment, DICOM format input (DICOM, Digital Imaging and Communications in Medicine) is supported, the T stage result of gastric cancer (probability values for stages T1-T4), the Grad-CAM heat map, and the dynamic nomogram are output, and the model can also be integrated into a Picture Archiving and Communication System (PACS system) for convenient clinical application.

[0038] The intelligent decision-making method in this embodiment can completely abandon manual annotation, achieve end-to-end analysis, with a single-case time consumption of 12 seconds and a recognition accuracy of 95% for stage T4; the dynamic nomogram and heat map solve the black-box problem.

[0039] In some embodiments, acquiring CT image data and preprocessing the CT image data specifically includes: Extract the region of interest (RIO) from the CT image data, calculate the CT value of each voxel within the ROI, and map all CT values ​​to the [0,1] interval; CT scanners scan the human body with X-rays, generating raw image data in DICOM format. This data is essentially a three-dimensional matrix composed of countless voxels. Each voxel contains its spatial coordinates (x, y, z) and corresponding CT value (reflecting tissue density). In spleen CT image analysis, the Region of Interest (ROI) of the spleen must first be extracted from the raw CT data using manual or automatic segmentation methods (such as thresholding or deep learning segmentation). All voxels within the ROI then constitute the three-dimensional image data of the spleen. Each voxel corresponds to a specific three-dimensional spatial location within the spleen, and its stored CT value represents the degree of X-ray attenuation at that location. For example, the CT value of spleen parenchyma is typically between 30-60 HU (Henness units), while the CT value of adipose tissue is approximately -100 HU, and bone can reach +1000 HU. Therefore, voxels quantify the density characteristics of spleen parenchyma through their CT values.

[0040] If the CT values ​​of the spleen vary greatly among different cases (e.g., the CT value of the spleen in one case is 20-50 HU, while in another case it is 40-70 HU), unnormalized CT values ​​will cause the model to be biased towards high-value cases during learning, affecting generalization ability. In this embodiment, the original pixel values ​​(CT values) of the CT images are standardized. The CT value reflects the degree of attenuation of X-rays by tissue. The CT value of the spleen parenchyma is usually within a specific range (e.g., the CT value of a normal spleen on plain scan is about 30-60 HU), but different equipment and scanning parameters will lead to differences in the distribution of CT values. By using linear transformation (e.g., min-max normalization, Z-score normalization), the CT values ​​of the spleen parenchyma are mapped to a uniform range (e.g., [0,1] or a mean of 0 and a standard deviation of 1), eliminating grayscale bias between equipment.

[0041] Resampling of CT image data involves resampling the original voxels in the CT image data into cubes of uniform size in all directions. In this embodiment, the spatial resolution of the original CT images is standardized. The voxel size of the original CT images is inconsistent along the Z-axis (slice thickness) and XY-axis (plane) (e.g., the original voxel is 0.5×0.5×3mm³), resulting in uneven resolution in three-dimensional space. Through algorithms such as cubic spline interpolation, the original voxels are resampled into cubes (1×1×1mm³) with consistent dimensions in all directions, so that the images have the same spatial sampling density in three-dimensional space.

[0042] The same window width and window level are applied to the CT image data, and the CT image data is enhanced by random translation, rotation and mirror flipping to obtain standardized venous phase CT images.

[0043] Specifically, in this embodiment, the image preprocessing stage first standardizes and reconstructs the original DICOM data, including linear normalization of the spleen parenchyma CT values ​​(50±5 HU in plain scan) to eliminate differences between devices. Subsequently, cubic spline interpolation is used to resample the voxels to isotropic resolution (1x1x1mm). To optimize tumor boundary display, all images are uniformly set with a window width of 350 HU and a window level of 50 HU, and data enhancement is performed through random translation (±5%), rotation (±10%), and mirror flipping.

[0044] In this context, window width refers to the range of CT values ​​displayed in a CT image, reflecting the span from the minimum to the maximum CT value in the image. A larger window width can display a wider range of CT values, allowing simultaneous observation of various tissues with significant density differences, but image contrast will decrease; a smaller window width focuses on a narrower range of CT values, highlighting details of specific tissues and resulting in higher image contrast. For example, in this invention, the window width is set to 350 HU to map a certain range of CT values ​​onto the grayscale display of the image, ensuring appropriate grayscale differences between gastric tissue and tumors in the image, facilitating subsequent analysis. Window level is the center CT value within the window width range, determining the center position of the image's grayscale display. By adjusting the window level, the grayscale display of various tissues in the image can be altered. When the window level is set to a certain value, the corresponding CT value will be displayed in medium grayscale in the image; tissues with CT values ​​higher than this value will be displayed in brighter grayscale, and tissues with CT values ​​lower than this value will be displayed in darker grayscale. In this embodiment of the invention, the window level is set to 50HU, which means that 50HU is used as the center of grayscale display, so that the gastric tissue and tumor can be clearly distinguished in the image.

[0045] In this invention, setting the window width to 350 HU and the window level to 50 HU was determined through extensive experiments and research. This maximizes the display of gastric tissue and tumor features, allowing the improved ResNet152 network architecture to better extract effective information from the image, thereby improving the accuracy of gastric cancer T staging.

[0046] like Figure 2 As shown, in some embodiments, in step S12, the standardized venous phase CT image is input into an improved residual network. The improved residual network extracts features from the standardized venous phase CT image to obtain feature maps at different levels. The gastric cancer T-staging prediction result is obtained through the feature maps. Specifically, the steps include the following: S121 receives standardized venous phase CT images through the input layer; This embodiment targets the T-staging task for gastric cancer. It specifies that the input layer receives standardized 224×224 pixel venous phase CT images and sets a specific window width of 350HU and a window level of 50HU to highlight gastric tissue and tumor features, making the image data input to the network more conducive to extracting features related to gastric cancer.

[0047] S122: Preliminary feature extraction is performed on the standardized venous phase CT image through convolution kernel to obtain a preliminary feature map, and the resolution of the preliminary feature map is reduced by the max pooling layer to obtain shallow texture features. In this embodiment, a 7×7 convolution kernel is used for preliminary feature extraction. Compared with a small-sized convolution kernel, a large-sized convolution kernel can capture the overall structural information of the image in a wider field of view. A 3×3 max pooling layer is used to reduce the resolution of the feature map, reduce the amount of subsequent computation, and at the same time retain the key features of the image to prepare for subsequent processing.

[0048] S123, deep semantic features are extracted from the preliminary feature map through residual modules; each residual module includes a convolutional dimensionality reduction layer, a spatial convolutional layer, and a convolutional dimensionality increase layer, and batch normalization and ReLU activation functions are inserted after the convolutional dimensionality reduction layer, the spatial convolutional layer, and the convolutional dimensionality increase layer, respectively.

[0049] In this embodiment, 16 residual modules are connected, and a bottleneck structure is adopted for each residual module. That is, dimensionality reduction is first performed through 1×1 convolution to reduce the amount of computation and the number of parameters, thereby improving computational efficiency; then, 3×3 spatial convolution is performed to extract spatial features; finally, dimensionality is increased through 1×1 convolution to restore the number of channels. Batch normalization and ReLU activation function are inserted in each layer. Batch normalization can alleviate the gradient vanishing problem and accelerate model convergence, while the ReLU activation function increases the nonlinear expressive power of the network.

[0050] S124 integrates shallow texture features and deep semantic features to obtain a fused feature map.

[0051] In this embodiment, a fused feature map is obtained by weightedly fusing and integrating shallow texture features and deep semantic features through an adaptive multi-scale feature fusion mechanism. This fused feature map contains comprehensive information about the image at different scales and levels, which is of great significance for predicting the T stage of gastric cancer.

[0052] Superficial texture features preserve information such as the edges and texture details of the tumor and surrounding tissues in CT images, while deep semantic features reflect higher-level information such as the degree of tumor invasion and its relationship with surrounding tissues. The integrated feature map organically combines these two types of information, containing both image details and an understanding of image semantics, and can more comprehensively express the characteristics of gastric cancer in CT images. For example, in the T staging of gastric cancer, superficial features may show the rough texture of the tumor surface, while deep features indicate whether the tumor has broken through the stomach wall; the fused feature map will simultaneously present these key pieces of information.

[0053] The fused feature map, used as input to subsequent network layers, enhances the model's sensitivity to tumor invasion depth. Compared to using shallow or deep features alone, the fused features provide the model with richer and more discriminative information, thereby improving the accuracy of gastric cancer T-stage prediction. The patent mentions that the model's internal test set AUC reaches 0.92, and its external multi-center validation AUC is ≥0.86. The fused feature map obtained through the adaptive multi-scale feature fusion mechanism plays a crucial role in helping the model better learn and distinguish gastric cancer features at different T stages.

[0054] When the model predicts the T stage of gastric cancer, the fused feature map provides higher-quality input to subsequent decision layers such as the fully connected layer. Based on this, the model's output of the T1-T4 four-class classification probabilities will be more accurate and reliable. Doctors can obtain more valuable references when making diagnoses based on these probabilities, thereby developing more appropriate treatment plans.

[0055] This embodiment introduces an adaptive multi-scale feature fusion mechanism, integrating shallow texture features and deep semantic features through weighted fusion after different convolutional layers. Shallow features contain rich detailed information, while deep features have stronger semantic understanding capabilities. The fusion of the two can enhance the model's sensitivity to tumor invasion depth and improve the accuracy of gastric cancer T-staging.

[0056] S125 maps the fused feature map to the gastric cancer T-stage prediction result through a global average pooling layer and a fully connected layer.

[0057] In this embodiment, the feature map is converted into a fixed-length vector through a global average pooling layer to avoid the risk of overfitting caused by the large number of parameters in the fully connected layer. Then, the T1-T4 four-class classification probabilities are output through the fully connected layer to achieve the prediction output of the T stage of gastric cancer.

[0058] In some embodiments, the fused feature map is mapped to a T-stage prediction result through a global average pooling layer and a fully connected layer, specifically including: The global average pooling layer calculates the average value of all pixels in each fused feature map, transforms the fused feature map into a vector of a specific length, and outputs it to the fully connected layer. Each neuron in the fully connected layer is connected to all elements of a vector of a specific length. Through weight matrix multiplication and bias addition, a linear transformation is performed on the vector of the specific length to establish the mapping relationship between the vector of the specific length and each category of T stage. The score of each element in the vector of the specific length corresponding to the T stage category is calculated. The T-stage scores are converted into probability values ​​using the Softmax activation function to obtain the T-stage prediction results for gastric cancer.

[0059] Specifically, after passing through 16 residual modules and an adaptive multi-scale feature fusion mechanism, the network obtains feature maps with different numbers of channels and spatial dimensions. The global average pooling layer operates on each feature map, performing average pooling on its spatial dimensions (height and width), i.e., calculating the average value of all pixels in each feature map. Assuming the feature map size is H×W×C (H is height, W is width, and C is the number of channels), after global average pooling, the output becomes 1×1×C, compressing each feature map into a single value, ultimately resulting in a vector of length C. In this process, the global average pooling layer effectively preserves the global information of the feature maps while reducing data dimensionality, decreasing the computational load of subsequent fully connected layers, and avoiding overfitting. For example, for gastric cancer CT images, the feature maps extracted by the preceding network layers, containing various information such as tumor texture and shape, are integrated by the global average pooling layer into a vector containing comprehensive feature information. This vector represents the average feature response of the entire image across different channels.

[0060] The vector output by the global average pooling layer serves as the input to the fully connected layer. Each neuron in the fully connected layer is connected to all elements of the input vector, and the input features are linearly transformed through a series of weight matrix multiplications and bias additions. In this model, the number of output nodes in the fully connected layer is set to 4, corresponding to the four stage categories T1-T4. Assuming the weight matrix of the fully connected layer is W with a dimension of 4×C, the input vector is x with a dimension of C×1, and the bias vector is b with a dimension of 4×1, then the output of the fully connected layer is y = Wx + b, resulting in a vector with a dimension of 4×1. Each element in this vector represents the score of the input CT image belonging to the corresponding T stage category. To convert these scores into probability values, a Softmax activation function is typically applied after the output of the fully connected layer. The Softmax activation function performs an exponential operation on each element and then normalizes the result so that the sum of all elements is 1, thus obtaining the probability value of the input CT image belonging to stage T1, T2, T3, or T4. For example, after the output vector is processed by the Softmax activation function, it may obtain [0.1, 0.3, 0.4, 0.2], which represent the probabilities of the CT image belonging to the T1, T2, T3, and T4 stages, respectively.

[0061] This embodiment utilizes the collaborative work of global average pooling layers and fully connected layers. The improved ResNet152 network architecture can effectively integrate and classify the complex features extracted from the previous layers, and finally output the predicted probability of gastric cancer T stage corresponding to CT images, providing doctors with quantitative diagnostic reference.

[0062] like Figure 3 As shown, in some embodiments, step S13 involves calculating the gradient of the predicted class for each feature map pair in the last convolutional layer of the residual network and obtaining a visual heatmap, specifically including: S131: Obtain the predicted class score from the score of the T stage category, backpropagate from the predicted class score, calculate the gradient of the predicted class score with respect to the feature map of the last convolutional layer, and perform global average pooling on the gradient in the spatial dimension to obtain the importance weight of each channel.

[0063] Specifically, after the model completes forward propagation, the output score y of the fully connected layer corresponding to the predicted category (e.g., T3 period) is obtained. c , where c represents the target category (a phase from T1 to T4). This score is the model's unnormalized prediction that the input CT image belongs to category c. From the target category score y... c Initially, backpropagation is performed, but during the propagation process, all gradients except for the last convolutional layer are set to zero. This allows us to focus on the impact of the last convolutional layer on the target class. The resulting feature map A of the last convolutional layer is then obtained. k The dimensions are H×W×K, where K is the number of channels with respect to y.c gradient This gradient represents the degree to which each feature map unit contributes to the target class score.

[0064] gradient Global average pooling is performed on the spatial dimension H×W to obtain the importance weight of each channel. : ; Where Z = H × W, This represents the importance of the k-th feature map to category c. A larger weight indicates a greater contribution of that feature map to the prediction result. All feature maps are weighted and fused according to their importance to the target category, highlighting regions that play a crucial role in the prediction.

[0065] S132, the feature map is weighted and summed by importance weights to obtain the initial version of the heatmap, and the ReLU activation function is applied to retain the positive contribution region of the initial version of the heatmap to obtain the Grad-CAM heatmap; Specifically, the importance weights With the corresponding feature map A k Multiply the results and sum them over all channels to obtain the initial version of the heatmap. : ; right Applying the ReLU activation function: ReLU ensures that only regions that positively contribute to the prediction are retained, while regions that negatively contribute (i.e., regions that may mislead the prediction) are suppressed. This step ensures that the heatmap only displays the meaningful regions that the model is interested in.

[0066] S133 upsamples the Grad-CAM heatmap to the size of the original CT image, converts the upsampled heatmap into a pseudo-color heatmap, and overlays the pseudo-color heatmap with the original CT image to generate a visualized heatmap.

[0067] In this embodiment, The heatmap is upsampled to the original CT image size (e.g., 224×224) using methods such as bilinear interpolation to align with the original image. The upsampled heatmap is then converted to a pseudo-color image using a color mapping scheme such as Jet. For example, red can represent areas of high interest from the model's attention mechanism (areas deemed important by the model), while blue can represent areas of low importance. The pseudo-color heatmap is then overlaid on the original CT image to generate the final visualization. The transparency can be adjusted during overlay to allow doctors to see both the details of the original image and the areas of interest from the model. Grad-CAM heatmaps are used to visualize the areas of high interest from the model, helping doctors quickly locate key lesion sites. Figure 4 As shown, representative enhanced CT images of T1-T4 stage gastric cancer are presented (the large image on the left in each group). Within the bounded areas of the original images, radiologists manually annotated the areas of focus (the area indicated by the X in the upper right corner of each image) and the areas focused on by the model's attention mechanism (the area indicated by the Y in the lower right corner of each image). It can be seen that the areas of focus for radiologists and the areas of focus for the model's attention mechanism have a high degree of overlap. This indicates that in unlabeled images, the model has a very strong ability to identify tumor lesions, and the areas of focus are largely consistent with those of human experts.

[0068] In this embodiment, through heatmaps, doctors can intuitively see the areas that the model considers most relevant to T-staging, helping to confirm the location and extent of tumor invasion. Heatmaps can also be used to determine whether the model is focusing on real lesions (such as tumors) rather than misjudging normal tissue. Heatmaps can indicate key areas such as tumor boundaries and metastases, providing a reference for subsequent quantitative analysis (such as tumor volume calculation). Heatmaps provide a visual interpretation of the model's prediction results, enhancing doctors' trust in the AI ​​system. Furthermore, the repeatedly appearing high-interest areas in the heatmap correspond to key imaging features related to gastric cancer progression, which is helpful for medical research.

[0069] In some embodiments, constructing a radiomics score based on key features specifically includes: The selected key features are standardized to obtain key features of the same scale; the weight of each key feature is determined according to the regression model; and the sum of the products of each key feature and its corresponding weight is obtained to obtain the radiomics score.

[0070] Specifically, the selected key features are standardized to eliminate differences in units and numerical ranges, ensuring all features are on the same scale for easier subsequent weighted calculations. Z-score standardization can be used as a standardization method. Based on the coefficients of each key feature obtained from the LASSO regression model, the weight of each feature is determined. The final radiomics score is obtained by multiplying each key feature by its corresponding weight and then summing the results.

[0071] In this embodiment, gastric cancer CT images contain a large amount of texture, morphological, and other information. Constructing a radiomics score allows for the selection of key features related to the T stage of gastric cancer from the 2248-dimensional depth features extracted before the ResNet152 fully connected layer using LASSO regression. These features are then weighted and combined to quantify the contribution of each image feature to the T stage. In this way, the originally complex image feature information is transformed into a comprehensive numerical value, clearly demonstrating the importance of different image features in staging.

[0072] The radiomics scoring system integrates radiomics features with clinical parameters and combines them with an ordered logistic regression model, comprehensively considering multiple factors influencing gastric cancer T-staging. Compared to relying on a single feature or a simple model, it can more comprehensively and accurately reflect the actual stage of gastric cancer, effectively improving the model's accuracy in predicting gastric cancer T-staging. The model achieved an AUC (Area Under the Curve of ROC) of 0.92 on the internal test set and an AUC ≥ 0.86 in external multicenter validation. The radiomics scoring played a crucial role in this achievement, significantly outperforming the average accuracy of 42.4% of traditional physician subjective interpretation.

[0073] Radiomics scoring visually demonstrates the impact of various variables (image features and clinical parameters) on staging in numerical form, making the previously complex decision-making process of deep learning models interpretable. Doctors can use this score to clearly understand the key factors the model relies on for staging, as well as the relative importance of these factors, effectively solving the "black box" problem of deep learning models and enhancing doctors' trust in the model's results. In clinical practice, radiomics scoring provides doctors with quantitative diagnostic references, helping them to more scientifically assess patients' conditions and develop more appropriate treatment plans. Interactive nomogram HTML (Hypertext Markup Language) reports visually display the weight of each variable on staging, allowing doctors to quickly understand the impact of different factors on staging, providing strong support for surgical plan selection and auxiliary treatment decisions. Experiments show that the decision curve shows the greatest net benefit within the 0.3-0.7 threshold range, with a relative improvement of 40%, fully demonstrating the important value of radiomics scoring in clinical decision-making.

[0074] In some embodiments, generating a dynamic nomogram based on radiomics scores and clinical parameters specifically includes: The radiomics scores and clinical parameters are standardized and normalized to obtain a standardized dataset.

[0075] In this embodiment, the 2048-dimensional depth features extracted from the ResNet152 fully connected layer are used to construct a radiomics score by selecting key features through LASSO regression. This score is then integrated with clinical parameters (such as CEA and Lauren classification) to form a complete dataset. These data are standardized and normalized to ensure they are on the same scale, facilitating subsequent model computation.

[0076] Using the standardized dataset as the independent variable and the gastric cancer T-stage prediction results as the dependent variable, an ordered logistic regression model is constructed. The regression coefficient of each independent variable is calculated using the ordered logistic regression model, and the odds ratio of each independent variable is calculated based on the regression coefficient.

[0077] Specifically, processed radiomics scores and clinical parameters are used as independent variables, and gastric cancer T stage is used as the dependent variable (ordinal categorical variable, T1-T4 stages), constructing an ordinal logistic regression model. Through model training, regression coefficients for each independent variable (i.e., each radiomics feature and clinical parameter) are calculated. These regression coefficients reflect the direction and extent of each variable's influence on gastric cancer T stage; the larger the absolute value of the coefficient, the greater the influence of that variable on the stage; a positive coefficient indicates that an increase in the variable value tends to lead to an increase in stage, and vice versa. Based on the regression coefficients, the odds ratio for each variable is further calculated. The odds ratio more intuitively shows the multiple by which changes in each variable affect the probability of stage progression. For example, an odds ratio of 2 for a variable means that for every unit change in that variable, the probability of a stage progression increases by one level doubles.

[0078] In this embodiment, gastric cancer T staging is influenced by multiple factors, including radiomics features (such as tumor texture and morphology) and clinical parameters (such as CEA and Lauren classification). An ordered logistic regression model can incorporate these multidimensional parameters into a unified framework and analyze their impact on gastric cancer T staging. By quantifying the contribution of each parameter, the relative importance of different factors in the staging process is clarified, avoiding the limitations of relying solely on a single factor to determine staging, and achieving a more comprehensive and accurate assessment of gastric cancer T staging.

[0079] In clinical practice, doctors need clear and quantifiable reference information to formulate treatment plans and assess prognosis. The contribution weights of each parameter output by an ordered logistic regression model provide doctors with intuitive and quantifiable decision-making basis. Based on these weights, doctors can determine which factors have a greater impact on patient staging, thus enabling them to develop more scientifically personalized treatment plans. For example, if a certain radiomics feature has a high weight, it indicates that this feature is crucial in staging, and doctors will pay more attention to related imaging manifestations during diagnosis.

[0080] Deep learning models are often considered black boxes, with their decision-making processes difficult to understand. Combining ordered logistic regression models with deep learning, and quantifying the contribution of each parameter, clearly demonstrates the role of different input parameters in predicting T-staging of gastric cancer, making the model's decision-making logic transparent. This helps physicians understand the model's judgment criteria, increases their confidence in the model's predictions, and promotes the model's application and widespread adoption in clinical practice.

[0081] Quantifying the contributions of multidimensional parameters can help analyze the effectiveness of each parameter in the model. For parameters with small or irrelevant contributions, further research can be conducted to determine if they can be removed, thereby simplifying the model structure, reducing computational load, and improving model efficiency. For parameters with large contributions, further information mining or other methods can be considered to enhance their effects, thereby optimizing model performance and improving the accuracy of gastric cancer T staging.

[0082] This embodiment combines ordered logistic regression model to quantify the contribution of multi-dimensional parameters, which plays an important role in gastric cancer T staging model, such as integrating information, assisting decision-making, and improving model performance. It is a key link in achieving accurate staging and clinical application value.

[0083] Based on the odds ratio and regression coefficient, the scale range and position of each independent variable in the dynamic nomogram are determined, and a line segment corresponding to the gastric cancer T-staging prediction result is created for each independent variable. The mapping relationship between the scale value of each independent variable in the dynamic nomogram and the line segment corresponding to the gastric cancer T-staging prediction result is established to obtain a complete dynamic nomogram.

[0084] This embodiment visualizes clinical decision-making by outputting a radiomics score (Radscore) in a dynamic nomogram. As shown in Figure 5, a) displays the top 30 of the 290 key radiomics features selected after LASSO. b) shows the dynamic prediction nomogram, where the Radscore axis is scaled with the quartiles of the feature value distribution, and the clinical severity axis optimizes the display range based on the clinical decision threshold. The stratified Brier scores for T stages are P(T≤1)=0.195, P(TS2)-0.192, and P(TS3)0.203, respectively, with a combined Brier score of 0.197, indicating the model's reliability in distinguishing different T stages.

[0085] This embodiment uses HTML, CSS, and JavaScript technologies, combined with visualization libraries such as D3.js and Chart.js, to design the nomogram. Specifically, based on the dominance ratio and regression coefficient of each variable, the scale range and position of each variable in the nomogram are determined. Variables with greater influence are placed in more prominent positions, and the scale division is more refined. In the nomogram, a corresponding line segment is created for each variable, labeled with the variable name, scale value, and other information. Simultaneously, an overall prediction result line segment is set to display the predicted probability of gastric cancer T stage after combining all variables.

[0086] In this embodiment, the improved ResNet152 network architecture, after processing with 7×7 convolutional kernels, 3×3 max pooling layers, and 16 residual modules, finally outputs the T1-T4 four-class classification probabilities through a global average pooling layer and a fully connected layer. For example, for a certain CT image data, the model might output a probability of 0.1 for the sample belonging to stage T1, 0.3 for stage T2, 0.4 for stage T3, and 0.2 for stage T4. Radiomics scores and clinical parameters (such as CEA and Lauren classification) are input into an ordered logistic regression model. This model quantifies the influence of each variable on the probability of gastric cancer T stage by calculating regression coefficients and odds ratios. For example, if the model predicts a probability odds ratio of 1.5 for a sample changing from stage T2 to stage T3 for every unit increase in the CEA value, it indicates that changes in the CEA value affect the stage probability proportionally. Based on the influence weights of each variable on the staging probability obtained from the ordered logistic regression model, a mapping relationship is established between the scale values ​​of each variable and the predicted result segment (i.e., the probability of gastric cancer T stage). When a user adjusts the scale of a variable (e.g., the CEA value) in the nomogram, the system recalculates the predicted T stage probability after integrating all variables based on this mapping relationship and marks the new probability value on the predicted result segment. For example, if the user adjusts the CEA value, the model will recalculate the predicted T stage probability after integrating all variables based on the regression coefficient and odds ratio of CEA and mark it on the predicted result segment. In this way, doctors can intuitively see the changes in the probability of gastric cancer T stage under different variable values, assisting clinical decision-making.

[0087] In some embodiments, the method further includes: embedding a dynamic nomogram into an HTML page, and adding a title, legend, and annotations to the HTML page to generate an interactive nomogram HTML report.

[0088] In this embodiment, rich interactive functions can be added to the nomogram. When the user hovers the mouse over a variable's scale or marker, the mouse event is captured via JavaScript, and a pop-up tooltip displays detailed information about the variable, such as its meaning and the impact of the current value on the stage. This allows users to dynamically adjust the values ​​of each variable by dragging sliders or entering numerical values, and updates the probability values ​​and related graphical displays on the prediction result segments in real time, enabling users to intuitively understand the impact of different variable changes on the prediction of gastric cancer T-stage.

[0089] The colors, fonts, and line styles of nomograms can be customized to better suit the style and aesthetic requirements of medical charts, making them easier for doctors to read and understand. The designed dynamic nomograms can be embedded into HTML pages and integrated with gastric cancer T-staging results, Grad-CAM visualization heatmaps, and other content to form a complete interactive nomogram HTML report. This report can be displayed and interacted with correctly on different devices (such as computers, tablets, and mobile phones) and browsers, and supports integration into hospital PACS systems, facilitating its use by doctors in clinical work.

[0090] In some embodiments, the method further includes: The backbone network of the residual network is initialized using ImageNet pre-trained weights, and the deep learning-based multi-classifier model is trained by combining the cross-entropy loss function, optimizer, and dynamic early stopping.

[0091] Specifically, ImageNet is a massive dataset containing a vast amount of image data. Models pre-trained on this dataset have already learned rich general image features, such as edges, textures, and shapes. Applying these pre-trained weights to the backbone of the improved ResNet152 network architecture enables the network to possess certain feature extraction capabilities from the early stages of training, eliminating the need to learn basic features from scratch. This significantly accelerates the network's convergence speed and reduces the training time and computational resources required. Furthermore, with the help of pre-trained weights, the model can learn specific features related to gastric cancer more quickly on the gastric cancer T-staging task, contributing to improved final model performance.

[0092] After completing the network architecture construction steps, including receiving venous phase CT images of specific dimensions as input, initial dimensionality reduction using 7×7 convolutional kernels and 3×3 max pooling layers, connecting 16 residual modules and setting relevant structures and functions, and introducing an adaptive multi-scale feature fusion mechanism, the backbone network is initialized using ImageNet pre-trained weights. This involves loading the pre-trained weight file and assigning its parameters to the corresponding layers of the improved ResNet152 network. Only after initialization will the subsequent training phase, which optimizes parameters using the cross-entropy loss function and the Adam optimizer, begin. This embodiment leverages ImageNet pre-trained weights to endow the network with initial feature extraction capabilities, laying the foundation for efficient subsequent training.

[0093] In the gastric cancer T-staging model, the network outputs probability values ​​for four categories: T1-T4. The cross-entropy loss function evaluates the accuracy of the model's predictions by measuring the difference between the predicted probability distribution and the true category distribution. The more accurate the model's predictions, i.e., the closer the predicted probability distribution is to the true category distribution, the smaller the cross-entropy loss value. During training, the model adjusts its network parameters based on the loss value calculated by the cross-entropy loss function, continuously reducing the loss value and making the model's predictions increasingly closer to the actual gastric cancer T-staging. For example, if a CT image actually indicates T2 stage, the model initially predicts a low probability of T2. The cross-entropy loss function calculates a larger loss value, which then drives the model to adjust its parameters, increasing the probability of predicting T2 stage. This embodiment employs the Adam optimizer, an algorithm for parameter optimization that combines the advantages of the Adaptive Gradient Descent (AdaGrad) and Root Mean Square Propagation (RMSProp) algorithms. During training, the Adam optimizer adaptively adjusts the learning rate based on the gradient of different parameters. For parameters with large gradient changes, it decreases the learning rate to avoid over-updating; for parameters with small gradient changes, it increases the learning rate to accelerate parameter updates. In this way, the Adam optimizer can accelerate the model's convergence speed while ensuring training stability, enabling the model to find the optimal parameter combination more efficiently during training, thereby improving the model's predictive ability for gastric cancer T-staging. A learning rate of 0.001 is set, a value that has been experimentally verified to provide a good balance between training speed and stability in this model training.

[0094] In this embodiment, the cross-entropy loss function effectively measures the difference between the model's predicted probability distribution and the true class distribution. In the gastric cancer T-staging model, the model outputs probability values ​​for four categories: T1-T4. The cross-entropy loss function clearly reflects the closeness of the model's prediction to the actual staging, providing a clear direction for adjusting the model's parameters. The more accurate the model's prediction, the smaller the cross-entropy loss value. The model adjusts the network parameters based on this loss value, continuously reducing the loss and making the model's prediction increasingly closer to the actual gastric cancer T-staging. The Adam optimizer combines the advantages of the Adaptive Gradient Algorithm (AdaGrad) and the Root Mean Square Propagation Algorithm (RMSProp), adaptively adjusting the learning rate based on the gradient of different parameters. For parameters with large gradient changes, it reduces the learning rate to avoid over-updating parameters; for parameters with small gradient changes, it increases the learning rate to accelerate parameter updates. This ensures training stability while accelerating the model's convergence speed, enabling the model to find the optimal parameter combination more efficiently during training and improving the model's ability to predict gastric cancer T-staging.

[0095] Furthermore, during model training, continuous training may lead to overfitting, as the model may learn excessively from noise and special cases in the training data, resulting in poorer performance on new data. This dynamic early stopping mechanism in this embodiment continuously monitors the model's performance on the validation set (e.g., AUC). If the model shows no performance improvement on the validation set for 20 consecutive rounds, it indicates that the model may be overfitting or trapped in a local optimum, at which point training is terminated. This effectively prevents overfitting, ensuring the model has good generalization ability on new data, and enabling the trained model to reliably predict gastric cancer T-stage using CT image data from different patients and hospitals in practical applications.

[0096] In this embodiment, after initializing the backbone network with ImageNet pre-trained weights, the model is trained using cross-entropy loss function and Adam optimizer based on gastric cancer T-staging data. During this process, the model continuously adjusts the parameters of convolutional layers, residual modules, fully connected layers, and other structures in the network to reduce the cross-entropy loss value, thereby improving the accuracy of gastric cancer T-staging prediction.

[0097] This invention starts directly from the original CT images, eliminating the need for doctors to manually annotate tumor areas. It automatically extracts multi-scale image features using an improved ResNet152 network and combines this with radiomics data and clinical indicators to construct a multi-classification model that distinguishes between T1 and T4 stages. The model utilizes a weakly supervised learning strategy to locate tumor areas and generates heatmaps using Grad-CAM to visually demonstrate the decision-making process. Simultaneously, it integrates dynamic nomograms to quantify the impact of various clinical parameters.

[0098] Example 2 like Figure 6 As shown, this embodiment of the invention provides a deep learning-based intelligent decision-making system for T-staging of gastric cancer CT images, which includes the following modules: The data processing module is used to acquire CT image data and preprocess the CT image data to obtain standardized venous phase CT images; An improved network architecture module is used to input standardized venous phase CT images into an improved residual network. The improved residual network extracts features from the standardized venous phase CT images to obtain feature maps at different levels. The gastric cancer T-staging prediction results are obtained through the feature maps. A visualization heatmap generation module is used to calculate the gradient of each feature map pair for the predicted class in the last convolutional layer of the residual network and obtain a visualization heatmap. The dynamic nomogram generation module is used to filter out key features from the N-dimensional deep features extracted from the fully connected layer of the residual network through a regression model, obtain radiomics scores based on the key features, and generate dynamic nomograms based on the radiomics scores and clinical parameters. The output module is used to output the gastric cancer T-stage prediction results, a visual heatmap, and a dynamic nomogram.

[0099] This invention discloses a deep learning-based intelligent decision-making system for T-staging of gastric cancer CT images, used for fully automated intelligent decision-making in T-staging of gastric cancer. Based on an improved ResNet152 network architecture, its input layer receives a standardized 224×224 pixel venous phase CT image (window width 350HU, window level 50HU). Initial dimensionality reduction is performed using 7×7 convolutional kernels and 3×3 max-pooling layers, followed by 16 residual modules (including bottleneck structures: 1×1 convolutional dimensionality reduction, 3×3 spatial convolution, and 1×1 convolutional dimensionality increase). Batch normalization and ReLU activation functions are inserted at each layer to alleviate gradient vanishing. Finally, T1-T4 four-class classification probabilities are output through global average pooling layers and fully connected layers. The network introduces an adaptive multi-scale feature fusion mechanism, integrating shallow texture and deep semantic features to enhance sensitivity to tumor invasion depth.

[0100] In the data processing workflow, CT images undergo standardized preprocessing (linear normalization based on spleen parenchymal CT values, isotropic resampling to 1×1×1 mm³ voxels) and data augmentation through random translation, rotation, and mirror flipping to improve model robustness. Grad-CAM technology is used to backpropagate gradient information, enabling automatic tumor region localization without manual annotation. The model integrates radiomics features with clinical parameters (e.g., CEA, Lauren classification). 2048-dimensional depth features extracted from the ResNet152 fully connected layer are used to screen key features via LASSO regression to construct a radiomics score (Radscore). Combined with an ordered logistic regression model, the contribution of multi-dimensional parameters is quantified, generating a dynamic nomogram tool to intuitively analyze the weight of each variable on the staging.

[0101] During training, ImageNet pre-trained weights are used to initialize the backbone network, and parameters are optimized using the cross-entropy loss function and the Adam optimizer (learning rate 0.001). A dynamic early stopping mechanism is set (training is terminated after 20 consecutive rounds without improvement) to prevent overfitting.

[0102] The following is for reference. Figure 7The present invention illustrates an electronic device suitable for implementing embodiments of the present disclosure. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0103] like Figure 7 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 607 into a random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of the electronic device. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0104] Typically, the following devices can be connected to I / O interface 605: input devices 609 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 608 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 607 including, for example, magnetic tapes, hard disks, etc.; and communication devices 606. Communication device 606 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0105] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 606, or installed from storage device 607, or installed from ROM 602. When the computer program is executed by processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0106] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0107] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad-hoc networks), as well as any currently known or future-developed networks.

[0108] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0109] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the following steps: Acquire CT image data and preprocess the CT image data to obtain standardized venous phase CT images; The standardized venous phase CT image is input into an improved residual network. The improved residual network is used to extract features from the standardized venous phase CT image to obtain feature maps at different levels. The gastric cancer T-stage prediction result is obtained through the feature maps. Calculate the gradient of each feature map pair for the predicted class in the last convolutional layer of the residual network to obtain a visual heatmap; Key features are selected from the N-dimensional deep features extracted from the fully connected layer of the residual network using a regression model. Radiomics scores are obtained based on the key features. Dynamic nodal plots are generated based on the radiomics scores and clinical parameters. Outputs gastric cancer T-stage prediction results, visualization heatmap, and dynamic nomogram.

[0110] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk (an object-oriented programming language), and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0112] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0113] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0114] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0115] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0116] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0117] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A deep learning-based intelligent decision-making method for T-staging of gastric cancer CT images, characterized in that, Includes the following steps: Acquire CT image data and preprocess the CT image data to obtain standardized venous phase CT images; The standardized venous phase CT image is input into an improved residual network. The improved residual network is used to extract features from the standardized venous phase CT image to obtain feature maps at different levels. The gastric cancer T-stage prediction result is obtained through the feature maps. Calculate the gradient of each feature map pair for the predicted class in the last convolutional layer of the residual network to obtain a visual heatmap; Key features are selected from the N-dimensional deep features extracted from the fully connected layer of the residual network using a regression model. Radiomics scores are obtained based on the key features. Dynamic nodal plots are generated based on the radiomics scores and clinical parameters. Output the gastric cancer T-stage prediction results, the visualized heatmap, and the dynamic nomogram.

2. The method according to claim 1, characterized in that, The acquisition of CT image data and the preprocessing of the CT image data to obtain standardized venous phase CT images specifically include: Extract the region of interest (ROI) from the CT image data, calculate the CT value of each voxel within the ROI, and map all CT values ​​to the [0,1] interval; The CT image data is resampled, and the original voxels in the CT image data are resampled into cubes with the same size in all directions. The CT image data is processed using the same window width and window level, and enhanced by random translation, rotation, and mirror flipping to obtain standardized venous phase CT images.

3. The method according to claim 1, characterized in that, The process of inputting the standardized venous phase CT image into an improved residual network, extracting features from the standardized venous phase CT image using the improved residual network to obtain feature maps at different levels, and obtaining gastric cancer T-staging prediction results using the feature maps specifically includes: The standardized venous phase CT images are received through the input layer; Preliminary feature extraction is performed on the standardized venous phase CT image using convolutional kernels to obtain a preliminary feature map. Then, the resolution of the preliminary feature map is reduced by a max pooling layer to obtain shallow texture features. Deep semantic features are extracted from the preliminary feature map through several residual modules; wherein each residual module includes a convolutional dimensionality reduction layer, a spatial convolutional layer, and a convolutional dimensionality increase layer, and batch normalization and ReLU activation functions are inserted after the convolutional dimensionality reduction layer, the spatial convolutional layer, and the convolutional dimensionality increase layer, respectively; The shallow texture features and the deep semantic features are integrated to obtain a fused feature map; The fused feature map is mapped to the gastric cancer T-stage prediction result through a global average pooling layer and a fully connected layer.

4. The method according to claim 3, characterized in that, The method of mapping the fused feature map to gastric cancer T-stage prediction results through a global average pooling layer and a fully connected layer specifically includes: The average value of all pixels in each fused feature map is calculated through the global average pooling layer, and the fused feature map is transformed into a vector of a specific length and output to the fully connected layer. Each neuron in the fully connected layer is connected to all elements of the vector of a specific length. Through weight matrix multiplication and bias addition, the vector of a specific length is linearly transformed to establish the mapping relationship between the vector of a specific length and each category of T stage. The score of each element in the vector of a specific length corresponding to the T stage category is calculated. The scores of the T-stage categories are converted into probability values ​​using the Softmax activation function to obtain the T-stage prediction results for gastric cancer.

5. The method according to claim 4, characterized in that, The calculation of the gradient of each feature map pair in the last convolutional layer of the residual network to predict the class and obtain a visual heatmap specifically includes: The predicted category score is obtained from the score of the T stage category. Backpropagation starts from the predicted category score. The gradient of the predicted category score with respect to the feature map of the last convolutional layer is calculated. The gradient is then subjected to global average pooling in the spatial dimension to obtain the importance weight of each channel. The initial version of the heatmap is obtained by weighting and summing the feature maps according to the importance weights, and the positive contribution region of the initial version of the heatmap is retained by applying the ReLU activation function to obtain the Grad-CAM heatmap. The Grad-CAM heatmap is upsampled to the size of the original CT image, the upsampled heatmap is converted into a pseudo-color heatmap, and the pseudo-color heatmap is superimposed on the original CT image to generate a visualized heatmap.

6. The method according to claim 1, characterized in that, The acquisition of radiomics scores based on the key features specifically includes: The selected key features are standardized to obtain key features of the same scale. The weight of each key feature is determined based on the regression model; The radiomics score is obtained by multiplying each key feature by its corresponding weight and summing the results.

7. The method according to claim 1, characterized in that, The generation of a dynamic nomogram based on the radiomics score and the clinical parameters specifically includes: The radiomics scores and clinical parameters are standardized and normalized to obtain a standardized dataset; Using data from the standardized dataset as independent variables and the gastric cancer T-stage prediction results as dependent variables, an ordered logistic regression model was constructed. The regression coefficient of each independent variable is calculated using the ordered logistic regression model, and the odds ratio of each independent variable is calculated based on the regression coefficient. Based on the odds ratio and the regression coefficient, the scale range and position of each independent variable in the dynamic nomogram are determined, and a line segment corresponding to the gastric cancer T-staging prediction result is created for each independent variable. The mapping relationship between the scale value of each independent variable in the dynamic nomogram and the line segment corresponding to the gastric cancer T-staging prediction result is established to obtain a complete dynamic nomogram.

8. The method according to claim 7, characterized in that, The method further includes: embedding the dynamic nomogram into an HTML page, and adding a title, legend, and annotations to the HTML page to generate a dynamic interactive nomogram HTML report.

9. The method according to claim 1, characterized in that, The method further includes: The backbone network of the residual network is initialized using ImageNet pre-trained weights, and the deep learning-based multi-classifier model is trained by combining the cross-entropy loss function, optimizer, and dynamic early stopping.

10. A deep learning-based intelligent decision-making system for T-staging of gastric cancer CT images, characterized in that, Includes the following modules: The data processing module is used to acquire CT image data and preprocess the CT image data to obtain standardized venous phase CT images. An improved network architecture module is used to input the standardized venous phase CT image into an improved residual network, extract features from the standardized venous phase CT image through the improved residual network, obtain feature maps at different levels, and obtain gastric cancer T-staging prediction results through the feature maps; A visualization heatmap generation module is used to calculate the gradient of each feature map pair for the predicted class in the last convolutional layer of the residual network and obtain a visualization heatmap. The dynamic nomogram generation module is used to select key features from the N-dimensional deep features extracted from the fully connected layer of the residual network through a regression model, obtain radiomics scores based on the key features, and generate dynamic nomograms based on the radiomics scores and clinical parameters. The output module is used to output the gastric cancer T-stage prediction results, the visualized heatmap, and the dynamic nomogram.