CT image segmentation and classification system based on class activation graph guidance
By building a multi-task collaborative network based on a CT image segmentation and classification system guided by class activation maps, we solved the problems of difficult data acquisition and lack of multi-task collaborative mechanism in gastrointestinal tumor diagnosis, achieved accurate segmentation and classification of tumor areas, and improved the accuracy and efficiency of diagnosis.
Patent Information
- Application Number
- CN202510944082.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies in gastrointestinal tumor diagnosis face difficulties in data acquisition, high labeling costs, complex tumor representation, and uncertainty in imaging features. In addition, there is a lack of multi-task collaboration mechanisms, which leads to limited generalization capabilities of deep learning models and makes it difficult to achieve accurate tumor segmentation and classification.
A CT image segmentation and classification system based on class activation map guidance is adopted. Through the coarse segmentation subnetwork enhanced by boundary perception, the classification subnetwork guided by segmentation mask and the fine segmentation subnetwork enhanced by class activation map, a knowledge transfer mechanism is constructed between tasks to achieve accurate segmentation and classification of tumor areas.
It significantly improves the accuracy of tumor segmentation and classification, enhances the sensitivity of tumor area positioning and the ability to distinguish pathological characteristics, and realizes the early assessment and treatment of gastrointestinal tumors.
Smart Images

Figure CN120808026A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of medical image processing, and particularly discloses a CT image segmentation and classification system based on class activation map guidance. BACKGROUND
[0002] Gastrointestinal malignant tumors, as common malignant tumors in clinical practice, mainly include two types of gastric cancer and colorectal cancer. Among the gastrointestinal mesenchymal tumors, gastrointestinal stromal tumor (GIST) is the most common potential malignant mesenchymal tumor, and its biological behavior has significant heterogeneity. The tumor originates from Cajal interstitial cells, and the anatomical distribution has significant tendency: it is commonly found in the body of the stomach and the fundus of the stomach, accounting for 50-70%; followed by the small intestine, accounting for 20-30%, the colon and rectum accounts for about 10%, and the esophagus has the lowest incidence, accounting for only 3-5%. It is worth noting that GIST needs to be differentiated from benign mesenchymal tumors such as gastric leiomyoma and gastric schwannoma. Among them, gastric leiomyoma (GL) is the representative of benign interstitial tumors, which originates from the smooth muscle tissue of the intrinsic muscle layer or the muscularis mucosa, and accounts for the highest proportion in benign tumors of the stomach. Gastric schwannoma (GS) is a rare tumor derived from Schwann cells of the nerve sheath, and its incidence accounts for only 0.2% of all nerve sheath tumors. Due to the high similarity of GS in imaging performance to other mesenchymal tumors, the clinical misdiagnosis rate is high. In addition, ectopic pancreas (EP) is a congenital developmental abnormality disease, more than 70% of which occurs in the stomach and duodenum, and it is easy to be misdiagnosed as interstitial tumor or leiomyoma due to small lesions and lack of specific symptoms.
[0003] In clinical practice, the diagnosis of gastrointestinal tumors relies on the application of medical imaging technology. The current mainstream imaging methods include computed tomography (CT), magnetic resonance imaging (MRI), endoscopic ultrasound (EUS), positron emission computed tomography (PET), etc. Among them, computed tomography has become a routine imaging examination method for gastrointestinal tumor screening due to its high imaging density resolution, convenience, speed, non-invasiveness, etc. In clinical practice, doctors make diagnoses and develop treatment plans based on CT reports made by radiologists. Even experienced radiologists can only diagnose 7 chest CT scans per hour on average. However, the manual inspection process layer by layer is quite time-consuming and largely depends on the experience of radiologists. Developing a non-invasive and rapid method for accurate positioning of gastrointestinal tumor types and tumor regions helps in early evaluation of gastrointestinal tumors, so that such patients can receive treatment earlier and faster. Accurate segmentation of tumors not only helps in diagnosis, but also has value in evaluating the severity of the disease and prognosis. Therefore, an effective CT image computer-aided diagnosis (CAD) system is crucial for assisting doctors in selecting reasonable surgical methods and prognosis schemes.
[0004] The current medical image analysis field faces the following key challenges: First, data acquisition difficulty and annotation bottleneck. Medical image data has strong privacy attributes and high acquisition costs, resulting in a lack of available training samples. More importantly, segmentation model training relies heavily on pixel-level annotation data, which requires senior radiologists to outline each layer, consuming a lot of time and involving high labor costs. This double restriction directly leads to the difficulty of meeting the parameter optimization needs of deep learning models in terms of quality and quantity of training data. Second, the complexity of tumor representation and the uncertainty of imaging features. Different pathological types of tumors are highly similar in morphology, volume, and anatomical location. The same type of tumor has significant differences in different individuals. The partial volume effect, motion artifacts, and low signal-to-noise ratio commonly seen in CT imaging lead to discontinuous tumor boundaries, significantly increasing segmentation uncertainty. Third, the lack of multi-task coordination mechanism. Existing researches mainly use independent segmentation or classification task architecture, ignoring the inherent relevance of the two in the feature space. In fact, the morphological features of tumors are strongly correlated with pathological grading, and traditional single-task models cannot effectively utilize this cross-task semantic consistency, resulting in limited model generalization ability.
[0005] The goal of medical image segmentation is to divide the image into multiple regions with semantic meaning according to features such as gray level, color, texture, and shape, achieving the unity of feature homogeneity within regions and feature heterogeneity between regions. Traditional image segmentation methods can be divided into threshold-based segmentation methods, edge detection-based segmentation methods, and region growing-based segmentation methods, according to different feature utilization methods. Threshold segmentation methods distinguish between objects and backgrounds based on gray level differences, but struggle with complex medical images. Edge detection methods use edge operators such as Sobel and Canny to detect gray level changes and locate boundaries, but are easily affected by noise, leading to discontinuous edges. Region growing and splitting and merging methods expand regions based on pixel similarity, but are sensitive to initial seed points and easily disturbed by noise. Common machine learning methods such as K-means clustering and fuzzy C-means (FCM) process handcrafted features to achieve target segmentation.
[0006] In the field of medical image analysis, accurate image segmentation plays a crucial role in improving clinical diagnosis accuracy and reducing the risk of misdiagnosis. However, traditional medical image segmentation methods mainly rely on low-level pixel-level features, which often struggle to achieve ideal diagnostic performance in complex scenarios with low contrast between lesions and normal tissue and blurred boundaries. This limitation stems from the lack of understanding of deep semantic information in traditional methods, making it difficult to effectively model the complex anatomical structure relationships in medical images.
[0007] Based on the strong feature expression ability, deep learning has become the mainstream technology framework in the field of medical image segmentation. As a milestone work, the fully convolutional neural network (FCN) replaced the fully connected layer in the traditional CNN with a convolutional layer and an up-sampling module, and pioneered the end-to-end pixel-level prediction paradigm. On this basis, Ronneberger et al. proposed a network architecture (U-Net) for medical image segmentation. The network has a classic encoder-decoder structure, and the encoder is used to extract multi-level features of the image, and the decoder reconstructs the features. The jump connection between the encoder and the decoder supplements the information at each level, providing the feature extraction capability of the network.
[0008] The effective design of the feature extraction module is a key factor to improve the performance of the model. For example, ResUNet combines the residual learning mechanism of ResNet with the encoder-decoder structure of U-Net, and uses the residual block with gradient propagation advantage to enhance the feature extraction capability, which shows superior performance in retinal image segmentation task. Gibson et al. introduced the dense connection paradigm to reconstruct the traditional convolution module, and built a DenseU-Net that made breakthrough progress in image denoising. Its dense jump connection mechanism effectively promotes the reuse and fusion of multi-scale features. Zhou et al. proposed U-Net++ by establishing cross-level connection network from single layer to four layers, which gives the model the ability to autonomously judge the importance of feature level; but this method also leads to a significant increase in network parameters due to the introduction of dense connection mechanism, which affects the lightweight of the model.
[0009] One of the core challenges in the field of medical image segmentation is the scale difference of the target object, which poses special requirements for the design of the model's receptive field, as the scale of the receptive field directly affects the ability to capture contextual information. To address this issue, He et al. proposed Spatial Pyramid Pooling (SPP) to achieve multi-scale feature extraction by constructing a spatial pyramid structure, which divides the feature map into different spatial units of fine granularity, while preserving the accuracy of local features and fusing global context information. Inspired by SPP, Gu et al. designed a residual multi-kernel pooling module, which constructs multi-level feature representation through four different scale pooling kernels. Although this module effectively expands the receptive field, the up-sampling operation used in the feature decoding stage cannot recover the spatial detail information lost in the pooling process.
[0010] In summary, although significant progress has been made, existing methods are generally limited to single tasks and fail to effectively exploit the inherent consistency between classification and segmentation tasks to achieve collaborative enhancement. This indicates that there is still a lot of research space in the field of gastrointestinal tumor segmentation and classification.
[0011] Multi-task learning framework effectively explores the potential correlation between classification and segmentation tasks by sharing feature representation and parameter constraints. Existing multi-task collaborative optimization research can be divided into two technical routes: non-end-to-end serial collaboration and end-to-end parallel collaboration methods.
[0012] In the serial collaboration method of non-end-to-end serial collaboration, a phased optimization strategy is usually adopted to improve model performance through iterative feedback. In the serial architecture, each task configures an independent network module and is connected through information flow. Specifically, first, train the base model for the priority task, and use the primary feature map and prediction results as the enhanced input for the secondary task. For example, Jin et al. proposed a cascaded knowledge diffusion network for skin lesion diagnosis and segmentation, which realizes bidirectional knowledge transfer through two feature entanglement modules: segmentation features guide the classification network to focus on lesion areas, and classification context knowledge diffuses to the segmentation network. Diao et al. constructed a complementary mask-guided cascaded network, which dynamically adjusts the regional attention weight of the segmentation network through the class activation map generated by the classification task, achieving accurate positioning of the lesion. Wang et al. proposed EMTS-Net, which integrates offline dynamic class activation mapping and random multi-scale training strategy: the random multi-scale strategy eliminates image redundancy interference through multi-scale input, and the offline dynamic class activation mapping uses dynamic attention map to guide the segmentation network to focus on key areas, effectively improving the segmentation accuracy of small polyps. Sun et al. proposed a collaborative multi-task learning method that constructs a three-stream decoder including boundary flow, position flow, and segmentation flow to realize lesion segmentation, and uses the classification subnetwork to fuse multiple information such as lesion mask, boundary mask, and original image for benign and malignant discrimination. At the same time, a refinement segmentation subnetwork is constructed, and samples are adaptively selected according to the classification results to select the corresponding class segmentation head. Jin et al. proposed a mutual guidance model MB-DCNN for skin lesion segmentation and classification. First, use the coarse lesion mask generated by the coarse segmentation network to enhance the lesion positioning and discrimination ability of the classification network, and then transfer the improved positioning ability from the classification network to the enhanced segmentation network to achieve accurate skin lesion segmentation. Although this approach has achieved good performance, each model in the network needs to be trained separately, resulting in complex model training and high computational cost.
[0013] The end-to-end parallel collaborative method synchronously improves the classification and segmentation performance by constructing a joint optimization target. The parallel architecture usually adopts a shared encoder and multiple decoder structure. The encoding network extracts general features for multiple tasks, which are simultaneously input into the segmentation and classification task branches. Each branch extracts differentiated representations through task-specific feature transformation layers and outputs segmentation and classification results. This parallel feature interaction strategy effectively improves the feature expression ability of multi-task learning while maintaining task independence. For example, Wang et al. proposed an interpretable multi-task information bottleneck network MIB-Net, which strengthens the learning of potential representations of tumor regions through information bottleneck constraints and innovatively introduces a dual prior guidance strategy: using location prior to constrain the spatial attention distribution of the segmentation task, and using semantic prior to enhance the feature discriminability of the classification task. Zhu et al. designed a deep collaborative interaction network DSI-Net, which constructs a bidirectional feature transmission mechanism: the lesion location mining module refines the lesion area through spatial attention to optimize the classification features, and the class-guided feature generation module enhances the pixel-level representation ability of the segmentation task through prototype matching. He et al. proposed the MTL-CNN network, which innovatively introduces an edge information enhancement module and a local attention embedding module: the former supplements the lesion contour features through an edge prediction branch, and the latter transmits the spatial constraints of segmentation features to the classification task to realize the guidance of segmentation features to classification. The advantage of this method lies in that, on the one hand, it reduces model redundancy by sharing bottom-level features, and on the other hand, it improves the lesion representation discriminability of the model through task interaction. However, this method only implicitly models the segmentation and classification tasks through a multi-task network, which limits the performance of the multi-task model.
[0014] In summary, current multi-task collaborative algorithms based on deep learning for medical image segmentation and classification have achieved significant results, fully demonstrating the effectiveness of the collaboration between segmentation and classification tasks. However, existing multi-task methods have not fully explored the potential benefits of segmentation and classification results on gastrointestinal tumor classification and segmentation tasks, respectively. Therefore, it is necessary to further study higher-performance classification and segmentation collaborative algorithms to better adapt to the needs of gastrointestinal tumor classification and segmentation and further improve the practicality of the algorithm. SUMMARY
[0015] To solve the technical problems existing in the prior art, the application discloses a CT image segmentation and classification system based on class activation map guidance, provides a network for combined classification and segmentation of gastrointestinal tumors in CT images, the network constructs a boundary perception enhanced coarse segmentation subnetwork, a segmentation mask guided classification subnetwork and a class activation map enhanced fine segmentation subnetwork, the coarse segmentation subnetwork generates an initial segmentation mask as spatial prior, and guides the classification subnetwork to focus on feature learning of the tumor area; the fine segmentation subnetwork focuses on specific tumor areas by using the class activation map generated by the classification subnetwork, realizes iterative optimization of segmentation accuracy of the lesion boundary, and realizes performance improvement of the classification and segmentation tasks through collaborative optimization between networks.
[0016] To achieve the above object, the technical scheme adopted by the application is: a CT image segmentation and classification system based on class activation map guidance, comprising a boundary perception enhanced coarse segmentation subnetwork (EMGNet), a segmentation mask guided classification subnetwork (EMGNet (cls)) and a class activation map guided fine segmentation subnetwork (EMGNet (seg)), and a knowledge transfer mechanism between tasks is constructed through the three independent segmentation and classification networks to realize mutual promotion of performance between different tasks.
[0017] In the training strategy, the coarse segmentation subnetwork and the coarse segmentation subnetwork use the segmentation dataset with pixel-level labels for supervised learning, and the classification subnetwork only uses the image-level classification label for training. The network operation process presents a bidirectional interaction feature: first, the EMGNet generates a tumor coarse segmentation mask to provide spatial positioning prior for the subsequent classification network, effectively enhancing the tumor discrimination ability of the EMGNet (cls); then the fine-grained positioning information captured by the classification network in the form of class activation map is reversely transmitted to the EMGNet (seg), forming a closed-loop optimization mechanism. In the design of the boundary perception enhanced network, a collaborative module (EMG) of a boundary generator (EG) and a mask generator (MG) is constructed, the global interaction of deep encoding features is realized through a grouping shuffling attention (GSA) module, the boundary semantic information in the encoding stage and the global context features in the decoding stage are fully fused, and the initial segmentation accuracy is significantly improved.
[0018] The class activation map guided gastrointestinal tumor segmentation and classification method adopts a multi-task collaborative architecture, mainly composed of three functionally complementary sub-networks: EMGNet, EMGNet (cls) and EMGNet (seg). A phased optimization strategy is adopted: first, the EMGNet is used to preliminarily segment the tumor area, and the generated coarse segmentation mask is used as spatial prior information; then, the mask and the original image are input into EMGNet (cls), and the feature attention mechanism is used to guide the classification network to focus on the lesion area, thereby improving the discrimination ability of pathological features; on this basis, the gradient weighted class activation map is generated by using the trained EMGNet (cls), and the tumor positioning map with spatial discrimination is extracted by using the back propagation algorithm; finally, the original image and the corresponding tumor positioning map are input into EMGNet (seg) to realize fine segmentation of the lesion area.
[0019] The overall architecture of the boundary perception enhanced network adopts an encoder-decoder paradigm, mainly composed of a feature encoder, a segmentation decoder, a global spatial attention module and a mask-boundary guiding module. The basic segmentation framework is composed of an encoding-decoding network composed of a feature encoder and a segmentation decoder, wherein the GSA module enhances the global context interaction of deep features to improve feature utilization, and the EMG module enhances the perception ability of the network to tumor boundaries and regional features through a double-path supervision mechanism. The tumor segmentation result generated by the network can provide spatial prior information for the downstream classification network EMGNet (cls), effectively enhancing its tumor positioning and classification discrimination ability.
[0020] In EMGNet, the feature encoder adopts the feature extraction part of the ResNet34 network, removes the bottom global pooling layer and the fully connected layer, and retains the first five stages of feature extraction modules. Compared with other pre-trained models, ResNet can effectively extract deep semantic features while maintaining parameter efficiency. The output feature maps of the five stages of the encoder are denoted as , and the spatial resolution is , respectively. At the end of the encoder, the group shuffle attention module adopts a three-stage processing flow: first, the group convolution is used to reduce the computational complexity, then the external attention mechanism is used to realize group feature enhancement, and finally the channel shuffle operation is used to realize cross-group information interaction, thereby establishing global context association.
[0021] The segmentation decoder corresponding to the feature encoder includes four decoding blocks and a segmentation head, wherein the decoding block is up-sampled by deconvolution operation and then concatenated with the features of the jump connection in the channel level, and then the decoding feature is generated by the decoding block stacked by two 3x3 convolutions (including normalization and activation processing). Finally, the final decoding feature is output by 1x1 convolution and Sigmoid activation.
[0022] To address the spatial characteristic difference between the encoding-decoding features, the EMG module designs a dual-path supervision mechanism: in the encoding path, the boundary generator guides the network to focus on the boundary information of multi-level encoding features through deep supervision; in the decoding path, the mask generator optimizes the regional features of multi-level decoding features through deep supervision. Through the feature weighting fusion strategy, the module significantly improves the sensitivity of the network to tumor boundaries and the accuracy of regional positioning. The specific implementation is shown in equations (1) and (2):
[0023] (1)
[0024]
[0025] wherein, represents multi-level encoding features, represents multi-level decoding features, is a Sigmoid function, conv(·) represents a 1 × 1 convolution, and represents the element-wise multiplication operation, and represent the boundary-enhanced encoding features and the region-enhanced decoding features, respectively. Through adaptive calibration at the feature level, the boundary information and regional features are optimized cooperatively.
[0026] EMGNet (cls) adopts the Xception architecture as the feature extraction backbone network. According to the characteristics that shallow features retain detailed information while deep features are rich in semantic information, a multi-stage feature fusion strategy is proposed. This network effectively improves the tumor positioning and tumor type discrimination ability by integrating coarse segmentation mask information and multi-scale feature extraction mechanism.
[0027] The feature extraction process of the classification network contains four levels , and the resolution of the feature maps of each level is , , and , respectively. The scale alignment of the feature maps is achieved through adaptive up-sampling and convolution operation, and the fusion process can be represented as:
[0028]
[0029] wherein, represents the channel concatenation operation, and represent 3 × 3 convolution and up-sampling, respectively. Finally, a coarse segmentation mask guidance mechanism is introduced to strengthen the tumor region feature expression through residual attention mechanism. Specifically, the fused features are combined with the coarse segmentation mask Residual processing is performed, and the classification result is guided by the coarse segmentation mask by focusing on the features of the tumor region. The final classification result is output after a global average pooling layer, a fully connected layer, and a softmax function , and the specific formula is as follows:
[0030] (4)
[0031] In the formula, represents a global average pooling layer, represents a fully connected layer, represents an element-wise multiplication operation, represents an element-wise addition operation.
[0032] The Grad-CAM algorithm provides visual explanation for the classification network when making classification decisions. The CAM generated for the input image can indicate the degree of attention of the network to different regions in the image when making classification predictions. Under the guidance of the coarse segmentation mask, EMGNet (cls) pays more attention to the discriminative information in the tumor region, thereby guiding the fine segmentation of the tumor region.
[0033] The multi-task collaborative architecture based on class activation map guidance realizes the joint optimization of gastric tumor positioning and classification through a collaborative training mechanism. In the classification branch EMGNet (cls), the coarse segmentation mask generated by the boundary perception enhancement network EMGNet is injected into the classification subnetwork as prior knowledge, which significantly enhances the positioning sensitivity and class discrimination ability of the model to the tumor region. After sufficient training, the class activation map generated by EMGNet (cls) not only contains high-precision tumor spatial distribution information, but also encodes fine-grained pathological features, providing key information guidance for the downstream segmentation network EMGNet (seg).
[0034] EMGNet (seg) has a similar structure to EMGNet, with the only difference being the introduction of the corresponding CAM generated by the classification network EMGNet (cls), which provides positioning information for different types of tumors, improving the segmentation performance of EMGNet (seg). The input of EMGNet (seg) includes two parts; the original image data and the CAM generated by the classification branch, which realizes information complementation through feature fusion. In the feature preprocessing stage, the nearest neighbor interpolation algorithm is used for upsampling the CAM, which effectively avoids the positioning information decay that may be caused by traditional interpolation methods by maintaining the spatial correlation of the original activation values. After sigmoid function normalization processing, the CAM is converted to The spatial attention weight map of the interval guides the multi-level decoding features to focus on the region of interest. By establishing a residual connection between the original features and the attention weighted features, the advantages of CAM guided lesion positioning are maintained, and the tumor edge regions not covered by the initial CAM are effectively captured. This design strategy significantly improves the segmentation quality of the lesion boundary while maintaining the segmentation accuracy of the tumor main body region.
[0035] In the design of the segmentation network EMGNet, to address the common foreground-background pixel imbalance problem in medical images, a composite loss function is constructed by jointly optimizing the Dice loss and the binary cross-entropy loss.
[0036] Specifically, the Dice loss effectively alleviates the class imbalance problem by calculating the pixel-level similarity between the prediction results and the true segmentation mask; while the BCE loss strengthens the learning of edge details through pixel-by-pixel classification constraints. Based on the deep supervision strategy, the three lateral outputs and four edge outputs of the network are all subjected to upsampling operations and are spatially aligned with the clinical expert-labeled gold standard and its corresponding edge The total segmentation loss is composed of the following two parts:
[0037] (5)
[0038] In the formula, and respectively represent the upsampled side output feature map and the edge feature map, and the composite loss mechanism optimizes the main body region segmentation accuracy and edge positioning accuracy simultaneously through multi-scale supervision.
[0039] For the gastrointestinal tumor classification task, the cross-entropy loss function is used to measure the difference between the prediction results and the true labels, and the mathematical expression of the cross-entropy loss function is shown in equation (6):
[0040] (6)
[0041] In the formula, represents the total number of classes in the classification task, represents the true label of the sample belonging to the class, represents the probability of the model predicting that the sample belongs to the class, and by minimizing the distribution difference between the prediction probability and the true label, the network learns discriminative pathological feature representations.
[0042] The application proposes a bidirectional knowledge transfer framework based on class activation map guidance, which mines the internal correlation between anatomical structure features and pathological semantic features in gastrointestinal tumor images, establishes a dynamic feature enhancement channel across tasks, and realizes mutual improvement of segmentation and classification performance.
[0043] In the design of the EMGNet, a double-path feature generation mechanism is adopted: a GSA module is introduced to construct feature interaction of global information, a boundary generator is used to extract high-frequency boundary features in the encoding stage, and a mask generator is used to integrate global context information in the decoding stage, so that more accurate segmentation results are generated.
[0044] The application utilizes the internal correlation between segmentation and classification tasks to construct a bidirectional optimization path of segmentation-guided classification and classification-feedback segmentation. On the one hand, the coarse segmentation results containing tumor location information are used to enhance the tumor discrimination ability of the classification network; on the other hand, the class activation map generated by the classification network is converted into spatial position weight to guide the segmentation network to focus on the lesion area with discriminative ability, and a more refined segmentation mask is generated. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 It is the overall structure diagram of the gastrointestinal tumor segmentation and classification network guided by the class activation map.
[0046] Figure 2 It is the overall structure diagram of the boundary perception enhanced network.
[0047] Figure 3 It is the structure diagram of the mask-guided classification network EMGNet (cls).
[0048] Figure 4 It is the structure diagram of the class activation map guided enhanced network EMGNet (seg).
[0049] Figure 5 It is the comparison diagram of the segmentation performance of the system and the existing public method on GISTS.
[0050] Figure 6 It is the visualization diagram of the segmentation results of different algorithms on the GISTS dataset.
[0051] Figure 7 It is the comparison diagram of the confusion matrix of the classification network EMGNet (cls) proposed by the system and other classification methods, Figure 7 (a) is the comparison diagram of the confusion matrix with ResNet, Figure 7 (b) is the comparison diagram of the confusion matrix with SE-ResNet, Figure 7 (c) is the comparison diagram of the confusion matrix with DenseNet, Figure 7 (d) is the comparison diagram of the confusion matrix with EMGNet (cls).
[0052] Figure 8 A diagram showing the class activation map for the classification network ablation method. Figure 8 (a) is a display diagram of the class activation map of the CT image. Figure 8 (b) is a display diagram of the class activation map of the Xception map, and 8 (c) is a display diagram of the class activation map of EMGNet (cls). Figure 8 (d) is a display of the class activation map of the segmentation label. DETAILED DESCRIPTION
[0053] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0054] like Figure 1 As shown in the figure, the CT image segmentation and classification system guided by class activation maps includes a boundary-aware enhanced coarse segmentation subnetwork (EMGNet), a segmentation mask-guided subclassification network (EMGNet(cls)), and a class activation map-guided fine segmentation subnetwork (EMGNet(seg)). By constructing a knowledge transfer mechanism between tasks through three independent segmentation and classification networks, the performance of different tasks can be mutually promoted and improved.
[0055] In terms of training strategy, the coarse and coarse segmentation subnets utilize supervised learning from a segmentation dataset with pixel-level annotations, while the classification subnet is trained solely using image-level classification labels. The network's operational process exhibits bidirectional interaction: First, EMGNet generates a coarse tumor segmentation mask, providing a spatial localization prior for the subsequent classification network, effectively enhancing EMGNet (cls)'s tumor discrimination capability. The fine-grained localization information captured by the classification network via class activation maps is then passed back to EMGNet (seg), forming a closed-loop optimization mechanism. In the design of the boundary-aware enhancement network, a collaborative module (EMG) combining the boundary generator (EG) and the mask generator (MG) is constructed. The grouped shuffle attention (GSA) module enables global interaction of deep encoding features, fully integrating boundary semantic information from the encoding phase with global contextual features from the decoding phase, significantly improving initial segmentation accuracy.
[0056] The class activation map guided gastrointestinal tumor segmentation and classification method adopts a multi-task collaborative architecture, which is mainly composed of three functionally complementary sub-networks: EMGNet, EMGNet(cls) and EMGNet(seg). Figure 1As shown, a phased optimization strategy is adopted: first, the preliminary segmentation of the tumor region is performed by EMGNet, and the generated coarse segmentation mask is used as spatial prior information; then, the mask and the original image are input into EMGNet (cls), and the classification network is guided to focus on the lesion area through the feature attention mechanism, thereby improving the discrimination ability of pathological features; on this basis, the gradient weighted class activation map is generated by using the trained EMGNet (cls), and the tumor localization map with spatial discriminability is extracted through the back propagation algorithm; finally, the original image and the corresponding tumor localization map are input into EMGNet (seg), and the fine segmentation of the lesion area is realized.
[0057] The overall architecture of the boundary perception enhanced network is as shown in Figure 2 The core architecture adopts an encoder-decoder paradigm, and is mainly composed of four key components: a feature encoder, a segmentation decoder, a global spatial attention module, and a mask-boundary guiding module. The basic segmentation framework is composed of an encoding-decoding network composed of a feature encoder and a segmentation decoder, wherein the GSA module enhances the global context interaction of deep features to improve feature utilization, and the EMG module enhances the perception ability of the network to tumor boundaries and regional features through a double-path supervision mechanism. The tumor segmentation result generated by the network can provide spatial prior information for the downstream classification network EMGNet (cls), effectively enhancing its tumor localization and classification discrimination ability.
[0058] In EMGNet, the feature encoder adopts the feature extraction part of the ResNet34 network, removes the bottom global pooling layer and the fully connected layer, and retains the feature extraction modules of the first five stages. Compared with other pre-trained models, ResNet can effectively extract deep semantic features while maintaining parameter efficiency. The output feature maps of the five stages of the encoder are denoted as , and the spatial resolution is , respectively. At the end of the encoder, the group shuffle attention module adopts a three-stage processing flow: first, the group convolution is used to reduce the computational complexity, then the external attention mechanism is used to realize the intra-group feature enhancement, and finally the channel shuffle operation is used to realize the cross-group information interaction, thereby establishing the global context association.
[0059] The segmentation decoder corresponding to the feature encoder includes four decoding blocks and a segmentation head, wherein the decoding block is up-sampled through the deconvolution operation, and then the channel concatenation is performed with the features connected through the jump connection, and then the decoding features are generated through the decoding block stacked by two 3x3 convolutions (including normalization and activation processing). Finally, the final decoding features are output through the 1x1 convolution and the Sigmoid activation to output the segmentation prediction map.
[0060] To address the spatial characteristic difference of the encoding-decoding features, the EMG module designs a dual-path supervision mechanism: in the encoding path, the boundary generator guides the network to focus on the boundary information of multi-level encoding features through deep supervision; in the decoding path, the mask generator optimizes the region features of multi-level decoding features through deep supervision. Through the feature weighting fusion strategy, the module significantly improves the sensitivity of the network to tumor boundaries and the accuracy of region positioning. The specific implementation is shown in equations (1) and (2).
[0061] (1)
[0062]
[0063] In the formula, represents multi-level encoding features, represents multi-level decoding features, is a Sigmoid function, conv(·) represents a 1×1 convolution, and ⊙ represents element-wise multiplication, and represent boundary-enhanced encoding features and region-enhanced decoding features, respectively. Through adaptive calibration at the feature level, the boundary information and region features are optimized cooperatively.
[0064] EMGNet(cls) adopts the Xception architecture as the feature extraction backbone network, as shown in Figure 3 . According to the characteristics that shallow features retain detailed information while deep features are rich in semantic information, a multi-stage feature fusion strategy is proposed. This network effectively improves tumor positioning and tumor type discrimination ability by integrating coarse segmentation mask information and multi-scale feature extraction mechanism.
[0065] The feature extraction process of the classification network contains four levels , and the resolution of each level feature map is , , and of the input, respectively. The scale alignment of the feature map is achieved through adaptive up-sampling and convolution operation, and its fusion process can be represented as:
[0066]
[0067] In the formula, represents the channel concatenation operation, and represent 3×3 convolution and up-sampling, respectively. Finally, a coarse segmentation mask guidance mechanism is introduced to strengthen the tumor region feature expression through residual attention mechanism. Specifically, the fused feature is combined with the coarse segmentation mask Residual processing is performed, and the guidance of the coarse segmentation mask to the classification result is realized by focusing on the features of the tumor region. The final classification result is output after a global average pooling layer, a fully connected layer, and a softmax function . The specific formula is as follows:
[0068] (4)
[0069] In the formula, represents a global average pooling layer, represents a fully connected layer, represents an element-wise multiplication operation, represents an element-wise addition operation.
[0070] The Grad-CAM algorithm provides visual explanation for the classification network when making classification decisions. The CAM generated for the input image can indicate the degree of attention of the network to different regions in the image when making classification predictions. Under the guidance of the coarse segmentation mask, EMGNet (cls) pays more attention to the discriminative information in the tumor region, thereby guiding the fine segmentation of the tumor region.
[0071] The multi-task collaborative architecture based on class activation map guidance realizes the joint optimization of gastric tumor positioning and classification through a collaborative training mechanism. In the classification branch EMGNet (cls), the coarse segmentation mask generated by the boundary perception enhancement network EMGNet is injected into the classification subnetwork as prior knowledge, which significantly enhances the positioning sensitivity and class discrimination ability of the model to the tumor region. After sufficient training, the class activation map generated by EMGNet (cls) not only contains high-precision tumor spatial distribution information, but also encodes fine-grained pathological features, providing key information guidance for the downstream segmentation network EMGNet (seg). The network structure of EMGNet (seg) is as shown in Figure 4 .
[0072] EMGNet (seg) has a similar structure to EMGNet, with the only difference being the introduction of the corresponding CAM generated by the classification network EMGNet (cls), which provides positioning information for different types of tumors, improving the segmentation performance of EMGNet (seg). The input of EMGNet (seg) includes two parts: the original image data and the CAM generated by the classification branch, which realizes information complementation through feature fusion. In the feature preprocessing stage, the nearest neighbor interpolation algorithm is used for upsampling the CAM, which effectively avoids the positioning information decay that may be caused by traditional interpolation methods by maintaining the spatial correlation of the original activation values. After sigmoid function normalization processing, the CAM is converted into The spatial attention weight map of the interval guides the multi-level decoding features to focus on the region of interest. By establishing a residual connection between the original features and the attention weighted features, the advantages of CAM guided lesion positioning are maintained, and the tumor edge regions not covered by the initial CAM are effectively captured. This design strategy significantly improves the segmentation quality of the lesion boundary while maintaining the segmentation accuracy of the tumor main body region.
[0073] In the design of the segmentation network EMGNet, to address the common foreground-background pixel imbalance problem in medical images, a composite loss function is constructed by jointly optimizing the Dice loss and the binary cross-entropy loss.
[0074] Specifically, the Dice loss effectively alleviates the class imbalance problem by calculating the pixel-level similarity between the prediction results and the true segmentation mask; while the BCE loss strengthens the learning of edge details through pixel-by-pixel classification constraints. Based on the deep supervision strategy, the three lateral outputs and four edge outputs of the network are all subjected to upsampling operations and are spatially aligned with the clinical expert-labeled gold standard and its corresponding edge The total segmentation loss is composed of the following two parts:
[0075] (5)
[0076] where and represent the upsampled side output feature maps and edge feature maps respectively, and the composite loss mechanism optimizes the main body region segmentation accuracy and edge positioning accuracy simultaneously through multi-scale supervision.
[0077] For the gastrointestinal tumor classification task, the cross-entropy loss function is used to measure the difference between the prediction results and the true labels, and the mathematical expression of the cross-entropy loss function is shown in equation (6):
[0078] (6)
[0079] where denotes the total number of classes in the classification task, denotes the true label of the sample belonging to the th class, denotes the probability of the model predicting that the sample belongs to the th class, and by minimizing the distribution difference between the prediction probability and the true label, the network learns discriminative pathological feature representations.
[0080] In the design of the segmentation network EMGNet, to cope with the common foreground-background pixel imbalance problem in medical images, the application constructs a composite loss function by jointly optimizing the Dice loss and the binary cross-entropy loss.
[0081] Specifically, the Dice loss effectively alleviates the class imbalance problem by calculating the pixel-level similarity between the prediction result and the true segmentation mask; and the BCE loss strengthens the learning of edge details through pixel-by-pixel classification constraints. Based on the deep supervision strategy, the three lateral outputs and the four edge outputs of the network are all subjected to upsampling operations and are kept spatially aligned with the gold standard annotated by clinical experts and the corresponding edges . The total segmentation loss is composed of the following two parts:
[0082] (7)
[0083] In the formula, and respectively represent the upsampled side output feature map and the edge feature map. The composite loss mechanism optimizes the segmentation accuracy of the main body region and the edge positioning accuracy at the same time through multi-scale supervision.
[0084] For the gastrointestinal tumor classification task, the application uses the cross-entropy loss function to measure the difference between the prediction result and the true label. The mathematical expression of the cross-entropy loss function is shown in formula (8):
[0085] (8)
[0086] In the formula, denotes the total number of classes of the classification task, denotes the true label of the sample belonging to the th class, denotes the probability of the model predicting that the sample belongs to the th class. By minimizing the distribution difference between the prediction probability and the true label, the network is guided to learn discriminative pathological feature representations.
[0087] The application systematically evaluates the proposed multi-task collaborative network architecture on the collected GISTS dataset. The dataset contains 410 labeled pathological images, which are divided into 233 gastrointestinal stromal tumors, 45 leiomyomas, 110 schwannomas and 22 ectopic pancreases according to the lesion type. The dataset is randomly divided into 287 training samples and 123 test samples in a 7:3 ratio. To enhance the generalization ability of the model, various data augmentation strategies are implemented on the training set, including horizontal and vertical flipping, scaling and random cropping spatial transformation methods.
[0088] Both segmentation networks EMGNet and EMGNet(seg) adopt ResNet34 as the encoder of UNet architecture, and the classification network adopts Xception as the encoder. In the training process, the network parameters are updated using the SGD optimizer, and the training lasts for 300 epochs, the momentum is set to 0.9, the batch size is 8, and the initial learning rate is set to 5x10 -3 , and the image size is 512x512 pixels.
[0089] For the classification task, OA, CK, PRE, RE and F1 are used to evaluate the performance of the model; for the segmentation task, AC, JA, SE, DI and SP are used as quantitative evaluation criteria.
[0090] In order to verify the effectiveness of the method, the proposed multi-task cascade method is compared with the existing medical image segmentation. The quantitative index comparison results of EMGNet and DeepLabv3+, DenseASPP, CE-Net, CPFNet, PraNet, CA-Net, LDNet, MSCA-Net, DSI-Net and other methods are shown in Table 1. It is worth noting that in order to ensure the fairness of the experiment, all the compared methods use the same data preprocessing process, use the same optimizer, learning rate and other hyperparameters, and train and test the network under the same experimental configuration.
[0091]
[0092] In terms of core indicators for measuring the coincidence degree of segmented regions, the DI and JA of EMGNet(seg) are 0.7817 and 0.7011, respectively, which are 6.2% and 6.4% higher than the 0.7359 and JA of the suboptimal model MSCA-Net. The results show that in the presence of significant morphological heterogeneity in the gastrointestinal tumor region, the multi-task cascade architecture proposed in the application can more accurately capture the spatial features of the lesion boundary, and the segmentation result has a higher matching degree with the expert annotated segmentation result.
[0093] On the AC and SP indicators reflecting the overall segmentation accuracy, although limited by the inherent characteristics of gastrointestinal tumors, i.e. the average proportion of lesion area <5%, the method still maintains excellent levels of 0.9986 and 0.9995, which is statistically consistent with the mainstream method. It is worth noting that the SP indicators of all models are higher than 0.998, which confirms the maturity of each algorithm in background area recognition, and the method of the application realizes higher segmentation accuracy while maintaining the index. For the key lesion detection capability in clinical diagnosis, the method achieves a sensitivity of 0.7858 through the SE index, which is 1.2 percentage points higher than the existing optimal model PraNet. Through comprehensive multi-dimensional evaluation results, EMGNet(seg) maintains the accuracy of background recognition while focusing more on the segmentation optimization of tumor area through the innovative multi-task learning architecture.
[0094] The partial segmentation index comparison results of the method of the application and the existing most advanced medical image segmentation method on the gastrointestinal tumor segmentation dataset are shown in Figure 5 DI, JA and SE are selected to represent the segmentation result and label consistency index for comparison of different methods, wherein each pole represents a comparison method, and each circle of different colors represents a different segmentation index. As can be seen from the figure, the EMGNet(seg) proposed in the application has higher segmentation performance and more excellent stability in segmentation performance.
[0095] The segmentation results of different comparison methods are shown in Figure 6 From top to bottom, the segmentation results of the EMGNet(seg) proposed in the application and other comparison methods such as DeepLabv3+ are shown. Each row in the figure represents the segmentation result of a tumor of different size, and each table represents the segmentation result of different comparison methods. Among them, the green line represents the real tumor area marked by the doctor, and the red line represents the tumor area predicted by each method. As shown in the results in the figure, there is no obvious difference between the tumor area and the surrounding organs and tissues, which is not conducive to accurate segmentation of the tumor. As shown in the second column of DeepLabv3+ and CPFNet, there are obvious unsegmented areas, and the same situation also occurs in the fourth column of EMGNet and other methods, which shows that the tumor area is not completely segmented. In addition, DeepLabv3+, LDNet and PraNet in the first column respectively appear unsegmented, missegmented and over-segmented tumor areas. Although the segmentation results of each method have the problem of inaccurate segmentation area, thanks to the preliminary tumor positioning of the coarse segmentation network and the category guidance of the classification network, the segmentation result of the EMGNet(seg) proposed in the application is closest to the real segmentation graph in terms of tumor position, shape and boundary, which fully proves its excellent segmentation performance.
[0096]
[0097] To verify the superiority of the classification subnet EMGNet(cls), the classification performance of the proposed classification network EMGNet(cls) and ResNet, SE-ResNet, DenseNet and EfficientNet on the GISTS dataset is compared as shown in Table 2. Among them, the CK coefficient and OA of EMGNet(cls) are 0.8978 and 0.9431, which are better than other classification models, and are 24.8% and 11.5% higher than the suboptimal model (DenseNet), indicating that the classification consistency and global accuracy are optimal. The F1 of EMGNet(cls) for interstitial tumor is 0.9272, which is 6.7% higher than DenseNet, indicating the lowest risk of missing detection for key categories; the F1 for schwannoma is 0.8358, which is slightly lower than EfficientNet; the F1 for leiomyoma is 0.8571, which is optimal with DenseNet, reflecting that the model has stronger ability to distinguish similar features.
[0098] To prove the effectiveness of the text method, the confusion matrix of the classification network EMGNet(cls) and other classification methods proposed by the application is as shown in Table 4. Figure 7 EMGNet(cls) can accurately distinguish different types of tumors with an accuracy of 92.0%, 100%, 97.22% and 100%. In addition, 1.33% and 6.67% of interstitial tumors are classified as leiomyoma and schwannoma, which may be due to the class imbalance problem and low contrast with the background. Compared with other classification methods, EMGNet(cls) can achieve more accurate classification.
[0099] To verify the effectiveness of the three modules GSA, EG and MG of the coarse segmentation network EMGNet and their collaborative relationship, ResUNet34 is used as the benchmark model, the modules are gradually added and the influence on the segmentation performance is evaluated as shown in Table 3. From the results of the second and third rows, it is found that when only using the GSA module, DI, JA and SE are improved to 0.7480, 0.6666 and 0.7705, respectively, indicating that the GSA module improves the segmentation consistency through global context interaction; when only using the EMG module, SE is significantly improved, indicating that the optimization of the tumor region by the EMG module is more obvious; the DI and JA of GSA combined with EG are 0.6772 and 0.7606, respectively, which are significantly better than the single module, indicating that EG further utilizes the encoding layer detail features to optimize the edge segmentation based on the global features provided by GSA; the DI, JA and SE of EMGNet are 0.7768, 0.6950 and 0.7952, respectively, indicating that GSA, EG and MG cooperate to significantly improve the segmentation accuracy and robustness.
[0100]
[0101] The GSA module captures the global context through deep feature interaction, and provides structured semantic information for subsequent modules. The EG module focuses on the tumor edge detail features of the encoding layer, which is complementary to the global information of the GSA, and reduces the loss of details. The MG module guides the decoding layer to focus on the overall shape of the tumor, and combines the context of the GSA to enhance the regional consistency and avoid local misjudgment.
[0102] In order to verify the effectiveness of the proposed method, the performance comparison of the EMGNet (cls) classification network and its variants is shown in Table 4. The experimental setting includes four comparison models: the benchmark model uses the Xception network as the baseline; Model 1 introduces a rough segmentation mask for feature enhancement based on deep classification features; Model 2 supplements background features and establishes a dual-channel cascade structure based on Model 1; the proposed EMGNet (cls) realizes classification optimization through multi-level fusion feature cascade.
[0103] The experimental results show that the EMGNet (cls) proposed in the application significantly outperforms other comparison models in classification performance by using a multi-level fusion feature cascade scheme. The advantages mainly come from two aspects: first, the feature cascade strategy is used to introduce mask information while retaining the original deep features, effectively alleviating the deviation problem between the rough segmentation mask and the real label; second, the multi-level feature fusion mechanism fully integrates the morphological features of the tumor region and the context information of the surrounding tissues, among which the internal features and edge textures of the tumor provide the basis for the classification of gastrointestinal tumors.
[0104]
[0105] To further verify the effectiveness of the classification subnet EMGNet (cls), the application uses the class activation map visualization method to qualitatively evaluate the explainability of the network. As shown in Figure 8 The original CT image, the CAM results generated by the Baseline and EMGNet (cls), and the real segmentation mask are shown in each row from top to bottom. By comparing and analyzing, it can be seen that compared with the baseline model, the class activation heat map generated by EMGNet (cls) shows more accurate positioning ability in terms of tumor edge contour sketching and internal heterogeneity region feature response, and the spatial coincidence degree of its activation region with the fourth row segmentation mask is significantly higher. This provides prior knowledge guidance with anatomical significance for subsequent segmentation tasks, and effectively enhances the explainability of the deep learning model.
[0106] For the fusion strategy of class activation map in the fine segmentation network, the application carries out systematic comparative research, mainly evaluates the influence of four different feature guiding modes on the segmentation performance. As shown in Table 5, the experimental settings of different guiding modes are as follows: Model 1 adopts a channel cascading strategy, directly splices CAM and deep encoding features after 1x1 convolution fusion. Model 2 takes the Sigmoid normalized CAM as a spatial attention weight, and weights the deep encoding features. Model 3 applies the CAM spatial attention weight to the multi-level encoding features. The EMGNet(seg) scheme extends the CAM attention mechanism from the encoding end to the decoding end, and weights the multi-level decoding features through spatial attention. As can be seen from the results in the table, EMGNet(seg) improves DI to 0.7817, which is 1.1% higher than Model 3, JA reaches 0.7011, and SE index jumps to 0.7858. The performance advantage is due to the multi-level feature fusion mechanism in the decoding stage, which not only preserves the semantic information of deep features, but also enhances the detail expression ability of shallow features through attention guidance, achieving the best balance between global context modeling and local detail preservation.
[0107]
[0108] The influence of the coarse segmentation network EMGNet, the classification network EMGNet(cls) and the fine segmentation network EMGNet(seg) in the multi-task cooperation framework proposed in the application on the segmentation performance is analyzed, as shown in Table 6. Among them, EMGNet represents the coarse segmentation network, (cls) represents the classification network EMGNet(cls), and (seg) represents the fine segmentation network EMGNet(seg).
[0109] Three schemes are compared respectively: (1) only the coarse segmentation network, directly generating the final segmentation result; (2) first generate the classification result and the corresponding CAM by the classification network, and then guide the fine segmentation network EMGNet(seg) to generate the segmentation result by the CAM; (3) first generate the coarse segmentation mask by the coarse segmentation network; second, guide the classification network EMGNet(cls) to generate the classification result and the corresponding CAM by taking the coarse segmentation mask as prior information; finally, guide the fine segmentation network EMGNet(seg) to generate the final segmentation result by taking the CAM as prior information. DI and JA are improved to 0.7817 and 0.7011 respectively, and SE achieves the suboptimal 0.7858, the experimental results show that the multi-stage guiding strategy improves the segmentation performance while effectively suppressing false positives.
[0110]
[0111] The performance of the parallel collaborative network SFGNet and the multi-task collaborative network based on EMGNet proposed in the application is compared. As shown in Tables 7 and 8, overall, the segmentation and classification indicators of the multi-network collaborative architecture based on EMGNet are better than those of the parallel collaborative architecture SFGNet, and SFGNet achieves higher classification performance in specific categories. For example, the PRE of SFGNet for interstitial tumor is 0.945, which is better than 0.921 of EMGNet (cls), indicating that SFGNet has stronger discrimination ability for interstitial tumor.
[0112]
[0113]
[0114]
[0115] In addition to the comparison of segmentation and classification performance mentioned above, the important indicators of the model in actual application also include the parameter amount of the model and the test time. As shown in Table 9, although the segmentation and classification indicators of the parallel network SFGNet are slightly lower than those of the multi-task collaborative framework based on EMGNet, SFGNet has smaller parameter amount and faster test time, which are 39.2M and 6.94ms respectively. Due to the time efficiency limitation of CAM generation, the multi-task framework EMGNet needs to consume more time to obtain better segmentation and classification performance.
[0116] Comprehensive comparison of the performance of the multi-task network, SFGNet uses only a small amount of parameters to achieve good segmentation and classification performance and faster test time; the multi-task collaborative framework based on EMGNet fully utilizes the performance of each network for specific tasks, and achieves better segmentation and classification performance through task interaction between different networks.
[0117] The application provides a multi-task cooperative classification segmentation network based on class activation map guidance, which is used for segmentation and classification of gastrointestinal tumor CT images. The framework realizes cooperative optimization and performance improvement of classification and segmentation tasks by constructing three functionally complementary segmentation and classification networks, namely a boundary perception enhanced coarse segmentation network, a segmentation mask guided classification network and a class activation map enhanced fine segmentation network. The multi-task cooperative mechanism forms a closed-loop optimization process: the classification network enhances the feature discriminability with the help of the segmentation result, and the attention information generated by the classification network is used to improve the positioning accuracy of the segmentation network. Experimental verification shows that through the bidirectional information interaction, the multi-task model effectively overcomes the limitations of the traditional single-task model, and achieves significant performance improvement in both precise positioning and pathological discrimination of gastrointestinal tumors. In addition, the application compares two schemes proposed by the application, SFGNet has less parameter quantity and faster test time, and the multi-task cooperative network based on EMGNet achieves better performance.
[0118] The above merely describes preferred embodiments of the application and is not intended to limit the application, and any modifications, equivalent replacements and improvements made within the spirit and principle of the application shall fall within the scope of the application.
Claims
1. A CT image segmentation and classification system based on class activation map guidance, characterized by: Construct a class activation map-guided enhanced mask generation network, build a collaborative module between the boundary generator and the mask generator, realize the global interaction of deep encoding features through the group shuffle attention module, and fuse the boundary semantic information in the encoding stage with the global context features in the decoding stage; First, the coarse segmentation sub-network performs preliminary segmentation of the tumor area. The generated coarse segmentation mask is used as spatial prior information. The mask and the original image are then input into the sub-classification network. The feature attention mechanism guides the classification network to focus on the lesion area. The sub-classification network generates a gradient-weighted class activation map, and the back-propagation algorithm is used to extract a spatially discriminative tumor localization map. The original image and the corresponding tumor localization map are then input into the fine segmentation sub-network to achieve segmentation of the lesion area. The boundary-aware enhancement network consists of a feature encoder, a segmentation decoder, a global spatial attention module, and a mask-boundary guidance module. The basic segmentation framework consists of an encoder-decoder network consisting of a feature encoder and a segmentation decoder. The tumor segmentation results generated by this network provide spatial prior information for the sub-classification network. The sub-classification network uses the Xception architecture as the feature extraction backbone network and proposes a multi-stage feature fusion strategy to integrate coarse segmentation mask information and extract multi-scale features. The sub-classification network achieves feature map scale alignment through adaptive upsampling and convolution operations. A multi-task collaborative architecture guided by class activation maps achieves joint optimization of gastric tumor localization and classification through a collaborative training mechanism, and constructs a composite loss function by jointly optimizing Dice loss and binary cross entropy loss. The cross entropy loss function is used to measure the difference between the predicted results and the true labels. By minimizing the distribution difference between the predicted probability and the true label, the network is guided to learn discriminative pathological feature representations.
2. The CT image segmentation and classification system based on class activation map guidance according to claim 1, characterized in that: In the coarse segmentation sub-network, the feature encoder uses the feature extraction part of the ResNet34 network, removes the underlying global pooling layer and fully connected layer, and retains the feature extraction modules of the first five stages; In the coarse segmentation sub-network, the output feature maps of the five stages of the encoder are recorded as , whose spatial resolution is the input image ; The segmentation decoder corresponding to the feature encoder contains four decoding blocks and a segmentation head. The decoding block is upsampled by deconvolution operation and channel-concatenated with the jump connection features, and then two 3×3 convolution stacked decoding blocks are used to generate decoding features. , the decoded features are output as segmentation prediction maps after 1×1 convolution and Sigmoid activation; The collaborative module designs a dual-path supervision mechanism: in the encoding path, the boundary generator guides the network to focus on multi-level encoding features through deep supervision boundary information; in the decoding path, the mask generator optimizes the multi-level decoding features through deep supervision The regional characteristics of are shown in formulas (1) and (2): (1) ; Where, represents multi-level encoding features, represents multi-level decoding features, is the Sigmoid function, conv(·) represents a 1×1 convolution, ⊙ represents the element-wise multiplication operation, and They represent the boundary-enhanced encoding features and region-enhanced decoding features, respectively.
3. The CT image segmentation and classification system based on class activation map guidance according to claim 2, characterized in that: The sub-classification network achieves feature map scale alignment through adaptive upsampling and convolution operations. The fusion process is expressed as: ; Where, Indicates channel cascade operation, and Represents 3×3 convolution and upsampling respectively; The fusion features With coarse segmentation mask Perform residual processing and guide the classification results of the coarse segmentation mask by focusing on the characteristics of the tumor area. After the global average pooling layer, the fully connected layer and the softmax function are processed, the final classification results are output. , the specific formula is as follows: (4) Where, represents the global average pooling layer, represents the fully connected layer, represents the element-wise product operation, Represents element-wise addition operation.
4. The CT image segmentation and classification system based on class activation map guidance according to claim 3, characterized in that: The input of the fine segmentation sub-network is the original image data and the CAM generated by the classification branch, and the two achieve information complementarity through feature fusion; In the feature preprocessing stage, the nearest neighbor interpolation algorithm is used to upsample the CAM. After normalization by the sigmoid function, the CAM is converted into The spatial attention weight map of the interval guides the multi-level decoding features to focus on the region of interest and establishes a residual connection between the original features and the attention-weighted features.
5. The CT image segmentation and classification system based on class activation map guidance according to claim 4, characterized in that: The Dice loss calculates the pixel-level similarity between the predicted result and the true segmentation mask, and the BCE loss constrains pixel-by-pixel classification. Based on the deep supervision strategy, the network has three lateral outputs. and four edge outputs All are upsampled and the corresponding edges Keeping spatial alignment, the total segmentation loss consists of the following two parts: (5) Where, and Represent the upsampled side output feature map and edge feature map respectively.
6. The CT image segmentation and classification system based on class activation map guidance according to claim 5, characterized in that: The cross entropy loss function is used to measure the difference between the predicted result and the true label. The expression of the cross entropy loss function is shown in formula (6): (6) Where, represents the total number of categories for the classification task, Indicates that the sample belongs to The true label of the class, Indicates that the model predicts that the sample belongs to The probability of the class.
Citation Information
Cited By
Prostate magnetic resonance image segmentation method and system based on multi-level context aggregation
CN121305090A
Intelligent karst groundwater exploration image processing system
CN121962779A
Bidirectional position guided medical image segmentation and classification method
CN122199981A
Construction of brain glioma subregion segmentation model based on multi-modal edge feature fusion
CN122223030A
Difficulty airway evaluation system and device using MRI image in combination with deep learning algorithm
CN122244025A