Cross-resolution collaborative tumor segmentation method and device based on attention guidance
By adopting a cross-resolution collaborative network based on attention guidance in medical image segmentation, the problem of tumor segmentation difficulty in CT images is solved, and high-precision and automated tumor segmentation is achieved, reducing the burden of doctors' labeling, and improving the accuracy of segmentation results.
Patent Information
- Application Number
- CN202510014566.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-06-03
AI Technical Summary
Existing medical image segmentation technology is difficult to effectively process tumors in CT images, especially when the tumor has low contrast with surrounding tissue and is morphologically heterogeneous, and the automatic segmentation technology relies on complex models and high computing resources.
A cross-resolution collaborative network based on attention guidance is adopted, and through multi-attention fusion modules and cross-resolution fusion modules, combined with channel, space and global attention mechanisms, deep feature extraction and cross-resolution fusion of feature maps are carried out to achieve end-to-end accurate 3D detection and segmentation of tumors.
It realizes automated, end-to-end, and high-precision tumor segmentation, which reduces the burden of manual labeling by doctors, enhances the model's ability to capture tumor characteristics, alleviates oversegmentation, and improves the accuracy of segmentation results.
Smart Images

Figure CN120088472A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing and deep learning, and particularly relates to a tumor segmentation method and device based on attention-guided cross-resolution collaboration. Background Art
[0002] Medical image tumor segmentation is an important research direction in the field of medical image processing and plays a crucial role in medical diagnosis, treatment, and research; computer tomography (CT) is widely used in tumor diagnosis due to its high density resolution and fast imaging speed; therefore, accurate tumor segmentation of CT images can provide support for subsequent research such as metastasis prediction and prognosis prediction. However, manually annotating tumors layer by layer is very time-consuming and relies on the experience of radiologists, so automatically and accurately segmenting tumors in CT scans can effectively assist clinical diagnosis and result prediction. From a clinical perspective, there are mainly the following two difficulties in tumor segmentation: one is that the contrast between tumors and surrounding tissues is low and the boundaries are unclear; the other is that due to the heterogeneity of tumors, the morphology, size, and location of tumors vary among individuals. In recent years, deep learning technology has achieved great success in the field of medical image segmentation. In particular, deep learning models developed based on UNet, such as 3D-UNet, Attention UNet, and nnUNet, have achieved remarkable results in various 3D medical image segmentation tasks. However, it should be noted that these models are all "generalist models" and are not specifically designed for the tumor characteristics in medical images. The attention mechanism was initially designed to mimic human attention and encourage deep learning models to focus on specific regions or targets, and is now widely used in medical image segmentation tasks. For example, AnatomyNet utilizes channel attention in its 3D squeeze-and-excitation residual blocks (SE-Blocks) to dynamically adjust the weights of channels in the feature map and enhance the feature representation ability; Chen et al. combined self-attention and global spatial attention to learn the non-local interactions between encoder features and focused on leveraging key spatial information. However, these methods ignore the correlation between pixels. Peng et al. proposed RPAN, which generates a rescaling factor for each pixel and learns its importance by considering the interdependence between pixels in the feature map. In addition, Attention UNet introduced an attention gate, which uses the self-attention mechanism to generate an attention score map between the encoder and the decoder to help the model focus on important features. However, most of these methods only use a single form of attention mechanism, and there is still room for development in using multiple attention mechanisms to capture local and global information simultaneously to enhance the feature representation ability.
[0003] For medical image segmentation tasks, since the sizes and scales of tumors usually vary significantly, many studies have integrated information from different resolutions to assist segmentation. For example, Gerard et al. proposed a multi-resolution model that uses a cascaded network from low resolution to high resolution, eliminating the need to trade off between detailed information and global context. UNet++ enables the interaction of high-resolution and low-resolution information during the encoding process by using multi-layer connections and multiple sub-U-Net structures. Sun et al. proposed the SAUNet model, which constructs a shape feature assistance stream that operates in parallel with the backbone network and perceives and processes the contour and texture information of the target by continuously integrating the encoded feature information at different resolutions. Sinha et al. introduced a multi-scale guidance network that aggregates feature maps at different resolutions into global features, and then integrates local features and global information, using an attention mechanism to capture richer dependencies. However, when fusing feature maps at different resolutions, these methods do not fully consider the scale consistency between feature maps, which may lead to feature distortion or information loss.
[0004] UNet cannot fully exploit the valuable information from features at different scales. A common solution is to use a deep supervision strategy to generate independent segmentation results for each layer, enabling the gradient to propagate between different layers. This effectively utilizes multi-scale information and alleviates the problem of gradient vanishing; however, it also requires higher computational and storage requirements. Zhang et al. proposed a phased deep supervision method that integrates feature information from two different phases to improve features, avoiding generating multiple segmentation results and saving computational resources. Dou et al. proposed a phased attention refinement method that uses an attention mechanism to improve the decoded features and directly fuses them for cortical plate segmentation. Zhang et al. proposed a supervision module that aggregates sufficient boundary information in low-level encoded features with multi-scale decoded features to improve the final segmentation result. Gu et al. compressed the decoded feature maps into single-channel images and combined them, and then used an attention mechanism to further improve the fused features. When improving the above methods, channel compression of the feature maps is often selected before fusion. Although it alleviates the computational and storage pressure, compressing the channels of the feature maps may lead to the loss of valuable information. Summary of the Invention
[0005] The main objective of the present invention is to overcome the drawbacks and deficiencies of the prior art, and provide a tumor segmentation method and device based on attention-guided cross-resolution collaboration. By constructing an attention-guided cross-resolution collaboration network, it is used for end-to-end precise 3D detection and segmentation of tumors in CT images, facilitating a series of applications such as subsequent computer-aided diagnosis and treatment, computer-aided surgery, and clinical research, so as to reduce the burden on clinicians of manually annotating the lesion area and judging the segmentation layer by layer.
[0006] To achieve the above object, on the one hand, the present invention adopts a tumor segmentation method based on attention-guided cross-resolution collaboration, including the following steps:
[0007] Obtain the original CT image and its corresponding pixel-level label, resample the original CT image and crop it into patches of a set size;
[0008] Construct an attention-guided cross-resolution collaboration network, including multiple multi-attention fusion modules, multiple cross-resolution fusion modules, and a scale-aware activation module; the multiple multi-attention fusion modules are connected in sequence, and each multi-attention fusion module includes an encoding layer, three attention modules of channel, spatial, and global, a weighted fusion module, and a downsampling module; the multiple cross-resolution fusion modules are connected in sequence, and each cross-resolution fusion module includes a resolution feature extraction module, a resolution feature fusion module, a feature adjustment module, and a convolutional output module; the multiple multi-attention fusion modules are skip-connected to the multiple cross-resolution fusion modules, and each cross-resolution fusion module is connected to the corresponding skip-connected multi-attention fusion module and its adjacent multi-attention fusion module; the scale-aware activation module includes an upsampling fusion module, a scale attention module, and a classification module connected in sequence; the upsampling fusion module is connected to the multiple cross-resolution fusion modules; the number of multi-attention fusion modules is the same as that of cross-resolution fusion modules;
[0009] Use the backpropagation algorithm to train the attention-guided cross-resolution collaboration network, and the training process is as follows:
[0010] After simple feature extraction of the patch through the convolutional backbone, input it into multiple multi-attention fusion modules for deep feature extraction to obtain multiple encoded feature maps; input the multiple encoded feature maps into multiple cross-resolution fusion modules for feature map reduction and decoding to obtain multiple decoded feature maps; send the multiple decoded feature maps into the scale-aware activation module and use the pixel-level label corresponding to the original CT image for supervision to obtain the predicted segmentation result; calculate the loss function to update the network weights in the reverse direction and use the SGD optimizer to optimize the network parameters until the loss function converges or reaches the maximum number of training times to obtain the trained attention-guided cross-resolution collaboration network;
[0011] Collect the CT image to be detected and input it into the trained attention-guided cross-resolution collaboration network for tumor segmentation.
[0012] As a preferred technical solution, each multi-attention fusion module first extracts encoded features from the input encoded feature map through an encoding layer, and then calculates the attention scores under three attention mechanisms respectively; then, the attention scores under the three attention mechanisms, the encoded features, and the input encoded feature map are merged in a weighted fusion module, and then the encoded feature map of this multi-attention fusion module is obtained through a downsampling module;
[0013] The attention scores include channel-level attention scores, spatial-level attention scores, and global attention scores.
[0014] As a preferred technical solution, the channel-level attention score is expressed as:
[0015]
[0016] where is the channel-level attention score of the i-th multi-attention fusion module, σ is the Sigmoid function, CRC is the convolution-ReLU activation-convolution operation, Conv is the convolution operation, is the element-wise addition, Pavg c is the average pooling operation in the channel dimension, Pmax c is the max pooling operation in the channel dimension, N is the number of multi-attention fusion modules; F i-1 is the encoded feature map output by the (i - 1)-th multi-attention fusion module, and when i = 1, F 0 is the encoded feature map obtained by the patch for simple feature extraction through the convolution backbone;
[0017] The spatial-level attention score is expressed as:
[0018]
[0019] where is the spatial-level attention score of the i-th multi-attention fusion module, Pavg s is the average pooling operation on the spatial level, Pmax s is the max pooling operation on the spatial level, is the concatenation operation in the channel dimension;
[0020] The global attention score is expressed as:
[0021]
[0022] where is the global attention score of the i-th multi-attention fusion module;
[0023] The encoded features of the multi-attention fusion module are expressed as:
[0024]
[0025] Among them, represents element-wise multiplication, and Pmax is the downsampling module, that is, the max pooling operation.
[0026] As a preferred technical solution, each cross-resolution fusion module first uses a resolution feature extraction module to align the sizes and channels of the encoded feature maps of the multi-attention fusion module corresponding to its skip connection and its adjacent attention fusion modules and extract features to obtain a resolution attention map, and then inputs it into the resolution fusion module for fusion to obtain a cross-resolution fusion feature; the resolution attention map includes a high-resolution attention map and a low-resolution attention map;
[0027] Next, the input decoded feature map is adjusted in size and channels by the feature adjustment module, and after being fused with the cross-resolution fusion feature, the decoded feature map of this cross-resolution fusion module is obtained through the convolutional output module.
[0028] As a preferred technical solution, the high-resolution attention map is expressed as:
[0029]
[0030] Among them, is the high-resolution attention map of the cross-resolution fusion module corresponding to the processing of F i , F i is the encoded feature map of the i-th multi-attention fusion module, represents element-wise multiplication, σ is the Sigmoid function, Pmax is the max pooling operation, Conv is the convolutional operation, is element-wise addition, M is the number of cross-resolution fusion modules, F i-1 is the encoded feature map of the previous multi-attention fusion module of the i-th multi-attention fusion module. When i = 1, F 0 is the encoded feature map obtained by simple feature extraction of the patch through the convolutional main branch;
[0031] The low-resolution attention map is expressed as:
[0032]
[0033] Among them, F i l is the low-resolution attention map of the cross-resolution fusion module corresponding to the processing of F i , UP is the interpolation upsampling operation, and F i+1 is the encoded feature map of the next multi-attention fusion module of the i-th multi-attention fusion module;
[0034] The cross-resolution fusion feature is expressed as:
[0035]
[0036] Among them, F′ i is the cross-resolution fusion feature of the i-th cross-resolution fusion module, is the concatenation operation in the channel dimension, and ReLU is the ReLU activation function;
[0037] The decoded feature map of the cross-resolution fusion module is expressed as:
[0038]
[0039] Among them, D i is the decoded feature map of the i-th cross-resolution fusion module, D i-1 is the decoded feature map of the previous cross-resolution fusion module of the i-th cross-resolution fusion module. When i = 1, D 0 is the encoded feature map F of the M-th multi-attention fusion module M .
[0040] As a preferred technical solution, the process by which the scale-aware activation module obtains the predicted segmentation result is as follows:
[0041] Input multiple decoded feature maps into the scale-aware activation module, and use the upsampling fusion module to perform convolution and interpolation upsampling operations on the decoded feature maps respectively to unify the scale size and merge them to obtain a mixed feature map;
[0042] Use the scale attention module to process the mixed feature map, calculate the scale attention result and the spatial attention result on the channel dimension and the spatial level, and then obtain the mixed attention map;
[0043] Based on the mixed feature map and the mixed attention map, the classification module obtains the pixel-level predicted segmentation result under the supervision of the pixel-level label corresponding to the original CT image.
[0044] As a preferred technical solution, the mixed feature map is expressed as:
[0045]
[0046] Among them, M is the number of cross-resolution fusion modules, Conv is the convolution operation, UP is the interpolation upsampling operation, is the concatenation operation in the channel dimension, D i is the decoded feature map of the i-th cross-resolution fusion module;
[0047] The mixed attention map is expressed as:
[0048]
[0049] Among them, represents element-wise multiplication, σ is the Sigmoid function, Pavg s is the average pooling operation in the spatial dimension, Pmax s is the maximum pooling operation in the spatial dimension, MLP is a multi-layer perceptron, D′ ca is the scale attention result, expressed as:
[0050]
[0051] Among them, is element-wise addition, Pavg c is the average pooling operation in the channel dimension, Pmax c is the maximum pooling operation in the channel dimension, MLP s is a shared multi-layer perceptron;
[0052] The predicted segmentation result at the pixel level is expressed as:
[0053] Pred = SofMax(BN(Conv(D′ + D SA ))),
[0054] Among them, SofMax is the SofMax classification function, and BN is batch normalization.
[0055] As a preferred technical solution, the loss function L seg adopts cross-entropy loss L CE and Dice loss L DC , expressed as:
[0056] L seg = L CE + L DC ,
[0057] The cross-entropy loss L CE is expressed as:
[0058]
[0059] The Dice loss L DC is expressed as:
[0060]
[0061] Among them, is the number of categories, HWD is the size of the original CT image; I(·) is the indicator function, when the true label of voxel t is category n, the function value is 1, otherwise it is 0; Y t is the true label of voxel t, Pred t,n is the segmentation probability that the t-th voxel belongs to category n.
[0062] On the other hand, a tumor segmentation system based on attention-guided cross-resolution collaboration is provided, including a data processing module, a network construction module, a network training module, and a tumor segmentation module;
[0063] The data processing module is used to obtain the original CT image and its corresponding pixel-level label, resample the original CT image and crop it into patches of a set size;
[0064] The network construction module is used to construct an attention-guided cross-resolution collaboration network, including multiple multi-attention fusion modules, multiple cross-resolution fusion modules, and a scale-aware activation module; the multiple multi-attention fusion modules are connected in sequence, and each multi-attention fusion module includes an encoding layer, three attention modules of channel, spatial, and global, a weighted fusion module, and a downsampling module; the multiple cross-resolution fusion modules are connected in sequence, and each cross-resolution fusion module includes a resolution feature extraction module, a resolution feature fusion module, a feature adjustment module, and a convolutional output module; the multiple multi-attention fusion modules are skip-connected to the multiple cross-resolution fusion modules, and each cross-resolution fusion module is connected to the corresponding skip-connected multi-attention fusion module and its adjacent multi-attention fusion module; the scale-aware activation module includes an upsampling fusion module, a scale attention module, and a classification module connected in sequence; the upsampling fusion module is connected to the multiple cross-resolution fusion modules; the number of multi-attention fusion modules is the same as that of cross-resolution fusion modules;
[0065] The network training module is used to train the attention-guided cross-resolution collaboration network using the backpropagation algorithm. The training process is as follows:
[0066] After simple feature extraction of the patches through the convolutional backbone, input them into multiple multi-attention fusion modules for deep feature extraction to obtain multiple encoded feature maps; input the multiple encoded feature maps into multiple cross-resolution fusion modules for feature map reduction and decoding to obtain multiple decoded feature maps; send the multiple decoded feature maps into the scale-aware activation module and supervise them using the pixel-level label corresponding to the original CT image to obtain the predicted segmentation result; calculate the loss function, update the network weights in reverse, and use the SGD optimizer to optimize the network parameters until the loss function converges or reaches the maximum number of training times to obtain the trained attention-guided cross-resolution collaboration network;
[0067] The tumor segmentation module is used to collect the CT image to be detected and input it into the trained attention-guided cross-resolution collaboration network for tumor segmentation.
[0068] On the other hand, a computer-readable storage medium is provided, storing a program which, when executed by a processor, implements the above-mentioned tumor segmentation method based on attention-guided cross-resolution collaboration.
[0069] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0070] 1. Due to the particularity of medical images, generally only radiologists and those with medical clinical knowledge can judge tumors based on images and perform pixel-level annotation. However, based on the image features of tumors, the present invention specifically designs an attention-guided cross-resolution collaboration network, realizing automated, end-to-end, and high-precision tumor segmentation, greatly reducing the burden on doctors to annotate images layer by layer.
[0071] 2. The multi-attention fusion module in the attention-guided cross-resolution collaboration network of the present invention adopts three different levels of attention mechanisms, namely channel, spatial, and global, to comprehensively capture distal context information, thereby enhancing the model's focus on the segmentation target while reducing attention to irrelevant regions to alleviate the over-segmentation phenomenon in tumor segmentation.
[0072] 3. The cross-resolution fusion module in the attention-guided cross-resolution collaboration network of the present invention mines, fuses, and refines features by obtaining adjacent high-resolution detail information and low-resolution semantic information of feature maps at different levels to learn tumor features of different sizes and shapes, enhancing the network's semantic perception ability.
[0073] 4. The scale-aware activation module in the attention-guided cross-resolution collaboration network of the present invention adaptively extracts and integrates specific features from decoded feature maps at different scales, rather than simply performing weighted summation on features at each stage, to refine tumor features and promote more accurate segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0075] Figure 1 It is a flow framework diagram of the tumor segmentation method based on attention-guided cross-resolution collaboration in the embodiments of the present invention.
[0076] Figure 2 It is a structural schematic diagram of the tumor segmentation system based on attention-guided cross-resolution collaboration in the embodiments of the present invention.
[0077] Figure 3 This is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present invention. Detailed implementation manners
[0078] In order to enable those skilled in the art of this technology to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of this application.
[0079] Referring to "embodiments" in this application means that specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in this application may be combined with other embodiments.
[0080] As Figure 1 shown, the tumor segmentation method based on attention-guided cross-resolution collaboration in this embodiment includes the following steps:
[0081] S1. Obtain the original CT image and its corresponding pixel-level label, and resample and crop the original CT image into patches of a set size.
[0082] In this embodiment, first, the original CT image is resampled to a space of 0.8625 mm × 0.8625 mm × 2 mm, and then it is cropped into patches of 256 × 256 × 96 voxels; the corresponding pixel-level annotation label is a matrix of the same size as the original CT image, and the value of each element in the matrix is 0 and 1 respectively. Different values represent different semantic classes to which the corresponding pixels belong, where 0 represents the background and 1 represents the tumor.
[0083] S2. Construct an attention-guided cross-resolution collaborative network, which includes multiple multi-attention fusion modules (MAF), multiple cross-resolution fusion modules (CRF), and a scale-aware activation module (SAA). Among them, multiple MAF are connected in sequence. Each MAF includes an encoding layer, three attention modules of channel, spatial, and global, a weighted fusion module, and a downsampling module. Multiple CRF are connected in sequence. Each CRF includes a resolution feature extraction module, a resolution feature fusion module, a feature adjustment module, and a convolutional output module. Multiple MAF and multiple CRF are skip-connected. Each CRF is connected to the corresponding skip-connected MAF and its adjacent MAF. The scale-aware activation module SAA includes an upsampling fusion module, a scale attention module, and a classification module connected in sequence. The upsampling fusion module is connected to multiple CRF. The number of MAF and CRF is the same.
[0084] In this embodiment, as Figure 1 shown, it includes four MAF and four CRF. If numbered in order, the first MAF is skip-connected to the fourth CRF, and so on. The fourth MAF is skip-connected to the first CRF.
[0085] S3. Use the backpropagation algorithm to train the attention-guided cross-resolution collaborative network. The training process is as follows:
[0086] After simple feature extraction of the patch through the convolutional backbone, input it into multiple MAF for deep feature extraction to obtain multiple encoded feature maps. Input multiple encoded feature maps into multiple CRF for feature map reduction and decoding to obtain multiple decoded feature maps. Send multiple decoded feature maps into the scale-aware activation module and supervise them with the pixel-level labels corresponding to the original CT image to obtain the predicted segmentation result. Calculate the loss function to update the network weights in reverse and use the SGD optimizer to optimize the network parameters until the loss function converges or reaches the maximum number of training times to obtain the trained attention-guided cross-resolution collaborative network.
[0087] In this embodiment, the initial learning rate of the SGD optimizer is 0.01, the learning rate decay strategy adopts cosine annealing decay, the batch data size (Batch-size) in the training process is 2, and the iterative training is 200,000 times.
[0088] S4. Collect the CT image to be detected and input it into the trained attention-guided cross-resolution collaborative network for tumor segmentation.
[0089] Further, one of the difficulties in segmenting tumors based on CT images is the low contrast between tumors and normal tissues. Therefore, a multi-attention fusion module MAF is designed in the attention-guided cross-resolution collaborative network constructed in this application. It collaborates through three types of attention mechanisms to capture distal context information, enhance the encoding ability of the module to more comprehensively learn the characteristics of tumors. Specifically:
[0090] Each multi-attention fusion module MAF first extracts encoded features from the input encoded feature map through an encoding layer, and then calculates the attention scores under three attention mechanisms respectively, including channel-level attention scores, spatial-level attention scores, and global attention scores. Among them, the calculation process of the channel-level attention score is described as:
[0091]
[0092] Among them, is the channel-level attention score of the i-th multi-attention fusion module, σ is the Sigmoid function (in the figure ), CRC is the convolution-ReLU activation-convolution operation, Conv is the convolution operation, is the element-wise addition, Pavg c is the average pooling operation on the channel dimension, Pmax c is the max pooling operation on the channel dimension, N is the number of multi-attention fusion modules (in this embodiment, N = 4); F i-1 is the encoded feature map output by the (i - 1)-th multi-attention fusion module. When i = 1, F 0 is the encoded feature map obtained by the patch through simple feature extraction by the convolution stem branch.
[0093] The calculation process of the spatial-level attention score is described as:
[0094]
[0095] Among them, is the spatial-level attention score of the i-th multi-attention fusion module, Pavg s is the average pooling operation on the spatial level, Pmax s is the max pooling operation on the spatial level, is the concatenation operation on the channel dimension.
[0096] The calculation process of the global attention score is described as:
[0097]
[0098] Among them, is the global attention score of the i-th multi-attention fusion module.
[0099] Then, the attention scores, encoded features under the three attention mechanisms and the input encoded feature map are merged in the weighted fusion module. The merging is performed through residual connection to alleviate the problem of gradient disappearance, and then the encoded feature map of this multi-attention fusion module is obtained through the downsampling module, which is expressed as:
[0100]
[0101] Among them, represents element-wise multiplication, Pmax is the downsampling module, that is, the downsampling operation of convolution-max pooling.
[0102] Furthermore, during the encoding process, the feature map gradually changes from a high resolution carrying more detailed information to a low resolution carrying more high-level semantic information; in order to take into account both details and tumor semantic information, this application introduces multiple cross-resolution fusion modules CRF to learn tumor features of different sizes and shapes. Specifically:
[0103] Each cross-resolution fusion module CRF first uses the resolution feature extraction module to align the sizes and channels of the encoded feature maps of the multi-attention fusion module MAF corresponding to its skip connection and its adjacent attention fusion modules, so as to obtain as much detailed information as possible; then, the resolution attention map is extracted and input into the resolution fusion module for fusion to obtain the cross-resolution fusion feature; among them, the resolution attention map includes the high-resolution attention map and the low-resolution attention map.
[0104] For the high-resolution attention map, its acquisition process is described as:
[0105]
[0106] Among them, is the high-resolution attention map of the cross-resolution fusion module corresponding to the processing F i , F i is the encoded feature map of the i-th multi-attention fusion module, represents element-wise multiplication, σ is the Sigmoid function, Pmax is the max pooling operation, Conv is the convolution operation, is element-wise addition, M is the number of cross-resolution fusion modules (M = 4 in this embodiment), F i-1 is the encoded feature map of the previous multi-attention fusion module of the i-th multi-attention fusion module. When i = 1, F 0 is the encoded feature map obtained by the patch for simple feature extraction through the convolution stem.
[0107] For the low-resolution features that retain relatively more semantic information extracted, direct addition may lead to boundary confusion. Therefore, element-wise multiplication is used to perceive high-level context information. Thus, the process of obtaining the low-resolution attention map is described as follows:
[0108]
[0109] where F i l is the low-resolution attention map that has fused deep semantic information in the i-th cross-resolution fusion module, UP is the interpolation upsampling operation used to enlarge the size of the low-resolution features, and F i+1 is the encoded feature map of the next multi-attention fusion module in the i-th multi-attention fusion module.
[0110] For the cross-resolution fusion features, its fusion process is described as follows:
[0111]
[0112] where F′ i is the cross-resolution fusion feature of the i-th cross-resolution fusion module, is the concatenation operation in the channel dimension, and ReLU is the ReLU activation function.
[0113] Finally, the input decoded feature map is resized and adjusted in channels by the feature adjustment module, and after fusing with the cross-resolution fusion features, it passes through the convolutional output module to obtain the decoded feature map of this cross-resolution fusion module, denoted as:
[0114]
[0115] where D i is the decoded feature map of the i-th cross-resolution fusion module, D i-1 is the decoded feature map of the previous cross-resolution fusion module of the i-th cross-resolution fusion module. When i = 1, D 0 is the encoded feature map F of the M-th multi-attention fusion module M .
[0116] After cross-resolution fusion, not only the low-resolution detailed information of the tumor is obtained, but also more advanced semantic information is learned, fully mining and refining the tumor features, and ensuring the classification accuracy.
[0117] Furthermore, as a deep structure, the traditional U-shaped network only uses the decoding result of the last layer for segmentation prediction and loss calculation, which has the potential risk of gradient disappearance. Therefore, this application combines all the feature information generated by decoding for the final segmentation prediction. In order to selectively extract and merge useful information for more accurate segmentation prediction for decoding feature maps of different scales, this application designs a scale-aware activation module SAA, and its process of predicting segmentation is as follows:
[0118] First, input multiple decoding feature maps into the scale-aware activation module SAA, and use the upsampling fusion module to perform convolution and interpolation upsampling operations on the decoding feature maps respectively to unify the scale size and merge them to obtain a mixed feature map, which is described as:
[0119]
[0120] where M is the number of cross-resolution fusion modules, Conv is the convolution operation, UP is the interpolation upsampling operation, is the concatenation operation for the channel dimension, D i is the decoding feature map of the i-th cross-resolution fusion module.
[0121] The obtained mixed feature map D' has the same size as the decoding feature map D M of the M-th cross-resolution fusion module, but the number of channels of D' is M times that of D M ; therefore, use the scale attention module to process the mixed feature map, calculate the scale attention result and the spatial attention result on the channel dimension and the spatial level to obtain the mixed attention map, and refine the decoding feature by adaptively extracting and activating the semantically relevant information in each scale.
[0122] Specifically, the scale attention module first processes the mixed feature map D' based on the shared multi-layer perceptron (MLP s ) in two forms of information representation: global maximum and average pooling, and then integrates it to obtain the scale confidence score, and then combines it with the mixed feature map D' by weighting to obtain the scale attention result D′ sca , which is expressed as:
[0123]
[0124] where, is the element-wise addition, Pavg c is the average pooling operation on the channel dimension, Pmax c is the maximum pooling operation on the channel dimension, MLP s is the shared multi-layer perceptron;
[0125] Subsequently, the scale attention results are further processed at the spatial level to obtain the global spatial information representation and calculate the spatial confidence score to calculate the final result of the scale attention module, and then the hybrid attention map is obtained, which is expressed as:
[0126]
[0127] Among them, represents element-wise multiplication, σ is the Sigmoid function, Pavg s is the average pooling operation at the spatial level, Pmax s is the max pooling operation at the spatial level, and MLP is the multi-layer perceptron.
[0128] Finally, based on the hybrid feature map D' and the hybrid attention map D SA , the classification module obtains the pixel-level prediction segmentation result under the supervision of the pixel-level label corresponding to the original CT image, which is expressed as:
[0129] Pred = SofMax(BN(Conv(D′ + D SA ))),
[0130] Among them, SofMax is the SofMax classification function, and BN is batch normalization.
[0131] Furthermore, the loss function L in the training process of the attention-guided cross-resolution collaborative network in this application seg adopts the cross-entropy loss L CE and the Dice loss L DC , which is expressed as:
[0132] L seg = L CE + L DC ,
[0133] Among them, the cross-entropy loss L CE is expressed as:
[0134]
[0135] The Dice loss L DC is expressed as:
[0136]
[0137] In the formula, is the number of categories (i.e., the categories are tumor and background), HWD is the size of the original CT image; I(·) is the indicator function, when the true label of voxel t is category n, the function value is 1, otherwise it is 0; Y t is the true label of voxel t, Pred t ,nis the segmentation probability that the t-th voxel belongs to the category n.
[0138] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, some steps can be performed in other sequences or simultaneously.
[0139] Based on the same idea as the attention-guided cross-resolution collaborative tumor segmentation method in the above embodiments, the present invention also provides an attention-guided cross-resolution collaborative tumor segmentation system, which can be used to execute the above attention-guided cross-resolution collaborative tumor segmentation method. For the sake of convenience of description, in the structural schematic diagram of the attention-guided cross-resolution collaborative tumor segmentation system embodiment, only the parts related to the embodiments of the present invention are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and may include more or fewer components than those illustrated, or combine some components, or arrange different components.
[0140] As Figure 2 shown, another embodiment of the present invention provides an attention-guided cross-resolution collaborative tumor segmentation system, including a data processing module, a network construction module, a network training module, and a tumor segmentation module;
[0141] Among them, the data processing module is used to obtain the original CT image and its corresponding pixel-level label, resample the original CT image and crop it into patches of a set size;
[0142] The network construction module is used to construct an attention-guided cross-resolution collaborative network, including a plurality of multi-attention fusion modules, a plurality of cross-resolution fusion modules, and a scale-aware activation module; the plurality of multi-attention fusion modules are connected in sequence, and each multi-attention fusion module includes an encoding layer, a channel, a spatial, and a global attention module, a weighted fusion module, and a downsampling module; the plurality of cross-resolution fusion modules are connected in sequence, and each cross-resolution fusion module includes a resolution feature extraction module, a resolution feature fusion module, a feature adjustment module, and a convolutional output module; the plurality of multi-attention fusion modules are skip-connected to the plurality of cross-resolution fusion modules, and each cross-resolution fusion module is connected to the corresponding skip-connected multi-attention fusion module and its adjacent multi-attention fusion module; the scale-aware activation module includes an upsampling fusion module, a scale attention module, and a classification module connected in sequence; the upsampling fusion module is connected to the plurality of cross-resolution fusion modules; the number of multi-attention fusion modules is the same as the number of cross-resolution fusion modules;
[0143] The network training module is used to train the attention-guided cross-resolution collaborative network using the back-propagation algorithm. The training process is as follows:
[0144] After the patch is subjected to simple feature extraction through convolution, it is input into multiple multi-attention fusion modules for deep feature extraction to obtain multiple encoding feature maps; multiple encoding feature maps are input into multiple cross-resolution fusion modules for feature map restoration decoding to obtain multiple decoding feature maps; multiple decoding feature maps are sent into the scale-aware activation module and supervised using the pixel-level labels corresponding to the original CT image to obtain the predicted segmentation result; the loss function is calculated to reversely update the network weights and the SGD optimizer is used to optimize the network parameters until the loss function converges or the maximum number of training times is reached to obtain a trained attention-guided cross-resolution collaborative network;
[0145] The tumor segmentation module is used to collect the CT images to be detected and input them into the trained attention-guided cross-resolution collaborative network for tumor segmentation.
[0146] It should be noted that the attention-guided cross-resolution collaborative tumor segmentation system of the present invention corresponds one-to-one to the attention-guided cross-resolution collaborative tumor segmentation method of the present invention. The technical features and beneficial effects described in the above-mentioned embodiment of the attention-guided cross-resolution collaborative tumor segmentation method are applicable to the embodiment of the attention-guided cross-resolution collaborative tumor segmentation system. For specific contents, please refer to the description in the embodiment of the method of the present invention, which will not be repeated here. This is hereby declared.
[0147] In addition, in the implementation of the attention-guided cross-resolution collaborative tumor segmentation system in the above-mentioned embodiment, the logical division of each program module is only an example. In actual applications, the above-mentioned functions can be assigned to different program modules as needed, for example, for the convenience of corresponding hardware configuration requirements or software implementation. That is, the internal structure of the attention-guided cross-resolution collaborative tumor segmentation system is divided into different program modules to complete all or part of the functions described above.
[0148] like Figure 3 As shown, in one embodiment, a computer-readable storage medium is provided, in which a program is stored in a memory, and when the program is executed by a processor, the method for tumor segmentation based on attention-guided cross-resolution collaboration is implemented, specifically:
[0149] Get the original CT image and its corresponding pixel-level labels, resample the original CT image and crop it into patches of a set size;
[0150] Construct an attention-guided cross-resolution collaborative network, including multiple multi-attention fusion modules, multiple cross-resolution fusion modules, and a scale-aware activation module; among them, the multiple multi-attention fusion modules are connected in sequence, and each multi-attention fusion module includes an encoding layer, three attention modules of channel, spatial, and global, a weighted fusion module, and a downsampling module; the multiple cross-resolution fusion modules are connected in sequence, and each cross-resolution fusion module includes a resolution feature extraction module, a resolution feature fusion module, a feature adjustment module, and a convolutional output module; the multiple multi-attention fusion modules are skip-connected to the multiple cross-resolution fusion modules, and each cross-resolution fusion module is connected to the corresponding multi-attention fusion module with which it is skip-connected and its adjacent multi-attention fusion modules; the scale-aware activation module includes an upsampling fusion module, a scale attention module, and a classification module connected in sequence; the upsampling fusion module is connected to the multiple cross-resolution fusion modules; the number of multi-attention fusion modules is the same as that of the cross-resolution fusion modules;
[0151] Use the backpropagation algorithm to train the attention-guided cross-resolution collaborative network. The training process is as follows:
[0152] After simple feature extraction of the patch through the convolutional backbone, input it into multiple multi-attention fusion modules for deep feature extraction to obtain multiple encoded feature maps; input the multiple encoded feature maps into multiple cross-resolution fusion modules for feature map reduction and decoding to obtain multiple decoded feature maps; send the multiple decoded feature maps into the scale-aware activation module and supervise them with the pixel-level labels corresponding to the original CT image to obtain the predicted segmentation result; calculate the loss function to update the network weights backward and use the SGD optimizer to optimize the network parameters until the loss function converges or reaches the maximum number of training times to obtain the trained attention-guided cross-resolution collaborative network;
[0153] Collect the CT image to be detected and input it into the trained attention-guided cross-resolution collaborative network for tumor segmentation.
[0154] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0155] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0156] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent substitution methods and are all included in the protection scope of the present invention.
Claims
1. A tumor segmentation method based on attention-guided cross-resolution collaboration, characterized in that: The steps include: Get the original CT image and its corresponding pixel-level labels, resample the original CT image and crop it into patches of a set size; Construct an attention-guided cross-resolution collaborative network, including multiple multi-attention fusion modules, multiple cross-resolution fusion modules and a scale-aware activation module; the multiple multi-attention fusion modules are connected in sequence, and each multi-attention fusion module includes a coding layer, three attention modules of channel, space and global, a weighted fusion module and a downsampling module; the multiple cross-resolution fusion modules are connected in sequence, and each cross-resolution fusion module includes a resolution feature extraction module, a resolution feature fusion module, a feature adjustment module and a convolution output module; the multiple multi-attention fusion modules are jump-connected with multiple cross-resolution fusion modules, and each cross-resolution fusion module is connected with the corresponding jump-connected multi-attention fusion module and its adjacent multi-attention fusion module; The scale-aware activation module includes an upsampling fusion module, a scale attention module, and a classification module connected in sequence; the upsampling fusion module is connected to a plurality of cross-resolution fusion modules; The number of the multi-attention fusion modules is the same as the number of the cross-resolution fusion modules; The attention-guided cross-resolution collaborative network is trained using the back-propagation algorithm. The training process is as follows: After the patch is subjected to simple feature extraction through convolution, it is input into multiple multi-attention fusion modules for deep feature extraction to obtain multiple encoding feature maps; multiple encoding feature maps are input into multiple cross-resolution fusion modules for feature map restoration and decoding to obtain multiple decoding feature maps; Multiple decoded feature maps are fed into the scale-aware activation module and supervised using the pixel-level labels corresponding to the original CT image to obtain the predicted segmentation result; the loss function is calculated to reversely update the network weights and the network parameters are optimized using the SGD optimizer until the loss function converges or the maximum number of training times is reached to obtain a trained attention-guided cross-resolution collaborative network; The CT images to be detected are collected and input into the trained attention-guided cross-resolution collaborative network for tumor segmentation.
2. The attention-guided cross-resolution collaborative tumor segmentation method according to claim 1, characterized in that: Each multi-attention fusion module first extracts the encoding features from the input encoding feature map through the encoding layer, and then calculates the attention scores under the three attention mechanisms respectively; then the attention scores, encoding features under the three attention mechanisms and the input encoding feature map are merged in the weighted fusion module, and then the encoding feature map of the multi-attention fusion module is obtained through the downsampling module; The attention scores include channel-level attention scores, spatial-level attention scores, and global attention scores.
3. The attention-guided cross-resolution collaborative tumor segmentation method according to claim 2, characterized in that: The channel-level attention score is expressed as: in, is the channel-level attention score of the i-th multi-attention fusion module, σ is the Sigmoid function, CRC is the convolution-ReLU activation-convolution operation, Conv is the convolution operation, ⊕ is the element-level addition, Pavg c is the average pooling operation in the channel dimension, Pmax c is the maximum pooling operation in the channel dimension, N is the number of multi-attention fusion modules; F i-1 is the encoded feature map output by the i-1th multi-attention fusion module. When i=1, F0 is the encoded feature map of the patch through simple feature extraction through convolution. The spatial level attention score is expressed as: in, is the spatial attention score of the i-th multi-attention fusion module, Pavg s is the average pooling operation at the spatial level, Pmax s is the maximum pooling operation at the spatial level, It is the join and merge operation of channel dimension; The global attention score is expressed as: in, is the global attention score of the i-th multi-attention fusion module; The encoding feature of the multi-attention fusion module is expressed as: in, represents element-level multiplication, and Pmax is the downsampling module, that is, the maximum pooling operation.
4. The attention-guided cross-resolution collaborative tumor segmentation method according to claim 1, characterized in that: Each cross-resolution fusion module first uses the resolution feature extraction module to align the size and channel of the encoded feature maps of the multi-attention fusion module and its adjacent attention fusion module with the corresponding jump connection, and extracts features to obtain the resolution attention map, which is then input into the resolution fusion module for fusion to obtain the cross-resolution fusion feature; The resolution attention map includes a high-resolution attention map and a low-resolution attention map; Next, the input decoding feature map is resized and channel adjusted by the feature adjustment module, fused with the cross-resolution fusion feature, and then passed through the convolution output module to obtain the decoding feature map of the cross-resolution fusion module.
5. The method for tumor segmentation based on attention-guided cross-resolution collaboration according to claim 4, characterized in that: The high-resolution attention map is expressed as: in, For the corresponding processing F i The high-resolution attention map of the cross-resolution fusion module, F i is the encoding feature map of the i-th multi-attention fusion module, represents element-wise multiplication, σ is the Sigmoid function, Pmax is the maximum pooling operation, Conv is the convolution operation, is element-level addition, M is the number of cross-resolution fusion modules, and F i-1 is the encoding feature map of the previous multi-attention fusion module of the i-th multi-attention fusion module. When i=1, F0 is the encoding feature map of the patch through simple feature extraction through convolution. The low-resolution attention map is represented as: in, For the corresponding processing F i The low-resolution attention map of the cross-resolution fusion module, UP is the interpolation upsampling operation, and F i+1 is the encoding feature map of the next multi-attention fusion module of the i-th multi-attention fusion module; The cross-resolution fusion feature is expressed as: Among them, F i ′ is the cross-resolution fusion feature of the i-th cross-resolution fusion module, is the connection and merging operation of the channel dimension, and ReLU is the ReLU activation function; The decoded feature map of the cross-resolution fusion module is expressed as: Among them, D i is the decoded feature map of the i-th cross-resolution fusion module, D i-1 is the decoded feature map of the previous cross-resolution fusion module of the i-th cross-resolution fusion module. When i=1, D0 is the encoded feature map F of the M-th multi-attention fusion module. M .
6. The attention-guided cross-resolution collaborative tumor segmentation method according to claim 1, characterized in that: The process of obtaining the predicted segmentation result by the scale-aware activation module is as follows: Multiple decoded feature maps are input into the scale-aware activation module, and the upsampling fusion module is used to perform convolution and interpolation upsampling operations on the decoded feature maps to unify the scale size and merge them to obtain a mixed feature map; The mixed feature map is processed using the scale attention module, and the scale attention results and spatial attention results are calculated at the channel dimension and spatial level to obtain the mixed attention map; Based on the hybrid feature map and the hybrid attention map, the classification module obtains the pixel-level predicted segmentation results under the supervision of the pixel-level labels corresponding to the original CT image.
7. The method for tumor segmentation based on attention-guided cross-resolution collaboration according to claim 6, characterized in that: The mixed feature map is expressed as: Among them, M is the number of cross-resolution fusion modules, Conv is the convolution operation, UP is the interpolation upsampling operation, is the connection and merging operation in the channel dimension, D i is the decoded feature map of the i-th cross-resolution fusion module; The mixed attention map is expressed as: in, represents element-wise multiplication, σ is the Sigmoid function, Pavg s is the average pooling operation at the spatial level, Pmax s is the maximum pooling operation at the spatial level, MLP is a multi-layer perceptron, and D s ′ ca is the scale attention result, expressed as: in, is element-wise addition, Pavg c is the average pooling operation in the channel dimension, Pmax c is the maximum pooling operation in the channel dimension, MLP s is a shared multi-layer perceptron; The pixel-level prediction segmentation result is expressed as: Pred=SofMax(BN(Conv(D′+D SA ))), Among them, SofMax is the SofMax classification function, and BN is batch normalization.
8. The attention-guided cross-resolution collaborative tumor segmentation method according to claim 1, characterized in that: The loss function L seg Using cross entropy loss L CE and Dice loss L DC , expressed as: L seg =L CE +L DC , The cross entropy loss L CE It is expressed as: The Dice loss L DC It is expressed as: in, is the number of categories, HWD is the size of the original CT image; I(·) is the indicator function, when the true label of voxel t is category n, the function value is 1, otherwise it is 0; Y t is the true label of voxel t, Pred t,n is the segmentation probability that the t-th voxel belongs to category n.
9. Attention-guided cross-resolution collaborative tumor segmentation system, characterized by: The system includes a data processing module, a network construction module, a network training module and a tumor segmentation module; The data processing module is used to obtain the original CT image and its corresponding pixel-level labels, resample the original CT image and crop it into patches of a set size; The network construction module is used to construct an attention-guided cross-resolution collaborative network, including multiple multi-attention fusion modules, multiple cross-resolution fusion modules and a scale-aware activation module; the multiple multi-attention fusion modules are connected in sequence, and each multi-attention fusion module includes a coding layer, three attention modules of channel, space and global, a weighted fusion module and a downsampling module; the multiple cross-resolution fusion modules are connected in sequence, and each cross-resolution fusion module includes a resolution feature extraction module, a resolution feature fusion module, a feature adjustment module and a convolution output module; the multiple multi-attention fusion modules are jump-connected with multiple cross-resolution fusion modules, and each cross-resolution fusion module is connected with the corresponding jump-connected multi-attention fusion module and its adjacent multi-attention fusion module; The scale-aware activation module includes an upsampling fusion module, a scale attention module, and a classification module connected in sequence; the upsampling fusion module is connected to a plurality of cross-resolution fusion modules; The number of the multi-attention fusion modules is the same as the number of the cross-resolution fusion modules; The network training module is used to train the attention-guided cross-resolution collaborative network using a back-propagation algorithm. The training process is as follows: After the patch is subjected to simple feature extraction through convolution, it is input into multiple multi-attention fusion modules for deep feature extraction to obtain multiple encoding feature maps; multiple encoding feature maps are input into multiple cross-resolution fusion modules for feature map restoration and decoding to obtain multiple decoding feature maps; Multiple decoded feature maps are fed into the scale-aware activation module and supervised using the pixel-level labels corresponding to the original CT image to obtain the predicted segmentation result; the loss function is calculated to reversely update the network weights and the network parameters are optimized using the SGD optimizer until the loss function converges or the maximum number of training times is reached to obtain a trained attention-guided cross-resolution collaborative network; The tumor segmentation module is used to collect the CT image to be detected and input it into the trained attention-guided cross-resolution collaborative network to perform tumor segmentation.
10. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the attention-guided cross-resolution collaborative tumor segmentation method described in any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Transfer prediction method and system based on resampling and three-dimensional attention convolution
CN121190814A