CT (Computed Tomography) image auxiliary segmentation system with double-flow channel and attention comparison fusion
Patent Information
- Application Number
- CN202510458484.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-29
Smart Images

Figure CN120388027A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing. More specifically, the present invention relates to an automatic segmentation method for aortic dissection CT images based on deep learning and attention mechanism. Background Art
[0002] Aortic dissection is a high-risk aortic vascular disease, which is characterized by the tearing of the aortic intima, resulting in blood flowing between the inner and outer layers of the aortic wall, forming a dissection. Timely computed tomography (CT) is an important tool for the diagnosis and treatment of aortic dissection. Traditional medical image processing methods often rely on doctors' experience, and the segmentation process is time-consuming and vulnerable to subjective factors.
[0003] In recent years, with the development of deep learning technology, image segmentation methods based on deep neural networks have shown significant advantages in medical image analysis. Especially in the segmentation tasks of complex structures, deep learning can automatically learn features, reduce human intervention, improve the accuracy and consistency of segmentation. By segmenting the lesion area through a deep learning model, it can help medical staff observe the lesion morphology more clearly and improve work efficiency.
[0004] Currently, the deep learning segmentation model for aortic dissection CT images is not yet mature, and existing technologies often cannot effectively cope with problems such as large differences in lesion morphology and unclear boundaries in the images. Summary of the Invention
[0005] The purpose of the present invention is to solve the problem that the existing aortic dissection segmentation system has insufficient adaptation to the diverse morphology of the lesion area and blurred boundaries, and to provide a CT image assisted segmentation system containing a dual-stream channel and attention comparison fusion, which is characterized by including:
[0006] a): A preprocessing module: used for slicing the CT image to generate two-dimensional slice images containing the lesion area. The processed images can be used as training samples independent of the source patient to input into the model to support the training and recognition of the model. At the same time, the labels of the two-dimensional lesion images are saved separately in the form of binary label images containing only two elements, black and white. The label images correspond one-to-one with the slice images;
[0007] b): Intra-class feature extraction module, which is used to perform intra-class feature extraction operations on the cross-sectional CT images of aortic dissection transmitted into the model; the intra-class feature extraction module includes four layers of intra-class feature encoders and four layers of intra-class feature decoders. The intra-class feature encoder is used to reduce the scale and extract features of the upper-layer output feature map (the input of the first-layer intra-class feature encoder is the cross-sectional image of the aortic dissection CT). The intra-class feature decoder is used to enlarge the scale of the output feature map of the lower-layer intra-class feature decoder (the input of the fourth-layer intra-class feature decoder feature map is the output of the fourth-layer intra-class feature encoder), and perform the fusion and feature screening of the enlarged feature map and the input feature map of the same-layer intra-class feature encoder; the image input to the deep learning model passes through the intra-class feature extraction module to obtain a feature map of intra-class features with the same size as the input image;
[0008] c): Edge-anchoring feature extraction module, which is used to perform edge-anchoring feature extraction operations on the images transmitted into the model; the edge-anchoring feature extraction module includes four layers of edge-anchoring feature encoders and four layers of edge-anchoring feature decoders. The connection, layout of each encoder and decoder, the structure of the decoder, and the size of the input and output feature maps of each encoder and decoder module are the same as those of the intra-class feature extraction module; for the edge-anchoring feature encoder, specifically, it includes a feature extraction sub-module, a downsampling sub-module, a feature superposition operation, and an activation function; the feature map input to this module passes through the feature extraction sub-module and the downsampling sub-module respectively to generate two intermediate feature maps, and after feature superposition, it enters the activation function to obtain the output of the encoder module, so that the feature map output by the edge-anchoring feature encoder contains the overall context information feature information; all the activation functions in this module are relu functions;
[0009] d): Attention comparison and fusion module, which is used to fuse the intra-class features and edge-anchoring features and generate the final segmentation result; the attention comparison and fusion module has two input ends, which respectively input the intra-class features and edge-anchoring features; among them, the attention comparison and fusion module includes a feature superposition operation, an attention sub-module, and a result generation sub-module. The two types of features generated by the two feature extraction modules first perform feature superposition in the attention comparison and fusion module, so that the two types of features are fused and compared at the same time, highlighting the different parts in the two feature maps, that is, the controversial edge information part in the two types of feature maps. The generated intermediate feature map and the feature map of the intra-class features are subjected to attention comparison and fusion guided by the edge features through the attention sub-module, and then through the result generation sub-module to generate the final segmentation result;
[0010] The serial connection method of each module is as follows: starting from the CT image, first generate a two-dimensional medical image containing the lesion area through the preprocessing module, and then input the image into the intra-class feature extraction module and the edge-anchored feature extraction module at the same time to generate two corresponding types of features. These two types of features are used as the two input ends of the attention comparison and fusion module to fuse the features through this module, and finally generate the segmentation result.
[0011] For the aortic dissection CT image dataset used in model training, the preprocessing process is as follows:
[0012] For the three-dimensional aortic dissection CT image of each patient, according to the markings of professional physicians, obtain the cross-sectional position with the largest lesion area, and then, taking this position as the reference, extend two unit positions in each of the two directions perpendicular to the cross-section. The unit length standard is the minimum unit for the CT image to move along the cross-section. In this way, a total of five cross-sectional positions are obtained;
[0013] Perform cross-sectional cutting operations on the three-dimensional aortic dissection CT image of each patient at five cross-sectional positions in the cross-sectional direction, and save the cross-sections as.JPG type images. Five interface images containing the lesion area are obtained from the CT images of each patient. At the same time, the markings related to the lesions in the CT images of each patient are cut along the cross-section at the five cross-sectional positions of the patient and saved as.PNG images. The background in the marked image is black, and the lesion area is white.
[0014] Uniformly process the cross-sectional images and the corresponding marked images obtained from the three-dimensional aortic dissection CT images of each patient into (512, 512) pixel sizes, perform numbering operations, and eliminate the connection between each image and the patient. Each medical image and its corresponding marked image have the same name.
[0015] The intra-class feature extraction module is built using U-net; the feature extraction sub-module of the edge-anchored feature extraction module includes two alternating convolution operations and normalization operations, and a activation function connected in series after the first normalization operation. The convolution kernel sizes of the two convolution operations are (3, 3) and the padding is 1, but the stride of the first convolution operation is 2, so that the size of the feature map is reduced by half; the activation function is the relu function; the downsampling sub-module of the edge-anchored feature extraction module includes a convolution operation with a convolution kernel size of (1, 1) and a stride of 2 and a normalization operation.
[0016] The attention sub-module of the attention comparison and fusion module includes two feature weighting operations, one feature superposition operation, one activation function, and one convolutional sub-module. Among them, the convolutional sub-module includes one convolutional operation with a kernel size of (1,1), one normalization operation, and one activation function, and each part is connected in series in turn. For the two inputs of the attention comparison and fusion module, they are first multiplied by a learnable weight matrix, the results are feature-superposed, and then the input to the convolutional sub-module is generated through the activation function, and the final output of the attention sub-module is generated by the convolutional sub-module. Among them, the activation function of the convolutional sub-module is the Sigmoid function, and the rest are relu functions. The result generation sub-module of the attention comparison and fusion module includes one convolutional operation with a kernel size of (1,1).
[0017] The size of the medical images used is (512,512). In the intra-class feature extraction module, the side length of the image matrix becomes 256, 128, 64, and 32 in turn after passing through multiple encoders, and then returns to 512 in size after passing through multiple decoders. In the attention feature fusion module, the sizes of all intermediate process feature maps are (512,512). Brief Description of the Drawings
[0018] Figure 1 It is a structural diagram of a deep neural network.
[0019] Figure 2 It is a connection structure diagram of the encoder and decoder of the feature extraction module.
[0020] Figure 3 It is a structural diagram of the encoder of the edge-anchored feature extraction module.
[0021] Figure 4 It is a structural diagram of the attention comparison and fusion module. Detailed Embodiments
[0022] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation examples described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not form a conflict with each other.
[0023] The data form of this example is the CT images of patients with aortic dissection. Taking the first use of the system as an example, according to Figure 1 the model structure shown, the following steps are included:
[0024] Data collection and preprocessing: The dataset used for model training is the CT scan images of 204 patients with dissecting aneurysm collected by Yantai Yuhuangding Hospital, which are marked by professional physicians with rich clinical experience on CT tomographic scans. First, the CT image data is sliced through a preprocessing module to generate two-dimensional cross-sectional images containing the lesion area, and the image size is adjusted to (512, 512), so that the data is transformed into a form that can be input into the deep learning model. At the same time, the generated image data is renamed to eliminate the dependence relationship between each generated image and the patient source. The names of the lesion area images and their corresponding labeled images are the same, and the image set and the label set are placed in two folders respectively.
[0025] Model training and result output: The processed image data is input into the deep learning model, and at the same time, it enters the intra-class feature extraction module and the edge anchor feature extraction module for two parallel feature extraction processes to extract two different features. Then, the two types of features are used as two inputs to enter the attention comparison and fusion module, which is used to fuse the intra-class features and the edge anchor features and generate the final segmentation result. The presentation form of the segmentation result is a binary image corresponding to the input CT cross-sectional image containing the aortic dissection lesion area, which is the same as the expression method of the labeled image in the preprocessing stage. The output image size is consistent with the input image, the background is displayed in black, and the lesion area is displayed in white. After overlapping with the CT cross-sectional image, the part corresponding to the white area in the CT cross-sectional image is the automatically segmented lesion area.
[0026] For each component module of the model of the present invention, the specific description is as follows:
[0027] The intra-class feature extraction module is implemented by the Unet network structure, which includes four layers of intra-class feature encoders and four layers of intra-class feature decoders. The connection and distribution of each encoder and decoder are as Figure 2 shown. The input medical images pass through four layers of intra-class feature encoders and four layers of intra-class feature decoders in sequence. When passing through the encoder, the size of the feature map becomes 0.5 times the size of the upper-layer feature map after passing through each layer. After a complete encoding process, the size of the feature map becomes (32, 32). And when passing through the decoder in sequence, the size of the feature map gradually expands to (512, 512). The encoder has one input end and one output end, and the decoder has two input ends and one output end. In addition to the output of the previous layer decoder or the last layer encoder as the input, the decoder will also use the intermediate result with the same size as the decoder output during the encoding process as another input. Through the intra-class feature extraction module, the model extracts the intra-class features related to the morphological features of the model.
[0028] The edge-anchored feature extraction module consists of four layers of edge-anchored feature encoders and four layers of edge-anchored feature decoders. The technical feature of this module compared with the intra-class feature extraction module lies in the re-design of the encoder. The encoder structure of the edge-anchored feature extraction module is as shown in Figure 3 Figure [0000067], which includes a feature extraction sub-module, a down-sampling sub-module, a feature stacking operation, and an activation function. Among them, the feature extraction sub-module is used for feature extraction of edge-anchored features. Its structure includes two alternating convolution operations and normalization operations, and an activation function concatenated after the first normalization operation. The convolution kernel sizes of the two convolution operations are (3,3) and the padding is 1, but the stride of the first convolution operation is 2, reducing the feature map size by half; the feature stacking operation is to add the corresponding pixels of the two feature maps, and the activation function is the relu function; the down-sampling sub-module is used to endow the edge-anchored features with more generalized overall features. Its structure includes one convolution operation and one normalization operation, the convolution kernel size is (3,3), the padding is 1, and the stride is 2. The distribution and connection methods of the decoder, encoder, and decoder of the edge-anchored feature extraction module are the same as those of the intra-class feature extraction module. The feature extraction path and scale change of medical images in the edge-anchored feature extraction module are the same as those of the intra-class feature extraction module, and the final output is the edge-anchored features containing the edge features of the lesion area.
[0029] The structure of the attention comparison and fusion module is as shown in Figure 4 Figure [0000070]. The two types of features are first subjected to feature stacking so as to extract the edge information by comparing the two features. Then, the result of the feature stacking is input into the attention sub-module together with the intra-class features. This module assigns weights to the two inputs to balance the influence of the features related to the lesion morphology and edge on the final segmentation result. The output of the attention sub-module passes through the result generation module and is transformed into the final segmentation result. The attention sub-module consists of two feature weighting operations, one feature stacking operation, an activation function, and a convolution sub-module. Among them, the convolution sub-module includes one convolution operation with a convolution kernel size of (1,1), one normalization operation, and an activation function, and each part is concatenated in sequence; the feature weighting operation is to multiply the feature map by the weight matrix, the feature stacking operation is to add the corresponding pixels of the two feature maps, the activation function of the convolution sub-module is the Sigmoid function, and the rest are relu functions. The result generation sub-module of the attention comparison and fusion module includes one convolution operation with a convolution kernel size of (1,1).
[0030] The output result of the present invention is a binary image showing the aortic dissection lesion area in the patient's CT image. During the training phase, professional physicians' markings of the same form of the lesion area were obtained. Therefore, the Dice coefficient, mean intersection over union (mIoU), and accuracy (Accuracy) can be used as indicators to evaluate the performance of the deep learning model, and it is compared with multiple classic deep learning models.
[0031] The Dice coefficient is an index to measure the similarity between two samples; mIoU is used to evaluate the overlapping degree between the predicted result and the ground truth label; the accuracy rate represents the proportion of correctly predicted pixels in the total pixels. By comparing the magnitudes of these three metrics, the consistency between the segmentation result of the model of the present invention for the lesion area and the marking of the lesion area by professional physicians can be measured, reflecting the scientificity and feasibility of the model of the present invention for realizing the automatic segmentation of aortic dissection CT images. At the same time, by comparing the scores of the three evaluation metrics of the model of the present invention and other classic modules for the same data set, the superiority of the segmentation effect of the model of the present invention can be proved.
[0032] The comparison results between the model of the present invention and other models are shown in Table 1. In terms of the Dice coefficient, mIoU, and accuracy rate, the present invention can more accurately extract the boundary and morphological information of the lesion area, achieving a higher degree of overlap with the actual lesion area; by optimizing the feature extraction and fusion methods, the present invention significantly improves the segmentation accuracy of each category, ensuring the fine-grained segmentation effect of different lesion areas; by adopting the attention comparison fusion module, the noise interference is effectively suppressed, and the correctness of the overall pixel classification is improved, fully proving the superiority of the model of the present invention in processing aortic dissection CT images.
[0033] In addition to the above comparison with the segmentation results of multiple classic models to determine the superiority of the dual-stream channel, in order to verify the effectiveness of the edge-anchored feature encoder and the attention comparison fusion module in the present invention for optimizing the lesion segmentation effect of the model, ablation experiments are respectively carried out on the edge-anchored feature encoder and the attention comparison fusion module. The results of the two ablation experiments are shown in Table 2, where Modle1 is the model obtained by replacing the encoder of the edge-anchored feature extraction module with ordinary convolution operations, and Model2 is the model obtained by replacing the attention comparison fusion module with feature stacking operations.
[0034] Through the technical solution of the present invention, the difficulties in the segmentation of complex aortic dissection lesion CT images in the prior art can be effectively solved. Especially when dealing with lesion areas with variable morphologies, the intra-class features and edge features of the lesions can be accurately extracted. Through the dual-stream channel structure and the design of the edge-anchored feature extraction anchor box and the attention comparison fusion module, the present invention can significantly improve the accuracy and robustness of image segmentation, and has important practical application value, especially in the field of medical image analysis.
[0035] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0036] Table 1 Comparison of Model Effects
[0037] Dice mIoU Acc Unet 93.48 93.86 93.43 Unet++ 94.39 94.58 94.39 Att-Unet 93.74 94.86 93.73 res-Unet++ 95.94 96.35 95.92 Swin-UNet 96.81 96.88 96.67 The present invention 97.62 97.76 97.67
[0038] Table 2 Ablation Experiment
[0039] Dice mIoU Acc Model1 97.09 97.07 97.11 Model2 97.50 97.57 97.50 The present invention 97.62 97.76 97.67 References
[0040] [1] Ronneberger O, Fischer P, Brox T. U-net: Convolutional networks for biomedical image segmentation[C] / / Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18. Springer International Publishing, 2015: 234-241.
[0041] [2]Diakogiannis F I, Waldner F, Caccetta P, et al. ResUNet-a: A deep learning framework for semantic segmentation of remotely sensed data[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2020, 162: 94-114.
[0042] [3]Ravichandran S R, Nataraj B, Huang S, et al. 3D inception U-Net for aorta segmentation using computed tomography cardiac angiography[C] / / 2019 IEEE EMBS International Conference on Biomedical&Health Informatics (BHI). IEEE, 2019: 1-4.
[0043] [4] Cao H, Wang Y, Chen J, et al. Swin-unet: Unet-like pure transformer for medical image segmentation[C] / / European conference on computer vision. Cham: Springer Nature Switzerland, 2022: 205-218.
[0044] [5] Liu Z, Lin Y, Cao Y, et al. Swin transformer: Hierarchical vision transformer using shifted windows[C] / / Proceedings of the IEEE / CVF International conference on computer vision. 2021: 10012-10022.
[0045] [6] Vaswani A. Attention is all you need[J]. Advances in Neural Information Processing Systems, 2017。
Claims
1. A CT image assisted segmentation system containing a dual-stream channel and attention comparison fusion, characterized in that, include: a) Preprocessing module: This module processes slices of CT images to generate two-dimensional slice images containing lesion areas. The processed images can be used as training samples independent of the source patient and input into the model to support model training and recognition. The two-dimensional lesion images are labeled as binary labeled images containing only black and white elements, and the labeled images correspond one-to-one with the slice images. b): Intra-class feature extraction module, used as a feature extraction channel to generate feature maps for the aortic dissection CT cross-sectional images transmitted into the model; the intra-class feature extraction module includes four layers of intra-class feature encoders and four layers of intra-class feature decoders; the input of the top-level intra-class feature encoder is the CT cross-sectional layer image, and the input of the intra-class feature encoders of other layers is the output feature map of the upper-level intra-class feature encoder; the input of the bottom-level intra-class feature decoder is the input feature map and output feature map of the bottom-level intra-class feature encoder, and the input of the intra-class feature decoders of other layers is the output feature map of the lower-level intra-class feature decoder and the input of the intra-class feature encoder of the same layer; the image input to the deep learning model passes through the intra-class feature extraction module to obtain a feature map of the intra-class features that is consistent with the size of the input image; c) An edge-anchored feature extraction module, used as another feature extraction channel to generate feature maps for the aortic dissection CT cross-sectional images transmitted into the model; the edge-anchored feature extraction module includes four layers of edge-anchored feature encoders and four layers of edge-anchored feature decoders. The connection and layout of each encoder and decoder, the structure of the decoder, and the input and output feature map sizes of each encoder and decoder are consistent with those of the intra-class feature extraction module; the edge-anchored feature encoder specifically includes a feature extraction submodule, a downsampling submodule, a feature superposition operation, and an excitation function; the feature map input to this submodule passes through the feature extraction submodule and the downsampling submodule respectively to generate two intermediate feature maps, which enter the excitation function after feature superposition to obtain the output of the encoder module; all excitation functions in this module are relu functions; d): Attention comparison fusion module, used to fuse intra-class features and edge anchor features and generate the final segmentation results; The attention comparison fusion module has two input ports, one for intra-class features and the other for edge anchor features. The attention comparison fusion module includes a feature superposition operation, an attention submodule, and a result generation submodule. The two types of features generated by the two feature extraction modules are first superimposed in the attention comparison fusion module, and then passed through the result generation submodule to generate the final segmentation result. The modules are connected in series as follows: starting from the CT image, a two-dimensional medical image containing the lesion area is first generated through the preprocessing module, and then the image is simultaneously input into the intra-class feature extraction module and the edge anchor feature extraction module to generate two corresponding types of features respectively. These two types of features are used as the two input ends of the attention comparison fusion module to fuse the features through this module, and finally generate the segmentation result.
2. The CT image assisted segmentation system with dual-stream channel and attention comparison fusion according to claim 1, characterized in that: It includes an edge-anchored feature extraction module and an attention comparison fusion module. Connecting the edge-anchored feature extraction module in parallel with Unet can improve technical indicators such as the Dice coefficient, mean intersection ratio, and accuracy. Changing the parallel connection mode to the attention comparison fusion module can further improve technical indicators such as the Dice coefficient, mean intersection ratio, and accuracy.
3. The CT image assisted segmentation system with dual-stream channel and attention comparison fusion according to claim 1, characterized in that: The aortic dissection CT image dataset used for model training has the following preprocessing process: For each patient's 3D aortic dissection CT image, the section position with the largest lesion area in the cross section was obtained according to the professional physician's markings. Then, based on this position, the section positions were extended two units in each direction perpendicular to the cross section. The length of the unit was the smallest unit of movement of the CT image along the cross section. A total of five section positions were obtained. The three-dimensional aortic dissection CT image of each patient was cut into sections at five sections in the cross-sectional direction and saved as .JPG images. Five interface images containing the lesion area were obtained from each patient's CT image. The lesion area at the five sections in each patient's CT image was also saved as .PNG images, with the background in the image marked as black and the lesion area as white. The cross-sectional images and corresponding labeled images obtained from the three-dimensional aortic dissection CT images of each patient were uniformly processed to a resolution of (512,512) and numbered to eliminate the connection between each image and the patient. Each medical image and its corresponding labeled image had the same name except for the file suffix.
4. The CT image assisted segmentation system with dual-stream channel and attention comparison fusion according to claim 1, characterized in that: The intra-class feature extraction module is built using U-net.
5. The CT image-assisted positioning system with dual-stream channel and attention comparison fusion according to claim 1, characterized in that: The feature extraction submodule of the edge anchor feature extraction module includes two convolution operations and normalization operations. The convolution operations and normalization operations are connected alternately, and an excitation function is connected in series after the first normalization operation. The convolution kernel size of the two convolution operations is (3,3) and the padding is 1. The step size of the first convolution operation is 2, which reduces the size of the feature map by half, and the step size of the second convolution operation is 1; the excitation function is the ReLU function.
6. The CT image assisted segmentation system with dual-stream channel and attention comparison fusion according to claim 1, characterized in that: The downsampling submodule of the edge anchor feature extraction module includes a convolution operation with a kernel size of (1,1) and a stride of 2 and a normalization operation.
7. The CT image assisted segmentation system with dual-stream channel and attention comparison fusion according to claim 1 is characterized by: The attention sub-module of the attention comparison and fusion module includes two feature weighting operations, one feature stacking operation, one activation function, and one convolutional sub-module; among them, the convolutional sub-module includes one convolutional operation with a kernel size of (1,1), one normalization operation, and one activation function, and each part is connected in series in turn; for the two input feature maps of the attention comparison and fusion module, they are first multiplied by a learnable weight matrix, and then the results are feature-stacked. The result of the feature stacking passes through an activation function to generate the input to the convolutional sub-module, and the convolutional sub-module generates the final output of the attention sub-module; among them, the activation function of the convolutional sub-module is the Sigmoid function, and the rest are relu functions.
8. The CT image assisted segmentation system with dual-stream channels and attention comparison and fusion according to claim 1, wherein: The result generation sub-module of the attention comparison and fusion module includes one convolutional operation with a kernel size of (1,1).
9. The CT image assisted segmentation system with dual-stream channels and attention comparison and fusion according to claim 1, wherein: The size of the medical image used is (512,512). In the intra-class feature extraction module, the side length of the image matrix becomes 256, 128, 64, and 32 in turn after passing through multiple encoders, and then returns to 512 in size after passing through multiple decoders. In the attention feature fusion module, the sizes of all intermediate process feature maps are (512,512).