Medical image segmentation method based on hybrid convolutional neural network and converter
By adopting a hybrid convolutional neural network and transformer method in medical image segmentation, the problem of difficulty in capturing local and global features at the same time through the design of channel feature correlation matrix and cross-type spatial feature fusion module is solved, and medical image segmentation with higher accuracy is achieved.
Patent Information
- Application Number
- CN202510001447.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-02
AI Technical Summary
The prior art is difficult to capture local feature information and global feature information simultaneously in medical image segmentation, resulting in poor segmentation effect.
Using a medical image segmentation method based on a hybrid convolutional neural network and a transformer, the interactive fusion of channel information and spatial information between the global feature map and the local feature map is realized by constructing a channel feature correlation matrix and an inter-space feature fusion module.
The image segmentation model captures global feature information and local feature information, and improves the accuracy of medical image segmentation, especially when processing low-quality medical images.
Smart Images

Figure CN120088268A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and more particularly, to a medical image segmentation method based on a hybrid convolutional neural network and a transformer. Background Art
[0002] plays an important role in the field. In the past few decades, semantic segmentation technology based on deep learning has attracted extensive attention from researchers due to its higher efficiency than manual annotation. The essence of semantic segmentation is to classify pixel values to achieve pixel-level annotation of complex lesion areas (such as brain tumors, melanomas, and various cancerous areas) in medical images (Azad et al., 2024; Asgari Taghanaki et al., 2021).
[0003] Semantic segmentation models based on deep convolutional neural networks have been widely applied to various vision tasks, and the U-shaped structure is particularly popular in the medical field. The U-shaped structure usually consists of an encoder and a decoder. The encoder captures semantic and context information through successive convolutional layers and downsampling; the decoder then reconstructs the output mask through step-by-step upsampling (Zhou et al., 2019). Although deeper convolutional layers and more downsampling can expand the receptive field of the model, it may lead to the loss of context information. However, the U-shaped model can recover the lost context information through skip connections. However, due to the convolution mechanism of the convolutional layer, the limited receptive field and the inability to model long-range dependencies are still difficult to solve (Yuan et al., 2023; Heidari et al., 2023).
[0004] Vision Transformers (ViTs) improve the global receptive field of the model by dividing the image into small pixel patches (Patches) and modeling the relationships between the small pixel patches. However, ViTs still have deficiencies in capturing low-level features (Heidari et al., 2023).
[0005] Generally speaking, the contributions of convolutional neural networks in medical image segmentation are their lightweight design and the effective capture of local features to achieve efficient segmentation. However, since the design of convolutional neural networks is based on sliding windows for feature extraction and cannot pay attention to the information correlation outside the sliding windows, convolutional neural networks cannot model global information in medical image segmentation. On the other hand, transformers are very good at capturing global information because they model feature patches through the self-attention mechanism. However, due to this mechanism, transformers cannot model local information within features. Therefore, it is difficult to achieve the modeling of both local and global feature information and effective medical image segmentation by relying solely on one of these technologies. Summary of the Invention
[0006] The technical problem to be solved by the present invention is how to improve the ability to capture local and global feature information and achieve accurate medical image segmentation.
[0007] The present invention provides a medical image segmentation method based on a hybrid convolutional neural network and a transformer, including: Step 1. Obtain the medical image to be segmented; Step 2. Construct an image segmentation model, and input the medical image to be segmented into the image segmentation model. The image segmentation model includes a preprocessing layer, a hybrid encoder layer, and a decoder layer; The preprocessing layer is used to perform image segmentation processing and local feature extraction on the medical image; The hybrid encoder layer is connected to the preprocessing layer and is used to extract a global feature map and a local feature map from the segmented image and local features, construct a channel feature correlation matrix based on the channel features of the global feature map and the local feature map, perform interactive fusion of channel information on the global feature map and the local feature map based on the channel feature correlation matrix, and then perform spatial information interactive fusion on the globally and locally feature maps after the interactive fusion; The decoder layer is used to splice and upsample the local features output by the preprocessing layer and the features after the spatial information interactive fusion performed by the hybrid encoder layer, and output the target segmentation image.
[0008] Compared with the prior art, the present application has the following advantages: Based on the channel features between the local feature map and the global feature map, the present invention constructs a channel feature correlation matrix, and based on the channel feature correlation matrix, performs interactive fusion on the channel information between the global feature map and the local feature map. Then, performs interactive and fusion on the spatial information of the globally and locally feature maps after the interactive fusion, so that both the local feature map and the spatial feature map have local feature information and global feature information, thereby enhancing the ability of the image segmentation model to capture global feature information and local feature information, and improving the ability of the image segmentation model to reconstruct the mask with high accuracy.
[0009] In a possible implementation manner, the hybrid encoder layer includes a first fusion layer, a second fusion layer, a third fusion layer, and a fourth fusion layer connected in sequence. The decoder layer includes a fifth decoder layer, a fourth decoder layer, a third decoder layer, a second decoder layer, and a first decoder layer connected in sequence. The first fusion layer performs interactive fusion of communication information and then transmits it to the second fusion layer. The second fusion layer performs interactive fusion of channel information and then transmits it to the third fusion layer. The third fusion layer performs interactive fusion of channel information and then transmits it to the fourth fusion layer. The fourth fusion layer performs interactive fusion of spatial information and then enters the fifth decoder layer for upsampling operation. The third fusion layer performs interactive fusion of spatial information and then is concatenated with the features output by the fifth decoder layer and enters the fourth decoder layer for upsampling operation. The second fusion layer performs interactive fusion of spatial information and then is concatenated with the features output by the fourth decoder layer and enters the third decoder layer for upsampling operation. The first fusion layer performs interactive fusion of spatial information and then is concatenated with the features output by the third decoder layer and enters the second decoder layer for upsampling operation. The local features output by the preprocessing layer are concatenated with the features output by the second decoder layer and enter the first decoder layer for double upsampling operation and then output to obtain the target segmentation image.
[0010] Compared with the prior art, although using one fusion layer can make the local feature map have global information and the global feature map have local information, in order to facilitate the decoder layer to reconstruct the mask, four fusion layers are used to perform interactive and fusion of the channel information and spatial information between the local feature map and the global feature map, so that the image segmentation model has a more accurate local feature capture ability and global feature capture ability, thereby being able to segment low-quality medical images.
[0011] In a possible implementation, the second decoder layer, the third decoder layer, the fourth decoder layer, and the fifth decoder layer have the same network structure, each including a first convolutional block, a second convolutional block, and a transposed convolutional block connected in sequence. The first decoder layer includes two layers of CNN decoder layers and a 1×1 convolutional block connected in sequence, and the network structures of the two layers of CNN decoder layers are the same as those of the second decoder layer to the fifth decoder layer.
[0012] Compared with the prior art, the present invention jointly performs upsampling operations through the double convolutional blocks and the transposed convolutional block of the decoder layer, which helps to restore the local feature map and the spatial feature map in terms of spatial dimensions, and at the same time ensures the lightweight of the image segmentation model.
[0013] In a possible implementation, the network structures of the first convolutional block and the second convolutional block each include a 3×3 convolutional block, a BN block, a ReLu function block, a 3×3 convolutional block, a BN block, and a ReLu function block, which perform 3×3 convolutional operations, normalization operations, and activation operations twice in sequence.
[0014] In a possible implementation, the network structures of the first fusion layer, the second fusion layer, the third fusion layer, and the fourth fusion layer are the same, each including a transformer module, a convolutional neural network module, a cross-domain channel attention module, and a cross-type spatial feature fusion module; The transformer module is used to extract the global feature map; The convolutional neural network module is used to extract the local feature map; The cross-domain channel attention module is respectively connected to the transformer module and the convolutional neural network module, and is used to construct a channel feature correlation matrix, and perform interactive fusion of channel information on the global feature map and the local feature map based on the channel feature correlation matrix; The cross-type spatial feature fusion module is connected to the cross-domain channel attention module, and is used to perform interactive fusion of spatial information on the global feature map and the local feature map after interactive fusion.
[0015] In a possible implementation, the cross-domain channel attention module includes: A first branch, connected to the transformer module, and used to mine the channel information inside the global feature map; A second branch, connected to the convolutional neural network module, and used to mine the channel information inside the local feature map; An outer product block, respectively connected to the first branch and the second branch, and used to construct a channel correlation matrix based on the channel information inside the global feature map and the local feature map; A local softmax block, connected to the outer product block, and used to adjust the dimension of the local channel attention in the channel correlation matrix; The global softmax block, connected to the outer product block, is used to adjust the dimension of the global channel attention in the channel correlation matrix; The global subspace block, connected to the transformer module and the global softmax block respectively, performs a one-mode tensor product of the global feature map and the global channel attention output by the global softmax block to obtain the global subspace; The local subspace block, connected to the convolutional neural network module and the local softmax block respectively, performs a one-mode tensor product of the local feature map and the local channel attention output by the local softmax block to obtain the local subspace; The global feature fusion block, connected to the transformer module and the local subspace block respectively, is used to fuse the local features in the local feature map into the global feature map; The local feature fusion block, connected to the convolutional neural network module and the global subspace block respectively, is used to fuse the global features in the global feature map into the local feature map.
[0016] In a possible implementation manner, the first branch includes a first adaptive average pooling block, a first linear compression block, a first ReLu activation function block, a first linear excitation block, and a first Sigmoid compression function block connected in sequence; The first adaptive average pooling block compresses the global feature map channel by channel to obtain more lightweight global channel-level statistical information, and the expression is: , where represents the global feature map, , represents the global channel-level statistical information, ; The first linear compression block compresses the global channel-level statistical information and maps the global channel-level statistical information to ; after non-linearly mapping through the first ReLu activation function block, the first linear excitation block expands to , and finally the first Sigmoid compression function block compresses the mapped global channel-level statistical information; the expression of the above process is: ; represents the linear compression function for the channel-level statistical information, represents the linear excitation function for the channel-level statistical information, represents the Sigmoid function function, represents the global channel attention; The second branch includes a second adaptive average pooling block, a second linear excitation block, a second ReLu activation function block, a second linear compression block, and a second Sigmoid compression function block connected in sequence. Among them, the second adaptive average pooling block compresses the local feature map channel by channel into more lightweight local channel-level statistical information. The expression is: , where represents the local feature map, , represents the local channel-level statistical information, ; The second linear excitation block excites the local channel-level statistical information, mapping the local channel-level statistical information from to ; After non-linear mapping through the second ReLu activation function block, the second linear compression block maps the local channel-level statistical information from to ; Finally, the mapped local channel-level statistical information is compressed by the Sigmoid compression function block to between 0 and 1 to prevent probability overflow; The expression for the above process is: ; In the formula, represents the linear compression function for channel-level statistical information, represents the linear excitation function for channel-level statistical information, represents the Sigmoid function function; represents the local channel attention.
[0017] Compared with the prior art, the cross-domain channel attention module first performs a spatial transformation on the global feature map and the local feature map, turning their per-channel features into one-dimensional channel feature statistics, and then interacts with the global channel statistical information and the local channel statistical information through linear excitation and linear compression respectively, effectively reducing the number of parameters while mining the internal channel correlation; Then, a cross-channel correlation between the global channel statistical information and the local channel statistical information is constructed by the outer product block method, and softmax calculations are performed respectively in the convolutional dimension and the transformer dimension. Finally, the global feature map and the local feature map are multiplied by the channel correlation matrix respectively, realizing the attenuation and increase of the number of channels, and completing the mutual mapping and interaction between the local feature map and the global feature map.
[0018] In a possible implementation manner, the expression for the outer product block to construct the channel correlation matrix is: , where , where \(T\) is the transpose of the matrix; The expression for the local softmax block to adjust the dimension of local channel attention is: ; In the formula, denotes the subspace of The expression for the global softmax block to adjust the dimension of global channel attention is: ; In the formula, denotes the subspace of The expression for the global feature fusion block to fuse local features in the local feature map into the global feature map is: ; In the formula, denotes the global feature fusion map; The expression for the local feature fusion block to fuse global features in the global feature map into the local feature map is: , in the formula, denotes the local feature fusion map.
[0019] In a possible implementation, the cross - type spatial feature fusion module includes a \(3\times3\) convolution block, a \(5\times5\) convolution block, and a \(3\times3\) output convolution block. The output of the \(3\times3\) convolution block is skip - connected to the local feature fusion map and then connected to the \(3\times3\) output convolution block. The output of the \(5\times5\) convolution block is skip - connected to the global feature fusion map and then connected to the \(3\times3\) output convolution block, where: The \(3\times3\) convolution block is connected to the output end of the global feature fusion block, and is used to perform on the global feature fusion map , converting the channel feature dimension of the global feature fusion map from to . The expression for the output of the \(3\times3\) convolution block to skip - connect is: ; The \(5\times5\) convolution block is connected to the output end of the local feature fusion block, and is used to perform on the local feature fusion map , converting the channel feature dimension of the local feature fusion map from to ; The output of the 5×5 convolution block is skip-connected to the global feature fusion map The expression is: ; The 3×3 output convolution block concatenates the input and on the channel dimension. The expression is: ; In the formula, represents concatenation on the channel dimension; represents the feature concatenation map.
[0020] Compared with the prior art, the cross-type feature fusion module performs 5⨯5 convolution on the local feature fusion map after cross-fusion of the cross-domain channel attention module to capture a larger receptive field, and performs 3⨯3 convolution on the global feature fusion map to capture local features, and constructs the final feature map through addition and concatenation operations; to avoid excessive channel features received by the decoder layer and redundant information calculation, the cross-type feature fusion module adds a final 3⨯3 data convolution block to compress the feature channels, and the method realizes the gradual fusion of spatial information through two crosses and effectively reduces the huge difference in spatial features.
[0021] In a possible implementation manner, the medical image segmentation method further includes step 3. Obtaining a plurality of data sets containing medical images, and dividing the data sets into a training set and a test set; Step 4. Setting a loss function, training the image segmentation model based on the training set and the inspection function; and then testing the image segmentation model based on the test set; The loss function is an adopted balanced joint loss function, expressed as: ; In the formula, is a weight factor; represents the Dice loss function, represents the cross-entropy loss function. Description of the Drawings
[0022] Figure 1 is the framework diagram of the image segmentation model of the present invention; Figure 2 is the framework diagram of each fusion layer of the present invention; Figure 3 is the framework diagram of the second decoder layer or the third decoder layer or the fourth decoder layer or the fifth decoder layer of the present invention; Figure 4 is the framework diagram of the first decoder layer of the present invention; Figure 5This is the segmentation result diagram of breast ultrasound images in Experiment D1 of this specific embodiment; Figure 6 This is the segmentation result diagram of skin images in Experiment D2 of this specific embodiment; Figure 7 This is the segmentation result diagram of intestinal polyp endoscopic images in Experiment D3 of this specific embodiment; Figure 8 This is the segmentation result diagram of multi-organ CT images in Experiment D4 of this specific embodiment; Figure 9 This is the segmentation result diagram of brain tumor MRI images in Experiment D5 of this specific embodiment; Figure 10 This is the analysis diagram of GPU resource usage in the experiment of this specific embodiment; Figure 11 This is the analysis diagram of the average inference speed in the experiment of this specific embodiment. Detailed implementation manners
[0023] First of all, those skilled in the art should understand that these implementation manners are only used to explain the technical principles of the embodiments of the present application, and are not intended to limit the protection scope of the embodiments of the present application. Those skilled in the art can make adjustments according to needs to adapt to specific application scenarios.
[0024] In the description of the embodiments of the present application, it should be noted that unless otherwise clearly specified and limited, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific situations.
[0025] In the embodiments of the present application, unless otherwise clearly specified and limited, the first feature being "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on" the second feature can be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "below" and "beneath" the second feature can be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.
[0026] The present invention Figure 1 and Figure 2 The English characters in it are translated as follows: "feature map multiplication" represents multiplying feature maps; "feature map concatenation" represents concatenating feature maps; "pixel-wise addition" represents adding pixels element by element; "sigmoid" represents the sigmoid function; "compression" represents linear compression; "excitation" represents linear excitation; "adaptive average pooling" represents adaptive average pooling; "upsampling" represents upsampling; "skip connection" represents skip connection.
[0027] Figure 10 and Figure 11 In the horizontal coordinate, "GUP memory usage" represents CPU memory usage, and in the vertical coordinate, "average dice score" represents the average dice score.
[0028] The following further elaborates on this application in conjunction with the accompanying drawings and specific embodiments.
[0029] See Figures 1 to 4 As shown, an embodiment of this application discloses a medical image segmentation method based on a hybrid convolutional neural network and a transformer, including: Step 1. Obtain the medical image to be segmented; in this specific embodiment, a medical image with a size of is collected. Step 2. Construct an image segmentation model, and input the medical image to be segmented into the image segmentation model. The image segmentation model includes a preprocessing layer, a hybrid encoder layer, and a decoder layer; where: The preprocessing layer is used to perform image segmentation processing and local feature extraction on the medical image; in this specific embodiment, a convolutional neural network (CNN) module is used for local feature extraction. The convolutional neural network module uses ResNet34 as a framework to perform local feature extraction on the medical image, obtaining a local feature map with a size of ; the image segmentation module is used to perform image segmentation processing on the input medical image. The hybrid encoder layer is connected to the preprocessing layer, and is used to extract a global feature map and a local feature map from the segmented image and local features, construct a channel feature correlation matrix based on the channel features of the global feature map and the local feature map, perform interactive fusion of channel information on the global feature map and the local feature map based on the channel feature correlation matrix, and then perform spatial information interactive fusion on the globally and locally feature maps after interactive fusion; the network structure of the hybrid encoder layer includes a first fusion layer, a second fusion layer, a third fusion layer, and a fourth fusion layer connected in sequence. The decoder layer is used to splice and upsample the local features output by the preprocessing layer and the features after the spatial information interaction and fusion of the hybrid encoder layer, and output the target segmentation image; the decoder layer includes a fifth decoder layer, a fourth decoder layer, a third decoder layer, a second decoder layer, and a first decoder layer connected in sequence; the second decoder layer to the fifth decoder layer have the same network structure, all using a CNN decoder layer, and all include a first convolutional block, a second convolutional block, and a transposed convolutional block connected in sequence. The first decoder layer includes two CNN decoder layers and a 1×1 convolutional block connected in sequence, and the network structures of the two CNN decoder layers are the same as those of the second decoder layer to the fifth decoder layer.
[0030] Compared with the prior art, the present invention jointly performs an upsample operation through the double convolutional block and the transposed convolutional block of the decoder layer, which helps to restore the local feature map and the spatial feature map in terms of spatial dimensions, and at the same time ensures the lightweight of the image segmentation model.
[0031] The network structures of the first convolutional block and the second convolutional block both include a 3×3 convolutional block, a BN block, a ReLu function block, a 3×3 convolutional block, a BN block, and a ReLu function block connected in sequence for two 3×3 convolutional operations, normalization operations, and activation operations.
[0032] The data processing relationships between the first fusion layer to the fourth fusion layer and between the first decoder layer to the fifth decoder layer include: The first fusion layer performs interactive fusion of communication information and then transmits it to the second fusion layer. The second fusion layer performs interactive fusion of channel information and then transmits it to the third fusion layer. The third fusion layer performs interactive fusion of channel information and then transmits it to the fourth fusion layer. The fourth fusion layer performs interactive fusion of spatial information to obtain a feature map with a size of , and then enters the fifth decoder layer for an upsample operation. The third fusion layer performs interactive fusion of spatial information to obtain a feature map with a size of , and then is spliced with the features output by the fifth decoder layer and enters the fourth decoder layer for an upsample operation. The second fusion layer performs interactive fusion of spatial information to obtain a feature mosaic map with a size of , and then is spliced with the features output by the fourth decoder layer and enters the third decoder layer for an upsample operation. The first fusion layer performs interactive fusion of spatial information to obtain a feature mosaic map with a size of 4, and then is spliced with the features output by the third decoder layer and enters the second decoder layer for an upsample operation. The local features output by the preprocessing layer are spliced with the features output by the second decoder layer and enter the first decoder layer for a double upsample operation and then output to obtain the target segmentation image.
[0033] The network structures of the first fusion layer, the second fusion layer, the third fusion layer, and the fourth fusion layer are the same, and each includes a Transformer module, a convolutional neural network module, a cross-domain channel attention module, and a cross-type spatial feature fusion module; specifically, it includes: The Transformer module is used to extract the global feature map; the cross-domain channel attention module (CFCA module) is connected to the Transformer module and the convolutional neural network module respectively, and is used to construct a channel feature correlation matrix, and perform interactive fusion of channel information on the global feature map and the local feature map based on the channel feature correlation matrix; the cross-type spatial feature fusion module (XFF module) is connected to the cross-domain channel attention module, and is used to perform interactive fusion of spatial information on the globally and locally feature maps after interactive fusion; The convolutional neural network module (CNN module) is used to extract the local feature map; the convolutional neural network module ResNet34 (He et al., 2016) of this specific embodiment is used as a framework to extract the local feature map ; , represents the number of channels of the convolutional neural network module, respectively represent the height and width of the feature map; The cross-type spatial feature fusion module is connected to the cross-domain channel attention module, and is used to perform interactive fusion of spatial information on the globally and locally feature maps after interactive fusion.
[0034] The specific network structure and data processing method of the cross-domain channel attention module specifically include: The first branch, connected to the Transformer module, is used to mine the channel information inside the global feature map; the first branch includes a first adaptive average pooling block (Adaptive Average Pooling, AAP), a first linear compression block (Linear), a first ReLu activation function block (ReLu), a first linear excitation block (Linear), and a first Sigmoid compression function block connected in sequence; The first adaptive average pooling block (Adaptive Average Pooling, AAP) compresses the global feature map channel by channel to obtain lighter global channel-level statistical information, and the expression is: , where, represents the global feature map, , represents the global channel-level statistical information, ; The first linear compression block (Linear) compresses the global channel-level statistical information and maps the global channel-level statistical information to ; After non - linear mapping through the first ReLu activation function block, the first linear excitation block expands to , and finally, the first Sigmoid compression function block is used to compress the mapped global channel - level statistical information; the expression for the above process is: ; represents the linear compression function for channel - level statistical information, represents the linear excitation function for channel - level statistical information, represents the Sigmoid function function, represents the global channel attention.
[0035] The second branch, connected to the convolutional neural network module, is used to mine the channel information inside the local feature map; the second branch includes a second Adaptive Average Pooling block (AAP), a second Linear excitation block, a second ReLu activation function block, a second Linear compression block, and a second Sigmoid compression function block connected in sequence. Among them, the second Adaptive Average Pooling block compresses the local feature map for each channel feature map into more lightweight local channel - level statistical information, and the expression is: , where represents the local feature map, , represents the local channel - level statistical information, ; The second linear excitation block stimulates the local channel - level statistical information, mapping the local channel - level statistical information from to ; After non - linear mapping through the second ReLu activation function block, the second linear compression block maps the local channel - level statistical information from to ; Finally, the mapped local channel - level statistical information is compressed by the Sigmoid compression function block to between 0 and 1 to prevent probability overflow; the expression for the above process is: ; In the formula, represents the linear compression function for channel - level statistical information, represents the linear excitation function for channel - level statistical information, represents the Sigmoid function function; Represents local channel attention.
[0036] The outer product block is connected to the first branch and the second branch respectively, and constructs a channel correlation matrix based on the channel information inside the global feature map and the local feature map; the expression of the channel correlation matrix is: , where in the formula, , The T in is the transpose of the matrix; The local softmax block is connected to the outer product block and is used to adjust the dimension of the local channel attention in the channel correlation matrix; The global softmax block is connected to the outer product block and is used to adjust the dimension of the global channel attention in the channel correlation matrix; The global subspace block is connected to the transformer module and the global softmax block respectively, and performs a one-mode tensor product of the global feature map and the global channel attention output by the global softmax block to obtain a global subspace; the expression is: ; where in the formula, represents the subspace of.
[0037] The local subspace block is connected to the convolutional neural network module and the local softmax block respectively, and performs a one-mode tensor product of the local feature map and the local channel attention output by the local softmax block to obtain a local subspace; the expression is: ; where in the formula, represents the subspace of.
[0038] The global feature fusion block is connected to the transformer module and the local subspace block respectively, and is used to fuse the local features in the local feature map into the global feature map; the expression is: ; where in the formula, represents the global feature fusion map.
[0039] The local feature fusion block is connected to the convolutional neural network module and the global subspace block respectively, and is used to fuse the global features in the global feature map into the local feature map; the expression is: , where in the formula, represents the local feature fusion map.
[0040] The specific network structure and its image processing process of the cross-type spatial feature fusion module in this specific embodiment include: The cross - type spatial feature fusion module includes a 3×3 convolution block, a 5×5 convolution block, and a 3×3 output convolution block. The output of the 3×3 convolution block is skip - connected to the local feature fusion map and then connected to the 3×3 output convolution block. The output of the 5×5 convolution block is skip - connected to the global feature fusion map and then connected to the 3×3 output convolution block. Among them: The 3×3 convolution block is connected to the output end of the global feature fusion block, and is used to perform on the global feature fusion map to convert the channel feature dimension of the global feature fusion map from to . The expression of the skip - connection of the output of the 3×3 convolution block is: The 5×5 convolution block is connected to the output end of the local feature fusion block, and is used to perform on the local feature fusion map to convert the channel feature dimension of the local feature fusion map from to . The expression of the skip - connection of the output of the 5×5 convolution block is: The 3×3 output convolution block concatenates the input and in the channel dimension. The expression is: ; In the formula, represents concatenation in the channel dimension; represents the feature concatenation map.
[0041] Step 3. Obtain multiple datasets containing medical images and divide the datasets into a training set and a test set; in this specific embodiment, 8 datasets are obtained, including: BUSI (Al-Dhabyani et al., 2020), Dataset B (Yap et al., 2017), ISIC2016 (Gutman et al., 2016), PH2 (Mendonça et al., 2013), KvasirSeg (Jha et al., 2020), CVC-ClinicDB (Jha et al., 2019), Synapse multi-organ segmentation dataset (Landman et al., 2015) and Brain-MRI (Buda et al., 2019). The above 8 datasets cover five modalities: ultrasound imaging (US), dermoscopy imaging, computed tomography (CT), colonoscopy, and magnetic resonance imaging (MRI).
[0042] Step 4. Set the loss function and train the image segmentation model based on the training set and the inspection function; then test the image segmentation model based on the test set.
[0043] The loss function is an adopted balanced combined loss function, expressed as: ; In the formula, is the weight factor; represents the Dice loss function, represents the cross-entropy loss function. In order to balance the accuracy of pixel-level classification and the optimization of the global region, in this specific embodiment, is set; in order to ensure that during the training process, and both contribute equally, thus avoiding the image segmentation model from over-focusing on pixel-level classification; when the image segmentation model tends to prioritize the consistency of the global segmentation region and may ignore fine-grained pixel-level classification; when the image segmentation model may perform better in pixel-level classification but fails to fully optimize the consistency of the global segmentation region.
[0044] Experiment A. Experimental settings To alleviate overfitting and improve the generalization ability of the model, various data augmentation techniques were applied in this experiment, including: random cropping at a ratio of 0.5, random horizontal flipping at a probability of 0.5, random vertical flipping at a probability of 0.5, and random rotation of ±15 degrees at a probability of 0.6. We standardized the images with a mean of [0.485, 0.456, 0.406] and a standard deviation of [0.229, 0.224, 0.225]. These data augmentation strategies were applied to all datasets except the Synapse dataset.
[0045] This experiment was conducted under the PyTorch framework, and all models were trained and tested on NVIDIA A5000 GPUs. A random seed of 42 was set for all models, including the worker initialization of the data loader, data extraction, and data partitioning. The total number of training epochs was set to 130, including 10 warm-up epochs and 120 training epochs. The AdamW optimizer was used with a weight decay of , and the betas parameters were (0.9, 0.999). The initial learning rate was set to 0.0003, and the "Poly" learning rate strategy was adopted with a power . Subsequently, the code will be released on Github for further exploration.
[0046] B. Datasets The tests were conducted on 8 datasets, which included a total of 5 modalities. The specific data volumes for training, validation, testing, and the image scaling sizes are summarized in Table 1 below:
[0047] Table 1 Dataset Partition C. Evaluation Metrics and Comparison Methods In the experiment, multiple evaluation metrics were adopted to strictly evaluate the performance of the model in different modalities, including the Dice coefficient, Jaccard index, and 95th percentile Hausdorff distance (HD95). The specific descriptions of the key evaluation metrics are as follows:
[0048] Dice coefficient: The Dice similarity coefficient is used to measure the overlap between the predicted segmentation and the ground truth segmentation, and is particularly effective in dealing with class imbalance problems. A higher Dice score indicates a higher similarity between the predicted segmentation and the ground truth segmentation.
[0049] Jaccard Index: Evaluates the similarity between the predicted segmentation and the ground truth segmentation. Compared to Dice, it penalizes false positives and false negatives more severely. A higher Jaccard score indicates better segmentation accuracy.
[0050] HD95: The 95th percentile Hausdorff distance (HD95) quantifies the spatial distance between the boundaries of the predicted segmentation and the ground truth segmentation, focusing on the largest deviations and ignoring extreme outliers. A lower HD95 value indicates that the boundaries of the predicted segmentation are closer to the boundaries of the ground truth segmentation.
[0051] To evaluate the efficiency of the model in practical applications and its computational resource requirements, GPU memory usage and frames per second (FPS) were used as the main evaluation metrics in the experiment. The number of parameters was not used as a standard metric because it does not directly reflect the actual storage requirements or running efficiency of the model. For example, the sparsity of the weight matrix may result in a higher number of parameters, but the memory footprint may remain negligible. Additionally, while FLOPs is an important metric for measuring computational complexity, it is not necessarily related to the actual inference speed. The actual performance often depends on the degree of model optimization and the support of the underlying hardware. Therefore, using GPU memory usage and FPS as evaluation metrics can more intuitively and accurately reflect the performance of the model in practical applications, making this method more persuasive and relevant in practical applications.
[0052] To ensure comprehensive benchmarking, in the experiment, the image segmentation model designed in this invention was compared with a variety of state-of-the-art (SOTA) methods, covering CNN-based models and hybrid CNN-Transformer architectures. For all possible CNN models and hybrid CNN-Transformer models, their pre-trained weights were called as much as possible to maintain the rigor of the experiment and ensure fair comparison. CNN-based models include U-Net (Ronneberger et al., 2015), Attention U-Net (Oktay et al., 2018), ResUnet (Diakogiannis et al., 2020), FATnet (Wu et al., 2022), DCSAUnet (Xu et al., 2023), M2Snet (Zhao et al., 2023), CMUNeXt-Large (Tang et al., 2024) and I2U-Net-Large (Dai et al., 2024), while hybrid CNN-Transformer models include MISSFormer (Huang et al., 2021), TransUnet (Chen et al., 2021), HiFormer (Base) (Heidari et al., 2023), H2Former (He et al., 2023a) and BEFUnet (Manzari et al., 2024).
[0053] D. Experimental Results D1. Ultrasonic Image Segmentation Challenge Breast ultrasound images usually have characteristics such as uniform color gamut intensity distribution, blurred boundaries, and irregular tumor morphology, which may indirectly affect the performance of the model (Zhang et al., 2024). Therefore, this poses a major challenge for the model to effectively capture global features.
[0054] Table 2 Quantitative Results Table for Breast Ultrasound Image Segmentation The quantitative results of breast ultrasound image segmentation are shown in Table 2 and Figure 5As shown, the image segmentation model designed by the present invention performs best in terms of Dice coefficient, Jaccard index, and HD95 metric on the BUSI dataset and the B dataset. As shown in the table, on the BUSI dataset, the image segmentation model of the present invention exceeds the SOTA model H2Former (He et al., 2023a) by 1.31% in terms of Dice coefficient, exceeds by 1.81% in terms of Jaccard index, and achieves a lower HD95 value of 7.48. At the same time, on the Dataset B dataset, the image segmentation model of the present invention exceeds the SOTA model HiFormer-Base (Heidari et al., 2023) by 2.37% in terms of Dice coefficient, exceeds by 3.01% in terms of Jaccard index, and the HD95 is 3.47. It should be noted that the Dataset B dataset is a small dataset, and this evaluation also tests whether the state-of-the-art (SOTA) model can still achieve accurate segmentation in the case of insufficient medical image data.
[0055] To evaluate the generalization ability of the model, a domain transfer experiment was also conducted, in which the model was trained on the relatively large BUSI dataset and tested on the Dataset B dataset. The results show that the domain transfer performance of the image segmentation model of the present invention exceeds that of the model directly trained on the Dataset B dataset in all metrics. As shown in Table 2, we observe that M2Snet (Zhao et al., 2023), TransUnet (Chen et al., 2021), and the image segmentation model of the present invention have significant improvements in all metrics, while the performance of other models remains unchanged or decreases. This indicates that there are still obvious differences in data distribution between the two datasets, and other models encounter problems due to overly high or low model complexity. These problems may significantly limit the application of these models in actual medical image segmentation tasks. In the domain transfer experiment, our model exceeded the SOTA model TransUnet, achieving a Dice coefficient of 89.52, a Jaccard of 81.81, and an HD95 of 4.01.
[0056] D2. Dermoscopic Image Segmentation Challenge Compared with ultrasound images, dermoscopic images have higher resolution and less noise, resulting in better image quality and more obvious color features. In the experiment, we used the relatively large dataset ISIC-2016 (Gutman et al., 2016) and a smaller dataset PH2 (Mendonça et al., 2013) to evaluate the segmentation performance of our model. In this experiment, we continued to evaluate the generalization ability of the model through the domain transfer scenario.
[0057] Although both datasets focus on melanoma segmentation, the PH2 dataset contains a greater variety of non-melanoma samples, such as 80 common moles and atypical moles. This setting challenges the generalization ability of the model and its performance in segmenting abnormal data.
[0058] Table 3 Quantitative results of skin image segmentation As shown in Table 3 and Figure 6 As shown, on the ISIC-2016 dataset (Gutman et al., 2016), most models demonstrated strong segmentation performance, indicating relatively low data complexity. After analysis, we observed that CNN-based models performed comparably to hybrid models, suggesting that the clear boundaries and distinct color contrasts in this dataset are particularly beneficial for the CNN architecture.
[0059] On the ISIC2016 dataset (Gutman et al., 2016), the image segmentation model of the present invention achieved state-of-the-art performance, with Dice, Jaccard, and HD95 scores of 92.20, 86.55, and 3.06, respectively. Meanwhile, on the PH2 dataset (Mendonc¸a et al., 2013) containing more sample types, the image segmentation model of the present invention outperformed existing methods, achieving Dice, Jaccard, and HD95 scores of 95.14, 90.85, and 0.82, respectively. In the domain transfer experiment, the image segmentation model of the present invention ranked third on average in all metrics, demonstrating excellent generalization ability and robust segmentation performance when dealing with abnormal data. This highlights the effectiveness of the model in addressing cross-domain challenges in medical image segmentation.
[0060] D3. Intestinal polyp endoscopic image segmentation challenge Intestinal polyp endoscopic images exhibit significant variability in polyp shape, size, color, location, and texture, which poses a great challenge for the model to accurately capture semantic features and boundary recognition. In this study, we evaluated the segmentation performance of the model on two datasets, namely Kvasir-SEG (Jha et al., 2020) and CVC-ClinicDB (Zhou et al., 2019), where the number of samples provided by Kvasir-SEG is approximately twice that of CVC-ClinicDB. The comparison and domain transfer results are summarized in Table 4 and Figure 7 as follows.
[0061] Table 4 Domain transfer results of intestinal polyp endoscopic images In Table 4, the segmentation performance of the image segmentation model of the present invention on polyp images is excellent. On the Kvasir-SEG dataset (Jha et al., 2020), we outperformed the SOTA model by 1.93% in the Dice metric, by 2.99% in the Jaccard metric, and the lowest HD95 value was 5.73. On the CVC-ClinicDB dataset, we exceeded M2Snet in all metrics, achieving Dice, Jaccard, and HD95 values of 93.86, 88.71, and 1.77 respectively. In the domain transfer experiment, we first trained the model on the Kvasir-SEG dataset (Jha et al., 2020) and then tested it on the CVC-ClinicDB dataset (Zhou et al., 2019). The results showed that the performance of all models in the domain transfer experiment was lower than that of training directly on the CVC-ClinicDB dataset (Zhou et al., 2019), which was due to the differences in the datasets themselves. However, the image segmentation model of the present invention still maintained SOTA performance even after domain transfer, fully demonstrating its excellent generalization ability.
[0062] D4. Multi-organ CT Image Segmentation Challenge The Synapse dataset (Landman et al., 2015) was selected for this challenge to evaluate the performance of the model in multi-class segmentation tasks. The significant morphological differences between organs and tissues, as well as the fact that the data was sourced from 3D scans (not every CT image contains all organs), posed great challenges for the model to learn spatial relationships and context information. Table 5 shows the performance of our model on the 8-organ segmentation task of the Synapse dataset, Figure 8 and presents some visualization results.
[0063] Table 5 Organ Segmentation Task Data The results show that the average Dice score of the image segmentation model of the present invention on 8 organs is 2.03% higher than that of H2Former (He et al., 2023a), and the average HD is 8.90. In the segmentation challenge of 8 organs, the image segmentation model of the present invention outperforms the SOTA in the segmentation of the right kidney, liver and stomach, achieving Dice scores of 91.63, 95.41 and 84.96 respectively. In addition, the image segmentation model of the present invention also achieves the second-best performance in the segmentation of the spleen, left kidney, gallbladder and pancreas. The performance of the image segmentation model of the present invention on the aorta is also highly competitive. Therefore, through the multi-class segmentation challenge, the image segmentation model of the present invention demonstrates the ability to handle complex variations and exhibits strong generalization ability. By combining the feature maps of CNN and Transformer, the model's ability to learn context information has been significantly improved.
[0064] D5. Brain Tumor MRI Image Segmentation Challenge Irregular shapes, heterogeneity, and low contrast remain significant challenges in brain tumor MRI image segmentation. In this study, we used a brain MRI segmentation dataset to evaluate the model's ability to capture context and semantic information. The experimental results are shown in Table 6, and some visualization results are presented in Figure 9 .
[0065] The image segmentation model of the present invention outperforms the existing state-of-the-art (SOTA) in multiple metrics, including Dice, Jaccard, recall, pixel accuracy, and HD95. Our model is 0.59% higher than HiFormer-Base (Heidari et al., 2023) in terms of the Dice score, reaching a score of 88.18, and 3.57% higher than H2Former (He et al., 2023a). In terms of the Jaccard index, the image segmentation model of the present invention exceeds HiFormer-Base (Heidari et al., 2023) by 0.86%. In terms of pixel accuracy and HD95, the image segmentation model of the present invention also outperforms the existing state-of-the-art technologies, achieving scores of 99.53 and 1.89 respectively. The image segmentation model of the present invention is also highly competitive with other SOTA models in terms of precision and recall.
[0066] Table 6 Brain Tumor MRI Image Segmentation Data Table E. Ablation Experiments In the ablation experiments, first, the combination of a CNN encoder and a decoder was evaluated, where the CNN encoder used ResNet34 (He et al., 2016) as the backbone network, and its segmentation performance on eight datasets was evaluated. In addition, the performance of using Swin Transformer V2 (Liu et al., 2022) as the encoder paired with the decoder of the present invention was also tested. The results showed that the convolutional neural network module performed better than the Transformer module on the ultrasound datasets, while the Transformer was more excellent in the polyp segmentation task.
[0067] Next, a dual-encoder structure was tried, using a simple convolutional layer to fuse the feature maps. However, the results showed that the performance of this method was not as good as using a single encoder. This finding highlighted the significant differences between the convolutional neural network module and the Transformer module in terms of spatial features and channel features, and simple convolutional operations could not effectively eliminate these differences.
[0068] To handle the differences in channel features, especially when the number of feature maps was inconsistent, a selection mechanism was introduced in the experiments to screen and map the features. Specifically, a matrix was designed to map the channel features according to the features of the two encoders. In the model architecture of the present invention, the local features extracted by the convolutional neural network module were fused with the global features extracted by the Transformer after channel mapping as the input to the next convolutional neural network module. Similarly, the global feature fusion was fused with the local features after channel selection as the input to the next Transformer module. This method enabled the convolutional neural network module to obtain global features with a larger receptive field, while providing more detailed local features for the Transformer module.
[0069] In addition, the present invention integrated a cross-style feature fusion module into the model to effectively fuse the spatial features. Through iterative convolutional operations and feature fusion, the significant differences in spatial features were gradually alleviated. As shown in Table 7, the image segmentation model of the present invention achieved significant improvements in the Dice and Jaccard metrics and showed high competitiveness in the HD95 metric. These results strongly verified the effectiveness of the proposed CFCA and XFF modules.
[0070] Table 7 Data table of ablation experiments F. Performance analysis The number of individual parameters is insufficient to comprehensively capture the actual computational load of the model on the GPU. Therefore, we use the peak GPU memory usage to provide a more comprehensive perspective on GPU resource consumption. Similarly, although FLOPs (floating-point operations per second) can measure computational complexity, they do not reliably reflect the performance of the model in actual inference. For a more accurate evaluation, we report the FPS (frames per second) to directly represent the inference speed.
[0071] As Figure 10 shown, the image segmentation model of the present invention maintains a moderate level of GPU resource usage, significantly lower than the memory requirements of HiFormer-Base (Heidari et al., 2023), CMUNeXtLarge (Tang et al., 2024), DCSAUnet (Xu et al., 2023), BEFUnet (Manzari et al., 2024), and H2Former (He et al., 2023a). Despite the low memory footprint, the image segmentation model of the present invention still achieved the highest average Dice score on eight datasets compared with other state-of-the-art (SOTA) models, highlighting its efficiency in resource utilization without compromising segmentation accuracy.
[0072] In Figure 11 it, upon further observing the average inference speed of the image segmentation model of the present invention on 1600 images, it outperformed multiple hybrid models, including H2Former (He et al., 2023a), HiFormer-Base (Heidari et al., 2023), TransUnet (Chen et al., 2021), BEFUnet (Manzari et al., 2024), and CNN-based models such as I2U-Net-Large (Dai et al., 2024), DCSAUnet (Xu et al., 2023), ResUnet (Diakogiannis et al., 2020), and CMUNeXtLarge (Tang et al., 2024). Notably, this speed advantage combined with the highest average Dice score highlights the superiority of the image segmentation model of the present invention in terms of segmentation effect and inference efficiency. These findings indicate that the image segmentation model of the present invention achieves an optimal balance between GPU efficiency and competitive segmentation performance.
[0073] G. Summary The present invention proposes a novel hybrid CNN-Transformer architecture. By introducing a Cross-Domain Channel Attention (CFCA) module after the convolutional neural network module and the transformer module, the Cross-Domain Channel Attention (CFCA) module utilizes lightweight cross-channel attention calculation to map feature interactions between the convolutional neural network module and the transformer module. The Cross-Domain Channel Attention (CFCA) module enables local features to be integrated into global features while ensuring that the convolutional neural network module can access global feature information. In addition, the present invention also proposes a Spatial Feature Fusion (XFF) module, which efficiently performs two local and global feature fusions to provide key outputs for skip connections. This design significantly enhances the model's ability to reconstruct masks with high accuracy. Extensive experimental results on eight datasets and five modalities show that our model performs excellently in terms of segmentation performance and generalization ability.
[0074] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inner", "outer", etc., indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application.
[0075] In the description of the present application, the descriptions referring to terms such as "one embodiment", "some embodiments", "in this embodiment", "specific examples", or "some examples", etc., mean that the specific features, mechanisms, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0076] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0077] Compared with the prior art, the cross-domain channel attention module first performs a spatial transformation on the global feature map and the local feature map, turning their per-channel features into one-dimensional channel feature statistics. Then, through linear excitation and linear compression, it respectively interacts with the global channel statistics and the local channel statistics, effectively reducing the number of parameters while mining the internal channel correlation. Next, it constructs the cross-channel correlation between the global channel statistics and the local channel statistics through a one-mode tensor product method, and performs softmax calculations in the convolutional dimension and the transformer dimension respectively. Finally, it multiplies the global feature map and the local feature map with the channel correlation matrix respectively, achieving the attenuation and increase of the number of channels, and completing the mutual mapping and interaction between the local feature map and the global feature map.
[0078] In a possible implementation, the second decoder layer, the third decoder layer, the fourth decoder layer, and the fifth decoder layer have the same network structure, all including a first convolutional block, a second convolutional block, and a transposed convolutional block connected in sequence. The first decoder layer includes two layers of CNN decoder layers and a 1×1 convolutional block connected in sequence, and the network structures of the two layers of CNN decoder layers are the same as those of the second decoder layer to the fifth decoder layer. The network structures of the first convolutional block and the second convolutional block both include two 3×3 convolutional operations, normalization operations, and activation operations connected in sequence, which are a 3×3 convolutional block, a BN block, a ReLu function block, a 3×3 convolutional block, a BN block, and a ReLu function block.
[0079] Compared with the prior art, the cross-type feature fusion module performs 5⨯5 convolution on the locally feature-fused map after cross-fusion by the cross-domain channel attention module to capture a larger receptive field, and performs 3⨯3 convolution on the global feature-fused map to capture local features. The final feature map is constructed through addition and splicing operations. To avoid excessive channel features received by the decoder layer and the calculation of redundant information, the cross-type feature fusion module adds a final 3⨯3 data convolutional block to compress the feature channels. The method realizes the gradual fusion of spatial information through two crosses and effectively reduces the huge difference in spatial features.
[0080] In the description of the embodiments of the present application, it should be noted that in the description of the present application, the terms indicating the direction or positional relationship such as "inner" and "outer" are based on the direction or positional relationship shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of the present application.
[0081] In the description of the present application, the descriptions referring to terms such as "one embodiment", "some embodiments", "in this embodiment", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0082] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A medical image segmentation method based on a hybrid convolutional neural network and a transformer, characterized in that: include: Step 1, obtaining a medical image to be segmented; Step 2, constructing an image segmentation model, inputting the medical image to be segmented into the image segmentation model, wherein the image segmentation model includes a preprocessing layer, a hybrid encoder layer and a decoder layer; The preprocessing layer is used to perform image segmentation processing and local feature extraction on the medical image; The hybrid encoder layer is connected to the preprocessing layer, and is used to extract a global feature map and a local feature map from the segmented image and the local features, and construct a channel feature correlation matrix based on the channel features of the global feature map and the local feature map, perform interactive fusion of channel information on the global feature map and the local feature map based on the channel feature correlation matrix, and then perform interactive fusion of spatial information on the interactively fused global feature map and the local feature map; The decoder layer is used to perform splicing and up-sampling operations on the local features output by the preprocessing layer and the features after the hybrid encoder layer performs spatial information interactive fusion, and output a target segmentation image.
2. The medical image segmentation method based on a hybrid convolutional neural network and a transformer according to claim 1, characterized in that: The hybrid encoder layer includes a first fusion layer, a second fusion layer, a third fusion layer and a fourth fusion layer connected in sequence, and the decoder layer includes a fifth decoder layer, a fourth decoder layer, a third decoder layer, a second decoder layer and a first decoder layer connected in sequence. The first fusion layer performs interactive fusion of communication information and transmits it to the second fusion layer, the second fusion layer performs interactive fusion of channel information and transmits it to the third fusion layer, the third fusion layer performs interactive fusion of channel information and transmits it to the fourth fusion layer, the fourth fusion layer performs interactive fusion of spatial information and enters the fifth decoder layer for upsampling operation, the third fusion layer performs interactive fusion of spatial information and splices with the features output by the fifth decoder layer and then enters the fourth decoder layer for upsampling operation, the second fusion layer performs interactive fusion of spatial information and splices with the features output by the fourth decoder layer and then enters the third decoder layer for upsampling operation, the first fusion layer performs interactive fusion of spatial information and splices with the features output by the third decoder layer and then enters the second decoder layer for upsampling operation, the local features output by the preprocessing layer are spliced with the features output by the second decoder layer and then enter the first decoder layer for double upsampling operation and then output to obtain the target segmented image.
3. The medical image segmentation method based on a hybrid convolutional neural network and a transformer according to claim 2, characterized in that: The second decoder layer, the third decoder layer, the fourth decoder layer and the fifth decoder layer have the same network structure, and all include a first convolution block, a second convolution block and a deconvolution block connected in sequence. The first decoder layer includes two layers of CNN decoder layers and a 1×1 convolution block connected in sequence. The network structure of the two layers of CNN decoder layers is the same as the network structure of the second decoder layer to the fifth decoder layer.
4. The medical image segmentation method based on a hybrid convolutional neural network and a transformer according to claim 3, characterized in that: The network structures of the first convolution block and the second convolution block both include a 3×3 convolution block, a BN block, a ReLu function block, a 3×3 convolution block, a BN block, and a ReLu function block connected in sequence.
5. The medical image segmentation method based on hybrid convolutional neural network and transformer according to claim 2, characterized in that: The network structures of the first fusion layer, the second fusion layer, the third fusion layer and the fourth fusion layer are the same, and all include a transformer module, a convolutional neural network module, a cross-domain channel attention module and a cross-spatial feature fusion module; The transformer module is used to extract a global feature map; The convolutional neural network module is used to extract local feature maps; The cross-domain channel attention module is connected to the transformer module and the convolutional neural network module respectively, and is used to construct a channel feature correlation matrix, and perform interactive fusion of channel information on the global feature map and the local feature map based on the channel feature correlation matrix; The cross-spatial feature fusion module is connected to the cross-domain channel attention module to perform interactive fusion of spatial information on the interactively fused global feature map and the local feature map.
6. The medical image segmentation method based on hybrid convolutional neural network and transformer according to claim 5, characterized in that: The cross-domain channel attention module includes: The first branch is connected to the transformer module and is used to mine the channel information inside the global feature map; The second branch is connected to the convolutional neural network module and is used to mine the channel information inside the local feature map; The outer product block is connected to the first branch and the second branch respectively, and constructs a channel correlation matrix based on the channel information inside the global feature map and the local feature map; The local softmax block, connected to the outer product block, is used to adjust the dimension of the local channel attention in the channel correlation matrix; The global softmax block, connected to the outer product block, is used to adjust the dimension of the global channel attention in the channel correlation matrix; The global subspace block is connected to the transformer module and the global softmax block respectively, and the global feature map is multiplied by the global channel attention output by the global softmax block to obtain the global subspace; The local subspace block is connected to the convolutional neural network module and the local softmax block respectively, and the local subspace is obtained by performing a one-module tensor product of the local feature map and the local channel attention output by the local softmax block; A global feature fusion block, connected to the transformer module and the local subspace block, is used to fuse local features in the local feature map into the global feature map; The local feature fusion block is connected to the convolutional neural network module and the global subspace block respectively, and is used to fuse the global features in the global feature map into the local feature map.
7. The medical image segmentation method based on a hybrid convolutional neural network and a transformer according to claim 6, characterized in that: The first branch includes a first adaptive average pooling block, a first linear compression block, a first ReLu activation function block, a first linear excitation block and a first Sigmoid compression function block connected in sequence; The first adaptive average pooling block compresses the global feature map channel by channel to obtain more lightweight global channel-level statistical information, which is expressed as: , where represents the global feature map, , Represents global channel-level statistics, ; The first linear compression block compresses the global channel-level statistical information to convert the global channel-level statistical information into Map to After the first ReLu activation function block performs nonlinear mapping, the first linear excitation block will Expand to , and finally the first Sigmoid compression function block is used to compress the mapped global channel-level statistical information; the expression of the above process is: ; represents the linear compression function of channel-level statistical information, represents the linear activation function for channel-level statistical information, Represents Sigmoid function function, represents global channel attention; The second branch includes a second adaptive average pooling block, a second linear excitation block, a second ReLu activation function block, a second linear compression block, and a second Sigmoid compression function block connected in sequence, wherein the second adaptive average pooling block compresses the local feature map channel by channel feature map into more lightweight local channel-level statistical information, and the expression is: , where Represented as a local feature map, , represents local channel-level statistics, ; The second linear excitation block excites the local channel-level statistical information, converting the local channel-level statistical information from Map to After the second ReLu activation function block performs nonlinear mapping, the second linear compression block converts the local channel-level statistical information from Map to ; Finally, the mapped local channel-level statistical information is compressed to between 0 and 1 through the Sigmoid compression function block to prevent probability overflow; the expression of the above process is: ; In the formula, represents the linear compression function of channel-level statistical information, represents the linear activation function for channel-level statistical information, Represents Sigmoid function function; represents local channel attention.
8. The medical image segmentation method based on a hybrid convolutional neural network and a transformer according to claim 7, characterized in that: The expression of the channel correlation matrix constructed by the outer product block is: , where , T in the pair Transpose of a matrix; The expression for adjusting the dimension of the local channel attention by the local softmax block is: ; In the formula, express subspace of ; The expression for adjusting the dimension of global channel attention by the global softmax block is: ; In the formula, express subspace of ; The expression of the global feature fusion block fusing the local features in the local feature map into the global feature map is: ; In the formula, Represents the global feature fusion map; The expression of the local feature fusion block fusing the global features in the global feature map into the local feature map is: ; In the formula, Represents the local feature fusion map.
9. The medical image segmentation method based on hybrid convolutional neural network and transformer according to claim 8, characterized in that: The cross-space feature fusion module includes a 3×3 convolution block, a 5×5 convolution block and a 3×3 output convolution block. The output of the 3×3 convolution block jumps to connect the local feature fusion map. Then connected with the 3×3 output convolution block, the output of the 5×5 convolution block jumps to the global feature fusion map Then connected with a 3×3 output convolution block, where: The 3×3 convolution block is connected to the output end of the global feature fusion block to fusion the global feature map. conduct , the global feature fusion map The channel feature dimension changes from Convert to , the output of the 3×3 convolutional block is skip-connected The expression is: ; The 5×5 convolution block is connected to the output end of the local feature fusion block to fusion the local feature map. conduct , the local feature fusion map The channel feature dimension changes from Convert to ; The output of the 5×5 convolutional block jumps to connect the global feature fusion map The expression is: ; The 3×3 output convolution block is based on the input and The concatenation is performed on the channel dimension, and the expression is: ; In the formula, Indicates splicing in the channel dimension; Represents a feature concatenation graph.
10. The medical image segmentation method based on hybrid convolutional neural network and transformer according to claim 1, characterized in that: The medical image segmentation method further comprises: Step 3, obtaining multiple data sets containing medical images, and dividing the data sets into a training set and a test set; Step 4, set the loss function, train the image segmentation model based on the training set and the patrol function; then test the image segmentation model based on the test set; The loss function is a balanced joint loss function, expressed as: ; In the formula, is the weight factor; represents the Dice loss function, represents the cross entropy loss function.
Citation Information
Patent Citations
Ultrasonic image quantification method based on interactive fusion Transform
CN114863111A
Medical image segmentation method based on CNN-Transform parallel encoder
CN118297961A
Medical image segmentation method based on global and local feature joint learning and multi-scale feature fusion
CN118840548A
System and method for efficiently amalgamated CNN-transformer architecture for mobile vision applications
US20240193404A1
System and method for 3D medical image segmentation
US20240362788A1