An Image Segmentation Method Based on the Global Perception U-Net Model
The Global-Aware U-Net model with CFGC and CSCP modules addresses the challenge of balancing computational efficiency and feature capture in medical image segmentation, achieving superior performance by optimizing feature fusion and reducing semantic gaps.
Patent Information
- Application Number
- CN202510661781.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-05-22
AI Technical Summary
Existing medical image segmentation technology is difficult to effectively focus on local and global features while reducing the computational complexity, especially when dealing with organs or tissues with high overlap in different scales and backgrounds, the segmentation accuracy is insufficient.
Using the global perception U-Net model, by setting up CFGC modules and CSCP modules in each layer of the encoder and decoder, the layer by layer fusion of features and global feature representation is realized. The CFGC modules and CSCP modules are used to restore and cross-channel fusion layer by layer to optimize the semantic gap between the codecs.
While reducing the computational complexity, it improves the accuracy and stability of medical image segmentation, can effectively pay attention to local and global features, reduce semantic ambiguity, and improves the segmentation effect.
Smart Images

Figure CN120182309B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular to an image segmentation method based on a global perception U-Net model. Background Art
[0002] In the field of medical image segmentation, most organs, tissues, or lesion sites need to be extracted for medical analysis and research. Although they belong to the same type of organs or tissues, there are sometimes problems of inconsistent scales among different individuals, and some have a high degree of overlap with the background area, which leads to high requirements for the segmentation accuracy of the network in the medical field.
[0003] To solve the above problems, most of the existing technologies focus on improving the network's analysis ability for objects of different scales. For example, by stacking multiple convolutional layers to obtain features of different scales and more fully achieve feature fusion. However, stacking too many convolutional layers, especially large kernels, will not only lead to a sharp increase in the number of parameters and computational complexity, but even lead to a decline in the stability of the network; moreover, simply stacking convolutional layers cannot make the segmentation network get rid of the disadvantage that convolution pays more attention to local region information. On the contrary, Transformer gives full play to its advantage of focusing on global information in the field of medical image segmentation by virtue of its advantage of modeling long-range dependencies, but it brings a relatively high computational complexity, which will lead to high-cost model deployment in practical applications.
[0004] In view of this, the present invention aims to provide an image segmentation method based on a global perception U-Net model to solve the above-mentioned related problems. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that the prior art in the field of medical image segmentation is difficult to achieve the technical problem of simultaneously paying attention to local features and global features while reducing the computational complexity. The purpose is to provide an image segmentation method based on a global perception U-Net model. By setting a CFGC module with global perception collaborative fusion in each layer of the encoder and decoder, it can not only have the feature expression ability at different scales, but also strengthen the feature representation based on the global features at different scales. By setting a CSCP module and a CFGC module in the decoder, and making the layers of the encoder jump-connect to the CSCP modules of the corresponding layers of the decoder one by one, the CFGC module and the CSCP module can be used to recover features layer by layer to achieve effective fusion in the process of feature splicing of each layer of the encoder and decoder. At the same time, through the CSCP module that gradually fuses cross-channels, it can directly fuse the features in the encoder and decoder stages, optimize the semantic gap between the encoder and decoder without changing the jump connection structure, and will not cause additional semantic ambiguity. At the same time, the CSCP module with different divided channels is used to divide the feature images from the output of the encoder and the output of the transposed convolution module respectively, and gradually cross-splice and fuse the divided deep features and shallow features. And through different global average pooling units, global average pooling processing is performed on different feature images to obtain different feature weight vectors. By combining the feature weight vectors with the spliced and fused feature images, it is possible to simultaneously pay attention to the important semantic information in the encoder stage and the decoder channels.
[0006] The present invention is realized through the following technical solutions:
[0007] An image segmentation method based on a global perception U-Net model, the method includes:
[0008] Obtain the medical image to be segmented, input the medical image to be segmented into the global perception U-Net model for image segmentation processing, and obtain the segmented image segmentation result;
[0009] Among them, the global perception U-Net model includes an encoder, a decoder, a CFGC module and a CSCP module. Both the encoder and the decoder adopt a five-layer structure connected in sequence. The first layer of the encoder adopts a first convolution module and a max-pooling module connected in sequence. The remaining four layers of the encoder both adopt a CFGC module and a max-pooling module connected in sequence. The last layer of the decoder adopts a second convolution module. The remaining four layers of the decoder both adopt a transposed convolution module, a CSCP module and a CFGC module connected in sequence. The first four layers of the encoder are jump-connected to the first four layers of the decoder respectively; the CFGC module includes H a dimensional feature selection unit and W a dimensional feature selection unit, which is used to select the feature information in the H and W dimensions at different scales, and based onH The dimension feature selection unit and W The dimension feature selection unit enhances the global features in the way of matrix multiplication; the CSCP module includes a feature image division channel, a feature splicing and fusion unit, and a global average pooling unit, which are used to perform channel-wise feature averaging division on the feature images in the encoder and decoder stages, splice the divided features in the encoder and decoder stages respectively, and complete step-by-step feature fusion based on the channel dependence relationship between the feature images in the encoder and decoder stages.
[0010] Furthermore, the first convolution module in the first layer of the encoder is jump-connected to the CSCP module in the fourth layer of the decoder, the CFGC module in the second layer of the encoder is jump-connected to the CSCP module in the third layer of the decoder, the CFGC module in the third layer of the encoder is jump-connected to the CSCP module in the second layer of the decoder, and the CFGC module in the fourth layer of the encoder is jump-connected to the CSCP module in the first layer of the decoder.
[0011] Furthermore, the CFGC module in the last layer of the encoder is connected to the transposed convolution module in the first layer of the decoder.
[0012] Furthermore, the CFGC module includes a first depth convolution unit, a second depth convolution unit, a first feature splicing and fusion unit, a first H dimension feature selection unit, a first W dimension feature selection unit, a second H dimension feature selection unit, a second W dimension feature selection unit, a second feature splicing and fusion unit, a third feature splicing and fusion unit, a third standard convolution unit, a fourth standard convolution unit, a fifth standard convolution unit, a first feature shape reshaping unit, a second feature shape reshaping unit, a first global feature calculation unit, a first activation function, a second feature calculation unit, and a sixth standard convolution unit;
[0013] The input ends of the first depth convolution unit and the second depth convolution unit jointly serve as the input end of the CFGC module;
[0014] The output end of the first depth convolution unit is respectively connected to the input ends of the first feature splicing and fusion unit, the first H dimension feature selection unit, and the first W dimension feature selection unit. The output end of the second depth convolution unit is respectively connected to the input ends of the first feature splicing and fusion unit, the second H dimension feature selection unit, and the second W dimension feature selection unit;
[0015] The first H dimension feature selection unit and the second HThe output ends of the dimensional feature selection units are all connected to the input ends of the second feature splicing and fusion unit. The first W dimensional feature selection unit and the second W dimensional feature selection unit's output ends are all connected to the input ends of the third feature splicing and fusion unit;
[0016] The output end of the second feature splicing and fusion unit is connected to the input end of the third standard convolution unit. The output end of the first feature splicing and fusion unit is connected to the input end of the fourth standard convolution unit. The output end of the third feature splicing and fusion unit is connected to the input end of the fifth standard convolution unit;
[0017] The output end of the third standard convolution unit is connected to the input end of the first feature shape reshaping unit. The output end of the fifth standard convolution unit is connected to the input end of the second feature shape reshaping unit;
[0018] The output ends of the first feature shape reshaping unit and the second feature shape reshaping unit are both connected to the input end of the first global feature calculation unit. The output end of the first global feature calculation unit is connected to the first activation function. The output ends of the first activation function and the fourth standard convolution unit are both connected to the input end of the second feature calculation unit. The output end of the second feature calculation unit is connected to the input end of the sixth standard convolution unit. The output end of the sixth standard convolution unit serves as the output end of the CFGC module.
[0019] Furthermore, the CSCP module includes a first feature image division channel, a second feature image division channel, a fourth feature splicing and fusion unit, a fifth feature splicing and fusion unit, a seventh standard convolution unit, an eighth standard convolution unit, a third feature product calculation unit, a fourth feature product calculation unit, a sixth feature splicing and fusion unit, a ninth standard convolution unit, a first global average pooling unit, and a second global average pooling unit;
[0020] The input end of the first feature image division channel and the input end of the first global average pooling unit jointly serve as the first input end of the CSCP. The input end of the second feature image division channel and the input end of the second global average pooling unit jointly serve as the second input end of the CSCP. The output end of the first feature image division channel is respectively connected to the input ends of the fourth feature splicing and fusion unit and the fifth feature splicing and fusion unit. The output end of the second feature image division channel is respectively connected to the input ends of the fourth feature splicing and fusion unit and the fifth feature splicing and fusion unit. The output end of the fourth feature splicing and fusion unit is connected to the input end of the seventh standard convolution unit. The output end of the fifth feature splicing and fusion unit is connected to the input end of the eighth standard convolution unit;
[0021] The output end of the seventh standard convolution unit and the output end of the first global average pooling unit are both connected to the input end of the third feature product calculation unit. The output end of the eighth standard convolution unit and the output end of the second global average pooling unit are both connected to the input end of the fourth feature product calculation unit;
[0022] The output end of the third feature product calculation unit and the output end of the fourth feature product calculation unit are both connected to the input end of the sixth feature splicing and fusion unit. The output end of the sixth feature splicing and fusion unit is connected to the input end of the ninth standard convolution unit. The output end of the ninth standard convolution unit serves as the output end of the CSCP module.
[0023] Furthermore, the first depth convolution unit and the second depth convolution unit are both used to receive the first feature image input to the CFGC module, and respectively perform depth convolution processing on the first feature image at different scales to obtain the second feature image and the third feature image;
[0024] The first feature splicing and fusion unit is used to splice the second feature image and the third feature image, and fuse the spliced features through the fourth standard convolution unit to obtain the fourth feature image;
[0025] First H dimensional feature selection unit is used to select and splice the maximum feature and the average feature of the second feature image in the H dimension to obtain the first H dimensional feature; The first W dimensional feature selection unit is used to select and splice the maximum feature and the average feature of the second feature image in the W dimension to obtain the first W dimensional feature; The second H dimensional feature selection unit is used to select and splice the maximum feature and the average feature of the third feature image in the H dimension to obtain the second H dimensional feature, and the second W dimensional feature selection unit is used to select and splice the maximum feature and the average feature of the third feature image in the W dimension to obtain the second W dimensional feature;
[0026] The second feature splicing and fusion unit is used to splice the first H dimensional feature and the second H dimensional feature to obtain the fifth feature image; The third feature splicing and fusion unit is used to splice the first W dimensional feature and the second W dimensional feature to obtain the sixth feature image;
[0027] The third standard convolutional unit is used to perform standard convolution processing on the fifth feature image, and the first feature shape reshaping unit is used to reshape the feature shape of the fifth feature image after convolution processing by the third standard convolutional unit to obtain the first reshaped feature image; the fifth standard convolutional unit is used to perform standard convolution processing on the sixth feature image, and the second feature shape reshaping unit is used to reshape the feature shape of the sixth feature image after convolution processing by the fifth standard convolutional unit to obtain the second reshaped feature image;
[0028] The first global feature calculation unit is used to perform matrix calculation on the first reshaped feature image and the second reshaped feature image to obtain a global feature representation matrix; the first activation function is used to activate the global feature representation matrix to obtain a global response matrix;
[0029] The second feature calculation unit is used to perform Hadamard product calculation on the fourth feature image and the global response matrix to obtain a seventh feature image;
[0030] The seventh feature image is input into the sixth standard convolutional unit for convolution fusion to obtain the output feature image of the CFGC module.
[0031] Furthermore, the first feature image channel division and the first global average pooling unit are used to receive the eighth feature image input to the CSCP module, the second feature image channel division and the second global average pooling unit are used to receive the ninth feature image input to the CSCP module, the first feature image channel division is used to perform average division on the eighth feature image in the channel to obtain the tenth feature image and the eleventh feature image, the first global average pooling unit is used to perform global average pooling on the eighth feature image to obtain the first encoder global weight vector, the second feature image channel division is used to perform average division on the ninth feature image in the channel to obtain the twelfth feature image and the thirteenth feature image, and the second global average pooling unit is used to perform global average pooling on the ninth feature image to obtain the second decoder global weight vector;
[0032] The fourth feature splicing and fusion unit is used to splice the tenth feature image and the twelfth feature image, and perform convolution fusion on the spliced feature image through the seventh standard convolutional unit to obtain a fourteenth feature image;
[0033] The fifth feature splicing and fusion unit is used to splice the eleventh feature image and the thirteenth feature image, and perform convolution fusion on the spliced feature image through the eighth standard convolutional unit to obtain a fifteenth feature image;
[0034] The third feature multiplication calculation unit is used to perform multiplication calculation on the first encoder global weight vector and the fourteenth feature image to obtain the sixteenth feature image, and the fourth feature multiplication calculation unit is used to perform multiplication calculation on the second decoder global weight vector and the fifteenth feature image to obtain the seventeenth feature image;
[0035] The sixth feature splicing and fusion unit is used to splice the features of the fifteenth feature image and the sixteenth feature image, and use the ninth standard convolutional unit to perform feature fusion on the spliced feature image to obtain the output feature image of the CSCP module.
[0036] Furthermore, the convolution kernels of the first depth convolution unit, the third standard convolution unit, the fifth standard convolution unit, and the sixth standard convolution unit are all 3×3, the convolution kernel of the second depth convolution unit is 7×7, and the convolution kernel of the fourth standard convolution unit is 1×1.
[0037] Furthermore, the convolution kernels of the seventh standard convolution unit, the eighth standard convolution unit, and the ninth standard convolution unit are all 1×1.
[0038] Furthermore, the first activation function uses activation function.
[0039] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0040] 1. In the present invention, by setting CFGC modules with global perception collaborative fusion in each layer of the encoder and decoder, it is possible to not only have the feature expression ability at different scales, but also strengthen the feature representation based on the global features at different scales; by setting CSCP modules and CFGC modules in the decoder, and making the CSCP modules of each layer of the encoder correspond to each layer of the decoder in a one-to-one skip connection, it is possible to use the CFGC modules and CSCP modules to perform feature recovery layer by layer, and achieve effective fusion during the feature splicing process of each layer of the encoder and decoder; at the same time, through the CSCP module that gradually fuses cross-channel, it is possible to directly fuse the features in the encoder and decoder stages, optimize the semantic gap between the encoder and decoder without changing the skip connection structure, and will not cause additional semantic ambiguity; at the same time, using different divided channels of the CSCP module to respectively perform feature division on the feature image output from the encoder and the feature image output from the transposed convolution module, and gradually cross-splicing and fusing the divided deep features and shallow features, and performing global average pooling processing on different feature images through different global average pooling units to obtain different feature weight vectors, and combining the feature weight vectors with the spliced and fused feature images, it is possible to simultaneously focus on the important semantic information in the encoder stage and the decoder channels.
[0041] 2. In the present invention, by using two different deep convolutional units in the CFGC module with global perception collaborative fusion, local features can be captured from different scales, and at the same time, these two scale features are selected, thus making up for the limitation that convolution pays more attention to local region features. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings. In the drawings:
[0043] Figure 1 is a schematic structural diagram of the X-UNet model in an image segmentation method based on a global perception U-Net model in this embodiment;
[0044] Figure 2 is a schematic structural diagram of the CFGC module in an image segmentation method based on a global perception U-Net model in this embodiment;
[0045] Figure 3 is a schematic diagram of the operation process of the CSCP module in an image segmentation method based on a global perception U-Net model in this embodiment;
[0046] Figure 4 is a schematic structural diagram of the CSCP module in an image segmentation method based on a global perception U-Net model in this embodiment;
[0047] Figure 5 is a schematic diagram of the operation process of the CSCP module in an image segmentation method based on a global perception U-Net model in this embodiment;
[0048] Figure 6 is a schematic diagram of model optimization in an image segmentation method based on a global perception U-Net model in this embodiment;
[0049] Figure 7 is a schematic diagram of the comparison result in an image segmentation method based on a global perception U-Net model in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0051] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0052] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0053] Embodiment 1
[0054] See Figure 1 As shown, in this embodiment, an image segmentation method based on a global perception U-Net model is provided. The method includes:
[0055] Obtain the medical image to be segmented, input the medical image to be segmented into the global perception U-Net model for image segmentation processing, and obtain the segmented image segmentation result;
[0056] Among them, the global perception U-Net model (X-UNet model) includes an encoder, a decoder, a CFGC module, and a CSCP module. Both the encoder and the decoder adopt a five-layer structure connected in sequence. The first layer of the encoder adopts a first convolutional module and a max pooling module connected in sequence. The remaining four layers of the encoder both adopt a CFGC module and a max pooling module connected in sequence. The last layer of the decoder adopts a second convolutional module. The remaining four layers of the decoder all adopt a transposed convolutional module, a CSCP module, and a CFGC module connected in sequence. The first four layers of the encoder are respectively connected to the first four layers of the decoder by skip connections. Specifically: the first convolutional module of the first layer of the encoder is connected to the CSCP module of the fourth layer of the decoder by a skip connection; the CFGC module of the second layer of the encoder is connected to the CSCP module of the third layer of the decoder by a skip connection; the CFGC module of the third layer of the encoder is connected to the CSCP module of the second layer of the decoder by a skip connection; the CFGC module of the fourth layer of the encoder is connected to the CSCP module of the first layer of the decoder by a skip connection.
[0057] It should be noted that in this embodiment, the global perception U-Net model is similar to the existing U-Net model architecture and is obtained by improving the existing U-Net model. The global perception U-Net model is a U-shaped encoder-decoder architecture. At the same time, it should be noted that the number of filters in each layer of the encoder and decoder in this embodiment are 32, 64, 128, 256, and 512 respectively from the first layer to the fifth layer. And to improve the stability during training and introduce non-linearity, a BN unit and a ReLU activation function are set after all convolutional units of the global perception U-Net model in this embodiment. At the same time, both the max pooling module and the transposed convolutional module adopt a max pooling layer with a convolution kernel of 2×2 and a transposed convolutional layer with a convolution kernel of 2×2 in the conventional U-Net model. This technical solution is a conventional technical means in this field and will not be elaborated here.
[0058] At the same time, it should be noted that since the channel dimension of the first layer of the encoder is relatively low and the performance of depth convolution will decline in a relatively low channel dimension, therefore, the first layer of the encoder in this embodiment adopts a first convolutional module with a convolution kernel of 3×3 to extract primary semantic information; and then a max pooling module with a kernel size of 2×2 is used to complete the first downsampling operation.
[0059] The working process of the global perception U-Net model is as follows: First, a medical image to be segmented with a size of 256×256×3 (H×W×C) is input into the first convolutional module of the first layer of the encoder, and a conventional convolutional operation with a convolutional kernel of 3×3, a stride of 1, and a padding of 3 is performed to output an image with a size of 256×256×32; then the output image is respectively input into the max-pooling module of the first layer of the encoder and the CSCP module of the fourth layer of the decoder. The max-pooling module of the first layer of the encoder performs a max-pooling operation with a kernel of 2×2, a stride of 2, and a padding of 1 on it, and outputs an image with a size of 128×128×32 to the CFGC module of the second layer of the encoder;
[0060] After the CFGC module of the second layer of the encoder processes the received image, it outputs an image with a size of 128×128×64; then the output image is respectively input into the max-pooling module of the second layer of the encoder and the CSCP module of the third layer of the decoder. The max-pooling module of the second layer of the encoder performs a max-pooling operation with a kernel of 2×2, a stride of 2, and a padding of 1 on it, and outputs an image with a size of 64×64×64 to the CFGC module of the third layer of the encoder;
[0061] After the CFGC module of the third layer of the encoder processes the received image, it outputs an image with a size of 64×64×128; then the output image is respectively input into the max-pooling module of the third layer of the encoder and the CSCP module of the second layer of the decoder. The max-pooling module of the third layer of the encoder performs a max-pooling operation with a kernel of 2×2, a stride of 2, and a padding of 1 on it, and outputs an image with a size of 32×32×128 to the CFGC module of the fourth layer of the encoder;
[0062] After the CFGC module of the fourth layer of the encoder processes the received image, it outputs an image with a size of 32×32×256; then the output image is respectively input into the max-pooling module of the fourth layer of the encoder and the CSCP module of the first layer of the decoder. The max-pooling module of the fourth layer of the encoder performs a max-pooling operation with a kernel of 2×2, a stride of 2, and a padding of 1 on it, and outputs an image with a size of 16×16×256 to the CFGC module of the fifth layer of the encoder;
[0063] After the CFGC module of the fifth layer of the encoder processes the received image, it outputs an image with a size of 16×16×512, and then the output image is input into the transposed convolutional module of the first layer of the decoder;
[0064] The transposed convolution module in the first layer of the decoder performs deconvolution on the received image with a convolution kernel of 2×2 and a stride of 2, and outputs an image with a size of 32×32×256. Then, the output image is input into the CSCP module in the first layer of the decoder. After processing the image input by the transposed convolution module and the image output by the fourth layer of the encoder, the CSCP module in the first layer of the decoder outputs an image with a size of 32×32×512. Then, the output image is input into the CFGC module in the first layer of the decoder. After processing the input image, the CFGC module in the first layer of the decoder outputs an image with a size of 32×32×256. Then, the output image is input into the transposed convolution module in the second layer of the decoder.
[0065] The transposed convolution module in the second layer of the decoder performs deconvolution on the received image with a convolution kernel of 2×2 and a stride of 2, and outputs an image with a size of 64×64×128. Then, the output image is input into the CSCP module in the second layer of the decoder. After processing the image input by the transposed convolution module and the image output by the third layer of the encoder, the CSCP module in the second layer of the decoder outputs an image with a size of 64×64×256. Then, the output image is input into the CFGC module in the second layer of the decoder. After processing the input image, the CFGC module in the second layer of the decoder outputs an image with a size of 64×64×128. Then, the output image is input into the transposed convolution module in the third layer of the decoder.
[0066] The transposed convolution module in the third layer of the decoder performs deconvolution on the received image with a convolution kernel of 2×2 and a stride of 2, and outputs an image with a size of 128×128×64. Then, the output image is input into the CSCP module in the third layer of the decoder. After processing the image input by the transposed convolution module and the image output by the second layer of the encoder, the CSCP module in the third layer of the decoder outputs an image with a size of 128×128×128. Then, the output image is input into the CFGC module in the third layer of the decoder. After processing the input image, the CFGC module in the third layer of the decoder outputs an image with a size of 128×128×64. Then, the output image is input into the transposed convolution module in the fourth layer of the decoder.
[0067] The transposed convolution module in the fourth layer of the decoder performs a deconvolution on the received image with a convolution kernel of 2×2 and a stride of 2, and outputs an image with a size of 256×256×32. Then, the output image is input into the CSCP module in the fourth layer of the decoder; after processing the image input by the transposed convolution module and the image output by the first layer of the encoder, the CSCP module in the fourth layer of the decoder outputs an image with a size of 256×256×64, and then the output image is input into the CFGC module in the fourth layer of the decoder; after processing the input image, the CFGC module in the fourth layer of the decoder outputs an image with a size of 256×256×32, and then the output image is input into the second convolution module in the fifth layer of the decoder;
[0068] The second convolution module in the fifth layer of the decoder performs a 1×1 convolution operation on the received image and outputs an image with a size of 256×256×1.
[0069] Specifically, in this embodiment, by setting the CFGC module with global perception collaborative fusion in each layer of the encoder and decoder, it is possible to not only have the feature expression ability at different scales, but also strengthen the feature representation based on the global features at different scales; by setting the CSCP module and the CFGC module in the decoder, and making the layers of the encoder correspond to the CSCP modules of each layer of the decoder in a one-to-one skip connection, it is possible to use the CFGC module and the CSCP module to gradually recover the features layer by layer, and achieve effective fusion in the process of feature stitching of each layer of the encoder and decoder; at the same time, through the CSCP module that gradually fuses cross-channel, it is possible to directly fuse the features in the encoder and decoder stages, optimize the semantic gap between the encoder and decoder without changing the skip connection structure, and will not cause additional semantic ambiguity; at the same time, using different divided channels of the CSCP module to respectively perform feature division on the feature images output from the encoder and the transposed convolution module, and gradually cross-stitch and fuse the divided deep features and shallow features, and performing global average pooling processing on different feature images through different global average pooling units to obtain different feature weight vectors, and combining the feature weight vectors with the stitched and fused feature images, it is possible to simultaneously focus on the important semantic information in the encoder stage and the decoder channels.
[0070] Furthermore, referring to Figure 2 as shown, the CFGC module includes a first depth convolution unit, a second depth convolution unit, a first feature stitching and fusion unit, a first H dimensional feature selection unit, a first W dimensional feature selection unit, a second H dimensional feature selection unit, a second WA dimension feature selection unit, a second feature splicing and fusion unit, a third feature splicing and fusion unit, a third standard convolutional unit, a fourth standard convolutional unit, a fifth standard convolutional unit, a first feature shape reshaping unit, a second feature shape reshaping unit, a first global feature calculation unit, a first activation function, a second feature calculation unit, and a sixth standard convolutional unit;
[0071] The input ends of the first depth convolutional unit and the second depth convolutional unit together serve as the input end of the CFGC module;
[0072] The output end of the first depth convolutional unit is respectively connected to the input ends of the first feature splicing and fusion unit, the first H dimension feature selection unit and the first W dimension feature selection unit, and the output end of the second depth convolutional unit is respectively connected to the input ends of the first feature splicing and fusion unit, the second H dimension feature selection unit and the second W dimension feature selection unit;
[0073] The first H dimension feature selection unit and the second H dimension feature selection unit's output ends are both connected to the input end of the second feature splicing and fusion unit, and the first W dimension feature selection unit and the second W dimension feature selection unit's output ends are both connected to the input end of the third feature splicing and fusion unit;
[0074] The output end of the second feature splicing and fusion unit is connected to the input end of the third standard convolutional unit, the output end of the first feature splicing and fusion unit is connected to the input end of the fourth standard convolutional unit, and the output end of the third feature splicing and fusion unit is connected to the input end of the fifth standard convolutional unit;
[0075] The output end of the third standard convolutional unit is connected to the input end of the first feature shape reshaping unit, and the output end of the fifth standard convolutional unit is connected to the input end of the second feature shape reshaping unit;
[0076] The output ends of the first feature shape reshaping unit and the second feature shape reshaping unit are both connected to the input end of the first global feature calculation unit, the output end of the first global feature calculation unit is connected to the first activation function, the output ends of the first activation function and the fourth standard convolutional unit are both connected to the input end of the second feature calculation unit, the output end of the second feature calculation unit is connected to the input end of the sixth standard convolutional unit, and the output end of the sixth standard convolutional unit serves as the output end of the CFGC module.
[0077] It should be noted that in this embodiment, the convolution kernels of the first depth convolution unit, the third standard convolution unit, the fifth standard convolution unit, and the sixth standard convolution unit are all 3×3, the convolution kernel of the second depth convolution unit is 7×7, and the convolution kernel of the fourth standard convolution unit is 1×1; at the same time, the first depth convolution unit and the second depth convolution unit are both depth convolutions, and the third standard convolution unit, the fourth standard convolution unit, the fifth standard convolution unit, and the sixth standard convolution unit are all standard convolutions. It should be noted that both standard convolution and depth convolution are well-known common knowledge in the art and will not be elaborated here; the first activation function uses activation function.
[0078] Specifically, in this embodiment, the CFGC module not only has the ability to extract multi-scale features, but also perceives global feature information in a lightweight manner, successfully achieving effective attention to both local and global information; and through H dimensional feature selection unit and W dimensional feature selection unit, global feature enhancement is effectively achieved in the H and W dimensions.
[0079] Furthermore, as shown in Figure 4 , the CSCP module includes a first feature image division channel, a second feature image division channel, a fourth feature splicing and fusion unit, a fifth feature splicing and fusion unit, a seventh standard convolution unit, an eighth standard convolution unit, a third feature product calculation unit, a fourth feature product calculation unit, a sixth feature splicing and fusion unit, a ninth standard convolution unit, a first global average pooling unit, and a second global average pooling unit;
[0080] The input end of the first feature image division channel and the input end of the first global average pooling unit together serve as the first input end of the CSCP, the input end of the second feature image division channel and the input end of the second global average pooling unit together serve as the second input end of the CSCP, the output end of the first feature image division channel is respectively connected to the input ends of the fourth feature splicing and fusion unit and the fifth feature splicing and fusion unit, the output end of the second feature image division channel is respectively connected to the input ends of the fourth feature splicing and fusion unit and the fifth feature splicing and fusion unit, the output end of the fourth feature splicing and fusion unit is connected to the input end of the seventh standard convolution unit, and the output end of the fifth feature splicing and fusion unit is connected to the input end of the eighth standard convolution unit;
[0081] The output end of the seventh standard convolution unit and the output end of the first global average pooling unit are both connected to the input end of the third feature product calculation unit, and the output end of the eighth standard convolution unit and the output end of the second global average pooling unit are both connected to the input end of the fourth feature product calculation unit;
[0082] The output ends of the third feature product calculation unit and the fourth feature product calculation unit are both connected to the input end of the sixth feature splicing and fusion unit. The output end of the sixth feature splicing and fusion unit is connected to the input end of the ninth standard convolution unit. The output end of the ninth standard convolution unit serves as the output end of the CSCP module.
[0083] It should be noted that, in this embodiment, the convolution kernels of the seventh standard convolution unit, the eighth standard convolution unit, and the ninth standard convolution unit are all 1×1; at the same time, the seventh standard convolution unit, the eighth standard convolution unit, and the ninth standard convolution unit are all standard convolutions.
[0084] At the same time, it should be noted that the excellent performance of U-Net largely owes to the design of the skip connections between the encoder and the decoder, which enables the network to preserve the semantic information at different stages lost during the pooling process, while allowing the shallow features in the encoder stage and the deep semantic information in the decoder stage to perform semantic interaction, thus better restoring the details and boundary features of the target. However, most of the recent existing technologies seem to overly emphasize optimizing the features in the encoder stage by improving the structure of the skip connections. They often ignore the subsequent fusion process after splicing with the decoder stage. U-Net and most of its variants still use simple splicing operations to fuse the features of the encoder and decoder stages. As far as we know, there still exists a semantic gap between the features in the encoder stage optimized by various methods and the features in the decoder stage. Simple splicing operations may not meet the requirements of high segmentation accuracy. Therefore, we provide the CSCP module to directly fuse the features of the encoder and decoder stages, optimize the semantic gap between the encoder and decoder from another perspective, and reduce semantic ambiguity.
[0085] Furthermore, referring to Figure 3 as shown, both the first depth convolution unit and the second depth convolution unit are used to receive the first feature image input to the CFGC module, and respectively perform depth convolution processing on the first feature image at different scales to obtain the second feature image and the third feature image. Specifically: ; , where , respectively represent the second feature image and the third feature image, , ; represents a depth convolution with a convolution kernel size of ; represents a depth convolution with a convolution kernel size of ; represents the input first feature image, .
[0086] The first feature splicing and fusion unit is used to splice the second feature image and the third feature image, and fuse the features after splicing through the fourth standard convolution unit to obtain the fourth feature image. Specifically: , where represents the fourth feature image, ; represents a standard convolution with a convolution kernel size of 1×1; represents splicing feature maps along the channel dimension.
[0087] The first H dimensional feature selection unit is used to select and splice the maximum feature and the average feature of the second feature image in the H dimension to obtain the first H dimensional feature; The first W dimensional feature selection unit is used to select and splice the maximum feature and the average feature of the second feature image in the W dimension to obtain the first W dimensional feature; The second H dimensional feature selection unit is used to select and splice the maximum feature and the average feature of the third feature image in the H dimension to obtain the second H dimensional feature, and the second W dimensional feature selection unit is used to select and splice the maximum feature and the average feature of the third feature image in the W dimension to obtain the second W dimensional feature. Specifically: ; ; ; , where , respectively represent the output first H dimensional (height dimension) feature and the second H dimensional (height dimension) feature, , ; respectively represent the output first W dimensional (width dimension) feature and the second W dimensional (height dimension) feature, , ; represents the average selection of layer-by-layer features in the dimension; represents the maximum selection of layer-by-layer features in the dimension, represents W or H dimension.
[0088] The second feature splicing and fusion unit is used to splice the firstH The dimensional feature and the second H dimensional feature are used to obtain the fifth feature image; the third feature splicing and fusion unit is used to splice the first W dimensional feature and the second W dimensional feature to obtain the sixth feature image, specifically: ; , where represents the fifth feature image, ; represents the sixth feature image, ;
[0089] The third standard convolution unit is used to perform standard convolution processing on the fifth feature image, and the first feature shape reshaping unit is used to reshape the feature shape of the fifth feature image after the convolution processing of the third standard convolution unit to obtain the first reshaped feature image; the fifth standard convolution unit is used to perform standard convolution processing on the sixth feature image, and the second feature shape reshaping unit is used to reshape the feature shape of the sixth feature image after the convolution processing of the fifth standard convolution unit to obtain the second reshaped feature image, specifically: ; , where represents feature shape reshaping; represents a standard convolution with a convolution kernel size of 3×3; represents the first reshaped feature image; represents the second reshaped feature image;
[0090] The first global feature calculation unit is used to perform matrix calculation on the first reshaped feature image and the second reshaped feature image to obtain a global feature representation matrix, specifically: , where represents the influence of the th position on the th position, represents the total amount of pixels, ; respectively represent the pixel points at the position of the first reshaped feature image and the pixel points at the position of the second reshaped feature image;
[0091] The first activation function is used to activate the global feature representation matrix to obtain a global response matrix, specifically: , where represents the global response matrix, ; represents the global response matrix.
[0092] The second feature calculation unit is used to perform Hadamard product calculation on the fourth feature image and the global response matrix to obtain the seventh feature image, specifically: , where represents the fourth feature image; represents the global response matrix; represents the seventh feature image, .
[0093] The seventh feature image is input into the sixth standard convolutional unit for convolutional fusion to obtain the output feature image of the CFGC module, specifically: , where represents the output feature image of the CFGC module, .
[0094] Furthermore, as shown in Figure 5 , the first feature image channel division and the first global average pooling unit are used to receive the eighth feature image input to the CSCP module , ; the second feature image channel division and the second global average pooling unit are used to receive the ninth feature image input to the CSCP module , ; the first feature image channel division is used to perform an average division of the eighth feature image on the channels to obtain the tenth feature image and the eleventh feature image , , ; the first global average pooling unit is used to perform global average pooling on the eighth feature image to obtain the first encoder global weight vector, specifically: , where represents the first encoder global weight vector; represents global average pooling;
[0095] The second feature image channel division is used to perform an average division of the ninth feature image on the channels to obtain the twelfth feature image and the thirteenth feature image , , ; the second global average pooling unit is used to perform global average pooling on the ninth feature image to obtain the second decoder global weight vector, specifically: , where represents the first encoder global weight vector;
[0096] The fourth feature splicing and fusion unit is used to splice the tenth feature image and the twelfth feature image, and perform convolution fusion on the spliced feature image through the seventh standard convolution unit to obtain the fourteenth feature image. The fifth feature splicing and fusion unit is used to splice the eleventh feature image and the thirteenth feature image, and perform convolution fusion on the spliced feature image through the eighth standard convolution unit to obtain the fifteenth feature image. Specifically: ; , where represents a standard convolution with a kernel size of 1×1; respectively represent the fourteenth feature image and the fifteenth feature image.
[0097] The third feature product calculation unit is used to calculate the product of the first encoder global weight vector and the fourteenth feature image to obtain the sixteenth feature image. The fourth feature product calculation unit is used to calculate the product of the second decoder global weight vector and the fifteenth feature image to obtain the seventeenth feature image. Specifically: ; , where represents the sixteenth feature image; represents the seventeenth feature image;
[0098] The sixth feature splicing and fusion unit is used to splice the fifteenth feature image and the sixteenth feature image, and perform feature fusion on the spliced feature image by using the ninth standard convolution unit to obtain the output feature image of the CSCP module. Specifically: , where represents the output feature image of the CSCP module.
[0099] Specifically, as shown in Figure 6 , in this embodiment, the model parameters of the X-UNet model are also updated by using the binary cross-entropy loss function in combination with the Adam optimizer. This technical solution is a conventional technical means and will not be elaborated here.
[0100] It should be noted that in this embodiment, in order to verify the effect of the X-UNet model provided by this method, we completed all comparative experiments and ablation studies. In addition, we selected a random partitioning strategy for all data sets in this embodiment, 80% of which were used for training and the remaining 20% for testing. At the same time, in order to ensure the rigor and fairness of the comparative experiment and ensure the repeatability of data segmentation, 3 random number seeds were used to divide the data set. During the training process, we used cosine annealing to decay the learning rate. In the CVC-ClinicDB, Kvasir-SEG and PH2 data sets, the initial learning rate value was 0.0001, and in BUSI it was 0.0002, and the minimum value of the cosine annealing decay learning rate in the four data sets was 0.00001. In the CVC-ClinicDB, Kvasir-SEG and BUSI data sets, the initial image size was set to 256×256, and in the PH2 data set, the image size was set to 320×320, so as to explore the segmentation performance of the network at different image resolutions. Batch size and epoch are uniformly set to 4 and 100 respectively, and BCEDiceLoss loss function and Adam optimizer are used in the training process. Most importantly, all algorithm models used for comparison in this embodiment (including this technical solution) have not undergone any pre-training.
[0101] IoU (Intersection over Union), Dice coefficient (DICE), and ACC (Accuracy) are classic evaluation indicators commonly used in medical image segmentation. IoU is the ratio of the intersection over union between the predicted result and the true result, and Dice is always used to capture the overlap between two samples. The formula is as follows: ; , where A and B represent the sets of two sets of sampling points. IoU and Dice are both used to measure the similarity between two samples. Obviously, the higher the value, the better the performance of the network. Acc is the proportion of the number of correct samples in the prediction results to the total number of predictions. The higher the value of Acc, the stronger the ability of the network to predict the correct true positive targets. The formula is as follows: ,in, , , and They represent true positive, false positive, true negative, and false negative, respectively.
[0102] Comparative experimental results show:
[0103] 1) The detailed results on the Kvasir-SEG dataset can be seen in Table 1. Compared with other advanced baseline methods, our X-UNet (78.65%) is leading in the three basic evaluation metrics of IoU, Dice, and Acc while having fewer computational costs and parameters. Compared with the most classical UNet (75.55%), the IoU has increased by 3.1%, proving the superiority of our model. At the same time, the TransAttUnet (78.45%) based on Transfomer achieved sub-optimal performance with the most computational costs, and its IoU is 0.2% lower than ours. And the IoU of the model FDFUNet (78.37%) designed based on CNN is 0.28% lower than that of X-UNet. In the case of comparable computational costs, the number of parameters of FDFUNet is more than three times that of ours, which is due to the advantage of depth convolution adopted by our focus on lightweight design. Obviously, X-UNet has achieved a good balance among computational costs, parameters, and performance.
[0104] Table 1: Experimental results of X-UNet and other baseline methods on the Kvasir-SEG dataset
[0105]
[0106] Among them, GFLOPS is the abbreviation of Giga Floating-point Operations Per Second, which refers to the index of one billion floating-point operations per second; Params represents the variable-length parameter index.
[0107] 2) All experimental settings are consistent with the Kvasir-SEG dataset. The detailed experimental results on the CVC-ClinicDB dataset can be referred to Table 2 below. Obviously, our X-UNet (87.90%) also achieved a leading level in the segmentation experiment of another polyp dataset. The IoU has increased by 3.02% compared with UNet (84.88%). It is worth noting that the performance differences between FDFUNet (87.81%) and TransAttUNet (87.66%) and our X-UNet in the CVC-ClinicDB dataset are relatively small, being 0.09% and 0.24% lower than ours respectively. However, their numbers of parameters significantly exceed ours. Especially, TransAttUnet has a significantly leading computational cost due to the adoption of the multi-head attention mechanism to establish long-range dependencies. Compared with them, X-UNet can save more computational resources while globally modeling long-range dependencies through its CFGC module, and has higher practical application value.
[0108] Table 2: Experimental Results of X-UNet and Other Baseline Methods on the CVC-ClinicDB Dataset
[0109]
[0110] 3) Different from the above two datasets, in the BUSI dataset, we explored the network performance at different learning rates, and the images in the BUSI dataset are ultrasound images with noise. More detailed experimental results can be referred to in Table 3 below. Even under the condition of ultrasound images, our X-UNet (69.08%) still maintained the best segmentation performance. Compared with UNet (63.25%), the IoU was 5.83% higher. CANet (68.52%), which performed mediocrely in the previous two polyp datasets, finally achieved sub-optimal performance in the BUSI dataset, with an IoU 0.56% lower than ours. Different from the previous polyp datasets, FDFUNet (68.12%) showed a significant performance decline in the environment of ultrasound images, with an IoU 0.96% lower than ours. This indicates that our network model can also achieve good performance in the ultrasound image dataset, demonstrating a certain degree of generality and adaptability of X-UNet.
[0111] Table 3: Experimental Results of X-UNet and Other Baseline Methods on the BUSI Dataset
[0112]
[0113] 4) Other experimental settings are consistent with the Kvasir-SEG and CVC-ClinicDB datasets, but the resolution of the input images in the PH2 dataset is set to 320×320. More detailed experimental data can be referred to in Table 4 below. It can be clearly seen that although the performance of our X-UNet (91.49%) did not significantly exceed other baseline networks, with an IoU 1.34% higher than the original UNet (90.15%), X-UNet still led in the three metrics of IoU, Dice, and Acc with fewer parameters and computational costs. TransAttUnet (90.27%), which achieved good performance in the previous datasets, showed a significant decline in the PH2 dataset, probably because the images used for training in the PH2 dataset are too scarce, making it difficult to learn the data features. This indicates that our X-UNet can also have good performance when the training images are scarce, and at the same time verifies the relatively strong generality of X-UNet on different types of datasets.
[0114] Table 4: Experimental Results of X-UNet and Other Baseline Methods on the PH2 Dataset
[0115]
[0116] In Figure 7 this section, we completed some visualization experimental results of X-UNet and other standard baseline networks. Obviously, compared with other advanced baseline networks, our X-UNet has a more accurate perception ability for the boundary information of the lesion area, resulting in the final prediction result of X-UNet being the most consistent with the ground truth. This is likely to rely on the feature enhancement based on global response in the CFGC module to establish long-range dependencies, leading to the sensitivity to the boundaries of the lesion area.
[0117] In summary, in this invention, a lightweight and efficient X-UNet for medical image segmentation is proposed based on the global perception collaborative fusion module (CFGC module) and the cross-channel step-by-step fusion module (CSCO module). To model long-range dependencies globally, we designed the CFGC module, which not only preserves feature information at different scales but also effectively realizes global feature enhancement in the H and W dimensions. To better reduce semantic ambiguity between the encoder and decoder, we proposed the CSCP module to directly act on the features in the encoder-decoder fusion stage to promote more sufficient fusion. Compared with traditional CNNs, X-UNet is not limited by the disadvantage that convolution overly focuses on local regions; compared with most Transformer-based networks, X-UNet models long-range dependencies with less computational cost and fewer parameters while protecting the 2D structure of the image, captures precious spatial information that multi-head attention cannot focus on through convolution, and pays attention to both local semantic interaction and global feature enhancement. Finally, our experiments prove that X-UNet achieves leading-level performance with less computational cost and fewer parameters compared with other advanced algorithms on multiple datasets of different scales.
[0118] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An image segmentation method based on a global perception U-Net model, characterized in that, The method includes: Obtaining a medical image to be segmented, inputting the medical image to be segmented into a global perception U-Net model for image segmentation processing, and obtaining a segmented image segmentation result; Among them, the global perception U-Net model includes an encoder, a decoder, a CFGC module, and a CSCP module. Both the encoder and the decoder adopt a five-layer structure connected in sequence. The first layer of the encoder adopts a first convolutional module and a max-pooling module connected in sequence. The remaining four layers of the encoder all adopt a CFGC module and a max-pooling module connected in sequence. The last layer of the decoder adopts a second convolutional module. The remaining four layers of the decoder all adopt a transposed convolutional module, a CSCP module, and a CFGC module connected in sequence. The first four layers of the encoder are respectively skip-connected to the first four layers of the decoder; the CFGC module includes H a dimensional feature selection unit and W a dimensional feature selection unit, which is used to select the feature information in the H and W dimensions at different scales, and based on H a dimensional feature selection unit and W a dimensional feature selection unit to enhance the global features in a matrix multiplication manner; The CSCP module includes a feature image division channel, a feature splicing and fusion unit, and a global average pooling unit, and is used for performing channel-wise feature average division on the feature images in the encoder and decoder stages, splicing the divided features in the encoder and decoder stages respectively, and completing step-by-step feature fusion based on the channel dependence relationship between the feature images in the encoder and decoder stages.
2. The image segmentation method based on the global perception U-Net model according to claim 1, characterized in that The first convolution module in the first layer of the encoder is skip-connected to the CSCP module in the fourth layer of the decoder, the CFGC module in the second layer of the encoder is skip-connected to the CSCP module in the third layer of the decoder, the CFGC module in the third layer of the encoder is skip-connected to the CSCP module in the second layer of the decoder, and the CFGC module in the fourth layer of the encoder is skip-connected to the CSCP module in the first layer of the decoder.
3. The image segmentation method based on the global perception U-Net model according to claim 2, wherein, The CFGC module in the last layer of the encoder is connected to the transposed convolution module in the first layer of the decoder.
4. A method for image segmentation based on a global perception U-Net model according to claim 1, characterized in that, The CFGC module includes a first depth convolution unit, a second depth convolution unit, a first feature splicing and fusion unit, a first H dimensional feature selection unit, a first W dimensional feature selection unit, a second H dimensional feature selection unit, a second W dimensional feature selection unit, a second feature splicing and fusion unit, a third feature splicing and fusion unit, a third standard convolution unit, a fourth standard convolution unit, a fifth standard convolution unit, a first feature shape reshaping unit, a second feature shape reshaping unit, a first global feature calculation unit, a first activation function, a second feature calculation unit, and a sixth standard convolution unit; The input ends of the first depth convolution unit and the second depth convolution unit jointly serve as the input end of the CFGC module; The output ends of the first depth convolution unit are respectively connected to the input ends of the first feature splicing and fusion unit, the first H dimensional feature selection unit and the first W dimensional feature selection unit. The output ends of the second depth convolution unit are respectively connected to the input ends of the first feature splicing and fusion unit, the second H dimensional feature selection unit and the second W dimensional feature selection unit; First H The output ends of the first H dimensional feature selection unit and the second W dimensional feature selection unit are both connected to the input end of the second feature splicing and fusion unit. The output ends of the first W dimensional feature selection unit and the second dimensional feature selection unit are both connected to the input end of the third feature splicing and fusion unit; The output end of the second feature splicing and fusion unit is connected to the input end of the third standard convolution unit, the output end of the first feature splicing and fusion unit is connected to the input end of the fourth standard convolution unit, and the output end of the third feature splicing and fusion unit is connected to the input end of the fifth standard convolution unit; The output end of the third standard convolution unit is connected to the input end of the first feature shape reshaping unit, and the output end of the fifth standard convolution unit is connected to the input end of the second feature shape reshaping unit; The output ends of the first feature shape reshaping unit and the second feature shape reshaping unit are both connected to the input end of the first global feature calculation unit, the output end of the first global feature calculation unit is connected to the first activation function, the output ends of the first activation function and the fourth standard convolution unit are both connected to the input end of the second feature calculation unit, the output end of the second feature calculation unit is connected to the input end of the sixth standard convolution unit, and the output end of the sixth standard convolution unit serves as the output end of the CFGC module.
5. A method for image segmentation based on a global perception U-Net model according to claim 1, characterized in that The CSCP module includes a first feature image division channel, a second feature image division channel, a fourth feature splicing and fusion unit, a fifth feature splicing and fusion unit, a seventh standard convolution unit, an eighth standard convolution unit, a third feature product calculation unit, a fourth feature product calculation unit, a sixth feature splicing and fusion unit, a ninth standard convolution unit, a first global average pooling unit, and a second global average pooling unit; The input end of the first feature image division channel and the input end of the first global average pooling unit together serve as the first input end of the CSCP. The input end of the second feature image division channel and the input end of the second global average pooling unit together serve as the second input end of the CSCP. The output end of the first feature image division channel is respectively connected to the input ends of the fourth feature splicing and fusion unit and the fifth feature splicing and fusion unit. The output end of the second feature image division channel is respectively connected to the input ends of the fourth feature splicing and fusion unit and the fifth feature splicing and fusion unit. The output end of the fourth feature splicing and fusion unit is connected to the input end of the seventh standard convolution unit. The output end of the fifth feature splicing and fusion unit is connected to the input end of the eighth standard convolution unit; The output end of the seventh standard convolution unit and the output end of the first global average pooling unit are both connected to the input end of the third feature product calculation unit. The output end of the eighth standard convolution unit and the output end of the second global average pooling unit are both connected to the input end of the fourth feature product calculation unit; The output end of the third feature product calculation unit and the output end of the fourth feature product calculation unit are both connected to the input end of the sixth feature splicing and fusion unit. The output end of the sixth feature splicing and fusion unit is connected to the input end of the ninth standard convolution unit. The output end of the ninth standard convolution unit serves as the output end of the CSCP module.
6. A method for image segmentation based on a global perception U-Net model according to claim 4, characterized in that, The first depth convolution unit and the second depth convolution unit are both used to receive the first feature image input to the CFGC module and respectively perform depth convolution processing on the first feature image at different scales to obtain the second feature image and the third feature image; The first feature splicing and fusion unit is used to splice the second feature image and the third feature image and fuse the spliced features through the fourth standard convolution unit to obtain the fourth feature image; First H The dimensional feature selection unit is used to select and splice the maximum feature and the average feature of the second feature image in the H dimension to obtain the first H dimensional feature; First W The dimensional feature selection unit is used to select and splice the maximum feature and the average feature of the second feature image in the W dimension to obtain the first W dimensional feature; Second H The second-dimensional feature selection unit is used to H select and splice the maximum feature and the average feature of the third feature image in the H dimension to obtain the second W The second-dimensional feature selection unit is used to W select and splice the maximum feature and the average feature of the third feature image in the W dimension to obtain the second dimensional feature; The second feature splicing and fusion unit is used to splice the first H dimensional feature and the second H dimensional feature to obtain a fifth feature image; The third feature splicing and fusion unit is used to splice the first W dimensional feature and the second W dimensional feature to obtain a sixth feature image; The third standard convolution unit is used to perform standard convolution processing on the fifth feature image. The first feature shape reshaping unit is used to reshape the feature shape of the fifth feature image after convolution processing by the third standard convolution unit to obtain the first reshaped feature image; The fifth standard convolution unit is used to perform standard convolution processing on the sixth feature image. The second feature shape reshaping unit is used to reshape the feature shape of the sixth feature image after convolution processing by the fifth standard convolution unit to obtain the second reshaped feature image; The first global feature calculation unit is used to perform matrix calculation on the first reshaped feature image and the second reshaped feature image to obtain the global feature representation matrix. The first activation function is used to activate the global feature representation matrix to obtain the global response matrix; The second feature calculation unit is used to perform Hadamard product calculation on the fourth feature image and the global response matrix to obtain the seventh feature image; The seventh feature image is input into the sixth standard convolution unit for convolution fusion to obtain the output feature image of the CFGC module.
7. A method for image segmentation based on a global perception U-Net model according to claim 5, characterized in that, The first feature image partitioning channel and the first global average pooling unit are used to receive the eighth feature image input to the CSCP module. The second feature image partitioning channel and the second global average pooling unit are used to receive the ninth feature image input to the CSCP module. The first feature image partitioning channel is used to perform an average partition on the eighth feature image in the channels, obtaining a tenth feature image and an eleventh feature image. The first global average pooling unit is used to perform global average pooling on the eighth feature image, obtaining a first encoder global weight vector. The second feature image partitioning channel is used to perform an average partition on the ninth feature image in the channels, obtaining a twelfth feature image and a thirteenth feature image. The second global average pooling unit is used to perform global average pooling on the ninth feature image, obtaining a second decoder global weight vector; The fourth feature splicing and fusion unit is used to perform feature splicing on the tenth feature image and the twelfth feature image, and perform convolution fusion on the spliced feature image through the seventh standard convolution unit to obtain a fourteenth feature image; The fifth feature splicing and fusion unit is used to perform feature splicing on the eleventh feature image and the thirteenth feature image, and perform convolution fusion on the spliced feature image through the eighth standard convolution unit to obtain a fifteenth feature image; The third feature product calculation unit is used to perform product calculation on the first encoder global weight vector and the fourteenth feature image to obtain a sixteenth feature image. The fourth feature product calculation unit is used to perform product calculation on the second decoder global weight vector and the fifteenth feature image to obtain a seventeenth feature image; The sixth feature splicing and fusion unit is used to perform feature splicing on the fifteenth feature image and the sixteenth feature image, and perform feature fusion on the spliced feature image using the ninth standard convolution unit to obtain the output feature image of the CSCP module.
8. A method for image segmentation based on a global perception U-Net model according to claim 4, characterized in that The convolution kernels of the first depth convolution unit, the third standard convolution unit, the fifth standard convolution unit, and the sixth standard convolution unit are all 3×3. The convolution kernel of the second depth convolution unit is 7×7. The convolution kernel of the fourth standard convolution unit is 1×1.
9. A method for image segmentation based on a global perception U-Net model according to claim 5, characterized in that, The convolution kernels of the seventh standard convolution unit, the eighth standard convolution unit, and the ninth standard convolution unit are all 1×1.
10. A method for image segmentation based on a global perception U-Net model according to claim 4, characterized in that, The first activation function uses activation function.
Citation Information
Patent Citations
Building extraction method in remote sensing image, electronic equipment and storage medium
CN115345866A
Liver and tumor image segmentation method
CN115345889A