A medical image segmentation system and method

By combining an encoder-decoder structure and a differential gating unit, the problem of interrupted layers in the segmentation of medical images with significant anisotropy is solved, achieving high-precision and lightweight 3D medical image segmentation, which is suitable for deployment in clinical devices.

CN122335692APending Publication Date: 2026-07-03HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2026-03-24
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing medical image segmentation methods struggle to accurately segment medical images with significant anisotropic characteristics, especially in the Z-axis direction where tomography is prone to occur.

Method used

An encoder-decoder structure is adopted, which combines differential gating units and sub-pixel aggregation units. Inter-layer feature fusion is performed through differential perception layer and neighborhood fusion layer. Upsampling is performed using frequency domain feature extraction and coordinate attention mechanism to construct a deep network containing large kernel inverse residual units, so as to achieve continuous segmentation of anatomical structure between slices.

Benefits of technology

It improves the segmentation accuracy of medical images with significant anisotropic features, dynamically suppresses interlayer misalignment noise, generates high-fidelity segmentation prediction maps, and significantly enhances the geometric continuity of the target region.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122335692A_ABST
    Figure CN122335692A_ABST
Patent Text Reader

Abstract

This invention discloses a medical image segmentation system and method, belonging to the field of medical image processing technology. The invention segments each slice in a three-dimensional medical image to be segmented, taking into account the anatomical continuity between slices during segmentation. For each slice I along the Z-axis, its preceding and following adjacent slices in the Z-axis direction are extracted, constructing a slice sequence containing three consecutive adjacent slices as input to the medical image segmentation system. Then, feature extraction is performed on slice I based on an encoder-decoder structure. A differential gating unit is introduced at the jump connection between the encoder unit and the corresponding decoder unit in the decoder to fuse inter-layer information of the slices into the feature extraction process of slice I, achieving accurate representation of slice features and avoiding "faults" in the Z-axis direction. This improves the accuracy of segmenting medical images with significant anisotropic features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image processing technology, and more specifically, relates to a medical image segmentation system and method. Background Technology

[0002] Medical image segmentation is a crucial step in assisting clinical diagnosis, treatment planning, and efficacy evaluation. With the development of medical imaging equipment, three-dimensional volumetric data, such as CT and MRI, has become a mainstream diagnostic tool for doctors. However, in practical clinical applications, existing segmentation methods face significant challenges.

[0003] Many medical image datasets exhibit significant anisotropy, meaning that inter-slice resolution is much lower than intra-slice resolution. While some existing image segmentation methods, such as traditional 2D networks like U-Net and SegNet, can process images slice by slice, they completely ignore the continuity of anatomical structures between slices, resulting in "discontinuities" in the segmentation results along the Z-axis. Therefore, existing image segmentation methods often struggle to accurately segment medical images with significant anisotropy. Summary of the Invention

[0004] In view of the above-mentioned defects or improvement needs of the existing technology, the present invention provides a medical image segmentation system and method to solve the technical problem that existing image segmentation methods often have difficulty in accurately segmenting medical images with significant anisotropic characteristics.

[0005] To achieve the above objectives, in a first aspect, the present invention provides a medical image segmentation system, comprising: The encoder includes M cascaded coding units and downsampling units located between two adjacent coding units, for receiving the first... slice sequence The input, its first Level coding units are used to extract the slice sequence. A sequence of visual feature maps at various scales, including: , and No. Visual feature maps at various scales , and ; ; ; The first segment of the three-dimensional medical image to be segmented along the Z-axis direction. One slice; ; The number of slices along the Z-axis of the 3D medical image to be segmented; slices With slices Same, slice With slices same; The fusion module includes: The first differential gating unit; the first Each differential gating unit is used for calculation. and and The total difference is used to obtain the total difference feature map; based on the channel attention mechanism and the total difference feature map, the total difference feature map is used to... and Perform fusion and overlay the fused feature maps onto... The above yields an inter-layer feature fusion map. ; ; The decoder comprises M cascaded decoding units and upsampling units located between adjacent decoding units; the first-stage decoding unit is used to extract visual feature maps. The global context features are used to obtain the first decoded feature map; the second... Level decoding units are used to decode feature maps Fusion map with interlayer features The fused feature map is decoded to obtain the first... Decoding feature maps; ; For the first The decoded feature map is the feature map after being upsampled by the corresponding upsampling unit; Segmentation head, used for segmentation based on the first Decoding feature map pairs Segmentation is performed to obtain The segmentation prediction map.

[0006] More preferably, the differential gating unit includes: Differential sensing layer, used for pixel-by-pixel calculation and and The differences between them are used to obtain a difference feature map. and ,right and The summation is performed pixel by pixel to obtain the total difference feature map; The neighborhood fusion layer includes a channel attention network, which, based on the total difference feature map, obtains the gate weight values ​​at each pixel location in the total difference feature map, performs normalization processing, and then obtains the normalized gate weight values ​​at each pixel location. and Each pixel value in the average feature map is multiplied by the normalized gating weight value at the corresponding pixel location to achieve... and The images are then merged and superimposed onto the image. The above yields an inter-layer feature fusion map. .

[0007] More preferably, the first-stage decoding unit includes: Frequency domain feature extraction unit, used to extract visual feature maps. The frequency domain features are obtained to create a frequency domain feature map; Spatial feature extraction unit, used to extract visual feature maps. The spatial characteristics are used to obtain a spatial feature map; The global fusion unit is used to multiply the frequency domain feature map and the spatial domain feature map pixel by pixel to obtain the dot product feature map; and to combine the dot product feature map with the visual feature map. The first decoded feature map is obtained by adding pixels one by one.

[0008] More preferably, the frequency domain feature extraction unit is used for visual feature maps. Perform a Fourier transform to obtain the feature map. ; feature map After element-wise multiplication with the learnable complex weights, an inverse Fourier transform is performed, and the real part of the result is used as the frequency domain feature map. The spatial feature extraction unit is used for visual feature maps The spatial feature map is obtained by sequentially performing depthwise separable convolution, nonlinear activation, and normalization operations.

[0009] More preferably, the upsampling unit is a sub-pixel aggregation unit, comprising: Point convolutional layers are used to process the input feature map. Perform a convolution operation to decode the input feature map. The number of channels was expanded from C to The extended feature map is obtained. ; These are the input feature maps. The number of channels, height, and width; This is the preset upsampling factor; Sub-pixel rearrangement layers are used to expand the feature map. By rearranging the channels into spatial representations, a transform feature map is obtained. ; Depthwise separable convolutional layers are used to transform feature maps. Perform smoothing; The coordinate attention calibration layer is used to calibrate the smoothed transformation feature map based on the coordinate attention mechanism to obtain the calibration feature map. Feature fusion layer, used to transform feature maps The decoded feature map is obtained by adding it pixel-by-pixel to the calibration feature map. The upsampled feature map.

[0010] More preferably, the above-mentioned encoding unit and the second to Mth level decoding units all include: a large-core inverse residual unit; Large kernel inverse residual units are used for input images After expanding the number of channels, the image is divided equally along the channel dimension to obtain the first image and the second image; the first image and the second image are then respectively processed by convolution kernels with a kernel size of [missing value]. The depth of the separable convolutional layer and the size of the convolutional kernel are The input image is processed by a depthwise separable convolutional layer to obtain a first feature map and a second feature map. The first and second feature maps are then concatenated along the channel dimension, and layer normalization is performed to obtain a normalized feature map. The normalized feature map is then scaled to match the number of channels of the input image. The number of channels is the same; the normalized feature maps after scale transformation are then subjected to adaptive scaling and DropPath operation, and compared with the input image. The features are added pixel by pixel, and the resulting feature map is used as the output of the corresponding encoding or decoding unit; where... .

[0011] More preferably, the above-mentioned medical image segmentation system further includes: a post-processing module for processing the slices. The segmentation prediction maps are stacked along the Z-axis to obtain the segmentation result of the three-dimensional medical image to be segmented.

[0012] Secondly, the present invention provides a medical image segmentation method, comprising: The first slice sequence Input into a medical image segmentation system to obtain The segmentation prediction map; where, The first segment of the three-dimensional medical image to be segmented along the Z-axis direction. One slice; ; The number of slices along the Z-axis of the 3D medical image to be segmented; slices With slices Same, slice With slices same; slice The segmentation prediction maps are stacked along the Z-axis to obtain the segmentation results of the three-dimensional medical image; The medical image segmentation system is the medical image segmentation system provided in the first aspect of this invention.

[0013] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the medical image segmentation method provided in the second aspect of the present invention.

[0014] Fourthly, the present invention also provides a computer-readable storage medium comprising a stored computer program, wherein the computer program, when executed by a processor, controls the device containing the storage medium to perform the medical image segmentation method provided in the second aspect of the present invention.

[0015] Fifthly, the invention also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the medical image segmentation method provided in the second aspect of the invention.

[0016] In summary, the above-described technical solutions conceived in this invention can achieve the following beneficial effects: 1. This invention provides a medical image segmentation system that segments each slice in a three-dimensional medical image to be segmented. During segmentation, the anatomical continuity between slices in the three-dimensional medical image is considered. For each slice I along the Z-axis in the three-dimensional medical image, its preceding and following adjacent slices in the Z-axis direction are extracted, constructing a slice sequence containing three consecutive adjacent slices as input to the medical image segmentation system. Then, feature extraction is performed on slice I based on an encoder-decoder structure. A differential gating unit is introduced at the jump connection between the encoding unit of the encoder and the corresponding decoding unit in the decoder to fuse the inter-layer information of the slices into the feature extraction process of slice I, achieving accurate expression of slice features and avoiding "faults" in the Z-axis direction. This improves the accuracy of segmenting medical images with significant anisotropic features.

[0017] 2. Furthermore, in the medical image segmentation system provided by the present invention, the differential gating unit controls the influence of adjacent slices on slice I by controlling the information difference between each slice I and its two adjacent adjacent slices, thereby dynamically suppressing inter-layer misalignment noise and further improving the accuracy of slice feature extraction, thereby further improving the accuracy of medical image segmentation with significant anisotropic features.

[0018] 3. Furthermore, in the medical image segmentation system provided by this invention, the first-level decoding unit is used to extract slices. Visual feature map obtained by encoder The global context features are used to obtain the first decoded feature map; the frequency domain feature extraction unit in the first-level decoding unit transforms the expensive global interaction in the spatial domain into efficient point-by-point multiplication in the frequency domain, and then restores it through inverse Fourier transform, which effectively preserves long-range global context information while greatly reducing the amount of computation.

[0019] 4. Furthermore, in the medical image segmentation system provided by the present invention, the upsampling unit is a sub-pixel aggregation unit; the sub-pixel aggregation unit uses sub-pixel convolution instead of transposed convolution and performs position calibration through a coordinate attention mechanism, thereby eliminating the checkerboard effect from a mechanism perspective, significantly enhancing the geometric continuity of the target region, and generating a high-fidelity segmentation prediction map.

[0020] 5. Furthermore, in the medical image segmentation system provided by the present invention, the encoding unit and the second to Mth level decoding units all include large kernel inverse residual units, which perform feature extraction on both small and large receptive fields, paying attention to both details and overall features, and thus obtaining more accurate feature representations. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the structure of a medical image segmentation system provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of the large-core inverted residual unit provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a spectrum mixer provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the differential gating unit provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of the sub-pixel aggregation unit provided in an embodiment of the present invention; Figure 6(a) is a bar chart comparing the performance of the medical image segmentation system proposed in this embodiment of the invention with existing mainstream methods in terms of parameter quantity; Figure 6(b) is a bar chart comparing the computational performance of the medical image segmentation system proposed in this embodiment of the invention with existing mainstream methods; Figure 6(c) is a bar chart comparing the performance of the medical image segmentation system proposed in the embodiment of the present invention with that of existing mainstream methods in terms of segmentation accuracy; Figure 7 This is a visual comparison of the segmentation performance of the medical image segmentation system proposed in this embodiment of the invention on the MSD brain tumor dataset. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0023] To achieve the above objectives, in a first aspect, the present invention provides a medical image segmentation system, comprising: The encoder includes M cascaded coding units and downsampling units located between two adjacent coding units, for receiving the first... slice sequence The input, its first Level coding units are used to extract the slice sequence. A sequence of visual feature maps at various scales, including: , and No. Visual feature maps at various scales , and ; ; ; The first one along the Z-axis in a three-dimensional medical image One slice; , and These are three adjacent slices along the Z-axis in the three-dimensional medical image to be segmented; ; The number of slices along the Z-axis of the 3D medical image to be segmented; slices With slices Same, slice With slices same; The fusion module includes: The first differential gating unit; the first Each differential gating unit is used for calculation. and and The total difference is used to obtain the total difference feature map; based on the channel attention mechanism and the total difference feature map, the total difference feature map is used to... and Perform fusion and overlay the fused feature maps onto... The above yields an inter-layer feature fusion map. ; ; The decoder comprises M cascaded decoding units and upsampling units located between adjacent decoding units; the first-stage decoding unit is used to extract visual feature maps. The global context features are used to obtain the first decoded feature map; the second... Level decoding units are used to decode feature maps Fusion map with interlayer features The fused feature map is decoded to obtain the first... Decoding feature maps; ; For the first The decoded feature map is the feature map after being upsampled by the corresponding upsampling unit; Segmentation head, used for segmentation based on the first Decoding feature map pairs Segmentation is performed to obtain The segmentation prediction map.

[0024] It should be noted that the above and Differences, and and The difference can be measured using L1 distance, L2 distance, Mahalanobis distance, etc., and no particular method is specified here. L1 distance is preferred.

[0025] Preferably, in one optional implementation, the differential gating unit includes: Differential sensing layer, used for pixel-by-pixel calculation and and The differences between them are used to obtain a difference feature map. and ,right and The summation is performed pixel by pixel to obtain the total difference feature map; The neighborhood fusion layer includes a channel attention network, which, based on the total difference feature map, obtains the gate weight values ​​at each pixel location in the total difference feature map, performs normalization processing, and then obtains the normalized gate weight values ​​at each pixel location. and Each pixel value in the average feature map is multiplied by the normalized gating weight value at the corresponding pixel location to achieve... and The images are then merged and superimposed onto the image. The above yields an inter-layer feature fusion map. .

[0026] It should be noted that the first-level decoding unit can employ self-attention-based models such as Transformer, GCViT, and Swin Transformer, or multi-scale models such as the Spatial Pyramid Pooling (ASPP) module and the Pyramid Pooling (PPM) module; no limitation is made here. Preferably, in one optional implementation, the first-level decoding unit is a spectrum mixer; the spectrum mixer includes: Frequency domain feature extraction unit, used to extract visual feature maps. The frequency domain features are obtained to create a frequency domain feature map; Spatial feature extraction unit, used to extract visual feature maps. The spatial characteristics are used to obtain a spatial feature map; The global fusion unit is used to multiply the frequency domain feature map and the spatial domain feature map pixel by pixel to obtain the dot product feature map; and to combine the dot product feature map with the visual feature map. The first decoded feature map is obtained by adding pixels one by one.

[0027] Preferably, the frequency domain feature extraction unit is used for visual feature maps. Perform a Fourier transform to obtain the feature map. ; feature map After element-wise multiplication with the learnable complex weights, an inverse Fourier transform is performed, and the real part of the result is used as the frequency domain feature map. The spatial feature extraction unit is used for visual feature maps The spatial feature map is obtained by sequentially performing depthwise separable convolution, nonlinear activation, and normalization operations, which is used to characterize the confidence of the location in the global features of the frequency domain.

[0028] The frequency domain feature extraction unit transforms the expensive global interaction in the spatial domain into efficient point-by-point multiplication in the frequency domain based on the convolution theorem, and then restores it through inverse Fourier transform (IFFT). This module effectively preserves long-range global context information while significantly reducing the amount of computation (graphics memory usage).

[0029] It should be noted that the above upsampling unit can be implemented using existing upsampling algorithms, such as nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, and other interpolation-based algorithms. It can also employ dynamic upsampling algorithms such as transposed convolution, unpooling, and content-aware reconstruction (CARAFE), without limitation. However, these algorithms generally use transposed convolution for upsampling, which can easily produce a checkerboard effect during feature recovery, leading to blurred and jagged boundaries of the target region (such as lesions or organs), affecting the geometric accuracy of segmentation. Therefore, preferably, in one optional implementation, the above upsampling unit is a sub-pixel aggregation unit, including: Point convolutional layers are used to process the input feature map. Perform a convolution operation to decode the input feature map. The number of channels was expanded from C to The extended feature map is obtained. ; These are the input feature maps. The number of channels, height, and width; The preset upsampling factor is an integer greater than or equal to 2 in this embodiment, preferably 2.

[0030] Sub-pixel rearrangement layers are used to expand the feature map. By rearranging the channels into spatial representations, a transform feature map is obtained. ; Depthwise separable convolutional layers are used to transform feature maps. Perform smoothing; The coordinate attention calibration layer is used to calibrate the smoothed transformation feature map based on the coordinate attention mechanism to obtain the calibration feature map. Feature fusion layer, used to transform feature maps The decoded feature map is obtained by adding it pixel-by-pixel to the calibration feature map. The upsampled feature map.

[0031] It should be noted that the aforementioned encoding units can be CNNs, residual modules, encoding units in Transformers, etc., and are not limited here. The aforementioned second to Mth level decoding units can be CNNs, decoding units in Transformers, NCC models, etc., and are not limited here.

[0032] Preferably, in an optional implementation, the above-mentioned encoding unit and the second to Mth level decoding units all include: a large kernel inverse residual unit; Large kernel inverse residual units are used for input images After expanding the number of channels, the image is divided equally along the channel dimension to obtain the first image and the second image; the first image and the second image are then respectively processed by convolution kernels with a kernel size of [missing value]. The depth of the separable convolutional layer and the size of the convolutional kernel are The input image is processed by a depthwise separable convolutional layer to obtain a first feature map and a second feature map. The first and second feature maps are then concatenated along the channel dimension, and layer normalization is performed to obtain a normalized feature map. The normalized feature map is then scaled to match the number of channels of the input image. The number of channels is the same; the normalized feature map after scale transformation is then subjected to adaptive scaling and DropPath operation (i.e., random depth operation) and compared with the input image. The features are added pixel by pixel, and the resulting feature map is used as the output of the corresponding encoding or decoding unit; where... Preferably, .

[0033] It should be noted that the large-kernel inverse residual unit extracts features by using both small and large receptive fields, which pays attention to both details and overall features, thus obtaining a more accurate feature representation.

[0034] It should be noted that the segmentation head mentioned above can be any type of segmentation head, such as a linear classification head based on 1×1 convolution, a multilayer perceptron (MLP) classification head, or a query-based Transformer decoding head, etc. No limitation is made here. Taking the segmentation as an example, the segmentation head will... Decoding the feature map maps to the final The segmentation prediction map, where each pixel value indicates whether the pixel's location belongs to the target region, thus achieving [the desired segmentation]. Segmentation of the target region in the process.

[0035] In one optional implementation, the above-described medical image segmentation system further includes: a post-processing module for processing the slices. The segmentation prediction maps are stacked along the Z-axis to obtain the segmentation results of the three-dimensional medical image.

[0036] It should be noted that the training method for the above-mentioned medical image segmentation system can employ conventional training methods, such as end-to-end training. In one optional implementation, the training method for the above-mentioned medical image segmentation system includes: Each slice sequence of each 3D medical image sample in the training set is input into the medical image segmentation system to obtain a segmentation prediction map for each slice. The medical image segmentation system is trained by minimizing the difference loss between the obtained segmentation prediction map and the corresponding ground truth label (the mask image of the target region in the slice).

[0037] It should be noted that the target area mentioned above is determined based on the specific type of the medical image sample. It can be a lesion area or an organ area, and there is no limitation here. For example, when the medical image sample is a brain tumor image sample, the target area mentioned above is the brain tumor area; when the medical image sample is a lung tumor image sample, the target area mentioned above is the lung tumor area.

[0038] To further illustrate the medical image segmentation system provided in the first aspect of the present invention, a detailed description is provided below with reference to a specific embodiment: This embodiment proposes an anisotropic medical image segmentation system (Freq-2.5D-UNet) based on frequency domain awareness and neighborhood-guided gating. This system combines a 2.5D slice input strategy, frequency domain global modeling, and a neighborhood-guided gating mechanism to achieve lightweight and high-precision 3D medical image segmentation.

[0039] like Figure 1 The diagram shown is a schematic representation of the medical image segmentation system provided in this embodiment, and its processing flow is as follows: a. Obtain the volumetric data of the 3D medical image to be segmented. For each slice along the Z-axis in the 3D medical image (i.e., a cross-sectional image perpendicular to the Z-axis), extract its preceding and following adjacent slices along the Z-axis, constructing a 2.5D slice sequence containing three adjacent slices as input to the medical image segmentation system. Specifically, for the... slice With this slice For the center slice, construct a structure containing the center slice. and adjacent slices , The 2.5D slice sequence is used as input to the medical image segmentation system, and the input size is adjusted to 224×224. In this embodiment, the number of slices along the Z-axis of the three-dimensional medical image to be segmented is... The value is 3.

[0040] b. Input the slice sequence into the encoder. The encoder consists of four cascaded coding units and downsampling units located between adjacent coding units; the coding units include: Large Kernel Inverse Residual Units (LK-IBB), which are used to extract multi-scale spatial features in the two-dimensional plane of the slice. Figure 2 As shown, the large kernel inverse residual unit (LK-IBB) employs 3×3 depthwise separable convolution and... Depthwise separable convolution extracts spatial features, and when combined with inverted residual structures, it significantly expands the effective receptive field in the two-dimensional plane while maintaining extremely low parameter counts compared to standard convolution, enabling it to better capture the shape features of large-scale organs.

[0041] The decoder includes four cascaded decoding units and an upsampling unit located between adjacent decoding units; the first-stage decoding unit is used to extract slices. Visual feature map obtained by encoder The global context features are used to obtain the first decoded feature map, which includes: a spectrum mixer (SFM) and a Fourier transform (FFT) to transform the visual feature map. Transforming to the frequency domain and performing global interactive processing only on the frequency domain components results in a time complexity of O(NlogN), effectively capturing long-range context, as detailed below. Figure 3 As shown.

[0042] c. At the jump connection between the encoding and decoding units, feature fusion is performed using Differential Gated Units (NGDG); such as Figure 4 As shown. The specific processing procedure is as follows: (1) Differential sensing: Calculate the central slice Visual feature map With adjacent slices , Visual feature map , The pixel-level L1 distance between them is obtained and Difference feature map between and and Difference feature map between ,right and The summation is performed pixel by pixel to obtain the total difference feature map; (2) Gating generation: The total differential feature map is sequentially input into the convolution kernel with a size of [missing value]. In the convolutional layers and channel attention network, the gate weights at each pixel location of the total difference feature map are obtained. After normalization, the normalized gate weights at each pixel location are obtained, forming a weight matrix G; the th... Line 1 The elements in the column represent the positions in the total difference feature map. Normalized gating weight values ​​at the location; In this embodiment, the channel attention network includes: cascaded dimensionality-reducing convolutional layers, global average pooling layers, fully connected layers, and dimensionality-increasing convolutional layers; the above normalization processing preferably adopts activation function normalization processing, such as using the Sigmoid activation function.

[0043] (3) Neighborhood integration: and Average feature map Each pixel value in the algorithm is multiplied by the normalized gating weight value at the corresponding pixel location to achieve... and The images are then merged and superimposed onto the image. The above yields an inter-layer feature fusion map. The specific formula is as follows: .

[0044] The above operations utilize the obtained normalized gating weights to dynamically filter and fuse inter-layer features, achieving [the desired effect]. The weighted calibration controls the influence of adjacent slices in the previous layer and adjacent slices in the next layer on the central slice by the information difference between adjacent slices and the central slice, thereby dynamically suppressing interlayer misalignment noise.

[0045] d. Input the feature image fused by the differential gating unit into the encoder, use the sub-pixel aggregation unit (SPA) for progressive upsampling, and work with the decoding unit to gradually restore details and generate slices. The segmentation prediction map. For example... Figure 5 As shown, the aforementioned sub-pixel aggregation unit (SPA) includes: Point convolutional layers are used to process the input feature map. Perform a convolution operation to decode the input feature map. The number of channels was expanded from C to The extended feature map is obtained. ; These are the input feature maps. The number of channels, height, and width; This is the preset upsampling factor; Sub-pixel rearrangement layers are used to expand the feature map. Perform a scaling transformation to obtain the transformed feature map. ; Depthwise separable convolutional layers are used to transform feature maps. Perform smoothing; The coordinate attention calibration layer is used to calibrate the smoothed transformation feature map based on the coordinate attention mechanism to obtain the calibration feature map. Feature fusion layer, used to transform feature maps The decoded feature map is obtained by adding it pixel-by-pixel to the calibration feature map. The upsampled feature map.

[0046] In summary, the subpixel aggregation unit (SPA) replaces transposed convolution with subpixel convolution and performs position calibration through a coordinate attention mechanism. This design fundamentally eliminates the checkerboard effect, significantly enhances the geometric continuity of the target region (such as organ regions or lesion regions), and can generate high-fidelity segmentation prediction maps.

[0047] e. Segment the data using the segmentation head based on the decoder output to obtain... The segmentation prediction map.

[0048] f. Stack and recombine the segmentation prediction maps of all slices along the Z-axis to obtain the complete 3D segmentation result.

[0049] In summary, the medical image segmentation system provided in this embodiment is a lightweight medical image segmentation method that can efficiently process anisotropic data and achieve global modeling with accurate boundaries at low cost. Specifically, this embodiment constructs a deep network including Large Kernel Inverse Residual Unit (LK-IBB), Spectral Mixer (SFM), Differential Gated Unit (NGDG), and Subpixel Aggregation Unit (SPA), and designs a 2.5D segmentation architecture that combines fine-grained perception within layers with robust fusion between layers. Specifically, the LK-IBB module uses large-size convolutional kernels to adaptively balance the extraction of local textures and large-scale anatomical structures; the SFM module uses Fast Fourier Transform at the bottleneck layer to transform spatial interactions into frequency domain computation, achieving global long-range dependency modeling with low computational cost; the NGDG unit generates consistency weights at skip connections by calculating the differential features of adjacent slices, dynamically calibrating inter-layer feature fusion and effectively suppressing inter-layer misalignment noise; the SPA unit employs subpixel convolution and coordinate attention mechanisms in the decoding stage to eliminate the checkerboard effect and enhance boundary continuity. This embodiment only requires inputting adjacent slice sequences to achieve high-precision, continuous segmentation on three-dimensional medical images that is superior to mainstream 3D networks, and is easy to deploy on clinical end-devices.

[0050] The performance of the medical image segmentation system obtained in this embodiment will be tested below.

[0051] 1) Experimental setup and dataset Validation was performed using the brain tumor (Task01) task from the Medical Segmentation Decathlon (MSD) dataset. The training and test sets were split in a 4:1 ratio, and the AdamW optimizer was used with an initial learning rate of 1×10⁻⁶. 3 The training consists of 300 rounds.

[0052] 2) Module validity verification (ablation experiment) To verify the contribution of each module to segmentation performance, tests were conducted under the same experimental environment, with modules added incrementally. The Dice coefficient was used to measure segmentation accuracy. Results showed that the Dice coefficient of the basic baseline model was 63.27%; after introducing the large kernel inverse residual unit (LK-IBB), the Dice coefficient of this embodiment increased to 68.41%; after introducing the spectrum mixer (SFM), the Dice coefficient of this embodiment increased to 73.69%; and after further introducing the sub-pixel aggregation unit (SPA) and the differential gating unit (NGDG), the Dice coefficient of this embodiment finally reached 81.23%, and the HD95 distance decreased to 8.13 mm. This indicates that each module proposed in this embodiment can effectively improve segmentation accuracy.

[0053] Comprehensive performance comparison test: The method of this embodiment is compared with mainstream methods such as SwinUNETR, nnU-Net, and CSA-Net. Figure 6(a) shows a bar chart comparing the performance of the medical image segmentation system proposed in this embodiment with existing mainstream methods in terms of parameter quantity; Figure 6(b) shows a bar chart comparing the performance of the medical image segmentation system proposed in this embodiment with existing mainstream methods in terms of computational quantity; Figure 6(c) shows a bar chart comparing the performance of the medical image segmentation system proposed in this embodiment with existing mainstream methods in terms of segmentation accuracy. As can be seen from the figures, the Dice coefficient (81.23%) of the medical image segmentation system proposed in this embodiment on the MSD brain tumor task is better than that of SwinUNETR (80.6%) and CSA-Net. More importantly, the medical image segmentation system proposed in this embodiment has only 1.42M parameters and only 6.40G of floating-point operations (FLOPs), which is significantly lighter than SwinUNETR (62.19M parameters), and the inference speed is improved by about 4 times.

[0054] Visual verification: Figure 7 The figure shows a visual comparison of the segmentation performance of the medical image segmentation system proposed in this embodiment of the invention on the MSD brain tumor dataset. As can be seen from the figure, the medical image segmentation system proposed in this embodiment of the invention can accurately segment the complete tumor region and maintain good geometric continuity at the tumor boundary, without obvious jagged artifacts or interlayer discontinuities, demonstrating the robustness of the medical image segmentation system proposed in this embodiment of the invention on anisotropic data.

[0055] Secondly, the present invention provides a medical image segmentation method, comprising: The first slice sequence Input into a medical image segmentation system to obtain The segmentation prediction map; where, The first segment of the three-dimensional medical image to be segmented along the Z-axis direction. One slice; ; The number of slices along the Z-axis of the 3D medical image to be segmented; slices With slices Same, slice With slices same; slice The segmentation prediction maps are stacked along the Z-axis to obtain the segmentation result of the three-dimensional medical image to be segmented; The medical image segmentation system is the medical image segmentation system provided in the first aspect of this invention.

[0056] The related technical solutions are the same as those provided in the first aspect of this invention for the medical image segmentation system, and are not limited here.

[0057] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the medical image segmentation method provided in the second aspect of the present invention.

[0058] The related technical solutions are the same as the medical image segmentation method provided in the second aspect of this invention, and are not limited here.

[0059] Fourthly, the present invention also provides a computer-readable storage medium comprising a stored computer program, wherein the computer program, when executed by a processor, controls the device containing the storage medium to perform the medical image segmentation method provided in the second aspect of the present invention.

[0060] The related technical solutions are the same as the medical image segmentation method provided in the second aspect of this invention, and are not limited here.

[0061] Fifthly, the invention also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the medical image segmentation method provided in the second aspect of the invention.

[0062] The related technical solutions are the same as the medical image segmentation method provided in the second aspect of this invention, and are not limited here.

[0063] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A medical image segmentation system, characterized in that, include: The encoder includes M cascaded coding units and downsampling units located between two adjacent coding units, for receiving the first... slice sequence The input, its first Level coding units are used to extract the slice sequence. A sequence of visual feature maps at various scales, including: , and No. Visual feature maps at various scales , and ; ; ; The first segment of the three-dimensional medical image to be segmented along the Z-axis direction. One slice; ; The number of slices along the Z-axis of the 3D medical image to be segmented; slices With slices Same, slice With slices same; The fusion module includes: The first differential gating unit; the first Each differential gating unit is used for calculation. and and The total difference is used to obtain the total difference feature map; based on the channel attention mechanism and the total difference feature map, the total difference feature map is used to... and Perform fusion and overlay the fused feature maps onto... The above yields an inter-layer feature fusion map. ; ; The decoder comprises M cascaded decoding units and upsampling units located between adjacent decoding units; the first-stage decoding unit is used to extract visual feature maps. The global context features are used to obtain the first decoded feature map; the second... Level decoding units are used to decode feature maps Fusion map with interlayer features The fused feature map is decoded to obtain the first... Decoding feature maps; ; For the first The decoded feature map is the feature map after being upsampled by the corresponding upsampling unit; Segmentation head, used for segmentation based on the first Decoding feature map pairs Segmentation is performed to obtain The segmentation prediction map.

2. The medical image segmentation system according to claim 1, characterized in that, The differential gating unit includes: Differential sensing layer, used for pixel-by-pixel calculation and and The differences between them are used to obtain a difference feature map. and ,right and The summation is performed pixel by pixel to obtain the total difference feature map; The neighborhood fusion layer includes a channel attention network, which, based on the total difference feature map, obtains the gate weight values ​​at each pixel location in the total difference feature map, performs normalization processing, and then obtains the normalized gate weight values ​​at each pixel location. and Each pixel value in the average feature map is multiplied by the normalized gating weight value at the corresponding pixel location to achieve... and The images are then merged and superimposed onto the image. The above yields an inter-layer feature fusion map. .

3. The medical image segmentation system according to claim 1, characterized in that, The first-level decoding unit includes: Frequency domain feature extraction unit, used to extract visual feature maps. The frequency domain features are obtained to create a frequency domain feature map; Spatial feature extraction unit, used to extract visual feature maps. The spatial characteristics are used to obtain a spatial feature map; The global fusion unit is used to multiply the frequency domain feature map and the spatial domain feature map pixel by pixel to obtain a dot product feature map; and to combine the dot product feature map with the visual feature map. The first decoded feature map is obtained by adding pixels one by one.

4. The medical image segmentation system according to claim 3, characterized in that, The frequency domain feature extraction unit is used to process visual feature maps. Perform a Fourier transform to obtain the feature map. ; the feature map After element-wise multiplication with the learnable complex weights, an inverse Fourier transform is performed, and the real part of the result is used as the frequency domain feature map. The spatial feature extraction unit is used to process visual feature maps. The spatial feature map is obtained by sequentially performing depthwise separable convolution, nonlinear activation, and normalization operations.

5. The medical image segmentation system according to any one of claims 1-4, characterized in that, The upsampling unit is a sub-pixel aggregation unit, comprising: Point convolutional layers are used to process the input feature map. Perform a convolution operation to decode the input feature map. The number of channels was expanded from C to The extended feature map is obtained. ; These are the input feature maps. The number of channels, height, and width; This is the preset upsampling factor; Sub-pixel rearrangement layers are used to expand the feature map. By rearranging the channels into spatial representations, a transform feature map is obtained. ; Depthwise separable convolutional layers are used to transform feature maps. Perform smoothing; The coordinate attention calibration layer is used to calibrate the smoothed transformation feature map based on the coordinate attention mechanism to obtain the calibration feature map. Feature fusion layer, used to transform feature maps The decoded feature map is obtained by adding it pixel-by-pixel to the calibration feature map. The upsampled feature map.

6. The medical image segmentation system according to any one of claims 1-4, characterized in that, The encoding unit and the second to Mth level decoding units all include: a large-core inverse residual unit; The large kernel inverse residual unit is used to process the input image. After expanding the number of channels, the image is divided equally along the channel dimension to obtain the first image and the second image; the first image and the second image are then respectively processed by convolution kernels with a kernel size of [missing value]. The depth of the separable convolutional layer and the size of the convolutional kernel are The input image is processed by a depthwise separable convolutional layer to obtain a first feature map and a second feature map. The first and second feature maps are then concatenated along the channel dimension, and layer normalization is performed to obtain a normalized feature map. The normalized feature map is then scaled to match the number of channels of the input image. The number of channels is the same; the normalized feature maps after scale transformation are then subjected to adaptive scaling and DropPath operation, and compared with the input image. The features are added pixel by pixel, and the resulting feature map is used as the output of the corresponding encoding or decoding unit; where... .

7. The medical image segmentation system according to any one of claims 1-4, characterized in that, Also includes: The post-processing module is used to slice... The segmentation prediction maps are stacked along the Z-axis to obtain the segmentation result of the three-dimensional medical image to be segmented.

8. A medical image segmentation method, characterized in that, include: The first slice sequence Input into a medical image segmentation system to obtain The segmentation prediction map; where, The first segment of the three-dimensional medical image to be segmented along the Z-axis direction. One slice; ; The number of slices along the Z-axis of the 3D medical image to be segmented; slices With slices Same, slice With slices same; slice The segmentation prediction maps are stacked along the Z-axis to obtain the segmentation results of the three-dimensional medical image; The medical image segmentation system is the medical image segmentation system according to any one of claims 1-7.

9. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the medical image segmentation method of claim 8.

10. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, implements the medical image segmentation method of claim 8.