Efficient image segmentation system based on improved UNet model

By improving the lightweight encoder, dynamic feature fusion and efficient post-processing technology of the UNet model, the lack of computing efficiency and accuracy of the existing UNet model is solved, and efficient image segmentation is realized, suitable for mobile and embedded devices.

CN120339619APending Publication Date: 2025-07-18NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510475306.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In image segmentation, the existing UNet models have problems such as large calculation volume, increased parameter scale and calculation time, difficult to meet real-time requirements, limited cross-scale feature fusion efficiency, artifacts introduced by decoder upsampling, extensive post-processing, single model compression technology and obvious accuracy losses, which limit their efficient deployment on mobile or embedded devices.

Method used

The improved encoder module, multi-scale feature fusion module, dynamic decoder module and post-processing output module are adopted, and combined with deep separation convolution, hollow convolution, dynamic weight coefficient, sub-pixel convolution, model compression and other technologies are combined to achieve lightweight, dynamic feature fusion and efficient post-processing.

Benefits of technology

It improves the computing efficiency and accuracy of image segmentation, supports real-time processing of high-resolution images, adapts to the needs of fine segmentation in complex scenarios, avoids artifact problems, optimizes edge details and semantic consistency, and is suitable for mobile and embedded device deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339619A_ABST
    Figure CN120339619A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and image processing, in particular to an efficient image segmentation system based on an improved UNet model, which comprises an input preprocessing module, an improved encoder module, a multi-scale feature fusion module, a dynamic decoder module and a post-processing output module. The output end of the input preprocessing module is in one-way connection with the input end of the improved encoder module, and the output end of the improved encoder module is in two-way data transmission with the input end of the multi-scale feature fusion module and the initial layer of the dynamic decoder module. The output end of the multi-scale feature fusion module is in cross-scale connection with the middle layer of the dynamic decoder module, and the tail end output layer of the dynamic decoder module is cascaded with the input end of the post-processing output module. Through a lightweight encoder, dynamic feature fusion, sub-pixel convolution, model compression and the like, efficient image segmentation is realized, precision and efficiency are improved, edge details are optimized, and multi-scene lightweight deployment is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and particularly to an efficient image segmentation system based on an improved UNet model. Background Technique

[0002] Image segmentation, as a key technology in the field of computer vision, aims to divide an image into regions with specific semantics, and has important application values in fields such as medical image diagnosis, remote sensing data analysis, and industrial inspection. Early methods relied on manually designed features, such as threshold segmentation, edge detection, etc. Limited by the feature expression ability, it was difficult to handle complex scenarios. With the rise of deep learning, segmentation models based on convolutional neural networks (CNNs) have developed rapidly. Among them, the UNet model, with its encoder-decoder structure and cross-layer connection design, effectively integrates deep semantic information and shallow spatial details, significantly improving the segmentation accuracy and becoming the mainstream framework in this field. Subsequently, researchers have carried out improvements around UNet, including reducing the computational complexity of the encoder through lightweight convolution operations, introducing multi-scale feature fusion strategies to enhance the utilization of context information, optimizing the upsampling method of the decoder to improve the resolution, etc., promoting the continuous progress of image segmentation technology in terms of efficiency and accuracy.

[0003] However, there are still many problems to be solved in the existing technology: in the encoder design, the traditional convolutional unit has a large amount of computation. Especially when processing high-resolution images, the parameter scale and computational time-consuming increase significantly, making it difficult to meet the real-time requirements; during the cross-scale feature fusion process, the weight allocation of different-level features usually adopts a fixed mode, lacking dynamic perception of feature importance, resulting in limited fusion efficiency of semantic information and spatial details. In the decoder upsampling link, the commonly used transposed convolution or bilinear interpolation methods are prone to introducing artifacts, affecting the accuracy of edge details; in the post-processing stage, the edge optimization and semantic correction of the segmentation results are relatively rough, making it difficult to meet the fine segmentation requirements in complex scenarios. In addition, model compression technology has problems such as a single channel pruning strategy and obvious accuracy loss during quantization in practical applications, restricting the efficient deployment of segmentation models on mobile or embedded devices.

[0004] Therefore, there is an urgent need to design an image segmentation system that can balance computational efficiency and segmentation accuracy, and has a lightweight encoder structure, a dynamic feature fusion mechanism, a precise upsampling method, and an efficient post-processing strategy. Summary of the Invention

[0005] The purpose of the present invention is to overcome the above problems and provide an efficient image segmentation system based on an improved UNet model. To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] An efficient image segmentation system based on an improved UNet model, including an input preprocessing module, an improved encoder module, a multi-scale feature fusion module, a dynamic decoder module, and a post-processing output module; the output end of the input preprocessing module is unidirectionally connected to the input end of the improved encoder module, and the output end of the improved encoder module transmits data bidirectionally to the input end of the multi-scale feature fusion module and the initial layer of the dynamic decoder module respectively. The output end of the multi-scale feature fusion module is cross-scale connected to the middle layer of the dynamic decoder module, and the end output layer of the dynamic decoder module is cascaded with the input end of the post-processing output module.

[0007] Further, the improved encoder module is composed of N levels of cascaded lightweight convolution units, and the mathematical expression of each level of unit is:

[0008] F l = BNDSConv(DilatedConv(X l-1 , r l ), k l ))

[0009] Among them, F l is the output feature map of the l-th layer, DSConv is the depthwise separable convolution operation, DilatedConv is the dilated convolution operation with a dilation rate of r l , X l-1 is the input feature map of the previous layer, and k l is the convolution kernel size.

[0010] Further, the cross-scale connection of the multi-scale feature fusion module is realized through a bidirectional feature pyramid, and its fusion formula is:

[0011] H l = α l ·Conv(F l )+(1 - α l )·UpSample(D L-l )

[0012] Among them, H l is the feature map after fusion of the l-th layer, F l is the output feature of the l-th layer of the encoder, D L-l is the feature map of the L-l-th layer of the decoder, αl is the dynamic weight coefficient, which is generated by a learnable parameter through the Softmax function and satisfies

[0013] Further, the cross-scale connection of the middle layer of the dynamic decoder module adopts an adaptive channel selection mechanism, and its channel screening threshold is:

[0014] θ = μ(S c ) + 0.5σ(Sc )

[0015] Among them, μ(S c ) is the mean of the channel saliency scores, and σ(S c ) is the standard deviation. The screened feature map doubles the resolution through the sub-pixel convolutional layer.

[0016] Furthermore, the cascaded input end of the post-processing output module receives the output feature map of the dynamic decoder, and generates the final segmentation result through the series path of the edge refinement sub-module and the semantic correction sub-module.

[0017] Furthermore, the calculation of the dynamic weight coefficient α l introduces a spatial attention mechanism, specifically: perform global average pooling on the encoder feature map F l and the decoder feature map D L-l respectively to generate the channel description vectors and After mapping through the fully connected layer, generate the weight ratio through the Sigmoid function:

[0018]

[0019] Among them, W is the trainable parameter matrix, and [;] represents the vector concatenation operation.

[0020] Furthermore, the one-way connection transmission path of the input preprocessing module sequentially performs adaptive histogram equalization and noise suppression filtering operations, and its mathematical representation is:

[0021]

[0022] Among them, CLAHE is the contrast-limited adaptive histogram equalization operation, and K Gaussian is the Gaussian filter kernel with σ = 1.5, represents the convolution operation.

[0023] Furthermore, in the series path of the edge refinement sub-module, the spatial kernel standard deviation σ s and the color kernel standard deviation σ c are adaptively adjusted according to the image gradient, satisfying:

[0024] σ s = 0.1·max(G(x,y)), σ c = 0.3·median(|I(x,y) - μ I |)

[0025] Among them, G(x,y) is the Sobel gradient magnitude, and μ I is the mean of the image grayscale.

[0026] Furthermore, it also includes a model compression module. The input end of the model compression module is feedback-connected to the output end of the improved encoder module and the feature map output end of the encoder. Pruning is performed based on the channel importance score, and convolutional kernels with channel importance scores lower than the threshold τ = 0.2·max(S c ) are deleted through structured pruning, and the remaining parameters are quantized to 8 bits.

[0027] Furthermore, the step-by-step upsampling process of the dynamic decoder module is implemented through a sub-pixel convolutional layer, and its resolution doubling formula is:

[0028] Y = PixelShuffle(Conv(X))

[0029] where X is the input feature map, PixelShuffle is the pixel rearrangement operation, and Y is the output high-resolution feature map.

[0030] The advantages of the present invention are as follows:

[0031] 1. By adopting a lightweight convolutional unit that combines depthwise separable convolution and dilated convolution in the improved encoder module and integrating a model compression module for structured pruning and 8-bit quantization, the present invention reduces the calculation parameters and model complexity while retaining the multi-scale feature extraction ability, achieving a significant improvement in the calculation efficiency during the image segmentation process, enabling the system to efficiently process high-resolution images, meeting the real-time application requirements, and supporting lightweight deployment on mobile or embedded devices.

[0032] 2. By introducing a spatial attention mechanism through the multi-scale feature fusion module to generate dynamic weight coefficients, adaptively weighting and fusing the deep semantic features of the encoder and the shallow detail features of the decoder, and using the adaptive channel selection mechanism of the dynamic decoder module to screen important channel features, the present invention realizes the accurate perception and efficient fusion of the importance of different-level features, effectively enhancing the feature expression ability and improving the accuracy of image segmentation in complex scenarios, especially outstanding in edge detail processing and semantic consistency optimization.

[0033] 3. By using a sub-pixel convolutional layer in the dynamic decoder module to replace the traditional upsampling method, the present invention avoids the artifact problems introduced by deconvolution or bilinear interpolation. Combining with a bilateral filter that adaptively adjusts parameters based on the image gradient in the post-processing output module for edge refinement, and semantic correction of the segmentation result by the semantic correction sub-module, the present invention realizes the high-quality restoration of the resolution of the segmentation result and the accurate optimization of edge details, enabling the final segmented image to reach a higher level in terms of edge sharpness, semantic accuracy, etc. Description of the Drawings

[0034] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application, making other features, objectives, and advantages of this application more obvious. The schematic embodiments and their descriptions of this application are used to explain this application and do not constitute an improper limitation of this application.

[0035] In the accompanying drawings:

[0036] Figure 1 It is a module interaction diagram of an efficient image segmentation system based on an improved UNet model in Embodiment 1.

[0037] Figure 2 It is a timing diagram of an efficient image segmentation system based on an improved UNet model in Embodiment 1. Detailed implementation manners

[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0039] The present invention will be introduced in detail and specifically below through specific embodiments to better understand the present invention. However, the following embodiments do not limit the protection scope of the present invention.

[0040] Embodiment 1

[0041] As Figure 1-2 shown, an efficient image segmentation system based on an improved UNet model includes an input preprocessing module, an improved encoder module, a multi-scale feature fusion module, a dynamic decoder module, and a post-processing output module; the output end of the input preprocessing module is unidirectionally connected to the input end of the improved encoder module, the output end of the improved encoder module respectively transmits data bidirectionally to the input end of the multi-scale feature fusion module and the initial layer of the dynamic decoder module, the output end of the multi-scale feature fusion module is cross-scale connected to the middle layer of the dynamic decoder module, and the end output layer of the dynamic decoder module is cascaded with the input end of the post-processing output module.

[0042] Each module works collaboratively through a specific data transmission connection method. After the input preprocessing module preliminarily processes the image, it is transmitted to the improved encoder module. The feature data output by the encoder is respectively transmitted to the multi-scale feature fusion module and the dynamic decoder module. The multi-scale feature fusion module is cross-scale connected to the middle layer of the decoder, and the output layer at the end of the decoder is cascaded with the post-processing output module. This module division and connection method constructs a complete image segmentation processing flow, enabling each module to cooperate and achieve efficient segmentation of the image, laying a system architecture foundation for the realization of the specific functions of subsequent modules.

[0043] Further, the improved encoder module is composed of N levels of cascaded lightweight convolution units, and the mathematical expression of each level of unit is:

[0044] F l = BN(DSConv(DilatedConv(X l-1 , r l ), k l ))

[0045] Among them, F l is the output feature map of the l-th layer, DSConv is the depthwise separable convolution operation, DilatedConv is the dilated convolution operation with dilation rate r l , X l-1 is the input feature map of the previous layer, and k l is the convolution kernel size.

[0046] The improved encoder module is composed of N levels of cascaded lightweight convolution units. Each level of unit first performs a dilated convolution operation with dilation rate r l to expand the receptive field and extract multi-scale features, then performs a depthwise separable convolution operation to reduce the number of calculation parameters and lower the model complexity, and finally performs a batch normalization operation to normalize the features, making the output feature map more stable. The design of the lightweight convolution unit effectively reduces the amount of calculation while ensuring the feature extraction ability, improves the running efficiency of the encoder, and meets the requirements of efficient image segmentation.

[0047] Further, the cross-scale connection of the multi-scale feature fusion module is realized through a bidirectional feature pyramid, and its fusion formula is:

[0048] H l = α l ·Conv(F l )+(1 - α l )·UpSample(D L-l )

[0049] Among them, H l is the feature map after fusion of the l-th layer, F lis the output feature of the l-th layer of the encoder, D L-l is the feature map of the (L - l)-th layer of the decoder, α l is the dynamic weight coefficient, which is generated by a learnable parameter through the Softmax function and satisfies

[0050] The multi-scale feature fusion module realizes cross-scale connection through a bidirectional feature pyramid, performs convolutional processing on the output feature of the l-th layer of the encoder, performs upsampling processing on the feature map of the (L - l)-th layer of the decoder, and then uses the dynamic weight coefficient α l to perform weighted fusion on the two. The dynamic weight coefficient is generated by a learnable parameter through the Softmax function and satisfies the condition that the sum of the weights is 1. This fusion method can adaptively combine the deep semantic features of the encoder and the shallow detail features of the decoder, make full use of the feature information of different scales, improve the expression ability and fusion effect of the features, and thus improve the accuracy of image segmentation.

[0051] Furthermore, the intermediate layer cross-scale connection of the dynamic decoder module adopts an adaptive channel selection mechanism, and its channel screening threshold is:

[0052] θ = μ(S c ) + 0.5σ(S c )

[0053] where μ(S c ) is the mean of the channel saliency scores, σ(S c ) is the standard deviation, and the screened feature map is doubled in resolution through a sub-pixel convolutional layer.

[0054] The intermediate layer cross-scale connection of the dynamic decoder module adopts an adaptive channel selection mechanism. The channel screening threshold is calculated based on the mean of the channel saliency scores and 0.5 times the standard deviation. Important channels are screened out through this threshold, and the screened feature map is doubled in resolution through a sub-pixel convolutional layer, retaining key features while increasing the resolution. The adaptive channel selection mechanism can remove redundant channel information and reduce the computational amount. The sub-pixel convolutional layer avoids the artifact problems that may be brought by traditional upsampling methods, improving the resolution and quality of the feature map during the decoding process.

[0055] Furthermore, the cascaded input end of the post-processing output module receives the output feature map of the dynamic decoder and generates the final segmentation result through the series path of the edge refinement sub-module and the semantic correction sub-module.

[0056] The post - processing output module receives the output feature map of the dynamic decoder, and processes the feature map through the cascaded path of the edge refinement sub - module and the semantic correction sub - module. The edge refinement sub - module is used to optimize the edge details of the segmentation result, and the semantic correction sub - module is used to adjust the semantic consistency of the segmentation result, and finally generates an accurate final segmentation result. The post - processing module further optimizes the feature map output by the decoder, improving the edge clarity and semantic accuracy of the segmentation result, making the segmentation result more refined and in line with the actual semantics.

[0057] Furthermore, the calculation of the dynamic weight coefficient α l introduces a spatial attention mechanism, specifically: performing global average pooling on the encoder feature map F l and the decoder feature map D L-l respectively to generate channel description vectors and After being mapped through a fully - connected layer and passed through the Sigmoid function to generate a weight ratio:

[0058]

[0059] where W is a trainable parameter matrix, and [;] represents the vector concatenation operation.

[0060] When calculating the dynamic weight coefficient α l a spatial attention mechanism is introduced. Global average pooling is performed on the encoder feature map and the decoder feature map respectively to generate channel description vectors. After concatenating these two vectors and mapping through a fully - connected layer, a weight ratio is generated through the Sigmoid function. This weight ratio is used for weighted fusion of the encoder and decoder features during multi - scale feature fusion. The introduction of the spatial attention mechanism enables the model to dynamically adjust the fusion weights according to the spatial context information of the feature map, pay more attention to the regional features important for the segmentation task, improve the pertinence and effectiveness of feature fusion, and thus improve the performance of image segmentation.

[0061] Furthermore, the unidirectional connection transmission path of the input pre - processing module sequentially performs adaptive histogram equalization and noise suppression filtering operations, and its mathematical representation is:

[0062]

[0063] where CLAHE is the contrast - limited adaptive histogram equalization operation, K Gaussian is a Gaussian filter kernel with σ = 1.5, represents the convolution operation.

[0064] The unidirectional connection transmission path of the input preprocessing module sequentially performs adaptive histogram equalization and noise suppression filtering operations. First, the local contrast of the image is enhanced through the contrast-limited adaptive histogram equalization operation, and then a Gaussian filter with a standard deviation of 1.5 is used to perform convolution operations on the processed image to suppress noise. The preprocessing operation improves the quality of the input image, enhances the contrast of the image, reduces noise interference, provides higher-quality input data for subsequent feature extraction and segmentation processing, and helps improve the segmentation effect of the entire system.

[0065] Further, in the cascaded path of the edge refinement sub-module, the spatial kernel standard deviation σ s and the color kernel standard deviation σ c of the bilateral filter are adaptively adjusted according to the image gradient, satisfying:

[0066] σ s = 0.1·max(G(x,y)), σ c = 0.3·median(|I(x,y) - μ I |)

[0067] where G(x,y) is the Sobel gradient magnitude, and μ I is the average image gray value.

[0068] In the cascaded path of the edge refinement sub-module, the spatial kernel standard deviation σ s and the color kernel standard deviation σ c of the bilateral filter are adaptively adjusted according to the image gradient. σ s is obtained by multiplying the maximum value of the Sobel gradient magnitude by 0.1, and σ c is obtained by multiplying the median of the absolute value of the difference between the average image gray value and the pixel gray value by 0.3. The image is filtered by adaptively adjusting the parameters. Dynamically adjusting the parameters of the bilateral filter according to the image gradient can better retain the edge details of the image during the filtering process, avoid edge blurring, and improve the accuracy and clarity of the edges in the segmentation result.

[0069] Further, it also includes a model compression module. The input end of the model compression module is feedback-connected to the feature map output end of the encoder from the output end of the improved encoder module. Pruning is performed based on the channel importance score, and convolutional kernels with channel importance scores lower than the threshold τ = 0.2·max(S c ) are deleted through structured pruning, and the remaining parameters are quantized to 8 bits.

[0070] Through pruning and quantization operations, the model compression module reduces the number of parameters and the computational amount of the model, and improves the running speed and efficiency of the model on the premise of maintaining the model performance, making it more suitable for deployment and application on resource-constrained devices.

[0071] Furthermore, the step-by-step upsampling process of the dynamic decoder module is implemented through a sub-pixel convolutional layer, and its resolution doubling formula is:

[0072] Y = PixelShuffle(Conv(X))

[0073] where X is the input feature map, PixelShuffle is the pixel rearrangement operation, and Y is the output high-resolution feature map.

[0074] The step-by-step upsampling process of the dynamic decoder module is implemented through a sub-pixel convolutional layer. First, a convolutional operation is performed on the input feature map, and then the resolution of the feature map is doubled through the pixel rearrangement operation to obtain a high-resolution output feature map. The use of the sub-pixel convolutional layer avoids the image blurring and artifact problems that may be brought about by traditional upsampling methods, and can more effectively improve the resolution of the feature map, providing a guarantee for generating high-quality segmentation results.

[0075] The specific embodiments of the present invention have been described in detail above, but they are only examples, and the present invention is not equivalent to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions made to the present invention are also within the scope of the present invention. Therefore, all equivalent transformations and modifications made without departing from the spirit and scope of the present invention should be covered within the scope of the present invention.

Claims

1. An efficient image segmentation system based on an improved UNet model, characterized in that: It includes an input preprocessing module, an improved encoder module, a multi-scale feature fusion module, a dynamic decoder module, and a post-processing output module; the output end of the input preprocessing module is unidirectionally connected to the input end of the improved encoder module, the output end of the improved encoder module bi-directionally transmits data to the input end of the multi-scale feature fusion module and the initial layer of the dynamic decoder module respectively, the output end of the multi-scale feature fusion module is cross-scale connected to the middle layer of the dynamic decoder module, and the end output layer of the dynamic decoder module is cascaded with the input end of the post-processing output module.

2. The efficient image segmentation system based on the improved UNet model according to claim 1, characterized in that: The improved encoder module is composed of N levels of cascaded lightweight convolution units, and the mathematical expression of each level of unit is: F l = BN(DSConv(DilatedConv(X l-1 , r l ), k l )) Among them, F l is the output feature map of the l-th layer, DSConv is the depthwise separable convolution operation, and DilatedConv is the dilated convolution operation with dilation rate r l ; X l-1 is the input feature map of the previous layer, and k l is the convolution kernel size.

3. The efficient image segmentation system based on the improved UNet model according to claim 2, characterized in that: The cross-scale connection of the multi-scale feature fusion module is realized through a bidirectional feature pyramid, and its fusion formula is: H l = α l · Conv(F l ) + (1 - α l ) · UpSample(D L-l ) Among them, H l is the feature map after fusion of the l-th layer, F l is the output feature of the l-th layer of the encoder, D L-l is the feature map of the (L-l)-th layer of the decoder, α l is the dynamic weight coefficient, which is generated by the Softmax function through learnable parameters and satisfies 4. The efficient image segmentation system based on the improved UNet model according to claim 3, characterized in that: The cross-scale connection of the middle layer of the dynamic decoder module adopts an adaptive channel selection mechanism, and its channel screening threshold is: θ = μ(S c ) + 0.5σ(S c ) Among them, μ(S c ) is the mean of the channel saliency scores, and σ(S c ) is the standard deviation. The resolution of the filtered feature map is doubled through a sub-pixel convolutional layer.

5. The efficient image segmentation system based on the improved UNet model according to claim 4, characterized in that: The cascaded input end of the post-processing output module receives the output feature map of the dynamic decoder, and generates the final segmentation result through the serial path of the edge refinement sub-module and the semantic correction sub-module.

6. The efficient image segmentation system based on the improved UNet model according to claim 5, wherein: The dynamic weight coefficient α l is calculated by introducing a spatial attention mechanism, specifically: for the encoder feature map F l and the decoder feature map D L-l perform global average pooling respectively to generate channel description vectors and After being mapped by a fully connected layer, generate a weight ratio through the Sigmoid function: Among them, W is a trainable parameter matrix, and [;] represents the vector splicing operation.

7. The efficient image segmentation system based on the improved UNet model according to claim 6, characterized in that: The unidirectional connection transmission path of the input preprocessing module sequentially performs adaptive histogram equalization and noise suppression filtering operations, and its mathematical representation is: Among them, CLAHE is the Contrast Limited Adaptive Histogram Equalization operation, and K Gaussian is a Gaussian filter kernel with σ = 1.5, denotes the convolution operation.

8. The efficient image segmentation system based on the improved UNet model according to claim 7, characterized in that: In the series path of the edge thinning sub-module, the spatial kernel standard deviation σ of the bilateral filter s and the color kernel standard deviation σ c are adaptively adjusted according to the image gradient and satisfy: σ s = 0.1·max(G(x,y)), σ c = 0.3·median(|I(x,y) - μ I |) Among them, G(x, y) is the Sobel gradient magnitude, and μ I is the average image grayscale value.

9. The efficient image segmentation system based on the improved UNet model according to claim 8, characterized in that: It also includes a model compression module. The input end of the model compression module is feedback-connected to the output end of the improved encoder module to the feature map output end of the encoder, performs pruning based on the channel importance score, deletes the convolution kernels with channel importance scores lower than the threshold τ = 0.2·max(Sc), and performs 8-bit quantization on the remaining parameters.

10. The efficient image segmentation system based on the improved UNet model according to claim 9, characterized in that: The step-by-step upsampling process of the dynamic decoder module is realized through a sub-pixel convolution layer, and its resolution doubling formula is: Y = PixelShuffle(Conv(X)) Among them, X is the input feature map, PixelShuffle is the pixel rearrangement operation, and Y is the output high-resolution feature map.