Image segmentation method, system, device and medium

By compressing the network architecture and adding specific modules in the U3Plus network model, the LM-U3Plus image segmentation model is formed, which solves the problem of high computational complexity of the existing image segmentation model and achieves more efficient feature extraction and image segmentation accuracy.

CN120047421AActive Publication Date: 2025-05-27FUYANG NORMAL UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510137479.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-27
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

The existing image segmentation model has high computational complexity and is difficult to meet the requirements of real-time and resource efficiency in industrial applications.

Method used

Based on the U3Plus network model, the image segmentation model LM-U3Plus is formed by compressing the network architecture into four layers, and adding hybrid channel convolution RFC module, feature fusion module MFF and ED-VSS Layers module. This model performs feature extraction through hybrid channel convolution and depth separation convolution, combining feature fusion and blocking processing to reduce the computational complexity.

Benefits of technology

It effectively reduces the computational complexity and parameter quantity of the model, improves feature extraction efficiency and accuracy, improves the real-time and resource efficiency of image segmentation, and maintains high segmentation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047421A_ABST
    Figure CN120047421A_ABST
Patent Text Reader

Abstract

The invention provides an image segmentation method, system and device and a medium, and belongs to the technical field of image segmentation, and the method comprises the following steps: obtaining an industrial scene image data set; the method comprises the following steps: taking a U3Plus network model as a basic model, compressing U3Plus into a four-layer network architecture, and sequentially adding a hybrid channel convolution RFC module, a feature fusion module MFF and an ED-VSS Layers module into the compressed network architecture to obtain an image segmentation model LM-U3Plus; inputting an industrial scene image data set into the LM-U3Plus, and performing feature extraction on the input low-dimensional features and high-dimensional features by using a hybrid channel convolution RFC module; and fusing the extracted features by using a feature fusion module MFF, and carrying out block processing on the fused features by using an ED-VSS Layers module and outputting to obtain a segmented image. The image segmentation efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation, and particularly relates to an image segmentation method, system, device and medium. Background Art

[0002] With the continuous advancement of intelligent manufacturing and industrial automation, image processing technology, especially image segmentation technology, has gradually become a crucial part in the industrial field. In industrial application scenarios such as quality inspection, workpiece recognition, and defect detection, image segmentation technology can automatically and accurately extract the target area from the image, providing an important basis for subsequent decision-making and operations. However, existing image segmentation models usually have a high computational complexity and are difficult to meet the strict requirements for real-time performance and resource efficiency in industrial applications. Especially in an industrial environment, computational resources are usually limited. Therefore, how to design a lightweight model that can not only ensure segmentation accuracy but also effectively reduce the computational overhead has become an important topic in current research.

[0003] The existing mainstream image segmentation model, the CNN network model, has achieved remarkable results in tasks such as semantic segmentation and instance segmentation. Models such as SegNet, FCN, and U-Net have strong performance, but when dealing with complex industrial scenarios, due to a large number of parameters, the computational complexity is high, and thus it is still difficult to meet the requirements of industrial scenarios in terms of real-time performance and resource requirements. Summary of the Invention

[0004] In order to overcome the deficiencies of the above-mentioned existing technologies, the present invention provides an image segmentation method, which includes the following steps:

[0005] Obtain an industrial scenario image dataset;

[0006] Based on the U3Plus network model, compress the U3Plus into a four-layer network architecture, and sequentially add a hybrid channel convolution RFC module, a feature fusion module MFF, and an ED-VSS Layers module to the compressed network architecture to obtain an image segmentation model LM-U3Plus;

[0007] Input the images in the industrial scenario image dataset into LM-U3Plus, use the hybrid channel convolution RFC module to respectively perform feature extraction on the input low-dimensional features and high-dimensional features; use the feature fusion module MFF to fuse the extracted features, and use the ED-VSS Layers module to perform block processing on the fused features and output, to obtain the segmented image.

[0008] Preferably, the hybrid channel convolution RFC module includes multi-channel convolution and depthwise separable convolution, uses the multi-channel convolution to perform feature extraction on the low-dimensional features, and uses the depthwise separable convolution to extract features from the high-dimensional features.

[0009] Preferably, a channel attention mechanism ECA module is added after the multi-channel convolution and depthwise separable convolution of the RFC module. The channel attention mechanism ECA module performs multi-scale fusion on the extracted features and performs layer-by-layer channel weighting operations on the features of each different scale.

[0010] Preferably, the feature fusion module MFF is used to fuse the extracted features. Specifically, MFF adaptively weights and fuses and splices the multi-level features after the layer-by-layer channel weighting operation, and then performs feature fusion on the spliced features through 1×1 and 3×3 convolutions.

[0011] Preferably, the ED-VSS Layers module is used to perform block processing on the fused features. Specifically, ED-VSS Layers divides the fused features into four parts according to the channel dimension. Each part extracts features through the VSS Block respectively, and then fuses the four extracted parts of features through the ECA module. GroupNorm is used to normalize the fused features. Finally, residual connection is used to retain the original input features, and a weighting coefficient is introduced to dynamically regulate the residual path to obtain the segmented image.

[0012] The present invention also provides an image segmentation system, including:

[0013] An image acquisition module, configured to acquire an industrial scenario image dataset;

[0014] A model construction module, configured to use the U3Plus network model as the basic model, compress the U3Plus into a four-layer network architecture, and sequentially add a hybrid channel convolution RFC module, a feature fusion module MFF, and an ED-VSS Layers module to the compressed network architecture to obtain an image segmentation model LM-U3Plus;

[0015] An image segmentation module, configured to input the images in the industrial scenario image dataset into the LM-U3Plus, use the hybrid channel convolution RFC module to respectively extract features from the input low-dimensional features and high-dimensional features; use the feature fusion module MFF to fuse the extracted features, use the ED-VSS Layers module to perform block processing on the fused features and output to obtain the segmented image.

[0016] The present invention also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is configured to run the computer program in the memory to execute the image segmentation method.

[0017] The present invention also provides a computer-readable storage medium storing a computer program, which is adapted to be loaded by a processor to execute the image segmentation method.

[0018] The image segmentation method, system, device and medium provided by the present invention have the following beneficial effects:

[0019] In the present invention, when the images in the industrial scenario image dataset are input into the image segmentation model LM-U3Plus and the input images are processed using a network architecture compressed into four layers, the computational complexity and the number of parameters of LM-U3Plus can be effectively reduced; by adopting RFC for feature extraction, the feature extraction efficiency and accuracy can be improved; by using the feature fusion module MFF for multi-level feature fusion, the feature information of different levels can be efficiently integrated; by using ED-VSS Layers to lightweight the VSS Block module, the computational performance and feature extraction accuracy of the model can be optimized. When the image segmentation model LM-U3Plus constructed by the present invention processes complex industrial image segmentation, the number of parameters can be greatly reduced and the segmentation accuracy can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present invention and their design schemes, the accompanying drawings required for the present embodiments will be briefly introduced below. The accompanying drawings in the following description are only partial embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a flowchart of the image segmentation method according to the embodiment of the present invention;

[0022] Figure 2 It is a network structure diagram of U3Plus;

[0023] Figure 3 It is a structure diagram of the ECA module;

[0024] Figure 4 It is a structure diagram of the VSS Block;

[0025] Figure 5 It is a diagram of the selective scanning mechanism;

[0026] Figure 6 It is a structure design diagram of LM-U3Plus;

[0027] Figure 7 It is a diagram of the hybrid convolution scheme;

[0028] Figure 8 It is a schematic diagram of applying the ECA module to the RFC module;

[0029] Figure 9 It is the structure diagram of the RFC module;

[0030] Figure 10 It is the structure diagram of the MFF module;

[0031] Figure 11 It is the structure design diagram of the ED-VSS Layers;

[0032] Figure 12 It is the diagram of the empty belt segmentation;

[0033] Figure 13 It is the belt material flow segmentation;

[0034] Figure 14 It is the special case segmentation. Specific implementation manners

[0035] In order to enable those skilled in the art to better understand the technical solution of the present invention and be able to implement it, the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0036] Embodiment

[0037] The present invention provides an image segmentation method, specifically as Figure 1 shown, including the following steps:

[0038] Step 1: Obtain an industrial scenario image dataset.

[0039] The dataset adopted by the present invention is a self-collected industrial scenario dataset, and the scenario is industrial logistics transportation. There are a total of 4500 pictures, including 4000 training set pictures and 500 validation set pictures. The GPU model used in the experiment is NVIDIA GeForce RTX 2080Ti, the video memory size is 11GB, and the cuda version is 11.8.

[0040] Step 2: Based on the U3Plus network model, compress the U3Plus into a four-layer network architecture, and sequentially add a hybrid channel convolution RFC module, a feature fusion module MFF, and an ED-VSS Layers module to the compressed network architecture to obtain an image segmentation model LM-U3Plus.

[0041] (1) The present invention designs the LM-U3Plus network model based on the U3Plus network model. The specific structure of the U3Plus network model is as follows:

[0042] The U3Plus network model is an improved version of the U-Net network model. Its structure includes an ECA module and a Visual State Space (VSS) Block module, as specifically shown in Figure 2 the following figure.

[0043] By effectively modeling the dependencies between channels, the ECA module enhances the network's ability to focus on important features while maintaining the lightweight characteristics of the model. The structure is shown in Figure 3 the following figure.

[0044] The ECA module generates channel attention weights through adaptive average pooling (GAP), 1D convolution, and the Sigmoid activation function, and applies them to the input feature map to enhance the expressive power of the features. Its core idea is to improve the feature representation ability by adaptively adjusting the weights of each channel. Different from the traditional channel attention mechanism, the ECA module realizes the weight learning between channels by introducing one-dimensional convolution, without the complex calculations of the fully connected layer. Its advantage is to maintain a low computational cost while improving the effectiveness of feature extraction, thereby enhancing the performance of the model in image processing tasks.

[0045] The Visual State Space (VSS) Block module focuses on enhancing the processing ability of multi-resolution features through flexible feature scaling and splicing. It introduces variable skip connections to dynamically adjust the information flow between feature maps of different scales, enabling the network to more effectively fuse features from different levels while controlling memory occupancy and computational volume. The design of the VSS Block takes into account the adaptive fusion of features, enabling the model to more effectively capture important information in image segmentation tasks. The specific structure is shown in Figure 4 the following figure.

[0046] The core part of the VSS Block module is the SS2D module, which implements a new self-attention mechanism to process the input features through a linear layer and depthwise separable convolution (DW-Conv). Specifically, the SS2D module performs a linear projection on the input features and extracts local features through DW-Conv. This structure not only effectively reduces the computational volume but also maintains the expressive power of the features. Among them, the SS2D module adopts a selective scanning mechanism, and its operating mechanism is shown in Figure 5 the following figure.

[0047] The selective scanning mechanism of SS2D consists of three parts: Scan Expanding, S6 block, and Scan Merging. The Scan Expanding operation unfolds the input image into sequences along four different directions (from top left to bottom right, from bottom right to top left, from top right to bottom left, and from bottom left to top right). Then, these sequences are subjected to feature extraction by the S6 block to ensure that information from all directions is scanned, thereby capturing different features. The Scan Merging operation sums and merges the features from each direction to restore the image to the input size.

[0048] In the VSS Block, the input features are first normalized and then subjected to feature extraction by the self-attention module. The self-attention mechanism enables the network to adaptively adjust the weights of the features according to the context information of the input features, thereby enhancing the influence of important features. Finally, through the DropPath operation, the module can randomly discard certain paths to improve the generalization ability of the model.

[0049] (2) The specific process of improving the U3Plus network model to LM-U3Plus is as follows:

[0050] In the U3Plus base model, although its design has considered the extraction and fusion of multi-scale features, there is a lack of an effective fusion strategy for processing low-level features. This problem may lead to the loss of low-level feature information in some complex scenarios, thereby affecting the overall segmentation performance of the model. To improve the feature expression ability of the model at different scales, LM-U3Plus introduces a full-scale feature fusion strategy to improve the U3Plus network model. At the same time, the model is compressed to reduce the number of model parameters, converting the network model from the original five-layer architecture to a four-layer architecture. The model design is as Figure 6 shown.

[0051] In the U3Plus base network, the convolution module adopts multi-channel hybrid convolution. Although this design can achieve high segmentation accuracy, its computational complexity and the number of parameters also increase, bringing a large burden to practical applications. Specifically, first, depthwise separable convolution is attempted for single-channel convolution. This convolution method effectively reduces the computational complexity by splitting the standard convolution into depthwise convolution and pointwise convolution. However, in practical applications, although using depthwise separable convolution has computational advantages, it may lead to a decrease in segmentation accuracy in some cases, especially when dealing with low-dimensional features, and details and information may be lost as a result.

[0052] To solve this problem, a mixed-channel convolution scheme is introduced in LM-U3Plus to better balance accuracy and computational efficiency. The specific scheme is asFigure 7 as shown

[0053] In this solution, for low-dimensional features, multi-channel convolution is still used to ensure sufficient extraction of feature information and maintain high segmentation accuracy. For high-dimensional features, depthwise separable convolution is adopted, which can not only reduce the computational amount but also have little impact on the segmentation effect. Through this design, the model can better utilize the advantages of low-dimensional features in detail extraction and improve the computational efficiency in high-dimensional feature processing at the same time.

[0054] To further improve the model performance, an ECA module is added after the convolutional layer. This module adaptively generates weights for each channel through 1D convolution. However, when assigning weights to channels, the original ECA module does not consider the problem that features in different layers may have different representation capabilities, resulting in insufficient capture of global information during feature processing, especially between deep and shallow features. To address the above problems, LM-U3Plus optimizes the original ECA module by introducing multi-scale feature extraction and layer-by-layer channel weighting into the original ECA module, and applies the optimized ECA module to the RFC module, as Figure 8 as shown

[0055] To fully capture the feature information at each level, the ECA module is mainly optimized by multi-scale feature extraction and layer-by-layer channel weighting using multiple convolutional kernels. The specific operations are as follows:

[0056] (1) Multi-scale feature extraction.

[0057] To fuse multi-scale features, multiple convolutional kernels with different sizes (k = 3, 5, 7) are introduced. This multi-scale processing method can capture more levels of features, enabling the network to better process different sizes of receptive fields and feature information, thereby improving the model's expressive ability.

[0058] (2) Layer-by-layer channel weighting.

[0059] To further improve the accuracy of channel selection, layer-by-layer feature weighting is added to each feature extraction operation with different scales, that is, features extracted by different convolutional kernels will get different weights, allowing the model to adaptively adjust the contribution of each channel at different scales. The output result of each convolutional kernel will be multiplied by its corresponding channel weighting coefficient, and the multi-scale features after weighting will finally be fused together.

[0060] Overall, by adding the ECA module to the U3Plus network, not only the feature extraction ability of the model is enhanced, but also a new idea is provided for the design of lightweight industrial segmentation models. This improvement effectively balances the computational complexity and segmentation performance of the model, laying a solid foundation for achieving efficient and accurate image segmentation tasks. Aiming at the problem of large span between the four layers of the LM-U3Plus network, the RFC module further adopts the method of forcibly retaining the original input features to avoid the loss of original input feature information caused by large feature span. Specifically, global average pooling is performed on the original input image and spliced to each layer of features respectively, so that each layer of features retains the original input information. The overall structure of the RFC module is as shown in Figure 9 as follows.

[0061] For the U3Plus network model, simple splicing fusion is used for the fusion of multi-layer features. However, features from different levels have different attentions to information. If simple splicing is used, the attention to important information may be reduced. To solve this problem, the LM-U3Plus uses the MFF module to process multi-level fusion features. Specifically, for multi-level features, adaptive weights are used for fusion splicing, and then the spliced features are subjected to feature fusion through 1×1 and 3×3 convolutions. The MFF module enables the model to improve the attention to important information and reduce the interference of irrelevant information during the feature fusion process. The specific design is as shown in Figure 10 as follows.

[0062] To improve the processing ability of the network model for multi-scale features, the VSS Block module is introduced to optimize the fusion features. However, although the original VSS Block module can effectively optimize the fusion features, its computational complexity and number of parameters are large, which is not suitable for lightweight industrial image segmentation models. To solve this problem, the LM-U3Plus uses ED-VSS Layers obtained by improving the VSS Block to process the fusion features, effectively reducing the number of parameters and computational complexity. The specific design of ED-VSS Layers is as shown in Figure 11 as follows.

[0063] ED-VSS Layers divides the input features into four parts according to the channel dimension, and each part is subjected to feature extraction through the VSSBlock respectively. The block processing reduces the computational overhead of each sub-module and improves the localization effect of feature extraction at the same time. The four parts of the extracted features are fused through the optimized ECA module to enhance the correlation between channels. Group normalization GroupNorm is used to normalize the fused features to improve the consistency and stability of feature expression. Finally, residual connections are used to retain the original input features, and a learnable weighting coefficient is introduced to dynamically regulate the residual path, further improving the flexibility of the model.

[0064] The ED-VSS Layers significantly reduces the computational complexity and the number of parameters of the module by block processing the features. For the same input feature X(1, 256, 256, 256), the original VSS Block module requires 13.17 GFLOPs of floating-point operation counts (FLOPs) and 200.2K parameters, while the ED-VSS Layers only requires 3.52 GFLOPs of FLOPs (a 3.7-fold decrease compared to the original), and the number of parameters is only 13.2k (a 15.2-fold decrease compared to the original).

[0065] Step 3: Input the images in the industrial scenario image dataset into the LM-U3Plus. Use the hybrid-channel convolutional RFC module to extract features from the input low-dimensional and high-dimensional features respectively; use the feature fusion module MFF to fuse the extracted features, and use the ED-VSS Layers module to perform block processing on the fused features and output, obtaining the segmented images.

[0066] To verify the method of the present invention, the present invention uses the mean intersection over union (Miou), accuracy, and the number of parameters (Params) as evaluation metrics to evaluate the performance of the LM-U3Plus in the image segmentation task.

[0067] (1) Mean intersection over union (Miou).

[0068] The mean intersection over union is a commonly used accuracy metric in semantic segmentation. It calculates the ratio of the intersection to the union between the ground truth label and the predicted result, and then takes the mean of the IoU for all classes. The value of Miou ranges from 0 to 1, and the closer it is to 1, the higher the similarity between the predicted result and the ground truth label. Its calculation formula is Equation (1).

[0069]

[0070] Among them, n is the total number of classes, and IoU i is the intersection over union of the i-th class. The calculation formula of IoU i is Equation (2), where N is the intersection area between the ground truth label and the predicted result, and M is the union area between the ground truth label and the predicted result.

[0071]

[0072] (2) Accuracy.

[0073] Accuracy is one of the most commonly used evaluation metrics in classification problems. In image segmentation, it represents the ratio of the number of correctly classified pixels to the total number of pixels. Its calculation formula is Equation (3):

[0074]

[0075] Among them, X is the total number of correctly classified pixels, and Y is the total number of pixels. When calculating specifically, all predicted values can be traversed, and the predicted values are compared with the label values (true values) of the dataset. When the two are equal, the count is 1, and then the sum is divided by the total number of pixels.

[0076] (3) Parameter Count.

[0077] The parameter count is one of the important indicators to measure the complexity of the model. Especially in the field of deep learning, it represents the total number of learnable parameters in the model and is usually associated with the storage space requirements and computational complexity of the model. In the semantic segmentation task, the parameter count directly affects the inference speed and resource occupancy of the model. The total parameter count is the sum of the parameter counts of different layers. The calculation method for the convolutional layer is shown in Equation (4):

[0078] Params = (C in × K h × K w × C out ) + Bias (4);

[0079] Among them, C in is the number of input channels, K h and K w are the height and width of the convolutional kernel, C out is the number of output channels, and Bias is the parameter count of the bias term (usually equal to C out ). The calculation method for the fully connected layer is shown in Equation (5):

[0080] Params = (N in × N out ) + Bias (5);

[0081] Among them, N in is the number of input units, N out is the number of output units, and Bias is the parameter count of the bias term (usually equal to N out ). If depthwise separable convolutions are used, the calculation methods for the Depthwise part (convolving each channel separately) and the Pointwise part (1×1 convolution) are shown in Equation (6) and Equation (7) respectively:

[0082] Params = C in × K h × K w (6);

[0083] Params = C in × C out (7);

[0084] The total number of parameters of the model is the sum of the number of parameters of all layers, and its calculation method is shown in Equation (8):

[0085]

[0086] where L is the total number of layers of the network, and Params l is the number of parameters of the l-th layer of the network.

[0087] The size of the number of parameters reflects the complexity of the model. A smaller number of parameters usually means a lighter model, while a larger number of parameters may bring higher expressive power but also increase the computational overhead.

[0088] By combining the number of parameters with other performance metrics (such as Miou and Accuracy) for analysis, it can help evaluate the trade-off of the model, that is, how to reduce the complexity of the model while maintaining the accuracy of the model.

[0089] Through a series of experiments, the performance of multiple network models was compared, and the specific values are shown in Table 1.

[0090] Table 1 Index Comparison Table

[0091] ModelName Params Accuracy Miou DANet 4948,7027 97.43 91.36 VM-UNet 4427,3883 96.89 89.32 Deeplabv3 4034,7555 96.59 89.44 UNet 3105,4356 95.99 86.17 U3Plus 3001,3793 97.18 89.75 SegNet 2944,4739 96.8 88.53 PspNet 237,3955 94.15 81.13

[0092] Among these models, the U3Plus network is the most balanced. With a parameter number of 30.01M, it achieves a good segmentation effect, and its Accuracy and Miou reach 97.18 and 89.75 respectively. Compared with DANet, VM-UNet, Deeplabv3, and UNet, the U3Plus network has fewer parameters. Compared with SegNet and PspNet, the U3Plus network has higher accuracy. Therefore, the present invention uses U3Plus as the basic model.

[0093] LM-U3Plus is optimized by resetting the model structure (X), introducing the AFC (Y), MFF (Z), and ED-VSS Layers (H) modules. The experimental results of the gradual optimization are shown in Table 2.

[0094] Table 2 Optimization Model Effect

[0095] ModelName Params Accuracy Miou U3Plus 3001,3793 97.18 89.75 X 734,6495 97.25 90.30 Y 731,5049 97.39 90.82 Z 797,0921 97.62 91.65 H 770,9945 97.83 92.55

[0096] As can be seen from Table 2, by compressing the model structure and improving the feature distribution, while reducing the number of model parameters, the original accuracy is maintained. The introduction of the AFC, MFF, and ED-VSS Layers modules enables LM-U3Plus to maintain light weight while significantly improving the accuracy, and its Miou has increased by nearly 3%.

[0097] The comparison of the indicators of LM-U3Plus and other image segmentation models is shown in Table 3.

[0098] Table 3 Comparison of performance indicators of multiple models

[0099]

[0100] Experimental results show that LM-U3Plus not only has a lower number of parameters, but also has a significant performance advantage. Compared with the original basic model U3Plus, the number of parameters has dropped by nearly 75%, and the accuracy of the model has increased by 3% on Miou. This is a huge improvement for industrial segmentation, especially for some industrial scenarios that require pixel-level fine segmentation. Compared with the DANet model, the number of parameters has dropped by nearly 84%, while the accuracy has exceeded DANet. Compared with the lightweight model PspNet, the accuracy and Miou have increased by 3.68% and 11.42% respectively, while the number of parameters has only increased by 5.33M. This enables the model to meet the requirements of lightweight and high precision in industrial scenarios at the same time.

[0101] The optimized model is applied to real industrial scenes to accurately segment the belt and material flow in the logistics transportation process. It mainly focuses on three types of images: empty belt segmentation, normal belt and material flow segmentation, and special situation (lighting, shadow, smoke) segmentation. The partial segmentation effects of multiple models are compared, such as Figure 12 , Figure 13 and Figure 14 As shown, from Figure 12 , Figure 13 and Figure 14 It can be seen that whether it is empty belt image segmentation, normal image segmentation, or special image segmentation, LM-U3Plus has demonstrated good segmentation performance. It can complete accurate segmentation for different types of images. The edges of the segmentation map are clear and the boundary segmentation is accurate, which fully meets the requirements of industrial segmentation.

[0102] The present invention also provides an image segmentation system, comprising:

[0103] Image acquisition module, used to acquire industrial scene image datasets;

[0104] The model building module is used to compress the U3Plus network model into a four-layer network architecture based on the U3Plus network model, and sequentially add the mixed channel convolution RFC module, the feature fusion module MFF and the ED-VSSLayers module to the compressed network architecture to obtain the image segmentation model LM-U3Plus;

[0105] An image segmentation module is used to input the images in the industrial scene image dataset into LM-U3Plus, and use the hybrid-channel convolutional RFC module to extract features from the input low-dimensional features and high-dimensional features respectively; use the feature fusion module MFF to fuse the extracted features, and use the ED-VSS Layers module to perform block processing on the fused features and output, so as to obtain the segmented image.

[0106] The present invention also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the image segmentation method.

[0107] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded by a processor to execute the image segmentation method.

[0108] The above embodiments are only preferred specific embodiments of the present invention, and the protection scope of the present invention is not limited thereto. Any simple changes or equivalent replacements of technical solutions that can be obviously obtained by those skilled in the art within the technical scope disclosed by the present invention all belong to the protection scope of the present invention.

Claims

1. An image segmentation method, characterized in that: The steps include: Obtain industrial scene image datasets; Taking the U3Plus network model as the basic model, U3Plus is compressed into a four-layer network architecture. The mixed channel convolution RFC module, feature fusion module MFF and ED-VSS Layers module are added to the compressed network architecture in sequence to obtain the image segmentation model LM-U3Plus. The images in the industrial scene image dataset are input into LM-U3Plus, and the mixed channel convolution RFC module is used to extract the input low-dimensional features and high-dimensional features respectively; the extracted features are fused using the feature fusion module MFF, and the fused features are divided into blocks and output using the ED-VSS Layers module to obtain the segmented image.

2. The image segmentation method according to claim 1, characterized in that: The hybrid channel convolution RFC module includes multi-channel convolution and depth separation convolution. The multi-channel convolution is used to extract low-dimensional features, and the depth separation convolution is used to extract high-dimensional features.

3. The image segmentation method according to claim 2, characterized in that: A channel attention mechanism ECA module is added after the multi-channel convolution and depth separation convolution of the RFC module. The channel attention mechanism ECA module performs multi-scale fusion on the extracted features and performs layer-by-layer channel weighting operations on each feature of different scales.

4. The image segmentation method according to claim 3, characterized in that: The feature fusion module MFF is used to fuse the extracted features. Specifically, MFF uses adaptive weights to fuse and splice multi-level features after layer-by-layer channel weighting operations, and then performs feature fusion on the spliced ​​features through 1×1 and 3×3 convolutions.

5. The image segmentation method according to claim 1, characterized in that: The ED-VSS Layers module is used to perform block processing on the fused features. Specifically, ED-VSS Layers divides the fused features into four parts according to the channel dimension, and each part is subjected to feature extraction through VSS Block. The four extracted features are then fused through the ECA module, and the fused features are normalized using GroupNorm. Finally, the residual connection is used to retain the original input features, and a weighting coefficient is introduced to dynamically regulate the residual path to obtain the segmented image.

6. An image segmentation system, characterized in that: include: Image acquisition module, used to acquire industrial scene image datasets; The model building module is used to compress the U3Plus network model into a four-layer network architecture based on the U3Plus network model, and sequentially add the mixed channel convolution RFC module, the feature fusion module MFF and the ED-VSS Layers module to the compressed network architecture to obtain the image segmentation model LM-U3Plus; The image segmentation module is used to input the images in the industrial scene image dataset into LM-U3Plus, and use the mixed channel convolution RFC module to extract the input low-dimensional features and high-dimensional features respectively; use the feature fusion module MFF to fuse the extracted features, and use the ED-VSS Layers module to block process and output the fused features to obtain the segmented image.

7. A computer device, characterized in that: It comprises a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute the image segmentation method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor to execute the image segmentation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Joint segmentation method for optic cups and optic disks in fundus medical images

    CN112598650A

  • Lung X-ray image segmentation method and system, computer equipment and storage medium

    CN112651979A

  • U-Net-based blood vessel image segmentation method, device and equipment

    CN113205524A

  • Pathological cell image segmentation method and system based on adaptive fusion module and cross-stage AU-Net network

    CN116128890A

  • HL-UNet image segmentation model and heart dynamic magnetic resonance imaging segmentation method

    CN118470036A