Image segmentation method, device and medium

By adopting a twin network of single-stage encoder decoder and image extraction guide feature module in image segmentation, the problems of boundary distortion and model complexity in the prior art are solved, and more efficient image segmentation and feature extraction are achieved.

CN119205835BActive Publication Date: 2025-05-13SHENYANG URBAN CONSTR COLLEGE +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411437795.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-05-13
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

In the prior art, the cutout frame based on deep learning focuses too much on the extraction of foreground features, ignores the utilization of background features, resulting in boundary distortion, and the model architecture is complex, the number of parameters is large, and the training time is long.

Method used

The single-stage encoder decoder and the twin network of the image extraction guide feature module are used for image segmentation. Features are extracted by combining deep separation convolution and partial convolution, and the ECA module is used to extract channel attention features as upsampled guidance information.

Benefits of technology

It reduces the complexity of the model, improves the running speed of the model, enhances feature extraction capabilities, improves the recognition accuracy of image edges and details, and improves the accuracy of segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119205835B_ABST
    Figure CN119205835B_ABST
Patent Text Reader

Abstract

The present invention provides an image segmentation method, device and medium, the method comprising: obtaining an image to be processed; inputting the image to be processed into a trained twin network to obtain a segmented foreground image and background image; wherein the twin network at least comprises a single-stage encoder-decoder and an image extraction guide feature module. The trained twin network in the method comprises a single-stage encoder-decoder and an image extraction guide feature module, which improves the traditional twin network, wherein a single-stage encoder-decoder is used to reduce the model complexity and increase the model running speed, and an image extraction guide feature module is used to extract key feature information from the image to be processed, suppressing irrelevant information, thereby improving the distinction between the foreground and the background, and further improving the accuracy of segmentation. In short, the method can more effectively capture subtle feature differences in the image, not only enhancing the ability of feature extraction, but also improving the recognition accuracy of image edges and details.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically provides an image segmentation method, device and medium. Background Art

[0002] Matting is an image processing method that requires the foreground target to be accurately extracted from the composite image. The deep learning-based matting framework in the prior art focuses too much on the extraction of foreground features and ignores the use of background features, resulting in boundary distortion. For example, the Late Fusion Matting method generates probability maps of the foreground and background in the middle layer of the network, and then merges them into the final foreground target image. Although it utilizes background information, it does not fully utilize the significant features between the foreground and background. Secondly, the number of parameters in the model architecture is too large, including the use of GAN as a generative model training. For example, HATT and HaMaGAN both use the GAN training mode, which undoubtedly increases the model parameters and increases the time required for training.

[0003] Therefore, the field needs a new image segmentation solution to solve the above problems. Summary of the invention

[0004] In order to overcome the above-mentioned defects, the present invention is proposed to provide a solution or at least partially solve the problem of overly complex model architecture design and excessive number of model parameters.

[0005] In the first aspect, the present invention provides an image segmentation method, comprising: obtaining an image to be processed; inputting the image to be processed into a trained twin network to obtain a segmented foreground image and a background image; wherein the trained twin network includes at least a single-stage encoder-decoder and an image extraction guided feature module.

[0006] In a technical solution of the above-mentioned image segmentation method, the process of inputting the image to be processed into the trained twin network to obtain the segmented foreground image and background image at least includes: inputting the image to be processed into the image extraction guide feature module to obtain up-sampled guidance information; acquiring down-sampled features and up-sampled features based on the single-stage encoder-decoder; inputting the up-sampled features and the up-sampled guidance information into the single-stage encoder-decoder to obtain the segmented foreground image and background image.

[0007] In a technical solution of the above-mentioned image segmentation method, the image extraction guidance feature module includes at least: an ECA module; the process of inputting the image to be processed into the image extraction guidance feature module, and obtaining the upsampled guidance information includes at least: using the channel attention feature extracted based on the ECA module as the upsampled guidance information.

[0008] In a technical solution of the above-mentioned image segmentation method, the single-stage encoder-decoder includes an encoder, and the process of obtaining downsampling features based on the single-stage encoder-decoder includes: obtaining the downsampling features of the image to be processed based on the encoder, wherein a combination of depth-separable convolution and partial convolution is adopted when obtaining the downsampling features.

[0009] In a technical solution of the above-mentioned image segmentation method, the single-stage encoder-decoder includes a decoder, and the process of obtaining upsampling features based on the single-stage encoder-decoder includes: weighting the upsampled guidance information to obtain the weight of the corresponding guidance information; fusing the downsampled features, the upsampled guidance information and the weights and inputting them into the two decoders to obtain a first foreground upsampled feature and a first background upsampled feature, respectively.

[0010] In a technical solution of the above-mentioned image segmentation method, the decoder includes: a first residual module; the single-stage encoder-decoder includes a decoder, and the process of obtaining upsampling features based on the single-stage encoder-decoder also includes: extracting first residual features based on the first residual module; fusing the first residual features with the first foreground upsampling features and the first background upsampling features respectively to obtain second foreground upsampling features and second background upsampling features; and performing feature extraction and separation on the second foreground upsampling features and the second background upsampling features.

[0011] In a technical solution of the above-mentioned image segmentation method, the single-stage encoder-decoder also includes a dilated spatial convolution pooling pyramid pooling module; after obtaining the down-sampled features of the image to be processed based on the encoder, it includes: inputting the down-sampled features into the dilated spatial convolution pooling pyramid pooling module, and fusing them with the weights of the up-sampled guidance information as the input of the decoder.

[0012] In a second aspect, an electronic device includes a processor and a storage device, wherein the storage device is suitable for storing multiple program codes, and the program codes are suitable for being loaded and run by the processor to execute the image segmentation method of any one of the technical solutions of the above-mentioned image segmentation method.

[0013] In a third aspect, the present invention provides a computer-readable storage medium storing a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute the image segmentation method of any one of the technical solutions of the above-mentioned image segmentation method.

[0014] The above one or more technical solutions of the present invention have at least one or more of the following beneficial effects:

[0015] In the technical scheme for implementing the present invention, the present invention provides an image segmentation method, including: obtaining an image to be processed; inputting the image to be processed into a trained twin network to obtain a segmented foreground image and background image; wherein the twin network at least includes a single-stage encoder-decoder and an image extraction guide feature module. Compared with the prior art, the beneficial effects of the image segmentation method provided by the present invention are as follows: the trained twin network in the scheme includes a single-stage encoder-decoder and an image extraction guide feature module, which improves the traditional twin network, wherein a single-stage encoder-decoder is used to reduce the model complexity and increase the model running speed, and an image extraction guide feature module is used to extract key feature information from the image to be processed, suppressing irrelevant information, thereby improving the distinction between the foreground and the background, and further improving the accuracy of segmentation. This method can more effectively capture subtle feature differences in images, not only enhancing the ability of feature extraction, but also improving the recognition accuracy of image edges and details. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The disclosure of the present invention will become more easily understood with reference to the accompanying drawings. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present invention. In addition, similar numbers in the figures are used to represent similar components, among which:

[0017] Figure 1 is a schematic flow chart of main steps of an image segmentation method according to an embodiment of the present invention;

[0018] Figure 2 It is a flowchart of inputting an image to be processed into a trained twin network to obtain a segmented foreground image and background image according to an embodiment of the present invention. DETAILED DESCRIPTION

[0019] Some embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0020] Embodiment 1

[0021] See attached Figure 1 , Figure 1 FIG. 1 is a flow chart of the main steps of an image segmentation method according to an embodiment of the present invention. Figure 1 As shown, the image segmentation method in the embodiment of the present invention mainly includes the following steps S1 to S2.

[0022] Step S1, obtaining an image to be processed;

[0023] In this embodiment, the image to be processed may be various types of pictures, such as human portraits, animal images, etc.

[0024] Step S2: input the image to be processed into the trained twin network to obtain segmented foreground image and background image; wherein the trained twin network includes at least a single-stage encoder-decoder and an image extraction guide feature module.

[0025] In this embodiment, the single-stage encoder-decoder is a neural network architecture for image segmentation, which completes the encoding, decoding and foreground and background segmentation tasks of image features in one continuous processing stage. The single-stage encoder-decoder has the characteristics of efficient and fast image processing, and can complete the image segmentation task in one stage. The main function of the image extraction guided feature module is to extract key feature information from the input image to be processed. This feature information can guide the twin network to better segment the foreground and background. It may extract representative features in the image, such as color, texture, shape, etc., through specific algorithms and techniques, so that the twin network can more accurately identify the foreground and background areas.

[0026] In one embodiment, Figure 2 As shown, step S2, inputting the image to be processed into the trained twin network to obtain the segmented foreground image and background image, comprises at least:

[0027] Step S21, inputting the image to be processed into the image extraction guidance feature module to obtain upsampled guidance information;

[0028] In this embodiment, the guidance information may include color distribution, texture features, object edge clues, size, etc. of the image. Such information will help the decoder to more accurately restore the details of the image when upsampling, especially to make the distinction between foreground and background clearer.

[0029] Step S22, obtaining down-sampling features and up-sampling features based on the single-stage encoder-decoder;

[0030] In this embodiment, the encoder part in the single-stage encoder-decoder first downsamples the image to be processed. The downsampled features are obtained by gradually reducing the image and contain abstract information of different levels of the image. The upsampled features are obtained by gradually enlarging and restoring the features obtained in the downsampling process.

[0031] Step S23: input the up-sampled features and the up-sampled guidance information into the single-stage encoder-decoder to obtain the segmented foreground image and background image.

[0032] In this embodiment, the guidance information extracts key feature information from the original image and inputs it into the single-stage encoder-decoder as the guidance information for upsampling. This can help the model better understand the content and structure of the original image, so that the details and features of the image can be more accurately restored when upsampling, and the fitting accuracy of the model is improved. In addition, the edges and details of the image are crucial for accurately segmenting the foreground and background. The guidance information can help the model better capture the subtle feature differences in the image when upsampling, especially at the edges and details of the image. This helps to improve the model's recognition accuracy of the edges and details of the image, thereby obtaining more accurate foreground and background images.

[0033] In one embodiment, the image extraction guidance feature module includes at least: an ECA module; the process of inputting the image to be processed into the image extraction guidance feature module, and obtaining the upsampled guidance information includes at least: using the channel attention feature extracted based on the ECA module as the upsampled guidance information.

[0034] In this embodiment, the ECA module (Efficient Channel Attention module) is used to extract the channel attention features of the image. The channel attention mechanism allows the model to pay more attention to the important information carried by different channels in the image. For example, some color channels may play a key role in distinguishing the foreground and background. The channel attention features extracted based on the ECA module are used as guiding information for upsampling, which means that these channel attention features will play a guiding role in the subsequent image segmentation process after certain processing (upsampling operation). Upsampling can match the resolution of the guidance information with the resolution of the image to be processed or other related feature maps, so as to better integrate and guide the segmentation process.

[0035] In one embodiment, the single-stage encoder-decoder includes an encoder, and the process of obtaining down-sampling features based on the single-stage encoder-decoder includes:

[0036] Based on the encoder, down-sampled features of the image to be processed are obtained, wherein a combination of depthwise separable convolution and partial convolution is adopted when obtaining the down-sampled features.

[0037] In this embodiment, when acquiring downsampled features, the module uses a combination of deep separable convolution and partial convolution to extract weights. The weights extracted in this way can better represent the importance of the downsampled features and provide more valuable information for subsequent processing. Specifically, a deep convolution operation is performed on the input features to extract the spatial features of each channel. This step can preliminarily reduce the amount of calculation and extract local spatial information. The output of the deep convolution is used as the input of the partial convolution. During the partial convolution process, for the pixels at each position, a convolution operation is performed according to the number of valid pixels, and the number of valid pixels is updated. This can handle missing or irregular areas in the image and make the model more robust. The output of the partial convolution is subjected to a point-by-point convolution operation, and the spatial features extracted by the deep convolution are channel-fused to achieve cross-channel information interaction.

[0038] In one embodiment, the single-stage encoder-decoder includes a decoder, and the process of obtaining up-sampling features based on the single-stage encoder-decoder includes: weighting the up-sampled guidance information to obtain the weight of the corresponding guidance information; fusing the down-sampled features, the up-sampled guidance information and the weights and inputting them into the two decoders to obtain a first foreground up-sampling feature and a first background up-sampling feature, respectively.

[0039] In this embodiment, the guidance information usually includes specific clues related to the foreground and background, such as size, color distribution, texture features, object edge information, etc. This information is very important for accurately distinguishing the foreground and background in the process of restoring the image. For example, the guidance information is feature information extracted by size. For the fine details of the foreground object, the feature information of smaller size may be more critical, so it can be given a higher weight in weighted processing. For the overall structure of the background, the feature information of larger size may be more valuable, and accordingly there will be different weight distributions. In this way, more accurate adjustments can be made according to the contribution of different size features to the upsampling of the foreground and background. Using the guidance information as the weight of the upsampling layered data means that in the upsampling process of each layer, the degree of restoration of different parts is adjusted according to the importance of the guidance information. For example, if the guidance information of a certain area indicates that it is likely to be the foreground, then in the upsampling process of the area, more features related to the foreground will be given a greater weight, so that the restored result is more inclined to the features of the foreground. In this way, the upsampling result can more accurately reflect the true situation of the foreground and background in the target image, and improve the accuracy of image segmentation.

[0040] In one embodiment, the decoder includes: a first residual module; the single-stage encoder-decoder includes a decoder, and the process of obtaining upsampling features based on the single-stage encoder-decoder also includes: extracting first residual features based on the first residual module; fusing the first residual features with the first foreground upsampling features and the first background upsampling features respectively to obtain second foreground upsampling features and second background upsampling features; and performing feature extraction and separation on the second foreground upsampling features and the second background upsampling features.

[0041] In this embodiment, when obtaining the upsampling feature, the first residual feature is first extracted based on the first residual module. The residual feature usually reflects the difference or change information in the image compared with the original feature. In this process, the first residual module can extract residual information that can supplement and enhance the existing features from the input data through a specific operation or neural network structure. Specifically, the first residual module is a 1×1 convolution module added to the upsampling part of ResBlock. The function of ResBlock itself is to solve the gradient vanishing problem in deep neural networks by adding input and output to learn residual mapping. After adding a 1×1 convolution module to the upsampling part, this module can extract residual features. The first residual feature is fused with the first foreground upsampling feature and the first background upsampling feature respectively. This fusion process can combine the information in the residual feature with the existing foreground and background upsampling features, thereby enriching the feature representation of the foreground and background.

[0042] Furthermore, ResBlock is also used in the downsampling process.

[0043] Furthermore, the second foreground upsampled features and the second background upsampled features are extracted and separated, and the loss is calculated. The loss function has two functions: one is to increase the feature distance between the foreground and the background, and the other is to make the distance between the output image and the actual image close, so that the elements with feature confusion can be effectively separated.

[0044] Specifically, a combination of MSE, MAD and SSIM losses can be used. The difference value and similarity between the target image and the real image are used to guide and constrain the model. The final total loss function design is shown in formula (1):

[0045] L total =L ssim +L mse +L mad (1)

[0046] Where L mssim , L mse , L mad The definition is shown in formula (2):

[0047]

[0048] In the above formula, p represents the pixel index, N represents the number of pixels in the unknown area, and the symbol μ x and σ x Represent the mean and standard deviation of image x respectively. The cutout model is evaluated at three levels: absolute distance, relative distance, and image similarity between image x and image y, and they are linearly combined to guide model training. The design of this loss function not only enables the model to accurately capture the subtle differences between foreground and background at the pixel level, but also ensures the overall similarity between the generated image and the real image.

[0049] In one embodiment, the single-stage encoder-decoder also includes a dilated spatial convolution pooling pyramid pooling module; after obtaining the down-sampled features of the image to be processed based on the encoder, it includes: inputting the down-sampled features into the dilated spatial convolution pooling pyramid pooling module, and fusing them with the weights of the up-sampled guidance information as the input of the decoder.

[0050] In this embodiment, the down-sampled features are input into the hollow space convolutional pooling pyramid pooling module (ASPP module) and fused with the up-sampled guide information weight as the input of the decoder, which can significantly improve the performance of the single-stage encoder-decoder. This method can make full use of multi-scale features and guide information, so that the model can better adapt to different image contents and task requirements, thereby improving the accuracy and robustness of image segmentation. Compared with the multi-stage method, the single-stage encoder-decoder combined with the ASPP module and weight fusion can complete the image segmentation task in one stage, simplifying the structure of the model. This helps to reduce the amount of calculation and the number of parameters, and improve the operating efficiency and scalability of the model.

[0051] In one embodiment, in the process of acquiring up-sampling features, a depthwise separable convolution processing method is adopted.

[0052] In this embodiment, in the process of obtaining up-sampled features, deep separable convolution can be used to enhance the expressiveness of features. By performing deep separable convolution operations on low-resolution features, richer feature information can be extracted, providing a better basis for upsampling. In addition, during the upsampling process, low-resolution features need to be restored to high resolution, which usually requires a lot of calculations. Deep separable convolution can reduce the amount of calculation while maintaining a good upsampling effect and improving the performance of the model.

[0053] Embodiment 2

[0054] The present invention also provides an electronic device. In an embodiment of a device according to the present invention, the device includes a processor and a storage device, the storage device can be configured to store a program for executing the image segmentation method of the above method embodiment, and the processor can be configured to execute the program in the storage device, which includes but is not limited to the program for executing the image segmentation method of the above method embodiment. For ease of explanation, only the parts related to the embodiment of the present invention are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present invention. The control device can be a control device device formed by various electronic devices.

[0055] Embodiment 3

[0056] The present invention also provides a computer-readable storage medium. In a computer-readable storage medium embodiment according to the present invention, the computer-readable storage medium can be configured to store a program for executing the image segmentation method of the above method embodiment, and the program can be loaded and run by a processor to implement the above image segmentation method. For ease of explanation, only the parts related to the embodiment of the present invention are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present invention. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present invention is a non-temporary computer-readable storage medium.

[0057] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the original technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. An image segmentation method, characterized in that: include: Get the image to be processed; Input the image to be processed into the trained twin network to obtain the segmented foreground image and background image; The trained twin network at least includes a single-stage encoder-decoder and an image extraction guided feature module; Inputting the image to be processed into the trained twin network to obtain the segmented foreground image and background image specifically includes: inputting the image to be processed into the image extraction guide feature module to obtain up-sampled guide information; obtaining down-sampled features and up-sampled features based on the single-stage encoder-decoder; inputting the up-sampled features and the up-sampled guide information into the single-stage encoder-decoder to obtain the segmented foreground image and background image; The image extraction guide feature module at least includes: an ECA module; the process of inputting the image to be processed into the image extraction guide feature module to obtain the up-sampled guide information at least includes: using the channel attention feature extracted based on the ECA module as the up-sampled guide information; The single-stage encoder-decoder comprises an encoder and a decoder, and the process of obtaining down-sampling features based on the single-stage encoder-decoder comprises: obtaining down-sampling features of the image to be processed based on the encoder; The process of obtaining up-sampled features based on the single-stage encoder-decoder includes: weighting the up-sampled guidance information to obtain the weight of the corresponding guidance information; inputting the down-sampled features and the up-sampled guidance information after weight fusion into the two decoders respectively to obtain the first foreground up-sampled features and the first background up-sampled features respectively.

2. The method according to claim 1, characterized in that When obtaining down-sampling features, a combination of depth-wise separable convolution and partial convolution is used.

3. The method according to claim 1, characterized in that The decoder includes: a first residual module; the process of obtaining up-sampling features based on the single-stage encoder-decoder also includes: Extracting a first residual feature based on the first residual module; The first residual feature is respectively combined with the first foreground upsampling feature and the first background upsampling feature to obtain a second foreground upsampling feature and a second background upsampling feature; Perform feature extraction and separation on the second foreground upsampled features and the second background upsampled features.

4. The method according to claim 1, characterized in that: The single-stage encoder-decoder also includes a dilated spatial convolution pooling pyramid pooling module; after obtaining the down-sampled features of the image to be processed based on the encoder, the method includes: inputting the down-sampled features into the dilated spatial convolution pooling pyramid pooling module as the input of one of the decoders, and performing up-sampled guidance information after weight fusion as the input of another decoder.

5. The method according to claim 1, characterized in that In the process of obtaining upsampling features, depth-wise separable convolution processing is adopted.

6. An electronic device comprising a processor and a storage device, the storage device being adapted to store a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to perform the image segmentation method according to any one of claims 1 to 5.

7. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to perform the image segmentation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image recognition method based on multi-scale feature fusion, computer device and computer readable storage medium

    CN117830703A

  • Cell nucleus segmentation method of cascade coding segmentation network based on large model guidance

    CN118366153A