Frequency domain guided superpixel segmentation method and system

By using a frequency-domain-guided superpixel segmentation method that integrates spatial and frequency domain features, the problem of boundary preservation in complex scenes is solved, and a more efficient superpixel segmentation effect is achieved.

CN117079280BActive Publication Date: 2025-11-11HAINAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311120581.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-31
Publication Date
2025-11-11
Estimated Expiration
2043-08-31

AI Technical Summary

Technical Problem

Existing superpixel segmentation methods struggle to preserve detailed boundaries in complex scenes and ignore unavoidable environmental constraints in practical applications, resulting in performance limitations.

Method used

A frequency-domain guided superpixel segmentation method is adopted, which generates superpixels with sharp boundaries by fusing spatial and frequency domain depth features and using a frequency domain information extractor and dense hybrid dilated convolutional blocks.

Benefits of technology

It improves the performance of superpixel segmentation in complex scenes, maintains clearer boundaries and better semantic information, and enhances the segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117079280B_ABST
    Figure CN117079280B_ABST
Patent Text Reader

Abstract

This invention discloses a frequency-domain guided superpixel segmentation method and system, belonging to the field of superpixel segmentation technology. The method includes acquiring an image to be segmented, inputting the image into a preset superpixel segmentation network for processing, and generating a pixel-superpixel association map containing semantically aware superpixels. The superpixel segmentation network includes a frequency domain information extractor, a densely mixed dilated convolutional block, and an association implantation module. The frequency domain information extractor acquires a frequency domain feature map based on the image to be segmented; the densely mixed dilated convolutional block captures high-level semantic information of the image to be segmented and generates an output feature map; the association implantation module acquires a spatial domain feature map based on the output feature map, and fuses the spatial domain feature map and the frequency domain feature map to obtain depth features, which are then applied to the softmax function to obtain the superpixel segmentation result. This method can generate superpixels with sharp boundaries for complex scenes, solving the problem that superpixel segmentation ignores image degradation and uncertain environmental constraints in practical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of superpixel segmentation technology, and in particular to a frequency-domain guided superpixel segmentation method and system. Background Technology

[0002] The statements in this section merely refer to the background art related to this invention and do not necessarily constitute prior art.

[0003] Superpixel segmentation is a fundamental work that aims to reduce the size of the original elements by oversegmenting an image by grouping pixels with similar low-level attributes such as boundaries and color, thereby generating a large number of perceptually meaningful units. Benefiting from more efficient representation of image content, superpixels have been widely used in various computer vision tasks, such as image enhancement, image colorization, saliency detection, scene recognition, and semantic segmentation. In practical applications of various computer vision tasks, many photos are taken in complex scenes, which are often produced under conditions of blurred boundaries and low lighting due to unavoidable environmental and technical limitations. Generally, complex scenes mainly include: hazy scenes, rainy scenes, underwater scenes, and low-light scenes. How to solve the problem that existing image processing algorithms struggle to achieve satisfactory performance in these scenes remains a significant challenge.

[0004] In recent years, various superpixel segmentation methods have been proposed and widely applied. Achanta et al. proposed the Simple Linear Iterative Clustering (SLIC) method, which efficiently generates superpixels based on the k-means algorithm. This clustering-based method has been widely used in various superpixel-based applications due to its fast runtime and good performance. To address the difficulty of preserving detailed boundaries, Li et al. proposed Spatial Constrained Subspace Clustering (SCS). They treated superpixel generation as a subspace clustering problem and developed a locally constrained subspace clustering model that can solve the problems of boundary confusion and segmentation errors. However, this method is limited in performance because it uses hand-designed features instead of deep features. Hand-designed features are based on prior knowledge and assumptions and may not capture all relevant information.

[0005] Recently, Wang et al. proposed a deep learning-based superpixel segmentation method (AINet), in which they developed an Association Implantation Module (AIM) to embed grid features into neighboring pixels, allowing AINet to explicitly perceive the relationship between a pixel and its surrounding grid cells. They also proposed a boundary-aware loss to identify boundary pixels and improve boundary accuracy. However, existing deep learning-based superpixel segmentation methods are designed for high-quality natural images, neglecting the inevitable image degradation and uncertain environmental constraints in real-world applications, which may lead to performance limitations. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a frequency-domain guided superpixel segmentation method, system, electronic device, and computer-readable storage medium that generates superpixels with sharp boundaries for complex scenes by fusing spatial and frequency domain depth features.

[0007] In a first aspect, the present invention provides a superpixel segmentation method based on frequency domain guidance;

[0008] A frequency-domain guided superpixel segmentation method includes:

[0009] Obtain the image to be segmented;

[0010] The image to be segmented is input into a preset superpixel segmentation network for processing to generate a pixel-to-superpixel association map containing semantically aware superpixels;

[0011] The superpixel segmentation network includes a frequency domain information extractor, a densely mixed dilated convolutional block, and an association implantation module. The frequency domain information extractor is used to obtain a frequency domain feature map based on the image to be segmented. The densely mixed dilated convolutional block is used to capture high-level semantic information of the image to be segmented and generate an output feature map. The association implantation module is used to obtain a spatial domain feature map based on the output feature map, so as to fuse the spatial domain feature map and the frequency domain feature map to obtain depth features, and apply them to the softmax function to obtain a pixel-superpixel association mapping map containing semantically aware superpixels.

[0012] Furthermore, obtaining the frequency domain feature map based on the image to be segmented includes:

[0013] Generate a guide map based on the image to be segmented;

[0014] The image to be segmented is converted into Fourier coefficients by Fast Fourier Transform, and the Fourier coefficients are divided into real and imaginary parts.

[0015] Based on the real and imaginary parts, real and imaginary features are obtained, and the real and imaginary features are connected in the channel dimension to generate a bilateral mesh through a three-dimensional convolutional layer;

[0016] Based on the guided graph, query the bilateral grid and obtain the frequency domain feature map through slicing operations.

[0017] Preferably, obtaining the real part features and imaginary part features based on the real part and imaginary part includes:

[0018] The weights of the real and imaginary parts are adaptively adjusted in parallel by a channel attention module with skip connections, and then processed in blocks and stretched into a labeled embedding.

[0019] The labels corresponding to the real and imaginary parts are embedded and processed in parallel by a channel mixer to obtain the real and imaginary features respectively.

[0020] Furthermore, the frequency domain information extractor includes a channel attention module, a channel mixer, a three-dimensional convolutional layer, a convolutional layer, and a slicing layer;

[0021] The convolutional layer is used to generate a guide map based on the image to be segmented;

[0022] The channel attention module is used to adaptively adjust the real or imaginary parts of the Fourier coefficients and jump-connect them with the Fourier coefficients. The channel mixer then obtains the real or imaginary features. The Fourier coefficients are generated from the image to be segmented by a fast Fourier transform.

[0023] The three-dimensional convolutional layer is used to generate a two-sided mesh based on the real and imaginary features;

[0024] The slicing layer is used to query the bilateral grid through the guide graph and generate a frequency domain feature map through slicing operations.

[0025] Furthermore, the step of capturing high-level semantic information of the image to be segmented and generating an output feature map includes:

[0026] The input feature maps are fed into the first convolutional branch, the second convolutional branch, the third convolutional branch, and the global average pooling branch, respectively, to obtain the first sub-output feature map, the second sub-output feature map, the third sub-output feature map, and the fourth sub-output feature map;

[0027] The first sub-output feature map, the second sub-output feature map, the third sub-feature map and the fourth sub-feature map are concatenated and fused along the channel dimension to obtain the output feature map;

[0028] The input feature map is obtained by processing the image to be segmented through multiple convolutional layers.

[0029] Furthermore, the densely mixed dilated convolutional block includes a first convolutional branch, a second convolutional branch, a third convolutional branch, a global average pooling branch, and a Conv-BN-ReLU layer;

[0030] The first convolutional branch is used to obtain a first sub-output feature map based on the input feature map; the second convolutional branch is used to obtain a second sub-output feature map based on the input feature map; the third convolutional branch is used to obtain a third sub-output feature map based on the input feature map; the global average pooling branch is used to obtain a fourth sub-output feature map based on the input feature map; the Conv-BN-ReLU layer is used to fuse the first, second, third, and fourth sub-output feature maps concatenated along the channel dimension to obtain an output feature map.

[0031] Furthermore, during the training of the superpixel network, a loss function is used to optimize the superpixel network, and the loss function is expressed as:

[0032] L=Lsem+λLcompact+αLB

[0033] Where Lsem is the semantic loss, Lcompact is the reconstruction loss, LB is the boundary-aware loss, and λ and α are the weight parameters.

[0034] Secondly, the present invention provides a frequency-domain guided superpixel segmentation system;

[0035] A frequency-domain guided superpixel segmentation system includes:

[0036] The acquisition module is configured to acquire the image to be segmented.

[0037] The superpixel segmentation module is configured to: input the image to be segmented into a preset superpixel segmentation network for processing, so as to generate a pixel-superpixel association map containing semantically aware superpixels;

[0038] The superpixel segmentation network includes a frequency domain information extractor, a densely mixed dilated convolutional block, and an association implantation module. The frequency domain information extractor is used to obtain a frequency domain feature map based on the image to be segmented. The densely mixed dilated convolutional block is used to capture high-level semantic information of the image to be segmented and generate an output feature map. The association implantation module is used to obtain a spatial domain feature map based on the output feature map, so as to fuse the spatial domain feature map and the frequency domain feature map to obtain depth features, and apply them to the softmax function to obtain a pixel-superpixel association mapping map containing semantically aware superpixels.

[0039] Thirdly, the present invention provides an electronic device;

[0040] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps of the frequency domain-guided superpixel segmentation method described above.

[0041] Fourthly, the present invention provides a computer-readable storage medium;

[0042] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the steps of the frequency-domain guided superpixel segmentation method described above.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] 1. The technical solution provided by this invention proposes an improved frequency domain information extractor to extract frequency domain information with sharp boundary features in order to utilize the frequency domain information of an image. The design of this frequency domain extractor fully considers the high-frequency information in the image, which corresponds to the boundaries and details in the image, and can effectively improve the segmentation performance in complex scenes.

[0045] 2. The technical solution provided by this invention introduces a dense hybrid dilated convolution (DHAC) block to preserve the semantic information of superpixels in order to avoid over-sharpening features from damaging the semantic information of superpixels. This block can capture wider and deeper semantic information in the spatial domain, thereby enhancing the performance of the superpixel segmentation network.

[0046] 3. The technical solution provided by this invention proposes a frequency domain-guided superpixel segmentation network (FSNet), which generates superpixels with sharp boundaries for complex scenes by fusing spatial and frequency domain depth features. This feature fusion strategy helps improve the performance of superpixel segmentation, enabling the generated superpixels to better maintain the clarity of their boundaries. Attached Figure Description

[0047] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0048] Figure 1 This is a schematic diagram of the architecture of the superpixel segmentation network provided in an embodiment of the present invention;

[0049] Figure 2 This is a schematic diagram of the architecture of the frequency domain information extractor provided in an embodiment of the present invention;

[0050] Figure 3 A schematic diagram of the architecture of a densely dilated convolutional block provided in an embodiment of the present invention;

[0051] Figure 4 This is a schematic diagram illustrating a quantitative comparison between the method provided in this embodiment of the invention and other state-of-the-art methods, wherein (a) is a comparison diagram of boundary recall and precision, (b) is a comparison diagram of segmentation accuracy, and (c) is a comparison diagram of undersegmentation errors;

[0052] Figure 5 A qualitative comparison chart of the method provided in this embodiment of the invention with other state-of-the-art methods on the BSDS500 dataset;

[0053] Figure 6 A qualitative comparison chart of the method provided in this embodiment of the invention with other state-of-the-art methods on the Foggy Cityscapes dataset;

[0054] Figure 7 A qualitative comparison chart of the method provided in this embodiment of the invention with other state-of-the-art methods on the Rain Cityscapes dataset;

[0055] Figure 8 A qualitative comparison chart of the method provided in this embodiment of the invention with other state-of-the-art methods on the SUIM dataset;

[0056] Figure 9 A qualitative comparison chart of the method provided in this embodiment of the invention with other state-of-the-art methods on the LLRGBD dataset;

[0057] Figure 10 A comparison chart of ablation studies on the BSDS500 dataset between the method provided in this embodiment of the invention and other state-of-the-art methods;

[0058] Figure 11 A comparison chart of the inference efficiency of the method provided in this embodiment of the invention with other state-of-the-art methods in salient target detection;

[0059] Figure 12 The graph shows a quantitative comparison of the method provided in this embodiment of the invention with other state-of-the-art methods in the detection of salient targets. Detailed Implementation

[0060] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0061] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. Furthermore, it should be understood that the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0062] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0063] Example 1

[0064] Existing superpixel segmentation techniques suffer from difficulties in preserving detailed boundaries, and are designed for high-quality natural images, neglecting unavoidable environmental constraints in practical applications, resulting in limited performance in real-world applications. Therefore, this invention provides a frequency-domain guided superpixel segmentation method that extracts and fuses depth features of the input image in the frequency and spatial domains respectively to generate semantically perceptive superpixels with sharp boundaries.

[0065] Next, combined Figures 1-12 This embodiment discloses a frequency-domain guided superpixel segmentation method, which includes the following steps:

[0066] S1. Obtain the image to be segmented.

[0067] S2. Input the image to be segmented into a preset superpixel segmentation network for processing to generate a pixel-to-superpixel association map containing semantically aware superpixels.

[0068] Among them, such as Figure 1 As shown, the Superpixel Segmentation Network (FSNet) includes a Conv-BN-ReLU layer, a frequency domain information extractor, a densely mixed dilated convolutional block, a Deconv-ReLU layer, and an association implantation module. The frequency domain information extractor is used to obtain a frequency domain feature map based on the image to be segmented. The Conv-BN-ReLU layer is used to capture the color and semantic information of the image to be segmented, and then the densely mixed dilated convolutional block captures wider and deeper semantic information to preserve semantic information. Next, the association implantation module (AIM) is used to enable the model to explicitly capture the relationship between pixels and their surrounding grids, thereby further improving performance. Finally, by applying the depth features obtained by fusing the frequency domain feature image and the spatial domain feature image to the softmax function, a pixel-to-superpixel association map containing semantically aware superpixels is obtained.

[0069] Furthermore, the specific steps for processing the image to be segmented into a preset superpixel segmentation network to generate a pixel-to-superpixel association map containing semantically aware superpixels include:

[0070] S201. Input the image to be segmented into the frequency domain information extractor for processing to obtain the frequency domain feature map. Simultaneously execute S202.

[0071] The frequency domain information extractor includes a channel attention module, a channel mixer, a 3D convolutional layer, a convolutional layer, and a slice layer.

[0072] Specifically, such as Figure 2 As shown, the specific steps of inputting the image to be segmented into the frequency domain information extractor for processing include:

[0073] S2011. Generate a guide map based on the image to be segmented, I; specifically, input the image to be segmented, I, into two concatenated convolutional layers with kernel size of 1×1 to generate the guide map. Simultaneously execute S2012.

[0074] S2012. Convert the image I to be segmented into Fourier coefficients using Fast Fourier Transform, and divide the Fourier coefficients into real parts rea∈R. (C×H×W) and the imaginary part ima∈R (C×H×W) Where C, H, and W represent the number of channels, height, and width of the image to be segmented, respectively.

[0075] S2013. The image to be segmented is converted into Fourier coefficients using Fast Fourier Transform, and the Fourier coefficients are divided into real and imaginary parts. A channel attention module with skip connections is used to adaptively adjust the weights of the real and imaginary parts in parallel, and the data is processed in blocks and stretched into labeled embeddings. The labeled embeddings corresponding to the real and imaginary parts are then input into the corresponding channel mixers in parallel for processing to obtain real and imaginary features respectively. The real and imaginary features are then concatenated along the channel dimension, and a two-sided mesh is generated through a 3D convolutional layer.

[0076] Specifically, the adaptive adjustment of the weights of rea and ima by the channel attention module (CAM) with skip connections can be expressed in the following form:

[0077] rea′=rea+CA(rea) (1)

[0078] ima′=ima+CA(ima) (2)

[0079] Next, rea and ima are divided into smaller blocks and stretched into labeled embeddings, denoted as Trea∈R. (((H / P)×(W / P))×(C×P×P)) and Tima∈R (((H / P)×(W / P))×(C×P×P)) .

[0080] Subsequently, frequency domain information is extracted from the real part of Frea and the imaginary part of Fima using a pair of channel mixers to obtain the real part feature Orea and the imaginary part feature Oima, which are represented as:

[0081] T′=Conv(PR(Conv(PR(Conv(LN(T))))))+T) (3)

[0082] O=Cono(PR(FC(PR(FC(LN(T)))))+T′) (4)

[0083] Where O is the real feature Orea or the imaginary feature Oima, T is Trea or Tima, F is the input feature map, LN is layer normalization, Conv is a convolutional layer with a 1×1 kernel size, FC is a fully connected layer, and PR is the Parametric ReLU (PReLU) activation function.

[0084] Finally, the real feature Orea and the imaginary feature Oima are connected in the channel dimension and a two-sided mesh is generated through a three-dimensional convolutional layer.

[0085] S2014. After slicing, the guided graph is used to query the bilateral grid, and the frequency domain feature map is obtained through slicing operations. S202. The image to be segmented is processed by five concatenated Conv-BN-ReLU layers to extract the color and semantic information of the image to be segmented in the spatial domain and generate the input feature map; the input feature map is then processed by a densely mixed dilated convolutional block to obtain the output feature map.

[0086] The densely mixed dilated convolutional block includes three convolutional branches (i.e., the first convolutional branch, the second convolutional branch, and the third convolutional branch), a global average pooling branch, and a Conv-BN-ReLU layer.

[0087] Specifically, to capture high-level semantic information, an intuitive and simple approach is to stack more dilated convolutional layers to expand the receptive field. However, improperly set dilation rates can lead to the loss of local information, resulting in a "rasterization effect" in semantic segmentation tasks, and can also negatively impact the model's learning in other computer vision tasks, as dilated convolutions (or dilated convolutions) introduce "holes" (zero values). In the DHAC proposed in this embodiment, the dilation rate is designed following the rules proposed in Hybrid Dilated Convolution (HDC).

[0088] Specifically, assuming there are n convolutional layers with a kernel size of K×K, the hole ratio is set as follows:

[0089] That is, [r1,...,ri,...,rn].

[0090] The maximum distance between two non-zero values ​​is defined as:

[0091] M i =max[M i+1 -2r i M i+1 -2(M i+1 -r i ),r i (5)

[0092] Based on this definition, three convolutional layers are formed, with a kernel size of 3×3 and a porosity set to [1, 2, 5] as a group. The three convolutional branches in DHAC contain groups of 1, 2, and 3 such convolutional layers, respectively. The groups between different branches share weights to reduce the number of model parameters.

[0093] In addition, a global average pooling (GAP) branch is added in this embodiment to obtain global contextual information, thereby enabling the model to better understand the entire image.

[0094] Finally, the features extracted from all branches are concatenated along the channel dimension and fused by a "Conv-BN-ReLU" layer.

[0095] Furthermore, the input feature map is processed by a densely mixed dilated convolutional block to obtain the output feature map. The specific process is as follows:

[0096] (1) Input the input feature maps into the first convolution branch, the second convolution branch, the third convolution branch and the global average pooling branch in parallel to obtain the first sub-output feature map, the second sub-output feature map, the third sub-output feature map and the fourth sub-output feature map.

[0097] (2) The first sub-output feature map, the second sub-output feature map, the third sub-feature map and the fourth sub-feature map are concatenated and fused in the channel dimension to obtain the output feature map.

[0098] Specifically, assume F∈R C×H×W Given an input feature map, the output feature map U∈R contains rich high-level semantic information. C ×H×W It can be expressed as follows:

[0099] F′=Concat(G1(F),G2(G1(F)),G3(G2(G1(F)))) (6)

[0100] U=ReLU(BN(Conv(Concat(GAP(F),F′)))) (7)

[0101] Where G represents dilated convolutional groups, Concat is a connection operation, Conv is a convolutional layer with a kernel size of 1×1, BN is batch normalization, and ReLU is the ReLU activation function, as shown above.

[0102] The Conv-BN-ReLU layer skips connections with the corresponding Deconv-BN-ReLU layer, aiming to combine lower-level (low-resolution) and higher-level (high-resolution) feature maps. This helps the network retain detailed information of the original input while allowing the network to utilize the abstract representation of high-level features; it helps to better capture information at different levels, thereby improving the network's performance and expressiveness.

[0103] The output feature map is then input into four concatenated Deconv-BN-RELU layers. The Deconv-BN-RELU layers perform deconvolution operations on the output feature map to expand it, thereby increasing the spatial resolution so that feature representation and processing can be performed in a higher resolution feature space. Finally, the output and input are correlated and implanted into the module.

[0104] The output feature map is then fed into two cascaded convolutional layers. These layers perform convolution operations on the output feature map to capture more abstract feature information, and the final output is fed into the association implantation module. S203: The association implantation module processes the input feature map to obtain pixel-level embeddings (i.e., spatial domain feature maps) with superpixel-level context. This output processes the input feature map to retain information from each pixel while also incorporating information from surrounding superpixels, thus providing a broader context.

[0105] S204. The spatial domain feature map and the frequency domain feature map are fused to obtain deep features, which are then applied to the softmax function to obtain a pixel-to-superpixel correlation map containing semantically aware superpixels.

[0106] Specifically, the spatial domain feature map and the frequency domain feature map are concatenated along the channel dimension and then fed into two concatenated convolutional layers for processing to obtain deep features.

[0107] Furthermore, during the training of the superpixel network, a loss function is used to optimize the superpixel network, which is expressed as:

[0108] L=Lsem+λLcompact+αLB (8)

[0109] Where Lsem is the semantic loss, Lcompact is the reconstruction loss, LB is the boundary-aware loss, and λ and α are the weight parameters.

[0110] (1) Semantic loss:

[0111] Semantic loss is used to encourage the network to group pixels and superpixels with similar colors and align them with semantic boundaries. This loss is represented by the cross-entropy loss function CE as follows:

[0112] Lsem=CE(S,S*) (9) where S is the ground-truth one-hot encoded semantic label, and S* is the reconstructed semantic label.

[0113] (2) Reconstruction loss:

[0114] To make superpixels more spatially compact, a reconstruction loss is introduced to promote lower spatial variance in superpixel clusters, which can be expressed as follows:

[0115]

[0116] (3) Boundary-aware loss:

[0117] Boundary-aware loss is used to optimize the segmentation results of boundaries, thereby enhancing the discriminative power of semantic features of different boundaries.

[0118] Assume P∈R K×K×C It is a local block sampled from the boundary pixel embedding map, where C is the number of channels and K is the block size.

[0119] To make features within the same category more similar, while features in different categories are more dissimilar, features within the same category are divided into two groups f. 1 ,f 2 ,g 1 ,g 2 Then, the classification-based loss can be expressed as follows:

[0120]

[0121] Here, Avg represents the average representation of the features, while sim(,) is the similarity measure between two features.

[0122] Finally, by considering all local blocks B, the boundary-aware loss can be expressed as follows:

[0123]

[0124] Next, in order to qualitatively and quantitatively demonstrate the superiority of the superpixel segmentation method proposed in this embodiment, this method will be compared with some state-of-the-art methods, including SSN, FCN, SLIC, SNIC, ERS, LSC, and AINet.

[0125] (1) Quantitative comparison:

[0126] Quantitative comparisons of the proposed FSNet with other state-of-the-art methods on the BSDS500, Foggy Cityscapes, RainCityscapes, SUIM, and LLRGBD datasets have been published. Figure 4The results are shown in the diagram. As can be seen, the proposed FSNet achieves the highest BR-BP scores on all datasets, and also achieves the highest ASA and UE scores on BSDS500, Foggy Cityscapes, Rain Cityscapes, and LLRGBD, while its performance on SUIM is comparable to FCN. Taking 150 superpixels as an example, the method in this embodiment achieves a minimum percentage gain (calculated based on the highest score of other methods) of 0.2% and 3.1% in ASA and UE on BSDS500, respectively, while the maximum percentage gains are 1.9% and 31.5%, respectively.

[0127] (2) Qualitative comparison:

[0128] As Figure 5 , Figure 6 , Figure 7 , Figure 8 and Figure 9 The qualitative comparison results show that, although the proposed FSNet performs best in visual performance in scenes with complex boundaries, it preserves more detail and complete object boundaries. For example, in Foggy Cityscapes, Rain Cityscapes, and SUIM, the method of this embodiment better follows the boundaries of warning signs, lane lines, and fish. Overall, the proposed FSNet outperforms the compared methods in most cases, demonstrating that the method of this embodiment achieves state-of-the-art performance on the evaluation criteria and achieves best visual performance in some complex scenes.

[0129] (3) Ablation studies:

[0130] To validate the individual contributions of each module in the proposed network, including the improved Frequency Domain Information Extractor (IFIE) and dense Hybrid Dilated Convolutional (DHAC) blocks, an ablation study was conducted on the BSDS500 dataset to investigate their effectiveness in depth.

[0131] Specifically, IFIE represents the complete model, excluding dense mixed dilated convolutional blocks; DHAC represents the complete model, excluding the improved frequency domain information extractor.

[0132] In the ablation study, AINT was used as the baseline method. Figure 10The BR-BP scores of the ablation studies are shown. The BR-BP score of the ablation model using IFIE is slightly improved compared to the baseline method. This is because the over-sharpened features extracted by IFIE may impair the semantic information of superpixels, resulting in a larger gap between the superpixel segmentation results and the baseline. The ablation model using DHAC also improves the BR-BP score, indicating that DHAC helps the model capture broader and deeper semantic information, thereby improving performance. The optimal BR-BP score is achieved when IFIE and DHAC are combined.

[0133] In summary, the proposed network demonstrates that each module can collaborate effectively, and the results of the ablation study validate the effectiveness of the proposed strategy.

[0134] (4) Reasoning efficiency:

[0135] Since inference speed is just as important as performance, the inference efficiency of several deep learning-based methods was also compared on a PC equipped with an NVIDIA 1050Ti GPU. The comparison results of various methods are available in [link to comparison]. Figure 11 The results show that AINET achieves the best inference efficiency, while our FSNet is slightly slower than AINET.

[0136] However, both FCN and FSNet are significantly more efficient at inference than SSN because the k-means clustering algorithm is time-consuming for SSN. While FSNet's inference efficiency is not the highest, its qualitative and quantitative performance are superior to AINet, meaning that FSNet achieves a good balance between performance and inference efficiency, making a sound trade-off.

[0137] (5) Application of salient target detection

[0138] Satisfactory object detection (SOD) plays a crucial role in various computer vision tasks, including semantic segmentation, video summarization, and object discovery. To leverage the advantages of superpixels, such as providing accurate object boundaries to improve SOD performance, Zhu et al. proposed a superpixel-based principled optimization framework for SOD. They observed that the connectivity between object regions and image boundaries is relatively weak, less so than that between object regions and background boundaries. Therefore, they designed a metric called boundary connectivity to quantify the degree of connectivity between a region and image boundaries, and used this metric for SOD. However, since computing boundary connectivity is challenging, they devised a "soft" method based on superpixel segmentation to indirectly compute this metric.

[0139] In this embodiment, the superpixels of this method and some state-of-the-art methods are used to replace the default superpixel segmentation methods in the above methods, including SSN and AINet, to further verify whether this method can perform better than other methods in downstream tasks.

[0140] In our experiments, we used the SUIM and UFO-120 datasets for evaluation. UFO-120 is a large and challenging underwater salient object detection dataset containing both high-resolution and low-resolution versions of images. We evaluated our models using the SUIM test set and the UFO-120 (low-resolution) set. The SUIM images are 640×480 pixels, while the UFO-120 images are 320×240 pixels. I also used Mean Absolute Error (MAE) to quantitatively compare each method; this metric measures the pixel-level average difference between the ground truth and the normalized saliency prediction map. The quantitative comparison results are available in [link to comparison]. Figure 12 As shown in the results, it is clear that this method outperforms all the compared methods in terms of MAE score.

[0141] Example 2

[0142] This embodiment discloses a frequency-domain guided superpixel segmentation system, including:

[0143] The acquisition module is configured to acquire the image to be segmented.

[0144] The superpixel segmentation module is configured to: input the image to be segmented into a preset superpixel segmentation network for processing, so as to generate a pixel-superpixel association map containing semantically aware superpixels;

[0145] The superpixel segmentation network includes a frequency domain information extractor, a densely mixed dilated convolutional block, and an association implantation module. The frequency domain information extractor is used to obtain a frequency domain feature map based on the image to be segmented. The densely mixed dilated convolutional block is used to capture high-level semantic information of the image to be segmented and generate an output feature map. The association implantation module is used to obtain a spatial domain feature map based on the output feature map, so as to fuse the spatial domain feature map and the frequency domain feature map to obtain depth features, and apply them to the softmax function to obtain a pixel-superpixel association mapping map containing semantically aware superpixels.

[0146] It should be noted that the acquisition module and superpixel segmentation module described above correspond to the steps in Embodiment 1. The examples and application scenarios implemented by these modules and their corresponding steps are the same, but they are not limited to the content disclosed in Embodiment 1. It should also be noted that these modules, as part of a system, can be executed in a computer system, such as a set of computer-executable instructions.

[0147] Example 3

[0148] Embodiment 3 of the present invention provides an electronic device, including a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When the computer instructions are executed by the processor, they complete the steps of the above-described frequency domain-guided superpixel segmentation method.

[0149] Example 4

[0150] Embodiment 4 of the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the steps of the above-described frequency-domain guided superpixel segmentation method.

[0151] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0152] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0153] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment, whereby a series of operational steps are performed to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0154] The descriptions of each embodiment in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0155] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A frequency-domain guided superpixel segmentation method, characterized in that, include: Obtain the image to be segmented; The image to be segmented is input into a preset superpixel segmentation network for processing to generate a pixel-to-superpixel association map containing semantically aware superpixels; The superpixel segmentation network includes a frequency domain information extractor, a densely mixed dilated convolutional block, and an association implantation module. The frequency domain information extractor is used to obtain a frequency domain feature map based on the image to be segmented. The densely mixed dilated convolutional block is used to capture high-level semantic information of the image to be segmented and generate an output feature map. The association implantation module is used to obtain a spatial domain feature map based on the output feature map, so as to fuse the spatial domain feature map and the frequency domain feature map to obtain depth features, and apply them to the softmax function to obtain a pixel-superpixel association mapping map containing semantically aware superpixels. The frequency domain information extractor includes a channel attention module, a channel mixer, a three-dimensional convolutional layer, a convolutional layer, and a slicing layer; The convolutional layer is used to generate a guide map based on the image to be segmented; The channel attention module is used to adaptively adjust the real or imaginary parts of the Fourier coefficients and jump-connect them with the Fourier coefficients. The channel mixer then obtains the real or imaginary features. The Fourier coefficients are generated from the image to be segmented by a fast Fourier transform. The three-dimensional convolutional layer is used to generate a two-sided mesh based on the real and imaginary features; The slicing layer is used to query the bilateral grid through the guide graph and generate a frequency domain feature map through slicing operations.

2. The frequency-domain guided superpixel segmentation method as described in claim 1, characterized in that, The step of obtaining the frequency domain feature map based on the image to be segmented includes: Generate a guide map based on the image to be segmented; The image to be segmented is converted into Fourier coefficients by Fast Fourier Transform, and the Fourier coefficients are divided into real and imaginary parts. Based on the real and imaginary parts, real and imaginary features are obtained, and the real and imaginary features are connected in the channel dimension to generate a bilateral mesh through a three-dimensional convolutional layer; Based on the guided graph, query the bilateral grid and obtain the frequency domain feature map through slicing operations.

3. The frequency-domain guided superpixel segmentation method as described in claim 2, characterized in that, The step of obtaining the real part features and imaginary part features based on the real part and imaginary part includes: The weights of the real and imaginary parts are adaptively adjusted in parallel by a channel attention module with skip connections, and then processed in blocks and stretched into a labeled embedding. The labels corresponding to the real and imaginary parts are embedded and processed in parallel by a channel mixer to obtain the real and imaginary features respectively.

4. The frequency-domain guided superpixel segmentation method as described in claim 1, characterized in that, The process of capturing high-level semantic information of the image to be segmented and generating an output feature map includes: The input feature maps are fed into the first convolutional branch, the second convolutional branch, the third convolutional branch, and the global average pooling branch, respectively, to obtain the first sub-output feature map, the second sub-output feature map, the third sub-output feature map, and the fourth sub-output feature map; The first sub-output feature map, the second sub-output feature map, the third sub-feature map and the fourth sub-feature map are concatenated and fused along the channel dimension to obtain the output feature map; The input feature map is obtained by processing the image to be segmented through multiple convolutional layers.

5. The frequency-domain guided superpixel segmentation method as described in claim 1, characterized in that, The densely mixed dilated convolutional block includes a first convolutional branch, a second convolutional branch, a third convolutional branch, a global average pooling branch, and a Conv-BN-ReLU layer; The first convolutional branch is used to obtain a first sub-output feature map based on the input feature map; the second convolutional branch is used to obtain a second sub-output feature map based on the input feature map; the third convolutional branch is used to obtain a third sub-output feature map based on the input feature map; the global average pooling branch is used to obtain a fourth sub-output feature map based on the input feature map; and the Conv-BN-ReLU is used to fuse the first, second, third, and fourth sub-output feature maps, which are concatenated along the channel dimension, to obtain an output feature map.

6. The frequency-domain guided superpixel segmentation method as described in claim 1, characterized in that, When training the superpixel network, a loss function is used to optimize the superpixel network, and the loss function is expressed as: L = Lsem +λLcompact + αLB Where Lsem is the semantic loss, Lcompact is the reconstruction loss, LB is the boundary-aware loss, and λ and α are the weight parameters.

7. A frequency-domain guided superpixel segmentation system, characterized in that, include: The acquisition module is configured to acquire the image to be segmented. The superpixel segmentation module is configured to: input the image to be segmented into a preset superpixel segmentation network for processing, so as to generate a pixel-superpixel association map containing semantically aware superpixels; The superpixel segmentation network includes a frequency domain information extractor, a densely mixed dilated convolutional block, and an association implantation module. The frequency domain information extractor is used to obtain a frequency domain feature map based on the image to be segmented. The densely mixed dilated convolutional block is used to capture high-level semantic information of the image to be segmented and generate an output feature map. The association implantation module is used to obtain a spatial domain feature map based on the output feature map, so as to fuse the spatial domain feature map and the frequency domain feature map to obtain depth features, and apply them to the softmax function to obtain a pixel-superpixel association mapping map containing semantically aware superpixels. The frequency domain information extractor includes a channel attention module, a channel mixer, a three-dimensional convolutional layer, a convolutional layer, and a slicing layer; The convolutional layer is used to generate a guide map based on the image to be segmented; The channel attention module is used to adaptively adjust the real or imaginary parts of the Fourier coefficients and jump-connect them with the Fourier coefficients. The channel mixer then obtains the real or imaginary features. The Fourier coefficients are generated from the image to be segmented by a fast Fourier transform. The three-dimensional convolutional layer is used to generate a two-sided mesh based on the real and imaginary features; The slicing layer is used to query the bilateral grid through the guide graph and generate a frequency domain feature map through slicing operations.

8. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Image segmentation method and device and terminal equipment

    CN111199547A

  • Real-time superpixel segmentation method and system based on recurrent neural network

    CN112926596A