Unsupervised super-pixel segmentation method and system assisted by collaboration between atrous pyramid and attention mechanism
Through the hollow pyramid synergistic attention mechanism, the problems of insufficient generalization ability and high complexity of the existing superpixel segmentation method are solved, and adaptive depth feature extraction and superpixel generation are realized, which improves the accuracy and meticulousness of the segmentation results.
Patent Information
- Application Number
- PCT/CN2023/133843
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-23
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-30
AI Technical Summary
The existing superpixel segmentation methods have insufficient generalization capabilities, high complexity, difficulty in obtaining effective feature information, and lack of adaptive capabilities.
The hollow pyramid synergistic attention mechanism is adopted to introduce the spatial relationship between pixel points through image preprocessing, and the attention mechanism is used to enhance attention to important feature channels. The depth features are extracted in combination with the hollow space pyramid pooling, and the parameter update is realized through the loss function and the Adam optimizer, and finally the adaptive superpixel generation is realized.
The complexity of the superpixel segmentation algorithm is reduced, the generalization and adaptability of the model are improved, and the generated superpixel segmentation results are more accurate and detailed.
Smart Images

Figure CN2023133843_30052025_PF_FP_ABST
Abstract
Description
A diaphanous pyramid collaborative attention mechanism assisted unsupervised superpixel segmentation method and system Technical Field
[0001] The present invention belongs to the field of digital image processing technology, and in particular relates to a hollow pyramid collaborative attention mechanism assisted unsupervised superpixel segmentation method and system. Background Art
[0002] Superpixel segmentation is an effective means to improve the efficiency and accuracy of image processing. Deep learning supervised superpixel segmentation methods rely on large amounts of labeled data, and are subject to data bias that leads to inaccurate segmentation results and insufficient generalization of segmentation models. Traditional superpixel segmentation methods based on energy optimization, watershed, graph, and clustering rely on appropriate parameter selection and are unable to adaptively determine the number of superpixels based on the image's inherent characteristics. They are sensitive to noise and suffer from high computational complexity with large-size images. Unsupervised superpixel segmentation methods, on the other hand, are not restricted by labeled data and do not require manual adjustment of model parameters and structures. They have the advantages of better generalization, avoidance of noise interference, and no high computational complexity due to changes in image size. Therefore, they are an important alternative to supervised and traditional superpixel segmentation methods. However, achieving superpixel segmentation using unsupervised methods requires solving two key problems: effective deep feature extraction and superpixel generation.
[0003] Commonly used feature extraction methods are mainly divided into manually designed features and deep features extracted by convolutional neural networks. The former usually converts the RGB channels of the image into LAB channel representations, and then combines the position information of the pixels in the image to generate five-dimensional features for subsequent superpixel generation. The latter automatically learns deep features through the constructed convolutional model. Both methods can extract useful features from images. Manually designed features rely on the researcher's domain knowledge and experience of the problem, while convolutional neural networks extract deep features relying on the construction of the convolutional model and the selection of feature extraction dimensions. However, manually designed features have the problem of insufficient features for superpixel segmentation of complex images, while convolutional neural networks extract deep features relying on the constructed convolutional model. Simple convolutional models cannot extract useful features, and complex convolutional models have the problem of high training difficulty.
[0004] Superpixel generation primarily involves leveraging the available information from feature extraction to segment the original image into contiguous, compact, and similarly characterized small regions. Commonly used methods include watershed, graph, and clustering. In watershed techniques, dark regions are typically considered valleys and brighter regions are considered ridges. The height of the ridge is defined by the grayscale value or gradient of a specific pixel, enabling superpixel generation. However, this approach suffers from the problem of being untrainable. Graph techniques treat pixels as nodes and edge weights as the similarity between adjacent pixels. However, this approach is difficult to implement for large images, and superpixel segmentation performance is highly dependent on parameters such as the merging rule, similarity metric, and the number of superpixels. Clustering techniques can rapidly generate superpixels without the need for additional labels. Trainable clustering algorithms (typically differentiable K-means clustering) have been developed, but they require multiple iterations to achieve good results, increasing the computational cost and time overhead of the overall approach. Other clustering techniques generally suffer from high computational complexity. Technical issues
[0005] Through the above analysis, the problems and defects of the existing technology are: the existing superpixel segmentation method has insufficient generalization ability, high complexity, difficulty in obtaining effective feature information and lack of adaptability. Technical Solutions
[0006] In response to the problems existing in the prior art, the present invention provides a hollow pyramid collaborative attention mechanism to assist unsupervised superpixel segmentation method and system.
[0007] The present invention is implemented as follows: first, the image is preprocessed to introduce the spatial relationship between pixels, and the attention mechanism is applied to the superpixel segmentation task for the first time to enable the model to strengthen its attention to important feature channels and suppress the response to irrelevant channels; the hollow space pyramid pooling is used to reduce parameters while expanding the receptive field, and the optimizer is used in combination with the constructed loss function to realize parameter update and the extraction of the final effective depth features. The argmax function is applied to the extracted effective depth features to convert the superpixel segmentation task into a classification problem, and the final adaptive superpixel generation is achieved by adding a size restriction condition. In this way, the use of clustering algorithms can be avoided and effective depth features can be extracted with a small number of parameters, thereby greatly reducing the complexity of the superpixel segmentation algorithm. The proposed superpixel segmentation method is unsupervised and therefore has strong transferability.
[0008] Furthermore, the attention mechanism cooperates with the dilated spatial pyramid pooling to promote the unsupervised superpixel segmentation method, which includes the following steps:
[0009] Step 1: Combine the RGB channel information of the image with the position information of the pixel points to convert the three-dimensional features into five-dimensional features. First, assign the image to be segmented by superpixel to the variables in the image preprocessing, and change the variables expressed in array form to tensor form. Use the permute function to rearrange the dimensions into And change the data type to floating point type, add a batch dimension externally through the None operation to get the variable image, whose shape is ; Use the torch.arange function to generate height and width sequences, combine the torch.meshgrid function to convert the two sequences into two coordinate grids and stack them using the torch.stack function, and compare the results to the tensor The final image preprocessing result is obtained by concatenating and normalizing in the channel dimension.
[0010] Step 2: Use the attention mechanism to build a channel attention module; apply the image preprocessing results to the point-by-point convolution layer to obtain a shape of The tensor is then processed using channel global average pooling to obtain aggregate features. The result is automatically calculated with a kernel size of The attention mechanism's processing results are obtained by performing element-wise multiplication of the tensor obtained by the point-by-point convolution and the fast one-dimensional convolution. The weights of the point-by-point convolution and fast one-dimensional convolution involved are initialized using the Kaiming initialization method, and the bias is initialized with a constant value of 0. For the instance normalization layer, the normalized weights are initialized with a constant value of 1 to ensure appropriate initial parameter values during training.
[0011] Step 3: Use dilated spatial pyramid pooling to process the results of the attention mechanism and extract deep features suitable for superpixel segmentation; the tensors obtained by the channel attention module are processed by convolution with a size of , fill size is 0, sampling rate is 1 Convolutional layer; the convolution kernel size is , padding size is 2, sampling rate is 2 Convolutional layer; the convolution kernel size is , padding size of 4, sampling rate of 4 Convolutional layer; the convolution kernel size is , padding size is 6, sampling rate is 6 Convolutional layer; resize the input tensor to , and then perform a step of 1 Convolution and application of ReLU activation function Operation; set the output channels of all operations to 16, and use instance normalization to normalize each output channel, and concatenate the results of each output channel in the channel dimension to obtain a shape of The concatenated tensor uses a convolution kernel size of The convolution layer, instance normalization layer and ReLU function are processed to obtain the final shape The deep feature tensor of is obtained; and the weights and biases of the convolutional layer in the dilated spatial pyramid pooling and the instance normalization layer are initialized using the same initialization method as the attention mechanism.
[0012] Step 4: Construct the loss function. First, construct the clustering loss term. At the same time, use the spatial smoothing loss term to quantify the difference between adjacent pixels. Then construct the reconstruction loss term. Separate the deep features obtained by the void space pyramid pooling into a shape of The tensor is used to calculate the clustering loss term and the spatial smoothing loss term, and the other shape is The tensor is used to calculate the reconstruction loss term; the softmax function is used to calculate the tensor The channel dimension is processed to convert it into the corresponding category probability. The negative log-likelihood of each pixel is calculated and the average of all pixel loss values is taken to obtain the final loss value and the average probability estimate of each sample for each category is used to construct the clustering loss term. The softmax function is used again to transform the tensor The channel dimension of is processed to obtain the corresponding category probability and regard it as a probability map. The difference between each element of the probability map and the variable image in the W dimension and its adjacent element to the right and the difference between each element of the H dimension and its adjacent element below are calculated to obtain the gradient of the probability map and the variable image in the horizontal and vertical directions. The four calculated gradients are used to construct the spatial smoothing loss term. The tensor and the variable image are used to calculate the mean square error function in PyTorch to measure the difference between the two tensors, thereby determining the reconstruction loss term in the loss function.
[0013] Step 5: Update the model parameters by setting the learning rate and number of iterations of the Adam optimizer to find the model parameters that minimize the loss function; iterate through the defined optimize function and use the Adam optimizer to update the model parameters, set the number of iterations to 500, and the learning rate to The coefficient defined in the clustering loss term in the loss function is set to a constant value of 2, and the weights of the spatial smoothing loss term and the reconstruction loss term are set to constant values of 2 and 10, respectively, completing the construction of the overall loss function. This allows the model to automatically find appropriate parameters to minimize the overall loss function within the set number of iterations.
[0014] Step 6: Use the argmax function to obtain the maximum value of the channel dimension, and convert the final effective depth feature into an optimal superpixel label index corresponding to each pixel point; convert the argmax function processing result into a two-dimensional array and complete the adaptive superpixel segmentation in the CPU according to the limited conditions.
[0015] Furthermore, the image is preprocessed using the following formula in step 1:
[0016]
[0017] In the formula , , Indicates the color channel value of the k-th pixel in the image converted from integer to floating point, where , Indicates the row and column number of the k-th pixel in the image. Represent the five-dimensional features of the k-th pixel in the image and apply them to subsequent processing.
[0018] Furthermore, the construction method of step 2 specifically includes:
[0019] Step 21: Apply the preprocessed five-dimensional features to a point-by-point convolutional layer, and linearly combine and transform the five-dimensional features to achieve eight-dimensional feature output while keeping the height and width unchanged;
[0020] Step 22: Calculate the channel global average pooling pair The aggregated features are obtained by processing
[0021]
[0022] Step 23, calculate the aggregated features through fast one-dimensional convolution with kernel size L:
[0023]
[0024] Where, is the sigmoid function, is a fast one-dimensional convolution with kernel size L, is the learned channel weight; the kernel size L is automatically calculated based on the number of channels:
[0025]
[0026] In the formula , , Defined as the odd integer closest to t.
[0027] Step 24: The aggregated features are processed by a fast one-dimensional convolution with a kernel size of L and a sigmoid function, and the result is expressed as , and with Perform element-wise multiplication to obtain the attention mechanism processing result expressed as , the shape is consistent with the eight-dimensional features of the input attention mechanism;
[0028] .
[0029] Furthermore, the eight-dimensional feature output is expressed as And directly apply channel global average pooling, where H represents the image height, W represents the image width, and C represents the number of channels of the image. In this case, C=8.
[0030] Furthermore, the construction of the channel attention module in step 3 specifically includes:
[0031] Step 31, the depth feature obtained after the void space pyramid pooling process is expressed as , where H and W remain unchanged and still correspond to the height and width of the image, ; The calculation process of the intermediate convolutional layer is as follows:
[0032]
[0033] Where, Indicates that the convolution kernel size is , convolutional layer with padding size 0 and sampling rate 1; Indicates that the convolution kernel size is , convolutional layer with padding size 2 and sampling rate 2; Indicates that the convolution kernel size is , convolutional layer with padding size 4 and sampling rate 4; Indicates that the convolution kernel size is , convolutional layers with padding size 6 and sampling rate 6; Indicates that the size of the input tensor is adjusted to ; Perform a step of 1 Convolution is performed with 8 input channels and 16 output channels. Each output channel is normalized using instance normalization, and finally a ReLU activation function is applied to introduce nonlinearity. 、 、 、 as well as is the intermediate tensor;
[0034] Step 32, deep features The calculation is as follows:
[0035]
[0036] Where, Indicates that 、 、 、 and Concatenate in the channel dimension to form a larger tensor, 、 、 、 and The output channels are all 16, and the number of channels of the spliced tensor is 80; Indicates that the concatenated tensor is Convolution, the output channel is set to 128, which is suitable for superpixel segmentation requirements; instance normalization is used to normalize each output channel; finally, the ReLU activation function is applied to introduce nonlinearity.
[0037] Furthermore, step 4 constructs the loss function, which specifically includes:
[0038] Step 41, the overall loss function consists of three parts: clustering loss term, spatial smoothing loss term and reconstruction loss term:
[0039]
[0040] Where, Represents the overall loss function; represents the clustering loss term; represents the spatial smoothing loss term; represents the reconstruction loss term; and are all constant coefficients and 、 .
[0041] In step 42, the clustering loss term is calculated as follows:
[0042]
[0043] Where, ; Represents the average value of the class probability vector over all pixels; Represents the category probability vector of the pixel located at row i and column j, which is calculated as follows:
[0044]
[0045] Step 43, The first three channels in the channel dimension in are separated and used for the subsequent reconstruction loss calculation, and the features of the remaining 125 channels are used To express, confirm Rounding, Indicates that the softmax function is used to The channel dimension is converted into the category probability of the corresponding pixel;
[0046] Step 44, the calculation of the spatial smoothing loss term is divided into two parts: Direction and The smoothness loss in the direction is to rearrange the original input image to obtain the image ; By calculating the absolute value of the probability difference and the image The exponential function of the squared difference of the gradient is defined, and the average spatial smoothness loss of all pixels is calculated as follows:
[0047]
[0048] Where, Indicates that the channel and clustering losses are consistent; Expressed as Pixel probability difference in direction; Expressed as The difference in pixel intensity in the direction; Expressed as Pixel probability difference in direction; Expressed as The pixel intensity difference in the direction; the specific calculation method is as follows:
[0049]
[0050] In step 45, the reconstruction loss term is calculated as follows:
[0051]
[0052] When calculating the clustering loss and spatial smoothing loss, the first three channels separated from the deep features are used for image reconstruction and expressed as ; Indicates the use of the 2-norm.
[0053] Furthermore, in step 44, the original input image dimension is ,image The dimension is .
[0054] Furthermore, in step five, the depth feature is extracted under the minimized model parameters. The depth feature is the effective depth feature, and the effective depth feature obtained by the last separation is used for superpixel generation.
[0055] Furthermore, the size constraint for superpixel generation in step 6 is calculated as follows:
[0056]
[0057] Where, represents the average size of an ideal superpixel; Represents the total number of pixels in the image; and are the minimum and maximum thresholds for limiting superpixels, respectively, and are used to filter superpixels that are too small or too large so that the size of the final generated superpixels is as uniform as possible; the hyperparameters in the formula are .
[0058] Another object of the present invention is to provide an attention mechanism that cooperates with void space pyramid pooling to promote unsupervised superpixel segmentation method, and an attention mechanism that cooperates with void space pyramid pooling to promote unsupervised superpixel segmentation system, the system comprising:
[0059] Image preprocessing module, which is used to combine the image RGB channel information with the pixel position information to convert the three-dimensional features into five-dimensional features;
[0060] Attention mechanism module, used to build channel attention module using attention mechanism;
[0061] The dilated spatial pyramid pooling module is used to process the results of the attention mechanism using dilated spatial pyramid pooling to extract deep features suitable for superpixel segmentation;
[0062] The loss function construction module is used to construct the loss function. First, the clustering loss term is constructed; at the same time, the spatial smoothing loss term is used to quantify the difference between adjacent pixels; and then the reconstruction loss term is constructed.
[0063] The parameter update module is used to update the model parameters by setting the learning rate and number of iterations of the Adam optimizer;
[0064] The superpixel segmentation module is used to convert the results of the argmax function into a two-dimensional array and complete adaptive superpixel segmentation based on the limited conditions in the CPU. Beneficial effects
[0065] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0066] First, in view of the technical problems existing in the above-mentioned prior art and the difficulty of solving these problems, we closely combine the technical solutions to be protected by the present invention and the results and data during the research and development process, and analyze in detail and in depth how the technical solutions of the present invention solve the technical problems and some creative technical effects brought about by solving the problems. The specific description is as follows:
[0067] This paper provides an unsupervised superpixel segmentation method that uses an attention mechanism in conjunction with dilated spatial pyramid pooling. This method extracts effective deep features through image preprocessing, an attention mechanism in conjunction with dilated spatial pyramid pooling, and updates the loss function and Adam optimizer parameters. It also transforms superpixel segmentation into a classification problem, enabling adaptive superpixel generation. This method addresses the problems of existing superpixel segmentation methods, such as insufficient generalization, high complexity, difficulty in obtaining effective feature information, and a lack of adaptability.
[0068] Second, considering the technical solution as a whole or from the perspective of the product, the technical effects and advantages of the technical solution to be protected by the present invention are described in detail as follows:
[0069] Image preprocessing: Using the coordinates of image pixel locations as additional features improves the performance of superpixel segmentation methods, allowing the model to better capture the spatial structure and relationships between pixels in the image. The row and column numbers corresponding to each pixel in the image generate coordinate information. This generated coordinate information is appended to each pixel value in the original image to create a new image with five channels. This new image with five-dimensional features is fed into a deep feature extraction network, which can better understand the image and extract deep features more suitable for superpixel segmentation.
[0070] Attention mechanism: In this superpixel segmentation method, the coverage of local cross-channel interactions is determined by adaptively selecting the size of the one-dimensional convolution kernel. Efficient channel attention is achieved with only a small number of parameters, which can improve the superpixel segmentation performance while balancing the complexity it brings.
[0071] Atrous Spatial Pyramid Pooling: In superpixel segmentation tasks, each pixel must be classified to determine which superpixel class it belongs to. Traditional convolutional layers can only capture information within a limited range, necessitating the modeling of information at different scales and contexts. Atrous Spatial Pyramid Pooling expands the receptive field by introducing a dilation rate to the input image, capturing information across a wider range without increasing the number of parameters. The resulting atrous spatial pyramid pooling extracts deep features suitable for superpixel segmentation tasks from images of varying complexity.
[0072] Loss function: To ensure that the extracted deep features come from the image itself, a reconstruction loss term is introduced into the loss function so that the deep feature extraction is forced to generate an output that matches the original image, thereby restoring the original input image. The constructed clustering loss term is an entropy-based clustering cost, similar to the mutual information term of regularized information maximization. Clustering is achieved by maximizing the mutual information of pixels under the deep feature representation, and the complexity of the superpixel segmentation method is controlled by the regularization term to avoid overfitting. The constructed spatial smoothness loss term is the main prior for image processing tasks. It can quantify the differences between adjacent pixels. While ensuring that the superpixel segmentation method generates smooth output, it also greatly improves the generalization ability of the superpixel segmentation method and prevents overfitting.
[0073] Parameter Update and Superpixel Generation: To update model parameters, the Adam optimizer is used for gradient descent optimization. By calculating the first- and second-order moment estimates of the gradient, independent adaptive learning rates are designed for different parameters. Furthermore, parameter updates are unaffected by gradient scaling, overcoming the significant noisy gradients and making model parameter updates simpler and more efficient. The argmax function is used to convert the final effective depth features into a single optimal superpixel label index for each pixel. This process transforms the superpixel segmentation problem into a classification problem. Combined with constraints, adaptive superpixel segmentation can be quickly completed, avoiding the use of complex clustering algorithms.
[0074] Third, the expected benefits and commercial value of the technical solution of the present invention after transformation are: it can be applied to remote sensing image processing, greatly accelerating its processing speed.
[0075] The technical solution of this invention overcomes technical bias: Traditional superpixel segmentation methods mostly use clustering algorithms. This invention transforms superpixel segmentation into a classification task, thus avoiding the use of clustering algorithms. This makes the superpixel segmentation method highly adaptable to changes in image size.
[0076] Fourth, the significant technological advancements brought about by the unsupervised superpixel segmentation method provided by the present invention are mainly reflected in the following aspects:
[0077] 1) Enhanced feature extraction capabilities:
[0078] By combining the attention mechanism and dilated spatial pyramid pooling, this method can more effectively extract deep features of images. The introduction of the attention mechanism enables the model to pay more attention to important feature channels, thereby improving the accuracy of feature extraction.
[0079] 2) Improved superpixel segmentation quality:
[0080] Traditional superpixel segmentation methods ignore some detailed information, while this method can better preserve the details and structure of the image by combining spatial relationships and depth features, thereby generating more accurate and detailed superpixel segmentation results.
[0081] 3) Improved flexibility and adaptability:
[0082] This method can flexibly adapt to images of different types and qualities by adaptively generating superpixels, which is especially important when dealing with diverse image datasets.
[0083] 4) Parameter optimization and computational efficiency:
[0084] Combining the Adam optimizer with a specific loss function allows for more efficient optimization of model parameters, reducing computational cost and time. Furthermore, dilated spatial pyramid pooling expands the receptive field while reducing the number of parameters, further improving computational efficiency.
[0085] 5) Wide application potential:
[0086] This method is not only applicable to standard image processing tasks, but can also be extended to other fields such as medical image analysis, machine vision, image recognition, etc., showing wide application potential.
[0087] In summary, the method provided by the present invention has achieved significant technological progress in the field of superpixel segmentation through its innovative technology combination and optimization strategy, improved segmentation quality, expanded the scope of application, and improved processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0089] FIG1 is a flow chart of a method for promoting unsupervised superpixel segmentation by using an attention mechanism in collaboration with atrous spatial pyramid pooling according to an embodiment of the present invention;
[0090] FIG2 is a structural diagram of a system for promoting unsupervised superpixel segmentation by using an attention mechanism in collaboration with atrous spatial pyramid pooling according to an embodiment of the present invention;
[0091] FIG3 is a diagram of an attention mechanism constructed according to an embodiment of the present invention;
[0092] FIG4 is a diagram of a hollow space pyramid pooling constructed according to an embodiment of the present invention;
[0093] FIG5 is a diagram showing superpixel segmentation effects provided by an embodiment of the present invention;
[0094] Figure 6 is a comparison chart of the effects of the present invention; A, true label; B, SLIC; C, Algorithm 2; D, superpixel segmentation result of the present invention. Modes for Carrying Out the Invention
[0095] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0096] The present invention provides an unsupervised superpixel segmentation method based on the attention mechanism and dilated spatial pyramid pooling. The method mainly includes six steps, specifically:
[0097] 1) Image preprocessing and feature conversion:
[0098] The RGB channel information of the image is combined with the position information of the pixel points to convert the three-dimensional features into five-dimensional features.
[0099] 2) Construct channel attention module:
[0100] Through the attention mechanism, the model's attention to important feature channels is enhanced, and responses to irrelevant channels are suppressed.
[0101] 3) Applying atrous spatial pyramid pooling:
[0102] Dilated spatial pyramid pooling is applied to the results of attention mechanism processing to extract deep features suitable for superpixel segmentation.
[0103] 4) Construct and apply loss function:
[0104] Construct clustering loss, spatial smoothness loss (quantifying the difference between adjacent pixels), and reconstruction loss.
[0105] 5) Model parameter update:
[0106] The model parameters are updated by setting the learning rate and number of iterations of the Adam optimizer to find the model parameters that minimize the loss function.
[0107] 6) Superpixel label generation and adaptive superpixel segmentation:
[0108] Use the argmax function to obtain the maximum value of the channel dimension and convert the effective depth features into the superpixel label index of each pixel. Convert the argmax result into a two-dimensional array and perform adaptive superpixel segmentation based on the constraints on the CPU.
[0109] The following are two specific embodiments and implementations of the present invention. These embodiments typically include specific applications of the method, such as specific types of image datasets (e.g., natural scene images, medical images, etc.) or specific application scenarios (e.g., image segmentation, object tracking, etc.). However, since your description does not provide sufficient details to identify these embodiments, I will provide two examples:
[0110] Application Example 1: Superpixel Segmentation of Natural Scene Images
[0111] In this embodiment, the method can be applied to natural scene images. By analyzing features such as color and texture, the model can effectively segment the image into groups of superpixels with similar characteristics. This is very useful in image editing and enhancement applications.
[0112] Application Example 2: Medical Image Analysis
[0113] In medical images, such as MRI or CT scans, superpixel segmentation can be used to identify and differentiate different tissue types or lesions. This approach can help doctors diagnose and plan treatments more accurately.
[0114] The specific implementation will involve selecting an appropriate image dataset, adjusting model parameters (such as learning rate, number of iterations, configuration of pooling layers, etc.), and post-processing steps such as merging or segmenting superpixel groups to improve segmentation accuracy and efficiency.
[0115] In response to the problems existing in the prior art, the present invention provides a hollow pyramid collaborative attention mechanism to assist unsupervised superpixel segmentation method and system.
[0116] As shown in FIG1 , an embodiment of the present invention provides an attention mechanism in conjunction with dilated spatial pyramid pooling to promote unsupervised superpixel segmentation, including the following steps:
[0117] Step 1: Image preprocessing
[0118] In order to ensure the extraction of effective depth features, the image needs to be preprocessed, and the image RGB channel information is directly combined with the pixel position information to realize the conversion of three-dimensional features into five-dimensional features, ensuring the efficiency of the overall superpixel segmentation method.
[0119]
[0120] In the above formula , , Indicates the color channel value of the k-th pixel in the image converted from integer to floating-point type, where the type conversion is mainly to ensure the consistency of subsequent data types and the accuracy of the overall method. , Indicates the row and column number of the k-th pixel in the image. Represent the five-dimensional features of the k-th pixel in the image and apply them to subsequent processing.
[0121] Step 2: Attention Mechanism
[0122] In order to enable the overall superpixel segmentation method to dynamically adjust pixel attention according to different input images, a channel attention module was constructed that does not require local cross-channel interaction of dimensionality reduction, adaptively determines the size of one-dimensional convolution kernel, and is efficient.
[0123] First, the preprocessed five-dimensional features are applied to a point-by-point convolution layer. While keeping the height and width unchanged, the five-dimensional features are linearly combined and transformed to achieve eight-dimensional feature output. The eight-dimensional feature output is represented as And directly apply channel global average pooling, where H represents the image height, W represents the image width, and C represents the number of channels of the image. In this case, C=8.
[0124] Calculate channel global average pooling pair The aggregated features are obtained by processing
[0125]
[0126] Calculate the aggregated features through fast one-dimensional convolution with kernel size L to achieve channel attention learning for local cross-channel interaction
[0127]
[0128] Where, is the sigmoid function, is a fast one-dimensional convolution with kernel size L, is the learned channel weight.
[0129] Automatically calculate the kernel size L based on the number of channels
[0130]
[0131] Where, , , Defined as the nearest odd integer to t.
[0132] The aggregated features are processed by fast one-dimensional convolution with kernel size L and sigmoid function, and the result is expressed as , and with Perform element-wise multiplication to obtain the attention mechanism processing result expressed as Its shape is consistent with the eight-dimensional feature input to the attention mechanism.
[0133]
[0134] Step 3: Atrous Space Pyramid Pooling
[0135] The results of the attention mechanism are processed using dilated spatial pyramid pooling to extract deep features suitable for superpixel segmentation.
[0136] The depth feature obtained after the void space pyramid pooling is expressed as , where H and W remain unchanged and still correspond to the height and width of the image. The calculation process of the intermediate convolutional layer is as follows:
[0137]
[0138] Indicates that the convolution kernel size is , convolutional layer with padding size 0 and sampling rate 1; Indicates that the convolution kernel size is , convolutional layer with padding size 2 and sampling rate 2; Indicates that the convolution kernel size is , convolutional layer with padding size 4 and sampling rate 4; Indicates that the convolution kernel size is , convolutional layers with padding size 6 and sampling rate 6; Indicates that the size of the input tensor is adjusted to , and then perform a step of 1 Convolution is performed with 8 input channels and 16 output channels. Each output channel is normalized using instance normalization, and the ReLU activation function is applied to introduce nonlinearity. 、 、 and The input channels are all set to 8, the output channels are all set to 16, and each convolution layer uses instance normalization to normalize each output channel. 、 、 、 as well as is the intermediate tensor.
[0139] Deep Features The calculation is as follows:
[0140]
[0141] Where, Indicates that 、 、 、 and Concatenate in the channel dimension to form a larger tensor, 、 、 、 and The output channels are all 16, so the number of channels of the spliced tensor is 80; Indicates that the concatenated tensor is Convolution, the output channel is set to 128 to meet the requirements of super pixel segmentation; Similarly, instance normalization is used to normalize each output channel; finally, the ReLU activation function is applied to introduce nonlinearity.
[0142] Step 4: Construct loss function
[0143] In order to encourage deterministic superpixel allocation and make the size of each superpixel as uniform as possible, a clustering loss term is constructed; at the same time, a spatial smoothing loss term is introduced to quantify the differences between adjacent pixels; by constructing a reconstruction loss term, it can be ensured that the extracted deep features come from the image itself.
[0144] The overall loss function consists of three parts: clustering loss, spatial smoothing loss and reconstruction loss.
[0145]
[0146] Where, Represents the overall loss function; represents the clustering loss term; represents the spatial smoothing loss term; represents the reconstruction loss term; and are all constant coefficients and 、 .
[0147] The clustering loss term is calculated as follows:
[0148]
[0149] in, ; Represents the average value of the class probability vector over all pixels; Represents the category probability vector of the pixel located at row i and column j, which is calculated as follows:
[0150]
[0151] First, The first three channels in the channel dimension in are separated and used for the subsequent reconstruction loss calculation, and the features of the remaining 125 channels are used To indicate that Rounding, Indicates that the softmax function is used to The channel dimension is converted into the category probability of the corresponding pixel.
[0152] The computation of spatial smoothness loss is divided into two parts: Direction and The smoothness loss in the direction is to rearrange the original input image to obtain the image (The original input image dimension is ,image The dimension is ). By calculating the absolute value of the probability difference and the image The exponential function of the squared gradient difference is defined, and the average spatial smoothness loss of all pixels is calculated. The specific calculation method is as follows:
[0153]
[0154] Where, Indicates that the channel and clustering losses are consistent; Expressed as Pixel probability difference in direction; Expressed as The difference in pixel intensity in the direction; Expressed as Pixel probability difference in direction; Expressed as The pixel intensity difference in the direction. The specific calculation method is as follows:
[0155]
[0156] The reconstruction loss term is calculated as follows:
[0157]
[0158] Where, when calculating the clustering loss and spatial smoothing loss, the first three channels separated from the deep features are used for image reconstruction and expressed as ; Indicates the use of the 2-norm.
[0159] Step 5: Parameter update and superpixel generation
[0160] By setting the learning rate and number of iterations of the Adam optimizer, the model parameters are updated, that is, the model parameters that satisfy the minimum loss function are found. The deep features extracted under this parameter are regarded as effective deep features. Since different deep features are obtained each time, the deep features obtained in the last separation are The effective depth features are used to generate superpixels. The argmax function is used to obtain the maximum value of the channel dimension, so that each pixel corresponds to the most accurate superpixel label index. The result of the argmax function is converted into a two-dimensional array and adaptive superpixel segmentation is performed on the CPU according to the constraints. The size constraint for superpixel generation is calculated as follows:
[0161]
[0162] Where, represents the average size of an ideal superpixel; Represents the total number of pixels in the image; and are the minimum and maximum thresholds for limiting superpixels, respectively, and are used to filter superpixels that are too small or too large so that the size of the final generated superpixels is as uniform as possible; the hyperparameters in the formula are .
[0163] As shown in FIG2 , an embodiment of the present invention provides an attention mechanism that cooperates with atrous spatial pyramid pooling to promote unsupervised superpixel segmentation. The system includes:
[0164] Image preprocessing module, which is used to combine the image RGB channel information with the pixel position information to convert the three-dimensional features into five-dimensional features;
[0165] Attention mechanism module, used to build channel attention module using attention mechanism;
[0166] The dilated spatial pyramid pooling module is used to process the results of the attention mechanism using dilated spatial pyramid pooling to extract deep features suitable for superpixel segmentation;
[0167] The loss function construction module is used to construct the loss function. First, the clustering loss term is constructed; at the same time, the spatial smoothing loss term is used to quantify the difference between adjacent pixels; and then the reconstruction loss term is constructed.
[0168] The parameter update module is used to update the model parameters by setting the learning rate and number of iterations of the Adam optimizer;
[0169] The superpixel segmentation module is used to convert the results of the argmax function into a two-dimensional array and complete adaptive superpixel segmentation based on the limited conditions in the CPU.
[0170] Example:
[0171] The deep learning framework used in this embodiment is PyTorch, and the programming language is Python.
[0172] Step 1: Assign the acquired 0.8m resolution color image from GF-2 to the variable img and check if CUDA is available. If so, run the subsequent code on the GPU; if not, run it on the CPU. The acquired image is applied to the image preprocessing function written. By rearranging the dimensions, converting the data type to floating point, and adding a batch dimension externally with the None operation, the shape is obtained. tensor of ;
[0173] Read the last two dimensions of the tensor to get the height and width values of the image, use the torch.arange function to generate a sequence of height and width, use the torch.meshgrid function to convert the two sequences into two coordinate grids and stack them using the torch.stack function; connect the image and the coordinate grid and perform the normalization operation to get a tensor shape of results.
[0174] Step 2: Apply the image preprocessing result to the point-by-point convolution layer to obtain a shape of The tensor is then processed using channel global average pooling to obtain aggregate features. The result is automatically calculated with a kernel size of The fast one-dimensional convolution and sigmoid function are processed, and the element-wise convolution tensor is multiplied to obtain the attention mechanism (as shown in Figure 3). The tensor shape of the processing result is also ;
[0175] The Kaiming initialization method is used to initialize the weights of point-by-point convolution and fast one-dimensional convolution, and the bias is initialized with a constant value of 0. For the instance normalization layer, the normalized weights are initialized with a constant value of 1 to ensure appropriate initial parameter values during training.
[0176] Step 3: Apply the tensor obtained by the attention mechanism to the void space pyramid pooling (as shown in Figure 4), so that it passes through the convolution size of , fill size is 0, sampling rate is 1 Convolutional layer; the convolution kernel size is , padding size is 2, sampling rate is 2 Convolutional layer; the convolution kernel size is , padding size is 4, sampling rate is 4 Convolutional layer; the convolution kernel size is , padding size is 6, sampling rate is 6 Convolutional layer; the input tensor is resized to , and then perform a step of 1 Convolution, applying ReLU activation function to introduce nonlinearity operate;
[0177] Set the output channels of the above operations to 16, so that the shape after splicing is The concatenated tensor uses a convolution kernel size of The convolution layer, instance normalization layer and ReLU are processed to obtain the final shape The deep feature tensor of is obtained; each output channel is normalized using instance normalization, and finally a ReLU activation function is applied to introduce nonlinearity; the weights and biases of the convolutional layer and the instance normalization layer in the dilated spatial pyramid pooling are initialized using the same initialization method as the attention mechanism.
[0178] Step 4: Separate the extracted depth features into a shape of A tensor of is used to calculate the reconstruction loss term, and another one of shape is The tensor is used to calculate the clustering loss term and the spatial smoothing loss term; it is iterated through the defined optimize function, and the Adam optimizer is used to update the model parameters to minimize the loss function; the number of iterations is set to 500 times, and the learning rate is set to ; When the number of iterations is reached, the argmax function is used to obtain the maximum value of the channel dimension, so that each pixel corresponds to the most accurate superpixel label index;
[0179] After obtaining the label index, the channel dimension is removed through the squeeze function, the gradient is no longer calculated, and the result is transferred from the GPU to the CPU for processing; the PyTorch tensor is converted into a Numpy array for subsequent superpixel segmentation, and finally the _enforce_label_connectivity_cython function is constructed to adaptively implement superpixel segmentation according to the limited superpixel size condition.
[0180] This embodiment processes a facial image obtained from the scipy library. FIG5 is a diagram showing the superpixel segmentation effect of the embodiment.
[0181] As can be seen from the above examples, the present invention realizes unsupervised superpixel segmentation, and the segmentation accuracy test result reaches 94.7%. The method provided by the present invention flexibly and efficiently realizes effective depth feature extraction and superpixel generation adaptively based on the characteristics of the image itself without providing real labels and the number of superpixels. Unsupervised superpixel segmentation can also be achieved quickly and accurately under the conditions of different image sizes and image complexities. It has the advantages of low complexity, adaptability and strong generalization ability, and provides effective support for improving image processing efficiency and accuracy.
[0182] When applied to classification tasks, most researchers use graph convolution to extract features of irregularly distributed objects, as traditional convolution cannot effectively extract features. However, obtaining the appropriate graph structure required for graph convolution is difficult. This superpixel segmentation method can be used to generate superpixels from intermediate features during network training to adaptively generate homogeneous regions, obtain graph structure, and further generate spatial descriptors as graph nodes. By considering the relationship between descriptors to obtain an adjacency matrix, it meets the necessary conditions for graph convolution.
[0183] When applied to segmentation tasks, if the image size being processed is too large, the computational overhead will increase and the segmentation algorithm will be easily interfered by noise. Traditional pixel-level segmentation methods are also prone to over-segmentation problems. The provided superpixel algorithm can solve these problems well, because superpixels can reduce the number of pixels in the image. Compared with segmenting each pixel, segmenting superpixels can significantly reduce the computational cost of segmentation while maintaining the image structure. Since superpixels tend to merge similar pixels together, they can provide spatial consistency over a larger range, making the segmentation results more continuous while reducing the occurrence of over-segmentation problems. By aggregating local similarities, the influence of individual pixels is reduced, so various noises in the image can be suppressed to a certain extent.
[0184] During experiments, the present invention was compared with the publicly available SLIC algorithm and a superpixel segmentation algorithm using a convolutional neural network with regularized information maximization (hereinafter referred to as Algorithm 2). The parameters used remained consistent with those used by the algorithm developers. The following are the true labels and superpixel segmentation results generated by the segmentation algorithm, respectively. As shown in Figure 6, A, true labels; B, SLIC; C, Algorithm 2; D, superpixel segmentation results of the present invention.
[0185] From the above results, it can be seen that the SLIC algorithm and Algorithm 2 have poor adaptability for superpixel segmentation of large-size images. Among them, the SLIC algorithm has certain under-segmentation problems and poor boundary adhesion, but its algorithm operation efficiency is higher than the segmentation method of the present invention; Algorithm 2 has over-segmentation problems and cannot generate good superpixels at the image boundaries. Although it is unsupervised like the segmentation method of the present invention, its algorithm time complexity is about 13 times that of the segmentation method of the present invention, and the segmentation accuracy generated by the two superpixel segmentation algorithms used for comparison is lower than the superpixel segmentation method of the present invention.
[0186] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. An unsupervised superpixel segmentation method that promotes the cooperation of attention mechanism and dilated spatial pyramid pooling, characterized in that, firstly, preprocess the image to introduce the spatial relationship between pixel points, and for the first time apply the attention mechanism to the superpixel segmentation task to make the model strengthen the attention to important feature channels and suppress the response to irrelevant channels; use dilated spatial pyramid pooling to expand the receptive field while reducing parameters, and combine the constructed loss function to use an optimizer to achieve parameter update and the extraction of final effective depth features. Applying the argmax function to the extracted effective depth features can transform the superpixel segmentation task into a classification problem, and finally achieve adaptive superpixel generation by adding size limit conditions.
2. The unsupervised superpixel segmentation method that promotes the cooperation of attention mechanism and dilated spatial pyramid pooling according to claim 1, characterized in that, the unsupervised superpixel segmentation method that promotes the cooperation of attention mechanism and dilated spatial pyramid pooling includes the following steps: Step 1, combine the RGB channel information of the image with the position information of pixel points to transform the three-dimensional features into five-dimensional features; Step 2, use the attention mechanism to construct a channel attention module; Step 3, use dilated spatial pyramid pooling to process the result of the attention mechanism and extract depth features suitable for superpixel segmentation; Step 4, construct a loss function, first construct a clustering loss term; at the same time, use a spatial smoothness loss term to quantify the difference between adjacent pixels; then construct a reconstruction loss term; Step 5, update the model parameters by setting the learning rate and the number of iterations of the Adam optimizer to find the model parameters that minimize the loss function; Step 6, use the argmax function to obtain the maximum value in the channel dimension, and transform the final effective depth features into a most suitable superpixel label index corresponding to each pixel point; convert the result processed by the argmax function into a two-dimensional array and complete adaptive superpixel segmentation according to the limit conditions in the CPU.
3. The unsupervised superpixel segmentation method that promotes the cooperation of attention mechanism and dilated spatial pyramid pooling according to claim 2, characterized in that, in Step 1, the following formula is used for image preprocessing: In the formula , , Denote the color channel value of the k-th pixel point in the image converted from integer type to floating point type, where , Indicates the row and column numbers where the k-th pixel point in the image is located. represents the five-dimensional feature of the k-th pixel point in the image and is applied to subsequent processing.
4. The unsupervised superpixel segmentation method that promotes the cooperation of attention mechanism and dilated spatial pyramid pooling according to claim 2, characterized in that, the construction method of Step 2 specifically includes: Step 21, apply the preprocessed five-dimensional features to a pointwise convolutional layer, and perform linear combination and transformation on the five-dimensional features while keeping the height and width unchanged to achieve the output of eight-dimensional features; Step 22, calculate the channel global average pooling pair Processing results in aggregated features Step 23, calculate the fast one-dimensional convolution of the aggregated features with a kernel size of L: In the formula, is the sigmoid function, For fast one-dimensional convolution with a kernel size of L, is the learned channel weight; the kernel size L is automatically calculated according to the number of channels: wherein , , is defined as the odd number closest to t. Step 24, representing the result of processing the aggregated feature through a fast one-dimensional convolution with a kernel size of L and the sigmoid function as , and with Performing an element-wise product to obtain the result of the attention mechanism, denoted as , the shape is consistent with the eight-dimensional features of the input attention mechanism; 。 5. The unsupervised superpixel segmentation method that promotes the cooperation of attention mechanism and dilated spatial pyramid pooling according to claim 4, characterized in that, The output of the eight-dimensional features in step 21 is expressed as and is directly applied to channel global average pooling, where H represents the image height, W represents the image width, and C represents the number of channels of the image. Here, C = 8.
6. The unsupervised superpixel segmentation method that promotes the cooperation of attention mechanism and dilated spatial pyramid pooling according to claim 2, characterized in that, In step three, constructing the channel attention module specifically includes: Step 31, represent the depth features obtained after the hole spatial pyramid pooling process as , where H and W remain unchanged and still correspond to the height and width of the image, ; The calculation process of the intermediate convolutional layer is as follows: In the formula, Indicates that the convolution kernel size is , a convolutional layer with a padding size of 0 and a sampling rate of 1; Indicates that the convolution kernel size is , a convolutional layer with a padding size of 2 and a sampling rate of 2; Indicates that the convolution kernel size is , a convolutional layer with a padding size of 4 and a sampling rate of 4; Indicates that the convolution kernel size is , a convolutional layer with a padding size of 6 and a sampling rate of 6; Indicates that the size of the input tensor is adjusted using an adaptive average pooling layer to ; with a step size of 1 Perform convolution with 8 input channels and 16 output channels, normalize each output channel using instance normalization, and finally apply the ReLU activation function to introduce non-linearity; 、 、 、 and is the intermediate tensor; Step 32, depth features The calculation is as follows: In the formula, Indicates that 、 、 、 And Concatenate along the channel dimension to form a larger tensor, 、 、 、 And The output channels are all 16, and the number of channels of the concatenated tensor is 80; in the formula Indicates that the concatenated tensor is to be Convolution, with the output channels set to 128 suitable for the requirements of superpixel segmentation; Instance normalization is used to normalize each output channel; Finally, the ReLU activation function is applied to introduce non-linearity.
7. The attention mechanism collaborative atrous spatial pyramid pooling for promoting unsupervised superpixel segmentation method according to claim 2, characterized in that, In step four, constructing the loss function specifically includes: Step 41, the overall loss function consists of three parts: a clustering loss term, a spatial smoothness loss term, and a reconstruction loss term: In the formula, Represents the overall loss function; Represents the clustering loss term; Represents the spatial smoothing loss term; Represents the reconstruction loss term; With are all constant coefficients and 、 。 Step 42, the calculation method of the clustering loss term is as follows: In the formula, ; Represents the average of the class probability vectors over all pixels; represents the class probability vector of the pixel located at row i and column j, and the calculation method is as follows: Step 43, place The first three channels in the channel dimension are separated for subsequent reconstruction loss term calculation, and the features of the remaining 125 channels are used with For representation, determine Rounding Indicating that the The channel dimension of is transformed into the class probability of the corresponding pixel; Step 44, calculating the spatial smoothing loss term is divided into two parts: Direction and For the smoothness loss in the ; by calculating the absolute value of the probability difference and the image The exponential function of the squared difference of the gradients is defined, and the average spatial smoothness term loss of all pixels is calculated. The specific calculation method is as follows: In the formula, Indicates that the channel and clustering losses are consistent; Expressed as Pixel probability difference in the direction; Denoted as The pixel intensity difference in the Denoted as Pixel probability difference in the Denoted as The pixel intensity difference in the direction; The specific calculation method is as follows: Step 45, the calculation method of the reconstruction loss term is as follows: When calculating the clustering loss term and the spatial smoothness loss term, the first three channels separated from the deep features are used for image reconstruction and are denoted as ; represents the selection of the 2-norm; The dimension of the original input image in step 44 is , image Dimension is 。 8. The attention mechanism collaborative atrous spatial pyramid pooling for promoting unsupervised superpixel segmentation method according to claim 2, characterized in that, In step five, the depth features are extracted under the minimized model parameters. These depth features are effective depth features, and the effective depth features separated last time are used for superpixel generation.
9. The attention mechanism collaborative atrous spatial pyramid pooling for promoting unsupervised superpixel segmentation method according to claim 2, characterized in that, In step six, the size limitation condition for superpixel generation is calculated as follows: In the formula, Represents the average size of an ideal superpixel; Represents the total number of pixels in the image; With are the thresholds for limiting the minimum and maximum of the superpixels, respectively, used to filter out superpixels that are too small or too large so that the sizes of the finally generated superpixels are as uniform as possible; the hyperparameters in the formula 。 10. The attention mechanism collaborative atrous spatial pyramid pooling for promoting unsupervised superpixel segmentation system of the attention mechanism collaborative atrous spatial pyramid pooling for promoting unsupervised superpixel segmentation method according to any one of claims 1 to 9, characterized in that, This system includes: An image preprocessing module, used to combine the RGB channel information of the image with the position information of the pixel points, and transform the three-dimensional features into five-dimensional features; An attention mechanism module, used to construct a channel attention module using the attention mechanism; An atrous spatial pyramid pooling module, used to process the results of the attention mechanism using atrous spatial pyramid pooling and extract depth features suitable for superpixel segmentation; A loss function construction module, used to construct a loss function, first constructing a clustering loss term; At the same time, the spatial smoothness loss term is used to quantify the difference between adjacent pixels; Then construct the reconstruction loss term; A parameter update module, used to update the parameters of the model by setting the learning rate and the number of iterations of the Adam optimizer; A superpixel segmentation module, used to convert the processing result of the argmax function into a two-dimensional array and complete adaptive superpixel segmentation on the CPU according to the limitation conditions.
Citation Information
Patent Citations
Full-view adaptive segmentation network configuration method based on lump differentiation classification
CN112241954A
Farmland crop identification method based on fusion of semantic segmentation and superpixel segmentation
CN114067219A
Non-contact belt tearing detection system and method based on image segmentation
CN114772208A
Hyperspectral classification identification method based on superpixel segmentation
CN115937685A
Method for salient object segmentation of image by aggregating multi-linear exemplar regressors
US20180204088A1
Cited By
Oil displacement rate prediction method based on oil reservoir displacement etching image
CN115204456A
An oil displacement rate prediction method based on reservoir displacement etching image
CN115204456B
Deepwater image enhancement method based on multi-color space coupling
CN120410953A
Railway scene image enhancement method based on image feature aggregation model
CN120510403A
Deep steganography image secret information blind extraction method based on self-supervised learning
CN120543358A