Image segmentation method and device, equipment, storage medium and computer program product
By grouping and fusing the pre-defined feature channels of remote sensing images and combining them with a spatial attention weight matrix, the problem of insufficient segmentation accuracy in remote sensing images is solved, and higher segmentation accuracy is achieved.
Patent Information
- Application Number
- CN202510894565.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies lack sufficient segmentation accuracy when dealing with remote sensing images where objects are scattered and have large scale variations, and fail to fully consider the differences in spatial information across different channels.
By grouping the preset feature channels of remote sensing images into channels, fusing the feature images within the target channel group, determining the spatial attention weight matrix, generating a weighted feature map, and finally performing image segmentation.
It improves the accuracy of remote sensing image segmentation by taking into account the differences in spatial information of different channels, thereby enhancing segmentation precision.
Smart Images

Figure CN120876848A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image segmentation technology, and in particular to an image segmentation method, apparatus, device, storage medium, and computer program product. Background Technology
[0002] Image segmentation is a key technology in computer vision. It subdivides an image into multiple specific parts or objects to facilitate subsequent analysis and understanding. Its main purpose is to separate target objects or regions from the background in an image, thereby achieving more accurate recognition and understanding of the image content. Image segmentation technology can be applied in various fields, such as medical image analysis, remote sensing image processing, autonomous driving, industrial inspection, video surveillance, and facial recognition.
[0003] Traditionally, image segmentation typically involves calculating attention weights across all channels to weight the feature maps, and then using these weighted feature maps for segmentation. However, when dealing with remote sensing images where objects are scattered and have large scale variations, this approach fails to adequately consider the differences in spatial information across different channels, resulting in insufficient accuracy in remote sensing image segmentation. Summary of the Invention
[0004] The main objective of this application is to provide an image segmentation method that addresses the technical problem of insufficient segmentation accuracy in existing technologies when dealing with remote sensing images.
[0005] To achieve the above objectives, this application proposes an image segmentation method, the method comprising:
[0006] Based on at least two preset feature channels, feature extraction is performed on the remote sensing image to be segmented to obtain the feature image corresponding to each preset feature channel;
[0007] Each preset feature channel is grouped to obtain a corresponding target channel group, and the feature images of each preset feature channel in the target channel group are fused to obtain the target feature image of each target channel group.
[0008] Determine the spatial attention weight matrix for each target feature image, and obtain a weighted feature map based on each target feature image and the corresponding spatial attention weight matrix;
[0009] The weighted feature maps are stitched together to obtain a comprehensive feature map, and the remote sensing image to be segmented is then segmented based on the comprehensive feature map.
[0010] In one embodiment, the step of grouping each of the preset feature channels to obtain the corresponding target channel group includes:
[0011] The correlation between each preset feature channel is calculated to obtain the channel correlation between each preset feature channel.
[0012] Based on the channel correlation, at least one correlation threshold is determined, and each of the preset feature channels is grouped based on the correlation threshold and the channel correlation to obtain the corresponding target channel group.
[0013] In one embodiment, the step of determining the spatial attention weight matrix of each target feature image and obtaining a weighted feature map based on each target feature image and the corresponding spatial attention weight matrix includes:
[0014] Global average pooling is performed on each of the target feature images to obtain the target feature matrix corresponding to each target feature image;
[0015] The target feature matrix is normalized based on a preset normalization coefficient to obtain the spatial attention weight matrix corresponding to each target feature matrix. The preset normalization coefficient is obtained by adjusting the initial normalization coefficient using sample tasks and sample datasets.
[0016] The weighted feature map corresponding to each target feature image is obtained by multiplying each target feature matrix and the corresponding spatial attention weight matrix element by element.
[0017] In one embodiment, the step of performing global average pooling on each of the target feature images to obtain the target feature matrix corresponding to each target feature image includes:
[0018] Global average pooling is performed on each of the target feature images to obtain a global average pooling result, and global max pooling is performed on each of the target feature images to obtain a global max pooling result.
[0019] The target feature matrix corresponding to each target feature image is obtained based on the preset balance parameters, the global average pooling result, and the global max pooling result.
[0020] In one embodiment, the step of concatenating the weighted feature maps to obtain a comprehensive feature map includes:
[0021] Obtain the dimension parameters corresponding to each of the weighted feature maps, and determine the splicing axis based on the dimension parameters and the number of channels in each of the target channel groups;
[0022] The weighted feature maps are spliced together according to the splicing axis to obtain a comprehensive feature map.
[0023] In one embodiment, the step of performing image segmentation on the remote sensing image to be segmented based on the comprehensive feature map includes:
[0024] The comprehensive feature map is decoded based on a preset query vector and a preset decoder to obtain a segmentation mask;
[0025] The segmentation mask is classified into categories, and the remote sensing image to be segmented is segmented based on the category classification results and the segmentation mask.
[0026] Furthermore, to achieve the above objectives, this application also proposes an image segmentation apparatus, the apparatus comprising:
[0027] The feature extraction module is used to extract features from the remote sensing image to be segmented based on at least two preset feature channels, and obtain the feature image corresponding to each preset feature channel;
[0028] The feature fusion module is used to group each of the preset feature channels to obtain the corresponding target channel group, and to fuse the feature images of each of the preset feature channels in the target channel group to obtain the target feature image of each target channel group.
[0029] The feature weighting module is used to determine the spatial attention weight matrix of each target feature image, and to obtain a weighted feature map based on each target feature map and the corresponding spatial attention weight matrix.
[0030] The image segmentation module is used to stitch together the weighted feature maps to obtain a comprehensive feature map, and to perform image segmentation on the remote sensing image to be segmented based on the comprehensive feature map.
[0031] In addition, to achieve the above objectives, this application also proposes an image segmentation apparatus, the apparatus comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the image segmentation method as described above.
[0032] In addition, to achieve the above objectives, this application also proposes a storage medium that is a computer-readable storage medium, on which a computer program is stored, which, when executed by a processor, implements the steps of the image segmentation method described above.
[0033] In addition, to achieve the above objectives, this application also proposes a computer program product comprising a computer program that, when executed by a processor, implements the steps of the image segmentation method described above.
[0034] This application proposes an image segmentation method, apparatus, device, storage medium, and computer program product. The method includes: extracting features from a remote sensing image to be segmented based on at least two preset feature channels to obtain feature images corresponding to each preset feature channel; grouping each preset feature channel to obtain corresponding target channel groups, and fusing the feature images of each preset feature channel within the target channel group to obtain target feature images of each target channel group; determining the spatial attention weight matrix of each target feature image, and obtaining weighted feature maps based on each target feature image and the corresponding spatial attention weight matrix; stitching together the weighted feature maps to obtain a comprehensive feature map, and performing image segmentation on the remote sensing image to be segmented based on the comprehensive feature map. This application, when segmenting remote sensing images, groups preset feature channels, fuses the feature images within the target channel groups to obtain target feature images for each target channel group, weights each target feature image, and finally fuses the weighted target feature images according to preset feature channels. The fused feature map is then used for image segmentation. Compared to existing methods that use full-channel attention calculation to obtain attention weights, this approach takes into account the differences in spatial information across different channels, thus improving the accuracy of remote sensing image segmentation. Attached Figure Description
[0035] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0036] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a flowchart of the first embodiment of the image segmentation method proposed in this application;
[0038] Figure 2 This is a diagram illustrating the overall architecture of the image segmentation method proposed in this application.
[0039] Figure 3 This is a structural diagram of the GS-FPN module in the image segmentation method proposed in the embodiments of this application;
[0040] Figure 4 This is a structural diagram of the SENet model in the image segmentation method proposed in the embodiments of this application;
[0041] Figure 5This is a flowchart of a second embodiment of the image segmentation method proposed in this application;
[0042] Figure 6 This is a flowchart of the third embodiment of the image segmentation method proposed in this application;
[0043] Figure 7 A diagram of an image segmentation apparatus provided in an embodiment of this application;
[0044] Figure 8 This is a schematic diagram of the structure of an image segmentation device suitable for implementing the embodiments of this application.
[0045] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0046] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not intended to limit this application.
[0047] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0048] It should be noted that all directional indicators (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.
[0049] Image segmentation is a key technology in computer vision, as it breaks down an image into multiple specific parts or objects to facilitate subsequent analysis and understanding. Its main purpose is to separate target objects or regions from the background, thereby achieving more accurate recognition and understanding of the image content. Image segmentation technology can be applied in various fields, such as medical image analysis, remote sensing image processing, autonomous driving, industrial inspection, video surveillance, and facial recognition.
[0050] Traditionally, image segmentation typically involves calculating attention weights across all channels to weight the feature maps, and then using these weighted feature maps for segmentation. However, when dealing with remote sensing images where objects are scattered and have large scale variations, this approach fails to adequately consider the differences in spatial information across different channels, resulting in insufficient accuracy in remote sensing image segmentation.
[0051] Therefore, in order to solve the technical problem of insufficient segmentation accuracy of existing technologies when dealing with remote sensing images, this embodiment proposes an image segmentation method, which includes: extracting features from the remote sensing image to be segmented based on at least two preset feature channels to obtain feature images corresponding to each preset feature channel; grouping each preset feature channel to obtain corresponding target channel groups, and fusing the feature images of each preset feature channel within the target channel group to obtain target feature images of each target channel group; determining the spatial attention weight matrix of each target feature image, and obtaining weighted feature maps based on each target feature image and the corresponding spatial attention weight matrix; stitching the weighted feature maps together to obtain a comprehensive feature map, and performing image segmentation on the remote sensing image to be segmented based on the comprehensive feature map. In this embodiment, when segmenting the remote sensing image to be segmented, preset feature channels are grouped into channels, and the feature images within the target channel groups are fused to obtain target feature images for each target channel group. Each target feature image is then weighted, and finally, the weighted target feature images are fused according to the preset feature channels. The fused feature images are then used for image segmentation. Compared to the existing method of using full-channel attention calculation to obtain attention weights, this method takes into account the differences in spatial information of different channels, thus improving the accuracy of remote sensing image segmentation.
[0052] For ease of understanding, the following is combined with Figures 1 to 8 The image segmentation method provided in the embodiments of this application, as well as the image segmentation method, apparatus, device, storage medium, and computer program product provided in the following embodiments, will be described in detail.
[0053] This application provides an image segmentation method, referring to... Figure 1 , Figure 1 This is a flowchart of the first embodiment of the image segmentation method proposed in this application.
[0054] like Figure 1 As shown, the method includes:
[0055] Step S10: Extract features from the remote sensing image to be segmented based on at least two preset feature channels to obtain feature images corresponding to each preset feature channel.
[0056] It should be noted that the executing entity in this embodiment can be a multifunctional machine device with image segmentation capabilities, such as an image segmentation device, or a device capable of performing the aforementioned functions. This embodiment uses an image segmentation device (hereinafter referred to as the device) for description.
[0057] It should also be noted that the aforementioned preset feature channels can be channels pre-defined during image processing that can capture specific features of the image, such as spectral information of different bands. For example, remote sensing images typically contain red, green, blue, and near-infrared bands, which can be used as preset feature channels. The aforementioned feature images can be images that retain certain specific feature information from the original image after feature extraction, and are typically used for subsequent image analysis and processing tasks. In specific implementations, the aforementioned device performs feature extraction on the remote sensing image to be segmented using at least two preset feature channels to obtain feature images corresponding to each preset feature channel.
[0058] For ease of understanding, the following examples are provided, but they do not constitute a specific limitation on this embodiment. For instance, remote sensing images typically contain red, green, blue, and near-infrared bands. The aforementioned device can utilize these bands as preset feature channels. Suppose that trees in the remote sensing image have high reflectivity in the near-infrared band, while water has high reflectivity in the blue band. The aforementioned device extracts features from the near-infrared and blue bands respectively, generating corresponding feature images.
[0059] Step S20: Group each of the preset feature channels to obtain the corresponding target channel group, and fuse the feature images of each preset feature channel in the target channel group to obtain the target feature image of each target channel group.
[0060] It should be noted that the aforementioned target feature image can be an image obtained by integrating information from multiple preset feature channels and then performing feature fusion. The aforementioned channel grouping can be a process of grouping multiple preset feature channels according to certain rules or feature similarity. The aforementioned device can employ one or more rules, such as grouping by feature scale, by feature semantic level, by correlation between channels, by task requirements, or by the spatial distribution of features. The aforementioned feature fusion can be a process of comprehensively processing information from multiple feature images to generate a more representative and comprehensive feature image.
[0061] In its implementation, the aforementioned device first extracts features from the remote sensing image based on at least two preset feature channels during image segmentation, obtaining feature images corresponding to each channel. The device then groups the preset feature channels and fuses the feature images of each preset feature channel within each target channel group to generate a target feature image that better reflects the target features.
[0062] For ease of understanding, the following examples illustrate the concept, but do not limit the scope of this embodiment. Taking remote sensing image processing as an example, remote sensing images typically contain multiple spectral bands, such as red, green, blue, and near-infrared bands. The aforementioned device can use the red, green, and blue bands as a set of preset feature channels because they collectively provide rich color information; simultaneously, it can use the near-infrared band and other specific vegetation index bands as another set of preset feature channels because they are more sensitive to vegetation feature identification. After completing the channel grouping, the aforementioned device will fuse the feature images within each channel group. For the color information group, the aforementioned device can use color feature extraction algorithms to extract feature images from the red, green, and blue bands respectively, and then fuse these three feature images using methods such as weighted averaging or principal component analysis to generate a target feature image with comprehensive color features. For the vegetation feature group, the aforementioned device may fuse the feature images of the near-infrared band with the feature images of the vegetation index band, and use algorithms to enhance vegetation features, such as Normalized Difference Vegetation Index (NDVI) calculation, to obtain a target feature image that highlights vegetation information.
[0063] Furthermore, in order to accurately group channels and thus achieve accurate image segmentation, the step of grouping each of the preset feature channels to obtain the corresponding target channel group includes:
[0064] Step S21: Perform correlation calculation on each of the preset feature channels to obtain the channel correlation between each of the preset feature channels.
[0065] It should be noted that the above-mentioned channel correlation can be an indicator of the degree of information association between different feature channels. High correlation means that the information overlap between channels is high.
[0066] Step S22: Determine at least one correlation threshold based on the channel correlation, and group each of the preset feature channels based on the correlation threshold and the channel correlation to obtain the corresponding target channel group.
[0067] It should be noted that the correlation threshold can be a critical value used to distinguish between high and low correlations between channels. In this embodiment, the device can determine one or more values based on the obtained channel correlation to group preset channels. The target channel group can be a channel group obtained by grouping based on the correlation threshold, where the channels within each group have high correlations.
[0068] In its specific implementation, when processing remote sensing images, the aforementioned device performs correlation calculations on each preset feature channel to obtain the channel correlation between them. After obtaining the channel correlation, the device determines at least one correlation threshold based on the channel correlation, and groups each preset feature channel according to the correlation threshold and the channel correlation to obtain the corresponding target channel group.
[0069] Step S30: Determine the spatial attention weight matrix of each target feature image, and obtain a weighted feature map based on each target feature image and the corresponding spatial attention weight matrix.
[0070] It should be noted that the target feature image mentioned above can be a feature image obtained after feature extraction and fusion, used for subsequent image segmentation. The spatial attention weight matrix mentioned above can be a matrix with the same size as the feature map, where each element represents the importance weight of the corresponding spatial location. The weighted feature map mentioned above can be a feature map obtained by multiplying the feature map element-wise with the spatial attention weight matrix, highlighting the features of important spatial locations.
[0071] In its specific implementation, the device first calculates the spatial attention weight matrix of each target feature image during image segmentation, and then generates a weighted feature map by multiplying each target feature image element by element with its corresponding spatial attention weight matrix.
[0072] Step S40: The weighted feature maps are stitched together to obtain a comprehensive feature map, and the remote sensing image to be segmented is segmented based on the comprehensive feature map.
[0073] It should be noted that the aforementioned comprehensive feature map can be obtained by stitching together multiple weighted feature maps. In specific implementation, the device stitches together the weighted feature maps to obtain a comprehensive feature map, and then performs image segmentation on the remote sensing image to be segmented based on this map. Specifically, the device first stitches together the weighted feature maps, a process that merges multiple feature maps with different numbers of channels into a comprehensive feature map containing more information. The stitching operation is usually performed along the channel dimension to ensure that the height and width of the feature map are consistent. For example, assuming there are three weighted feature maps with shapes of 256×256×2, 256×256×3, and 256×256×1, the device can stitch them together along the channel dimension to form a comprehensive feature map of 256×256×6. After obtaining the comprehensive feature map, the device inputs it into a segmentation network, such as a fully convolutional network (FCN) or U-Net. Through the forward propagation of the network, a segmentation result of the same size as the input image is finally generated, with each pixel representing the corresponding category.
[0074] refer to Figure 2 , Figure 2 This is a diagram illustrating the overall architecture of the image segmentation method proposed in this application. In one example, to achieve more accurate image segmentation, the device extracts features of multiple dimensions from the input remote sensing image to be segmented using the backbone module, and performs the aforementioned steps on the feature images of each size to obtain a comprehensive feature map of each size (i.e., ...). Figure 2After the GGSA and FPN modules (as described in the original text), the first three generated feature maps can be sequentially fed into the Transformer Decoder module via the pixel decoder module for decoding, thereby generating a mask and category corresponding to each query. Then, the mask and the last feature map from the pixel decoder are multiplied to obtain the foreground feature map. Each foreground feature map is multiplied by the category to obtain the final segmentation result. In the feature extraction process, a Feature Pyramid Network (FPN) can be used by introducing top-down and lateral connections into the network to extract rich semantic information from feature maps at different levels.
[0075] In addition, it should be noted that the reference Figure 3 , Figure 3 This is a structural diagram of the GS-FPN module in the image segmentation method proposed in this application. This embodiment combines the Global Grouped Space Attention (GGSA) mechanism and the FPN mechanism to form the GS-FPN module.
[0076] like Figure 3 As shown, for a given input feature pyramid in Let i represent the feature map of the i-th stage. This method generates an enhanced feature pyramid with more informational representation by fusing feature maps from different stages. To improve the performance of downstream tasks, feature maps or feature map The resolution is reduced to 1 / 2 of the input image. i This module uses data processed by the backbone module. The feature map is taken as input and passed through multiple GGSA modules to obtain... Then on and The feature map is upsampled and its channel number is adjusted using a 1×1 convolution. and The number of channels, then with and By merging, we obtain and For feature maps and This embodiment uses a skip connection method, connecting it with... and Add them together to get and Its formula is expressed as follows:
[0077]
[0078] The GS-FPN module combines features from different levels to form a multi-scale feature pyramid, which allows high-level features to have richer detailed information and low-level features to have stronger semantic information. It can also help generate more stable and hierarchical feature inputs, thereby improving the performance of the Transformer structure.
[0079] refer to Figure 2 as well as Figure 4 , Figure 4 This is a structural diagram of the SENet model in the image segmentation method proposed in the embodiments of this application. Figure 3 As shown, in order to improve the feature representation capability of convolutional neural networks and thus enhance the accuracy of image segmentation, this embodiment, after obtaining feature maps of different levels through the GS-FPN module, also utilizes the Squeeze-and-Excitation Networks (SENet) module to dynamically adjust the channel weights of the feature maps by explicitly modeling the relationships between channels. Specifically, the SENet module achieves dynamic adjustment of the channel weights of the feature maps through the following three steps.
[0080] Compression (Squeeze) (i.e.) Figure 4 F in aq In this stage, SENet uses global average pooling to aggregate the global features of each channel across the spatial dimension of the feature map, generating a compact channel-level global description. This process enables the network to capture the contribution of each channel to the global features. For a given feature map... The Squeeze formula is as follows:
[0081]
[0082] Where H and W represent the height and width of the feature map. This is the matrix after Squeeze.
[0083] Excitation (i.e.) Figure 4 F in ex In this stage (·, W), the dependencies between channels are modeled explicitly using a bottleneck structure consisting of two fully connected layers and a ReLU nonlinear activation function, generating a set of weights. These weights represent the importance of each channel and are normalized to the range [0,1] using a Sigmoid activation function. The excitation operation formula is shown below:
[0084]
[0085] Where δ is the ReLU activation function and σ is the Sigmoid activation function. and These are the weight parameters of the two fully connected layers, r is a scaling factor, and C represents the number of channels.
[0086] Recalibration in this stage uses the weights from the activation stage to recalibrate the input feature map. The recalibration formula is shown below:
[0087]
[0088] in This represents the calibrated feature map, s C This represents the weight of the c-th channel.
[0089] In addition, refer to Figure 2 Since feature maps at different time points may contain different details, it is necessary to fuse these feature maps to enhance feature representation capabilities. To this end, this embodiment proposes an interactive multi-scale three-time-channel fusion module (TCFM). Based on this TCFM module, for feature maps... Where B represents the batch size, C represents the number of channels, and H and W represent the height and width of the feature map, respectively. The TCFM module first sums the input feature maps. Then, through a Squeeze operation, it aggregates the spatial dimensions of the feature maps to obtain the global features for each channel. Next, through an Excitation operation, it models the global features for each channel, generating a set of weights. Finally, through a Recalibration operation, it recalibrates the weights onto the input feature map to obtain the final feature map. The TCFM module introduces a channel attention mechanism, allowing the module to focus only on the relationships between channels during computation, thereby enhancing the network's ability to focus on important features and improving feature representation capabilities. Its formula is shown below:
[0090]
[0091] Among them is Dimensionality reduction weight matrix (r is the reduction rate), W1 and W2 are learnable parameters, GAP represents global average pooling, C is the number of channels, σ represents the Sigmoid activation function, and split represents the grouping operation of the input feature map.
[0092] The TCFM module fuses three quarter-feature maps from three different time points—the backbone processed by GGSA and SENet, the pixel decoder's quarter-feature map, and the pixel decoder's quarter-feature map processed by SENet—into a more robust feature representation through an adaptive weighting mechanism. Compared to the original model, this multi-time-point channel fusion strategy captures richer feature information. By introducing an adaptive weighting mechanism, TCFM dynamically adjusts the feature map at each time point, ensuring that each feature is appropriately represented in the final fusion. This weighting mechanism is implemented through fully connected layers and a sigmoid activation function, enabling the learning of complex nonlinear feature interactions. This module uses only simple global average pooling and a lightweight fully connected network, ensuring efficient feature fusion without introducing excessive computational overhead and parameter count. This improvement effectively focuses attention on more precise locations, thereby further optimizing the model's segmentation performance.
[0093] In this embodiment, when segmenting a remote sensing image to be segmented, preset feature channels are grouped into channels, and the feature images within the target channel groups are fused to obtain target feature images for each target channel group. Each target feature image is then weighted, and finally, the weighted target feature images are fused according to the preset feature channels. The fused feature images are then used for image segmentation. Compared to using full-channel attention calculation to obtain attention weights, this method takes into account the differences in spatial information of different channels, thus improving the accuracy of remote sensing image segmentation.
[0094] Based on the first embodiment, in the second embodiment, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 5 , Figure 5 This is a flowchart of a second embodiment of the image segmentation method proposed in this application. Further, to achieve more accurate image segmentation, the step of determining the spatial attention weight matrix of each target feature image and obtaining a weighted feature map based on each target feature image and the corresponding spatial attention weight matrix includes:
[0095] Step S31: Perform global average pooling on each of the target feature images to obtain the target feature matrix corresponding to each target feature image;
[0096] It should be noted that the global average pooling mentioned above can be a feature dimensionality reduction technique, which compresses the feature map into a feature vector of a fixed size through global average pooling.
[0097] In its specific implementation, the device performs average pooling on each channel of each target feature image, calculates the average value of pixel values at all spatial locations, obtains a one-dimensional feature vector, and combines the feature vectors to form a target feature matrix.
[0098] Furthermore, in order to obtain more accurate target feature maps, the step of performing global average pooling on each of the target feature images to obtain the target feature matrix corresponding to each target feature image includes:
[0099] Step S311: Perform global average pooling on each of the target feature images to obtain the global average pooling result, and perform global max pooling on each of the target feature images to obtain the global max pooling result.
[0100] It should be noted that the global max pooling result mentioned above can be the maximum pixel value found for each channel of the feature map. In the specific implementation, the device first performs global average pooling on each target feature image to obtain the global average pooling result corresponding to each target feature image, and then performs global max pooling on each target feature image to obtain the global max pooling result corresponding to each target feature image.
[0101] For ease of understanding, the following example is used, but it does not impose specific limitations on this embodiment. Assume the above device processes a target feature image with a size of 2×2×2, and the feature value F is as follows:
[0102]
[0103] The above device performs global average pooling on each channel to calculate the average pixel value across all spatial locations for each channel:
[0104] Channel1: (0.1+0.2+0.3+0.4) / 4=0.25;
[0105] Channel2: (0.5+0.6+0.7+0.8) / 4=0.65;
[0106] The obtained global average pooling result Mavg is:
[0107] Mavg = [0.25, 0.65];
[0108] The above device performs global max pooling on each channel to calculate the maximum pixel value across all spatial locations for each channel:
[0109] Channel'1: max(0.1, 0.2, 0.3, 0.4)=0.4;
[0110] Channel'2: max(0.5, 0.6, 0.7, 0.8)=0.8;
[0111] The global max pooling result Mmax is:
[0112] Mmax = [0.4, 0.8].
[0113] Step S312: Obtain the target feature matrix corresponding to each target feature image based on the preset balance parameters, the global average pooling result, and the global max pooling result.
[0114] It should be noted that the aforementioned target feature matrix can be a matrix obtained by fusing the results of global average pooling and global max pooling. The aforementioned preset balance parameter can control the weights of the global average pooling result and the global max pooling result in the fusion process. In specific implementation, after the device performs global average pooling and global max pooling on the target feature image, it will perform weighted fusion of the two results based on the preset balance parameter using a weighted fusion formula to generate the target feature matrix. The weighted fusion formula is Mfinal = alpha × Mavg + (1 - alpha) × Mmax, where alpha is the balance parameter.
[0115] For ease of understanding, the following examples illustrate the concept, but do not impose specific limitations on this embodiment. For instance, if the balancing parameter alpha is set to 0.7, the global average pooling result Mavg is [0.25, 0.65], and the global max pooling result Mmax is [0.4, 0.8], then the target feature matrix Mfinal is:
[0116] Mfinal = 0.7 × Mavg + 0.3 × Mmax
[0117] =[0.70.25+0.30.4,0.70.65+0.30.8]=[0.225,0.625。
[0118] Step S32: Normalize the target feature matrix based on the preset normalization coefficient to obtain the spatial attention weight matrix corresponding to each target feature matrix. The preset normalization coefficient is obtained by adjusting the initial normalization coefficient using sample tasks and sample datasets.
[0119] It should be noted that the aforementioned preset normalization coefficients can be parameters used to adjust the weight distribution during the normalization process, obtained by adjusting the initial normalization coefficients using sample tasks and sample datasets. The aforementioned sample tasks and sample datasets can be tasks and datasets used for training and adjusting model parameters. The aforementioned initial normalization coefficients can be preset initial parameters used to adjust the weight distribution during the normalization process.
[0120] In its implementation, the aforementioned device adjusts the initial normalization coefficients using sample tasks and sample datasets to optimize segmentation performance. After obtaining the preset normalization coefficients, the target feature matrix is normalized based on these coefficients to obtain the spatial attention weight matrix.
[0121] Furthermore, it should be noted that matrix normalization typically uses the Softmax function. The formula for the Softmax function is:
[0122]
[0123] Among them, z i is the original output score corresponding to category i, and K is the total number of categories. The normalization coefficient is an exponential function used to amplify numerical differences. The denominator is the sum of the exponents of all categories, used to normalize the output, ensuring their sum equals 1, thus forming a probability distribution. The normalization coefficient is adjusted based on the sample task and dataset, with an initial value typically of 1.0, and dynamically adjusted during training based on the performance on the validation set. For example, if the segmentation performance on the validation set is poor, the normalization coefficient may need to be adjusted to balance the weight distribution. A smaller normalization coefficient results in a more concentrated weight distribution, highlighting salient features; a larger normalization coefficient results in a smoother weight distribution, avoiding overfitting. The aforementioned device improves image segmentation accuracy by continuously adjusting the normalization coefficient to optimize the generation of the spatial attention weight matrix.
[0124] For ease of understanding, the following example illustrates the concept, but does not impose any specific limitations on this embodiment. For instance, when processing a target feature matrix M = [0.25, 0.65], assuming a preset normalization coefficient τ = 0.5, the normalized spatial attention weight matrix A is calculated as follows:
[0125] A = Softax([0.5, 1.3]);
[0126] Calculate the exponent value for each element:
[0127] e 0.5 ≈1.6487, e 1.3 ≈3.6693;
[0128] Summation:
[0129] S = 1.6487 + 3.6693 = 5.3180;
[0130] The normalized weight matrix is as follows:
[0131]
[0132] In this example, the device improves the accuracy of image segmentation by adjusting the normalization coefficients so that the spatial attention weight matrix can better highlight important features.
[0133] Step S33: Multiply each of the target feature matrices and the corresponding spatial attention weight matrices element by element to obtain the weighted feature map corresponding to each target feature image.
[0134] In its implementation, the device performs element-wise multiplication of each target feature matrix with its corresponding spatial attention weight matrix. By multiplying the target feature matrix with the spatial attention weight matrix element-wise, a weighted feature map is generated, in which each feature value is multiplied by its corresponding weight, thereby highlighting important features and suppressing unimportant features.
[0135] For ease of understanding, the following examples are provided, but they do not impose specific limitations on this embodiment. For example, assuming the target feature matrix M = [0.25, 0.65] and the spatial attention weight matrix A = [0.31, 0.69], the above device calculates the weighted feature map W by multiplying element by element:
[0136] W=M⊙A=[0.25×0.31,0.65×0.69]≈[0.0775,0.4485].
[0137] Based on the first and second embodiments, in the third embodiment, the content that is the same as or similar to that in Embodiments 1 and 2 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 6 , Figure 6 This is a flowchart of the third embodiment of the image segmentation method proposed in this application. Further, the step of concatenating the weighted feature maps to obtain a comprehensive feature map includes:
[0138] Step S41: Obtain the dimension parameters corresponding to each weighted feature map, and determine the splicing axis based on the dimension parameters and the number of channels in each target channel group.
[0139] It should be noted that the aforementioned dimensional parameters can be the size parameters of the feature map, including height, width, and number of channels, used to describe the size and channel information of the feature map. The aforementioned stitching axis can be the dimension axis selected during the stitching operation, used to determine along which dimension the feature maps are stitched.
[0140] In its implementation, when stitching weighted feature maps, the device first obtains the dimensional parameters corresponding to each weighted feature map, and determines the stitching axis based on these dimensional parameters and the number of channels in each target channel group. The dimensional parameters include the height, width and number of channels of the feature map. The device checks these parameters to ensure that all feature maps have consistent dimensions on the selected stitching axis, thus avoiding stitching errors.
[0141] For ease of understanding, the following examples are provided, but they do not impose specific limitations on this embodiment. Assume the above device has the following three weighted feature maps:
[0142] The dimension parameters of feature map W1 are 256×256×3;
[0143] The dimension parameters of feature map W2 are 256×256×5;
[0144] The dimension parameters of feature map W3 are 256×256×2;
[0145] The device first obtains the dimensional parameters of each feature map, finding that their height and width are both 256×256, but the number of channels are 3, 5, and 2 respectively. Based on this information, the device determines the splicing axis as the channel axis (usually the third dimension) and splices the feature maps along this axis to obtain a composite feature map with dimensions of 256×256×10.
[0146] Step S42: Segment the weighted feature maps according to the splicing axis to obtain a comprehensive feature map.
[0147] In its implementation, the aforementioned device stitches together the weighted feature maps according to a stitching axis to obtain a composite feature map. The device merges feature maps along the selected stitching axis, ensuring that the dimensions of the merged feature map on that axis are added together, while other dimensions remain unchanged. For example, if the stitching axis is a channel axis, the device adds the number of channels in each feature map to generate a composite feature map.
[0148] Further, the step of performing image segmentation on the remote sensing image to be segmented based on the comprehensive feature map includes:
[0149] Step S43: Decode the comprehensive feature map based on the preset query vector and preset decoder to obtain the segmentation mask;
[0150] Step S44: Classify the segmentation mask by category, and perform image segmentation on the remote sensing image to be segmented based on the category classification result and the segmentation mask.
[0151] It should be noted that the aforementioned preset query vector can be a fixed set of vectors used to guide the decoder to focus on different regions or objects in the image. The aforementioned preset decoder can be a neural network module used to convert the synthesized feature map into a segmentation mask. The aforementioned segmentation mask can be a two-dimensional matrix of the same size as the input image, used to indicate the category of each pixel in the image.
[0152] In its implementation, the aforementioned device decodes the comprehensive feature map based on a preset query vector and a preset decoder to obtain a segmentation mask. First, the device inputs the comprehensive feature map into the preset decoder. The decoder generates a segmentation mask corresponding to each query vector through the interaction between the query vector and the feature map. The device then performs pixel-by-pixel classification on each segmentation mask to generate a category classification result. Finally, based on the category classification result and the segmentation mask, the device performs image segmentation on the original remote sensing image to generate the final segmentation result, where each pixel is labeled with its corresponding category.
[0153] For ease of understanding, the following example illustrates the concept, but does not limit the scope of this embodiment. Assume the device processes a composite feature map with a shape of 256×256×64. A preset decoder contains 100 preset query vectors, each with a dimension of 64. The device inputs the composite feature map into the preset decoder, which generates 100 segmentation masks through the interaction between the query vectors and the feature map. Each segmentation mask has a shape of 256×256 and indicates the category of each pixel in the image. The device then performs pixel-by-pixel classification on each segmentation mask, generating category classification results. For example, each pixel value in the segmentation mask represents the probability that the pixel belongs to a certain category. Based on these probabilities, the device assigns each pixel to a specific category. Finally, based on the category classification results and the segmentation masks, the device performs image segmentation on the original remote sensing image, generating the final segmentation result. For example, the segmentation result may include multiple categories, such as trees, buildings, and water bodies, each corresponding to a different color or label.
[0154] This embodiment also provides a first embodiment of an image segmentation device, please refer to... Figure 7 , Figure 7 This is a diagram of an image segmentation apparatus provided in an embodiment of this application. The image segmentation apparatus includes:
[0155] The feature extraction module is used to extract features from the remote sensing image to be segmented based on at least two preset feature channels, and obtain the feature image corresponding to each preset feature channel;
[0156] The feature fusion module is used to group each of the preset feature channels to obtain the corresponding target channel group, and to fuse the feature images of each of the preset feature channels in the target channel group to obtain the target feature image of each target channel group.
[0157] The feature weighting module is used to determine the spatial attention weight matrix of each target feature image, and to obtain a weighted feature map based on each target feature map and the corresponding spatial attention weight matrix.
[0158] The image segmentation module is used to stitch together the weighted feature maps to obtain a comprehensive feature map, and to perform image segmentation on the remote sensing image to be segmented based on the comprehensive feature map;
[0159] The feature fusion module further performs correlation calculation on each of the preset feature channels to obtain the channel correlation between each of the preset feature channels; determines at least one correlation threshold based on the channel correlation, and groups each of the preset feature channels based on the correlation threshold and the channel correlation to obtain the corresponding target channel group.
[0160] Referring to the first embodiment of the image segmentation device, this embodiment also proposes a second embodiment of the image segmentation device. The contents that are the same as or similar to those in the first embodiment of the image segmentation device can be referred to the above description, and will not be repeated hereafter.
[0161] The feature weighting module is further configured to perform global average pooling on each of the target feature images to obtain a target feature matrix corresponding to each target feature image; normalize the target feature matrix based on a preset normalization coefficient to obtain a spatial attention weight matrix corresponding to each target feature matrix, wherein the preset normalization coefficient is obtained by adjusting the initial normalization coefficient using sample tasks and sample datasets; and perform element-wise multiplication of each target feature matrix and the corresponding spatial attention weight matrix to obtain a weighted feature map corresponding to each target feature image.
[0162] The feature weighting module is further configured to perform global average pooling on each of the target feature images to obtain a global average pooling result, and to perform global max pooling on each of the target feature images to obtain a global max pooling result; and to obtain a target feature matrix corresponding to each of the target feature images based on a preset balance parameter, the global average pooling result, and the global max pooling result.
[0163] Referring to the first embodiment and the second embodiment of the image segmentation device, this embodiment also proposes a third embodiment of the image segmentation device. The contents that are the same as or similar to the first embodiment and the second embodiment of the image segmentation device can be referred to the above description, and will not be repeated hereafter.
[0164] The image segmentation module is further configured to obtain the dimension parameters corresponding to each of the weighted feature maps, determine the splicing axis based on the dimension parameters and the number of channels in each of the target channel groups, and splice each of the weighted feature maps according to the splicing axis to obtain a comprehensive feature map;
[0165] The image segmentation module is further configured to decode the comprehensive feature map based on a preset query vector and a preset decoder to obtain a segmentation mask; classify the segmentation mask by category; and perform image segmentation on the remote sensing image to be segmented based on the category classification result and the segmentation mask.
[0166] The image segmentation apparatus provided in this embodiment employs the image segmentation method described in the above embodiments, which can solve the technical problem of insufficient segmentation accuracy in existing technologies when dealing with remote sensing images. Compared with the prior art, the beneficial effects of the image segmentation apparatus provided in this embodiment are the same as those of the image segmentation method described in the above embodiments, and other technical features in the image segmentation apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0167] This embodiment provides an image segmentation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the image segmentation method in the first embodiment described above.
[0168] The following is for reference. Figure 8 , Figure 8 This is a schematic diagram of the structure of an image segmentation device suitable for implementing the embodiments of this application. The image segmentation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 8 The image segmentation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0169] like Figure 8As shown, the image segmentation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the image segmentation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the image segmentation device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows image segmentation devices with various systems, it should be understood that it is not required to implement or possess all of the systems shown. More or fewer systems may be implemented alternatively.
[0170] Specifically, according to this embodiment, the process described above with reference to the flowchart can be implemented as a computer software program. For example, this embodiment includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the disclosed embodiments of this embodiment.
[0171] The image segmentation device provided in this embodiment, employing the image segmentation method described in the above embodiments, can solve the technical problem of insufficient segmentation accuracy in existing technologies when dealing with remote sensing images. Compared with the prior art, the beneficial effects of the image segmentation device provided in this embodiment are the same as those of the image segmentation method described in the above embodiments, and other technical features in this image segmentation device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0172] It should be understood that the various parts disclosed in this embodiment can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0173] The above description is merely a specific implementation of this embodiment, but the protection scope of this embodiment is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this embodiment should be included within the protection scope of this embodiment. Therefore, the protection scope of this embodiment should be determined by the protection scope of the claims.
[0174] This embodiment provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the image segmentation method described in the above embodiment.
[0175] The computer-readable storage medium provided in this embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0176] The aforementioned computer-readable storage medium may be included in the image segmentation device; or it may exist independently and not be assembled into the image segmentation device.
[0177] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the image segmentation device, cause the image segmentation device to perform image segmentation.
[0178] Computer program code for performing the operations of this embodiment can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0179] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this embodiment. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0180] The modules described in this embodiment can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0181] The readable storage medium provided in this embodiment is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-described image segmentation method, and can solve the technical problem of insufficient segmentation accuracy in the prior art when dealing with remote sensing images. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this embodiment are the same as the beneficial effects of the image segmentation method provided in the above embodiments, and will not be repeated here.
[0182] The above descriptions are only some embodiments and do not limit the patent scope of this embodiment. All equivalent structural transformations made based on the technical concept of this application and the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included within the patent protection scope of this application.
Claims
1. An image segmentation method, characterized in that, The method includes: Based on at least two preset feature channels, feature extraction is performed on the remote sensing image to be segmented to obtain the feature image corresponding to each preset feature channel; Each preset feature channel is grouped to obtain a corresponding target channel group, and the feature images of each preset feature channel in the target channel group are fused to obtain the target feature image of each target channel group. Determine the spatial attention weight matrix for each target feature image, and obtain a weighted feature map based on each target feature image and the corresponding spatial attention weight matrix; The weighted feature maps are stitched together to obtain a comprehensive feature map, and the remote sensing image to be segmented is then segmented based on the comprehensive feature map.
2. The method as described in claim 1, characterized in that, The step of grouping each of the preset feature channels to obtain the corresponding target channel group includes: The correlation between each of the preset feature channels is calculated to obtain the channel correlation between each of the preset feature channels. Based on the channel correlation, at least one correlation threshold is determined, and each of the preset feature channels is grouped based on the correlation threshold and the channel correlation to obtain the corresponding target channel group.
3. The method as described in claim 1, characterized in that, The step of determining the spatial attention weight matrix of each target feature image and obtaining a weighted feature map based on each target feature image and the corresponding spatial attention weight matrix includes: Global average pooling is performed on each of the target feature images to obtain the target feature matrix corresponding to each target feature image; The target feature matrix is normalized based on a preset normalization coefficient to obtain the spatial attention weight matrix corresponding to each target feature matrix. The preset normalization coefficient is obtained by adjusting the initial normalization coefficient using sample tasks and sample datasets. The weighted feature map corresponding to each target feature image is obtained by multiplying each target feature matrix and the corresponding spatial attention weight matrix element by element.
4. The method as described in claim 3, characterized in that, The step of performing global average pooling on each of the target feature images to obtain the target feature matrix corresponding to each target feature image includes: Global average pooling is performed on each of the target feature images to obtain a global average pooling result, and global max pooling is performed on each of the target feature images to obtain a global max pooling result. The target feature matrix corresponding to each target feature image is obtained based on the preset balance parameters, the global average pooling result, and the global max pooling result.
5. The method as described in claim 1, characterized in that, The step of concatenating the weighted feature maps to obtain a comprehensive feature map includes: Obtain the dimension parameters corresponding to each of the weighted feature maps, and determine the splicing axis based on the dimension parameters and the number of channels in each of the target channel groups; The weighted feature maps are spliced together according to the splicing axis to obtain a comprehensive feature map.
6. The method as described in claim 1, characterized in that, The step of performing image segmentation on the remote sensing image to be segmented based on the comprehensive feature map includes: The comprehensive feature map is decoded based on a preset query vector and a preset decoder to obtain a segmentation mask; The segmentation mask is classified into categories, and the remote sensing image to be segmented is segmented based on the category classification results and the segmentation mask.
7. An image segmentation apparatus, characterized in that, The device includes: The feature extraction module is used to extract features from the remote sensing image to be segmented based on at least two preset feature channels, and obtain the feature image corresponding to each preset feature channel; The feature fusion module is used to group each of the preset feature channels to obtain the corresponding target channel group, and to fuse the feature images of each of the preset feature channels in the target channel group to obtain the target feature image of each target channel group. The feature weighting module is used to determine the spatial attention weight matrix of each target feature image, and to obtain a weighted feature map based on each target feature map and the corresponding spatial attention weight matrix. The image segmentation module is used to stitch together the weighted feature maps to obtain a comprehensive feature map, and to perform image segmentation on the remote sensing image to be segmented based on the comprehensive feature map.
8. An image segmentation device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the image segmentation method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the image segmentation method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the image segmentation method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image segmentation method, system and device and storage medium
CN115170934A
Remote sensing image road segmentation method fusing multi-scale features and double attention mechanism
CN117078943A
Farmland drainage ditch remote sensing image semantic segmentation method based on improved VM-UNet model
CN119625327A
Remote sensing image target segmentation method and network based on expansion multi-scale fusion
CN119992100A
Weakly supervised semantic segmentation method and apparatus based on attention mask
WO2025060272A1
Cited By
Image enhancement method and device, equipment and storage medium
CN121391658A