SAR image segmentation model training method, image segmentation method, device and equipment
By combining edge enhancement and superpixel segmentation with convolutional neural networks and graph convolutional networks, the problem of insufficient boundary recognition in SAR image segmentation is solved, and efficient boundary clarity and regional-level semantic association are achieved, which is suitable for the rapid processing of high-resolution SAR images.
Patent Information
- Application Number
- CN202511006055.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-22
AI Technical Summary
Existing SAR image segmentation methods are insufficient in terms of boundary recognition and regional consistency, and have problems such as edge blur, transition smoothness and semantic aliasing. In addition, CNN has difficulty in modeling complex non-local or regional-level semantic relationships between objects.
Through edge enhancement processing, an edge intensity map is generated, superpixel segmentation is performed, and feature extraction and semantic reasoning are performed by combining the initial convolutional neural network and the graph convolutional network. A regional level graph structure is established, and model training is performed to improve segmentation accuracy and boundary clarity.
It significantly improves the structural clarity and integrity of the segmentation results, enhances the perception of target edges, improves the classification consistency and robustness under complex backgrounds, reduces the computational burden and memory consumption, and is suitable for the rapid processing of high-resolution SAR images.
Smart Images

Figure CN120510389B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a SAR image segmentation model training method, an image segmentation method, a device and equipment. Background Art
[0002] Currently, SAR image segmentation usually adopts pixel-level semantic segmentation methods based on Convolutional Neural Network (CNN). This type of method usually includes: extracting multi-scale image features through the CNN backbone network, and then using a pixel-level classifier to perform semantic prediction on each pixel.
[0003] However, due to the significant speckle noise and unstructured texture in SAR images, traditional pixel-level methods still lack the ability to identify boundaries and maintain regional consistency, often suffering from edge blurring, transition smoothing, and semantic aliasing. Furthermore, CNNs are essentially local perception models based on Euclidean structures, making them difficult to model more complex non-local or regional semantic relationships between objects. Summary of the Invention
[0004] The purpose of the present invention is to address the deficiencies in the above-mentioned prior art and provide a SAR image segmentation model training method, image segmentation method, device and equipment to effectively improve segmentation accuracy and boundary clarity.
[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows:
[0006] In a first aspect, an embodiment of the present invention provides a SAR image segmentation model training method, the method comprising:
[0007] Acquire a sample radar image and a segmentation label map of the sample radar image, where the segmentation label map is used to indicate a semantic category of each pixel point in the sample radar image;
[0008] performing edge enhancement processing on the sample radar image to determine a sample edge intensity map;
[0009] Performing superpixel segmentation according to the sample edge intensity map to obtain a sample superpixel region image including a plurality of superpixel regions;
[0010] Performing feature extraction and semantic segmentation on the sample radar image using an initial convolutional neural network to determine a sample category response map of the sample radar image;
[0011] Creating a graph structure based on the sample category response map and the sample superpixel region image, and performing semantic reasoning using an initial graph convolutional network to obtain a sample segmentation image;
[0012] The initial convolutional neural network and the initial graph convolutional network are trained according to the sample segmentation image and the segmentation label map to obtain a SAR image segmentation model including a target convolutional neural network and a target graph convolutional network.
[0013] Optionally, performing edge enhancement processing on the sample radar image to determine a sample edge intensity map includes:
[0014] Construct rectangular double window templates with the central axis at multiple angles;
[0015] Performing pixel-by-pixel convolution on the sample radar image using each rectangular double-window template to obtain two pixel-by-pixel convolution response values in each angular direction;
[0016] Determining a pixel-by-pixel structural intensity response in each angular direction according to the two pixel-by-pixel convolution response values in each angular direction;
[0017] The sample edge intensity map is determined according to the pixel-by-pixel structural intensity responses in the multiple angular directions.
[0018] Optionally, performing superpixel segmentation according to the sample edge intensity map to obtain a sample superpixel region image including a plurality of superpixel regions includes:
[0019] generating a potential energy map according to the sample edge strength map, wherein a pixel point with a higher potential energy value in the potential energy map has a greater probability of being at an edge position;
[0020] Performing a connected domain analysis based on the edge strength value of each pixel point in the sample edge strength map to determine that the sample radar image includes a seed marker map of multiple connected sub-regions;
[0021] Based on the seed marker image, a watershed algorithm is used to perform regional diffusion, and the segmentation boundary is determined by the potential energy map to generate an initial regional image;
[0022] Boundaries of multiple super-pixel regions of the initial region image are repaired, and the multiple super-pixel regions are numbered to obtain the sample super-pixel region image, where each super-pixel region has a unique label value.
[0023] Optionally, performing boundary repair on the multiple superpixel regions includes:
[0024] For target pixels that do not generate label values, number the target pixels according to the numbers of the neighboring pixels; and / or,
[0025] Morphological dilation and boundary interpolation algorithms are used to smooth the boundaries of the multiple superpixel regions.
[0026] Optionally, the using an initial convolutional neural network to perform feature extraction and semantic segmentation on the sample radar image to determine a sample category response map of the sample radar image includes:
[0027] Performing multi-stage feature extraction on the sample radar image using multiple encoding modules in the initial convolutional neural network and an attention mechanism module connected to the multiple encoding modules to generate a multi-layer feature map;
[0028] The deepest feature map is upsampled step by step, and is connected with the multi-layer feature map at the same scale during the upsampling process to obtain an intermediate feature map;
[0029] The convolution prediction model in the initial convolutional neural network is used to predict the intermediate feature map to determine the sample category response map.
[0030] Optionally, the step of creating a graph structure based on the sample category response graph and the sample superpixel region image, and performing semantic reasoning using an initial graph convolutional network to obtain a sample segmentation image includes:
[0031] Determining the category features of each superpixel region in the sample superpixel region image according to the category features of each pixel point in the sample category response map;
[0032] Calculating feature similarity between any two superpixel regions based on the category features of the multiple superpixel regions;
[0033] Connecting the multiple superpixel regions based on feature similarity between any two superpixel regions to generate a region-level graph structure;
[0034] Using the initial graph convolutional network to perform semantic reasoning on the region-level graph structure to obtain a region-level classification result;
[0035] Pixel demapping and spatial resolution restoration are performed on the region-level classification results to obtain the sample segmentation image.
[0036] In a second aspect, an embodiment of the present invention further provides a SAR image segmentation method, the method comprising:
[0037] Acquire target radar images;
[0038] Performing edge enhancement processing on the target radar image to determine a target edge intensity map;
[0039] Performing superpixel segmentation on the target edge intensity map to obtain a target superpixel region image including a plurality of superpixel regions;
[0040] Using a target convolutional neural network in a SAR image segmentation model to perform feature extraction and semantic segmentation on the target radar image, and determining a target category response map of the target radar image;
[0041] Creating a graph structure based on the target category response graph and the target superpixel region image, and performing semantic reasoning using a target graph convolutional network in the SAR image segmentation model to obtain an image segmentation result of the target radar image;
[0042] Wherein, the SAR image segmentation model is trained using the method according to any one of claims 1 to 6.
[0043] In a third aspect, an embodiment of the present invention further provides a SAR image segmentation model training device, the device comprising:
[0044] a sample acquisition module, configured to acquire a sample radar image and a segmentation label map of the sample radar image, wherein the segmentation label map is used to indicate a semantic category of each pixel point of the sample radar image;
[0045] A sample edge enhancement module is used to perform edge enhancement processing on the sample radar image to determine a sample edge intensity map;
[0046] a sample pixel segmentation module, configured to perform superpixel segmentation on the sample edge intensity map to obtain a sample superpixel region image comprising a plurality of superpixel regions;
[0047] a sample category response module, configured to perform feature extraction and semantic segmentation on the sample radar image using an initial convolutional neural network, and determine a sample category response map of the sample radar image;
[0048] A sample image segmentation module is used to create a graph structure based on the sample category response map and the sample superpixel region image, and use an initial graph convolutional network to perform semantic reasoning to obtain a sample segmentation image;
[0049] A model training module is used to train the initial convolutional neural network and the initial graph convolutional network according to the sample segmentation image and the segmentation label map to obtain a SAR image segmentation model including a target convolutional neural network and a target graph convolutional network.
[0050] Optionally, the sample edge enhancement module is specifically used to construct a rectangular double-window template with a central axis in multiple angular directions; use each rectangular double-window template to perform pixel-by-pixel convolution on the sample radar image to obtain two pixel-by-pixel convolution response values in each angular direction; determine the pixel-by-pixel structural intensity response in each angular direction based on the two pixel-by-pixel convolution response values in each angular direction; and determine the sample edge intensity map based on the pixel-by-pixel structural intensity responses in the multiple angular directions.
[0051] Optionally, the sample pixel segmentation module is specifically used to generate a potential energy map based on the sample edge intensity map, wherein the higher the potential energy value of the pixel point in the potential energy map, the greater the probability of being at the edge position; perform connected domain analysis based on the edge intensity value of each pixel point in the sample edge intensity map to determine that the sample radar image includes a seed labeling map of multiple connected sub-regions; based on the seed labeling map, use a watershed algorithm to perform regional diffusion, and use the potential energy map to determine the segmentation boundary to generate an initial region image; perform boundary repair on multiple super-pixel regions of the initial region image, and number the multiple super-pixel regions to obtain the sample super-pixel region image, wherein each super-pixel region has a unique label value.
[0052] Optionally, the sample pixel segmentation module is further used to number the target pixel points for which no label value is generated according to the numbering of the neighboring pixel points; and / or to use morphological dilation and boundary interpolation algorithms to smooth the boundaries of the multiple superpixel areas.
[0053] Optionally, the sample category response module is specifically used to use multiple encoding modules in the initial convolutional neural network and an attention mechanism module connected to the multiple encoding modules to perform multi-stage feature extraction on the sample radar image to generate a multi-layer feature map; upsample the deepest feature map step by step, and connect it with the multi-layer feature map at the same scale during the upsampling process to obtain an intermediate feature map; use the convolution prediction model in the initial convolutional neural network to predict the intermediate feature map to determine the sample category response map.
[0054] Optionally, the sample image segmentation module is specifically used to determine the category characteristics of each superpixel area in the sample superpixel area image according to the category characteristics of each pixel point in the sample category response map; calculate the feature similarity of any two superpixel areas according to the category characteristics of the multiple superpixel areas; connect the multiple superpixel areas according to the feature similarity of the any two superpixel areas to generate a region-level graph structure; use the initial graph convolution network to perform semantic reasoning on the region-level graph structure to obtain a region-level classification result; perform pixel demapping and spatial resolution restoration on the region-level classification result to obtain the sample segmented image.
[0055] In a fourth aspect, an embodiment of the present invention further provides a SAR image segmentation device, the device comprising:
[0056] Target acquisition module, used to acquire target radar images;
[0057] A target edge enhancement module is used to perform edge enhancement processing on the target radar image to determine a target edge intensity map;
[0058] a target pixel segmentation module, configured to perform superpixel segmentation on the target edge intensity map to obtain a target superpixel region image comprising a plurality of superpixel regions;
[0059] A target category response module is used to perform feature extraction and semantic segmentation on the target radar image using a target convolutional neural network in a SAR image segmentation model to determine a target category response map of the target radar image;
[0060] a target image segmentation module, configured to create a graph structure based on the target category response graph and the target superpixel region image, and perform semantic reasoning using the target graph convolutional network in the SAR image segmentation model to obtain an image segmentation result of the target radar image;
[0061] Wherein, the SAR image segmentation model is trained using the device according to claim 8.
[0062] Fifth, an embodiment of the present invention also provides an electronic device, comprising a processor, a memory, and a communication bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory through the communication bus, and the processor executes the machine-readable instructions to implement any of the methods described above.
[0063] Sixth, an embodiment of the present invention further provides a storage medium, wherein a computer program is stored on the storage medium, and when the computer program is executed by a processor, any of the above methods is executed.
[0064] The beneficial effects of the present invention are:
[0065] The SAR image segmentation model training method, image segmentation method, device and equipment provided by the present invention introduce an edge-guided superpixel partitioning mechanism, which enhances the perception of target edges while ensuring spatial coherence, effectively maintains boundary information, and significantly improves the structural clarity and integrity of the segmentation results; adopts a regional graph structure and GCN for regional-level reasoning, which can effectively capture cross-regional semantic associations and contextual information, and improve the classification consistency and robustness in complex backgrounds; graph reasoning is only performed at the regional node level, which greatly reduces the computational burden and memory consumption, improves the operating efficiency of the model, and is particularly suitable for the rapid processing of high-resolution SAR images. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0067] Figure 1 This is a schematic diagram of the structure of a typical pixel-level semantic segmentation model;
[0068] Figure 2 The process of the SAR image segmentation model training method provided by the embodiment of the present invention Figure 1 ;
[0069] Figure 3 The process of the SAR image segmentation model training method provided by the embodiment of the present invention Figure 2 ;
[0070] Figure 4 A schematic diagram of a rectangular double-window structure provided by an embodiment of the present invention;
[0071] Figure 5 The original SAR image and edge intensity map provided by the embodiment of the present invention;
[0072] Figure 6 The process of the SAR image segmentation model training method provided by the embodiment of the present invention Figure 3 ;
[0073] Figure 7 The process of the SAR image segmentation model training method provided by the embodiment of the present invention Figure 4 ;
[0074] Figure 8The process of the SAR image segmentation model training method provided by the embodiment of the present invention Figure 5 ;
[0075] Figure 9 Flowchart of the SAR image segmentation method provided by an embodiment of the present invention;
[0076] Figure 10 A flowchart of the image segmentation process provided by an embodiment of the present invention;
[0077] Figure 11 Schematic diagram of the structure of a SAR image segmentation model training device according to an embodiment of the present invention;
[0078] Figure 12 Schematic diagram of the structure of a SAR image segmentation device according to an embodiment of the present invention;
[0079] Figure 13 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0080] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0081] The following describes the embodiments of the present invention by means of specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention. In the absence of conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0082] The accompanying drawings are for illustrative purposes only and are schematic diagrams rather than actual drawings, and should not be construed as limiting the present invention. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the accompanying drawings.
[0083] Please refer to Figure 1 , which is a schematic diagram of the structure of a typical pixel-level semantic segmentation model, such as Figure 1As shown in the figure, taking DeepLabV3+ as an example, it consists of two parts: encoder and decoder. The encoder includes a feature extraction network (such as Backbone) and an atrous spatial pyramid pooling module (ASPP), and the decoder is used for detail recovery and pixel classification.
[0084] This type of method achieves end-to-end segmentation through pixel-level feature extraction and classification. Although it has achieved good results in natural images, due to the characteristics of SAR images such as high noise, unstructured texture, and unclear edges, traditional CNN methods are prone to the following problems during the image segmentation process: blurred boundaries and fragmented regions, especially in areas with complex boundaries of objects; weak spatial context modeling capabilities, relying only on local convolution, and difficulty in understanding cross-regional semantics; large computational complexity and low efficiency when processing high-resolution images.
[0085] Based on the problems existing in the above-mentioned prior art, the present invention intends to provide a SAR image segmentation model training method, image segmentation method, device and equipment. By generating an edge intensity map, edge intensity guidance information is introduced in the graph structure stage to improve the consistency between regional division and actual boundaries. Combined with the graph convolutional neural network, regional-level feature optimization is achieved, effectively improving boundary retention ability and anti-forging resistance.
[0086] The specific implementation of the SAR image segmentation model training method, image segmentation method, device and equipment provided by the present invention is described below with reference to the embodiments.
[0087] Please refer to Figure 2 , which is the process of the SAR image segmentation model training method provided by the embodiment of the present invention Figure 1 ,like Figure 2 As shown, the method may further include:
[0088] Step 101: Obtain a sample radar image and a segmentation label map of the sample radar image. The segmentation label map is used to indicate the semantic category of each pixel point in the sample radar image.
[0089] Specifically, a radar image training set is obtained using Synthetic Aperture Radar (SAR). Radar images from the training set are obtained as sample radar images. Semantic segmentation and labeling are performed on the sample radar images to determine the segmentation label map (ground truth) of the sample radar images. The segmentation label map assigns a corresponding semantic category to each pixel point.
[0090] Step 102: Perform edge enhancement processing on the sample radar image to determine a sample edge intensity map.
[0091] Specifically, a preset edge extraction method is used to extract the edges of the sample radar image, determine the structural edge information in the sample radar image, and output a sample edge strength map (ESM). The sample edge strength map is used to guide the superpixel segmentation, so that the superpixel division is more consistent with the real boundary, thereby achieving the coordinated optimization of boundary preservation and spatial consistency in the graph structure construction result, effectively improving the structural clarity of the final segmented image.
[0092] Step 103: Perform superpixel segmentation according to the sample edge intensity map to obtain a sample superpixel region image including multiple superpixel regions.
[0093] Specifically, superpixel segmentation is performed on the sample radar image based on the color, texture and other characteristics of each pixel point in the sample radar image, and the sample radar image is divided into multiple superpixel regions. In the process of superpixel segmentation of the sample radar image, the edge in the sample edge intensity map is used as the segmentation boundary to obtain a sample superpixel region image including multiple superpixel regions.
[0094] By dividing the sample radar image into regions with structural consistency through superpixel segmentation, the regional-level semantic expression ability can be effectively enhanced. In addition, by combining the edge intensity map, the edge information is fully utilized, which significantly improves the alignment between the superpixel and the boundary of the ground object, avoiding the blurred processing of the boundary by the pixel-level method.
[0095] Step 104: Use the initial convolutional neural network to perform feature extraction and semantic segmentation on the sample radar image to determine a sample category response map of the sample radar image.
[0096] Specifically, the initial convolutional neural network may include a feature extraction module and a convolution prediction model. The feature extraction model is used to extract features from the sample radar image to obtain a sample feature image. The convolution prediction model is used to perform pixel-by-pixel category prediction on the sample feature image to obtain a sample category response map.
[0097] In some embodiments, the sample category response map is a multi-channel tensor map, where each channel corresponds to a semantic category and each pixel position has a category probability prediction vector.
[0098] For example, multiple channels correspond to category 1, category 2, and category 3 respectively. The category probability prediction vectors of a pixel position in the three channels are (0.12, 0.84, 0.04), and the semantic category of the pixel position is determined to be category 2.
[0099] In some embodiments, the feature extraction module may be a CNN using a residual network ResNet50 as the backbone, and the convolution prediction model may be a lightweight convolution prediction model.
[0100] Furthermore, the feature extraction model can also be replaced by other convolutional neural network structures with similar extraction capabilities, such as ResNet-101, EfficientNet, MobileNet, etc.
[0101] Step 105: Create a graph structure based on the sample category response map and the sample superpixel region image, and use the initial graph convolutional network for semantic reasoning to obtain a sample segmentation image.
[0102] Specifically, each superpixel region in the sample superpixel region image is regarded as a graph node, and the category probability prediction vector of the pixels contained in each superpixel region is determined according to the sample category response graph. According to the category probability prediction vector of the pixels contained in each superpixel region, the initial feature vector of each graph node is determined.
[0103] A graph structure is created based on the initial feature vectors of multiple graph nodes corresponding to multiple superpixel regions. The graph structure is used to represent the semantic similarity relationship between different superpixel regions. The initial graph convolutional network (GCN) is used to perform semantic reasoning on the graph structure to obtain a sample segmentation image.
[0104] A regional graph structure is established based on the sample superpixel region image, and GCN is introduced for regional feature propagation and semantic reasoning, so that the model has cross-region context modeling capabilities and enhances the model's recognition ability of similar structures and semantic commonalities in complex scenes, thereby improving the consistency and accuracy of classification in background clutter or boundary adjacent areas.
[0105] In addition, a regional-level graph structure is established based on the sample superpixel region image, which greatly reduces the number of nodes. The graph convolution operation is also limited to the regional level, which greatly reduces the number of model parameters and computational complexity. While maintaining the high-resolution image processing accuracy, it improves the inference speed and system efficiency. It is particularly suitable for real-time or quasi-real-time segmentation tasks of large-scale SAR images.
[0106] Step 106: Train the initial convolutional neural network and the initial graph convolutional network based on the sample segmentation image and the segmentation label map to obtain a SAR image segmentation model including a target convolutional neural network and a target graph convolutional network.
[0107] Specifically, the classification results of each pixel in the sample segmentation image are compared with the classification labels of each pixel in the segmentation label map, and the prediction error is measured using the cross entropy loss function. The connection parameters of the initial convolutional neural network and the initial graph convolutional network are optimized according to the prediction error. After the number of training times reaches the prediction time or the prediction error converges, the training of the initial convolutional neural network and the initial graph convolutional network is completed to obtain the target convolutional neural network and the target graph convolutional network. The SAR image segmentation model includes the target convolutional neural network and the target graph convolutional network.
[0108] In some embodiments, the training process uses the Adam optimizer to simultaneously optimize all parameters of the initial convolutional neural network and the initial graph convolutional network to achieve end-to-end joint learning.
[0109] This embodiment integrates CNN and GCN into a unified network architecture. Through an end-to-end joint training mechanism, the feature extraction stage and the graph reasoning stage promote each other, enhancing boundary continuity and clarity. During the training process, the regional graph structure guides the CNN features to focus more on semantically consistent areas, and the CNN features provide the GCN with more discriminative input node representations. The GCN completes semantic propagation on the graph structure, effectively solving the regional relationship problem that CNN cannot understand, thereby realizing the co-evolution of feature expression and semantic modeling, and improving the robustness and training convergence of the overall model.
[0110] The SAR image segmentation model training method provided by the above embodiment of the present invention introduces an edge-guided superpixel partitioning mechanism, which enhances the perception of target edges while ensuring spatial coherence, effectively preserves boundary information, and significantly improves the structural clarity and integrity of the segmentation results; adopts a regional graph structure and GCN for regional-level reasoning, which can effectively capture cross-regional semantic associations and contextual information, and improve the classification consistency and robustness in complex backgrounds; graph reasoning is only performed at the regional node level, which greatly reduces the computational burden and memory consumption, improves the operating efficiency of the model, and is particularly suitable for the rapid processing of high-resolution SAR images.
[0111] The present invention constructs an overall framework from feature extraction, superpixel segmentation to regional graph reasoning, supports joint training and multi-module collaborative optimization, guides structural perception through superpixels, and optimizes regional semantics through graph convolution, achieving stable model convergence and steadily improved effects. The structural perception mechanism and regional reasoning enable the model to maintain good segmentation performance under the influence of noise.
[0112] The present invention not only excels in segmentation accuracy and boundary preservation, but also achieves a good balance between computational resource consumption and model generalization capability. It is suitable for practical SAR image semantic segmentation tasks and has high application value and promotion potential.
[0113] In one possible implementation, see Figure 3 , which is the process of the SAR image segmentation model training method provided by the embodiment of the present invention Figure 2 ,like Figure 3 As shown, the above step 102 may include:
[0114] Step 201: construct a rectangular double-window template with a central axis at multiple angles.
[0115] Step 202: Use each rectangular double-window template to perform pixel-by-pixel convolution on the sample radar image to obtain two pixel-by-pixel convolution response values in each angular direction.
[0116] Step 203: Determine the pixel-by-pixel structural intensity response in each angular direction according to the two pixel-by-pixel convolution response values in each angular direction.
[0117] Step 204: Determine a sample edge intensity map based on the pixel-by-pixel structural intensity responses in multiple angular directions.
[0118] In this embodiment, based on the rectangular bi-window strategy, structural contrast operations are performed on the sample radar image from multiple directions to extract structural edge information in the sample radar image and output a sample edge intensity map to provide guidance for subsequent structure-aware segmentation.
[0119] Specifically, the specific steps for extracting structural edge information of a sample radar image based on the rectangular double-window structure strategy are as follows:
[0120] 1. Constructing a local coordinate grid: At each image pixel position, construct a local grid window centered at the point. The window size is a preset parameter winSize. For example, winSize=10 in this embodiment.
[0121] 2. Set the angle set: uniformly sample P angle directions in the range [0, π) to simulate the structural response in each angle direction. For example, P in this embodiment is 16.
[0122] 3. Construct a rotating rectangular double window template: At each angle direction θ, construct a pair of rectangular double window structures. For an example, please refer to Figure 4 , is a schematic diagram of a rectangular double window structure provided by an embodiment of the present invention, such as Figure 4 As shown, the rectangular double window structure has one window located above the center (upper window) and the other located below (lower window), and the two windows are symmetrical about the center point. The size of each window is w×l, and the center position is offset from the center point by d. For example, in this embodiment, w=2, l=10, and d=2.
[0123] 4. Perform a double-window convolution operation on each pixel of the sample radar image: Convolve each pair of rectangular double-window structures with each pixel of the sample radar image. Two convolution response values are obtained for each pixel, indicating the degree of structural contrast of the local area in the angular direction.
[0124] 5. Calculate the response ratio: At each angle, calculate the convolution ratio of the upper / lower windows and its reciprocal, and take the smaller value of the convolution ratio and its reciprocal as the structural strength response at that angle.
[0125] 6. Directional minimization fusion: Take the minimum value of the structural intensity response in all angular directions pixel by pixel to generate the final edge intensity map. The edge intensity map has a higher response value at the edge of the image or where the structure suddenly changes, and a lower value inside the area.
[0126] For examples, please refer to Figure 5 , which are the original SAR image and edge intensity map provided by the embodiment of the present invention, Figure 5 The original SAR image (a) is subjected to edge extraction through the above steps 201 to 204 to obtain Figure 5 (b) Edge strength map.
[0127] The SAR image segmentation model training method provided by the above embodiment of the present invention uses a rectangular double-window strategy to extract an edge intensity map, and uses this as a guide to perform superpixel segmentation on the sample radar image. It fully utilizes edge information, has higher superpixel division accuracy, and significantly improves the alignment between superpixels and ground object boundaries. While ensuring spatial coherence, it enhances the perception ability of target edges, achieves effective preservation of boundary information, and significantly improves the structural clarity and integrity of the segmentation results.
[0128] In one possible implementation, see Figure 6 , which is the process of the SAR image segmentation model training method provided by the embodiment of the present invention Figure 3 ,like Figure 6 As shown, step 103 may include:
[0129] Step 301: Generate a potential energy map based on the sample edge strength map. In the potential energy map, the higher the potential energy value of a pixel point, the greater the probability of the pixel point being at the edge position.
[0130] Specifically, the sample edge intensity map is a real-valued image, and each pixel value reflects the probability that the pixel is located at the edge. The smaller the value, the higher the probability of being at the edge. In order to adapt to the watershed algorithm, a potential energy map of 1-ESM value is constructed. The edge position has a low ESM value and a high potential energy value G, forming a "high potential energy barrier" to inhibit regional expansion.
[0131] Step 302: Perform connected domain analysis based on the edge strength value of each pixel in the sample edge strength map to determine a seed marker map including multiple connected sub-regions.
[0132] Specifically, set the threshold T , select the ESM value greater than T Pixels with ESM values greater than 0 are considered as candidate pixels inside the region, that is, these points are considered not on the edge and are suitable as the center of the region. T The connected component labeling is performed on the candidate pixels, all connected sub-regions are extracted, and a unique label is assigned to each connected sub-region in the sample radar image to obtain the initial seed label map as the marker input of the watershed algorithm.
[0133] Step 303: Based on the seed marker image, a watershed algorithm is used to perform regional diffusion, and a potential energy map is used to determine the segmentation boundary to generate an initial region image.
[0134] Specifically, the potential energy map is used as the "topographic map" input of the watershed algorithm, and the marker generated in step 402 is used as the initial water source point. The standard watershed algorithm is executed to simulate the process of "expanding outward from the seed area". When water expands from different seeds, it will be blocked when encountering high potential energy areas (i.e., edges), forming natural boundaries; low potential energy areas are filled first, thereby realizing structure-based region growing. The output result is a label map for region segmentation of the entire sample segmentation image, and each label represents a superpixel area.
[0135] Step 304: perform boundary repair on multiple super-pixel regions of the initial region image, and number the multiple super-pixel regions to obtain a sample super-pixel region image, where each super-pixel region has a unique label value.
[0136] Specifically, in order to further improve the effect of the superpixel area, the boundaries of multiple superpixel areas of the initial area image are repaired. The repair includes filling the unlabeled areas and / or correcting and smoothing the boundaries of the superpixel areas. After the repair is completed, the multiple superpixel areas are numbered consecutively, so that the label values are numbered consecutively starting from 0 to obtain a sample superpixel area image.
[0137] In some embodiments, step 304 may include:
[0138] For target pixels that do not generate label values, the target pixels are numbered according to the numbers of the neighboring pixels; and / or, the boundaries of multiple superpixel regions are smoothed using morphological dilation and boundary interpolation algorithms.
[0139] Specifically, the watershed algorithm may generate unlabeled pixels at the edge of the image or in the edge noise area. Using the majority neighborhood voting strategy, the target pixel is numbered based on the numbers of the neighboring pixels of the target pixel that has not generated a label value. The number of the neighboring pixel with the highest frequency is determined and the target pixel is numbered.
[0140] Morphological dilation is used to fill irregular gaps at the boundaries of superpixel regions and enhance boundary continuity. Specifically, based on the labels of each superpixel region in the initial region image, the boundary pixels of the initial region image are determined, and a binary boundary mask is constructed. The superpixel interior is 0 and the boundary pixels are 1. The boundary mask is morphologically dilated to expand the boundary outward by 1-2 pixels. For the expanded boundary region, the labels are reallocated according to the main color similarity or spatial distance of adjacent superpixels to ensure that the expanded region is reasonably attributed.
[0141] The boundary interpolation algorithm extracts boundary chain codes from the expanded superpixel label map, identifies curvature mutation points on the boundary chain codes, and smoothes the boundaries of curvature mutations through linear interpolation or spline interpolation.
[0142] The SAR image segmentation model training method provided by the above embodiment of the present invention, by introducing an edge-guided superpixel partitioning mechanism, enhances the perception of target edges while ensuring spatial coherence, achieves effective preservation of boundary information, and significantly improves the structural clarity and integrity of the segmentation results.
[0143] In one possible implementation, see Figure 7 , which is the process of the SAR image segmentation model training method provided by the embodiment of the present invention Figure 4 ,like Figure 7 As shown, step 104 may include:
[0144] Step 401: Use multiple encoding modules in the initial convolutional neural network and an attention mechanism module connected to the multiple encoding modules to perform multi-stage feature extraction on the sample radar image to generate a multi-layer feature map.
[0145] Step 402: Upsample the deepest feature map step by step, and connect it with the multi-layer feature map at the same scale during the upsampling process to obtain an intermediate feature map.
[0146] Step 403: Use the convolution prediction model in the initial convolutional neural network to predict the intermediate feature map and determine the sample category response map.
[0147] Specifically, the initial convolutional neural network may include: multiple encoding modules, a Convolutional Block Attention Module (CBAM) connected to the multiple encoding modules, and a convolutional prediction model. For example, the multiple encoding modules may include 5.
[0148] Multiple encoding modules are used to extract edge features, texture information, and semantic representations layer by layer from shallow to deep. Since the sample radar image has four-channel input, the first-layer convolution kernel is initialized based on the ImageNet pre-trained weights, and the red channel weights are copied to the fourth channel to adapt to the channel structure. After each encoding module, a CBAM module is introduced to perform weighted optimization on the feature distribution of the channel dimension and spatial dimension respectively, to improve the network's ability to discriminate key areas and semantic boundaries, and output multi-layer feature maps.
[0149] For multi-layer feature maps, a combination of progressive decoding and skip connection mechanisms is used for multi-scale fusion, enabling the comprehensive utilization of low-level edge details and high-level semantic information to form an intermediate feature map with strong fusion and expression capabilities. The progressive decoding and skip connection mechanisms involve upsampling the deepest feature map and fusing it with a feature map of the same size. The fused feature map is then further upsampled and fusing with the corresponding feature map of the same size until it is concatenated with the feature map of the last layer to form the intermediate feature map.
[0150] The intermediate feature map is upsampled to the original image resolution by bilinear interpolation and input into a lightweight convolution prediction module to output the sample category response map (logits) for each pixel position.
[0151] In some embodiments, the CBAM module improves channel and spatial perception capabilities, and other attention mechanism modules, such as the SE (Squeeze-and-Excitation) module, the Non-local module, or the Transformer Encoder module, can also be used to achieve similar effects.
[0152] The SAR image segmentation model training method provided by the above-mentioned embodiment of the present invention introduces a channel-spatial attention module into the CNN backbone feature extraction network. This module performs weighted optimization on the feature distribution in the channel and spatial dimensions, guiding feature extraction to focus more on areas with significant differences between the target and background in the sample radar image. This improves the quality of feature expression, thereby enhancing the network's ability to discriminate between key areas and semantic boundaries, enhancing the ability to express features at different scales, and suppressing noise in complex backgrounds. Furthermore, layer-by-layer decoding and skip connections are combined to restore spatial details, effectively avoiding detail loss while maintaining the ability to express abstract semantic features.
[0153] In one possible implementation, see Figure 8 , which is the process of the SAR image segmentation model training method provided by the embodiment of the present invention Figure 5 ,like Figure 8 As shown, step 105 may include:
[0154] Step 501: Determine the category features of each superpixel region in the sample superpixel region image based on the category features of each pixel point in the sample category response map.
[0155] Step 502: Calculate the feature similarity between any two superpixel regions based on the category features of the multiple superpixel regions.
[0156] Step 503: Based on the feature similarity between any two super-pixel regions, multiple super-pixel regions are connected to generate a region-level graph structure.
[0157] Specifically, the category features of each pixel point included in each superpixel region are averaged and pooled to obtain the category features of each superpixel region. According to the type features of multiple superpixel regions, the feature similarity of any two superpixel regions is calculated based on the k-nearest neighbors (k-NN) strategy of feature similarity to construct a region-level graph structure.
[0158] In some embodiments, the Euclidean distance may be used to measure the semantic distance between any two superpixel regions as the feature similarity of any two superpixel regions, wherein the smaller the semantic distance, the higher the feature similarity.
[0159] According to the feature similarity of any two superpixel regions, k adjacent superpixel regions with the smallest semantic distance are selected for each superpixel region to establish undirected edge connections. The resulting connection relationship constitutes a region-level graph structure, which can be expressed using an unweighted adjacency matrix. Indicates that N is the total number of superpixel regions obtained by dividing the sample superpixel region image, and the elements in the matrix A ij Representation node i With node j Whether there is an edge connection (1 means connected, 0 means not connected).
[0160] The above-mentioned k-NN construction strategy can realize irregular and non-local graph connection, significantly improving the semantic adaptability and expressiveness of the graph structure. It is particularly suitable for scenarios with problems such as blurred boundaries, texture aliasing, or speckle noise in synthetic aperture radar images, and helps to avoid the information mis-diffusion phenomenon caused by rule-based connection methods.
[0161] In some embodiments, cross-region information modeling can be performed based on graph cuts, region merging, or Transformer global attention.
[0162] Step 504: Use the initial graph convolutional network to perform semantic reasoning on the region-level graph structure to obtain a region-level classification result.
[0163] Specifically, after completing the construction of the regional-level graph structure, GCN is used to propagate and update the semantic features of the superpixel region to achieve regional-level semantic reasoning and information integration.
[0164] The original GCN consists of multiple graph convolutional layers (GCNConv), nonlinear activation functions, and residual connections. Each node input feature consists of the concatenation of its corresponding region semantic vector and normalized spatial coordinates. This serves as the input node feature of the graph neural network. Through layer-by-layer graph convolution operations, feature interaction between adjacent nodes in the graph structure is achieved, thereby enhancing semantic consistency between similar regions.
[0165] To prevent overfitting and enhance the model's generalization, a dropout mechanism was introduced into the graph convolutional layer. Furthermore, a residual connection structure was designed between the input and output to preserve the original regional features. After graph neural network inference, the region-level classification results are obtained, which serve as the optimized regional representation of each superpixel region. This is then further restored to a pixel-level response map to improve the accuracy and boundary continuity of the overall segmentation result.
[0166] In some embodiments, GCN can be used as a replacement for other graph neural network structures, such as GraphSAGE, Graph Attention Network (GAT), or Simplified Graph Convolution (SGC).
[0167] Step 505: Perform pixel demapping and spatial resolution restoration on the region-level classification results to obtain a sample segmentation image.
[0168] Specifically, after the initial GCN outputs the region-level classification result, a pixel inverse mapping strategy is used to restore the prediction result of each superpixel region to the original image space to achieve a fine pixel-level segmentation map output. First, the attribution relationship of each pixel in the superpixel region division is recorded, that is, a superpixel index map consistent with the image size is constructed. Subsequently, based on the index relationship, the region-level classification result output by the initial GCN is copied and assigned to all pixel positions contained in the superpixel region, thereby obtaining a coarse pixel-level response map of the same size as the original image. Finally, in order to ensure the consistency of the spatial size of the response map and the feature map output by the initial CNN, the inverse mapping result is further spatially aligned and size-uniformed using bilinear interpolation. This process ensures that the region-level semantic enhancement results can be smoothly and effectively restored to the pixel level, and fused or compared with other branch outputs.
[0169] The SAR image segmentation model training method provided by the above embodiment of the present invention treats each superpixel region as a node in the graph, establishes graph edge connections through the k-nearest neighbor strategy, and forms node features by averaging or weighted aggregation of pixel features extracted by CNN, thereby realizing semantic abstraction from the pixel to the regional level. The graph structure is simpler and the semantics are more concentrated, avoiding the dimensionality disaster and noise propagation problems of the pixel graph method. GCN is used to perform feature propagation on the regional-level graph structure to model the spatial-semantic relationship across regions, realizing the collaborative optimization of CNN and GCN at the feature level, which not only ensures pixel accuracy but also improves regional consistency.
[0170] Based on the SAR image segmentation model trained by the above SAR image segmentation model training method, the following describes the method of image segmentation using the SAR image segmentation model. Figure 9 , which is a flow chart of the SAR image segmentation method provided by an embodiment of the present invention, such as Figure 9 As shown, the method may include:
[0171] Step 601: Acquire a target radar image.
[0172] Step 602: Perform edge enhancement processing on the target radar image to determine a target edge intensity map.
[0173] Step 603: Perform superpixel segmentation on the target edge intensity map to obtain a target superpixel region image including multiple superpixel regions.
[0174] Step 604: Use the target convolutional neural network in the SAR image segmentation model to perform feature extraction and semantic segmentation on the target radar image to determine a target category response map of the target radar image.
[0175] Step 605: Create a graph structure based on the target category response graph and the target superpixel region image, and use the target graph convolutional network in the SAR image segmentation model to perform semantic reasoning to obtain an image segmentation result of the target radar image.
[0176] Among them, the SAR image segmentation model is trained using the above method.
[0177] Please refer to Figure 10 , is a flowchart of the image segmentation process provided by an embodiment of the present invention, such as Figure 10 As shown in the figure, for the input SAR image, on the one hand, the edge intensity map is calculated, and the superpixel region image is calculated based on the edge intensity map. On the other hand, CNN is used to extract multi-scale pixel features and predict the category response map. The regional-level graph structure is created based on the superpixel region image and the category response map. GCN is used to perform regional-level semantic reasoning on the regional-level graph structure, and the regional prediction structure is mapped back to the pixel-level space. The pixel-by-pixel cross entropy loss is calculated with the label map. End-to-end training optimization is performed based on the cross entropy loss, and the parameters of CNN and GCN are updated. After the training is completed, the regional prediction structure is mapped back to the pixel-level space to output the SAR image segmentation result.
[0178] Please refer to Figure 11 , which is a structural diagram of a SAR image segmentation model training device according to an embodiment of the present invention. Figure 11 As shown, the device may include:
[0179] A sample acquisition module 701 is used to acquire a sample radar image and a segmentation label map of the sample radar image, where the segmentation label map is used to indicate the semantic category of each pixel point in the sample radar image;
[0180] The sample edge enhancement module 702 is used to perform edge enhancement processing on the sample radar image and determine a sample edge intensity map;
[0181] The sample pixel segmentation module 703 is used to perform super-pixel segmentation on the sample edge intensity map to obtain a sample super-pixel region image including multiple super-pixel regions;
[0182] A sample category response module 704 is configured to perform feature extraction and semantic segmentation on the sample radar image using an initial convolutional neural network to determine a sample category response map of the sample radar image;
[0183] The sample image segmentation module 705 is used to create a graph structure based on the sample category response map and the sample superpixel region image, and use the initial graph convolution network to perform semantic reasoning to obtain a sample segmentation image;
[0184] The model training module 706 is used to train the initial convolutional neural network and the initial graph convolutional network according to the sample segmentation image and the segmentation label map to obtain a SAR image segmentation model including a target convolutional neural network and a target graph convolutional network.
[0185] Optionally, the sample edge enhancement module 702 is specifically used to construct a rectangular double-window template with a central axis in multiple angular directions; use each rectangular double-window template to perform pixel-by-pixel convolution on the sample radar image to obtain two pixel-by-pixel convolution response values in each angular direction; determine the pixel-by-pixel structure intensity response in each angular direction based on the two pixel-by-pixel convolution response values in each angular direction; and determine the sample edge intensity map based on the pixel-by-pixel structure intensity responses in multiple angular directions.
[0186] Optionally, the sample pixel segmentation module 703 is specifically used to generate a potential energy map based on the sample edge intensity map, where the higher the potential energy value of the pixel point in the potential energy map, the greater the probability of being at the edge position; perform connected domain analysis based on the edge intensity value of each pixel point in the sample edge intensity map to determine that the sample radar image includes a seed labeling map of multiple connected sub-regions; based on the seed labeling map, use a watershed algorithm to perform regional diffusion, and use the potential energy map to determine the segmentation boundary to generate an initial region image; perform boundary repair on multiple super-pixel regions of the initial region image, and number the multiple super-pixel regions to obtain a sample super-pixel region image, where each super-pixel region has a unique label value.
[0187] Optionally, the sample pixel segmentation module 703 is further used to number the target pixel points for which no label value is generated according to the numbering of the neighboring pixel points; and / or, use morphological dilation and boundary interpolation algorithms to smooth the boundaries of multiple superpixel areas.
[0188] Optionally, the sample category response module 704 is specifically configured to use multiple encoding modules in the initial convolutional neural network and an attention mechanism module connected to the multiple encoding modules to perform multi-stage feature extraction on the sample radar image to generate a multi-layer feature map; upsample the deepest feature map step by step, and connect it with the multi-layer feature map at the same scale during the upsampling process to obtain an intermediate feature map; use the convolution prediction model in the initial convolutional neural network to predict the intermediate feature map to determine the sample category response map.
[0189] Optionally, the sample image segmentation module 705 is specifically used to determine the category characteristics of each superpixel region in the sample superpixel region image based on the category characteristics of each pixel point in the sample category response map; calculate the feature similarity of any two superpixel regions based on the category characteristics of multiple superpixel regions; connect multiple superpixel regions based on the feature similarity of any two superpixel regions to generate a region-level graph structure; use the initial graph convolution network to perform semantic reasoning on the region-level graph structure to obtain a region-level classification result; perform pixel de-mapping and spatial resolution restoration on the region-level classification result to obtain a sample segmented image.
[0190] Please refer to Figure 12 , which is a structural diagram of a SAR image segmentation device according to an embodiment of the present invention. Figure 12 As shown, the device may include:
[0191] Target acquisition module 801, used to acquire target radar images;
[0192] The target edge enhancement module 802 is used to perform edge enhancement processing on the target radar image and determine a target edge intensity map;
[0193] The target pixel segmentation module 803 is used to perform super-pixel segmentation on the target edge intensity map to obtain a target super-pixel region image including multiple super-pixel regions;
[0194] A target category response module 804 is configured to perform feature extraction and semantic segmentation on the target radar image using a target convolutional neural network in the SAR image segmentation model, and determine a target category response map of the target radar image;
[0195] The target image segmentation module 805 is used to create a graph structure based on the target category response map and the target superpixel region image, and use the target graph convolution network in the SAR image segmentation model to perform semantic reasoning to obtain the image segmentation result of the target radar image;
[0196] The SAR image segmentation model is trained using the above-mentioned device.
[0197] In a possible implementation, the embodiment of the present invention further provides an electronic device, please refer to Figure 13 , is a schematic diagram of an electronic device provided by an embodiment of the present invention, such as Figure 13 As shown, the electronic device may include a processor 901, a memory 902 and a communication bus 903, wherein the memory 902 stores machine-readable instructions executable by the processor 901. When the electronic device is running, the processor 901 communicates with the memory 902 through the communication bus 903, and the processor 901 executes the machine-readable instructions to implement the method of the above embodiment.
[0198] In a possible implementation, an embodiment of the present invention further provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method of the above embodiment is executed.
[0199] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A SAR image segmentation model training method, characterized in that: The method comprises: Acquire a sample radar image and a segmentation label map of the sample radar image, where the segmentation label map is used to indicate a semantic category of each pixel point in the sample radar image; performing edge enhancement processing on the sample radar image to determine a sample edge intensity map; Performing superpixel segmentation according to the sample edge intensity map to obtain a sample superpixel region image including a plurality of superpixel regions; Performing feature extraction and semantic segmentation on the sample radar image using an initial convolutional neural network to determine a sample category response map of the sample radar image; Creating a graph structure based on the sample category response map and the sample superpixel region image, and performing semantic reasoning using an initial graph convolutional network to obtain a sample segmentation image; Training the initial convolutional neural network and the initial graph convolutional network according to the sample segmentation image and the segmentation label map to obtain a SAR image segmentation model including a target convolutional neural network and a target graph convolutional network; The superpixel segmentation is performed according to the sample edge intensity map to obtain a sample superpixel region image including a plurality of superpixel regions, including: generating a potential energy map according to the sample edge strength map, wherein a pixel point with a higher potential energy value in the potential energy map has a greater probability of being at an edge position; Performing a connected domain analysis based on the edge strength value of each pixel point in the sample edge strength map to determine that the sample radar image includes a seed marker map of multiple connected sub-regions; Based on the seed marker image, a watershed algorithm is used to perform regional diffusion, and the segmentation boundary is determined by the potential energy map to generate an initial regional image; Performing boundary repair on multiple super-pixel regions of the initial region image and numbering the multiple super-pixel regions to obtain the sample super-pixel region image, where each super-pixel region has a unique label value; The step of creating a graph structure based on the sample category response graph and the sample superpixel region image and performing semantic reasoning using an initial graph convolutional network to obtain a sample segmentation image includes: Determining the category features of each superpixel region in the sample superpixel region image according to the category features of each pixel point in the sample category response map; Calculating feature similarity between any two superpixel regions based on the category features of the multiple superpixel regions; Connecting the multiple superpixel regions based on feature similarity between any two superpixel regions to generate a region-level graph structure; Using the initial graph convolutional network to perform semantic reasoning on the region-level graph structure to obtain a region-level classification result; Pixel demapping and spatial resolution restoration are performed on the region-level classification results to obtain the sample segmentation image.
2. The method according to claim 1, wherein The performing edge enhancement processing on the sample radar image to determine a sample edge intensity map includes: Construct rectangular double window templates with the central axis at multiple angles; Performing pixel-by-pixel convolution on the sample radar image using each rectangular double-window template to obtain two pixel-by-pixel convolution response values in each angular direction; Determining a pixel-by-pixel structural intensity response in each angular direction according to the two pixel-by-pixel convolution response values in each angular direction; The sample edge intensity map is determined according to the pixel-by-pixel structural intensity responses in the multiple angular directions.
3. The method according to claim 1, wherein The performing boundary repair on the plurality of super-pixel regions includes: For target pixels that do not generate label values, number the target pixels according to the numbers of the neighboring pixels; and / or, Morphological dilation and boundary interpolation algorithms are used to smooth the boundaries of the multiple superpixel regions.
4. The method according to claim 1, wherein The using an initial convolutional neural network to perform feature extraction and semantic segmentation on the sample radar image to determine a sample category response map of the sample radar image includes: Performing multi-stage feature extraction on the sample radar image using multiple encoding modules in the initial convolutional neural network and an attention mechanism module connected to the multiple encoding modules to generate a multi-layer feature map; The deepest feature map is upsampled step by step, and is connected with the multi-layer feature map at the same scale during the upsampling process to obtain an intermediate feature map; The convolution prediction model in the initial convolutional neural network is used to predict the intermediate feature map to determine the sample category response map.
5. A SAR image segmentation method, characterized in that: The method comprises: Acquire target radar images; Performing edge enhancement processing on the target radar image to determine a target edge intensity map; Performing superpixel segmentation on the target edge intensity map to obtain a target superpixel region image including a plurality of superpixel regions; Using a target convolutional neural network in a SAR image segmentation model to perform feature extraction and semantic segmentation on the target radar image, and determining a target category response map of the target radar image; Creating a graph structure based on the target category response graph and the target superpixel region image, and performing semantic reasoning using a target graph convolutional network in the SAR image segmentation model to obtain an image segmentation result of the target radar image; Wherein, the SAR image segmentation model is trained using the method according to any one of claims 1 to 4.
6. A SAR image segmentation model training device, characterized in that: The device comprises: a sample acquisition module, configured to acquire a sample radar image and a segmentation label map of the sample radar image, wherein the segmentation label map is used to indicate a semantic category of each pixel point of the sample radar image; A sample edge enhancement module is used to perform edge enhancement processing on the sample radar image to determine a sample edge intensity map; a sample pixel segmentation module, configured to perform superpixel segmentation on the sample edge intensity map to obtain a sample superpixel region image comprising a plurality of superpixel regions; a sample category response module, configured to perform feature extraction and semantic segmentation on the sample radar image using an initial convolutional neural network, and determine a sample category response map of the sample radar image; A sample image segmentation module is used to create a graph structure based on the sample category response map and the sample superpixel region image, and use an initial graph convolutional network to perform semantic reasoning to obtain a sample segmentation image; A model training module is used to train the initial convolutional neural network and the initial graph convolutional network according to the sample segmentation image and the segmentation label map to obtain a SAR image segmentation model including a target convolutional neural network and a target graph convolutional network; The sample pixel segmentation module is specifically configured to generate a potential energy map based on the sample edge intensity map, wherein a pixel point with a higher potential energy value in the potential energy map has a greater probability of being at an edge position; perform a connected domain analysis based on the edge intensity value of each pixel point in the sample edge intensity map to determine that the sample radar image includes a seed labeling map of multiple connected sub-regions; perform regional diffusion using a watershed algorithm based on the seed labeling map, and determine the segmentation boundary using the potential energy map to generate an initial region image; perform boundary repair on multiple super-pixel regions of the initial region image, and number the multiple super-pixel regions to obtain the sample super-pixel region image, wherein each super-pixel region has a unique label value; The sample image segmentation module is specifically used to determine the category characteristics of each superpixel area in the sample superpixel area image based on the category characteristics of each pixel point in the sample category response map; calculate the feature similarity of any two superpixel areas based on the category characteristics of the multiple superpixel areas; connect the multiple superpixel areas based on the feature similarity of the any two superpixel areas to generate a region-level graph structure; use the initial graph convolution network to perform semantic reasoning on the region-level graph structure to obtain a region-level classification result; perform pixel demapping and spatial resolution restoration on the region-level classification result to obtain the sample segmentation image.
7. A SAR image segmentation device, characterized in that: The device comprises: Target acquisition module, used to acquire target radar images; A target edge enhancement module is used to perform edge enhancement processing on the target radar image to determine a target edge intensity map; a target pixel segmentation module, configured to perform superpixel segmentation on the target edge intensity map to obtain a target superpixel region image comprising a plurality of superpixel regions; A target category response module is used to perform feature extraction and semantic segmentation on the target radar image using a target convolutional neural network in a SAR image segmentation model to determine a target category response map of the target radar image; a target image segmentation module, configured to create a graph structure based on the target category response graph and the target superpixel region image, and perform semantic reasoning using the target graph convolutional network in the SAR image segmentation model to obtain an image segmentation result of the target radar image; Wherein, the SAR image segmentation model is trained using the device according to claim 6.
8. An electronic device, characterized in that: The electronic device comprises a processor, a memory and a communication bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the communication bus, and the processor executes the machine-readable instructions to implement the method according to any one of claims 1 to 4, or the method according to claim 5.
Citation Information
Patent Citations
Polarized SAR image classification method based on superpixels and image convolution network
CN113298129A
End-to-end polarimetric SAR image superpixel segmentation method, system and device based on full convolutional neural network, and medium
CN119762494A