A Real-time Segmentation Method for SAR Images Based on Lovász Loss and Lightweight Bilateral Network
The Lovasz loss and lightweight bilateral network approach addresses the precision and real-time challenges in SAR image segmentation, enhancing classification accuracy and reducing computational demands.
Patent Information
- Application Number
- CN202211610172.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-12
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-12
AI Technical Summary
The existing SAR image segmentation method has insufficient classification accuracy and speed. The traditional convolutional neural network has a high computing burden, excessive boundaries are smooth, and the traditional bilateral network uses cross-entropy loss function to lead to optimization target deviation.
The LOVASZ loss function and lightweight bilateral network are adopted, and the BiSeNet network is trained through pre-training and fine-tuning, combining the Lovasz loss and cross-entropy loss function as optimization goals, and the learning rate is dynamically adjusted to achieve lightweight real-time semantic segmentation.
On the basis of ensuring real-time performance, the classification accuracy of SAR images is improved, and the accuracy of semantic segmentation is obtained, filling the gap in the optimization goals of traditional methods.
Smart Images

Figure CN116051828B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of SAR image segmentation, and in particular relates to a real-time SAR image segmentation method based on Lovasz loss and a lightweight bilateral network. Background Technique
[0002] Polarimetric SAR image segmentation is one of the research hotspots in the field of SAR image interpretation, which can achieve pixel-level classification of SAR images and has been widely applied in urban segmentation, road extraction, crop classification, etc. Traditional classification methods mainly rely on the statistical distribution of polarimetric SAR data and the scattering characteristics of targets. However, such methods can only extract some low-level features and cannot make full and effective use of SAR image information.
[0003] To meet the requirements of classification accuracy and speed, deep learning methods represented by convolutional neural networks (CNNs) have been widely introduced into the segmentation of SAR images. However, the sliding window method adopted by CNNs performs repeated sampling on the same pixel multiple times, resulting in a huge computational and storage burden. At the same time, since a sampling window often contains different types of pixels, this region-to-pixel method can only predict the central pixel, leading to over-smoothing of the boundary and uncertain segmentation results.
[0004] To overcome the shortcomings of the region-to-pixel method, the fully convolutional neural network (FCN) proposed by Long realizes the segmentation of images of any size by replacing the fully connected layer of CNNs with convolutional layers. Subsequently, more research on SAR semantic segmentation networks has been carried out and applied. Among them, the bilateral network (BiSeNet) decouples spatial information and receptive fields using a spatial path and a context path, and further improves the accuracy using a feature fusion module (FFM) and an attention refinement module (ARM).
[0005] However, traditional bilateral networks all adopt the cross-entropy loss function to optimize the loss curve, which will cause certain deviations in the final result when directly used as the optimization target. Summary of the Invention
[0006] The purpose of the present invention is to provide a real-time SAR image segmentation method based on Lovasz loss and a lightweight bilateral network, which takes IOU (intersection over union) as the optimization target and can achieve higher classification accuracy on the basis of ensuring real-time performance (lightweight).
[0007] The present invention adopts the following technical solutions: A real-time SAR image segmentation method based on Lovasz loss and a lightweight bilateral network, comprising the following steps:
[0008] The BiSeNet network is pre-trained and fine-tuned in sequence; among them, during the pre-training process, the cross-entropy loss function is used as the loss function of the BiSeNet network, and during the fine-tuning training process, the combined loss function is used as the loss function of the BiSeNet network, and the combined loss function includes the Lovasz loss function and the cross-entropy loss function;
[0009] Based on the BiSeNet network after fine-tuning training, the polarimetric SAR image is segmented.
[0010] Furthermore, the Lovasz loss function is:
[0011]
[0012] where Loss(f)2 is the Lovasz loss function, C is the number of pixel categories in the SAR image, c is the category ordinal number of the pixels in the SAR image, is the extended loss for each category.
[0013] Furthermore, the cross-entropy loss function is:
[0014]
[0015] where Loss(f)1 is the cross-entropy loss function, p is the number of pixels input into the BiSeNet network, is the true label of the i-th pixel, is the classification result of the BiSeNet network for the i-th pixel as the probability.
[0016] Furthermore, the combined loss function is:
[0017]
[0018] where Loss(f)3 is the combined loss function.
[0019] Furthermore, during the pre-training / fine-tuning training process, it includes:
[0020] Alternately pre-train / fine-tune the BiSeNet network and validate it, and record the network weights and the corresponding validation losses after each pre-training / fine-tuning training;
[0021] When the number of validation times reaches the first threshold, select the network weights corresponding to the minimum validation loss as the weights of the BiSeNet network.
[0022] Furthermore, during the pre-training / fine-tuning training process, when the number of consecutive pre-training / fine-tuning training times reaches the second threshold and the corresponding validation loss does not decrease, end the pre-training / fine-tuning training process.
[0023] Furthermore, during the pre-training / fine-tuning process, it also includes:
[0024] Dynamically adjusting the learning rate of the BiSeNet network.
[0025] Another technical solution of the present invention: A real-time SAR image segmentation device based on Lovasz loss and lightweight bilateral network, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method of the above-mentioned real-time SAR image segmentation device based on Lovasz loss and lightweight bilateral network.
[0026] The beneficial effects of the present invention are: Based on the BiSeNet network, the present invention uses the commonly used index union intersection in semantic segmentation as the optimization target, introduces the extended Lovasz loss for it, realizes real-time semantic segmentation of the lightweight bilateral network based on Lovasz loss, and obtains higher classification accuracy. Description of the Drawings
[0027] Figure 1 It is a schematic diagram of the architecture of the BiSeNet network in the embodiment of the present invention;
[0028] Figure 2 It is a specific architecture diagram of the ARM module in the embodiment of the present invention;
[0029] Figure 3 It is a specific architecture diagram of the FFM module in the embodiment of the present invention;
[0030] Figure 4 It is a schematic diagram of the amplitude image in a polarization state and the corresponding color ground truth in the verification embodiment of the present invention;
[0031] Figure 5 It is a schematic diagram of the comparison result of the classification result with different semantic networks in the verification embodiment of the present invention. Detailed Embodiments
[0032] The present invention will be described in detail below with reference to the drawings and specific embodiments.
[0033] IOU is an important evaluation index in semantic segmentation, which can effectively reflect the accuracy of segmentation. After being extended by Lovasz, it can be used as an optimization target to improve the network accuracy. At the same time, the lightweight bilateral segmentation network can meet the requirements of real-time performance while obtaining high accuracy. The bilateral segmentation network can meet the requirements of real-time performance while obtaining high accuracy.
[0034] The present invention takes the weighted Lovasz loss function and cross-entropy loss function as the final optimization objective, and establishes a lightweight bilateral real-time semantic segmentation network. Compared with the BiSeNet network using the traditional cross-entropy loss, the network introduced in the present invention has higher classification accuracy. The improved loss function proposed by the present invention can better fill the gap in the optimization objective in the field of polarimetric SAR semantic segmentation, taking into account both real-time performance and high accuracy.
[0035] The present invention discloses a method for real-time segmentation of SAR images based on LOVASZ loss and a lightweight bilateral network, comprising the following steps: pre-training and fine-tuning the BiSeNet network in sequence; wherein, during the pre-training process, the loss function of the BiSeNet network adopts the cross-entropy loss function, and during the fine-tuning training process, the loss function of the BiSeNet network adopts a combined loss function, and the combined loss function includes the Lovasz loss function and the cross-entropy loss function; segmenting the polarimetric SAR image based on the BiSeNet network (bilateral segmentation network) after fine-tuning training.
[0036] Based on the BiSeNet network, the present invention takes the commonly used metric in semantic segmentation, intersection over union (IOU), as the optimization objective, introduces the extended Lovasz loss for it, realizes the real-time semantic segmentation of the lightweight bilateral network based on the Lovasz loss, and obtains higher classification accuracy.
[0037] Specifically, in the embodiment of the present invention, the BiSeNet network needs to be trained. First, the input data to be classified is extracted from the polarimetric SAR scattering matrix and normalized. Then, data augmentation (such as random cropping, horizontal flipping, and vertical flipping, etc.) is performed on the normalized input data, and it is divided into a training set, a test set, and a validation set.
[0038] Next, the BiSeNet network is trained using the data in the training set. In the embodiment of the present invention, the maximum number of training times is designed to be 200. At the same time, during the training process, in the network training optimization, if the loss of the training set does not decrease for several rounds, the learning rate is reduced at this time, that is, the learning rate of the BiSeNet network is dynamically adjusted.
[0039] Specifically, in the embodiment of the present invention, the specific architecture of the network is as Figure 1As shown in the figure, it includes four modules: a spatial path, a context path, an ARM (attention refinement module), and an FFM (feature fusion module). The spatial path is composed of alternating convolutional layers (Conv), normalization layers (BN), and activation functions (ReLU), and a total of three corresponding convolutional processing modules are included. The convolution operation is a convolution with a convolution kernel of 3, a stride of 2, and a padding of 1. Each convolution operation reduces the size of the feature map to 1 / 2 of the original. The context path contains four downsampling layers. At the end of the 16-fold and 32-fold downsampling, the ARM module is used to refine the features, and the FFM is used to concatenate the output of the spatial path with the context path.
[0040] The spatial path consists of three blocks, and each block is composed of a convolutional layer, a BN layer, and a ReLU layer with [stride = 2, kernel size = 3, padding = 1]. The convolutional layer extracts the underlying features of the image through the convolution operation of the convolution kernel; the batch normalization layer normalizes the output of the convolutional layer, controls the gradient explosion to prevent the gradient disappearance, and prevents overfitting; the ReLU layer introduces non-linear features to enhance the feature learning ability of the network. The final output of the spatial path is a feature map with an eighth of the size of the original image, retaining rich spatial details.
[0041] The context path adopts the Resnet network structure, which includes 17 convolutional layers and a fully connected layer. In addition, global average pooling is used at the end of the context path to provide global context information for the receptive field. In practical applications, the lightweight ResNet18 network is used for multiple downsamplings to obtain a considerable receptive field, and finally the output feature map is combined with the ResNet18 global pooling output.
[0042] As Figure 2 shown in the figure, it is the specific architecture of the ARM module. The input of the ARM is the original feature map after downsampling, which is composed of four parts: global pooling, 1×1 convolution, normalization (BN) layer, and activation function layer. The original feature map and the feature map processed by the ARM are multiplied to obtain a feature map with enhanced features. The global pooling is used to obtain global context information, compress the length and width, and retain the channel information; 1×1 convolution is used for cross-channel feature information integration to reduce the convolution kernel parameters. Then the BN layer is used to normalize the input features. Then the sigmoid function is used to activate and introduce more non-linear features, thereby enhancing the feature expression ability.
[0043] As Figure 3As shown in the figure, it is the specific architecture of the FFM module. The two ARM modules in the context path, after dilation (to the same size) and summation operation, together with the output of the spatial path, serve as the input of the FFM to correctly fuse features. The FFM module first concatenates the outputs of the spatial path and the context path, and then sequentially passes through a convolutional layer, a BN layer, and a ReLu layer. The result is processed through two paths. One path goes through global pooling, 1×1 convolution, ReLU, 1×1 convolution, and a sigmoid function and is multiplicatively fused with the result of the previous step, while the other path directly outputs. The results of the two paths are processed by an adder to obtain the final feature map.
[0044] In the training of the embodiment of the present invention, a dynamic learning rate is adopted. The specific strategy is that if the classification loss does not decrease during the training of 10 batches, the learning rate will be decreased to 0.7 times the current learning rate, and the minimum learning rate is set to 10 -7 to prevent the optimization speed from being too slow due to an overly low learning rate.
[0045] More specifically, in this embodiment, 200 groups of training are first carried out. The learning rate starts from 0.001, and the minimum learning rate is 10 -7 After each training, the validation set is used for verification. After the training is completed, the weights corresponding to the minimum loss of the validation set are selected to set the network.
[0046] As a specific implementation manner, during the pre-training / fine-tuning training process, it includes: alternately performing pre-training / fine-tuning training and verification on the BiSeNet network, and recording the network weights and the corresponding validation losses after each pre-training / fine-tuning training; when the number of verification times reaches the first threshold (such as 200 times as mentioned above), the network weights corresponding to the minimum validation loss are selected as the weights of the BiSeNet network.
[0047] In addition, during the pre-training / fine-tuning training process, when the number of consecutive pre-training / fine-tuning training times reaches the second threshold (such as 20 times) and the corresponding validation loss does not decrease, the pre-training / fine-tuning training process is ended.
[0048] The cross-entropy loss function in the above pre-training process is:
[0049]
[0050] where Loss(f)1 is the cross-entropy loss function, p is the number of pixels input into the BiSeNet network, is the true label of the i-th pixel, is the classification result of the BiSeNet network for the i-th pixel as the probability.
[0051] After completing the above pre-training process, the optimal weights are selected through the validation set as the parameters of the pre-trained network, and the fine-tuning stage is entered.
[0052] In the fine-tuning stage, the Lovasz loss function and the cross-entropy loss function (i.e., the combined loss function) are used to further optimize the pre-trained network with a new loss function. The initial learning rate is set to 0.0005 and runs to the minimum learning rate of 10 -6 。
[0053] Regarding the Lovasz loss function, it is calculated based on the Jaccard loss. Specifically, the Jacaad index (i.e., IOU):
[0054]
[0055] where y * represents the true label of the pixel, represents the predicted label of the pixel by the BiSeNet network, c is the category ordinal number of the pixel in the SAR image, c ∈ C, and C is the number of categories of pixels in the SAR image.
[0056] Then the Jaccard loss:
[0057]
[0058] Since this metric is discontinuous, Lovasz extension is performed on it for better optimization, and the Lovasz loss function is obtained after extension:
[0059]
[0060] where Loss(f)2 is the Lovasz loss function, is the extended loss for each category, and the derivation of the extended loss is as follows:
[0061]
[0062]
[0063] where x i (c) is the probability that pixel i is predicted as the c-th category, and m i (c) is the loss caused by this pixel. When calculating the loss for the c-th category, c is fixed. For all pixels belonging to the c-th category, that is, c = y i (c), the greater the prediction probability, the better the network accuracy. Therefore, the loss is 1 - x i (c). For other pixels that do not belong to the c-th category, the greater the prediction probability, the greater the network accuracy. Therefore, the loss is x i (c). m(c) represents the m of all pixels i(c) The composed vector, i.e., the error of all pixels for the c-th category. g i m(c) = m c ({π1,..., π i}) - m c ({π1,..., π i-1}) where {π1,..., π i} represents arranging the elements in descending order. After arranging m(c) in descending order, m c ({π1,..., π i}) - m c ({π1,..., π i-1}) represents the element in the i-th position in the descending order. By The Jaccard loss becomes a continuous function, which is convenient for calculation.
[0064] In summary, the joint loss function can be obtained as follows:
[0065]
[0066] where Loss(f)3 is the joint loss function.
[0067] Thus, the overall loss function in the overall training process of the BiSeNet network can be expressed as follows:
[0068]
[0069] Then, use the test set data for testing. After performing the "softmax" operation on the output during the testing process, assign the category with the highest classification value to the pixel.
[0070] To illustrate the effectiveness of the method of the present invention, the present invention also conducts the following verification examples.
[0071] Specifically, the experimental environment and content: Experimental environment: Python 3.9, Win10; Experimental configuration: Intel Xeon Platinum 8269CY CPU, 128GB RAM, and NVIDIA GeForce RTX3080 GPU.
[0072] Experimental content: 1. Comparison of networks with different loss function strategies; 2. Comparison with other semantic segmentation networks.
[0073] Experimental data: The AIR-PolSAR-Seg dataset is used for verification. The AIR-PolSAR-Seg dataset provides fully polarized data obtained by the Gaofen-3 satellite. The spatial resolution is 8m, and six typical terrain categories, namely residential area, industrial area, natural area, land use area, water area, and others, are labeled at the pixel level. Such asFigure 4 As shown, an amplitude image in a polarized state and the corresponding color ground truth are given. There is one for each polarization, a total of 4, and one of them is shown in the figure. The color ground truth is the ground annotation, and different ground classes are represented by different colors.
[0074] Experimental steps: The modulus values of the 4 polarization channels (HH, HV, VH, VV) of the SAR image data are used as input data. The experiment is carried out according to the above steps.
[0075] Evaluation metrics: The experiment evaluated the accuracy of each scene in the dataset. Three comparison metrics were used for evaluation, including MPA, MIOU, and FWI. Among them, MPA is mainly used for classification tasks and can evaluate the class accuracy of the classification results. FWI is a frequency-based weighting of the IOU of each class and is more stable. MIOU is the most important and commonly used metric in all segmentation tasks. Therefore, the results of this embodiment improve MIOU while maintaining class accuracy and FWI.
[0076]
[0077]
[0078]
[0079] The calculation methods of each metric are as shown above, where P ij represents the number of pixels belonging to class i but recognized as class j, and K represents the total number of classes. MPA is used to evaluate the pixel segmentation performance of all classes, and MIOU and FWI are used to evaluate the overall segmentation performance of all classes.
[0080] Analysis of experimental results:
[0081] Experimental content 1:
[0082] In the fine-tuning stage, the results of the network under different loss functions are given in Table 1. Among them, 3-CE, 1-Lovasz, 3-Lovasz, and 3-CE + 3-Lovasz respectively represent that the three outputs (the main output and the outputs of two ARM modules) use cross-entropy loss (BiSeNet), only the main output and apply Lovasz loss, the three outputs use Lovasz loss, and the three outputs use a weighted loss of cross-entropy and Lovasz (the method proposed in the present invention).
[0083] It can be clearly seen from Table 1 that the MIOU (60.65%) obtained by the method proposed in the present invention is higher than the MIOU (56.51%) of BiSeNet and also higher than the MIOU (59.67%) of the network trained only with Lovasz. While taking FWI into account, the method proposed in the present invention obtains the best MIOU metric and an acceptable MPA.
[0084] Table 1 Evaluation Index Results of AIR-PolSAR-Seg Segmentation Using Different Loss Functions
[0085]
[0086]
[0087] Experimental Content 2:
[0088] Taking one group as an example, the classification results of different semantic segmentation networks for AIR-PolSAR-Seg are as Figure 5 shown, where (a) is the sample pseudo-color image, (b) is the ground truth category, (c) is the segmentation result of ENetV2, (d) is the segmentation result of ERFNet, (e) is the segmentation result of DeepLabv3, (f) is the segmentation result of BiSeNet, and (g) is the segmentation result of the method of the present invention.
[0089] Table 2 Three Metrics and Time of AIR-PolSAR-Seg Segmentation Based on Different Semantic Segmentation Networks
[0090]
[0091] According to Table 2, through the quantitative analysis of the indicators and time costs of different networks for AIR-PolSAR-Seg, the MPA, MIOU, and FWI evaluations of the method proposed in the present invention are far better than the other methods listed, especially MIOU. The MIOU obtained by the method proposed in the present invention is 60.65%, which is 31.38% higher than the real-time semantic segmentation network EnetV2 and 28.23% higher than the high-precision semantic segmentation network DeepLabv3. At the same time, LoSARNet has the highest recognition speed, exceeding the two real-time semantic segmentation networks EnetV2 and ERFNet. This shows that the method proposed in the present invention can maintain real-time performance while completing high-precision segmentation tasks and is suitable for the actual application scenarios of SAR images.
[0092] In short, compared with the traditional BiSeNet in the full-polarization dataset AIR-PolSAR-Seg, the method proposed in the present invention has higher classification accuracy. The improved loss function can better fill the gaps in segmentation and obtain more comprehensive classification results.
[0093] The present invention also discloses a SAR image real-time segmentation device based on Lovasz loss and a lightweight bilateral network, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned method of a SAR image real-time segmentation device based on Lovasz loss and a lightweight bilateral network.
[0094] It should be noted that for the information interaction, execution process, etc. among the above-mentioned devices, since they are based on the same concept as the method embodiment of the present invention, for their specific functions and the technical effects brought about, reference can be specifically made to the method embodiment part, and details will not be elaborated here.
[0095] The device can be a computing device such as a desktop computer, a notebook, a palm computer, a radar, and a cloud server. The device may include but is not limited to a processor and a memory. Those skilled in the art can understand that it may include more or fewer components, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0096] The so-called processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0097] In some embodiments, the memory may be an internal storage unit of the extraction device, such as the hard disk or memory of the extraction device. In other embodiments, the memory may also be an external storage device of the extraction device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the extraction device. Further, the memory may also include both the internal storage unit and the external storage device of the extraction device. The memory is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory may also be used to temporarily store data that has been output or will be output.
Claims
1. A real-time segmentation method for SAR images based on Lovász loss and lightweight bilateral network, characterized in that, It includes the following steps: Pre-train and fine-tune the BiSeNet network in sequence; wherein, during the pre-training process, the loss function of the BiSeNet network adopts the cross-entropy loss function, and during the fine-tuning training process, the loss function of the BiSeNet network adopts a joint loss function, and the joint loss function includes the Lovasz loss function and the cross-entropy loss function; Segment the polarimetric SAR image based on the BiSeNet network after fine-tuning training; The Lovasz loss function is: Among them, Loss(f)2 is the Lovasz loss function, C is the number of pixel categories in the SAR image, and c is the category ordinal number of the pixels in the SAR image. is the extended loss for each category; The cross-entropy loss function is: Among them, Loss(f)1 is the cross-entropy loss function, p is the number of pixels input into the BiSeNet network, is the true label of the i-th pixel, is the classification result of the BiSeNet network for the i-th pixel as the probability; The joint loss function is: Wherein, Loss(f)3 is the joint loss function.
2. The real-time SAR image segmentation method based on Lovasz loss and lightweight bilateral network according to claim 1, characterized in that During the pre-training / fine-tuning training process, it includes: Alternately pre-train / fine-tune the BiSeNet network and verify it, and record the network weights and the corresponding verification losses after each pre-training / fine-tuning training; When the number of verifications reaches the first threshold, select the network weights corresponding to the minimum verification loss as the weights of the BiSeNet network.
3. The real-time SAR image segmentation method based on Lovasz loss and lightweight bilateral network according to claim 2, characterized in that During the pre-training / fine-tuning training process, when the number of consecutive pre-training / fine-tuning training reaches the second threshold and the corresponding verification loss does not decrease, end the pre-training / fine-tuning training process.
4. The real-time SAR image segmentation method based on Lovasz loss and lightweight bilateral network according to claim 3, characterized in that, During the pre-training / fine-tuning training process, it also includes: Dynamically adjust the learning rate of the BiSeNet network.
5. A real-time SAR image segmentation device based on Lovász loss and lightweight bilateral network, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it realizes the method of a SAR image real-time segmentation device based on the LOVASZ loss and the lightweight bilateral network according to any one of claims 1-4 above.