Multi-modal aircraft target segmentation method based on improved loss
By introducing improved loss function and Shuffle-Attention module into the aircraft target segmentation method, combined with the Channel Shuffle method, the problems of unsatisfactory segmentation effect, large parameters and slow inference speed in the prior art are solved, and better segmentation effect and performance are achieved.
Patent Information
- Application Number
- CN202510276349.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-24
AI Technical Summary
The existing aircraft target segmentation method is not ideal in terms of segmentation effect, with large parameters and slow inference speed, especially when the background is similar to the target, the segmentation performance is affected.
A multimodal aircraft target segmentation method based on improved losses is proposed. By constructing an aircraft target semantic segmentation network and introducing the Shuffle-Attention module, combined with the Channel Shuffle method, the key point loss function is designed to utilize geometric information to improve the segmentation effect and inference speed.
By improving the loss function and network structure, the segmentation effect and performance of aircraft target segmentation is significantly improved, the situation where the background is mistaken for the target area is reduced, the performance of the algorithm in complex backgrounds is improved, and the parameter quantity and calculation cost are reduced.
Smart Images

Figure CN120198666A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image segmentation, and particularly relates to an aircraft target segmentation method. Background Art
[0002] In the field of computer vision, semantic segmentation of aircraft targets is an important task. Its main purpose is to identify and distinguish the categories of each pixel in the image, and accurately identify the aircraft and its different parts, so as to segment the aircraft target from the complex background image. With the popularization of remote sensing technology and drone photography, it has become easier to obtain high-resolution aerial images, which provides broad application prospects for the semantic segmentation of aircraft targets. Fine aircraft target segmentation not only helps to identify aircraft models, but also can be used for tasks such as analyzing air traffic flow and supporting damage detection. Especially for small targets, semantic segmentation can better complete the target extraction task because it conforms more closely to the target shape. For example, the Chinese patent application with publication number CN118537550A discloses a method for segmenting aircraft surface features based on contour constraint optimization. By collecting image data and performing operations such as feature extraction, contour constraint extraction, target prediction, and mask optimization, the edge segmentation effect is optimized. The Chinese patent application with publication number CN118608912A discloses a method for detecting foreign objects on aircraft runways based on image segmentation. By designing multi-scale feature learning fusion and random fusion feature training, the effect of enhancing foreign object detection positioning and algorithm robustness is achieved. The paper "Dual-Threshold Segmentation Method for Infrared Image Sequences of Aircraft Targets" (Tu Jianping, Peng Yingning, Acta Armamentarii) proposes a dual-threshold segmentation method for infrared image sequences of aircraft targets. By determining the threshold range and searching for the optimal value using the maximum entropy and genetic algorithms, real-time segmentation of the aircraft is achieved.
[0003] Although the existing target segmentation methods have achieved remarkable results, there are still the following technical problems: First, the loss functions designed in the existing solutions usually directly calculate the loss between the predicted mask and the ground truth mask. For example, the MSE loss is the difference between the entire predicted image and the ground truth mask image, while the Dice loss is the intersection over union ratio of the predicted mask region and the ground truth mask region. Although these methods indirectly consider and optimize the differences in the position, shape, etc. between the predicted mask and the ground truth mask region, they fail to make full use of the geometric information contained in the mask labels to guide the training, resulting in a relatively slow training convergence process and difficulty in better learning the boundary region, and the segmentation effect is not ideal.
[0004] Second, when the background is similar to the target, there is a certain regular deviation between the predicted target region and the ground truth mask region. For example, the predicted mask is likely to identify part of the background as the target region, which affects the segmentation performance of the algorithm.
[0005] Thirdly, existing segmentation methods usually consider less performance such as the inference speed of the algorithm, resulting in a large number of parameters and slow inference speed. Summary of the Invention
[0006] The present invention proposes a multi-modal aircraft target segmentation method based on an improved loss, and its purpose is to solve the problems of unsatisfactory segmentation effect, many parameters, and slow inference speed of existing target segmentation methods.
[0007] The technical solution of the present invention is as follows: A multi-modal aircraft target segmentation method based on an improved loss, the steps include: Step S1, construct an aircraft target semantic segmentation network, the input of the aircraft target semantic segmentation network is a multi-modal image to be segmented, and the output is a predicted mask; Step S2, construct a key point loss function, and train the aircraft target semantic segmentation network based on the key point loss function; Step S3, input the multi-modal image to be segmented into the trained aircraft target semantic segmentation network to obtain a predicted mask, and finally segment the target from the image according to the predicted mask.
[0008] As a further improvement of the multi-modal aircraft target segmentation method based on the improved loss: the aircraft target semantic segmentation network is based on the U-Net network structure, and its input layer splices the input multi-modal image according to the channel dimension to obtain a multi-modal feature map, and adds a Shuffle-Attention module after each convolutional layer of the encoder and decoder.
[0009] As a further improvement of the multi-modal aircraft target segmentation method based on the improved loss: the Shuffle-Attention module combines a channel shuffle layer with an attention mechanism, first shuffles the channels, and then generates an attention mask through a 1x1 convolution.
[0010] As a further improvement of the multi-modal aircraft target segmentation method based on the improved loss: in step 2, the training is carried out in batches, and each batch contains e training samples; For each batch: first, input the multi-modal images contained in each training sample in this batch into the aircraft target semantic segmentation network respectively to obtain corresponding predicted masks, then calculate the total loss function of each sample based on the predicted masks and the corresponding true masks of each training sample, and then optimize the parameters of the aircraft target semantic segmentation network based on the average value of the total loss functions of all training samples in this batch to complete the training of this batch; Through the above method, the training of the specified number of batches is completed.
[0011] As a further improvement of the multi-modal aircraft target segmentation method based on the improved loss: during training, the network parameters are updated according to the total loss function, and the total loss function is calculated as follows: ; where is the basic loss function of the current training sample, is the key point loss function corresponding to the th true mask region in the training sample, is the number of true mask regions in the training sample; in the true mask of the training sample, each true mask region corresponds to a target to be segmented; is the weight coefficient.
[0012] As a further improvement of the multi-modal aircraft target segmentation method based on the improved loss, the basic loss function is calculated as follows: ; where is the number of pixels in the training sample, is the pixel value corresponding to the th pixel in the predicted mask, that is, the predicted probability value, is the pixel value corresponding to the th pixel in the true mask.
[0013] As a further improvement of the multi-modal aircraft target segmentation method based on the improved loss, the calculation method of the key point loss function is as follows: First, determine the segmentation region in the predicted mask according to the segmentation threshold: regard the pixels greater than the segmentation threshold in the predicted mask as the attention region, and divide the attention region according to the image connection relationship to obtain several independent segmentation regions; Then, for each true mask region, calculate the intersection over union (IoU) between the true mask region and each segmentation region respectively, and take the segmentation region with the largest IoU as the predicted mask region corresponding to the true mask region; Then, calculate the key point loss function for each group of true mask regions and predicted mask regions.
[0014] As a further improvement of the multi-modal aircraft target segmentation method based on the improved loss, for the key point loss function corresponding to the th true mask region, its calculation method is as follows: Obtain the boundary point sequence ,..., , is the boundary point sequence The number of boundary points in the middle; Let the predicted mask region corresponding to the th real mask region be the th predicted mask region, and obtain the boundary point sequence of this predicted mask region ,..., , is the number of boundary points in the boundary point sequence ; Randomly select boundary points from the boundary point sequence as key points; for each key point, calculate the shortest distance between this key point and the boundary point sequence as the loss of this key point; for the th key point, its corresponding loss is: ; Then the key point loss function corresponding to the th real mask region is: .
[0015] As a further improvement of the multi-modal aircraft target segmentation method based on the improved loss: The boundary points of the real mask region refer to the pixel points that simultaneously satisfy the following two conditions: Condition a, its own pixel value is 1; Condition b, there are both pixel points with pixel value 1 and pixel points with pixel value 0 among the 8 adjacent pixels around this pixel; The boundary points of the predicted mask region refer to the pixel points that simultaneously satisfy the following two conditions: Condition c, its own pixel value is greater than the segmentation threshold; Condition d, there are both pixel points with pixel value greater than the segmentation threshold and pixel points with pixel value less than or equal to the segmentation threshold among the 8 adjacent pixels around this pixel.
[0016] As a further improvement of the multi-modal aircraft target segmentation method based on the improved loss: the segmentation threshold θ = 0.5.
[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention designs a key-point loss function combined with geometric information. Based on the Frechet distance, this loss function proposes a measurement method for measuring the similarity between the predicted mask and the boundary of the true mask. This method can make full use of the geometric information contained in the true mask, and can better capture and maintain the shape characteristics of the boundary curve, prompting the boundary features learned by the model to be closer to the true boundary, effectively reducing the situation where the background area is misrecognized as the target area, and improving the segmentation effect and performance of the algorithm in complex backgrounds.
[0018] 2. Considering that the time cost of calculating the Frechet distance for all points on the entire closed curve is relatively high, the present invention calculates the distance between boundary points by means of random sampling based on key points. By randomly sampling the boundary points of the true mask as key points and simplifying the complex curve into a set of key points, the computational complexity is significantly reduced. Moreover, by adjusting the sampling density, a trade-off can be made between computational accuracy and computational cost, making the method more flexible.
[0019] 3. The present invention also proposes a lightweight semantic segmentation network for aircraft targets. Combining the Channel Shuffle method, it realizes information fusion between channels at a relatively low cost by reorganizing the channels of the feature map. This network can not only maintain the expressive ability of the feature map through the information flow between different groups of convolutional layers, but also reduce the use of pointwise convolutions (1x1 convolutions), thereby significantly reducing the number of parameters and computational volume while enabling the model to better understand the context information of the image, and improving the inference speed of the model.
[0020] 4. The present invention can simultaneously achieve the segmentation of multiple targets by calculating the losses of multiple mask regions respectively. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic diagram showing the relationship between a certain true mask region r and a certain predicted mask region s in the specific implementation manner. SPECIFIC IMPLEMENTATION MANNER
[0022] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0023] A multi-modal aircraft target segmentation method based on an improved loss, the steps include: Step S1. Construct a semantic segmentation network for aircraft targets. The input of the semantic segmentation network for aircraft targets is a multi-modal image to be segmented, and the output is a predicted mask.
[0024] Considering that the method of the present invention needs to run on edge devices, the network adopts a lightweight design. Specifically, the aircraft target semantic segmentation network is based on the U-Net network structure. Its input layer splices the input infrared and visible light images in the channel dimension to obtain a multi-modal feature map, and a Shuffle-Attention module is added after each convolutional layer of the encoder and decoder. The Shuffle-Attention module combines the channel shuffle layer (ChannelShuffle) with the attention mechanism, first shuffles the channels, and then generates an attention mask through a 1x1 convolution. Its introduction can not only achieve efficient information fusion between the channels of the feature map, thereby reducing the number of parameters and improving the inference speed, but also improve the network's adaptive ability to weight allocation for important channels, and more specifically learn the weights of important channels. The structure of the aircraft target semantic segmentation network is shown in the following table.
[0025] Table 1: Structure levels and main parameters of the aircraft target semantic segmentation network.
[0026] Level Type Output Size Number of Channels Stride Description 1 Input(Infrared); Input(Visible Light) 512×512×3;512×512×3 - - Original image input, both are 512×512×3 2 Conv2D; Conv2D 256×256×24;256×256×24; 24 2 Convolve the two inputs separately using 3x3 depthwise separable convolution 3 Channel Shuffle; ChannelShuffle 256×256×24;256×256×24; 24 - Apply channel shuffle to the features extracted from the two inputs respectively 4 Concatenate 256×256×48 48 - Concatenate the features extracted from the two inputs 5 Conv2D 128×128×48 48 2 3x3 depthwise separable convolution 6 Channel Shuffle 128×128×48 48 - Channel shuffle 7 Conv2D 64×64×96 96 2 3x3 depthwise separable convolution 8 Channel Shuffle 64×64×96 96 - Channel shuffle 9 Conv2D 32×32×192 192 2 3x3 depthwise separable convolution 10 Channel Shuffle 32×32×192 192 - Channel shuffle 11 DeConv2D 64×64×96 96 2 4x4 transposed convolution 12 Channel Shuffle 64×64×96 96 - Channel shuffle 13 DeConv2D 128×128×48 48 2 4x4 transposed convolution 14 Channel Shuffle 128×128×48 48 - Channel shuffle 15 DeConv2D 256×256×24 24 2 4x4 transposed convolution 16 Channel Shuffle 256×256×24 24 - Channel shuffle 17 DeConv2D 512×512×1 1 2 4x4 transposed convolution, with sigmoid as the activation function ; Step S2, construct a key point loss function, and train the aircraft target semantic segmentation network based on the key point loss function.
[0027] Before training, first construct a training set and a test set: collect aligned visible light images and infrared images of aircraft targets, uniformly crop and resize the images to 512×512×3, and manually annotate the segmentation mask to obtain the true mask label. Among them, the pixel value of the pixel points belonging to the aircraft target is set to 1, and other areas are 0. The <visible light image, infrared image, true mask label> after annotation is used as a sample to form a data set, and then randomly divided into training samples and test samples according to a ratio of 7:3 to obtain a training set and a test set.
[0028] During training, a batch training method is adopted, and each batch contains e training samples. For each batch: first input the multi-modal images contained in each training sample in this batch into the aircraft target semantic segmentation network to obtain the corresponding predicted mask, then calculate the total loss function of each sample based on the predicted mask of each training sample and the corresponding true mask, and then optimize the parameters of the aircraft target semantic segmentation network based on the average value of the total loss functions of all training samples in this batch to complete the training of this batch. Through the above method, the training of the specified number of batches is completed.
[0029] In this embodiment, the learning rate during training is set to 0.0002, the number of training times is 500, and the batch size is e = 64.
[0030] The total loss function is calculated as follows: ; Among them, is the basic loss function of the current training sample, is the key point loss function corresponding to the th true mask region in the training sample, is the number of true mask regions in the training sample. In the true mask of the training sample, each true mask region corresponds to a target to be segmented respectively. is the weight coefficient. In this embodiment, .
[0031] (1) Basic loss function The calculation method of ; Among them, is the number of pixels in the training sample, is the pixel value corresponding to the th pixel in the predicted mask, that is, the predicted probability value, is the pixel value corresponding to the th pixel in the true mask.
[0032] (2) The calculation method of the key point loss function is: First, determine the segmentation region in the predicted mask according to the segmentation threshold: regard the part greater than the segmentation threshold in the predicted mask as the attention region, and divide the attention region according to the image connection relationship to obtain several independent segmentation regions (that is, the connected parts are regarded as one segmentation region).
[0033] In this embodiment, the segmentation threshold θ = 0.5.
[0034] Then, for each true mask region, calculate the intersection over union (IoU) between the true mask region and each segmentation region respectively, and regard the segmentation region with the largest IoU as the predicted mask region corresponding to the true mask region.
[0035] For the key point loss function corresponding to the th true mask region, its calculation method is: Obtain the boundary point sequence ,..., of this true mask region, is the number of boundary points in the boundary point sequence . The boundary points of the true mask region refer to the pixel points that satisfy the following two conditions at the same time: Condition a: Its own pixel value is 1; Condition b: Among the 8 adjacent pixels around this pixel, there are both pixel points with pixel value 1 and pixel points with pixel value 0.
[0036] Let the predicted mask region corresponding to the th true mask region be the th predicted mask region, and obtain the sequence of boundary points of this predicted mask region ,..., . Let be the number of boundary points in the sequence of boundary points . The boundary points of the predicted mask region refer to the pixel points that satisfy the following two conditions simultaneously: Condition c: Its own pixel value is greater than the segmentation threshold;
[0037] Randomly select boundary points from the sequence of boundary points as key points. The value of can be determined according to the calculation accuracy and calculation cost. For each key point, calculate the shortest distance between this key point and the sequence of boundary points as the loss of this key point. For the th key point, its corresponding loss is: ; Then the key point loss function corresponding to the th true mask region is: : .
[0038] Figure 1 In Figure 1 and , taking two key points as examples, the corresponding shortest distances and are calculated respectively, and these two distances are their corresponding losses.
[0039] Step S3: Input the multi-modal image to be segmented into the trained aircraft target semantic segmentation network to obtain a predicted mask, and finally segment the target from the image according to the predicted mask.
[0040] It should be noted that for those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. The scope of the present invention is defined by the claims rather than the above description.
Claims
1. A multimodal aircraft target segmentation method based on improved loss, characterized in that the steps Including: Step S1: Construct an aircraft target semantic segmentation network. The input of the aircraft target semantic segmentation network is a multi-modal image to be segmented, and the output is a prediction mask. Step S2: Construct a key point loss function and train the aircraft target semantic segmentation network based on the key point loss function. Step S3: Input the multi-modal image to be segmented into the trained aircraft target semantic segmentation network to obtain a prediction mask, and finally segment the target from the image according to the prediction mask.
2. The multimodal aircraft target segmentation method based on improved loss according to claim 1, characterized in that: The aircraft target semantic segmentation network is based on the U-Net network structure. Its input layer concatenates the input multi-modal images along the channel dimension to obtain a multi-modal feature map, and adds a Shuffle-Attention module after each convolutional layer in the encoder and decoder.
3. The multimodal aircraft target segmentation method based on improved loss according to claim 2, wherein: The Shuffle-Attention module combines a channel shuffle layer with an attention mechanism. First, it shuffles the channels, and then generates an attention mask through a 1x1 convolution.
4. The multimodal aircraft target segmentation method based on improved loss according to claim 1, characterized in that: In step 2, the training is carried out in batches, and each batch contains e training samples. For each batch: First, input the multi-modal images contained in each training sample in this batch into the aircraft target semantic segmentation network respectively to obtain the corresponding prediction masks. Then, calculate the total loss function of each sample based on the prediction masks and the corresponding ground truth masks of each training sample. Then, optimize the parameters of the aircraft target semantic segmentation network based on the average value of the total loss functions of all training samples in this batch to complete the training of this batch. Through the above method, the training of the specified number of batches is completed.
5. The multi-modal aircraft target segmentation method based on improved loss according to claim 1, characterized in that: During training, the network parameters are updated according to the total loss function, and the total loss function is calculated as follows: ; Among them, is the basic loss function of the current training sample, is the key point loss function corresponding to the th true mask region in the training sample, is the number of true mask regions in the training sample; in the true mask of the training sample, each true mask region corresponds to a target to be segmented respectively; is the weight coefficient.
6. The multimodal aircraft target segmentation method based on improved loss according to claim 5, wherein, Basic loss function is calculated as follows: ; Among them, is the number of pixels in the training sample, is the pixel value corresponding to the th pixel in the prediction mask, that is, the predicted probability value, is the pixel value corresponding to the th pixel in the ground truth mask.
7. The multimodal aircraft target segmentation method based on improved loss according to claim 5, characterized in that, The calculation method of the key point loss function is as follows: First, determine the segmentation regions in the prediction mask according to the segmentation threshold: regard the values greater than the segmentation threshold in the prediction mask as the attention regions, and divide the attention regions according to the image connection relationship to obtain several independent segmentation regions. Then, for each ground truth mask region, calculate the intersection over union (IoU) between this ground truth mask region and each segmentation region respectively, and regard the segmentation region with the largest IoU as the prediction mask region corresponding to this ground truth mask region. Then calculate the key point loss function for each group of ground truth mask regions and prediction mask regions.
8. The multimodal aircraft target segmentation method based on improved loss according to claim 7, characterized in that For the keypoint loss function corresponding to the th true mask region, its calculation method is as follows: Obtain the sequence of boundary points of the true mask region ,..., , is the sequence of boundary points is the number of boundary points in the sequence Let the predicted mask region corresponding to the th true mask region be the th predicted mask region, and obtain the sequence of boundary points of this predicted mask region ,..., , is the number of boundary points in the sequence of boundary points ; Randomly select from the boundary point sequence and use these boundary points as key points; For each key point, calculate the shortest distance between the key point and the boundary point sequence as the loss of the key point; for the th key point, its corresponding loss is: ; Then the keypoint loss function corresponding to the th true mask area is: 。 9. The multi-modal aircraft target segmentation method based on improved loss according to claim 8, characterized in that: The boundary points of the ground truth mask region refer to the pixel points that satisfy the following two conditions simultaneously: Condition a: Its own pixel value is 1. Condition b: Among the 8 adjacent pixels around this pixel, there are both pixel points with a pixel value of 1 and pixel points with a pixel value of 0. The boundary points of the prediction mask region refer to the pixel points that satisfy the following two conditions simultaneously: Condition c: Its own pixel value is greater than the segmentation threshold. Condition d: Among the 8 adjacent pixels around this pixel, there are both pixel points with a pixel value greater than the segmentation threshold and pixel points with a pixel value less than or equal to the segmentation threshold.
10. The multimodal aircraft target segmentation method based on improved loss according to any one of claims 7 to 9, characterized in that: The segmentation threshold θ = 0.5.
Citation Information
Patent Citations
Aircraft surface feature segmentation method based on contour constraint optimization
CN118537550A
Airplane runway foreign matter detection method based on image segmentation
CN118608912A