A method and device for aircraft detection in optical remote sensing images guided by regional saliency

Through the regional significance guidance method, the significant map of the airport area is extracted and processed, and the problem of difficult balance between the detection rate and false alarm rate in the peripheral area of ​​the airport is solved, high-performance aircraft target detection is achieved, and the accuracy and reliability of the detection results are ensured.

CN113743185BActive Publication Date: 2025-05-13BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110644392.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-09
Publication Date
2025-05-13
Estimated Expiration
2041-06-09

AI Technical Summary

Technical Problem

The prior art is difficult to achieve the balance between aircraft target detection and false alarm rate in the periphery of the airport. At the same time, data enhancement operations in the training stage cause the center point of the sample labeling box to shift, affecting the characteristics learned by the network.

Method used

Using a regional significance-guided method, the significant graph of the airport area is extracted through a feature ensemble deep learning network, and converted it into convex polygons, divided into three levels of sub-regions of concern, and the confidence of the aircraft target is treated using a dual-threshold weighting method.

Benefits of technology

It improves the aircraft target detection rate in the periphery of the airport, reduces the false alarm rate, and ensures the accuracy of sample labeling frames, and improves the network's ability to learn the most prominent features of aircraft targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113743185B_ABST
    Figure CN113743185B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for detecting aircraft in an optical remote sensing image guided by regional saliency, the method comprising scanning an original image, extracting an airport area and detecting an aircraft target; converting a saliency map of the airport area into a convex polygon of the airport area; dividing the original image into sub-areas with three levels of attention based on the convex polygon of the airport area, and using a double threshold weighting method to perform weighted processing on the confidence of the detected aircraft target. According to the scheme of the present invention, the aircraft target in the airport scene of a large field of view and high resolution optical remote sensing image can be well detected, and good results are also achieved when facing complex conditions such as scale changes and dense distribution of aircraft targets, and the false alarm rate in the peripheral area of ​​the airport can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of target detection, and in particular to a method and a device for detecting aircraft in optical remote sensing images guided by regional saliency. Background Art

[0002] Target detection is a basic problem in remote sensing image interpretation tasks, that is, to determine whether there is a target of interest in the image and its location through computer vision algorithms. It is the basis for tasks such as target tracking and trend analysis. Early detection methods were based on traditional image processing methods, describing the target through manually designed features, and using traditional machine learning methods to complete classification. Since the widespread application of deep learning technology, convolutional neural networks have demonstrated powerful automatic feature extraction capabilities and the potential for continuous learning and evolution, which has significantly improved target detection capabilities and stability.

[0003] According to the generation method of target candidate regions, target detection methods based on deep learning can be divided into two categories: anchor box methods and non-anchor box methods. The anchor box method uses a set of rectangular candidate boxes with different scales and ratios to generate candidate regions for the target. The convolutional neural network determines whether the area in the box contains the target. If it contains the target, its category is determined and the candidate box is regressed to a more precise position. The non-anchor box method does not have an explicit candidate region generation process. Instead, it models the target with the help of key points and feature lines, and uses encoding and decoding to complete the regression of the target border. It can break the constraints of the anchor box on the target size and ratio, which is conducive to dealing with the scale diversity of aircraft targets in remote sensing. At the same time, the detection process is more concise, which helps to improve the detection speed.

[0004] In airport scene images, most airplanes are distributed in areas such as airport runways and aprons. The background of the airport's peripheral areas is relatively complex, and false alarms are prone to occur during detection; however, there may be scattered airplanes in these areas. Some existing methods reduce the false alarm rate by extracting areas such as airport runways and aprons and detecting aircraft targets in them, but they are prone to missed detections in the airport's peripheral areas. In the training stage of the aircraft target detection method, data augmentation operations such as scaling and cropping may truncate the sample annotation box, causing the center point of the sample to shift, which is not conducive to the backbone network learning the most significant features of the sample. Summary of the invention

[0005] In order to solve the above technical problems, the present invention proposes a method and device for aircraft detection in optical remote sensing images guided by regional saliency. The method and device are used to solve the technical problems in the prior art that the detection rate and false alarm rate in the airport peripheral area cannot be balanced, and the data enhancement operation in the training stage causes the center point of the sample annotation box to be offset.

[0006] According to a first aspect of the present invention, a method for detecting aircraft in optical remote sensing images guided by regional saliency is provided, the method comprising the following steps:

[0007] Step S101: Scan the original image, extract the airport area and detect the aircraft target;

[0008] The airport area extraction comprises: slicing the original image according to multi-scale factors, classifying the slices based on a trained feature integration deep learning network; predicting slices in which the airport area exists in the original image according to the classification results, generating a salient map of the airport area in the original image in a Gaussian weighted manner; processing the salient map as an extraction result of the airport area;

[0009] The aircraft target detection includes: inputting the original image into the backbone network, correcting the sample annotation frame according to geometric prior knowledge, extracting features and detecting the aircraft target;

[0010] Step S102: converting the saliency map of the airport area into a convex polygon of the airport area;

[0011] Step S103: Based on the convex polygon of the airport area, the original image is divided into sub-areas with three levels of attention, and the confidence of the detected aircraft target is weighted using a double threshold weighting method.

[0012] According to a second aspect of the present invention, a regional saliency guided optical remote sensing image aircraft detection device is provided, the device comprising:

[0013] Extraction module: configured to scan the original image, extract the airport area and detect the aircraft target;

[0014] The airport area extraction comprises: slicing the original image according to multi-scale factors, classifying the slices based on a trained feature integration deep learning network; predicting slices in which the airport area exists in the original image according to the classification results, generating a salient map of the airport area in the original image in a Gaussian weighted manner; processing the salient map as an extraction result of the airport area;

[0015] The aircraft target detection includes: inputting the original image into the backbone network, correcting the sample annotation frame according to geometric prior knowledge, extracting features and detecting the aircraft target;

[0016] A conversion module: configured to convert the saliency map of the airport area into a convex polygon of the airport area;

[0017] Confidence weighted processing module: configured to divide the original image into sub-areas with three levels of attention based on the convex polygon of the airport area, and use a double threshold weighted method to perform weighted processing on the confidence of the detected aircraft target.

[0018] According to a third aspect of the present invention, there is provided a regional saliency guided optical remote sensing image aircraft detection system, comprising:

[0019] A processor, which is used to execute multiple instructions;

[0020] A memory for storing a plurality of instructions;

[0021] The plurality of instructions are used to be stored by the memory, and loaded and executed by the processor for executing the aforementioned regional saliency-guided optical remote sensing image aircraft detection method.

[0022] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the aforementioned regional significance-guided optical remote sensing image aircraft detection method.

[0023] According to the above scheme of the present invention, the airport salient area is first extracted by a feature integration deep learning network, and then the improved anchor-free frame method is used to realize the detection of aircraft targets. Finally, the airport extraction and target detection are combined to complete the high-performance detection of aircraft targets in the airport scene of the large-field-of-view high-resolution optical remote sensing image, and reduce the false alarm rate in the peripheral area of ​​the airport. The above scheme of the present invention can well detect the aircraft targets in the airport scene of the large-field-of-view high-resolution optical remote sensing image, and also achieves good results when facing complex conditions such as the scale change and dense distribution of the aircraft targets, and can reduce the false alarm rate in the peripheral area of ​​the airport. This method uses a relatively lightweight backbone network to achieve high-performance detection of aircraft targets with a trade-off between speed and accuracy, and has good practical application value.

[0024] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention, and the present invention is described by providing the following accompanying drawings. In the accompanying drawings:

[0026] Figure 1 A flow chart of a method for detecting aircraft in optical remote sensing images guided by regional saliency according to an embodiment of the present invention;

[0027] Figure 2This is a schematic diagram of extracting an airport area according to one embodiment of the present invention;

[0028] Figure 3 A schematic diagram of detecting an aircraft target according to an embodiment of the present invention;

[0029] Figure 4 A schematic diagram of obtaining an airport area and detecting an aircraft target area in a serial and parallel manner according to an embodiment of the present invention;

[0030] Figure 5 A schematic diagram of a composition method of a feature integration deep learning network according to an embodiment of the present invention;

[0031] Figure 6 A schematic diagram of a method for generating a saliency map of an airport area according to an embodiment of the present invention;

[0032] Figure 7 A schematic diagram of correcting a sample annotation frame using geometric prior knowledge according to an embodiment of the present invention;

[0033] Figure 8 A schematic diagram of converting a saliency map of an airport area into a convex polygon of the airport area according to an embodiment of the present invention;

[0034] Fig. 9 A schematic diagram of dividing three regions of different interest levels and weighting confidence levels according to an embodiment of the present invention;

[0035] Fig.10 This is a structural block diagram of a regional saliency-guided optical remote sensing image aircraft detection device according to one embodiment of the present invention. DETAILED DESCRIPTION

[0036] First combine Figure 1 The following is a flow chart of a method for detecting aircraft in optical remote sensing images using regional saliency guidance according to an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:

[0037] Step S101: Scan the original image, extract the airport area and detect the aircraft target;

[0038] The airport area extraction comprises: slicing the original image according to multi-scale factors, classifying the slices based on a trained feature integration deep learning network; predicting slices in which the airport area exists in the original image according to the classification results, generating a salient map of the airport area in the original image in a Gaussian weighted manner; processing the salient map as an extraction result of the airport area;

[0039] The aircraft target detection includes: inputting the original image into the backbone network, correcting the sample annotation frame according to geometric prior knowledge, extracting features and detecting the aircraft target;

[0040] Step S102: converting the salient map of the airport area into a convex polygon of the airport area;

[0041] Step S103: Based on the convex polygon of the airport area, the original image is divided into sub-areas with three levels of attention, and the confidence of the detected aircraft target is weighted using a double threshold weighting method.

[0042] like Figure 2 As shown, the airport area extraction: the original image is sliced ​​according to multi-scale factors, and the slices are classified based on the trained feature integration deep learning network; the slices containing the airport area in the original image are predicted according to the classification results, and a salient map containing the airport area in the original image is generated by using Gaussian weighting; the salient map of the airport area is processed as the extraction result of the airport area, including:

[0043] Step S201: The original image S A According to the multi-scaling factor α 1 ,…,α L-1 ,α L Split and downsample with overlap to get slices S of the same size 1 ,S 2 ,…,S N ;

[0044] Step S202: Slice {S 1 ,S 2 ,…,S N}∈S A The feature integration deep learning network predicts whether it belongs to the airport area. If it does, the slice classification result T n Set to 1, otherwise T n Set to 0;

[0045] Step S203: According to the classification result T n , predicting the area where the airport exists in the original image, and generating a saliency map of the airport area using a Gaussian weighting method;

[0046] Step S204: using a threshold segmentation method to generate a binary image for the salient image of the airport area, and the binary image is used as the extraction result of the airport area.

[0047] In this embodiment, the multi-scale factor α L The L in is set to 3, the overlap rate is set to 50%, and the threshold segmentation method is the Otsu method.

[0048] The composition and principle of the feature integration deep learning network are as follows: extract the global features of each slice by the first ResNet-50 network, perform convolution operations of multiple convolution sizes on the output of the last convolution layer conv5_x of the first ResNet-50 network, and obtain a feature pyramid with several scales, wherein the several scales of the feature pyramid correspond to the local regions of different scales of the original slice, that is, generate several multi-scale local regions; use the feature pyramid network to screen the generated several multi-scale local regions, and screen out several candidate local regions; input the several candidate local regions into the second ResNet-50 network to extract local features and calculate the confidence of classification. Select several local regions with the highest confidence, integrate the corresponding local region features with the global features of the slice, obtain a feature vector, and predict the classification result of any slice.

[0049] like Figure 5 As shown, in this embodiment, the composition of the feature integration deep learning network is as follows:

[0050] (1) Slice S n The size of the slice is normalized to 3×448×448, and the ResNet-50 network is used to extract the global features of the slice.

[0051] (2) Perform 1×1 and 3×3 convolutions on the output of the last layer conv5_x of the first ResNet-50 network to obtain feature pyramids of sizes {14×14, 7×7, 4×4}, which correspond to local regions of sizes {48×48, 96×96, 192×192} in the original slices. At these three scales, three ratios {1:1, 2:3, 3:2} are used to generate A multi-scale local regions.

[0052] (3) Use the feature pyramid network to score the information content I of these multi-scale local regions, use non-maximum suppression (NMS) to screen these A local regions, select the M candidate local regions with the highest information content I, and normalize their sizes to 3×224×224.

[0053] (4) The M candidate local regions {R 1 ,R 2 ,…,R M} is sent to the second ResNet-50 network to extract local features and calculate the classification confidence {C(R 1 ),C(R 2 ),…,C(R M )}. Select the K local regions with the highest confidence C {R 1,R 2 ,…,R K}, integrate its features with the global features of the slice to obtain a 1×2048 feature vector f(S n ,R 1 ,R 2 ,…,R K ), and predict the slice S n The classification result T n .

[0054] The training goal of the feature integration deep learning network is that the candidate local regions with higher information content I generated by the feature pyramid network should have higher confidence C during classification, which can be described as follows.

[0055] For {R 1 ,R 2 ,…,R M}∈S n

[0056] If C(R 1 )>C(R 2 )>…>C(R M ), then I(R 1 )>I(R 2 )>…>I(R M )

[0057] To achieve this training goal, the following training strategies are adopted:

[0058] (1) The candidate local regions generated by the feature pyramid network are sorted in descending order according to the confidence C. Use the binary (Ri, R j ) represents a set of candidate local regions, where the values ​​of i and j satisfy C(R i ) and C(R j ) satisfies the value of C(R i )≥C(R j ). The loss function shown in the following formula is used to constrain the selection process of the candidate region, where the function f(x) = max{1-x,0}

[0059] remember

[0060] Among them, M is a hyperparameter, is a positive natural number, x is the independent variable of the function f(x), R i is the i-th candidate local region, R j is the jth candidate local region, I(R i ) is the information score of the i-th candidate local region, I(R j ) is the information score of the jth candidate local region.

[0061] (2) Based on the confidence C output by the second ResNet-50 network, the entire slice S is calculated using the following formula n And the M local regions R selected by the feature pyramid network i The cross entropy loss of -log C(S n ) represents the entire slice S n The cross entropy loss is Represents the sum of the cross entropy losses in the local area.

[0062]

[0063] (3) Record slice S n And the K local regions R with the highest confidence i The feature vector obtained by concatenating the features of n ,R 1 ,R 2 ,…,R K ), the function A(·) represents the feature integration and classification prediction module of the network, then the classification result T n =A(S n ,R 1 ,R 2 ,…,R K ), use the following formula to calculate the classification result T n The cross entropy loss.

[0064] L A = -log A(S n ,R 1 ,R 2 ,…,R K )

[0065] (4) The loss of the feature integration deep learning network can be expressed as the sum of the above three losses. The joint loss is calculated using the following formula, and back propagation is performed to complete the update of the parameters of the feature integration deep learning network.

[0066] L joint =L I +λ·L C +μ·L A

[0067] Among them, λ and μ are constants, L joint The loss of the deep learning network is integrated for the features.

[0068] In this embodiment, the empirical values ​​of the hyperparameters in the above process are M=6, K=3, λ=μ=1.

[0069] In summary, during the training process of the feature integration deep learning network, the loss function Ljoint For L joint =L I +λ·L C +μ·L A (Formula 1), where

[0070]

[0071]

[0072] L A = -log A(S n ,R 1 ,R 2 ,…,R K ) (Formula 4)

[0073] Among them, C is the confidence of the classification, I is the information score given by the feature pyramid network, i is the first control variable traversing the candidate local area, j is the second control variable traversing the candidate local area, M is the hyperparameter, R i is the i-th candidate local region, R j is the jth candidate local region, I(R i ) is the information score of the i-th candidate local region, I(R j ) is the information score of the jth candidate local region; S n is a slice, M is the number of candidate local regions screened by the feature pyramid network, -log C(S n ) represents the entire slice S n The cross entropy loss is Represents the sum of the cross entropy losses of the local area; K is the local area R with the highest confidence i The function A(·) represents the feature integration and classification prediction module of the network; λ and μ are both constants.

[0074] like Figure 6 As shown, the slices of the original image where the airport area exists are predicted according to the classification results, and a salient map of the original image where the airport area exists is generated using a Gaussian weighting method, where S n ′ is slice S n The corresponding saliency map, then S n ' can be calculated by the following formula:

[0075]

[0076] Where T n For slice S n The classification result is, if it belongs to an airport, then T n =1, if not, then T n =0; αn For slice S n The corresponding multi-scale factor is used to restore the slice to the size before downsampling; x and y are the saliency map S n ′, the horizontal and vertical coordinates, the variance σ n S output by the deep learning network integrated with the features n The confidence level C n Positively correlated; x n and n Slice S n The horizontal and vertical coordinates of the pixel point in the original image; S A For the original image, calculate each slice S n The salient map S n ′, we can sum them up to get the original image S using the following formula A The salient map S A ′, N is the number of slices segmented from the original image.

[0077]

[0078] In this embodiment, the feature integration deep learning network is used to extract the airport area in the image, and the most discriminative overall and local features in the scene are adaptively selected for fusion through the attention mechanism, and the saliency map of the airport area is generated using Gaussian weighting to simulate human visual characteristics.

[0079] like Figure 3 As shown, in this embodiment, the backbone network is composed of DLA34 and DCN.

[0080] The original image is input into the backbone network, the sample annotation frame is corrected according to geometric prior knowledge, features are extracted and aircraft targets are detected, wherein:

[0081] The original image is input, and after the backbone network extracts features, three sets of feature maps are generated: (1) a single-channel center point heat map, which indicates the location of the center point of the aircraft target on the original image; (2) a dual-channel offset map, which indicates the rounding error when the points on the feature map are mapped back to the original image. The two channels represent the horizontal and vertical offsets respectively; (3) a dual-channel width-height map, which indicates the size information of the aircraft target corresponding to the points on the feature map. The two channels represent the width and height of the target respectively. After the three sets of feature maps are processed by the interpretation network, the location, offset, width and height information of the target on the original image are obtained. Then, a horizontal rectangular box is used to encircle the area where the aircraft target is located as the result of aircraft target detection.

[0082] Furthermore, the horizontal rectangular frame is corrected using a sample annotation frame correction method based on geometric prior knowledge, including:

[0083] To identify aircraft targets, it is necessary to train an aircraft target detection model during the training phase and perform data augmentation on the input images. When performing operations involving coordinate transformation such as scaling and cropping the images, there may be a phenomenon where samples in the image are too close to the edge and are truncated. In this embodiment, with the help of geometric prior knowledge, a square is used to approximately complete the truncated annotation box and delete the annotation boxes with too low quality.

[0084] As Figure 7 shown, there are three positional relationships between the annotation box in the data-augmented image and the four sides of the original image: (1) Two or more sides overlap, that is, the annotation box is at the four corners of the image; (2) One side overlaps, that is, the annotation box is close to one side of the image; (3) There is no overlap, that is, the annotation box is inside the image. The first two categories need to be processed, while the third category does not need to be processed.

[0085] For the positional relationship (1), if max(w, h) / min(w, h) ≤ 1.5 and max(w / W, h / H) ≥ 0.4, that is, the ratio of the long side to the short side of the annotation box is not greater than 1.5 and the length or width is not less than 0.4 of the corresponding dimension of the image, then the annotation box is close to a square and occupies a large area of the image, and it should be retained; if this condition is not met, then this annotation box is deleted, where w is the length of the long side of the annotation box and h is the length of the short side of the annotation box. Figure 7 (1) The upper left annotation box meets the condition and is retained, while other annotation boxes are deleted due to non-compliant aspect ratios or too small occupied areas.

[0086] For the positional relationship (2), there are two cases. Denote the data-augmented image as and a certain sample annotation box in it as S, and the coordinates of the upper left and lower right points are (x 1 , y 1 ) and (x 2 , y 2 ), where, is the set of real numbers, C is the number of channels of the original image, H is the height of the original image, and W is the width of the original image.

[0087] Case 1: If one side of the sample annotation box S overlaps with the height direction of I, that is, x 1 = 0 or x 2 = W - 1:

[0088] If w < h, use the following formula to calculate the center point p of the corrected annotation box :

[0089]

[0090] If p x∈ [0, W - 1], that is, if p is above I, then replace S with S′, whose width and height are both h, as Figure 7 shown by the annotation box ① in (2), the dotted line is the corrected result S′ of S; if this condition is not met, then p is outside I, and S should be deleted, as Figure 7 shown by the annotation box ② in (2)

[0091] If w ≥ h, then directly retain S, as shown by the annotation box ③.

[0092] Case 2: If one side of S overlaps with the width direction of I, that is, y 1 = 0 or y 2 = H - 1:

[0093] If h < w, use the following formula to calculate the center point p of.

[0094]

[0095] If p y ∈ [0, H - 1], that is, if p is above I, then replace S with S′, whose width and height are both w; if this condition is not met, then p is outside I, and S should be deleted.

[0096] If h ≥ w, then directly retain S.

[0097] In this embodiment, an improved anchor - free method is used to implement aircraft target detection. With the help of a backbone network for multi - scale feature depth fusion, and according to geometric prior knowledge, the sample annotation box is corrected, so that the network can learn the most discriminative multi - scale features of the aircraft target.

[0098] As Figure 4 shown, to scan the original image, extract the airport area and detect the aircraft target, it can be done in a serial or parallel manner, scanning the entire image, and completing the airport area extraction and aircraft target area detection sequentially or simultaneously.

[0099] As Figure 8 shown, the step S102: converting the saliency map of the airport area into a convex polygon of the airport area includes:

[0100] Step S1021: Perform adaptive binarization processing on the saliency map S of the airport area to obtain the binary map S bin ;

[0101] Step S1022: Perform contour extraction on the binary map S bin to obtain a polygon P composed of corner points;

[0102] Step S1023: Perform convex hull detection on the polygon P to obtain the convex polygon P′.

[0103] This conversion method can convert the continuous result composed of pixel points into a discrete result composed of polygon vertices, reducing the amount and complexity of calculations, so that the area of ​​the extracted airport area is enlarged and the edges become smooth.

[0104] The step S103: based on the convex polygon of the airport area, the original image is divided into sub-areas of three levels of attention, and the confidence of the detected aircraft target is weighted using a double threshold weighting method, wherein:

[0105] like Fig. 9 As shown, according to the first distance determination threshold d Th1 The convex polygon P′ of the airport area is expanded into three types of areas: inside the airport, around the airport, and outside the airport. The center point of a target obj is denoted as p obj , (1) If p obj ∈P′, then obj is inside the airport; (2) if the distance d min (p obj ,P′)<=d Th1 , then obj is around the airport; (3) If the distance is d min (p obj ,P′)>d Th1 , then obj is outside the airport. In addition, select the confidence level higher than s Th The target is the high confidence target group obj h , target group obj h The set of the center points of each target in is denoted as P h ={p h1 ,p h2 ,…,p hn}, where s Th is the high confidence threshold.

[0106] For each detected target obj, according to its center point p obj With the airport area P′ and the high confidence target group obj h The position relationship is divided into four cases and its confidence is weighted:

[0107] If p obj ∈P′, the confidence remains unchanged;

[0108] If d min (p obj ,P′)>d Th1 And d min (p obj ,P h )<=d Th2 , multiply its confidence by the weighting coefficient λ 2 ;

[0109] If dmin (p obj ,P′)<=d Th1 And d min (p obj ,P h )>d Th2 , multiply its confidence by the weighting coefficient λ 2 ;

[0110] If d min (p obj ,P′)>d Th1 And d min (p obj ,P h )>d Th2 , multiply its confidence by the weighting coefficient λ 1 ;

[0111] d Th2 is the second distance determination threshold.

[0112] in:

[0113] The first distance determination threshold d Th1 Depends on the average of the percentage of real airport areas in each image in the dataset If the length of the image diagonal is l, we can use To calculate, It can be obtained by statistically analyzing the data set to be inspected (when there are airport area annotations) or making a rough estimate (when there are no airport area annotations).

[0114] The second distance determination threshold d Th2 It is determined by the image slice size when detecting the aircraft target. If the slice width is l′, d Th2 =l′.

[0115] Weighting coefficient λ 1 ∈(0,1), determines the penalty for the confidence of targets far from the airport. In this embodiment, we take λ 1 =0.4.

[0116] Weighting coefficient λ 2 ∈(0,1) and λ 2 >λ 1 (The penalty is weaker than λ 1 ), in this embodiment, take λ 2 =0.75.

[0117] High confidence threshold Th , that is, the confidence lower bound of the high confidence target group. In this embodiment, s Th=0.8. In this embodiment, the extracted airport saliency map guides the detection of aircraft targets, and the entire image is divided into three sub-regions of different attention levels according to the airport extraction results. The airport region convex polygon corresponds to the innermost sub-region, and two sub-regions are expanded outward. The confidence of the detection result is post-processed using a double threshold weighting method.

[0118] The embodiment of the present invention further provides a regional saliency guided optical remote sensing image aircraft detection device, such as Fig.10 As shown, the device comprises:

[0119] Extraction module: configured to scan the original image, extract the airport area and detect the aircraft target;

[0120] The airport area extraction comprises: slicing the original image according to multi-scale factors, classifying the slices based on a trained feature integration deep learning network; predicting slices in which the airport area exists in the original image according to the classification results, generating a salient map of the airport area in the original image in a Gaussian weighted manner; processing the salient map as an extraction result of the airport area;

[0121] The aircraft target detection includes: inputting the original image into the backbone network, correcting the sample annotation frame according to geometric prior knowledge, extracting features and detecting the aircraft target;

[0122] A conversion module: configured to convert the saliency map of the airport area into a convex polygon of the airport area;

[0123] Confidence weighted processing module: configured to divide the original image into sub-areas with three levels of attention based on the convex polygon of the airport area, and use a double threshold weighted method to perform weighted processing on the confidence of the detected aircraft target.

[0124] The embodiment of the present invention further provides a regional saliency guided optical remote sensing image aircraft detection system, comprising:

[0125] A processor, which is used to execute multiple instructions;

[0126] A memory for storing a plurality of instructions;

[0127] The plurality of instructions are used to be stored by the memory, and loaded and executed by the processor for executing the aforementioned regional saliency-guided optical remote sensing image aircraft detection method.

[0128] The embodiment of the present invention further provides a computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the above-mentioned regional saliency-guided optical remote sensing image aircraft detection method.

[0129] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0130] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0131] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0132] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0133] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a physical server, or a network cloud server, etc., and the Ubuntu operating system must be installed) to perform some steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.

[0134] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention still falls within the scope of the technical solution of the present invention.

Claims

1. A method for aircraft detection in optical remote sensing images guided by regional saliency, characterized in that: The method comprises the following steps: Step S101: Scan the original image, extract the airport area and detect the aircraft target; Extracting the airport area includes: slicing the original image according to multi-scale factors, and classifying the slices based on a trained feature integration deep learning network; predicting the slices in which the airport area exists in the original image according to the classification results, and generating a salient map of the airport area in the original image in a Gaussian weighted manner; processing the salient map as the extraction result of the airport area; Detecting aircraft targets, including: inputting the original image into the backbone network, correcting the sample annotation box based on geometric prior knowledge, extracting features and detecting aircraft targets; Step S102: converting the salient map of the airport area into a convex polygon of the airport area; Step S103: Based on the convex polygon of the airport area, the original image is divided into sub-areas with three levels of attention, and the confidence of the detected aircraft target is weighted using a double threshold weighting method; The step S103 includes: According to the first distance determination threshold d Th1 The convex polygon P′ of the airport area is expanded into three types of areas: inside the airport, around the airport, and outside the airport; the center point of a target obj is denoted as p obj , (1) If p obj ∈P′, then obj is inside the airport; (2) if the distance d min (p obj ,P′)<=d Th1 , then obj is around the airport; (3) If the distance is d min (p obj ,P′)>d Th1 , then obj is outside the airport; Select confidence level higher than s Th The target is the high confidence target group obj h , high confidence target group obj h The set of center points of each target in is denoted as P h ={p h1 ,p h2 ,…,p hn }, where s Th is the high confidence threshold; For each detected target obj, according to its center point p obj The convex polygon P′ of the airport area and the high confidence target group obj h The position relationship is divided into four cases and its confidence is weighted: If p obj ∈P′, the confidence remains unchanged; If d min (p obj ,P′)>d Th1 And d min (p obj ,P h )<=d Th2 , multiply its confidence by the weighting coefficient λ2; If d min (p obj ,P′)<=d Th1 And d min (p obj ,P h )>d Th2 , multiply its confidence by the weighting coefficient λ2; If d min (p obj ,P′)>d Th1 And d min (p obj ,P h )>d Th2 , multiply its confidence by the weighting coefficient λ1; d Th2 is the second distance determination threshold, and the weighting coefficients λ1 and λ2 are both empirical values.

2. The method for detecting aircraft in optical remote sensing images guided by regional saliency as claimed in claim 1, characterized in that: The airport area extraction includes: slicing the original image according to multi-scale factors, and classifying the slices based on a trained feature integration deep learning network; predicting the slices in which the airport area exists in the original image according to the classification results, and generating a salient map of the airport area in the original image in a Gaussian weighted manner; processing the salient map as the extraction result of the airport area, including: Step S201: The original image S A According to the multi-scaling factors α1,…,α L-1 ,α L Split and downsample with overlap to get slices S1, S2, ..., S of the same size N ; Step S202: Slice {S1, S2, ..., S N }∈S A The feature integration deep learning network predicts whether it belongs to the airport area. If it does, the slice classification result T n Set to 1, otherwise T n Set to 0; Step S203: According to the classification result T n , predicting the area where the airport exists in the original image, and generating a saliency map of the area where the airport exists in the original image by Gaussian weighting; Step S204: using a threshold segmentation method to generate a binary image for the salient image of the airport area, and the binary image is used as the extraction result of the airport area.

3. The method for detecting aircraft in optical remote sensing images guided by regional saliency as claimed in claim 2, characterized in that: During the training process of the feature integration deep learning network, the loss function L joint For L joint =L I +λ·L C +μ·L A (Formula 1), where L A = -logA(S n ,R1,R2,…,R K ) (Formula 4) Among them, C is the confidence of the classification, I is the information score given by the feature pyramid network, i is the first control variable traversing the candidate local area, j is the second control variable traversing the candidate local area, M is the hyperparameter, R i is the i-th candidate local region, R j is the jth candidate local region, I(R i ) is the information score of the i-th candidate local region, I(R j ) is the information score of the jth candidate local region; S n is a slice, M is the number of candidate local regions screened by the feature pyramid network, -logC(S n ) represents the entire slice S n The cross entropy loss is Represents the sum of the cross entropy losses of the local area; K is the local area R with the highest confidence i The function A(·) represents the feature integration and classification prediction module of the network; λ and μ are both constants.

4. The method for detecting aircraft in optical remote sensing images guided by regional saliency as claimed in claim 3, characterized in that: The slices of the original image where the airport area exists are predicted according to the classification results, and a salient map of the original image where the airport area exists is generated in a Gaussian weighted manner, wherein: Remember S n ′ is slice S n The corresponding saliency map, then S n ' can be calculated by the following formula: Where T n For slice S n The classification result is, if it belongs to an airport, then T n =1, if not, then T n =0; α n For slice S n The corresponding multi-scale factor is used to restore the slice to the size before downsampling; x and y are the saliency map S n ′, the horizontal and vertical coordinates, the variance σ n S output by the deep learning network integrated with the features n The confidence level C n Positively correlated; x n and n Slice S n The horizontal and vertical coordinates of the pixel point in the original image; S A is the original image; Calculate each slice S n The salient map S n ′, calculate The original image S can be obtained A The salient map S A ′, N is the number of slices segmented from the original image.

5. The method for detecting aircraft in optical remote sensing images guided by regional saliency as claimed in claim 4, characterized in that: The original image is input into the backbone network, the sample annotation frame is corrected according to the geometric prior knowledge, the features are extracted and the aircraft target is detected, wherein: The original image is input, and after the backbone network extracts features, three sets of feature maps are generated: a single-channel center point heat map, which indicates the location where the center point of the aircraft target may exist on the original image; a dual-channel offset map, which indicates the rounding error of the points on the feature map mapped back to the original image, and the two channels represent the lateral and longitudinal offsets respectively; a dual-channel width-height map, which indicates the size information of the aircraft target corresponding to the points on the feature map, and the two channels represent the width and height of the target respectively.

6. The method for aircraft detection in optical remote sensing images guided by regional saliency as claimed in claim 5, characterized in that: The sample annotation box is corrected based on geometric prior knowledge, including: The input image is data augmented. The annotation box in the data augmented image has three positional relationships with the four edges of the original image: (1) two or more edges overlap, that is, the annotation box is at the four corners of the image; (2) one edge overlaps, that is, the annotation box is close to one side of the image; (3) there is no overlap, that is, the annotation box is inside the image; For position relationship (1), if the ratio of the long side to the short side of the annotation box is not greater than 1.5 and the ratio of the length or width to the corresponding dimension of the original image is not less than 0.4, then the annotation box should be retained; if the condition is not met, then the annotation box is deleted, w is the length of the long side of the annotation box, and h is the length of the short side of the annotation box; For the position relationship (2), the image after data enhancement is recorded as One of the sample annotation boxes is The coordinates of the upper left and lower right corners of S are (x1, y1) and (x2, y2), respectively, where: is a real number set, C is the number of channels of the original image, H is the height of the original image, and W is the width of the original image. There are two cases for processing: Case 1: If one edge of the sample annotation box S overlaps with the height direction of I, If w<h, use the following formula to calculate the corrected annotation box The center point p: p=(p x ,p y ), p y =(y1+y2) / 2 If p is satisfied x ∈[0,W-1], then use S′ to replace S, whose width and height are both h; if this condition is not met, delete S; If w ≥ h, then directly retain S; Case 2: If one edge of S overlaps with the width direction of I, If h<w, use the following formula to calculate The center point p: p=(p x ,p y ),p x =(x1+x2) / 2, If p is satisfied y ∈[0,H-1], then use S′ to replace S, whose width and height are both w; if this condition is not met, delete S; If h≥w, then directly retain S; For position relationship (3), no processing is required.

7. A regional saliency guided optical remote sensing image aircraft detection device, used to execute the method according to any one of claims 1 to 6, characterized in that: The device comprises: Extraction module: configured to scan the original image, extract the airport area and detect the aircraft target; Extracting the airport area includes: slicing the original image according to multi-scale factors, and classifying the slices based on a trained feature integration deep learning network; predicting the slices in which the airport area exists in the original image according to the classification results, and generating a salient map of the airport area in the original image in a Gaussian weighted manner; processing the salient map as the extraction result of the airport area; Detecting aircraft targets, including: inputting the original image into the backbone network, correcting the sample annotation box based on geometric prior knowledge, extracting features and detecting aircraft targets; A conversion module: configured to convert the saliency map of the airport area into a convex polygon of the airport area; Confidence weighted processing module: configured to divide the original image into sub-areas with three levels of attention based on the convex polygon of the airport area, and use a double threshold weighted method to perform weighted processing on the confidence of the detected aircraft target.

8. A regional saliency guided optical remote sensing image aircraft detection system, comprising: A processor, which is used to execute multiple instructions; A memory for storing a plurality of instructions; The plurality of instructions are used to be stored by the memory, and loaded and executed by the processor according to the regional significance guided optical remote sensing image aircraft detection method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the regional significance-guided optical remote sensing image aircraft detection method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • A remote sensing image scene classification method based on the fusion of depth features and saliency features

    CN109165682A

  • Hierarchical identification method for remote sensing image aircraft targets

    CN109614936A