Variable-Scale Remote Sensing Image Target Detection Method Based on Ternary Feature Fusion
Through the variable-scale remote sensing image object detection method based on ternary feature fusion, the problem of insufficient detection accuracy of small and medium-sized objects in the prior art is solved, more accurate multi-scale object detection is achieved, and the semantic representation ability of deep features is enhanced.
Patent Information
- Application Number
- CN202211347000.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-10-31
AI Technical Summary
The existing remote sensing object detection methods still have shortcomings in the detection accuracy of small objects, especially in complex backgrounds, which are difficult to accurately identify small objects, and the receptive field of the low-resolution feature map does not match the size of the small object, resulting in poor network generalization performance.
A variable-scale remote sensing image object detection method based on ternary feature fusion is adopted. By constructing a variable-scale remote sensing image object detection model, a ternary feature fusion network is used to integrate feature information of different scales, and the semantic representation ability of deep features is enhanced through the feature self-expanding network, and finally the prediction box marked with the target position is output through the detection head network.
It improves the detection accuracy of small targets, enhances the detection ability of multi-scale targets, reduces background response, retains key local details, and adapts to the problem of large-scale scale changes in remote sensing images.
Smart Images

Figure CN115641484B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image target recognition, and particularly relates to a variable-scale remote sensing image target detection method based on ternary feature fusion. Background Technique
[0002] Remote sensing is a high-tech observation technology developed on the basis of simulating the human visual system, and it has extremely broad application prospects in both military and civilian fields. With the development of science and technology, in recent years, the maturity of optical remote sensing image technology has also promoted the further development of remote sensing target detection technology. Remote sensing target detection is to predict a series of bounding boxes with class labels from complex aerial images, which is a basic problem in optical remote sensing image understanding and has important application value. At present, remote sensing target detection has been widely applied in fields such as urban environmental monitoring, land use planning, forest fire monitoring, and traffic flow management. However, it is still a challenge to quickly and effectively utilize remote sensing image information to achieve more accurate multi-scale target detection. The missed detection of small objects and the false detection of non-target objects are problems that need to be solved urgently. To better complete the recognition task, a multi-scale target detection method that can more accurately identify small objects in complex backgrounds is needed.
[0003] Currently, most of the state-of-the-art remote sensing target detection methods are based on deep learning technology. Compared with traditional handcrafted feature-based methods, deep learning-based target detection methods can learn both low-level detailed features and high-level semantic features of images simultaneously. Therefore, the extracted image features are often more representative.
[0004] Deep learning-based target detection algorithms can be divided into two categories: single-stage and two-stage according to different detection processes. Generally, two-stage networks have high accuracy but low efficiency, and the R-CNN series is a classic two-stage target detection network. In contrast, single-stage networks have faster inference speeds, and SSD, YOLO, and RetinaNet are typical single-stage networks. The YOLOv3 algorithm mainly draws on the design ideas of residuals and multi-scales, and uses a fully convolutional network as the backbone network to extract three different-scale feature maps, and their theoretical receptive fields have different semantic information granularities. In addition, the YOLOv3 adopts an anchor box mechanism to preset three groups of prior boxes for the three-scale feature maps respectively, which are used for the prediction of large, medium, and small targets.
[0005] Compared with natural images, the target scales in remote sensing object detection vary greatly. At the same time, there are a larger number of small and medium-sized targets in remote sensing images. The most common and effective method to promote small object detection is to use high-resolution images or feature maps. However, this often brings a huge computational cost. Although remote sensing object detection has made great progress in recent years, the detection accuracy of small objects still remains at a low level. On the one hand, the features highlighting small objects disappear due to the downsampling operation of the convolutional neural network backbone; on the other hand, the theoretical receptive field on the low-resolution feature map may not match the actual size of small objects, which makes the network generalization performance not ideal. Thus, when performing object recognition, there is an offset between the target position and the true position. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention provides a variable-scale remote sensing image object detection method based on ternary feature fusion. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0007] The present invention provides a variable-scale remote sensing image object detection method based on ternary feature fusion, including:
[0008] Step 1: Obtain a dataset composed of remote sensing images and the remote sensing image to be measured;
[0009] Among them, the dataset is divided into a training set and a test set, and each remote sensing image in the dataset contains at least one anchor box for annotating the target position;
[0010] Step 2: Generate multiple prior boxes for predicting the position where the target is located according to each anchor box;
[0011] Step 3: Obtain a variable-scale remote sensing image object detection model constructed based on ternary feature fusion;
[0012] Step 4: Use each remote sensing image in the dataset as a training sample, and iteratively train the variable-scale remote sensing image object detection model based on all training samples and the prior boxes to obtain a trained variable-scale remote sensing image object detection model;
[0013] Step 5: Use the trained variable-scale remote sensing image object detection model to detect the remote sensing image to be measured, and obtain a prediction box annotating the specific position of the target;
[0014] Step 6: Decode the prediction box to obtain the true coordinate position of the target.
[0015] Advantages of the present invention:
[0016] The present invention has the following advantages compared with the prior art:
[0017] 1. Since the present invention provides a variable-scale remote sensing image target detection method based on ternary feature fusion, constructs a variable-scale remote sensing image target detection model based on ternary features, and uses the variable-scale remote sensing image target detection model to detect the targets in remote sensing images, it helps to retain more key local detail information during the network downsampling process and suppress the response of the background. The present invention can solve the problem of prominent redundant information in the shallow feature map and weakening of small targets, and improve the detection accuracy of small targets.
[0018] 2. Since the present invention uses a feature self-expansion network to process the feature map, it can more flexibly handle the large-scale scale change problem in remote sensing images, adaptively expand the receptive field, enhance the semantic information representation ability of the deep feature map while not losing detail features, and improve the detection accuracy of multi-scale targets.
[0019] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a schematic flow chart of a variable-scale remote sensing image target detection method based on ternary feature fusion proposed by the present invention;
[0021] Figure 2 is the overall structure diagram of the variable-scale remote sensing image target detection model proposed by the present invention;
[0022] Figure 3 is a schematic diagram of the ternary feature fusion network of the present invention;
[0023] Figure 4 is a schematic diagram of the local feature saliency enhancement network of the present invention;
[0024] Figure 5 is a schematic diagram of the feature self-expansion network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The following further describes the present invention in detail with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0026] Embodiment 1
[0027] As Figure 1 shown, a variable-scale remote sensing image target detection method based on ternary feature fusion provided by the present invention includes:
[0028] Step 1: Obtain a dataset composed of remote sensing images and the remote sensing image to be measured;
[0029] Among them, the dataset is divided into a training set and a test set, and each remote sensing image in the dataset contains at least one anchor box for annotating the target position;
[0030] It should be noted that: the dataset and the remote sensing image to be measured are respectively preprocessed so that the remote sensing images in the dataset and the remote sensing image to be measured are both of a fixed size. The preprocessing can adopt operations such as scaling, random rotation, and mirror flipping.
[0031] Step 2: Generate multiple prior boxes at the position where the predicted target is located according to each anchor box;
[0032] Step 3: Obtain a variable-scale remote sensing image target detection model constructed based on ternary features;
[0033] Reference Figure 2 , the variable-scale remote sensing image target detection model of the present invention includes: a feature extraction backbone network, a ternary feature fusion network, a feature self-expansion network, and a detection head network connected in sequence;
[0034] The feature extraction backbone network is composed of multiple convolutional layers. As Figure 3 shown, the ternary feature fusion network includes an upsampling network, a local feature saliency enhancement network, and a feature fusion network.
[0035] The feature extraction backbone network is used to perform multiple downsampling operations to extract multiple-scale feature images from the input image and input them into the upsampling network, the local feature saliency enhancement network, and the feature fusion network; the upsampling network is used to transfer deep semantic information to the shallow network through upsampling operations on the input scale feature image and input the deep semantic information into the feature fusion network; the local feature saliency enhancement network is used to transfer shallow key local detail information to the deep network through downsampling operations on the input scale feature image and input the shallow key local detail information into the feature fusion network; the feature fusion network is used to perform a ternary splicing operation on the deep semantic information, the shallow key local detail information, and the scale feature input by the ternary feature fusion network to obtain a fusion feature map retaining local key features and deep semantic features, and input the fusion feature map into the feature self-expansion network respectively; the feature self-expansion network is used to generate an expansion coefficient according to the category bias and spatial bias of the fusion feature map, and fine-tune the fusion feature map according to the expansion coefficient to obtain discriminative features, and input the discriminative features into the detection head network; the detection head network is used to set multiple prior boxes for each pixel in the feature map corresponding to each discriminative feature, calculate the relevant parameters corresponding to each prior box, generate predicted values for the set prior boxes according to the relevant parameters, correct the corresponding prior boxes according to the predicted values to obtain multiple predicted boxes, retain the predicted boxes with the confidence ranking within the preset number, and use the NSM algorithm to filter out the predicted boxes with the overlap degree exceeding the overlap degree threshold to obtain the predicted boxes at the position where the predicted target is located.
[0036] Step 4: Use each remote sensing image in the dataset as a training sample, and iteratively train the variable-scale remote sensing image target detection model based on all training samples and the prior boxes to obtain a trained variable-scale remote sensing image target detection model;
[0037] Step 5: Use the trained variable-scale remote sensing image target detection model to detect the to-be-detected remote sensing image to obtain a prediction box marking the specific position of the target;
[0038] Step 6: Decode the prediction box to obtain the true coordinate position of the target.
[0039] The present invention provides a variable-scale remote sensing image target detection method based on ternary feature fusion. After training the variable-scale remote sensing image target detection model is completed, for the to-be-detected remote sensing image, a multi-scale feature map is extracted from the to-be-detected remote sensing image through the backbone network in the trained variable-scale remote sensing image target detection model. Then, a ternary feature fusion network is used to integrate feature information from different scales. Then, a feature self-expansion network is used to adjust the scale of the feature map to enhance the semantic representation ability of deep features. Finally, a prediction box marking the target position is output through the detection head network, so as to decode and obtain the specific position of the target in the to-be-detected remote sensing image. The present invention can improve the accuracy of predicting the specific position of the target in the remote sensing image.
[0040] Embodiment 2
[0041] As an optional embodiment of the present invention, the step 2 includes:
[0042] Step 2a: Randomly select a preset number of anchor boxes from all the anchor boxes as cluster centers;
[0043] Step 2b: Calculate the distance between each anchor box and other anchor boxes;
[0044] Step 2c: Determine the nearest cluster center of each anchor box and assign the anchor box to the cluster;
[0045] Step 2d: Update the cluster centers again according to the distance between the anchor boxes;
[0046] Step 2e: Use the updated cluster centers as the prior boxes for the position where the predicted target is located.
[0047] In the present invention, the image input size is set to 416×416. The width and height of the prior boxes are obtained by the K-means clustering algorithm, and nine prior boxes are set in the present invention. The size setting of the prior boxes follows the following rules:
[0048] (1) Randomly select 9 from all the annotated anchor boxes (bboxes) in the dataset as the centers of the clusters (c1, c2, …, c9).
[0049] (2) Calculate the distance between each bboxes and each cluster center. For the distance between samples, the following formula is used for quantification:
[0050] d(bboxs i , c j ) = 1 - IOU(bboxs i , c j )
[0051] Among them, IOU represents the intersection over union of two elements.
[0052] (3) Calculate the distance between each bboxes and the nearest cluster center, and assign it to the cluster closest to it.
[0053] (4) Calculate the sample average distance of the bboxes in each cluster according to the following formula to regenerate the cluster center:
[0054]
[0055] n represents the total number of sample objects in the dataset, and d(bboxs i , c j ) is the IOU distance between the sample bboxes i and c j .
[0056] (5) Repeat steps (3) and (4) until each cluster center no longer changes, and the clustering is completed.
[0057] Example 3
[0058] As an optional embodiment of the present invention, the feature extraction backbone network performs 5 upsampling operations to obtain 5 scale feature images: f0, f1, f2, f3, and f4, and f1, f2, f3, and f4 are retained for subsequent processing.
[0059] For the scale feature map f s , input the scale feature map f s into the feature fusion network, input the scale feature map F s+1 into the upsampling network, and feed the feature map f s-1 to the local feature saliency enhancement network;
[0060] The upsampling network is used to upsample the scale feature map f s+1 to obtain the feature map f′ s+1 , and make f′ s+1 and the scale feature map f sWith consistent dimensions, the feature map f is obtained u ;
[0061] The local feature saliency enhancement network is used to s-1 downsample the scale feature map F s to obtain a feature map f with the same dimensions as the scale feature map f d ;
[0062] The feature fusion network is used to u perform a ternary concatenation operation on the feature map f d and the scale feature map f s to obtain a fused feature map p that retains local key features and deep semantic features s and input the fused feature map p s into the corresponding feature self-expansion network;
[0063] where s is a positive integer from 2 to 4.
[0064] Each layer of the scale feature map f s is respectively concatenated with f u and f d to perform a ternary concatenation operation. After processing, a feature map p that retains both local key features and deep semantic features is obtained s . The specific feature fusion method is calculated according to the following formula.
[0065]
[0066] Here, Convs(·) represents the channel downsampling operation, which is implemented as a 1×1 convolution with a downsampling ratio of 2. Upsample(·) represents the upsampling operation, which is implemented as nearest neighbor interpolation. LFSE(·) represents the local feature saliency enhancement module.
[0067] Example 4
[0068] As an optional example of the present invention, as Figure 4 shown, the local feature saliency enhancement network includes: a local enhancement network and a spatial re-scoring network;
[0069] The local enhancement network is used to
[0070] map the scale feature map f s-1 to different subspaces through two independent embedding functions to obtain feature maps H and G of the two subspaces;
[0071] Use a sliding window Φ with a fixed stride to divide the feature maps H and G into individual image patches and
[0072] It should be noted that the local feature saliency enhancement module of the present invention focuses on extracting local detailed features, which are mainly divided into two steps, namely the local enhancement module and the spatial re-scoring module. For the local enhancement module, taking the input feature map f s-1 as an example. First, it is fed into two independent embedding functions to map the features to different subspaces, and the feature maps with stronger representation ability are denoted as H and G. In this module, the two embedding functions are implemented as 1×1 convolutions with a stride of 1. Then, using a sliding window Φ with a size of 3×3 and a stride of 2, the feature maps H and G are further segmented into separate small-sized patches channel by channel. Thus, there is H m representing the m-th channel of H.
[0073] For the image patch perform a max pooling operation, and use the image features obtained after the max pooling operation as the query feature map;
[0074] perform a transpose operation on the image patch, and use the feature map obtained after the transpose operation as the feature map to be queried;
[0075] multiply the feature map to be queried with the query feature map to obtain the dependency matrix M of the pixel points in the local neighborhood and the high-response points;
[0076] calculate the attention weight matrix W according to the dependency matrix M;
[0077] use the attention weight matrix W as the dynamic kernel parameter, and perform local feature enhancement operation on the dynamic kernel parameter to obtain the output feature map of local enhancement;
[0078] restore the optimal data distribution of the output feature map of local enhancement through the BN layer to obtain the new feature map L;
[0079] transfer the new feature map L and the input scale feature map f s-1 to the spatial re-scoring network;
[0080] It should be noted that: take the obtained output feature perform a max pooling operation and use it as the query feature, perform a transpose operation and use it as the feature map to be queried. Multiply the two obtained feature maps to get the dependency matrix M of each pixel point in the local neighborhood and the high-response points. M is then used to calculate the attention weight matrix W in the following way, and W will participate in the local feature enhancement operation as the dynamic kernel parameter. The specific operation process can be expressed by the following formula.
[0081]
[0082]
[0083]
[0084] where represents the element in the \(i\)-th row and \(j\)-th column within the sliding window of the \(q\)-th channel, \(x\)-th row, and \(y\)-th column of \(M\). Similarly, and represent the elements within the sliding window of the \(q\)-th channel, \(x\)-th row, and \(y\)-th column of \(W\) and \(F\) respectively. \(K(\cdot)\) is a kernel function with a fixed parameter form, such as a Gaussian distribution, which participates in the operation as relative position encoding. For the output feature map of the local enhancement module, a BN layer is used to restore its optimal data distribution to obtain a new feature map \(L\). Subsequently, the input feature \(f\) of the local enhancement module s-1 and the new feature map \(L\) are passed to the spatial rescoring network.
[0085] The spatial rescoring network is used to
[0086] perform a channel-level max pooling operation on the new feature map \(L\) to obtain a feature map \(L'\);
[0087] perform a convolution operation on the scale feature map \(f\) s-1 to obtain a feature map \(F'\);
[0088] flatten the feature map \(L'\) and the feature map \(F'\) from two dimensions to one-dimensional vectors respectively, and multiply the transposes of the two one-dimensional vectors to obtain a correlation matrix \(A\) g ;
[0089] input the correlation matrix \(A\) g into the Softmax layer to obtain a spatial attention weight matrix \(A\) q ;
[0090] multiply the spatial attention weight matrix \(A\) q by the feature map \(L'\) to obtain a spatial attention mask for the new feature map \(L\), and obtain a mask \(mask\) according to this spatial attention mask l ;
[0091] According to the mask \(mask\) l and the new feature map \(L\), obtain a feature map \(f\) s with the same size as the scale feature map \(f\) d ;
[0092] input the feature map \(f\) d into the feature fusion network.
[0093] It should be noted that: The spatial rescoring module first performs a channel-level max pooling operation on \(L\) to obtain a feature map \(L'\), and for the input initial feature map \(f\) sAfter performing convolution operations using a convolutional layer with a kernel size of 7×7 and a stride of 3, the feature map F′ is obtained. For the two resulting 1-channel feature maps, after using the flatten operation to flatten the two feature maps from two dimensions into one-dimensional vectors, L′ is multiplied by the transpose of F′ to obtain the correlation matrix A g . Then A g is input into the Softmax layer to obtain the spatial attention weight matrix A q , and its process can be described as follows.
[0094]
[0095]
[0096] Among them represents the element in the q-th channel, the x-th row and y-th column of the correlation matrix A g . (A q ) xy represents the element in the q-th channel, the x-th row and y-th column of the spatial attention weight matrix A.
[0097] Multiply the obtained attention weight matrix A by the feature map L to obtain a spatial attention mask for the output feature L, and subtract it from a matrix filled with all 1s of the same size to obtain the mask mask l . Multiply mask l by L using the HadamardProduct operation, then perform the Flement-wise sum operation with L, and use the BN layer to restore its optimal data distribution. Finally, use a 1×1 convolution to map the features to the initial feature space, and the final feature map is denoted as f d .
[0098] Example Five
[0099] As an optional embodiment of the present invention, the variable-scale remote sensing image target detection model includes three parallel feature self-expansion networks, as Figure 5 shown. Each feature self-expansion network includes a channel content perception module, a weighted average network, and a spatial self-expansion pool network connected in sequence;
[0100] The feature fusion network is used to input the fused feature maps into the corresponding feature self-expansion networks respectively;
[0101] Each feature self-expansion network is used to generate an expansion coefficient according to the category bias and spatial bias of the input fused feature map, and fine-tune the fused feature map according to the expansion coefficient to obtain discriminative features, and input the discriminative features into the detection head network.
[0102] It should be noted that: the output feature maps p2, p3, and p4 of the ternary feature fusion network are respectively input into three parallel feature self-expansion network modules. Assuming there is a feature map input into the feature self-expansion network module, denoted as p s , first, the channel content perception module calculates the class response and spatial distribution characteristics of each channel to generate a set of expansion coefficients, and then the dilation spatial pooling fine-tunes the receptive field of the feature map according to the coefficients to generate more discriminative features.
[0103] Embodiment Six
[0104] As an optional embodiment of the present invention, the channel content perception module includes a parallel class bias perception module and a spatial semantic perception module;
[0105] The class bias perception module is used to
[0106] compress the spatial information of the fusion feature map p s into a set of feature vectors p;
[0107] According to the set of feature vectors p, calculate the class bias score of the fusion feature map p s and input the class bias score into the connected weighted average network;
[0108] The spatial semantic perception module is used to
[0109] downsample the fusion feature map p s to obtain the downsampled feature map s x ;
[0110] Compress the spatial information of the feature map s x into a set of feature vectors and activate them using the Relu function to obtain the spatial context scoring vector s;
[0111] Take the spatial context scoring vector s as the spatial bias score and input it into the connected weighted average network;
[0112] The weighted average network is used to
[0113] perform weighted averaging on the spatial bias score and the class bias score to obtain the vector ΔN;
[0114] Take the vector ΔN as the channel expansion coefficient and input it into the connected spatial self-expansion pooling network;
[0115] The spatial self-expansion pooling network is used to
[0116] fine-tune the fusion feature map by approximating the channel expansion coefficient through linear interpolation to obtain the discriminative feature map P out ;
[0117] Input the discriminative feature map P out into the detection head network.
[0118] It should be noted that: the channel content perception network is composed of two parallel branches, namely the category bias perception branch and the spatial semantic perception branch. For the category bias perception branch, first, global average pooling is used to compress the spatial information of the input feature p s into a set of feature vectors p, and this feature vector p represents a compact global information set. The calculation method of the elements in the nth channel of p is as follows.
[0119]
[0120] where H and W represent the height and width of the feature map. represents the element in the i-th row and j-th column of the nth channel of p s .
[0121] Then it is input into two fully connected layers with Relu activation functions to capture the dependencies between channels. The category bias score c = σ(W2δ(W1p)), where σ represents a Relu activation function, δ represents a LeakyRelu activation function, and W1 and W2 represent fully connected operations.
[0122] For the spatial semantic perception branch, first, two groups of grouped convolution blocks with a convolution kernel size of 3×3, a dilation rate of 2, and a stride of 1 are used for downsampling to obtain the feature s x , and then global average pooling is used to summarize the spatial information into a set of feature vectors, and the output is activated using the ReLU function to obtain a spatial context scoring vector s. The spatial context scoring vector s is used as the spatial bias score. The elements in the nth channel of s where σ represents the ReLU activation function, and f g (·) represents a 1×1 convolution.
[0123] Subsequently, weighted averaging is performed on the spatial bias score and the category bias score to obtain a set of vectors ΔN, ΔN = αc+(1 - α)s, and it is fed as the channel dilation coefficient to the dilated spatial pool, where α is a learnable parameter, and its initial value is defined as 0.5.
[0124] The vector N input to the dilated spatial pool is learned through convolution and fully connected layers and is usually not an integer. The present invention uses linear interpolation to approximate the dilation coefficient to alleviate the forced offset at all positions. Therefore, the final output can be calculated as follows.
[0125]
[0126] where Represent the output feature P out of the i-th channel. is the initial pooling kernel size of the i-th channel, set to 1. γ is a scaling factor, default set to 1. ψ(·) represents the max pooling operation, which has two parameters, namely the input feature and the pooling kernel size. represents the floor operation, represents the ceiling operation.
[0127] Embodiment VII
[0128] As an optional embodiment of the present invention, the detection head network is used to
[0129] take the prior box in step 2 as the prior box of the pixel unit, and for each pixel unit in the feature map corresponding to each discriminative feature, set 3 prior boxes for this pixel unit;
[0130] calculate the relevant parameters corresponding to the 3 prior boxes;
[0131] Among them, the relevant parameters include: the probability scores of each category, the center point coordinates, the width and height values, and the confidence scores;
[0132] generate the predicted values of the set prior boxes according to the relevant parameters;
[0133] correct the corresponding prior boxes according to the predicted values to obtain multiple predicted boxes.
[0134] It should be noted that: for the feature map output by the feature self-expansion network, assuming its size is w×h, there are w×h pixel units in total. Each pixel unit is preset with 3 groups of prior boxes, and the specific parameters of the prior boxes have been calculated in step 2. Each prior box corresponds to parameters such as the probability scores of each category, the center point coordinates, the width and height values, and the confidence scores. Therefore, each unit generates 3×(C + 4 + 1) predicted values, where C represents the number of target categories. The present invention uses a 1×1 convolution as the detection head network to directly perform classification and regression prediction, and the output feature dimension is set to WH(3×(C + 4 + 1)).
[0135] For the prediction results, first determine the prediction box category and its confidence value according to the category scores. Then filter out some prediction boxes with lower scores according to the confidence threshold, retain the 100 prediction boxes with the highest confidence values according to the confidence value ranking, and use the NMS algorithm to filter out the bounding boxes with large overlap. Here, GIOU is used to measure the coincidence degree between the prediction boxes, and its formula is shown as follows.
[0136]
[0137]
[0138] Among them, A represents the prediction box A, B represents the prediction box B, and C represents the area of the smallest closed area of the two boxes A and B. When the intersection over union is greater than 0.5, it is regarded as the same sample. Finally, the remaining prediction boxes after being processed by the NMS algorithm are the detection results, and then decoding is performed to obtain the true coordinate information.
[0139] Embodiment VIII
[0140] As an optional embodiment of the present invention, step 4 includes:
[0141] Taking each remote sensing image in the dataset as a training sample;
[0142] Based on all training samples and the prior boxes, iteratively training the variable-scale remote sensing image target detection model to adjust the internal parameters of the variable-scale remote sensing image target detection model in the direction of reducing the loss function until the maximum number of iterations is reached, and obtaining a trained variable-scale remote sensing image target detection model;
[0143] Among them, the loss function is composed of position error, confidence error, classification error, and attention error.
[0144] The loss function consists of four parts, defined as the weighted sum of position error (lbox), confidence error (lobj), classification error (lcls), and attention error (latt):
[0145] loss = lbox + lobj + lcls + latt
[0146]
[0147]
[0148]
[0149] In the position error, confidence error, and classification error, λ coord , λ class and λ noobj , λ obj are balance coefficients, default set to 1. S 2 represents the size of the predicted feature map. The predicted feature maps of three scales are 13×13, 26×26, and 52×52 pixels respectively. B represents the number of prior boxes, and classes represents the number of target categories labeled in the dataset. x i , y i , w i , h i represent the center point coordinates and width and height values of the prediction box at the i-th pixel in sequence. When the center point is at the j-th prior box at the i-th pixel and is a positive sample, The value of is set to 1, otherwise it is 0. When the jth prior box with the center point at pixel i is a negative sample, The value is set to 1, otherwise it is 0. and p i (c) represent the true value and predicted value of the category respectively. and c i Represent the true value and predicted value of confidence respectively, The value of is determined by whether the prior box at the i-th pixel position is responsible for containing the positive sample. If it is included, then otherwise,
[0150]
[0151] In the attention loss, λ att is the balance factor, which is set to 1 by default. c Represents mask l The attention prediction value of the c-th element of . For mask l The true value of the element at position c, if position c is included in the annotated anchor box, then otherwise,
[0152] During the training process, the prior box that matches the anchor box annotated with the training image is responsible for predicting the target. There are two matching principles. First, to ensure that each annotated true value box must have a corresponding prior box, any true value box in the image will be matched with the prior box with the largest intersection over union ratio. This prior box is called a positive sample. If a prior box has no corresponding true value box or the maximum IoU is less than the preset threshold, it is called a negative sample.
[0153] The trained model is used to detect the test data set to obtain the detection accuracy mAP of each category in the test data set. After reaching the accuracy, the trained variable-scale remote sensing image target detection model is used to detect the remote sensing image to be tested to obtain the prediction box of the specific location of the marked target. During the prediction, each network in the variable-scale remote sensing image target detection model can be set according to the following table:
[0154] Backbone network (Darknet53) parameter settings
[0155]
[0156] Parameter setting of local feature saliency enhancement network (taking f2 as an example)
[0157]
[0158]
[0159] Parameter settings of the feature self-expanding network (taking p2 as an example):
[0160]
[0161] The following is a simulation to illustrate the specific process of the detection target of the present invention.
[0162] 1. Simulation conditions:
[0163] The hardware platform is: HP-Z840 workstation, Intel(R) Xeon(R) E5-2630-CPU, with a main frequency of 2.40 GHz, RTX2080TI-11GB-GPU, 64GB RAM.
[0164] The software platform is: Python, PyTorch deep learning framework.
[0165] 2. Simulation content and results:
[0166] The simulation experiment of the present invention uses the NWPU VHR-10 remote sensing image dataset. The NWPU VHR-10 dataset is a challenging remote sensing target detection dataset annotated by a certain university. The dataset contains 800 images, of which 650 images contain objects. Among them, the total number of labeled images is 650, including 10 categories: airplane, oil tank, port, playground, car, baseball field, tennis court, basketball court, ship, and bridge. The dataset is randomly divided into a training set and a test set in a ratio of 6:4. The trained detection model is used to detect the remote sensing image test dataset and compared with the traditional target detection model.
[0167] Performance comparison between the present invention and the classical Yolov3 method
[0168] Method The present invention (AP) Yolov3 method (AP) Airplane 99.8 98.6 Ship 87.0 84.4 Oil tank 98.7 98.0 Baseball field 96.5 96.5 Tennis court 96.7 92.4 Basketball court 88.4 66.5 Playground 99.1 98.2 Port 97.1 96.2 Bridge 86.8 78.4 Sedan 92.2 84.7 mAP 94.2 89.4
[0169] As can be seen from the above table, compared with the traditional method, the method of the present invention has a greater improvement in detection accuracy, and the detection effects of small targets (such as cars and ships) and densely distributed targets (such as oil tanks) have been improved.
[0170] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0171] Although the present application has been described in conjunction with various embodiments, those skilled in the art will recognize other variations of the disclosed embodiments upon reviewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.
[0172] The above is a further detailed description of the present invention in connection with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited only to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as falling within the scope of protection of the present invention.
Claims
1. A variable-scale remote sensing image target detection method based on ternary feature fusion, characterized in that Including: Step 1: Obtain a dataset composed of remote sensing images and the remote sensing image to be measured; Among them, the dataset is divided into a training set and a test set, and each remote sensing image in the dataset contains at least one anchor box for annotating the target position; Step 2: Generate multiple prior boxes for predicting the position of the target based on each anchor box; Step 3: Obtain a variable-scale remote sensing image target detection model constructed based on triple feature fusion; Step 4: Use each remote sensing image in the dataset as a training sample, and iteratively train the variable-scale remote sensing image target detection model based on all training samples and the prior boxes to obtain a trained variable-scale remote sensing image target detection model; Step 5: Use the trained variable-scale remote sensing image target detection model to detect the remote sensing image to be measured, and obtain a prediction box annotating the specific position of the target; Step 6: Decode the prediction box to obtain the true coordinate position of the target; The variable-scale remote sensing image target detection model in Step 3 includes: a feature extraction backbone network, a triple feature fusion network, a feature self-expansion network, and a detection head network connected in sequence; The feature extraction backbone network is composed of multiple convolutional layers, and the triple feature fusion network includes an upsampling network, a local feature saliency enhancement network, and a feature fusion network; Among them, the feature extraction backbone network is used to perform multiple downsampling operations to extract multiple-scale feature images from the input image and input them into the upsampling network, the local feature saliency enhancement network, and the feature fusion network; the upsampling network is used to transfer deep semantic information to the shallow network through upsampling operations on the input scale feature images and input the deep semantic information into the feature fusion network; the local feature saliency enhancement network is used to transfer shallow key local detail information to the deep network through downsampling operations on the input scale feature images and input the shallow key local detail information into the feature fusion network; the feature fusion network is used to perform a triple splicing operation on the deep semantic information, the shallow key local detail information, and the scale feature input by the triple feature fusion network to obtain a fusion feature map retaining local key features and deep semantic features, and input the fusion feature map into the feature self-expansion network respectively; the feature self-expansion network is used to generate an expansion coefficient according to the category bias and spatial bias of the fusion feature map, and fine-tune the fusion feature map according to the expansion coefficient to obtain discriminative features, and input the discriminative features into the detection head network; the detection head network is used to set multiple prior boxes for each pixel in the feature map corresponding to each discriminative feature, calculate the relevant parameters corresponding to each prior box, generate the predicted values of the set prior boxes according to the relevant parameters, correct the corresponding prior boxes according to the predicted values to obtain multiple prediction boxes, retain the prediction boxes with the confidence ranking within the preset number in the prediction boxes, and use the NSM algorithm to filter out the prediction boxes with the overlap degree exceeding the overlap degree threshold to obtain the prediction boxes of the position where the target is located.
2. The method for detecting targets in variable-scale remote sensing images based on ternary feature fusion according to claim 1, wherein Before step 2, the variable-scale remote sensing target detection method based on ternary feature fusion further includes: Preprocessing the dataset and the to-be-detected remote sensing image respectively, so that the remote sensing images in the dataset and the to-be-detected remote sensing image are both of a fixed size.
3. The method for detecting targets in variable-scale remote sensing images based on ternary feature fusion according to claim 1, wherein Step 2 includes: Step 2a: Randomly select a preset number of anchor boxes from all the anchor boxes as cluster centers; Step 2b: Calculate the distances between each anchor box and other anchor boxes; Step 2c: Determine the nearest cluster center for each anchor box and assign the anchor box to the cluster; Step 2d: Update the cluster centers again according to the distances between the anchor boxes; Step 2e: Use the updated cluster centers as the prior boxes for the positions where the prediction targets are located.
4. The method for detecting targets in variable-scale remote sensing images based on ternary feature fusion according to claim 1, wherein, The feature extraction backbone network performs 5 upsampling operations to obtain 5 scale feature images: f0, f1, f2, f3, and f4. For the scale feature map f s , the scale feature map f s is input into the feature fusion network, the scale feature map f s+1 is input into the upsampling network, and the feature map f s-1 is fed into the local feature saliency enhancement network; The upsampling network is used to upsample the scale feature map f s+1 to obtain the feature map f' s+1 , and make f' s+1 the same size as the scale feature map f s to obtain the feature map f u ; The local feature saliency enhancement network is used to downsample the scale feature map f s-1 to obtain a feature map f s with the same size as the scale feature map f d ; The feature fusion network is used to perform a triple splicing operation on the feature map f u , the feature map f d , and the scale feature map f s to obtain a fused feature map p s that retains local key features and deep semantic features, and inputs the fused feature map p s into the corresponding feature self-expansion network; Where s is a positive integer from 2 to 4.
5. The method for detecting targets in variable-scale remote sensing images based on ternary feature fusion according to claim 4, wherein The local feature saliency enhancement network includes: a local enhancement network and a spatial re-scoring network; The local enhancement network is used for The scale feature map f is mapped to different subspaces through two independent embedding functions s-1 to obtain the feature maps H and G of the two subspaces; Using a sliding window Φ with a fixed step size, split the feature maps H and G into individual image patches channel by channel and Perform a max pooling operation on the image patch and use the image features obtained after the max pooling operation as the query feature map; Performing a transpose operation on the image patch and using the feature map obtained after the transpose operation as the to-be-query feature map; Performing a dot product on the to-be-query feature map and the query feature map to obtain a dependency matrix M of pixel points and high-response points in the local area; Calculating an attention weight matrix W according to the dependency matrix M; Using the attention weight matrix W as a dynamic kernel parameter and performing a local feature enhancement operation on the dynamic kernel parameter to obtain a locally enhanced output feature map; Restoring the optimal data distribution of the locally enhanced output feature map through a BN layer to obtain a new feature map L; Transfer the new feature map L and the input scale feature map f s-1 to the spatial rescoring network; The spatial re-scoring network is used for Performing a channel-level maximum pooling operation on the new feature map L to obtain a feature map L'; Perform a convolution operation on the scale feature map f s-1 to obtain the feature map F'; Flatten the feature map L' and the feature map F' from two dimensions into one-dimensional vectors respectively, and multiply the transposes of the one-dimensional vectors of the two to obtain the correlation matrix A g ; Input the correlation matrix A g into the Softmax layer to obtain the spatial attention weight matrix A q ; Multiply the spatial attention weight matrix A q by the feature map L' to obtain a spatial attention mask for the new feature map L, and obtain the mask mask according to this spatial attention mask l ; According to the mask l and the new feature map L, a feature map f s with the same size as the scale feature map f d is obtained; Input the feature map f d into the feature fusion network.
6. The variable-scale remote sensing image target detection method based on ternary feature fusion according to claim 4, characterized in that, The variable-scale remote sensing image target detection model contains three parallel feature self-expansion networks, and each feature self-expansion network includes a channel content perception module, a weighted average network, and a spatial self-expansion pooling network connected in sequence; The feature fusion network is used for inputting the fused feature maps into the corresponding feature self-expansion networks respectively; Each feature self-expansion network is used for generating an expansion coefficient according to the class bias and spatial bias of the input fused feature map, and fine-tuning the fused feature map according to the expansion coefficient to obtain a discriminative feature, and inputting the discriminative feature into the detection head network.
7. The method for detecting targets in variable-scale remote sensing images based on ternary feature fusion according to claim 6, wherein The channel content perception module includes a parallel class bias perception module and a spatial semantic perception module; The class bias perception module is used for Compress the spatial information of the fused feature map p s into a set of feature vectors p; Calculate the class bias score of the fused feature map p according to the set of feature vectors p, and input the class bias score into the connected weighted average network; s The spatial semantic perception module is used for Downsample the fused feature map p s to obtain the downsampled feature map s x ; Compress the spatial information of the feature map s x into a set of feature vectors, and use the Relu function for activation to obtain the spatial context scoring vector s; Using the spatial context scoring vector s as the spatial bias score and inputting it into the connected weighted average network; The weighted average network is used for Performing a weighted average on the spatial bias score and the class bias score to obtain a vector ΔN; Using the vector ΔN as the channel expansion coefficient and inputting it into the connected spatial self-expansion pooling network; The spatial self-expansion pooling network is used for The fused feature map is fine-tuned by approximating the channel expansion coefficient through linear interpolation to obtain the discriminative feature map P out ; Input the discriminative feature map P out into the detection head network.
8. The method for detecting targets in variable-scale remote sensing images based on ternary feature fusion according to claim 4, wherein The detection head network is used for Using the prior boxes in step 2 as the prior boxes for pixel units, and setting 3 prior boxes for each pixel unit in the feature map corresponding to each discriminative feature; Calculating the relevant parameters corresponding to the 3 prior boxes; Among them, the relevant parameters include: probability scores of various categories, center point coordinates, width and height values, and confidence scores; Generate predicted values of prior boxes according to the relevant parameters; Rectify the corresponding prior boxes according to the predicted values to obtain multiple predicted boxes.
9. The method for detecting targets in a variable-scale remote sensing image based on ternary feature fusion according to claim 1, wherein Step 4 includes: Take each remote sensing image in the dataset as a training sample; Based on all training samples and the prior boxes, iteratively train the variable-scale remote sensing image target detection model to adjust the internal parameters of the variable-scale remote sensing image target detection model in the direction of reducing the loss function until the maximum number of iterations is reached, and obtain a trained variable-scale remote sensing image target detection model; Among them, the loss function is composed of position error, confidence error, classification error, and attention error.
Citation Information
Patent Citations
Lightweight remote sensing target detection method based on SE-YOLOv3
CN112396002A
Variable-scale target detection method based on multistage feature adaptive fusion
CN112733942A