Multi-scale Object Detection Method Based on Joint Recursive Feature Pyramid

By introducing joint recursive feature pyramids and joint feedback processors, the feature pyramid structure is optimized, and the semantic gap and information loss problems in multi-scale object detection are solved, and efficient multi-scale object detection is achieved.

CN115527095BActive Publication Date: 2025-07-04XIDIAN UNIV

Patent Information

Application Number
CN202211339440.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-29
Publication Date
2025-07-04
Estimated Expiration
2042-10-29

AI Technical Summary

Technical Problem

The existing multi-scale object detection algorithm has the problem of low detection accuracy in complex scenarios, especially due to the semantic gap and information loss between the layers of the feature pyramid, the detection performance cannot be achieved optimally.

Method used

The method based on joint recursive feature pyramid is adopted, and by introducing a loop mechanism and a joint feedback processor, combining channel and spatial attention modules, the feature pyramid structure is optimized, and the unified processing and fusion of feature information is realized, the semantic gap is reduced, and the detection performance is improved.

Benefits of technology

It significantly improves the accuracy of multi-scale object detection, reduces inference speed and video memory usage, and improves detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115527095B_ABST
    Figure CN115527095B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-scale object detection method based on a joint recursive feature pyramid, which mainly solves the problem of low multi-scale object detection accuracy in complex scenarios in the prior art. The implementation scheme is as follows: 1) Read the data of the object detection database and preprocess the image data; 2) Use the ResNet convolutional neural network as the backbone network to extract the features of the image; 3) Construct a feature pyramid according to the extracted image features; 4) Construct a joint feedback processor composed of a channel attention module and a spatial attention module in series; 5) Use the joint feedback processor to process the pyramid features of each layer to complete feature fusion; 6) Repeat steps 3) to 5) twice to obtain multi-scale features; 7) Input the multi-scale features into the existing detection head to complete multi-scale detection. The present invention significantly improves the accuracy of multi-scale object detection in complex scenarios and can be used in intelligent transportation, intelligent security, and remote sensing image processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a multi-scale object detection method based on cyclic features, which can be used in fields such as traffic, security, and medical treatment. Background Art

[0002] Object detection is one of the basic tasks in the field of computer vision. It is widely applied in fields such as traffic, security, and medical treatment, and has extremely high application value. The task of object detection includes two items: locating the position of the object in the image and predicting the category of the object in the image. Among them, due to the different sizes of the objects themselves and the distances from the camera, the scales of the objects presented in the image usually have large differences, resulting in a decline in detection performance.

[0003] In recent years, the problem of multi-scale object detection has received extensive attention. Existing algorithms adopt the method of constructing a feature pyramid, that is, a specific layer in the backbone network is output separately, and a feature pyramid is constructed through downsampling and feature fusion to obtain features with high resolution and rich semantic information. In addition, some scholars improve the detection effect by introducing a cyclic mechanism into the feature pyramid and introducing switchable dilated convolutions into the backbone network.

[0004] There is a large semantic gap between the levels in the traditional feature pyramid structure. The top-down downsampling feature fusion method directly cannot well transfer high-level semantic information to the low level, and only information loss occurs in the highest layer without fusing features from a higher layer. Therefore, the multi-scale information extraction ability is insufficient. For this reason, some variant methods based on the feature pyramid structure have been proposed in the prior art.

[0005] Since the strategy of constructing feature maps with different spatial resolutions layer by layer in the feature pyramid can significantly improve the detection performance of the model for targets of different scales, the object detection algorithms based on the feature pyramid and its variants are the mainstream methods for multi-scale object detection. Ghiasi et al. used an automatic search algorithm with the feature maps to be fused as the search space and searched out a feature pyramid structure. However, the structures searched out by such automatic search algorithms often have a high dependence on the dataset, usually showing good performance on a specific dataset but mediocre performance on other datasets. Qiao et al. first introduced a cyclic mechanism into the object detection task and proposed a cyclic feature pyramid structure, and designed switchable dilated convolutions for the backbone network. However, they ignored the inherent semantic gap between the layers of the pyramid, resulting in the performance of this method not reaching the best, and the switchable dilated convolutions have a slow inference speed and high video memory occupancy. Guo et al. considered the problem that only information loss exists in the highest layer of the feature pyramid, designed a residual feature enhancement module to complement the features of the highest layer of the feature pyramid, and also designed an adaptive spatial fusion module for fusing the layers of the feature pyramid, and the fused features are then used to predict the object category and regress the object position, significantly improving the multi-scale information extraction ability of the detector. However, this method simply fuses the features of each layer and then performs prediction and regression, ignoring the inherent semantic gap between the layers, so the performance cannot reach the best. Liu et al. believed that the information propagation path in the traditional feature pyramid is too long, so they optimized the connection path in the feature pyramid to make the underlying features conducive to object localization flow to the higher layer faster, so as to improve the multi-scale object detection ability of the detector. Although this method optimized the information propagation path in the feature pyramid, they still ignored the inherent semantic gap between the layers, so the performance cannot reach the best. Summary of the Invention

[0006] Aiming at the deficiencies of the above-mentioned existing technologies, considering the information loss in the highest layer and the inherent semantic gap between the layers, the purpose of the present invention is to propose a multi-scale object detection method based on a joint recursive feature pyramid, by introducing a cyclic mechanism and integrating its advantages with the feature pyramid, so as to achieve the optimal detection performance without the need for a special convolutional layer of switchable convolutions.

[0007] To achieve the above purpose, the implementation steps of the technical solution of the present invention are as follows:

[0008] (1) Read the data of the object detection database, adjust, flip and normalize the images of the training data in turn, adjust and normalize the images of the test data in turn, and set the normalization mean and standard deviation of the three RGB channels, and finally obtain the tensor data corresponding to the images;

[0009] (2) Use the ResNet convolutional neural network including 5 cascaded convolutional blocks as the backbone network, and input the preprocessed image tensor data in (1) into this convolutional neural network to obtain the image features extracted by the 5 convolutional blocks respectively, denoted as C1, C2, C3, C4, and C5;

[0010] (3) Construct a feature pyramid based on the image features extracted by the ResNet convolutional neural network:

[0011] 3a) Respectively pass the image features C2, C3, C4, and C5 extracted by the ResNet convolutional neural network through 4 convolutional layers with a kernel size of 1×1 and a stride of 1, so that the number of channels of the C2 feature remains 256, the number of channels of the C3 feature drops from 512 to 256, the number of channels of the C4 feature drops from 1024 to 256, and the number of channels of the C5 feature drops from 2048 to 256, finally obtaining 4 layers of backbone dimensionality-reduced features C2′, C3′, C4′, and C5′;

[0012] 3b) Perform a top-down feature fusion operation on each layer of the backbone dimensionality-reduced features obtained in 3a) to form a feature pyramid structure composed of P2, P3, P4, and P5 pyramid features;

[0013] (4) Construct a joint feedback processor composed of a channel attention module and a spatial attention module connected in series;

[0014] (5) Use the joint feedback processor to process each layer of pyramid features obtained in step (3) to complete feature fusion:

[0015] 5a) Input the 4 layers of pyramid features P2, P3, P4, and P5 into the channel attention module to obtain the channel attention feature M C ;

[0016] 5b) Input the channel attention feature M obtained in 5a) C into the spatial attention module to obtain the spatial attention feature M S ;

[0017] 5c) Split the spatial attention feature M S into 4 feature maps, and downsample these 4 feature maps to the same size as the output features C i of each convolutional block in the backbone network;

[0018] 5d) Respectively pass the upsampled feature maps through 4 convolutional layers with a kernel size of 1×1 and a stride of 1 to increase the number of channels to 256, 512, 1024, and 2048 respectively, obtaining the feature maps M i to be fused with the backbone network, and then each feature map M i and the output features C of each convolutional block in the backbone networki Add corresponding values to complete feature fusion;

[0019] (6) Repeat steps (3) to (5) twice to obtain the final multi-scale features P2′, P3′, P4′, and P5′, and input them into the existing detection head network to output the predicted target position parameters (x, y, w, h) and the confidence c of the corresponding target category. Among them, (x, y) are the coordinates of the upper left corner of the target bounding box in the image, w is the width of the target bounding box, and h is the height of the target bounding box, completing the detection of multi-scale targets.

[0020] The present invention has the following advantages compared with the prior art:

[0021] First, since the present invention introduces a joint feedback processor on the basis of the cyclic feature pyramid to uniformly process the feedback features of the feature pyramid, it can enable the top-layer features of the feature pyramid to have information flow supplementation, and the semantic gap between layers can be reduced, and the multi-scale information extraction ability of the detector is improved, thereby improving the network detection effect;

[0022] Second, since the present invention does not require special convolution operations such as switchable dilated convolutions to increase the receptive field, the inference speed of the method of the present invention is significantly improved compared with other cyclic methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is the implementation flowchart of the present invention;

[0024] Figure 2 is the schematic diagram of the joint recursive feature pyramid in the present invention;

[0025] Figure 3 is the schematic diagram of the joint feedback processor in the present invention;

[0026] Figure 4 is the simulation result diagram of detecting ship targets in optical remote sensing images using the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] The embodiments and effects of the present invention will be further described below with reference to the accompanying drawings.

[0028] Refer to Figure 1 , the implementation steps of this embodiment are as follows:

[0029] Step 1, read the data of the target detection database and preprocess the image data.

[0030] The data of the target detection database includes the data in the training stage and the data in the test stage. In this step, the image data in these two stages are preprocessed as follows:

[0031] 1.1) Data preprocessing in the training stage:

[0032] First, scale the size of the input image to 800×800, then randomly adjust the brightness, contrast, saturation, and hue of the image with a probability of 0.5, and then randomly flip the image with a probability of 0.5;

[0033] Normalize the image using the mean-standard deviation normalization method. Among them, set the normalization means of the RGB three channels to [123.675, 116.28, 103.53] respectively, and set the standard deviations of the three channels to [58.395, 57.12, 57.375] respectively. Finally, obtain the tensor data corresponding to the image in this stage;

[0034] 1.2) Data preprocessing in the testing stage:

[0035] Scale the size of the input image to 800×800;

[0036] Normalize the image using the mean-standard deviation normalization method. Among them, set the normalization means of the RGB three channels to [123.675, 116.28, 103.53] respectively,

[0037] Set the standard deviations of the three channels to [58.395, 57.12, 57.375] respectively. Finally, obtain the tensor data corresponding to the image in this stage.

[0038] Step 2, use the ResNet convolutional neural network as the backbone network to extract the features of the image.

[0039] The ResNet convolutional neural network has 5 cascaded convolutional blocks. Each convolutional block contains several groups of convolutional groups. Each group of convolutional groups contains a convolutional layer, a batch normalization layer, and a ReLu activation function. The backbone network used in the present invention involves three versions of ResNet-50, ResNet-101, and ResNet-152. Input the image tensor preprocessed in Step 1 into the ResNet convolutional neural network to extract the image features. The image features extracted by the 5 convolutional blocks are respectively denoted as C1, C2, C3, C4, and C5. The backbone network structure and the respectively extracted image features are shown in Table 1.

[0040] Table 1: ResNet convolutional neural network structure and extracted image features

[0041]

[0042] Step 3, construct a feature pyramid according to the image features extracted by the ResNet convolutional neural network.

[0043] Refer to Figure 2 , the specific implementation of this step is as follows:

[0044] 3.1) The image features C2, C3, C4, and C5 extracted by the ResNet convolutional neural network are respectively passed through 4 convolutional layers with a kernel size of 1×1 and a stride of 1, so that the number of channels of the C2 feature remains 256, the number of channels of the C3 feature drops from 512 to 256, the number of channels of the C4 feature drops from 1024 to 256, and the number of channels of the C5 feature drops from 2048 to 256. Finally, 4 layers of backbone dimensionality reduction features C2′, C3′, C4′, and C5′ are obtained;

[0045] 3.2) Perform a top-down feature fusion operation on the backbone dimensionality reduction features of each layer obtained in 3.1):

[0046] 3.2.1) Denote the highest-layer backbone dimensionality reduction feature as the highest-layer pyramid feature P5. After performing a 2-fold upsampling operation on P5 and directly adding it to the sub-highest-layer backbone dimensionality reduction feature, the sub-highest-layer pyramid feature P4 is obtained;

[0047] 3.2.2) After performing a 2-fold upsampling operation on the sub-highest-layer pyramid feature P4 and directly adding it to the sub-lowest-layer backbone dimensionality reduction feature, the sub-lowest-layer pyramid feature P3 is obtained;

[0048] 3.2.3) After performing a 2-fold upsampling operation on the sub-lowest-layer pyramid feature P3 and directly adding it to the lowest-layer backbone dimensionality reduction feature, the lowest-layer pyramid feature P2 is obtained;

[0049] 3.3) Arrange the above pyramid features P2, P3, P4, and P5 from bottom to top to form a feature pyramid structure.

[0050] Step 4, construct a joint feedback processor.

[0051] 4.1) Select a channel attention module that sequentially includes an upsampling layer, a feature concatenation layer, a global average pooling layer, a fully connected layer, and a Sigmoid function to extract channel attention features. Among them, the expression of the Sigmoid function is:

[0052]

[0053] 4.2) Select a spatial attention module that sequentially includes an average pooling layer, a max pooling layer, a convolutional layer, and a Sigmoid function to extract spatial attention features;

[0054] 4.3) Connect the channel attention module and the spatial attention module in series to form a joint feedback processor.

[0055] Step 5, use the joint feedback processor to process the pyramid features of each layer obtained in Step 3 to complete feature fusion.

[0056] Refer toFigure 3 , the specific implementation of this step is as follows:

[0057] 5.1) Input the four-layer pyramid features P2, P3, P4, and P5 into the channel attention module to obtain the channel attention feature M C :

[0058] 5.1.1) Upsample the pyramid features P2, P3, P4, and P5 respectively to obtain their corresponding upsampled features X2, X3, X4, and X5. The sizes of these corresponding features are all 200×200, and the number of channels is 256;

[0059] 5.1.2) Concatenate the upsampled corresponding pyramid features X2, X3, X4, and X5 into a total channel feature M cat1 , whose size is 200×200 and the number of channels is 1024;

[0060] 5.1.3) Compress the total channel feature M cat1 through a global average pooling layer into an average pooling compression vector V with a length of 1024 gap ;

[0061] 5.1.4) Pass the average pooling compression vector V gap through a set of fully connected layers, batch normalization layers, and a ReLu activation function, and compress it again to obtain a channel re-compression vector V with a length of 256 fc1 ;

[0062] 5.1.5) Pass the channel re-compression vector V fc1 through another fully connected layer to release the number of channels, and obtain a channel release vector with a length of 1024

[0063] 5.1.6) Use the Sigmoid function to normalize the channel release vector V fc2 to obtain a normalized vector V with a length of 1024 norm ;

[0064] 5.1.7) Take the dot product of the total channel feature M cat1 and the normalized vector V norm to obtain the channel attention feature M C :

[0065] M C = M cat1 · V norm

[0066] where the channel attention feature M C has a size of 200×200 and the number of channels is 1024.

[0067] 5.2) Feed the channel attention feature M obtained in 5.1) C into the spatial attention module to obtain the spatial attention feature M S :

[0068] 5.2.1) Pass the channel attention feature M C through a max - pooling layer and an average - pooling layer respectively to obtain the max - pooling feature M max and the average - pooling feature M avg , where the sizes of the max - pooling feature and the average - pooling feature are both 200×200, and the number of channels is 1;

[0069] 5.2.2) Concatenate the max - pooling feature M max and the average - pooling feature M avg into a total spatial feature M cat2 , whose size is 200×200 and the number of channels is 2;

[0070] 5.2.3) After passing the total spatial feature M cat2 through a convolutional layer with a kernel size of 7×7 and a stride of 1, obtain a new feature M un , whose size is 200×200 and the number of channels is 1;

[0071] 5.2.4) Use the Sigmoid function to normalize the new feature M un to obtain the normalized feature M norm , whose size is 200×200 and the number of channels is 1;

[0072] 5.2.5) Perform the Hadamard product on the channel attention feature M C and the normalized feature M norm to obtain the spatial attention feature M S :

[0073]

[0074] where the symbol represents the Hadamard product, and the size of the spatial attention feature M S is 200×200 and the number of channels is 1024;

[0075] 5.3) Split the spatial attention feature M S into 4 feature maps, and downsample these 4 feature maps respectively to the same size as the output feature C i of each convolutional block in the backbone network;

[0076] 5.4) Input the upsampled feature maps into four convolutional layers with a kernel size of 1×1 and a stride of 1 respectively, and increase the number of channels to 256, 512, 1024, and 2048 respectively to obtain the feature maps M to be fused with the backbone network. i , and then for each feature map M i and the output feature C of each convolutional block in the backbone network i perform element-wise addition to complete feature fusion.

[0077] Step 6, complete the detection of multi-scale targets.

[0078] 6.1) Repeat steps 3 to 5 twice to obtain the final multi-scale features P2′, P3′, P4′, and P5′;

[0079] 6.1) Input the multi-scale features P2′, P3′, P4′, and P5′ into the existing detection head network, and output the predicted target position parameters (x, y, w, h) and the confidence c of the corresponding target category. Among them, (x, y) are the coordinates of the upper left corner of the target bounding box in the image, w is the width of the target bounding box, and h is the height of the target bounding box, to complete the detection of multi-scale targets.

[0080] The following further describes the effect of the present invention in combination with simulation experiments.

[0081] 1. Experimental conditions:

[0082] The computer processor used is an Intel(R) Core(TM) i7 CPU@3.5GHz, the running memory is 128G, and the graphics card is an NVIDIA TITAN X GPU with a video memory of 12GB.

[0083] The operating system is 64-bit Ubuntu 18.04 (LTS), and the deep learning framework used is PyTorch (version 1.8.0).

[0084] All network trainings use the backpropagation algorithm to calculate the residuals of each layer, and use the stochastic gradient descent algorithm with momentum term and weight decay term to update the network parameters, where the momentum term is 0.9 and the weight decay term is 0.0001.

[0085] The experiments use the HRSC2016 optical remote sensing ship detection database, the self-built database HRSC2016-MS, and the DIOR large-scale optical remote sensing target detection database for evaluation, and the evaluation indicators are mAP, AP S , AP M and AP L . Among them, mAP is the mean average precision under the 50% intersection over union threshold, AP S is the average precision of targets with a size smaller than 32×32, APM The average precision for targets with sizes greater than or equal to 32×32 and less than 96×96 is AP. L The average precision for targets with sizes greater than 96×96.

[0086] The HRSC2016 database is the only currently open-source optical remote sensing ship detection database, containing 1,070 optical remote sensing images with spatial resolutions of 2 meters and 0.4 meters. The image sizes range from 300×300 to 1500×900, and most of the images are larger than 1000×1000, containing 2,976 ship instances.

[0087] The self-built database HRSC2016-MS is an optical remote sensing ship detection database expanded and re-annotated on the basis of the HRSC2016 database, containing 1,680 optical remote sensing images and 7,655 ship instances.

[0088] The DIOR database is a relatively large-scale optical remote sensing target detection database at present, containing 23,463 optical remote sensing images, covering 20 target categories with a total of 192,472 target instances.

[0089] 2. Experimental content:

[0090] Experiment 1: Under the above experimental conditions, the ship targets in the HRSC2016 and HRSC2016-MS databases are detected using the method of the present invention and 13 existing methods. The detection results are shown in Table 2.

[0091] Table 2 Detection results of the present invention and 13 existing methods on the HRSC2016 and HRSC2016-MS databases

[0092]

[0093]

[0094] The 13 existing methods in Table 2 are as follows:

[0095] SSD: A single-stage multi-bounding box target detection algorithm proposed by Liu et al.

[0096] YOLOF: A target detection algorithm based on a single-level feature map proposed by Chen et al.

[0097] RetinaNet: A single-stage target detection algorithm based on Focal Loss proposed by Lin et al.

[0098] NAS-FPN: A target detection algorithm with a pyramid feature structure searched in a certain search space based on a neural network architecture search algorithm proposed by Ghiasi et al.

[0099] FCOS: A fully convolutional single-stage object detection algorithm proposed by Tian et al.;

[0100] PANet: A two-stage object detection algorithm based on path aggregation feature pyramid proposed by Liu et al.;

[0101] Faster R-CNN: A real-time two-stage object detection algorithm based on region proposal network proposed by Ren et al.;

[0102] Mask R-CNN: An algorithm for object instance segmentation and object detection that adds a mask prediction branch to Faster R-CNN by He et al.;

[0103] Cascade R-CNN: A two-stage object detection algorithm based on cascade R-CNN proposed by Cai et al.;

[0104] DetectoRS: An object detection algorithm based on recurrent feature pyramid structure proposed by Qiao et al.;

[0105] Libra R-CNN: An object detection algorithm based on balanced intersection over union sampling, balanced feature pyramid, and balanced L1 loss function proposed by Pang et al.;

[0106] YOLOX: A high-performance single-stage fast object detection algorithm proposed by Ge et al. by integrating various design techniques;

[0107] HTC: A hybrid task cascade model proposed by Chen et al. for object detection and object instance segmentation tasks based on Mask R-CNN and Cascade R-CNN.

[0108] Among them, the subjective results of the ship detection method of the present invention on the HRSC2016-MS database are as Figure 4 shown, and small, medium, and large multi-scale ship targets in the optical remote sensing images can all be accurately detected to obtain the corresponding bounding boxes.

[0109] From Figure 4 the subjective results shown in and the objective results shown in Table 2, it can be seen that the method of the present invention has achieved the best detection effect on both the HRSC2016 and HRSC2016-MS databases, proving the effectiveness of the method of the present invention.

[0110] Experiment 2: Under the above conditions, the joint recursive feature pyramid structure proposed in the present invention and 5 existing feature pyramid structures are used as the neck structure and combined with the baseline method for comparison on the HRSC2016-MS database. The baseline method is the HTC model with the neck structure and semantic prediction branch removed. The results are shown in Table 3.

[0111] Table 3 Comparison results of the joint recursive feature pyramid of the present invention and the existing five feature pyramid structures in the HRSC2016-MS database

[0112]

[0113] The methods in Table 3 are described as follows:

[0114] Baseline: Baseline method, specifically the HTC model without the neck structure and semantic prediction branch;

[0115] Baseline+FPN: A method that combines the traditional feature pyramid as a neck structure with the baseline method;

[0116] Baseline+PAFPN: A method that combines the path aggregation feature pyramid as a neck structure with the baseline method;

[0117] Baseline+BFP: Balanced feature pyramid as a neck structure combined with the baseline method;

[0118] Baseline+BiFPN: A method that combines a two-stream feature pyramid as a neck structure with a baseline method;

[0119] Baseline+RFP: a method that combines the recurrent feature pyramid as a neck structure with the baseline method;

[0120] Baseline+JRFP: The proposed joint recursive feature pyramid as a neck structure is combined with the baseline method.

[0121] From the results shown in Table 3, it can be seen that the joint recursive feature pyramid proposed by the method of the present invention has achieved the best detection effect as the neck structure on the HRSC2016-MS database, and has achieved the best detection effect at three scales: large, medium and small, which further proves the effectiveness of the method of the present invention.

[0122] Experiment 3: Under the above conditions, the method of the present invention and 15 existing methods are used to perform target detection on the large-scale optical remote sensing database DIOR. The results are shown in Table 4.

[0123] Table 4 Detection results of the method of the present invention and 15 existing methods on the DIOR database

[0124] Method Mean Average Precision R-CNN 37.7 RICNN 44.2 RICAOD 50.9 RIFD-CNN 56.1 SSD 58.6 Faster R-CNN 63.1 Mask R-CNN 63.5 CornerNet 64.9 RetinaNet 65.7 Cascade R-CNN 70.3 YOLOv3 71.0 PANet 71.1 DetectoRS 71.8 HTC 72.6 AFPN 72.6 The method of the present invention 76.9

[0125] The seven methods not mentioned above in Table 4 are introduced as follows:

[0126] R-CNN: An object detection algorithm based on region convolution proposed by Girshick et al.;

[0127] RICNN: An object detection algorithm for high-resolution optical remote sensing images based on rotation-invariant convolution proposed by Cheng et al.;

[0128] RICAOD: A remote sensing image object detection algorithm based on a rotation-insensitive region proposal network and a local context feature fusion network proposed by Li et al.;

[0129] RIFD-CNN: A remote sensing image object detection algorithm based on rotation invariance and Fisher discriminant convolution proposed by Cheng et al.;

[0130] CornerNet: An object detection algorithm based on an hourglass network proposed by Law et al.;

[0131] YOLOv3: The third version of the YOLO series of single-stage fast object detection algorithms proposed by Joseph et al.;

[0132] AFPN: A remote sensing image object detection algorithm based on a perceptual feature pyramid structure proposed by Cheng et al.

[0133] From the results shown in Table 4, it can be seen that the method of the present invention has achieved the best detection effect on the large-scale optical remote sensing database DIOR database, further proving the effectiveness of the method of the present invention.

Claims

1. A multi-scale object detection method based on a joint recursive feature pyramid, characterized by comprising the following steps: (1) Read the data of the object detection database, adjust, flip, and normalize the images of the training data in sequence, adjust and normalize the images of the test data in sequence, and set the normalization means and standard deviations of the three RGB channels, and finally obtain the tensor data corresponding to the images; (2) Use a ResNet convolutional neural network including 5 cascaded convolutional blocks as the backbone network, input the image tensor data preprocessed in (1) into the convolutional neural network, and obtain the image features extracted by the 5 convolutional blocks respectively, denoted as C1, C2, C3, C4, and C5; (3) Construct a feature pyramid according to the image features extracted by the ResNet convolutional neural network: 3a) Respectively pass the image features C2, C3, C4, and C5 extracted by the ResNet convolutional neural network through 4 convolutional layers with a kernel size of 1×1 and a stride of 1, so that the number of channels of the C2 feature remains 256, the number of channels of the C3 feature drops from 512 to 256, the number of channels of the C4 feature drops from 1024 to 256, and the number of channels of the C5 feature drops from 2048 to 256, and finally obtain 4 layers of backbone downsampled features C2′, C3′, C4′, and C5′; 3b) Perform a top-down feature fusion operation on each layer of the backbone downsampled features obtained in 3a) to form a feature pyramid structure composed of P2, P3, P4, and P5 pyramid features; (4) Construct a joint feedback processor composed of a channel attention module and a spatial attention module connected in series; (5) Use the joint feedback processor to process each layer of pyramid features obtained in step (3) to complete feature fusion; 5a) Input the four layers of pyramid features P2, P3, P4, and P5 into the channel attention module to obtain the channel attention feature M C ; 5b) Input the channel attention feature M obtained in 5a) C into the spatial attention module to obtain the spatial attention feature M S ; 5c) Split the spatial attention feature M S into four feature maps, and downsample these four feature maps to the same size as the output features C of each convolutional block in the backbone network i respectively; 5d) Input the upsampled feature maps into 4 convolutional layers with a kernel size of 1×1 and a stride of 1 respectively, and increase the number of channels to 256, 512, 1024, and 2048 respectively to obtain the feature maps M to be fused with the backbone network. i , and then each feature map M i is added to the output feature C of each convolutional block in the backbone network i correspondingly to complete the feature fusion; (6) Repeat steps (3) to (5) twice to obtain the final multi-scale features P2′, P3′, P4′, and P5′, input them into the existing detection head network, and output the predicted target position parameters (x, y, w, h) and the confidence c of the corresponding target category, where (x, y) are the coordinates of the upper left corner of the target bounding box in the image, w is the width of the target bounding box, and h is the height of the target bounding box, and complete the detection of multi-scale objects.

2. The method according to claim 1, characterized in that, In step (1), the images in the training stage and the test stage are adjusted, flipped, and normalized in sequence, and the means and standard deviations of the three RGB channels are set to achieve the following: 1a) Data preprocessing in the training stage: Scale the size of the input image to 800×800, and randomly adjust the brightness, contrast, saturation, and hue of the image with a probability of 0.5; Then randomly flip with a probability of 0.5, and normalize the image using the mean standard deviation normalization method; Set the normalization means of the three RGB channels to [123.675, 116.28, 103.53] respectively, and set the standard deviations of the three channels to [58.395, 57.12, 57.375] respectively, and finally obtain the tensor data corresponding to the images in this stage; 1b) Data preprocessing in the test stage: Scale the size of the input image to 800×800, and then normalize the image using the mean standard deviation normalization method; Set the normalized means of the three RGB channels to [123.675, 116.28, 103.53] respectively, and set the standard deviations of the three channels to [58.395, 57.12, 57.375] respectively. Finally, obtain the tensor data corresponding to the image at this stage.

3. The method according to claim 1, wherein The five cascaded convolutional blocks of the ResNet convolutional neural network in step (2) have the same structure. Each convolutional block contains several groups of convolutional groups, and each group of convolutional groups contains a convolutional layer, a batch normalization layer, and a ReLu activation function.

4. The method according to claim 1, wherein In step (3b), perform a top-down feature fusion operation on the features of each layer obtained in 3a) as follows: 3b1) Denote the highest-level backbone downsampled feature C5′ as the highest-level pyramid feature P5. After performing a 2-fold upsampling operation on P5, directly add it to the second-highest-level backbone downsampled feature C4′ to obtain the second-highest-level pyramid feature P4; 3b2) After performing a 2-fold upsampling operation on the second-highest-level pyramid feature P4, directly add it to the second-lowest-level backbone downsampled feature C3′ to obtain the second-lowest-level pyramid feature P3; 3b3) After performing a 2-fold upsampling operation on the second-lowest-level pyramid feature P3, directly add it to the lowest-level backbone downsampled feature C2′ to obtain the lowest-level pyramid feature P2; 3b4) Arrange the above pyramid features P2, P3, P4, and P5 from bottom to top, to form a feature pyramid structure.

5. The method according to claim 1, wherein The structures of the channel attention module and the spatial attention module in step (4) are as follows: The channel attention module sequentially includes operations of upsampling, feature splicing, global average pooling layer, fully connected layer, and Sigmoid function. This module is used to extract channel attention features; The spatial attention module sequentially includes an average pooling layer, a max pooling layer, a convolutional layer, and a Sigmoid function. This module is used to extract spatial attention features.

6. The method according to claim 1, wherein In step 5a), the four pyramid features of P2, P3, P4, and P5 are input into the channel attention module to obtain the channel attention feature M C , which is implemented as follows: 5a1) Perform upsampling on the pyramid features P2, P3, P4, and P5 respectively to obtain their corresponding upsampled features X2, X3, X4, and X5. The sizes of these corresponding features are all 200×200, and the number of channels is all 256; 5a2) Concatenate the upsampled pyramid corresponding features X2, X3, X4, and X5 into a total channel feature M cat1 , with a size of 200×200 and 1024 channels; 5a3) Compress the total channel feature M cat1 through a global average pooling layer into an average pooling compression vector V with a length of 1024 gap ; 5a4) Average pool the compressed vector V gap Pass it through a set of fully connected layers, batch normalization layers, and a ReLu activation function, and compress it again to obtain a channel re-compressed vector V of length 256 fc1 ; 5a5) Recompress the channel into vector V fc1 Release the number of channels through another fully connected layer to obtain a channel release vector with a length of 1024 5a6) Normalize the channel release vector V using the Sigmoid function fc2 to obtain a normalized vector V of length 1024 norm ; 5a7) Dot the total channel feature M cat1 with the normalized vector V norm to obtain the channel attention feature M C : M C = M cat1 ·V norm Among them, the channel attention feature M C has a size of 200×200 and 1024 channels.

7. The method according to claim 1, characterized in that, In step 5b), the channel attention feature M C is input into the spatial attention module to obtain the spatial attention feature M S , which is implemented as follows: 5b1) Pass the channel attention feature M C through a max pooling layer and an average pooling layer respectively to obtain the max pooling feature M max and the average pooling feature M avg , where the sizes of the max pooling feature and the average pooling feature are both 200×200, and the number of channels is both 1; 5b2) Concatenate the maximum pooling feature M max and the average pooling feature M avg to form a total spatial feature M cat2 with a size of 200×200 and 2 channels; 5b3) The overall spatial feature M cat2 After passing through a convolutional layer with a kernel size of 7×7 and a stride of 1, a new feature M is obtained un , with a size of 200×200 and 1 channel; 5b4) Use the Sigmoid function to normalize the new feature M un to obtain the normalized feature M norm , whose size is 200×200 and the number of channels is 1; 5b5) Multiply the channel attention feature M C with the normalized feature M norm using Hadamard product to obtain the spatial attention feature M S : Among them, the symbol represents the Hadamard product, and the spatial attention feature M S has a size of 200×200 and 1024 channels.

Citation Information

Patent Citations

  • Remote sensing image vehicle target detection method based on multi-scale attention mechanism

    CN111738110A

  • Lightweight license plate detection and recognition method based on multi-scale attention mechanism

    CN112308092A

Cited By

  • Multi-pedestrian target detection method based on YOLOCDG network in complex scene

    CN121330721A