Intelligent Detection Method for Dim and Small Targets in Electro-optical Imaging under Complex Environments
By using deep convolution, residual module and step-type feature fusion methods in convolutional neural networks, the problem of low accuracy of weak object detection in complex environments is solved, and efficient feature information extraction and target recognition are achieved.
Patent Information
- Application Number
- CN202310196144.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-03
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-03-03
AI Technical Summary
When existing neural network algorithms perform weak target detection in complex environments, there are problems such as high false alarm rate, high missed detection rate, and lack of generalization ability. Especially in photoelectric imaging, the target size is small and the feature information is weak and difficult to identify, resulting in low detection accuracy.
Feature extraction is performed using depth convolution with a convolution kernel size of 5×5 and 3×3 convolution with a step size of 2. Multi-scale feature fusion is performed by combining the residual module layer stacked by four layers and the step-type structure. The lightweight convolution module ShuffleNet2 is used for feature supplementation, and weak object detection is performed through the classification detection judge.
It improves detection accuracy in complex environments, reduces error detection rate, reduces model complexity, and improves the richness and detection efficiency of feature information.
Smart Images

Figure CN116310359B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and particularly relates to an intelligent detection method for small and weak targets in optoelectronic imaging in a complex environment. Background Technique
[0002] In a variety of application scenarios, there is a need to use algorithm software to simulate the human eye to intelligently detect, identify and classify targets in image files and perform more complex tasks on this basis. In the long process of biological evolution, the visual system with the human eye as the information input and the human brain as the information processor is very powerful, with not only a fast processing speed but also a strong anti-interference ability. The early development of computer vision was slow, and it was quite difficult to reach the level of the human visual system. With the continuous development of the field of machine learning research, after the introduction of deep learning methods into computer vision, computers have achieved better results than the human eye in many visual tasks.
[0003] Classification detection and segmentation tasks constitute the basic tasks of computer vision. Therefore, target detection and recognition, as the pre-task of advanced visual tasks, have a long research history. The detection and recognition tasks not only need to determine whether a target is included but also need to describe the detailed coordinate information of the target in the picture in the form of a rectangular prediction box. After completing the detection of the position of a small and weak target, the target category is classified and judged and the result is given. In recent years, with the continuous progress of computer vision-related theories and technologies, the detection of small and weak targets in complex environments has gradually become a new research hotspot in this field. However, existing neural network algorithms still have problems such as false alarms and missed detections when detecting small and weak targets in complex environments.
[0004] Traditional object detection algorithms mainly use manually set artificial design features and use classifiers to make judgments under sliding windows. Traditional algorithms are mainly divided into two categories: one based on spatial filtering algorithms and one based on the human visual system. However, traditional methods require a large amount of expert knowledge and labor costs to design templates or detection rules. At the same time, there are also problems such as large computational complexity, poor generalization performance, high false alarm rates and high missed detection rates of algorithms in complex environments. With the emergence of deep learning technology, due to its strong feature extraction and information abstraction capabilities, it has gradually been migrated to deep neural networks for weak object detection. In particular, convolutional neural networks have a low false alarm rate and low missed detection rate for detecting small and weak targets in complex environments, and more and more scholars use this deep learning method for task research. Convolutional neural networks for object detection can be divided into two-stage networks and single-stage networks at the detector installation stage. Representatives of two-stage object detectors are RCNN (Regions with CNN features) and Faster RCNN obtained by subsequent optimization. Representatives of single-stage object detectors are SSD (Single Shot MultiBox Detector) and YOLO. However, the current object detection algorithms still have deficiencies in the detection of small and weak targets in complex environments, such as lack of generalization ability, high missed detection rates of detectors, and poor recognition effects in complex environments. Summary of the Invention
[0005] Aiming at the technical problems of difficult recognition of small and weak targets in optoelectronic imaging in complex environments due to small target sizes, weak feature information, and complex backgrounds resulting in low accuracy, the present invention provides an intelligent detection method for small and weak targets in optoelectronic imaging in complex environments.
[0006] The technical solution adopted by the present invention is as follows:
[0007] An intelligent detection method for small and weak targets in optoelectronic imaging in complex environments, the method comprising the following steps:
[0008] Step S1: Perform size normalization processing on the image to be detected to obtain the desired image size;
[0009] Step S2: Perform grouped convolution on the image processed in step S1 using a deep convolution with a convolution kernel size of 5×5 to obtain a first feature map;
[0010] Step S3: Perform the first downsampling through a convolution with a stride of 2 and a convolution kernel size of 3×3 to obtain a second feature map;
[0011] Step S4: Perform multi-scale feature extraction on the second feature map through four stacked residual module layers;
[0012] The residual module layer is specifically as follows: The second feature map generated in step S3 serves as the input feature map of the first-layer residual module layer, and the input feature maps of the second, third, and fourth-layer residual module layers are the output feature maps of the previous residual module layer; that is, the output feature map of the previous layer serves as the input feature map of the next layer.
[0013] The network structure of each layer of the residual module layer is the same. The lightweight convolutional structure ShuffleNet2 is used as the basic convolutional module. After the input feature map of the residual module layer passes through the basic convolutional module, it is concatenated with the input feature map of the current residual module layer by channels to obtain the output feature map of the current residual module layer.
[0014] Step S5: Perform multi-scale fusion processing on the output feature maps of the second, third, and fourth-layer residual module layers using a stepped structure feature fusion method to obtain a fused feature map.
[0015] Among them, the stepped structure feature fusion method is as follows:
[0016] Define the output feature maps of the second, third, and fourth-layer residual module layers as the first-scale feature map, the second-scale feature map, and the third-scale feature map respectively, and define the feature map dimensions of the output feature maps of the second, third, and fourth-layer residual module layers as the first scale, the second scale, and the third scale respectively.
[0017] Transform the feature map dimension of the third-scale feature map into the second scale, and then perform channel concatenation with the second-scale feature map to obtain the first concatenation result; and transform the feature map dimension of the first concatenation result into the second scale to obtain the first concatenation result of the second scale.
[0018] Transform the feature map dimension of the second-scale feature map into the first scale, and then perform channel concatenation with the first-scale feature map to obtain the second concatenation result; and transform the feature map dimension of the second concatenation result into the first scale to obtain the second concatenation result of the first scale.
[0019] Transform the feature map dimension of the first concatenation result of the second scale into the first scale, and then perform channel concatenation with the second concatenation result of the first scale to obtain the third concatenation result, and transform the feature map dimension of the third concatenation result into the first scale to obtain the third concatenation result of the first scale.
[0020] Perform global average pooling operation on the third-scale feature map, and transform the feature map dimension of the global average pooling operation result into the first scale, and then perform feature fusion with the third concatenation result of the first scale to obtain a fused feature map.
[0021] Step S6: Perform weak target detection processing on the fused feature map through a classification and detection discriminator to obtain the detection processing result of the weak target.
[0022] The weak and small target refers to a target whose size is less than or equal to a specified size.
[0023] Further, step S6 includes:
[0024] Step S601: Extract candidate boxes from the fused feature map, and obtain multiple final candidate boxes after screening and non-maximum suppression processing on the extracted candidate boxes;
[0025] Step S602: Based on each of the final candidate boxes obtained in step S601, extract the features corresponding to each final candidate box from the fused feature map, and perform a pooling operation on the features of the candidate boxes to obtain candidate box features that meet the expected feature size;
[0026] Step S603: Detect weak and small targets for the candidate box features based on the classification detection decision maker.
[0027] Further, the classification detection decision maker is composed of a fully connected layer and a softmax function layer.
[0028] Further, in step S1, the expected image size is: 512×512.
[0029] The technical solution provided by the present invention at least brings the following beneficial effects:
[0030] Regarding the weak and small target detection task in optoelectronic imaging in a complex environment, due to the technical problems of small target size, weak feature information, difficult recognition, and low accuracy caused by complex background, the present invention improves the detection network so that the network model algorithm of the present invention has the characteristic of richer feature information. While effectively improving the accuracy performance, the present invention reduces the false detection rate and lowers the model complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0032] Figure 1 It is a schematic diagram of U-shaped structure feature fusion in an embodiment of the present invention;
[0033] Figure 2 It is a schematic diagram of stepped structure feature fusion in an embodiment of the present invention;
[0034] Figure 3 It is a schematic diagram of the residual module structure in an embodiment of the present invention;
[0035] Figure 4 It is a schematic structural diagram of the same-scale residual module in an embodiment of the present invention;
[0036] Figure 5 It is a schematic process diagram of an intelligent detection method for small and weak targets in optoelectronic imaging in a complex environment provided by an embodiment of the present invention. Specific embodiments
[0037] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below in conjunction with the accompanying drawings.
[0038] The problem of small and weak target detection has always been an important research issue in the fields of computer vision and artificial intelligence. It is not only a prerequisite for advanced vision tasks but also can be widely applied to real scenarios such as satellite remote sensing, military detection, and airport scene detection. With the in-depth research of deep learning, especially neural network algorithms, research methods based on neural networks are increasingly applied to the detection of small and weak targets in complex environments. Aiming at the characteristics of less and easily lost feature information of small and weak targets and large image disturbances in complex environments, improving the eigenvalue extraction and fusion methods and introducing the residual idea will improve the detection efficiency.
[0039] The task requirements for identifying small and weak targets in optoelectronic imaging in complex environments are common in all aspects of social production and life. There will be small objects on the airport runway, such as hats, screws, washers, nails, and fuses. If these small problems can be intelligently detected through surveillance images, the work pressure of airport staff will be effectively reduced. For autonomous driving, intelligent detection from surveillance images will effectively judge the road congestion level and the probability of traffic accidents. For factory raw materials, foreign objects and defects can be identified through the small and weak target detection algorithm in complex environments. For military reconnaissance, if the enemy's deployment information and movements can be accurately detected through satellite remote sensing images, the combat efficiency of our army can be greatly improved and the personnel casualties can be reduced. For public security management, human behavior analysis can be carried out through surveillance perspective data, and early warnings can be issued for dangerous behaviors such as possible stampedes and violent attacks to improve public safety. At the same time, the intelligent detection and recognition technology of small and weak targets is also an important support for tracking and predicting the routes of possible criminal suspects.
[0040] A complex environment refers to the difficulties of target overlap and target occlusion. Optoelectronic imaging refers to an electronic information image obtained by an optoelectronic imaging device. A small and weak target refers to a target with a size smaller than a specified size. For example, in the case of a 256*256 pixel size, the target size does not exceed 9*9 pixel size.
[0041] As a possible implementation, in the embodiments of the present invention, based on the Faster RCNN framework and improved by adopting a method of fusing same-scale and multi-scale features, that is, using a residual structure and channel splicing on the convolutional layer to fuse feature maps of the same size at the same level, thereby improving the detection accuracy. And replacing the original convolutional module with a lightweight convolutional module ShuffleNetV2 to improve the algorithm efficiency.
[0042] In the Faster RCNN network structure, first, feature maps are extracted from the image to be detected through a convolutional network. The convolutional network includes a convolutional layer (conv layer), a Relu activation layer (relu layer), and a pooling layer (pooling layer). There is a Relu activation layer after each convolutional layer, with a total of 13 convolutional layers and 13 Relu activation layers, and 4 pooling layers. The first two pooling layers are: one pooling layer is set after every two convolutional layers and Relu activation layers. The last two pooling layers are: one pooling layer is set after every three convolutional layers and Relu activation layers. According to the convolutional and pooling formulas, the size of the feature map remains unchanged after passing through each conv layer and relu layer; after passing through each pooling layer, the width and height of the feature map become half of the previous ones. For example, the size of the feature map generated after a picture of M*N is extracted through the network is (M / 16)*(N / 16).
[0043] Then, the obtained feature maps are input into the RPN (Region Proposal Network) to extract candidate boxes. This step is the main difference between the two-stage algorithm and the single-stage algorithm. After inputting into the FPN network, candidate boxes are obtained, and they are classified into two categories using SVM, and multiple (for example, 2000) candidate boxes with the highest scores are obtained through screening and non-maximum suppression.
[0044] Next, features corresponding to the candidate boxes are screened from the feature maps, and then pooling operations are performed on these features to make their sizes meet the expectations. ROI pooling has a preset width and height, indicating that each proposal feature should be unified into a feature map of such a size. When processing, first map the coordinates of the preselected boxes based on the M*N scale back to the (M / 16)*(N / 16) scale. Then divide the corresponding area of each preselected box into grids according to the preset size. Then perform max pooling on each part of the grid, and after processing, output. The output vectors have the same size.
[0045] Finally, based on the generated multi-dimensional feature vectors, object recognition, classification, and prediction of the position offset of the bounding boxes are completed for the candidate boxes. That is, the specific categories of all preselected boxes are classified through the fully connected layer and softmax. Usually, there are multiple categories, and the bounding boxes of the preselected boxes are regressed to obtain a final classification with higher accuracy.
[0046] When performing multi-scale feature image fusion, the traditional approach uses a top-down U-shaped structure as Figure 1 shown. Among them, 8s, 16s, and 32s represent feature maps of three different scales. The scale of 8s is 64*64*128, the scale of 16s is 32*32*256, and the scale of 32s is 16*16*512. Among them, 128, 256, and 512 represent the corresponding number of channels. That is, first, the feature map of 32s is transformed into 32*32*256 and then fused with the feature map of 16s. Then, the fusion result is transformed into 64*64*128, and finally, it is fused with the feature map of 8s to obtain a fused feature map of 64*64*128. The feature maps formed by the deep network lack sufficient spatial position information for the recognition of small and weak targets. Therefore, it is necessary to use the feature information of the shallow network for supplementation. However, the traditional connection structure is relatively simple, and the feature information of the shallower layer participates less in aggregation, which is not enough to complete the supplementation of the spatial position information of small and weak targets. At the same time, when performing fusion, the information abstraction degree of the shallow features is low. In summary, for the detection problem of small and weak targets, it is necessary to improve the multi-scale feature fusion method of the U-shaped structure.
[0047] As Figure 2 shown, the embodiment of the present invention proposes a stepped structure feature fusion method, and the specific process is as follows:
[0048] Step 1: Along the direction of indication ①, use a 1*1 convolution to reduce the number of channels of the 32s feature map by half, and then upsample it by a factor of two to make its size become 32*32*256. Then, along the direction of indication ②, splice it with the 16s feature map in channels to obtain a feature map of 32*32*512, and then perform a 1*1 convolution to halve the number of channels and change the size to 32*32*256.
[0049] Step 2: Along the direction of indication ③ and the direction of indication ④, fuse the 8s feature map and the feature map formed by downsampling the 16s feature map to obtain a new 8s feature map. The new 16s feature map obtained through the direction of indication ② is fused with the 8s feature map in the direction of indication ⑥ along the direction of indication ⑤ to obtain a new 8s feature map.
[0050] Step 3: Along the direction of indication ⑦, perform a global average pooling operation on the 32s feature map, and then along the direction of indication ⑧, perform a vector expansion on the feature map after the pooling operation to generate an 8s feature map. Fuse the 8s feature map generated after the vector expansion with the 8s feature map generated in Step 2 to obtain the final feature map, whose size is 64*64*128.
[0051] The feature fusion method of the present invention is from deep to shallow, and finally outputs a feature map of the expected size. The biggest feature of the fusion method proposed by the present invention is that on the basis of the U-shaped structure, the fusion of features between two adjacent layers is added. Such a non-linear design makes the feature fusion more sufficient, the features of the generated feature map are richer, and it can better represent the complete information of the picture. At the same time, the present invention also performs global average pooling on the initial shallow feature map (16*16*512), then performs vector expansion, and adds it to the fusion process. The calculation of this step is relatively simple, but by enhancing the global information, the receptive field is improved.
[0052] In order to reduce the model calculation amount, the present invention selects the ShuffleNetV2 lightweight convolution module. ShullfeNetV2 is improved on the basis of ShuffleNetV1. ShuffleNetV1 proposed grouped convolution and channel shuffle algorithm to optimize the convolutional neural network module, so as to reduce the amount of addition, subtraction, multiplication and division operations required during convolution.
[0053] The principle of grouped convolution used in ShuffkeNetV1 is similar to that of depth convolution. The depth convolution algorithm is to design separate convolution kernels and perform convolution on each feature channel. The grouped convolution algorithm is to divide the channels of the feature map according to the set value, divide the channels into multiple groups, and then let the convolution kernel process the features of each group. When the number of channels in each group is set to 1, the effects of the two algorithms are equivalent.
[0054] ShuffleNetV2 mainly increases the ratio of the algorithm calculation operation to the memory access operation, and improves the parallel ability of the model, specifically reflected in:
[0055] (1) Add a channel split at the beginning to divide the input picture feature channels into two groups, and cancel the subsequent grouped convolution operation.
[0056] (2) Replace the element-wise addition (adding the corresponding feature maps) with channel concatenation.
[0057] (3) Move the channel shuffle operation to the back of the channel concatenation and merge it with the channel split.
[0058] Compared with other common lightweight convolution modules, ShuffleNetV2 not only has better running performance but also achieves better accuracy in the ImageNet field research, although it is slightly inferior to MobileNetV2 in terms of computational complexity. Since the processing task of the present invention is to improve the detection accuracy of weak and small targets in electro-optical imaging under complex environments, the present invention uses it to lightweight the convolution module and complete the algorithm optimization.
[0059] To address the problem of gradient vanishing for weak and small targets, the present invention designs the same-scale residual module of the present invention. The main improvement of this module is to fuse the feature maps at the same scale through the method of residual module and channel splicing. By this method, the problems of gradient vanishing and model degradation of the algorithm model are alleviated, and the richness of the extracted features is enhanced.
[0060] The specific structure of the residual module is as Figure 3 shown. The input feature map first passes through a basic convolution module with a stride of 1 to obtain the output feature map of the basic convolution module. Then, the output feature map of the basic convolution module is concatenated with the input feature map in channels to obtain the output feature map of the residual module. The feature fusion structure of the same-scale residual module combines the dense connection network and the residual network. It can combine the features of shallow and deep convolutional layers in a dense connection manner, thereby enhancing the richness of the extracted features. At the same time, due to the existence of the residual module, the backpropagation can be effectively improved, and the problem of gradient vanishing occurring in the deeper convolution process can be alleviated to a certain extent. The structure of the same-scale residual module is as Figure 4 shown. Its input is the generated feature map; the rounded rectangle is the residual module; the intersection of multiple arrows represents the channel-spliced feature map, and the number of channels of the feature map is kept unchanged by means of convolution kernel compression. That is, the feature map first passes through the residual module, and then the feature map is concatenated with the output of the residual module in channels. This design can efficiently complete tasks such as feature learning, feature reuse, and feature selection through the residual structure and channel splicing, and can increase the number of channels while reducing the size of the feature map, maintaining the richness and diversity of features.
[0061] For small targets in complex environments, such as small targets on the ground, since the ground targets are relatively small, in order to preserve the original features of the ground small targets as much as possible, in the embodiments of the present invention, a relatively large pixel size of 512*512 can be used as the input size of the picture. The network structure uses a residual module layer that incorporates the ideas of residual modules and channel splicing to obtain feature maps, and finally uses a multi-scale fusion layer to fuse the different-sized feature maps generated by different dense fusion layers. At the same time, in order to increase the receptive field and reduce the computational amount, the picture is continuously downsampled during the transmission in the convolutional network, and the size of the feature map gradually decreases. Therefore, in order to reduce the loss of features during this process and maintain the richness and diversity of features, each time downsampling is performed, the size of the feature map is reduced by half, and the number of channels is doubled.
[0062] See Figure 5 , in the embodiments of the present invention, the implementation of the intelligent detection method for small and weak targets in optoelectronic imaging in complex environments includes the following steps:
[0063] Step S1: Use the normalization processing method for the picture to be predicted to unify the input size. In this embodiment, the input size of the picture is 512*512 pixels.
[0064] Step S2: For the input picture, first use a 5*5 depth convolution to perform convolution on the 3 input channels (RGB three channels) in groups to obtain a feature map of 512*512*3. Using a large-size convolution kernel in this step is to obtain a relatively large receptive field, and using grouped convolution is to reduce the computational amount.
[0065] Step S3: Perform the first downsampling through a 3*3 convolution with a stride of 2 to obtain a feature map of 256*256*32. A conventional 3*3 convolution is used in this step because, as the most commonly used convolution kernel size, it can effectively extract features while increasing the receptive field with relatively less computational amount and higher computational efficiency. Adding this convolution layer in the shallow layer of the network can ensure the quality of feature extraction by the network and improve the stability of the network.
[0066] Step S4: Mainly obtain feature maps of different scales through the residual module layer (4 layers). This module layer mainly uses the lightweight convolution structure in ShuffleNet2 as the basic convolution module, and at the same time incorporates the ideas of residual modules and channel splicing.
[0067] Step S5: Use the stepped feature fusion method proposed by the present invention to perform feature fusion on the obtained multi-scale feature maps, and finally obtain a feature map of 64*64*128.
[0068] Step S6: On the finally obtained feature map, use the classification detection discriminator for detection and classification.
[0069] In view of the problem that the feature information of small and weak targets is prone to be missing in the deep network due to their small size, the present invention supplements the feature information of the shallow network through feature fusion, thereby improving the detection accuracy of small and weak targets.
[0070] To further verify the detection performance of the method of the present invention, detection performance analysis was carried out on the VISDRONE dataset and the VEDAI dataset. Among them, the VISDRONE dataset contains more than six thousand training datasets, more than five hundred validation datasets, and more than one thousand test datasets. The dataset mainly collects pictures containing people and vehicles from the perspective of drones, and there are a total of ten target categories. The targets in the dataset are relatively concentrated and the environment is relatively complex. The VEDAI dataset is an aerial image obtained by satellite with extremely small target sizes and relatively complex backgrounds. The size of all the pictures in the dataset is 1024*1024, and it contains one thousand training datasets, more than eight hundred validation datasets and one hundred test datasets. The detection targets of the images are 11 different types of vehicles, and the image backgrounds are relatively rich, covering vehicle targets in scenarios such as residential areas, urban streets, and highways. Due to the relatively large original size of the pictures, the target sizes are extremely small from the satellite perspective, making the detection difficult. Most of the target sizes in the VEDAI dataset and the VISDRONE dataset are below 20*20. There is almost no situation where the size of the target to be detected exceeds 60*60 in the dataset, and the dataset as a whole meets the requirements of small and weak targets. At the same time, both of these datasets have problems of target occlusion and target overlap, and the backgrounds are relatively complex, meeting the requirements of the task of the present invention.
[0071] The comparison results of the accuracy of the method of the present invention with the Faster RCNN and SSD algorithm models are shown in Table 1. The AP evaluation index of the present invention (Average Precision, which is the average accuracy rate for the dataset, calculated as the area under the Precision-recall curve and used to measure the quality of the trained object detection model for each target category) is the average value of the mAP (average of APs for multiple categories) at IoU (Intersection over Union) thresholds of 0.50:0.05:0.95. It can be seen from the data in the table that the present invention has a relatively good improvement in detection accuracy, and this result proves the effectiveness and efficiency of the present invention.
[0072] Table 1 Comparison table of the accuracy of algorithm models
[0073] model VEDAI VISDRONE Faster RCNN 16.7 8.6 SSD512 22.8 9.1 the present invention 23.4 12.2
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0075] The above are only some embodiments of the present invention. For those of ordinary skill in the art, without departing from the inventive concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. An intelligent detection method for weak and small targets in optoelectronic imaging in complex environments, characterized in that, It includes the following steps: Step S1: Perform size normalization on the image to be detected to obtain the desired image size; Step S2: Perform grouped convolution on the image processed in Step S1 using a depth convolution with a convolution kernel size of 5×5 to obtain a first feature map; Step S3: Perform the first downsampling through a convolution with a stride of 2 and a convolution kernel size of 3×3 to obtain a second feature map; Step S4: Perform multi-scale feature extraction on the second feature map through four stacked residual module layers; The residual module layer is specifically: the second feature map generated in Step S3 is used as the input feature map of the first residual module layer, and the input feature maps of the second, third, and fourth residual module layers are the output feature maps of the previous residual module layer; that is, the output feature map of the previous layer is used as the input feature map of the next layer; The network structure of each layer of the residual module layer is the same, and the lightweight convolution structure ShuffleNet2 is used as the basic convolution module. After the input feature map of the residual module layer passes through the basic convolution module, it is then concatenated with the input feature map of the current residual module layer by channels to obtain the output feature map of the current residual module layer; Step S5: Perform multi-scale fusion processing on the output feature maps of the second, third, and fourth residual module layers using a stepped structure feature fusion method to obtain a fused feature map; Among them, the stepped structure feature fusion method is: Define the output feature maps of the second, third, and fourth residual module layers as the first-scale feature map, the second-scale feature map, and the third-scale feature map respectively, and define the feature map dimensions of the output feature maps of the second, third, and fourth residual module layers as the first scale, the second scale, and the third scale respectively; Transform the feature map dimension of the third-scale feature map into the second scale, and then perform channel concatenation with the second-scale feature map to obtain a first concatenation result; and transform the feature map dimension of the first concatenation result into the second scale to obtain the first concatenation result of the second scale; Transform the feature map dimension of the second-scale feature map into the first scale, and then perform channel concatenation with the first-scale feature map to obtain a second concatenation result; and transform the feature map dimension of the second concatenation result into the first scale to obtain the second concatenation result of the first scale; Transform the feature map dimension of the first concatenation result of the second scale into the first scale, and then perform channel concatenation with the second concatenation result of the first scale to obtain a third concatenation result, and transform the feature map dimension of the third concatenation result into the first scale to obtain the third concatenation result of the first scale; Perform global average pooling operation on the third-scale feature map, and transform the feature map dimension of the global average pooling operation result into the first scale, and then perform feature fusion with the third concatenation result of the first scale to obtain a fused feature map; Step S6: Perform small target detection processing on the fused feature map through a classification detection discriminator to obtain the detection processing result of the small target; The small target refers to a target whose size is less than or equal to the specified size.
2. The method according to claim 1, characterized in that, Step S6 includes: Step S601, extract candidate boxes from the fused feature map, and obtain multiple final candidate boxes after screening and non-maximum suppression processing on the extracted candidate boxes; Step S602: Based on each of the finally obtained candidate bounding boxes in Step S601, extract the features corresponding to each of the finally obtained candidate bounding boxes from the fused feature map, and perform a pooling operation on the features of the candidate bounding boxes to obtain candidate bounding box features that meet the expected feature size; Step S603: Based on the classification and detection discriminator, perform small target detection on the candidate bounding box features.
3. The method according to claim 1, characterized in that The classification and detection discriminator is composed of a fully connected layer and a softmax function layer.
4. The method according to claim 1, wherein In Step S1, the expected image size is: 512×512.
Citation Information
Patent Citations
Target detection method
CN112232232A
Remote sensing image target detection method based on multi-scale feature fusion and feature enhancement
CN114708511A