Method for object detection and recognition based on convolutional neural network
By improving the VGG16 network and introducing attention mechanism methods, the Faster R-CNN algorithm's shortcomings in small object detection are solved, and more efficient feature extraction and detection accuracy are achieved.
Patent Information
- Application Number
- CN202310060521.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-01-13
AI Technical Summary
Faster R-CNN algorithm lacks performance in small target objects, and traditional candidate area algorithms lead to increased operation time and insufficient real-time and accuracy.
By improving the backbone network VGG16, the attention mechanism is introduced, the attention map is generated and the pooling operation is performed, and the object detection and recognition is performed in combination with border regression.
It improves the detection accuracy and speed of small target objects, enhances the robustness of feature extraction, retains important information, and reduces unnecessary interference.
Smart Images

Figure CN116152510B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing applications, and particularly relates to a method for target detection and recognition based on a convolutional neural network. Background Art
[0002] A convolutional neural network is a class of feedforward neural networks that contain convolutional computations and have a deep structure, and is one of the representative algorithms of deep learning. A convolutional neural network has the ability of feature learning and can perform translation-invariant classification on input information according to its hierarchical structure, so it is also called a "translation-invariant artificial neural network". With the proposal of the deep learning theory and the improvement of numerical computing devices, convolutional neural networks have developed rapidly and have been applied to fields such as computer vision and natural language processing.
[0003] In the field of target detection, starting from R-CNN, many breakthroughs have been made by introducing convolutional neural networks, but it has never been able to get rid of the limitations of traditional candidate region algorithms (such as Selective Search). Using the Selective Search algorithm to determine candidate regions greatly increases the operation time of the Fast R-CNN algorithm, making the Fast R-CNN network structure model unable to meet the requirements in terms of real-time performance. To solve the bottleneck of candidate region extraction and further share convolutional operations, Ren Shaoqing et al. proposed Faster R-CNN in 2016, which effectively solved the above problems.
[0004] The Faster R-CNN network consists of two modules, the RPN network and the Fast R-CNN network. The RPN network mainly predicts and extracts possible target candidate regions and performs preliminary localization of the target. The Fast R-CNN network maps the region proposal boxes to the feature map, performs class judgment and regression correction on the regions of interest, so as to distinguish the target from the background and achieve the purpose of screening and refining the target region box. The two modules share the feature map and convolutional weights extracted from the whole image, and predict the target boundary and score at each position. This network model integrates generating candidate regions, classification and regression into a network framework, has invariance to scale and illumination, etc., and the detection accuracy of the algorithm is relatively ideal. This network model replaces the selective search method with a region proposal network, accurately extracts candidate region features on the feature map, realizes true end-to-end computing, shortens the candidate box extraction time, greatly improves the detection speed and accuracy, and realizes fast real-time target detection training and testing. The Faster R-CNN algorithm directly performs subsequent operations such as predicting the target box on the last feature map of the feature extraction network, ignoring the relevance of the semantic information before and after the target, and lacking in the detection performance of small target objects. Summary of the Invention
[0005] In view of the technical problem that the above Faster R-CNN algorithm lacks detection performance for small target objects, the present invention proposes a method for target detection and recognition based on a convolutional neural network, which is simple in method, convenient to operate, and can improve the detection performance for small target objects.
[0006] To achieve the above object, the technical solution adopted by the present invention is that the present invention provides a method for target detection and recognition based on a convolutional neural network, including the following steps:
[0007] a. First, input the image to be recognized into the improved backbone network VGG16 for feature extraction to obtain the corresponding feature map;
[0008] b. Input the obtained feature map into the RPN network and the attention mechanism respectively. Among them, the feature map is input into the RPN network to select regions of interest and generate candidate regions; the feature map is input into the attention mechanism for weighted operation to generate an attention map;
[0009] c. Then, input the output results of the RPN network and the attention mechanism into the RoI pooling layer, and then perform pooling operations on the candidate boxes of the feature map;
[0010] d. Determine and classify the categories of feature information of the feature map after pooling operation in the fully connected layer;
[0011] e. Finally, use bounding box regression to refine and adjust the positions of the candidate boxes again to achieve target detection and recognition;
[0012] Among them, the operation method of the improved backbone network VGG16 is as follows:
[0013] a1. First, use a convolutional kernel of 1×1×256 to change the number of channels of the shallow features to 256 and also change the number of channels of the deep features to 256, and then perform deconvolution operation;
[0014] a2. Then, after the deconvolution operation, use an additive fusion function to perform feature layer fusion on the shallow features and the deep features to obtain the fused feature layer;
[0015] a3. Pass the fused feature layer through a 3×3 convolutional kernel to generate a new feature layer, and then obtain the corresponding feature map.
[0016] Preferably, in the step a2, batch normalization processing is added to each subsequent step of each convolutional layer, and a summation operation is performed on modules with the same scale to obtain the fused feature layer.
[0017] Preferably, in the step c, the attention mechanism is used as follows: First, the feature map is compressed using the compression formula, and the feature values of each channel are added up and averaged to make the number of channels of the input feature map consistent with the output dimension.
[0018] Preferably, the compression formula is:
[0019]
[0020] where z c represents the value after compression of the c-th channel, u c is the c-th variable of the transformation output U, U ∈ {u1, u2,..., u c}, and H×W represents the spatial dimension.
[0021] Compared with the prior art, the advantages and positive effects of the present invention are that
[0022] The present invention provides a method for target detection and recognition based on a convolutional neural network. By improving the backbone network VGG16, the features of different layers are fused, so that the feature maps at each scale contain rich semantic information. The forward and backward transmitted features are superimposed to enable each convolutional layer to obtain feature maps with different characteristics, enhancing the robustness of the entire feature extraction process. At the same time, the attention mechanism is introduced in the RPN network stage to retain the important information after fusion and discard the interference of redundant information, improving the accuracy of the entire algorithm and the feature extraction effect for small target objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 is the structure diagram of the VGG16 network;
[0025] Figure 2 is the improved feature fusion method;
[0026] Figure 3 is the way of feature fusion;
[0027] Figure 4 The application of the attention mechanism in the convolutional neural network;
[0028] Figure 5 is the target detection and recognition effect diagram of the traditional Faster R-CNN algorithm;
[0029] Figure 6 The effect diagram of the method target detection and recognition provided for Example 1. Specific implementation manners
[0030] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described below in conjunction with the accompanying drawings and embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0031] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the present invention is not limited by the limitations of the specific embodiments disclosed in the following specification.
[0032] Example 1, this embodiment provides a method for target detection and recognition based on a convolutional neural network. The specific operation steps are as follows:
[0033] First, the graph to be recognized is input into the improved backbone network VGG16 for feature extraction to obtain the corresponding feature map. As Figure 1 shown, the VGG network improves the learning ability of the entire network by repeatedly stacking convolutional kernels. The size of the convolutional kernel in the VGG16 network is 3×3, which is the smallest scale convolutional kernel that can extract features. Among them, the stride and padding are both 1.
[0034] In this embodiment, as Figure 2 shown, taking the feature fusion of conv2_2 and conv4_3 as an example, the improvement content of the backbone network VGG16 will be specifically described. First, use a convolutional kernel of 1×1×256 to change the number of channels of conv2_2 to 256 and the number of channels of conv4_3 to 256 as well. Then, perform a deconvolution operation, so that conv2_2 and conv4_3 have the same size when performing feature fusion. The deconvolution upsamples the feature map with a low resolution to a feature map with a higher resolution for feature fusion, and more abundant semantic information is obtained after fusion. After the deconvolution operation, use an additive fusion function to perform feature layer fusion on conv2_2 and conv4_3. The fused feature layer then passes through a 3×3 convolutional kernel to generate the P2 feature layer, effectively avoiding the problem of information aliasing caused by upsampling. The fusion methods of the remaining feature layers are the same. After feature fusion, new feature layers P2, P3, P4, and P5 are formed, fusing shallow features and deep features to generate new feature layers.
[0035] In this embodiment, as Figure 3As shown, deconvolution integrates the shallow feature map with the feature information of the deconvolution layer, which is the deconvolution feature fusion method in this paper. Batch normalization is added after each convolutional layer, and the modules with the same scale are summed to better utilize the feature information of each layer.
[0036] Then, the fused feature map needs to be input into the RPN network. In this embodiment, in order to retain the important information after fusion and eliminate the interference of redundant information. As Figure 5 shown, the feature map obtained by processing through the feature extraction network is respectively input into the RPN and the attention mechanism. The RPN is used to select the regions of interest to generate candidate regions; it is input into the attention mechanism for weighted operation to generate an attention map, and then the results of both parts are input into the RoI pooling layer for subsequent processing.
[0037] The method of using the attention mechanism is to first compress the feature map using the compression formula, sum the feature values of each channel and take the average value, that is, perform global average pooling. During the entire pooling process, the number of channels of the input feature map is the same as the output dimension. The compression formula is:
[0038]
[0039] where z c represents the value after compression of the c-th channel, u c is the c-th variable of the transformation output U, U ∈ {u1, u2,..., u c}, and H×W represents the spatial dimension.
[0040] In this way, the output results of both the RPN network and the attention mechanism are input into the RoI pooling layer, and then the pooling operation is performed on the candidate boxes of the feature map. Next, the pooled feature map is used to determine and classify the category of the feature information in the fully connected layer; finally, the position of the candidate box is refined and adjusted again using bounding box regression, and the object detection and recognition can be achieved.
[0041] In this way, by improving the backbone network VGG16, the features of different layers are fused, so that the feature maps at each scale contain rich semantic information. The forward and backward transfer features are used for superposition, so that the feature maps with different characteristics are obtained by each convolutional layer, and the robustness of the entire feature extraction process is enhanced. The present invention fuses the shallow features and deep features in the network without increasing the time required for the detection process. First, the features are selectively fused multiple times, then the fused features are hierarchically predicted, and finally all the prediction results are fused. By fusing the shallow features and deep features of the network, deconvolution is used to perform upsampling on the deep convolution. After obtaining the feature maps of the same size, additive operations are performed on the corresponding channels to obtain multi-scale feature layers. The dimension size of each layer output is 256. The RPN is used to predict the regions of interest, and the accuracy of the entire algorithm is improved, and the feature extraction effect of small target objects becomes better.
[0042] Method detection: As Figure 5 , Figure 6 shown, where Figure 5 indicates that before the algorithm improvement, pedestrians blocked by the green belt could not be detected; Figure 6 After adding the improved algorithm of feature fusion provided in this embodiment, the occluded pedestrian targets that could not be detected in the original algorithm can be detected. The Faster R-CNN object detection algorithm after feature fusion makes more full use of the feature layer information, making the discrimination of the RPN for foreground targets more accurate. From the comparison between Figure 5 and Figure 6 , it can be found that the improved algorithm has better effect than the original object detection algorithm. The improved object detection algorithm fuses high-level semantic information and low-level detail information, so that the small target feature information can also be better retained, so that more small targets can be detected. Therefore, the improved feature fusion detection algorithm can effectively improve the detection accuracy of small targets.
[0043] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.
Claims
1. A method for object detection and recognition based on a convolutional neural network, characterized in that, It includes the following steps: a. First, input the image to be recognized into the improved backbone network VGG16 for feature extraction to obtain the corresponding feature map; b. Input the obtained feature map into the RPN network and the attention mechanism respectively. Among them, the feature map is input into the RPN network to select the region of interest and generate candidate regions; the feature map is input into the attention mechanism for weighted operation to generate an attention map; c. Then, input the output results of the RPN network and the attention mechanism into the RoI pooling layer, and then perform pooling operations on the candidate boxes of the feature map; d. Determine and classify the category of feature information for the feature map after pooling operation in the fully connected layer; e. Finally, use bounding box regression to refine and adjust the position of the candidate box again to achieve object detection and recognition; Among them, the operation method of the improved backbone network VGG16 is: a1. First, use a convolutional kernel of 1×1×256 to change the number of channels of the shallow features to 256 and the number of channels of the deep features to 256, and then perform deconvolution operation; a2. Then, after the deconvolution operation, use an additive fusion function to perform feature layer fusion on the shallow features and the deep features to obtain the fused feature layer; a3. Pass the fused feature layer through a 3×3 convolutional kernel to generate a new feature layer, and then obtain the corresponding feature map.
2. The method for object detection and recognition based on a convolutional neural network according to claim 1, characterized in that In step a2, batch normalization processing is added to each subsequent step of each convolutional layer, and a summation operation is performed on modules with the same scale to obtain the fused feature layer.
3. The method for target detection and recognition based on a convolutional neural network according to claim 2, characterized in that, In step c, the usage method of the attention mechanism is: first compress the feature map using the compression formula, add up the feature values of each channel, take the average value, and make the number of channels of the input feature map consistent with the output dimension.
4. The method for target detection and recognition based on a convolutional neural network according to claim 3, characterized in that The compression formula is: where z c represents the compressed value of the c-th channel, and u c is the c-th variable of the transformed output U, U ∈ {u1, u2,..., u c}, and H×W represents the spatial dimension.
Citation Information
Patent Citations
Target detection method, target detection model and target detection system based on cascade detector
CN109886286A
Training method based on image-instance alignment network and cross-domain target detection method
CN114693983A