A cascaded regression object detection method, device, and computer-readable storage medium
Through the cascading regression object detection method, combined with the cascading area suggestion module and the two-way regressor module, the two-stage object detection algorithm is optimized, which solves the problems of detection accuracy and time consumption, and realizes efficient object detection.
Patent Information
- Application Number
- CN202111092255.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-09-17
AI Technical Summary
The existing two-stage object detection algorithms are difficult to meet actual needs in terms of performance and computing resource consumption, and it is difficult to reduce detection time while maintaining high detection accuracy.
The cascading regression object detection method is adopted, and the cascading area suggestion module and two-way regressor module are combined with the distance loss function and the loss function to optimize the detection process, reduce the algorithm detection time, and maintain high accuracy.
While maintaining high detection accuracy, it significantly reduces detection time and algorithm resource consumption, and improves the performance of the detector.
Smart Images

Figure CN114241250B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and particularly to a cascade regression object detection method, apparatus and computer-readable storage medium. Background Art
[0002] In recent years, with the improvement of computer computing performance, the development of artificial intelligence has not only made great progress, but also gradually shown a phenomenon of large-scale application in all walks of life. As an important research field of artificial intelligence, computer vision has a huge promoting effect on the innovation of production and lifestyle. For example, driverless cars representing future travel modes, intelligent manufacturing that improves factory production efficiency, intelligent security for early warning, etc., computer vision is the technical foundation for their realization.
[0003] In image classification, object detection, semantic segmentation, and instance segmentation are three basic visual recognition tasks in the field of computer vision. Object detection not only needs to identify the object category, but also needs to predict the position of the object using a rectangular box. Object detection is a basic module that must be used in many scenarios, such as face recognition, pedestrian detection, video analysis, and sign detection. Semantic segmentation assigns specific class labels to each pixel, thus providing a more informative image description. Different from object detection, semantic segmentation does not distinguish multiple objects of the same class. Instance segmentation is a combination of object detection and semantic segmentation, which needs to identify different objects and assign a per-pixel class mask to each object. In fact, instance segmentation can be considered a special case of object detection. It does not use a rectangular box to locate an object, but uses per-pixel localization.
[0004] Object detection algorithms are generally divided into two-stage algorithms and one-stage algorithms. A two-stage object detector usually consists of two parts. First is the candidate region generator. During the process of generating candidate regions, the generator tries to find regions in the image where objects may exist. The main purpose is to find regions with a higher recall rate, so that all objects in the image can be matched with a candidate region as much as possible. Then is the detector that classifies and regresses the candidate regions. Two-stage object detectors have the advantage of high detection accuracy, but often cannot meet the requirements of real-time detection.
[0005] One-stage object detectors are usually simpler in architecture. Different from the two-stage object detection algorithm that divides the detection process into two stages: candidate box generation and region classification, it considers each pixel on the feature map as potentially having an object, and tries to assign a rectangular box and multiple confidence levels to this pixel respectively to determine the accurate position and probability value of the object. One-stage object detectors have a high detection speed, but often come with poor detection accuracy.
[0006] Generally, the precision requirement is greater than the speed requirement. Therefore, the two-stage object detection algorithm has a wider applicability. However, the existing two-stage object detection algorithms have the following deficiencies: the algorithms with higher performance often require high computing power resources and detection time, making it difficult to meet the actual requirements. Therefore, there is an urgent need for a lightweight and high-performance two-stage object detection method. Summary of the Invention
[0007] Aiming at the problem of poor detection performance in the prior art, the present invention provides a cascade regression object detection method, device and computer-readable storage medium, which improves the performance of the two-stage object detector through the idea of cascade regression, and reduces the detection time of the algorithm while maintaining high detection accuracy.
[0008] The following are the technical solutions of the present invention.
[0009] A cascade regression object detection method includes the following steps:
[0010] Obtain the image to be detected, perform pixel standardization and scale it to the same size;
[0011] Input the adjusted image to be detected into a trained deep convolutional neural network model, which is trained by training images with annotation information and includes a backbone network for feature extraction, a cascaded region proposal module for step-by-step adjustment of the preset box, a cascaded two-way regressor module for fine-tuning the preset box, a distance loss function and a loss function for optimizing the above modules;
[0012] Use the trained deep convolutional neural network model to detect the image to be detected and obtain the detection result.
[0013] Preferably, the backbone network extracts features of the image through several feature prediction layers with different resolutions; the backbone network includes one of the ResNet series networks or VGG series networks, and an additional convolutional network added on this basis.
[0014] Preferably, the cascaded region proposal module classifies and regresses according to the extracted features, and includes a cascaded region proposal network, a feature fusion module and a classification and regression module; the cascaded region proposal network fine-tunes the size and position of the preset box in two steps to generate region proposals; the feature fusion module fuses features of different scales together; the classification and regression module extracts the features of the candidate region N×N according to the region proposals provided by the cascaded region proposal network for classification and regression.
[0015] Preferably, the cascaded two-way regressor module uses neural networks with a fully connected architecture and a fully convolutional architecture as regression branches respectively to regress the same candidate region and adjust the size and position of the candidate region, including:
[0016] The first stage: Using the position of the preset box as the candidate region, extracting the target features and cropping them to the specified size, inputting them into two regression branches respectively for prediction, adjusting the size and position of the candidate region according to the prediction results and extracting two sets of features, and using a convolutional neural network to fuse the two sets of features;
[0017] The second stage: Inputting the features fused in the first stage into two regression branches respectively for prediction, adjusting the size and position of the candidate region according to the prediction results and extracting two sets of features, and using a convolutional neural network to fuse the two sets of features;
[0018] The third stage: Inputting the features fused in the second stage into two regression branches respectively for prediction, adjusting the size and position of the candidate region according to the prediction results and extracting two sets of features, and using a convolutional neural network to fuse the two sets of features to obtain the size and position of the finally predicted candidate region.
[0019] Preferably, in each stage of regression, the distance loss function optimizes the classification using the cross-entropy function and optimizes the regression using the Smooth L1 function to minimize the output result gap between the two regression branches.
[0020] Preferably, the definition of the distance loss function includes:
[0021]
[0022] Among them, the Smooth L1 function is as follows:
[0023]
[0024] Among them, x t is the feature information fused according to the coordinates predicted by the two regression branches in the previous stage, f b represents the regression branch of the fully connected layer architecture, f d represents the regression branch of the fully convolutional architecture, and i represents the center coordinates, width and height of the i-th candidate region.
[0025] Preferably, the loss function includes: the cascaded region proposal module loss function and the cascaded two-way regressor loss function:
[0026]
[0027]
[0028] Among them, Loss1 is the cascaded region proposal module loss function, Loss2 is the loss function of the cascaded two-way regressor, where i represents the subscript of the preset anchor box, and respectively represent the true category and the position offset vector of the preset anchor box with subscript i, c i and t i are the multi-class prediction probability and the coordinate detection result in the second stage. N1, N2, and N3 respectively represent the number of positive samples in the detection processes of the first stage, the second stage, and the third stage of the cascaded regression. L m is the multi-class cross-entropy loss for judging the object category, L r is the Smooth-L1 loss function; the total loss Loss is the weighted sum of the loss function of the cascaded region proposal module and the loss function of the cascaded two-way regressor.
[0029] Preferably, the detection of the image to be detected by using the trained network model includes:
[0030] Input the trained network model for detection using Q test images;
[0031] Save the detection results R = {R1, R2, …, R q , …, R Q} by category;
[0032] Calculate the intersection over union of the rotated rectangular boxes, perform non-maximum suppression, and only retain the detection boxes with larger scores and small overlapping areas as the final detection results.
[0033] Preferably, the non-maximum suppression includes:
[0034] Re-sort the prediction scores of each detection box in the same category in the initial detection result R q in descending order. The sorted result is R' q = {R' c1 , R' c2 , …, R' cf , …, R' cF}, where R' cf is the detection result on the j-th class after sorting;
[0035] For any detection box b in R' cf , calculate the intersection over union between it and all detection boxes with prediction scores less than the current score;
[0036] If the intersection over union of two detection boxes exceeds the threshold t iou , then discard the detection box bs with a lower score.
[0037] The present invention also includes a cascaded regression object detection device, including:
[0038] An acquisition module for acquiring the image to be detected and / or the training image;
[0039] A preprocessing module for performing pixel standardization and scaling operations on the detected images and / or training images obtained by the acquisition module; a neural network module for configuring a deep convolutional neural network, including a backbone network for feature extraction, a cascaded region proposal module for stepwise adjusting the preset boxes, a cascaded two-way regressor module for fine-tuning the preset boxes, a distance loss function and a loss function for optimizing the above modules;
[0040] The neural network module is further configured to perform training by means of the training images and detect the images to be detected to obtain detection results.
[0041] The present invention further includes a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the above method. The present invention uses a popular convolutional neural network algorithm to solve the object detection problem. In the cascaded region proposal module, a strategy of stepwise adjusting the position and size of the preset boxes is adopted to provide more accurate initialization information for the detection probe, thereby improving the performance of the detector. The cascaded region proposal network can provide more accurate preset boxes for the detector by stepwise fine-tuning the size and position of the preset boxes, avoiding overfitting. Since the size and position of the preset boxes are fine-tuned on the global feature map, all calculations for all preset boxes share the global feature map, avoiding a large amount of repetition and resource consumption caused by using the cascaded strategy at the detection head.
[0042] In the cascaded two-way regressor module, each branch adopts a different architecture to predict the position of the target, and then extracts the target region features according to the positions predicted by the two architecture branches. After convolution fusion, the features are sent to the next stage for classification and regression. During training, the fully connected and fully convolutional heads are respectively used to regress the same target. The fully connected regressor has stronger spatial sensitivity, and the fully convolutional regressor can capture more refined target-level information. The advantages of the two architectures can be fully utilized, and complementary functions can be achieved.
[0043] This application adopts the gradient backpropagation algorithm and, with the help of the loss function calculated by the neural network, calculates the update gradients of all learnable parameters in the neural network through the chain rule during the neural network training process, thereby completing the update of the neural network parameters and realizing the end-to-end training process.
[0044] According to the principle of digital image processing, this application performs various data augmentations on the training images, including image flipping, color space conversion, image scaling, etc., improving the utilization rate of the training images, increasing the diversity of samples, reducing the need for data annotation to a certain extent, and enhancing the robustness and generalization ability of the model.
[0045] This application uses non-maximum suppression as a post-processing method for the detection results in the image to be detected, effectively reducing redundant detection results in the image. Description of the Drawings
[0046] Figure 1 It is a schematic diagram of the network structure according to an embodiment of the present invention;
[0047] Figure 2 It is a schematic diagram of the candidate region feature fusion according to an embodiment of the present invention;
[0048] Figure 3 It is a schematic diagram of the non-maximum suppression processing flow of the detection structure according to an embodiment of the present invention;
[0049] Figure 4 It is a schematic diagram of the training and detection process according to an embodiment of the present invention. Detailed Embodiments
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in combination with the embodiments. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0051] The terms "first", "second", "third", etc. (if any) in the specification and claims of the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein.
[0052] It should be understood that in various embodiments of the present invention, the magnitude of the sequence numbers of the various processes does not mean the order of execution, and the order of execution of the various processes should be determined by their functions and internal logics, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0053] It should be understood that in the present invention, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0054] The following will detail the technical solutions of the present invention with specific embodiments. The embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0055] Example:
[0056] A cascaded regression object detection method runs on a device with a processor and a computer-readable storage medium, including:
[0057] S1. Obtain the training set of the detection pictures;
[0058] Obtain the training set of the detection pictures. The training set includes M training pictures as X = {X1, X2, …, X m , …, X M}, where X m represents the m-th training picture;
[0059] The M training pictures are selected with one-to-one corresponding M labels as Y = {Y1, Y2, …, Y m , …, Y M}, where Y m represents the m-th training picture;
[0060] The M labels include the category and coordinate information of N target objects in the corresponding pictures as Y m = {P m,1 , B m,1 , P m,2 , B m,2 , …, P m,n , B m,n , …, P m,N , B m,N}, where P m,n represents the category of the n-th target object in the m-th picture, and P m,n ∈ {C0, C1, C2, …, C j , …, C J}, C represents the total category, C j represents the j-th category, C0 represents the background category, J is the total number of categories, B m,n represents the coordinates of the n-th object in the m-th picture, and B m,n = {w m,n , h m,n , cx m,n , cy m,n , θ m,n}, respectively representing the width w m,n of the rectangular frame of the labeled object, height h m,n , abscissa cx m,n of the center point, ordinate cy m,n of the center point, and rotation angle θ m,n .
[0061] S2. Standardize the training pictures and test pictures of the training set;
[0062] In some specific embodiments, the specific steps for image standardization of the training images and test images in the training set include:
[0063] According to the preset pixel mean Pixel mean and pixel standard deviation Pixel std , standardize the images in the training set X at the pixel level;
[0064] Uniformly scale the images in the training set X to a size of 320×320. It should be noted that after the image is scaled, the labeled positions of the objects in the image also need to be adjusted accordingly, otherwise there will be a mismatch. The image can also be scaled to 512×512 or 640×640. Higher-resolution images can improve the detection accuracy, but will reduce the detection speed. Keep the image size consistent so that the image dimensions are the same and meet the input conditions of the network.
[0065] The standardization formula for any pixel point of the image is:
[0066] Pixel x =(Pixel x -Pixel mean ) / Pixelx std ;
[0067] where Pixel mean is the pixel mean, and Pixel std is the pixel standard deviation.
[0068] S3. Design a deep convolutional neural network structure and train the network model using the method of step-by-step fine-tuning;
[0069] In some specific embodiments, the deep convolutional neural network structure includes: a backbone network for extracting image features, cascaded region proposal modules, feature fusion, a loss function for the cascaded region proposal modules, cascaded two-way regression modules, a distance loss function, and a loss function for the cascaded two-way regression modules.
[0070] Obtain the backbone for extracting basic features;
[0071] Specifically, use the ResNet series network or the VGG series network as the backbone network to extract basic features, where the ResNet network includes ResNet50, ResNet101, and ResNet152; the VGG network includes VGG-16 and VGG-19. After selecting the backbone network, it is necessary to add an additional convolutional network on the basis of the backbone network to obtain a feature map with a lower resolution, which has a smaller spatial resolution but a higher degree of feature abstraction, a larger receptive field, and can detect large objects in the picture. Therefore, 5 scales of features are extracted from the backbone network and the additionally added convolutional network for network prediction.
[0072] Initialize the parameters of the backbone network and the additionally added network:
[0073] M weight = MP weight ;
[0074] MA weight = Gaussian(0,1);
[0075] where M weight and MA weight are the parameters of the backbone network and the additional convolutional network respectively; MP represents the pre-training result of the backbone network M on the dataset, and MP weight represents the parameters of the pre-trained network; Gaussian(0,1) means that the weight parameters of the additional convolutional network MA satisfy the Gaussian distribution with a mean of 0 and a variance of 1.
[0076] In some specific embodiments, the backbone network uses ResNet50, and its pre-trained model comes from the classification model on the ImageNet dataset, and the learning rate of the first two residual structures of ResNet50 is set to 0 so that it does not participate in the training, which can reduce the risk of overfitting in the network training process.
[0077] On the basis of the backbone network and the additionally added convolutional network, first build a cascaded region proposal network to perform binary classification and position and size adjustment on the target region;
[0078] Taking VGG-16 as the backbone network as an example, take three layers in the VGG network and an additionally added layer of network as the global feature prediction layer. The resolutions of these four global feature prediction layers gradually decrease, and the receptive fields gradually increase. Therefore, they can be responsible for detecting objects of different scales respectively. Taking each pixel position on the global feature map as the center, each position is called an anchor point, and K preset boxes with fixed scales and aspect ratios are generated. Each preset box is responsible for matching with potential targets. As Figure 1As shown in the figure, the cascaded region proposal network consists of two sets of binary classification and regression modules. The first binary classification and regression module uses the original global feature map of the backbone network for prediction, performs simple binary classification on each preset box, and outputs the offset relative to the original position of the preset box. After calculating the overlap rate between the preset box and the ground truth box, some simple negative samples are discarded using a predefined threshold. Then, the cross-entropy function and the Smooth L1 function are used to calculate the deviation to optimize the first binary classification and regression. The preset box after fine-tuning the position and size by the first binary classification and regression module is used as the input of the second binary classification and regression module. Different from Cascade R-CNN, where the features output by the backbone network are used as the input in each stage, in order to obtain more discriminative information, the second binary classification and regression module uses a feature fusion block to fuse the original features to meet the needs of the second binary classification and regression module. The features used by the two binary classification and regression modules have the same dimension. The second binary classification and regression module uses the fused features to predict 4 offsets and 2 confidence scores for each fine-tuned preset box, discards simple samples according to the confidence scores, and then uses the non-maximum suppression algorithm to filter out the preset boxes with a high overlap rate. Through two-step cascaded regression operations, the region proposal network provides more accurate initialization information for the subsequent detection head.
[0079] The cascaded region proposal network serves a two-stage object detection algorithm. Since the cascaded region proposal network provides more accurate feature information for the regressor, it can improve the performance of the object detector. Different from Cascade R-CNN, in order to obtain more accurate initialization information, the regressors in all stages need to calculate the features of the candidate regions separately, consuming a large amount of computational resources. In the cascaded region proposal network, all candidate regions share the predicted feature map, requiring only a small amount of computational resources.
[0080] In the cascaded region proposal module, feature fusion is added to enhance the object perception ability;
[0081] Feature fusion is used at the connection between the first binary classification and regression and the second binary classification and regression. The feature fusion consists of a convolutional layer, an activation layer, and a deconvolutional layer, where the deconvolution is mainly responsible for enlarging the resolution of the deep feature map.
[0082] The feature fusion module is mainly used to transfer the rich semantic information of the deep network to the shallow network to improve the detection accuracy of the algorithm. Additionally, if the two binary classification and regression modules share the prediction feature map, the results of the two cascaded regressions will not bring good effects. To match the dimensions of the inputs of different network layers, deconvolution is used to increase the resolution of the feature map, and then pixel-wise summation is used for feature fusion. Finally, a convolutional layer is added to generate discriminative features. Feature fusion is also used at the connection between the second binary classification and regression module and the detection head to further enhance the information of the feature map.
[0083] In the cascaded region proposal network, binary classification of foreground and background is performed on the candidate regions, and at the same time, the size and position of the preset boxes are fine-tuned, and the network is optimized through the loss function.
[0084] The loss function of the cascaded region proposal network consists of the loss functions of the two binary classification and regression modules. In the cascaded region proposal network, a binary classification label (i.e., whether it is an object) is assigned to each preset box, and at the same time, the preset box is fine-tuned.
[0085]
[0086] Among them, i represents the index number of the preset box in a batch, l i is the category number of the real object corresponding to the preset box i, g i is the position and size of the real object corresponding to the preset box i. p i is the confidence that the preset box contains an object, x i is the position for fine-tuning the preset box i. c i is the multi-class label prediction for the preset box i, t i is the predicted offset. L b is the cross-entropy loss based on binary classification.
[0087] In the cascaded two-way regression module, the preset boxes provided by the cascaded region proposal module are used as the initial input information, and the target is accurately and precisely located through feature fusion.
[0088] The cascaded two-way regressor is mainly responsible for accurately locating the target. In the algorithm, the cascaded region proposal network initially adjusts the size and position of the candidate regions, and then uses them as the input for the next-stage two-way regressor to accurately locate the target. Specifically, the cascaded two-way regressor consists of three stages with the same function, and each stage has two regression branches with different architectures (the regression branch based on the fully connected architecture and the regression branch based on the fully convolutional architecture). The two regression branches respectively locate the same object. For example Figure 1As shown, in the first stage, according to the candidate region location information provided by the cascaded region proposal network, target features are extracted and cropped to a specified size, serving as the inputs for the two regression branches respectively. In the second stage, according to the location information predicted by the two regression branches in the previous stage for the same object, different features of the same object are extracted. The two sets of features are scaled to the same size and added pixel by pixel. Then, a convolutional layer is used to fuse their features. Since the network with a fully convolutional architecture can capture the spatial information of the target, and the network with a fully connected architecture can capture the target-level information, there are biases in the predictions of the two regression branches for the same object. By extracting the target features according to the location information predicted by the two architectures and then fusing the two sets of features, the limitations of a certain network for localization can be compensated, thereby improving the network performance. The fused features are used as the input for the network in the third stage, and the remaining operations are basically the same as those in the second stage. For the predicted values of the two regression estimators in each stage, the Smooth L1 function is used for optimization to make the predicted values of the two functions for the same target as close as possible. In each stage, the cross-entropy function is used for optimization for classification, and the Smooth L1 function is used for optimization for regression.
[0089] In the distance loss function, the distance loss function is mainly used to constrain the fine-tuning of the size and location of the same target by the two regression branches;
[0090] In the cascaded two-regressor module, the regression branches of the two architectures respectively adopt Smooth L1 to optimize the predicted offsets. Since the two regression branches predict the same candidate region, the ability of the two regression branches to work together can be considered enhanced. Therefore, the Smooth L1 function is used to optimize the offsets predicted by the two regression branches. This is because the two regression branches regress the same target, and the Smooth L1 function can make the offsets predicted by the two regression branches as close as possible, thereby improving the accuracy of the detector. The introduced optimization function is named distance loss. The expression of the distance loss is as follows:
[0091]
[0092] In the above formula, the detailed expression of the Smooth L1 function is as follows:
[0093]
[0094] where x t is the feature information fused according to the coordinates predicted by the two regression branches in the previous stage. f b represents the regression branch of the fully connected layer architecture, and f d represents the regression branch of the fully convolutional architecture. i represents the center coordinates, width, and height of the i-th candidate region.
[0095] The Smooth L1 is often used to optimize the gap between the offset predicted by the regressor and the annotation. Instead of calculating the gap between the predicted value of the regressor and the true annotation, the distance loss function calculates the gap between the predicted offsets of the two-way regressor, making the IoU (Intersection over Union) of the regression results of the two-way regression branches for the same candidate region as close to 1 as possible. It should be noted that the distance loss function does not directly introduce the annotation data because the two-way regression branches have been separately constrained by the annotation data, so that the distance loss will not cause the detector to deviate from the true result. Therefore, the distance loss function is theoretically feasible.
[0096] Based on the basic feature extraction network and the additional convolutional network, the classification network CLS1 and the localization network LOC1 of the first stage are built, and CLS1 and LOC1 are each composed of F convolutional layers. Among them, the classification network CLS1 and the localization network LOC1 are respectively expressed as CLS1 = {CLS 11 , CLS 12 , …, CLS 1f , CLS 1F}, LOC1 = {LOC 11 , LOC 12 , …, LOC 1f , LOC 1F}, where F is the number of feature maps jointly generated by the basic feature extraction network M and the additional convolutional network MA. CLS 1f and LOC 1f respectively represent the classification and localization networks on the f-th feature map, and are expressed as follows:
[0097] CLS 1f = Conv(channel 1f , 2, stride h1 , stride w1 );
[0098] LOC 1f = Conv(channel 1f , 5, stride h1 , stride w1 );
[0099] Among them, Conv represents a single convolutional layer. The input channel number channel 1f represents the channel number of the f-th feature map obtained by the basic feature extraction network and the additional convolutional network; 2 represents the convolutional output channel number of CLS 1f , representing that only the binary classification discrimination between foreground and background is performed at this time. 5 represents LOC1f The number of output channels of the convolution represents that the parameters for coordinate regression at this time are 5, corresponding to the object coordinates B described above m,n ; stride h1 and stride w1 are the height and width of the convolution kernel.
[0100] In some specific embodiments, the number of feature maps jointly generated by ResNet50 and the additional convolution network is 4, that is, F is 4, and the corresponding channel numbers of its feature maps are {512, 1024, 2048, 512}. stride h1 and stride w1 are both 3.
[0101] Based on the basic feature extraction network M and the additional convolution network MA, a feature pyramid network is constructed, and F first-stage feature maps FEA1 are generated, and further high-resolution second-stage feature maps FEA2 are generated;
[0102] Specifically, generating F first-stage feature maps FEA1 is represented as FEA1 = {FEA 11 , FEA 12 , …, FEA 1f , …, FEA 1F}, the width and height of the first-stage feature maps are respectively represented as W1 = {W 11 , W 12 , …, W 1f , …, W 1F} and H1 = {H 11 , H 12 , …, H 1f , …, H 1F}, where W 1f and H 1f respectively represent the width and height of the f-th feature map in the first stage.
[0103] When 1 ≤ i ≤ F - 1, it satisfies W 1i = 2 × W 1i+1 , H 1i = 2 × H 1i+1 . In FEA1, the high-level feature maps have a small spatial resolution but rich semantic information, and the low-level feature maps have a large spatial resolution, so they have more refined local features. The feature pyramid can transfer the semantic information of the high-level feature maps to the low-level, so as to combine the advantages of both and obtain feature maps with higher resolution and rich semantic information. The feature maps generated by the feature pyramid are denoted as FEA2, which is called the second-stage feature map, FEA2 = {FEA 21 , FEA 22 , …, FEA 2f,…,FEA 2F}, where FEA 2f represents the f-th feature map in the first stage. The number of feature maps of FEA2 is the same as that of FEA1, and FEA 2f has the same width and height as FEA 1f .
[0104] Among them, the feature pyramid network includes a feature map conversion network TS and a feature map scaling network INP. The feature map conversion network can be expressed as TS = {TS1, TS2, …, TS f , …, TS F}, and TS is also composed of F parts, where TS f represents the f-th feature map conversion network; INP = {INP1, INP2, …, INP f , …, INP F}, and INP is composed of F - 1 parts, where INP f represents the feature map scaling network between the f-th feature map and the (f + 1)-th feature map. The width and height of the feature map passing through the feature map scaling network will become twice the original.
[0105] Among them, during the construction of the feature map pyramid, the highest-level feature map is first processed separately;
[0106] Then, the processing is carried out in order of increasing spatial resolution of the feature maps:
[0107] FEA 2F = TS F (FEA 1F );
[0108] t = TSi(FEA 1i );
[0109] FEA 2i = t + INPi(FEA 2i+1 );
[0110] Among them, t is the intermediate feature map during the construction of the feature pyramid; where the value range of i is {F - 1, F - 2, …, 1}, and the feature pyramid network includes a feature map conversion network TS and a feature map scaling network INP.
[0111] Specifically, the intermediate feature map during the construction of the feature pyramid does not perform the final detection step. FEA 2F only needs to be executed once, while formula t and formula FEA 2i need to be executed F - 1 times in total.
[0112] In some specific embodiments, the feature map transformation network is implemented by a Res2net structure, which performs residual-based transformation and connection between different channels of a feature map, enhancing the feature extraction ability; the feature map scaling network is completed by a feature map interpolation function in the PyTorch library.
[0113] As Figure 2 shown, after the candidate region is processed by the double regression branch, two regression results are obtained; for the two regression results, the features of the candidate region N×N are extracted respectively for classification and regression, and after the features are added and averaged, the fused region features are obtained. The scaled feature map is merged with the previous feature map through a channel concatenation operation, and only then is it sent into the feature map transformation network to generate a new feature map. The feature map transformation network contains a total of 5 identical structures, and the feature map scaling network contains 4 identical structures. Each of the identical structures has independent trainable parameters.
[0114] Re-register the classification network CLS1 and the convolutional region CR1 of the convolutional network in the first stage with the coordinate detection result LR1 of the first-stage localization network LOC1;
[0115] Among them, the coordinate detection result can be expressed as LR1 = {w1, h1, cx1, cy1, θ1}, representing the width, height, center point coordinates, and rotation angle detected based on a preset anchor box.
[0116] Taking the result of a 3×3 convolution operation in the origin in a two-dimensional space as an example:
[0117]
[0118] CR2 = Rotate(Scale(Shift(CR1, LR1)))
[0119]
[0120] At this time, CR1 is a 3×3 rectangular region, SP1 represents the set of sampling point coordinates in the convolutional region CR1, with a total of 9 positions; Rotate, Scale, and Shift represent translation, scaling, and rotation operations on the convolutional region CR1 respectively according to the detection result LR1, and CR2 is the resulting new convolutional region; in the formula, SP2 is the set of sampling points of the new convolutional region CR2, and {p1, p2, p3, p4, p5, p6, p7, p8, p9} are the corresponding 9 sampling point coordinates.
[0121] Based on the feature map FEA2 and the reallocated convolutional region CR2 in the second stage, perform classification and localization in the second stage to obtain the classification network CLS2 and the localization network LOC2 in the second stage;
[0122] Are respectively represented as:
[0123] CLS2 = {CLS 21 , CLS 22 , …, CLS 2f , CLS 2F},
[0124] LOC2 = {LOC 21 , LOC 22 , …, LOC 2f , LOC 2F},
[0125] CLS 2f and LOC 2f respectively represent the classification and localization networks on the f - th feature map, and are represented as follows: CLS 2f = Conv(channel 2f , J, stride h2 , stride w2 );
[0126] LOC 2f = Conv(channel 2f , 5, stride h2 , stride w2 );
[0127] Among them, Conv represents a single convolutional layer, the input channel number channel 1f represents the channel number of the f - th feature map obtained by the basic feature extraction network and the additional convolutional network; 2 represents the convolutional output channel number of CLS 1f , representing that only binary classification discrimination between foreground and background is performed at this time, 5 represents the convolutional output channel number of LOC 1f , representing that the number of parameters for coordinate regression is 5 at this time, corresponding to the object coordinates B m,n described above; stride h1 and stride w1 are the height and width of the convolutional kernel.
[0128] Among them, channel 2f represents the channel number of the f - th feature map FEA 2f in the second stage, Conv represents a convolutional layer, J is the convolutional output channel number of CLS 2f , and is also the total number of object categories in the training and test images. Compared with CLS 1f , at this time, instead of performing binary classification, specific category judgment of the object is carried out; LOC 2fThe number of convolutional output channels is 5, which is used to detect the coordinates of the object and is the same as that of LOC 1f The structure is the same. The difference is that at this time, it is no longer the position detection based on the preset anchor boxes, but the result LR1 of the first-stage position detection is used as the basis to further refine the detection of the object position.
[0129] In some specific embodiments, the number of second-stage feature maps is 4, and their corresponding channel numbers are {256, 256, 256, 256}. stride h2 and stride w2 are both 3.
[0130] Define the loss function during the detection process of the first stage and the second stage;
[0131] Specifically, the loss function includes the binary classification and regression losses of the first-stage detection and the multi-classification and regression losses of the second stage; use the training set to train the network and obtain the final network model. The loss function is to calculate the final classification and position detection results of the network and the true category and position in the picture annotation information to obtain a value. The larger this value is, the worse the network performance is, and vice versa, the better the network performance is. The purpose of training is to reduce this loss value.
[0132] The loss function is:
[0133]
[0134]
[0135] Among them, i represents the subscript of the preset anchor box, p i and x i respectively represent the binary classification prediction probability and the coordinate detection result of the first stage; and respectively represent the true category and the position offset vector of the preset anchor box with subscript i, c i and t i are the multi-classification prediction probability and the coordinate detection result of the second stage, and N1 and N2 respectively represent the number of positive samples in the detection processes of the first stage and the second stage. L b is the binary cross-entropy loss for judging whether the object is a foreground or a background, L m is the multi-class cross-entropy loss for judging the category of the object, L r is the Smooth-L1 loss function.
[0136] In some specific embodiments, for all the preset anchor boxes, first determine whether they belong to positive samples or negative samples by calculating with the labeled positions in the input image; all the anchor boxes participate in the calculation of the classification loss, but only the anchor boxes belonging to positive samples participate in the calculation of the position loss, because for the anchor boxes belonging to negative samples, that is, the background category, their position information is not important.
[0137] And the total loss finally used to optimize the objective function is defined as the weighted sum of the losses in two stages:
[0138] Loss = λ1Loss1 + λ2Loss2;
[0139] Where λ1 and λ2 are weighting coefficients. Specifically, both λ1 and λ2 are 1.
[0140] Train with the training set to obtain the final network model.
[0141] S4. Test the test image according to the network model, calculate the intersection over union (IoU) and perform non-maximum suppression to obtain the final detection result.
[0142] According to the trained network model, use the sample T = {T1, T2, …, T q , …, T Q} of Q test images for testing. During testing, only send the image into the network for forward propagation to obtain the class scores and regression coordinates at each anchor point position in the image, discard the regions discriminated as background and the regions with scores less than the set score threshold t score and input them into the network model;
[0143] And save the detection results R = {R1, R2, …, R q , …, R Q} by category, where R q represents the detection result of the q-th test image, and R q = {R c1 , R c2 , …, R cj , …, R cJ}, where R cj represents all the detection results of the current test image in the j-th class; In some specific embodiments, the score threshold t score is 0.5, and discard all the results with predicted scores below 0.5 with low confidence.
[0144] The steps of image testing are as follows:
[0145] Normalize the test image at the pixel level;
[0146] Scale the test image to the same size as the image used for training;
[0147] Change the network model to test mode, stop calculating the loss and backpropagating the gradient for the detection results, and only perform the forward propagation process;
[0148] Obtain the initial detection result R of the q-th test image currently q .
[0149] In some specific embodiments, the initial detection result is the multi-classification and position detection result of the second stage, and the detection result of the first stage is only used for the forward propagation process of the network and is not used as the final detection result.
[0150] Finally, calculate the intersection over union of the rotated bounding boxes for the initial detection result R, perform non-maximum suppression, and only retain the detection boxes with larger scores and small overlapping areas as the final detection result.
[0151] As Figure 3 shown, the steps of the non-maximum suppression are as follows:
[0152] For the initial detection result R q re-sort the predicted scores of each detection box in the same class in descending order, and the sorted result is R' q ={R' c1 ,R' c2 ,…,R' cf ,…,R' cF}, where R' cf is the detection result on the j-th class after sorting, and each time retain the result with the largest current score as the current processing object;
[0153] For any detection box b in R' cf , calculate the intersection over union between it and all detection boxes with predicted scores less than the current score. The formula for the intersection over union is:
[0154] T = area b + area bs ;
[0155] I = inter w × inter h ;
[0156] U = T - I;
[0157] IOU = I / U;
[0158] where, area b represents the area of the detection box b, area bsRepresents the area of any detection box bs with a score less than b, inter w and inter h respectively represent the width and height of the intersection area of two detection boxes;
[0159] If the intersection over union of the areas of two detection boxes is less than the threshold, the corresponding detection result is discarded. If it exceeds the threshold t iou , then the detection box bs with the lower score is discarded;
[0160] From the currently remaining detection results, select the detection result with the highest score except the current detection object as the current processing object until the current processing object is the last detection result; end the process and output all the retained detection results.
[0161] Based on the understanding of the above solution, the protection scope of this application should not be limited to the information expressed literally, but should also include the connection relationship and execution logic in the method steps. Therefore, this embodiment can also be presented in other expression forms. For example Figure 4 in, this embodiment can also be expressed as the following steps: input an image, perform enhancement and scaling operations on the image, and synchronously adjust the position information of the object annotations in the image; obtain the first-stage feature map through a feature extraction network and an additional convolutional network, and perform first-stage binary classification and position detection; obtain the second-stage feature map through a feature pyramid network and perform second-stage binary classification and position detection according to the detection results of the first stage; based on the detection results of the second stage, perform third-stage multi-classification and position detection as the final result; if it is a training process, calculate the losses of the detection results of the first stage, the second stage, and the third stage, and update the network parameters through the gradient direction propagation algorithm; if it is not a training process, use the classification and position detection results as the initial detection results and perform result post-processing through the non-maximum suppression method.
[0162] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of a specific device is divided into different functional modules to complete all or part of the functions described above.
[0163] In the embodiments provided in this application, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the embodiments of the structures described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another structure, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of structures or units can be in electrical, mechanical or other forms.
[0164] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0165] In addition, each functional unit in the embodiments of this application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0166] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks or optical discs and other various media that can store program codes.
[0167] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A cascaded regression object detection method, characterized in that It includes the following steps: Obtain the image to be detected, perform pixel standardization and scale it to the same size; Input the adjusted image to be detected into the trained deep convolutional neural network model. The deep convolutional neural network model is trained by training images with annotation information and includes a backbone network for feature extraction, a cascaded region proposal module for step-by-step adjustment of preset boxes, a cascaded two-way regressor module for fine-tuning preset boxes, a distance loss function and a loss function for optimizing the above modules. In each stage of regression, the distance loss function uses the cross-entropy function to optimize classification and uses the Smooth L1 function to optimize regression, which is used to minimize the output result gap between the two regression branches. The definition of the distance loss function includes: ; Among them, Smooth L1 The function is as follows: ; Among them is the feature information fused by the coordinates predicted by the two regression branches in the previous stage, represents the regression branch of the fully connected layer architecture, represents the regression branch of the fully convolutional architecture, and i ∈ {x, y, w, h} represents the center coordinates, width, and height of the i-th candidate region; The loss function includes: the loss function of the cascaded region proposal module and the loss function of the cascaded two-way regressor: ; ; Among them, Loss1 is the loss function of the cascaded region proposal module, and Loss2 is the loss function of the cascaded two-way regressor. Here, i represents the subscript of the preset anchor box. and respectively represent the binary classification prediction probability and the coordinate detection result in the first stage. is the preset box corresponding to the class number of the real object. is the preset box corresponding to the position and size of the real object. and are the multi-classification prediction probability and the coordinate detection result in the second stage. N arm1 、N arm2 and N det respectively represent the number of positive samples in the detection processes of the first stage, the second stage, and the third stage of the cascaded regression. is the multi-classification cross-entropy loss for judging the object category. is the Smooth L1 loss function; the total loss is the weighted sum of the loss function of the cascaded region proposal module and the loss function of the cascaded two-way regressor; the trained deep convolutional neural network model is used to detect the picture to be detected to obtain the detection result.
2. The cascade regression object detection method according to claim 1, characterized in that The backbone network extracts features from the image through several feature prediction layers with different resolutions; the backbone network includes one of the ResNet series networks or the VGG series networks, and an additional convolutional network added on this basis.
3. The cascade regression object detection method according to claim 2, wherein, The cascaded region proposal module performs classification and regression based on the extracted features, including a cascaded region proposal network, a feature fusion module, and a classification and regression module; the cascaded region proposal network fine-tunes the size and position of the preset box in two steps to generate region proposals; The feature fusion module fuses features of different scales together; the classification and regression module extracts the features of the candidate region N×N according to the region proposals provided by the cascaded region proposal network for classification and regression.
4. The cascade regression object detection method according to claim 3, wherein The cascaded two-way regressor module uses neural networks with fully connected architectures and fully convolutional architectures as regression branches respectively to perform regression on the same candidate region, and adjusts the size and position of the candidate region, including: The first stage: Use the position of the preset box as the candidate region, extract the target features and crop them to the specified size, input them into the two regression branches respectively for prediction, adjust the size and position of the candidate region according to the prediction results and extract two sets of features, and use a convolutional neural network to fuse the two sets of features; The second stage: Input the features fused in the first stage into the two regression branches respectively for prediction, adjust the size and position of the candidate region according to the prediction results and extract two sets of features, and use a convolutional neural network to fuse the two sets of features; The third stage: Input the features fused in the second stage into the two regression branches respectively for prediction, adjust the size and position of the candidate region according to the prediction results and extract two sets of features, and use a convolutional neural network to fuse the two sets of features to obtain the size and position of the finally predicted candidate region.
5. The cascade regression object detection method according to claim 1, characterized in that Using the trained network model to detect the image to be detected includes: Use Detect the input of the trained network model using Zhang's test pictures Save the detection results by category; Calculate the intersection over union of the rotated rectangular boxes, perform non-maximum suppression, and only retain the detection boxes with larger scores and smaller overlapping areas as the final detection results.
6. The cascade regression object detection method according to claim 5, characterized in that The non-maximum suppression includes: For the initial detection results the predicted scores of each detection box in the same category in the initial detection results are respectively re-sorted in descending order, and the sorted result is , where is the detection result of the th category after sorting; For any detection box b in, calculate the intersection over union (IoU) between it and all detection boxes with prediction scores less than the current score; If the intersection over union of the areas of two detection boxes exceeds the threshold , then discard the detection box bs with the lower score.
7. A cascaded regression object detection device, characterized in that, Includes: An acquisition module for acquiring the image to be detected and / or training images; A preprocessing module for performing pixel standardization and scaling operations on the detection images and / or training images acquired by the acquisition module; A neural network module for configuring a deep convolutional neural network, including a backbone network for feature extraction, a cascaded region proposal module for step-by-step adjustment of preset boxes, a cascaded two-way regressor module for fine-tuning preset boxes, a distance loss function and a loss function for optimizing the above modules; in each stage of regression, the distance loss function uses the cross-entropy function to optimize classification and uses the Smooth L1 function to optimize regression, which is used to minimize the output result gap between the two-way regression branches; The definition of the distance loss function includes: ; Among them, Smooth L1 The function is as follows: ; Among them is the feature information fused by the coordinates predicted by the two regression branches in the previous stage, represents the regression branch of the fully connected layer architecture, represents the regression branch of the fully convolutional architecture, and i ∈ {x, y, w, h} represents the center coordinates, width, and height of the i-th candidate region; The loss function includes: the loss function of the cascaded region proposal module and the loss function of the cascaded two-way regressor: ; ; Among them, Loss1 is the loss function of the cascaded region proposal module, and Loss2 is the loss function of the cascaded two-way regressor. Here, i represents the subscript of the preset anchor box. and respectively represent the binary classification prediction probability and the coordinate detection result in the first stage. is the preset box corresponding to the class number of the real object. is the preset box corresponding to the position and size of the real object. and are the multi-classification prediction probability and the coordinate detection result in the second stage. N arm1 、N arm2 and N det respectively represent the number of positive samples in the detection processes of the first stage, the second stage, and the third stage of the cascaded regression. is the multi-classification cross-entropy loss for judging the object category. is the Smooth L1 loss function; the total loss is the weighted sum of the loss function of the cascaded region proposal module and the loss function of the cascaded two-way regressor. The neural network module is also used to train with the training images and detect the image to be detected to obtain the detection results.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the instructions are executed by a processor, the method described in any one of claims 1 to 6 is implemented.