Object Detection and Precise Positioning Method Based on Arbitrary Quadrilateral Regression
By constructing an object detection network model based on arbitrary quadrilateral regression, combining multiple attention mechanisms and multi-task cascade structures, the target's quadrilateral vertex coordinates are directly predicted, which solves the problem of precise positioning of rotating multi-angle targets and improves the robustness and accuracy of detection.
Patent Information
- Application Number
- CN202211365117.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-03
AI Technical Summary
When facing a rotating multi-angle target, it is difficult to accurately obtain any quadrilateral position of the target, and it is not robust enough in complex environments to meet the precise positioning requirements of specific detection tasks.
A target detection network model based on arbitrary quadrilateral regression is constructed, combining multiple attention mechanisms and multi-task multi-stage hybrid cascade structures, and directly predict the target's quadrilateral vertex coordinates through feature interaction fusion and local mapping for precise positioning.
It realizes high-precision positioning of the target in complex environments, can effectively correct the target angle, improves the robustness and accuracy of detection, and is suitable for different lighting and scale changes.
Smart Images

Figure CN115719414B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and object detection technology, and relates to an object detection and precise positioning method based on arbitrary quadrilateral regression. Background Art
[0002] Object detection is one of the basic tasks in the field of computer vision. In recent years, with the development of deep learning technology, object detection algorithms have shifted from traditional algorithms based on handcrafted features to detection technologies based on deep neural networks.
[0003] With the in-depth study of computer vision classification and recognition tasks, the research on detection algorithms based on convolutional neural networks has gradually expanded from improving the general object detection accuracy and detection speed of algorithms to object detection in specific fields. The objects in some scenarios generally have multi-angles with arbitrary rotations, and the method of ordinary positive box detection cannot meet the requirements. For example, remote sensing object detection, shelf commodity detection, text detection in natural scenes, and human body or object detection under top-shot fisheye lenses. Compared with general object detection, specific object detection has a more specific research background, and its research content often focuses on these special backgrounds.
[0004] Taking the instrument detection and recognition task in industrial applications as an example, currently, deep learning methods are widely applied to instrument detection and positioning. By adopting the position prediction method in general object detection tasks, the minimum circumscribed rectangle of the instrument is located, and the effect is much higher than the traditional detection and positioning method. However, due to the particularity of the instrument detection task, it is necessary to accurately obtain the dial position of the instrument and perform skew correction on it to facilitate subsequent tasks such as reading and recognition. Only locating the minimum rectangle surrounding the instrument has great limitations and may have an adverse impact on subsequent tasks, and more accurate positioning results are required.
[0005] The difficulties in using computer vision technology for object detection and precise positioning are mainly in two aspects: First, the object is likely to be deformed due to the tilt of the angle, and it is necessary to perform skew correction on the target area. Therefore, the algorithm should have the ability to predict the position of an arbitrary quadrilateral. The general object detection technology only predicts the minimum bounding box of the object, and for specific detection and positioning tasks, the positioning effect is difficult to meet the requirements of subsequent tasks, and it is impossible to further correct the object image in the frontal view angle only by using the position information of the minimum bounding box surrounding the object. In addition, even if instance segmentation technology is used to obtain the position mask of the object, there is a problem of how to reduce the error introduced when performing perspective transformation using the position mask; Second, the application scenarios are different, the indoor and outdoor lighting conditions are different, the environment where the object is located is complex with a large amount of interference information, the distance between the imaging device and the object is different, the scale of the object will change greatly with a large number of small objects, and the object types are diverse and the shapes are changeable. These adverse conditions pose challenges to the robustness of the method. Summary of the Invention
[0006] Technical Problem to be Solved
[0007] To avoid the deficiencies of the prior art, the present invention proposes a method for target detection and precise positioning based on arbitrary quadrilateral regression. First, the position of the quadrilateral is marked on the target image data, and then preprocessing operations are performed on the image, mainly including the division of the image data training set, validation set, and test set, as well as specific image enhancement processing; then, a meter detection network model based on arbitrary quadrilateral regression is constructed, including a feature extraction network module, an FPN module, an RPN module, an ROI Align pooling layer, a classification regression branch based on a fully connected layer, and a key point detection branch Grid Head based on a fully convolutional network; then, the enhanced image data set is used to train the network model to obtain a trained network; finally, the trained network is used to process the target image to be detected to obtain the final detection result. The present invention has the ability to detect arbitrary quadrilaterals, and can conveniently obtain the front view of the target key area using the predicted quadrilateral position information, providing convenience for subsequent processing; in addition, the network structure is adjusted and optimized, the feature size of the input key point detection branch is enlarged and local area mapping is performed according to the coordinates, multiple feature fusion is performed based on a multi-attention mechanism, and a multi-task multi-stage hybrid cascade structure and information interaction between branches are used to improve the algorithm performance from multiple angles and improve the detection accuracy and robustness of the method under various adverse conditions.
[0008] Technical Solution
[0009] A method for target detection and precise positioning based on arbitrary quadrilateral regression, characterized by the following steps:
[0010] Step 1: Construct a target detection network model based on arbitrary quadrilateral regression. This model is built based on the Faster RCNN network model, and a key point detection branch Grid Head with feature interaction and fusion based on a multi-attention mechanism is connected to the output ends of the ROI Align pooling layer and the bounding box regression branch of the Faster RCNN network.
[0011] The key point detection branch Grid Head is built based on a fully convolutional network, including a convolutional sequence for feature extraction, a module for increasing local feature mapping, a feature interaction and fusion module, a deconvolution layer for changing the feature size, and a hybrid cascade structure; the convolutional sequence is used to extract features from the input image features to be detected, after feature extraction, the features are enlarged and locally mapped, and then a feature fusion module based on a multiple attention mechanism is used to perform multi-level fusion processing on the extracted features, the fused output feature map is input into a multi-layer deconvolution layer to output a heat map for extracting the key point coordinates, and a hybrid cascade structure with multiple tasks and multiple stages is combined with information interaction to further refine the bounding box regression results, and the coordinates of the four vertices Grid Point of the arbitrary quadrilateral of the key area of the target to be detected are obtained by converting the finally obtained heat map;
[0012] Step 2: Independently collect and organize the target picture data under the monitoring device. After dividing the image training set, validation set and test set, corresponding data augmentation means are performed on each target image, and the images before and after the augmentation processing together constitute the target image data set;
[0013] Step 3: Using the training set and validation set in the image data set obtained in Step 2 as the input, the target detection network model based on arbitrary quadrilateral regression constructed in Step 1 is trained by the stochastic gradient descent method to obtain a trained network model, and the performance of the obtained network model is evaluated using the test set;
[0014] Step 4: Input the target image to be detected into the network model trained in Step 3, and output the class information and the vertex coordinates of the arbitrary quadrilateral of the target key area, and further accurately locate on the basis of completing the target detection.
[0015] The specific process of increasing the local feature mapping in Step 1 is as follows:
[0016] For the target to be detected, all Grid Points share the same feature expression area. To solve the problem of the feature expression area, the mapping relationship between the key point position coordinates predicted by the heat map and the position coordinates of the point corresponding to the original image is changed. The process is as follows:
[0017] First, expand the width and height of the input feature map of Grid Head to twice the original, increase the area mapped by the feature map on the original image, and include the Grid Point inside the candidate box generated by the RPN network;
[0018] Then, locally map the enlarged feature map according to the position where the Grid Point is located. For each GridPoint, the new output represents the entire feature Figure 4One - tenth of the area, the heatmaps corresponding to the four Grid Points are generated from different regions of the complete features, rather than all key points sharing the same feature expression region;
[0019] After processing, the expression of each Grid Point can be approximately regarded as a normalization process, which improves the positioning accuracy without increasing the computational complexity.
[0020] The specific process of feature interaction and fusion based on the multi - attention mechanism in step 1 is as follows:
[0021] The convolutional sequence for feature extraction consists of multiple convolutional layers, which extracts features from the input image features F to be detected din The extracted features are denoted as F d , when extracting features, first increase the features and perform local mapping, and then use the feature fusion module based on the multi - attention mechanism to perform multi - level fusion processing on the extracted features F d The specific process is as follows:
[0022] Divide the features F d into M groups on average according to the channels. The feature map corresponding to the i - th Grid Point is denoted as F di , and the feature map corresponding to the j - th point in the source point set S i is denoted as F dj , i = 1, 2,..., M, M is the number of Grid Points, j = 1, 2,..., K i , K i is the number of source points contained in the source point set S i ;
[0023] The so - called source points are the points in the Grid grid that are at a distance of 1 from the i - th Grid Point, and all source points form a source point set;
[0024] Then, pass the feature map F dj through a convolutional layer to obtain the corresponding new feature map to be fused, denoted as T d:j→i (F dj ); Next, add and fuse the feature map F di with the fused feature map T d:j→i (F dj ) according to the following formula for i = 1, 2,..., M to obtain the fused feature map F′ di :
[0025]
[0026] Then, perform secondary addition and fusion processing on the feature map F′ di according to the following formula to obtain the second - level fused feature map F″di :
[0027]
[0028] Among them, T' j→i (F' dj ) represents the new secondary feature map to be fused obtained from the feature map F' dj through the convolutional layer. The convolutional layer structure is the same as the convolutional layer structure in the previously obtained feature map T d:j→i (F dj ), where i = 1, 2,..., M and j = 1, 2,..., K i ;
[0029] For the multi-level features {F di , F', dj , F' di} obtained by secondary fusion, each level of feature is respectively represented as a four-dimensional tensor F ∈ R L×H×W×C . Among them, L represents the number of layers of the feature, W and H are respectively the height and width of the feature, and C is the number of channels; define S = H × W to obtain a three-dimensional tensor of L × S × C, and apply the attention mechanism to learn from the three dimensions of feature level, space, and task, using three consecutive attentions:
[0030] W(F) = π C (π S (π L (F)·F)·F)·F
[0031] π L , π S , and π C respectively represent different attention methods in the L, S, and C dimensions, and use the attention mechanism separately for the three dimensions of the feature: the hierarchical attention module is only used in the hierarchical dimension, and it learns the relative importance of each semantic level, enhancing the features of the target at the appropriate level; the spatial attention module is used in the S = H × W dimension to learn the internal discriminative representation at each spatial position; the task attention module is used in the channel dimension, and according to the different responses of the convolutional kernel to the object, it guides different feature channels to perform different tasks, making the features more suitable for key point learning;
[0032] The module containing three perceptual attentions is stacked in series to form a complete attention module, and the multi-level features {F di , F', dj , F'' di} are fused through a unified attention mechanism to obtain a heat map for predicting the coordinates of four vertices; then the fused output feature map is input into a multi-layer deconvolutional layer to output the final heat map for extracting the key point coordinates.
[0033] The specific process of the multi-task and multi-stage hybrid cascade structure and information interaction between branches in step 1 is as follows:
[0034] Combine bounding box regression and Grid Point prediction in a multi-task manner, abandon the parallel structure and alternate. In each stage, first execute the bounding box regression branch, and then hand over the regressed bounding box to the Grid Head to predict the Grid Point. At the same time, add a connection between the Grid Heads of adjacent stages. The features G i of are embedded through convolution and then input into the next stage G i+1 , G i+1 can obtain both the original features and the features of the previous stage, and integrate cascading and multi-tasking in each stage to process and improve the information flow.
[0035] The specific process of converting the heat map in step 2 into vertex coordinates is as follows:
[0036] Map the heat map coordinates back to the original image and calculate according to the following formula:
[0037]
[0038]
[0039] Among them, (I x , I y ) represents the vertex position coordinates of the target to be detected in the image, (P x , P y ) represents the vertex position coordinates of the bounding box generated by the RPN module, (H x , H y ) represents the position of the finally predicted point in the feature heat map, (w p , h p ) represents the width and height of the bounding box generated by the RPN model, (w o , h o ) represents the width and height of the heat map.
[0040] The specific process of training the network model in step 3 is as follows:
[0041] The loss function of the network is calculated according to the following formula:
[0042] Loss = L cls + L reg
[0043] Among them, Loss represents the total loss of the network, L cls represents the sum of the classification loss of the RPN module and the classification loss in the classification regression detection head, L regIt represents the sum of the location regression loss of the RPN module, the bounding box regression loss in the detection head, and the key point regression loss in the Grid Head; the classification loss, the location regression loss of the RPN module, and the bounding box regression loss are consistent with those in Faster RCNN; the key point regression loss L grid is the cross-entropy loss between the heatmap and the label map in the GridHead, which is calculated as follows:
[0044] L grid = L grid未融合 + L grid已融合
[0045] where, L grid未融合 represents the cross-entropy loss corresponding to the un-fused heatmap, and L grid已融合 represents the cross-entropy loss corresponding to the finally fused heatmap, which are calculated according to the following formulas respectively:
[0046]
[0047]
[0048] where, M is the number of GridPoints, N is the number of pixels of the heatmap, t k,l represents the value of the k-th pixel in the finally fused feature heatmap corresponding to the l-th GridPoint, and t′ k,l represents the value of the k-th pixel in the un-fused feature heatmap corresponding to the l-th Grid Point, t k,l and t′ k,l range from 0 to 1, represents the value of the k-th pixel in the label map corresponding to the un-fused feature heatmap corresponding to the l-th GridPoint, with a value range of 0 and 1. When the pixel is 1, it means that the corresponding area is the predicted Grid Point area, and when the pixel is 0, it means that the corresponding area is not the predicted Grid Point area.
[0049] Beneficial effects
[0050] A method for object detection and precise positioning based on arbitrary quadrilateral regression proposed by the present invention. First, perform image preprocessing. After dividing the image data into training set, validation set and test set, according to the characteristics of the dataset, adopt corresponding data augmentation means, such as adding random cropping based on object coordinates, brightness perturbation and brightness histogram equalization. Secondly, construct a neural network model. After the backbone network extracts features, construct a key point detection branch, and directly predict the positions of four key points in the target area through heatmap regression, so that the neural network has the ability to directly predict an arbitrary quadrilateral, thereby precisely locating the key area of the target. Finally, improve and optimize the algorithm model, expand the feature size and perform local mapping, perform feature fusion based on the attention mechanism, and use a multi-task multi-stage hybrid cascade structure and information interaction between branches to further improve the object detection and positioning accuracy.
[0051] In the target image dataset collected by the present invention, the size of the target and the lighting conditions, etc. cannot cover all situations. Through data augmentation means such as random cropping enhancement based on object coordinates, brightness perturbation and brightness histogram equalization, the scale and brightness change range of the data pictures can be increased, thereby increasing the data diversity, which is beneficial to training a network model with stronger generalization ability; since the method of predicting heatmp is adopted to predict the vertices of an arbitrary quadrilateral, the vertex coordinates can be flexibly selected, and the front view of the target key area can be directly obtained through perspective transformation, excluding interference information during correction, providing a good premise for subsequent operations; by further improving and optimizing the network structure and using a Grid Head network module for feature interaction and fusion based on multiple attentions, the detection and positioning accuracy and robustness of the network are improved. The present invention can effectively locate the position of the target after imaging at different imaging angles, and still maintain a high object detection accuracy in the face of various adverse conditions such as lighting interference and target scale change, which helps to promote the development of current object detection in special fields using computer vision technology. Brief Description of the Drawings
[0052] Figure 1 is the flowchart of the object detection and precise positioning method based on arbitrary quadrilateral regression of the present invention;
[0053] Figure 2 is the structural diagram of the Grid Head network module for feature interaction and fusion based on multiple attentions;
[0054] Figure 3 is the schematic diagram of the distribution of 4 Grid Points during the quadrilateral position prediction in the present invention;
[0055] Figure 4 is the result image of applying the method of the present invention to instrument detection;
[0056] In the figure, (a) - the detection result image of arrester 1; (b) - the dial image obtained by perspective transformation of the detection result of arrester 1; (c) - the detection result image of arrester 2; (d) - the dial image obtained by perspective transformation of the detection result of arrester 2.
[0057] Figure 5 is the result image of applying the method of the present invention to license plate detection.
[0058] In the figure, (a) - the detection result image of license plate 1; (b) - the detection result image of license plate 2 in snowy days; (c) - the license plate image obtained by perspective transformation of the detection result of license plate 1; (d) - the license plate image obtained by perspective transformation of the detection result of license plate 2. Specific embodiments
[0059] The present invention will be further described below in conjunction with embodiments and drawings:
[0060] As Figure 1 shown, the present invention provides a target detection and precise positioning method based on arbitrary quadrilateral regression. Taking the detection and positioning of instruments as an example, the specific implementation process is as follows:
[0061] 1. Image data preprocessing
[0062] After dividing the training set, validation set and test set of instrument image data, according to the characteristics of the data, corresponding data augmentation means are carried out on each instrument image respectively. Random cropping and stitching augmentation based on the instrument coordinates are adopted, random angle rotation and direction flipping are added, brightness perturbation and brightness histogram equalization are applied, and the scale and brightness change range of the data pictures are increased to further increase the diversity of the data. The images before and after the augmentation process together constitute the instrument image data set.
[0063] 2. Construct an instrument detection network model based on arbitrary quadrilateral regression
[0064] In order to accurately locate the key area of the instrument, the backbone network needs to ensure both efficient feature extraction operations and fast processing speed. The ResNet-50 network is selected as the backbone network to complete the feature extraction. Therefore, the instrument detection network model based on arbitrary quadrilateral regression constructed by the present invention includes a feature extraction network module based on ResNet-50, an FPN module, an RPN module, a ROI Align pooling layer, a classification regression branch based on a fully connected layer, and a key point detection branch Grid Head based on a fully convolutional network.
[0065] The ResNet-50 neural network is composed of multiple convolutional layers, pooling layers and residual structures. The instrument picture to be detected is input into the ResNet-50 network module, and the features of the instrument picture to be detected are output.
[0066] The FPN module is a Feature Pyramid Network. It takes the feature maps extracted by the backbone network as input and outputs multiple feature maps of different sizes to cope with the scale changes of the instrument.
[0067] The RPN module includes multiple convolutional layers, which perform preliminary localization processing on the features of the instrument image to be detected and output the coordinates of the rectangular bounding box of the instrument.
[0068] The ROI Align pooling layer calculates the pixel values at non-integer positions using bilinear interpolation, realizes the normalization of feature maps of different sizes, and outputs feature maps of different sizes into feature maps of the same size. That is, according to the coordinates of the rectangular bounding box of the instrument, the features of the instrument image to be detected corresponding to it are pooled into the same size.
[0069] The pooled features of the instrument image to be detected are input into the Grid Head, and the marking information of the four vertex positions of the quadrilateral of the instrument to be detected is output.
[0070] 3. Grid Head Module Based on Multi-attention Feature Interaction and Fusion
[0071] The structure of the Grid Head network module based on multi-attention feature interaction and fusion is as Figure 2 shown, including a convolutional sequence network module, a multi-attention feature interaction and fusion module, and a deconvolution layer.
[0072] The feature extraction network module uses 8 convolutional layers to extract features from the input features F of the image to be detected din and the extracted features are denoted as F d . When extracting features, the features are first enlarged and locally mapped, as follows:
[0073] For the target to be detected, all Grid Points share the same feature expression region. To solve the problem of the feature expression region, the mapping relationship between the key point position coordinates predicted by the heat map and the position coordinates of the point corresponding to the original image is changed. First, the width and height of the feature map input to the Grid Head are doubled, increasing the area mapped by the feature map on the original image, so that the Grid Points are as much as possible included inside the candidate boxes generated by the RPN network; then, the enlarged feature map is locally mapped according to the position of the Grid Points. For each Grid Point, the new output represents the entire feature Figure 4One - tenth of the area. The heatmaps corresponding to the four Grid Points are generated from different regions of the complete features, rather than all key points sharing the same feature expression region. After such processing, the expression of each GridPoint can be approximately regarded as a normalization process, which improves the positioning accuracy without increasing the computational complexity;
[0074] Then, use the feature interaction and fusion module to perform multi - level fusion processing on the extracted feature F d The specific process is as follows:
[0075] Divide the feature F d into M groups on average according to the channels. The feature map corresponding to the i - th Grid Point is denoted as F di , and the feature map corresponding to the j - th point in the source point set S i is denoted as F dj , where i = 1, 2, …, M, M is the number of Grid Points, j = 1, 2, …, K i , K i is the number of source points contained in the source point set S i . The source point is the point in the Grid grid that is at a distance of 1 from the i - th Grid Point. All source points form the source point set. Then, pass the feature map F dj through a convolutional layer with 2 convolutional kernels of 3×3 to obtain the corresponding new feature map to be fused, denoted as T d:j→i (F dj ); Next, add and fuse the feature map F di with the fused feature map T d:j→i (Fd j) ) according to the following formula, i = 1, 2, …, M, to obtain the fused feature map F′ di :
[0076]
[0077] Then, perform a secondary addition and fusion process on the feature map F′ di according to the following formula to obtain the second - level fused feature map F″ di :
[0078]
[0079] where T′ j→i (F′ dj ) represents the new second - level feature map to be fused obtained by passing the feature map F′ dj through 2 convolutional layers. The convolutional layer structure here is the same as the convolutional layer structure in obtaining the feature map T d:j→i (F dj ), i = 1, 2, …, M, j = 1, 2, …, Ki .
[0080] For the multi-level features {F di , F′ dj , F″ di} obtained by secondary fusion, each level of feature can be respectively represented as a four-dimensional tensor F∈R L×H×W×C , where L represents the number of feature layers, W and H are the height and width of the feature respectively, and C is the number of channels. Define S = H×W to obtain a three-dimensional tensor of L×S×C, and apply the attention mechanism to learn from three dimensions: feature level, space, and task, using three consecutive attentions:
[0081] W(F)=π C (π S (π L (F)·F)·F)·F (3)
[0082] π L , π S , π S respectively represent three different attention methods in the L, S, and C dimensions, each responsible for one part. The attention mechanism is used separately for the three dimensions of the feature: the level attention module is only used in the level dimension, which learns the relative importance of each semantic level and enhances the features of the target at the appropriate level; the spatial attention module is used in the S = H×W dimension to learn the intrinsic discriminative representation at each spatial position; the task attention module is used in the channel dimension, and according to the different responses of the convolution kernel to the object, it guides different feature channels to perform different tasks, making the features more suitable for key point learning.
[0083] The module containing three perceptual attentions is stacked in series to form a complete attention module. The multi-level features {F di , F′ dj , F″ di} are fused through a unified attention mechanism to obtain a 4-channel feature map, which is applied to the heat map for predicting the coordinates of four vertices. Then, the fused output feature map is input into a multi-layer deconvolution layer to output the final heat map for extracting the key point coordinates.
[0084] To further improve the accuracy of Grid Point detection, the idea of cascade is introduced, and a multi-task multi-stage hybrid cascade detection head is constructed. The bounding box regression and Grid Point prediction are combined in a multi-task manner, abandoning the parallel structure and alternating execution. In each stage, the box regression branch is executed first, and then the regressed box is handed over to the key point detection branch to predict the Grid Point. At the same time, a connection is added between the Grid Heads of adjacent stages. The feature G iThe features are embedded through a 1×1 convolution and then input into the subsequent stage G. i+1 G i+1 can obtain both the original features and the features of the previous stage, and integrates cascading and multi-tasks at each stage to process and improve the information flow.
[0085] 3. Network model training
[0086] Using the images in the image dataset obtained in step 1 as input, the instrument detection network model based on arbitrary quadrilateral regression constructed in step 2 is trained by the stochastic gradient descent method to obtain a trained network model; among them, the loss function of the network is calculated according to the following formula:
[0087] Loss = L cls + L reg (4)
[0088] where Loss represents the total loss of the network, L cls represents the sum of the classification loss of the RPN module and the classification loss of the RCNN module, and L reg represents the sum of the location regression loss of the RPN module, the box regression branch loss and the key point regression loss. The classification loss, the location regression loss of the RPN module and the box regression loss are the same as those in Faster RCNN; the key point regression loss L grid is the cross-entropy loss between the heatmap and the label map in the key point detection branch GridHead, and is calculated according to the following formula:
[0089] L grid = L grid未融合 + L grid已融合 (5)
[0090] where L grid未融合 represents the cross-entropy loss corresponding to the un-fused heatmap feature map, and L grid已融合 represents the cross-entropy loss of the finally fused heatmap feature map, and are calculated according to the following formulas respectively:
[0091]
[0092]
[0093] where M is the number of Grid Points, N is the number of pixels of the heatmap feature map, and t k,t represents the value of the k-th pixel in the finally fused heatmap feature map corresponding to the l-th GridPoint, and t′ k,l represents the value of the k-th pixel in the un-fused heatmap feature map corresponding to the l-th Grid Point, and t k,land t' k,l The value range is from 0 to 1, represents the value of the k-th pixel in the label map corresponding to the unfused heatmap feature map of the l-th GridPoint, and the value range is 0 and 1. When the pixel is 1, it means that the corresponding area is the predicted GridPoint area, and when the pixel is 0, it means that the corresponding area is not the predicted GridPoint area.
[0094] 4. Instrument detection
[0095] Input the instrument image to be detected into the network model trained in step 3, and output the predicted heatmap feature map. Convert the generated heatmap into the quadrilateral vertex positions of the instrument to be detected. The schematic diagram of Grid Point and the quadrilateral vertex positions is as Figure 3 shown as follows:
[0096]
[0097]
[0098] Among them, (I x , I y ) represents the vertex position coordinates of the instrument to be detected in the image, (P x , P y ) represents the vertex position coordinates of the bounding box generated by the RPN module, (H x , H y ) represents the position of the final heatmap prediction point in the heatmap feature map, (w p , h p ) represents the width and height of the bounding box generated by the RPN model, (w o , h o ) represents the width and height of the heatmap.
[0099] To verify the effectiveness of the method of the present invention, under the hardware environment: CPU: i9-9900, memory: 16G, hard disk: 1T, independent graphics card: NVIDIA GeForce RTX 2080ti, 11G, and the system environment is Ubuntu18.0.4, software python3.7, opencv3.4, and Pytorch1.3 are used for simulation experiments. The dataset used in the experiment is a self-built instrument dataset, Figure 4 The instrument detection result image obtained by detecting using the method of the present invention is given. To verify that the method of the present invention is applicable to different application scenarios, the same processing process is carried out on the license plate public dataset CCPD, Figure 5The license plate detection result image is given. It can be seen that by using the Gird Point prediction method, the key target area can be accurately located and corrected, which is applicable to different application scenarios, and the network model can still achieve a high positioning accuracy under adverse conditions such as different illuminations, different perspectives, and different scales.
Claims
1. A method for object detection and precise positioning based on arbitrary quadrilateral regression, characterized in that The steps are as follows: Step 1: Construct an object detection network model based on arbitrary quadrilateral regression. This model is built based on the Faster RCNN network model, and a key point detection branch Grid Head based on multi-attention mechanism feature interaction and fusion is connected to the output ends of the ROI Align pooling layer and the bounding box regression branch of the Faster RCNN network; The key point detection branch Grid Head is built based on a fully convolutional network, including a convolutional sequence for feature extraction, a module for increasing feature local mapping, a feature interaction and fusion module, a transposed convolutional layer for changing the feature size, and a hybrid cascade structure; the convolutional sequence is used to extract features from the input image features to be detected. After feature extraction, the features are increased and locally mapped, and then the feature fusion module based on the multi-attention mechanism is used to perform multi-level fusion processing on the extracted features. The fused output feature map is input into a multi-layer transposed convolutional layer to output a heat map for extracting the key point coordinates. The hybrid cascade structure with multi-task and multi-stage and information interaction is combined with the bounding box regression result to further refine, and the coordinates of the four vertices Grid Point of the arbitrary quadrilateral of the key area of the object to be detected are obtained by converting the finally obtained heat map; Step 2: Automatically collect and organize the target picture data under the monitoring device. After dividing the image training set, validation set, and test set, corresponding data augmentation means are performed on each target image respectively. The images before and after the augmentation processing together constitute the target image dataset; Step 3: Using the training set and validation set in the image dataset obtained in Step 2 as the input, the object detection network model based on arbitrary quadrilateral regression constructed in Step 1 is trained by the stochastic gradient descent method to obtain a trained network model, and the performance of the obtained network model is evaluated using the test set; Step 4: Input the target image to be detected into the network model trained in Step 3, and output the class information and the vertex coordinates of the arbitrary quadrilateral of the key area of the target, and further accurately locate on the basis of completing object detection.
2. The object detection and precise positioning method based on arbitrary quadrilateral regression according to claim 1, characterized in that: The specific process of increasing feature local mapping in Step 1 is as follows: For the object to be detected, all Grid Points share the same feature expression area. To solve the problem of the feature expression area, the mapping relationship between the key point position coordinates predicted by the heat map and the position coordinates of this point corresponding to the original image is changed. The process is as follows: First, expand the width and height of the feature map input to Grid Head to twice the original, increase the area mapped by the feature map on the original image, and include Grid Point inside the candidate box generated by the RPN network; Then, locally map the enlarged feature map according to the position where Grid Point is located. For each Grid Point, the new output represents one-fourth of the entire feature map area. The heat maps corresponding to the four Grid Points are generated from different areas of the complete feature, rather than all key points sharing the same feature expression area; After processing, the expression of each Grid Point can be approximately regarded as a normalization process, which improves the positioning accuracy without increasing the computational complexity.
3. The object detection and precise positioning method based on arbitrary quadrilateral regression according to claim 1, characterized in that: The specific process of feature interaction and fusion based on the multiple attention mechanism in step 1 is as follows: The convolutional sequence for feature extraction consists of multiple convolutional layers, and performs feature extraction on the input image features F to be detected din The extracted features are denoted as F d , during feature extraction, the features are first enlarged and locally mapped, and then the feature fusion module based on the multi-attention mechanism is used to perform multi-level fusion processing on the extracted features F d The specific process is as follows: Feature F d is evenly divided into M groups according to the channels, and the feature map corresponding to the i-th Grid Point is denoted as F di , and the feature map corresponding to the j-th point in the source point set S i is denoted as F dj , where i = 1, 2, …, M, M is the number of Grid Points, and j = 1, 2, …, K i , and K i is the number of source points included in the source point set S i ; The source points are the points in the Grid grid that are at a distance of 1 from the i-th Grid Point, and all the source points form a source point set; Then, the feature map F dj passes through the convolutional layer to obtain the corresponding new feature map to be fused, denoted as T d:j→i (F dj ); Next, the feature map F di and the fused feature map T d:j→i (F dj ) are added and fused according to the following formula, i = 1, 2, …, M, to obtain the fused feature map F′ di : Then, for the feature map F' di perform a secondary addition fusion process according to the following formula to obtain the secondary fusion feature map F'' di : Among them, T' j→i (F' dj ) represents the new secondary feature map to be fused obtained from the feature map F' dj through the convolutional layer. The convolutional layer structure is the same as the convolutional layer structure in the feature map T d:j→i (F dj ), where i = 1, 2, …, M and j = 1, 2, …, K i ; For the multi-level features {F di , F′ dj , F″ di} obtained by secondary fusion, each level of feature is respectively represented as a four-dimensional tensor F ∈ R L×H×W×c , where L represents the number of layers of the feature, W and H are the height and width of the feature respectively, and C is the number of channels; define S = H × W to obtain a three-dimensional tensor of L × S × C, and apply the attention mechanism to learn from three dimensions: feature level, space, and task, using three consecutive attentions: W(F) = π C (π S (π L (F)·F)·F)·F π L and π S and π C respectively represent different attention methods in the L, S, and C dimensions, using the attention mechanism separately for the three dimensions of features: the hierarchical attention module is only used in the hierarchical dimension, where it learns the relative importance of each semantic level and enhances the features of the target at the appropriate level; the spatial attention module is used in the S = H×W dimension to learn the intrinsic discriminative representations at each spatial position; the task attention module is used in the channel dimension. Based on the different responses of the convolutional kernels to objects, it guides different feature channels to perform different tasks, making the features more suitable for key point learning; A complete attention module is formed by cascading and stacking three modules that include three types of perceptual attention. Multilevel features {F di , F′ dj , F″ di} are fused through a unified attention mechanism to obtain a heatmap for predicting the coordinates of four vertices; then the fused output feature map is input into a multi-layer transposed convolutional layer to output the final heatmap for extracting the coordinates of key points.
4. A method for object detection and precise positioning based on arbitrary quadrilateral regression according to claim 1, characterized in that: The specific process of the multi-task and multi-stage hybrid cascade structure and information interaction between branches in step 1 is as follows: The bounding box regression and grid point prediction are combined in a multi-task manner, and the parallel structure is abandoned and alternated. In each stage, the bounding box regression branch is executed first, and then the regressed bounding box is handed over to the Grid Head to predict the grid point. At the same time, a connection is added between the Grid Heads of adjacent stages. The feature G i of is embedded through convolution and then input into the next stage G i+1 , G i+1 can obtain both the original feature and the feature of the previous stage, and integrate cascading and multi-tasking in each stage to process and improve the information flow.
5. A method for object detection and precise positioning based on arbitrary quadrilateral regression according to claim 1, characterized in that: The specific process of converting the heat map in step 2 into vertex coordinates is as follows: The heat map coordinates are mapped back to the original image and calculated according to the following formula: Among them, (I x , I y ) represents the vertex position coordinates of the target to be detected in the image, (P x , P y ) represents the vertex position coordinates of the bounding box generated by the RPN module, (H x , H y ) represents the position of the finally predicted point in the feature heat map, (w p , h p ) represents the width and height of the bounding box generated by the RPN model, (w o , h o ) represents the width and height of the heat map.
6. The object detection and precise positioning method based on arbitrary quadrilateral regression according to claim 1, wherein: The specific process of training the network model in step 3 is as follows: The loss function of the network is calculated according to the following formula: Loss=L cls +L reg Among them, Loss represents the total loss of the network, and L cls represents the sum of the classification loss of the RPN module and the classification loss in the classification regression detection head. L reg represents the sum of the location regression loss of the RPN module, the bounding box regression loss in the detection head, and the key point regression loss in the Grid Head; the classification loss, the location regression loss of the RPN module, and the bounding box regression loss are the same as those in Faster RCNN; the key point regression loss L grid is the cross-entropy loss between the heatmap and the label map in the Grid Head, and is calculated according to the following formula: L grid = L gri d 未融合 + L grid已融合 Among them, L grid未融合 represents the cross-entropy loss corresponding to the un-fused heatmap, and L grid已融合 represents the cross-entropy loss corresponding to the finally fused heatmap, which are calculated according to the following formulas respectively: Where M is the number of Grid Points, N is the number of pixels of the heat map, and t k,l represents the value of the k-th pixel in the finally fused feature heat map corresponding to the l-th Grid Point, and t' k,l represents the value of the k-th pixel in the unfused feature heat map corresponding to the l-th Grid Point, t k,l and t' k,l range from 0 to 1, represents the value of the k-th pixel in the label map corresponding to the unfused feature heat map corresponding to the l-th Grid Point, with a range of 0 and 1. A pixel value of 1 indicates the corresponding area is the predicted Grid Point area, and a pixel value of 0 indicates the corresponding area is not the predicted Grid Point area.
Citation Information
Patent Citations
Target area detection model training method, system and device, and medium
CN112949766A
Instrument detection method based on one-shot mechanism
CN113792721A