An Adaptive Focusing and Positioning Target Detection Method
By introducing residual aliasing module, the sum of the multiplication of the target basis and coefficients, and the center point error of the mask guidance in the traffic sign detection, the problem of insufficient traffic sign recognition accuracy and aliasing between the targets in the prior art is solved, and a higher recognition accuracy and more stable detection results are achieved.
Patent Information
- Application Number
- CN202210501677.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-05-09
AI Technical Summary
The existing traffic sign detection method deals with the similarity and varied size between targets, and there is insufficient recognition accuracy and aliasing between targets.
Adaptive focus positioning object detection method is adopted, and the residual aliasing module, the sum of the multiplication of the target basis and coefficients, and the center point error guided by masks are introduced into the object detection model, the recognition accuracy is improved and aliasing between the targets is reduced.
The accuracy of traffic sign recognition is improved, the influence of low-quality target base is weakened, and the detection results are deteriorated due to unvalued errors in segmentation.
Smart Images

Figure CN114926812B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and particularly to an adaptive focusing and positioning target detection method. Background Art
[0002] Current traffic sign detection can generally be divided into two categories. One is the traditional method of recognizing traffic signs based on features such as the shape and color of traffic signs; the other is using deep learning to recognize traffic signs. The deep learning method can be further divided into single-stage and two-stage methods according to the method it adopts. The two-stage method first uses a neural network to generate candidate boxes of various sizes and scales at each pixel point, and then determines whether the target is contained in the candidate box and the category of the target. The single-stage method directly obtains the parameters of the target rectangular box and the category label of the corresponding target through the neural network. The accuracy of the two-stage target detection method is generally higher than that of the single-stage method, but this comes at the cost of computational complexity.
[0003] Compared with other target detection tasks, the traffic sign detection task does not need to consider the occlusion situation, but the similarity of different traffic signs and the variable sizes of traffic signs caused by different distances will lead to difficulties in recognition. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an adaptive focusing and positioning target detection method, which can improve the accuracy of target detection.
[0005] The technical solution adopted by the present invention to solve its technical problem is: to provide an adaptive focusing and positioning target detection method, including the following steps:
[0006] Receiving an image to be recognized;
[0007] Inputting the image to be recognized into a target detection model to obtain the position and category of the target in the image to be recognized; wherein, the target detection model includes:
[0008] A feature extraction layer for extracting features of the image to be recognized;
[0009] A target category prediction layer for performing a block operation on each layer of features extracted by the feature extraction layer, and performing category prediction and coefficient prediction on each block;
[0010] A target location prediction layer for generating a mask tuple according to the features extracted by the feature extraction layer, and multiplying the mask tuple by the coefficients obtained by the target category prediction layer and then summing to obtain a target mask.
[0011] An aliasing residual structure is added between adjacent residual structures in the feature extraction layer.
[0012] The aliasing residual structure includes a first processing unit, a second processing unit, and a second ReLU activation function layer. The input of the first processing unit is the feature information of this layer. The output of the first processing unit and the low-level feature information are used as the inputs of the second processing unit at the same time. The output of the second processing unit and the feature information of this layer are used as the inputs of the second ReLU activation function layer. The first processing unit includes a 3×3 convolutional layer, a first batch normalization layer, and a first ReLU activation function layer connected in sequence. The second processing unit includes a 1×1 convolutional layer and a second batch normalization layer connected in sequence.
[0013] The number of channels of the mask tuple is the same as the vector dimension of the coefficient obtained by the target category prediction layer.
[0014] The total loss function of the target detection model includes a category loss function, a segmentation loss function, and a center point loss function. Among them, the center point loss function is a mask-guided center point error function.
[0015] The center point loss function is: Among them, m represents the number of positive samples; k represents the annotation index of the S×S block from left to right and from top to bottom, then i = [k / S], j = k mod S; I represents the indicator function, which is 1 when p i,j > 0, otherwise 0; represents the predicted center point position, c i represents the true center point position.
[0016] Beneficial effects
[0017] Due to the adoption of the above technical solutions, compared with the prior art, the present invention has the following advantages and positive effects: The present invention introduces a residual aliasing module to maintain low-level information to cope with the similarity between targets in the traffic sign recognition task, thereby improving the recognition accuracy. The present invention obtains an independent single target positioning result through the sum of the product of the target basis and the coefficient, which can improve the quality of the target basis and weaken the influence of low-quality target bases, thereby improving the aliasing phenomenon between targets existing in the current method. The present invention avoids the situation where the detection result deteriorates due to the neglected error in segmentation by introducing a mask-guided center point error. Brief description of the drawings
[0018] Figure 1 is a flowchart of an embodiment of the present invention;
[0019] Figure 2 is a network architecture diagram of the target detection model in an embodiment of the present invention;
[0020] Figure 3 is a schematic structural diagram of the feature extraction layer in an embodiment of the present invention;
[0021] Figure 4 It is a schematic diagram of the aliasing residual structure in the embodiment of the present invention. Specific embodiments
[0022] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0023] The embodiment of the present invention relates to an adaptive focusing and positioning target detection method, as Figure 1 shown, including the following steps: receiving an image to be recognized; inputting the image to be recognized into a target detection model to obtain the position and category of the target in the image to be recognized. This embodiment improves the target detection model, including: (1) introducing a residual aliasing module to maintain low-level information to improve the recognition accuracy; (2) obtaining an independent single target positioning result by the sum of the target basis multiplied by the coefficient, weakening the influence of low-quality target bases, thereby improving the aliasing phenomenon between targets existing in the current method; (3) introducing a mask-guided center point error to avoid the situation where the detection result deteriorates due to the error that is not taken seriously in the segmentation. Therefore, the detection method of this embodiment can be applied to the detection of traffic signs. The present invention will be further described below taking SOLOv2 as an example.
[0024] The main idea of SOLOv2 is to divide the picture into S*S blocks, and predict the category of each block (the one without a target is the background class). Each block also predicts an N-dimensional coefficient vector. SOLOv2 calls the coefficient vectors predicted by all blocks the kernel. Convolving the kernel with the mask branch generates the target mask. For large targets that span multiple blocks, SOLOv2 stipulates that the block responsible for predicting the target is determined by the center position of the target.
[0025] The existing network structure of SOLOv2 includes two parts, namely: a feature extraction part and a decoding part.
[0026] The feature extraction part is mainly composed of resnet50 plus FPN. Resnet consists of residual modules and downsampling operations. Resnet continuously extracts features through residual modules, and obtains a new layer of features with a larger receptive field through downsampling operations. FPN then integrates the features with a large receptive field into the low-level features.
[0027] The decoding part can be divided into two branches. One branch performs a chunking operation on each layer of the FPN output. Each chunk will perform class prediction and coefficient prediction. All the coefficients of each layer are the kernels. The other branch will generate a mask with the same number of channels as the dimension of the coefficient vector. The masks of each target are generated by convolving the kernel and the mask. One advantage of putting the kernel prediction and class prediction in the same branch is that the results of class prediction can be used to filter the predicted coefficients, filtering out some background-class chunks.
[0028] As Figure 2 shown, the object detection model in this embodiment includes: a feature extraction layer for extracting the features of the image to be recognized; an object class prediction layer for performing a chunking operation on each layer of features extracted by the feature extraction layer and performing class prediction and coefficient prediction on each chunk; an object localization prediction layer for generating a mask tuple according to the features extracted by the feature extraction layer and summing the product of the mask tuple and the coefficients obtained by the object class prediction layer to obtain an object mask. This embodiment improves the feature extraction layer and the object localization prediction layer in the network structure of the existing SOLOv2.
[0029] In this embodiment, the feature extraction layer is improved for Resnet. As Figure 3 shown, the feature extraction layer adds an aliased residual structure between adjacent residual structures to strengthen the transmission of detailed features, and maintains low-level information through the introduced residual aliasing structure to improve the recognition accuracy.
[0030] As Figure 4 shown, the aliased residual structure includes a first processing unit, a second processing unit, and a second ReLU activation function layer. The input of the first processing unit is the feature information of this layer. The output of the first processing unit and the low-level feature information are used as the inputs of the second processing unit at the same time. The output of the second processing unit and the feature information of this layer are used as the inputs of the second ReLU activation function layer. The first processing unit includes a 3×3 convolutional layer, a first batch normalization layer, and a first ReLU activation function layer connected in sequence. The second processing unit includes a 1×1 convolutional layer and a second batch normalization layer connected in sequence. This structure strengthens the detailed features of the input features by introducing low-level features into the current module and aliasing them with the input features passing through the first processing unit, and this structure is used between different layers of the encoder, so that the detailed features are retained and transmitted.
[0031] In the target location prediction layer of this embodiment, a mask tuple with the same number of channels as the vector dimension of the coefficient obtained by the target category prediction layer is generated based on the features extracted by the feature extraction layer, and the target mask is obtained by multiplying the mask tuple by the coefficient obtained by the target category prediction layer and then summing. The model tuple in this embodiment is similar to the concept of a basis in a vector space, so it can be regarded as a target basis. Just in a vector space, once the basis is determined, the vector has a unique representation. However, in a neural network, the number of channels is fixed while the number of targets in the input image is not, so it cannot be guaranteed that the masks of each channel are linearly independent or different. The inventor found that for an image, most channels do not help much with the detection result. Therefore, in this embodiment, the target mask is obtained by multiplying the mask tuple by the coefficient obtained by the target category prediction layer and then summing, which can improve the quality of the target basis, weaken the influence of low-quality target bases, and thus improve the aliasing phenomenon between targets existing in the existing SOLOv2 network.
[0032] In terms of the loss function, the target detection model in this embodiment introduces a mask-guided center point error. Therefore, the total loss function of the target detection model in this embodiment consists of a category loss function, a segmentation loss function, and a center point loss function, where the center point loss function is: where m represents the number of positive samples; k represents the annotation index for the S×S blocks from left to right and top to bottom, then i = [k / S], j = k mod S; I represents the indicator function, which is 1 when p i,j > 0 and 0 otherwise; represents the predicted center point position, c i represents the true center point position, c = (u c , v c ), and u, v represent pixel positions. By introducing the mask-guided center point error, the situation where the detection result deteriorates due to the errors that are not taken seriously in the segmentation is avoided.
[0033] It is not difficult to find that the present invention introduces a residual aliasing module to maintain low-level information to cope with the similarity between targets in the traffic sign recognition task, thereby improving the recognition accuracy. The present invention obtains an independent single target location result by the sum of the product of the target basis and the coefficient, which can improve the quality of the target basis, weaken the influence of low-quality target bases, and thus improve the aliasing phenomenon between targets existing in the current method. The present invention avoids the situation where the detection result deteriorates due to the errors that are not taken seriously in the segmentation by introducing the mask-guided center point error.
Claims
1. An adaptive focusing and positioning target detection method, characterized in that, It includes the following steps: Receive the image to be recognized; Input the image to be recognized into the target detection model to obtain the position and category of the target in the image to be recognized; Among them, the target detection model includes: A feature extraction layer for extracting the features of the image to be recognized; A target category prediction layer for performing a chunking operation on each layer of features extracted by the feature extraction layer, and performing category prediction and coefficient prediction on each block; A target location prediction layer for generating a mask tuple according to the features extracted by the feature extraction layer, and multiplying the mask tuple by the coefficients obtained by the target category prediction layer and then summing to obtain a target mask.
2. The adaptive focusing and positioning target detection method according to claim 1, wherein The feature extraction layer adds an aliasing residual structure between adjacent residual structures.
3. The adaptive focusing and positioning target detection method according to claim 2, wherein The aliasing residual structure includes a first processing unit, a second processing unit, and a second ReLU activation function layer. The input of the first processing unit is the feature information of this layer. The output of the first processing unit and the low-level feature information are used as the inputs of the second processing unit at the same time. The output of the second processing unit and the feature information of this layer are used as the inputs of the second ReLU activation function layer. The first processing unit includes a 3×3 convolutional layer, a first batch normalization layer, and a first ReLU activation function layer connected in sequence. The second processing unit includes a 1×1 convolutional layer and a second batch normalization layer connected in sequence.
4. The adaptive focusing and positioning target detection method according to claim 1, wherein The number of channels of the mask tuple is the same as the vector dimension of the coefficients obtained by the target category prediction layer.
5. The adaptive focusing and positioning target detection method according to claim 1, wherein, The total loss function of the target detection model includes a category loss function, a segmentation loss function, and a center point loss function. Among them, the center point loss function is a mask-guided center point error function.
6. The adaptive focusing and positioning target detection method according to claim 5, wherein, The center point loss function is as follows: where m represents the number of positive samples; k represents the annotation index for the S×S blocks from left to right and top to bottom, then i = [k / S], j = k mod S; I represents the indicator function, which is 1 when p i,j > 0 and 0 otherwise; represents the predicted center point position, and c i represents the true center point position.
Citation Information
Patent Citations
Text position positioning method and system and model training method and system
CN110414499A
Obstacle detection method for railway vehicle, computer equipment and storage medium
CN113936268A