A robust Chinese license plate detection and correction method in uncontrolled environment

By improving on the YOLOv5 framework, a robust Chinese license plate detection method was designed, which solved the problem of inaccurate positioning and slow speed of license plate detection in uncontrollable environments, and achieved fast and accurate license plate detection and correction, which improved the accuracy of identification.

CN114463611BActive Publication Date: 2025-05-13BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111557327.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-18
Publication Date
2025-05-13
Estimated Expiration
2041-12-18

AI Technical Summary

Technical Problem

In an uncontrollable environment, license plate detection has problems such as inaccurate positioning and slow speed, especially when the license plate has severe tilt or deformation, it is difficult for existing methods to accurately locate the license plate, which affects the accuracy of subsequent identification.

Method used

The robust Chinese license plate detection method based on the YOLOv5 framework is adopted to improve the accuracy and speed of the detection by establishing a license plate detection data set, input image preprocessing, designing improved network structure, including depth feature extraction and license plate coordinate position regression, and license plate correction.

Benefits of technology

It realizes rapid and accurate detection of license plates in an uncontrollable environment, can handle any inclined or deformed license plates, and improves the accuracy of subsequent license plate recognition, and has strong generalization and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114463611B_ABST
    Figure CN114463611B_ABST
Patent Text Reader

Abstract

A robust Chinese license plate detection and correction method in an uncontrolled environment belongs to the field of image processing. At present, most of the license plate detection methods use matrix frame positioning. In an uncontrolled environment, if the license plate is severely tilted or deformed, it will lead to inaccurate license plate positioning, that is, there are more backgrounds in the located license plate area or the positioning is incomplete, which will interfere with the subsequent license plate recognition and affect the recognition accuracy. The Chinese license plate detection method proposed by the present invention can improve the feature extraction ability of the model by introducing ACON, RBN and deformable convolution, improve the detection head and design the corresponding coordinate regression formula, which can accurately locate any tilted license plate and obtain ideal detection results in various complex uncontrolled environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and specifically relates to Chinese license plate detection, deep learning and other technologies. Background Art

[0002] The license plate number reflects the information of the vehicle and the owner. Accurately identifying the license plate number is a key step in intelligent transportation, and the accuracy of license plate detection greatly affects the accuracy of license plate recognition. At present, license plate detection and recognition have been widely used in some controllable environments, such as parking lots, highway toll intersections, etc. At present, most of the license plate detection methods use matrix frame positioning. In an uncontrolled environment, if the license plate is severely tilted or deformed, it will lead to inaccurate license plate positioning, that is, the located license plate area has a lot of background or the positioning is incomplete, which will interfere with the subsequent license plate recognition and affect the recognition accuracy.

[0003] Xu et al. constructed a lightweight network RPnet, which gives the coordinates of the license plate through regression in the last layer of the license plate positioning network. This method has a fast detection speed, but does not support multiple license plate detection, and the location of the license plate will be output even if there is no license plate network in the input image.

[0004] Silva et al. divided the license plate detection into two steps. First, the vehicle is detected by using YOLOV2 (You Only Look Once), and then the license plate is detected by WPOD Network. The detection head will output the affine transformation coefficient for subsequent license plate correction. This method can locate and correct the license plate and can detect multiple license plates, but the speed is relatively slow. Summary of the invention

[0005] Aiming at the problems of inaccurate positioning and slow speed in license plate detection under uncontrolled environment, this paper proposes a robust Chinese license plate detection method under uncontrolled environment. This method is implemented based on the YOLOv5 framework and mainly includes four steps: establishing a license plate detection dataset, input image preprocessing, network structure design, and license plate correction.

[0006] Step 1: Construction of license plate detection dataset

[0007] The performance of convolutional neural networks is based on a large amount of training data. In order to train the license plate detection network model, a license plate dataset needs to be established. The license plate dataset should contain license plate images under different environmental conditions to improve the robustness of detection.

[0008] Step 2: Input image preprocessing

[0009] Before sending the image to the network, it needs to be preprocessed, which mainly includes two steps:

[0010] (1) Normalization of input image size. Due to different acquisition devices, the size of license plate images is often inconsistent. Therefore, it is necessary to normalize the input license plate image size and adjust it to a uniform size through methods such as bilinear interpolation.

[0011] (2) Normalization of input image pixel values: Normalize all pixel values ​​in the input image to between 0 and 1 to make the network converge more easily.

[0012] Step 3: Network structure design

[0013] Step 3.1: Overall network architecture

[0014] The Chinese license plate detection network designed by the present invention is based on the YOLOv5 architecture. The original YOLOV5 network outputs a rectangular license plate, while the present invention outputs the coordinates of the four vertices of the license plate. The entire license plate detection framework mainly consists of two parts, namely deep feature extraction and license plate coordinate position regression.

[0015] Deep feature extraction

[0016] In order to ensure the speed and accuracy of detection, the present invention improves the backbone network of YOLOv5. The depth and width of the YOLOv5 backbone network, that is, the number of convolution layers and channels, are reduced. In addition, in order to enhance the backbone network's ability to extract and express features, the present invention replaces the BN (Batch Normalization) layer in the backbone network with the RBN (Representative Batch Normalization) layer. RBN can combine the individual features of each sample with the statistical features of each batch of samples, which can better adapt to the data; in addition, the activation function of the backbone network is replaced with ACON (Activate Or Not). The ACON activation function can adaptively choose whether to activate neurons, which can improve the performance of the network; deformable convolution is added to the lower layer of the backbone network, and the deformable convolution can better focus on the area around the feature points.

[0017] License plate coordinate position regression

[0018] The present invention improves the detection head of YOLOv5, and by changing the number of convolution channels of the detection head, the network can output the four vertex coordinate values ​​of the license plate. That is, the number of output elements of each anchor frame is increased by 8, and these 8 values ​​are the vertex coordinate values ​​of the license plate, and the coordinate value of the license plate is determined by regression.

[0019] Step 4: License Plate Correction

[0020] Since the license plate may be tilted or distorted, correcting the license plate is beneficial to the subsequent license plate recognition. According to the detected vertex coordinates of the license plate, the license plate image can be tilted by calculating the perspective transformation matrix.

[0021] Compared with the existing license plate detection method, the present invention has the following obvious advantages and effects:

[0022] 1. Fast detection speed and high accuracy;

[0023] 2. The detection result is the coordinates of the four vertices of the license plate, which can locate the license plate of any tilt and length, facilitating the subsequent correction of the tilted license plate;

[0024] 3. It has strong generalization and robustness, and can be applied to various complex and uncontrollable scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 Overall block diagram of license plate detection method

[0026] Figure 2 Backbone network structure

[0027] Figure 3 License Plate Correction Example DETAILED DESCRIPTION

[0028] The specific implementation of the present invention is described in detail below with reference to the accompanying drawings.

[0029] The overall block diagram of the Chinese license plate detection method proposed in this invention consists of four parts: input preprocessing, deep feature extraction, license plate coordinate position regression, and license plate correction. Figure 1 .

[0030] The implementation details of each step are as follows:

[0031] Step 1: Create a license plate detection dataset

[0032] The present invention obtains 100,000 license plate images by downloading from the Internet, collecting on-site data, and using existing data sets, and manually annotates the license plate areas therein to construct a license plate detection data set for training a deep convolutional neural network model.

[0033] Step 2: Input license plate preprocessing

[0034] Step 2.1: Normalize the input image size

[0035] Set the input image height to input h , width is input w , the actual height of the image is img h , the actual width is img w,If the size of the image is adjusted directly by ,downsampling and other methods, the proportion of the license plate in the image may change, affecting the ,accuracy of detection, so the bilinear interpolation plus padding method is ,adjusted to the image size to keep the aspect ratio of the license plate unchanged.

[0036] First, calculate the resizing factor, which is calculated as follows:

[0037]

[0038]

[0039] In formulas (1) and (2), r w Represents the width adjustment factor, r h Indicates the height adjustment factor.

[0040] Then, the image size after bilinear interpolation is calculated by the following formula:

[0041]

[0042]

[0043] Finally, the image with the size w′×h′ after bilinear interpolation is resized to the input by padding. w ×input h size.

[0044] Step 2.2: Normalize the pixel values ​​of the input image

[0045] Since the maximum value of each color channel of the license plate image is 255, the present invention normalizes the pixel value to between -1 and 1 through formula (5), and the calculation formula is as follows:

[0046]

[0047] Among them, x px is the original pixel value, is the normalized value.

[0048] Step 3: Overall network architecture

[0049] The license plate detection network architecture is mainly divided into two parts: deep feature extraction and license plate coordinate position regression.

[0050] Step 3.1: Deep feature extraction

[0051] As is known to all, the feature maps of different layers of convolutional neural networks have different sizes. In order to meet the detection requirements of license plates of different sizes, it is often necessary to perform detection on feature maps of different layers. In the present invention, the input image can obtain feature maps of three scales through a deep feature extraction network. The detection head performs detection on these three scale feature maps respectively, and after fusing the three detection results, the final license plate position is obtained.

[0052] (1) Backbone network

[0053] The backbone network structure of the present invention is as follows Figure 2 The parameters of each layer in the structure are shown in Table 1. The input image size of this part is (b, 3, input h , input w ), where b is the number of samples input to the network. The feature map sizes of CSP6_1, CSP7_1, and CSP8_1 are (b, 128, input h / 8,input w / 8)、(b,256,input h / 16,input w / 16) and (b,128,input h / 32,input w / 32). The present invention performs license plate detection on these feature maps respectively, and obtains the final license plate position after fusing the detection results.

[0054] Table 1 Parameters of each layer in the backbone network

[0055] Network Layer Kernel size Input Channels Output Channel Activation Function standardization Filling size Step Length Focus 3×3 12 32 ACON RBN 1 1 DCRA1 3×3 32 64 ACON RBN 1 2 DCSP1_1 - 64 64 ACON RBN - - DCRA2 3×3 64 128 ACON RBN 1 2 CSP2_3 - 128 128 ACON RBN - - CRA3 3×3 128 256 ACON RBN 1 2 CSP3_3 - 256 256 ACON RBN - - CRA4 3×3 256 512 ACON RBN 1 2 SPP - 512 512 ACON RBN - - CSP4_1 - 512 512 ACON RBN - - CRA5 1×1 512 256 ACON RBN 0 1 Unsample - - - - - - - Concat - - - - - - - CSP5_1 - 512 256 ACON RBN - - CRA6 1×1 256 128 ACON RBN 0 1 Concat - - - - - - - CSP6_1 - 256 128 ACON RBN - - CRA7 3×3 128 128 ACON RBN 1 2 Concat - - - - - - - CSP7_1 - 256 256 ACON RBN - - CRA8 3×3 256 256 ACON RBN 1 2 Concat - - - - - - - CSP8_1 - 512 512 ACON RBN - -

[0056] In Table 1, Unsample represents the upsampling layer; Concat represents the feature concatenation layer; SPP (Spatial Pyramid Pooling) represents the spatial pyramid pooling layer; CRA represents the layer composed of ordinary convolution, RBN and ACON, and the number after CRA represents the layer number; DCRA represents the layer composed of deformable convolution, RBN and ACON, and the number after DCRA represents the layer number; the first number of CSP1_1 represents the layer number 1, and the second number represents that the layer has 1 residual component, and the same applies to the others. DCSP is a CSP layer composed of deformable convolution. The parameters of each layer in CSP1_1 are shown in Table 2.

[0057] Table 2 Parameters of each layer in CSP1_1

[0058] Network Layer Kernel size Input Channels Output Channel Activation Function standardization Filling size Step Length Conv1 1×1 64 32 ACON RBN 0 1 Conv2 1×1 64 32 ACON RBN 0 1 Conv3 1×1 64 64 ACON RBN 0 1 Res uint - 32 32 ACON RBN - -

[0059] In Table 2, Conv is a common convolution, and the number after Conv represents the layer number; Res uint is a residual component, and the parameters of each layer are shown in Table 3.

[0060] Table 3 Parameters of each layer of Res uint in CSP1_1

[0061] Network Layer Kernel size Input Channels Output Channel Activation Function standardization Filling size Step Length Conv1 1×1 32 32 ACON RBN 0 1 Conv2 3×3 32 32 ACON RBN 1 1

[0062] (2) ACON activation function

[0063] The ACON activation function can adaptively choose whether to activate neurons. By replacing the activation function of the original network, the performance of the network can be improved.

[0064] The most common form of the ACON series activation function is ACON-C, which is expressed as follows:

[0065] ACON C =(p1-p2)x·σ(β(p1-p2)x)+p2x#(6)

[0066] Among them, x is the input of the activation function, σ is the Sigmoid function, and p1 and p2 are learnable parameters.

[0067] The expression of β is as follows:

[0068]

[0069] Among them, β is also a learnable parameter, C represents the number of channels of the input feature map, H and W represent the height and width of the input feature map respectively. c, h, and wd represent the channel index, height index, and width index respectively. The network is trained for 15 rounds, and the p1, p2, and β values ​​corresponding to the round with the highest accuracy are taken as the final values ​​of p1, p2, and β.

[0070] (3) RBN

[0071] The BN layer can accelerate the convergence of the model and reduce the possibility of gradient vanishing and explosion. However, it relies more on the mean and variance of the sample, ignoring the differences between each instance in the normalization process. RBN combines the unique features of each sample with the statistical features of each batch of samples, which can better adapt to the data. The following is an introduction to the algorithm flow of RBN.

[0072] First, center the input:

[0073] X cm =X+w m ⊙K m #(8)

[0074] Among them, X is the input feature, X cmis the feature after center calibration, w m is a learnable variable, K m Represent the characteristics of each instance, and then perform standardization:

[0075] X m =X cm -E(X cm )#(9)

[0076]

[0077] Among them, X m For X cm With X cm The difference between the means, E represents the mean, Var represents the variance, X s is the standardized feature, ∈ is a very small number with a value between 0 and 10 -8 To prevent zero variance, we then s Do a zoom calibration:

[0078] X cs =X s *R(w v ⊙K s +w b )#(11)

[0079] Among them, ⊙ is the dot product operator, R() is the restricted function, and w v 、w b is a learnable parameter. The network is trained for 15 rounds and the w corresponding to the round with the highest accuracy is taken. v 、w b The value is w v 、w b The final value of X cs Represents the features after scaling and calibration, and finally cs Do stretching and offsetting:

[0080] Y=γ*X cs +β′#(12)

[0081] Among them, Y is the output of RBN, γ and β′ are learnable parameters, the network is trained for 15 rounds, and the γ and β′ values ​​corresponding to the round with the highest accuracy are taken as the final values ​​of γ and β′.

[0082] (4) Deformable Convolution

[0083] The present invention adds deformable convolution to the lower layer of the backbone network, which can better focus on the area around the feature points, thereby improving the detection accuracy.

[0084] Let L represent the receptive field of the convolution kernel, and the number of elements N in L is the number of convolution kernel parameters. For example, L = [(-1,-1), (-1,0), ..., (0,1), (1,1)] represents the receptive field of a 3 × 3 convolution kernel, and N is 9. For each position p of the feature map 0 ,have:

[0085]

[0086] Among them, x is the input of the deformable convolution, p n is an element in L, y(p 0 ) is the position p 0 The result of convolution calculation using deformable convolution, Δp n is the offset, and w is the weight of the convolution kernel.

[0087] (5) Detection head

[0088] The backbone network outputs feature maps of three scales. When detecting license plates, three convolutional layers are used to perform convolution operations with the three scale feature maps respectively, and then the detection results of the three parts are spliced ​​together as the final detection output. The above three convolutional layers constitute the detection head, and the parameters of each layer are shown in Table 4. In addition, the number of output elements of each anchor frame is increased by 8. These 8 values ​​are the vertex coordinates of the license plate, and the coordinates of the license plate are determined by regression.

[0089] Table 4 Parameters of each layer of the detection head

[0090] Network Layer Kernel size Input Channels Output Channel Activation Function standardization Filling size Step Length Conv1 3×3 128 42 - - 1 1 Conv2 3×3 256 42 - - 1 1 Conv3 3×3 512 42 - - 1 1

[0091] Step 3.2: License plate coordinate position regression

[0092] The regression expression of the license plate coordinates is as follows:

[0093] x cd =((0.5-σ(px cd ))*4*aw+gridx)*stride#(14)

[0094] y cd =((0.5-σ(py cd ))*4*ah+gridy)*stride#(15)

[0095] In formulas (14) and (15), px cd ,py cdis the output value of the feature point, σ is the Sigmoid activation function, aw is the width of the anchor box relative to the current feature map, ah is the height of the anchor box relative to the current feature map, gridx and gridy are the horizontal and vertical coordinates of the current feature point, and stride is the multiple of the input feature map size relative to the current feature map size. σ(px cd ) is between 0 and 1. Since the vertices of the license plate are distributed in different directions of the current feature point, the offset is not necessarily a positive number, so 0.5 is subtracted from the activated value to make its range (-0.5, 0.5). In addition, the distance between the vertex of the license plate and the current feature point is not necessarily less than 0.5, so the value of the previous step is multiplied by 4 times the size of the anchor box, and finally the coordinates of the license plate in the current feature map are mapped to the input image.

[0096] Step 4: License Plate Correction

[0097] According to the detected vertex coordinates of the license plate, the license plate image can be tilted by calculating the perspective transformation matrix. The correction formula is as follows:

[0098]

[0099] Among them, x cd ,y cd is the coordinate before transformation, X′ cd , Y′ cd , Z′ cd is the transformed three-dimensional space coordinate, m ij (i, j = 1, 2, 3) are the matrix parameters of the perspective transformation.

[0100] The three-dimensional space coordinates are converted into two-dimensional coordinates by the following formula.

[0101]

[0102] x′ cd ,y′ cd is the converted two-dimensional coordinate. The corrected license plate image is more conducive to the subsequent license plate recognition, for example Figure 3 shown.

[0103] The Chinese license plate detection method proposed in the present invention can improve the feature extraction ability of the model by introducing ACON, RBN and deformable convolution, improve the detection head and design the corresponding coordinate regression formula, which can accurately locate arbitrarily tilted license plates and obtain ideal detection results in various complex uncontrollable environments.

Claims

1. A robust Chinese license plate detection and correction method in an uncontrolled environment, characterized by: Step 1: Construction of license plate detection dataset It is necessary to establish a license plate dataset; the license plate dataset should contain license plate images under different environmental conditions; Step 2: Input image preprocessing Before sending the image to the network, it needs to be preprocessed, which includes two steps: (1) Normalization of input image size; (2) Normalization of the pixel values ​​of the input image: Normalize all pixel values ​​in the input image to between 0 and 1 Step 3: Network structure design Step 3.1: Overall network architecture The Chinese license plate detection network is built on the basis of the YOLOv5 architecture, and the output is the coordinates of the four vertices of the license plate. The entire license plate detection framework consists of two parts, namely deep feature extraction and license plate coordinate position regression. Deep feature extraction Improvements were made to the backbone network of YOLOv5; the depth and width of the backbone network of YOLOv5, i.e. the number of convolutional layers and channels, were reduced; the BN (Batch Normalization) layer in the backbone network was replaced with the RBN (Representative Batch Normalization) layer; in addition, the activation function of the backbone network was replaced with ACON (Activate Or Not), and deformable convolution was added to the lower layers of the backbone network; 1) Backbone network The input image size of this part is (b, 3, input h , input w ), where b is the number of samples input to the network; the feature map sizes of the CSP6_1 layer, CSP7_1 layer, and CSP8_1 layer are (b, 128, input h / 8,input w / 8)、(b,256,input h / 16,input w / 16) and (b,128,input h / 32,input w / 32); perform license plate detection on these feature maps respectively, and fuse the detection results to obtain the final license plate position; 2) The activation function is the ACON activation function; The most widely used form of the ACON family of activation functions is ACON-C; License plate coordinate position regression By changing the number of convolution channels of the detection head, the network can output the four vertex coordinate values ​​of the license plate; that is, the number of output elements of each anchor frame is increased by 8, and these 8 values ​​are the vertex coordinate values ​​of the license plate, and the coordinate value of the license plate is determined by regression; Step 4: License Plate Correction According to the detected vertex coordinates of the license plate, the license plate image can be tilted by calculating the perspective transformation matrix.

2. The method according to claim 1, characterized in that: The implementation details of each step are as follows: Step 1: Create a license plate detection dataset Step 2: Input license plate preprocessing Step 2.1: Normalize the input image size The image size is adjusted by bilinear interpolation and padding to keep the aspect ratio of the license plate unchanged; First, calculate the resizing factor, which is calculated as follows: In formulas (1) and (2), r w Represents the width adjustment factor, r h represents the height adjustment factor; Then, the image size after bilinear interpolation is calculated by the following formula: Finally, the image with the size w′×h′ after bilinear interpolation is resized to the input by padding. w ×input h size; Step 2.2: Normalize the pixel values ​​of the input image Since the maximum value of each color channel of the license plate image is 255, the pixel value is normalized to between -1 and 1 using formula (5), as follows: Among them, x px is the original pixel value, is the normalized value; Step 3: Overall network architecture The license plate detection network architecture is mainly divided into two parts: deep feature extraction and license plate coordinate position regression; Deep feature extraction The input image can obtain feature maps of three scales through the deep feature extraction network. The detection head performs detection on these three scale feature maps respectively, and after fusing the three detection results, the final license plate position is obtained; 1) Backbone network The parameters of each layer in the backbone network structure are shown in Table 1. License plate detection will be performed on these feature maps respectively, and the final license plate position will be obtained after the detection results are fused. Table 1 Parameters of each layer in the backbone network In Table 1, Unsample represents the upsampling layer; Concat represents the feature concatenation layer; SPP (Spatial Pyramid Pooling) represents the spatial pyramid pooling layer; CRA represents the layer composed of ordinary convolution, RBN and ACON, and the number after CRA represents the layer number; DCRA represents the layer composed of deformable convolution, RBN and ACON, and the number after DCRA represents the layer number; the first number of CSP1_1 represents the layer number 1, and the second number represents that the layer has 1 residual component, and the same applies to the others; DCSP represents the CSP layer composed of deformable convolution; the parameters of each layer in CSP1_1 are shown in Table 2; Table 2 Parameters of each layer in CSP1_1 In Table 2, Conv is a common convolution, and the number after Conv represents the layer number; Resuint is a residual component, and the parameters of each layer are shown in Table 3; Table 3 Parameters of each layer of Resuint in CSP1_1 2) ACON activation function The most common form of the ACON series activation function is ACON-C, which is expressed as follows: ACON C =(p1-p2)x·σ(β(p1-p2)x)+p2x#(6) Among them, x is the input of the activation function, σ is the Sigmoid function, and p1 and p2 are learnable parameters; The expression of β is as follows: Among them, β is also a learnable parameter, C represents the number of channels of the input feature map, H and W represent the height and width of the input feature map respectively; c, h, wd represent the channel index, height index and width index respectively; the network is trained for 15 rounds, and the p1, p2, β values ​​corresponding to the round with the highest accuracy are taken as the final values ​​of p1, p2, β; 3) RBN First, center the input: X cm =X+w m ⊙K m #(8) Among them, X is the input feature, X cm is the feature after center calibration, w m is a learnable variable, K m Represent the characteristics of each instance, and then perform standardization: X m =X cm -E(X cm )#(9) Among them, X m For X cm With X cm The difference between the means, E represents the mean, Var represents the variance, X s is the standardized feature, ∈ is a very small number with a value between 0 and 10 -8 To prevent zero variance, we then s Do a zoom calibration: X cs =X s *R(w v ⊙K s +w b )#(11) Among them, ⊙ is the dot product operator, R() is the restricted function, and w v 、w b is a learnable parameter. The network is trained for 15 rounds and the w corresponding to the round with the highest accuracy is taken. v 、w b The value is w v 、w b The final value of X cs Represents the features after scaling and calibration, and finally cs Do stretching and offsetting: Y=γ*X cs +β′#(12) Among them, Y is the output of RBN, γ and β′ are learnable parameters, the network is trained for 15 rounds, and the γ and β′ values ​​corresponding to the round with the highest accuracy are taken as the final values ​​of γ and β′; 4) Deformable Convolution Deformable convolutions are added to the lower layers of the backbone network; Let L represent the receptive field of the convolution kernel, and the number of elements N in L is the number of convolution kernel parameters; for each position p0 of the feature map, we have: Among them, x is the input of the deformable convolution, p n is an element in L, y(p0) is the result of the convolution calculation using deformable convolution at position p0, Δp n is the offset, w is the weight of the convolution kernel; 5) Detection head The backbone network outputs feature maps of three scales. When detecting license plates, three convolutional layers are used to perform convolution operations with the three scale feature maps respectively, and then the detection results of the three parts are spliced ​​together as the final detection output; the above three convolutional layers constitute the detection head, and the parameters of each layer are shown in Table 4; in addition, the number of output elements of each anchor frame is increased by 8, and these 8 values ​​are the vertex coordinates of the license plate, and the coordinates of the license plate are determined by regression; Table 4 Parameters of each layer of the detection head Step 3.2: License plate coordinate position regression The regression expression of the license plate coordinates is as follows: x cd =((0.5-σ(px cd ))*4*aw+gridx)*stride#(14) y cd =((0.5-σ(py cd ))*4*ah+gridy)*stride#(15) In formulas (14) and (15), py cd ,py cd is the output value of the feature point, σ is the Sigmoid activation function, aw is the width of the anchor box relative to the current feature map, ah is the height of the anchor box relative to the current feature map, gridx and gridy are the horizontal and vertical coordinates of the current feature point, and stride is the multiple of the input feature map size relative to the current feature map size; σ(px cd ) takes a value between 0 and 1. Since the vertices of the license plate are distributed in different directions of the current feature point, the offset is not necessarily a positive number, so 0.5 is subtracted from the activated value to make its range (-0.5, 0.5); the distance between the vertex of the license plate and the current feature point is not necessarily less than 0.5, so the value of the previous step is multiplied by 4 times the size of the anchor box, and finally the coordinates of the license plate in the current feature map are mapped to the input image; Step 4: License Plate Correction According to the detected vertex coordinates of the license plate, the license plate image is tilted and corrected by calculating the perspective transformation matrix. The correction formula is as follows: Among them, x cd ,y cd is the coordinate before transformation, X′ cd , Y′ cd , Z′ cd is the transformed three-dimensional space coordinate, m ij (i, j = 1, 2, 3) are the matrix parameters of perspective transformation; The three-dimensional space coordinates are converted into two-dimensional coordinates by the following formula; x′ cd ,y′ cd is the transformed two-dimensional coordinate.

Citation Information

Patent Citations

  • Large-angle license plate inclination correction method based on end-to-end neural network

    CN110059683A

  • Efficient license plate positioning method of convolutional neural network

    CN111310773A