License Plate Recognition Method, Model Training Method, Electronic Device and Readable Storage Medium

By introducing Focal_ctc_loss loss function and adding license plate color recognition function in the LPRnet model, the problem of low Chinese license plate recognition accuracy is solved, and higher Chinese recognition accuracy and license plate color recognition effect are achieved.

CN113920511BActive Publication Date: 2025-07-11SUNELL TECH CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111156055.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-07-11
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

When dealing with Chinese license plate recognition, the existing LPRnet model has sample imbalance problem, resulting in low Chinese identification accuracy. Especially in Chinese license plates, the proportion of English and numbers far exceeds that of Chinese, affecting the recognition effect.

Method used

By introducing the Focal_ctc_loss loss function, combining the equilibrium factor α and γ, the network parameters of the LPRnet model are adjusted, and the license plate color recognition function is added to the head part of the model. Two output branches are used to perform license plate recognition and license plate color recognition respectively, and the combination of the full connection layer + Reshape + tile layer is replaced by a convolutional layer to improve the convergence and recognition accuracy of the model.

Benefits of technology

The LPRnet model has improved the recognition accuracy of Chinese license plates, and added the license plate color recognition function on the basis of license plate recognition, with the recognition accuracy higher than that of traditional HSV algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113920511B_ABST
    Figure CN113920511B_ABST
Patent Text Reader

Abstract

The embodiments of this application are applicable to the field of machine learning, and disclose a license plate recognition method, a model training method, an electronic device and a readable storage medium, which are used to improve the Chinese recognition accuracy of the LPRnet model when recognizing Chinese license plates. In the model training stage of the LPRnet model, the embodiments of this application use Focal_ctc_loss = α * (1 - p) γ * ctc_loss to balance the losses between Chinese samples and English and digital samples, where p = e ‑ctc_loss , α and γ are balance factors, α is used to balance the proportion of positive and negative samples, γ is used to adjust the rate of reduction of the weight of simple samples, and ctc_loss is the connectionist temporal classification loss value; after obtaining the trained LPRnet model, the trained LPRnet model is used to recognize the input license plate image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of machine learning, and particularly relates to a license plate recognition method, a model training method, an electronic device, and a computer-readable storage medium. Background Art

[0002] Currently, license plate recognition generally includes steps such as license plate location, license plate character segmentation, and license plate character recognition. Specifically, first, an image including the license plate is captured by a camera; then, the license plate area is located from the image including the license plate, and the license plate image is extracted; then, the extracted license plate image is subjected to character segmentation, and finally, the segmented characters are recognized to obtain the license plate recognition result.

[0003] License plate character recognition is a process of effectively confirming Chinese characters, letters, and numbers on the license plate based on accurate license plate location. This solution depends on the segmentation effect.

[0004] To better recognize license plates, the LPRnet model has been proposed. The LPRnet model does not require prior segmentation of the characters to be recognized and supports high-quality recognition of variable-length license plates that are template- and character-independent.

[0005] However, in Chinese license plates, the first digit of a regular license plate is the province, the second digit is English, and the subsequent digits are all combinations of English and numbers. This makes the proportion of English and numbers in the license plate sample distribution far exceed that of Chinese, that is, there is a problem of imbalance between Chinese samples and English and number samples. For example, for a license plate: Yue B*HX89, there is only one Chinese character in this license plate, but it includes multiple English letters and numbers. The problem of sample imbalance will cause the performance of the LPRnet model to be slightly weaker in terms of the accuracy of recognizing Chinese provinces. Summary of the Invention

[0006] The embodiments of this application provide a license plate recognition method, a model training method, an electronic device, and a computer-readable storage medium, which can improve the Chinese recognition accuracy of the LPRnet model for Chinese license plates.

[0007] In a first aspect, the embodiments of this application provide a model training method, including:

[0008] Obtain a training data set, where the training data set includes at least one license plate image and the label of each license plate image, and the label includes a license plate label;

[0009] Input the license plate image into a pre-constructed LPRnet model to obtain the output result of the LPRnet model. The output result includes the license plate recognition result output by the first output branch. The LPRnet model includes a first output branch;

[0010] Calculate the connection timing classification loss value according to the license plate recognition result and the license plate label;

[0011] Calculate the first loss value through the formula Focal_ctc_loss = α * (1 - p) γ * ctc_loss; where p = e -ctc_loss , α and γ are balance factors, and ctc_loss is the connection timing classification loss value;

[0012] Adjust the network parameters of the LPRnet model according to the first loss value;

[0013] Iteratively train multiple times to obtain the trained LPRnet model.

[0014] As can be seen from the above, in the model training stage of this application embodiment, Focal_ctc_loss = α * (1 - p) γ * ctc_loss is used to balance the losses between Chinese samples and English and digital samples, so that the trained LPRnet model has a higher recognition accuracy when recognizing Chinese characters on license plates.

[0015] Specifically, based on ctc_loss, Focal_ctc_loss introduces balance factors α and γ. Among them, α is used to balance the imbalance of positive and negative sample ratios, and γ is used to adjust the rate of reduction of the weights of simple samples.

[0016] In some possible implementation manners of the first aspect, the above label further includes a license plate color label, and the above output result further includes the license plate color output by the second output branch;

[0017] The LPRnet model includes a backbone network and a head part, and the head part includes a first output branch and a second output branch;

[0018] The above method further includes: calculating a second loss value according to the license plate color and the license plate color label;

[0019] The above adjusting the network parameters of the LPRnet model according to the first loss value includes: calculating a total loss value according to the first loss value and the second loss value; adjusting the network parameters of the LPRnet model according to the total loss value.

[0020] In this implementation manner, the head part of the LPRnet model includes a first output branch and a second output branch, that is, it includes two output branches, and these two output branches are respectively used to output the license plate color prediction result and the license plate recognition prediction result.

[0021] However, the head part of the existing LPRnet model has only one output branch and can only obtain the license plate recognition prediction result. For license plate color, only traditional recognition algorithms (such as the HSV algorithm) can be used.

[0022] In contrast, in the embodiments of the present application, by improving the head part, on the basis of license plate recognition, the function of license plate color recognition is added, so that the LPRnet model can output both the license plate recognition result and the license plate color recognition result. In addition, the head part includes two output branches, which is more conducive to the convergence of the network model and the improvement of the recognition accuracy.

[0023] In some possible implementation manners of the first aspect, the head part includes a first convolutional layer, a first activation function, a Concat layer, a first output branch, and a second output branch;

[0024] The Concat layer is used to connect the output of the backbone network and the output of the first activation function; both the first output branch and the second output branch are connected to the Concat layer;

[0025] The first output branch includes a second convolutional layer, a second activation function, and a first Transpose layer connected in sequence;

[0026] The second output branch includes a third convolutional layer, a third activation function, a classification layer, and a second Transpose layer connected in sequence.

[0027] In some possible implementation manners of the first aspect, calculating the total loss value according to the first loss value and the second loss value includes:

[0028] Calculating the total loss value through the formula Loss = Focal_ctc_loss + 2bce_loss;

[0029] where bce_loss is the second loss value and Focal_ctc_loss is the first loss value.

[0030] In this implementation manner, after adding the loss value of the classification cross-entropy loss function (Binary Cross Entropy loss, bceloss) and Focal_ctc_loss in a ratio of 2:1 and backpropagating it into the LPRnet model to update the gradient, compared with other ratios (such as 1:1), the LPRnet model trained by the former has higher recognition accuracy.

[0031] In some possible implementation manners of the first aspect, α is 0.95 and γ is 1. In this way, the recognition accuracy of the trained LPRnet model is higher.

[0032] Second aspect, an embodiment of the present application provides a license plate recognition method, including:

[0033] Obtain a license plate image to be recognized;

[0034] Input the license plate image to be recognized into a pre-trained LPRnet model to obtain the output result of the LPRnet model. The output result includes the license plate recognition result output by the first output branch. The LPRnet model includes a first output branch;

[0035] The above LPRnet model is a model trained using the model training method of any item in the first aspect above.

[0036] In some possible implementation manners of the second aspect, the LPRnet model includes a backbone network and a head part. The head part includes a first output branch and a second output branch;

[0037] The output result further includes the license plate color output by the second output branch.

[0038] In some possible implementation manners of the second aspect, the head part includes a first convolutional layer, a first activation function, a Concat layer, a first output branch, and a second output branch;

[0039] The Concat layer is used to connect the output of the backbone network and the output of the first activation function; both the first output branch and the second output branch are connected to the Concat layer;

[0040] The first output branch includes a second convolutional layer, a second activation function, and a first Transpose layer connected in sequence;

[0041] The second output branch includes a third convolutional layer, a third activation function, a classification layer, and a second Transpose layer connected in sequence.

[0042] Third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any item of the first aspect or the second aspect above is implemented.

[0043] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the method described in any item of the first aspect or the second aspect above is implemented.

[0044] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device is caused to execute the method described in any one of the above first aspect or second aspect.

[0045] It can be understood that the beneficial effects of the above second aspect to fifth aspect can be referred to the relevant descriptions in the above first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 Schematic diagram of the network structure of the existing head part provided by the embodiment of the present application;

[0048] Figure 2 Schematic diagram of the improved network structure of the head part provided by the embodiment of the present application;

[0049] Figure 3 Schematic block diagram of a process of the model training method provided by the embodiment of the present application;

[0050] Figure 4 Another schematic block diagram of the process of the model training method provided by the embodiment of the present application;

[0051] Figure 5 Schematic diagram of a process of the license plate recognition method provided by the embodiment of the present application;

[0052] Figure 6 Schematic block diagram of the structure of the model training device provided by the embodiment of the present application;

[0053] Figure 7 Schematic block diagram of the structure of the license plate recognition device provided by the embodiment of the present application;

[0054] Figure 8 Schematic diagram of the structure of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application.

[0056] The LPRnet model involved in the embodiments of the present application will be introduced below.

[0057] In the embodiments of this application, the model structure of the existing LPRnet model is improved to obtain the improved LPRnet model. To better introduce the differences between the improved LPRnet model and the existing LPRnet model, the model structure of the existing LPRnet model will be introduced first, and then the improved LPRnet model will be introduced.

[0058] (1) Existing LPRnet model.

[0059] The LPRnet model includes a backbone network and a head part. The backbone network is used to extract features from the input image, that is, to extract information from the image for use by the subsequent network.

[0060] The head part refers to the network connected behind the backbone network, which is used to make predictions using the features extracted by the backbone network, obtain the prediction results, and use the prediction results as the output of the model.

[0061] Exemplarily, the network structure of the backbone network can be as shown in Table 1 below:

[0062] Table 1

[0063] Layer Type Parameter Input Layer 94x24 RGB Image Convolution Layer #64 3x3 stride 1 Max Pooling Layer #64 3x3 stride 1 Small Basic Block #128 3x3 stride 1 Max Pooling Layer #64 3x3 stride(2,1) Small Basic Block #256 3x3 stride 1 Small Basic Block #256 3x3 stride 1 Max Pooling Layer #64 3x3 stride(2,1) Dropout Layer 0.5 ratio Convolution Layer #256 4x1 stride 1 Dropout Layer 0.5 ratio Convolution Layer #class_number 1x13 stride 1

[0064] As can be seen from Table 1 above, the backbone network includes an input layer, a convolutional layer, a max pooling layer, a small basic block, a max pooling layer, a small basic block, a small basic block, a max pooling layer, a Dropout layer, a convolutional layer, a Droupout layer, and a convolutional layer connected in sequence.

[0065] Exemplarily, the network structure of the head part can be as Figure 1 shown. Figure 1 Only a part of the backbone network part in [] is shown, where the convolutional layer of the backbone network is the convolutional layer corresponding to class_number 1x13 stride1 in Table 1 above, and class_number refers to the total number of license plate characters.

[0066] Figure 1 The head part in [] includes an InnerProduct layer, a Reshape function, a Tile layer, a Concat layer, a convolutional layer, a Permute layer, and a Reshape function. The InnerProduct layer is a fully connected layer. At this time, there is only one network output in the head part, and this output is specifically the license plate recognition result.

[0067] (2) Improved LPRnet model.

[0068] In some embodiments, the backbone network of the LPRnet model is improved. Specifically, in order to reduce the size of the intermediate feature maps and the computational cost of the total inference, and to reduce the network parameters, the convolutional kernels of the second max pooling layer and the third max pooling layer in the backbone network are changed from the original 2*1 to 2*2.

[0069] In order to increase the receptive field and the number of parameters, so that the feature representation is more obvious, the convolutional kernel of the second last convolutional layer in the backbone network is changed from the original 4*1 to 4*4.

[0070] Exemplarily, the backbone network structure of the improved LPRnet model can be as shown in Table 2 below.

[0071] Table 2

[0072] Layer Type Parameter Input Layer 94x24 RGB Image Convolution Layer #64 3x3 stride 1 Max Pooling Layer #64 3x3 stride 1 Small Basic Block #128 3x3 stride 1 Max Pooling Layer #64 3x3 stride(2,2) Small Basic Block #256 3x3 stride 1 Small Basic Block #256 3x3 stride 1 Max Pooling Layer #64 3x3 stride(2,2) Dropout Layer 0.5 ratio Convolution Layer #256 4x4 stride 1 Dropout Layer 0.5 ratio Convolution Layer #class_number 1x13 stride 1

[0073] It can be understood that when improving the backbone network, any one of the following operations or any combination can be performed: changing the convolutional kernels of the second max pooling layer and the third max pooling layer in the backbone network from the original 2*1 to 2*2; changing the convolutional kernel of the second last convolutional layer in the backbone network from the original 4*1 to 4*4.

[0074] In some embodiments, an additional output branch is added to the head part for outputting the license plate color prediction result. In this way, the LPRnet model can output both the license plate recognition result and the license plate color prediction result.

[0075] Exemplarily, referring to Figure 2 the schematic diagram of the improved head part shown, the head part includes a first output branch and a second output branch. The first output branch includes a convolutional layer, a Relu function, and a Transpose layer connected in sequence. The first output branch outputs the license plate recognition result, Figure 2 where 70 in 1x1x18x70 represents the total number of all characters.

[0076] The second output branch includes a convolutional layer, a Relu function, a Sigmoid function, and a Transpose layer connected in sequence. The second output branch outputs the license plate color, Figure 2 where 5 in 1x1x1x5 represents 5 license plate colors.

[0077] The convolutional layers in the first output branch and the second output branch perform feature extraction respectively, and the weights are not shared, which can improve the accuracy of their respective feature extraction.

[0078] The head part includes two output branches, which is more conducive to the convergence of the network model and the improvement of the recognition accuracy.

[0079] in addition, Figure 2 The head part in the network is also replaced by a convolutional layer. Figure 1 The head part is a combination of fully connected layers + Reshape + tile layers. This makes it easier to deploy the LPRnet model on the device, and is more conducive to the transplantation and deployment of the model on the device. For devices that do not support the Tile layer, by replacing the fully connected layers + Reshape + tile layers with convolutional layers, the LPRnet model can also be deployed on devices that do not support the Tile layer.

[0080] Of course, in other embodiments, the convolution layer may be directly used to replace the tile layer.

[0081] After testing, the model accuracy was not lost when the convolutional layer was used to replace the combination of the fully connected layer + Reshape + tile layer.

[0082] After introducing the LPRnet model, the training process of the LPRnet model is introduced below.

[0083] See also Figure 3 , is a schematic flow chart of a model training method provided in an embodiment of the present application, and the method may include the following steps:

[0084] Step S301, obtaining a training data set, the training data set including at least one license plate image and a label of each license plate image, the label including a license plate label.

[0085] It is understandable that a corresponding license plate label is added to each license plate image to represent the license plate in the license plate image.

[0086] For example, if a license plate image includes the license plate of Yue B*HX89, the license plate label added to the license plate image is Yue B*HX89.

[0087] Step S302: input the license plate image into the pre-built LPRnet model to obtain the output result of the LPRnet model, the output result includes the license plate recognition result output by the first output branch, and the LPRnet model includes the first output branch.

[0088] It should be noted that the above-mentioned LPRnet model can be an existing LPRnet model; it can also be an improved LPRnet model, and the improved part can be the backbone network and / or the head part. The model structure of the LPRnet model here is not limited.

[0089] For example, the above LPRnet model is an improved LPRnet model. The specific improvement is: the convolution kernels of the second maximum pooling layer and the third maximum pooling layer in the backbone network are changed from the original 2*1 to 2*2.

[0090] For example, the above LPRnet model is an improved LPRnet model, and the specific improvement is: replacing the fully connected layer + Reshape + tile layer in the head part with a convolutional layer.

[0091] It can be understood that if the output branch of the head part of the LPRnet model is not improved, that is, the model has only one output branch, the above first output branch refers to the output branch of the model. If the head part of the LPRnet model includes two output branches, the above first output branch refers to one of the two output branches.

[0092] Step S303: Calculate the Connectionist Temporal Classification loss (CTC loss) according to the license plate recognition result and the license plate label.

[0093] Exemplarily, the size of the license plate recognition result output by the LPRnet model is (N, T, C). N refers to the number of input images, and T refers to the length of the output sequence. C includes the total number of Chinese provinces, English letters, and digital categories.

[0094] Since Focal_ctc_loss assumes that in the worst case, there is at least one blank label before and after each true label to separate duplicates, it is generally taken as 2*n + 1 in length of the license plate.

[0095] In the embodiments of the present application, since the length n of the Chinese license plate may be 8 digits or 7 digits. Therefore, T can be set to 18.

[0096] After transposing (N, T, C) to (T, N, C), then perform log_softmax on it, that is, Log(softmax(x)), to prevent data overflow.

[0097] The Softmax function can be shown as follows:

[0098]

[0099] Among them, represents the input of the j-th neuron in the L-th layer (usually the last layer), represents the output of the j-th neuron in the L-th layer, represents the sum of the inputs of all neurons in the L-th layer.

[0100] The calculation of the CTC loss can adopt the following formula:

[0101]

[0102] Among them, πt is an element in sequence O, y is the probability that the model outputs each character at all times, with a shape of T*C, where T is the time step and C is the number of character categories, equal to all characters + blank, and it is the probability that the model outputs πt at time t.

[0103] It should be noted that the calculation process of CTC loss is well-known to those skilled in the art. The above is only an example and will not be elaborated here.

[0104] Step S304: Calculate the first loss value through the formula Focal_ctc_loss = α * (1 - p) γ *ctc_loss; where p = e -ctc_loss , and α and γ are balancing factors, and ctc_loss is the connectionist temporal classification loss value.

[0105] It should be noted that based on ctc loss, Focal_ctc_loss introduces a balancing factor α and γ. Adding the balancing factor α is used to balance the problem of the unbalanced ratio of positive and negative samples itself; adding the balancing factor γ is used to adjust the rate at which the weights of simple samples decrease. When γ is 0, it is ctc loss; when γ increases, the influence of the adjustment factor also increases. In some embodiments, α is 0.95 and γ is 1.

[0106] Step S305: Adjust the network parameters of the LPRnet model according to the first loss value.

[0107] It should be noted that if the LPRnet model has only one output branch, only according to the first loss value Focal_ctc_loss, it is backpropagated into the model to update the gradient to adjust the model network parameters. If the LPRnet model has two output branches, the network parameters are updated according to the first loss value and the second loss value.

[0108] Step S306: Iteratively train multiple times to obtain the trained LPRnet model.

[0109] Specifically, loop through the above steps S302 to S305 until the loss of the model tends to be stable, then it can be considered that the model training is completed, and the trained LPRnet model is obtained.

[0110] It can be seen that due to using Focal_ctc_loss during the model training stage to balance the problem of unbalanced samples, the trained LPRnet model has a higher Chinese character recognition accuracy when recognizing Chinese license plates.

[0111] Based on the above embodiments, in some other embodiments, the above-mentioned label further includes a license plate color label, and the above output result further includes the license plate color output by the second output branch.

[0112] At this time, the LPRnet model includes a backbone network and a head part. The head part includes a first output branch and a second output branch. That is, the LPRnet model includes two output branches, one for predicting the license plate color and one for license plate recognition.

[0113] In some embodiments, the head part includes a first convolutional layer, a first activation function, a Concat layer, a first output branch, and a second output branch; the Concat layer is used to connect the output of the backbone network and the output of the first activation function; both the first output branch and the second output branch are connected to the Concat layer; the first output branch includes a second convolutional layer, a second activation function, and a first Transpose layer connected in sequence; the second output branch includes a third convolutional layer, a third activation function, a classification layer (such as Figure 2 the sigmoid function) and a second Transpose layer connected in sequence. At this time, the first output branch and the second output branch can be as Figure 2 shown.

[0114] See Figure 4 Another flow schematic block diagram of the model training method shown. This method may include the following steps:

[0115] Step S401, obtain a training data set. The training data set includes at least one license plate image and the label of each license plate image. The label includes a license plate label and a license plate color label.

[0116] Step S402, input the license plate image into the pre-constructed LPRnet model to obtain the output result of the LPRnet model. The output result includes the license plate recognition result output by the first output branch and the license plate color output by the second output branch. The LPRnet model includes a first output branch and a second output branch.

[0117] Step S403, calculate the connection temporal classification loss value according to the license plate recognition result and the license plate label.

[0118] Step S404, calculate the first loss value through the formula Focal_ctc_loss = α * (1 - p) γ * ctc_loss; where p = e -ctc_loss , α and γ are balance factors, and ctc_loss is the connection temporal classification loss value.

[0119] It should be noted that for the same parts of the embodiments of the present application as those in the above embodiments, please refer to the above, and details will not be repeated here.

[0120] Step S405: Calculate the second loss value according to the license plate color and the license plate color label.

[0121] Step S406: Calculate the total loss value according to the first loss value and the second loss value.

[0122] In some embodiments, the total loss value is calculated by the formula Loss = Focal_ctc_loss + 2bce_loss; where bce_loss is the second loss value and Focal_ctc_loss is the first loss value.

[0123] At this time, the multi-class cross-entropy (Binary Cross Entropy, BCE) loss function is adopted to calculate the loss between the predicted license plate color and the license plate color label.

[0124] Exemplarily, the calculation process of the BCE loss can be as follows: First, calculate through the sigmoid function, and then input the output of the sigmoid function into the loss calculation function.

[0125] The sigmoid function can be as follows:

[0126]

[0127] where x is the input.

[0128] The loss calculation function can be as follows:

[0129]

[0130] where L is the above-mentioned second loss value, that is, the loss value of the license plate color. y i is the category to which the i-th sample belongs, and p i is the predicted value of the i-th sample.

[0131] It should be noted that the weight ratio of Focal_ctc_loss and bce_loss is 1:2, and the accuracy is higher than that with a weight ratio of 1:1.

[0132] Of course, the ratio between the two loss values is arbitrary and is not limited here.

[0133] It should also be noted that for license plate color recognition, the license plate color recognition method in the embodiments of the present application has higher recognition accuracy than the traditional HSV algorithm. In addition, the embodiments of the present application adopt the multi-class cross-entropy function with the sigmoid activation function as the cost calculation function. Compared with the multi-class cross-entropy loss using softmax, the former has higher accuracy and faster running speed on the device.

[0134] Step S407: Adjust the network parameters of the LPRnet model according to the total loss value.

[0135] Step S408: Iteratively train multiple times to obtain a trained LPRnet model.

[0136] It can be seen that the embodiment of the present application not only balances the problem of sample imbalance through Focal_ctc_loss to improve the Chinese recognition accuracy of the model, but also changes the head part of the LPRnet model, enabling the LPRnet model to add a license plate color recognition function on the basis of license plate recognition. Moreover, the license plate color recognition function has higher accuracy than the traditional HSV recognition method.

[0137] It should be noted that in some other embodiments, the head part of the LPRnet model may include a first output branch and a second output branch, and these two output branches may be specifically as Figure 2 shown. At this time, when training the LPRnet model, for the loss of the license plate recognition result, instead of using the above-mentioned Focal_ctc_loss, ctc_loss may be used. And after calculating the loss of the license plate recognition result and the loss of the license plate color, calculate the total loss value according to these two loss values, and then use the total loss value to iteratively update the model network parameters.

[0138] After introducing the training process of the LPRnet model, the application stage of the LPRnet model will be exemplarily introduced below.

[0139] Refer to Figure 5 , which is a schematic flowchart of a license plate recognition method provided by the embodiment of the present application. The method may include the following steps:

[0140] Step S501: Obtain a license plate image to be recognized.

[0141] Exemplarily, after capturing a license plate image through a camera, scale the captured license plate image to 94*24; then, normalize the scaled license plate image, subtract the mean of 127.5, and divide by the variance of 128, which is beneficial to reducing the influence of illumination and other brightness. In this way, the above-mentioned license plate image to be recognized is obtained.

[0142] Step S502: Input the license plate image to be recognized into the pre-trained LPRnet model to obtain the output result of the LPRnet model. The output result includes the license plate recognition result output by the first output branch. The LPRnet model includes a first output branch.

[0143] At this time, the above LPRnet model is the model trained by using the model training method of the above-mentioned various embodiments.

[0144] It should be noted that the LPRnet model can be an existing model or an improved model. When it is an existing model, the output of the model is only the license plate recognition result. At this time, due to balancing the problem of sample imbalance through Focal_ctc_loss during the model training stage, the Chinese recognition accuracy of the model is higher.

[0145] If it is an improved model and the model includes two output branches, the output of the model has two outputs, namely the license plate recognition result and the license plate color respectively.

[0146] At this time, the LPRnet model includes a backbone network and a head part. The head part includes a first output branch and a second output branch; the output result also includes the license plate color output by the second output branch.

[0147] Furthermore, the head part includes a first convolutional layer, a first activation function, a Concat layer, a first output branch and a second output branch; the Concat layer is used to connect the output of the backbone network and the output of the first activation function; both the first output branch and the second output branch are connected to the Concat layer; the first output branch includes a second convolutional layer, a second activation function and a first Transpose layer connected in sequence; the second output branch includes a third convolutional layer, a third activation function, a classification layer and a second Transpose layer connected in sequence.

[0148] Exemplarily, after obtaining the license plate image to be recognized, the license plate image to be recognized is input into the improved LPRnet model, and outputs of two sizes of N*1*1*5 and N,T,C are obtained. Among them, 5 in the first group of outputs represents 5 different colors of the license plate, N represents the number of input images, and T represents the length of the output sequence. T is set to 18; C includes the total number of Chinese provinces, English letters and digital categories.

[0149] Then, for the output license plate color result, the predicted value with the highest score is selected as the color of the license plate; for the data of N,T,C size, in each array of size (T,C), the one with the largest C score is selected as the output letter, then the value of N*T is obtained, and in each data with a length of T, the positions of the spaces are excluded, and finally the variable-length data is obtained as the license plate recognition result.

[0150] It should be noted that for the improvement of the model structure of the LPRnet model and the specific effects after improvement, reference can be made to the content about the model above, which will not be elaborated here.

[0151] It can be seen that in the embodiment of the present application, the problem of sample imbalance is balanced through Focal_ctc_loss during model training to improve the Chinese recognition accuracy of the model.

[0152] Furthermore, the head part of the LPRnet model is also changed, enabling the LPRnet model to add a license plate color recognition function on the basis of license plate recognition, and the license plate color recognition function has higher accuracy than the traditional HSV recognition method.

[0153] It should be noted that in some other embodiments, if the LPRnet model includes two output branches and ctc_loss is used instead of Focal_ctc_loss during the model training phase. At this time, when using the trained LPRnet model for license plate recognition, the LPRnet model adds a license plate color recognition function on the basis of license plate recognition, and the license plate color recognition function has higher accuracy than the traditional HSV recognition method.

[0154] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0155] Corresponding to the model training method described in the above embodiments, Figure 6 The structural block diagram of the model training device provided by the embodiments of the present application is shown. For the convenience of description, only the parts related to the embodiments of the present application are shown.

[0156] Referring to Figure 6 , the device includes:

[0157] An acquisition module 61, configured to acquire a training data set, where the training data set includes at least one license plate image and the label of each license plate image, and the label includes a license plate label;

[0158] An input module 62, configured to input the license plate image into a pre-constructed LPRnet model to obtain an output result of the LPRnet model. The output result includes a license plate recognition result output by a first output branch, and the LPRnet model includes a first output branch;

[0159] A loss value calculation module 63, configured to calculate a connectionist temporal classification loss value according to the license plate recognition result and the license plate label;

[0160] A first loss value calculation module 64, configured to calculate a first loss value through the formula Focal_ctc_loss = α * (1 - p) γ * ctc_loss; where p = e -ctc_loss , α and γ are balance factors, and ctc_loss is the connectionist temporal classification loss value;

[0161] An adjustment module 65, configured to adjust the network parameters of the LPRnet model according to the first loss value;

[0162] An iterative module 66 for iteratively training multiple times to obtain a trained LPRnet model.

[0163] Specifically, based on the ctc_loss, Focal_ctc_loss introduces the balance factors α and γ. Among them, α is used to balance the imbalance of positive and negative sample ratios, and γ is used to adjust the rate of reduction of the weights of easy samples.

[0164] In some possible implementation manners, the above tags further include license plate color tags, and the above output result further includes the license plate color output by the second output branch.

[0165] The LPRnet model includes a backbone network and a head part, and the head part includes a first output branch and a second output branch;

[0166] The above device further includes: a second loss value calculation module for calculating a second loss value according to the license plate color and the license plate color tag;

[0167] The above adjustment module 65 is specifically configured to: calculate a total loss value according to the first loss value and the second loss value; and adjust the network parameters of the LPRnet model according to the total loss value.

[0168] In some possible implementation manners, the head part includes a first convolutional layer, a first activation function, a Concat layer, a first output branch, and a second output branch; the Concat layer is used to connect the output of the backbone network and the output of the first activation function; both the first output branch and the second output branch are connected to the Concat layer; the first output branch includes a second convolutional layer, a second activation function, and a first Transpose layer connected in sequence; the second output branch includes a third convolutional layer, a third activation function, a classification layer, and a second Transpose layer connected in sequence.

[0169] In some possible implementation manners, the adjustment module 65 is specifically configured to: calculate the total loss value through the formula Loss = Focal_ctc_loss + 2bce_loss; where bce_loss is the second loss value and Focal_ctc_loss is the first loss value.

[0170] In some possible implementation manners, α is 0.95 and γ is 1.

[0171] It should be noted that the information interaction, execution process, etc. among the above devices / modules, due to being based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.

[0172] Corresponding to the license plate recognition method described in the above embodimentFigure 7 The block diagram of the license plate recognition device provided by the embodiment of the present application is shown. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.

[0173] Refer to Figure 7 , the device includes:

[0174] A license plate image acquisition module 71, configured to acquire a license plate image to be recognized;

[0175] A license plate recognition module 72, configured to input the license plate image to be recognized into a pre-trained LPRnet model, and obtain the output result of the LPRnet model. The output result includes the license plate recognition result output by the first output branch. The LPRnet model includes a first output branch;

[0176] The above LPRnet model is a model trained by using the model training method of any item in the first aspect as described above.

[0177] In some possible implementation manners, the LPRnet model includes a backbone network and a head part. The head part includes a first output branch and a second output branch; the output result further includes the license plate color output by the second output branch.

[0178] In some possible implementation manners, the head part includes a first convolutional layer, a first activation function, a Concat layer, a first output branch, and a second output branch; the Concat layer is used to connect the output of the backbone network and the output of the first activation function; both the first output branch and the second output branch are connected to the Concat layer; the first output branch includes a second convolutional layer, a second activation function, and a first Transpose layer connected in sequence; the second output branch includes a third convolutional layer, a third activation function, a classification layer, and a second Transpose layer connected in sequence.

[0179] It should be noted that for the information interaction, execution process, etc. between the above-mentioned device / modules, since they are based on the same concept as the method embodiment of the present application, their specific functions and the technical effects brought, for details, please refer to the method embodiment part, and will not be elaborated here.

[0180] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, the electronic device 8 of this embodiment includes: at least one processor 80 ( Figure 8 only one is shown in the figure), a memory 81, and a computer program 82 stored in the memory 81 and executable on the at least one processor 80. When the processor 80 executes the computer program 82, the steps in any of the above-mentioned embodiment of the charging pile recognition method are implemented.

[0181] The electronic device may include, but is not limited to, a processor 80 and a memory 81. Those skilled in the art can understand that Figure 8 This is merely an example of the electronic device 8 and does not constitute a limitation on the electronic device 8. It may include more or fewer components than those shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0182] The so-called processor 80 may be a central processing unit (CPU), and the processor 80 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0183] In some embodiments, the memory 81 may be an internal storage unit of the electronic device 8, such as the hard disk or memory of the electronic device 8. In other embodiments, the memory 81 may also be an external storage device of the electronic device 8, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the electronic device 8. Further, the memory 81 may also include both the internal storage unit and the external storage device of the electronic device 8. The memory 81 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program, etc. The memory 81 may also be used to temporarily store the data that has been output or will be output.

[0184] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0185] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0186] The embodiment of this application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0187] The embodiment of this application provides a computer program product, when the computer program product runs on an electronic device, it enables the electronic device to implement the steps in the above-mentioned method embodiments when executed.

[0188] It should be understood that, as used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their combinations.

[0189] It should also be understood that the term "and / or" as used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0190] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0191] The reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0192] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0193] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0194] In the embodiments provided in the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0195] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0196] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A model training method, characterized in that, Including: Obtain a training data set, where the training data set includes at least one license plate image and the label of each license plate image, and the label includes a license plate label; Input the license plate image into a pre-constructed LPRnet model to obtain the output result of the LPRnet model; the output result includes the license plate recognition result output by the first output branch, and the LPRnet model includes the first output branch; or, the output result includes the license plate recognition result output by the first output branch and the license plate color output by the second output branch, and the LPRnet model includes the first output branch and the second output branch; Calculate the connection temporal classification loss value according to the license plate recognition result and the license plate label; Through the formula Focal_ctc _ loss = α*(1 - p) γ *ctc _ loss calculates the first loss value; where p = e -ctc_loss , α and γ are balance factors, and ctc_loss is the connection timing classification loss value; Adjust the network parameters of the LPRnet model according to the first loss value, or adjust the network parameters of the LPRnet model according to the first loss value and the second loss value, where the second loss value is determined according to the license plate color; Iteratively train multiple times to obtain a trained LPRnet model.

2. The method according to claim 1, characterized in that, The label includes a license plate color label, and the output result includes the license plate recognition result output by the first output branch and the license plate color output by the second output branch; The LPRnet model includes a backbone network and a head part, and the head part includes the first output branch and the second output branch; Adjusting the network parameters of the LPRnet model according to the first loss value and the second loss value includes: Calculate the second loss value according to the license plate color and the license plate color label; Calculate the total loss value according to the first loss value and the second loss value; Adjust the network parameters of the LPRnet model according to the total loss value.

3. The method according to claim 2, characterized in that, The head part includes a first convolutional layer, a first activation function, a Concat layer, the first output branch, and the second output branch; The Concat layer is used to connect the output of the backbone network and the output of the first activation function; both the first output branch and the second output branch are connected to the Concat layer; The first output branch includes a second convolutional layer, a second activation function, and a first Transpose layer connected in sequence; The second output branch includes a third convolutional layer, a third activation function, a classification layer, and a second Transpose layer connected in sequence.

4. The method according to claim 2 or 3, characterized in that, Calculating the total loss value according to the first loss value and the second loss value includes: Calculate the total loss value through the formula Loss = Focal_ctc_loss + 2bce_loss; Where bce_loss is the second loss value, and Focal_ctc_loss is the first loss value.

5. The method according to claim 1, wherein The α is 0.95 and the γ is 1.

6. A license plate recognition method, characterized in that, Including: Obtain a license plate image to be recognized; Input the license plate image to be recognized into the pre-trained LPRnet model to obtain the output result of the LPRnet model. The output result includes the license plate recognition result output by the first output branch. The LPRnet model includes the first output branch; or, the output result includes the license plate recognition result output by the first output branch and the license plate color output by the second output branch. The LPRnet model includes the first output branch and the second output branch; The LPRnet model is a model trained using the model training method described in any one of claims 1 to 5.

7. The method according to claim 6, wherein The LPRnet model includes a backbone network and a head part. The head part includes the first output branch and the second output branch.

8. The method according to claim 7, wherein The head part includes a first convolutional layer, a first activation function, a Concat layer, the first output branch, and the second output branch; The Concat layer is used to connect the output of the backbone network and the output of the first activation function; both the first output branch and the second output branch are connected to the Concat layer; The first output branch includes a second convolutional layer, a second activation function, and a first Transpose layer connected in sequence; The second output branch includes a third convolutional layer, a third activation function, a classification layer, and a second Transpose layer connected in sequence.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method described in any one of claims 1 to 5 or 6 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method described in any one of claims 1 to 5 or 6 to 8.

Citation Information

Patent Citations

  • Two-place double-license-plate detection and recognition method and system based on deep learning

    CN111666938A

  • License plate detection and recognition method and handheld terminal

    CN112215233A