An end-to-end license plate correction and recognition method

Through the end-to-end deep learning model, the license plate correction and recognition steps are integrated, which solves the problems of difficult license plate recognition performance and complex license plate correction steps in the prior art, and realizes an efficient and simplified license plate correction and recognition process.

CN114067300BActive Publication Date: 2025-06-24ANHUI TSINGLINK INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110713952.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-25
Publication Date
2025-06-24
Estimated Expiration
2041-06-25

AI Technical Summary

Technical Problem

When the existing license plate recognition method deals with difficult license plates such as severe distortion and large angle tilt, the recognition performance is significantly reduced, and the license plate correction steps are complicated, requiring additional corner marking and high computing power requirements.

Method used

End-to-end license plate correction and recognition methods are adopted, and license plate correction and recognition steps are integrated through a deep learning model. The basic network is used to extract license plate features, the license plate correction head is corrected, and the license plate character recognition head is recognized, and the basic network is shared to simplify model design.

Benefits of technology

While ensuring the accuracy of difficult license plate recognition, the license plate correction steps are simplified, the additional corner marking and high computing power requirements are avoided, and the robustness and recognition efficiency of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067300B_ABST
    Figure CN114067300B_ABST
Patent Text Reader

Abstract

The present invention discloses an end-to-end license plate correction and recognition method, belonging to the technical field of image processing, which includes: obtaining a license plate image and using it as the input of the license plate correction and recognition fusion model. The license plate correction and recognition fusion model includes a backbone network, a license plate correction head, and a license plate character recognition head, and the license plate correction head and the license plate character recognition head share the backbone network; the backbone network performs multi-scale low-level and high-level feature extraction and fusion on the license plate image to obtain a license plate feature map F; the license plate correction head corrects the license plate based on the license plate feature map F to obtain a corrected license plate; the license plate character recognition head recognizes the license plate characters based on the license plate feature map F. The present invention integrates the license plate correction and license plate recognition steps into an end-to-end deep learning model, which simplifies the license plate correction step while improving the recognition accuracy of difficult license plates.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to an end-to-end license plate correction and recognition method. Background Art

[0002] License plate recognition is one of the core technologies of smart cities and has extensive applications in scenarios such as entrance and exit billing, illegal capture, and vehicle tracking.

[0003] In recent years, license plate recognition algorithms have developed well, and deep learning methods have been comprehensively popularized, resulting in a significant improvement in recognition accuracy. However, due to the lack of rotational invariance in convolutional networks, and at the same time, in the design of common license plate recognition networks, the default license plate direction is horizontal, such as the CTC series methods, the existing license plate recognition methods will significantly reduce the recognition performance when dealing with difficult license plates such as severe distortion and large-angle inclination. And the general solution is to introduce an additional license plate correction step.

[0004] Currently, the mainstream algorithms for license plate correction include two types: one is the method based on traditional image processing such as color segmentation, edge detection, and line detection to determine the correction matrix of the license plate image; this method has a good processing effect on simple samples but has poor robustness to difficult samples. The other is the method based on deep learning, which directly predicts the correction matrix of the license plate image through a convolutional network; this method has good performance and strong generalization ability, but requires a large number of license plate corner point annotation images and will increase additional computing power requirements. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies in the above background art and simplify the license plate frame correction step while ensuring the recognition accuracy of difficult license plates.

[0006] To achieve the above purpose, an end-to-end license plate correction and recognition method is adopted, including:

[0007] Obtain a license plate image and use it as the input of the license plate correction and recognition fusion model. The license plate correction and recognition fusion model includes a backbone network, a license plate correction head, and a license plate character recognition head, and the license plate correction head and the license plate character recognition head share the backbone network;

[0008] The backbone network performs multi-scale low-level and high-level feature extraction and fusion on the license plate image to obtain a license plate feature map F;

[0009] The license plate correction head corrects the license plate based on the license plate feature map F to obtain a corrected license plate;

[0010] The license plate character recognition head recognizes the license plate characters based on the license plate feature map F.

[0011] Further, the backbone network includes a downsampling layer, a feature pyramid layer, an upsampling layer, and two convolutional blocks CBL. The downsampling layer is used to extract license plate feature maps of different sizes. The feature pyramid layer is used to enhance the license plate feature maps of different sizes extracted by the downsampling layer. One convolutional block CBL enhances the feature maps output by the feature pyramid layer and serves as the input to the upsampling layer. The upsampling layer restores the license plate features based on the output of the feature pyramid layer and the output of one convolutional block CBL. The other convolutional block CBL enhances the feature maps output by the upsampling layer to obtain the license plate feature map F.

[0012] Further, the license plate correction head includes a license plate generator and a license plate discriminator. The license plate generator is used to simulate and generate real corrected license plate images. The license plate discriminator is used to determine whether the corrected license plates generated by the license plate generator are real corrected license plate images or simulated corrected license plate images.

[0013] Further, the license plate character recognition head includes a residual network, a Resize function, a Softmax function, and a CTC decoding network, where:

[0014] The residual network is used to compress the features of the license plate feature map F to obtain compressed features;

[0015] The compressed features are subjected to a channel merging Resize operation column by column to obtain an ordered feature map;

[0016] The ordered feature map is compressed column by column to N dimensions using a fully connected structure, and a Softmax operation is used to activate the confidence of each column to obtain the execution degree of each recognized character at each column position;

[0017] The execution degree of each recognized character at each column position is decoded by CTC to obtain the license plate character recognition result.

[0018] Further, before obtaining the license plate image and using it as the input of the license plate correction and recognition fusion model, it further includes:

[0019] Obtaining a license plate data set, where the data set includes real license plate images, labeled license plate number tags, and corrected images generated by simulation;

[0020] Using the license plate data set to train the license plate correction and recognition fusion model to learn the model parameters.

[0021] Further, it further includes: using an identification loss function Loss ctc and a generation loss function Loss g as well as a discriminant loss function Loss dOptimize the model parameters, where the recognition loss function is used to optimize the model parameters of the backbone network and the license plate character recognition head, the generation loss function is used to optimize the model parameters of the backbone network and the license plate generator, and the discriminant loss function is used to optimize the model parameters of the license plate discriminator.

[0022] Further, the recognition loss function Loss ctc is used to measure the error between the license plate character label L corresponding to the real license plate image I input to the model and the predicted label L obtained by model inference. The formula is as follows: pred

[0023]

[0024] where t is the label length, and p t represents the probability that the t-th dimension of the predicted label is under the condition of the real label L. The product of t = [1, 14] dimensions is the input label L, and the probability of obtaining the predicted label L pred by inference.

[0025] Further, the generation loss function Loss g is used to measure the error between the license plate correction image T generated by the model pred and the simulated correction image T. The formula is as follows:

[0026] Loss g = ||T - T pred || - logY pred

[0027] where ||T - T pred || is the L2 loss, used as a regularization term, and -logY pred is the cross-entropy loss, and Y pred represents the category predicted by the discriminator for the input image.

[0028] Further, the discriminant loss function Loss d is used to measure whether the license plate discriminator determines that the input image is a correction image generated by the model. The formula is as follows:

[0029] Loss d = -YlogY pred - (1 - Y)log(1 - Y pred )

[0030] where Y pred represents the category predicted by the discriminator for the input image, Y is the discriminator label, Y = 0 indicates that the discriminator believes the input image is a generated image, and Y = 1 indicates that the discriminator believes the input image is a simulated correction image.​

[0031] Furthermore, the optimization loss Loss of the entire network of the fusion model sum is expressed by the following formula:

[0032] Loss sum = Loss d + Loss g + αLoss ctc

[0033] where α is the license plate recognition loss magnification factor.

[0034] Compared with the prior art, the present invention has the following technical effects: By integrating the license plate correction and license plate recognition steps into a deep learning model, the present invention can correct the license plate image and recognize the license plate information simultaneously. This method does not rely on additional license plate corner annotation images, simplifies the license plate correction step while ensuring the recognition accuracy of difficult license plates. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The following describes in detail the specific embodiments of the present invention with reference to the accompanying drawings:

[0036] Figure 1 is a flowchart of an end-to-end license plate correction and recognition method;

[0037] Figure 2 is a structural diagram of the fusion model;

[0038] Figure 3 is a structural diagram of the backbone network;

[0039] Figure 4 is a structural diagram of the license plate correction head;

[0040] Figure 5 is a structural diagram of the license plate character recognition head;

[0041] Figure 6 is a schematic diagram of image correction;

[0042] Figure 7 is a schematic diagram of model parameter optimization. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] To further illustrate the features of the present invention, please refer to the following detailed description and drawings of the present invention. The accompanying drawings are for reference and illustration only and are not intended to limit the scope of protection of the present invention.

[0044] As Figure 1 shown, this embodiment discloses an end-to-end license plate correction and recognition method, including the following steps S1 to S4:

[0045] S1. Obtain a license plate image and use it as the input of the license plate correction and recognition fusion model. The license plate correction and recognition fusion model includes a backbone network, a license plate correction head, and a license plate character recognition head. The license plate correction head and the license plate character recognition head share the backbone network;

[0046] S2. The backbone network performs multi-scale low-level and high-level feature extraction and fusion on the license plate image to obtain a license plate feature map F;

[0047] S3. The license plate correction head corrects the license plate based on the license plate feature map F to obtain a corrected license plate;

[0048] S4. The license plate character recognition head recognizes license plate characters based on the license plate feature map F.

[0049] As a further preferred technical solution, in this embodiment, the design structure of the license plate correction and recognition fusion model is as Figure 2 shown. The fusion model consists of a backbone network (Backbone) and two output heads (license plate correction head GAN Head and license plate character recognition head CTC Head). The input of the network is the license plate image, and the outputs of the two output heads are the corrected image and the license plate number recognition result respectively.

[0050] The backbone network includes a downsampling layer, a feature pyramid layer, an upsampling layer, and two convolutional blocks CBL. The downsampling layer is used to extract license plate feature maps of different sizes. The feature pyramid layer is used to enhance the license plate feature maps of different sizes extracted by the downsampling layer. One convolutional block CBL enhances the feature map output by the feature pyramid layer and uses it as the input of the upsampling layer. The upsampling layer restores the license plate features based on the output of the feature pyramid layer and the output of one convolutional block CBL. The other convolutional block CBL enhances the feature map output by the upsampling layer to obtain the license plate feature map F.

[0051] As Figure 3 shown, the backbone network in this embodiment mainly includes 4 downsampling layers (DownLayer), 3 feature pyramid layers (FPN), 3 upsampling layers (UpLayer), and 2 convolutional blocks CBL. In this embodiment, for the convenience of description, the feature map size is represented in the format of W×H×C (i.e., width × height × number of channels), and C in and C out are used to represent the input and output channels of the network respectively.

[0052] Figure 3 The convolutional block CBL shown in in includes three basic operations: convolution Conv, batch normalization BN, and activation function ReLU. CBL(k, s, c out , c in, the number of output channels is c out CBL operation of

[0053] The working process of the backbone network is as follows: four downsampling layers are used to sequentially extract four license plate feature maps of different sizes. Specifically, the four downsampling layers, namely DownLayer i , i = 1, 2, 3, 4, are each composed of two convolutional block CBL structures and a max pooling layer Maxpooling with a stride of 2, and are used to gradually extract the license plate features of the input license plate image. In this embodiment, for each downsampling layer DownLayer i , its number of input channels are 3, 16, 32, 64 respectively, and the corresponding number of output channels are 16, 32, 64, 128 respectively. Then the two CBL structures of each downsampling layer are specifically: and Therefore, after the input license plate image with a resolution of 440×140 undergoes four downsampling operations, license plate feature maps with sizes of 220×70×16, 110×35×32, 55×18×64, and 28×9×128 are obtained respectively, denoted as

[0054] The feature pyramid layer is used to enhance the extracted license plate feature maps. Specifically, three feature pyramid layers, namely FPN j , j = 1, 2, 3, are each composed of an upsampling structure and an addition operation. In this embodiment, let the output feature map of each feature pyramid layer FPN j be and their sizes are 220×70×8, 110×35×16, and 55×18×32 respectively. Then as Figure 3 shown, for each feature pyramid layer FPN j , its input is in two parts, namely the output j+1 of the previous feature pyramid layer FPN and the license plate feature map of the downsampling layer of the same scale Specifically, in the FPN j structure, the input is processed by 2-fold linear interpolation (UpSample) and channel reduction (CBL), and the input is processed by channel reduction. After the two processed feature maps are summed element by element, an added feature map is obtained. Finally, 1 CBL operation is used to enhance this feature. Among them, as Figure 3 shown, since the topmost FPn3 has no previous feature pyramid layer, after the is enhanced by the CBL operation, it is used as the alternative input of FPN3.

[0055] The upsampling layer is used to gradually restore the license plate features. Specifically, as Figure 3 shown, the upsampling structure includes 3 upsampling layers UpLayer k , k = 1, 2, 3. For UpLayer k , its two inputs are the outputs of the previous UpLayer k+1 and the feature pyramid layer FPN j of the same scale. Specifically, in the upsampling layer structure UpLayer k , the input of UpLayer k+1 is restored by deconvolution (Dconv) with a scale factor of 2 and the number of channels is reduced, the input of UpLayer k+1 has its number of channels reduced, and the two processed feature maps are concatenated (Concate) and enhanced using the CBL structure. Finally, a license plate feature map F with a size of 220×70×16 is obtained.

[0056] It should be noted that the backbone network in this embodiment can effectively learn multi-scale low-level and high-level features. Among them, the downsampling layer can effectively extract difficult license plate features, the feature pyramid can effectively enhance multi-scale license plate features, and finally the features of each scale are fused through the upsampling layer, which can effectively restore and correct the license plate while retaining the license plate character features.

[0057] As a further preferred technical solution, the license plate correction head includes a license plate generator and a license plate discriminator. The license plate generator is used to simulate and generate real corrected license plate images, and the license plate discriminator is used to determine whether the corrected license plate generated by the license plate generator is a real corrected license plate image or a simulated corrected license plate image.

[0058] As Figure 4 shown, the inputs of both the license plate correction head and the license plate character recognition head are the license plate feature map F output by the backbone network. Among them, the license plate correction head is used to correct the real license plate image to obtain a standardized forward license plate; the license plate character recognition head is used to recognize license plate characters, that is, the license plate number.

[0059] Specifically, as Figure 4 shown, the license plate correction head, that is, GAN-Head, consists of a license plate generator and a license plate discriminator. Among them, the license plate generator is used to simulate and generate a more real corrected license plate as much as possible, and the license plate discriminator is used to determine whether the generated corrected license plate is a real image or a simulated image as much as possible. Usually, during the actual end-to-end prediction of the corrected license plate (i.e., forward propagation), the license plate correction head only needs to use the license plate generator to generate the corrected image, while the license plate discriminator is usually only used when optimizing the parameters of the fusion model (i.e., backpropagation training the model).

[0060] Furthermore, the license plate generator structure is as follows Figure 4 As shown, for the license plate feature map F with an input size of 220×70×16, first use a convolutional operation Conv1×1 with a convolutional kernel size of 1×1 and a stride of 1 to refine the features; then use bilinear interpolation UpSample by a factor of 2 to double the size of the feature map to 440×140×16; then use a convolutional operation with a convolutional kernel size of 3×3 and a stride of 1 to enhance the enlarged feature map; finally, use a convolutional operation Conv1×1 with a convolutional kernel size of 1×1 and a stride of 1 to restore the feature map to 3 channels, that is, generate the corrected image. The above convolutional operations all include batch normalization and ReLU activation.

[0061] Furthermore, the license plate discriminator uses a pre-trained residual network Resnet. In this embodiment, the network is Resnet18, and the number of output classes is 2, respectively indicating that the input license plate is a real image or a simulated image.

[0062] As a further preferred technical solution, the license plate character recognition head includes a residual network, a Resize function, a Softmax function, and a CTC decoding network, where:

[0063] The residual network is used to compress the features of the license plate feature map F to obtain compressed features;

[0064] Perform a channel merging Resize operation on the compressed features by column to obtain an ordered feature map;

[0065] Use a fully connected structure to compress the ordered feature map by column to N dimensions, and use the Softmax operation to activate the confidence of each column to obtain the execution degree of each recognized character at each column position;

[0066] Perform CTC decoding on the execution degree of each recognized character at each column position to obtain the license plate character recognition result.

[0067] Specifically, as Figure 5As shown, the license plate recognition head, namely CTC-Head, is used to recognize the corresponding license plate characters (i.e., license plate numbers). For the license plate feature map F with an input size of 220×70×16, first, a conventional residual network, such as Resnet18, is used to further compress the features to a size of 14×5×128; then, the compressed features are channel merged (Resize) by column (i.e., W) to obtain an ordered feature map with a size of 14×640; next, a fully connected structure is used to compress the ordered feature map by column to N dimensions, and Softmax is used to activate the confidence of each column to obtain the execution degree of each recognized character at each column position. In this embodiment, there are 65 types of characters in total, so N = 65; finally, CTC decoding is performed on the above results. Specifically, the 14 N-dimensional Softmax encodings obtained are subjected to one-hot processing to obtain a character sequence with a length of 14. After CTC duplicate removal, a license plate character with a length of 7 is finally obtained, that is, the license plate character recognition is completed.

[0068] It should be noted that in this embodiment, the license plate correction head and the license plate character recognition head can simultaneously learn correction information and character information from the license plate feature map F. The multi-task network structure can effectively promote learning from each other while completing a single correction and recognition task. That is, since the two task heads share the basic network, when the license plate correction head learns correction information, the classification information of the license plate recognition head can effectively assist in updating the parameters of the license plate correction head, ensuring that the corrected license plate characters generated by it are the same as the characters of the corresponding real license plate image, avoiding the uncontrollability of the generated image information by the traditional generative adversarial network (GAN) (the traditional GAN network cannot specify the desired license plate number when generating a license plate); at the same time, when the license plate recognition head learns recognition information, due to parameter sharing, the license plate correction head can effectively provide positive license plate information to the network, enabling the license plate recognition head to effectively use the correction information when recognizing difficult license plates such as severely distorted and large-angle tilted license plates, simplifying the difficult license plate recognition task into a conventional license plate recognition task, and effectively improving the recognition accuracy of difficult license plates.

[0069] As a further preferred technical solution, before the above step S1: obtaining a license plate image and using it as the input of the license plate correction and recognition fusion model, it further includes:

[0070] Obtaining a license plate data set, where the data set includes real license plate images, labeled license plate number tags, and corrected images generated by simulation;

[0071] Using the license plate data set to train the license plate correction and recognition fusion model to learn model parameters.

[0072] Specifically, the real license plate image I is an image of the circumscribed rectangle area of a single complete license plate. In this embodiment, the image is scaled to a resolution of 440×140;

[0073] The license plate number label L is the license plate number corresponding to each real license plate image (such as Wan A12**5). In this embodiment, the length of the license plate number label is 7;

[0074] The corrected image T is the corrected image corresponding to each real license plate image. In this embodiment, the resolution is 220×70; specifically, as Figure 6 shown, the corrected image is generated by selecting a background template and license plate characters according to the released license plate standard and the corresponding license plate number label above through image processing technology and simulating them in proportion.

[0075] Compared with the traditional license plate correction method based on perspective transformation or affine transformation, by generating the corrected image in this simulation way, there is no need to perform license plate corner point annotation, which can effectively reduce the annotation cost and the annotation error caused by manual annotation, and at the same time avoid the computing power consumption of the correction matrix and perspective transformation.

[0076] As a further preferred technical solution, by adopting the recognition loss function Loss ctc 、the generation loss function Loss g and the discriminant loss function Loss d to optimize the model parameters, where the recognition loss function is used to optimize the model parameters of the backbone network and the license plate character recognition head, the generation loss function is used to optimize the model parameters of the backbone network and the license plate generator, and the discriminant loss function is used to optimize the model parameters of the license plate discriminator.

[0077] The recognition loss function Loss ctc is used to measure the error between the license plate character label L corresponding to the real license plate image I input to the model and the predicted label L pred inferred by the model. The expression formula is as follows:

[0078]

[0079] where t is the label length, and p t represents the probability that the t-th dimension of the predicted label is under the condition of the real label L. The product of t = [1, 14] dimensions is the input label L, and the probability of inferring the predicted label L pred .

[0080] The generation loss Loss g is used to measure the license plate corrected image T generated by the model predThe error between the simulated corrected image T. Discriminative loss Loss d Used to measure whether the license plate discriminator classifies the input image (T pred or T) as the corrected image generated by the model. Specifically, the above Loss g and Loss d is a loss for alternate verification. When optimizing the model, by minimizing the generative loss Loss g to ensure the generation of a corrected image with better quality, and by minimizing the discriminative loss Loss d to enable the discriminative model to distinguish as much as possible between the predicted corrected image and the simulated corrected image. This way of alternate verification makes the generated corrected image more realistic and effective.

[0081] Specifically, let Y be the discriminator label. Y = 0 indicates that the discriminator believes the input image is a generated image, and Y = 1 indicates that the discriminator believes the input image is a simulated corrected image. When training the generator, we assume that the discriminator parameters are good enough, and only update the parameters P b and P g . Then, when training the generator, the aim is to minimize Loss g such that the discriminator cannot distinguish whether the generated corrected image is generated by the model or simulated, thus making the generated corrected image as realistic as possible. The generative loss function Loss g is expressed as follows:

[0082] Loss g = ||T - T pred || - logY pred

[0083] where ||T - T pred || is the L2 loss, which serves as a regularization term to ensure the stability of the initial training of the GAN; -logY pred is the cross-entropy loss, and Y pred represents the class predicted by the discriminator for the input image, and is used to optimize the above model parameters P b and P g .

[0084] The discriminative loss function Loss d is expressed as follows:

[0085] Loss d = -YlogY pred - (1 - Y)log(1 - Y pred )

[0086] where Y predIt represents the category predicted by the discriminator for the input image. By minimizing the above formula as much as possible, the purpose is to optimize the discriminator, so that the discriminator can distinguish whether the input image is the corrected image generated by the model or the simulated corrected image as much as possible. The training of the discriminator and the generator is alternately iterated and gradually optimized with each other, and finally the generated corrected image can be made as good as possible.

[0087] As a further preferred technical solution, the optimization loss Loss of the entire network sum can be expressed by the following formula:

[0088] Loss sum = Loss d + Loss g + αLoss ctc

[0089] Among them, α is the license plate recognition loss magnification factor. In this embodiment, α = 2, which is used to enhance the license plate recognition accuracy.

[0090] Generally speaking, when training and optimizing the end-to-end model, this embodiment uses Loss ctc to optimize the license plate recognition accuracy, and uses Loss g and Loss d to optimize the license plate correction accuracy. At the same time, in order to ensure the training stability of the GAN-Head, an additional L2 loss is added to ensure the rapid training of the GAN-Head to stability in the initial stage. The two task heads of the entire end-to-end network promote each other. The CTC-Head makes the license plate characters generated by the GAN-Head consistent with the input real image, solving the problem that the traditional GAN network cannot set the attributes of the generated image (that is, the generated license plate characters are randomly transformed and cannot be consistent with the character labels); while the GAN-Head enables the network to automatically correct difficult license plates, ensuring that the CTC-Head can obtain effective positive license plate features, solving the problem of low recognition accuracy of the existing CTC algorithm for difficult license plates.

[0091] As a further preferred technical solution, when using the trained end-to-end model (including the backbone network, license plate correction head, and license plate character recognition head) to perform license plate correction or license plate recognition tasks, for license plate recognition-related tasks or projects, the license plate correction head can be removed, and only the backbone network and the license plate recognition head are retained for license plate recognition; for license plate correction-related tasks or projects, the license plate recognition head can be removed, and only the backbone network and the license plate correction head are retained for license plate correction. Using this method can effectively reduce the forward inference time of the model and accelerate the recognition or correction speed. Specifically, during the training of the end-to-end model, since the corrected images and character labels are used for supervised training simultaneously, the learned optimal model parameters can contain both correction information and character recognition information. Therefore, when performing forward inference, removing the additional task head will not affect the correction or recognition accuracy and can be used for model simplification and acceleration.

[0092] The present invention has the following remarkable effects:

[0093] (1) Simplify the license plate recognition steps: Compared with the two-stage method (license plate correction - license plate recognition) of the prior art, the correction and recognition steps are integrated into a single deep learning model, simplifying the algorithm parameter adjustment and model deployment problems.

[0094] (2) Design an end-to-end license plate correction and recognition model: Based on the existing generative adversarial concept, two additional supervision branches are designed respectively to generate license plate correction images corresponding to the input difficult license plate images, and while ensuring the invariance of license plate character information, predict the results of the corrected license plate images and license plate recognition.

[0095] (3) Solve the problem of recognizing difficult license plates such as severe distortion and large-angle tilt: By using the method of deep learning to integrate the license plate correction and license plate recognition algorithms, through the mutual promotion of multi-tasks, the correction and recognition of difficult license plates are effectively solved. At the same time, the multi-task integration method avoids the loss of recognition accuracy of the two-stage method.

[0096] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. An end-to-end license plate correction and recognition method, characterized in that Including: Obtain a license plate image and use it as the input of the license plate rectification and recognition fusion model. The license plate rectification and recognition fusion model includes a backbone network, a license plate rectification head, and a license plate character recognition head. The license plate rectification head and the license plate character recognition head share the backbone network; The backbone network performs multi-scale low-level and high-level feature extraction and fusion on the license plate image to obtain a license plate feature map F; The license plate rectification head rectifies the license plate based on the license plate feature map F to obtain a rectified license plate; The license plate character recognition head recognizes license plate characters based on the license plate feature map F; The backbone network includes a downsampling layer, a feature pyramid layer, an upsampling layer, and two convolutional blocks CBL. The downsampling layer is used to extract license plate feature maps of different sizes. The feature pyramid layer is used to enhance the license plate feature maps of different sizes extracted by the downsampling layer. One convolutional block CBL enhances the feature map output by the feature pyramid layer and uses it as the input of the upsampling layer. The upsampling layer restores the license plate features based on the output of the feature pyramid layer and the output of one convolutional block CBL. The other convolutional block CBL enhances the feature map output by the upsampling layer to obtain the license plate feature map F; The license plate rectification head includes a license plate generator and a license plate discriminator. The license plate generator is used to simulate and generate a real rectified license plate image. The license plate discriminator is used to determine whether the rectified license plate generated by the license plate generator is a real rectified license plate image or a simulated rectified license plate image; The license plate character recognition head includes a residual network, a Resize function, a Softmax function, and a CTC decoding network, where: The residual network is used to compress the features of the license plate feature map F to obtain compressed features; Perform a channel merging Resize operation on the compressed features by column to obtain an ordered feature map; Use a fully connected structure to compress the ordered feature map by column to N dimensions, and use the Softmax operation to activate the confidence of each column to obtain the confidence of each recognized character at each column position; Perform CTC decoding on the confidence of each recognized character at each column position to obtain the license plate character recognition result.

2. The end-to-end license plate correction and recognition method according to claim 1, wherein Before obtaining the license plate image and using it as the input of the license plate rectification and recognition fusion model, it further includes: Obtain a license plate data set, which includes real license plate images, labeled license plate number tags, and rectified images generated by simulation; Use the license plate data set to train the license plate rectification and recognition fusion model to learn model parameters.

3. The end-to-end license plate correction and recognition method according to claim 2, wherein, It also includes: Adopt an identification loss function and a generation loss function as well as a discriminant loss function to optimize the model parameters. Among them, the identification loss function is used to optimize the model parameters of the backbone network and the license plate character recognition head, the generation loss function is used to optimize the model parameters of the backbone network and the license plate generator, and the discriminant loss function is used to optimize the model parameters of the license plate discriminator.

4. The end-to-end license plate correction and recognition method according to claim 3, wherein The recognition loss function is used to measure the license plate character label L corresponding to the real license plate image I input to the model and the predicted label obtained by model inference. The formula is as follows: where t is the tag length, denotes the probability that the t-th dimension of the predicted tag is under the condition of the true tag L, with t = [1, 14] dimensions The product of which is the input tag L, and the predicted tag is inferred with a probability.

5. The end-to-end license plate correction and recognition method according to claim 3, characterized in that, The generated loss function is used to measure the error between the license plate correction image generated by the model and the simulated correction image T, and the formula is as follows: Among them, is L 2 losses, as a regularization term, is the cross-entropy loss, indicating the category predicted by the discriminator for the input image.

6. The end-to-end license plate correction and recognition method according to claim 3, wherein, The discriminant loss function is used to measure whether the license plate discriminator determines that the input image is the corrected image generated by the model, and the formula is as follows: Among them, represents the category predicted by the discriminator for the input image. Y is the discriminator label. Y = 0 indicates that the discriminator believes the input image is a generated image, and Y = 1 indicates that the discriminator believes the input image is a simulation-corrected image.

7. The end-to-end license plate correction and recognition method according to claim 3, wherein, The optimization loss of the entire network of the fusion model The formula is expressed as follows: Among them, is the license plate recognition loss magnification factor.

Citation Information

Patent Citations

  • License plate recognition method and device, storage medium and terminal

    CN112446383A

  • Deep neural network for fine recognition of vehicle attributes, and training method thereof

    WO2019169816A1