A robust Chinese license plate recognition method in uncontrolled environment
Through the CRNN architecture combined with VGG16, PSA, DropBlock and Bi-LSTM technologies, the robustness and performance problems of the license plate recognition system in an uncontrollable environment are solved, and a high accuracy and rapid identification of license plate recognition method is achieved.
Patent Information
- Application Number
- CN202111262855.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-28
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-10-28
AI Technical Summary
In an uncontrollable environment, the license plate recognition system has poor robustness and unstable performance, resulting in a decrease in recognition accuracy.
The license plate recognition method based on the CRNN architecture is adopted, and the VGG16 backbone network is combined with the pyramid split attention module (PSA) and the DropBlock layer for deep feature extraction, and character sequence recognition is combined with Bi-LSTM and CTC algorithm.
It realizes high accuracy and rapid recognition in an uncontrollable environment, supports arbitrary fixed-length license plate character sequence recognition, and is highly generalized and robust.
Smart Images

Figure CN114639090B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and specifically relates to Chinese character recognition, deep learning and other technologies. Background Art
[0002] As the only sign of a vehicle, the license plate not only reflects the information of the vehicle, but also the information of the owner. License plate recognition technology can be applied to various fields. For example, in terms of traffic safety, when there are incidents such as speeding, wrong-way driving, illegal parking, and traffic accidents; in terms of daily life, at expressway toll roads, community entrances and exits, and parking lot entrances and exits, accurate license plate information recognition is required.
[0003] At present, Automatic License Plate Recognition (ALPR) technology has made great progress and has been widely used in controllable environments, such as parking lot entrances and exits, highway toll intersections, etc. ALPR collects vehicle images through cameras, and then uses image processing, machine learning and other technologies to automatically detect and recognize license plates, which greatly saves labor costs and improves license plate recognition efficiency. ALPR has a fast recognition speed and high accuracy, but this is achieved under controllable environmental conditions: the license plate angle is basically fixed, the vehicle image is clear, the light intensity is uniform, the distance between the vehicle and the camera is basically fixed, etc. In uncontrolled environments, such as tunnels and intersections, the performance of ALPR is greatly affected by factors such as complex environmental backgrounds, drastic changes in lighting conditions, and unfixed license plate angles and scales. In recent years, deep learning technology has developed rapidly, and many experts and scholars have proposed license plate recognition methods based on convolutional neural networks (CNNs), relying on big data to solve the problem of accurate license plate recognition in uncontrolled environments.
[0004] Xu Z et al. extracted multi-scale features of the target through CNN, fused the multi-scale features and sent them to 7 classifiers to complete license plate recognition. This method uses a shallow CNN network, so the recognition speed is fast. However, since there are only seven classifiers, it does not support the recognition of indefinite length sequences, such as the eight-digit license plate number of new energy vehicles.
[0005] Zherzdev et al. proposed the LPRNet license plate recognition method. This method uses the STN (SpatialTransformer Networks) module to perform spatial transformation on the input image, adds global information to enhance the intermediate feature map, and finally uses the CTC (Connectionist Temporal Classification) loss to solve the problem of variable-length sequence alignment. This method supports variable-length sequence recognition and has a faster speed, but its generalization is weak. Summary of the invention
[0006] Aiming at the problems of poor robustness and low performance of license plate recognition in uncontrolled environment, the present invention proposes a robust Chinese license plate recognition method in uncontrolled environment. The method is implemented based on CRNN (Convolutional Recurrent Neural Network) framework and mainly includes three steps: establishing data set, input image preprocessing and network structure design.
[0007] Step 1: Construction of training dataset
[0008] Deep learning requires a large amount of training data to achieve ideal performance. The present invention uses the CRNN architecture to recognize license plates. In order to train the network model, a license plate dataset needs to be established. The license plate dataset should contain license plate images from different provinces, different types, and different environmental conditions in my country.
[0009] Step 2: Input image preprocessing
[0010] Before inputting into the network, the license plate image needs to be preprocessed, which mainly includes two steps:
[0011] (1) License plate image size adjustment: According to the size of the Chinese license plate and the image resolution, the present invention adjusts the input image size to a uniform size of (w, 48), where w is the width of the image.
[0012] (2) Normalization of license plate images: Normalize all pixel values of the input image to between -1 and 1 to make the network easier to converge.
[0013] Step 3: Network structure design
[0014] Step 3.1: Overall network architecture
[0015] The Chinese license plate recognition network architecture designed by the present invention mainly includes three parts, namely deep feature extraction, character sequence recognition and CTC.
[0016] Deep feature extraction
[0017] In order to ensure the recognition speed, the present invention uses VGG16 as the backbone network to extract deep features. The network consists of 7 convolutional layers and 4 pooling layers. In order to improve the network's ability to extract and express features, the present invention adds a pyramid split attention module (PSA) to the network. PSA divides the input tensor into S groups from the channel, and each group is convolved with a convolution kernel of different sizes to obtain receptive fields of different scales and extract information of different scales. The weighted value of each group of feature channels is then calculated through the channel attention mechanism, and finally the weighted values of the S groups are Softmax normalized and weighted. The PSA attention mechanism can enhance important features useful for recognition and suppress irrelevant features.
[0018] In addition, the present invention also adds a DropBlock layer to the network, which randomly discards part of the image to improve the robustness of the model. When implemented, some inputs are set to zero with a certain probability to accelerate the network convergence speed.
[0019] Character sequence recognition
[0020] The present invention adopts Bi-LSTM (Bidirectional Long Short-Term Memory) to recognize character sequences. Compared with other RNN (Recurrent Neural Networks) structures, Bi-LSTM can handle long-term dependency problems and reduce the possibility of gradient vanishing and gradient explosion. In addition, the present invention adds a Dropout module to the character sequence recognition part, which is used in the fully connected layer to improve the model representation ability by randomly inactivating neurons. The Dropout module has only one parameter p, which represents the probability of node zeroing.
[0021] ·CTC
[0022] Since the lengths of Chinese license plate characters vary, the present invention uses a CTC algorithm to solve the alignment problem of license plate characters of indefinite length. CTC can convert a series of character sequences obtained from the sequence recognition layer into a final character sequence.
[0023] Compared with the existing license plate recognition method, the present invention has the following obvious advantages and beneficial effects:
[0024] 1. Fast recognition speed and high accuracy;
[0025] 2. It can realize the recognition of any fixed-length license plate character sequence and meet the recognition requirements of license plates of different lengths;
[0026] 3. It has strong generalization and robustness, and can be applied to various complex and uncontrollable scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 Overall block diagram of license plate recognition method
[0028] Figure 2 PSA structure diagram
[0029] Figure 3 Channel attention structure diagram
[0030] Figure 4 Bi-LSTM structure diagram DETAILED DESCRIPTION
[0031] The specific implementation of the present invention is described in detail below with reference to the accompanying drawings.
[0032] The overall block diagram of the Chinese license plate recognition method proposed in this invention consists of four parts: input preprocessing, deep feature extraction, character sequence recognition and CTC. Figure 1 .
[0033] The implementation details of each step are as follows:
[0034] Step 1: Create a training dataset
[0035] The training dataset mainly consists of two parts:
[0036] (1) Existing public datasets
[0037] The CCPD dataset is the largest publicly available Chinese license plate dataset. The dataset was released by the research team of USTC and contains about 300,000 license plate images. Each license plate label consists of Chinese characters, English characters, and numbers, with a total of 67 character categories. The dataset contains a variety of license plate images, such as tilted license plates, blurred license plates, snowy license plates, etc. However, most of the license plates start with Anhui A, and lack other provinces and new energy license plates.
[0038] (2) Computer-generated datasets.
[0039] To solve the problem of sample imbalance in the CCPD dataset, the present invention uses computer-generated license plates to expand the dataset.
[0040] Finally, a dataset of 1.4 million license plate images was constructed for training the network structure.
[0041] Step 2: Input preprocessing
[0042] Step 2.1: Input image resizing
[0043] According to the requirements, the input image size is uniformly adjusted to (w, 48) by downsampling and other methods. The height of the input image must be 48 (sequence recognition layer input requirement), and the width can be adjusted arbitrarily. Considering that the aspect ratio of ordinary license plates is about 3 to 1, the width w is set to 144 in this invention.
[0044] Step 2.2: Normalize the input image
[0045] First, divide all pixel values by 255 to limit the pixel values to between 0 and 1, and then use the following formula to normalize them:
[0046]
[0047] In the formula, x is the value of the original pixel value divided by 255, μ is the mean of the sample values, σ is the standard deviation of the sample values, and μ and σ are both set to 0.5. * The value is between -1 and 1.
[0048] Step 3: Overall network architecture
[0049] The overall network architecture of the present invention is mainly divided into three parts, namely deep feature extraction, character sequence recognition and CTC. The deep feature extraction part is mainly composed of VGG16, PSA and DropBlock; the character sequence recognition part is mainly composed of Bi-LSTM and FC layers. The input image is converted into a feature sequence through the feature extraction layer, the sequence recognition layer recognizes the feature sequence as a string, and CTC converts the string into the final output.
[0050] Step 3.1: Deep Feature Extraction Layer
[0051] The input size of the license plate image of the feature extraction layer is (b, 3, 48, w), where b is the number of samples sent to the network, 3 is the number of channels, 48 is the height of the license plate image, and w is the width of the license plate image, which is 144; the output size is (b, 512, 1, 36). After the feature extraction layer, the number of feature channels output is 512 and the feature map size is (36, 1).
[0052] (1) Backbone network
[0053] The backbone network is mainly composed of 7 convolutional layers and 4 pooling layers. The first convolutional layer has 3 input channels, that is, the three RGB channel images of the license plate image are all fed into the network, the output channels are 64, the convolution kernel size is (3,3), the step size is (1,1), and the edge padding size is 1; the second layer is the maximum pooling layer, the pooling size is (2,2), the step size is (2,2); the third layer is the convolutional layer, the output channels are 128, the convolution kernel size is (3,3), the step size is (1,1), and the edge padding size is 1; the fourth layer is the maximum pooling layer, the step size is (2,2), and the pooling size is (2,2); the fifth layer is the convolutional layer, the output channels are 256, the convolution kernel size is (3,3), the step size is (1,1), and the edge padding size is 1; the sixth layer is the convolutional layer, the output channels are 256, the convolution kernel size is (3,3), the step size is (1,1), and the edge padding size is 1. The size of the convolution kernel is (3,3), the stride is (1,1), and the edge padding size is 1; the 7th layer is a maximum pooling layer with a stride of (2,1) and a pooling size of (2,1); the 8th layer is a convolution layer with an output channel of 512, a convolution kernel size of (3,3), a stride of (1,1), and an edge padding size of 1; the 9th layer is a batch normalization layer; the 10th layer is a convolution layer with an output channel of 512, a convolution kernel size of (3,3), a stride of (1,1), and an edge padding size of 1; the 11th layer is a batch normalization layer; the 12th layer is a maximum pooling layer with a pooling size of (2,1) and a stride of (2,1); the 13th layer is a convolution layer with an output channel of 512, a convolution kernel size of (3,1), a stride of (1,1), and an edge padding size of 1.
[0054] The 3rd and 4th maximum pooling layers use a 1×2 rectangular pooling frame instead of the 2×2 square pooling frame of VGG16. License plate characters are all rectangles with short width and long height. Using a rectangular pooling frame is more conducive to license plate character recognition. The specific structure is shown in Table 1.
[0055] Table 1 Backbone network structure
[0056] Input (b,3,48,W) Convolution #maps:64,k:(3,3),s:1,p:1 MaxPooling Window:(2,2),s:2 Convolution #maps:128,k:(3,3),s:1,p:1 MaxPooling Window:(2,2),s:2 Convolution #maps:256,k:(3,3),s:1,p:1 Convolution #maps:256,k:(3,3),s:1,p:1 MaxPooling Window:(2,1),s:(2,1),p:0 Convolution #maps:512,k:(3,3),s:1,p:1 BatchNormalization - Convolution #maps:512,k:(3,3),s:1,p:1 BatchNormalization - MaxPooling Window:(2,1),s:(2,1),p:0 Convolution #maps:512,k:(3,1),s:1,p:0
[0057] In the table, Convolution represents the convolution layer, MaxPooling represents the maximum pooling layer, BatchNormalization represents the batch normalization layer; #maps is the number of output channels, k is the convolution kernel size, s is the convolution kernel moving step, p is the edge padding size, and Window is the pooling size.
[0058] (2) Pyramid Split Attention Mechanism
[0059] The present invention adds PSA to the first convolutional layer of the backbone network, and its structure is as follows: Figure 2As shown in the figure. PSA divides the input tensor into S groups from the channel, and convolves each group with convolution kernels of different sizes to obtain receptive fields of different scales and extract information of different scales. The weighted value of each group of channels is then calculated through the channel attention mechanism, and finally the weighted values of the S groups are Softmax normalized and weighted. Here S is set to 4, and considering that the size of the input feature map is (144, 48), the convolution kernel size is set to (1, 3, 5, 7). Through such operations, PSA can fuse contextual information of different scales and improve the expressiveness of features.
[0060] The channel attention structure is as follows Figure 3 As shown in Figure 2, the output is the weighted value of each group of channels. The channel attention mechanism can guide the network to focus on features that are useful for recognition and suppress useless features, thereby improving recognition performance.
[0061] (3) DropBlock layer
[0062] The DropBlock layer is located after the attention layer. It is a regularization method for the convolutional layer, which is to set the input feature map to zero with a certain probability in units of blocks. The DropBlock module has two main parameters, Block_size and γ. Block_size represents the size of the block that is set to zero, and γ represents the probability of setting it to zero. Due to the limitation of the size of the license plate characters, Block_size cannot be too large, otherwise the feature value of the entire character area will be set to 0. Taking all factors into consideration, the present invention sets Block_size to 3, which is more appropriate, and sets γ to 0.1. By setting a part of the adjacent area to zero, the network will focus on learning the features of other positions of the character to achieve correct character recognition, thereby showing better generalization.
[0063] Step 3.2: Character sequence recognition
[0064] The input feature map size of character sequence recognition is (T, b, 512), and the output feature map size is (T, b, num_c). For the input feature map, T represents the width of the feature map, b represents the number of samples sent to the network, and 512 represents the number of channels; for the output feature map, T represents the number of characters output, and num_c represents the number of character categories, which is specified as 68 here (including the blank character "-" for a total of 68 categories).
[0065] The first layer of the character sequence recognition network is a bidirectional LSTM layer with 256 nodes; the second layer is a linear layer with 512 nodes; the third layer is a Dropout layer; the fourth layer is a bidirectional LSTM layer with 256 nodes; the fifth layer is a linear layer with the number of nodes equal to the number of license plate character categories, i.e. 68.
[0066] The network structure of character sequence recognition is shown in Table 2.
[0067] Table 2 Character sequence recognition network structure
[0068] Input (T,b,512) Bidirectional_LSTM #hidden units:256 Linear #hidden units:512 Dropout - Bidirectional_LSTM #hidden units:256 Linear #out units:68
[0069] In the table, Bidirectional_LSTM is bidirectional LSTM, Linear is linear layer, Dropout is node dropout layer, #hidden units is the number of hidden nodes, and out units is the number of output layer nodes.
[0070] (1) Bi-LSTM
[0071] The structure of Bi-LSTM is as follows Figure 4 As shown in the figure, it is composed of a forward LSTM and a backward LSTM. LSTM itself can handle long-term dependency problems and reduce the possibility of gradient vanishing and gradient explosion. Bi-LSTM combines forward and backward information, which is more conducive to character recognition.
[0072] (2) Dropout
[0073] Dropout is used in the training process of the fully connected layer, that is, randomly setting some nodes to zero and stopping the update of their gradients. Since the probability of nodes being set to zero each time is different, the network model becomes more diverse and simpler, so it can effectively solve the overfitting problem. In addition, since many nodes are set to zero, the amount of network calculations is reduced, saving the training time of the model.
[0074] To apply Dropout, you only need to specify the parameter p, which represents the probability of setting a node to zero. According to experience, it is more appropriate to set p to 0.3.
[0075] Step 3.3: CTC
[0076] The output feature map size of the sequence recognition layer is (T, b, num_c), where T represents the number of characters that should be output for each sample, and the value here is 36. In fact, the number of characters in the license plate is 7 or 8, so it is necessary to match the output T characters with the actual label characters.
[0077] Many of the T characters are repeated, and the repeated characters may correspond to the same label character. How to determine repeated characters and eliminate repeated characters are two key issues that need to be solved. The CTC algorithm solves the above problem by introducing blank characters (Blank). Here, "-" is used to represent the blank character. If there is a blank character between two identical characters, then they are not repeated. Otherwise, they are repeated and need to be merged into one. This character deduplication transformation is represented by B. The specific example is as follows:
[0078] B(π1)=B(--stta-t--e)=state#(2)
[0079] B(π2)=B(sst-aaa-tee-)=state#(3)
[0080] In the formula, "--sttat--e" and "sst-aaa-tee-" are two strings corresponding to the output of the neural network. Each string can be regarded as a string path, represented by π. π1 (path 1) is transformed by B to obtain "state" (the same applies to π2). "-" represents a blank character, and "s, t, a, e" represent English letters.
[0081] The probability of a path is:
[0082]
[0083] In the above formula, x is the input, represents the probability of each character, and T represents the number of characters output by the neural network. Assuming the label is l, if the path π t After B transformation, we can get label l, then the path π t is a correct path. The CTC algorithm maximizes the sum of the probabilities of all correct paths, that is,
[0084]
[0085] The CTC loss function is:
[0086] CTC(x)=-log(l|x)#(6)
[0087] By minimizing the CTC loss function, the network model can recognize characters of variable length.
[0088] The Chinese license plate recognition method proposed in the present invention can improve the generalization and robustness of the model by introducing two regularization methods (Dropout, DropBlock) and the PSA attention mechanism, so that the model can obtain ideal recognition results in various complex and uncontrollable environments.
Claims
1. A robust Chinese license plate recognition method in an uncontrolled environment, characterized by: It consists of four parts: input preprocessing, deep feature extraction, character sequence recognition, and CTC; The implementation details of each step are as follows: Step 1: Create a training dataset Step 2: Input preprocessing Step 2.1: Input image resizing Step 2.2: Normalize the input image Step 3: Overall network architecture The overall network architecture is divided into three parts: deep feature extraction, character sequence recognition, and CTC. The deep feature extraction part is composed of VGG16, PSA, and DropBlock; the character sequence recognition part is composed of Bi-LSTM and FC layers. The input image is converted into a feature sequence through the feature extraction layer, the sequence recognition layer recognizes the feature sequence as a string, and CTC converts the string into the final output. Step 3.1: Deep Feature Extraction Layer The input size of the license plate image of the feature extraction layer is (b, 3, 48, w), where b is the number of samples sent to the network, 3 is the number of channels, 48 is the height of the license plate image, and w is the width of the license plate image, which is 144; the output size is (b, 512, 1, 36). After the feature extraction layer, the number of feature channels output is 512, and the feature map size is (36, 1); (1) Backbone network The backbone network consists of 7 convolutional layers and 4 pooling layers; (2) Pyramid Split Attention Mechanism After PSA is added to the first convolutional layer of the backbone network, PSA divides the input tensor into S groups from the channel, and convolves each group with convolution kernels of different sizes to obtain receptive fields of different scales and extract information of different scales; then the weighted value of each group of channels is calculated through the channel attention mechanism, and finally the weighted values of the S groups are Softmax normalized and weighted; here S is set to 4, considering that the size of the input feature map is (144, 48), so the convolution kernel size is set to (1, 3, 5, 7); through such operations, PSA integrates contextual information of different scales and improves the expressiveness of features; The output of the channel attention structure is the weighted value of each group of channels. The channel attention mechanism guides the network to focus on features that are useful for recognition and suppress useless features, thereby improving recognition performance; (3) DropBlock layer The DropBlock layer is located after the attention layer. It is a regularization method for the convolutional layer, which sets the input feature map to zero with a certain probability in blocks. The DropBlock module has two parameters, Block_size and γ. Block_size indicates the size of the zeroed block, and γ indicates the probability of zeroing. Block_size is set to 3 and γ is set to 0.
1. By setting a part of the adjacent area to zero, the network will focus on learning the features of other positions of the character to achieve correct character recognition, thereby showing better generalization. Step 3.2: Character sequence recognition The first layer of the character sequence recognition network is a bidirectional LSTM layer with 256 nodes; the second layer is a linear layer with 512 nodes; the third layer is a Dropout layer; the fourth layer is a bidirectional LSTM layer with 256 nodes; the fifth layer is a linear layer with the number of nodes equal to the number of license plate character categories, i.e. 68; To apply Dropout, you only need to specify the parameter p, which represents the probability of setting a node to zero, and p is set to 0.3; Step 3.3: Use the CTC algorithm to convert the series of character sequences obtained from the sequence recognition layer into the final character sequence.
2. The method according to claim 1, characterized in that Preprocessing specifically includes: Step 2.1: Input image resizing By downsampling, the input image size is uniformly adjusted to (w, 48); the height of the input image must be 48, and the width w is set to 144; Step 2.2: Normalize the input image Divide all pixel values by 255 to limit the pixel values to between 0 and 1, and then use the following formula to normalize: In the formula, x is the value of the original pixel value divided by 255, μ is the mean of the sample values, σ is the standard deviation of the sample values, and μ and σ are both set to 0.5; x after normalization is * The value is between -1 and 1.
3. The method according to claim 1, characterized in that The backbone network specifically includes: The number of input channels of the first convolution layer is 3, that is, the three RGB channel images of the license plate image are all sent to the network, the number of output channels is 64, the convolution kernel size is (3,3), the step size is (1,1), and the edge padding size is 1; the second layer is the maximum pooling layer, the pooling size is (2,2), and the step size is (2,2); the third layer is the convolution layer, the number of output channels is 128, the convolution kernel size is (3,3), the step size is (1,1), and the edge padding size is 1; the fourth layer is the maximum pooling layer, the step size is (2,2), and the pooling size is (2,2); the fifth layer is the convolution layer, the number of output channels is 256, the convolution kernel size is (3,3), the step size is (1,1), and the edge padding size is 1; the sixth layer is the convolution layer, the number of output channels is 256, and the convolution kernel size is The size of the convolution kernel is (3,3), the stride is (1,1), and the edge padding size is 1; the 7th layer is a maximum pooling layer, the stride is (2,1), and the pooling size is (2,1); the 8th layer is a convolution layer, the number of output channels is 512, the convolution kernel size is (3,3), the stride is (1,1), and the edge padding size is 1; the 9th layer is a batch normalization layer; the 10th layer is a convolution layer, the number of output channels is 512, the convolution kernel size is (3,3), the stride is (1,1), and the edge padding size is 1; the 11th layer is a batch normalization layer; the 12th layer is a maximum pooling layer, the pooling size is (2,1), the stride is (2,1); the 13th layer is a convolution layer, the number of output channels is 512, the convolution kernel size is (3,1), the stride is (1,1), and the edge padding size is 1; The 3rd and 4th maximum pooling layers use a 1×2 rectangular pooling frame instead of the 2×2 square pooling frame of VGG16. License plate characters are all rectangles with a short width and a long height.
4. The method according to claim 1, characterized in that: The input feature map size of character sequence recognition in step 3.2 is (T, b, 512), and the output feature map size is (T, b, num_c); for the input feature map, T represents the width of the feature map, b is the number of samples sent to the network, and 512 is the number of channels; for the output feature map, T represents the number of output characters, and num_c represents the number of character categories, which is specified as 68 here, including the space character "-", for a total of 68 categories.
5. The method according to claim 1, characterized in that Step 3.3 is as follows: the output feature map size of the sequence recognition layer is (T, b, num_c), where T represents the number of characters that should be output for each sample, and the value here is 36; in fact, the number of characters in the license plate is 7 or 8, so it is necessary to match the output T characters with the actual label characters; The CTC algorithm uses "-" to represent a blank character; if there is a blank character between two identical characters, then they are not duplicates, otherwise they are duplicates and need to be merged into one. This character deduplication transformation is represented by B, as follows: B(π1)=B(--stta-t--e)=state#(2) B(π2)=B(sst-aaa-tee-)=state#(3) Where "--stta-t--e" and "sst-aaa-tee-" are two strings corresponding to the output of the neural network. Each string is regarded as a string path, represented by π; π1 is path 1 transformed by B to obtain "state"; π2 is the same; "-" represents a blank character, and "s, t, a, e" represent English letters; The probability of a path is: In the above formula, x is the input, represents the probability of each character, T represents the number of characters output by the neural network; assuming the label is l, if the path π t After B transformation, we get label l, then path π t is a correct path; the CTC algorithm maximizes the probability of all correct paths, that is, The CTC loss function is: CTC(x)=-log(l|x)#(6) By minimizing the CTC loss function, the network model can recognize characters of variable length.
Citation Information
Patent Citations
License plate correction and recognition method based on deep learning
CN111598089A
Detection and recognition method for amplified number of license plate
CN112906699A