License plate recognition method based on multi-granularity space attention
By introducing a multi-grained spatial attention mechanism into the license plate recognition technology, the attention to the license plate area is dynamically adjusted, and the problem of unstable recognition effect when dealing with double-travel license plates is solved, achieving higher recognition accuracy and adaptability.
Patent Information
- Application Number
- CN202510098628.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
AI Technical Summary
The existing license plate recognition technology is difficult to effectively adapt to different license plate structures when dealing with diverse license plate styles, especially double-travel license plates, resulting in unstable identification effect.
The license plate recognition method based on multi-grained spatial attention is adopted. By introducing an attention mechanism, the model's attention to different areas of the license plate is dynamically adjusted to improve the accuracy of character positioning and segmentation.
It significantly improves the recognition accuracy and robustness of single-line and double-line license plates, enhances the model's adaptability in complex license plate environments, and reduces character omissions and duplications.
Smart Images

Figure CN119992531A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a license plate recognition method based on multi-granularity spatial attention. Background Art
[0002] As an important part of intelligent transportation system, license plate recognition technology is widely used in vehicle management, traffic monitoring, electronic toll collection and other fields. Traditional license plate recognition methods mainly include image preprocessing, license plate positioning, character segmentation and character recognition. Early license plate positioning usually relies on traditional image processing technologies such as edge detection, color features and shape features. These methods have low recognition accuracy in complex environments such as lighting changes, license plate damage or blur.
[0003] With the rapid development of computer vision and machine learning technologies, license plate recognition methods based on machine learning have gradually become a research hotspot. Support vector machines (SVMs) and neural networks have been introduced into the character recognition stage, significantly improving the accuracy of recognition. However, these methods still face certain challenges when dealing with a variety of license plate styles, especially those containing single-row and double-row layouts. Due to the dense arrangement and complex structure of double-row license plates, traditional methods are prone to errors in character segmentation and arrangement recognition, affecting the overall recognition effect.
[0004] In recent years, deep learning technology has made significant progress in the field of license plate recognition. Convolutional neural networks are widely used in license plate location and character recognition tasks, which can automatically extract high-level features from license plate images and improve the robustness and accuracy of recognition. In particular, the end-to-end recognition method based on region proposal networks and fully convolutional networks simplifies the traditional multi-step processing flow and further improves processing efficiency.
[0005] Although deep learning methods perform well in single-row license plate recognition, they still have certain limitations when dealing with all types of license plates, especially temporary license plates and license plates with two-row layouts. The character arrangement of two-row license plates is complex and the spatial layout is changeable. Existing recognition models may have difficulty adapting to different license plate structures in terms of feature extraction and character positioning. Therefore, there is an urgent need for a recognition mechanism that can effectively capture the spatial relationship of characters in license plates and adapt to multiple layouts.
[0006] Although the existing license plate recognition technology has improved the accuracy and efficiency of recognition to a certain extent, it still has some shortcomings, especially when processing all types of license plates, especially single-row, temporary and double-row license plates, facing the following main shortcomings:
[0007] 1. Lack of adaptability to multiple layouts
[0008] Although existing deep learning models have made significant progress in single-row license plate recognition, they often have difficulty adapting to double-row license plates due to the complexity of character arrangement and variability of spatial layout, resulting in unstable recognition results. Existing methods lack an effective mechanism to dynamically adjust the focus on different license plate areas and have difficulty handling complex relationships between characters.
[0009] 2. It is impossible to accurately locate license plates with different widths in the upper and lower lines.
[0010] Some current recognition methods based on two-dimensional attention will cause the attention of the upper row to shift downward when recognizing license plates with different widths of the upper and lower rows, resulting in problems in recognizing the two characters in the upper row.
[0011] In view of the above problems, the present invention proposes a license plate recognition method based on a multi-granularity spatial attention mechanism. The method introduces an attention mechanism to dynamically adjust the model's attention to different areas in the license plate, thereby more accurately identifying single-row and double-row license plates. This not only improves the accuracy and robustness of recognition, but also enhances the model's adaptability in complex license plate environments. Summary of the invention
[0012] In order to solve the technical problems raised in the background technology, the present invention provides a license plate recognition method based on multi-granularity spatial attention.
[0013] The present invention is implemented by the following technical solution: a license plate recognition method based on multi-granularity spatial attention, comprising the following steps:
[0014] First, a license plate recognition system based on multi-granularity spatial attention is constructed; then the license plate is recognized by the license plate recognition system;
[0015] The license plate recognition system comprises:
[0016] The backbone network module consists of a multi-layer convolutional network, which is used for feature extraction and extracts global feature maps of multiple scales from the input image to represent different field of view features;
[0017] Multi-granularity spatial attention mechanism module; it adopts a position-guided spatial attention mechanism, which can effectively focus on the key information area in the license plate image and fuse the attention feature maps generated by multiple attentions at different scales into a global feature map;
[0018] The multi-head attention decoding mechanism module takes as input the fused global feature map and outputs the final license plate character classification.
[0019] Specifically, the input image of the license plate recognition system based on multi-granularity spatial attention is the image information output after being processed by the DBNet license plate detection model.
[0020] Specifically, the multi-granularity spatial attention mechanism module adopts a position-guided spatial attention mechanism, which can effectively focus on the key information area in the license plate image; and the working steps are as follows:
[0021] Feature extraction: extract high-level features of the input image through two layers of convolution and pooling operations;
[0022] Fully connected layer processing: After flattening the extracted feature map, it is processed through two layers of fully connected layers to generate attention weights;
[0023] Deconvolution and fusion: The deconvolution operation upsamples the feature map and fuses it with the initial convolution feature to generate the final attention map; as shown in the formula:
[0024] f attn =α attn *f
[0025] Where f is the original feature map, α attn is the attention weight generated by the attention network;
[0026] The multi-granularity is reflected in the fact that we apply the attention module to the last layer feature map extracted by the backbone network and the previous layer feature map, and then fuse the two attention feature maps. This can make the obtained attention weight positioning more accurate and avoid the problem of attention sinking of the upper characters.
[0027] Specifically, the multi-head attention decoding mechanism;
[0028] In order to further improve the expressiveness of the attention mechanism, a multi-head attention mechanism is introduced; the specific implementation is as follows: the generated attention map is divided into K parts, each part corresponds to an attention head; each attention head is weighted separately, and the final character classification is generated in parallel;
[0029] In this paper, for each target character y i Use a linear decoder; first, for y i Calculate a context vector c i ∈R 128 ,
[0030] Then use the Softmax function to get the value from c i Prediction y i The probability distribution of
[0031] P(y i |x)=(ρ i1 ,ρ i2 ,ρ i3 ,…,ρ in );
[0032]
[0033] Among them, ρ ij =P(y i =j|x), It is a trainable parameter for Softmax; when the probability distribution P(y i |x), decode the character at position i to y i = argmax j ρ ij , if y i = <end>, all characters following the sequence are discarded.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] This paper proposes a license plate recognition method based on a position-guided multi-granularity spatial attention mechanism, which has the following advantages:
[0036] Enhanced character positioning and segmentation capabilities: By introducing the spatial attention mechanism, the model can dynamically adjust the focus on different areas of the license plate, thereby improving the accuracy of character positioning, especially significantly enhancing the processing capability of double-row license plates in complex environments, reducing character omissions and duplications.
[0037] Good adaptability to multiple layouts: The present invention can adapt to the different layouts of single-row and double-row license plates through a multi-granularity spatial attention mechanism, flexibly adjust the model's processing strategy for different character arrangements, and significantly improve the adaptability and recognition effect to a variety of license plate structures.
[0038] In summary, the present invention effectively overcomes the shortcomings of existing license plate recognition technology in character positioning, feature extraction, multi-layout adaptability and computational efficiency by introducing a position-guided multi-granularity spatial attention mechanism, and significantly improves the recognition accuracy of single-row, temporary and double-row license plates and the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A schematic diagram of the overall structure of the license plate recognition method based on multi-granularity spatial attention provided in Example 1 of the present invention;
[0040] Figure 2 This is a schematic diagram of the structure of the license plate detection model based on DBNet used in the present invention;
[0041] Figure 3 A schematic diagram showing the attention generated by the position of each character of a double-row license plate during the recognition process of the method proposed by the present invention;
[0042] Figure 4 A schematic diagram showing the attention generated at each character position of a common single-row license plate during the recognition process of the method proposed in the present invention;
[0043] Figure 5 A schematic diagram showing the attention generated at each character position of a special license plate and a single-row license plate during the recognition process of the method proposed in the present invention;
[0044] Figure 6 This is a schematic diagram of attention generated on the positions of each character of a special license plate or a temporary license plate during the recognition process of the method proposed in the present invention. DETAILED DESCRIPTION
[0045] The present invention is further described below in conjunction with the accompanying drawings and specific implementation methods. It should be noted that, under the premise of no conflict, the various embodiments or technical features described below can be arbitrarily combined to form a new embodiment.
[0046] Embodiment 1:
[0047] Reference Figure 1-Figure 6 ,This scheme proposes a license plate recognition method based on multi-granularity spatial attention,including the following steps:
[0048] First, a license plate recognition system based on multi-granularity spatial attention is constructed; then the license plate is recognized by the license plate recognition system;
[0049] The license plate recognition system comprises:
[0050] The backbone network module consists of a multi-layer convolutional network, which is used for feature extraction and extracts global feature maps of multiple scales from the input image to represent different field of view features;
[0051] Multi-granularity spatial attention mechanism module; it adopts a position-guided spatial attention mechanism, which can effectively focus on the key information area in the license plate image and fuse multiple attention feature maps of different scales into a global feature map;
[0052] The multi-head attention decoding mechanism module takes as input the fused global feature map and outputs the final license plate character classification.
[0053] Specifically, the input image of the license plate recognition system based on multi-granularity spatial attention is the image information output after being processed by the DBNet license plate detection model.
[0054] Specifically, the multi-granularity spatial attention mechanism module adopts a position-guided spatial attention mechanism, which can effectively focus on the key information area in the license plate image; and the working steps are as follows:
[0055] Feature extraction: extract high-level features of the input image through two layers of convolution and pooling operations;
[0056] Fully connected layer processing: After flattening the extracted feature map, it is processed through two layers of fully connected layers to generate attention weights;
[0057] Deconvolution and fusion: The deconvolution operation upsamples the feature map and fuses it with the initial convolution feature to generate the final attention map; as shown in the formula:
[0058] f attn =α attn *f
[0059] Where f is the original feature map, α attn is the attention weight generated by the attention network;
[0060] The multi-granularity is reflected in the fact that we apply the attention module to the last layer feature map extracted by the backbone network and the previous layer feature map, and then fuse the two attention feature maps. This can make the obtained attention weight positioning more accurate and avoid the problem of attention sinking of the upper characters.
[0061] Specifically, the multi-head attention decoding mechanism;
[0062] In order to further improve the expressiveness of the attention mechanism, a multi-head attention mechanism is introduced; the specific implementation is as follows: the generated attention map is divided into K parts, each part corresponds to an attention head; each attention head is weighted separately, and the final character classification is generated in parallel;
[0063] In this paper, for each target character y i Use a linear decoder; first, for y i Calculate a context vector c i ∈R 128 ,
[0064] Then use the Softmax function to get the value from c i Prediction y i The probability distribution of
[0065] P(y i |x)=(ρ i1 ,ρ i2 ,ρ i3 ,…,ρ in );
[0066]
[0067] Among them, ρ ij =P(y i =j|x), It is a trainable parameter for Softmax; when the probability distribution P(y i |x), decode the character at position i to y i = argmax j ρ ij , if y i = <end>, all characters following the sequence are discarded.
[0068] The license plate recognition system proposed by the present invention includes a position encoding unit;
[0069] The position encoding module is designed to add position information to the input feature map (i.e. the license plate image processed by the DBNet license plate detection model) so that the model can perceive the feature differences at different spatial locations. The specific implementation is as follows:
[0070] Position encoding matrix generation: Use sine and cosine functions to generate a position encoding matrix with a length of max_len and a dimension of d_model.
[0071] Position encoding addition: Add the generated position encoding matrix to the input feature map to obtain a feature representation containing position information.
[0072] The attention mechanism module proposed in this solution adopts a two-dimensional spatial attention mechanism based on position guidance, which can effectively focus on the key information area in the license plate image. The specific implementation steps are as follows:
[0073] 1. Feature extraction: Extract high-level features of the input image through two layers of convolution and pooling operations.
[0074] 2. Fully connected layer processing: After flattening the extracted feature map, it is processed through two layers of fully connected layers to generate attention weights.
[0075] 3. Position encoding application: The generated attention weights are position encoded to enhance the representation of spatial position information.
[0076] 4. Deconvolution and fusion: The deconvolution operation upsamples the feature map and fuses it with the initial convolution feature to generate the final attention map. It includes:
[0077] Convolutional layer: Through continuous convolutional layers and pooling layers, high-level feature information is gradually extracted.
[0078] Deconvolution layer: The low-resolution feature map is upsampled to high resolution through deconvolution operation so as to be fused with the initial feature.
[0079] 5. Fully connected layer and attention weight calculation
[0080] The fully connected layer is used to process the feature map extracted by convolution and generate attention weights. The specific steps include:
[0081] 1. Feature flattening: Flatten the convolutional feature map into a form suitable for the input of the fully connected layer.
[0082] 2. Double-layer fully connected processing: Learn and generate attention weights through two fully connected layers.
[0083] 3. Attention weight reconstruction: The generated attention weights are reconstructed into the same spatial dimension as the original feature map for subsequent attention weighted operations.
[0084] 6. Application of Positional Encoding
[0085] After the attention weights are generated, position encoding is applied to enhance the representation of spatial position information. The specific steps are as follows:
[0086] 1. Feature rearrangement: Rearrange the attention weights to adapt to the input requirements of position encoding.
[0087] 2. Position encoding addition: Add position information to the attention weight through the position encoding module.
[0088] 3. Feature reconstruction: rearrange the encoded attention weights back to the original spatial dimensions.
[0089] 7. Residual Connection
[0090] In order to alleviate the gradient vanishing problem in deep networks and improve the training efficiency and performance of the model, residual connections are introduced. The specific implementation is:
[0091] After the deconvolution operation, the upsampled feature map is added to the initial convolution feature to form a residual connection.
[0092] Through residual connections, the model can learn feature representation more effectively and improve the overall recognition performance.
[0093] 8. Multi-head attention mechanism
[0094] In order to further improve the expressiveness of the attention mechanism, a multi-head attention mechanism is introduced. The specific implementation is:
[0095] The generated attention map is divided into K parts, each part corresponds to an attention head.
[0096] Each attention head is weighted separately, and finally the output of each head is fused to generate the final attention output.
[0097] Output processing is performed using a multi-head attention decoding mechanism module, and the specific steps include:
[0098] 1. Feature rearrangement: Rearrange the convolutional features to adapt to the subsequent recognition model input.
[0099] 2. Attention weighting: The convolutional features are weighted through attention weights to highlight the license plate character area.
[0100] Feature fusion: The weighted features are fused with the initial features to generate the final feature representation for character recognition.
[0101] This technical solution is suitable for Chinese license plate recognition in various complex environments, including recognition requirements for different lighting, different angles, single-line, temporary license plates and double-line license plates. Through the two-dimensional spatial attention mechanism guided by position, the system can effectively focus on the license plate area and improve the accuracy and robustness of recognition. With this solution, the license plates of vehicles in the scene can be effectively recognized, and one model can recognize all types of Chinese license plates. At the same time, it is fast and efficient, avoiding the problem of misrecognition caused by similar Chinese characters.
[0102] The above-mentioned embodiments are only preferred embodiments of the present invention and cannot be used to limit the scope of protection of the present invention. Any non-substantial changes and substitutions made by technicians in this field on the basis of the present invention shall fall within the scope of protection required by the present invention.< / end> < / end>
Claims
1. A license plate recognition method based on multi-granularity spatial attention, characterized in that: The steps include: First, a license plate recognition system based on multi-granularity spatial attention is constructed; then, the license plate is recognized by the license plate recognition system; The license plate recognition system comprises: The backbone network module consists of a multi-layer convolutional network, which is used for feature extraction and extracts global feature maps of multiple scales from the input image to represent different field of view features; Multi-granularity spatial attention mechanism module; it adopts a position-guided spatial attention mechanism, which can effectively focus on the key information area in the license plate image, generate attention-weighted feature maps from global feature maps of multiple scales, and finally fuse the attention maps of different scales to form a global feature map; The multi-head attention decoding mechanism module takes as input the fused global feature map and outputs the final license plate character classification.
2. The license plate recognition method based on multi-granularity spatial attention as claimed in claim 1, characterized in that: The input image of the license plate recognition system based on multi-granularity spatial attention is the image information output after being processed by the DBNet license plate detection model.
3. The license plate recognition method based on multi-granularity spatial attention as claimed in claim 1, characterized in that: The multi-granularity spatial attention mechanism module adopts a position-guided spatial attention mechanism, which can effectively focus on the key information area in the license plate image; and the working steps are as follows: Feature extraction: extract high-level features of the input image through two layers of convolution and pooling operations; Fully connected layer processing: After flattening the extracted feature map, it is processed through two layers of fully connected layers to generate attention weights; Deconvolution and fusion: The deconvolution operation upsamples the feature map and fuses it with the initial convolution feature to generate the final attention map; as shown in the formula: f attn =a attn *f Where f is the original feature map, α attn is the attention weight generated by the attention network.
4. The license plate recognition method based on multi-granularity spatial attention as claimed in claim 1, characterized in that: The multi-head attention decoding mechanism is specifically implemented as follows: the generated attention map is divided into K parts, each part corresponds to an attention head; each attention head is weighted separately, and the final character classification is generated in parallel; In this paper, for each target character y i Use a linear decoder; first, for y i Calculate a context vector c i ∈R 128 , Then use the Softmax function to get the value from c i Prediction y i The probability distribution of P(y i |x)=(ρ i1 ,r i2 ,r i3 ,…,r in ) Among them, ρ ij =P(y i =j|x), It is a trainable parameter for Softmax; when the probability distribution P(y i |x), decode the character at position i to y i = argmax j ρ ij , if y i = <end> , all characters following the sequence are discarded.< / end>