Lightweight license plate recognition method and electronic device
Patent Information
- Application Number
- CN202611037656.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]本发明的目的在于克服现有技术中所存在的车牌识别方法在边缘设备上几何适应性不足、计算效率低下及解码空间过大的缺陷,提供一种轻量化车牌识别方法及电子设备
1.本发明提供一种轻量化车牌识别方法,通过不同层级的可变形卷积层自适应提取车牌图像的多尺度特征图,利用特征金字塔网络进行多尺度融合以丰富空间表征,再经因子化自注意力层对特征序列进行低秩近似建模以降低计算复杂度,最后依据车牌结构化先验约束进行解码以压缩解码空间,从而在保证识别精度的同时显著提升处理效率;
Smart Images

Figure CN122821533A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent transportation technology, and in particular to a lightweight license plate recognition method and electronic device. Background Technology
[0002] License plate recognition is one of the core technologies in intelligent transportation systems, widely used in parking management, traffic monitoring, and electronic toll collection. In recent years, with the popularization of edge computing devices, achieving high-precision and real-time license plate recognition on resource-constrained mobile or embedded devices has become an urgent need for technological development.
[0003] Existing license plate recognition technologies are mainly divided into two categories: two-stage methods and single-stage methods. Two-stage methods, such as those based on a cascaded architecture of detection and recognition, suffer from high cumulative latency due to multiple processing steps and the problem of cascaded error amplification, making it difficult to meet the real-time requirements of edge devices. Single-stage methods, such as those based on an end-to-end architecture of convolutional recurrent neural networks (CRNNs), can increase inference speed to over 60 frames per second, but their recognition accuracy remains unsatisfactory when dealing with the extreme aspect ratios of Chinese license plates (1:3 to 1:4) and complex formats such as those for new energy vehicles.
[0004] In terms of specific technical approaches, methods based on convolutional neural networks (CNNs) use standard convolutional kernels for feature extraction. However, the fixed sampling grid of standard convolutional kernels struggles to adaptively cover the long edge regions of license plates, leading to the loss of edge information, especially under scenarios with large angles of tilt or partial occlusion, resulting in significant performance degradation. While methods based on visual Transformers can model the global context through self-attention mechanisms, their attention computational complexity increases quadratically with sequence length, resulting in high inference latency on edge devices and making real-time processing difficult. Lightweight network designs, such as depthwise separable convolutions, can reduce the number of parameters and computational cost, but lack robustness under complex lighting or geometric deformation conditions.
[0005] Furthermore, existing methods generally neglect the structured prior knowledge of Chinese license plates, such as the finite set of province abbreviations, fixed character lengths, and rules for combining letters and numbers. This prior knowledge could be used to constrain the decoding space and improve recognition accuracy, but existing technologies lack effective means of utilizing it, resulting in an excessively large decoding space for the connected temporal classification decoder (CTC decoder), which easily leads to illegal character combinations in the output.
[0006] In summary, how to achieve both high accuracy and real-time license plate recognition on resource-constrained edge devices, and enhance robustness to complex conditions such as extreme aspect ratios, large-angle tilts, occlusion, and dirt, is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of existing license plate recognition methods, such as insufficient geometric adaptability on edge devices, low computational efficiency, and excessive decoding space, and to provide a lightweight license plate recognition method and electronic device.
[0008] In a first aspect, the present invention provides a lightweight license plate recognition method, comprising:
[0009] The target license plate image is acquired, and multi-scale feature maps of the target license plate image are extracted through deformable convolutional layers of different levels; The multi-scale feature maps are fused using a feature pyramid network to obtain a fused feature map; By using a factorized self-attention layer, low-rank approximate attention modeling is performed on the feature sequence corresponding to the fused feature map to generate a context-enhanced feature sequence. The context-enhanced feature sequence is decoded based on the license plate structured prior constraints, and the recognition result of the target license plate image is output.
[0010] Preferably, the multi-scale feature map of the target license plate image is extracted through deformable convolutional layers of different levels, including: Obtain the learnable offset and modulation scalar corresponding to each sampling point in the deformable convolutional layer; The sampling grid of the convolution kernel is adjusted according to the learnable offset so that the sampling position of the convolution kernel is adaptively focused on the long edge of the license plate region in the target license plate image; The feature weights of the corresponding sampling points are adjusted according to the modulation scalar to enhance the feature contribution of the focusing position.
[0011] Furthermore, the learnable offset of the deformable convolutional layer is trained using an elliptical Gaussian mask: the aspect ratio of the elliptical Gaussian mask is set to match the aspect ratio of the target license plate; the optimal offset is solved by maximizing the feature response integral within the area covered by the elliptical Gaussian mask, so that the sampling points of the convolutional kernel are adaptively stretched and focused on the long edge region of the target license plate.
[0012] Furthermore, the deformable convolutional layer is a DCNv3 layer.
[0013] Preferably, all convolutional layers in the feature pyramid network from bottom to top are DCNv3 layers; the lateral connections use only a single 1×1 convolution for channel alignment; and the final output resolution is a single fused feature map with a fixed resolution of 160×32 to adapt to the aspect ratio of the target license plate.
[0014] Preferably, the low-rank approximate attention modeling of the feature sequence corresponding to the fused feature map is performed through the factorized self-attention layer, including: The query matrix, key matrix, and value matrix corresponding to the feature sequence are projected from the original dimension to a low-rank subspace, where the dimension of the low-rank subspace is smaller than the original dimension. In the low-rank subspace, an approximate attention weight matrix between the query matrix and the key matrix is calculated based on the Nyström approximation method. The computational complexity of the approximate attention weight matrix is proportional to the first power of the length of the feature sequence and the square root of the feature dimension. The value matrix is weighted and aggregated using the approximate attention weight matrix to generate the context-enhanced feature sequence.
[0015] Furthermore, the factorized self-attention layer uses rotational position encoding to encode the relative position information of the characters.
[0016] Preferably, decoding the context-enhanced feature sequence based on the license plate structured prior constraints includes: The structured prior constraints include at least constraints on the province abbreviation, city code, alphanumeric order, and number of digits of the license plate. Construct a prefix tree mask based on the structured prior constraints; During the path search process based on the connection-time classification decoder, the prefix tree mask is used to filter candidate decoding paths and block illegal paths that do not conform to the structured prior constraints. Based on the selected candidate decoding paths, the recognition result of the target license plate image is output.
[0017] Furthermore, during the training phase, the license plate structured prior constraints are transformed into structured constraint losses, and the structured constraint losses are combined with the connection time-series classification sequence alignment loss and the character segmentation auxiliary cross-entropy loss to form a weighted multi-task loss function for end-to-end joint training.
[0018] In a second aspect, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the program, implements a lightweight license plate recognition method as described in any of the first aspects.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention provides a lightweight license plate recognition method, which adaptively extracts multi-scale feature maps of license plate images through deformable convolutional layers of different levels, uses a feature pyramid network for multi-scale fusion to enrich spatial representation, then performs low-rank approximation modeling of the feature sequence through a factorized self-attention layer to reduce computational complexity, and finally performs decoding based on the structured prior constraints of the license plate to compress the decoding space, thereby significantly improving processing efficiency while ensuring recognition accuracy. 2. This invention provides an electronic device that, by embedding the aforementioned lightweight license plate recognition method in memory as a computer program and executing it through a processor, effectively integrates deformable convolutional geometric adaptation, feature pyramid multi-scale fusion, factorized self-attention low-rank approximation modeling, and structured prior constraint decoding into an edge computing hardware platform. This reduces the computational overhead of complex deep learning models on resource-constrained devices and improves the real-time response capability and recognition accuracy of license plate recognition. Attached Figure Description
[0020] Figure 1 This is a flowchart of a lightweight license plate recognition method in one embodiment; Figure 2 This is a schematic diagram of the overall architecture of a lightweight license plate recognition method in the embodiment; Figure 3 This is a visualization diagram of the DCNv3 convolution kernel offset in the embodiment; Figure 4 This is a schematic diagram of the FSA mechanism in the embodiment; Figure 5 This is a schematic diagram of CTC decoding for prior constraints on Chinese license plates in the embodiment. Figure 6 This is a sample data diagram from the example. Figure 7 This is a flowchart illustrating the character segmentation process in the embodiment. Figure 8 This is a comparison image of license plate image enhancement effects under different lighting conditions in the embodiments. Figure 9 This is a visual diagram illustrating the failure modes under the occlusion scenario in the embodiment. Figure 10 This is a diagram analyzing the degree of character adhesion in the embodiment; Figure 11 This is a visualization of the character segmentation process in the embodiment; Figure 12 This is a comparison chart of the license plate recognition system results in the embodiment. Detailed Implementation
[0021] The present invention will now be described in further detail with reference to specific embodiments. However, this should not be construed as limiting the scope of the present invention to the following embodiments; all technologies implemented based on the content of the present invention fall within the scope of the present invention.
[0022] Unless otherwise specified, the terms "upper," "lower," "left," "right," "center," "inner," and "outer," etc., used in the description of specific embodiments of the present invention to indicate orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings, or the orientation or positional relationship in which the product / equipment / device is usually placed during use. These terms are merely for the purpose of facilitating the description of the present invention or simplifying the description in specific embodiments, and for enabling those skilled in the art to quickly understand the solution, and do not indicate or imply that a particular device / component / element must have a specific orientation, or be constructed and operated in a specific positional relationship. Therefore, they should not be construed as limitations on the present invention.
[0023] Furthermore, the use of terms such as "horizontal," "vertical," "suspended," "parallel," and "coaxial" does not imply that the corresponding device / component / element must be absolutely horizontal, vertical, suspended, parallel, or coaxial. Slight tilt or deviation is permissible, as long as it does not affect the normal function of the relevant component. For example, "horizontal" simply means that its direction is more horizontal relative to "vertical," not that the structure must be perfectly horizontal; a slight tilt is acceptable. "Coaxial" means that two components are arranged as coaxially as possible, allowing them to move coaxially or approximately coaxially when their relative positions change. Alternatively, it can be simplified to mean that the corresponding device / component / element, when arranged in "horizontal," "vertical," "suspended," "parallel," or "coaxial" directions, can have an error / deviation of ±10% relative to the corresponding direction, more preferably within ±8%, more preferably within ±6%, more preferably within ±5%, and more preferably within ±4%. For example, the deviation in the "coaxial" direction is controlled within 0.2-1mm, preferably within 0.2-0.5mm. As long as the corresponding device / component / element is within the error / deviation range, it can still achieve its function in the solution of the present invention.
[0024] Furthermore, the use of terms such as "first," "second," and "third" in terminology is merely for distinguishing descriptions of identical or similar components and should not be interpreted as emphasizing or implying the relative importance of a particular component.
[0025] Furthermore, in the description of the embodiments of the present invention, "several", "more than", and "a number of" represent at least two. The number can be any number, such as two, three, four, five, six, seven, eight, or nine, and can even exceed nine.
[0026] Furthermore, in the description of the technical solution of this invention, unless otherwise explicitly specified / limited / restricted, the terms "set up," "install," "connect," "link," "provided with," "laid out," and "arranged" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to connection methods commonly used in the art, such as welding, riveting, bolting, and threaded connections. Such connections can be mechanical, electrical, or communication connections; they can be direct connections or indirect connections through an intermediate medium; and they can refer to the internal communication between two components.
[0027] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0028] First, provide a detailed description of the overall framework.
[0029] The lightweight license plate recognition method proposed in this invention is based on an end-to-end trainable architecture and achieves efficient license plate recognition by leveraging the deep collaboration of deformable convolution and factorized self-attention. Figure 1 This is a flowchart illustrating a lightweight license plate recognition method provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps: First, acquire the target license plate image; then, extract multi-scale feature maps of the target license plate image through deformable convolutional layers of different levels; next, fuse the multi-scale feature maps through a feature pyramid network to obtain a fused feature map; then, perform low-rank approximate attention modeling on the feature sequence corresponding to the fused feature map through a factorized self-attention layer to generate a context-enhanced feature sequence; finally, decode the context-enhanced feature sequence according to the license plate structured prior constraints to output the recognition result of the target license plate image.
[0030] In a specific implementation, such as Figure 2 As shown, this invention uses DCNv3-MobileNetV3-Small as the backbone network. Through multi-scale feature extraction and fusion with a Feature Pyramid Network (FPN), rich spatial feature representations are obtained. Then, a factorized self-attention module is used for sequence modeling, and finally, a Connectionized Temporal Classification (CTC) decoder outputs the license plate number and type prediction. The network has 4.2M parameters, a computational complexity of 2.3G FLOPs, and an input resolution of 160×32. It can achieve a real-time inference performance of 67 frames per second on a Jetson Orin Nano edge device, fully meeting the efficiency requirements of practical deployment.
[0031] The core design philosophy of this architecture lies in the organic integration of three key technical paths. First, through the adaptive learning mechanism of deformable convolutional network (DCN) offsets, the sampling points of the convolutional kernel can be dynamically adjusted to adapt to the geometric changes of the license plate, significantly improving the accuracy of edge feature extraction and solving the problem that standard convolutional kernels cannot cover the long edges of the license plate. Second, by employing factorized self-attention technology, the computational complexity of self-attention is reduced from quadratic to near linear, significantly reducing computational overhead while maintaining sequence modeling capabilities, providing theoretical support for edge device deployment. Finally, a prefix tree mask is constructed using structured prior knowledge of the license plate to filter candidate paths during CTC decoding, effectively narrowing the decoding space and suppressing the generation of illegal character combinations, thereby improving recognition accuracy and decoding efficiency. This multi-path collaborative design not only achieves a balance between accuracy and efficiency but also lays a solid foundation for subsequent theoretical analysis and experimental verification.
[0032] The following description, in conjunction with the accompanying drawings, further details the implementation of the key modules involved in the above method.
[0033] Second, a detailed description of variable convolution is provided. (1) Adaptive mechanism of deformable convolution kernel Standard 3×3 convolution kernels struggle to effectively cover the long edge features of license plates, especially given the 1:3 to 1:4 aspect ratio of Chinese license plates, where traditional convolution operations easily lead to the loss of edge information. To overcome these problems, this invention introduces deformable convolution. By learning a learnable offset for the convolution kernel, the sampling position can be adaptively adjusted, thereby better adapting to the geometric characteristics of the license plate.
[0034] Specifically, this invention replaces all 3×3 depthwise separable convolution operations in the backbone network with DCNv3 operations. Although this replacement results in a slight increase in computation, it significantly improves the accuracy of edge feature extraction. In a preferred implementation, the DCNv3 operation is performed as follows: learning a learnable offset and modulation scalar corresponding to each sampling point in the deformable convolutional layer; adjusting the sampling grid of the convolutional kernel according to the learnable offset so that the sampling position of the convolutional kernel adaptively focuses on the long edge of the license plate region in the target license plate image; and adjusting the feature weights of the corresponding sampling points according to the modulation scalar to enhance the feature contribution of the focused position. This design enables the deformable convolutional kernel to dynamically change its sampling position according to the image content, effectively improving the model's feature extraction capability under extreme aspect ratio license plates and complex lighting conditions while maintaining computational efficiency.
[0035] For a given input feature map ( Let H represent the set of real numbers, W represent the feature map height, and C represent the feature map width. The output of deformable convolution is defined as:
[0036] in, This represents the feature value output at position p in the output feature map by the deformable convolution, where K is the total number of sampling points of the convolution kernel. To output the pixel coordinates on the feature map, For the first Convolution weights for each sampling point For the standard convolution kernel, the first Preset offset for each sampling point For the first Learnable offset of each sampling point For the first The modulation scalar of each sampling point.
[0037] For aspect ratio The learnable offset of DCNv3 in the license plate area The following geometric fitness conditions must be met:
[0038] in, Indicates the offset Find the expected optimal solution that maximizes the integral. For license plate candidate areas, For image intensity, To output the pixel coordinates on the feature map, The general offset variable that the algorithm tries during the search process ( (This is the optimal offset result found by the algorithm) It is an elliptical Gaussian mask with an aspect ratio of . Taking the elliptical model of the license plate area as an example, assuming the major axis is... The minor axis is ,but DCN learns the optimal offset by maximizing the feature response. To achieve this, the network uses gradient descent to optimize the learnable offset, and the gradient calculation process is as follows:
[0039] in, L For the total loss function, This represents the feature value output at position p in the output feature map by the deformable convolution. For license plate candidate areas, To output the pixel coordinates on the feature map, For the first Convolution weights for each sampling point For the standard convolution kernel, the first Preset offset for each sampling point For the first Learnable offset of each sampling point For the first The modulation scalar at each sampling point The feature map spatial gradient is obtained through bilinear interpolation.
[0040] Compared to standard convolution, DCNv3 adds an offset learning branch, increasing the number of parameters. for:
[0041] in, K =9 is The number of sampling points in the convolution kernel, where 2 represents the offset in the x and y directions, 1 represents the modulation scalar, and C is the number of channels. Therefore, for typical channel numbers C, such as 128 or 256, the parameter increase in DCNv3 accounts for only a very small portion of the total parameters in the backbone network. This indicates that while significantly improving geometric adaptability, DCNv3 does not introduce excessive additional parameters, maintaining the network's lightweight characteristics and facilitating deployment on edge devices.
[0042] Deployed on a Jetson Orin Nano edge device, DCNv3 is implemented using CUDA 11.4 and cuDNN 8.6, calling the deform-conv2d interface from the ATen operator library. To accelerate inference, TensorRT 8.5 is used for operator fusion, merging offset calculation and convolution operations into a single CUDA kernel function to reduce memory access overhead. During the training phase, offset... Initialized as a zero vector, modulated scalar It is initialized to 1, and its learning rate is independently set to 0.01, which is one-tenth of the learning rate of the main network weights, in order to stabilize the geometric adaptation process in the early stage of training.
[0043] Offset distribution characteristics under different aspect ratios, such as Figure 3 As shown in (e), this figure visualizes the difference in sampling point offset distribution of DCNv3 under two typical license plate aspect ratios of 1:3 and 1:4. When the aspect ratio... At that time, the standard deviation of the offset in the major axis direction reached 1.87 pixels, which was significantly higher than the 0.62 pixels in the minor axis direction. This confirms that the convolution kernel can adaptively stretch to cover longer edge regions.
[0044] To further verify the geometric adaptability of deformable convolution, this invention compares the edge feature extraction performance under different mask shapes. As shown in Table 1, the elliptical Gaussian mask outperforms the rectangular uniform mask and the Gaussian circular mask in terms of edge intersection-over-union ratio (IoU), edge response peak value, and end-to-end average accuracy (E2E-mAP). Specifically, the elliptical Gaussian mask achieves an edge IoU of 89.2%, an improvement of 7.1 percentage points compared to the rectangular uniform mask; and an end-to-end average accuracy of 83.7%, an improvement of 0.8 percentage points compared to the rectangular uniform mask. Furthermore, ablation experiments (DCNv3-Rectangle) further confirm that the geometric adaptability of the rectangular uniform mask is weaker than that of the elliptical Gaussian mask. These results indicate that deformable convolution based on an elliptical model is more suitable for license plate recognition tasks with extreme aspect ratios.
[0045] Table 1. Comparison of geometric adaptability of different mask shapes
[0046] Note: Jetson Orin Nano, input 160×32. Elliptical Gaussian mask improves E2E-mAP by 0.8pp compared to rectangular mask; DCNv3-Rectangle is ablation experiment using a rectangular uniform mask.
[0047] (2) Deformable convolution mechanism and architecture optimization To verify DCNv3's adaptive focusing capability on license plate edges from an internal model perspective, this invention further employs Grad-CAM for visualization analysis. The specific process is as follows: Forward inference yields the license plate category score. Extract the output feature map of the 3rd DCNv3 layer. ;calculate The channel weights are obtained through global average pooling. Weighted summation yields spatial saliency map The spatial saliency map L is upsampled to the input resolution and binarized with a threshold of 0.5. The ratio of the number of activated pixels within a 5-pixel region of the license plate edge to the number of activated pixels in the non-edge region of the same license plate is defined as the "edge sampling density ratio". Statistical results on 1000 CCPD (Chinese City Parking Dataset) test images show that this ratio has an average of 2.3 and a standard deviation of 0.17, indicating that the learnable offset does indeed guide the convolutional kernel to cluster towards the edges. Figure 3 The visualization results in (c) are consistent.
[0048] After verifying the geometric adaptability and edge-focusing capability of deformable convolution, this invention further evaluates the overall performance of the DCNv3-MobileNetV3-Small backbone network. Six representative lightweight networks were selected for comparative experiments, covering three mainstream design paradigms: CNN, Transformer, and hybrid architectures. All experiments were conducted on a Jetson Orin Nano edge device to ensure the results are relevant for practical deployment. As shown in Table 2, DCNv3-MobileNetV3-Small has the fewest parameters among the eight comparative models, reducing them by 20.8% compared to the second-best, EfficientNet-B0; it is the fastest, improving speed by 21.8% compared to MobileNetV3-Large; it has the highest accuracy, significantly outperforming other models; and it has the highest energy efficiency, improving performance by 35.2% compared to the second-best model. Progressive ablation experiments show (i.e., DCNv3 removal in Table 2) that removing DCNv3 reduces E2E-mAP by 1.8%, replacing only the intermediate layer DCNv3 can partially restore performance, and the fully configured DCNv3 can achieve the optimal 83.7% E2E-mAP, verifying the key role of geometrically adaptive design in license plate recognition tasks.
[0049] Table 2 Backbone Network Performance Comparison (Input Resolution 160) 32)
[0050] Third, a detailed description of factorized self-attention sequence modeling is provided. This invention addresses the characteristic of short sequence lengths in license plate recognition tasks, i.e., the typical sequence length... This paper proposes a Factorized Self-Attention (FSA) mechanism. While maintaining modeling capabilities, this mechanism reduces the computational complexity of self-attention from quadratic to approximately linear levels, effectively filling the theoretical gap in existing low-rank approximation methods, such as Linformer and Nystromformer, for short sequence scenarios.
[0051] In license plate recognition scenarios, the computational complexity of standard self-attention is O(n log n). , where. When the sequence length Feature Dimension At that time, its computational workload reached 131,584 FLOPs, of which more than 99.2% of the computation came from The matrix multiplication corresponding to each term becomes a performance bottleneck on edge devices. To address this issue, this invention reduces the complexity to a minimum using a low-rank approximation. Theoretically, it can achieve an acceleration of about 3.17 times, and in actual tests, it achieved an inference acceleration of 1.9 times.
[0052] The approximation error of the factorization self-attention mechanism proposed in this invention satisfies the following relationship:
[0053] in, For querying the matrix, The key matrix, For value matrices, For standard self-attention output, For factorized self-attention output, d is the feature dimension at each position. The attention matrix of the th A singular value, The Frobenius norm is represented. This error bound indicates that when the singular values of the attention matrix decay rapidly, the approximation error can be effectively controlled within a small range. In actual testing, the approximation error of this invention is less than 2.1%.
[0054] Specifically, the forward computation process of the factorized self-attention layer is as follows: Input query matrix Key matrix Sum matrix ,in Represents the set of real numbers. For sequence length, This invention sets a low-rank dimension as the feature dimension. Nyström sampling points In a preferred implementation, when , hour, , .
[0055] The factorized self-attention layer performs the following operations in sequence: The first step, low-rank projection: This involves mapping high-dimensional features to a low-rank subspace using a linear transformation, preserving key semantic information. The computational complexity of this step is O(n log n). .
[0056] The second step is to compress attention computation: Attention weights are computed in a low-rank subspace, reducing the computational complexity of traditional self-attention from... Reduce to .
[0057] The third step, the Nyström approximation: This involves performing a low-rank decomposition of the attention matrix and sampling... A representative anchor point further reduces the computational complexity to .
[0058] The fourth step is weight aggregation: the attention weights obtained from the low-rank approximation are used to perform a weighted summation of the value matrix to generate the final attention output. .
[0059] The theoretical total complexity of the above algorithm is: .when When, the dominant term is Compared to standard self-attention The computational workload was reduced by approximately Times. Under typical parameter configurations for license plate recognition ( , The actual computational cost is approximately [amount missing] times that of standard self-attention. The number of floating-point operations is reduced to approximately Sub-floating-point operations reduce memory usage by approximately 85%.
[0060] Furthermore, the factorized self-attention layer preferably employs Rotary Position Embedding (RoPE) to encode the relative position information of characters in the sequence. RoPE embeds relative position information through a rotation matrix, achieving superior performance in both character accuracy and position accuracy compared to absolute position encoding and two-dimensional relative position encoding, while having a smaller impact on inference speed.
[0061] Figure 4 The visualization results of the factorized self-attention mechanism proposed in this invention are shown. For example... Figure 4 As shown, standard self-attention requires the computation of complete... The computational complexity of the attention matrix varies with the sequence length. It grows quadratically. In contrast, the FSA mechanism proposed in this invention decomposes the query matrix through low-rank decomposition. Project to The low-rank subspace of the dimensionality significantly reduces computational complexity while maintaining expressive power, making it particularly suitable for medium-length sequence modeling tasks such as license plate recognition.
[0062] Figure 4 (a) and Figure 4 (b) compares the attention distribution and attention matrix of standard self-attention and FSA. It can be seen that FSA effectively preserves the main structure of the attention matrix through low-rank approximation, and its approximation error is controlled within 2.1%, which has a negligible impact on the final recognition accuracy. Figure 4 (c) further illustrates different low-rank dimensions Performance-efficiency trade-off curves for various values. Experimental results show that when... At that time, the approximation error rose sharply to over 8%, leading to a significant decrease in recognition accuracy; when At this point, the improvement in inference speed slows down, and the frame rate gain decays to below 5%. Therefore, this invention preferably sets a low-rank dimension. To achieve the best balance between accuracy and efficiency.
[0063] Furthermore, this invention also compared the impact of different positional encoding methods on sequence modeling performance. As shown in Table 3, with an input resolution of 160×32 and a sequence length of... Under certain conditions, four configurations were tested: 1D absolute position encoding, 2D relative position encoding, Rotated Position Encoding (RoPE), and no position encoding. Experimental results show that Rotated Position Encoding outperforms other encoding methods in both character accuracy and position accuracy, while achieving an inference speed of 67 FPS and a memory footprint of 115MB, achieving the best balance between accuracy and efficiency. In contrast, the character accuracy of no position encoding is only 96.9%, and the position accuracy is only 94.1%, indicating that position information plays a crucial role in the recognition of license plate character sequences. Therefore, this invention preferably adopts Rotated Position Encoding as the default position encoding scheme for the factorized self-attention layer.
[0064] Table 3 Comparison of Position Encoding Performance
[0065] Fourth, a detailed description of structured prior constraint decoding. This invention further utilizes structured prior knowledge of Chinese license plates to constrain and optimize the Connectionized Temporal Classification (CTC) decoding process. Chinese license plates possess distinct structured features, including a fixed length (7-8 characters) and limited provincial codes (abbreviations of provincial administrative regions). Existing methods generally neglect this prior information, leading to an excessively large CTC decoding space, which easily generates illegal character combinations and reduces recognition accuracy. Therefore, this invention proposes a License Plate Prior Loss (LP-Prior) function, which effectively compresses the decoding space by embedding structured prior knowledge into the decoding process. Total Loss Function Defined as:
[0066] in, For standard CTC sequence alignment loss, For prior constraint loss, This represents the balance coefficient. It should be noted that, for ease of analysis of the mechanism of prior constraint loss, only the loss term directly related to the prior constraints is shown here. In actual end-to-end training, the total loss function will also include a character segmentation auxiliary loss term. The full definition will be given in the training strategy section later. The prior constraint loss filters the decoding path by constructing a prefix tree-based mask (Trie mask). Theoretical analysis shows that the LP-Prior loss can improve the lower bound of CTC decoding accuracy:
[0067] In the formula, , This represents the total decoding space size when there are no constraints. The effective decoding space size to conform to the license plate structure rules.
[0068] Balance coefficient The optimal value can be solved using the Lagrange multiplier method. After introducing the L2 regularization term, the objective function is:
[0069] in This is the regularization coefficient, used to prevent excessively strong prior constraints from causing model convergence difficulties. Taking the derivative and setting it to zero, we get:
[0070] Calculated Experiments show that when At that time, the model achieves the optimal balance between convergence speed and recognition accuracy, see Figure 5 (c) Therefore, the present invention preferably sets... .
[0071] This invention calculates the effective path space digit by digit based on the actual rules of Chinese license plates. Specific constraints are as follows: A total of 8 digits; the first digit is the abbreviation of the provincial-level administrative region, including Hong Kong, Macau, and Taiwan; the second digit is 24 uppercase English letters excluding O and I; the third to eighth digits are set according to the differences between new energy vehicle license plates and ordinary license plates: the last 6 digits of new energy vehicle license plates allow a mixture of letters and numbers, while the last 5 digits of ordinary license plates are a mixture of letters and numbers, with letters appearing at most twice and not consecutively, while easily confused characters O and I are removed. Based on the above rules, the effective path space size can be obtained. The total unconstrained decoding space size Therefore, the calculation yields... Expected prior loss Through the aforementioned structural constraints, this invention compresses the CTC decoding space by approximately [amount missing]. The theoretical accuracy improvement is approximately 0.8%.
[0072] In terms of prefix tree construction, this invention uses a Trie data structure to store all valid paths. The time complexity of the construction process is O(n log n). The space complexity is ,in Where is the sequence length. In actual decoding, a Trie tree is used for path searching, which reduces the decoding complexity to [value missing]. This significantly improves decoding efficiency.
[0073] Figure 5 This demonstrates how prior constraints on Chinese license plates can effectively reduce the CTC decoding space. Standard CTC decoding requires searching all possible path combinations, with a computational complexity of O(n log n). ,in For sequence length, For character set size, The sequence length is denoted by . By introducing structured prior knowledge of Chinese license plates, the decoding space is restricted to a subset of valid paths conforming to the license plate rules, significantly reducing search complexity while improving recognition accuracy.
[0074] Balance coefficient Impact on model performance, such as Figure 5 As shown in (c) and Table 4. Figure 5 (c) visualized exist The impact curves on provincial accuracy (PA) and character accuracy (CA) within the interval. When At that time, CTC decoding space compression was insufficient, and the accuracy rate of provinces was below 96%; when At that time, excessive constraints caused the CTC path search to prematurely fall into a local optimum, resulting in a 0.5 percentage point decrease in character accuracy. Within the optimal working range... Within the frame rate, both PA and CA reach their peak values simultaneously, while the frame rate remains stable.
[0075] Table 4 shows the different A comparison of specific performance values. Experimental data shows that when... At that time, the province accuracy reached 96.5%, the character accuracy reached 98.2%, the frame rate was 67 FPS, and the number of epochs required for training convergence was 10, resulting in optimal overall performance. Therefore, the present invention preferably sets... .
[0076] Table 4. Sensitivity Analysis of LP-Prior Loss (License Plate Prior) Weights
[0077] Seventh, Experimental Verification and Performance Analysis (1) Dataset and training strategy To comprehensively evaluate the performance of the lightweight license plate recognition method proposed in this invention, a large-scale dataset containing 342,110 images was constructed. This dataset accurately covers seven license plate types and systematically incorporates extreme condition samples to verify the robustness boundaries of the model. Specifically, the dataset covers four time periods: early morning, noon, evening, and night; 12 meteorological conditions including sunny, cloudy, rainy, and foggy; and seven typical lighting scenarios: natural light, streetlights, vehicle lights, backlighting, sidelighting, low light, and strong light. In particular, 16.1% of all samples are challenging samples, totaling 55,080 images, including 15,340 occluded samples with an occlusion area of 20% to 70%, 12,280 soiled samples such as mud and dust, 18,540 samples with a large-angle tilt and a shooting angle of ±45 degrees, and 8,920 low-resolution samples with a width of less than 80 pixels. The above dataset structure provides sufficient assurance for verifying the superiority of the method of this invention in terms of diversity and environmental robustness. The specific distribution of various license plate samples is as follows: Figure 6 As shown.
[0078] This invention employs an end-to-end joint training approach to optimize lightweight neural network parameters. Total loss function. Defined as weighted multi-task loss:
[0079] in, The sequence alignment loss for Connectionist Temporal Classification (CTC) is used to align the entire character sequence, and its weight is fixed at 1.0. To address the structured prior constraint loss for license plates, a prefix tree mask is constructed based on the structured prior of Chinese license plates. A 0 / 1 mask is applied to the CTC decoding path to forcibly block illegal paths. Hyperparameters are also included. The preferred setting is 0.5; As an auxiliary loss for character segmentation, pixel-level cross-entropy loss is calculated for the character segmentation branch to improve edge localization accuracy. (Hyperparameter) The preferred setting is 0.3.
[0080] like Figure 7 As shown, this invention introduces a character segmentation auxiliary task during the training process. The specific process includes: image preprocessing of the original license plate image, including color space conversion, size normalization, and adaptive thresholding; detecting valley positions through vertical projection analysis to determine character boundaries; after character segmentation and extraction, the results are processed into two branches: one for single-character recognition, and the other for result fusion and verification together with contextual constraints and structured verification, ultimately outputting the license plate type and character sequence. The loss of this auxiliary branch... This refers to the pixel-level cross-entropy of the segmentation mask, which can effectively improve the model's accuracy in locating character edges.
[0081] To verify the robustness of the method of the present invention in different lighting environments, three subsets were divided from the constructed outdoor dataset of 340,000 images: low light, normal light, and strong light, with the sample ratio of each subset being approximately 3:4:3. Figure 8 The image enhancement effects under three lighting conditions are compared. As shown in Table 5, under normal lighting conditions, the end-to-end average accuracy (E2E-mAP) of the proposed method reaches 83.7%, the character segmentation accuracy is 98.9%, and the province segmentation accuracy is 96.5%. Under low lighting conditions, the end-to-end average accuracy is 83.1%, the character segmentation accuracy is 98.7%, and the province segmentation accuracy is 96.2%, a decrease of only 0.6 percentage points compared to normal lighting conditions; under strong lighting conditions, the end-to-end average accuracy is 83.4%, the character segmentation accuracy is 98.8%, and the province segmentation accuracy is 96.4%, a decrease of only 0.3 percentage points. These results demonstrate that the proposed multi-task loss function and geometric prior constraints exhibit excellent stability under extreme lighting conditions. The entire training process consists of 120 epochs, with the minimum learning rate decaying to... .
[0082] Table 5 Final performance for different illumination subsets
[0083] Note: Data are the mean ± standard deviation of three independent experiments. To verify the performance boundary of the method of the present invention in severely degraded scenarios, the test set was divided into multiple subsets according to the challenge type, including occlusion (20% to 70% area occlusion), dirt (such as mud and dust), large-angle tilt (with a shooting angle of ±45°), and low resolution (width less than 80 pixels). The results are shown in Table 6.
[0084] Table 6 Performance degradation analysis under different scenarios
[0085] As shown in Table 6, occlusion causes a 5.5 percentage point decrease in the end-to-end average accuracy, which is the main failure mode of the method of this invention. Figure 9 The visualization results show that when the license plate edge is occluded, DCNv3's offset prediction can still focus on the visible edge, but the factorized self-attention sequence modeling produces misjudgments due to the lack of character position information. In contrast, the performance degradation is smaller in the smudged scene because it preserves the complete geometric structure. The performance degradation is minimal for low-resolution samples, thanks to the low-rank characteristic of factorized self-attention's ability to suppress noise.
[0086] This invention also includes challenging tests such as illumination variations and character adhesion to verify the robustness of the character segmentation algorithm. Under normal, low, and strong illumination conditions, the character segmentation accuracy of the method in this invention reaches 98.9%, 98.7%, and 98.8%, respectively, demonstrating good illumination adaptability (as shown in Table 5). To address the character adhesion problem, this invention employs an improved projection analysis algorithm. Through adaptive threshold valley detection and contextual constraints, it achieves a segmentation accuracy of 95.2% in moderately adhered scenarios, significantly outperforming the traditional fixed-threshold vertical projection segmentation method. Figure 10 The results of the character adhesion analysis are presented, further confirming the effectiveness and robustness of the method of the present invention in complex and challenging scenarios.
[0087] (2) Ablation test To systematically evaluate the independent contributions and interactions of each core module in this invention, a progressive ablation experiment was designed in this embodiment. The baseline model used was MobileNetV3-Small with a CTC decoder, which did not include DCNv3, factorized self-attention FSA, or license plate structured prior LP-Prior. Based on this, single modules and combined modules were gradually added, and the experimental results are shown in Table 7.
[0088] Table 7. Progressive ablation experiments and synergistic effect analysis
[0089] Note: Synergistic gain = Actual gain - Sum of independent gains of each component, the formula is:
[0090] As shown in Table 7, adding DCNv3, FSA, or LP-Prior individually improves the end-to-end average accuracy (E2E-mAP) by 3.2, 3.9, and 2.6 percentage points, respectively. When DCNv3 and FSA are used together, the improvement reaches 7.1 percentage points, producing a synergistic gain of 0.7 percentage points; when DCNv3 and LP-Prior are used together, the improvement is 6.0 percentage points, producing a synergistic gain of 0.2 percentage points; and when all three are used together, the improvement reaches 8.5 percentage points, with a synergistic gain of 0.8 percentage points. These results indicate that in this invention, deformable convolution, factorized self-attention, and structured prior constraints form a positive feedback loop, rather than a simple superposition. Specifically, DCNv3 guides the convolutional kernel sampling points to the long edge of the license plate through adaptive offset, significantly improving the integrity of edge features and providing higher-quality feature input for FSA. FSA suppresses background noise in the low-rank subspace, reducing the attention entropy of sequence modeling, thus helping LP-Prior construct a more compact prefix tree mask. LP-Prior compresses the CTC decoding space by approximately 10^16 times, strengthening the constraint on DCNv3's edge focusing capability during backpropagation, thereby forming a perception-decision-constraint closed loop. This nonlinear coupling mechanism verifies the rationality and superiority of the architecture design of this invention.
[0091] (3) Character segmentation visualization To further verify the effectiveness of the character segmentation algorithm, this invention visualizes the complete character segmentation process. For example... Figure 11 As shown, the process sequentially includes: acquisition of the original license plate image, image preprocessing (including color space conversion, size normalization, and adaptive thresholding), vertical projection analysis, valley detection to determine character boundaries, and final character segmentation and extraction. This algorithm accurately identifies the boundaries between characters through projection analysis, maintaining a 95.2% segmentation accuracy even under complex lighting and character adhesion conditions, laying a solid foundation for subsequent single-character recognition. Compared with traditional end-to-end recognition methods, the segmentation strategy adopted in this invention significantly improves edge localization accuracy and overall recognition robustness.
[0092] (4) Platform deployment verification To verify the generalization and deployment capability of the lightweight license plate recognition method proposed in this invention in a real-world engineering environment, this embodiment was tested on edge computing platforms with different computing power levels and power consumption constraints. Ten license plate images containing complex scenarios such as occlusion, damage, and multiple license plates were randomly selected. A 100% consistency rate of recognition results was achieved on all test platforms, confirming the model robustness and hardware adaptability of the method of this invention.
[0093] The recognition results of some test samples are as follows Figure 12 As shown. By Figure 12It is known that the method of the present invention can accurately output license plate numbers in various complex scenarios, including blue plates, yellow plates, new energy green plates, Hong Kong and Macao entry and exit license plates, and has strong resistance to interference such as uneven lighting, tilt and partial occlusion.
[0094] The above deployment verification results show that the method of the present invention has good real-time performance and robustness on resource-constrained edge devices, and can meet the application requirements of practical intelligent transportation systems.
[0095] In summary, this invention addresses the technical challenges of real-time license plate recognition on edge devices by proposing a lightweight license plate recognition method and electronic device based on the collaborative design of deformable convolution and factorized self-attention. Through the collaborative design of deformable convolution, factorized self-attention, and structured prior knowledge of the license plate, an effective balance between recognition accuracy and processing efficiency is achieved.
[0096] Theoretically, this invention establishes the geometric adaptation theorem for DCNv3 under typical license plate aspect ratios of 1:3 to 1:4, proving that the lower bound of edge feature coverage can reach 89%; it also verifies the effectiveness of low-rank factorized attention in short sequences. In this scenario, the approximation error bound is less than 2.1%, and the computational complexity is reduced from... Effectively reduced to At the application level, the LP-Prior loss function based on prefix tree masking compresses the CTC decoding space by approximately This significantly reduces the generation of illegal decoding paths. On a Jetson Orin Nano edge device, the method of this invention achieves a real-time inference speed of 67 FPS with 4.2M parameters and an end-to-end average accuracy of 83.7%, which is 7.69 percentage points higher than mainstream methods, providing a feasible system-level solution for license plate recognition in resource-constrained scenarios.
[0097] Eighth, the electronic device in the embodiments of the present invention will be described.
[0098] Based on the same inventive concept, embodiments of the present invention also provide an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the method described in any of the above-mentioned embodiments.
[0099] For example, the electronic device may be a central core processor, network controller, base station equipment, edge server, or independent network resource management server in a power distribution network communication network. The processor may include one or more processing units, such as a neural network processing unit (NPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a digital signal processor (DSP), a baseband processor, etc. Different processing units may be independent devices or integrated into one or more processors. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.
[0100] The memory can be used to store executable program code, including instructions. Internal memory may include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function, etc. The data storage area may store data created during the use of the electronic device. Furthermore, internal memory may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFs), etc. The processor executes various functional applications and data processing of the electronic device by running instructions stored in the internal memory and / or instructions stored in memory located within the processor.
[0101] The license plate image information used in this invention is only for algorithm verification and technology demonstration. All data comes from public datasets or has been anonymized, and complies with the relevant regulations on reasonable use.
[0102] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A lightweight license plate recognition method, characterized in that, include: The target license plate image is acquired, and multi-scale feature maps of the target license plate image are extracted through deformable convolutional layers of different levels; The multi-scale feature maps are fused using a feature pyramid network to obtain a fused feature map; By using a factorized self-attention layer, low-rank approximate attention modeling is performed on the feature sequence corresponding to the fused feature map to generate a context-enhanced feature sequence. The context-enhanced feature sequence is decoded based on the license plate structured prior constraints, and the recognition result of the target license plate image is output.
2. The lightweight license plate recognition method according to claim 1, characterized in that, Multi-scale feature maps of the target license plate image are extracted using deformable convolutional layers of different levels, including: Obtain the learnable offset and modulation scalar corresponding to each sampling point in the deformable convolutional layer; The sampling grid of the convolution kernel is adjusted according to the learnable offset so that the sampling position of the convolution kernel is adaptively focused on the long edge of the license plate region in the target license plate image; The feature weights of the corresponding sampling points are adjusted according to the modulation scalar to enhance the feature contribution of the focusing position.
3. The lightweight license plate recognition method according to claim 2, characterized in that, The learnable offset of the deformable convolutional layer is trained using an elliptical Gaussian mask: the aspect ratio of the elliptical Gaussian mask is set to match the aspect ratio of the target license plate; the optimal offset is solved by maximizing the feature response integral within the area covered by the elliptical Gaussian mask, so that the sampling points of the convolutional kernel are adaptively stretched and focused on the long edge region of the target license plate.
4. The lightweight license plate recognition method according to claim 3, characterized in that, The deformable convolutional layer is a DCNv3 layer.
5. The lightweight license plate recognition method according to claim 1, characterized in that, All convolutional layers in the feature pyramid network from bottom to top are DCNv3 layers; the lateral connections use only a single 1×1 convolution for channel alignment; the final output is a single fused feature map with a fixed resolution of 160×32 to adapt to the aspect ratio of the target license plate.
6. The lightweight license plate recognition method according to claim 1, characterized in that, The factorized self-attention layer performs low-rank approximate attention modeling on the feature sequence corresponding to the fused feature map, including: The query matrix, key matrix, and value matrix corresponding to the feature sequence are projected from the original dimension to a low-rank subspace, where the dimension of the low-rank subspace is smaller than the original dimension. In the low-rank subspace, an approximate attention weight matrix between the query matrix and the key matrix is calculated based on the Nyström approximation method. The computational complexity of the approximate attention weight matrix is proportional to the first power of the length of the feature sequence and the square root of the feature dimension. The value matrix is weighted and aggregated using the approximate attention weight matrix to generate the context-enhanced feature sequence.
7. The lightweight license plate recognition method according to claim 6, characterized in that, The factorized self-attention layer uses rotational position encoding to encode the relative position information of characters.
8. The lightweight license plate recognition method according to claim 1, characterized in that, Decoding the context-enhanced feature sequence based on license plate structured prior constraints includes: The structured prior constraints include at least constraints on the province abbreviation, city code, alphanumeric order, and number of digits of the license plate. Construct a prefix tree mask based on the structured prior constraints; During the path search process based on the connection-time classification decoder, the prefix tree mask is used to filter candidate decoding paths and block illegal paths that do not conform to the structured prior constraints. Based on the selected candidate decoding paths, the recognition result of the target license plate image is output.
9. The lightweight license plate recognition method according to claim 8, characterized in that, It also includes transforming the license plate structured prior constraints into structured constraint losses during the training phase, and combining the structured constraint losses with the connection time-series classification sequence alignment loss and the character segmentation auxiliary cross-entropy loss to form a weighted multi-task loss function for end-to-end joint training.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements a lightweight license plate recognition method as described in any one of claims 1 to 9.