Intelligent parking lot license plate recognition system based on deep learning

Through multimodal sensor data fusion and environmental adaptation strategies, the recognition accuracy and stability problems of existing license plate recognition systems in complex environments have been solved, and efficient license plate recognition in all weather and all scenarios has been achieved.

CN120766263AInactive Publication Date: 2025-10-10SHENZHEN CHIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510899331.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing license plate recognition systems lack recognition accuracy, system stability, and environmental adaptability in complex and changeable actual parking environments, especially under severe weather and lighting conditions, making it difficult to meet practical application requirements in all weather and all scenarios.

Method used

An intelligent parking lot license plate recognition system based on deep learning is adopted. Through multimodal sensor data fusion (visible light images, infrared images and millimeter wave radar data), combined with super-resolution reconstruction, feature extraction and deconstruction, environmental perception and assessment, feature fusion and recognition modules, dynamic weight allocation and complementary enhancement fusion are achieved to improve the system's recognition performance in extreme environments.

Benefits of technology

It maintains stable recognition performance under various extreme environmental conditions, improves recognition accuracy, achieves 24-hour reliable operation, improves the robustness and adaptability of the system in complex environments, and meets practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766263A_ABST
    Figure CN120766263A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of license plate recognition, and discloses an intelligent parking lot license plate recognition system based on deep learning, and the system comprises a data obtaining and preprocessing module which obtains multi-modal data and carries out the preprocessing of the multi-modal data; the super-resolution reconstruction module is used for carrying out modal specific super-resolution reconstruction on the preprocessed multi-modal data; the feature extraction and deconstruction module is used for extracting features from the multi-modal data after super-resolution reconstruction and deconstructing the extracted features; the environment perception and evaluation module is used for carrying out reliability evaluation on the multi-modal data after feature deconstruction based on the current environment condition and generating a dynamic weight distribution strategy; the feature fusion and recognition module is used for carrying out complementary enhancement fusion on the deconstructed multi-modal features and carrying out license plate recognition; according to the invention, through multi-modal sensor fusion and an environment adaptive strategy, stable identification performance can be maintained under various extreme environment conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of license plate recognition, and more specifically, to an intelligent parking lot license plate recognition system based on deep learning. Background Art

[0002] With the acceleration of urbanization and the continued growth of car ownership, intelligent parking management systems have become an essential component of modern urban traffic management. License plate recognition, as the core technology of intelligent parking systems, directly impacts parking efficiency and user experience through its accuracy and stability.

[0003] Currently, existing license plate recognition technologies primarily rely on a single visual sensor, typically using a visible light camera for image acquisition. These sensors then use image processing and pattern recognition algorithms to locate, segment, and identify the license plate. While these technologies can achieve high recognition accuracy under ideal conditions, they face numerous technical challenges in practical applications.

[0004] The main problems of existing technologies include: The limitation of single modal data. Traditional license plate recognition systems mainly rely on image information obtained by visible light cameras. In adverse weather conditions (such as rain, snow, fog, dust, etc.) and extreme lighting environments (such as strong direct sunlight, weak light at night, backlight, etc.), the image quality is severely degraded, the contrast of the license plate area is reduced, and the edges of the characters are blurred, resulting in a decrease in recognition accuracy; multimodal data fusion technology is immature. Although studies have tried to combine multiple sensor data to improve recognition performance, due to differences in feature space, data format and physical meaning of different modal data (visible light images, infrared images, millimeter wave radar data, etc.), existing fusion methods mostly use simple feature splicing or weighted averaging strategies, which are difficult to fully explore the complementary information between different modalities and may even introduce noise interference, affecting the overall recognition effect; the problem of insufficient environmental adaptability. License plate recognition systems usually adopt fixed model structures and parameter configurations, and lack the ability to dynamically adjust processing strategies according to real-time environmental conditions. When environmental conditions change, the system cannot adaptively optimize the recognition algorithm, resulting in unstable recognition performance in diverse and complex environments; sensor data utilization strategies lack intelligence. Existing systems usually adopt a fixed weight distribution method when processing multi-sensor data. They are unable to dynamically select the optimal sensor combination or adjust the importance weight of each modal data based on current environmental conditions and the quality of each sensor data, limiting the system's recognition capabilities in complex scenarios; image enhancement technology is not very targeted. Traditional image enhancement methods (such as histogram equalization, sharpening filtering, denoising algorithms, etc.) are mainly aimed at general image processing, lack special optimization of text features in the license plate area, and cannot effectively improve the recognizability of license plate characters.

[0005] These technical problems result in that the existing license plate recognition system has obvious insufficient recognition accuracy, system stability and environmental adaptability in complex and changeable actual parking lot environment, especially in bad weather and light conditions, and is difficult to meet the actual application requirements of all-weather and all-scenarios, thereby restricting the popularization and application of intelligent parking lot system.

[0006] Therefore, it is urgent to develop an intelligent license plate recognition technology which can effectively fuse multi-modal sensor data, has environmental self-adaptive ability and can maintain high recognition accuracy in various complex environmental conditions. SUMMARY

[0007] The application provides an intelligent parking lot license plate recognition system based on deep learning, and solves the technical problems of low recognition accuracy, insufficient environmental adaptability and insufficient multi-modal data fusion of a single modal sensor in a complex environmental condition in the related art.

[0008] The application provides an intelligent parking lot license plate recognition system based on deep learning, comprising: A data acquisition and preprocessing module acquires multi-modal data including visible light images, infrared images and millimeter wave radar data, and pre-processes the multi-modal data; A super-resolution reconstruction module performs modal-specific super-resolution reconstruction on the pre-processed multi-modal data; A feature extraction and deconstruction module extracts features from the super-resolution reconstructed multi-modal data, and deconstructs the extracted features into shared information and modal-specific information; An environment perception and evaluation module performs reliability evaluation on the multi-modal data after feature deconstruction based on the current environmental condition, and generates a dynamic weight distribution strategy; A feature fusion and recognition module performs complementary enhancement fusion on the deconstructed multi-modal features according to the dynamic weight distribution strategy, and performs license plate recognition.

[0009] Further, the pre-processing step of the multi-modal data comprises: The multi-modal data collected by different sensors are spatio-temporally aligned; The multi-modal data are preliminarily filtered for noise; According to the physical characteristics of the multi-modal data, modal-specific enhancement processing is performed.

[0010] Further, the modal-specific super-resolution reconstruction step comprises: A license plate text guided super-resolution network is applied to the visible light image, and the license plate text guided super-resolution network comprises a text attention module for adaptively enhancing the license plate character region; Applying a temperature difference preserving super-resolution network to infrared images, the temperature difference preserving super-resolution network maintains the temperature difference of the object boundary through the temperature gradient preserving branch; Apply a density-enhanced upsampling algorithm to millimeter-wave radar data to improve spatial resolution.

[0011] Furthermore, the step of performing reliability assessment on the multimodal data after feature deconstruction based on the current environmental conditions includes: Build an environmental condition assessment module to detect the environmental parameters of the current scene, including weather conditions and lighting conditions; For multimodal data, extract key parameters that reflect the quality of each modal data; Based on environmental conditions and quality parameters, calculate the reliability score of each modal data in the current environment; According to the reliability scores of multimodal data, a dynamic weight allocation strategy for feature fusion is generated through softmax normalization.

[0012] Furthermore, the complementary enhancement fusion step includes: Apply a two-stream attention network to extract and enhance complementary information between different modalities; The conditional style transfer technology is applied to the fused features to map the license plate features in complex environments to the feature distribution under standard conditions.

[0013] Furthermore, the conditional style transfer technology adopts an architecture based on adaptive instance normalization, which includes a content encoder, a style encoder and a decoder.

[0014] Furthermore, the license plate recognition step includes: Detect and locate the license plate area using a region proposal network based on fused features; For the located license plate area, an attention-enhanced sequence recognition model is applied for character recognition; Verify and optimize the recognition results, apply prior knowledge of license plate formats to check legitimacy, and combine timing information for consistency optimization.

[0015] Furthermore, the character recognition adopts a dual-path decoding strategy of hybrid CTCAttention, which simultaneously utilizes connection temporal classification and attention mechanism to improve recognition accuracy.

[0016] Furthermore, the environmental condition assessment module adopts a two-stage structure: The first stage uses convolutional neural networks to extract environmental features; The second stage uses a multi-task classification head to predict weather type and light intensity.

[0017] The present invention provides a computer storage medium comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned deep learning-based intelligent parking lot license plate recognition system.

[0018] The beneficial effects of the present invention are: through multimodal sensor fusion and environmental adaptation strategy, the system can maintain stable recognition performance in various extreme environmental conditions (such as heavy fog, heavy rain, strong backlight, low light at night, etc.), improve recognition accuracy, and achieve 24-hour reliable operation; This invention fully utilizes the complementary advantages of different sensors. For example, visible light provides high-resolution detail information, infrared provides temperature distribution information unaffected by light, and millimeter waves provide structural information with strong penetrating power. Through feature deconstruction and reconstruction and complementary information enhancement, it overcomes the limitations of a single modality and improves the robustness of the system in harsh environments. Based on the modal reliability assessment mechanism of environmental perception, the present invention can automatically adjust the weight distribution of each modality according to the real-time environmental conditions, intelligently select the most reliable information source in the current environment, so that the system performance remains stable when the environment changes and the adaptability is improved; Through the feature deconstruction and reconstruction mechanism, the present invention can effectively separate and utilize the shared information and private information of each modality, overcome the redundancy and inconsistency problems caused by direct feature splicing, achieve more efficient feature expression and utilization, and improve computing resource utilization; While maintaining high recognition accuracy, the present invention is optimized for actual parking lot application scenarios. The processing speed meets real-time requirements, the hardware requirements are moderate, and it is easy to deploy and maintain. It improves the reliability and user experience of the smart parking system and is suitable for actual application needs in various complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flow chart of the smart parking lot license plate recognition system based on deep learning in the present invention; Figure 2 This is a line graph comparing the license plate recognition accuracy of the solution of the present invention and the traditional single visible light solution under different environmental conditions; Figure 3 This is a bar chart comparing the system response time of the solution of the present invention and the traditional single visible light solution under different environmental conditions; Figure 4 It is a radar chart of the reliability scores of different modal sensors under various environmental conditions; Figure 5 It is a scatter plot of feature quality assessment before and after feature deconstruction and reconstruction; Figure 6It is a pie chart of the computing resource allocation of the multimodal fusion license plate recognition system. DETAILED DESCRIPTION

[0020] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0021] At least one embodiment of the present invention discloses an intelligent parking lot license plate recognition system based on deep learning, such as Figure 1 Shown, including: The data acquisition and preprocessing module acquires multimodal data including visible light images, infrared images, and millimeter-wave radar data, and preprocesses the multimodal data; This step acquires and preprocesses multimodal data in the complex environment of a parking lot, and specifically includes the following sub-steps: Step 1.1, multimodal data acquisition; A multimodal sensor suite, including visible light cameras, infrared cameras, and millimeter-wave radar, deployed at the entrance and exit of the parking lot, simultaneously collects multimodal data from the area vehicles pass through. The visible light cameras capture color images with a resolution of at least 1920×1080 pixels; the infrared cameras collect thermal imaging data corresponding to the visible light camera's field of view; and the millimeter-wave radar collects reflection intensity and distance information in the 76-81 GHz frequency range.

[0022] Step 1.2, spatiotemporal alignment processing; Align the collected multimodal data in time and space to ensure synchronization and consistency across different sensor data. Temporal alignment is achieved through hardware triggering or software timestamp calibration, while spatial alignment is achieved through sensor calibration and coordinate transformation, generating multimodal data in a unified spatiotemporal reference frame.

[0023] Step 1.3, preliminary noise filtering; Perform preliminary noise filtering and preprocessing on each modality's data. Adaptive median filtering is applied to visible light images to remove noise; temperature equalization is applied to infrared images to enhance contrast; and spatial filtering is applied to millimeter-wave radar data to reduce clutter interference.

[0024] Step 1.4, physical model-guided modality-specific enhancement; Based on the physical characteristics of different modalities, corresponding enhancement algorithms are applied. Visible light images use a dark channel prior dehazing algorithm and adaptive gamma correction; infrared images use histogram normalization and edge-preserving filtering; and millimeter wave data uses a distance-adaptive enhancement algorithm to improve the consistency of near-field and far-field data.

[0025] Through the above sub-steps, this step outputs preprocessed and preliminarily enhanced multimodal data, laying the foundation for subsequent feature extraction and fusion.

[0026] The super-resolution reconstruction module performs modality-specific super-resolution reconstruction on the pre-processed multimodal data; Perform specific super-resolution reconstruction on the pre-processed multimodal data to improve the spatial resolution and detail clarity of each modality data. This includes the following sub-steps: Step 2.1: License plate text-guided visible light super-resolution reconstruction; For visible light images, according to the embodiments of the present application, a license plate text feature-guided super-resolution network is constructed. The network includes a feature extraction module, a text attention module, and a reconstruction module.

[0027] The feature extraction module consists of 5 residual blocks, each of which contains two 3×3 convolutional layers and a ReLU activation function for extracting multi-level features. The text attention module is one of the innovations of this application. It focuses on the text part in the license plate area through adaptive learning. This module receives the output of the feature extraction module, first reduces the dimension through 1×1 convolution, then applies spatial attention calculation, generates an attention weight map, and finally performs weighted fusion with the original features. It should be noted that the text attention weight map generated by the sigmoid activation function is weighted multiplied with the original features to highlight the important features of the license plate text area. The calculation formula of the text attention weight map is: ; in, is the text attention weight map; Input for super-resolution feature map; is the sigmoid activation function; is the second layer weight matrix of the text attention module; is the rectified linear unit activation function; is the weight matrix of the first layer of the text attention module; is the first layer bias term; is the second layer bias term.

[0028] In some implementations, in addition to the single-layer attention mechanism described above, a hierarchical multi-scale attention structure can be employed. This structure comprises multiple attention branches with different receptive field sizes, each focusing on character features at different scales. The outputs of these branches are then weighted and combined. For example, attention branches with 3×3, 5×5, and 7×7 convolution kernels can be used to capture fine-grained, medium-grained, and coarse-grained text features, respectively.

[0029] Optionally, the text attention module can be enhanced by incorporating external license plate prior knowledge. By incorporating supervisory signals about license plate character positions during training, the network can learn more accurate text region localization capabilities. This prior knowledge can be provided by labeling character position heatmaps in the training set.

[0030] In practical applications at parking entrances, it's important to note that this license plate text-guided attention algorithm can effectively address issues where parts of the license plate are blurred due to lighting variations. For example, when a vehicle enters a parking lot in strong backlight, the license plate area may be overexposed, making it difficult to recognize some characters. In this case, the text attention module automatically determines which areas contain recognizable characters and specifically enhances the feature representation of these areas, ultimately reconstructing a clearer license plate text.

[0031] Step 2.2, temperature difference-preserving infrared super-resolution reconstruction; For infrared images, this application provides a temperature difference preserving super-resolution network. This network improves spatial resolution while maintaining the original temperature distribution characteristics of thermal imaging, avoiding the temperature information distortion that may be caused by conventional super-resolution methods.

[0032] This network employs an encoder-decoder structure. The encoder consists of three downsampling blocks, each consisting of a 4×4 convolutional layer (with a stride of 2) and a LeakyReLU activation function. The decoder consists of three upsampling blocks, each consisting of a 4×4 deconvolutional layer and a ReLU activation function. Specifically, a temperature gradient-preserving branch is added between the encoder and decoder. This branch extracts temperature edge information using a gradient operator and maintains this gradient information during reconstruction. The network maintains temperature differences at object boundaries through residual learning and a temperature gradient-preserving loss function. Furthermore, the loss function consists of a mean squared error loss and a temperature gradient-preserving loss, weighted together by a balancing coefficient.

[0033] Optionally, in some embodiments, the temperature gradient preservation loss can be further subdivided into a directional gradient loss and a magnitude gradient loss to more comprehensively preserve temperature edge characteristics. The directional gradient loss focuses on the directional consistency of temperature changes, while the magnitude gradient loss focuses on the intensity consistency of temperature changes. The combination of the two can better preserve the structural characteristics of the heat map.

[0034] This temperature-difference-preserving super-resolution network demonstrates significant advantages in parking lot applications at night or in rainy and foggy environments. For example, at night, license plate images captured by infrared cameras often have low resolution, but there is a temperature difference between the license plate and the surrounding environment. By preserving the temperature gradient, this network improves spatial resolution while retaining the clear temperature boundary at the edge of the license plate, enabling subsequent processing to more accurately segment and identify the license plate area.

[0035] Step 2.3, density-enhanced mmWave upsampling; For millimeter-wave radar data with low spatial resolution, this application uses a density-enhanced upsampling algorithm to improve its spatial accuracy. This algorithm combines point cloud density priors with surface continuity constraints, and through iterative nearest point interpolation and local surface fitting, optimizes the spatial distribution of sparse millimeter-wave data to generate a higher-density point cloud representation.

[0036] Step 2.4, multi-scale feature fusion optimization; The super-resolution results of each modality are optimized through multi-scale feature fusion. Through a cross-scale feature pyramid network, detailed information at different scales is integrated to further improve reconstruction quality. This network uses bidirectional feature transfer from top to bottom and bottom to top, forming a closed-loop feedback structure to optimize the expressive power of features at each level.

[0037] Through this step, the spatial resolution and detail expression capabilities of each modal data are improved, especially the text clarity and edge sharpness in the license plate area are enhanced, providing high-quality input data for subsequent accurate recognition.

[0038] Feature extraction and deconstruction module, which extracts features from the multimodal data after super-resolution reconstruction and deconstructs the extracted features into shared information and modality-specific information; Perform feature extraction and deconstruction and reconstruction on the super-resolution reconstructed multimodal data to achieve effective decoupling and expression of inter-modal information. This includes the following sub-steps: Step 3.1, modality-specific feature extraction; According to the embodiments of this application, specialized feature extraction networks are used to extract modality-specific features for each modality. For visible light images, a modified residual network is used to extract texture, color, and shape features; for infrared images, a thermal feature extraction network is used to extract temperature distribution and thermal gradient features; and for millimeter wave data, a point cloud convolutional network is used to extract spatial structure and distance features. The extracted feature set includes feature representations for all modalities.

[0039] In some implementations, an adaptive parameter adjustment strategy can be employed for feature extraction under different environmental conditions. For example, for a feature extraction network for visible light images, the parameters of the convolutional layer can be dynamically adjusted based on the lighting conditions, increasing the weights of shallow features in low-light conditions and deep semantic features in normal lighting conditions. This adaptive adjustment can be achieved using a conditional batch normalization layer or a feature modulation layer.

[0040] Step 3.2, shared feature space mapping; The features of each modality are projected into a shared feature space through a nonlinear mapping function, establishing a correspondence between the features of different modalities. In addition, the nonlinear mapping function transforms the original features of each modality into a unified shared feature space.

[0041] These nonlinear mapping functions are actually multilayer perceptron (MLP) structures, with each modality corresponding to an independent mapping network. Each MLP consists of three fully connected layers, with 256, 128, and 64 neurons, respectively. The intermediate layers use the ReLU activation function, while the output layer does not use an activation function to maintain the linear nature of the mapping space. It should be noted that experiments have demonstrated that this structure can effectively map heterogeneous features from different modalities into a unified shared feature space.

[0042] Alternatively, to enhance the expressive power of the shared feature space, contrastive learning can be used for training. By constructing data from different modalities of the same license plate as positive pairs and data from different license plates as negative pairs, and optimizing feature representation using InfoNCE loss or triplet loss, the representations of the same license plate in different modalities become closer, while the representations of different license plates become more distinct.

[0043] Step 3.3, feature deconstruction and separation; The feature deconstruction network decomposes features in the shared space into shared information and modality-specific information. Shared information reflects the essential characteristics of the license plate (such as character shape and layout), which can be captured by all modalities. Modality-specific information reflects information unique to a specific sensor (such as material and temperature). Therefore, the features of each modality can be represented as a combination of shared and private information.

[0044] The feature deconstruction network is one of the core innovations of this application. This network leverages the concept of a variational autoencoder (VAE) but is modified to decouple multimodal features. The network consists of three main components: a shared encoder, a private encoder, and a decoupling regularizer. The shared encoder extracts common information from all modal features; the private encoder extracts information unique to each modality; and the decoupling regularizer ensures independence between shared and private information through orthogonal constraints.

[0045] In some implementations, adversarial learning strategies can be employed to further enhance feature decoupling. By introducing a discriminator network, it becomes difficult to distinguish the original modality type from the shared features, thereby ensuring that the shared features contain only modality-independent information. At the same time, private features should be accurately identified as modality-specific by the discriminator, ensuring that modality-specific information is correctly extracted.

[0046] Optionally, in order to deal with the imbalance problem of different modal information (e.g., some modal information is large while other modal information is small), an adaptive weight mechanism can be introduced to dynamically adjust the contribution weight of each modality in the shared information according to its information entropy.

[0047] The feature deconstruction network demonstrates its advantages in real-world applications in rainy and foggy conditions. For example, when a vehicle enters a parking lot during heavy rain, the image captured by the visible light camera may be blurred by raindrops, but the thermal distribution captured by the infrared camera and the spatial structure information acquired by the millimeter-wave radar remain unaffected. In this scenario, the feature deconstruction network is able to extract shared essential license plate features (such as character layout) from all three modalities while preserving the sharp details provided by the infrared and millimeter-wave images, effectively compensating for the shortcomings of the visible light image.

[0048] Step 3.4, feature reconstruction and consistency optimization; Based on the deconstructed shared and private information, the features of each modality are reconstructed, and the consistency of the feature representation is optimized using the reconstruction loss. The reconstruction process recombines the shared and private information through the reconstruction function to generate the reconstructed feature representation.

[0049] ; in, is the reconstruction loss function; is a modal set; is the modal index; For the The original features of each mode; For the Reconstruction features of each modality; represents the L2 norm; is the consistency constraint weight parameter; For the shared information extracted from different modalities; For the shared information extracted from multiple modalities.

[0050] In addition, the reconstruction loss function contains two parts: the first term is the reconstruction error, which measures the difference between the reconstructed features and the original features; the second term is the shared information consistency constraint, which ensures that the shared information extracted by different modalities remains consistent. By optimizing this loss function, the network can learn better feature reconstruction capabilities.

[0051] Optionally, additional robustness constraints can be added under extreme environmental conditions. For example, by artificially adding noise, blur, or occlusion to the training data, and then requiring the network to recover the original features using multi-modal information, the robustness of the system under harsh environmental conditions can be improved.

[0052] In addition, experiments show that this feature reconstruction and consistency optimization method can effectively improve the complementarity and consistency of information between modalities under different environmental conditions.

[0053] Through this step, effective decoupling and expression of multi-modal features are achieved, and essential license plate features shared by each modality and modality-specific supplementary information are extracted, laying a foundation for subsequent modality fusion.

[0054] The environmental perception and evaluation module performs reliability evaluation on the multi-modal data after feature reconstruction based on the current environmental conditions, and generates a dynamic weight distribution strategy; According to the embodiments of the present application, a modality reliability evaluation mechanism for environmental perception is constructed, and the reliability of each modality data is dynamically evaluated according to the current scene conditions, which includes the following sub-steps: Step 4.1, environmental condition detection; An environmental condition evaluation module is constructed to detect the environmental parameters of the current scene in real time, including weather conditions (sunny, rainy, snowy, foggy, etc.), lighting conditions (daytime, dusk, nighttime, strong backlight, etc.), and other environmental factors.

[0055] The environmental condition evaluation module adopts a two-stage structure. The first stage uses a lightweight convolutional neural network to extract environmental features, which consists of 4 convolutional blocks, each containing a convolutional layer, a batch normalization layer, and a ReLU activation function, finally generating a 128-dimensional environmental feature vector. The second stage uses a multi-task classification head to simultaneously predict the weather type (5 categories: sunny, rainy, snowy, foggy, nighttime) and the lighting intensity (3 categories: normal, weak light, strong light / backlight).

[0056] Optionally, in some embodiments, environmental condition detection can employ a temporal information enhancement strategy. By maintaining a sequence of environmental features within a short time window and analyzing environmental change trends using a recurrent neural network or a temporal convolutional network, this improves adaptability to rapidly changing environments (such as a vehicle entering an indoor parking lot from an outdoor location). It should be understood that this temporal enhancement method can capture dynamic information about environmental changes by analyzing environmental features over multiple consecutive frames, thereby improving the stability of environmental assessments.

[0057] This module comprehensively judges the current environment type by analyzing the color histogram features (hue distribution, saturation, brightness mean and variance) of visible light images, texture features (contrast and homogeneity extracted through gray-level co-occurrence matrix), and temperature distribution (average temperature, temperature range, number of hot spots) of infrared images.

[0058] In practical applications, it's important to note that this module can accurately distinguish various complex environmental conditions. For example, in foggy weather, visible light images typically exhibit low contrast, high brightness, and blurred edges, along with a relatively uniform temperature distribution. On rainy days, however, images exhibit dynamic texture changes and irregular light spots. By analyzing these combined features, the Environmental Condition Assessment Module can accurately identify the current environmental type, providing a basis for subsequent modal reliability assessment.

[0059] Step 4.2, modal mass parameter extraction; For each modality, key parameters reflecting its quality are extracted. For visible light images, the signal-to-noise ratio (SNR), sharpness, and visibility are calculated; for infrared images, the temperature contrast and temperature gradient intensity are calculated; and for millimeter wave data, the signal strength and point cloud density are calculated.

[0060] These quality parameters are calculated using a no-reference quality assessment method that does not rely on standard reference images and is suitable for real-time online evaluation. For example, the clarity of visible light images is quantified by the variance of the Laplacian operator response; visibility is calculated using the image's contrast and brightness features; and the temperature contrast of infrared images is assessed using the entropy of the heatmap histogram.

[0061] In some implementations, quality parameter extraction can also be performed separately for the license plate area and non-license plate areas. Using the rough license plate area location information provided by the license plate detection module, the system can evaluate the quality parameters of the license plate area and background area separately, more accurately determining the feasibility of license plate recognition. Furthermore, the overall quality score can be calculated using a weighted average of the regions, with the license plate area typically given a higher weight (approximately 0.7-0.8).

[0062] Step 4.3, modal reliability score calculation; Based on environmental conditions and quality parameters, the reliability score of each modality under the current environment is calculated. According to the embodiments of the present application, this calculation process considers factors such as signal-to-noise ratio, visibility, and current weather conditions, and obtains the final reliability score through weighted method.

[0063] In specific implementation, the function adopts a combination of rule-based fuzzy logic system and learning neural network. The fuzzy logic part defines a series of rules according to expert knowledge, such as "if the weather is foggy and the visibility is low, then the reliability of visible light camera is low"; the learning part learns the complex mapping relationship between each factor and reliability through a multi-layer perception network according to a large amount of labeled data.

[0064] Optionally, the reliability score calculation can also consider historical performance. The system can maintain an environmental condition-performance mapping database to record the historical recognition performance of each modality under different environments, and use these historical data as a reference basis for scoring. This adaptive learning strategy can continuously optimize the accuracy of reliability evaluation as the system running time increases.

[0065] In actual parking lot applications, this reliability score mechanism shows excellent environmental adaptability. For example, in foggy weather conditions, the system can automatically reduce the weight of the visible light modality (reliability score about 0.3) and increase the weight of the millimeter wave radar (reliability score about 0.6), effectively ensuring the stability of the recognition system in low-visibility environments.

[0066] Step 4.4, dynamic weight distribution strategy generation; According to the reliability score of each modality, a dynamic weight distribution strategy for subsequent feature fusion is generated. The weight calculation adopts softmax normalization method to convert the reliability score of each modality into a weight distribution with a total sum of 1.

[0067] ; wherein, is the weight of the th modality; is an exponential function; is a temperature parameter for controlling the smoothness of the weight distribution; and are the reliability scores of the th and th modality, respectively; is the set of modalities; denotes the summation operation.

[0068] In some implementations, to improve system robustness, upper and lower bounds can be set for modal weights to prevent any modal weight from being too high or too low. Furthermore, by limiting weights to a certain range (typically with a lower bound of 0.1 and an upper bound of 0.8), the system can consistently utilize information from multiple modalities, maintaining basic fusion capabilities even when some modalities perform poorly. These weights need to be normalized again to ensure they sum to 1.

[0069] Furthermore, this dynamic weighting strategy adjusts the importance of each modality in real time based on environmental changes. For example, during clear daylight conditions, visible light cameras receive the highest weight (approximately 0.6); at night or in rainy weather, infrared cameras receive a higher weight (approximately 0.5); and in extreme conditions such as heavy fog or snow, millimeter-wave radars receive the dominant weight (up to 0.7).

[0070] Through this step, the system can intelligently evaluate the reliability of each modal data according to the real-time environmental conditions and generate a corresponding weight distribution strategy to provide a basis for subsequent feature fusion, ensuring that the most reliable information source can be selected under different environmental conditions.

[0071] The feature fusion and recognition module performs complementary enhancement fusion on the deconstructed multimodal features based on the dynamic weight allocation strategy and performs license plate recognition; Based on the achievements of the data acquisition and preprocessing module, the super-resolution reconstruction module, the feature extraction and deconstruction module, and the environmental perception and assessment module, the complementary enhancement fusion of multimodal features and the final license plate recognition are achieved. Specifically, the following sub-steps are included: Step 5.1, complementary information enhancement fusion; Based on the weight allocation strategy derived from modal reliability assessment and the feature representations derived from deconstruction and reconstruction, embodiments of this application implement complementary information-enhanced fusion of multimodal features. Unlike traditional simple weighted fusion, this method focuses on extracting and enhancing complementary information between different modalities, rather than simply superimposing them.

[0072] Another innovative feature of this application is the complementary information enhancement module. This module adopts a dual-stream attention network structure, consisting of a channel attention branch and a spatial attention branch. The channel attention branch learns the importance relationship between different channels (corresponding to different features); the spatial attention branch learns the importance of different spatial positions. The outputs of the two branches are fused through a gating mechanism. Specifically, to extract complementary information between modal pairs, a feature difference map is first calculated, then channel attention and spatial attention are applied separately for processing, and finally a complementary feature representation is obtained through weighted combination.

[0073] Optionally, in some implementations, complementary information enhancement can incorporate a multi-level cross-attention mechanism, enabling information exchange between different modalities at multiple feature levels. This multi-level cross-attention can be applied at different stages of the feature extraction process (e.g., shallow, mid, and deep layers) to capture complementary information at varying granularities. It should be understood that this multi-level fusion approach fully exploits the complementarity of features at different levels, forming a more comprehensive fused representation.

[0074] In practical applications, the Complementary Information Enhancement module effectively leverages the strengths of different modalities. For example, at a parking lot entrance at dusk, a visible light camera may overexpose the license plate due to strong backlight, while an infrared camera may lack information due to the subtle temperature difference. However, millimeter-wave radar is unaffected by illumination. In this case, the Complementary Information Enhancement module automatically detects the complementarity between the license plate position information in visible light and the structural information in millimeter-wave, as well as the complementarity between partially visible character features in infrared and color features in visible light, thereby generating a more accurate fused representation.

[0075] Step 5.2: Conditional style transfer adaptation; Conditional style transfer is applied to the fused features to map the license plate features under complex environments to the feature distribution under standard conditions, thereby improving the robustness of subsequent recognition. It should be noted that conditional style transfer achieves environmental adaptation by adjusting the style distribution of features while preserving content information.

[0076] The conditional style transfer network uses an architecture based on AdaIN (Adaptive Instance Normalization) and consists of three parts: a content encoder, a style encoder, and a decoder. The content encoder extracts structural information from the fused features; the style encoder extracts style information from the target conditional sample; and the decoder combines the content and style information to generate style-transferred features. The network works by first extracting content features and target style features, then adjusting the statistical properties of the content features through AdaIN operations, and finally generating adapted features through the decoder.

[0077] Optionally, in some implementations, conditional style transfer can utilize a multi-target condition library to establish specialized target style templates for different types of license plates (e.g., standard blue plates, new energy vehicle plates, special vehicle plates, etc.). The system first performs a rough classification of the license plates and then selects the corresponding target condition for style transfer, further improving the recognition accuracy of specific license plate types.

[0078] In practice, the system learns an "ideal" feature distribution from a dataset of license plate images under standard lighting conditions, using this as the target condition. When the system operates in complex environments, the conditional style transfer network maps the detected license plate features to this ideal condition, significantly reducing the impact of environmental changes on recognition performance.

[0079] Step 5.3, license plate detection and positioning; Based on the adapted fusion features, a region proposal network is applied to detect and locate the license plate region, generating candidate regions containing the license plate. This network enhances the detection capability of license plates of different sizes through a multi-scale feature pyramid and outputs the license plate bounding box coordinates and confidence score.

[0080] In some implementations, license plate detection can employ a cascaded detection strategy, first performing coarse localization and then fine localization. The coarse localization phase uses a lightweight network to quickly screen possible license plate regions, while the fine localization phase uses a more complex network to precisely locate and correct the angle of the candidate regions, improving detection accuracy and speed.

[0081] Step 5.4, character-level recognition and post-processing; Based on the located license plate area, a deep character recognition network is applied according to the embodiments of this application to identify the license plate characters. This network uses an attention-enhanced sequence recognition model to recognize the license plate string as a whole, avoiding the errors that may be caused by traditional character segmentation. In addition, the decoding function processes the feature representation of the license plate area through the attention mechanism and outputs the final license plate string.

[0082] Alternatively, for character recognition, a dual-path decoding strategy combining CTC and attention can be employed, leveraging the strengths of both CTC (Connectionized Temporal Classification) and the attention mechanism to improve recognition accuracy by integrating the predictions of both methods. Furthermore, the system can perform a weighted combination of the outputs of the CTC and attention decoders, typically with a balanced weight (0.5 for each).

[0083] Step 5.5, verification and optimization of recognition results; Verify and optimize the recognition results, apply prior knowledge of license plate formats (such as license plate format rules in various regions) to check the legitimacy, and combine timing information (such as multi-frame recognition results of the same vehicle in a short period of time) to optimize consistency and improve the final recognition accuracy.

[0084] In some implementations, a memory network-based temporal consistency optimization can be introduced. The system maintains a short-term memory bank that stores recently passed vehicles and their recognition results. When consecutive frames that may be the same vehicle are detected, the historical information in the memory bank is used for voting or probabilistic fusion to reduce the impact of single-frame recognition errors.

[0085] Through this step, multimodal data undergoes complementary enhancement fusion and conditional style transfer, ultimately achieving highly reliable license plate recognition in all weather and all environments.

[0086] A computer storage medium includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the above-mentioned deep learning-based intelligent parking lot license plate recognition system.

[0087] Here, the present invention provides an implementation example: This implementation has been deployed and tested in the underground parking lot of a large commercial complex. The parking lot has approximately 1,200 parking spaces, spread across three underground floors, with four main entrances and exits. These entrances and exits are located in different parts of the building and face different environmental challenges: the north entrance is often affected by strong backlight; the east entrance is close to the main road of the shopping mall, with heavy and fast traffic; the west entrance is in a semi-open area and is significantly affected by inclement weather; and the south entrance is a ramp with license plates tilted at multiple angles. The following are specific application cases of this system: The underground parking system deploys the following hardware facilities: one visible light camera (resolution 2560×1440 pixels), one infrared camera (resolution 1280×720 pixels), and one millimeter-wave radar (resolution 64×64 point cloud) are installed at each entrance and exit; four edge computing servers are located in the processing box at each entrance, configured with an 8-core CPU, 16GB of memory, and a GPU accelerator card; one central management server is located in the parking lot monitoring room, responsible for managing the vehicle information database and system status monitoring.

[0088] The system workflow is as follows: When a vehicle enters the parking lot entrance area, the sensor array is triggered and begins collecting multimodal data; The edge server processes data in real time and performs multimodal fusion license plate recognition; The identification results are transmitted to the central management system for vehicle record and charging management.

[0089] In a typical rainy evening scenario, the system execution process is as follows: Multimodal data acquisition and preprocessing: When a white car entered the west entrance, the system triggered multimodal sensors to simultaneously collect data. The visible light image exhibited the blur and speckle characteristics typical of rainy days, with low contrast in the license plate area. The infrared image clearly captured the heat distribution in the engine compartment and exhaust system, but the thermal signature in the license plate area was less distinct. The millimeter-wave radar data provided the vehicle's precise outline and distance information.

[0090] The system first performs spatiotemporal alignment of the data from each modality: using a sensor calibration matrix, the data from different sensors is mapped into a unified coordinate system, with a time synchronization error of less than 10ms. To address harsh rainy conditions, the system automatically adjusts data preprocessing parameters: an adaptive rain removal algorithm and local contrast enhancement are applied to visible light images; thermal noise suppression and edge enhancement are applied to infrared images; and raindrop reflection filtering is applied to millimeter wave data.

[0091] Modality-specific super-resolution reconstruction: The system performs modality-specific super-resolution processing on low-quality raw data. For visible light images, the license plate text-guided attention module automatically locates the characters within the license plate area. Although some characters are blurred by raindrops, the attention mechanism enhances the clarity of these areas. For infrared images, even if the license plate's temperature characteristics are not obvious, the system still enhances the thermal gradient difference between the license plate edge and the surrounding environment through a temperature difference preservation algorithm. For millimeter wave data, density-enhanced upsampling improves point cloud resolution, refining the outline of the license plate area.

[0092] Multimodal feature extraction and deconstruction and reconstruction: The system extracts specific features from each modality after super-reconstruction: the visible light image extracts the basic shape of the license plate and some visible characters; the infrared image extracts the edge contour of the license plate; and the millimeter wave image extracts the position and shape characteristics of the license plate. Using a feature deconstruction network, the system decomposes these heterogeneous features into shared information (such as the license plate's position, size, and orientation) and modality-specific information (such as the color of visible light and the temperature distribution of infrared). In this rainy day example, some characters in the visible light image are obscured by raindrops, making the information incomplete. However, the infrared and millimeter wave images provide complementary information. The feature reconstruction process successfully extracts a complete feature representation of the license plate from all three modalities.

[0093] Modal reliability assessment of environmental perception: The system's environmental condition assessment module analyzes the color histogram and texture features of visible light images to identify the current combined environmental conditions of "rain and dusk." Based on this environmental identification result and the data quality parameters of each modality, the system calculates a reliability score for each modality: 0.35 for visible light, 0.45 for infrared, and 0.65 for millimeter wave. After applying softmax normalization, the fusion weights for each modality are: 0.25 for visible light, 0.32 for infrared, and 0.43 for millimeter wave. This reflects the relatively higher reliability of millimeter wave radar data in rainy conditions.

[0094] Complementary enhanced feature fusion and recognition: The system fuses features from various modalities based on reliability weights and the principle of complementary information. The complementary information enhancement module detects that certain strokes of the license plate character "苏" are missing in the visible light image due to raindrops, while the edges of this area are intact in the infrared image. The system automatically extracts and enhances this complementary information. The fused features are mapped to the feature distribution under standard lighting conditions through conditional style transfer. Finally, the character recognition module successfully recognizes the complete license plate number "苏A12345" based on the fused features, despite the difficulty in recognizing some characters in the original visible light image due to raindrops.

[0095] This implementation method has achieved technical results during the actual deployment and testing process in the above-mentioned parking lot, especially in terms of adaptability to complex environments and recognition accuracy: All-weather environmental adaptability: The system's license plate recognition success rate under different environmental conditions is as follows: Sunny daytime: single visible light solution 98.7%, this solution 99.3%; Clear night: single visible light scheme 87.5%, this scheme 97.8%; During rainy daytime: 82.3% for single visible light scheme and 96.1% for this scheme; Rainy nights: 75.8% for the single visible light solution and 95.2% for this solution; During foggy daytime: 68.4% for the single visible light scheme and 92.7% for this scheme; Foggy nights: Single visible light solution 59.2%, this solution 89.5%.

[0096] The data clearly demonstrates that this solution maintains a high recognition success rate in a variety of harsh environments, particularly in extreme conditions (such as foggy nights), where it achieves a 30.3 percentage point improvement over traditional single-visible light solutions, demonstrating its superiority. The system achieves truly stable recognition capabilities in all weather conditions and environments.

[0097] System response time and real-time performance: The system's end-to-end processing time (from triggering to outputting recognition results) under different environmental conditions is as follows: Standard environment (sunny daytime): single visible light solution 78ms, this solution 92ms; Complex environment (rainy night): Single visible light solution 135ms, this solution 108ms; Extreme environment (foggy night): Single visible light solution 212ms, this solution 125ms.

[0098] Data shows that while this solution's processing time is slightly longer than that of a single visible light approach under standard conditions (primarily due to the additional computational overhead of multimodal fusion), it is significantly faster in complex and extreme environments. This is because traditional methods require multiple retries and complex image enhancement in harsh environments, while this solution, through multimodal complementarity and environmental adaptation, can more directly achieve reliable results. In all scenarios, this solution's response time remains under 130ms, fully meeting the requirements of real-time recognition.

[0099] In summary, the technical advantages demonstrated by this solution in the complex and changeable actual parking environment, especially the improvement in environmental adaptability and recognition accuracy, provide strong support for the all-weather reliable operation of the parking management system.

[0100] like Figures 2 to 6 As shown, there are respectively a line graph comparing the license plate recognition accuracy of the solution of the present invention and the traditional single visible light solution under different environmental conditions; a bar graph comparing the system response time of the solution of the present invention and the traditional single visible light solution under different environmental conditions; a radar chart of the reliability scores of different modal sensors under various environmental conditions (all scores have been normalized to the range of 0-1); a scatter plot of feature quality assessment before and after feature deconstruction and reconstruction (the horizontal axis is feature separation, the vertical axis is feature reconstruction error, all values ​​have been normalized); and a pie chart of the computing resource allocation of the multimodal fusion license plate recognition system.

[0101] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. The intelligent parking lot license plate recognition system based on deep learning is characterized by: include: The data acquisition and preprocessing module acquires multimodal data including visible light images, infrared images, and millimeter-wave radar data, and preprocesses the multimodal data; The super-resolution reconstruction module performs modality-specific super-resolution reconstruction on the pre-processed multimodal data; Feature extraction and deconstruction module, which extracts features from the multimodal data after super-resolution reconstruction and deconstructs the extracted features into shared information and modality-specific information; The environmental perception and assessment module performs reliability assessment on the multimodal data after feature deconstruction based on the current environmental conditions and generates a dynamic weight allocation strategy; The feature fusion and recognition module performs complementary enhancement fusion on the deconstructed multimodal features based on the dynamic weight allocation strategy, and performs license plate recognition.

2. The deep learning-based intelligent parking lot license plate recognition system according to claim 1 is characterized in that: The step of preprocessing the multimodal data includes: Perform spatiotemporal alignment of multimodal data collected by different sensors; Perform preliminary noise filtering on multimodal data; Perform modality-specific enhancement processing based on the physical characteristics of multimodal data.

3. The deep learning-based intelligent parking lot license plate recognition system according to claim 1 is characterized in that: The modality-specific super-resolution reconstruction steps include: Applying a license plate text-guided super-resolution network to visible light images. The license plate text-guided super-resolution network includes a text attention module for adaptively enhancing the license plate character area. Applying a temperature difference preserving super-resolution network to infrared images, the temperature difference preserving super-resolution network maintains the temperature difference of the object boundary through the temperature gradient preserving branch; Apply a density-enhanced upsampling algorithm to millimeter-wave radar data to improve spatial resolution.

4. The deep learning-based intelligent parking lot license plate recognition system according to claim 1 is characterized in that: The step of performing reliability assessment on the multimodal data after feature deconstruction based on the current environmental conditions includes: Build an environmental condition assessment module to detect the environmental parameters of the current scene, including weather conditions and lighting conditions; For multimodal data, extract key parameters that reflect the quality of each modal data; Based on environmental conditions and quality parameters, calculate the reliability score of each modal data in the current environment; According to the reliability scores of multimodal data, a dynamic weight allocation strategy for feature fusion is generated through softmax normalization.

5. The deep learning-based intelligent parking lot license plate recognition system according to claim 1 is characterized in that: The complementary enhancement fusion step comprises: Apply a two-stream attention network to extract and enhance complementary information between different modalities; The conditional style transfer technology is applied to the fused features to map the license plate features in complex environments to the feature distribution under standard conditions.

6. The deep learning-based intelligent parking lot license plate recognition system according to claim 5 is characterized in that: The conditional style transfer technology adopts an architecture based on adaptive instance normalization, which includes a content encoder, a style encoder and a decoder.

7. The deep learning-based intelligent parking lot license plate recognition system according to claim 1 is characterized in that: The steps of license plate recognition include: Detect and locate the license plate area using a region proposal network based on fused features; For the located license plate area, an attention-enhanced sequence recognition model is applied for character recognition; Verify and optimize the recognition results, apply prior knowledge of license plate formats to check legitimacy, and combine timing information for consistency optimization.

8. The deep learning-based intelligent parking lot license plate recognition system according to claim 7 is characterized in that: The character recognition adopts a hybrid CTCAttention dual-path decoding strategy, which simultaneously utilizes connection temporal classification and attention mechanism to improve recognition accuracy.

9. The deep learning-based intelligent parking lot license plate recognition system according to claim 4 is characterized in that: The environmental condition assessment module adopts a two-stage structure: The first stage uses convolutional neural networks to extract environmental features; The second stage uses a multi-task classification head to predict weather type and light intensity.

10. A computer storage medium, characterized in that It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the deep learning-based intelligent parking lot license plate recognition system according to any one of claims 1 to 9.