Method and system for detecting container number of gate general truck based on deep learning
Through the gate universal truck box number detection method based on deep learning, high-precision box number recognition is achieved in complex environments using the detection network and self-attention mechanism, the problem of low recognition accuracy in complex environments in the existing technology is solved, and efficient and accurate box number detection and recognition is achieved.
Patent Information
- Application Number
- CN202510313277.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-17
AI Technical Summary
The existing truck box number identification technology has low recognition accuracy in complex environments (different lighting, angles, occlusions, etc.), and the traditional methods are inefficient, which cannot meet the efficient and accurate needs of modern logistics and intelligent transportation systems.
The general truck box number detection method based on deep learning is adopted to collect images through the gate camera, and the detection network is used to realize the precise positioning of the box number area. Combined with the self-attention mechanism and multi-scale feature fusion, improve the robustness of character recognition, and eliminate fuzzy frames through real-time calculation and inference through video streams, and optimize the calculation diagram in the inference process.
It significantly improves the accuracy and stability of box number detection and identification, adapts to complex environments, meets industrial-grade application needs, reduces computing redundancy, and improves inference speed and system real-timeness.
Smart Images

Figure CN120164202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a method and system for detecting the general truck box number at the gate based on deep learning. Background Art
[0002] With the rapid development of the logistics industry, the automatic recognition of truck box numbers has become one of the key technologies to improve logistics efficiency and intelligent transportation management. Traditional truck box number recognition methods mainly rely on manual entry or traditional image processing algorithms, and these methods have obvious deficiencies in terms of processing efficiency, adaptability, and environmental robustness. Modern logistics and intelligent transportation systems have an urgent need for efficient and accurate box number recognition technology, especially in the gate passage management, where the improvement of automation can significantly reduce manual intervention and improve the logistics operation efficiency.
[0003] The existing truck box number recognition technologies mainly include the following: 1. Manual entry, where the truck box number is observed manually and entered manually. This method is inefficient, error-prone, and cannot meet the large-scale logistics requirements; 2. Traditional image processing algorithms, which are based on image processing technologies such as edge detection and template matching for box number recognition. These methods may be effective under ideal conditions, but perform poorly in complex environments (such as different lighting, angles, occlusions, etc.), and even when simply combined with deep learning algorithms, there are situations where the recognition accuracy is low when facing different types of trucks and box numbers at different angles.
[0004] The patent application document CN110852324A discloses a method for detecting container box numbers based on a deep neural network, which includes the following steps: First, obtain the RGB image of the rear side of the container, and input the container picture containing the box number into the trained character segmentation neural network model to obtain a picture set of box number character segmentation. Then, perform perspective transformation and binarization processing on the obtained picture set, and then input it into the trained character recognition deep neural network to obtain the text information in each text box. Finally, screen and combine the obtained text information to obtain the accurate container box number. However, this patent cannot completely solve the existing technical problems and cannot meet the requirements of the present invention. Summary of the Invention
[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method and system for detecting the general truck box number at the gate based on deep learning.
[0006] The method for detecting the general truck box number at the gate based on deep learning provided by the present invention includes:
[0007] Step 1: Collect truck images under different lighting conditions, angles, and weather through the gate camera, label the position of the box number area and the corresponding character tags, and use corresponding simulation means for enhancement according to different weather conditions;
[0008] Step 2: Use a detection network based on deep learning to achieve precise positioning of the box number area;
[0009] Step 3: After detecting the box number area, perform affine correction, noise suppression, and contrast enhancement on it;
[0010] Step 4: Use a recognition network based on deep learning to recognize the characters in the box number area, improve the recognition ability of long strings through the self-attention mechanism, introduce multi-scale feature fusion in the feature extraction stage to enhance the robustness to characters of different sizes and angles, and combine the self-supervised pre-training strategy to use unlabeled data to enhance the generalization ability of the model;
[0011] Step 5: After completing the detection and recognition of the box number, use the video stream for real-time calculation and inference, judge the moving objects online, screen out and remove the blurred frames, adopt the pipeline parallel processing method, and subsequently optimize the computational graph in the inference process, and use operator fusion and tensor optimization to reduce computational redundancy.
[0012] Preferably, the step 2 includes:
[0013] Input the preprocessed image into a pre-trained backbone network, which extracts multi-scale feature maps through multiple convolutional operations, and then further processes the feature maps to output two key prediction maps: one is the probability map P(x,y), which represents the probability that each pixel in the image belongs to the box number area; the other is the threshold map T(x,y), which predicts an adaptive threshold for each pixel for subsequent discrimination between text and background; introduce a differentiable binarization module to convert the prediction into a binarized result, and its core formula is:
[0014] B(x,y) = σ(α(P(x,y) - T(x,y)))
[0015] where σ represents the Sigmoid function; α is a smoothing factor that controls the sharpness of the binarization process; when P(x,y) > T(x,y), the value of B(x,y) output by this function is close to 1, indicating that this pixel belongs to the box number area; conversely, a value close to 0 indicates the background area;
[0016] To improve the detection accuracy, introduce a feature enhancement mechanism, fuse multi-level features through an adaptive feature pyramid network, extract a series of feature maps {F1, F2... F N} from different levels of the backbone network, and introduce AFPN to adaptively fuse the features of each layer;
[0017] First, all feature maps are adjusted to the same spatial size through upsampling operations, and then weighted fusion is performed through the learned adaptive weights ω i The formula is as follows:
[0018]
[0019] where U represents the upsampling operation, N represents the number of network layers, and F i is the feature map of the i-th layer, and the weight ω i is calculated through the attention mechanism and normalized using Softmax:
[0020]
[0021] where a i and a j respectively represent the attention scores obtained after the features of the i-th and j-th layers pass through a small convolutional or fully connected layer;
[0022] To better capture the edge information of the box number area, an edge detection operator is used on the original input image to calculate the edge map:
[0023] E lap = Lap(F(x, y))
[0024] where Lap represents calculating the edge features using the Laplacian operator, and then F AFPN and E lap are combined to obtain the final enhanced feature map:
[0025] F fuse = F AFPN + μE lap
[0026] where μ is a hyperparameter for adjusting the weight of the edge features. During the fusion process, the edge features are weighted into the multi-scale features at a preset ratio, making the model more sensitive to edge information;
[0027] To further optimize the detection process, a loss function L = L cls + λ1L reg + λ2L iou is introduced, where L cls represents the classification loss, L reg is the bounding box regression loss, and L iou is the IoU loss, which is used to improve the localization accuracy; λ1 and λ2 respectively represent the hyperparameters for adjusting the weights of each part of the loss.
[0028] Preferably, step 3 includes:
[0029] The adaptive morphological processing method is adopted to make the boundaries of box number characters clearer. At the same time, the box number area is automatically segmented by the method based on morphological clustering, and the interfering characters are removed;
[0030] Aiming at the perspective distortion problem existing in the box number area, the homography transformation is adopted for correction, and the transformation formula is as follows:
[0031]
[0032] Among them, H is the homography matrix, which is calculated by four-point transformation, so as to correct the inclined box number area into a standard perspective; and respectively represent the position of the original image and the position after transformation.
[0033] Preferably, the step 4 includes:
[0034] The SVTR model is adopted as the basic recognition network. This model first extracts multi-scale features from the preprocessed image through the convolutional or Patch Embedding layer, that is, a set of feature maps {F1, F2... F N} are obtained at different levels. The feature enhancement pyramid is used for feature fusion. Next, the enhanced feature map is converted into a feature sequence suitable for the Transformer module to process and input into the attention-guided sequence modeling module. Let this feature sequence be X = {x1, x2... x n}, and the self-attention mechanism is used for global modeling. The calculation formula is:
[0035]
[0036] Among them, Q = XW Q , K = XW k and V = XW V are the linear transformations of the query, key, and value respectively. d k is the dimension of the key. W Q , W k , W V are the weight matrices used to calculate the query vector, key vector, and value vector respectively. By performing global context modeling, the ability to capture the dependence relationship between long string characters is improved;
[0037] In order to further improve the generalization performance of the model, the self-supervised pre-training strategy is combined. At this stage, the feature representation is optimized by pre-training on unlabeled data. The contrast loss function is expressed as:
[0038]
[0039] Among them, z i , z jIt is the feature representation of different enhanced versions of the same character or region. sim represents the cosine similarity, τ represents the hyperparameter, K represents the number of negative samples, and k represents the index variable;
[0040] The loss function in the recognition process adopts the CTC loss function, which is defined as follows:
[0041]
[0042] Among them, p(|π t X input ) represents the probability of the target character sequence π input under the condition of the given input sequence X t . t represents a certain sequence moment, and Β represents the function that maps the path π to the final character sequence y.
[0043] Preferably, the step 5 includes:
[0044] Using the optical flow method to judge the motion situation of the target object in the picture, adopting the Laplace transform method to calculate the image sharpness score. When a blurred frame is detected, combining the motion estimation result to judge whether it is caused by fast motion, and removing the low-quality frames caused by motion blur;
[0045] Decoupling the detection, preprocessing and recognition parts, and adopting the pipeline parallel processing method; splitting the whole processing flow into independent stages, including frame acquisition, preprocessing, target detection, feature extraction and final recognition. Each stage runs independently on different computing threads or computing units, and transfers data through the shared memory mechanism; adopting the task scheduling strategy to hand over the high-computation tasks to the GPU for processing, while leaving the lightweight tasks to the CPU;
[0046] In the inference stage, using operator fusion to merge multiple consecutive operators in the computation graph into one operator, and adopting tensor optimization technology for data storage and caching.
[0047] According to the gantry general freight car box number detection system based on deep learning provided by the present invention, it includes:
[0048] Module M1: Collecting freight car images under different lighting conditions, different angles and different weathers through the gantry camera, and annotating the box number area position and corresponding character labels, and adopting corresponding simulation means for enhancement according to different weathers;
[0049] Module M2: Implementing precise positioning of the box number area by using a detection network based on deep learning;
[0050] Module M3: After detecting the box number area, performing affine correction, noise suppression and contrast enhancement on it;
[0051] Module M4: Use a recognition network based on deep learning to perform character recognition on the box number area. Improve the recognition ability of long strings through the self-attention mechanism. Introduce multi-scale feature fusion in the feature extraction stage to enhance the robustness to characters of different sizes and angles. Combine the self-supervised pre-training strategy to utilize unlabeled data to enhance the generalization ability of the model;
[0052] Module M5: After completing the box number detection and recognition, use the video stream for real-time calculation and inference, judge the moving objects online, screen out and eliminate the blurred frames, adopt the pipeline parallel processing method, and subsequently reduce the calculation redundancy by optimizing the computational graph in the inference process and using operator fusion and tensor optimization.
[0053] Preferably, the module M2 includes:
[0054] Input the preprocessed image into the pre-trained backbone network. This network extracts multi-scale feature maps through multi-layer convolution operations. Then, further process the feature maps and output two key prediction maps: one is the probability map P(x,y), which represents the probability that each pixel in the image belongs to the box number area; the other is the threshold map T(x,y), which predicts an adaptive threshold for each pixel and is used to distinguish text from the background in the subsequent process. Introduce the differentiable binarization module to convert the prediction into a binarized result, and its core formula is:
[0055] B(x,y) = σ(α(P(x,y) - T(x,y)))
[0056] where σ represents the Sigmoid function; α is the smoothing factor that controls the sharpness of the binarization process; when P(x,y) > T(x,y), the value of B(x,y) output by this function is close to 1, indicating that this pixel belongs to the box number area; conversely, a value close to 0 indicates the background area;
[0057] To improve the detection accuracy, introduce the feature enhancement mechanism, fuse multi-level features through the adaptive feature pyramid network, extract a series of feature maps {F1, F2... F N} from different levels of the backbone network, and introduce AFPN to adaptively fuse the features of each layer;
[0058] First, adjust all the feature maps to the same spatial size through the upsampling operation, and then perform weighted fusion through the learned adaptive weights ω i The formula is expressed as:
[0059]
[0060] where U represents the upsampling operation, N represents the number of network layers, F i is the feature map of the i-th layer, and the weight ω iCalculated through the attention mechanism, using Softmax normalization:
[0061]
[0062] Among them, a i and a j respectively represent the attention scores obtained after the features of the i-th and j-th layers pass through a small convolutional or fully connected layer;
[0063] In order to better capture the edge information of the box number area, an edge detection operator is used on the original input image to calculate the edge map:
[0064] E lap = Lap(F(x,y))
[0065] Among them, Lap represents the Laplacian operator to calculate the edge features, and then F AFPN and E lap are combined to obtain the final enhanced feature map:
[0066] F fuse = F AFPN + μE lap
[0067] Among them, μ is a hyperparameter for adjusting the weight of the edge features. During the fusion process, the edge features are weighted into the multi-scale features at a preset ratio, so that the model is more sensitive to the edge information;
[0068] To further optimize the detection process, the loss function L = L cls + λ1L reg + λ2L iou is introduced, where L cls represents the classification loss, L reg is the bounding box regression loss, and L iou is the IoU loss, which is used to improve the localization accuracy; λ1 and λ2 respectively represent the hyperparameters for adjusting the weights of each part of the loss.
[0069] Preferably, the module M3 includes:
[0070] Adopt an adaptive morphological processing method to make the boundaries of the box number characters clearer, and at the same time use a morphological clustering-based method to automatically segment the box number area and remove interfering characters;
[0071] For the perspective distortion problem existing in the box number area, a homography transformation is used for correction, and the transformation formula is as follows:
[0072]
[0073] Among them, H is the homography matrix, which is calculated through four-point transformation to correct the tilted box number area into a standard perspective; and respectively represent the position of the original image and the position after transformation.
[0074] Preferably, the module M4 includes:
[0075] The SVTR model is used as the basic recognition network. This model first extracts multi-scale features from the preprocessed image through a convolutional or Patch Embedding layer, that is, a set of feature maps {F1, F2... F N} are obtained at different levels. Feature enhancement pyramids are used for feature fusion. Next, the enhanced feature maps are converted into feature sequences suitable for processing by the Transformer module and input into the attention-guided sequence modeling module. Let this feature sequence be X = {x1, x2... x n}. The self-attention mechanism is used for global modeling, and the calculation formula is:
[0076]
[0077] where Q = XW Q , K = XW k and V = XW V are the linear transformations of the query, key, and value respectively, d k is the dimension of the key, and W Q , W k , W V are the weight matrices used to calculate the query vector, key vector, and value vector respectively. By performing global context modeling, the ability to capture the dependency relationship between characters in long strings is improved;
[0078] In order to further improve the generalization performance of the model, a self-supervised pre-training strategy is combined. At this stage, through pre-training on unlabeled data, the feature representation is optimized, and the contrast loss function is expressed as:
[0079]
[0080] where z i , z j are the feature representations of different enhanced versions of the same character or region, sim represents the cosine similarity, τ represents the hyperparameter, K represents the number of negative samples, and k represents the index variable;
[0081] The loss function in the recognition process uses the CTC loss function, which is defined as follows:
[0082]
[0083] where p(|πt X input ) represents the probability of the target character sequence π input given the input sequence X t ; t represents a certain sequence moment, and Β represents the function that maps the path π to the final character sequence y.
[0084] Preferably, the module M5 includes:
[0085] Use the optical flow method to judge the motion of the target object in the picture, adopt the Laplace transform method to calculate the image sharpness score. When a blurred frame is detected, combine the motion estimation result to judge whether it is caused by fast motion, and remove the low-quality frames caused by motion blur;
[0086] Decouple the detection, preprocessing and recognition parts, and adopt the pipeline parallel processing method; split the entire processing process into independent stages, including frame acquisition, preprocessing, target detection, feature extraction and final recognition. Each stage runs independently on different computing threads or computing units, and transfers data through the shared memory mechanism; adopt the task scheduling strategy to hand over the tasks with high computational complexity to the GPU for processing, while leaving the lightweight tasks to the CPU;
[0087] In the inference stage, use operator fusion to merge multiple consecutive operators in the computation graph into one operator, and adopt tensor optimization technology for data storage and caching.
[0088] Compared with the prior art, the present invention has the following beneficial effects:
[0089] (1) By adopting the Adaptive Feature Pyramid Network (AFPN), it can fuse multi-level features, significantly improving the accuracy of box number detection at different sizes and angles, and is especially suitable for complex scenarios (such as partial occlusion of the box number, low-contrast environment, etc.); introducing the attention mechanism to optimize the channel feature allocation, effectively suppressing the interference of background noise, and improving the accuracy of box number positioning; combining the edge-aware feature guidance strategy to improve the detection quality of the edges of the box number area, making the framing of the box number more accurate;
[0090] (2) The present invention uses the homography transformation for perspective correction, so that the detected box number area is input to the recognition network under the standard perspective, reducing the recognition errors caused by perspective changes; combining the adaptive morphological processing method to enhance the clarity of character boundaries and improve the separability of characters, making the subsequent character recognition more accurate;
[0091] (3) Through multi-scale feature fusion (MFF), the adaptability of the recognition network to characters of different sizes and angles is enhanced, and the recognition accuracy of the container number characters is improved; attention-guided sequence modeling is introduced to optimize the performance of long string recognition, enabling multiple characters in the container number to maintain the correct order and improving recognition stability; the CTC (Connectionist Temporal Classification) loss function is adopted to eliminate the need for character alignment and still accurately recognize even when the container number characters are irregularly arranged.
[0092] (4) By optimizing the computational graph of the detection and recognition networks, redundant calculations are reduced and the inference speed is increased, enabling the entire detection and recognition process to be completed in milliseconds, meeting the requirements of industrial applications; combined with the edge computing deployment scheme, this method can run on local devices without relying on cloud computing, reducing bandwidth costs and improving the real-time performance and reliability of the system.
[0093] (5) Through large-scale data augmentation (such as random affine transformation, contrast adjustment, noise perturbation, etc.), the model can adapt to various environments (low light, complex background, shooting at different angles, etc.) to ensure the stability of detection and recognition; the method based on morphological clustering is used to automatically segment the container number area, which can effectively remove irrelevant interfering characters and reduce the probability of misrecognition. Description of the Drawings
[0094] Other features, purposes, and advantages of the present invention will become more apparent by reading the detailed description of the non-restrictive embodiments with reference to the following drawings:
[0095] Figure 1 It is the overall model training flowchart;
[0096] Figure 2 It is the progressive adaptive feature fusion process. Detailed Embodiments
[0097] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0098] Embodiment
[0099] The present invention provides a method for detecting the general freight car box number at the gate based on deep learning. Its core process includes five main steps: data collection and enhancement, box number location model construction, box number area extraction and preprocessing, box number recognition model construction, and end-to-end process integration. The specific implementation is as follows:
[0100] Step 1: Data collection. Collect freight car images under different lighting conditions (day / night), different angles (front view / side view), and different weather conditions (sunny / rainy) through the gate camera, and label the location of the box number area and the corresponding character tags. Data preprocessing and image enhancement. For complex weather, corresponding simulation means are used for enhancement to improve the richness of data. For example, different camera perspectives are simulated using geometric transformations, and the changes in lighting are simulated using different brightness and contrast, and then noise is introduced to simulate complex weather changes such as snow and rain.
[0101] Step 2: Box number location model construction. Use a detection network based on deep learning to achieve precise location of the box number area.
[0102] The preprocessed image is input into a pre-trained backbone network (such as ResNet). This network extracts multi-scale feature maps through multiple convolutional operations. These feature maps can capture low-level edge information and high-level semantic information in the image. Next, based on feature extraction, the detection head of the DB++ model further processes the feature maps and outputs two key prediction maps: one is the probability map P(x, y), which represents the probability that each pixel in the image belongs to the box number area; the other is the threshold map T(x, y), which predicts an adaptive threshold for each pixel and is used to distinguish text from the background in the follow-up. To convert these "soft" predictions into "hard" binary results, DB++ introduces a differentiable binary module, and its core formula is:
[0103] B(x,y)=σ(α(P(x,y)-T(x,y)))
[0104] Where, σ represents the Sigmoid function; α is a smoothing factor that controls the sharpness of the binary process. When P(x, y)>T(x, y), the value of B(x, y) output by this function is close to 1, indicating that this pixel is very likely to belong to the box number area; conversely, a value close to 0 indicates the background area.
[0105] To improve the detection accuracy, a feature enhancement mechanism is introduced. Multilevel features are fused through an Adaptive Feature Pyramid Network (AFPN), and a series of feature maps {F1, F2... F N}, since the box number area may exist at different scales, in order to make full use of this multi-scale information, AFPN is introduced to adaptively fuse the features of each layer. Specifically, all feature maps are first adjusted to the same spatial size through upsampling operations, and then weighted fusion is performed using the learned adaptive weights ω i for weighted fusion, which is expressed by the formula:
[0106]
[0107] where U represents the upsampling operation, N represents the number of network layers, and F i is the feature map of the i-th layer, and the weight ω i is calculated through the attention mechanism and often uses Softmax normalization:
[0108]
[0109] where a i , a j respectively represent the attention scores obtained after the features of the i-th and j-th layers pass through a small convolutional or fully connected layer. This process enables the network to automatically adjust the importance of the features of each layer, thereby obtaining an enhanced feature map and enabling the network to take into account the detection accuracy of large and small box numbers. In addition, combined with an edge-aware feature guidance strategy, the edge feature extraction of the box number area is optimized to make the box number framing more accurate.
[0110] To better capture the edge information of the box number area, we use an edge detection operator (such as the Laplacian operator) on the original input image to calculate the edge map:
[0111] E lap = Lap(F(x, y))
[0112] where Lap represents calculating the edge features using the Laplacian operator, and then combining F AFPN and E lap to obtain the final enhanced feature map:
[0113] F fuse = F AFPN + μE lap
[0114] where μ is a hyperparameter that adjusts the weight of the edge features. During the fusion process, the edge features are weighted into the multi-scale features at a certain ratio, making the model more sensitive to edge information.
[0115] To further optimize the detection process, the loss function L = L cls + λ1L reg + λ2L iou is introduced, where L clsRepresents the classification loss, L reg Is the bounding box regression loss, L iou Is the IoU loss, used to improve the localization accuracy; λ1 and λ2 respectively represent the hyperparameters for adjusting the weights of each part of the loss.
[0116] Step 3: Extraction and preprocessing of the box number area. After detecting the box number area, perform affine correction, noise suppression, and contrast enhancement on it to reduce the influence of environmental factors. Aiming at the problem of large differences in the font shapes of box numbers, an adaptive morphological processing method is adopted to make the boundaries of box number characters clearer, facilitating subsequent recognition tasks. At the same time, use a method based on morphological clustering to automatically segment the box number area and remove possible interfering characters. Aiming at the perspective distortion problem existing in the box number area, use the Homography Transformation for correction, and the transformation formula is as follows:
[0117]
[0118] Among them, H is the homography matrix, calculated through four-point transformation, so as to correct the tilted box number area to a standard perspective; and Respectively represent the position of the original image and the position after transformation.
[0119] Step 4: Construction of the box number recognition model. Use a recognition network based on deep learning to perform character recognition on the extracted box number area. In order to improve the recognition accuracy, introduce a character-level feature enhancement strategy, including Attention-guided Sequence Modeling, to improve the recognition ability of long strings through the self-attention mechanism. In addition, introduce Multi-scale Feature Fusion (MFF) in the feature extraction stage to enhance the robustness to characters of different sizes and angles. Combine the self-supervised pre-training strategy to use unlabeled data to enhance the generalization ability of the model, thereby improving the recognition accuracy.
[0120] Use the SVTR (Scene Text Visual Transformer) model as the basic recognition network. This model first extracts multi-scale features from the preprocessed image I through convolutional or Patch Embedding layers, that is, obtains a set of feature maps {F1, F2…F N}, use a feature enhancement pyramid for feature fusion, and then convert the enhanced feature maps into a feature sequence suitable for processing by the Transformer module and input it into the Attention-guided Sequence Modeling module. Let this feature sequence be X = {x1, x2…x n}, global modeling is carried out using the self-attention mechanism, and the calculation formula is:
[0121]
[0122] where Q = XW Q , K = XW k and V = XW V are the linear transformations of the query, key, and value respectively, d k is the dimension of the key, and W Q , W k , W V are the weight matrices for calculating the query vector, key vector, and value vector respectively; by performing global context modeling, the ability to capture the dependencies between characters in long strings is improved.
[0123] To further improve the generalization performance of the model, we combine self-supervised pre-training strategies. At this stage, by pre-training on a large amount of unlabeled data (such as using contrastive learning strategies), the feature representations are optimized. The contrastive loss function can be expressed as:
[0124]
[0125] where z i , z j are the feature representations of different augmented versions of the same character or region, sim represents the cosine similarity, τ represents the hyperparameter, K represents the number of negative samples, and k represents the index variable;
[0126] The loss function during the recognition process uses the CTC (Connectionist Temporal Classification) loss function, which is defined as follows:
[0127]
[0128] where p(|π t X input ) represents the probability of the target character sequence π input given the input sequence X t ; t represents a certain sequence moment, Β represents the function that maps the path π to the final character sequence y, and CTC can effectively solve the problem of unaligned annotations.
[0129] Step 5, End-to-end process integration: After completing the container number detection and recognition, all modules are integrated into an end-to-end inference process to implement an efficient and automated container number detection and recognition system. This method uses a video stream for real-time calculation and inference, and judges moving objects online, filters out and eliminates blurred frames, and improves the detection efficiency. We decouple the detection, preprocessing, and recognition modules and adopt a pipeline parallel processing method to reduce the computational pressure of the CPU on a single task, improve the overall operation efficiency, and subsequently reduce computational redundancy and improve the inference efficiency by optimizing the computational graph during the inference process using means such as Operator Fusion and Tensor Optimization.
[0130] This method uses a video stream for real-time calculation and inference, judges moving objects, filters out and eliminates blurred frames, and improves the detection efficiency. Specifically, we first use the Optical Flow method to judge the movement of the target object in the picture. Then, the Laplace transform method is used to calculate the image sharpness score. When a blurred frame is detected, it is judged whether it is caused by fast movement in combination with the motion estimation result, and low-quality frames caused by motion blur are eliminated to ensure the image quality input to the model and improve the detection stability.
[0131] Next, we decouple the detection, preprocessing, and recognition modules and adopt a pipeline parallel processing method to reduce the computational pressure of the CPU on a single task and improve the overall operation efficiency. Specifically, we split the entire processing flow into independent stages, including frame acquisition, preprocessing (denoising, enhancement), target detection, feature extraction, and final recognition. Each stage runs independently on different computational threads or computational units (such as GPUs, NPUs), and data is passed through the Shared Memory mechanism to avoid the computational overhead caused by data copying. In addition, we adopt a task scheduling strategy to hand high-computation tasks (such as deep neural network inference) to the GPU for processing, while lightweight tasks (such as data preprocessing) are left to the CPU, thereby optimizing the utilization of computational resources, avoiding overloading a single computational unit, and improving the overall throughput.
[0132] In the inference stage, we reduce computational redundancy and improve the inference efficiency by optimizing the computational graph. First, we use operator fusion to merge multiple consecutive operators (such as convolution, normalization, activation function) in the computational graph into one operator, reducing the computational overhead and the number of memory accesses. Second, Tensor Optimization techniques are adopted, such as reducing tensor dimension redundancy, using an efficient storage format such as NHWC to optimize memory access, and an intelligent caching strategy, to reduce the time overhead of data movement.
[0133] In addition, to enable the entire detection and recognition process to run efficiently locally, we adopt an edge computing deployment solution to perform real-time computing and inference on edge devices, reducing the dependence on cloud computing resources and improving the response speed. First, detection, preprocessing, and feature extraction are executed locally on the edge device, and some computing tasks are offloaded to the edge side to reduce the cloud computing burden. Then, the TensorRT framework is used in combination with model pruning and knowledge distillation methods to reduce computing resource consumption and ensure that low-power devices can also run efficient object detection tasks.
[0134] Figure 1 It represents the overall model training process. First, the data at the gate is collected, and then the corresponding enhancement is performed on the image data. The effectiveness of the method is demonstrated through text localization and recognition models, and finally, in an end-to-end manner, it shows that the method works with the least resources in implementation.
[0135] Figure 2 It represents progressive adaptive feature fusion. Multi-scale features (low-level, middle-level, high-level) are extracted through the backbone network, and an adaptive feature fusion module is used to dynamically integrate features at different levels to generate a multi-scale feature pyramid (P3, P4, P5). The feature expression ability is further enhanced through multi-scale feature interaction. Finally, the bounding boxes of the objects and the pixel-level segmentation results are output through the object detection head and the instance segmentation head respectively, realizing efficient multi-task object recognition and segmentation.
[0136] In the specific embodiments of the present invention, multiple embodiments are provided for different application environments and requirements. For example, in a truck license plate recognition scenario at a certain gate, an industrial camera with 5 million pixels is selected, the installation height is about 3.5 meters, and the tilt angle is controlled within 15° to ensure that the license plate number is clearly visible. The data acquisition device includes multiple fixed-position monitoring cameras to cover different perspectives and lighting conditions. In the license plate detection stage, a detection network based on deep learning is combined with an adaptive feature pyramid network (AFPN) for multi-scale feature fusion to improve the detection accuracy of license plate numbers of different sizes. In the process of license plate area extraction and preprocessing, homography transformation is used for distortion correction, and morphological methods are combined to enhance character boundaries to ensure the accuracy of subsequent recognition. In the character recognition stage, a deep learning recognition network optimized by CTC loss in a deep learning-based text recognition network is used to achieve high-precision recognition of deformed and blurred characters. Further optimized embodiments include: in low-light environments, the recognition stability is improved through contrast enhancement and noise reduction techniques. The present invention proposes an automated method based on deep learning for the detection and recognition of truck license plates in complex environments, and its feasibility is verified through multiple embodiments. The use of high-resolution industrial cameras and different types of data acquisition methods ensures adaptability to various environments. Overall, the present invention improves the detection and recognition accuracy, enhances the robustness and real-time performance of the system through an innovative deep learning architecture and optimized calculation processes, and has strong industrial application value.
[0137] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the systems, devices, and their respective modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structure within the hardware component; the modules for implementing various functions can also be regarded as either software programs for implementing the method or the structure within the hardware component.
[0138] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.
Claims
1. A method for detecting universal truck box numbers at a gate based on deep learning, characterized in that: include: Step 1: Use the gate camera to collect truck images under different lighting conditions, angles and weather conditions, mark the container number area location and corresponding character labels, and use corresponding simulation methods to enhance the images according to different weather conditions; Step 2: Use a deep learning-based detection network to accurately locate the box number area; Step 3: After detecting the box number area, perform affine correction, noise suppression and contrast enhancement on it; Step 4: Use a deep learning-based recognition network to perform character recognition in the box number area, improve the recognition ability of long character strings through the self-attention mechanism, introduce multi-scale feature fusion in the feature extraction stage to improve the robustness of characters of different sizes and angles, combine the self-supervised pre-training strategy, and use unlabeled data to enhance the generalization ability of the model; Step 5: After completing the box number detection and recognition, use the video stream for real-time calculation and reasoning, judge the moving objects online, filter out the blurred frames and remove them, use pipeline parallel processing, and then optimize the calculation graph in the reasoning process, use operator fusion and tensor optimization to reduce calculation redundancy.
2. The method for detecting universal truck box numbers at a gate based on deep learning according to claim 1 is characterized in that: The step 2 comprises: The preprocessed image is input into the pretrained backbone network, which extracts multi-scale feature maps through multi-layer convolution operations, and then further processes the feature maps to output two key prediction maps: one is the probability map P(x, y), which indicates the probability that each pixel in the image belongs to the box area; the other is the threshold map T(x, y), which predicts an adaptive threshold for each pixel for subsequent distinction between text and background; the differentiable binarization module is introduced to convert the prediction into a binary result, and its core formula is: B(x,y)=σ(α(P(x,y)-T(x,y))) Among them, σ represents the Sigmoid function; α is the smoothing factor, which controls the sharpness of the binarization process; when P(x,y)>T(x,y), the value of B(x,y) output by this function is close to 1, indicating that the pixel belongs to the box area; otherwise, the value close to 0 indicates the background area; In order to improve the detection accuracy, a feature enhancement mechanism is introduced. Multi-level features are fused through an adaptive feature pyramid network, and a series of feature maps {F1, F2…F N }, AFPN is introduced to adaptively fuse features of each layer; First, all feature maps are adjusted to the same spatial size through upsampling operations, and then the learned adaptive weights ω are used i Perform weighted fusion, the formula is expressed as: Among them, U represents the upsampling operation, N represents the number of network layers, and F i is the feature map of the i-th layer, and the weight ω i Calculated by the attention mechanism and normalized by Softmax: Among them, a i 、a j Respectively represent the attention scores of the i-th and j-th layer features after a small convolution or fully connected layer; In order to better capture the edge information of the box area, the edge detection operator is used on the original input image to calculate the edge map: E laP =Lap(F(x,y)) Among them, Lap represents the Laplace operator to calculate the edge features, and then F AFPN and E lap Combine them to get the final enhanced feature map: F fuse =F AFPN +μE lap Among them, μ is a hyperparameter for adjusting the weight of edge features. During the fusion process, edge features are weighted to multi-scale features at a preset ratio, making the model more sensitive to edge information. In order to further optimize the detection process, the loss function L = L cls +λ1L reg +λ2L iou , where L cls represents the classification loss, L reg is the bounding box regression loss, L iou is the IoU loss, which is used to improve the positioning accuracy; λ1 and λ2 represent the hyperparameters for adjusting the loss weights of each part.
3. The method for detecting universal truck box numbers at a gate based on deep learning according to claim 1 is characterized in that: The step 3 comprises: Adaptive morphological processing method is used to make the boundary of box number characters clearer. At the same time, the method based on morphological clustering is used to automatically segment the box number area and remove interfering characters. To solve the perspective distortion problem in the box number area, homography transformation is used for correction. The transformation formula is as follows: Among them, H is the homography matrix, which is calculated by four-point transformation, so as to correct the tilted box area to the standard perspective; and Represent the position of the original image and the transformed position respectively.
4. The method for detecting a universal truck box number at a gate based on deep learning according to claim 1, characterized in that: The step 4 comprises: The SVTR model is used as the basic recognition network. The model first extracts multi-scale features from the preprocessed image through convolution or Patch Embedding layer, that is, a set of feature maps {F1, F2…F N }, feature enhancement pyramid is used for feature fusion, and then the enhanced feature map is converted into a feature sequence suitable for processing by the Transformer module and input into the attention-guided sequence modeling module. Let the feature sequence be X = {x1, x2…x n }, use the self-attention mechanism for global modeling, and the calculation formula is: Where Q = XW Q , K = XW k and V=XW V are linear transformations of query, key and value respectively, d k is the dimension of the key, W Q , W k , W V They are weight matrices used to calculate query vectors, key vectors, and value vectors respectively; by modeling global context, the ability to capture dependencies between characters in long strings is improved; In order to further improve the generalization performance of the model, combined with the self-supervised pre-training strategy, at this stage, by pre-training on unlabeled data, the feature representation is optimized, and the contrast loss function is expressed as: Among them, z i ,z j It is the feature representation of different enhanced versions of the same character or region, sim represents cosine similarity, τ represents hyperparameter, K represents the number of negative samples, and k represents index variable; The loss function in the recognition process adopts the CTC loss function, which is defined as follows: Among them, p(|π t X input ) means that given an input sequence X input In the case of t ; t represents a certain sequence moment, and Β represents the function that maps the path π to the final character sequence y.
5. The method for detecting universal truck box numbers at a gate based on deep learning according to claim 1 is characterized in that: The step 5 comprises: The optical flow method is used to determine the motion of the target object in the picture, and the Laplace transform method is used to calculate the image clarity score. When a blurred frame is detected, the motion estimation result is combined to determine whether it is caused by fast motion, and the low-quality frames caused by motion blur are eliminated; Decouple the detection, preprocessing and recognition parts and adopt pipeline parallel processing; split the entire processing flow into independent stages, including frame acquisition, preprocessing, target detection, feature extraction and final recognition. Each stage runs independently on different computing threads or computing units and transfers data through a shared memory mechanism; adopt a task scheduling strategy so that high-computation tasks are handled by the GPU, while lightweight tasks are left to the CPU; In the inference phase, operator fusion is used to merge multiple consecutive operators in the computation graph into one operator, and tensor optimization technology is used for data storage and caching.
6. A general truck box number detection system at a gate based on deep learning, characterized in that: include: Module M1: The gate camera collects images of trucks under different lighting conditions, angles and weather conditions, and marks the container number area location and corresponding character labels, and uses corresponding simulation methods to enhance the images according to different weather conditions; Module M2: Use a deep learning-based detection network to achieve accurate positioning of the box number area; Module M3: After detecting the box number area, it performs affine correction, noise suppression and contrast enhancement; Module M4: Use a deep learning-based recognition network to perform character recognition in the box number area, improve the recognition ability of long character strings through the self-attention mechanism, introduce multi-scale feature fusion in the feature extraction stage to improve the robustness of characters of different sizes and angles, combine the self-supervised pre-training strategy, and use unlabeled data to enhance the generalization ability of the model; Module M5: After completing the box number detection and recognition, the video stream is used for real-time calculation and reasoning, the moving objects are judged online, the blurred frames are screened out and eliminated, the pipeline parallel processing method is adopted, and the calculation graph in the subsequent reasoning process is optimized, and the calculation redundancy is reduced by using operator fusion and tensor optimization.
7. The deep learning-based universal truck box number detection system at a gate according to claim 6 is characterized in that: The module M2 comprises: The preprocessed image is input into the pretrained backbone network, which extracts multi-scale feature maps through multi-layer convolution operations, and then further processes the feature maps to output two key prediction maps: one is the probability map P(x, y), which indicates the probability that each pixel in the image belongs to the box area; the other is the threshold map T(x, y), which predicts an adaptive threshold for each pixel for subsequent distinction between text and background; the differentiable binarization module is introduced to convert the prediction into a binary result, and its core formula is: B(x,y)=σ(α(P(x,y)-T(x,y))) Among them, σ represents the Sigmoid function; α is the smoothing factor, which controls the sharpness of the binarization process; when P(x,y)>T(x,y), the value of B(x,y) output by this function is close to 1, indicating that the pixel belongs to the box area; otherwise, the value close to 0 indicates the background area; In order to improve the detection accuracy, a feature enhancement mechanism is introduced. Multi-level features are fused through an adaptive feature pyramid network, and a series of feature maps {F1, F2…F N }, AFPN is introduced to adaptively fuse features of each layer; First, all feature maps are adjusted to the same spatial size through upsampling operations, and then the learned adaptive weights ω are used i Perform weighted fusion, the formula is expressed as: Among them, U represents the upsampling operation, N represents the number of network layers, and F i is the feature map of the i-th layer, and the weight ω i Calculated by the attention mechanism and normalized by Softmax: Among them, a i 、a j Respectively represent the attention scores of the i-th and j-th layer features after a small convolution or fully connected layer; In order to better capture the edge information of the box area, the edge detection operator is used on the original input image to calculate the edge map: E laP =Lap(F(x,y)) Among them, Lap represents the Laplace operator to calculate the edge features, and then F AFPN and E lap Combine them to get the final enhanced feature map: F fuse =F AFPN +μE lap Among them, μ is a hyperparameter for adjusting the weight of edge features. During the fusion process, edge features are weighted to multi-scale features at a preset ratio, making the model more sensitive to edge information. In order to further optimize the detection process, the loss function L = L cls +λ1L reg +λ2L iou , where L cls represents the classification loss, L reg is the bounding box regression loss, L iou is the IoU loss, which is used to improve the positioning accuracy; λ1 and λ2 represent the hyperparameters for adjusting the loss weights of each part.
8. The deep learning-based gate universal truck box number detection system according to claim 6 is characterized in that: The module M3 comprises: Adaptive morphological processing method is used to make the boundary of box number characters clearer. At the same time, the method based on morphological clustering is used to automatically segment the box number area and remove interfering characters. To solve the perspective distortion problem in the box number area, homography transformation is used for correction. The transformation formula is as follows: Among them, H is the homography matrix, which is calculated by four-point transformation, so as to correct the tilted box area to the standard perspective; and Represent the position of the original image and the transformed position respectively.
9. The deep learning-based gate universal truck box number detection system according to claim 6 is characterized in that: The module M4 comprises: The SVTR model is used as the basic recognition network. The model first extracts multi-scale features from the preprocessed image through convolution or Patch Embedding layer, that is, a set of feature maps {F1, F2…F N }, feature enhancement pyramid is used for feature fusion, and then the enhanced feature map is converted into a feature sequence suitable for processing by the Transformer module and input into the attention-guided sequence modeling module. Let the feature sequence be X = {x1, x2…x n }, use the self-attention mechanism for global modeling, and the calculation formula is: Where Q = XW Q , K = XW k and V=XW V are linear transformations of query, key and value respectively, d k is the dimension of the key, W Q , W k , W V They are weight matrices used to calculate query vectors, key vectors, and value vectors respectively; by modeling global context, the ability to capture dependencies between characters in long strings is improved; In order to further improve the generalization performance of the model, combined with the self-supervised pre-training strategy, at this stage, by pre-training on unlabeled data, the feature representation is optimized, and the contrast loss function is expressed as: Among them, z i ,z j It is the feature representation of different enhanced versions of the same character or region, sim represents cosine similarity, τ represents hyperparameter, K represents the number of negative samples, and k represents index variable; The loss function in the recognition process adopts the CTC loss function, which is defined as follows: Among them, p(|π t X input ) means that given an input sequence X input In the case of t ; t represents a certain sequence moment, and Β represents the function that maps the path π to the final character sequence y.
10. The deep learning-based gate universal truck box number detection system according to claim 6, characterized in that: The module M5 comprises: The optical flow method is used to determine the motion of the target object in the picture, and the Laplace transform method is used to calculate the image clarity score. When a blurred frame is detected, the motion estimation result is combined to determine whether it is caused by fast motion, and the low-quality frames caused by motion blur are eliminated; Decouple the detection, preprocessing and recognition parts and adopt pipeline parallel processing; split the entire processing flow into independent stages, including frame acquisition, preprocessing, target detection, feature extraction and final recognition. Each stage runs independently on different computing threads or computing units and transfers data through a shared memory mechanism; adopt a task scheduling strategy so that high-computation tasks are handled by the GPU, while lightweight tasks are left to the CPU; In the inference phase, operator fusion is used to merge multiple consecutive operators in the computation graph into one operator, and tensor optimization technology is used for data storage and caching.
Citation Information
Patent Citations
Container number detection method based on deep neural network
CN110852324A