Vision-based container truck top number and container number identification method and device

Through visual recognition technology combined with object detection and optical character recognition, the efficiency and accuracy problems in the identification of truck roof numbers and container numbers are solved, and high-precision recognition in complex scenarios is achieved, labor costs are reduced, and logistics transportation safety and efficiency are improved.

CN120340009APending Publication Date: 2025-07-18WUHAN CHUANFENG SOFTWARE TECH CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510317832.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has problems such as manual identification efficiency bottlenecks, limitations of single technology identification and poor adaptability in complex scenarios in the identification of truck roof numbers and container numbers, which leads to low recognition accuracy and difficult to meet real-time performance. The lack of spatial relationship modeling between the front and containers leads to a high matching error rate when multiple vehicles are paralleled.

Method used

The vision-based truck roof number and container number recognition method is adopted, combined with object detection and optical character recognition technology, through dynamic range expansion, defog treatment, ROI extraction, improved object detection network, perspective correction and super-resolution reconstruction, combined with the two-stage identification network and LSTM model, spatial position constraints and motion trajectory consistency judgment are established to achieve high-precision recognition.

Benefits of technology

It improves the accuracy and efficiency of the identification of truck roof numbers and container numbers, adapts to identification in complex scenarios, reduces labor costs, improves the safety and efficiency of logistics and transportation, and supports real-time identification and precise matching in multi-target scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340009A_ABST
    Figure CN120340009A_ABST
Patent Text Reader

Abstract

The invention discloses a vision-based container truck top number and container number identification method and device, a storage medium and electronic equipment. The method comprises the following steps: acquiring a roof image and a container image, and preprocessing the roof image and the container image; inputting the preprocessed roof image and container image into a target detection network to obtain a vehicle head target area and a container target area; performing character area precision processing on the vehicle head target area and the container target area to obtain a processed vehicle head target area and a processed container target area; and performing optical character recognition on the processed vehicle head target area and the container target area to obtain a container number and a vehicle top number. According to the invention, the identification accuracy and identification efficiency of the top number and the container number of the container truck in a complex scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image and character recognition, and particularly to a vision-based method, device, storage medium, and electronic device for identifying the roof number of a container truck and the container number. Background Art

[0002] In the fields of port logistics and container transportation, accurately identifying the roof number of a container truck and the container number is a key technology for realizing automated scheduling, cargo tracking, and intelligent gate management. The current industry mainly has the following technical defects: Bottleneck in manual recognition efficiency: Traditional methods rely on manual visual recording, and operators need to climb for inspection or check at close range. According to statistics, the customs clearance time of a single container truck increases by an average of 3 - 5 minutes, and there is a misrecognition rate of 15% - 20% caused by visual fatigue, and the error rate is even higher in harsh environments such as rainy and foggy weather.

[0003] Limitations of single technology recognition: Existing automated solutions mostly use independent OCR technology for direct recognition, but do not effectively solve the problem of target positioning. Experimental data shows that when directly performing OCR processing on the entire image, due to interference from vehicle body advertisements, background text, etc., the proportion of invalid recognition areas is as high as over 65%, and due to the difference in the distance of the target from the lens, the character resolution fluctuates (40 - 200 pixel heights), and the average recognition accuracy is less than 70%.

[0004] Poor adaptability to complex scenarios: The container truck operation environment has strong backlighting (illuminance > 10^5 lux), rain and snow occlusion, license plate fouling, etc. Interference. The positioning failure rate of traditional image processing methods (such as edge detection + projection segmentation) in low-contrast scenarios (gray difference < 30) exceeds 40%, and the processing delay of traditional CNN-based object detection models (such as Faster R-CNN) at 1080p resolution reaches 300 - 500 ms, which is difficult to meet the real-time requirements.

[0005] The industry has tried to improve the recognition effect through engineering means such as increasing fill lights and restricting vehicle speed (< 5 km / h), but this has led to an increase in equipment cost by more than 200% and a reduction in traffic efficiency. How to achieve high-precision real-time recognition in complex dynamic scenarios has become a technical difficulty restricting the intelligent upgrade of the industry. Summary of the Invention

[0006] Embodiments of this application provide a vision-based method, device, storage medium, and electronic device for identifying the roof number of a container truck and the container number, which can improve the accuracy and recognition efficiency of identifying the roof number of a container truck and the container number in complex scenarios.

[0007] Embodiments of this application provide a vision-based method for identifying the roof number of a container truck and the container number, including: Obtain the roof image and the container image, and preprocess the roof image and the container image; Input the preprocessed roof image and container image into the target detection network to obtain the cab target area and the container target area; Perform fine processing on the character areas of the cab target area and the container target area to obtain the processed cab target area and container target area; Perform optical character recognition on the processed cab target area and container target area to obtain the container number and the roof number.

[0008] Further, in the above method for recognizing the container roof number and the container number based on vision, where preprocessing the roof image and the container image includes: Perform dynamic range expansion on the roof image and the container image; Perform defogging processing on the roof image and the container image after dynamic range expansion; Perform ROI extraction on the roof image and the container image after defogging processing to divide the dynamic monitoring area.

[0009] Further, in the above method for recognizing the container roof number and the container number based on vision, where the target detection network includes an improved backbone network, an improved neck network, and a head network; The improved backbone network includes a Focus module, a CSP module, and an improved convolutional layer; The Focus module is used to perform slicing operations on the input image to obtain sliced images; The CSP module is used to divide the sliced image into two parts, one part passes through the improved convolutional layer, and the other part skips the improved convolutional layer, and the features of the two parts are spliced; The improved convolutional layer is a Ghost module, and the Ghost module is used to perform cheap operations on the input image; The improved neck network is implemented by embedding a CBAM attention module.

[0010] Further, in the above method for recognizing the container roof number and the container number based on vision, where performing fine processing on the character areas of the cab target area and the container target area to obtain the processed cab target area and container target area includes: Perform perspective correction on the cab target area and the container target area to obtain the corrected cab target area and container target area; Perform super-resolution reconstruction on the corrected cab target area and container target area to obtain the reconstructed corrected cab target area and container target area.

[0011] Further, in the above method for identifying the container number and the truck roof number based on vision, the perspective correction of the truck head target area and the container target area to obtain the corrected truck head target area and container target area includes: Extracting the corner points of the target area, calculating the homography matrix according to the corner points, performing image transformation based on the homography matrix, and correcting the target area with an inclination greater than a preset angle to a horizontal state; The super-resolution reconstruction of the corrected truck head target area and container target area to obtain the reconstructed corrected truck head target area and container target area includes: Judging whether the character height of the corrected truck head target area and container target area is less than a preset pixel value; If it is less than, input the corresponding target area into the super-resolution model for super-resolution reconstruction.

[0012] Further, in the above method for identifying the container number and the truck roof number based on vision, the optical character recognition of the processed truck head target area and container target area to obtain the container number and the truck roof number includes: Inputting the processed truck head target area and container target area into a two-stage recognition network to obtain the container number and the truck roof number; Among them, the two-stage recognition network includes a document detection module and a text recognition module; The text detection module is used to perform text segmentation on the processed truck head target area and container target area; The text recognition module is used to detect and recognize the segmented text.

[0013] Further, in the above method for identifying the container number and the truck roof number based on vision, after the step of performing optical character recognition on the processed truck head target area and container target area to obtain the container number and the truck roof number, it includes: Verifying the container number and predicting the missing characters of the truck roof number based on the LSTM model.

[0014] Further, in the above method for identifying the container number and the truck roof number based on vision, after the step of performing optical character recognition on the processed truck head target area and container target area to obtain the container number and the truck roof number, it further includes: Identifying the entire truck, regarding the truck head and the container as components of the entire truck, and determining whether the container and the truck head belong to the same truck through spatial position constraints and the consistency of the movement trajectory.

[0015] Further, in the above method for identifying the container number and the truck roof number based on vision, the determination of whether the container and the truck head belong to the same container truck through spatial position constraint and motion trajectory consistency includes: Establish a two-dimensional plane coordinate system with a fixed point in the passing area of the container truck as the coordinate origin, and determine whether the relative positions of the truck head and the container in the coordinate system meet the requirements; Calculate the motion trajectories of the truck head and the container. If the average deviation between the motion trajectory of the truck head and that of the container is less than a preset distance, and the included angle of the motion directions is less than a preset angle, it is determined that the motion trajectories of the truck head and the container are consistent.

[0016] An embodiment of the present application further provides an apparatus for identifying the container number and the truck roof number based on vision, including: An acquisition and preprocessing module, configured to acquire a roof image and a container image, and preprocess the roof image and the container image; A target detection module, configured to input the preprocessed roof image and container image into a target detection network to obtain a truck head target area and a container target area; A character area fine processing module, configured to perform character area fine processing on the truck head target area and the container target area to obtain a processed truck head target area and a container target area; A character recognition module, configured to perform optical character recognition on the processed truck head target area and container target area to obtain the container number and the truck roof number. An embodiment of the present application further provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded by a processor to execute any one of the above methods for identifying the container number and the truck roof number based on vision.

[0017] An embodiment of the present application further provides an electronic device, including a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used for the steps in any one of the above methods for identifying the container number and the truck roof number based on vision.

[0018] The method, apparatus, storage medium, and electronic device for identifying the container number and the truck roof number based on vision provided by the present application combine target detection and optical character recognition technologies to identify the truck head and the container number, improve the recognition accuracy of the container truck roof number and the container number, and can adapt to the recognition in various complex scenarios. The present application optimizes the target detection model into a lightweight model to improve the efficiency of target detection. The present application also performs joint recognition of the truck head and the container through an association matching technology to improve the recognition accuracy in a multi-target scenario. Description of the Drawings

[0019] Combined with the accompanying drawings, through a detailed description of the specific embodiments of the present application, the technical solutions and other beneficial effects of the present application will become obvious.

[0020] Figure 1 It is a flowchart of a method for identifying the roof number and container number of a container truck based on vision provided by an embodiment of the present application.

[0021] Figure 2 It is a schematic structural diagram of a device for identifying the roof number and container number of a container truck based on vision provided by an embodiment of the present application.

[0022] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific Embodiments

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0024] The current industry mainly has the following technical defects: Bottleneck in manual recognition efficiency: Traditional methods rely on manual visual records, and operators need to climb for inspection or check at close range. According to statistics, the customs clearance time of a single container truck increases by an average of 3 - 5 minutes, and there is a misrecognition rate of 15% - 20% caused by visual fatigue, and the error rate is even higher in harsh environments such as rainy and foggy weather.

[0025] Limitations of single - technology recognition: Existing automation solutions mostly use independent OCR technology for direct recognition, but do not effectively solve the problem of target positioning. Experimental data shows that when directly performing OCR processing on the entire image, due to interference from vehicle body advertisements, background text, etc., the proportion of invalid recognition areas is as high as more than 65%, and due to the difference in the distance of the target from the lens, the character resolution fluctuates (40 - 200 pixel heights), and the average recognition accuracy is less than 70%.

[0026] Poor adaptability to complex scenarios: The container truck operation environment has strong backlighting (illuminance > 10^5 lux), rain and snow occlusion, license plate fouling, etc. Interference. The positioning failure rate of traditional image - processing methods (such as edge detection + projection segmentation) in low - contrast scenarios (gray - level difference < 30) exceeds 40%, and the processing delay of traditional CNN - based object detection models (such as Faster R - CNN) at 1080p resolution reaches 300 - 500 ms, which is difficult to meet the real - time requirements.

[0027] Problem of missing correlation matching: The existing technology lacks the modeling of the spatial relationship between the truck head and the container, resulting in more than 30% of the matching errors between the container trucks when multiple trucks run in parallel. Especially when there is visual occlusion between the truck head and the container (occlusion area > 40%), the accuracy rate of the traditional matching algorithm based on position hypothesis drops sharply to less than 50%.

[0028] The root cause of the above problems is that the technical closed-loop of "detection - positioning - correlation - recognition" has not been constructed. The specific manifestations are as follows: (1) In the target detection stage, lightweight models are not used for optimization, making it difficult to balance accuracy and speed; (2) Before OCR recognition, there is a lack of character region correction based on perspective transformation, and when the inclination exceeds 15°, the recognition error rate increases by 3 times; (3) A box-truck matching model based on spatial geometric constraints has not been established, resulting in correlation errors in multi-target scenarios.

[0029] To solve the above problems, the embodiments of the present application provide a method, device, storage medium, and electronic device for identifying the truck roof number and container number based on vision. The identification device for the truck roof number and container number based on vision provided by the embodiments of the present application can be integrated in an electronic device, and the electronic device can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a microprocessing box, or other devices, etc.

[0030] Please refer to Figure 1 , Figure 1 which is the flowchart of the method for identifying the truck roof number and container number based on vision provided by the embodiments of the present application. It is applied to an electronic device, and the method for identifying the truck roof number and container number based on vision includes the following steps: S1, obtain the truck roof image and the container image, and preprocess the truck roof image and the container image.

[0031] First, set up the hardware configuration on the container truck: Use 10 high-definition cameras with 3 million pixels. 4 box number recognition probes are installed on both sides of the saddle beam of the large truck about 3.7 meters above the ground, covering the passing area of the container truck with a 70° pitch angle. 4 truck roof number recognition probes are installed on the 4 door legs of the large truck 15 - 17 meters above the ground, vertically downward covering the truck head stop area. 2 locking box number recognition probes are installed on the cross beam of the large truck about 22 meters above the ground, covering the aerial locking container area with a 70° pitch angle, and cooperating with an LED stroboscopic fill light (trigger frequency synchronized with the camera frame rate) to solve the backlight problem.

[0032] The preprocessing of the truck roof image and the container image in step S1 includes the following steps: S11, perform dynamic range expansion on the truck roof image and the container image.

[0033] The CLAHE (Contrast Limited Adaptive Histogram Equalization) algorithm is adopted to expand the gray level of the low-illumination areas (<50 lux) in the roof image and the container image from 0 - 80 to 0 - 150.

[0034] S12, Perform defogging processing on the roof image and the container image after dynamic range expansion.

[0035] Based on the dark channel prior model, optimize the transmission rate map of the roof image and the container image to increase the SSIM (structural similarity index) index of the vehicle number area by more than 0.3 S13, Extract the ROI (Region of Interest) from the defogged roof image and container image to divide the dynamic monitoring area.

[0036] The dynamic detection area is delimited by the background difference method, and subsequent detection is directly performed on the dynamic monitoring area. This method can reduce the amount of invalid calculations by 60%.

[0037] In one embodiment, the width of the dynamic detection area is 6 - 12 meters, and the height is 1.5 - 4 meters.

[0038] S2, Input the preprocessed roof image and container image into the target detection network to obtain the head target area and the container target area.

[0039] In one embodiment, the target detection network can be an improved YOLOv5s network, and the target detection network includes an improved backbone network, an improved neck network, and a head network.

[0040] The improved backbone network includes a Focus module, a CSP module, and an improved convolutional layer; The Focus module is used to perform slicing operations on the input image to obtain sliced images; The CSP module is used to divide the sliced image into two parts, one part passes through the improved convolutional layer, the other part skips the improved convolutional layer, and the features of the two parts are spliced; The improved convolutional layer is a Ghost module, and the Ghost module is used to perform cheap operations on the input image; The improved neck network is implemented by embedding a CBAM (Convolutional Block Attention Module) attention module.

[0041] The preprocessed roof image and container image are input into the object detection network to obtain the head target area and the container target area, which specifically includes: inputting the preprocessed roof image and container image into the improved backbone network, and through slicing, feature extraction and splicing, the roof feature map and the container feature map are obtained; inputting the roof feature map and the container feature map into the improved neck network for lightweight feature fusion, and finally inputting the fused features into the head network for detection to obtain the head target area and the container target area. Among them, the head target area (taking the area where the roof number is located as the detection target) includes the bounding box coordinates (x1, y1, x2, y2) and confidence of the head area, and the container target area (detecting the box number area on the side and end face of the container) includes the minimum bounding rectangle and confidence.

[0042] The object detection network generally includes the following improvements: lightweight design: replacing the C3 module in the Backbone (backbone network) with the Ghost module, reducing the number of model parameters from 7.2M to 3.8M; attention mechanism: embedding the CBAM attention module in the Neck part, increasing the head detection AP@0.5 from 89.2% to 93.6%; multi-scale training: adaptively adjusting the input image size (from 640×640 to 1280×1280) to cope with the target scale change. By performing lightweight optimization on the object detection network, the accuracy and speed of object detection are balanced.

[0043] S3. Perform fine processing on the character areas of the head target area and the container target area to obtain the processed head target area and container target area.

[0044] In one embodiment, step S3 includes the following steps: S31. Perform perspective correction on the head target area and the container target area to obtain the corrected head target area and container target area.

[0045] Step S31 includes: S311. Extract the corner points of the target area, calculate the homography matrix according to the corner points, and perform image transformation based on the homography matrix to correct the target area with an inclination greater than the preset angle to the horizontal state.

[0046] Specifically, the Shi-Tomasi algorithm is used to extract the corner points of the head target area and the container target area.

[0047] In one embodiment, the preset angle can be 15°.

[0048] S32. Perform super-resolution reconstruction on the corrected head target area and the container target area to obtain the reconstructed and corrected head target area and container target area.

[0049] Step S32 includes: S321, determining whether the character heights of the corrected vehicle head target area and the container target area are less than a preset pixel value.

[0050] In one embodiment, the preset pixel value can be 50 pixels.

[0051] S322, if less, input the corresponding target area into the super-resolution model for super-resolution reconstruction.

[0052] Specifically, the super-resolution model can use the ESRGAN network (Enhanced Super-Resolution Generative Adversarial Network) to magnify the low-resolution target by 2 times.

[0053] Furthermore, after super-resolution reconstruction, calculate the PSNR (Peak Signal-to-Noise Ratio) of the image, and evaluate the quality of super-resolution reconstruction through PSNR.

[0054] S4, perform optical character recognition on the processed vehicle head target area and container target area to obtain the container number and the roof number. Input the processed vehicle head target area and container target area into a two-stage recognition network to obtain the container number and the roof number.

[0055] Among them, the two-stage recognition network includes a document detection module and a text recognition module; The text detection module is used to perform text segmentation on the processed vehicle head target area and container target area.

[0056] Specifically, the text detection module adopts the DB (Differentiable Binarization) algorithm, which combines deformable convolution and adaptive thresholding to achieve high-precision text area detection.

[0057] The text recognition module is used to detect and recognize the segmented text.

[0058] Specifically, the text recognition module uses the SVTR-Lite network, which supports multi-language mixed recognition, is especially suitable for complex scenarios such as container numbers, and has an accuracy rate of 98.7% on the container number validation set.

[0059] Furthermore, after step S4, it further includes: S5, verify the container number and predict the missing characters of the roof number based on the LSTM model.

[0060] Specifically, apply the ISO 6364 standard check code algorithm to verify the container number and automatically filter out illegal numbers. Predict the missing characters in the roof number based on the LSTM model (Long Short-Term Memory model) (context window length = 7).

[0061] Further, after step S4, it further includes: S5. Identify the entire truck-trailer combination, regarding the cab and the container as components of the whole truck-trailer combination, and determine whether the container and the cab belong to the same truck-trailer combination through spatial position constraints and the consistency of the movement trajectories.

[0062] In one embodiment, step S5 includes the following steps: S51. Establish a two-dimensional plane coordinate system with a fixed point in the passing area of the truck-trailer combination as the coordinate origin, and determine whether the relative positions of the cab and the container in the coordinate system meet the requirements.

[0063] For example, the geometric center of the container needs to be located in the area that extends 3 - 10 meters backward with the geometric center of the cab as the reference and has a lateral offset of no more than 2 meters.

[0064] S52. Calculate the movement trajectories of the cab and the container. If the average deviation between the movement trajectory of the cab and that of the container is less than a preset distance, and the included angle of the movement directions is less than a preset angle, then it is determined that the movement trajectories of the cab and the container are consistent.

[0065] For example, in 10 consecutive frame detections, calculate the movement trajectories of the cab and the container. If the average deviation of their movement trajectories is less than 0.5 meters, and the included angle of the movement directions is less than 30°, then it is determined that the movement trajectories are consistent.

[0066] Further, real-time optimization is achieved through hardware acceleration. For example, through TensorRT quantization, the inference speeds of YOLOv5 and PaddleOCR are increased to 100 FPS (NVIDIA RTX3060 platform).

[0067] The method for identifying the roof number and container number of a truck-trailer combination based on vision provided by the present invention has the following technical effects: High accuracy: By combining YOLOv5 object detection and PaddleOCR optical character recognition technology, the recognition accuracy of the vehicle head and container number can reach over 98%, far exceeding the 80% accuracy of traditional technologies. Fast recognition: It can complete recognition in a short time, with an average recognition time of only 0.5 seconds, shortening by 1.5 seconds compared to traditional methods, greatly improving work efficiency. In actual application scenarios, through testing different models of container trucks and containers, the average recognition time of traditional methods is 2 seconds, and this invention shortens it to 0.5 seconds. Improved recognition accuracy: On the standard test set, the recognition accuracy of the roof number reaches 99.2%, and the recognition accuracy of the container number reaches 98.5% (a 28% improvement compared to existing technologies); Breakthrough in real-time performance: The single-frame processing time ≤ 200ms, supporting real-time recognition at a vehicle speed ≤ 30km / h; Adaptation to complex scenarios: In working conditions such as rain and fog (visibility < 50m), strong backlight (illumination difference > 3EV), etc., the system robustness is improved by 40%. Strong adaptability: It can adapt to different lighting, weather, and complex background environments, and still maintain a high recognition rate under harsh conditions such as low light and rainy days, broadening the application scenarios. In low light (illuminance below 50lux) and rainy day environments, the recognition accuracy of this invention can still reach over 90%, while the recognition accuracy of traditional technologies will drop below 50%.

[0068] Social effects: Improve the safety of logistics transportation: Traditional recognition methods rely on manual labor, are prone to recognition errors, may lead to incorrect or delayed cargo transportation, and even cause safety accidents. The high-precision recognition of this invention effectively reduces such risks and ensures the accurate and safe transportation of goods. Enhance the intelligent level of the logistics industry: Promote the development of the logistics industry towards intelligence and automation, reduce manual intervention, optimize logistics processes, improve logistics efficiency, and help build an efficient and intelligent logistics system.

[0069] Economic effects: Reducing labor costs: In the past, manual identification required a large amount of labor. The automated identification of the present invention reduces the labor input. Taking a medium-sized logistics park as an example, it can save about 2 million yuan in labor costs annually. Suppose originally 50 employees were needed for manual identification, with an annual comprehensive cost (salary, welfare, etc.) of 80,000 yuan per person, totaling 4 million yuan; after adopting the present invention, only 10 employees are needed for system maintenance and exception handling, with an annual comprehensive cost (salary, welfare, etc.) of 200,000 yuan per person. The labor cost is reduced to 10×20 = 2 million yuan, saving 2 million yuan annually. Improving logistics efficiency and increasing economic benefits: Fast and accurate identification significantly shortens the residence time of container trucks in ports and logistics parks, improving transportation efficiency. Suppose a port processes 100,000 container truck trips annually. After adopting the present invention, each trip saves an average of 2 hours, and it can increase economic benefits by about 4 million yuan annually. Calculated based on a profit of 200 yuan generated per container truck trip, saving 2 hours per trip allows for more goods to be transported. The number of additional trips per year is 100,000×2÷24≈8,333 trips (calculated based on 24-hour operation per day), increasing the profit by about 8,333×200 = 1.6666 million yuan; at the same time, the saved time can be used for equipment maintenance, etc., reducing equipment wear and maintenance costs, saving about 2.3334 million yuan annually, totaling about 4 million yuan.

[0070] According to the method described in the above embodiments, this embodiment will further describe from the perspective of a vision-based identification device for container truck roof numbers and container numbers. The vision-based identification device for container truck roof numbers and container numbers can be specifically implemented as an independent entity, or integrated in an electronic device, which can be a device such as a terminal, a server, etc. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.

[0071] Please refer to Figure 2 , Figure 2 Specifically describes the vision-based identification device for container truck roof numbers and container numbers provided in the embodiments of the present application, which is applied to an electronic device. The vision-based identification device for container truck roof numbers and container numbers can include: An acquisition and preprocessing module, configured to acquire a roof image and a container image, and preprocess the roof image and the container image; A target detection module, configured to input the preprocessed roof image and container image into a target detection network to obtain a vehicle head target area and a container target area; A character region fine - processing module, which is used to perform fine - processing on the vehicle - head target region and the container target region to obtain the processed vehicle - head target region and container target region; A character recognition module, which is used to perform optical character recognition on the processed vehicle - head target region and container target region to obtain the container number and the roof number.

[0072] In specific implementation, each of the above - mentioned modules and / or units can be implemented as an independent entity, or can be combined arbitrarily and implemented as the same or several entities. For the specific implementation of each of the above - mentioned modules and / or units, please refer to the previous method embodiments. For the specific beneficial effects that can be achieved, please also refer to the beneficial effects in the previous method embodiments and will not be elaborated here.

[0073] In addition, the embodiment of the present application also provides an electronic device, which can be a device such as a computer or a tablet computer. The electronic device can implement the steps in any of the embodiments of the method for identifying the container number and the roof number of a container truck based on vision provided by the embodiment of the present application. Therefore, it can achieve the beneficial effects that can be achieved by any of the methods for identifying the container number and the roof number of a container truck based on vision provided by the embodiment of the present invention. For details, please refer to the previous embodiments and will not be elaborated here.

[0074] Figure 3 The specific structural block diagram of the electronic device provided by the embodiment of the present invention is shown. The electronic device can be used to implement the method for identifying the container number and the roof number of a container truck based on vision provided in the above - mentioned embodiment. The electronic device 500 can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a micro - processing box, or other devices, etc.

[0075] The RF circuit 510 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 510 may include various existing circuit elements for performing these functions. For example, an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, and so on. The RF circuit 510 can communicate with various networks such as the Internet, an enterprise intranet, a wireless network, or communicate with other devices through a wireless network. The above-mentioned wireless network may include a cellular phone network, a wireless local area network, or a metropolitan area network. The above-mentioned wireless network can use various communication standards, protocols, and technologies, including but not limited to the Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messages, and any other suitable communication protocols, and may even include those protocols that have not been developed yet.

[0076] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, to implement functions such as taking pictures with the front camera, processing the captured images, and switching the display colors of the display content on the display screen. The memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 520 may further include a memory remotely disposed relative to the processor 580, and these remote memories can be connected to the electronic device 500 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.

[0077] The input unit 530 can be used to receive input digital or character information, and generate a keyboard and a mouse related to user settings and function controls. The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, and these graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode).

[0078] The audio circuit 560, speaker 561, and microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can transmit the electrical signal converted from the received audio data to the speaker 561, and the speaker 561 converts it into a sound signal for output; on the other hand, the microphone 562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 560 and then converted into audio data. After the audio data is output to the processor 580 for processing, it is sent to another terminal, for example, through the RF circuit 510, or the audio data is output to the memory 520 for further processing. The audio circuit 560 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device 500.

[0079] The electronic device 500 can help the user receive requests, send information, etc. through the transmission module 570 (such as a Wi-Fi module), and it provides the user with wireless broadband Internet access. Although the transmission module 570 is shown in the figure, it can be understood that it does not belong to the essential components of the electronic device 500 and can be omitted completely within the scope of not changing the essence of the invention according to needs.

[0080] The processor 580 is the control center of the electronic device 500, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 520, and by calling the data stored in the memory 520, it performs various functions of the electronic device 500 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 580 either.

[0081] The electronic device 500 also includes a power source 590 (such as a battery) for supplying power to each component. In some embodiments, the power source can be logically connected to the processor 580 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power source 590 may also include any components such as one or more DC or AC power sources, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0082] Although not shown, the electronic device 500 also includes a camera (such as a front camera and a rear camera), a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: Obtain an image of the vehicle roof and an image of the container, and preprocess the image of the vehicle roof and the image of the container; Input the preprocessed image of the vehicle roof and the image of the container into a target detection network to obtain a vehicle head target area and a container target area; Perform character area refinement processing on the vehicle head target area and the container target area to obtain a processed vehicle head target area and a container target area; Perform optical character recognition on the processed vehicle head target area and the container target area to obtain the container number and the vehicle roof number.

[0083] In specific implementation, the above-mentioned various modules can be implemented as independent entities, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above-mentioned various modules, reference can be made to the foregoing method embodiments, which will not be elaborated here.

[0084] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. To this end, an embodiment of the present invention provides a storage medium in which multiple instructions are stored. The instructions can be loaded by a processor to execute the steps of any one of the embodiments of the method for identifying the container number and the truck roof number based on vision provided by the embodiments of the present invention.

[0085] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0086] Since the instructions stored in the storage medium can execute the steps of any one of the embodiments of the method for identifying the container number and the truck roof number based on vision provided by the embodiments of the present invention, the beneficial effects achievable by any of the methods for identifying the container number and the truck roof number based on vision provided by the embodiments of the present invention can be realized. For details, see the previous embodiments and will not be repeated here.

[0087] The above has introduced in detail a method, device, storage medium, and electronic device for identifying the container number and the truck roof number based on vision provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A vision-based recognition method for the roof number of a container truck and the container number, characterized in that The method includes: Obtain a roof image and a container image, and preprocess the roof image and the container image; Input the preprocessed roof image and container image into a target detection network to obtain a vehicle head target region and a container target region; Perform character region fine processing on the vehicle head target region and the container target region to obtain the processed vehicle head target region and container target region; Perform optical character recognition on the processed vehicle head target region and container target region to obtain the container number and the roof number.

2. The method for identifying the container top number and the container number based on vision according to claim 1, wherein Preprocessing the roof image and the container image includes: Perform dynamic range expansion on the roof image and the container image; Perform defogging processing on the roof image and container image after dynamic range expansion; Perform ROI extraction on the roof image and container image after defogging processing to divide the dynamic monitoring area.

3. The vision-based recognition method for the container yard crane roof number and container number according to claim 1, characterized in that The target detection network includes an improved backbone network, an improved neck network, and a head network; The improved backbone network includes a Focus module, a CSP module, and an improved convolutional layer; The Focus module is used to perform slicing operations on the input image to obtain sliced images; The CSP module is used to divide the sliced image into two parts, one part passes through the improved convolutional layer, and the other part skips the improved convolutional layer, and splices the features of the two parts; The improved convolutional layer is a Ghost module, and the Ghost module is used to perform cheap operations on the input image; The improved neck network is implemented by embedding a CBAM attention module.

4. The method for identifying the container top number and the container number based on vision according to claim 1, wherein Performing character region fine processing on the vehicle head target region and the container target region to obtain the processed vehicle head target region and container target region includes: Perform perspective correction on the vehicle head target region and the container target region to obtain the corrected vehicle head target region and container target region; Perform super-resolution reconstruction on the corrected vehicle head target region and container target region to obtain the reconstructed and corrected vehicle head target region and container target region.

5. The method for identifying the container yard crane roof number and the container number based on vision according to claim 1, wherein Performing perspective correction on the vehicle head target region and the container target region to obtain the corrected vehicle head target region and container target region includes: Extract the corner points of the target region, calculate the homography matrix according to the corner points, perform image transformation based on the homography matrix, and correct the target region with an inclination greater than the preset angle to a horizontal state; Performing super-resolution reconstruction on the corrected vehicle head target region and container target region to obtain the reconstructed and corrected vehicle head target region and container target region includes: Judge whether the character height of the corrected vehicle head target region and container target region is less than the preset pixel value; If it is less, input the corresponding target region into the super-resolution model for super-resolution reconstruction.

6. The method for identifying the container yard crane roof number and the container number based on vision according to claim 1, wherein Performing optical character recognition on the processed vehicle head target region and container target region to obtain the container number and the roof number includes: Input the processed vehicle head target region and container target region into a two-stage recognition network to obtain the container number and the roof number; Among them, the two-stage recognition network includes a document detection module and a text recognition module; The text detection module is used for text segmentation of the processed front truck target area and the container target area; The text recognition module is used for detecting and recognizing the segmented text.

7. The method for identifying the container top number and the container number based on vision according to claim 1, characterized in that After the step of performing optical character recognition on the processed front truck target area and the container target area to obtain the container number and the roof number, it includes: Verifying the container number and predicting the missing characters of the roof number based on the LSTM model.

8. The method for identifying the container top number and the container number based on vision according to claim 1, wherein After the step of performing optical character recognition on the processed front truck target area and the container target area to obtain the container number and the roof number, it further includes: Recognizing the entire truck, regarding the front truck and the container as components of the whole truck, and determining whether the container and the front truck belong to the same truck through spatial position constraints and motion trajectory consistency.

9. The method for identifying the container yard crane roof number and the container number based on vision according to claim 1, wherein, The determining whether the container and the front truck belong to the same truck through spatial position constraints and motion trajectory consistency includes: Establishing a two-dimensional plane coordinate system with a fixed point in the truck passing area as the coordinate origin, and judging whether the relative positions of the front truck and the container in the coordinate system meet the requirements; Calculating the motion trajectories of the front truck and the container. If the average deviation between the motion trajectory of the front truck and the motion trajectory of the container is less than a preset distance, and the included angle of the motion directions is less than a preset angle, then it is determined that the motion trajectories of the front truck and the container are consistent.

10. A vision-based recognition device for the roof number of a container truck and the container number, characterized in that, It includes: An acquisition and preprocessing module, which is used for acquiring the roof image and the container image, and preprocessing the roof image and the container image; A target detection module, which is used for inputting the preprocessed roof image and container image into the target detection network to obtain the front truck target area and the container target area; A character region refinement processing module, which is used for performing character region refinement processing on the front truck target area and the container target area to obtain the processed front truck target area and container target area; A character recognition module, which is used for performing optical character recognition on the processed front truck target area and container target area to obtain the container number and the roof number.

Citation Information

Patent Citations

  • A container number recognition method based on large-angle perspective deformation

    CN109190625A

  • Container wharf container number and vehicle number identification method and system

    CN110378332A

  • To-be-recognized image correction method and device

    CN112686959A

  • License plate recognition method and device, electronic equipment and storage medium

    CN112990197A

  • New energy vehicle abnormal parking big data detection method for smart city

    CN114332513A