A two-stage container number recognition method

Through the two-stage container number identification method, combined with image preprocessing and complex network structure, the accuracy problem of container number identification under light and angle changes is solved, achieving higher recognition accuracy and stability.

CN119810840BActive Publication Date: 2025-07-08YANTAI PORT CONTAINER TERMINAL CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411868047.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-07-08
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

The prior art has limited generalization ability due to light, angle and background changes in container number recognition, and simple image enhancement methods, which affect the recognition accuracy.

Method used

The two-stage container box number recognition method is adopted, including image preprocessing, box number area detection and segmentation, and box number recognition algorithm. The image enhancement is used to use histogram equalization, affine transformation, and brightness adjustment, and feature extraction and segmentation are extracted and segmented by backbone network, neck network and detection head, and the channel attention and deep supervision module are used to improve the recognition accuracy.

Benefits of technology

It improves the accuracy and stability of container number identification, can adapt to complex scenarios in different environments, and enhances the feature extraction and recognition capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810840B_ABST
    Figure CN119810840B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of container number recognition, and particularly relates to a two-stage container number recognition method. The method includes: acquiring a container image, and performing image preprocessing on the container image; detecting and segmenting a container number area based on the preprocessed image; and recognizing the container number based on the recognized and segmented container number area. The present invention adopts a two-stage method, which can effectively improve the accuracy and stability of container number recognition, bringing a new breakthrough to the container number recognition technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of container number recognition, and particularly relates to a two-stage container number recognition method. Background Art

[0002] In the field of port logistics, container number recognition is a crucial technology, which is of great significance for improving logistics efficiency and reducing human errors. The Chinese invention patent with the publication number CN117809310A discloses a port container number recognition method and system based on machine learning, which mainly obtains container number information by directly recognizing container images. However, in practical applications, this method faces a series of challenges, resulting in the need to improve its recognition accuracy.

[0003] First of all, this patent uses a clustering method for the positioning of the container number area. Although the clustering method can achieve the preliminary positioning of the number area to a certain extent, its generalization ability is relatively limited, and it is difficult to adapt to container images under different lighting, angles, backgrounds and other conditions.

[0004] In addition, in terms of image enhancement, the method of this patent is relatively simple. It only adjusts the size of the image by upsampling and downsampling, without further optimizing the attributes of the image such as contrast, position, shape, brightness, etc. This simple image enhancement method limits the processing ability of the model for complex images and further affects the recognition accuracy.

[0005] In view of the above problems, the present invention proposes a two-stage container number recognition method. Summary of the Invention

[0006] In order to overcome the problems in the prior art, the present invention proposes a two-stage container number recognition method.

[0007] The technical solution of the present invention to solve the above technical problems is as follows:

[0008] The present invention provides a two-stage container number recognition method, including the following steps: Step 100: Obtain a container image and perform image preprocessing on the container image; Step 200: Detect and segment the number area based on the preprocessed image; Step 300: Recognize the number based on the recognized and segmented number area.

[0009] In the step 100, performing image preprocessing on the container image includes performing data enhancement on the collected container image, including histogram equalization, affine transformation, and brightness adjustment.

[0010] Further, in the step 200, a box number area detection algorithm is used to detect and locate the box number area, where the box number area recognition algorithm includes a backbone network, a neck network, and a detection head connected in sequence;

[0011] Input the preprocessed container image into the backbone network to extract key features; input the extracted key features into the neck network for feature fusion and optimization; input the fused features into the detection head for prediction and location of the box number area; segment and extract the box number area according to the prediction result.

[0012] Further, in the step 300, a box number recognition algorithm is used to recognize the segmented box number area, and the box number recognition algorithm includes a backbone network, a channel attention module, a neck network, a deep supervision module, and a detection head;

[0013] The backbone network, the channel attention module, the neck network, and the detection head are connected in sequence, and the channel attention module is connected after the neck network;

[0014] Input the segmented box number area into the backbone network, the backbone network outputs the features of the box number area, the channel attention module performs weighted processing on the features, and the neck network performs fusion and optimization on the weighted features; the deep supervision module performs upsampling and additional supervision on the output of the neck network; the detection head outputs the recognition result.

[0015] Further, the channel attention module includes a pooling layer and a fully connected layer; perform global pooling operation on the original feature map output by the backbone network to comprehensively capture the feature information of each channel; use the fully connected layer to deeply calculate the specific weights of each channel for the detection task; subsequently, multiply the calculated attention weights element-wise with the original feature map output by the backbone network to generate a weighted feature representation.

[0016] Further, the deep supervision module includes upsampling and a deep supervision detection head; in the deep supervision module, perform upsampling processing on the features after fusion and optimization by the neck network, and perform preliminary detection through the deep supervision detection head, and the label supervises and guides the prediction result of the deep supervision detection head.

[0017] Further, use MPDIoU as the loss function in the training stage.

[0018] Further, the backbone network includes a CBS module, a VCB module, an RCB module, and an RBF module;

[0019] The image features are extracted and saved using the CBS module and the VCB module; the RCB module helps to learn deeper feature representations by introducing residual connections; the RBF module is used to further improve the feature extraction ability of the model and the comprehensiveness of spatial information capture.

[0020] Further, the VCB module includes a variable convolution layer, a BN layer, and a SiLu activation function connected in sequence.

[0021] Further, the CBS module includes a convolution layer, a BN layer, and a SiLu activation function.

[0022] Compared with the prior art, the present invention has the following technical effects:

[0023] The present invention first performs image preprocessing on the container image, and then uses the box number area detection algorithm to detect and segment the box number area in the image. Among them, the backbone network is responsible for extracting the key features of the image, the neck network performs efficient fusion processing on these features, and the detection head predicts the specific position of the box number area based on the fused features. Finally, the box number recognition algorithm is used to accurately recognize the segmented box number area. Among them, after the backbone network extracts features, the channel attention enhancement model is used to improve the accuracy of box number recognition, and after the neck network fuses features, the deep supervision module is used to enhance the feature extraction ability of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is the flow chart of the present invention;

[0026] Figure 2 is the flow chart of the box number area detection algorithm of the present invention;

[0027] Figure 3 is the flow chart of the box number recognition algorithm of the present invention;

[0028] Figure 4 is the model diagram of the backbone network of the present invention;

[0029] Figure 5 is the model diagram of the neck network of the present invention;

[0030] Figure 6 is the model diagram of each module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes the implementation manner, structure, features and effects of the technical solution proposed according to the present invention in detail in conjunction with the accompanying drawings and preferred embodiments. Specific features, structures or characteristics in one or more embodiments may be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0032] The present invention discloses a two-stage container number recognition method. First, image enhancement such as histogram equalization, affine transformation, and brightness adjustment is performed on the container image; then, a license plate area detection algorithm is used to perform target detection on the enhanced image, and the container number area is recognized and segmented; finally, a license plate recognition algorithm is used to recognize the license plate number in the license plate area.

[0033] In one embodiment of the present invention, referring to Figures 1 - 6 , a two-stage container number recognition method is provided, including the following steps: Step 100: Obtain a container image and perform image preprocessing on the container image; Step 200: Based on the preprocessed image, detect and segment the license plate area; Step 300: Based on the recognized and segmented license plate area, recognize the license plate number.

[0034] The following details each of the above steps:

[0035] Step 100: Obtain a container image and perform image preprocessing on the container image.

[0036] Due to the influence of various factors such as the shooting environment, lighting conditions, and shooting angle, the obtained container images often have problems such as insufficient contrast, position offset, inconsistent size, shape distortion, and uneven brightness. These problems not only increase the difficulty of image recognition but also may lead to a decline in the performance of the subsequent recognition model in actual applications.

[0037] To improve the recognition ability of the model for container images in multiple environments, this embodiment adopts image data enhancement technologies such as histogram equalization, affine transformation, and brightness adjustment to adjust the attributes of the container image such as contrast, position, size, shape, and brightness, thereby achieving effective enhancement of the image.

[0038] Histogram equalization is a commonly used image contrast enhancement technology. By adjusting the gray histogram of the image, it makes the gray value distribution of the image more uniform, thereby enhancing the contrast of the image. In container image recognition, histogram equalization can effectively improve the problem of insufficient image contrast caused by uneven lighting or differences in shooting equipment, making the license plate area clearer and facilitating subsequent feature extraction and recognition.

[0039] Affine transformation is a two-dimensional coordinate transformation, including operations such as translation, rotation, scaling, and shearing. In container image recognition, affine transformation can be used to adjust the position, size, and shape of the image. By simulating container images at different shooting angles and distances, affine transformation can increase the diversity of training data, enabling the model to better adapt to complex scenarios in practical applications. At the same time, affine transformation can also be used to correct image distortion problems caused by shooting angles, improving the accuracy of recognition.

[0040] Brightness adjustment is a simple image enhancement technique that changes the brightness of an image by adjusting the brightness value. In container image recognition, brightness adjustment can be used to address image recognition problems under different lighting conditions. By increasing or decreasing the brightness value of the image, the box number area can maintain a clear visual effect under different lighting conditions, thereby improving the stability and accuracy of recognition.

[0041] Step 200: Based on the preprocessed image, detect and segment the box number area.

[0042] Use the box number area detection algorithm to detect and locate the box number area. The box number area recognition algorithm includes a backbone network, a neck network, and a detection head. The backbone network is used to extract key features of the preprocessed image. The neck network performs efficient fusion processing on the key features. The detection head is used to predict the specific position of the box number area based on the fused features.

[0043] Refer to Figure 2 , input the preprocessed container image into the backbone network to extract key features; input the extracted key features into the neck network for feature fusion and optimization; input the fused features into the detection head for prediction and positioning of the box number area; according to the prediction results, segment and extract the box number area for subsequent character recognition and other processing.

[0044] Refer to Figure 4 , the backbone network includes CBS modules, VCB modules, RCB modules, and RBF modules. First, use multiple CBS modules and VCB modules to extract and save image features; subsequently, through the RCB module with a residual structure, ensure that the model can stably learn during training and avoid gradient explosion; finally, use the RBF module to further improve the model's feature extraction ability and the comprehensiveness of spatial information capture.

[0045] Refer to Figure 6, the CBS module includes a convolutional layer Conv, a BN layer, and a SiLu activation function, aiming to efficiently extract features from the preprocessed container images as input. Among them, the SiLU activation function has smooth non-linear characteristics, which can provide different output responses in different input ranges, thereby improving the expressive ability of the model and making the network easier to converge during the training process.

[0046] The VCB module includes a deformable convolutional layer, a BN layer, and a SiLu activation function. Among them, deformable convolution enables the sampling positions of standard convolution to be flexibly adjusted by introducing two-dimensional offsets. This local, dense, and adaptive deformation mechanism fully depends on the input features and can better adapt to the spatial changes of the input data at the shallow stage of the network, thus greatly improving the accuracy and robustness of feature extraction. Therefore, the VCB module can achieve adaptive feature extraction, normalization processing, and non-linear transformation of the input data.

[0047] The RCB module adopts a residual structure, effectively avoiding the problem of gradient explosion that may occur during model training. The RBF module, through a multi-branch structure composed of convolutional layers with different sizes of convolutional kernels and dilated convolutional layers with different dilation rates, realizes the weighted fusion of global and local features, enhancing the feature extraction ability of the model and the comprehensiveness of spatial information capture.

[0048] Specifically, the backbone network includes a first CBS module, a first VCB module, a second CBS module, a first RCB module, a second VCB module, a second RCB module, a third CBS module, an RBF module, a fourth RCB module, and a fourth CBS module connected in sequence. The first CBS module is used to extract initial features from the input container images. The first VCB module performs adaptive feature extraction and transformation on the initial features. The second CBS module is used to further extract features. The first RCB module adopts a residual structure to avoid gradient explosion. The second VCB module performs adaptive feature extraction again. The second RCB module continues to adopt a residual structure for stable learning. The third CBS module extracts deeper features. The RBF module realizes the weighted fusion of global and local features. The fourth RCB module further stabilizes learning. The fourth CBS module outputs the finally extracted features for use by the neck network and the detection head. Among them, a first feature map Map1 is output after the first RCB module, a second feature map Map2 is output after the second RCB module, and a third feature map Map3 is output after the fourth CBS module.

[0049] Refer to Figure 5 , the neck network fuses different-level features generated by the backbone network through RCB modules and CBS modules, enhancing the diversity and robustness of the features. At the same time, dimensionality reduction and dimensionality increase operations are also performed on the features to meet the detection requirements of the subsequent detection head.

[0050] The neck network uses methods such as cascading and upsampling to compensate for the low-level feature information that may be lost due to convolution operations in the backbone network. Specifically, the neck network adjusts the deep feature information to the same size as the low-level feature information through upsampling technology and performs feature fusion. This process not only fuses multi-level features extracted from the backbone network but also enhances the feature representation, significantly improving the accuracy and robustness of the network for object detection.

[0051] Specifically, the third feature map Map3 first adjusts its resolution through upsampling technology to match the resolution of the second feature map Map2; the fused feature map Fusion_Map2_3 is used as the input for the subsequent module; Fusion_Map2_3 passes through the fifth RCB module and the fifth CBS module in sequence, and these two modules are used to enhance the feature representation and further extract features respectively, and the processed feature map is called Processed_Map_After_5RCB_5CBS; the Processed_Map_After_5RCB_5CBS adjusts its resolution through upsampling technology to match the resolution of the first feature map Map1, and the fused feature map Fusion_Map1_Processed is used as the input for the sixth RCB module; after being processed by the sixth RCB module, Fusion_Map1_Processed outputs the feature map F1.

[0052] The feature map output by the sixth RCB module is processed by the sixth CBS module and then undergoes feature fusion with Fusion_Map1_Processed, and the fused feature map is used as the input for the seventh RCB module; after being processed by the seventh RCB module, the fused feature map outputs the feature map F2;

[0053] The feature map output by the seventh RCB module is processed by the seventh CBS module and then undergoes feature fusion with the original third feature map Map3, and the fused feature map is used as the input for the eighth RCB module. After being processed by the eighth RCB module, the fused feature map finally outputs the feature map F3.

[0054] Refer to Figure 6 , the detection head part consists of multiple detection layers, and each detection layer consists of a series of convolutional layers, normalization layers, and activation function layers. These layers further process and transform the input feature map to extract the key information for object detection. Finally, the feature map is mapped to the output space of object detection through the detection head, and information such as the position coordinates of the box number, class prediction probability, and confidence is output.

[0055] MPDIoU is used as the loss function L during the training phase MPDIoU , and the formula is as follows:

[0056]

[0057] Among them, IoU represents the traditional intersection over union, d1 represents the distance between the upper left corner of the predicted bounding box and the actual annotated bounding box, d2 represents the distance between the lower right corner of the predicted bounding box and the actual annotated bounding box, w represents the width of the bounding box, and h represents the height of the bounding box. By introducing these additional distance metrics, MPDIoU can more accurately measure the similarity between bounding boxes and improve the accuracy of the model.

[0058] Step 300: Based on the identified and segmented box number region, identify the box number.

[0059] Use a box number recognition algorithm to accurately identify the segmented box number region. The box number recognition algorithm includes a backbone network, a channel attention module, a neck network, a deep supervision module, and a detection head. A channel attention module is added between the backbone network and the neck network to improve the accuracy of box number recognition; after the neck network fuses features, a deep supervision module is used to enhance the network's feature extraction ability.

[0060] The backbone network is used to extract key features, and a channel attention module is added between the backbone network and the neck network. The channel attention module first performs global average pooling operations to comprehensively capture the feature information of each channel; then, a fully connected layer is used to deeply calculate the specific weights of each channel for the detection task; subsequently, these calculated attention weights are multiplied element by element with the original feature map output by the backbone network to generate a weighted feature representation; this process realizes the dynamic adjustment of the weights of each channel, enabling the detection network to focus more precisely on the box number region, thereby significantly improving the accuracy of box number detection.

[0061] The neck network is used to further process and fuse the features processed by the channel attention module. The neck network fuses different-level features generated by the backbone network through RCB modules and CBS modules to enhance the diversity and robustness of the features. At the same time, dimensionality reduction and dimensionality increase operations are also performed on the features to meet the requirements of the subsequent deep supervision module.

[0062] Refer to Figure 6 , the deep supervision module includes a transposed convolution, a BN layer, an activation function layer, and a deep supervision detection head. In the deep supervision module, operations such as upsampling are performed on the features optimized by the neck network fusion, and preliminary detection is carried out through the deep supervision detection head. The label directly supervises and guides the prediction results of the deep supervision detection head, realizing the direct supervision of the label on the deep features. The label is a model-guided label, which is the position of the box number annotated by professionals using the labelme software. This enhances the ability of the backbone network in feature extraction, thereby improving the accuracy and efficiency of the entire box number recognition algorithm.

[0063] The detection head is the output part of the box number recognition algorithm. The detection head consists of multiple detection layers, and each detection layer is composed of a 3x3 convolutional layer, a normalization layer, and an activation function layer. These layers further fuse the input feature maps to extract the key information for object detection. Finally, the feature maps are mapped to the output space of object detection through a 1x1 convolution, and information such as the position coordinates of the box number, the class prediction probability, and the confidence level are output. The MPDIoU is used as the loss function in the training stage.

[0064] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A two-stage container number recognition method, characterized in that, It includes the following steps: Step 100: Obtain a container image and perform image preprocessing on the container image; Step 200: Detect and segment the container number area based on the preprocessed image; Step 300: Identify the container number based on the identified and segmented container number area; In step 200, a container number area detection algorithm is used to detect and locate the container number area. Among them, the container number area detection algorithm includes a backbone network, a neck network, and a detection head connected in sequence; In step 300, a container number recognition algorithm is used to recognize the segmented container number area. The container number recognition algorithm includes a backbone network, a channel attention module, a neck network, a deep supervision module, and a detection head; the backbone network, the channel attention module, the neck network, and the detection head are connected in sequence, and the channel attention module is connected after the neck network; The backbone network includes a CBS module, a VCB module, an RCB module, and an RBF module; the CBS module and the VCB module are used to extract and save image features; the RCB module helps to learn deeper feature representations by introducing residual connections; the RBF module is used to further improve the feature extraction ability of the model and the comprehensiveness of spatial information capture; among them, the VCB module includes a variable convolutional layer, a BN layer, and a SiLu activation function connected in sequence; the CBS module includes a convolutional layer, a BN layer, and a SiLu activation function; Use MPDIoU as the loss function during the training phase , and the formula is as follows: ; Among them, IoU represents the traditional intersection over union, d 1 represents the distance between the upper left corner of the predicted bounding box and the actual annotation bounding box, d 2 represents the distance between the lower right corner of the predicted bounding box and the actual annotation bounding box, w represents the width of the bounding box, h represents the height of the bounding box.

2. The two-stage container number recognition method according to claim 1, characterized in that, In step 100, image preprocessing of the container image includes data augmentation of the collected container image, including histogram equalization, affine transformation, and brightness adjustment.

3. The two-stage container number recognition method according to claim 1, wherein Input the preprocessed container image into the backbone network in step 200 to extract key features; input the extracted key features into the neck network for feature fusion and optimization; input the fused features into the detection head for prediction and location of the container number area; segment and extract the container number area according to the prediction result.

4. The two-stage container number recognition method according to claim 1, characterized in that, Input the segmented container number area into the backbone network in step 300. The backbone network in step 300 outputs the features of the container number area. The channel attention module performs weighted processing on the features. The neck network fuses and optimizes the weighted features; the deep supervision module performs upsampling and additional supervision on the output of the neck network; the detection head outputs the recognition result.

5. A two-stage container number recognition method according to claim 4, characterized in that, The channel attention module includes a pooling layer and a fully connected layer; perform global pooling on the original feature map output by the backbone network to comprehensively capture the feature information of each channel; use the fully connected layer to deeply calculate the specific weight of each channel for the detection task; Subsequently, multiply the calculated attention weights element by element with the original feature map output by the backbone network to generate a weighted feature representation.

6. The two-stage container number recognition method according to claim 5, wherein, The deep supervision module includes upsampling and a deep supervision detection head; in the deep supervision module, perform upsampling on the features after fusion and optimization by the neck network, and perform preliminary detection through the deep supervision detection head. The label supervises and guides the prediction result of the deep supervision detection head.

Citation Information

Patent Citations

  • Port container number identification method and system based on machine learning

    CN117809310A

  • Container weak and small serial number target detection and identification method based on deep learning

    CN117253154A

  • Feature correlation-based osteosarcoma CT image lesion area detection method and system

    CN118587217A