A Container Number Detection and Recognition Method Based on Semantic Segmentation and Transformer

By adopting semantic segmentation and Transformer methods in container number detection, combined with SwinTransformer and multi-head attention mechanism, the limitations of detection accuracy and speed in traditional methods are solved, and efficient and accurate container number identification is achieved.

CN115565165BActive Publication Date: 2025-06-13FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211103809.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-09
Publication Date
2025-06-13
Estimated Expiration
2042-09-09

AI Technical Summary

Technical Problem

Due to the complex background and noise of the picture, traditional container number detection methods have limitations in detection accuracy and speed, making it difficult to achieve efficient and accurate identification.

Method used

The container number detection and identification method based on semantic segmentation and Transformer is adopted, and the detection network is designed by SwinTransformer, combining semantic segmentation and multi-head attention mechanisms to achieve efficient segmentation and classification of container number areas, and the recognition accuracy is improved through TPS text correction technology.

Benefits of technology

It realizes efficient and accurate segmentation and identification of container box number areas, can effectively solve the problems of blurred characters, damaged and hyphenated characters in box number areas, and improves detection speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115565165B_ABST
    Figure CN115565165B_ABST
Patent Text Reader

Abstract

The present invention provides a method for detecting and recognizing container numbers based on semantic segmentation and Transformer, comprising the following steps: Step S1: Construction of a dataset for container number detection and recognition; Step S2: Construction and training of a box number detection network based on semantic segmentation; Step S3: Text correction of the detected box number area; Step S4: Construction and training of a container number recognition network; Step S5: Process of detecting and recognizing the container number using the trained detection and recognition networks. Applying this technical solution can efficiently and accurately segment the container number area in the image, classify three different types of box numbers, and obtain the corrected box number area image through TPS text correction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a container number detection and recognition method based on semantic segmentation and Transformer. Background Art

[0002] With the rapid development of international trade and social economy, the demand for the logistics and transportation industry in China is increasing day by day. As an important carrier for transporting goods, containers play a crucial role in the entire transportation system. To achieve the automation, informatization, and intelligence of large-scale container transportation and management, it is very necessary to design an efficient and accurate container number recognition system.

[0003] The traditional and commonly used container number detection methods include three types: edge detection-based, mathematical morphology-based, and maximally stable extremal regions (MSER)-based. These traditional methods mainly adopt manually designed features. After obtaining the image, the traditional method uses image grayscale conversion to obtain the grayscale image of the image; then uses histogram of oriented gradients and scale-invariant feature transform to obtain image features; then searches through sliding windows or methods based on image region similarity to generate the position information of the container number and the position information of the container number characters; in the recognition stage, methods such as decision trees, support vector machines, and template matching are used to recognize the detected container number region. Due to factors such as the complex background and noise of the image, the above traditional image processing methods inevitably have certain limitations in container number detection, and the detection speed is relatively low. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a container number detection and recognition method based on semantic segmentation and Transformer, which can efficiently and accurately segment the container number region in the image, classify three different types of container numbers, and obtain the corrected container number region image through TPS text correction.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A container number detection and recognition method based on semantic segmentation and Transformer, comprising the following steps:

[0006] Step S1: Construction of a container number detection and recognition data set;

[0007] Step S2: Construction and training of a container number detection network based on semantic segmentation;

[0008] Step S3: Text correction of the detected container number region;

[0009] Step S4: Construction and training of a container number recognition network;

[0010] Step S5: The process of detecting and recognizing the container number using the trained detection and recognition network.

[0011] In a preferred embodiment, step S1 is specifically as follows:

[0012] Step S11: Analyze the type of container number to be detected and recognized, and determine the pictures containing such information as training pictures;

[0013] Step S12: Collect the container data set at the actual port;

[0014] Step S13: Use the Labelme annotation software to annotate the container number area, and save the position information of the quadrilateral frame of the number area, classification information, and the container number annotation information in the form of a json file to obtain the initial container number data set;

[0015] Step S14: Perform data augmentation processing on the initial container number data set to obtain the container number data set.

[0016] In a preferred embodiment, the types of container numbers to be detected and recognized include horizontal, vertical, and double-row container numbers, where the horizontal and vertical container numbers are labeled with 4 coordinates, and the double-row ones are labeled with 6 coordinates.

[0017] In a preferred embodiment, step S2 is specifically as follows:

[0018] Step S21: Use the Swin Transformer as the backbone network to obtain the feature maps C1, C2, C3, and C4 with pixel sizes of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 respectively; Figure 1 / 4, 1 / 8, 1 / 16, 1 / 32 pixel sizes of the feature maps C1, C2, C3, C4;

[0019] Step S22: The feature maps C1, C2, C3, and C4 pass through the FPN structure to obtain the feature map F;

[0020] Step S23: The feature map F passes through the semantic segmentation network prediction head to obtain the feature maps for predicting three different types of container numbers;

[0021] Step S24: Use the three predicted maps obtained in step S23 and the designed loss function to calculate the loss value;

[0022] Step S25: Start training, and save the weight file after the training ends;

[0023] Step S26: Use the weight file saved in step S25 to detect the picture to be recognized, and obtain the container number coordinate data based on the semantic segmentation detection network.

[0024] In a preferred embodiment, the calculation of the loss value is specifically as follows:

[0025] Loss = lk , (k = 1, 2, 3)

[0026]

[0027] Among them, Loss represents the total loss, and Loss is the sum of the losses of 3 maps, l k (k = 1, 2, 3) represents the losses of predicting 3 maps of different box number types, and k represents the k-th map; y k represents the ground truth pixel segmentation map, f k (x) represents the predicted box number segmentation map, y ki represents the i-th pixel value of the ground truth heat map, f ki (x) represents the i-th pixel value of the predicted heat map, and n represents the total number of pixels; the losses of the 3 predicted box number segmentation heat maps are added together to obtain the total loss function of the detection network.

[0028] In a preferred embodiment, the method for correcting the box number area in step S3 is based on the thin plate spline interpolation method.

[0029] In a preferred embodiment, the training steps of the container number recognition network in step S4 are as follows:

[0030] Step S41: Build a network for container number recognition based on the multi-head attention mechanism;

[0031] Step S42: Perform preprocessing operations on the input container box number pictures and box number annotations;

[0032] Step S43: Design the loss function of the network, calculate the loss, and update the recognition network parameters according to the loss;

[0033] Step S44: Set the number of network training rounds, combine the box number annotations as supervision information, train the network until the end, and save the network framework and network parameters of the trained container box number recognition part.

[0034] In a preferred embodiment, in step S43, the container recognition network consists of 6 convolutional layers; 2 multi-head attention mechanism networks; 4 max pooling layers; and 3 fully connected layers.

[0035] In a preferred embodiment, in step S43, the calculation of the loss value is as follows:

[0036]

[0037] Among them, l rec represents the recognition loss value, w is the character sequence data after encoding the box number annotation, |w| is the number of characters in the training target sequence, here |w| = 11, w jis the j-th character of the character prediction sequence data, y j is the j-th character of the character true value sequence data.

[0038] Compared with the prior art, the present invention has the following beneficial effects: The present invention designs a container number detection network based on semantic segmentation with Swin Transformer as the backbone network. The detection network in this paper can efficiently and accurately segment the container number area in the image, classify three different types of container numbers, and obtain the corrected container number area image through TPS text correction. In the recognition stage of the container number, the network in this paper combines a convolutional neural network and a multi-head attention mechanism, and then designs a special container number sequence prediction network, effectively solving the problems of blurred, damaged, and connected characters in the container number area. Brief Description of the Drawings

[0039] Figure 1 is the structural flowchart of the preferred embodiment of the present invention.

[0040] Figure 2 is the effect diagram of collecting part of the data set in step S1 of the preferred embodiment of the present invention.

[0041] Figure 3 is an example of three different types of container numbers in step S13 of the preferred embodiment of the present invention.

[0042] Figure 4 is the container number detection network diagram in step S2 of the preferred example of the present invention.

[0043] Figure 5 is the TPS container number correction example diagram in step S3 of the preferred embodiment of the present invention.

[0044] Figure 6 is the container number recognition network diagram in step S4 of the preferred embodiment of the present invention.

[0045] Figure 7 is the CNN feature extraction network structure in step S4 of the preferred embodiment of the present invention.

[0046] Figure 8 is the overall container number detection and recognition example diagram in step S5 of the preferred embodiment of the present invention. Detailed Description of the Invention

[0047] The present invention will be further described below in conjunction with the drawings and embodiments.

[0048] It should be noted that the following detailed description is illustrative and is intended to provide further description of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0049] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0050] Please refer to Figure 1-8 , the present invention provides a container number detection and recognition method based on semantic segmentation and Transformer, including the following steps:

[0051] Step S1: Construction of a container number detection and recognition data set;

[0052] Step S2: Construction and training of a box number detection network based on semantic segmentation;

[0053] Step S3: Text correction of the detected box number area;

[0054] Step S4: Construction and training of a container number recognition network;

[0055] Step S5: The process of detecting and recognizing the container number using the trained detection and recognition networks.

[0056] In this embodiment, step S1 is specifically as follows:

[0057] Step S11: Analyze the types of container numbers to be detected and recognized, and determine the pictures containing such information as training pictures;

[0058] Step S12: As Figure 2 shown, collect the container data set on-site at the port;

[0059] Step S13: Use the LabeIme annotation software to annotate the container number area, and save the position information of the quadrilateral frame of the number area, classification information, and box number annotation information in the form of a json file to obtain the initial container number data set; as Figure 3 shown, the types of container numbers to be detected and recognized include horizontal, vertical, and double-row container numbers, where the horizontal and vertical box numbers are marked with 4 coordinates, and the double-row is marked with 6 coordinates.

[0060] Step S14: Perform data augmentation processing on the images of the initial container number data set, and perform dictionary encoding processing on the annotations of the data set to obtain the container number data set.

[0061] In this embodiment, step S2 is specifically as follows:

[0062] Step S21: AsFigure 4 As shown, using the Swin Transformer as the backbone network, feature maps C1, C2, C3, and C4 with pixel sizes of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 are obtained respectively. Figure 1 / 4, 1 / 8, 1 / 16, 1 / 32 pixel sizes of the feature maps C1, C2, C3, C4.

[0063] Step S22: The feature maps C1, C2, C3, and C4 pass through the FPN structure to obtain the feature map F;

[0064] Step S23: The feature map F passes through the semantic segmentation network prediction head to obtain the feature map for predicting three different box number types;

[0065] Step S24: Using the three predicted maps obtained in Step S23 and the designed loss function, calculate the loss value;

[0066] Step S25: Calculation of the loss value:

[0067] Loss=l k, (k = 1, 2, 3)

[0068]

[0069] Among them, Loss represents the total loss, Loss is the sum of the losses of the 3 maps, l k (k = 1, 2, 3) represents the loss of the 3 maps for predicting different box number types, k represents the k-th map; y k represents the ground truth pixel segmentation map, f k (x) represents the predicted box number segmentation map, y ki represents the i-th pixel value of the ground truth heat map, f ki (x) represents the i-th pixel value of the predicted heat map, n represents the total number of pixels; the losses of the 3 predicted box number segmentation heat maps are added together to obtain the total loss function of the detection network.

[0070] Step S26: Start training, and save the weight file after training;

[0071] Step S27: Use the weight file saved in Step S25 to detect the image to be recognized, and obtain the box number coordinate data based on the semantic segmentation detection network.

[0072] In this embodiment, Step S3 is specifically:

[0073] Step S31: As Figure 5 shown, in the input image, the accurately located quadrilateral frame is converted into evenly spaced input control points C'. Now, a rectangular area with appropriate length and width is set as the output image according to the detection frame coordinates, and K output control points are set;

[0074] Step S32: Linearly project the control point C of the output image to obtain the control point C' in the input image. The conversion relationship is shown in the following formula:

[0075]

[0076] Where:

[0077] φ(r) = r 2 log(r)

[0078] T = [C'O 2×3 ΔC -1

[0079] r = ||c i -c k ||, representing the Euclidean distance between the output control point c i and c k . The transformation matrix T is the target matrix to be solved; the calculation of △C is shown in the following formula:

[0080]

[0081] Where C + is a matrix composed of .

[0082] Step S33: For all pixel coordinates p' and p of the input image and the output image, the corresponding relationship also conforms to the TPS conversion relationship in Formula 6. The mapping relationship between p' and p is shown in the following formula:

[0083]

[0084] Step S34: Calculate the corresponding two-dimensional coordinate p' of the two-dimensional coordinate p in the input image in the predicted image according to the above coordinate conversion formula. The sampler calculates the value of the predicted image p through bilinear interpolation of the adjacent pixels of p', and finally obtains the corrected result image of the scene text.

[0085] In this embodiment, Step S4 is specifically as follows:

[0086] Step S41: As Figure 6 shown, build a network for container number recognition based on the multi-head attention mechanism;

[0087] Step S42: Perform preprocessing operations on the input container number picture and the container number annotation;

[0088] Step S43: Design the loss function of the network, calculate the loss and update the recognition network parameters according to the loss;

[0089] Step S44: Calculate the recognition loss function:

[0090]

[0091] Among them, l rec represents the recognition loss value, w is the character sequence data after encoding the box number annotation code, |w| is the number of characters in the training target sequence, where |w| = 11, and w j is the j-th character of the character prediction sequence data, and y j is the j-th character of the character true value sequence data.

[0092] Step S45: Set the number of network training rounds, combine the box number annotation as supervision information, train the network until it ends, and save the network framework and network parameters of the trained container box number recognition part.

[0093] In this embodiment, step S5 is specifically as follows:

[0094] Step S51: Concatenate the box number detection network and the box number recognition network together;

[0095] Step S52: Load the network weight data saved during training;

[0096] Step S53: Input the box number picture to be detected;

[0097] Step S54: As Figure 8 shown, obtain the recognition result of the container number in the picture;

[0098] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.

Claims

1. A method for detecting and recognizing container numbers based on semantic segmentation and Transformer, characterized in that, it includes the following steps: Step S1: Construction of a container number detection and recognition data set; Step S2: Construction and training of a box number detection network based on semantic segmentation; Step S3: Text correction of the detected box number area; Step S4: Construction and training of a container number recognition network; Step S5: The process of detecting and recognizing container numbers using the trained detection and recognition networks; The specific content of step S2 is as follows: Step S21: Use Swin Transformer as the backbone network to obtain feature maps C1, C2, C3, and C4 with pixel sizes of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image respectively; Step S22: Feature maps C1, C2, C3, and C4 pass through the FPN structure to obtain feature map F; Step S23: Feature map F passes through the prediction head of the semantic segmentation network to obtain feature maps for predicting three different box number types; Step S24: Use the three predicted maps obtained in step S23 and the designed loss function to calculate the loss value; Step S25: Start training, and save the weight file after training; Step S26: Use the weight file saved in step S25 to detect the image to be recognized, and obtain the box number coordinate data based on the semantic segmentation detection network; The training steps of the container number recognition network in step S4 are as follows: Step S41: Build a network for recognizing container numbers based on the multi-head attention mechanism; Step S42: Perform preprocessing operations on the input container number image and box number annotation; Step S43: Design the loss function of the network, calculate the loss, and update the recognition network parameters according to the loss; Step S44: Set the number of training rounds of the network, combine the box number annotation as the supervision information, train the network until the end, and save the network framework and network parameters of the trained container number recognition part; The specific content of step S5 is as follows: Step S51: Concatenate the box number detection network and the box number recognition network together; Step S52: Load the trained network weight data; Step S53: Input the box number image to be detected; Step S54: Obtain the recognition result of the container number in the image.

2. A method for detecting and recognizing container numbers based on semantic segmentation and Transformer according to claim 1, characterized in that, the specific content of step S1 is as follows: Step S11: Analyze the types of container numbers to be detected and recognized, and determine the images containing such information as training images; Step S12: Collect container data sets at the actual port; Step S13: Use the LabeIme annotation software to annotate the container number area, and save the quadrilateral frame position information, classification information, and box number annotation information of the number area in the form of a json file to obtain the initial container number data set; Step S14: Perform data augmentation processing on the initial container number data set to obtain the container number data set.

3. A method for detecting and recognizing container numbers based on semantic segmentation and Transformer according to claim 2, characterized in that, The types of container numbers to be detected and recognized include horizontal, vertical, and double-row container numbers, where the horizontal and vertical container numbers are marked with 4 coordinates, and the double-row ones are marked with 6 coordinates.

4. A method for detecting and recognizing container numbers based on semantic segmentation and Transformer according to claim 1, characterized in that, the calculation of the loss value is specifically as follows: Loss=l k ,(k=1,2,3) Among them, Loss represents the total loss, and Loss is the sum of the losses of 3 maps, l k , k = 1, 2, 3 represents the losses of predicting 3 different box number type maps, and k represents the k-th map; y k represents the ground truth pixel segmentation map, f k (x) represents the predicted box number segmentation map, y ki represents the i-th pixel value of the ground truth heat map, f ki (x) represents the i-th pixel value of the predicted heat map, and n represents the total number of pixel values; the loss functions of the 3 predicted box number segmentation heat maps are added together to obtain the total loss function of the detection network.

5. A method for detecting and recognizing container numbers based on semantic segmentation and Transformer according to claim 1, characterized in that, the method for correcting the container number area in step S3 is based on the thin plate spline interpolation method.

6. A method for detecting and recognizing container numbers based on semantic segmentation and Transformer according to claim 1, characterized in that, in the said step S43, the container recognition network consists of 6 convolutional layers; 2 multi-head attention mechanism networks; 4 max-pooling layers; and 3 fully connected layers.

7. A method for detecting and recognizing container numbers based on semantic segmentation and Transformer according to claim 1, characterized in that, the calculation of the loss value in the said step S43 is specifically as follows: Among them, l rec represents the recognition loss value, w is the character sequence data after encoding the box number annotation, |w| is the number of characters in the training target sequence, where |w| = 11, w j is the j-th character of the character prediction sequence data, y j is the j-th character of the character ground truth sequence data.