A port container number recognition method based on image enhancement

By combining Retinex decomposition and adversarial enhancement networks with convolutional neural networks, the problem of low accuracy in container image acquisition and recognition under low-light conditions at night is solved, achieving efficient and low-cost container number recognition and supporting intelligent cargo handling operations in ports.

CN115830585BActive Publication Date: 2026-04-21ZHEJIANG OCEAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG OCEAN UNIV
Filing Date
2022-12-05
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In low-light conditions at night, container image acquisition and recognition suffer from a high error rate, resulting in low efficiency and high cost of intelligent container tallying operations at ports. Existing technologies are unable to effectively improve recognition accuracy.

Method used

An adversarial enhancement network based on Retinex decomposition is used to process low-light container images. Combined with a convolutional neural network model, the container number characters are extracted and identified through image enhancement and preprocessing techniques.

Benefits of technology

It significantly improved the accuracy of container number recognition under low-light conditions at night, reduced the error rate, optimized energy consumption, and achieved efficient intelligent container handling operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115830585B_ABST
    Figure CN115830585B_ABST
Patent Text Reader

Abstract

This invention discloses a port container number recognition method based on image enhancement, comprising the following steps: collecting container images under low-light conditions at a cargo port to form an image library; constructing a Retinex decomposition-based adversarial enhancement network to enhance the image library and improve the enhancement effect; preprocessing the container numbers in the enhanced container images to extract characters and construct a container number character dataset, and iteratively training the constructed convolutional neural network model on the character dataset to improve the accuracy of character recognition; loading the actually collected low-light container images to be recognized into the trained image enhancement network to enhance the low-light images and obtain high-quality images; preprocessing the enhanced high-quality images to extract the container character images to be recognized, and inputting the obtained character images into the trained convolutional neural network model for container number recognition to obtain the final result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of container identification technology, specifically relating to a port container number identification method based on image enhancement. Background Technology

[0002] The ever-increasing number of containers has also spurred the gradual expansion of port operations. To avoid the heavy pressure caused by a large backlog of containers, 24-hour uninterrupted container tallying operations are necessary at various ports. Therefore, a modern intelligent port container tallying system is indispensable, and a crucial foundational task is the intelligent identification and collection of container number information.

[0003] However, for most small ports and general port areas in China, container tallying is still done manually, requiring multiple tally clerks to collect container number information on-site. This method has a high probability of errors during recording and subsequent verification. Furthermore, while failing to effectively improve work efficiency, it significantly increases operating costs, especially when performing container number identification and data collection in low-light conditions at night. To address areas with insufficient lighting or shadows, most ports use multiple high-wattage industrial searchlights to illuminate the data collection work and mitigate errors caused by low light levels. However, this method also increases energy consumption. Therefore, the problem of unclear image details and impaired image recognition due to low light at night is a current technical challenge that needs to be addressed.

[0004] This application primarily focuses on image enhancement processing for low-quality container images acquired under low-light conditions at night. By adjusting the overall color tone and corresponding color saturation of the image through algorithms, the dark areas of the image are brightened, while prominent areas are suppressed to a certain extent. Ultimately, this effectively improves the contrast and enhances the detailed features of the container number character area, enabling the container number information to be clearly presented for subsequent positioning and identification. This improves the accuracy of container number character recognition and provides a guarantee for intelligent container tallying operations in ports. Summary of the Invention

[0005] This invention provides a port container number recognition method based on image enhancement to solve technical problems such as errors in the acquisition, recognition, and positioning of container images in low-light conditions at night.

[0006] The specific technical solution of this application is as follows:

[0007] A port container number recognition method based on image enhancement includes the following steps:

[0008] Images of containers in low-light environments are collected at cargo ports to form an image library;

[0009] An adversarial enhancement network based on Retinex decomposition is constructed to perform image enhancement processing on an image database to improve the enhancement effect;

[0010] By preprocessing the container numbers in the enhanced container images, extracting characters to construct a container number character dataset, and then iteratively training a convolutional neural network model on the character dataset, the accuracy of character recognition is improved.

[0011] The actual acquired low-light container images to be identified are loaded into a trained image enhancement network to enhance the low-light images and obtain high-quality images.

[0012] The enhanced high-quality image is preprocessed to extract the character images of the boxes to be identified. The obtained character images are then input into a trained convolutional neural network model for box number recognition to obtain the final result.

[0013] Furthermore, the construction of the adversarial enhancement network for Retinex decomposition includes the following steps:

[0014] S1. The original low-light image is effectively decomposed in the decomposition network, and the ability to restore details during the decomposition process is guaranteed by multiple decomposition losses.

[0015] S2. Use a fusion-enhanced network to perform reinforcement learning on the decomposed results;

[0016] S3. The fusion enhancement result is compared with the original reference image by a discriminative network to obtain an enhancement result that is closer to the reference image.

[0017] Furthermore, the multinomial decomposition loss of the decomposition network in step S1 is L. RD =L ini +α*L wtv +β*L com +γ*L err The initial loss L ini Weighted total change loss L wtv Decomposition loss L com and reflection error loss L err The constants α, β, and γ are specific gravity parameters, as detailed below:

[0018] Initial loss L ini By calculating the decomposition of illumination P I With the estimated illumination P v The mean square error (MSE) between the two is used to achieve the decomposition of illumination P. I The initialization formula is as follows:

[0019]

[0020] Among them, In I For the input image lighting components, In V To estimate the illumination components of an image, Lb I For the reference image illumination components, Lb V To estimate the image illumination components;

[0021] Weighted total change loss L wtv Weighted total change loss L wtv This method utilizes the total variation of the image to limit the noise level, thereby smoothing the illumination components and avoiding halo artifacts. A weight matrix P is defined. W Specifically, it is expressed as follows:

[0022]

[0023] in, Let w(x) represent the total variation in the horizontal and vertical directions, and let w(x) be a 3×3 window centered at pixel x. Therefore, L wtv It can be represented as:

[0024]

[0025] Decomposition loss L com By calculating P I ·P R The mean square error (MSE) between P and P is used to ensure the accuracy of image decomposition, where P is the original image illumination. R The specific calculation process for reflected light is as follows:

[0026]

[0027] Where Input is the original low-light image, and Label is the illumination of the target image.

[0028] Reflection error loss L err :

[0029] In R For the reflected light component of the input image, Lb R The reflected light components of the reference image are calculated by In R With Lb R The mean square error (MSE) between the two is used to ensure the accuracy of the reflectance obtained from the image decomposition.

[0030] Furthermore, the fusion enhancement network in step S2 includes CRM, which maps the illumination components obtained from the decomposition of the low-light image to a coarse, good exposure effect. The specific process of CRM processing can be represented by the following formula:

[0031]

[0032] CRM (In) I Input) = e b(1-k) Input k ,

[0033] In the formula, ε represents a small constant, and a and b are camera parameters. To accommodate most exposures, they are set to -0.3293 and 1.1258 respectively. The final enhancement result obtained by the fusion enhancement network can be expressed as:

[0034] Out = FENet(Input, CRM(In) I Input), In R ),

[0035] The content loss parameters used in the VGG19 model are employed during the pre-training of this fusion augmentation network.

[0036]

[0037] The VGG19 network structure contains 19 hidden layers, consisting of 16 3×3 convolutional layers and 3 2×2 fully connected layers. The overall structure is very clear and simple. Moreover, the VGG network uses a combination of convolutional layers with multiple small filters, which produces better filtering results than using a single large filter.

[0038] Furthermore, in step S3, the discriminator network consists of 6 convolutional layers and 1 fully connected layer, and the Prelude function is used to activate the network. During the alternating training of the discriminator network, a novel adversarial loss RDGAN_d is used, which is calculated based on the enhancement result and the reflectivity component of the original reference image. This loss helps to restore the image color and details, and the formula is expressed as follows:

[0039]

[0040] The final adversarial enhancement loss can be composed of the content loss parameters and adversarial loss parameters from the pre-trained VGG19 model of the fused enhancement network:

[0041] L FE =L con +L RDGAN_g ,

[0042] Among them, D fake D refers to the probability that the input image is an augmented image. real This refers to the probability that the input image belongs to the reference image. R This represents the reflectance component obtained from the enhanced image after decomposition by the decomposition network.

[0043] Furthermore, image preprocessing includes image grayscale conversion, grayscale stretching, binarization, and morphological processing of the enhanced container image.

[0044] Furthermore, extracting the container number characters involves projecting the preprocessed binarized image horizontally and analyzing its grayscale distribution characteristics; by counting the number of white pixels in the y-direction, other container parameter information below the container number can be effectively distinguished.

[0045] Furthermore, the recognition of container number characters adopts a convolutional neural network model. The specific structure of the convolutional neural network model includes two convolutional layers, one max pooling layer, two fully connected layers, and a Dropout layer.

[0046] Furthermore, a GUI visualization system interface was designed using PyQt5 tools to integrate and display the container number identification process and results, presenting the completeness of the process.

[0047] Furthermore, the GUI visualization system interface includes the loaded original low-light container body image, the enhanced container body image, the located container number area, and the final container number character recognition result.

[0048] Compared with the prior art, the present invention has the following advantages:

[0049] (1) For the enhancement processing of low-light container images at night, this invention analyzes the method of using an adversarial enhancement network based on Retinex decomposition for low-light image enhancement. By building the enhancement network and inputting the actual collected night container images into the network, the corresponding processing results are obtained.

[0050] (2) Multiple image quality evaluation parameters were introduced to objectively analyze the results presented by the network model, and the corresponding brightness distribution histogram was combined. Through the verification of various data, the image processed by the adversarial enhancement network based on Retinex decomposition effectively avoided color distortion, preserved image details more well, and effectively suppressed the noise effect in the enhancement process.

[0051] (3) In order to reduce the difficulty of locating and obtaining the box number characters, a series of preprocessing operations were performed on the enhanced result image, including grayscale stretching, morphological processing, and binarization, to remove various noise and other interference factors, and to avoid the problem of characters sticking together.

[0052] (4) By understanding and analyzing the characteristics of the box number area and characters, the projection positioning segmentation method is used to obtain the box number characters. This process requires first projecting the overall image horizontally to locate the box number area; then projecting the image of the area vertically to capture the characteristics of the peaks and valleys in the histogram of the corresponding pixel distribution, so as to accurately segment and obtain the box number characters.

[0053] (5) In this invention, a convolutional neural network is built for the identification and acquisition of box number information, and iterative training is performed on it using the constructed box number character dataset. The accuracy obtained by training can reach 98.09%.

[0054] (6) In order to visualize the effects of each part and show the integrity of the whole process, the Qt Designer visual interface designer in PyQt5 was used to configure the controls called therein and design the corresponding GUI system interface.

[0055] (7) By extracting some actual nighttime port container images and inputting them into the final fusion system for multiple tests, it can be seen from the results that the system can achieve good enhancement for container images with insufficient nighttime lighting or shadow effects, and the accuracy of container number recognition after enhancement can be maintained above 95%. Attached Figure Description

[0056] Figure 1 This is a schematic diagram of the process of the present invention;

[0057] Figure 2 This is a diagram illustrating the overall framework of the adversarial enhancement network based on Retinex decomposition of the present invention.

[0058] Figure 3 The P of the present invention is used for initializing the estimated illumination component. v Display diagram;

[0059] Figure 4 The actual decomposition of the illumination component image obtained in this invention and the image without P v Compare the estimated decomposition results of the initial illumination;

[0060] Figure 5 This is the definition of the CRM function in this invention;

[0061] Figure 6 This is a diagram showing the CMR adjustment results of the present invention;

[0062] Figure 7 This is a comparison diagram of the enhanced result of the present invention and the original low-light image;

[0063] Figure 8The brightness histogram of the enhanced result of this invention and the original low-light image;

[0064] Figure 9 Images of containers as described in an embodiment of the present invention;

[0065] Figure 10 This is a horizontal projection effect diagram of the present invention;

[0066] Figure 11 This is a vertical projection effect diagram of the present invention;

[0067] Figure 12 This is a structural diagram of the Convolutional Neural Network (CNN) model of the present invention;

[0068] Figure 13 This is the box number character dataset of the present invention;

[0069] Figure 14 This is a table showing the accuracy and loss of the convolutional neural network training according to the present invention;

[0070] Figure 15 For the present invention Figure 14 The training loss curve and training accuracy curve;

[0071] Figure 16 This is a block diagram of the port container number identification system of the present invention;

[0072] Figure 17 This is a diagram showing the test results of the adversarial enhancement network of the present invention;

[0073] Figure 18 This is a system workflow diagram of the present invention;

[0074] Figure 19 This is the system test accuracy table of the present invention;

[0075] Figure 20 This is a graph showing the system test recognition accuracy of the present invention. Detailed Implementation

[0076] To enable those skilled in the art to better understand the specific solutions of the present invention, the following description, in conjunction with the accompanying drawings, further illustrates a port container number recognition method based on image enhancement.

[0077] like Figure 1 As shown, a port container number recognition method based on image enhancement includes the following steps:

[0078] S1. Collect images of containers in low-light environments at cargo ports to create an image library;

[0079] S2. Construct an adversarial enhancement network for Retinex decomposition to perform image enhancement processing on the image library to improve the enhancement effect;

[0080] S3. By preprocessing the container numbers in the enhanced container images, extracting characters to construct a container number character dataset, and iteratively training the constructed convolutional neural network model on the character dataset, the accuracy of character recognition is improved.

[0081] S4. Load the actual acquired low-light container images to be identified into the trained image enhancement network to enhance the low-light images and obtain high-quality images;

[0082] S5. Preprocess the enhanced high-quality image to extract the character images of the boxes to be identified. Input the obtained character images into the trained convolutional neural network model to identify the box numbers and obtain the final result.

[0083] Specifically, in step S1, the original low-light image is obtained: the original image can be an image collected in an environment with insufficient light, such as an image collected at night. This embodiment of the invention does not specifically limit this. For ease of description, the "original low-light image" will be referred to as "low-light image" or "original image" below.

[0084] like Figure 2 As shown, the Retinex-based adversarial enhancement network in step S2 integrates an adversarial learning framework based on Retinex theory, exhibiting good color and detail recovery performance. This Retinex-based adversarial enhancement network is an end-to-end adversarial learning network, with an overall framework comprising a generator network and a discriminator network. The generator network can be divided into a decomposition network (RDNet) and a fusion enhancement network (FENet). First, the decomposition network effectively decomposes the original low-light image, using multiple decomposition losses to ensure detail recovery during the decomposition process. Then, the fusion enhancement network performs enhancement learning on the decomposed result. The discriminator network aims to play a game between the fusion enhancement result and the original reference image to obtain an enhancement result that more closely resembles the reference image. Specifically:

[0085] (1) Decompose the network:

[0086] The original image and its corresponding reference image serve as input to the decomposition network, which outputs illumination and reflectance components. A luminance map is constructed by finding the channel with the maximum luminance value in the illumination component, significantly reducing the resolution space and computational cost. Therefore, the decomposition network operates on the HSV space of the original image. Furthermore, to fuse narrow breaks and fill gaps in the contour, a morphological closure operation is performed on the obtained V channel components to obtain the P-value used to initialize the estimation of the illumination components. v The entire network structure consists of multiple convolutional layers, and deconvolution operations are introduced to transform low-resolution images into high-resolution images. At the end of the decomposition network, nearest neighbor upsampling layers and 3×3 convolutional layers are used to replace deconvolutional layers, and PRelu is used as the activation function to consider negative regions.

[0087] The multivariate decomposition loss of the decomposition network is L RD =L ini +α*L wtv +β*L com +γ*L err The initial loss L ini Weighted total change loss L wtv Decomposition loss L com and reflection error loss L err The constants α, β, and γ are specific gravity parameters, as detailed below:

[0088] Initial loss L ini By calculating the decomposition of illumination P I With the estimated illumination P v The mean square error (MSE) between the two is used to achieve the decomposition of illumination P. I The initialization formula is as follows:

[0089]

[0090] Among them, In I For the input image lighting components, In V To estimate the illumination components of an image, Lb I For the reference image illumination components, Lb V To estimate the image illumination components;

[0091] Weighted total change loss L wtv Weighted total change loss L wtv This method utilizes the total variation of the image to limit the noise level, thereby smoothing the illumination components and avoiding halo artifacts. A weight matrix P is defined. W Specifically, it is expressed as follows:

[0092]

[0093] in, Let w(x) represent the total variation in the horizontal and vertical directions, and let w(x) be a 3×3 window centered at pixel x. Therefore, L wtv It can be represented as:

[0094]

[0095] Decomposition loss L com By calculating P I ·P R The mean square error (MSE) between P and P is used to ensure the accuracy of image decomposition, where P is the original image illumination. R The specific calculation process for reflected light is as follows:

[0096]

[0097] Where Input is the original low-light image and Label is the target image illumination;

[0098] Reflection error loss L err : In R For the reflected light component of the input image, Lb R The reflected light components of the reference image are calculated by In R With Lb R The mean square error (MSE) between the two is used to ensure the accuracy of the reflectance obtained from the image decomposition.

[0099] (2) Converged Enhanced Network:

[0100] The framework of the fusion enhancement network is largely the same as that of the decomposition network. The main differences lie in the number of convolutional layers and the absence of fully connected layers in the fusion enhancement network. To maintain the naturalness of the image, a CRM (Camera Response Model) is added to the fusion enhancement network. This model maps the illumination components obtained from the decomposition of the low-light image to a coarse, well-exposed effect. The specific process of CRM processing can be represented by the following formula:

[0101]

[0102] CRM (In) I Input) = e b(1-k) Input k

[0103] In the formula, ε represents a small constant, and a and b are camera parameters. To accommodate most exposures, they are set to -0.3293 and 1.1258 respectively. Therefore, the final enhancement result obtained by the fusion enhancement network can be expressed as:

[0104] Out = FENet(Input, CRM(In) I Input), In R )

[0105] The content loss parameters used in the VGG19 model are employed during the pre-training of this fusion augmentation network.

[0106]

[0107] The VGG19 network structure contains 19 hidden layers, consisting of 16 3×3 convolutional layers and 3 2×2 fully connected layers. The overall structure is very clear and simple. Moreover, the VGG network uses a combination of convolutional layers with multiple small filters, which produces better filtering results than using a single large filter.

[0108] (3) Discriminator Network:

[0109] The discriminator network consists of six convolutional layers and one fully connected layer, and uses the Prelude function for network activation. During the alternating training of the discriminator network, a novel adversarial loss, RDGAN_d, is employed. This loss is calculated based on the enhancement result and the reflectance component of the original reference image, which helps in the recovery of image color and detail. The formula is as follows:

[0110]

[0111] Among them, D fake D refers to the probability that the input image is an augmented image. real This refers to the probability that the input image belongs to the reference image. R This represents the reflectance component obtained from the enhanced image after decomposition by the decomposition network.

[0112] Therefore, the final adversarial enhancement loss can be composed of the content loss parameters and adversarial loss parameters from the pre-trained VGG19 model of the fused enhancement network: L FE =L con +L RDGAN_g .

[0113] The main purpose of the adversarial learning framework integrated with the Retinex decomposition-based adversarial enhancement network is to simultaneously feed low-light images and their corresponding reference images into the network. Through continuous interaction between the two, the overall adversarial enhancement network gradually learns the true distribution of the samples. To this end, we need to provide the training of the adversarial enhancement network model with an image dataset containing two types of samples: low-light images with different exposure levels and corresponding reference images under normal lighting conditions.

[0114] This embodiment uses a low-light box image actually captured at night as input to illustrate the entire enhancement process of the adversarial enhancement network. The specific enhancement process includes:

[0115] A1. Obtain the V channel image for initializing the illumination components: Extract the V channel components of the original image in HSV space, perform morphological closure operations on them to fuse narrow breaks, fill gaps on the contour, and extract and process the result (used to initialize the P-values ​​of the estimated illumination components). v )like Figure 3 As shown;

[0116] A2. Processing of the decomposed network: The P function used in A1 to initialize the estimated illumination components... v The original low-light image is used as input to the decomposition network. The network model processes the data to obtain the decomposition results, namely the reflectance component and the illumination component, as shown below. Figure 4 As shown, it can be seen that the actual light component image obtained by decomposition is compared with the image without P. v The estimated decomposition results of the initial illumination are compared, and P is used in the decomposition network. v The effect obtained by initializing the lighting is clearer, which can effectively avoid the problem of loss of detail when smoothing the lighting components later;

[0117] A3. CRM Exposure Adjustment: Before calling the fusion enhancement network model to obtain the enhancement result, the CRM function is needed to adjust the low-light image to a roughly good exposure effect. The definition of the function and the effect after adjustment are as follows: Figure 5 and Figure 6 As shown;

[0118] A4. Adversarial enhancement processing: The original low-light image, the decomposed reflectance components, and the CRM adjustment results are simultaneously input into the fusion enhancement network and discriminator network modules to achieve restoration adjustment and obtain the final enhancement result.

[0119] Image result analysis after the above enhancement process: The result of the adversarial enhancement is compared with the input low-light image as follows, such as... Figure 7 As shown: By comparing the enhanced result with the original low-light image, it can be seen that, based on Retinex...

[0120] The decomposed adversarial enhancement network can effectively adjust the brightness and contrast of an image, exhibiting good color restoration capabilities without significant color distortion. To further analyze exposure and brightness distribution, corresponding brightness histograms are plotted, such as... Figure 8As shown in the histogram, the brightness distribution is more balanced after enhancement. The grayscale value is roughly stable between 100 and 200, and there are almost no pixels around 0 and 255. This indicates that the image processed by this enhancement network does not have noise or overexposure issues, and the overall effect is good.

[0121] The image quality evaluation parameters include:

[0122] Image Mean: This is the most basic indicator parameter in image quality evaluation. It represents the average brightness of an image by calculating the average grayscale value of all pixels. The specific formula is as follows:

[0123]

[0124] Standard deviation: This represents the dispersion of gray values ​​relative to the image mean. A larger standard deviation indicates a more dispersed gray-level distribution, meaning greater image contrast; a smaller standard deviation indicates lower image contrast. The formula is as follows:

[0125]

[0126] Mean Squared Error (MSE): This is an objective evaluation metric that compares the difference between the processed image and the original reference image. The smaller the MSE value, the higher the detail similarity between the enhanced image and the original image, i.e., the better the relative quality of the image. Its mathematical formula is as follows:

[0127]

[0128] Peak Signal-to-Noise Ratio (PSNR): As the most important and widely used objective image evaluation metric, it is based on the error between corresponding pixels, that is, a parameter that evaluates image quality based on mean square error. PSNR is inversely proportional to image distortion; the higher the PSNR value, the lower the distortion and the better the image quality. The corresponding mathematical formula is as follows:

[0129]

[0130] Image entropy is a parameter used to evaluate an image using a quantitative standard. It is generally used to measure the amount of information in an image and reflects its richness. The larger the value, the more information the image contains. Assuming p(i) is the probability of the i-th gray level appearing in the image, the image entropy is maximized when all gray levels are equally distributed. At this point, the histogram distribution has been basically equalized. The formula for information entropy is as follows:

[0131]

[0132] Average Gradient: This is a parameter used to objectively reflect the sharpness of an image. In image processing technology, image sharpness is a very important indicator. The sharper the image, the larger the corresponding average gradient value, and the more comprehensive the image information received by the human visual system. Its mathematical formula is as follows:

[0133]

[0134] in, This represents the magnitude of the gradient in the horizontal direction of the image; This represents the magnitude of the gradient in the vertical direction of the image.

[0135] Specifically, in step S3, the container number recognition method includes image preprocessing, container number region localization, and segmentation of the container number characters, specifically including:

[0136] B1. Image preprocessing was performed, including image grayscale conversion, grayscale stretching, binarization, and morphological processing on the enhanced container image, as detailed below:

[0137] Image grayscale conversion transforms the original image from a color image in RGB space into a grayscale image, effectively reducing the storage space required for the image and improving the overall computing speed while preserving relatively complete image feature information.

[0138] Gray-scale stretching, through mapping calculations, stretches or compresses the gray-scale distribution range of the entire image, which can further improve the contrast of gray-scale images and make the images clearer;

[0139] Morphological processing utilizes image opening and top-hat operations. Opening operation involves erosion followed by dilation, typically used to eliminate subtle noise, smooth the edge contours of the target region, effectively separate narrow redundant connections, and ensure that the position and volume of the object do not change significantly. Top-hat operation is used to obtain the difference image between the original input grayscale image and the opening result image. Performing top-hat operation after opening results in a brighter region around the target than the original image, aiming to enhance the bright target object in a dark background image.

[0140] Image binarization, through an appropriate threshold, yields an image that clearly presents both overall and local features. The detailed information within the image is only related to the positions of black and white pixels with values ​​of 0 and 255, which is more conducive to subsequent analysis and processing of target regions in the image.

[0141] B2. Retrieve box number characters:

[0142] To improve the efficiency of port container management and ensure real-time tracking of goods during tallying and transportation, containers are typically printed with information such as container type markings, manufacturer logos, intended use, and container serial numbers. By analyzing a large number of collected container images, patterns and characteristics of their container number coding can be identified. Several container images are selected for illustration, such as... Figure 9 As shown, the container number coding rules follow the ISO 6346 coding standard and consist of 11 fixed-size printed characters. The first four characters are uppercase English letters, followed by six Arabic numerals, and the last character is a separate Arabic numeral, often distinguished from the first ten digits of the container registration code by adding an outer border. At the same time, the space corresponding to the container number area changes frequently, and the spacing between the English letters and numbers is basically the same.

[0143] B3. Box Number Location: The preprocessed binarized image is projected horizontally, and its grayscale distribution characteristics are analyzed. By counting the number of white pixels in the y-direction, it is possible to effectively distinguish other box parameters below the box number. The effect is as follows: Figure 10 As shown, for the horizontally arranged container numbers, the characters are mainly concentrated in one row, and the number of characters is significantly greater than in other character areas. Prior knowledge indicates that the container number is located at the top of the overall image. Therefore, by observing the histogram obtained from the horizontal projection, the region with the most concentrated white pixel distribution and the highest peak is selected as the candidate region. Furthermore, the boundaries of the container number character region are determined based on the characteristics of peaks and troughs: the row number corresponding to the first non-zero column vector and the row number corresponding to the next adjacent zero column vector. This completes the container number region localization.

[0144] B4. Box number character segmentation: By vertically projecting the segmented character region image, the number of white pixels in the x-direction is counted, and a corresponding distribution histogram can be obtained. The regions of each character will exhibit a peak-valley-peak pattern, such as... Figure 11 As can be seen, the image preprocessing in the early stage has effectively avoided the situation where characters stick together. Therefore, the gaps between each character can be well reflected in the histogram, which corresponds to the part where the number of white pixels is zero. By recording the column number corresponding to these items with zero pixels, the difference between the position of the next item and the position of the current item is determined, thereby determining the start and end column number corresponding to each character, and finally completing the segmentation and acquisition of each box number character.

[0145] B5. Box Number Character Recognition: A Convolutional Neural Network (CNN) model is used, specifically consisting of two convolutional layers, one max-pooling layer, two fully connected layers, and a Dropout layer, as shown below. Figure 12 As shown, the input to the recognition network is a binary image of a single segmented character, with a size of 32×40. The first convolutional layer has a kernel size of 3×3, 32 filters, and is activated using the ReLU function. The second convolutional layer increases the number of filters to 64 based on the first convolutional layer. The pooling layer uses max pooling with a size of 2×2 and a stride of 1. The first fully connected layer uses the ReLU activation function and has 128 filters. The second fully connected layer uses the Softmat function as the activation function, and the number of filters is determined by the number of recognition categories in the output, i.e., the number of digits and letters in the bin number.

[0146] Training the network of this invention requires pre-constructing a corresponding container number character dataset and a large number of container images collected from on-site snapshots. The container number characters are located and segmented using a projection localization and segmentation method. For severely damaged containers and those with side panels, manual localization and segmentation are used. Each character image is binarized and normalized to a size of 32×40, resulting in 3997 character images constituting the container number character dataset. A partial example is shown below. Figure 13 As shown; Figure 14-15 As shown, the constructed box number character dataset is fed into the built convolutional neural network for training. With the appropriate batch processing parameters and iteration count configured, the model loss is reduced from 2.41 to 0.07 during the training process, and the recognition accuracy is increased from 36.36% to 98.09%. Thus, a relatively mature convolutional neural network model can be obtained to complete the recognition of box number characters, and it is saved as a char_cnn.h5 model output for later direct use in the system.

[0147] like Figure 16 As shown, the image enhancement-based port container number recognition system was designed and built using Python 3.6 assembly language on the Tensorflow framework within the Anaconda3 compilation environment. Specifically, it includes:

[0148] In the development environment layer, it is relatively simple to write GUI graphical interfaces using PyQt5 because it has a large number of control styles, complete functional documentation, and high stability. In addition, PyQt5's ecosystem support features can convert the drawn .ui files into .py files for later integration and modification.

[0149] The base layer mainly contains two parts of data that this system needs to use: one is the pre-collected and gathered nighttime images of container bodies; the other is the container number character dataset that the recognition network needs to train.

[0150] The data processing layer is divided into two aspects: data loading and model training optimization. Data loading refers to loading the path information of the corresponding image from the specified dataset required for each model training. Model training optimization refers to training the adversarial enhancement network model and the convolutional neural network model used to recognize box number characters separately, and saving the obtained models for direct use in system research, so that the system can output high-quality enhancement effects and accurate recognition results.

[0151] The logic layer mainly loads the actual nighttime low-light container images into the trained augmentation network and recognition network for testing. First, the augmentation network enhances the low-light images, then a series of preprocessing operations are performed on the enhanced images, and finally the resulting character images are fed into the recognition network to identify the final result.

[0152] The presentation layer displays the original nighttime container image, the enhanced image, the segmented container number character area, and the final identified container number result in a system window using a GUI interface.

[0153] The overall image-enhanced port container number recognition system mainly includes an enhancement module, a recognition module, and a system visualization module.

[0154] After defining the necessary functional modules and functions for the overall network, a parser can be created during training, and the corresponding parameters can be configured by calling the `add_argument()` function to implement the training process. Once the final network model is obtained, the decomposition network model and the fusion enhancement network model are saved in their respective paths so that they can be directly called based on the path information when enhancing input nighttime low-light container images in the system. A nighttime low-light container image is input to the Retinex decomposition-based adversarial enhancement network. The decomposition network obtains the corresponding illumination and reflectance components. After obtaining the decomposition results, the introduced Camera Response Function (CRM) is used to adjust the obtained illumination components to a coarse, relatively natural, and good exposure result. Then, the fusion enhancement network (FENet) uses the input low-light image, the decomposed reflectance components, and the CRM processing results based on the illumination components as input to obtain the final enhanced image.

[0155] like Figure 17 The image shown is two images of a container at night in a low-light environment, which were actually collected at the port. These images were then input into an adversarial enhancement network for testing, and the corresponding enhancement results were obtained.

[0156] System GUI Interface Design: When designing the GUI visualization system interface using the configured PyQt5 tool in PyCharm, it can be implemented directly through Python or with the help of Qt Designer. To further enable direct calling of enhancement modules and recognition models and presentation of their results, the pyuic5 tool is needed to convert .ui files to .py files. The GUI visualization interface mainly includes the loaded original low-light nighttime image of the container body, the enhanced effect, the located container number area, and the final container number character recognition result. The functions of each button and the implementation effects of each part are as follows:

[0157] Load box image: Simultaneously present the original low-light box image and the enhanced image;

[0158] Locating Box Number: The segmented target box number area is displayed below the button;

[0159] Recognition Result: Displays the final box number character information obtained from the recognition.

[0160] like Figure 18 The diagram shows the overall system workflow. First, the imaging effect of the container body image acquired under low-light conditions at night is enhanced to remove the effects of low light while effectively preserving feature details and avoiding severe color distortion, thus presenting the information of the target area more clearly. Second, a series of preprocessing operations are performed on the enhanced image to remove the influence of various noise and other interference factors on the subsequent recognition effect. Next, the container number area and the container number characters to be recognized are effectively located and segmented, and normalized to facilitate the smooth progress of subsequent recognition work. Finally, the segmented individual character images are input into the recognition network to identify the container number information and display it on the system visualization interface.

[0161] In a specific embodiment, the system test extracted three images from the actual nighttime low-light container images for testing. These three original nighttime low-light container images were input into the system to complete the enhancement processing and container number recognition. The system's working process is presented through a visual interface, which clearly shows the implementation effect of each core functional module in the system. While loading the container images, a relatively clear enhancement result can be obtained, which can then accurately locate the container number area and finally obtain the required container number information.

[0162] like Figures 19-20 As shown, by identifying multiple box numbers and conducting multiple tests, the identification accuracy rate can be maintained at over 95%, providing a good guarantee for supporting the stable operation of the overall system.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A port container number recognition method based on image enhancement, characterized in that, Includes the following steps: Images of containers in low-light environments are collected at cargo ports to form an image library; An adversarial enhancement network based on Retinex decomposition is constructed to perform image enhancement processing on an image database to improve the enhancement effect; By preprocessing the container numbers in the enhanced container images, extracting characters to construct a container number character dataset, and then iteratively training a convolutional neural network model on the character dataset, the accuracy of character recognition is improved. The actual acquired low-light container images to be identified are loaded into a trained image enhancement network to enhance the low-light images and obtain high-quality images. The enhanced high-quality image is preprocessed to extract the character images of the boxes to be identified. The obtained character images are then input into a trained convolutional neural network model for box number recognition to obtain the final result. The construction of the adversarial enhancement network for Retinex decomposition includes the following steps: S1. The original low-light image is effectively decomposed in the decomposition network, and the ability to restore details during the decomposition process is guaranteed by multiple decomposition losses. S2. Use a fusion-enhanced network to perform reinforcement learning on the decomposed results; S3. The fusion enhancement result is compared with the original reference image by a discriminative network to obtain an enhancement result that is closer to the reference image; The multinomial decomposition loss of the decomposition network in step S1 is L. RD =L ini +α*L wtv +β*L com +γ*L err The initial loss L ini Weighted total change loss L wtv Decomposition loss L com and reflection error loss L err The constants α, β, and γ are specific gravity parameters, as detailed below: Initial loss L ini By calculating the decomposition of illumination P I With the estimated illumination P v The mean square error (MSE) between the two is used to achieve the decomposition of illumination P. I The initialization formula is as follows: Among them, In I For the input image lighting components, In V To estimate the illumination components of an image, Lb I For the reference image illumination components, Lb V To estimate the image illumination components; Weighted total change loss L wtv Weighted total change loss L wtv This method utilizes the total variation of the image to limit the noise level, thereby smoothing the illumination components and avoiding halo artifacts. A weight matrix P is defined. W Specifically, it is expressed as follows: in, Let w(x) represent the total variation in the horizontal and vertical directions, and let w(x) be a 3×3 window centered at pixel x. Therefore, L wtv It can be represented as: Decomposition loss L com By calculating P I ·P R The mean square error (MSE) between P and P is used to ensure the accuracy of image decomposition, where P is the original image illumination. R The specific calculation process for reflected light is as follows: Where Input is the original low-light image, and Label is the illumination of the target image. Reflection error loss In R For the reflected light component of the input image, Lb R The reflected light components of the reference image are calculated by In R With Lb R The mean square error (MSE) between the two is used to ensure the accuracy of the reflectance obtained from the image decomposition.

2. The port container number recognition method based on image enhancement according to claim 1, characterized in that, In step S2, the fusion enhancement network includes CRM, which maps the illumination components obtained from the decomposition of the low-light image to a coarse, good exposure effect. The specific process of CRM processing can be represented by the following formula: CRM(In I ,Input)=e b(1-k) Input k , In the formula, ε represents a small constant, and a and b are camera parameters. To accommodate most exposures, they are set to -0.3293 and 1.1258 respectively. The final enhancement result obtained by the fusion enhancement network can be expressed as: Out=FENet(Input,CRM(In I ,Input),In R ), The content loss parameters used in the VGG19 model are employed during the pre-training of this fusion augmentation network. The VGG19 network structure contains 19 hidden layers, including 16 3×3 convolutional layers and 3 2×2 fully connected layers.

3. The port container number recognition method based on image enhancement according to claim 1, characterized in that, In step S3, the discriminator network consists of 6 convolutional layers and 1 fully connected layer. The Prelude function is used to activate the network. During the alternating training of the discriminator network, a novel adversarial loss, RDGAN_d, is employed. This loss is calculated based on the enhancement result and the reflectance component of the original reference image, as expressed in the following formula: The final adversarial enhancement loss consists of the content loss parameters and adversarial loss parameters from the pre-trained VGG19 model of the fused enhancement network: L FE =L con +L RDGAN_g , where D fake D refers to the probability that the input image is an augmented image. real This refers to the probability that the input image belongs to the reference image. R This represents the reflectance component obtained from the enhanced image after decomposition by the decomposition network.

4. The port container number recognition method based on image enhancement according to claim 1, characterized in that, Image preprocessing includes performing image grayscale conversion, grayscale stretching, binarization, and morphological processing on the enhanced container image.

5. The port container number recognition method based on image enhancement according to claim 4, characterized in that, Extracting the box number characters involves horizontally projecting the preprocessed binarized image to locate the box number region; vertically projecting the obtained box number region image; and segmenting and obtaining the box number characters by analyzing the corresponding pixel distribution histogram.

6. The port container number recognition method based on image enhancement according to claim 5, characterized in that, The recognition of box number characters adopts a convolutional neural network model. The specific structure of the convolutional neural network model includes two convolutional layers, one max pooling layer, two fully connected layers, and a Dropout layer.

7. A port container number recognition method based on image enhancement according to any one of claims 1-6, characterized in that, A GUI visualization system interface was designed using PyQt5 to integrate and display the container number identification process and results, presenting the complete workflow.

8. The port container number recognition method based on image enhancement according to claim 7, characterized in that, The GUI visualization system interface includes the loaded original low-light container body image, the enhanced container body image, the located container number area, and the final container number character recognition result.

Citation Information

Patent Citations

  • Container number identification method and device, and electronic device

    CN107832767A