Container damage detection method and system

By converting RGB images into HSV images and combining UNet-Transformer and YOLOv5nano models for container damage detection, the problems of low accuracy and detection rates in the prior art are solved, and efficient damage detection under complex lighting conditions is achieved.

CN120107697APending Publication Date: 2025-06-06COSCO SHIPPING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510311298.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing container damage detection methods have shortcomings in accuracy and detection rates, especially when dealing with complex lighting conditions and background interference, it is difficult to accurately identify damaged areas and types.

Method used

The HSV format image conversion technology is used to convert RGB images into HSV images, and the first-stage damage area segmentation is performed using the UNet-Transformer model, and then the second-stage damage type detection is performed using the YOLOv5nano network model.

Benefits of technology

Through the combination of HSV format image conversion and deep learning model, the detection rate and accuracy of container damage detection are significantly improved, and the damaged areas and types can be accurately identified under complex lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107697A_ABST
    Figure CN120107697A_ABST
Patent Text Reader

Abstract

The invention provides a container damage detection method and system, and the method comprises the steps: firstly obtaining a plurality of to-be-detected container images and a historical container image for training, and converting each to-be-detected container image and the historical container image from an RGB color space to an HSV color space through a cvtColor function; a plurality of to-be-detected container images in an HSV format are obtained and input to a trained UNet-Transformer model for first-stage fine screening detection, a plurality of box body damage area segmentation images are output, the plurality of box body damage area segmentation images obtained through first-stage screening detection are respectively input to a trained YOLOv5nano network model for second-stage screening detection, and a plurality of to-be-detected container images obtained through second-stage screening detection are obtained. According to the method, multiple container pictures containing damage types are output to complete container damage detection, different color areas in the images can be segmented more accurately, all suspected damaged container body pictures can be detected, high robustness is achieved for complex scenes such as illumination change and background interference, and the detection rate and accuracy of container damage detection are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of container detection, and in particular to a container damage detection method and system. Background Art

[0002] As a standardized transport unit, containers are widely used in global logistics and supply chain management. They are used to transport goods by sea, land, air and other modes of transportation. In order to ensure the safety and transportation efficiency of the goods, the containers must be kept in good physical condition. However, during loading and unloading and transportation, containers may be damaged to varying degrees by collision, squeezing or natural environmental factors. The existing container damage detection algorithm directly passes the natural scene image into the single-stage detection algorithm. However, due to the many interference factors in the natural scene, the accuracy of the single-stage algorithm cannot meet the damage recognition requirements with high accuracy and detection rate.

[0003] In addition, there is a detection method that uses laser radar and image algorithm recognition technology to detect the damage of the container body, such as dents, breakage, and rust. However, the laser radar equipment is expensive, and the information collected is not as detailed and comprehensive as the camera. When the vehicle passes through the detection area, the image of the outside of the container body is detected and spliced. This method easily ignores the damage inside the container, and the outside detection is incomplete due to the fixed camera angle.

[0004] In view of the above problems, there is an urgent need for a detection method with high accuracy and detection rate and the ability to more accurately segment different color areas in the image to identify damaged boxes. Summary of the invention

[0005] In order to solve the problems of low accuracy and detection rate in the current container damage detection process, the present invention provides a container damage detection method, which can more accurately segment different color areas in an image, greatly improve the detection rate, and can identify the damage type of the container, greatly improving the accuracy. The present invention also relates to a container damage detection system.

[0006] The technical solution of the present invention is as follows:

[0007] A container damage detection method, characterized in that it comprises the following steps:

[0008] Image acquisition step: acquiring multiple container images to be tested and historical container images for training;

[0009] HSV format image conversion step: using the cvtColor function in the opencv library to convert each container image to be tested and each historical container image from the RGB color space to the HSV color space, and obtaining a plurality of container images to be tested in the HSV format and a plurality of historical container images in the HSV format; organizing the plurality of historical container images in the HSV format into a container image dataset in the HSV format for training;

[0010] The first stage damaged area segmentation step: constructing a UNet-Transformer model based on the U-Net semantic segmentation model and the Transformer model, using the container image data set in the HSV format as a training set sample to train the UNet-Transformer model, and obtaining a trained UNet-Transformer model; inputting the plurality of container images to be tested in the HSV format into the trained UNet-Transformer model for the first stage screening and detection, and outputting a plurality of segmented images of the damaged area of ​​the container;

[0011] The second stage damage type detection step: use the HSV format container image dataset as a training set sample to train the YOLOv5nano network model to obtain a trained YOLOv5nano network model; input the multiple container damaged area segmentation images output by the first stage screening detection into the trained YOLOv5nano network model for the second stage screening detection, and output multiple container images containing damage types to complete the container damage detection.

[0012] Preferably, in the HSV format image conversion step, converting the RGB color space into the HSV color space specifically includes:

[0013] The RGB color values ​​are normalized to obtain normalized RGB color values, and the maximum value of the normalized RGB color values ​​is used as the brightness. The saturation is calculated according to the brightness and the minimum value of the normalized RGB color values; the brightness is then compared with the red component, the green component, and the blue component of the normalized RGB color value. If the brightness is equal to the red component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the green component, and the blue component of the normalized RGB color value; if the brightness is equal to the green component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the red component, and the blue component of the normalized RGB color value; if the brightness is equal to the blue component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the red component, and the green component of the normalized RGB color value, so as to realize the conversion of the RGB color space into the HSV color space.

[0014] Preferably, in the HSV format image conversion step, organizing the plurality of HSV format historical container images into an HSV format container image dataset for training specifically comprises:

[0015] The multiple historical container images in the HSV format are annotated, normalized, and data augmented, wherein the annotation includes a bounding box annotation or a pixel-level mask annotation of a damaged area, and the data augmentation operation includes one or more of rotation, flipping, cropping, scaling, brightness adjustment, and contrast enhancement, so as to organize the images into a container image dataset in the HSV format for training.

[0016] Preferably, in the damaged area segmentation step of the first stage, the trained UNet-Transformer model includes an encoder, a bottleneck layer, a decoder, a bridge layer, a feature embedding layer, a Transformer layer and a feedforward network in sequence. The encoder extracts multiple feature maps from the input HSV format container image to be tested and reduces the size of the feature maps through multiple convolutional layers and maximum pooling operations. The bottleneck layer passes the multiple feature maps to the decoder. The decoder gradually restores the size of each feature map through upsampling, and uses jump connections to pass the feature maps in the encoder to the decoder to generate feature maps that retain detail information; the bridge layer resizes the feature map output by the decoder and inputs it to the feature embedding layer. The feature embedding layer uses pixel-level embedding technology to embed the resized feature map into the Transformer layer. The Transformer layer is used to capture global context information and restore the size of the feature map. The feedforward network processes the feature map of the restored size and outputs the segmentation result of each pixel in the feature map.

[0017] Preferably, in the damaged area segmentation step of the first stage, the cross entropy loss function is used as the loss function in the UNet-Transformer model training, and the adam optimizer is used to update the weight parameters of the UNet-Transformer model; during the UNet-Transformer model training process, a weight file with the smallest loss function value is obtained.

[0018] Preferably, in the damaged area segmentation step of the first stage, the image of the container to be tested in the HSV format is input into the trained UNet-Transformer model, the weight file is called, and each pixel point in the image of the container to be tested in the HSV format is classified by the softmax function to obtain the segmentation image of the damaged area of ​​the container and save it.

[0019] A container damage detection system, characterized by comprising an image acquisition module, an HSV format image conversion module, a first-stage damaged area segmentation module and a second-stage damage type detection module connected in sequence,

[0020] The image acquisition module acquires a plurality of container images to be tested and historical container images for training;

[0021] The HSV format image conversion module converts each container image to be tested and each historical container image from the RGB color space to the HSV color space using the cvtColor function in the opencv library, and obtains a plurality of container images to be tested in the HSV format and a plurality of historical container images in the HSV format; organizes the plurality of historical container images in the HSV format into a container image dataset in the HSV format for training;

[0022] The first-stage damaged area segmentation module builds a UNet-Transformer model based on the U-Net semantic segmentation model and the Transformer model, uses the container image dataset in the HSV format as a training set sample to train the UNet-Transformer model, and obtains a trained UNet-Transformer model; inputs the multiple HSV format images of the containers to be tested into the trained UNet-Transformer model for the first-stage screening and detection, and outputs multiple images of the damaged area segmentation of the container;

[0023] The second-stage damage type detection module uses the container image data set in the HSV format as a training set sample to train the YOLOv5nano network model to obtain a trained YOLOv5nano network model; multiple container damaged area segmentation images output by the first-stage screening detection are input into the trained YOLOv5nano network model for the second-stage screening detection, and multiple container images containing damage types are output to complete the container damage detection.

[0024] Preferably, in the HSV format image conversion module, converting the RGB color space into the HSV color space specifically includes:

[0025] The RGB color values ​​are normalized to obtain normalized RGB color values, and the maximum value of the normalized RGB color values ​​is used as the brightness. The saturation is calculated according to the brightness and the minimum value of the normalized RGB color values; the brightness is then compared with the red component, the green component, and the blue component of the normalized RGB color value. If the brightness is equal to the red component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the green component, and the blue component of the normalized RGB color value; if the brightness is equal to the green component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the red component, and the blue component of the normalized RGB color value; if the brightness is equal to the blue component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the red component, and the green component of the normalized RGB color value, so as to realize the conversion of the RGB color space into the HSV color space.

[0026] Preferably, in the HSV format image conversion module, organizing the plurality of HSV format historical container images into an HSV format container image dataset for training specifically comprises:

[0027] The multiple historical container images in the HSV format are annotated, normalized, and data augmented, wherein the annotation includes a bounding box annotation or a pixel-level mask annotation of a damaged area, and the data augmentation operation includes one or more of rotation, flipping, cropping, scaling, brightness adjustment, and contrast enhancement, so as to organize the images into a container image dataset in the HSV format for training.

[0028] Preferably, in the damaged area segmentation module of the first stage, the trained UNet-Transformer model includes an encoder, a bottleneck layer, a decoder, a bridge layer, a feature embedding layer, a Transformer layer and a feedforward network in sequence. The encoder extracts multiple feature maps from the input HSV format container image to be tested and reduces the size of the feature maps through multiple convolutional layers and maximum pooling operations. The bottleneck layer passes the multiple feature maps to the decoder. The decoder gradually restores the size of each feature map through upsampling, and uses jump connections to pass the feature maps in the encoder to the decoder to generate feature maps that retain detail information; the bridge layer resizes the feature map output by the decoder and inputs it to the feature embedding layer. The feature embedding layer uses pixel-level embedding technology to embed the resized feature map into the Transformer layer. The Transformer layer is used to capture global context information and restore the size of the feature map. The feedforward network processes the feature map of the restored size and outputs the segmentation result of each pixel in the feature map.

[0029] The beneficial effects of the present invention are:

[0030] The present invention provides a container damage detection method, which first obtains a plurality of container images to be tested and historical container images for training, and uses the cvtColor function in the opencv library to convert each container image to be tested and each historical container image from the RGB color space to the HSV color space to reduce the influence of light on the damaged part of the container body, thereby obtaining a plurality of container images to be tested in the HSV format and a plurality of historical container images in the HSV format, and organizes the plurality of historical container images in the HSV format into a container image data set in the HSV format for subsequent model training, and converts the RGB image into the HSV (hue, saturation, brightness) format. In the HSV format, the color features of the damaged area are more obvious, and the HSV color value can better separate the color information, saturation and brightness information, accurately segment different color areas in the image, and has stronger robustness for scenes with large lighting changes, and can more accurately detect the damaged area under complex lighting conditions. Then, after two stages of damaged area segmentation and damage type detection, different deep learning (UNet-Transformer model and YOLOv5nano network model) are used to identify damaged areas and damage types respectively. A UNet-Transformer model is constructed based on the U-Net semantic segmentation model and the Transformer model. The HSV format container image dataset that has been processed by format conversion, annotation, normalization and data enhancement is used as the training set sample. The UNet-Transformer model and the YOLOv5nano network model are trained respectively, and the trained UNet-Tr ansformer model and YOLOv5nano network model. In the first stage damaged area segmentation step, multiple images of the container to be tested in HSV format are input into the trained UNet-Transformer model for the first stage fine screening and detection, and multiple images of the damaged area of ​​the container are obtained, which greatly improves the detection rate; in the second stage damage type detection step, the multiple images of the damaged area of ​​the container obtained by the first stage screening and detection are finally input into the trained YOLOv5nano network model for the second stage screening and detection, and multiple container images containing damage types are obtained to complete the container damage detection, which greatly improves the accuracy.The present invention uses the HSV format to more accurately segment different color areas in the image, screen the pictures, detect all suspected damaged container photos, and has strong robustness to complex scenes such as lighting changes and background interference, which is suitable for container damage detection tasks in actual industrial environments; and uses a two-stage deep learning detection algorithm to identify possible damaged areas of the container in turn, and cuts out the pictures of the target area after the pictures are detected by the UNet-Transformer model, and then inputs them into the YOLOv5nano model for further detection and identification of the specific damage type, which can effectively improve the detection rate and accuracy of container damage detection. The UNet-Transformer model is built based on a deep learning algorithm, combining the semantic segmentation capability of U-Net and the global feature extraction capability of Transformer. The UNet-Transformer model is used to accurately segment the damaged area. The YOLOv5nano model is a lightweight deep learning model based on another deep learning algorithm. It adopts a single-stage target detection framework and can directly predict the location and category of the target from the input image. It can efficiently process the target recognition task in the image and is used to quickly classify and locate the damage type. The two-stage deep learning detection can not only ensure the detection accuracy, but also improve the detection efficiency.

[0031] Therefore, the container damage detection method provided by the present invention is essentially a container damage detection method based on deep learning and image analysis technology, wherein deep learning: through the Unet+Transformer model and the yolov5nano model, high-precision detection and classification of container damaged areas are achieved; image analysis technology: through RGB to HSV conversion, different color areas in the image are accurately segmented, and all suspected damaged container photos are detected, providing high-quality, clearer and more accurate input data for the deep learning model. The image analysis technology provides efficient preprocessing and preliminary analysis capabilities. The two-stage deep learning model uses powerful feature learning and pattern recognition capabilities, and combines deep learning with image analysis technology to achieve high detection rate, high accuracy, and high efficiency of container damage detection, and also reduces equipment costs, improves detection flexibility, and has significant technical advantages and application value.

[0032] The present invention also relates to a container damage detection system, which corresponds to the above-mentioned container damage detection method and can be understood as a system for implementing the above-mentioned container damage detection method, including an image acquisition module, an HSV format image conversion module, a first-stage damaged area segmentation module and a second-stage damaged type detection module connected in sequence, and each module works together to use the HSV format to more accurately segment different color areas in the image, screen the pictures, and detect all suspected damaged box photos, and then use the UNet-Transformer model to accurately segment the damaged area, which can effectively improve the detection rate of container damage detection, and by using the YOLOv5nano model to quickly classify and locate the damage type, the accuracy of container damage detection is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a flow chart of the container damage detection method of the present invention.

[0034] Figure 2 It is a schematic diagram of a box damage diagram converted from RGB format to HSV format according to the present invention.

[0035] Figure 3 It is a schematic diagram of the damaged box area screened by the UNet-Transformer model of the present invention.

[0036] Figure 4 It is a schematic diagram of the YOLOv5nano network model of the present invention for detecting the type of damage to the box. DETAILED DESCRIPTION

[0037] The present invention will be described below in conjunction with the accompanying drawings.

[0038] The present invention relates to a container damage detection method, the flow chart of which is as follows: Figure 1 As shown in the figure, image acquisition step (picture taken by mobile phone) → HSV format image conversion step (RGB to HSV) → first stage damaged area segmentation step (UNet-Transformer model) → second stage damaged type detection step (YOLOv5nano), for multiple container images to be tested, a two-stage deep learning detection algorithm is used, as shown in the figure. Figure 1The UNet-Transformer model and YOLOv5nano (i.e., YOLOv5nano network model) shown in the figure perform phased identification of containers. This solution can more accurately segment different color areas in the image for possible damaged areas of the container, effectively improving the detection rate and accuracy of container damage detection. First, the HSV format is used to more accurately segment different color areas in the image, and the pictures are screened and a low threshold is set to detect all suspected damaged container photos. The UNet-Transformer model is then used to accurately segment the damaged areas, which can effectively improve the detection rate of container damage detection. Finally, the YOLOv5nano model is used to quickly classify and locate the damage type, effectively improving the accuracy of container damage detection. Specifically, the container damage detection method includes the following steps in sequence:

[0039] 1. Image acquisition step: obtain multiple images of containers to be tested and historical container images for training, that is, use a mobile phone to take pictures of suspected damage on the surface of the container body and real damage pictures and divide them into images of containers to be tested and historical container images for training.

[0040] 2. HSV format image conversion step: Use the cvtColor function in the opencv library to convert each container image to be tested and each historical container image from the RGB color space to the HSV color space, and obtain multiple container images to be tested in the HSV format and multiple historical container images in the HSV format respectively; organize the multiple historical container images in the HSV format into a container image dataset in the HSV format for training. Preferably, the multiple historical container images in the HSV format are annotated, normalized and data enhanced, the annotation includes a bounding box annotation of the damaged area or a pixel-level mask annotation, and the data enhancement operation includes one or more of rotation, flipping, cropping, scaling, brightness adjustment, and contrast enhancement, so as to be organized into a container image dataset in the HSV format for training.

[0041] Specifically, if Figure 2 As shown in the figure, the HSV color mode can more accurately segment the damaged area in the image, which is convenient for the subsequent identification and analysis of damaged targets. Therefore, the cvtColor function in opencv is used to convert the RGB color value of each pixel in each container image to be tested and each historical container image into an HSV color value, that is, the RGB image is converted into the HSV format and stored as a jpg image. The specific steps for converting RGB color values ​​to HSV color values ​​are as follows:

[0042] For any pixel in the container image to be tested, its RGB color value is (R, G, B), and its HSV color value is (H, S, V), where R, G, B are the red component, green component, and blue component in the RGB color value, respectively; H, S, V are the brightness, saturation, and hue in the HSV color value, respectively. First, the R, G, and B values ​​need to be normalized to between 0 and 1:

[0043] R′=R / 255, G′=G / 255, B′=B / 255, where R′, G′ and B′ are normalized values.

[0044] Then calculate the brightness V, which is the maximum value of the normalized RGB values:

[0045] V = max(R, G, B) (1)

[0046] Then calculate the saturation S based on the difference between the maximum and minimum values ​​of the normalized RGB values, as shown in the following formula:

[0047]

[0048] Finally, the hue H is calculated based on the relative difference between the RGB components: the formula depends on which component is the maximum value, as shown below:

[0049]

[0050] If the calculated H value is less than 0, add 360 to it to ensure that the final H value is in the range of [0,360]. In addition, since Opencv needs to visualize HSV images, each value needs to be converted to between 0 and 255, that is, H = H / 2, S = S*255, V = V*255.

[0051] For example, suppose there is an RGB value of (255, 128, 64).

[0052] Normalized RGB values: R′=255 / 255=1, G′=128 / 255≈0.5, B′=64 / 255≈0.25.

[0053] Then calculate the brightness V: V = max (1, 0.5, 0.25) = 1;

[0054] Then calculate the saturation S: S = 1-0.25 / 1 = 0.75;

[0055] Calculate the hue H: Since V = R', H = 60*(0.5-0.25 / 1-0.25) = 60*(0.25 / 0.75) = 60*1 / 3 = 20

[0056] Therefore, a given RGB value of (255, 128, 64) would correspond to an HSV value of approximately (20, 0.75, 1).

[0057] 3. The first stage damaged area segmentation step: construct a UNet-Transformer model based on the U-Net semantic segmentation model and the Transformer model, use the HSV format container image data set as the training set sample to train the UNet-Transformer model, and obtain a trained UNet-Transformer model; input the multiple HSV format container images to be tested into the trained UNet-Transformer model for the first stage screening and detection, and output multiple damaged area segmentation images of the container.

[0058] Specifically, if Figure 3 As shown in the figure, firstly, a UNet-Transformer model is constructed based on the U-Net semantic segmentation model and the Transformer model, and the damaged area of ​​the container is locked through the UNet-Transformer model. The Unet semantic segmentation model adopts a U-shaped network structure and a jump connection mechanism to extract multi-scale features of the image, which can effectively capture feature information at different levels, thereby improving the accuracy of image segmentation and the ability to retain details; compared with other deep learning methods, the Unet semantic segmentation model performs well for small samples, can be trained on less labeled data, and achieves better segmentation results. The Transformer model is used to capture global context information and enhance the model's ability to recognize damaged areas. The advantage of the UNet-Transformer model is that it can simultaneously utilize the advantages of Transformer and Unet, which can capture global context information and retain detailed features.

[0059] The UNet-Transformer model includes an encoder, a bottleneck layer, a decoder, a bridge layer, a feature embedding layer, a Transformer layer, and a feedforward network in sequence. The UNet semantic segmentation model includes an encoder, a bottleneck layer, and a decoder, and the Transformer model includes a feature embedding layer, a Transformer layer, and a feedforward network. The input image first enters the encoder part of the UNet semantic segmentation model. The encoder extracts multiple feature maps from the input HSV format container image to be tested and reduces the size of the feature maps through multiple convolutional layers and maximum pooling operations. The bottleneck layer passes multiple feature maps to the decoder. The decoder gradually restores the size of each feature map through upsampling, and uses jump connections to pass the feature maps in the encoder to the decoder to generate feature maps that retain detail information. The bridge layer converts the size of the feature map output by the decoder and inputs it into the feature embedding layer of the Transformer model. The feature embedding layer uses pixel-level embedding technology to embed the feature map after the conversion into the Transformer layer. The Transformer layer is used to capture global context information and restore the size of the feature map. The feedforward network processes the feature map of the restored size and outputs the segmentation result of each pixel in the feature map. Among them, the Transformer layer captures global context information through its self-attention mechanism. The corresponding formula is as follows:

[0060]

[0061] In the above formula, Attention(Q, K, V) is the function of the self-attention mechanism, which receives three parameters: query vector (Q), key vector (K) and value vector (V), which are extracted from different parts of the input sequence, where R is a real number set; softmax is an activation function, d k is the dimension of the key vector. When calculating the attention weight, the result needs to be scaled to prevent the gradient from disappearing or exploding due to excessive values. A is the attention weight matrix.

[0062] Furthermore, this step can use the cross entropy loss function as the loss function in the UNet-Transformer model training, and use the adam optimizer to update the weight parameters of the UNet-Transformer model; in the UNet-Transformer model training process, obtain the weight file with the smallest loss function value, input the container image to be tested in the HSV format into the trained UNet-Transformer model, call the weight file, and classify each pixel point in the container image to be tested in the HSV format through the softmax function to obtain the segmentation image of the damaged area of ​​the box and save it, which can achieve high-precision pixel-level segmentation, accelerate the training process of the model, enable the model to achieve a lower loss value in fewer training rounds, and improve the training efficiency. The optimal weight file obtained during the training process ensures the robustness of the model under different lighting conditions and backgrounds, so that it can stably identify damaged areas.

[0063] 4. The second stage damage type detection step: use the container image dataset in the HSV format as the training set sample to train the YOLOv5nano network model to obtain the trained YOLOv5nano network model; input the multiple container damaged area segmentation images output by the first stage screening and detection into the trained YOLOv5nano network model for the second stage screening and detection, and output multiple container images containing damage types to complete the container damage detection.

[0064] Specifically, the container image dataset in HSV format is first used as a training set sample to train the YOLOv5nano network model to obtain a trained YOLOv5nano network model, and then the damaged areas in the multiple damaged area segmentation images output by the UNet-Transformer model are cropped into rectangular images, and respectively input into the trained yolov5nano network model for the second stage of damage type screening and detection, that is, the small model is used to perform further damage type detection on the cropped area box (cropped rectangular image), such as Figure 4 As shown, multiple container images containing damage types are output to complete the damage detection of the container.

[0065] The present invention also relates to a container damage detection system, which corresponds to the above container damage detection method and can be understood as a system for implementing the above method, including an image acquisition module, an HSV format image conversion module, a first-stage damaged area segmentation module and a second-stage damaged type detection module connected in sequence. Specifically,

[0066] The image acquisition module acquires a plurality of container images to be tested and historical container images for training;

[0067] The HSV format image conversion module converts each container image to be tested and each historical container image from the RGB color space to the HSV color space using the cvtColor function in the opencv library, and obtains a plurality of container images to be tested in the HSV format and a plurality of historical container images in the HSV format; organizes the plurality of historical container images in the HSV format into a container image dataset in the HSV format for training;

[0068] The first-stage damaged area segmentation module builds a UNet-Transformer model based on the U-Net semantic segmentation model and the Transformer model, uses the container image dataset in the HSV format as a training set sample to train the UNet-Transformer model, and obtains a trained UNet-Transformer model; inputs the multiple HSV format images of the containers to be tested into the trained UNet-Transformer model for the first-stage screening and detection, and outputs multiple images of the damaged area segmentation of the container;

[0069] The second-stage damage type detection module uses the container image data set in the HSV format as a training set sample to train the YOLOv5nano network model to obtain a trained YOLOv5nano network model; multiple container damaged area segmentation images output by the first-stage screening detection are input into the trained YOLOv5nano network model for the second-stage screening detection, and multiple container images containing damage types are output to complete the container damage detection.

[0070] Preferably, in the HSV format image conversion module, converting the RGB color space into the HSV color space specifically includes:

[0071] The RGB color values ​​are normalized to obtain normalized RGB color values, and the maximum value of the normalized RGB color values ​​is used as the brightness. The saturation is calculated according to the brightness and the minimum value of the normalized RGB color values; the brightness is then compared with the red component, the green component, and the blue component of the normalized RGB color value. If the brightness is equal to the red component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the green component, and the blue component of the normalized RGB color value; if the brightness is equal to the green component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the red component, and the blue component of the normalized RGB color value; if the brightness is equal to the blue component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the red component, and the green component of the normalized RGB color value, so as to realize the conversion of the RGB color space into the HSV color space.

[0072] Preferably, in the HSV format image conversion module, organizing the plurality of HSV format historical container images into an HSV format container image dataset for training specifically includes:

[0073] The multiple historical container images in the HSV format are annotated, normalized, and data augmented, wherein the annotation includes a bounding box annotation or a pixel-level mask annotation of a damaged area, and the data augmentation operation includes one or more of rotation, flipping, cropping, scaling, brightness adjustment, and contrast enhancement, so as to organize the images into a container image dataset in the HSV format for training.

[0074] Preferably, in the first-stage damaged area segmentation module, the trained UNet-Transformer model includes an encoder, a bottleneck layer, a decoder, a bridge layer, a feature embedding layer, a Transformer layer and a feedforward network in sequence. The encoder extracts multiple feature maps from the input HSV format container image to be tested and reduces the size of the feature maps through multiple convolutional layers and maximum pooling operations. The bottleneck layer passes the multiple feature maps to the decoder. The decoder gradually restores the size of each feature map through upsampling, and uses jump connections to pass the feature maps in the encoder to the decoder to generate feature maps that retain detail information; the bridge layer resizes the feature map output by the decoder and inputs it to the feature embedding layer. The feature embedding layer uses pixel-level embedding technology to embed the resized feature map into the Transformer layer. The Transformer layer is used to capture global context information and restore the size of the feature map. The feedforward network processes the feature map of the restored size and outputs the segmentation result of each pixel in the feature map.

[0075] The present invention provides an objective and scientific container damage detection method and system. By using the HSV format to more accurately segment different color areas in the image, the pictures are screened and all suspected damaged container photos are detected. Then, the UNet-Transformer model is used to accurately segment the damaged areas, which can effectively improve the detection rate of container damage detection. By using the YOLOv5nano model to quickly classify and locate the damage type, the accuracy of container damage detection is effectively improved.

[0076] It should be noted that the above-described specific implementations can enable those skilled in the art to more fully understand the invention, but do not limit the invention in any way. Therefore, although this specification has described the invention in detail with reference to the drawings and embodiments, those skilled in the art should understand that the invention can still be modified or replaced by equivalents. In short, all technical solutions and improvements that do not deviate from the spirit and scope of the invention should be included in the protection scope of the patent for the invention.

Claims

1. A container damage detection method, characterized in that: The following steps are involved: Image acquisition step: acquiring multiple container images to be tested and historical container images for training; HSV format image conversion step: using the cvtColor function in the opencv library to convert each container image to be tested and each historical container image from the RGB color space to the HSV color space, and respectively obtaining a plurality of container images to be tested in the HSV format and a plurality of historical container images in the HSV format; organizing the plurality of historical container images in the HSV format into a container image dataset in the HSV format for training; The first stage damaged area segmentation step: constructing a UNet-Transformer model based on the U-Net semantic segmentation model and the Transformer model, using the container image data set in the HSV format as a training set sample to train the UNet-Transformer model, and obtaining a trained UNet-Transformer model; inputting the plurality of container images to be tested in the HSV format into the trained UNet-Transformer model for the first stage screening and detection, and outputting a plurality of segmented images of the damaged area of ​​the container; The second stage damage type detection step: use the HSV format container image dataset as a training set sample to train the YOLOv5nano network model to obtain a trained YOLOv5nano network model; input the multiple container damaged area segmentation images output by the first stage screening detection into the trained YOLOv5nano network model for the second stage screening detection, and output multiple container images containing damage types to complete the container damage detection.

2. The container damage detection method according to claim 1, characterized in that: In the HSV format image conversion step, converting the RGB color space into the HSV color space specifically includes: The RGB color values ​​are normalized to obtain normalized RGB color values, and the maximum value of the normalized RGB color values ​​is used as the brightness. The saturation is calculated according to the brightness and the minimum value of the normalized RGB color values; the brightness is then compared with the red component, the green component, and the blue component of the normalized RGB color value. If the brightness is equal to the red component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the green component, and the blue component of the normalized RGB color value; if the brightness is equal to the green component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the red component, and the blue component of the normalized RGB color value; if the brightness is equal to the blue component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the red component, and the green component of the normalized RGB color value, so as to realize the conversion of the RGB color space into the HSV color space.

3. The container damage detection method according to claim 1, characterized in that: In the HSV format image conversion step, organizing the plurality of HSV format historical container images into an HSV format container image dataset for training specifically includes: The multiple historical container images in the HSV format are annotated, normalized, and data augmented, wherein the annotation includes a bounding box annotation or a pixel-level mask annotation of a damaged area, and the data augmentation operation includes one or more of rotation, flipping, cropping, scaling, brightness adjustment, and contrast enhancement, so as to organize the images into a container image dataset in the HSV format for training.

4. The container damage detection method according to any one of claims 1 to 3, characterized in that: In the damaged area segmentation step of the first stage, the trained UNet-Transformer model includes an encoder, a bottleneck layer, a decoder, a bridge layer, a feature embedding layer, a Transformer layer and a feedforward network in sequence. The encoder extracts multiple feature maps from the input HSV format container image to be tested and reduces the size of the feature maps through multiple convolutional layers and maximum pooling operations. The bottleneck layer passes the multiple feature maps to the decoder. The decoder gradually restores the size of each feature map through upsampling, and uses jump connections to pass the feature maps in the encoder to the decoder to generate feature maps that retain detail information; the bridge layer resizes the feature map output by the decoder and inputs it to the feature embedding layer. The feature embedding layer uses pixel-level embedding technology to embed the resized feature map into the Transformer layer. The Transformer layer is used to capture global context information and restore the size of the feature map. The feedforward network processes the feature map of the restored size and outputs the segmentation result of each pixel in the feature map.

5. The container damage detection method according to claim 1, characterized in that: In the damaged area segmentation step of the first stage, the cross entropy loss function is used as the loss function in the UNet-Transformer model training, and the adam optimizer is used to update the weight parameters of the UNet-Transformer model; during the UNet-Transformer model training process, a weight file with the smallest loss function value is obtained.

6. The container damage detection method according to claim 5, characterized in that: In the damaged area segmentation step of the first stage, the container image to be tested in HSV format is input into the trained UNet-Transformer model, the weight file is called, and each pixel point in the container image to be tested in HSV format is classified through the softmax function to obtain the segmentation image of the damaged area of ​​the container and save it.

7. A container damage detection system, characterized in that: It includes an image acquisition module, an HSV format image conversion module, a first-stage damaged area segmentation module, and a second-stage damaged type detection module, which are connected in sequence. The image acquisition module acquires a plurality of container images to be tested and historical container images for training; The HSV format image conversion module converts each container image to be tested and each historical container image from the RGB color space to the HSV color space using the cvtColor function in the opencv library, and obtains a plurality of container images to be tested in the HSV format and a plurality of historical container images in the HSV format; organizes the plurality of historical container images in the HSV format into a container image dataset in the HSV format for training; The first-stage damaged area segmentation module builds a UNet-Transformer model based on the U-Net semantic segmentation model and the Transformer model, uses the container image dataset in the HSV format as a training set sample to train the UNet-Transformer model, and obtains a trained UNet-Transformer model; inputs the multiple HSV format images of the containers to be tested into the trained UNet-Transformer model for the first-stage screening and detection, and outputs multiple images of the damaged area segmentation of the container; The second-stage damage type detection module uses the container image data set in the HSV format as a training set sample to train the YOLOv5nano network model to obtain a trained YOLOv5nano network model; multiple container damaged area segmentation images output by the first-stage screening detection are input into the trained YOLOv5nano network model for the second-stage screening detection, and multiple container images containing damage types are output to complete the container damage detection.

8. The container damage detection system according to claim 7, characterized in that: In the HSV format image conversion module, converting the RGB color space into the HSV color space specifically includes: The RGB color values ​​are normalized to obtain normalized RGB color values, and the maximum value of the normalized RGB color values ​​is used as the brightness. The saturation is calculated according to the brightness and the minimum value of the normalized RGB color values; the brightness is then compared with the red component, the green component, and the blue component of the normalized RGB color value. If the brightness is equal to the red component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the green component, and the blue component of the normalized RGB color value; if the brightness is equal to the green component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the red component, and the blue component of the normalized RGB color value; if the brightness is equal to the blue component, the hue is calculated according to the brightness, the minimum value of the normalized RGB color value, the red component, and the green component of the normalized RGB color value, so as to realize the conversion of the RGB color space into the HSV color space.

9. The container damage detection system according to claim 7, characterized in that: In the HSV format image conversion module, organizing the plurality of HSV format historical container images into an HSV format container image dataset for training specifically includes: The multiple historical container images in the HSV format are annotated, normalized, and data augmented, wherein the annotation includes a bounding box annotation or a pixel-level mask annotation of a damaged area, and the data augmentation operation includes one or more of rotation, flipping, cropping, scaling, brightness adjustment, and contrast enhancement, so as to organize the images into a container image dataset in the HSV format for training.

10. The container damage detection system according to any one of claims 7 to 9, characterized in that: In the first-stage damaged area segmentation module, the trained UNet-Transformer model includes an encoder, a bottleneck layer, a decoder, a bridge layer, a feature embedding layer, a Transformer layer and a feedforward network in sequence. The encoder extracts multiple feature maps from the input HSV format container image to be tested and reduces the size of the feature maps through multiple convolutional layers and maximum pooling operations. The bottleneck layer passes the multiple feature maps to the decoder. The decoder gradually restores the size of each feature map through upsampling, and uses jump connections to pass the feature maps in the encoder to the decoder to generate feature maps that retain detail information; the bridge layer resizes the feature map output by the decoder and inputs it to the feature embedding layer. The feature embedding layer uses pixel-level embedding technology to embed the resized feature map into the Transformer layer. The Transformer layer is used to capture global context information and restore the size of the feature map. The feedforward network processes the feature map of the restored size and outputs the segmentation result of each pixel in the feature map.