Container surface defect detection method based on multi-scale feature fusion

The container surface defect detection method based on multi-scale feature fusion, utilizing a lightweight backbone network and an adaptive attention module, solves the problems of high false detection rate and large computational load in container surface defect detection, achieving high-precision and fast detection results, and is suitable for port logistics environments.

CN120833340AActive Publication Date: 2025-10-24广东德智矩阵科技有限公司 +4

Patent Information

Application Number
CN202511340567.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-10-24
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Existing container surface defect detection technologies suffer from high false detection rates, large computational loads, and difficulty in achieving detection accuracy for defects of different sizes. Furthermore, traditional manual inspection is time-consuming and easily affected by the environment.

Method used

A multi-scale feature fusion detection method is adopted. Basic features are extracted through a lightweight backbone network, and combined with an adaptive attention module and a multi-branch detection head to perform multi-scale feature fusion and post-processing to obtain the defect type, location and size on the container surface.

Benefits of technology

It achieves high-precision and rapid detection of container surface defects, reduces computational complexity and the number of parameters, adapts to diverse scenarios, and improves port logistics efficiency and container management security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833340A_ABST
    Figure CN120833340A_ABST
Patent Text Reader

Abstract

The invention provides a container surface defect detection method based on multi-scale feature fusion, and the method comprises the steps: obtaining a first-resolution image of the surface of a container, carrying out the data preprocessing of the first-resolution image, and obtaining a preprocessed image; a lightweight backbone network is adopted to extract basic features of the preprocessed image; performing multi-scale feature fusion on the basic features to obtain fused features; an adaptive attention module is adopted to extract defect candidate region features of the fusion features so as to suppress background interference of container surface defects; detecting different scale defects of the defect candidate region features by adopting a multi-branch detection head to obtain multi-scale features; and performing post-processing on the multi-scale features to obtain a defect detection result. According to the method, high-precision detection of full-scale defects can be achieved, meanwhile, the model calculation complexity and parameter quantity are remarkably reduced, the method is suitable for industrial-grade application such as port logistics, and the requirements for high precision, high speed and high stability of container surface defect detection are met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and automatic detection technology, in particular to a container surface defect detection method based on multi-scale feature fusion. BACKGROUND

[0002] In the logistics transportation and international trade system, as the core carrying unit, the surface defects of the container, such as concave, rust, crack and deformation, directly affect the structural stability and cargo protection ability. With the increasing global container throughput year by year, the traditional manual inspection mode has significant shortcomings. Manual detection relies on experience and judgment, and the recognition rate of subtle rust and hidden cracks is low. Single container detection takes a long time, and long-term outdoor operation is easily affected by light and weather, with poor detection stability.

[0003] There are two core problems in the existing container surface defect detection technology. One is the method based on traditional machine vision, such as threshold segmentation and edge detection, which is sensitive to the complex background of the container surface, has a high defect false detection rate, and is difficult to adapt to diversified scenes. The second is the scheme based on a single deep learning model, such as CNN and YOLO series model, which can improve the detection accuracy, but faces two major bottlenecks. First, the container size is large, and high-resolution images need to be collected to cover the entire surface of the container, resulting in a sharp increase in model calculation, which cannot meet the real-time detection requirements. Second, the scale difference of different defects is large, and single-scale feature extraction cannot balance the complete recognition of large-scale defects and the accurate capture of small-scale defects, resulting in low detection accuracy.

[0004] Therefore, a method is needed that can achieve high precision, high speed and strong scene adaptability based on the fusion of deep learning architecture and multi-scale feature processing, so as to improve the logistics efficiency of the port and the safety of container management. SUMMARY

[0005] To overcome the problems in the related art, the purpose of the present application is to provide a container surface defect detection method based on multi-scale feature fusion, which can achieve high precision, high speed and strong scene adaptability based on the fusion of deep learning architecture and multi-scale feature processing, so as to improve the logistics efficiency of the port and the safety of container management.

[0006] A container surface defect detection method based on multi-scale feature fusion, comprising: Obtain a first resolution image of the container surface, and perform data preprocessing on the first resolution image to obtain a preprocessed image; wherein the data preprocessing includes defogging, noise reduction, light equalization and image stitching; The lightweight backbone network is used to extract basic features of the preprocessed image, and the basic features include shallow layer features, middle layer features and deep layer features. The basic features are subjected to multi-scale feature fusion to obtain fused features. An adaptive attention module is used to extract defect candidate region features of the fused features, so as to suppress the background interference of the container surface defects. A multi-branch detection head is used to detect defects of different scales of the defect candidate region features, to obtain multi-scale features. The multi-scale features are subjected to post-processing to obtain a defect detection result, wherein the defect detection result includes a defect type, a defect position, a defect confidence and a defect size.

[0007] In the preferred technical scheme of the present application, the first resolution image of the container surface is obtained by: A dynamic down-sampling module is designed, which is used for down-sampling according to texture complexity. The dynamic down-sampling module is used for regional image acquisition in a container detection scene to obtain the first resolution image of the container surface.

[0008] In the preferred technical scheme of the present application, the first resolution image of the container surface is obtained by using the dynamic down-sampling module for regional image acquisition in a container detection scene, which includes: A container scene image is acquired. The gradient entropy value of a local region of the container scene image is calculated. If the gradient entropy value is greater than a set threshold, the local region is subjected to 2 times down-sampling to obtain a first down-sampling result. If the gradient entropy value is less than or equal to the set threshold, the local region is subjected to 4 times down-sampling to obtain a second down-sampling result. The first down-sampling result and the second down-sampling result are combined to obtain the first resolution image of the container surface.

[0009] In the preferred technical scheme of the present application, the lightweight backbone network is used to extract the basic features of the preprocessed image, which includes: The depth convolution is used to extract the spatial features of each input channel of the preprocessed image. The point-by-point convolution is used to perform cross-channel information fusion on the spatial features to obtain cross-channel features. The channel attention gate mechanism is used to adaptively suppress the spatial features and the cross-channel features to obtain the basic features of the preprocessed image, wherein the channel attention gate mechanism is used to remove redundant channel features.

[0010] In the preferred technical solution of the present application, the multi-scale feature fusion is performed on the basic features to obtain fused features, which comprises: The feature fusion is performed on the shallow features and the deep features by using lateral connection and up-sampling to obtain small-scale fused features; The feature fusion is performed on the middle features and the deep features by using lateral connection and up-sampling to obtain medium-scale fused features; The feature fusion is performed on the middle features and the deep features by using lateral connection and down-sampling to obtain large-scale fused features; The small-scale fused features, the medium-scale fused features and the large-scale fused features are combined to form the fused features.

[0011] In the preferred technical solution of the present application, the defect candidate region features of the fused features are extracted by using an adaptive attention module, which comprises: The defect probability of the local region of the fused features is calculated, and a spatial mask is generated; The global average pooling is performed on the fused features, and a multi-layer MLP is used to generate channel weights; the channel weights are used to strengthen the channels related to defects and suppress the channels unrelated to defects; The attention calculation is performed on the fused features based on the spatial mask and the channel weights to obtain the defect candidate region features of the fused features.

[0012] In the preferred technical solution of the present application, before the different scale defects of the defect candidate region features are detected by using the multi-branch detection head to obtain multi-scale features, it further comprises: A detection branch is designed for the defect candidate region features; wherein the detection branch comprises a small-scale branch, a medium-scale branch and a large-scale branch; The detection branch is optimized by using a loss function to obtain a multi-branch detection head; the loss function comprises DIoU Loss and improved Focal Loss, the improved Focal Loss contains a class weight factor and a difficulty weight factor, the class weight factor is used to adjust the uniformity of the class distribution, and the difficulty weight factor is used to adjust the weight of easy and difficult classification samples.

[0013] In the preferred technical solution of the present application, the defect probability of the local region of the fused features is calculated, and a spatial mask is generated, which comprises: The gradient features of different local regions of a training sample set are counted; The gradient features are input into a feature discriminator for probability prediction to obtain the defect probability of different local regions; If the defect probability is greater than or equal to a defect probability threshold, the spatial mask of the corresponding local region is set to 1; If the defect probability is less than the defect probability threshold, a spatial mask of a corresponding local region is set to 0.

[0014] In the preferred technical solution of the present application, the channel attention gate mechanism is used to adaptively suppress the spatial features and the cross-channel features to obtain the base features of the preprocessed image, including: The base features of the preprocessed image are extracted by using the following formula: ; Wherein, F l is the l-th layer base feature of the preprocessed image, F l-1 is the (l-1)-th layer base feature of the preprocessed image, DepthSepConv is a depth separable convolution, ReLU is a nonlinear activation function, BN is a batch normalization, G l is a channel attention gate weight, W l is a weight of the depth separable convolution, is an output of a multiplexing channel of the base features of the preprocessed image.

[0015] In the preferred technical solution of the present application, the lateral connection and up-sampling are used to fuse the shallow features and the deep features to obtain small-scale fusion features, including: The deep features are up-sampled by 4 times to obtain first resolution features; The shallow features are convolved by using a 1x1 convolution kernel to obtain first channel dimension features; The first resolution features and the first channel dimension features are spliced to obtain first spliced features; The first spliced features are convolved by using a 3x3 convolution kernel to obtain small-scale fusion features.

[0016] The present application has the following beneficial effects: The container surface defect detection method based on multi-scale feature fusion provided by the present invention includes obtaining a first-resolution image of the container surface, performing data preprocessing on the first-resolution image to obtain a preprocessed image; wherein the data preprocessing includes dehazing, noise reduction, illumination equalization, and image stitching; using a lightweight backbone network to extract basic features of the preprocessed image; the basic features include shallow features, mid-level features, and deep features; performing multi-scale feature fusion on the basic features to obtain fused features, which are used to improve the method's ability to detect multi-scale defects on the container surface. An adaptive attention module is used to extract defect candidate region features from the fused features to suppress background interference from container surface defects. A multi-branch detection head is used to detect defects of different scales in the defect candidate region features to obtain multi-scale features; and the multi-scale features are post-processed to obtain defect detection results; wherein the defect detection results include defect type, defect location, defect confidence, and defect size. The method provided by the present invention uses multiple detection heads to detect defects of different scales on the basis of fusing multi-scale features. It does not rely on the feature map of a single scale and can accurately locate the boundaries of large-scale deformations and small-scale defects such as rust, thereby improving the recall rate of small-scale defects and reducing the missed detection rate of small-scale defects. Compared with traditional deep learning models and manual detection methods, the method provided by the present invention has higher container detection accuracy. The present invention adopts a lightweight backbone network to reduce the number of network parameters. On this basis, it combines the adaptive attention mechanism and dynamic downsampling optimization to achieve real-time reasoning of high-resolution container images, and can still maintain stable detection performance in extreme environments such as strong light and rainy days. The time consumption of single-box detection is significantly shortened, and low-power edge deployment is supported. The method provided by the present invention significantly reduces the computational complexity and parameter amount of the model while ensuring high detection accuracy and strong scene adaptability for defects of different scales. It is suitable for industrial-grade applications such as port logistics. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a flow chart of the container surface defect detection method based on multi-scale feature fusion of the present invention; Figure 2 This is a flow chart of the present invention for fusing shallow features, mid-level features, and deep features; Figure 3 1 is a diagram showing the detection results of small-scale defects of a container according to the present invention; Figure 4 1 is a diagram showing the detection results of large-scale defects of the container of the present invention. DETAILED DESCRIPTION

[0018] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although preferred embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present invention more thorough and complete and to fully convey the scope of the present invention to those skilled in the art.

[0019] Example 1 like Figure 1 As shown, this embodiment provides a container surface defect detection method based on multi-scale feature fusion, including: S1: Acquire a first-resolution image of the container surface and perform data preprocessing on the first-resolution image to obtain a preprocessed image; wherein the data preprocessing includes defogging, noise reduction, illumination balancing, and image stitching; S2: extracting basic features of the preprocessed image using a lightweight backbone network, wherein the basic features include shallow features, mid-layer features, and deep features; S3: Perform multi-scale feature fusion on the basic features to obtain fused features; S4: using an adaptive attention module to extract defect candidate region features from the fused features to suppress background interference from container surface defects; S5: using a multi-branch detection head to detect defects of different scales in the defect candidate region to obtain multi-scale features; S6: Post-processing the multi-scale features to obtain defect detection results; wherein the defect detection results include defect type, defect location, defect confidence and defect size.

[0020] The acquiring of a first-resolution image of the container surface includes: S11: Design a dynamic downsampling module, wherein the dynamic downsampling module is used to perform downsampling according to texture complexity; S12: Using a dynamic downsampling module to perform regional image acquisition in the container detection scene to obtain a first-resolution image of the container surface.

[0021] The dynamic downsampling module includes an image acquisition unit, a gradient entropy value calculation unit, a first downsampling unit, a second downsampling unit, and a downsampling result combination unit.

[0022] The method of using a dynamic downsampling module to collect images by region in a container detection scene to obtain a first-resolution image of the container surface includes: S121: collecting container scene images; S122: Calculating the gradient entropy value of the local area of ​​the container scene image; S123: If the gradient entropy value is greater than a set threshold, 2 times down-sampling is performed on the local region to obtain a first down-sampling result; S124: If the gradient entropy value is less than or equal to the set threshold, 4 times down-sampling is performed on the local region to obtain a second down-sampling result; S125: The first down-sampling result and the second down-sampling result are combined to obtain a first resolution image of the container surface.

[0023] The first resolution image is a digital image with high pixel density, i.e., a large number of pixels per unit area, and complete detail retention. The lateral resolution of the first resolution image is greater than or equal to 4000 pixels, or the pixel per inch (PPI) is greater than or equal to 300. The resolution of the first resolution image of the present embodiment is 8K, i.e., the width is about 8000 pixels, for example, 7680x4320 pixels, so as to ensure that each small area image has high clarity and rich details. Since the final resolution required by the first resolution image far exceeds the limit that a single 8K camera can achieve, it cannot be captured at one time, so in the container detection scene, image acquisition is performed in regions, and then the sub-images of each region are spliced to obtain the first resolution image of the container surface. In the process of regional image acquisition, the container scene image is first acquired, then the scene region of the scene image is divided into multiple regular local regions, and the camera or object is moved through a high-precision electric control platform to capture each local region multiple times. In the process of shooting, the gradient entropy value of the local region of the container scene image is calculated, 2 times down-sampling is used for the local region with a gradient entropy value higher than a set threshold to obtain a first down-sampling result, the local region with a gradient entropy value greater than the set threshold has a higher probability of having a container surface defect, and the first down-sampling result maximally retains the small-scale defect detail features of the container surface. 4 times down-sampling is used for the local region with a gradient entropy value lower than the set threshold to obtain a second down-sampling result. The local region with a gradient entropy value less than or equal to the set threshold has a lower probability of having a container surface defect, and a larger multiple of down-sampling can reduce the amount of invalid calculation. Finally, the first down-sampling result and the second down-sampling result are spliced to obtain the first resolution image of the container surface. In order to avoid the small-scale defect detail features of the container surface being diluted or lost in the deep network, a residual connection and feature reuse dual-channel mechanism are constructed in the network layer. The residual connection directly transmits the shallow features to the middle layer, skips part of the convolution operation to alleviate the gradient disappearance, and ensures that the small-scale defect features of the container surface can continuously participate in deep feature calculation.

[0024] After obtaining the first resolution image of the container surface, data preprocessing needs to be performed on the first resolution image, specifically including dehazing, noise reduction, light balance and image stitching of the first resolution image. The dehazing is to eliminate the interference of atmospheric haze, fog and dust in the image, which can cause image blurring, contrast reduction and color distortion, so as to restore the clear details and true colors of the image and improve the contrast and visibility of the image. The noise reduction is to suppress or remove random noise generated in the image acquisition and / or transmission process, so as to improve the quality of the image while preserving the true details of the image and reduce the interference of noise on subsequent analysis and human visual perception. The light balance is a process of correcting uneven brightness distribution in the image, for example, due to the angle of the light, the center of the image is bright and the periphery is dark, or there are reflections or shadows in some areas. After light balance processing, the light conditions of the entire image look consistent, providing a stable basis for subsequent image analysis and avoiding misjudgment of the same object due to different brightness. The image stitching is to seamlessly combine multiple images with overlapping areas into a single image with a wide viewing angle and high resolution, thereby breaking through the limitations of the field of view or resolution of a single image and creating a panoramic image or a super-resolution image.

[0025] The basic feature of the preprocessed image is extracted by using a lightweight backbone network, including: S21: spatial features of each input channel of the preprocessed image are extracted by using a deep convolution; S22: cross-channel information fusion is performed on the spatial features by using point-by-point convolution to obtain cross-channel features; S23: adaptive suppression is performed on the spatial features and the cross-channel features by using a channel attention gate mechanism to obtain the basic features of the preprocessed image; wherein the channel attention gate mechanism is used to remove redundant channel features.

[0026] The adaptive suppression of the spatial features and the cross-channel features by using the channel attention gate mechanism to obtain the basic features of the preprocessed image includes: The basic features of the preprocessed image are extracted by using the following formula: ; wherein F l is the lth layer basic feature of the preprocessed image, F l-1 is the (l-1)th layer basic feature of the preprocessed image, DepthSepConv is a depth separable convolution, ReLU is a nonlinear activation function, BN is batch normalization, G l is a channel attention gate weight, W l is a weight of the depth separable convolution, The output of the multiplexing channel for the base feature of the pre-processed image is denoted as * represents the channel-by-channel product. The 1x1 convolution is performed on the (l-2)th layer feature map of the EfficientNet-B0 backbone network to adjust the number of channels, and the output is obtained The output of the multiplexing channel refers to the same channel feature map participating in the subsequent calculation process multiple times without increasing the calculation amount.

[0027] The shallow and middle layers of the lightweight backbone network follow similar optimization ideas. The shallow layer introduces depth separable convolution and channel attention gate to dynamically suppress the channels of redundant textures on the container surface while reducing the calculation amount. The middle layer combines residual connection and feature multiplexing mechanism to avoid dilution of small defect features in deep transmission, allowing more efficient fusion of detail features and semantic features.

[0028] The lightweight backbone network is an improved EfficientNet-B0. The improved EfficientNet-B0 is based on the EfficientNet-B0 architecture, and achieves a balance between performance and efficiency through a three-layer structure. The improved EfficientNet-B0 is based on a reverse residual block, and the layer structure and connection method are optimized by neural architecture search (NAS). The reverse residual block first uses a 1x1 convolution kernel to expand the channel number of the input data, then uses a depth separable convolution to reduce the computational complexity, uses an SE attention mechanism to dynamically adjust the channel weight, and finally uses a residual connection to alleviate the problem of gradient disappearance caused by the Sigmoid activation function of the S-shaped curve. In the shallow layer of the network, a depth separable convolution is used to extract features from the preprocessed image, which can greatly reduce the model computation and parameter quantity. The depth separable convolution consists of two steps: depth convolution and pointwise convolution. The depth convolution is responsible for spatial feature extraction for each input channel of the preprocessed image, specifically, k convolution kernels are used to convolve k input channels, each kernel is responsible for one input channel, and finally k spatial features are output, where k≥2. The pointwise convolution is responsible for channel dimension operation, specifically, the k spatial features extracted by the depth convolution are taken as a group of feature vectors, and the group of feature vectors are input into multiple convolution kernels in turn to fuse the information of different channels in the channel dimension, and finally cross-channel features are obtained. Then, a channel attention gating mechanism is used to adaptively suppress the above spatial features and cross-channel features, thereby obtaining the basic features of the preprocessed image. The channel attention gating mechanism is a mechanism that allows the neural network to adaptively emphasize important feature channels and suppress unimportant or noisy channels. The calculation process of the channel attention gating weight is as follows: the dynamic weight vector is generated by the Sigmoid activation function, the redundant feature channels with repeated texture and high background proportion are adaptively suppressed to retain the key channels related to defects and reduce the parameter quantity of the model, ensuring the extraction accuracy of the basic features while reducing the model size. The dynamic weight vector includes multiple channel attention gating weights, and the sum of all channel attention gating weights is less than or equal to 1. The Sigmoid activation function normalizes the channel attention gating weight to the 0-1 interval, if the channel attention gating weight of a single channel approaches 0, the channel is suppressed; if the channel attention gating weight of a single channel approaches 1, the channel is retained. The Sigmoid activation function has the advantage of being derivable, which facilitates the direction propagation and continuously updates the parameters of the model. The results of the depth convolution in the depth separable convolution are obtained, and the texture and background information of the depth convolution results are extracted in any of the following two ways: (1) a first two-dimensional convolution kernel is used to perform two-dimensional convolution on the depth convolution results to obtain texture; a second two-dimensional convolution kernel is used to perform two-dimensional convolution on the depth convolution results to obtain background.The texture and the background are spliced to obtain the background and texture proportion of each channel. (2) The texture and the background are extracted by using the gray level co-occurrence matrix method, that is, the gray level co-occurrence probability in different distances and directions in the result of the depth convolution is counted to describe the texture, and the gray level co-occurrence probability is generated in the following manner: first, a direction and a step length in pixels are defined, and then a gray level co-occurrence matrix is constructed. The element M(i,j) in the gray level co-occurrence matrix represents the frequency of the pixels with gray levels i and j appearing in a point and a point spanned by the defined direction and step length, that is, the gray level co-occurrence probability. The average value of the eigenvalues in the four directions of 0°, 45°, 90° and 135° in the gray level co-occurrence matrix is calculated to obtain the average eigenvalue. According to the average eigenvalues of the intermediate results in different regions, the background and texture proportions of each channel are determined. The background and texture proportions are input into a Sigmoid activation function to calculate the mapping value. The range of the mapping value is between 0 and 1, and the channel attention gate weight of the corresponding channel is obtained by subtracting the mapping value from 1.

[0029] The container surface defect detection method based on multi-scale feature fusion of the embodiment comprises: acquiring a first resolution image of a container surface, performing data preprocessing on the first resolution image to obtain a preprocessed image; wherein the data preprocessing comprises defogging, noise reduction, illumination equalization and image stitching; extracting basic features of the preprocessed image by using a lightweight backbone network; the basic features comprise shallow features, middle features and deep features; performing multi-scale feature fusion on the basic features to obtain fusion features, which are used to improve the detection capability of the method for multi-scale defects of the container surface; extracting defect candidate region features of the fusion features by using an adaptive attention module to suppress the background interference of the container surface defects; detecting different scale defects of the defect candidate region features by using a multi-branch detection head to obtain multi-scale features; and performing post-processing on the multi-scale features to obtain a defect detection result; wherein the defect detection result comprises a defect type, a defect position, a defect confidence and a defect size. The method provided by the present application uses multiple detection heads to detect different scale defects on the basis of fusing multi-scale features, does not rely on a single scale feature map, can accurately locate the boundaries of large-scale deformations and small-scale defects such as rust, thereby improving the recall rate of small-scale defects and reducing the omission rate of small-scale defects. Compared with traditional deep learning models and artificial detection methods, the method provided by the present application has high container detection accuracy. The lightweight backbone network used in the present application can reduce the parameter quantity of the network, and on this basis, the adaptive attention mechanism and dynamic downsampling optimization are combined to realize real-time inference on high-resolution container images, and the detection performance remains stable in extreme environments such as strong light and rain, and the single-box detection time is significantly shortened, supporting low-power edge deployment. The method provided by the present application significantly reduces the computational complexity and parameter quantity of the model while ensuring high detection accuracy and strong scene adaptability of different scale defects, and is suitable for industrial-level applications such as port logistics.

[0030] Embodiment 2 As Figure 1 shown, the embodiment provides a container surface defect detection method based on multi-scale feature fusion, comprising: S1: obtaining a first resolution image of the container surface, performing data preprocessing on the first resolution image to obtain a preprocessed image; wherein the data preprocessing includes defogging, noise reduction, light balance and image stitching; S2: extracting the basic features of the preprocessed image using a lightweight backbone network, the basic features including shallow features, middle features and deep features; S3: performing multi-scale feature fusion on the basic features to obtain fusion features; S4: extracting defect candidate region features of the fusion features using an adaptive attention module to suppress background interference of the container surface defects; S5: detecting different scale defects of the defect candidate region features using a multi-branch detection head to obtain multi-scale features; S6: post-processing the multi-scale features to obtain defect detection results; wherein the defect detection results include defect type, defect location, defect confidence and defect size.

[0031] The multi-scale feature fusion of the basic features to obtain fusion features comprises: S31: using horizontal connection and up-sampling to fuse the shallow features and the deep features to obtain small-scale fusion features.

[0032] S32: using horizontal connection and up-sampling to fuse the middle features and the deep features to obtain middle-scale fusion features; S33: using horizontal connection and down-sampling to fuse the middle features and the deep features to obtain large-scale fusion features; S34: combining the small-scale fusion features, the middle-scale fusion features and the large-scale fusion features to obtain fusion features.

[0033] Using horizontal connection and up-sampling to fuse the shallow features and the deep features to obtain small-scale fusion features comprises: S311: performing 4 times up-sampling on the deep features to obtain first resolution features; S312: using 1x1 convolution kernel to convolve the shallow features to obtain first channel dimension features; S313: fusing the first resolution features and the first channel dimension features to obtain first splicing features; S314: Convolve the first spliced feature using a 3x3 convolution kernel to obtain a small-scale fusion feature.

[0034] In view of the problem of large difference in the size of the container defects, a cross-layer feature fusion structure is designed to realize full-scale defect coverage. First, the basic features of the preprocessed image are split, and the basic features are input into the lightweight backbone network of the present application. The basic features are split into shallow features, middle features and deep features. The shallow features focus on small-scale defects, such as fine rust and small cracks on the surface of the container, the middle features focus on medium-scale defects, such as local indentation and small area deformation on the surface of the container, and the deep features focus on large-scale defects, such as overall deformation and long cracks of the container body. Then, the shallow features, the middle features and the deep features are fused by horizontal connection and up-sampling operation, and three fusion features are generated, including small-scale fusion features, medium-scale fusion features and large-scale fusion features, which correspond to the small-scale branch, the medium-scale branch and the large-scale branch in the detection branch respectively. The horizontal connection of the present embodiment refers to feature splicing, that is, splicing feature maps of different scales or different branches according to spatial dimensions, which only increases the number of channels or keeps the number of channels unchanged without increasing the depth of the network.

[0035] The specific formula for extracting fusion features by horizontal connection and up-sampling is: ; ; ; Among them, is feature splicing, UpSample is up-sampling, DownSample is down-sampling, Conv is 1x1 convolution, F shallow is shallow feature, F middle is middle feature, F deep is deep feature, is small-scale fusion feature, is medium-scale fusion feature, is large-scale fusion feature.

[0036] Up-sampling refers to the process of increasing the spatial size of an image, for example, up-sampling a 500x500 image to a 600x600 image. In the process of up-sampling, the present embodiment uses the method of bilinear interpolation to enlarge the image. Bilinear interpolation is a high-efficiency and commonly used image enlargement algorithm. By performing linear interpolation twice in the horizontal and vertical directions, the value of the new pixel is calculated according to the weighted average value of the four original pixel points around the new pixel. Compared with the nearest neighbor interpolation, the image produced by bilinear interpolation has no obvious jaggies, and the visual effect is smoother and more natural. Moreover, the calculation amount is moderate, and the calculation speed is fast. Taking the small-scale fusion feature As an example, the extraction process of the deep feature F with resolution H / 16×W / 16 is firstly deep Perform 4 times upsampling to obtain the first resolution feature with resolution H / 4×W / 16, and then perform the shallow feature F with resolution H / 4×W / 4. shallow Perform 1×1 convolution to transform the shallow feature F shallow The number of channels is adjusted to be consistent with the first resolution feature to obtain the first channel dimension feature. Then the first resolution feature and the first channel dimension feature are concatenated to obtain the first spliced ​​feature. Finally, a 3×3 convolution kernel is used to fuse the first spliced ​​feature, and the final output is a small-scale fusion feature with a resolution of H / 4×W / 4 and a number of channels of 256. .

[0037] In extracting mid-scale fusion features When upsampling is first used to convert the deep feature F deep The image size is enlarged so that the deep feature F deep The image size and mid-level features F middle The image size is kept consistent, and the second resolution feature is obtained. Then the middle feature F middle Perform 1×1 convolution to transform the middle layer feature F middle The number of channels is adjusted to be consistent with the second resolution feature to obtain the second channel dimension feature. Then the second resolution feature and the second channel dimension feature are concatenated to obtain the second concatenated feature. Finally, a 3×3 convolution kernel is used to fuse the second concatenated feature to obtain the mid-scale fusion feature. .

[0038] With Shangcai Sample The opposite operation is downsampling, which is to reduce the image space size, for example, downsampling a 1000×1000 image to an 800×800 image. In this embodiment, the image is reduced by the maximum pooling method during the downsampling process, and the data size is reduced by taking the maximum value in the local area. When , we first use down sampling to reduce the middle-level features F middle The image size is reduced so that the middle layer feature F middle Image size and deep features F deep The image size is kept consistent, and the third resolution feature is obtained. Then the deep feature F deep Perform 1×1 convolution to transform the deep feature F deep The number of channels is adjusted to be consistent with the third resolution feature to obtain the third channel dimension feature. Then the third resolution feature and the third channel dimension feature are concatenated to obtain the third concatenated feature. Finally, a 3×3 convolution kernel is used to fuse the third concatenated feature to obtain a large-scale fusion feature. .

[0039] The multi-scale feature fusion of the base features in the embodiment obtains fusion features. First, the base features of the preprocessed image are divided into shallow features, middle features and deep features. The shallow features include fine rust and small cracks, the middle features include local depressions and small area deformations, and the deep features include overall deformations and long cracks. Then, the shallow features, the middle features and the deep features are fused by transverse connection and up-sampling operation, and three fusion features are generated, including small-scale fusion features, medium-scale fusion features and large-scale fusion features. By analyzing defects of different scales on the surface of the container, the detection accuracy and real-time performance of full-scale defects on the surface of the container can be significantly improved, and the recall rate of fine defects and the positioning accuracy of large-size defects are greatly improved. Compared with the traditional deep learning model and the artificial detection method, the method provided by the embodiment has better detection performance, reduces the model calculation complexity and the parameter amount on the basis of ensuring high detection accuracy and strong scene adaptability of full-scale defects.

[0040] Embodiment 3 As shown in Figure 1 The embodiment provides a container surface defect detection method based on multi-scale feature fusion, which comprises the following steps: S1: obtaining a first resolution image of a container surface, and performing data preprocessing on the first resolution image to obtain a preprocessed image; wherein the data preprocessing comprises defogging, noise reduction, illumination equalization and image stitching; S2: extracting base features of the preprocessed image by using a lightweight backbone network, wherein the base features comprise shallow features, middle features and deep features; S3: performing multi-scale feature fusion on the base features to obtain fusion features; S4: extracting defect candidate region features of the fusion features by using an adaptive attention module to suppress background interference of the container surface defects; S5: detecting different scale defects of the defect candidate region features by using a multi-branch detection head to obtain multi-scale features; S6: performing post-processing on the multi-scale features to obtain a defect detection result; wherein the defect detection result comprises a defect type, a defect position, a defect confidence and a defect size.

[0041] The adaptive attention module is used to extract defect candidate region features of the fusion features, which comprises the following steps: S41: calculating defect probabilities of local regions of the fusion features, and generating a spatial mask; S42: performing global average pooling on the fusion feature, and generating channel weights by using a multi-layer perceptron (MLP); the channel weights are used to strengthen channels related to defects and suppress channels irrelevant to defects; S43: performing attention calculation on the fusion feature based on the spatial mask and the channel weights to obtain defect candidate region features of the fusion feature.

[0042] Before the different scale defects in the defect candidate region features are detected by using the multi-branch detection head to obtain multi-scale features, the method further includes: S44: designing a detection branch for the defect candidate region features; the detection branch includes a small-scale branch, a medium-scale branch, and a large-scale branch; S45: optimizing the detection branch by using a loss function to obtain a multi-branch detection head; the loss function includes DIoU Loss and improved Focal Loss, the improved Focal Loss includes a class weight factor and a difficulty weight factor, the class weight factor is used to adjust the uniformity of the class distribution, and the difficulty weight factor is used to adjust the weights of easy-to-classify samples and difficult-to-classify samples.

[0043] The method of calculating the defect probability of the local region of the fusion feature and generating a spatial mask includes: S411: counting gradient features of different local regions of a training sample set; S412: inputting the gradient features into a feature discriminator to perform probability prediction to obtain defect probabilities of different local regions; S413: if the defect probability is greater than or equal to a defect probability threshold, setting a spatial mask of a corresponding local region to 1; S414: if the defect probability is less than the defect probability threshold, setting a spatial mask of a corresponding local region to 0.

[0044] The adaptive attention module is designed to solve the problem of background interference such as container surface stickers and stains, including spatial adaptive attention and channel adaptive attention. The spatial adaptive attention generates a spatial weight map to weight the features at different positions, highlighting the key areas and suppressing irrelevant backgrounds. Specifically, the gradient features of the defect area are calculated by training samples, the defect probability of the local area of the fusion features is calculated, the spatial mask is generated, and only the area with a probability greater than the threshold is retained for subsequent calculation, and the background suppression rate is high. The channel adaptive attention generates a channel weight vector to weight and sum the features of different channels, thereby enhancing useful channels while suppressing redundant channels. First, the global average pooling is performed on the fusion features, and the channel weight is output by 2-layer MLP (Multi-Layer Perceptron), which strengthens the channels related to defects, such as the color channel corresponding to rust and the edge channel corresponding to indentation, and suppresses irrelevant channels, such as the high-saturation color channel corresponding to stickers. The attention module uses a sparse matrix approach, retaining only a small number of effective calculation areas, reducing the computational load compared to traditional self-attention, and making the calculation real-time.

[0045] The multi-branch detection head is an improved YOLO detection head. First, the YOLO detection head is improved from the design of the detection branch. Based on small-scale fusion features, medium-scale fusion features, and large-scale fusion features, corresponding detection branches are designed, including small-scale branches, medium-scale branches, and large-scale branches. Through anchor box size adaptation and feature matching, accurate positioning of defects of different scales is achieved. The small-scale branch corresponds to the detection of defect candidate region features of small-scale fusion features, and the anchor box size of the small-scale branch is designed as 10x10 pixels, 20x20 pixels, and 30x30 pixels. The medium-scale branch corresponds to the detection of defect candidate region features of medium-scale fusion features, and the anchor box size of the medium-scale branch is designed as 40x40 pixels, 60x60 pixels, and 80x80 pixels. The large-scale branch corresponds to the detection of defect candidate region features of large-scale fusion features, and the anchor box size of the large-scale branch is designed as 100x100 pixels, 150x150 pixels, and 200x200 pixels.

[0046] Then the loss function is used to improve the YOLO detection head. The loss function is composed of DIoU Loss and improved FocalLoss, that is, DIoU Loss and improved FocalLoss are added to obtain the loss function. DIoU Loss is used to optimize the positioning accuracy of the bounding box, and improved FocalLoss is used to optimize the defect classification accuracy, especially for small-scale defects and sample imbalance problems. DIoU Loss is a loss function used for bounding box regression in target detection tasks, which is an improved version of the classic IoU Loss and GIoU Loss. DIoU Loss introduces the distance between the center points of the bounding boxes and the area of the minimum enclosing rectangle, which more comprehensively measures the difference between the predicted box and the true box.

[0047] The specific formula adopted by the DIoU Loss is: ; ; Wherein, A is the area of the prediction box, B is the area of the real box, IoU is the intersection over union, b is the center of the prediction box, b g is the center of the real box, is the Euclidean distance, c is the length of the diagonal of the minimum enclosing rectangle surrounding the prediction box and the real box, and L DIoU is the DIoU Loss.

[0048] Focal Loss is a loss function designed to solve the class imbalance problem. In order to solve the sample imbalance problem that the proportion of non-defect area samples is large in container defect detection, a class weight factor is introduced based on the traditional Focal Loss to further reduce the weight of easy-to-classify negative samples, and the improved Focal Loss is obtained. Through the differential setting of the class weight factor , the overfitting of the model to the non-defect samples is effectively avoided.

[0049] The specific formula adopted by the improved Focal Loss is: ; ; ; Wherein, p t is the prediction probability of the model for the t-class sample, γ is the difficulty weight factor, is the class weight factor, log is the logarithmic function with the natural constant as the base, and L Focal is the improved Focal Loss, pr t is the proportion of the number of t-class samples to the total number of samples. The class weight factor is determined by pr t . When the class weight factor is greater than or equal to 0.5, the class weight factor is negatively correlated with pr t , that is, the larger pr t is, the smaller the class weight factor is. When the class weight factor is less than 0.5, the class weight factor is positively correlated with pr t , that is, the larger pr t is, the larger the class weight factor The larger, the more easily classified the t-th class sample is, and the smaller, the more difficultly classified the t-th class sample is. The present application dynamically adjusts the class weight factor of different class samples according to the distribution law of all categories of samples, so as to offset the deviation caused by uneven data distribution. When p t is large, it indicates that the t-th class sample is easily classified, tends to 0, and the weight of the easily classified sample is reduced. When p t is small, it indicates that the t-th class sample is difficult to classify, tends to 1, and the weight of the difficult-to-classify sample is retained, so that the model pays more attention to the difficult-to-classify sample. α and γ jointly act, thereby significantly improving the recognition ability of the model for rare classes and difficult samples. In this embodiment, γ = 2, = 2pr t is taken as an example.

[0050] After improving the YOLO detection head through the design of the detection branch and the optimization of the loss function, the YOLO detection head is used to detect different scale defects of the fused features to obtain multi-scale features. Then, the multi-scale features are post-processed, which includes Non-Maximum Suppression (NMS) and defect coordinate mapping. The Non-Maximum Suppression is to retain the frame with the highest local confidence and remove all the remaining frames with high overlap with it, and output a prediction frame with the highest confidence. The defect coordinate mapping accurately maps the defect position located by the algorithm on the image to the physical position in the real world to match the actual container position. After the Non-Maximum Suppression and the defect coordinate mapping are performed on the multi-scale features, the defect detection result is obtained, which includes the defect type, the defect position, the defect confidence and the defect size. For the defect type, the classification label is defined in the model training stage, and the defect type includes mechanical damage, surface coating defect, sealing defect and accessory defect. The defect position is the specific position of the defect on the image or the container. Each defect detection frame corresponds to a defect confidence, and the defect confidence is the degree of confidence or reliability of the model on the predicted defect. The defect size refers to the size or area of the defect.

[0051] The adaptive attention module is used to extract the defect candidate region feature of the fusion feature, so as to suppress the background interference of the container surface defects, such as stickers and stains on the container surface. Different scale defects of the defect candidate region feature are detected by using a multi-branch detection head to obtain multi-scale features; the multi-scale features are post-processed to obtain a defect detection result; wherein the defect detection result includes a defect type, a defect position, a defect confidence and a defect size. The adaptive attention module solves the problem of background interference such as stickers and stains on the container surface by combining spatial adaptive attention and channel adaptive attention. The multi-branch detection head is improved from two aspects of detection branch design and loss function optimization, and the detection effect of full-scale coverage and high-precision classification is realized, so as to solve the core problems of large scale difference and sample imbalance of the container surface defects. Finally, after non-maximum suppression and defect coordinate mapping and other post-processing, the defect type, defect position, defect confidence and defect size of the container surface are obtained, and the detection accuracy of the container surface defects is greatly improved.

[0052] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that a process, device, article or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such a process, device, article or method. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, device, article or method including the element.

[0053] The above only describes the preferred embodiments of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A container surface defect detection method based on multi-scale feature fusion, characterized in that, The method comprises the following steps: acquiring a first resolution image of a container surface, pre-processing the first resolution image to obtain a pre-processed image; wherein the data pre-processing includes defogging, noise reduction, light balance and image stitching; extracting basic features of the pre-processed image by using a lightweight backbone network; the basic features include shallow features, middle features and deep features; performing multi-scale feature fusion on the basic features to obtain fused features; extracting defect candidate region features of the fused features by using an adaptive attention module to suppress background interference of the container surface defects; detecting different scale defects of the defect candidate region features by using a multi-branch detection head to obtain multi-scale features; performing post-processing on the multi-scale features to obtain a defect detection result; wherein the defect detection result includes defect type, defect position, defect confidence and defect size. 2.The container surface defect detection method based on multi-scale feature fusion according to claim 1, characterized in that, The method of acquiring a first resolution image of a container surface comprises: designing a dynamic downsampling module, wherein the dynamic downsampling module is used for downsampling according to texture complexity; acquiring a first resolution image of a container surface by using a dynamic downsampling module for regional image acquisition in a container detection scene. 3.The container surface defect detection method based on multi-scale feature fusion according to claim 2, characterized in that, The method of acquiring a first resolution image of a container surface by using a dynamic downsampling module for regional image acquisition in a container detection scene comprises: acquiring a container scene image; calculating gradient entropy values of local regions of the container scene image; if the gradient entropy value is greater than a set threshold, performing 2 times downsampling on the local region to obtain a first downsampling result; if the gradient entropy value is less than or equal to the set threshold, performing 4 times downsampling on the local region to obtain a second downsampling result; combining the first downsampling result and the second downsampling result to obtain a first resolution image of a container surface. 4.The container surface defect detection method based on multi-scale feature fusion according to claim 1, characterized in that, The method of extracting basic features of the pre-processed image by using a lightweight backbone network comprises: extracting spatial features of each input channel of the pre-processed image by using a deep convolution; performing cross-channel information fusion on the spatial features by using point-by-point convolution to obtain cross-channel features; performing adaptive suppression on the spatial features and the cross-channel features by using a channel attention gate mechanism to obtain the basic features of the pre-processed image; wherein the channel attention gate mechanism is used for removing redundant channel features. 5.The container surface defect detection method based on multi-scale feature fusion according to claim 1, characterized in that, The method of performing multi-scale feature fusion on the basic features to obtain fused features comprises: performing feature fusion on shallow features and deep features by using horizontal connection and upsampling to obtain small-scale fused features; performing feature fusion on middle features and deep features by using horizontal connection and upsampling to obtain medium-scale fused features; performing feature fusion on middle features and deep features by using horizontal connection and downsampling to obtain large-scale fused features; combining the small-scale fused features, the medium-scale fused features and the large-scale fused features to obtain the fused features. 6.The container surface defect detection method based on multi-scale feature fusion according to claim 1, characterized in that, The method of extracting defect candidate region features of the fused features by using an adaptive attention module comprises: calculating defect probabilities of local regions of the fused features and generating a spatial mask; The fusion features are globally average-pooled, and a multi-layer MLP is used to generate channel weights; the channel weights are used to strengthen channels related to defects and suppress channels unrelated to defects; Attention calculation is performed on the fusion features based on the spatial mask and the channel weights to obtain defect candidate region features of the fusion features. 7.The container surface defect detection method based on multi-scale feature fusion according to claim 1, characterized in that, Before the different scale defects of the defect candidate region features are detected using the multi-branch detection head to obtain multi-scale features, the method further includes: A detection branch is designed for the defect candidate region features; the detection branch includes a small-scale branch, a medium-scale branch, and a large-scale branch; The detection branch is optimized using a loss function to obtain a multi-branch detection head; the loss function includes DIoULoss and an improved Focal Loss, the improved Focal Loss includes a class weight factor and a difficulty weight factor, the class weight factor is used to adjust the uniformity of the class distribution, and the difficulty weight factor is used to adjust the weights of easy and difficult classification samples.

8. The method of claim 6, wherein the method of detecting surface defects of a container based on multi-scale feature fusion is characterized by, The defect probability of the local region of the fusion features is calculated, and a spatial mask is generated, including: The gradient features of different local regions of the training sample set are counted; The gradient features are input into a feature discriminator for probability prediction to obtain defect probabilities of different local regions; If the defect probability is greater than or equal to a defect probability threshold, the spatial mask of the corresponding local region is set to 1; If the defect probability is less than the defect probability threshold, the spatial mask of the corresponding local region is set to 0. 9.The container surface defect detection method based on multi-scale feature fusion according to claim 4, characterized in that, The channel attention gating mechanism is used to adaptively suppress the spatial features and the cross-channel features to obtain the base features of the preprocessed image, including: The base features of the preprocessed image are extracted using the following formula: ; wherein F l is the l-th layer base feature of the pre-processed image, F l-1 is the (l-1)-th layer base feature of the pre-processed image, DepthSepConv is a depth separable convolution, ReLU is a nonlinear activation function, BN is a batch normalization, G l is the channel attention gate weight, W l is the weight of the depth separable convolution, is the output of the multiplexing channel of the base feature of the pre-processed image.

10. The method of claim 5, wherein the method is based on multi-scale feature fusion for container surface defect detection. The lateral connection and up-sampling are used to fuse the shallow features and the deep features to obtain small-scale fusion features, including: The deep features are up-sampled by 4 to obtain first resolution features; A 1×1 convolution kernel is used to convolve the shallow features to obtain first channel dimension features; The first resolution features and the first channel dimension features are spliced to obtain first spliced features; A 3×3 convolution kernel is used to convolve the first spliced features to obtain small-scale fusion features.

Citation Information

Patent Citations

  • Light-weight industrial product surface defect detection method

    CN119313659A

  • Container damage detection method

    CN119359657A

  • Container surface damage detection method and device based on machine vision

    CN120525815A

  • Industrial surface defect detection method based on improved real-time target detection model

    CN120563514A

  • Cross-scale defect detection method based on deep learning

    US20230306577A1

Cited By

  • Container damage detection method based on improved UPAD-YOLO

    CN121074053A

  • An Improved UPAD-YOLO-Based Method for Container Damage Detection

    CN121074053B

  • Industrial product surface defect grading method

    CN122244058A