Metal surface micro-defect detection method and system based on attention enhancement
By introducing the perceptual attention module LCCA into the micro-defect detection method on metal surfaces, the problems of detail dilution and insufficient saliency in micro-defect detection are solved, achieving efficient and real-time micro-defect detection and improving the accuracy and stability of detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG JIAOTONG UNIV
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, the detection methods for micro-defects on metal surfaces suffer from dilution of detail information due to multi-scale feature pyramid downsampling, resulting in unstable localization of micro-defects. Furthermore, the weak contrast and complex texture interference on metal surfaces lead to insufficient defect saliency, resulting in both false detections and missed detections. In addition, the computational and latency introduced by complex enhancement modules are not conducive to online real-time deployment.
The perceptual attention module LCCA is introduced into the backbone and neck network of the neural network model. Through stepwise feature extraction and fusion, the feature expression is enhanced, and a high-resolution detection feature map is generated. Combined with depthwise separable convolution and local, coordinate and channel attention branches, the saliency and localization accuracy of weak defects are improved.
Without increasing computational burden, it improves the detection sensitivity and positioning accuracy of micro-defects, solves the problem that micro-defects on reflective metal surfaces are easily obscured by textures and highlights, and reduces the missed detection rate and false detection rate, making it suitable for real-time detection deployment on production lines.
Smart Images

Figure CN122492628A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection technology, and specifically to a method and system for detecting micro-defects on metal surfaces based on attention enhancement. Background Technology
[0002] With the increasing demands for product quality in industrial manufacturing, the detection of micro-defects on metal surfaces has become a key link in ensuring the quality control of production lines. There are many types of metal surface defects, among which micro-defects are difficult to detect reliably due to their small size and low contrast. At the same time, production lines have extremely high requirements for the real-time performance of the detection system, which needs to complete defect identification and localization within milliseconds. Traditional machine vision methods rely on manually designed features, which are difficult to adapt to the complex and ever-changing metal surface textures and lighting conditions. Deep learning methods, due to their powerful feature representation capabilities, have gradually become the mainstream technical approach.
[0003] Currently, existing methods employ a one-stage target detection network to meet the real-time requirements of production lines. They utilize multi-scale feature pyramids to perform defect prediction at different levels. Commonly used YOLO detectors perform predictions at three pyramid levels with step sizes of 8, 16, and 32. They leverage deeper semantic features to enhance recognition capabilities and fuse feature information at different scales through the feature pyramid structure, achieving a good balance between real-time performance and accuracy in conventional defect detection tasks.
[0004] In the aforementioned existing technologies, conventional prediction layers tend to favor low-resolution feature maps. The small number of pixels occupied by micro-defects are diluted during multiple downsampling processes, resulting in the loss of fine-grained texture and edge information, and unstable localization of micro-defects. At the same time, micro-defects appear as weak contrast and small-scale features on real metal surfaces, which are easily submerged by repetitive textures or specular highlights, resulting in insufficient defect saliency. Although the complex enhancement modules introduced to improve feature representation can improve accuracy, they bring additional computation and latency, which is not conducive to stable online deployment. Summary of the Invention
[0005] To address the shortcomings of existing technologies, the present invention aims to provide a method and system for detecting micro-defects on metal surfaces based on attention enhancement. This addresses the technical problems of existing technologies, such as the dilution of micro-defect detail information and unstable localization of micro-defects due to conventional multi-scale feature pyramid downsampling, insufficient defect saliency due to weak contrast, repetitive textures, and specular highlights on metal surfaces, resulting in both false positives and false negatives. Furthermore, the additional computation and latency introduced by complex enhancement modules to improve accuracy are not conducive to online real-time deployment.
[0006] This invention provides a method for detecting micro-defects on metal surfaces based on attention enhancement, comprising: Image acquisition steps: Acquire image data of the metal surface to be inspected; Model improvement steps: Add the perceptual attention module LCCA to the backbone and neck network of the neural network model to obtain the object detection model; Feature extraction steps: The metal surface image data is input into the backbone network of the target detection model for stepwise feature extraction, and then the feature is enhanced stepwise by the perceptual attention module LCCA to obtain feature images of different scales; Feature fusion steps: The feature images at different scales are input into the neck network of the target detection model for progressive upsampling and fusion, and then enhanced by the perceptual attention module LCCA to obtain an enhanced feature image. The enhanced feature image is further upsampled and fused to obtain a high-resolution detection feature map. The enhanced feature image is then downsampled and fused step by step to obtain a fused feature image. Target detection steps: Input the high-resolution detection feature map and the fused feature image into the detection head of the target detection model to perform target detection and obtain the detection results of micro-defects on the metal surface.
[0007] By adding the perceptual attention module LCCA to the backbone and neck network of the neural network model, and introducing this module for feature enhancement during the stepwise feature extraction and stepwise upsampling fusion process, the saliency of micro-defects on metal surfaces can be effectively improved, making weak, low-contrast defects stand out in feature representation. At the same time, more detailed information is retained by generating high-resolution detection feature maps, and multi-scale detection feature maps are obtained through stepwise downsampling fusion. Thus, without significantly increasing the computational burden, the detection sensitivity and localization accuracy of micro-defects are balanced, solving the problem that micro-defects on reflective metal surfaces are easily obscured by textures and highlights.
[0008] In some embodiments of the present invention, the feature extraction step specifically includes: After convolution processing or feature extraction of the metal surface image data, feature enhancement is performed through the perceptual attention module LCCA to obtain the feature image. ; The feature image After downsampling and feature extraction, the feature image is further enhanced by the perceptual attention module LCCA. ; The feature image After downsampling and feature extraction, the feature image is obtained. ; The feature image After downsampling and feature extraction, semantic feature extraction is then performed using the SPPF and C2fPSA modules to obtain the feature image. .
[0009] By analyzing the feature image during the feature extraction process and feature images Feature enhancement is performed separately using the LCCA (Learning Attention) module, while the feature images are... and feature images By employing downsampling and semantic feature extraction, local contrast and spatial location information can be enhanced in shallow and mid-level features, making the texture edges of weak defects clearer. At the same time, unnecessary computational overhead is avoided in deep features, maintaining the purity of high-level semantic features. This achieves a balance between enhancing the ability to represent micro-defects and maintaining network operating efficiency, which is beneficial for real-time detection deployment on production lines.
[0010] In some embodiments of the present invention, the feature fusion step specifically includes: Upsampling fusion step: The feature image Upsampled and compared with the feature image The images are then stitched together and then integrated using the C3k2 module to obtain the feature image. ; The feature image Upsampled and compared with the feature image After stitching, the features are integrated by the C3k2 module and then enhanced by the LCCA perceptual attention module to obtain the feature image. ; The feature image After upsampling, compared with the feature image The images are stitched together and then integrated using the C3k2 module to obtain a high-resolution detection feature image. .
[0011] By deep feature images After stepwise upsampling, the feature image Feature images Sequentially stitch and fuse them, and then fuse the feature images. By introducing the LCCA (Linguistic Computational Attention) module for enhancement, deep semantic information can be gradually propagated to shallower layers while targeted enhancement of defect candidate regions is performed at the mid-level feature level. This enhancement is then combined with the feature image... Fusion to generate high-resolution detection feature images This allows the details of micro-defects, which occupy only a small number of pixels, to be fully preserved at the prediction end, thereby significantly improving the recall rate of micro-defects and the localization accuracy of the detection box.
[0012] In some embodiments of the present invention, the feature fusion step further includes: Downsampling fusion step: The feature image After downsampling, compared with the feature image The images are then stitched together and then integrated using the C3k2 module to obtain the feature image. ; The feature image After downsampling, compared with the feature image The images are then stitched together and then integrated using the C3k2 module to obtain the feature image. .
[0013] Feature images obtained by upsampling and fusing Perform stepwise downsampling and compare with feature images and feature images By sequentially splicing and fusing features, a bottom-up feature aggregation path is constructed, which can transmit the enhanced information from the middle-layer fused features back to deeper layers, thus improving the deep feature image. and feature images To obtain richer contextual information, and at the same time due to feature images Having already been enhanced by the perception attention module, this downsampling fusion path can further improve the recognition ability of large-scale and medium-scale defects without introducing high-resolution noise, thus enhancing the adaptability of the detection model to defects of different sizes.
[0014] In some embodiments of the present invention, the target detection step specifically includes: The feature image The high-resolution detection feature image The feature image and the feature image The data is input into the detection head of the target detection model for target detection, and the detection results of micro-defects on the metal surface are obtained.
[0015] By feature image High-resolution detection feature images Feature images and feature images The common inputs are fed into the detection head for target detection, enabling the detection head to simultaneously obtain high-resolution detail features, mid-level enhancement features, and deep semantic features, among which the high-resolution detection feature image... Responsible for capturing the fine structure and feature images of minute defects. Provides attention-enhanced mid-level representations, feature images and feature images It is responsible for large-scale contextual reasoning, and the features of the four scales complement each other, so as to achieve full-scale coverage detection from small defects to large-area defects in complex textures and reflective backgrounds of metal surfaces, significantly reducing the false negative rate and false positive rate.
[0016] In some embodiments of the present invention, the perceptual attention module LCCA includes a local contrast branch, which is used to perform depth-separable convolutional feature extraction on the input feature image input to the perceptual attention module LCCA to obtain local contrast features.
[0017] By designing the local contrast branch in the perceptual attention module LCCA to extract features from the input feature image using depthwise separable convolution, local texture change information can be extracted efficiently while maintaining spatial resolution. Depthwise separable convolution decomposes standard convolution into channel-wise convolution and pointwise convolution, significantly reducing the number of parameters and computation. At the same time, the local contrast features output by this branch can highlight the edges and fine-grained structures of micro-defects, enhancing the contrast between the defect area and the background. This effectively improves the visual salience of weak defects in complex scenes such as reflective metal surfaces.
[0018] In some embodiments of the present invention, the perceptual attention module LCCA further includes a coordinate attention branch, which is used to perform aggregate encoding on the input feature image along the width direction and the height direction respectively to obtain width attention weights and height attention weights, and combine the width attention weights and the height attention weights to obtain coordinate attention weights.
[0019] By designing the coordinate attention branch in the perceptual attention module LCCA to aggregate and encode the input feature image along the width and height directions respectively, width attention weights and height attention weights can be generated. This is lighter than the traditional global attention mechanism, while retaining the ability to perceive direction. This allows the network to accurately capture the positional information of defects in the horizontal and vertical directions. The coordinate attention weights formed by combining the spatial weights of the two directions can selectively enhance or suppress the spatial position in the feature image, thereby more accurately locating micro-defects in complex texture backgrounds.
[0020] In some embodiments of the present invention, the perceptual attention module LCCA further includes a channel attention branch, which is used to perform global average pooling and convolution processing on the input feature image to obtain channel attention weights.
[0021] By designing the channel attention branch in the perceptual attention module LCCA to perform global average pooling and convolution on the input feature image, the spatial dimension can be compressed into a channel description vector. Then, the dependency relationship between channels can be established through one-dimensional convolution. Compared with the fully connected layer, this method has fewer parameters and can maintain the channel order information. The generated channel attention weights can adaptively recalibrate the importance of each channel, so that the feature channels that contribute more to the identification of micro-defects are amplified, while irrelevant or interfering channels are suppressed. This improves the purity and discrimination ability of defect features in scenarios with a lot of interference, such as reflective metal surfaces.
[0022] In some embodiments of the present invention, the perceptual attention module LCCA further includes a fusion convolution module, which is used to perform element-wise modulation on the input feature image based on the coordinate attention weights and the channel attention weights, obtain a modulated feature image, fuse the modulated feature image with the local contrast features, and then perform convolution processing to obtain an output feature image.
[0023] By designing the fusion convolution module in the perceptual attention module (LCCA) to modulate the input feature image element-wise based on coordinate attention weights and channel attention weights before fusing it with local contrast features, the synergistic enhancement of three different dimensions of features—spatial location, channel relationship, and local texture—is achieved. Coordinate attention provides orientation-sensitive position weights, channel attention provides adaptive recalibration of feature channels, and local contrast features provide edge and fine-grained information. After the three are fused and then processed by convolution, a feature representation with richer information and more significant defects can be generated, thereby comprehensively improving the detection performance of micro-defects with minimal computational overhead.
[0024] Some embodiments of the present invention further provide an attention-enhanced metal surface micro-defect detection system, comprising: The image acquisition module acquires image data of the metal surface to be inspected; The model improvement module adds the perceptual attention module LCCA to the backbone and neck network of the neural network model to obtain the target detection model; The feature extraction module inputs the metal surface image data into the backbone network of the target detection model for stepwise feature extraction, and then performs stepwise feature enhancement through the perception attention module LCCA to obtain feature images of different scales. The feature fusion module inputs the feature images at different scales into the neck network of the target detection model for progressive upsampling and fusion, and then performs feature enhancement through the perceptual attention module LCCA to obtain an enhanced feature image. The enhanced feature image is then further upsampled and fused to obtain a high-resolution detection feature map. Finally, the enhanced feature image is downsampled and fused stepwise to obtain a fused feature image. The target detection module inputs the high-resolution detection feature map and the fused feature image into the detection head of the target detection model to perform target detection and obtain the detection results of micro-defects on the metal surface.
[0025] By integrating the image acquisition module, model improvement module, feature extraction module, feature fusion module, and object detection module into a complete detection system, each module operates collaboratively in the logical order of hierarchical feature extraction, attention enhancement, multi-scale fusion, and high-resolution detection, enabling the method to be stably executed in a real hardware environment. The feature extraction module is responsible for extracting multi-scale features from image data, the feature fusion module completes upsampling and downsampling fusion and generates a high-resolution detection feature map, and the target detection module finally outputs the detection results. The data flow between the modules is clear and the functions are well-defined, thus transforming the proposed detection method into an engineering-deployable system solution, which is easy to integrate into the visual inspection equipment of the production line, improving the operational stability and deployment convenience of the inspection system. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 A schematic flowchart of an attention-enhanced metal surface micro-defect detection method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a perception attention module (LCCA) provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a target detection model provided in an embodiment of the present invention; Figure 4 A flowchart illustrating a feature fusion step provided in an embodiment of the present invention; Figure 5 A schematic diagram of a high-resolution detection branch provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of an attention-enhanced metal surface micro-defect detection system provided in an embodiment of the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application. It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. Furthermore, it should be understood that the terms “comprising” and “having”, and any variations thereof, are intended to cover a non-exclusive inclusion, for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus. Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other. The technical solution of the present invention will be described in detail below with reference to specific embodiments and accompanying drawings.
[0028] like Figure 1 As shown, the present invention provides a method for detecting micro-defects on metal surfaces based on attention enhancement, comprising: Image acquisition step S1: Acquire image data of the metal surface to be detected; Furthermore, surface images of metal sheets or metal parts on the production line are captured using industrial cameras. The images are RGB color images or grayscale images with a size of 640×640 pixels. During the acquisition process, a ring light source or a strip light source is set above the metal surface to provide uniform illumination and reduce the shading effect of specular highlights on the defect area. For continuous production line scenarios, a line scan camera or area scan camera is used in conjunction with an encoder triggering method to achieve online image acquisition synchronized with the production line speed, ensuring that the acquired metal surface image data has consistent resolution and lighting conditions. The acquired image data is transmitted to the image processing unit via a digital interface, serving as input data for subsequent feature extraction steps.
[0029] Model improvement step S2: Add the perceptual attention module LCCA to the backbone and neck network of the neural network model to obtain the target detection model; Furthermore, the neural network model uses YOLOv11 as the baseline model. The YOLOv11 model consists of an input terminal, a backbone network, a neck network, and a detection head. The backbone network consists of four feature extraction stages, which output feature maps with step sizes of 4, 8, 16, and 32, respectively. The perceptual attention module (LCCA) is embedded in the first feature extraction stage of the backbone network to enhance shallow features; and embedded in the second feature extraction stage to enhance mid-level features. The neck network performs multi-scale feature interaction through top-down and bottom-up feature fusion paths; in the neck network, the perceptual attention module LCCA is embedded in the top-down feature fusion path to enhance the fused mid-layer features; Through the above improvements, a target detection model is obtained. The target detection model also includes a newly added high-resolution detection head, which is used to receive feature maps with a step size of 4 for small defect prediction.
[0030] In some embodiments, such as Figure 2 As shown, the Perceptual Attention Module (LCCA) includes a Local Contrast Branch, which is used to extract depthwise separable convolutional features from the input feature image to the LCCA module to obtain local contrast features.
[0031] By designing the local contrast branch in the perceptual attention module LCCA to extract features from the input feature image using depthwise separable convolution, local texture change information can be extracted efficiently while maintaining spatial resolution. Depthwise separable convolution decomposes standard convolution into channel-wise convolution and pointwise convolution, significantly reducing the number of parameters and computation. At the same time, the local contrast features output by this branch can highlight the edges and fine-grained structures of micro-defects, enhancing the contrast between the defect area and the background. This effectively improves the visual salience of weak defects in complex scenes such as reflective metal surfaces.
[0032] In some embodiments, the perceptual attention module LCCA further includes a coordinate attention branch, which is used to perform aggregate encoding on the input feature image along the width and height directions respectively to obtain width attention weights and height attention weights, and to combine the width attention weights and height attention weights to obtain coordinate attention weights.
[0033] By designing the coordinate attention branch in the perceptual attention module LCCA to aggregate and encode the input feature image along the width and height directions respectively, width attention weights and height attention weights can be generated. This is lighter than the traditional global attention mechanism, while retaining the ability to perceive direction. This allows the network to accurately capture the positional information of defects in the horizontal and vertical directions. The coordinate attention weights formed by combining the spatial weights of the two directions can selectively enhance or suppress the spatial position in the feature image, thereby more accurately locating micro-defects in complex texture backgrounds.
[0034] In some embodiments, the perceptual attention module LCCA further includes a channel attention branch, which is used to perform global average pooling and convolution on the input feature image to obtain channel attention weights.
[0035] By designing the channel attention branch in the perceptual attention module LCCA to perform global average pooling and convolution on the input feature image, the spatial dimension can be compressed into a channel description vector. Then, the dependency relationship between channels can be established through one-dimensional convolution. Compared with the fully connected layer, this method has fewer parameters and can maintain the channel order information. The generated channel attention weights can adaptively recalibrate the importance of each channel, so that the feature channels that contribute more to the identification of micro-defects are amplified, while irrelevant or interfering channels are suppressed. This improves the purity and discrimination ability of defect features in scenarios with a lot of interference, such as reflective metal surfaces.
[0036] In some embodiments, the perceptual attention module LCCA further includes a fusion convolution module, which modulates the input feature image element by element based on coordinate attention weights and channel attention weights. After obtaining the modulated feature image, the modulated feature image is fused with the local contrast features and then convolved to obtain the output feature image.
[0037] By designing the fusion convolution module in the perceptual attention module (LCCA) to modulate the input feature image element-wise based on coordinate attention weights and channel attention weights before fusing it with local contrast features, the synergistic enhancement of three different dimensions of features—spatial location, channel relationship, and local texture—is achieved. Coordinate attention provides orientation-sensitive position weights, channel attention provides adaptive recalibration of feature channels, and local contrast features provide edge and fine-grained information. After the three are fused and then processed by convolution, a feature representation with richer information and more significant defects can be generated, thereby comprehensively improving the detection performance of micro-defects with minimal computational overhead.
[0038] Furthermore, the perceptual attention module LCCA includes a local contrast branch, a coordinate attention branch, a channel attention branch, and a fusion convolution module; Batch size is The number of channels is Height is Width is Input feature image After being input into the LCCA (Local Comparison Attention) module, the data is processed through the local contrast branch, coordinate attention branch, and channel attention branch, respectively. In the local contrast branch, the input feature image is processed by a 3×3 kernel, with the number of groups equal to the number of channels. Depth-separable convolution is used for feature extraction, so that spatial convolution is performed independently on each input channel of the input feature image to preserve the spatial resolution of the input feature image and extract local texture change information; After feature extraction via depthwise separable convolution, the features are then normalized and nonlinearly transformed by batch normalization layers and SiLU activation functions to obtain local contrast features. Local contrast features are used to highlight the edge structure and fine-grained texture of micro-defects on metal surfaces. Their expression is:
[0039] in, This indicates a depthwise separable convolution with a kernel size of 3×3. This indicates a batch normalization operation. Indicates the SiLU activation function; The input feature image is aggregated and encoded along the width and height directions in the coordinate attention branch, respectively; When aggregating along the width direction, the mean value of the height dimension is taken to obtain the feature vector along the width direction. Its expression is:
[0040] in, This is an operation to take the mean value for the height dimension; When aggregating along the height direction, the height direction feature vector is obtained by taking the mean value of the width dimension. Its expression is:
[0041] in, This is an operation to take the average value over the width dimension; Width-direction feature vector After a single 1×1 convolution for channel dimensionality reduction, followed by a nonlinear transformation using the SiLU activation function, and then a final 1×1 convolution for channel dimensionality enhancement, the width attention weights are obtained, expressed as follows:
[0042] in, This is a 1×1 convolution operation; Height-direction feature vector After a single 1×1 convolution for channel dimensionality reduction, followed by a nonlinear transformation using the SiLU activation function, and then a final 1×1 convolution for channel dimensionality enhancement, the high attention weights are obtained, expressed as follows:
[0043] Pay attention to width With high attention weight After feature concatenation, the coordinate attention weights are obtained by normalization and activation using the Sigmoid activation function. Its expression is:
[0044] in, For feature splicing operations; Use the Sigmoid activation function; In the channel attention branch, the input feature image first undergoes global average pooling in the height and width dimensions to compress the spatial dimensions into channel description vectors. After dimensional rearrangement, a one-dimensional convolution with a kernel size of 3 is used to establish local dependencies between channels. Finally, a sigmoid activation function is used to generate channel attention weights. Its expression is:
[0045] in, This is a one-dimensional convolution operation; This is a global average pooling operation; The fusion convolution module will be based on coordinate attention weights With channel attention weights For input feature image Element-wise modulation is performed to obtain a modulation feature image, thereby achieving dual-gated modulation in terms of spatial location and channel dimension; Modulation feature image and local contrast features Element-wise addition is performed to fuse local texture enhancement information; finally, channel blending and nonlinear transformation are applied using a 1×1 convolution, batch normalization layer, and SiLU activation function to obtain the output feature image. Its expression is:
[0046] Through the above processing, the perceptual attention module LCCA can simultaneously achieve orientation-sensitive spatial suppression and enhancement, channel adaptive recalibration, and superimposed local texture contrast enhancement with relatively small computational overhead, making micro-defects more easily highlighted against reflective texture backgrounds.
[0047] Feature extraction step S3: After the metal surface image data is input into the backbone network of the target detection model for stepwise feature extraction, it is then enhanced stepwise by the perceptual attention module LCCA to obtain feature images at different scales. In some embodiments, feature extraction step S3 specifically includes: After convolution processing or feature extraction of the metal surface image data, feature enhancement is performed through the perceptual attention module LCCA to obtain the feature image. ; Feature Image After downsampling and feature extraction, the feature image is further enhanced by the perceptual attention module LCCA. ; Feature Image After downsampling and feature extraction, the feature image is obtained. ; Feature Image After downsampling and feature extraction, semantic feature extraction is then performed using the SPPF and C2fPSA modules to obtain the feature image. .
[0048] By analyzing the feature image during the feature extraction process and feature images Feature enhancement is performed separately using the LCCA (Learning Attention) module, while the feature images are... and feature images By employing downsampling and semantic feature extraction, local contrast and spatial location information can be enhanced in shallow and mid-level features, making the texture edges of weak defects clearer. At the same time, unnecessary computational overhead is avoided in deep features, maintaining the purity of high-level semantic features. This achieves a balance between enhancing the ability to represent micro-defects and maintaining network operating efficiency, which is beneficial for real-time detection deployment on production lines.
[0049] Furthermore, such as Figure 3 As shown, the metal surface image data For a 640×640 resolution image, extract the metal surface image data. The data is input into the object detection model, first entering the backbone network of layers 0-12. The backbone network is divided into four stages, with stage 1 consisting of layers 0-3, containing metal surface image data. After convolution or feature extraction, local contrast enhancement and channel coordinate attention recalibration are performed in L3 using the perceptual attention module LCCA, resulting in a feature image with a stride of 4 and a resolution of 160×160. Feature Image The expression can be:
[0050] in, For feature extraction operations; Feature enhancement operations are performed via the LCCA (Learning-Carrying-Attention) module. Phase 2 consists of layers 4-6, which process the feature images. After performing downsampling convolution with a stride of 2 and feature extraction, local contrast enhancement and channel coordinate attention recalibration are performed in L6 via the perceptual attention module LCCA, resulting in a feature image with a stride of 8 and a resolution of 80×80. Its expression is:
[0051] in, This is a downsampling operation; Stage 3 consists of layers 7-8, which process the feature images. After performing downsampling convolution with a stride of 2 and feature extraction, a feature image with a stride of 16 and a resolution of 40×40 is obtained. Its expression is:
[0052] Stage 4 consists of layers 9-12, which process the feature images. After performing downsampling convolution with a stride of 2 and feature extraction, spatial pyramid pooling is applied sequentially through the SPPF module at L11, followed by self-attention enhancement through the C2fPSA module at L12 to obtain the feature image. Its expression is:
[0053] in, For spatial pyramid pooling operations via the SPPF module; This is for self-attention enhancement operations via the C2fPSA module. Feature fusion step S4: Input feature images of different scales into the neck network in the target detection model for progressive upsampling and fusion, and then perform feature enhancement through the perceptual attention module LCCA to obtain enhanced feature images. Continue to upsample and fuse the enhanced feature images to obtain high-resolution detection feature maps. Perform progressive downsampling and fusion on the enhanced feature images to obtain fused feature images. In some embodiments, such as Figure 4 As shown, the feature fusion step S4 is specifically as follows: Upsampling fusion step S41: Combine feature images Upsampling and feature image The images are then stitched together and then integrated using the C3k2 module to obtain the feature image. ; Feature Image Upsampling and feature image After concatenation, the features are integrated by the C3k2 module and then enhanced by the LCCA perceptual attention module to obtain the feature image. ; Feature Image After upsampling and feature image The images are stitched together and then integrated using the C3k2 module to obtain a high-resolution detection feature image. .
[0054] By deep feature images After stepwise upsampling, the feature image Feature images Sequentially stitch and fuse them, and then fuse the feature images. By introducing the LCCA (Linguistic Computational Attention) module for enhancement, deep semantic information can be gradually propagated to shallower layers while targeted enhancement of defect candidate regions is performed at the mid-level feature level. This enhancement is then combined with the feature image... Fusion to generate high-resolution detection feature images This allows the details of micro-defects, which occupy only a small number of pixels, to be fully preserved at the prediction end, thereby significantly improving the recall rate of micro-defects and the localization accuracy of the detection box.
[0055] Furthermore, the feature image output by the backbone network Feature images Feature images Feature images Features are fed into the neck network, which includes a top-down path in layers 13-19, a bottom-up path in layers 23-28, and a high-resolution detection branch. In the top-down path of the neck network, the feature image output by L12 is... The image enters the L13 upsampling module, undergoes a 2x upsampling, increasing the resolution from 20×20 to 40×40, and is then compared with the feature image. The image is stitched along the channel dimension, and the stitched result is then used for feature integration by the C3k2 module of L15 to obtain a feature image with a stride of 16 and a resolution of 40×40. Its expression is:
[0056] in, This is a 2x upsampling operation; For splicing operations; For feature integration operations performed via the C3k2 module; Feature Image The image enters the L16 upsampling module, undergoes a 2x upsampling, increasing the resolution from 40×40 to 80×80, and is then compared with the feature image output from L6. After concatenation along the channel dimension in L17, the concatenated result is integrated by the C3k2 module in L18. Then, in L19, the perceptual attention module LCCA performs local contrast enhancement and channel coordinate attention recalibration to obtain a feature image with a stride of 8 and a resolution of 80×80. Its expression is:
[0057] Feature Image The image enters the L20 upsampling module, undergoes a 2x upsampling, increasing the resolution from 80×80 to 160×160, and is then compared with the feature image. After being stitched together along the channel dimension by L21, the stitched result is then integrated using the C3k2 module of L22 to obtain a high-resolution detection feature image with a stride of 4 and a resolution of 160×160. High-resolution detection feature images The output of the high-resolution detection branch is the high-resolution detection feature image. It was not fed back to the bottom-up path; High-resolution detection feature images The expression is:
[0058] like Figure 5 As shown, the high-resolution detection branch is located at the end of the neck network and connected to the detection head, assuming a feature image from the neck network. for Feature images derived from the backbone network and enhanced by the LCCA feature enhancement module of the perceptual attention module. Features are First of all Perform a 2x upsampling and align to 160×160. Then and splicing and fusion at the channel dimension High-resolution detection feature images are obtained by feature integration using the C3k2 module. ; It should be noted that by employing a no-feedback strategy, the high-resolution detection branch is used only as input to the detection head for prediction, with no feedback to the bottom-up path. The bottom-up path draws from the feature image. start; The design philosophy of the no-feedback strategy is that while high-resolution features retain more details, they are also more likely to carry metal surface texture noise. If this noise is forcibly fed back into the bottom-up path, it may cause noise to spread and affect the stability of deep semantics. By using the above structural constraints, we can improve the accuracy of micro-defect recall and location while maintaining a simple fusion link and controllable inference latency, making it more suitable for online real-time detection deployment on production lines.
[0059] In some embodiments, the feature fusion step S4 further includes: Downsampling fusion step S42: Combine feature images After downsampling and feature image The images are then stitched together and then integrated using the C3k2 module to obtain the feature image. ; Feature Image After downsampling and feature image The images are then stitched together and then integrated using the C3k2 module to obtain the feature image. .
[0060] Feature images obtained by upsampling and fusing Perform stepwise downsampling and compare with feature images and feature images By sequentially splicing and fusing features, a bottom-up feature aggregation path is constructed, which can transmit the enhanced information from the middle-layer fused features back to deeper layers, thus improving the deep feature image. and feature images To obtain richer contextual information, and at the same time due to feature images Having already been enhanced by the perception attention module, this downsampling fusion path can further improve the recognition ability of large-scale and medium-scale defects without introducing high-resolution noise, thus enhancing the adaptability of the detection model to defects of different sizes.
[0061] Furthermore, the bottom-up path starts with L19, and in the bottom-up path, L19 outputs the feature image. As the initial input to the path, it enters the 3×3 convolutional downsampling module of L23 with a stride of 2. After a downsampling convolution with a stride of 2, its resolution is reduced from 80×80 to 40×40, and then compared with the aforementioned feature image. The image is stitched along the channel dimension, and the stitched result is then processed by the C3k2 module of L25 for feature integration, resulting in an enhanced feature image with a stride of 16 and a resolution of 40×40. Its expression is:
[0062] in, This is a downsampling convolution operation; Feature Image The image enters the 3×3 convolutional downsampling module of L26 with a stride of 2. After a downsampling convolution with a stride of 2, its resolution is reduced from 40×40 to 20×20, and then compared with the feature image output from L12. After concatenation along the channel dimension by L27, the concatenated result is used for feature integration by the C3k2 module of L28 to obtain a feature image with a stride of 32 and a resolution of 20×20. Its expression is:
[0063] It should be noted that the bottom-up path starts from the feature image. Initially, no high-resolution detection feature images were included. This avoids the propagation and diffusion of high-resolution shallow texture noise into deep semantic features, thus ensuring the stability of deep features.
[0064] Target detection step S5: Input the high-resolution detection feature map and the fused feature image into the detection head of the target detection model to perform target detection and obtain the detection results of micro-defects on the metal surface.
[0065] In some embodiments, the target detection step S5 specifically comprises: Feature Image High-resolution detection feature images Feature images and feature images The data is input into the detection head of the target detection model for target detection, and the detection results of micro-defects on the metal surface are obtained.
[0066] By feature image High-resolution detection feature images Feature images and feature images The common inputs are fed into the detection head for target detection, enabling the detection head to simultaneously obtain high-resolution detail features, mid-level enhancement features, and deep semantic features, among which the high-resolution detection feature image... Responsible for capturing the fine structure and feature images of minute defects. Provides attention-enhanced mid-level representations, feature images and feature images It is responsible for large-scale contextual reasoning, and the features of the four scales complement each other, so as to achieve full-scale coverage detection from small defects to large-area defects in complex textures and reflective backgrounds of metal surfaces, significantly reducing the false negative rate and false positive rate.
[0067] Furthermore, the high-resolution detection feature image output by L22 L19 output feature image L25 output feature image L28 output feature image The four-scale detection head of L29 is fed into the target, and the detection head processes the four-scale features to complete the target classification and bounding box regression prediction, and finally outputs the detection results. Among them, high-resolution detection feature images Corresponding to a 160×160 high-resolution detection input, feature image Corresponding to 80×80 detection input, feature image Corresponding to a 40×40 detection input, feature image Corresponds to 20×20 detection input.
[0068] Based on the above detection method, by adding the perceptual attention module LCCA to the backbone and neck network of the neural network model, and introducing this module for feature enhancement during the stepwise feature extraction and stepwise upsampling fusion process, the salience of micro-defects on the metal surface can be effectively improved, making weak and low-contrast defects stand out in feature expression. At the same time, by generating high-resolution detection feature maps, more detailed information is retained, and multi-scale detection feature maps are obtained through stepwise downsampling fusion. Thus, without significantly increasing the computational burden, the detection sensitivity and localization accuracy of micro-defects are taken into account, and the problem that micro-defects on reflective metal surfaces are easily buried by texture and highlights is solved.
[0069] Through the above detection methods, the perceptual attention module LCCA can enhance the salience of weak contrast defects and suppress background interference with almost no increase in the number of parameters. At the same time, the high-resolution detection branch improves the recall and localization quality of micro-defects through higher resolution predicted feature images. The combination of the two enables the system to achieve better overall performance in micro-defect detection tasks. Experimental results show that on the NEU-DET benchmark dataset, the target detection model with added perceptual attention module LCCA and high-resolution detection branch achieves a mean average accuracy (mAP50) of 75.7% across all classes, which is an improvement over the 74.3% of the YOLOv11 model. At the same time, it remains within a deployable range in terms of billions of floating-point operations per second (GFLOPs) and the number of parameters. Ablation experiments show that the combination of the perceptual attention module LCCA and the high-resolution detection branch can bring about an overall improvement in accuracy and detection performance, while keeping end-to-end latency and throughput within an acceptable range for real-time detection. Furthermore, in the analysis of test results on the GC10-DET dataset, compared with the YOLOv11 model, the target detection model with the addition of the perceptual attention module LCCA and the high-resolution detection branch showed a more adequate response to small, low-contrast defects. For defects that are small in size, have weak boundaries, or are in complex textured backgrounds, the target detection model can provide a predicted bounding box that is closer to the real defect area, reducing false detections caused by missed detections, bounding box offsets, and background interference. Especially in areas with gradual grayscale changes, surface reflections, or repetitive texture interference, the improved target detection model can still highlight defect targets relatively stably, demonstrating better detection robustness and localization consistency. In comparative experiments on the NEU-DET dataset, the target detection model showed a consistent trend of improvement in detecting various typical defects such as crazing, inclusion, patches, pitted surface, rolled-in scale, and scratches. In particular, when the defect scale is small, the texture noise is strong, or the contrast between the target and the background is weak, the improved model is more likely to output complete, compact, and more accurately positioned target boxes, and the overall prediction confidence is more stable. Based on the combined quantitative indicators and test results, it can be seen that by enhancing the expression of weak contrast features through the perception attention module LCCA and retaining high-resolution detail information through the high-resolution detection branch, the accuracy, stability and engineering deployability of micro-defect detection in complex metal surface scenarios are improved.
[0070] like Figure 6 As shown, this embodiment of the invention also provides an attention-enhanced metal surface micro-defect detection system, comprising: The image acquisition module acquires image data of the metal surface to be inspected; The model improvement module adds the perceptual attention module LCCA to the backbone and neck network of the neural network model to obtain the target detection model; The feature extraction module inputs the metal surface image data into the backbone network of the target detection model for stepwise feature extraction, and then performs stepwise feature enhancement through the perception attention module LCCA to obtain feature images of different scales. The feature fusion module inputs the feature images at different scales into the neck network of the target detection model for progressive upsampling and fusion, and then performs feature enhancement through the perceptual attention module LCCA to obtain an enhanced feature image. The enhanced feature image is then further upsampled and fused to obtain a high-resolution detection feature map. Finally, the enhanced feature image is downsampled and fused stepwise to obtain a fused feature image. The target detection module inputs the high-resolution detection feature map and the fused feature image into the detection head of the target detection model to perform target detection and obtain the detection results of micro-defects on the metal surface.
[0071] By integrating the image acquisition module, model improvement module, feature extraction module, feature fusion module, and object detection module into a complete detection system, each module operates collaboratively in the logical order of hierarchical feature extraction, attention enhancement, multi-scale fusion, and high-resolution detection, enabling the method to be stably executed in a real hardware environment. The feature extraction module is responsible for extracting multi-scale features from image data, the feature fusion module completes upsampling and downsampling fusion and generates a high-resolution detection feature map, and the target detection module finally outputs the detection results. The data flow between the modules is clear and the functions are well-defined, thus transforming the proposed detection method into an engineering-deployable system solution, which is easy to integrate into the visual inspection equipment of the production line, improving the operational stability and deployment convenience of the inspection system.
[0072] It should be noted that the above is a reference method and system for detecting micro-defects on metal surfaces based on attention enhancement, and the present invention is not limited thereto.
[0073] The embodiments of the present invention achieve the enhancement of the saliency of weak contrast defects through a lightweight contrast perception attention module without significantly increasing the computational burden, and retain the detailed information of micro-defects through a high-resolution detection branch. This solves the technical problems in the prior art, such as the dilution of micro-defect details, unstable localization, and insufficient saliency and coexistence of false positives and false negatives caused by weak contrast and reflection interference.
[0074] Finally, it should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can still be made to the specific implementation of the present invention or equivalent substitutions can be made to some technical features without departing from the spirit of the technical solutions of the present invention, and all such modifications and substitutions should be covered within the scope of the technical solutions claimed in the present invention.
Claims
1. A method for detecting micro-defects on metal surfaces based on attention enhancement, characterized in that, include: Image acquisition steps: Acquire image data of the metal surface to be inspected; Model improvement steps: Add the perceptual attention module LCCA to the backbone and neck network of the neural network model to obtain the object detection model; Feature extraction steps: The metal surface image data is input into the backbone network of the target detection model for stepwise feature extraction, and then the feature is enhanced stepwise by the perceptual attention module LCCA to obtain feature images of different scales; Feature fusion steps: The feature images at different scales are input into the neck network of the target detection model for progressive upsampling and fusion, and then enhanced by the perceptual attention module LCCA to obtain an enhanced feature image. The enhanced feature image is further upsampled and fused to obtain a high-resolution detection feature map. The enhanced feature image is then downsampled and fused step by step to obtain a fused feature image. Target detection steps: Input the high-resolution detection feature map and the fused feature image into the detection head of the target detection model to perform target detection and obtain the detection results of micro-defects on the metal surface.
2. The method for detecting micro-defects on metal surfaces based on attention enhancement according to claim 1, characterized in that, The feature extraction steps are as follows: After convolution processing or feature extraction of the metal surface image data, feature enhancement is performed through the perceptual attention module LCCA to obtain the feature image. ; The feature image After downsampling and feature extraction, the feature image is further enhanced by the perceptual attention module LCCA. ; The feature image After downsampling and feature extraction, the feature image is obtained. ; The feature image After downsampling and feature extraction, semantic feature extraction is then performed using the SPPF and C2fPSA modules to obtain the feature image. .
3. The method for detecting micro-defects on metal surfaces based on attention enhancement according to claim 2, characterized in that, The feature fusion step is specifically as follows: Upsampling fusion step: The feature image Upsampled and compared with the feature image The images are then stitched together and then integrated using the C3k2 module to obtain the feature image. ; The feature image Upsampled and compared with the feature image After stitching, the features are integrated by the C3k2 module and then enhanced by the LCCA perceptual attention module to obtain the feature image. ; The feature image After upsampling, compared with the feature image The images are stitched together and then integrated using the C3k2 module to obtain a high-resolution detection feature image. .
4. The attention-enhanced metal surface micro-defect detection method according to claim 3, characterized in that, The feature fusion step further includes: Downsampling fusion step: The feature image After downsampling, compared with the feature image The images are then stitched together and then integrated using the C3k2 module to obtain the feature image. ; The feature image After downsampling, compared with the feature image The images are then stitched together and then integrated using the C3k2 module to obtain the feature image. .
5. The attention-enhanced metal surface micro-defect detection method according to claim 4, characterized in that, The target detection steps are as follows: The feature image The high-resolution detection feature image The feature image and the feature image The data is input into the detection head of the target detection model for target detection, and the detection results of micro-defects on the metal surface are obtained.
6. The method for detecting micro-defects on metal surfaces based on attention enhancement according to any one of claims 1 to 5, characterized in that, The perceptual attention module LCCA includes a local contrast branch, which is used to extract depthwise separable convolutional features from the input feature image input to the perceptual attention module LCCA to obtain local contrast features.
7. The method for detecting micro-defects on metal surfaces based on attention enhancement according to claim 6, characterized in that, The perceptual attention module LCCA further includes a coordinate attention branch, which is used to perform aggregation encoding on the input feature image along the width and height directions respectively to obtain width attention weights and height attention weights, and then combine the width attention weights and the height attention weights to obtain coordinate attention weights.
8. The method for detecting micro-defects on metal surfaces based on attention enhancement according to claim 7, characterized in that, The perceptual attention module LCCA also includes a channel attention branch, which is used to perform global average pooling and convolution on the input feature image to obtain channel attention weights.
9. The method for detecting micro-defects on metal surfaces based on attention enhancement according to claim 8, characterized in that, The perceptual attention module LCCA further includes a fusion convolution module, which is used to modulate the input feature image element by element based on the coordinate attention weights and the channel attention weights. After obtaining the modulated feature image, the modulated feature image is fused with the local contrast features, and then convolutional processing is performed to obtain the output feature image.
10. A metal surface micro-defect detection system based on attention enhancement, characterized in that, include: The image acquisition module acquires image data of the metal surface to be inspected; The model improvement module adds the perceptual attention module LCCA to the backbone and neck network of the neural network model to obtain the target detection model; The feature extraction module inputs the metal surface image data into the backbone network of the target detection model for stepwise feature extraction, and then performs stepwise feature enhancement through the perception attention module LCCA to obtain feature images of different scales. The feature fusion module inputs the feature images at different scales into the neck network of the target detection model for progressive upsampling and fusion, and then performs feature enhancement through the perceptual attention module LCCA to obtain an enhanced feature image. The enhanced feature image is then further upsampled and fused to obtain a high-resolution detection feature map. Finally, the enhanced feature image is downsampled and fused stepwise to obtain a fused feature image. The target detection module inputs the high-resolution detection feature map and the fused feature image into the detection head of the target detection model to perform target detection and obtain the detection results of micro-defects on the metal surface.