Lightweight metal surface defect detection method and equipment
By improving the feature analysis network of the YOLOv8 model and combining dynamic convolution and ghost bottleneck structure, the problem of high false detection rate of metal surface defect detection model under noise and background interference is solved, and lightweight model and high-precision detection are achieved.
Patent Information
- Application Number
- CN202510951671.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-11-14
AI Technical Summary
Existing metal surface defect detection models struggle to distinguish between real defects and background interference under noise and background conditions, resulting in a high false detection rate. Furthermore, the model size needs to be optimized to accommodate edge device deployments.
An improved YOLOv8 model feature analysis module is adopted, combined with a dynamic convolution module and a ghost bottleneck structure. By preprocessing images to reduce noise, dynamically adjusting convolution parameters, and enhancing adaptability to defect morphology, the positioning accuracy is improved by optimizing the detection head.
While making the model lightweight, it improves the accuracy of metal surface defect detection, reduces the amount of computation and parameters, enhances the adaptability to complex backgrounds and multi-scale defects, and reduces missed detections and false detections.
Smart Images

Figure CN120953176A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of metal defect detection technology, and in particular to a lightweight metal surface defect detection method and equipment. Background Technology
[0002] With the rapid development of intelligent technology, the introduction of neural networks to replace or assist manual metal defect detection in the industrial manufacturing field has become a trend.
[0003] In particular, due to the presence of a large amount of noise (such as textures, scratches, uneven lighting, etc.) in the metal surface defect detection environment, the background interference is significant, making it difficult for existing metal surface detection models to distinguish between real defects and background interference, leading to an increased false detection rate (e.g., misclassifying metal surface textures as defects). Furthermore, in order to quickly deploy the model to intelligent metal surface defect detection equipment, the size of existing metal surface defect detection models also needs to be optimized.
[0004] Therefore, how to ensure the accuracy of metal surface defect detection while achieving lightweight model has become an urgent problem to be solved. Summary of the Invention
[0005] The main objective of this application is to provide a lightweight metal surface defect detection method and device, aiming to solve the technical problem of how to ensure the accuracy of metal surface defect detection while achieving model lightweighting.
[0006] To achieve the above objectives, this application proposes a method for detecting surface defects in lightweight metals, the method comprising:
[0007] The source metal surface image is preprocessed to obtain the processed metal image;
[0008] The processed metal image is subjected to multiple feature analysis by an improved feature analysis network to obtain the target defect features; the improved feature analysis network is a YOLOv8 model feature analysis module based on dynamic convolution module and ghost bottleneck structure improvement;
[0009] The abnormal coordinates corresponding to the target metal surface defect are determined based on the optimized detection head and the target defect features, and the target metal surface defect is marked in the processed metal image based on the abnormal coordinates.
[0010] In one embodiment, the improved feature analysis network is a YOLOv8 model feature analysis module improved based on dynamic convolution modules, ghost bottleneck structures, and average pooling downsampling convolution modules.
[0011] In one embodiment, the improved feature analysis network includes an improved backbone network and an improved neck network. The improved backbone network is obtained by replacing the standard convolutional modules in the backbone network of the traditional YOLOv8 network model with average pooling downsampling convolutional modules, and replacing each C2f module with a C2f-GhostDynamicConv module improved based on dynamic convolutional modules and ghost bottleneck structures. The improved neck network is obtained by replacing the standard convolutional modules in the neck network of the traditional YOLOv8 network model with average pooling downsampling convolutional modules, and replacing each C2f module with a C2f-GhostDynamicConv module improved based on dynamic convolutional modules and ghost bottleneck structures.
[0012] In one embodiment, prior to performing multiple feature analysis via the improved feature analysis network, the method further includes:
[0013] In the traditional YOLOv8 network model, the CBS module of each C2f module is replaced with the dynamic convolution module DynamicConv, and the bottleneck module Bottleneck of each C2f module is replaced with the ghost bottleneck structure GhostBottleneck to obtain the C2f-GhostDynamicConv module.
[0014] In one embodiment, the dynamic convolution module DynamicConv includes an attention generation layer, a multiple convolutional filter, a batch normalization layer, and an activation function. The attention generation layer is used to generate dynamic attention parameters corresponding to the first input feature map. The multiple convolutional filter is used to perform weighted attention convolution on the input feature map based on the dynamic attention parameters to obtain multiple defect features. The batch normalization layer is used to normalize the multiple defect features to obtain standard defect features. The activation function is used to smooth the standard defect features to generate a first output feature map.
[0015] In one embodiment, the average pooling downsampling convolutional module includes an average pooling processing layer, a convolutional feature layer, a max pooling feature layer, and a concatenation layer. The average pooling processing layer is used to perform average pooling on the second input feature map and to perform feature equalization along the channel dimension to generate a first average feature and a second average feature. The convolutional feature layer is used to extract features from the first average feature to obtain a first feature map. The max pooling feature layer is used to extract features from the second average feature using max pooling to obtain a second feature map. The concatenation layer is used to concatenate the first feature map and the second feature map along the channel dimension to generate a second output feature map.
[0016] In one embodiment, the optimized detection head includes an efficient multi-scale convolutional module, a merged convolutional layer, a bounding box convolutional layer, and a classification convolutional layer; the step of determining the anomaly coordinates corresponding to the target metal surface defect based on the optimized detection head and the target defect features includes:
[0017] The target defect features are analyzed using the efficient multi-scale convolution module to obtain multi-scale defect features.
[0018] The multi-scale defect features are merged by the merging convolutional layer to obtain fused defect features.
[0019] The bounding box loss value is obtained by performing convolution operations based on the fusion defect features through the bounding box convolutional layer.
[0020] The classification convolutional layer performs convolution operations based on the fused defect features to obtain a classification loss value; the anomaly coordinates corresponding to the target metal surface defect are determined based on the bounding box loss value and the classification loss value.
[0021] In one embodiment, the improved neck network is obtained by adding a coordinate attention module between the spatial pyramid pooling layer (SPPF) in the backbone network module of the conventional YOLOv8 network model and the splicing layer in the neck network of the conventional YOLOv8 network model.
[0022] In one embodiment, before preprocessing the source metal surface image to obtain the processed metal image, the method further includes:
[0023] Convert the model files corresponding to the improved feature analysis network and optimized detection head into target weight files;
[0024] Configure acceleration and optimization parameters according to target detection requirements;
[0025] Based on the acceleration optimization parameters, the target weight file is converted into an optimized inference engine, which is used to accelerate the inference process of the improved feature analysis network and the optimized detection head.
[0026] In addition, to achieve the above objectives, this application also proposes a lightweight metal surface defect detection device, the device including: a memory, a processor, and a lightweight metal surface defect detection program stored in the memory and capable of running on the processor, the lightweight metal surface defect detection program being configured to implement the steps of the lightweight metal surface defect detection method as described above.
[0027] This application provides a lightweight method and device for detecting metal surface defects. The method includes preprocessing a source metal surface image to obtain a processed metal image; performing multi-feature analysis on the processed metal image using an improved feature analysis network to obtain target defect features; the improved feature analysis network is a YOLOv8 model feature analysis module based on a dynamic convolution module and a ghost bottleneck structure; determining the anomaly coordinates corresponding to the target metal surface defect based on the optimized detection head and target defect features, and annotating the target metal surface defect in the processed metal image based on the anomaly coordinates. This application attempts to maintain detection performance while lightweighting and redesigning the network structure of the YOLOv8 model using efficient convolutional techniques. First, the image is preprocessed to reduce noise and background interference. Then, the dynamic convolution module in the improved feature analysis network adaptively generates convolution kernel weights to enhance adaptability to defect morphology by dynamically adjusting convolution parameters. Furthermore, efficient feature analysis is performed lightweightly using a ghost bottleneck structure, thereby enhancing adaptability to complex backgrounds and multi-scale defects, reducing false negatives and false positives, and lowering network computational load. Finally, by combining optimized detection head and target defect features, the localization accuracy is improved, enhancing the detection accuracy of minute defects, thereby effectively improving the precision of metal surface defect detection. Therefore, this application significantly reduces computation and parameter count through a unique feature generation process, while possessing higher-level feature extraction capabilities and better multi-scale adaptability. It can capture richer semantic and contextual information, effectively achieving both model lightweighting and ensuring the accuracy of metal surface defect detection. Attached Figure Description
[0028] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0029] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 This is a schematic flowchart of the first embodiment of the lightweight metal surface defect detection method of this application;
[0031] Figure 2 This is a schematic diagram of the Ghost module structure in the first embodiment of the lightweight metal surface defect detection method of this application;
[0032] Figure 3 This is a schematic diagram of the ghost bottleneck structure in the first embodiment of the lightweight metal surface defect detection method of this application.
[0033] Figure 4 This is a schematic diagram of the dynamic convolution module in the first embodiment of the lightweight metal surface defect detection method of this application;
[0034] Figure 5 This is a schematic diagram of the structural improvement of the C2f module in the first embodiment of the lightweight metal surface defect detection method of this application;
[0035] Figure 6 This is a schematic diagram of the average pooling downsampling convolution module in the first embodiment of the lightweight metal surface defect detection method of this application;
[0036] Figure 7 This is a schematic diagram of the coordinate attention module in the first embodiment of the lightweight metal surface defect detection method of this application;
[0037] Figure 8 This is a schematic flowchart of the second embodiment of the lightweight metal surface defect detection method of this application;
[0038] Figure 9 This is a schematic diagram of the improved detection head structure in the second embodiment of the lightweight metal surface defect detection method of this application;
[0039] Figure 10 This is a schematic diagram of the structure of the high-efficiency multi-scale convolution module in the second embodiment of the lightweight metal surface defect detection method of this application;
[0040] Figure 11 This is a schematic diagram of the GAC-YOLO model for lightweight metal surface defect detection, representing the second embodiment of the lightweight metal surface defect detection method of this application.
[0041] Figure 12 This is a schematic diagram of the model acceleration process in the second embodiment of the lightweight metal surface defect detection method of this application;
[0042] Figure 13 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the lightweight metal surface defect detection method in this application embodiment.
[0043] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0044] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0045] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0046] The main solution of this application is as follows: preprocessing the source metal surface image to obtain a processed metal image; performing multiple feature analysis on the processed metal image through an improved feature analysis network to obtain the target defect features; the improved feature analysis network is a YOLOv8 model feature analysis module based on dynamic convolution modules and ghost bottleneck structures; determining the abnormal coordinates corresponding to the target metal surface defect based on the optimized detection head and the target defect features, and annotating the target metal surface defect in the processed metal image based on the abnormal coordinates.
[0047] Currently, complex industrial environments present significant background interference, and metal surfaces exhibit both minute and large-scale defects. Existing defect detection models suffer from limitations in localization and target perception. Furthermore, the actual deployment of these models faces challenges such as limited computing resources, storage space constraints, and energy consumption. To achieve effective real-time automated metal defect detection and rapid deployment of models to intelligent metal surface defect detection equipment, the size of existing metal surface defect detection models also needs optimization.
[0048] To address this issue, this application attempts to redesign the YOLOv8 model's network structure with a lightweight approach, employing efficient convolutional techniques while maintaining detection performance. First, the application preprocesses images to reduce noise and background interference. Then, it improves the dynamic convolution module in the feature analysis network to adaptively generate convolutional kernel weights, enhancing adaptability to defect morphology through dynamic adjustment of convolutional parameters. Furthermore, it utilizes a lightweight ghost bottleneck structure for efficient feature analysis, thereby enhancing adaptability to complex backgrounds and multi-scale defects, reducing false negatives and false positives, and lowering network computational cost. Finally, it combines optimized detection head and target defect features to improve localization accuracy and enhance the detection accuracy of minute defects, thus effectively improving the precision of metal surface defect detection. Therefore, this application significantly reduces computational cost and parameter count through a unique feature generation process, while possessing higher-level feature extraction capabilities and better multi-scale adaptability. It can capture richer semantic and contextual information, effectively achieving both model lightweighting and high precision in metal surface defect detection.
[0049] It should be noted that the executing entity in this embodiment can be a lightweight metal surface defect detection system, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a lightweight metal surface defect detection device capable of performing the above functions. This embodiment does not specifically limit it in this way. The following uses a lightweight metal surface defect detection device (hereinafter referred to as the detection device) as the executing entity to describe this embodiment and the following embodiments.
[0050] Based on this, embodiments of this application provide a method for detecting defects on lightweight metal surfaces, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the lightweight metal surface defect detection method of this application.
[0051] In this embodiment, the method for detecting defects on the surface of lightweight metals includes steps S10 to S40:
[0052] Step S10: Preprocess the source metal surface image to obtain a processed metal image;
[0053] Step S20: Perform multiple feature analysis on the processed metal image through an improved feature analysis network to obtain the target defect features; the improved feature analysis network is a YOLOv8 model feature analysis module based on dynamic convolution modules and ghost bottleneck structures.
[0054] It is understood that the aforementioned source metal surface image can be a raw metal surface image directly acquired in an industrial setting, which may contain complex backgrounds, noise, and defects of different scales (such as micro-defects and large-sized defects) on the target metal surface. This embodiment preprocesses the source metal surface image by performing operations such as standardization, resizing (typically 640x640), and data augmentation (such as random flipping and brightness adjustment) to generate a processed metal image with improved image quality and a unified data format suitable for YOLOv8 model input. This facilitates subsequent high-precision defect detection based on the processed metal image.
[0055] It's important to understand that the YOLOv8 model, as one of the foundational models in the YOLO series, offers better performance in both accuracy and speed compared to traditional anchor-based detection methods due to its anchor-free strategy. Therefore, this embodiment preferentially uses the traditional YOLOv8 model as the foundational model for metal surface defect detection. However, in complex industrial environments with significant background interference and numerous minute and large-scale defects on metal surfaces, the traditional YOLOv8 model still suffers from issues such as localization errors and insufficient target perception.
[0056] To address this issue, this embodiment enhances the ability of the traditional YOLOv8 model to capture irregular defects and multi-scale features by using dynamic convolution modules and ghost bottleneck structures, thereby proposing the GAC-YOLO algorithm for high-precision metal surface defect detection.
[0057] It is understandable that existing models face challenges in deploying on edge devices, including limited computing resources, storage space constraints, and energy consumption. To achieve accurate detection of metal surface defects on edge devices with limited computing resources, this embodiment further optimizes the convolutional module in the YOLOv8 model feature analysis module to construct a lighter network. Therefore, in a feasible implementation, the improved feature analysis network in this embodiment is a YOLOv8 model feature analysis module improved based on dynamic convolutional modules, ghost bottleneck structures, and average pooling downsampling convolutional modules.
[0058] It is easy to understand that the feature analysis module of the traditional YOLOv8 model consists of a backbone network for basic feature extraction and a neck network for feature fusion. This results in redundant computation, leading to excessive model size and computational cost. To effectively improve model accuracy and optimize model size, this embodiment can optimize both simultaneously. Therefore, in a feasible implementation, the improved feature analysis network in this embodiment includes an improved backbone network and an improved neck network.
[0059] The improved backbone network is obtained by replacing the standard convolutional modules in the backbone network of the traditional YOLOv8 network model with average pooling downsampling convolutional modules, and replacing each C2f module with a C2f-GhostDynamicConv module improved based on dynamic convolutional modules and ghost bottleneck structure.
[0060] The improved neck network is obtained by replacing the standard convolutional modules in the neck network of the traditional YOLOv8 network model with average pooling downsampling convolutional modules, and replacing each C2f module with a C2f-GhostDynamicConv module based on dynamic convolutional modules and a ghost bottleneck structure.
[0061] It is important to understand that, in order to achieve accurate detection of metal surface defects on edge devices with limited computing resources, this embodiment proposes the aforementioned GhostDynamicConv module, which is an improvement based on dynamic convolution modules and a ghost bottleneck structure, to improve C2f and optimize the model's detection performance. Specifically, the dynamic convolution module can adaptively generate convolution kernel weights through input features, enhancing its adaptability to defect morphologies; the Ghost module included in the ghost bottleneck structure can generate redundant feature maps through low-cost linear operations, reducing computational load.
[0062] Therefore, the GhostDynamicConv module combines the advantages of the Ghost module and the dynamic convolution module. It selectively fuses multiple convolution kernels based on the characteristics of the input data and dynamically adjusts the weights, effectively improving the ability to identify and represent features related to metal surface defects. Simultaneously, the GhostDynamicConv module effectively addresses the challenges posed by small-target defects, defects with significant scale variations, and complex backgrounds in metal surface detection tasks. This module optimizes computational resources while improving processing efficiency under low-precision floating-point operations, thus significantly enhancing the feature extraction capabilities of the C2f module in the backbone and neck networks.
[0063] In one feasible implementation, this embodiment may further include step A1 before step S20:
[0064] Step A1: Replace the CBS module of each C2f module in the traditional YOLOv8 network model with the dynamic convolution module DynamicConv, and replace the bottleneck module Bottleneck of each C2f module with the ghost bottleneck structure GhostBottleneck to obtain the C2f-GhostDynamicConv module.
[0065] It is easy to understand that, as can be seen from the above analysis, the C2f-GhostDynamicConv module mainly improves each C2f module in the traditional YOLOv8 network model (i.e., each C2f module in the backbone network and each C2f module in the neck network of the traditional YOLOv8 network model) by incorporating the Ghost module and the dynamic convolution module.
[0066] The Ghost module can generate more defect-related feature maps with low-cost operations. It utilizes efficient methods such as depthwise separable convolution to significantly reduce the number of parameters and computational overhead. For its detailed structure, please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a schematic diagram of the Ghost module structure in the first embodiment of the lightweight metal surface defect detection method of this application.
[0067] like Figure 2 As shown, the Ghost module can first utilize a 1×1 convolution (i.e., Figure 2 In the context of conv), the input features (i.e. Figure 2 The input features are aggregated across channels to gather key information features, and then group convolution (i.e., ...) is used. Figure 2The method uses Φ1 to Φk in the diagram to generate new feature maps. This method decomposes traditional convolution into two steps: the first step uses regular convolution to generate feature maps with fewer channels, which has relatively low computational requirements; based on this, the second step uses low-cost operations such as group convolution to reduce the amount of computation and generate new feature maps. Finally, the feature maps generated in the two steps (i.e., the feature maps after processing Φ1 to Φk in the diagram and the Identity after processing 1×1 convolution) are merged into the final output.
[0068] Furthermore, this embodiment can utilize the advantages of the Ghost module to design the aforementioned GhostBottleneck structure, which is structurally similar to the basic residual block in ResNet. For specific effects, please refer to... Figure 3 , Figure 3 This is a schematic diagram of the ghost bottleneck structure in the first embodiment of the lightweight metal surface defect detection method of this application.
[0069] like Figure 3 As shown, the ghost bottleneck structure proposed in this embodiment typically consists of two cascaded Ghost modules (i.e., Figure 2 The network consists of two layers: a first-layer Ghost module and a second-layer Ghost module. The first Ghost module is responsible for channel expansion, increasing the number of feature channels to improve the model's expressive power. The second Ghost module performs channel compression, matching the output dimension to the skip connection path. The two modules establish a direct connection between input and output through skip connections. Each network layer can be followed by Batch Normalization (BN) and ReLU activation functions for non-linear processing. Figure 3 The ghost bottleneck structure shown in (a) is mainly used in scenarios with a step size of 1.
[0070] When handling cases where the stride is set to 2, feature downsampling can be achieved by embedding a depthwise separable convolution (DWConv) with a stride of 2 between the two Ghost modules. Figure 3 (b) shows the ghost bottleneck structure. In comparison, Figure 3 The ghost bottleneck structure shown in (a) can effectively improve computational efficiency.
[0071] Therefore, to address the high computational cost and inability to dynamically adapt to input features of the traditional C2f module's Bottleneck structure, this embodiment can replace the bottleneck module Bottleneck of each C2f module with, for example... Figure 3 The GhostBottleneck structure shown represents a lightweight structural modification of the C2f module.
[0072] Furthermore, to improve the feature extraction accuracy of the C2f module, this embodiment can replace the CBS module with the DynamicConv dynamic convolution module. In one feasible implementation, the DynamicConv dynamic convolution module includes an attention generation layer, multiple convolutional filters, a batch normalization layer, and an activation function.
[0073] The attention generation layer is used to generate dynamic attention parameters corresponding to the first input feature map;
[0074] The multiple convolutional filter is used to perform weighted attention convolution on the input feature map based on the dynamic attention parameters to obtain multiple defect features;
[0075] The batch normalization layer is used to normalize the multiple defect features to obtain standard defect features;
[0076] The activation function is used to smooth the standard defect features and generate a first output feature map.
[0077] It's important to understand that, to address the issue of traditional C2f modules' static convolutional kernels failing to adapt to changes in input features, leading to missed detections of complex defects, this embodiment introduces a dynamic convolutional module to enhance feature representation capabilities. The core idea is to dynamically generate or fine-tune convolutional kernels based on input features. Dynamic convolution can adaptively adjust kernel weights according to the characteristics of the input image, not only better capturing detailed defect features but also effectively handling environmental changes, angle variations, and occlusion issues, thus improving detection robustness in various industrial environments.
[0078] Understandably, referring to Figure 4 The construction of the dynamic convolution module will be explained. Figure 4 This is a schematic diagram of the dynamic convolution module in the first embodiment of the lightweight metal surface defect detection method of this application. Figure 4 As shown, the DynamicConv module can include an attention generation layer (such as...) Figure 4 Attention), multiple convolutional filters (such as...) Figure 4 k convolutional filters π1~π k ), batch normalization layer after dynamic convolution (e.g. Figure 4 The BN and activation functions shown are as follows: Figure 4 The activation shown can be a ReLU activation function. Figure 4 The first input feature shown Figure X It is the feature map input to the DynamicConv dynamic convolution module in the C2f-GhostDynamicConv module for feature analysis, and Figure 4 The first output feature map Y shown is the feature map after feature analysis by the DynamicConv module.
[0079] like Figure 4 As shown, in the dynamic convolution module, a set of k convolutional kernels with the same size and number of channels are first created for a certain layer, forming the aforementioned multiple convolutional filter. All inputs share the same set of convolutional kernels. These convolutional kernels are then processed by their respective attention weights generated by the attention generation layer. The parameters are combined to form the dynamic convolution kernel parameters of this layer, where "*" represents a multiplication operation. Figure 4 The attention box marked with a dashed line on the left side details the attention generation layer. The calculation process of (x) is as follows: First, a global average pooling layer (Avg pool) is used to process the first input feature. Figure X To capture global spatial features, the features are then mapped to k dimensions through two fully connected (FC) layers. Normalization is performed using ReLU and softmax to generate k attention weights, which are then assigned to the k convolutional kernels of the layer.
[0080] Therefore, DynamicConv dynamically adjusts the weights of each convolution kernel based on each input sample, and the final convolution kernel parameters are a weighted combination of the static convolution kernels. That is, dynamic convolution is based on dynamic perception, and its basic principle can be explained by the following formulas (1)-(3):
[0081]
[0082] In the formula, X represents the first input feature map, and Y represents the first output feature map. It can be observed that X requires two different operations: the first operation is to solve for the attention mechanism parameters used to generate the dynamic convolution kernel, and the second operation is the actual convolution operation.
[0083] Where g is an activation function (such as ReLU), used to perform nonlinear transformations on the results of dynamic convolution; The weight matrix of the k-th static convolution kernel is a predefined set of parameters used for the combination of dynamic convolution kernels; For the k-th static bias term, and Correspondingly, these are also predefined parameters; and For dynamically generated weights and biases, by and Through attention weight π k (X) is obtained by weighted combination, which varies with the input X to achieve dynamic convolution; π k (X) represents the attention weight of the k-th linear function, and represents the input X to the k-th convolution kernel. and bias The degree of dependence is calculated precisely as π by the attention mechanism. k (X) = Softmax(W ak T T(X)+b k This weight changes with different inputs x.
[0084] Among them, W ak b represents the weight parameters of the k-th branch in the attention mechanism, used to map the input feature T(X) to the weight space; k Here, T(X) represents the bias parameter of the k-th branch in the attention mechanism; T(X) represents the first input feature. Figure X The representation after some feature transformation (such as linear projection or pooling).
[0085] It should be understood that this embodiment ultimately proposes a C2f-GhostDynamicConv structure to replace the original C2f structure in the YOLOv8 model, based on the principle of dynamic convolution modules and combined with the design concept of ghost bottleneck structures. The newly designed structure can be referred to... Figure 5 , Figure 5 This is a schematic diagram of the structural improvement of the C2f module in the first embodiment of the lightweight metal surface defect detection method of this application.
[0086] like Figure 5 As shown, Figure 5 (b) The C2f-Ghost-DynamicConv Module shown in Figure 5 (a) Two important improvements were made to the traditional YOLOv8 network model's C2f Module, namely, the first step was to... Figure 5 (a) The traditional convolutional layer CBS module (ConvBNSILU) is replaced with Figure 5 (b) DynamicConv, through its mechanism of adaptively generating convolution kernels based on input features, effectively improves feature representation capabilities. Introducing dynamic convolution does not significantly increase computational cost, and it can dynamically adjust parameters under different input conditions to better adapt to varying target shapes.
[0087] The second step is to Figure 5 (a) The original bottleneck module Bottleneck, as shown, is replaced with Figure 5The GhostBottleneck structure shown in (a) utilizes the low-cost linear operations and depthwise separable convolutions provided by the Ghost module to generate redundant feature maps, significantly reducing the number of network parameters and computational overhead. This module branches the features output from the backbone network and feeds them into multiple layers of GhostBottleneck, then fuses the features from these branches, and finally enhances and integrates them through the final DynamicConv module. Through this series of designs, C2f-Ghost-DynamicConv effectively improves its adaptability to small target defects and complex backgrounds while maintaining detection accuracy, balancing lightweight design with high performance.
[0088] Furthermore, in a feasible implementation, in this embodiment, the average pooling downsampling convolution module includes an average pooling processing layer, a convolutional feature layer, a max pooling feature layer, and a splicing layer;
[0089] The average pooling layer is used to perform average pooling on the second input feature map and to perform feature equalization along the channel dimension to generate a first average feature and a second average feature.
[0090] The convolutional feature layer is used to extract features from the first average feature to obtain a first feature map.
[0091] The max-pooling feature layer is used to extract max-pooling features from the second average feature to obtain the second feature map;
[0092] The splicing layer is used to splice the first feature map and the second feature map along the channel dimension to generate a second output feature map.
[0093] Understandably, traditional YOLOv8 models often use convolutional layers for downsampling when extracting features. However, this module only uses convolutional operations, which cannot effectively capture the features of small targets, leading to the loss of original image information for some minor defects and resulting in missed detections. This embodiment replaces the last three traditional convolutional operations in the backbone network and the two traditional convolutional operations in the neck network of the traditional YOLOv8 network model with an Average Pooling Down Sampling (ADown) convolutional module.
[0094] It should be noted that, referring to Figure 6 It can be seen that, Figure 6 This is a schematic diagram of the average pooling downsampling convolution module in the first embodiment of the lightweight metal surface defect detection method of this application. The average pooling downsampling convolution module includes an average pooling processing layer, a convolution feature layer, a max pooling feature layer, and a splicing layer.
[0095] Among them, the downsampling module ADown can first process the average pooling layer in the average pooling layer (such as...) Figure 6 In AvgPool2d) for the second input feature map (e.g. Figure 6 The input c1) is subjected to a 2×2 average pooling operation with a stride of 1, which preserves more spatial details, can capture finer-grained features, and prevents the loss of small target features during downsampling; then the feature map size is reduced by half, and a separation layer (such as...) is used. Figure 6 The chunk is divided into two parts along the channel dimension, with the number of channels in each part halved, to generate the first average feature X1(c1 / 2) and the second average feature X2(c1 / 2), thereby reducing the number of parameters in a single path so that the two parts of features can be processed differently in the future, thus improving the feature expressive power.
[0096] Then, in the convolutional feature layer, the segmented first average feature X1 is directly convolved using a 3×3 convolution kernel (k=3) with Conv1 (stride of 2 (s=2), padding of 1 (p=1), used for local feature extraction) to extract local features and appropriately downsample to reduce computation, thus obtaining the first feature map. The second average feature X2 is first max-pooled in the max-pooling feature layer to retain the maximum value in the local region and highlight the salient features of small targets. Then, it is further extracted using a 1×1 convolution operation with Conv2 (stride of 1, padding of 0, to capture more abstract features), thus obtaining the second feature map.
[0097] Finally, the two sub-feature maps are concatenated along the channel dimension using the Contact layer to collaboratively utilize local and global features, generating a second output feature map (such as...). Figure 6 Output c2 in the network can provide richer context and detailed information for subsequent networks.
[0098] Throughout the process, the average pooling downsampling convolutional module enhances feature diversity through pooling dimensionality reduction, segmentation, and diversified convolutional operations, and then integrates features through concatenation to improve the output. This effectively addresses issues in metal surface defect detection, such as detail loss during downsampling, inability to capture important target features like edges, corners, and textures, and limitations on the model's understanding of larger receptive fields. Therefore, this embodiment uses the average pooling downsampling convolutional module instead of the traditional convolutional modules in the backbone and neck networks of the traditional YOLOv8 network model, which not only reduces the false negative rate of small defects but also effectively reduces the number of parameters.
[0099] It is easy to understand that, to compensate for the accuracy loss caused by the lightweight model improvement, this embodiment can also enhance the improved neck network's ability to extract global defect features by introducing a coordinate attention mechanism into the neck network of the traditional YOLOv8 network model. Furthermore, in a feasible implementation, the improved neck network is obtained by adding a coordinate attention module between the Spatial Pyramid Pooling Layer (SPPF) in the backbone network module of the traditional YOLOv8 network model and the splicing layer in the neck network of the traditional YOLOv8 network model.
[0100] Understandably, in complex industrial scenarios, defects are diverse, vary greatly in scale, and contain a lot of redundant information, all of which can impact detection accuracy. Attention mechanisms can improve the feature extraction capabilities of neural networks, allowing them to focus more on key features and reduce interference from other features. Therefore, this embodiment introduces a Coordinate Attention (CA) mechanism to capture cross-channel information, including orientation and position information. This helps the model more accurately locate and identify target defects, while reducing interference from other input features and improving target detection accuracy. The structure of the coordinate attention model can be found in [reference needed]. Figure 7 , Figure 7 This is a schematic diagram of the coordinate attention module in the first embodiment of the lightweight metal surface defect detection method of this application.
[0101] like Figure 7 As shown, the coordinate attention module can first perform global average pooling processing in the horizontal and vertical directions on the input features (Input, C×H×W) respectively (corresponding to...). Figure 7 The X Avg Pool and Y Avg Pool are used to generate a dual-path feature description vector with orientation-aware characteristics. This operation not only preserves the target's position information in space, but also provides a direction-aware foundation for subsequent attention calculations.
[0102] Subsequently, the feature vectors in the horizontal and vertical directions are concatenated (corresponding to...) Figure 7 After concat, a shared 1×1 convolution (corresponding to...) Figure 7 After processing with Conv2d, it is then batch-redirected to one level (corresponding to...) Figure 7 BatchNorm and non-linear activation functions (corresponding to) Figure 7 The non-linear method in the model enables cross-channel information interaction, which preserves the key features of spatial direction while integrating the dependencies between channels.
[0103] Subsequently, the intermediately generated feature map is further split into feature vectors in both horizontal and vertical directions (corresponding to...). Figure 7The feature vectors (with dimensions CxHx1 and Cx1xW) are each subjected to 1×1 convolutions to adjust their dimensions, and then each is used to generate independent attention weight maps through the Sigmoid function. This process enables the model to independently learn the distribution pattern of spatial importance for different directions.
[0104] Finally, the attention weight maps in the horizontal and vertical directions are multiplied element-wise with the original input feature map to dynamically enhance the response of the target region while suppressing background noise, thereby improving the model's accuracy in locating the target.
[0105] Therefore, given the large number of tiny defects in the metal surface defect dataset and the practical need for lightweight network design, this embodiment introduces the aforementioned lightweight CA attention mechanism module after the last layer of the backbone network module of the traditional YOLOv8 network model (i.e., between the Spatial Pyramid Pooling Layer (SPPF) in the backbone network module of the traditional YOLOv8 network model and the splicing layer in the neck network of the traditional YOLOv8 network model) to improve the detection performance of metal surface defects without increasing network complexity.
[0106] Step S30: Determine the abnormal coordinates corresponding to the target metal surface defect based on the optimized detection head and the target defect features, and mark the target metal surface defect in the processed metal image based on the abnormal coordinates.
[0107] It should be noted that the aforementioned target defect features can be key features (such as edges, textures, and shapes) extracted through an improved feature analysis network that characterize metal surface defects. The optimized detection head can be a YOLOv8 detection head that has been pre-trained iteratively to improve the detection accuracy for small targets and defects in complex backgrounds. Therefore, this embodiment can accurately detect the coordinates of target metal surface defects in the processed metal image, i.e., the aforementioned anomalous coordinates, based on high-precision target defect features and the optimized detection head. Boundary boxes are then drawn or positions are marked in the processed metal image based on these anomalous coordinates, allowing users to intuitively and efficiently observe metal surface defects through annotation.
[0108] In this embodiment, the C2f modules of the backbone and neck networks of the traditional YOLOv8 model are redesigned, and the C2f modules of the model are improved by combining dynamic convolution and ghost bottleneck structure to construct the C2f-GhostDynamicConv module. This module significantly reduces the amount of computation and parameters through a unique feature generation process, while having a higher level of feature extraction capability and better multi-scale adaptability. It can capture richer semantic and contextual information, and ensure the accuracy of metal surface defect detection while achieving model lightweighting.
[0109] In addition, this embodiment replaces the traditional convolutional modules of the backbone and neck networks of the traditional YOLOv8 model with Adown downsampling convolutional modules to improve the detection effect of small defects, while reducing the number of model parameters;
[0110] Finally, this embodiment can also enhance the ability to extract global defect features by introducing a coordinate attention mechanism into the backbone network, making up for the accuracy loss caused by the lightweight improvement, and further improving the accuracy of metal surface defect detection.
[0111] This embodiment provides a lightweight metal surface defect detection method, which includes: preprocessing a source metal surface image to obtain a processed metal image; performing multi-feature analysis on the processed metal image using an improved feature analysis network to obtain target defect features; the improved feature analysis network is a YOLOv8 model feature analysis module improved based on a dynamic convolution module, a ghost bottleneck structure, and an average pooling downsampling convolution module; determining the anomaly coordinates corresponding to the target metal surface defect based on the optimized detection head and the target defect features, and annotating the target metal surface defect in the processed metal image based on the anomaly coordinates. The improved feature analysis network is a YOLOv8 model feature analysis module improved based on a dynamic convolution module, a ghost bottleneck structure, and an average pooling downsampling convolution module. The improved feature analysis network includes an improved backbone network and an improved neck network. The improved backbone network is obtained by replacing the standard convolutional modules in the backbone network of the traditional YOLOv8 network model with average pooling downsampling convolutional modules, and replacing each C2f module with a C2f-GhostDynamicConv module improved based on dynamic convolutional modules and a ghost bottleneck structure. The improved neck network is obtained by replacing the standard convolutional modules in the neck network of the traditional YOLOv8 network model with average pooling downsampling convolutional modules, and replacing each C2f module with a C2f-GhostDynamicConv module improved based on dynamic convolutional modules and a ghost bottleneck structure. Additionally, this embodiment can also replace the CBS module of each C2f module in the traditional YOLOv8 network model with the dynamic convolutional module DynamicConv, and replace the bottleneck module Bottleneck of each C2f module with the ghost bottleneck structure GhostBottleneck to obtain the C2f-GhostDynamicConv module. This embodiment redesigns the C2f modules of the backbone and neck networks of the traditional YOLOv8 model, improving the C2f module by combining dynamic convolution and ghost bottleneck structures to construct the C2f-GhostDynamicConv module. This module significantly reduces the computational cost and parameter count through a unique feature generation process, while possessing higher-level feature extraction capabilities and better multi-scale adaptability. It can capture richer semantic and contextual information, ensuring the accuracy of metal surface defect detection while achieving model lightweighting. In addition, this embodiment replaces the traditional convolutional modules of the backbone and neck networks of the traditional YOLOv8 model with Adown downsampling convolutional modules to improve the detection effect of small defects, while reducing the number of model parameters.
[0112] This embodiment also discloses an improved neck network obtained by adding a coordinate attention module between the Spatial Pyramid Pooling Layer (SPPF) in the backbone network module of the traditional YOLOv8 network model and the splicing layer in the neck network of the traditional YOLOv8 network model. This embodiment can also enhance the ability to extract global defect features by introducing a coordinate attention mechanism into the backbone network, compensating for the accuracy loss caused by the lightweight improvement, and further improving the accuracy of metal surface defect detection.
[0113] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the first embodiment described above can be referred to the above description, and will not be repeated hereafter.
[0114] Based on the first embodiment, please refer to Figure 8 , Figure 8 This is a flowchart illustrating the second embodiment of the lightweight metal surface defect detection method of this application. The optimized detection head includes a high-efficiency multi-scale convolution module, a merged convolutional layer, a bounding box convolutional layer, and a classification convolutional layer. In this embodiment, step S30 further includes steps S31 to S35:
[0115] Step S31: Perform multi-scale feature analysis on the target defect features using the efficient multi-scale convolution module to obtain multi-scale defect features;
[0116] Step S32: The multi-scale defect features are merged through the merged convolutional layer to obtain fused defect features;
[0117] Step S33: Perform convolution operation on the bounding box convolutional layer based on the fusion defect features to obtain the bounding box loss value;
[0118] Step S34: The classification convolutional layer performs convolution operations based on the fusion defect features to obtain the classification loss value;
[0119] Step S35: Determine the anomaly coordinates corresponding to the target metal surface defect based on the bounding box loss value and the classification loss value.
[0120] It's easy to understand that the detector head is a crucial component of the YOLOv8 model, accounting for nearly half of the total network parameters. To further improve model efficiency, this embodiment modifies the detector head of the traditional YOLOv8 model. (Refer to...) Figure 9 It can be seen that, Figure 9 This is a schematic diagram of the improved detection head structure in the second embodiment of the lightweight metal surface defect detection method of this application. Figure 9As shown in (a), the original YOLOv8 model's detection head employs a decoupled head structure, using a parallel branching method. The input features first pass through two 3×3 convolutional layers (Conv(3,1)), and then a standard convolutional layer (Conv2d(1,1)) calculates the bounding box loss and classification loss respectively, separating the bounding box prediction from the classification task to achieve more targeted and effective object detection. This decoupling mechanism improves feature extraction accuracy and enhances the adaptability and robustness of YOLOv8n in various object detection challenges. However, the detection head has an extremely large number of parameters, resulting in a slow model processing speed.
[0121] Therefore, based on the idea of shared parameters, this embodiment proposes a... Figure 9 (b) shows a shared-parameter detection head that integrates efficient convolutions, maintaining the high feature extraction accuracy of the original detection head while reducing the number of parameters and computational load. Figure 9 As shown in (b), the optimized detection head can include an efficient multi-scale convolutional module, merged convolutional layers, bounding box convolutional layers, and classification convolutional layers. That is, the improved efficient detection head still adopts a decoupled head structure and a parallel branch feature processing method, but uses one efficient module to replace the four regular 3×3 convolutions in the original detection head, and finally calculates the output through a standard convolutional layer. Figure 9 (b) It can be seen that the efficient module consists of an efficient multi-scale convolutional module (EMSConvP) and a 1×1 merged convolutional layer Conv(1,1).
[0122] Depend on Figure 10 It can be seen that, Figure 10 This is a schematic diagram of the structure of the efficient multi-scale convolution module in the second embodiment of the lightweight metal surface defect detection method of this application. The efficient multi-scale convolution module EMSConvP introduced in this embodiment can perform channel splitting on the input features, using convolution filters of different sizes (1×1, 3×3, 5×5, 7×7) (i.e. Figure 10 The DWConv in the above is used to process the input features, namely the target defect features mentioned above.
[0123] Subsequently, the features from different groups after multi-channel convolution are compared with the original features (i.e., Figure 10 Identity in (in) is achieved through 1×1 merging convolutional layers (i.e. Figure 10 The 1x1PWConv) is merged, and the fused output features are output to the bounding box convolutional layer through a dual-branch structure to obtain the bounding box loss value; the output is then output to the classification convolutional layer to obtain the classification loss value. Finally, this embodiment can identify and classify metal surface defects based on the bounding box loss value and the classification loss value, and generate the corresponding anomaly coordinates.
[0124] In this embodiment, by replacing the four standard convolutions of the original detection head with an efficient module consisting of EMSConvP and 1×1 convolutions, the computational and parameter load is effectively reduced while maintaining the ability to process high-resolution feature maps. Furthermore, using convolution kernels of different sizes can improve the model's ability to detect defects of different sizes, resulting in better detection performance for defects with large scale variations and small target defects on metal surfaces. Therefore, this embodiment utilizes efficient convolution modules to develop a lighter detection head, minimizing accuracy loss while reducing model parameters and detection head size to better reduce the computational burden on the network and further improve model efficiency.
[0125] In summary, the GAC-YOLO model proposed in this embodiment, while maintaining detection performance, significantly improves the model's detection performance and efficiency by redesigning its network structure using efficient convolutional techniques. This allows the original model to be lightweight while minimizing the loss of detection accuracy, better meeting the needs of real-time detection and facilitating rapid deployment in intelligent metal surface defect detection equipment. (Refer to...) Figure 11 The model structure optimization process in this embodiment will be explained. Figure 11 This is a schematic diagram of the GAC-YOLO model for detecting lightweight metal surface defects, which is the second embodiment of the lightweight metal surface defect detection method of this application.
[0126] like Figure 11 As shown, in the GAC-YOLO model for detecting surface defects in lightweight metals, this embodiment can first, as in... Figure 11 (a) shows the improved backbone network and Figure 11 (b) Improved feature extraction in the Neck network based on the DynamicConv module and the GhostBottleneck structure.
[0127] Specifically, this embodiment can improve the C2f module C2f-GhostDynamicConv (i.e. Figure 11 In the C2f-GDConv module, such as Figure 11 As shown in (d), based on the Ghost convolutional module (i.e. Figure 11 (f) replaces the traditional Bottleneck structure of the original C2f module with the introduced Ghost Bottleneck structure.
[0128] Furthermore, the CBS module of the original C2f module was replaced with the DynamicConv module. This allows for better real-time performance in industrial applications while retaining the accuracy benefits of DCNv2.
[0129] At the same time, by Figure 11 (a) shows the improved backbone network and Figure 11 (b) Improved neck network structure: This embodiment can also replace the last three traditional convolution operations in the original backbone network and the two traditional convolution modules in the original neck network with the average pooling downsampling convolution module (ADown). This can more effectively capture the features of small targets, improve the phenomenon of missed detection caused by the loss of original image information due to some minor defects, and effectively reduce the amount of computation.
[0130] In addition, to compensate for the loss of precision caused by the lightweight improvements, Figure 11 (b) In the structure of the neck network, as... Figure 11 (a) shows that a coordinate attention module (CA) is added between the spatial pyramid pooling layer (SPPF) and the contact layer of the backbone network, thereby enhancing the ability to extract global defect features and improving the inherent disadvantage of smaller networks in feature extraction.
[0131] Furthermore, such as Figure 11 In the head structure shown in (c), this embodiment can be modified to include an optimized head improved by an efficient multi-scale convolution module. The efficient convolution redesign of the head significantly reduces the computational and parameter load, while improving the model's ability to detect defects of different sizes.
[0132] In summary, this embodiment improves the traditional C2f module by combining the Ghost module and the dynamic convolution module, reducing computational cost while enhancing feature extraction efficiency to some extent. Furthermore, it employs ADown downsampling convolution and efficient multi-scale convolution to reconstruct some convolutions and the detection head design, further reducing model parameters and computational complexity. The lightweight attention mechanism CA and optimized detection head enhance the model's ability to identify metal surface defects and backgrounds, as well as its target detection capabilities, compensating for the accuracy loss caused by the lightweight design. Through these model structure improvements, the proposed GAC-YOLO model not only retains its detection accuracy advantage but also achieves a significant leap in computational efficiency and industrial real-time performance, making it more suitable for deployment on devices with limited computing resources.
[0133] It should be understood that, in order to better deploy the aforementioned GAC-YOLO model on edge devices, this embodiment can reconstruct and optimize the network structure through an acceleration engine, thereby improving the inference speed of the model on edge devices. Furthermore, in a feasible implementation, prior to step S10, the lightweight metal surface defect detection method further includes steps A1 to A3:
[0134] Step A1: Convert the model file corresponding to the improved feature analysis network and optimized detection head into a target weight file;
[0135] Step A2: Configure acceleration and optimization parameters according to target detection requirements;
[0136] Step A3: Based on the acceleration optimization parameters, the target weight file is converted into an optimization inference engine, which is used to accelerate the inference process of the improved feature analysis network and the optimized detection head.
[0137] It should be noted that the optimized inference engine mentioned above in this embodiment can be based on NVIDIA's TensorRT acceleration engine. As a high-performance acceleration engine in the field of deep learning inference, TensorRT's technical advantages are mainly reflected in model structure optimization and computational acceleration. This tool can use key techniques such as structural reorganization, pruning, and layer fusion to perform topology optimization on neural networks, and improve inference performance by combining customized algorithms and efficient memory management strategies. TensorRT achieves effective improvement in computational efficiency through computational method optimization and the application of low-bit data formats. While model training typically relies on FP32 floating-point numbers to ensure accuracy, in actual deployment environments, the numerical representation range of FP16 is sufficient to meet the accuracy requirements of most inference tasks.
[0138] During optimization, TensorRT analyzes the deep learning model, merges mergeable layers, removes redundant information, and optimizes the output layer structure to improve computational efficiency. For example, in the vertical integration phase, Conv layers, Bias layers, and ReLU activation layers are merged to reduce the model's computational complexity. In the horizontal reorganization phase, TensorRT reduces computation and memory usage by merging layers with the same tensors and behaviors. Furthermore, the Concat layer optimization strategy directly passes input data to the next layer without requiring additional concatenation operations, thus reducing the computational overhead of data transfer.
[0139] It is easy to understand that, referring to Figure 12 The model acceleration process in this embodiment will be explained and described. Figure 12 This is a schematic diagram of the model acceleration process in the second embodiment of the lightweight metal surface defect detection method of this application, as shown below. Figure 12 As shown, in this embodiment, the model file in .pt format corresponding to the improved feature analysis network and optimized detection head trained based on the PyTorch framework can be obtained first, i.e., the pt weight file; then it can be converted into a weight file in .wts format to obtain the target weight file mentioned above.
[0140] Subsequently, based on the actual industrial inspection scenario, this embodiment can configure TensorRT optimization parameters (such as FP16 / INT8 quantization mode, maximum workspace configuration, etc.) for the metal surface defect inspection scenario requirements, that is, configure the above-mentioned acceleration optimization parameters to select the accuracy during the inspection process;
[0141] Finally, this embodiment can use TensorRT's optimization compiler to convert the preprocessed .wts format target weight file into a highly optimized .engine inference engine, i.e., the aforementioned optimized inference engine, to obtain... Figure 12 The TensorRT engine shown.
[0142] This optimized inference engine is used for accelerated inference, specifically speeding up the inference process of the improved feature analysis network and optimized detection head on deployed edge devices. This not only rapidly verifies the effectiveness of the GAC-YOLO model but also ultimately accelerates the inference of the GAC-YOLO algorithm on the NVIDIA GPU platform. Experimental results show that the TensorRT inference acceleration engine can significantly improve detection speed with only a minor loss of model precision, providing effective theoretical support for intelligent metal surface defect detection in real-world scenarios.
[0143] In this embodiment, in order to adapt the metal surface defect detection model to the lightweight deployment requirements of edge deployment devices and mobile devices, a lighter detection head can be developed using an efficient convolution module. This reduces model parameters and detection head size while minimizing accuracy loss, thereby reducing the computational burden on the network and further improving model efficiency.
[0144] Furthermore, to enable the GAC-YOLO model to be better deployed on edge devices, this embodiment can also reconstruct and optimize the network structure based on NVIDIA's TensorRT acceleration engine, improve the inference speed of the model on edge devices, verify the effectiveness of the GAC-YOLO model, and realize the inference and acceleration of the model.
[0145] This embodiment discloses an optimized detection head comprising an efficient multi-scale convolutional module, a merged convolutional layer, a bounding box convolutional layer, and a classification convolutional layer. The efficient multi-scale convolutional module performs multi-scale feature analysis on the target defect features to obtain multi-scale defect features. The merged convolutional layer merges these multi-scale defect features to obtain fused defect features. The bounding box convolutional layer performs convolution operations based on the fused defect features to obtain bounding box loss values. The classification convolutional layer performs convolution operations based on the fused defect features to obtain classification loss values. Based on the bounding box loss values and classification loss values, the anomaly coordinates corresponding to the target metal surface defect are determined. In this embodiment, to adapt the metal surface defect detection model to the lightweight deployment requirements of edge-deployed devices and mobile devices, an efficient convolutional module can be used to develop a lighter detection head. This reduces model parameters and detection head size while minimizing accuracy loss, thereby reducing the computational burden on the network and further improving model efficiency.
[0146] Furthermore, this embodiment discloses converting the model files corresponding to the improved feature analysis network and optimized detection head into target weight files; configuring acceleration optimization parameters according to target detection requirements; and converting the target weight files into an optimized inference engine based on the acceleration optimization parameters. The optimized inference engine is used to accelerate the inference process of the improved feature analysis network and optimized detection head. To enable the GAC-YOLO model to be better deployed in edge devices, this embodiment can also reconstruct and optimize the network structure based on NVIDIA's TensorRT acceleration engine, improving the inference speed of the model in edge devices, verifying the effectiveness of the GAC-YOLO model, and realizing the inference and acceleration of the model.
[0147] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the lightweight metal surface defect detection method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0148] This application provides a lightweight metal surface defect detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the lightweight metal surface defect detection method in Embodiment 1 above.
[0149] The following is for reference. Figure 13This document illustrates a structural schematic diagram of a lightweight metal surface defect detection device suitable for implementing embodiments of this application. The lightweight metal surface defect detection device in this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 13 The lightweight metal surface defect detection device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0150] like Figure 13 As shown, the lightweight metal surface defect detection device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the lightweight metal surface defect detection device. The processing unit 1001, the ROM 1002, and the RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the lightweight metal surface defect detection equipment to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows a lightweight metal surface defect detection equipment with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.
[0151] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment disclosed in this application includes a lightweight metal surface defect detection program product, comprising a lightweight metal surface defect detection program carried on a computer-readable medium, the lightweight metal surface defect detection program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the lightweight metal surface defect detection program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the lightweight metal surface defect detection program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0152] The lightweight metal surface defect detection device provided in this application, employing the lightweight metal surface defect detection method described in the above embodiments, can solve the technical problem of lightweight metal surface defect detection. Compared with the prior art, the beneficial effects of the lightweight metal surface defect detection device provided in this application are the same as those of the lightweight metal surface defect detection method provided in the above embodiments, and other technical features of this lightweight metal surface defect detection device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0153] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0154] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0155] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0156] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0157] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for detecting surface defects in lightweight metals, characterized in that, The method includes: The source metal surface image is preprocessed to obtain the processed metal image; The processed metal image is subjected to multiple feature analysis by an improved feature analysis network to obtain the target defect features; the improved feature analysis network is a YOLOv8 model feature analysis module based on dynamic convolution module and ghost bottleneck structure improvement; The abnormal coordinates corresponding to the target metal surface defect are determined based on the optimized detection head and the target defect features, and the target metal surface defect is marked in the processed metal image based on the abnormal coordinates.
2. The method for detecting surface defects in lightweight metals as described in claim 1, characterized in that, The improved feature analysis network is a YOLOv8 model feature analysis module improved based on dynamic convolution modules, ghost bottleneck structures, and average pooling downsampling convolution modules.
3. The method for detecting surface defects in lightweight metals as described in claim 2, characterized in that, The improved feature analysis network includes an improved backbone network and an improved neck network; The improved backbone network is obtained by replacing the standard convolutional modules in the backbone network of the traditional YOLOv8 network model with average pooling downsampling convolutional modules, and replacing each C2f module with a C2f-GhostDynamicConv module improved based on dynamic convolutional modules and ghost bottleneck structure. The improved neck network is obtained by replacing the standard convolutional modules in the neck network of the traditional YOLOv8 network model with average pooling downsampling convolutional modules, and replacing each C2f module with a C2f-GhostDynamicConv module based on dynamic convolutional modules and a ghost bottleneck structure.
4. The method for detecting surface defects in lightweight metals as described in claim 1, characterized in that, Before performing multi-feature analysis on the processed metal image using an improved feature analysis network to obtain the target defect features, the process includes: In the traditional YOLOv8 network model, the CBS module of each C2f module is replaced with the dynamic convolution module DynamicConv, and the bottleneck module Bottleneck of each C2f module is replaced with the ghost bottleneck structure GhostBottleneck to obtain the C2f-GhostDynamicConv module.
5. The method for detecting surface defects in lightweight metals as described in claim 4, characterized in that, The DynamicConv module includes an attention generation layer, a multiple convolutional filter, a batch normalization layer, and an activation function. The attention generation layer is used to generate dynamic attention parameters corresponding to the first input feature map; The multiple convolutional filter is used to perform weighted attention convolution on the input feature map based on the dynamic attention parameters to obtain multiple defect features; The batch normalization layer is used to normalize the multiple defect features to obtain standard defect features. The activation function is used to smooth the standard defect features and generate a first output feature map.
6. The method for detecting surface defects in lightweight metals as described in claim 2, characterized in that, The average pooling downsampling convolution module includes an average pooling processing layer, a convolutional feature layer, a max pooling feature layer, and a splicing layer. The average pooling layer is used to perform average pooling on the second input feature map and to perform feature equalization along the channel dimension to generate a first average feature and a second average feature. The convolutional feature layer is used to extract features from the first average feature to obtain a first feature map. The max-pooling feature layer is used to extract max-pooling features from the second average feature to obtain the second feature map; The splicing layer is used to splice the first feature map and the second feature map along the channel dimension to generate a second output feature map.
7. The method for detecting surface defects in lightweight metals as described in claim 1, characterized in that, The optimized detection head includes an efficient multi-scale convolutional module, a merged convolutional layer, a bounding box convolutional layer, and a classification convolutional layer; The step of determining the anomaly coordinates corresponding to the target metal surface defect based on the optimized detection head and the target defect features includes: The target defect features are analyzed using the efficient multi-scale convolution module to obtain multi-scale defect features. The multi-scale defect features are merged by the merging convolutional layer to obtain fused defect features. The bounding box loss value is obtained by performing convolution operations based on the fusion defect features through the bounding box convolutional layer. The classification loss value is obtained by performing convolution operations based on the fusion defect features through the classification convolutional layer. The anomaly coordinates corresponding to the target metal surface defect are determined based on the bounding box loss value and the classification loss value.
8. The method for detecting surface defects in lightweight metals as described in claim 3, characterized in that, The improved neck network is obtained by adding a coordinate attention module between the spatial pyramid pooling layer (SPPF) in the backbone network module of the traditional YOLOv8 network model and the splicing layer in the neck network of the traditional YOLOv8 network model.
9. The method for detecting surface defects in lightweight metals as described in claim 1, characterized in that, Before preprocessing the source metal surface image to obtain the processed metal image, the process further includes: Convert the model files corresponding to the improved feature analysis network and optimized detection head into target weight files; Configure acceleration and optimization parameters according to target detection requirements; Based on the acceleration optimization parameters, the target weight file is converted into an optimized inference engine, which is used to accelerate the inference process of the improved feature analysis network and the optimized detection head.
10. A lightweight metal surface defect detection device, characterized in that, The lightweight metal surface defect detection device includes: a memory, a processor, and a lightweight metal surface defect detection program stored in the memory and executable on the processor, the lightweight metal surface defect detection program being configured to implement the steps of the lightweight metal surface defect detection method as described in any one of claims 1 to 9.
Citation Information
Cited By
Steel surface defect detection method and system based on CDF-YOLO
CN121563898A
Anti-interference small target defect intelligent identification method and device in industrial scene
CN121707943A