A strip surface defect detection method based on a lightweight dual enhancement network
Through the feature fusion and mixed knowledge distillation method of lightweight dual enhancement network, the problems of large model size, slow speed and low accuracy in strip surface defect detection are solved, and efficient detection on embedded devices is achieved.
Patent Information
- Application Number
- CN202411568207.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-11-05
AI Technical Summary
The existing strip surface defect detection network model has large size, slow detection speed, and low lightweight network accuracy, making it difficult to balance detection accuracy and speed, especially in scenarios with limited application.
The lightweight dual enhancement network is adopted to form a backbone through the stacking of PConv modules, combined with multi-layer coordinated adaptive feature fusion module and decoupling head, and using a hybrid knowledge distillation strategy, we draw more effective knowledge from the conventional network to achieve efficient fusion and utilization of feature information.
A better balance between detection accuracy and speed provides a more applicable solution for embedded devices, reducing the amount of network parameters, improving detection speed and improving accuracy.
Smart Images

Figure CN119515817B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection, and in particular to a strip surface defect detection method based on a lightweight dual-enhanced network. Background Art
[0002] The detection of strip surface defects is a key link in high-quality steel production and also a key indicator for product maintenance. Traditional strip surface defect detection technologies are being replaced by machine vision-based detection methods, which have higher accuracy and lower computational costs. With the rapid development of deep learning, the detection accuracy of defect detection models based on neural networks has been continuously pushed to new heights. However, existing defect detection algorithms are usually based on large frameworks such as VGG, ResNet, and Transformer. The huge model parameters enable these conventional networks to achieve high detection accuracy, but they face challenges in balancing inference speed and model size, which limits their application scenarios. For example, FasterR-CNN based on Resnet50 has an FPS of only 19 when tested on a GTX4070, making it difficult to meet the real-time detection requirements of the industry.
[0003] The emergence of lightweight networks has greatly alleviated the limitation of hardware resources. Many researchers are committed to designing lightweight modules, and technologies such as depthwise separable convolution and channel shuffle have been proposed to fundamentally solve the model efficiency problem. In addition, knowledge distillation, as an effective lightweight means, has been introduced into dense detection tasks to improve the learning ability of lightweight student networks. By using a pre-trained large teacher model to guide the training process of the student model, a lightweight student model with good performance can be obtained. Compared with traditional networks, lightweight networks obtained by lightweight means have unique advantages in terms of computational resource consumption, running efficiency, and flexibility. However, due to the relatively small overall number of network parameters, the following problems will be brought to lightweight models: 1) How can limited feature information be efficiently utilized? 2) The overall performance of the model is limited due to the small number of parameters. How to break through the information quantity barrier to improve the performance bottleneck? Summary of the Invention
[0004] The purpose of the present invention is to provide a strip surface defect detection method based on a lightweight dual-enhanced network, which solves the problems of large model size and slow detection speed of conventional surface defect detection networks, improves the problem of low accuracy of lightweight surface defect detection networks, and can achieve a better balance between detection accuracy and speed, providing a more applicable solution for embedded devices and scenarios with limited computing power.
[0005] To achieve the above object, the present invention provides a strip surface defect detection method based on a lightweight dual-enhanced network, including the following steps:
[0006] S1. Read the image samples for strip surface defect detection, process them using data augmentation techniques, and incorporate them into the sample library.
[0007] S2. Use the defect image samples as the inputs to the conventional strip surface defect detection network and the lightweight dual enhancement network respectively. Extract the feature maps of the training images through the feature extraction module stacked by convolutional layers.
[0008] S3. Input the feature map information from different levels into the multi-layer coordinated adaptive feature fusion module to ensure the efficient fusion and utilization of feature information.
[0009] S4. Use the decoupled head to identify and locate the feature information.
[0010] S5. Use the distillation loss function to force the lightweight dual enhancement network to learn the intermediate hint layer features of the conventional strip surface defect detection network and the logits output of the classifier.
[0011] Preferably, in step S1, the NEU-DET public dataset is used as the dataset, which contains 1,800 images, representing six types of defects, with 300 images for each defect type. The size of each image is 200×200 pixels, and the dataset is divided into a training set and a test set in a ratio of 7:3.
[0012] Preferably, in step S2, the conventional strip surface defect detection network uses the C2f module stacked by ordinary convolutions as the Backbone, and introduces the multi-layer coordinated adaptive feature fusion module and the decoupled head to jointly form the Neck part.
[0013] Preferably, in step S2, the lightweight dual enhancement network uses the FasterNet module stacked with PConv as the core to form the backbone of the network, introduces the multi-layer coordinated adaptive feature fusion module and the decoupled head to jointly form the Neck part, and finally extracts more effective knowledge from the conventional strip surface defect detection network through the hybrid knowledge distillation strategy.
[0014] Preferably, in step S3, the multi-layer coordinated adaptive feature fusion module uses a top-down structure to integrate the high-level semantic feature maps, processes and adjusts the weights of the feature maps from different levels through the adaptive fusion ASFF module, and then uses a bottom-up path to fully fuse the localization information from the lower layer and the semantic information from the higher layer.
[0015] Preferably, in the adaptive fusion ASFF module, let the feature of the nth layer be x n , where n ∈ {2, 3, 4, 5}, and its adaptive feature fusion is expressed as
[0016]
[0017] Among them represents the feature vector at the position of the feature map (i, j) adjusted from different layers to the nth layer, represents the feature vector output along the corresponding channel;
[0018] respectively represent the adaptive weights of different levels, which are calculated using the Softmax function.
[0019] Preferably, in step S5, the hybrid distillation method adopts a two-channel strategy, that is, it consists of two parts of distillation losses based on the feature channel and the logit output channel, as follows:
[0020] S51. Represent the teacher network and the student network as T and S respectively, and the corresponding activation maps as y T and y S , then the distillation loss based on the feature channel can be expressed as follows:
[0021]
[0022] Among them, L(·) represents the cross-entropy loss, which is used to evaluate the difference in the channel distributions of the teacher network and the student network, and φ(·) represents converting the activation value into a probability distribution using the Softmax function, as follows:
[0023]
[0024] In the formula, c = 1, 2, ···, C represents the channel index, i represents the spatial position index of the channel, T is the temperature parameter, and it is expressed using the KL divergence as follows:
[0025]
[0026] Among them, x represents the sample. It can be seen from the above formula that when the probability knowledge of the student model and the teacher model is similar, the KL divergence value can be minimized;
[0027] S52. Divide the distillation loss based on the logit output into two parts: the classification distillation loss and the localization distillation loss. Among them, the classification distillation loss part uses the binary cross-entropy loss function in combination with the KL divergence, as follows:
[0028]
[0029] In the formula, L BCE represents the binary cross-entropy loss, represents the total classification distillation loss, ω = | y T '-y S '| denotes the loss weighting coefficient, which is used to measure the importance of sample x;
[0030] The EIoU loss function is used for the localization distillation loss. EIoU consists of three parts: IoU loss, distance loss, and width-height loss, which are specifically as follows:
[0031]
[0032] where w c and h c are the width and height of the smallest closed region covering the predicted bounding box and the ground truth bounding box, and ρ is the Euclidean distance between two points;
[0033] Similar to the classification loss, the formula of the localization loss function after introducing the weight coefficient is:
[0034]
[0035] In the formula, denotes the total localization distillation loss; therefore, the total distillation loss function based on the logit method is expressed as
[0036]
[0037] where α and β are hyperparameters for adjusting the loss weights of the classification loss and the localization loss respectively;
[0038] Therefore, the total hybrid distillation loss function is:
[0039]
[0040] where the distillation loss based on feature is expressed as It normalizes the feature map into a probability map using Softmax, and then obtains it by minimizing the KL divergence; the distillation loss based on logit consists of the binary classification cross-entropy loss and the EIOU localization loss The classification loss is calculated by converting the logit map into multiple binary classification maps The localization loss is calculated using the EIOU value between the predicted bounding boxes of the two models
[0041] Therefore, the strip surface defect detection method based on the lightweight dual enhancement network with the above structure adopted by the present invention has the following beneficial effects:
[0042] (1) The present invention uses the FasterNet module with PConv as the core to stack and form the backbone of the network, significantly reducing the number of network parameters and improving the detection speed, and solving the problems of large model size and slow detection speed of conventional surface defect detection networks.
[0043] (2) The present invention introduces a multi-layer coordinated adaptive feature fusion module to fully fuse the feature information from different layers and reduce redundant information, solving the problem of low accuracy of lightweight surface defect detection networks.
[0044] (3) The present invention can achieve a better balance between detection accuracy and speed, providing a more applicable solution for embedded devices and scenarios with limited computing power.
[0045] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0046] Figure 1 It is a schematic flow chart of a strip surface defect detection method based on a lightweight dual-enhanced network of the present invention;
[0047] Figure 2 It is a schematic overall flow chart of a strip surface defect detection network in a strip surface defect detection method based on a lightweight dual-enhanced network of the present invention;
[0048] Figure 3 It is a schematic diagram of a hybrid distillation strategy of a strip surface defect detection method based on a lightweight dual-enhanced network of the present invention. Detailed Embodiments
[0049] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0050] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those of ordinary skill in the field to which the present invention belongs. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "up", "down", "left", "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0051] Embodiment
[0052] As Figure 1 shown, the present invention provides a strip surface defect detection method based on a lightweight dual enhancement network, comprising the following steps:
[0053] S1. Read the image samples for strip surface defect detection, and after processing by data enhancement technology, incorporate them into the sample library;
[0054] S2. Respectively use the defect image samples as the inputs of a conventional strip surface defect detection network and a lightweight dual enhancement network, and extract the feature maps of the training images through a feature extraction module stacked by convolutional layers;
[0055] S3. Input the feature map information from different levels into a multi-layer coordinated adaptive feature fusion module to ensure the efficient fusion and utilization of feature information;
[0056] S4. Use a decoupled head to identify and locate the feature information;
[0057] S5. Use a distillation loss function to force the lightweight dual enhancement network to learn the intermediate hint layer features of the conventional strip surface defect detection network and the logits output of the classifier.
[0058] Among them, the multi-layer coordinated adaptive feature fusion module, as one of the key components of the network model, plays an internal enhancement role. Introduce a hybrid knowledge distillation method to transfer the knowledge of the conventional strip surface defect detection network to the lightweight dual enhancement network from the outside. By adopting this dual enhancement strategy, a lightweight dual enhancement network capable of efficiently detecting strip surface defects can be constructed. Compared with mainstream defect detection algorithms, it has a better balance among detection accuracy, speed, and lightweight.
[0059] In step S1, the dataset used is the NEU-DET public dataset. It contains 1800 images, representing six types of defects, with 300 images for each defect type, and the size of each image is 200×200 pixels. During the experiment, the dataset is divided into a training set and a test set in a ratio of 7:3.
[0060] Data enhancement technology is widely used in detection tasks due to its effectiveness. To improve the detection accuracy, the present invention applies data enhancement technology, such as geometric transformation and contrast enhancement, to better customize the dataset for training. After enhancement, the dataset contains 3000 images, among which 2100 are for training and 900 are for testing.
[0061] In step S2, for the conventional strip surface defect detection network: 1) Stack the C2f modules composed of ordinary convolutions as the Backbone; 2) Introduce a multi-layer coordinated adaptive feature fusion module and a decoupled head to jointly form the Neck part.
[0062] For the lightweight dual enhancement network: 1) Stack the FasterNet modules with PConv as the core to form the backbone of the network; 2) Introduce a multi-layer coordinated adaptive feature fusion module and a decoupled head to jointly form the Neck part; 3) Through the hybrid knowledge distillation strategy, draw more effective knowledge from the conventional strip surface defect detection network.
[0063] In step S3, design a multi-layer coordinated adaptive feature fusion module inside the network model to address challenges such as cross-layer cyclic feature information, effectively utilizing semantic and localization features, and reducing feature redundancy. The multi-layer coordinated adaptive feature fusion module uses a top-down structure to integrate high-level semantic feature maps. Then, the feature maps from different levels are processed by an adaptive spatial feature fusion (ASFF) module to adjust their information weights. Finally, a bottom-up path is used to fully fuse the localization information from the lower layers with the semantic information from the higher layers. This "spider web" structure can better fuse features and optimize the balance between accuracy and computational efficiency.
[0064] As Figure 2 shown, take the feature map channels (P2 - P5) at four levels as an example. First, input the feature maps of the C2 - C5 layers in the backbone network into the feature fusion module, and then use the upsampling in the opposite path of feature extraction to simply circulate the feature maps. Then introduce the adaptive spatial feature fusion (ASFF) module to fully fuse and utilize all channel information. Since the operation of ASFF is differentiable, the backpropagation algorithm can be used to automatically adjust the weights of feature fusion during the training process, thereby suppressing useless information and amplifying useful information. Finally, use the bottom-up path channel to enhance the high-level semantic features and low-level localization features.
[0065] For the ASFF module, let the feature of the nth layer be x n (n can be adjusted accordingly according to different network structures. In the legend, n ∈ {2, 3, 4, 5}). For features at different levels, the feature map sizes need to be adjusted to the same size. Taking the nth layer as an example, its adaptive feature fusion is expressed as
[0066]
[0067] where represents the feature vector at the (i, j) position of the feature map adjusted from different layers to the nth layer, represents the feature vector output along the corresponding channel. They represent the adaptive weights at different levels, which are calculated using the Softmax function. The auxiliary blocks shown in the figure are used to adjust the number of channels in different layers.
[0068] In step S5, the hybrid distillation method adopts a two-channel strategy, as Figure 3 shown. The outer pipeline in the figure represents feature-based distillation, and the inner pipeline represents logit-based distillation. Among them, the outer channels focus on feature knowledge extraction, while the inner channels pay more attention to logits output.
[0069] 1) For the feature extraction channels, the current general probability knowledge transfer method is adopted. First, the feature maps obtained by processing each channel through the activation function are normalized by the Softmax function to obtain the probability maps, and then the KL divergence of the Soft probability maps between the student model and the teacher model is minimized, so that the student model can learn the important feature representations on the teacher channels.
[0070] Specifically, the loss function of the feature channel distillation method is introduced first. The teacher network and the student network are represented as T and S respectively, and the corresponding activation maps are represented as y T and y S . Then the distillation loss along the channels can be expressed as follows:
[0071]
[0072] Among them, L(·) represents the cross-entropy loss, which is used to evaluate the difference in the channel distributions of the teacher network and the student network. φ(·) represents converting the activation values into a probability distribution using the Softmax function, as follows:
[0073]
[0074] In the formula, c = 1, 2, ···, C represents the channel index, i represents the spatial position index of the channel, and T is the temperature parameter. It can be expressed by the KL divergence as follows:
[0075]
[0076] Among them, x represents the sample. From the above formula, it can be seen that when the probability knowledge of the student model and the teacher model is similar, the KL divergence value can be minimized and more weights can be assigned.
[0077] 2) For the channels that focus on logit output, it is customary to divide them into two parts: classification distillation loss and localization distillation loss. This method can make up for the problem that the feature-based distillation method does not pay enough attention to the localization distillation loss.
[0078] ①For the classification distillation loss part of logit distillation, the currently common classification distillation method is adopted, that is, the binary cross-entropy loss function is used in combination with the KL divergence. The classification logic diagram is represented by the binary cross-entropy loss function as multiple binary classification diagrams, as follows:
[0079]
[0080] In the formula, L BCE represents the binary cross-entropy loss, represents the total classification distillation loss. And ω = |y T' -y S' | represents the loss weighting coefficient, which is used to measure the importance of the sample x.
[0081] ②In view of the good effect of IoU in the localization distillation loss, the present invention introduces a more advanced EIoU loss function for the localization distillation loss. EIOU solves the problem that when the aspect ratios of the predicted box and the ground truth box are the same, w and h cannot increase or decrease simultaneously, resulting in a pair of opposite numbers when calculating the gradient, by separating the influence factors of the aspect ratios of the predicted box and the ground truth box, and then calculating the length and width of the predicted box and the ground truth box respectively. EIoU includes three parts: IoU loss, distance loss, and width-height loss (i.e., overlapping area, center point distance, and aspect ratio). The width-height loss is used to minimize the difference between the width and height of the predicted target bounding box and the ground truth bounding box, so that it has a faster convergence speed and better localization results. Specifically as follows:
[0082]
[0083] Among them, w c and h c are the width and height of the smallest closed region covering the predicted bounding box and the ground truth bounding box. ρ is the Euclidean distance between two points.
[0084] Similar to the classification loss, the formula of the localization loss function after introducing the weight coefficient is:
[0085]
[0086] In the formula, represents the total localization distillation loss. Therefore, the total distillation loss function based on the logit method is expressed as
[0087]
[0088] In the formula, α and β are hyperparameters for adjusting the loss weights of the classification loss and the localization loss respectively. Therefore, the total hybrid distillation loss function is:
[0089]
[0090] In the experimental test, by default Among them, the distillation loss based on features is expressed as It uses Softmax to normalize the feature map into a probability map, and then obtains it by minimizing the KL divergence. (ii) The distillation loss based on logits consists of the binary classification cross-entropy loss and the EIOU localization loss It is composed of. The classification loss is calculated by converting the logit map into multiple binary classification maps The EIOU value between the predicted bounding boxes of the two models is used to calculate the localization loss
[0091] Therefore, the present invention adopts the above-mentioned strip surface defect detection method based on a lightweight dual-enhanced network. Internally, lightweight modules and efficient feature fusion modules are applied, while externally, hybrid knowledge distillation is utilized. In this framework, an additional teacher model will be constructed to assist the student model in hybrid knowledge refinement. The teacher model is called the conventional strip surface defect detection network, and the student model is called the lightweight dual-enhanced network. The conventional strip surface defect detection network is stacked with C2f modules composed of ordinary convolutions as the Backbone, and a multi-layer coordinated adaptive feature fusion module and a decoupled head are introduced to jointly form the Neck part. In order to obtain a detection model that can balance accuracy and speed, the lightweight dual-enhanced network: 1) adopts the stacking of FasterNet modules with PConv as the core to form the backbone of the network, which can greatly reduce the number of network parameters and improve the detection speed; 2) introduces a multi-layer coordinated adaptive feature fusion module to fully fuse the feature information from different layers and reduce redundant information; 3) adopts a hybrid knowledge distillation strategy to extract more effective knowledge from the teacher model.
[0092] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that: they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A strip surface defect detection method based on a lightweight dual-enhanced network, characterized in that It includes the following steps: S1. Read the image samples for strip surface defect detection, process them using data augmentation techniques, and incorporate them into the sample library; S2. Use the defective image samples as the inputs of the conventional strip surface defect detection network and the lightweight dual enhancement network respectively, and extract the feature maps of the training images through the feature extraction module stacked by convolutional layers; S3. Input the feature map information from different levels into the multi-layer coordinated adaptive feature fusion module to ensure the efficient fusion and utilization of feature information; In step S3, the multi-layer coordinated adaptive feature fusion module adopts a top-down structure to integrate the high-level semantic feature maps, processes and adjusts the weights of the feature maps from different levels through the adaptive fusion ASFF module, and then uses a bottom-up path to fully fuse the localization information from the lower layer and the semantic information from the higher layer; In the Adaptive and Selective Feature Fusion (ASFF) module, let the feature of the n-th layer be , where n ∈ {2, 3, 4, 5}, and its adaptive feature fusion is expressed as ; Among them represents the feature map adjusted from different layers to the nth layer at the position of the feature vector represents the feature vector output along the corresponding channel; , representing the adaptive weights at different levels, calculated using the Softmax function; S4. Use a decoupled head to identify and locate the feature information; S5. Use a distillation loss function to force the lightweight dual enhancement network to learn the intermediate hint layer features of the conventional strip surface defect detection network and the logits output of the classifier; In step S5, the hybrid distillation method adopts a two-channel strategy, mainly composed of two parts of distillation loss based on the feature channel and the logit output channel; In the distillation loss based on the feature channel, the teacher network and the student network are respectively denoted as T and S, and the corresponding activation maps are respectively denoted as and ; The distillation loss based on the logit output is divided into two parts: classification distillation loss and localization distillation loss. Among them, the classification distillation loss part uses the binary cross-entropy loss function in combination with the KL divergence, and the localization distillation loss uses the EIoU loss function.
2. The strip surface defect detection method based on a lightweight dual-enhanced network according to claim 1, characterized in that: In step S1, the NEU-DET public dataset is used as the dataset, which contains 1,800 images, representing six types of defects. Each defect type has 300 images. The size of each image is 200×200 pixels, and the dataset is divided into a training set and a test set in a ratio of 7:
3.
3. A strip surface defect detection method based on a lightweight dual-enhanced network according to claim 1, characterized in that: In step S2, the conventional strip surface defect detection network uses the C-2f module stacked by ordinary convolutions as the Backbone, and introduces the multi-layer coordinated adaptive feature fusion module and the decoupled head to jointly form the Neck part.
4. A strip surface defect detection method based on a lightweight dual-enhanced network according to claim 1, characterized in that: In step S2, the lightweight dual enhancement network uses the FasterNet module stacked with PConv as the backbone of the network, introduces the multi-layer coordinated adaptive feature fusion module and the decoupled head to jointly form the Neck part, and finally extracts more effective knowledge from the conventional strip surface defect detection network through the hybrid knowledge distillation strategy.
5. A strip surface defect detection method based on a lightweight dual-enhanced network according to claim 1, characterized in that: In step S5, the hybrid distillation method adopts a two-channel strategy, which is specifically as follows: S51. The distillation loss based on the feature channel is expressed as follows: ; Among them, represents the cross-entropy loss, which is used to evaluate the difference in the channel distributions between the teacher network and the student network. represents converting the activation values into a probability distribution using the Softmax function as follows: ; where represents the channel index, represents the spatial position index of the channel, is the temperature parameter, expressed by the KL divergence as follows: ; Among them is expressed as a sample. It can be seen from the above formula that when the probability knowledge of the student model is similar to that of the teacher model, the KL divergence value can be minimized; The details of the classification distillation loss part are as follows: ; ; In the formula, represents the binary cross-entropy loss, represents the total classification distillation loss, represents the loss weighting coefficient, which is used to measure the importance of the sample ; In the localization distillation loss, EIoU includes three parts: IoU loss, distance loss, and width-height loss, which are specifically as follows: ; where A and B are the regions of the predicted bounding box and the ground truth bounding box respectively, represents the intersection area of the two regions, represents the union area of the two regions; ; Among them, and are the width and height of the smallest closed region covering the predicted bounding box and the ground-truth bounding box, is the Euclidean distance between two points; Similar to the classification loss, the formula of the localization loss function after introducing the weight coefficient is: ; In the formula, represents the total localization distillation loss; thus, the total distillation loss function based on the logit method is expressed as ; where and are hyperparameters for adjusting the loss weights of the classification loss and the localization loss, respectively; Therefore, the total hybrid distillation loss function is: ; Among them, the distillation loss based on features is expressed as , which normalizes the feature map into a probability map using Softmax and then obtains it by minimizing the KL divergence; the distillation loss based on logits consists of the binary classification cross-entropy loss and the EIOU localization loss . The classification loss is calculated by converting the logit map into multiple binary classification maps, and the localization loss is calculated using the EIOU value between the predicted bounding boxes of the two models .
Citation Information
Patent Citations
Visual language navigation method based on double semantic comprehension and fusion
CN116429111A
Hot-rolled strip steel surface defect detection method based on double-detection-head D-YOLO, storage medium and equipment
CN118429302A