Target defect detection method and device

Through feature extraction, fusion and prediction networks in deep learning models, efficient and accurate detection of steel strip surface defects is achieved, and the problems of traditional manual inspection are solved.

CN119991581APending Publication Date: 2025-05-13709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510032639.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

During traditional industrial production, manual inspection targets (such as steel strip surfaces) have low defect efficiency, easy fatigue, and high bit error rate, making it difficult to meet the needs of industrial production.

Method used

The feature extraction network, feature fusion network and prediction network in the deep learning model are adopted to achieve target defect detection through feature extraction, fusion and prediction. The specific steps include obtaining feature maps of different scales by the target image input feature extraction network, performing feature fusion, and finally obtaining defect detection results through the prediction network.

Benefits of technology

It improves the efficiency and accuracy of steel strip surface defect detection, reduces the bit error rate, solves the problems of inefficiency and fatigue prone to manual inspection, and is suitable for batch inspection in industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991581A_ABST
    Figure CN119991581A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of defect detection in industrial production, and particularly discloses a target defect detection method and device. According to the invention, the feature extraction network in the deep learning model (target defect detection network) is adopted to extract the feature maps of various scales corresponding to the target image, and the feature fusion network in the target defect detection network is utilized to fuse the extracted feature maps of various scales to obtain a plurality of new feature maps. And finally, in combination with a prediction network in the target defect detection network, performing defect detection on the target image to obtain a target defect detection result. According to the method, the deep learning model is migrated from a laboratory to actual production application, the problems of low efficiency, easy fatigue, high error rate and the like of manual inspection of defects of targets (such as steel strip surfaces) in the traditional industrial production process are solved, and the development of the industrial production and manufacturing field is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of defect detection in industrial production, and more specifically, to a target defect detection method and device. Background Art

[0002] As an important product of the steelmaking industry, steel strips are widely used in shipbuilding, vehicle production, port construction and other fields. However, due to the limitations of current technical processes and steelmaking equipment, surface defects such as cracks and scratches often occur in the production process of steel strips. These defects not only affect the appearance of the steel strips, but also have a certain impact on the quality and performance of the steel strips. Therefore, during the production process of steel strips, it is necessary to conduct defect detection on their surfaces to ensure that the production process of the steel strips meets the standards.

[0003] At present, the commonly used detection method for target defects (such as steel strip surface) in industry is to inspect the steel strip surface by manual inspection. This method has a large demand for human resources and has practical problems, such as high false detection rate and low inspection efficiency, which cannot meet the industrial needs of mass production of steel strips. Summary of the invention

[0004] In view of the defects of the prior art, the purpose of this application is to provide a target defect detection method and device, which aims to solve the problems of low efficiency, easy fatigue, missed detection and high bit error rate in manual inspection of targets (such as steel strip surface) in traditional industrial production processes.

[0005] To achieve the above objectives, in a first aspect, the present application provides a target defect detection method, comprising:

[0006] Input the acquired target image into the feature extraction network in the target defect detection network to obtain feature maps of different scales;

[0007] Inputting feature maps of different scales into the feature fusion network in the target defect detection network, fusing the feature maps of different scales to obtain multiple new feature maps;

[0008] Multiple new feature maps are input into the prediction network in the target defect detection network to obtain the target defect detection results.

[0009] In some embodiments, a plurality of new feature maps are input into a prediction network in a target defect detection network to obtain a target defect detection result, including:

[0010] Input multiple new feature maps into the prediction network, and obtain multiple output features based on multiple permutation attention modules in the prediction network;

[0011] Based on multiple output features, target defect detection results are obtained.

[0012] In some embodiments, multiple new feature maps are input into the prediction network, and multiple output features are obtained based on multiple permutation attention modules in the prediction network, including:

[0013] For any new feature map:

[0014] The new feature map is input into the prediction network, and based on the permutation attention module in the prediction network, the new feature map is grouped based on the permutation attention module to obtain multiple sub-features;

[0015] Based on the channel attention module in the permutation attention module, obtain a first output vector corresponding to the plurality of sub-features;

[0016] Based on the spatial attention module in the permutation attention module, obtain a second output vector corresponding to the plurality of sub-features;

[0017] Perform feature concatenation on the first output vector and the second output vector to obtain a concatenated feature;

[0018] Perform feature aggregation on the spliced ​​features to obtain aggregated features;

[0019] Perform feature transposition on the aggregated features to obtain the output features.

[0020] In some embodiments, a method for acquiring a target defect detection network includes:

[0021] Obtain a dataset of sample images including different defect categories of the target;

[0022] Divide the sample image dataset into training set and test set;

[0023] Inputting the training set into the preset defect detection network for training until the preset defect detection network converges;

[0024] The test set is input into the converged preset defect detection network for optimization to obtain the target defect detection network.

[0025] In some embodiments, the conditions for pre-setting the defect detection network convergence include:

[0026] The total loss function of the preset defect detection network determined based on the confidence loss function, bounding box regression loss function and classification loss function tends to be stable.

[0027] In some embodiments, the collected target image is input into a feature extraction network in a target defect detection network to obtain feature maps of different scales, including:

[0028] Inputting the target image into the feature extraction network, and obtaining a feature map of a first scale based on a first module in the feature extraction network;

[0029] Inputting the feature map of the first scale into the second module in the feature extraction network, and obtaining the feature map of the second scale based on the second module;

[0030] The feature map of the second scale is input into the third module in the feature extraction network, and based on the third module, a feature map of the third scale is obtained.

[0031] In a second aspect, the present application provides a target defect detection device, comprising:

[0032] The first processing module is used to input the collected target image into the feature extraction network in the target defect detection network to obtain feature maps of different scales;

[0033] The second processing module is used to input feature maps of different scales into a feature fusion network in the target defect detection network, fuse the feature maps of different scales, and obtain multiple new feature maps;

[0034] The defect detection module is used to input multiple new feature maps into the prediction network in the target defect detection network to obtain the target defect detection results.

[0035] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is used to execute the target defect detection method described in the first aspect or any one embodiment of the first aspect.

[0036] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the target defect detection method described in the first aspect or any embodiments of the first aspect.

[0037] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the target defect detection method described in the first aspect or any embodiments of the first aspect.

[0038] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the prior art:

[0039] The target defect detection method and device provided in this application, by using the feature extraction network in the deep learning model (target defect detection network), extracts the feature maps of each scale corresponding to the target image, and uses the feature fusion network in the target defect detection network to fuse the extracted feature maps of each scale to obtain multiple new feature maps, and finally combines the prediction network in the target defect detection network to perform defect detection on the target image to obtain the target defect detection result. It realizes the migration of deep learning models from the laboratory to actual production applications, solves the pain points of low efficiency, easy fatigue, high bit error rate, etc. in the manual inspection of targets (such as steel strip surfaces) in traditional industrial production processes, and ensures the development of the industrial production and manufacturing field. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flowchart of a target defect detection method provided in an embodiment of the present application;

[0041] Figure 2 is a schematic diagram of the structure of a defect detection network provided in an embodiment of the present application;

[0042] Figure 3 It is a schematic diagram of the structure of the CBS module and the MP module provided in the embodiment of the present application;

[0043] Figure 4 It is a schematic diagram of the structure of the ELAN module and the ELAN-H module provided in the embodiments of the present application;

[0044] Figure 5 It is a structural diagram of the SPPCSPC module provided in the embodiment of the present application;

[0045] Figure 6 is a schematic diagram of a prediction network structure of a defect detection network provided in an embodiment of the present application;

[0046] Figure 7 is a schematic diagram of the structure of the replacement attention module provided in an embodiment of the present application;

[0047] Figure 8 This is a schematic diagram of the change of the loss function during the training process provided by the embodiment of the present application;

[0048] Fig. 9 It is a schematic diagram of the effects of different types of defect detection provided by the embodiments of the present application;

[0049] Fig.10 is a schematic diagram of the operation of the defect detection device provided in an embodiment of the present application;

[0050] Fig.11 is a schematic diagram of the structure of a target defect detection device provided in an embodiment of the present application;

[0051] Fig.12 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] The term "and / or" in this article is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The symbol " / " in this article indicates that the associated objects are in an or relationship, for example, A / B means A or B.

[0054] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of the objects. For example, a first processing module and a second processing module are used to distinguish different processing modules rather than to describe a specific order of the processing modules.

[0055] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0056] In the description of the embodiments of the present application, unless otherwise specified, “multiple” means two or more than two. For example, multiple output features refer to two or more output features, etc.

[0057] In the related technology, the commonly used defect detection algorithm based on image features is to automatically detect the surface of the steel strip by setting up a camera. This method has the characteristics of intelligence, high efficiency and low false detection, so it is gradually favored and recognized by the industry.

[0058] In the current academic and industrial circles, the defect detection algorithm based on image features is mainly based on deep learning technology, and the defect detection algorithm is introduced into the defect detection field through transfer learning. However, considering that the surface defects of steel strips usually have strong background interference, background classification, a wide variety of surface defects of steel strips, a high rate of false picks, and difficulty in collecting samples of surface defects of steel strips, it is difficult to obtain excellent detection results by simply transferring the target detection algorithm for learning.

[0059] Surface defects are prone to occur during the production of steel strips, which affect product quality and aesthetics. However, manual inspection and re-inspection are inefficient, prone to fatigue, have high bit error rates, and the existing defect detection algorithms have low accuracy, making them difficult to apply in practice. In response to the above-mentioned industry pain points, the present application embodiment proposes a target defect detection method and device to improve the inspection efficiency of products in steel strip production, which can stably, accurately and quickly identify surface defects of steel strips, and help improve the process compliance rate of products.

[0060] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.

[0061] See also Figure 1 , an embodiment of the present application provides a target defect detection method, which may include: step 110, step 120 and step 130.

[0062] Step 110 inputs the collected target image into a feature extraction network in a target defect detection network to obtain feature maps of different scales;

[0063] Inputting feature maps of different scales into the feature fusion network in the target defect detection network, fusing the feature maps of different scales to obtain multiple new feature maps;

[0064] Multiple new feature maps are input into the prediction network in the target defect detection network to obtain the target defect detection results.

[0065] In the embodiment of the present application, the target may be a steel strip, glass, etc. in industrial production, more specifically a steel strip surface, glass surface, etc. The target defect detection method provided in the embodiment of the present application is described below by taking the steel strip in industrial production as an example.

[0066] Construct an algorithm model for the problem of steel strip surface defect detection. The overall structure of the preset defect detection network constructed in the embodiment of the present application includes three sub-networks, namely: feature extraction network, feature fusion network and prediction network. The target defect detection network is trained by the preset defect detection network.

[0067] The target image may be acquired from image data containing surface defects of the steel strip captured by an imaging device, or may be acquired from offline data containing surface defects of the steel strip.

[0068] The target image is input into a target defect detection network, and based on a feature extraction network in the target defect detection network, feature maps of various scales corresponding to the target image are extracted.

[0069] By inputting the feature maps of each scale obtained above into the target defect detection network, and using the feature fusion network in the target defect detection network, the feature maps of each scale are fused in a bidirectional aggregation manner to obtain multiple new feature maps. The new feature maps fuse the deep semantic information and shallow position information of the target image.

[0070] The multiple new feature maps obtained above are input into the target defect detection network, and the prediction network in the target defect detection network is used to perform defect detection on the target image to obtain the target defect detection result.

[0071] In the embodiment of the present application, the types of surface defects of the steel strip that can be detected by the target defect detection network include cracks, inclusions, plaques, pitting surfaces, rolling scales and scratches.

[0072] The target defect detection method provided in the embodiment of the present application uses the feature extraction network in the deep learning model (target defect detection network) to extract the feature maps of each scale corresponding to the target image, and uses the feature fusion network in the target defect detection network to fuse the extracted feature maps of each scale to obtain multiple new feature maps, and finally combines the prediction network in the target defect detection network to perform defect detection on the target image to obtain the target defect detection result. It realizes the migration of deep learning models from the laboratory to actual production applications, solves the pain points of low efficiency, easy fatigue, high bit error rate, etc. in the manual inspection of targets (such as steel strip surfaces) in traditional industrial production processes, and ensures the development of the industrial production and manufacturing field.

[0073] Furthermore, in some embodiments, in the above steps, the method of acquiring the target defect detection network may include:

[0074] Obtain a dataset of sample images including different defect categories of the target;

[0075] Divide the sample image dataset into training set and test set;

[0076] Inputting the training set into the preset defect detection network for training until the preset defect detection network converges;

[0077] The test set is input into the converged preset defect detection network for optimization to obtain the target defect detection network.

[0078] In the embodiment of the present application, the types of surface defects of the steel strip are summarized during the production process of the steel strip. Six types of typical defects on the surface of the steel strip are summarized in the embodiment of the present application, namely: cracks, inclusions, plaques, pitting surfaces, rolling scales and scratches. Afterwards, the surface defect data of the steel strip is collected, screened, cropped and annotated through the data acquisition module to obtain a surface defect image of the steel strip, which is used as the target image. Specifically, the collected surface defect images of the steel strip are screened, low-quality image data are eliminated, and a surface defect detection data set of the steel strip is prepared as a sample image data set.

[0079] The collected sample image data set is manually annotated using Labelimg annotation software, and the annotation format follows the standard format of the target detection algorithm, such as the YOLO series algorithm. After the image annotation in the sample image data set is completed, the sample image data set is divided into a training set and a test set according to a certain ratio (for example, 8:2). In the embodiment of the present application, the sample image data set contains a total of 7450 image samples.

[0080] The algorithm model for the problem of steel strip surface defect detection is constructed through the model building and training modules. The overall structure of the preset defect detection network includes three sub-networks, namely: feature extraction network, feature fusion network and prediction network.

[0081] Construct a preset defect detection network for steel strip surface defect detection. The steps are as follows:

[0082] A bounding box regression loss function with an adaptive scale-aware factor is constructed to constrain the target defect detection network to focus on learning sample boxes of average quality in the sample image dataset during training. At the same time, auxiliary candidate boxes are used to improve the regression accuracy and speed of the preset defect detection network and enhance the generalization performance of the model.

[0083] A permutation attention module is embedded in the prediction network. The permutation attention module can divide the input feature into multiple sub-features in the channel dimension, and use the permutation unit in parallel to describe the feature dependencies of the sub-features in the spatial and channel dimensions. After that, all sub-features are aggregated, and the channel permutation operator is used to realize the information communication between different sub-features. By introducing the permutation attention module, the network can focus on learning the detailed features of the surface defects of the steel strip during the training process, suppress the complex background interference in the image, and then enhance the detection performance of the model and improve the accuracy of defect detection.

[0084] Using the labeled sample image dataset, in a data-driven way, the preset defect detection network is end-to-end trained using the training set until the preset defect detection network converges. After the training is completed, the performance and indicators of the preset defect detection network after convergence are evaluated using the test set, and the preset defect detection network after convergence is optimized to obtain the target defect detection network.

[0085] In an embodiment of the present application, during inference, the non-maximum suppression (NMS) operation in the preset defect detection network inference process is modified to an improved non-maximum suppression (Soft Non-Maximum Suppression, Soft-NMS), thereby achieving more accurate candidate box filtering and alleviating the occurrence of missed detection and false detection of targets due to erroneous suppression caused by the NMS operation.

[0086] In an embodiment of the present application, after obtaining the target defect detection network, the trained target defect detection network can be converted into the rknn format through the rknn-toolkit2 deployment tool and deployed to the RK3588 development board, which is equipped with an Ubuntu system.

[0087] Connect the development board to an imaging device such as a lens module, and call the target defect detection network to monitor the steel strip production process. When a surface defect of the steel strip is detected, the alarm module sends an alarm message to remind the staff to mark the problematic steel strip so that the steel strip that fails the initial inspection can be re-inspected and repaired in a unified manner later.

[0088] Furthermore, in some embodiments, step 110, inputting the acquired target image into a feature extraction network in the target defect detection network to obtain feature maps of different scales, may include:

[0089] Inputting the target image into the feature extraction network, and obtaining a feature map of a first scale based on a first module in the feature extraction network;

[0090] Inputting the feature map of the first scale into the second module in the feature extraction network, and obtaining the feature map of the second scale based on the second module;

[0091] The feature map of the second scale is input into the third module in the feature extraction network, and based on the third module, a feature map of the third scale is obtained.

[0092] Please see further Figure 2 ,The defect detection network consists of three parts: feature extraction network, feature fusion network and prediction network. ,The function of the feature extraction network is to extract features of the ,image sent to the defect detection network to obtain feature ,maps of various scales.

[0093] The feature extraction network consists of CBS module, ELAN module and MP module. Please refer to Figure 3 , where the CBS module contains a standard two-dimensional convolutional layer (Conv), a BN (Batch Normalization) layer, and a SILU activation function.

[0094] The ELAN module adopts a dual-branch architecture. The first branch contains a 1x1 standard convolution to change the number of channels of the input tensor. The second branch first uses a 1x1 standard convolution module to change the number of channels of the input tensor, and then passes through four 3x3 standard convolution modules for feature extraction. Finally, the feature vector is output after the two branches are superimposed.

[0095] The function of the MP module is downsampling. Its structure is similar to that of the ELAN module and also adopts a dual-branch architecture. The first branch contains a maximum pooling layer (Maxpool) and a 1x1 convolution layer, where the maximum pooling layer is used for downsampling operations and the convolution layer adjusts the number of channels of the input vector; the second branch uses a standard 1x1 convolution to adjust the number of channels of the input vector, and then downsamples it through a convolution layer with a stride of 2 and a convolution kernel of 3. Finally, the output vectors of the two branches are superimposed to complete the final downsampling operation.

[0096] The feature fusion network adopts the architecture of path aggregation network (PAN) to fuse the deep semantic information and shallow location information of the fusion target in a bidirectional aggregation manner.

[0097] The prediction network contains three output paths, which predict the output information of targets of different scales. The feature fusion network as a whole consists of SPPCSPC module, ELAN-H module, UP module and other structures. Please refer to Figure 4 and Figure 5 The SPPCSPC module adopts a three-branch structure, one of which performs pyramid pooling, and the other two branches perform standard convolution operations. Finally, the feature vectors output by the three branches are superimposed twice. The ELAN-H module is similar to the ELAN module in structure, differing only in the number of aggregation paths. The UP module performs upsampling by nearest neighbor interpolation.

[0098] This application strives to further improve the detection and identification effect of the defect detection network on the surface defects of steel strips without significantly increasing the computational overhead of the network model. After fully analyzing the structure of the network model, the following optimal implementation suggestions are proposed:

[0099] A lightweight permutation attention module is introduced into the prediction network of the defect detection network. Specifically, after the ELAN-H module in the three output branches of the feature fusion network, the permutation attention module is embedded to obtain a preset defect detection network, and the target defect detection network is obtained by training the preset defect detection network. In the embodiment of the present application, the overall structure of the prediction network in the target defect detection network after embedding is as follows: Figure 6 In the embodiment of the present application, the structure of the feature extraction network and the feature fusion network in the target defect detection network still adopts Figure 2 The feature extraction network and feature fusion network shown.

[0100] In the embodiment of the present application, after the target image is input into the target defect detection network, the first module in the feature extraction network extracts the feature map of the first scale through the first module. The first module specifically comprises Figure 2 The leftmost part of the feature extraction network shown is composed of the MP module and the ELAN module.

[0101] The feature map of the first scale is input to the second module, and the feature map of the second scale is extracted based on the second module. The second module is composed of Figure 2 The middle part of the feature extraction network shown is composed of the MP module and the ELAN module.

[0102] The feature map of the second scale is input to the third module, and the feature map of the third scale is extracted based on the third module. The third module is composed of Figure 2 The rightmost part of the feature extraction network shown is composed of the MP module and the ELAN module.

[0103] In the embodiment of the present application, the scales of the feature graphs of the three scales are, from large to small, a first scale, a second scale, and a third scale.

[0104] The feature maps of multiple scales are input into the feature fusion network in the target defect detection network, and the feature maps of different scales are fused based on the feature fusion network to obtain multiple new feature maps. In the embodiment of the present application, the feature maps input into the feature fusion network include three different scales, so the feature maps output by the feature fusion network also include three.

[0105] Furthermore, in some embodiments, step 130, inputting a plurality of new feature maps into a prediction network in a target defect detection network to obtain a target defect detection result, may include:

[0106] Input multiple new feature maps into the prediction network, and obtain multiple output features based on multiple permutation attention modules in the prediction network;

[0107] Based on multiple output features, target defect detection results are obtained.

[0108] In the embodiment of the present application, three new feature maps output by the feature fusion network in the target defect detection network are input into the prediction network in the target defect detection network, and multiple output features are obtained based on multiple permutation attention modules in the prediction network. After each output feature is processed by the RepConv layer, the target defect detection result is output through an Output layer. The target defect detection result may specifically include defect category, defect location and confidence.

[0109] Furthermore, in some embodiments, in the above steps, a plurality of new feature maps are input into the prediction network, and a plurality of output features are obtained based on a plurality of permutation attention modules in the prediction network, including:

[0110] For any new feature map:

[0111] The new feature map is input into the prediction network, and based on the permutation attention module in the prediction network, the new feature map is grouped based on the permutation attention module to obtain multiple sub-features;

[0112] Based on the channel attention module in the permutation attention module, obtain a first output vector corresponding to the plurality of sub-features;

[0113] Based on the spatial attention module in the permutation attention module, obtain a second output vector corresponding to the plurality of sub-features;

[0114] Perform feature concatenation on the first output vector and the second output vector to obtain a concatenated feature;

[0115] Perform feature aggregation on the spliced ​​features to obtain aggregated features;

[0116] Perform feature transposition on the aggregated features to obtain the output features.

[0117] Please see further Figure 7 , the overall structure of the permutation attention module is as follows Figure 7 As shown, for any new feature map input by the feature fusion network, it is firstly grouped along the channel direction. In the embodiment of the present application, the grouping parameter is set to 64. For each group of sub-features, different importance coefficients are generated by the spatial attention module and the channel attention module respectively, and then the two branches are connected. Finally, all sub-features are aggregated, and the channel permutation operator is used to realize information communication between different sub-features.

[0118] Furthermore, for the channel attention module, global average pooling is used to process the input sub-features to embed global information in the channel direction, and then a pair of parameters are used to scale and move the channel vector, and the vector is activated by the Sigmoid function to obtain the first output vector.

[0119] For the spatial attention module, group normalization (Group Norm) is used to generate spatial statistical information. Then, similar to the channel attention module, a pair of parameters is used to scale and move the spatial vector, and the vector is activated by the Sigmoid function to obtain the second output vector.

[0120] The output vector of the spatial attention module and the output vector of the channel attention module are concatenated in the channel dimension to obtain concatenated features, and the concatenated features are aggregated using the feature aggregation module to re-aggregate the input sub-features, and the information communication between the sub-features is realized based on the permutation operator in the feature transposition to ensure that the vector dimension of the input permutation attention module matches the vector dimension of the output permutation attention module to obtain the output feature.

[0121] The target defect detection method provided in the embodiment of the present application, by embedding a replacement attention module in the prediction network, enables the network to focus on key information such as the defect category and defect location of the learning target during the training process, suppresses the interference of the target such as the complex background of the steel strip surface, and enhances the model's detection ability for defective targets.

[0122] Furthermore, in some embodiments, in the above steps, the conditions for presetting the convergence of the defect detection network include:

[0123] The total loss function of the preset defect detection network determined based on the confidence loss function, bounding box regression loss function and classification loss function tends to be stable.

[0124] In the embodiment of the present application, the condition for the convergence of the preset defect detection network specifically includes that the total loss function of the preset defect detection network determined based on the confidence loss function, the bounding box regression loss function and the classification loss function tends to be stable. Specifically, the total loss function of the preset defect detection network constructed in the present application consists of three parts, as shown in Formula 1, L total is the total loss function of the preset defect detection network, L conf is the confidence loss function, L reg is the bounding box regression loss function, L cls is the classification loss function, N is the number of detection layers, λ1, λ2 and λ3 are the weights of the corresponding three subtasks. In the embodiment of the present application, their values ​​are 0.1, 0.05 and 0.125 respectively.

[0125]

[0126] Next, the loss functions of the three subtasks are explained in detail:

[0127] For the confidence loss function, the cross entropy loss function is used in the embodiment of the present application, and the formula is shown in Formula 2. Where y i represents the intersection-over-union (IOU) between the predicted box and the true labeled box, and its value range is [0, 1]; σ represents the Sigmoid function, x i represents the probability of the current defect category, n is the number of positive and negative samples in the sample image data set, and in the embodiment of the present application, images containing the above defect category and images not containing the above defect category are respectively used as positive and negative samples.

[0128]

[0129] For the classification loss function, the cross entropy loss function is also used, and the calculation formula is shown in Formula 3. Where x′ i Represents the defect category of the target predicted by the target defect detection network, y′ i represents the true value of the current defect category, σ represents the Sigmoid function, s i ×s i Indicates the number of grids at the current scale.

[0130]

[0131] For the bounding box regression loss function, the bounding box regression loss function with an adaptive scale-aware factor constructed in this application is used to constrain the defect detection network. During training, it focuses on learning sample boxes of average quality in the sample set, and cooperates with auxiliary candidate boxes to improve the network regression accuracy and speed and enhance the generalization performance of the model.

[0132] The calculation formula of the bounding box regression loss function is shown in Equation 4, where r is a non-monotonic function, δ and α are hyperparameters of the function, and in the embodiment of the present application, they are respectively 3 and 1.9. The calculation formula of β is shown in Formula 5. For L iou The average value of L base The calculation formula is shown in Equation 6, where x and y represent the horizontal and vertical coordinates of the center point of the prediction box. gt and gt Represents the center coordinates of the real annotation box, W g and H g is the width and height of the minimum bounding rectangle between the predicted box and the true box.

[0133] When calculating the intersection-over-union (IOU) in Formula 5, an adaptive scale perception factor is added to generate a training auxiliary box to dynamically adjust the process of bounding box regression. The specific method is shown in Formulas 7 to 12. Among them, and are the left and right boundaries of the real box, and is the upper and lower boundaries of the real box, w gt and h gt is the width and height of the real frame; b l and b r is the left and right boundaries of the prediction box, b t and b b are the upper and lower boundaries of the prediction box, w and h are the width and height of the prediction box; k is the adaptive scale perception factor, which is dynamically and adaptively adjusted during network training; inter is the intersection between the prediction box and the real box, and the calculation formula is shown in Formula 11. The IOU calculation formula is shown in Formula 12.

[0134]

[0135] In the embodiment of the present application, by constructing an intersection-and-union ratio loss function with an adaptive scale-aware factor, the outlier degree of the sample frame in the data set can be measured, and a smaller gradient gain can be assigned to high- and low-quality annotation frames, while a larger gradient gain can be provided to annotation frames of ordinary quality, thereby improving the convergence speed of the network and enhancing the generalization ability of the model; by introducing an adaptive scale-aware factor, for high intersection-and-union ratio samples in the data set, a smaller auxiliary frame is used to accelerate model learning; and for low intersection-and-union ratio samples, a larger auxiliary frame can be used to improve the performance of bounding box regression. The bounding box regression loss function with an adaptive scale-aware factor proposed in the embodiment of the present application can measure the outlier degree of the anchor frame in the data set, assign different gradient gains to sample frames of different qualities, and generate differentiated auxiliary frames for samples of different intersection-and-union ratios, thereby jointly promoting the improvement of bounding box regression accuracy.

[0136] After making the above modifications to the detection algorithm, the preset defect detection network was trained end-to-end using the pre-collected and annotated sample image dataset. First, the annotated image dataset was divided into a training set and a test set in a ratio of 8:2. The training set contains 5960 images and the test set contains 1490 images. During the experiment, a desktop computer with Ubuntu 20.04 system was selected as the algorithm development platform, and an NVIDIA GeForce RTX TiTAN (24G) graphics card was used to accelerate vector operations during network training and inference.

[0137] The defect detection algorithm uses Python as the development language. The CUDA version required for the experiment is 10.2. The pytorch development framework is 1.9.0, and the corresponding version of torchvision is 0.10.0. The total number of rounds of network training is 300, the initial learning rate is 0.01, the momentum is 0.937, and the batch size (batch_size) during training is 32. During the network training process, the corresponding bounding box regression loss, confidence loss, and classification loss change as shown below: Figure 8 shown.

[0138] After the training is completed, the constructed target defect detection network is obtained. Next, the target defect detection network is tested using a pre-divided test set and compared horizontally with today's mainstream detection algorithms. In order to further improve the model's detection effect on steel strip surface defects, this application uses Soft-NMS to replace the original NMS operation. Soft-NMS introduces a decreasing confidence penalty factor. The penalty factor is calculated based on the intersection-over-union (IOU) between the candidate box and other target boxes. By gradually reducing the confidence of the target box, the competitiveness of the overlapping boxes is gradually weakened. Soft-NMS uses a decreasing confidence method to better suppress the competitiveness of overlapping boxes, reduce the problems of missed target detection and too many candidate boxes, and promote further improvement of model detection accuracy.

[0139] The performance evaluation indicators used in this application are based on P, R, and map@0.5, and the calculation formulas of the indicators are shown in Formulas 13 to 15. In the formula, TP represents samples predicted to be positive but actually positive, FP represents samples predicted to be positive but actually negative, FN represents samples predicted to be negative but actually positive, P represents the precision, R represents the recall, AP represents the average precision, and map can be obtained by averaging the AP values ​​of all categories. Map@0.5 represents the average detection accuracy of all categories when the IOU is set to 0.5.

[0140]

[0141] AP=∫P(R)dR (15)

[0142] In order to more fully demonstrate the effectiveness of the method proposed in this application for steel strip surface defect detection, the detection effects of 6 different defect categories will be demonstrated respectively. The results are as follows: Fig. 9 As shown, Fig. 9 There are 6 different defect categories, including (a) cracks, (b) inclusions, (c) plaques, (d) pitting surfaces, (e) rolling scale, and (f) scratches. Fig. 9It can be seen that the method proposed in this application can generally detect the steel strip surface defect targets in the target image more accurately, which further proves the effectiveness of the proposed method.

[0143] Deploy the trained defect detection model to the RK3588 development board. Specifically:

[0144] After the network training is completed, a model file named "defect_detection.pt" will be generated and converted to ONNX format. Furthermore, the ONNX model is quantized int8 using the rknn-toolkit2 algorithm deployment tool. During the model quantization process, 200 images are selected from the data set as quantization calibration data. After the quantization is completed, a target defect detection network model in rknn format will be generated for forward inference on the development board.

[0145] The RK3588 development board is connected to the lens module and the alarm module to build a steel strip surface defect detection device, which is deployed in the steel strip production workshop. By calling the defect detection model on the board side, the steel strip production process is monitored in real time, and an alarm message is issued when a defect target is detected.

[0146] The operation process of the defect detection device constructed in the embodiment of the present application is as follows: Fig.10 As shown, the steel strip surface defect detection module receives the real-time monitoring data (such as images) collected by the data acquisition module during production, and performs defect detection on the images. When a defect target is detected, the steel strip surface defect detection module sends an alarm message to the alarm module; after the alarm module monitors the alarm message, it triggers the alarm, and reminds the production personnel through the flashing red warning light and the buzzer alarm that the steel strip may have surface defects and needs to be re-inspected and repaired.

[0147] Further, the alarm module involved in the embodiment of the present application is composed of an STM32 single-chip microcomputer, a buzzer and an LED indicator light. The STM32 single-chip microcomputer is connected to the RK3588 development board through a USB port as the main control chip of the alarm module, and the buzzer and the LED indicator light are controlled by the single-chip microcomputer through the general input and output port (GPIO) of the single-chip microcomputer to trigger the alarm prompt in time.

[0148] The embodiment of the present application designs a steel strip surface defect detection device, which migrates the deep learning model from the laboratory to actual production applications, solves the industry pain points such as low efficiency, easy fatigue, and many missed detections in manual inspection of surface defects in the traditional steel strip production process, and promotes the development of the steel strip production and manufacturing field.

[0149] The target defect detection device provided in the present application is described below. The target defect detection device described below and the target defect detection method described above can be referenced to each other.

[0150] See also Fig.11 An embodiment of the present application provides a target defect detection device, which may include: a first processing module 1110 , a second processing module 1120 and a defect detection module 1130 .

[0151] The first processing module 1110 is used to input the collected target image into the feature extraction network in the target defect detection network to obtain feature maps of different scales;

[0152] The second processing module 1120 is used to input the feature maps of different scales into the feature fusion network in the target defect detection network, fuse the feature maps of different scales, and obtain multiple new feature maps;

[0153] The defect detection module 1130 is used to input multiple new feature maps into the prediction network in the target defect detection network to obtain the target defect detection result.

[0154] The target defect detection device provided in the embodiment of the present application uses the feature extraction network in the deep learning model (target defect detection network) to extract the feature maps of each scale corresponding to the target image, and uses the feature fusion network in the target defect detection network to fuse the extracted feature maps of each scale to obtain multiple new feature maps, and finally combines the prediction network in the target defect detection network to perform defect detection on the target image to obtain the target defect detection result. It realizes the migration of deep learning models from the laboratory to actual production applications, solves the pain points of low efficiency, easy fatigue, high bit error rate and other problems in the manual inspection of target defects in the traditional industrial production process, and ensures the development of the industrial production and manufacturing field.

[0155] It can be understood that the detailed functional implementation of each of the above-mentioned units / modules can be found in the introduction of the aforementioned method embodiment, and will not be repeated here.

[0156] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method, which will not be repeated here.

[0157] Based on the method in the above embodiment, the present application embodiment provides an electronic device, see Fig.12The electronic device may include: a processor 1210, a communication interface 1220, a memory 1230 and a communication bus 1240, wherein the processor 1210, the communication interface 1220 and the memory 1230 communicate with each other via the communication bus 1240. The processor 1210 may call the logic instructions in the memory 1230 to execute the method in the above embodiment.

[0158] In addition, the logic instructions in the above-mentioned memory 1230 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application.

[0159] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.

[0160] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.

[0161] It is understandable that the processor in the embodiment of the present application may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0162] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0163] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from a website site, a computer, a server or a data center to another website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

[0164] It should be understood that the various numerical numbers involved in the embodiments of the present application are only used for the convenience of description and are not used to limit the scope of the embodiments of the present application.

[0165] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A target defect detection method, characterized in that: include: Input the acquired target image into the feature extraction network in the target defect detection network to obtain feature maps of different scales; Inputting the feature maps of different scales into a feature fusion network in the target defect detection network, fusing the feature maps of different scales to obtain a plurality of new feature maps; The multiple new feature maps are input into the prediction network in the target defect detection network to obtain the target defect detection result.

2. The target defect detection method according to claim 1, characterized in that: The step of inputting the multiple new feature maps into a prediction network in the target defect detection network to obtain a target defect detection result includes: Inputting the multiple new feature maps into the prediction network, and obtaining multiple output features based on multiple permutation attention modules in the prediction network; Based on the multiple output features, the target defect detection result is obtained.

3. The target defect detection method according to claim 2, characterized in that: The step of inputting the multiple new feature maps into the prediction network and obtaining multiple output features based on multiple permutation attention modules in the prediction network includes: For any new feature map: Inputting the new feature map into the prediction network, and based on a permutation attention module in the prediction network, performing feature grouping on the new feature map based on the permutation attention module to obtain a plurality of sub-features; Based on the channel attention module in the permutation attention module, obtaining a first output vector corresponding to the plurality of sub-features; Based on the spatial attention module in the permutation attention module, obtaining a second output vector corresponding to the plurality of sub-features; Performing feature splicing on the first output vector and the second output vector to obtain a splicing feature; Perform feature aggregation on the spliced ​​features to obtain aggregated features; Perform feature transposition on the aggregated features to obtain the output features.

4. The target defect detection method according to claim 1, characterized in that: The method for acquiring the target defect detection network includes: Obtain a dataset of sample images including different defect categories of the target; Dividing the sample image data set into a training set and a test set; Inputting the training set into a preset defect detection network for training until the preset defect detection network converges; The test set is input into the converged preset defect detection network for optimization to obtain the target defect detection network.

5. The target defect detection method according to claim 2, characterized in that: The conditions for the preset defect detection network to converge include: The total loss function of the preset defect detection network determined based on the confidence loss function, the bounding box regression loss function and the classification loss function tends to be stable.

6. The target defect detection method according to claim 1, characterized in that: The collected target image is input into the feature extraction network in the target defect detection network to obtain feature maps of different scales, including: Inputting the target image into the feature extraction network, and acquiring a feature map of a first scale based on a first module in the feature extraction network; Inputting the feature map of the first scale into a second module in the feature extraction network, and obtaining a feature map of a second scale based on the second module; The feature map of the second scale is input into a third module in the feature extraction network, and based on the third module, the feature map of the third scale is obtained.

7. A target defect detection device, characterized in that: include: The first processing module is used to input the collected target image into the feature extraction network in the target defect detection network to obtain feature maps of different scales; A second processing module is used to input the feature maps of different scales into a feature fusion network in the target defect detection network, fuse the feature maps of different scales, and obtain multiple new feature maps; The defect detection module is used to input the multiple new feature maps into the prediction network in the target defect detection network to obtain the target defect detection result.

8. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the target defect detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor is enabled to execute the target defect detection method according to any one of claims 1 to 6.

10. A computer program product, characterized in that When the computer program product runs on a processor, the processor is enabled to execute the target defect detection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Steel surface defect detection method based on one-stage target detection algorithm

    CN115496752A

  • Metallurgy crane trolley track surface defect detection system

    CN115797914A

  • Desert earthquake noise suppression method based on multi-scale attention interaction network

    CN115877461A

  • Unsupervised notebook appearance defect detection method based on multi-scale standardized flow

    CN116205876A

  • Defect detection method and device, electronic equipment and computer readable storage medium

    CN116485735A