Metal surface microdefect detection method and system based on deep learning

Through the metal surface micro-defect detection method based on deep learning, the attention combination module, enhanced feature extraction module and four-branch stacking module are used to solve the problem of low efficiency of existing detection methods, and efficient and accurate detection of metal surface micro-defects is achieved.

CN120125575AActive Publication Date: 2025-06-10CHINA JILIANG UNIV +1

Patent Information

Application Number
CN202510594144.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-10
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

Existing methods for detecting micro defects on metal surfaces are inefficient and it is difficult to quickly and accurately identify micro defects on metal surfaces.

Method used

Deep learning-based detection method is adopted, by acquiring metal surface images and preprocessing, the trained defect detection model is called for defect detection, and multi-scale features and deep local features are extracted using attention combination module, enhanced feature extraction module and four-branch stacking module to achieve efficient defect detection.

Benefits of technology

It improves the efficiency and accuracy of micro defect detection on metal surfaces, can quickly identify tiny defects of 10um or above, and significantly improves the accuracy and reliability of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125575A_ABST
    Figure CN120125575A_ABST
Patent Text Reader

Abstract

The invention relates to a metal surface micro-defect detection method and system based on deep learning. The method comprises the following steps: acquiring a surface image of to-be-detected metal and preprocessing the surface image; calling a trained defect detection model to carry out defect detection on the surface image to obtain a bounding box of a defect area in the surface image; wherein the trained defect detection model comprises an attention combination module, an enhanced feature extraction module and a four-branch stacking module; the attention combination module is used for extracting multi-scale features of the surface image; the enhanced feature extraction module extracts deep local features of the surface image by expanding a receptive field; the four-branch stacking module reduces the quantity of model parameters through network structure optimization; and determining the actual defect size of the to-be-detected metal based on the size information of the bounding box. By adopting the method, the detection efficiency of the metal surface microdefects can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of non-destructive testing technology, and particularly to a method and system for detecting micro-defects on the metal surface based on deep learning. Background Art

[0002] With the acceleration of the industrialization process, metal materials play an indispensable role in many industries due to their excellent mechanical properties and cost-effectiveness. However, during the production process, defects such as cracks, scratches, and inclusions often inevitably appear on the surface of metal materials. These defects will reduce the properties such as toughness and strength of the materials, and thus affect the safety performance and service life of metal products. Therefore, how to detect micro-defects on the metal surface non-destructively and efficiently during the production process is a key measure to improve product quality.

[0003] In the related art, common methods for detecting micro-defects on the metal surface include the metallographic method and the scanning electron microscope method. These two methods obtain material structure and defect information by observing detection photos, but there are problems such as difficulty in obtaining detection photos and complex operations, which affect the detection efficiency. Summary of the Invention

[0004] Based on this, it is necessary to provide a method and system for detecting micro-defects on the metal surface based on deep learning to solve the technical problem of low detection efficiency existing in the above methods.

[0005] In the first aspect, this application provides a method for detecting micro-defects on the metal surface based on deep learning. The method includes: Obtain the surface image of the metal to be detected and perform preprocessing; Call the trained defect detection model to perform defect detection on the surface image to obtain the bounding box of the defect area in the surface image; Determine the actual defect size of the metal to be detected based on the size information of the bounding box; Among them, the trained defect detection model includes an attention combination module, an enhanced feature extraction module, and a four-branch stacking module; the attention combination module is used to extract multi-scale features of the surface image; the enhanced feature extraction module extracts deep local features of the surface image by expanding the receptive field; the four-branch stacking module reduces the number of model parameters through network structure optimization.

[0006] In one embodiment, the obtaining the surface image of the metal to be detected and performing preprocessing is specifically as follows: Use an ultrasonic scanning microscope equipped with a 300MHz high-frequency probe and pure distilled water as a couplant to perform C-Scan scanning with a step of 1um to obtain a high-resolution metal surface image; Crop the collected high - resolution images into pictures of the same size, and clean the data of the cropped pictures, including removing unqualified cropped pictures and pictures containing noise introduced during the scanning process, etc.

[0007] In one embodiment, the trained defect detection model includes a backbone network, a neck network, and a detection head connected in sequence; The attention combination module is arranged in the backbone network and the neck network; The enhanced feature extraction module is arranged in the neck network; The four - branch stacking module is arranged in the neck network.

[0008] The attention combination module includes a multi - scale attention sub - module and a down - sampling feature extraction sub - module; The multi - scale attention sub - module is used for feature extraction by means of channel reshaping and grouping, encoding global information through parallel branches and recalibrating channel weights, and aggregating the output features of parallel branches across dimensions.

[0009] In one embodiment, the enhanced feature extraction module is obtained by adjusting the spatial pyramid pooling cross - stage partial connection module according to the principle of spatial pyramid fast pooling; it is used to extract deep local features through multi - scale pooling and cross - stage connection.

[0010] In one embodiment, the enhanced feature extraction module includes two branches; in the first branch, the input content is processed through multiple convolutional blocks to obtain a first output, three consecutive pooling operations are performed on the first output, the first output is concatenated with the output results obtained from each pooling operation, and then processed through two convolutional blocks to obtain the result of the first branch; In the second branch, the input content is processed through one convolutional block to obtain the result of the second branch; After concatenating the result of the first branch and the result of the second branch, and then processing through one convolutional block, the output result of the enhanced feature extraction module is obtained.

[0011] In one embodiment, the four - branch stacking module optimizes the network structure by adjusting the quantity and distribution of 1×1 convolutions and 3×3 convolutions, and reducing the concatenated branches.

[0012] In one embodiment, the four - branch stacking module includes two branches; In the first branch, the input content is processed through a 1×1 convolutional block to obtain a first output; In the second branch, the input content is processed by a 1×1 convolutional block to obtain a second output. The second output is processed by a 3×3 convolutional block to obtain a third output. The second output and the third output are concatenated and then input into a 1×1 convolutional block for processing to obtain a fourth output. The fourth output is processed by a 3×3 convolutional block to obtain a fifth output, and the fifth output is further processed by a 3×3 convolutional block to obtain a sixth output. The first output, the fourth output, the fifth output, and the sixth output are concatenated and then input into a 1×1 convolutional block for processing to obtain the output result of the four-branch stacking module.

[0013] In one embodiment, determining the actual defect size of the metal to be detected based on the size information of the bounding box includes: Obtaining a pre-determined size equation; the size equation is obtained by fitting the actual defect size of the sample surface image and the size information of the bounding box of the defect region in the sample surface image; Substituting the size information of the bounding box into the size equation to obtain the actual defect size of the metal to be detected.

[0014] In one embodiment, the defect detection model is trained in the following manner: Obtaining a sample data set; the sample data set includes sample surface images and the true bounding boxes of the sample surface images; Inputting the sample surface images into an initial defect detection model for defect detection to obtain predicted bounding boxes; Obtaining the intersection over union between the predicted bounding box and the true bounding box, and the shape similarity between the predicted bounding box and the true bounding box. Based on the intersection over union and the shape similarity, determining the loss value between the predicted bounding box and the true bounding box; Taking reducing the loss value as the training objective, training the initial defect detection model to obtain the trained defect detection model.

[0015] In a second aspect, the present application further provides a metal surface micro-defect detection system based on deep learning. The system includes: An acquisition unit that uses an ultrasonic scanning microscope to scan the surface of the metal to be detected, acquires the surface image of the metal to be detected, and performs preprocessing; The detection unit calls the trained defect detection model on the computer device to perform defect detection on the surface image, and obtains the bounding box of the defect area in the surface image; wherein, the trained defect detection model includes an attention combination module, an enhanced feature extraction module, and a four-branch stacking module; the attention combination module is used to extract multi-scale features of the surface image; the enhanced feature extraction module extracts deep local features of the surface image by expanding the receptive field; the four-branch stacking module reduces the number of model parameters through network structure optimization; The determination unit determines the actual defect size of the metal to be detected based on the size information of the bounding box.

[0016] The above-mentioned deep learning-based metal surface micro-defect detection method and system preprocess the surface image of the metal to be detected, and then call the trained defect detection model to perform defect detection on the surface image to obtain the bounding box of the defect area in the surface image; based on the size information of the bounding box, the actual defect size of the metal to be detected is determined. Through automated image processing and model inference, the metal surface defects can be quickly determined, improving the detection efficiency; further, an attention combination module, an enhanced feature extraction module, and a four-branch stacking module are constructed in the defect detection model; the attention combination module can improve the detection ability for defects of different sizes by performing multi-scale feature extraction; the enhanced feature extraction module can improve the characterization ability for complex defects by expanding the receptive field and extracting deep local features; the four-branch stacking module can reduce the number of model parameters through network structure optimization, making the model more suitable for deployment on resource-constrained devices. Thus, through the collaborative work of the attention combination module, the enhanced feature extraction module, and the four-branch stacking module, efficient, accurate, and lightweight defect detection can be achieved. Furthermore, the trained defect detection model has a good detection effect on metal surface micro-defects, and can achieve high-precision identification of micro-defects of 10um and above on the metal surface, effectively improving the accuracy and reliability of metal surface micro-defect detection. Description of the Drawings

[0017] Figure 1 It is a schematic flowchart of a deep learning-based metal surface micro-defect detection method in an embodiment; Figure 2 It is a schematic structural diagram of an ultrasonic scanning system in an embodiment; Figure 3 It is a schematic structural diagram of an improved YOLOv7 network in an embodiment; Figure 4 It is a schematic structural diagram of an efficient multi-scale attention sub-module in an embodiment; Figure 5 It is a structural diagram of an improved enhanced feature extraction module in an embodiment; Figure 6 Schematic diagram of the structure of a four-branch stacking module in an embodiment; Figure 7 Scatter plot fitting diagram of the diagonal length of the prediction box and the longest side length of the actual inclusion in an embodiment; Figure 8 Flow schematic diagram of a method for training a defect detection model in an embodiment; Figure 9 Comparison diagram of different processing results in an embodiment; Figure 10 Comparison diagram of the actual detection effects before and after the improvement of the YOLOv7 network in an embodiment; Figure 11 Block diagram of the structure of a metal surface micro-defect detection system based on deep learning in an embodiment; Figure 12 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0018] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0019] In one embodiment, as Figure 1 shown, a method for detecting metal surface micro-defects based on deep learning is provided. In this embodiment, the method is exemplified by being applied to a terminal. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is realized through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, and tablet computers. The server can be realized by an independent server or a server cluster composed of multiple servers. In this embodiment, the method includes the following steps: Step S110, obtaining a surface image of the metal to be detected and performing preprocessing.

[0020] In specific implementation, an ultrasonic microscopy detection technology can be used to collect the metal surface image. The ultrasonic microscopy detection technology is a kind of ultrasonic detection technology. This technology can obtain high-resolution images of micro-defects on the metal surface through a scanning imaging method, and can meet the basic requirements for detecting micro-defect targets through images.

[0021] Specifically, an ultrasonic scanning microscope can be used to collect the surface image of the metal to be detected. The ultrasonic scanning microscope can be equipped with a high-frequency probe (such as 300 MHz), use pure distilled water as a coupling agent, and perform C-Scan scanning with a preset step size (such as 1 um) to obtain a high-resolution metal surface image. Refer to Figure 2 As shown in the schematic diagram of the ultrasonic scanning system structure, the ultrasonic scanning microscope is realized by the propagation and reflection of high-frequency ultrasonic waves inside the metal to be detected. The ultrasonic waves are emitted by the transducer, focused by the acoustic lens and then transmitted to the metal to be detected. When the ultrasonic waves encounter the interface of different media, due to the acoustic impedance difference, reflected waves will be generated. These reflected waves are received by the transducer and converted into electrical signals. After the electrical signals are processed, they are transmitted to the computer digitally for signal and image processing.

[0022] When ultrasonic waves propagate in a material, its wave equation can be expressed as: (1) In the formula: is the amplitude of the ultrasonic wave; is the length and satisfies .

[0023] The reflected wave generated by the ultrasonic wave on the emission plane is expressed as: (2) If there are no defects inside the metal to be detected, then the ultrasonic waves can propagate in the metal to be detected all the time until they encounter the lower surface of the metal to be detected and are emitted. At this time, the bottom surface reflected wave is expressed as: (3) When there are defects in the metal to be detected and there is a difference in acoustic impedance between the defect and the metal to be detected, the ultrasonic waves will generate reflections, and the reflection signal is: (4) It can be known from the above formula that the reflected wave of the ultrasonic echo at the defect lags behind the reflected wave of the emission plane by , and is ahead of the reflected wave of the lower surface by . Therefore, by using the phase and amplitude information of each echo point to be presented as color or grayscale pixels on the screen, an ultrasonic scanning image can be obtained.

[0024] In specific implementation, the collected metal surface image can be preprocessed by cropping the image into pictures of the same size and cleaning the data of the cropped pictures, including removing unqualified cropped pictures and pictures containing noise points introduced during the scanning process.

[0025] Step S120, call the trained defect detection model to detect the defects in the surface image, and obtain the bounding box of the defect area in the surface image.

[0026] Among them, the defect detection model can be a model improved based on YOLOv7. Specifically, the improvement of YOLOv7 includes introducing an efficient multi-scale attention mechanism (EMA) and combining it with a downsampling feature extraction sub-module to form an attention combination module, improving the SPPCSPC module in the neck network to form an enhanced feature extraction module, and designing a four-branch stacking module to replace all ELAN-H modules in the neck network.

[0027] Among them, the attention combination module is used to extract multi-scale features of the surface image; the enhanced feature extraction module extracts deep local features of the surface image by expanding the receptive field; the four-branch stacking module reduces the model parameter quantity through network structure optimization.

[0028] Among them, the bounding box contains the position of the bounding box (such as the xy coordinates of the upper left corner) and the size of the bounding box (length and width), for example, expressed as bbox = (x, y, w, h).

[0029] In specific implementation, an improved YOLOv7 model can be pre-constructed as the initial defect detection model. At the same time, sample surface images of the sample metal are collected and the true bounding boxes of their defect regions are labeled to construct a sample data set. The initial defect detection model is trained and tested through this sample data set to obtain a trained defect detection model. After training is completed, the trained defect detection model can be called to perform defect detection on the surface image collected for the metal to be detected, and the bounding boxes of the detected defect regions are output.

[0030] Step S130, based on the size information of the bounding box, determine the actual defect size of the metal to be detected.

[0031] In specific implementation, the information of the bounding box of the defect region in the surface image output by the trained defect detection model includes the length and width of the bounding box, and the diagonal length of the bounding box is calculated through the length and width. According to this diagonal length, the side length of the longest side of the actual defect of the metal to be detected is determined. The size of the defect of the metal to be detected can be classified according to the side length of the longest side of the actual defect. For example, the preset multiple size ranges can be queried according to the actual defect size of the metal to be detected, and the target size range that matches the actual defect size is determined. Each size range has a corresponding level, so that the level corresponding to the target size range can be used as the defect level corresponding to the metal to be detected.

[0032] In the above metal surface micro-defect detection method based on deep learning, the surface image of the metal to be detected is obtained and preprocessed, and then the trained defect detection model is called to detect the defects in the surface image, and the bounding box of the defect area in the surface image is obtained; based on the size information of the bounding box, the actual defect size of the metal to be detected is determined. Through automated image processing and model inference, the metal surface defects can be quickly determined, improving the detection efficiency; further, an attention combination module, an enhanced feature extraction module, and a four-branch stacking module are constructed in the defect detection model; the attention combination module can improve the detection ability for defects of different sizes by performing multi-scale feature extraction; the enhanced feature extraction module can improve the characterization ability for complex defects by expanding the receptive field and extracting deep local features; the four-branch stacking module can reduce the number of model parameters through network structure optimization, making the model more suitable for deployment on resource-constrained devices. Thus, through the collaborative work of the attention combination module, the enhanced feature extraction module, and the four-branch stacking module, efficient, accurate, and lightweight defect detection can be achieved, and then the trained defect detection model has a good detection effect on metal surface micro-defects, and can achieve high-precision recognition of micro-defects of 10um and above on the metal surface, effectively improving the accuracy and reliability of metal surface micro-defect detection.

[0033] In an exemplary embodiment, the trained defect detection model includes a backbone network, a neck network, and a detection head connected in sequence; wherein, the attention combination module is arranged in the backbone network and the neck network; the enhanced feature extraction module is arranged in the neck network; the four-branch stacking module is arranged in the neck network.

[0034] Reference Figure 3 , is a schematic structural diagram of the defect detection model provided by this application, which is obtained by improving the YOLOv7 network. As Figure 3 shown, the defect detection model includes a backbone network, a neck network, and a detection head, and the three networks are connected in sequence. The output of the backbone network is the input of the neck network, and the output of the neck network is the input of the detection head. Figure 3 The attention combination module in is used to extract multi-scale features; the enhanced feature extraction module is used to expand the receptive field and extract deep local features of the surface image; the four-branch stacking module is used to reduce the number of model parameters through network structure optimization. Among them, there are multiple attention combination modules and four-branch stacking modules, and there is 1 enhanced feature extraction module. The attention combination module is arranged in the backbone network and the neck network, the enhanced feature extraction module is arranged in the neck network, and the four-branch stacking module is also arranged in the neck network.

[0035] In addition, Figure 3The CBS module in it consists of Conv (convolution) + BatchNorm (batch normalization) + SiLU (activation function), and is used to extract local features and perform non-linear transformation. The Efficient Layer Aggregation Network (ELAN) is an efficient feature aggregation module introduced in YOLOv7, which is used to enhance the multi-scale representation ability of features. The ELAN module fuses features at different levels through a multi-branch structure to improve the performance of the model. In the figure, concatenation is an operation for feature fusion, which concatenates multiple feature maps in the channel dimension. Upsampling is an operation for increasing the spatial resolution of the feature map, usually implemented by interpolation methods. Reparameterized convolution is a convolution optimization technique that combines multiple convolution operations into one convolution operation through reparameterization, thereby improving the inference speed.

[0036] In this embodiment, the attention combination module can improve the detection ability for defects of different sizes by performing multi-scale feature extraction; the enhanced feature extraction module can improve the characterization ability for complex defects by expanding the receptive field and extracting deep local features; the four-branch stacking module can reduce the number of model parameters and computational cost through network structure optimization, making the model more suitable for deployment on resource-constrained devices. Thus, through the collaborative work of the attention combination module, the enhanced feature extraction module, and the four-branch stacking module, efficient, accurate, and lightweight defect detection can be achieved.

[0037] In an exemplary embodiment, the above-mentioned attention combination module includes a multi-scale attention sub-module and a downsampling feature extraction sub-module. The multi-scale attention sub-module is used to extract features by means of channel reshaping and grouping, encoding global information through parallel branches and recalibrating channel weights, and cross-dimensional aggregation of the output features of parallel branches.

[0038] Specifically, the attention combination module consists of an Efficient Multi-scale Attention (EMA) sub-module and a downsampling feature extraction sub-module in sequence. Among them, EMA is a network architecture used to improve the feature extraction ability of convolutional neural networks. This architecture realizes multi-scale feature optimization through three key steps: first, the channel reorganization technology is used to group the feature maps to form a multi-granularity feature representation; then, through parallel spatial and channel attention branches, global context information and inter-channel dependencies are captured respectively; finally, a cross-dimensional interaction mechanism is introduced to dynamically fuse the feature responses of different branches. Compared with the traditional attention mechanism, EMA has significant advantages in computational efficiency and can achieve more accurate feature calibration.

[0039] Refer to Figure 4, which is a schematic structural diagram of the multi-scale attention sub-module shown in an embodiment. The multi-scale attention sub-module reshapes some channels into the batch dimension, groups the channel dimension into multiple sub-features, so that the spatial semantic features are well distributed within each feature group; and, in addition to encoding global information in the parallel branch to recalibrate the channel weights, it also aggregates the output features of the two parallel branches through cross-dimensional interaction to capture pixel-level pairwise relationships, thereby improving the feature extraction effect while avoiding the side effects of channel dimension reduction.

[0040] In this embodiment, an efficient multi-scale attention mechanism is introduced to construct the EMA sub-module, which is combined with the downsampling feature extraction sub-module to form an attention combination module, which can effectively capture the multi-scale features of the image without reducing the channel dimension and improve the feature extraction effect.

[0041] In an exemplary embodiment, the enhanced feature extraction module is obtained by adjusting the spatial pyramid pooling cross-stage partial connection module through the principle of spatial pyramid fast pooling, and is used to extract deep local features through multi-scale pooling and cross-stage connection.

[0042] Specifically, the enhanced feature extraction module is obtained by improving the spatial pyramid pooling cross-stage partial connection module (Spatial Pyramid Pooling Cross Stage Partial Connections, SPPCSPC module) by referring to the principle of spatial pyramid fast pooling (Spatial Pyramid Pooling Fast, SPPF).

[0043] More specifically, in an exemplary embodiment, the enhanced feature extraction module includes two branches; in the first branch, the input content is processed by multiple convolutional blocks to obtain a first output, and three consecutive pooling operations are performed on the first output. After splicing the first output with the output results obtained from each pooling operation, it is further processed by two convolutional blocks to obtain the result of the first branch; in the second branch, the input content is processed by one convolutional block to obtain the result of the second branch; after splicing the result of the first branch and the result of the second branch, it is further processed by one convolutional block to obtain the output result of the enhanced feature extraction module.

[0044] Reference Figure 5, is a schematic structural diagram of an enhanced feature extraction module shown in an embodiment. The enhanced feature extraction module includes two branches. The first branch is processed by 3 CBS convolutional blocks to obtain a first output (the convolutional kernels are 1×1, 3×3, and 1×1 respectively). The first output is successively subjected to 3 max-pooling operations with a pooling kernel size of 5. Each pooling operation outputs a result. After concatenating the first output with the results of the 3 pooling operations, it is further processed by 2 CBS convolutional blocks (the convolutional kernels are 1×1 and 3×3 respectively) to obtain the result of the first branch. The second branch is processed by 1 CBS convolutional block of 1×1 to obtain the result of the second branch. After concatenating the result of the first branch and the result of the second branch, it is further processed by a 1×1 CBS convolutional block to obtain the output result of the enhanced feature extraction module. The original SPPCSPC module respectively performs 3 max-pooling operations on the first output with pooling kernel sizes of 5, 9, and 13, and then concatenates the results output after each pooling operation with the first output; compared with this, the enhanced feature extraction module obtains more deep features by adjusting the pooling kernel size and deeper pooling, thereby improving the model training effect.

[0045] Among them, the CBS module is composed of a convolutional layer (Conv), a batch normalization layer (BatchNorm), and a SiLU (activation function) in sequence, and is used to extract local features and perform non-linear transformation.

[0046] In this embodiment, the SPPCSPC module in the neck network is improved by referring to the SPPF design idea to form an enhanced feature extraction module. By expanding the receptive field, more deep local features can be extracted, thereby realizing more accurate recognition of the surface image and improving the confidence of subsequent surface image defect detection.

[0047] In an exemplary embodiment, the four-branch stacking module optimizes the network structure by adjusting the number and distribution of 1×1 convolutions and 3×3 convolutions, and reducing the concatenated branches.

[0048] Specifically, the four-branch stacking module includes two branches; in the first branch, the input content is processed by a 1×1 convolutional block to obtain a first output; in the second branch, the input content is processed by a 1×1 convolutional block to obtain a second output, and the second output is processed by a 3×3 convolutional block to obtain a third output. After concatenating the second output and the third output, it is input into a 1×1 convolutional block for processing to obtain a fourth output; the fourth output is processed by a 3×3 convolutional block to obtain a fifth output, and the fifth output is further processed by a 3×3 convolutional block to obtain a sixth output; after concatenating the first output, the fourth output, the fifth output, and the sixth output and inputting them into a 1×1 convolutional block for processing, the output result of the four-branch stacking module is obtained.

[0049] ReferenceFigure 6 , which is a schematic structural diagram of a four-branch stacked module shown in an embodiment. The CBS module in the figure consists of Conv (convolution) + BatchNorm (batch normalization) + SiLU (activation function), and is used to extract local features and perform non-linear transformation. The four-branch stacked module specifically includes two branches. The first branch obtains the first output after passing through a 1×1 CBS convolution block; the second branch obtains the second output after passing through a 1×1 CBS convolution block. The second output obtains the third output after passing through a 3×3 CBS convolution block. After splicing the second output and the third output results, the fourth output is obtained after passing through a 1×1 CBS convolution block; then the fourth output obtains the fifth output after passing through a 3×3 CBS convolution block, and the fifth output obtains the sixth output after passing through a 3×3 CBS convolution block again; finally, the first output, the fourth output, the fifth output, and the sixth output results are spliced and then pass through a 1×1 CBS convolution block to obtain the final output result. Compared with the traditional ELAN-H module, the four-branch stacked module contains 3 3×3 convolutions and 3 1×1 convolutions, while the ELAN-H module has 4 3×3 convolutions and 2 1×1 convolutions. Among them, 3×3 and 1×1 refer to the convolution kernel size, which are essentially parameters to be trained. And the tail of the four-branch stacked module is the splicing of 4 branches, while the ELAN-H module is the splicing of 6 branches, reducing two branches. Therefore, the four-branch stacked module reduces the number of model parameters by adjusting the quantity and distribution of 1×1 convolutions and 3×3 convolutions, and reducing the splicing branches at the tail.

[0050] In this embodiment, by designing a four-branch stacked module to replace all ELAN-H modules in the neck network, the number of model parameters and the computational cost can be reduced while maintaining the training effect unchanged.

[0051] In an exemplary embodiment, the above step S130 determines the actual defect size of the metal to be detected based on the size information of the bounding box, including: obtaining a pre-determined size equation; the size equation is obtained by fitting the actual defect size of the sample surface image with the size information of the bounding box of the defect area in the sample surface image; substituting the size information of the bounding box into the size equation to obtain the actual defect size of the metal to be detected.

[0052] Specifically, considering the irregularity of the shape of metal surface defects and the different effects of defects of different sizes on the performance of metal materials, the longest side length of the defect, which is more practically significant, is used to classify the defects. The specific steps are as follows: based on the diagonal length of the predicted bounding box of several sample points and the longest side length of the actual defect, draw the corresponding scatter plot to confirm its linear correlation; then use the least squares method to obtain the linear fitting equation as the size equation, and further substitute the size information of the bounding box obtained from subsequent actual detections into this size equation to obtain the actual defect size of the metal to be detected.

[0053] The core idea of least squares linear fitting is to find a straight line that minimizes the sum of the squares of the perpendicular distances (i.e., errors) from all data points to this line. Assume that the fitting line of the diagonal length of the predicted bounding box and the longest side length of the actual inclusion is , then for each , its fitted value is , so the sum of squared errors is: (5) To find the values of when it reaches the minimum value and , respectively find the partial derivatives of with respect to and : (6) (7) Set the two partial derivatives to zero and solve to obtain the fitting line equation. Substitute the data to solve and get the fitting equation as , is the longest side length of the actual defect, is the diagonal length of the predicted bounding box, which is plotted in the scatter plot as shown in Figure 7 . The coefficient of determination of the fitting equation is 0.89, indicating a high goodness of fit of the equation. Select 10 sample points and use the fitting equation to calculate the defect size. The specific data is shown in the following table; the fitting results show that the effective classification of defect sizes can be achieved through the fitting equation.

[0054] Table 1 Fitting data of the selected sample points In this embodiment, considering the irregularity of the defect shape and the differences in the effects of defects of different sizes on the metal properties, the process of determining the actual defect size of the metal to be detected based on the longest side length of the defect as the bounding box size information can provide a more accurate evaluation method for defect detection of metal materials.

[0055] In an exemplary embodiment, as shown in Figure 8 , the defect detection model is trained in the following manner: Step S810, obtain a sample data set; the sample data set includes sample surface images and the true bounding boxes of the sample surface images; Step S820, input the sample surface images into the initial defect detection model for defect detection to obtain predicted bounding boxes; Step S830: Obtain the intersection over union (IoU) between the predicted bounding box and the ground truth bounding box, as well as the shape similarity between the predicted bounding box and the ground truth bounding box. Based on the IoU and the shape similarity, determine the loss value between the predicted bounding box and the ground truth bounding box. Step S840: Take reducing the loss value as the training objective, and train the initial defect detection model to obtain a trained defect detection model.

[0056] In specific implementation, a batch of sample surface images can be collected in advance. The sample surface images can be obtained by scanning the sample metal with an ultrasonic scanning microscope, or can be the surface images of the metal scanned historically. Each sample surface image is labeled. For example, it is labeled through labeling software, and the sample surface image and its labeling information are combined to form a sample data set. The sample data set is randomly divided into a training set, a test set, and a validation set according to a preset ratio (such as 8:1:1) for model training, testing, and validation respectively.

[0057] In each training process, the sample surface image is input into the initial defect detection model for defect detection, and the predicted bounding box detected for the defect area is output. Calculate the loss value according to this predicted bounding box and the labeled ground truth bounding box. Specifically, it includes obtaining the IoU between the predicted bounding box and the ground truth bounding box, as well as the shape similarity between the predicted bounding box and the ground truth bounding box. Substitute the IoU and the shape similarity into the loss function to calculate the loss value between the predicted bounding box and the ground truth bounding box. If the loss value is greater than the threshold, adjust the model parameters with reducing the loss value as the training objective to obtain a new defect detection model. Then, input the next sample surface image into the new defect detection model for detection, calculate the loss value, and return to the step of comparing with the threshold until the obtained loss value converges or reaches the preset number of training times, and then end the training to obtain a trained defect detection model.

[0058] Among them, the calculation formula of the loss function can be expressed as: (8) Among them, IoU represents the intersection over union between the predicted bounding box and the ground truth bounding box, 、 represents the shape similarity between the predicted bounding box and the ground truth bounding box.

[0059] The calculation formula of the intersection over union can be expressed as: (9) Among them, A is the ground truth bounding box, and B is the predicted bounding box.

[0060] The calculation process of is as follows: (10) (11) (12) Among them, , represents the center coordinates of the predicted bounding box, , represents the center coordinates of the ground truth bounding box, represents the diagonal length of the smallest bounding rectangle that can simultaneously contain the predicted bounding box and the ground truth bounding box. , respectively represent the width and height of the ground truth bounding box, is a scaling factor, which is related to the size of the target defect in the dataset.

[0061] The calculation formula of is as follows: (13) (14) Among them, , represent the width and height of the predicted bounding box.

[0062] In some embodiments, after obtaining the sample dataset, the sample dataset can also be preprocessed first. For example, the surface images of each sample in the sample dataset are cropped into images of the same size, and the cropped images are cleaned. The cleaning includes operations such as removing unqualified images after cropping and images containing noise introduced during the scanning process, etc. Thus, a preprocessed sample dataset is obtained, and the model is trained using the preprocessed sample dataset.

[0063] In this embodiment, during the training of the defect detection model, when calculating the loss between the predicted bounding box and the ground truth bounding box, not only the overlapping area of the bounding boxes is considered, but also the shape similarity metric is introduced, which can more accurately measure the matching degree between the predicted bounding box and the ground truth bounding box, make the regression of the bounding box more accurate, thus better capture the detailed features of small targets, improve the detection accuracy of small targets, and thereby achieve the micro-defect detection of the metal surface image to be detected.

[0064] In one embodiment, when training and testing a defect detection model, the Pytorch deep learning framework can be used to train and test the model. For example, the deep learning framework selects cuda11.3 + torch1.12.1 + torchvision0.13.1, and at the same time deploys a compilation environment of Python3.9.16. The size of the input surface image is 992×992. During training, the maximum learning rate is set to 0.001, and the minimum learning rate is set to 0.00001. During the entire training process, the cosine annealing learning rate decay method is used to gradually reduce the learning rate from 0.001 to 0.00001. The batch size is set to 16, and the Adam (Adaptive Moment Estimation) optimizer is used, with the momentum set to 0.9, and the number of training epochs is set to 300. During training, the mosaic data augmentation method (creating a new composite image by stitching multiple images together) is combined with the mixup data augmentation method (linearly combining two images) to improve the robustness of the model.

[0065] To evaluate the accuracy and stability of the model, the mean average precision (mAP) and frames per second (FPS) are selected as evaluation metrics. mAP is a comprehensive evaluation metric that combines the precision and recall of different classes and is used to measure the detection accuracy of the model. The higher the value, the higher the prediction accuracy. When calculating, the intersection over union (IoU) threshold is selected as 0.5. FPS represents the speed at which the model processes frames per second and is used to measure the recognition speed of the model. The higher the value, the faster the prediction speed.

[0066] Table 2 Comparison of experimental results As can be seen from Table 2, compared with the traditional YOLOv7 model, the improved YOLOv7 model of this application not only reduces the number of network parameters and computational volume, but also improves the detection accuracy from 95.3% to 98.2%, and the recognition speed is also increased by 5%, verifying the effectiveness of the improved model of this application.

[0067] Figure 9 The figure shows a comparison diagram of different processing results. Among them, (a) is the result of manual annotation, (b) is the result processed by the built-in evaluation system of the ultrasonic scanning microscope, (c) is the prediction result using the original YOLOv7 model, and (d) is the prediction result using the improved YOLOv7 model of this application. In figure (b), the built-in evaluation system of the ultrasonic scanning microscope uses a relatively simple threshold segmentation method for defect analysis. Affected by the brightness and darkness of the scanned images, its applicability is poor and it cannot accurately analyze defects. Compared with the prediction result of the traditional YOLOv7 model in figure (c), the prediction result of the improved YOLOv7 model in figure (d) not only accurately identifies all inclusions, but also improves the recognition confidence.

[0068] Figure 10 It is a comparison chart of the detection quality between the traditional YOLOv7 model and the improved YOLOv7 model. After numbering 50 new preprocessed pictures, they are respectively input into the improved YOLOv7 model and the traditional YOLOv7 model for prediction, and compared with the manually marked results. The number of misdetections and missed detections for every 10 pictures is counted to obtain Figure 10 the results shown. It can be seen from the figure that the actual detection effect of the improved YOLOv7 model is better than that of the original YOLOv7 model. Generally speaking, the detection performance of the improved YOLOv7 model is better.

[0069] It should be understood that although each step in the flowcharts involved in the above-described embodiments is shown in sequence according to the indication of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be completed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0070] Based on the same inventive concept, the embodiments of the present application also provide a metal surface micro-defect detection system for implementing the above-mentioned metal surface micro-defect detection method based on deep learning. The implementation solutions provided by this system to solve problems are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in one or more embodiments of the metal surface micro-defect detection system based on deep learning provided below can refer to the limitations on the metal surface micro-defect detection method based on deep learning in the above text, and will not be repeated here.

[0071] In one embodiment, as Figure 11 shown, a metal surface micro-defect detection system based on deep learning is provided, including: An acquisition unit 1101 scans the surface of the metal to be detected with an ultrasonic scanning microscope, acquires the surface image of the metal to be detected and performs preprocessing; The detection unit 1102 calls a trained defect detection model on the computer device to perform defect detection on the surface image, and obtains the bounding box of the defect area in the surface image. Among them, the trained defect detection model includes an attention combination module, an enhanced feature extraction module, and a four-branch stacking module. The attention combination module is used to extract multi-scale features of the surface image. The enhanced feature extraction module extracts deep local features of the surface image by expanding the receptive field. The four-branch stacking module reduces the number of model parameters through network structure optimization. The determination unit 1103 determines the actual defect size of the metal to be detected based on the size information of the bounding box.

[0072] In one embodiment, the trained defect detection model includes a backbone network, a neck network, and a detection head connected in sequence. The attention combination module is arranged in the backbone network and the neck network. The enhanced feature extraction module is arranged in the neck network. The four-branch stacking module is arranged in the neck network.

[0073] In one embodiment, the attention combination module includes a multi-scale attention sub-module and a down-sampling feature extraction sub-module. The multi-scale attention sub-module adopts a three-stage feature optimization mechanism to achieve efficient feature enhancement. First, through intelligent reshaping and grouping processing in the channel dimension, multi-granularity feature expressions are established. Secondly, a parallel branch structure is adopted to synchronously capture global context information, and adaptive re-weighting of feature channels is achieved based on the channel attention mechanism. Finally, through an innovative cross-dimensional feature aggregation strategy, the multi-scale features output by each branch are deeply fused, thereby constructing an enhanced feature representation with rich semantic information.

[0074] In one embodiment, the enhanced feature extraction module is obtained by adjusting the spatial pyramid pooling cross-stage partial connection module based on the principle of spatial pyramid fast pooling; it is used to extract deep local features through multi-scale pooling and cross-stage connection.

[0075] In one embodiment, the enhanced feature extraction module includes two branches. In the first branch, the input content is processed through multiple convolutional blocks to obtain a first output. Three consecutive pooling operations are performed on the first output. After splicing the first output with the output results obtained from each pooling operation, it is further processed through two convolutional blocks to obtain the result of the first branch. In the second branch, the input content is processed through one convolutional block to obtain the result of the second branch. After splicing the result of the first branch and the result of the second branch, it is further processed through one convolutional block to obtain the output result of the enhanced feature extraction module.

[0076] In one embodiment, the four-branch stacking module optimizes the network structure by adjusting the quantity and distribution of 1×1 convolutions and 3×3 convolutions, and reducing the splicing branches.

[0077] In one embodiment, the four-branch stacking module includes two branches; in the first branch, the input content is processed by a 1×1 convolution block to obtain a first output; in the second branch, the input content is processed by a 1×1 convolution block to obtain a second output, the second output is processed by a 3×3 convolution block to obtain a third output, the second output and the third output are spliced and then input into a 1×1 convolution block for processing to obtain a fourth output; the fourth output is processed by a 3×3 convolution block to obtain a fifth output, and the fifth output is further processed by a 3×3 convolution block to obtain a sixth output; the first output, the fourth output, the fifth output, and the sixth output are spliced and then input into a 1×1 convolution block for processing to obtain the output result of the four-branch stacking module.

[0078] In one embodiment, the computer device may be a server, and its internal structure diagram may be as Figure 12 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data during the metal surface micro-defect detection based on deep learning. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for detecting metal surface micro-defects based on deep learning.

[0079] Those skilled in the art can understand that Figure 12 the structure shown in

[0080] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0081] In one embodiment, the system further includes a model training module configured to obtain a sample data set; the sample data set includes sample surface images and the true bounding boxes of the sample surface images; input the sample surface images into an initial defect detection model for defect detection to obtain predicted bounding boxes; obtain the intersection over union (IoU) between the predicted bounding boxes and the true bounding boxes, as well as the shape similarity between the predicted bounding boxes and the true bounding boxes, and determine the loss value between the predicted bounding boxes and the true bounding boxes based on the IoU and the shape similarity; and train the initial defect detection model with the objective of reducing the loss value to obtain a trained defect detection model.

[0082] Each module in the above metal surface micro-defect detection system based on deep learning can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in the processor of a computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0083] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0084] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0085] The above-described embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of this application. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.

Claims

1. A metal surface micro-defect detection method based on deep learning, characterized in that: The method comprises: Acquire the surface image of the metal to be detected and perform preprocessing; Calling a trained defect detection model to perform defect detection on the surface image to obtain a boundary box of a defect area in the surface image; Determining the actual defect size of the metal to be detected based on the size information of the bounding box; Among them, the trained defect detection model includes an attention combination module, an enhanced feature extraction module, and a four-branch stacking module; the attention combination module is used to extract the multi-scale features of the surface image; the enhanced feature extraction module extracts the deep local features of the surface image by expanding the receptive field; the four-branch stacking module reduces the number of model parameters by optimizing the network structure.

2. The metal surface micro-defect detection method based on deep learning according to claim 1 is characterized in that: The specific steps of obtaining the surface image of the metal to be detected and preprocessing are as follows: Use an ultrasonic scanning microscope equipped with a 300MHz high-frequency probe and pure distilled water as a coupling agent to perform C-Scan scanning with a step of 1um to obtain high-resolution metal surface images; The collected high-resolution images are cropped into images of the same size, and the cropped image data is cleaned, including removing images that are unqualified after cropping and images that contain noise introduced during the scanning process.

3. The metal surface micro-defect detection method based on deep learning according to claim 1 is characterized in that: The trained defect detection model includes a backbone network, a neck network and a detection head connected in sequence; The attention combination module is arranged in the backbone network and the neck network; The enhanced feature extraction module is arranged in the neck network; The four-branch stacking module is arranged in the neck network.

4. The metal surface micro-defect detection method based on deep learning according to claim 2 is characterized in that: The attention combination module includes a multi-scale attention submodule and a downsampling feature extraction submodule; The multi-scale attention submodule adopts a three-stage feature optimization mechanism to achieve efficient feature enhancement. Firstly, a multi-granularity feature expression is established through intelligent reshaping and grouping processing of the channel dimension. Secondly, a parallel branch architecture is used to synchronously capture the global context information, and adaptive reweighting of feature channels is achieved based on the channel attention mechanism. Finally, an innovative cross-dimensional feature aggregation strategy is used to deeply fuse the multi-scale features output by each branch, thereby constructing an enhanced feature representation with rich semantic information.

5. The metal surface micro-defect detection method based on deep learning according to claim 2 is characterized in that: The enhanced feature extraction module is obtained by adjusting the spatial pyramid pooling cross-stage partial connection module through the spatial pyramid fast pooling principle; and is used to extract deep local features through multi-scale pooling and cross-stage connection.

6. The metal surface micro-defect detection method based on deep learning according to claim 4 is characterized in that: The enhanced feature extraction module includes two branches; in the first branch, the input content is processed by multiple convolution blocks to obtain a first output, the first output is subjected to three consecutive pooling operations, the first output is concatenated with the output result obtained by each pooling operation, and then processed by two convolution blocks to obtain the first branch result; In the second branch, the input content is processed by a convolution block to obtain a second branch result; After the first branch result and the second branch result are spliced, they are processed by a convolution block to obtain the output result of the enhanced feature extraction module.

7. The metal surface micro-defect detection method based on deep learning according to claim 2 is characterized in that: The four-branch stacking module optimizes the network structure by adjusting the number and distribution of 1×1 convolutions and 3×3 convolutions, and reducing splicing branches.

8. The metal surface micro-defect detection method based on deep learning according to claim 6 is characterized in that: The four-branch stacking module includes two branches; In the first branch, the input content is processed by a 1×1 convolution block to obtain the first output; In the second branch, the input content is processed by a 1×1 convolution block to obtain a second output, the second output is processed by a 3×3 convolution block to obtain a third output, the second output and the third output are concatenated and input into a 1×1 convolution block to obtain a fourth output; the fourth output is processed by a 3×3 convolution block to obtain a fifth output, and the fifth output is further processed by a 3×3 convolution block to obtain a sixth output; The first output, the fourth output, the fifth output and the sixth output are concatenated and input into a 1×1 convolution block for processing to obtain the output result of the four-branch stacking module.

9. A metal surface micro-defect detection system based on deep learning, characterized in that: The system comprises: An acquisition unit, which scans the metal surface to be detected with an ultrasonic scanning microscope, acquires the surface image of the metal to be detected and performs preprocessing; The detection unit calls a defect detection model trained on a computer device to perform defect detection on the surface image to obtain a boundary box of a defect area in the surface image; wherein the trained defect detection model includes an attention combination module, an enhanced feature extraction module, and a four-branch stacking module; the attention combination module is used to extract multi-scale features of the surface image; the enhanced feature extraction module extracts deep local features of the surface image by expanding the receptive field; the four-branch stacking module reduces the amount of model parameters by optimizing the network structure; A determination unit determines an actual defect size of the metal to be detected based on the size information of the boundary box.

Citation Information

Patent Citations

  • Metal surface defect detection method based on improved YOLOv7 model

    CN116721291A

  • Lightweight multi-scale aluminum profile surface defect detection method

    CN117911399A

  • PCB defect detection method based on global context attention mechanism

    CN117974544A

  • Steel surface defect detection method and system based on computer vision

    CN118037692A

  • Improved steel surface defect detection method

    CN118351109A

Cited By

  • Animal behavior anomaly detection method based on deep learning

    CN121191056A