Surface defect detection method and equipment based on multi-scale adaptive guidance and medium

Through the multi-scale adaptive guided surface defect detection method, multi-scale feature information is integrated to solve the problems of high missed detection rate and false detection rate of traditional deep learning in industrial inspection, and achieve higher-precision defect detection.

CN120672758AActive Publication Date: 2025-09-19NANCHANG HANGKONG UNIVERSITY

Patent Information

Application Number
CN202511178512.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-09-19
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Traditional deep learning technology has problems with high missed detection and false detection rates in industrial surface defect detection, especially when facing weak textures and multi-scale targets.

Method used

A surface defect detection method based on multi-scale adaptive guidance is adopted. Through the combination of multi-scale feature extraction sub-network, multi-scale adaptive guidance sub-network, feature fusion sub-network and detection head sub-network, multi-scale feature information is integrated to reduce the loss of weak texture defects and minor defects. The model is optimized through full-dimensional dynamic feature extraction and hybrid loss function.

Benefits of technology

The accuracy of surface defect detection is improved, the missed detection rate and false detection rate are reduced, and the detection capability of complex surface defects is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672758A_ABST
    Figure CN120672758A_ABST
Patent Text Reader

Abstract

The invention discloses a surface defect detection method and device based on multi-scale adaptive guidance and a medium, and relates to the technical field of surface defect detection.The method comprises the steps that an image of a to-be-detected object is obtained and input into a pre-trained surface defect detection network, and a surface defect detection result of the to-be-detected object is predicted; wherein a multi-scale adaptive guide sub-network is introduced between a multi-scale feature extraction sub-network and a feature fusion sub-network in the surface defect detection network so as to reduce the loss of weak and small target information, and a full-dimensional dynamic feature extraction module is introduced into the feature fusion sub-network so as to improve the detection accuracy. According to the method, the feature extraction capability of a multi-scale target and a complex-shape target is enhanced, and finally, a modulation factor capable of changing a weight parameter is introduced into a bounding box regression loss item, so that when defect detection is carried out based on a surface defect detection network, the precision of surface defect detection is improved, and the omission ratio and the false detection rate of industrial surface defect detection are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of surface defect detection, and in particular to a surface defect detection method, device and medium based on multi-scale adaptive guidance. Background Art

[0002] Surface defect detection, a technology used to identify and locate surface flaws in products, is widely used in industrial production and is crucial for ensuring product quality, reducing production costs, and improving efficiency. However, traditional manual visual inspection methods have numerous shortcomings. They are not only inefficient but also susceptible to subjective factors, resulting in unstable inspection results. With the continuous advancement of industrial automation and intelligentization, the industry urgently needs an efficient, stable, and automated technology to replace traditional manual inspection methods.

[0003] In recent years, deep learning-based industrial surface defect detection methods have been widely used in modern manufacturing due to their high precision, robustness, and automation capabilities. However, traditional deep learning techniques still suffer from high rates of missed detection and false detection when applied to defects with weak textures, multi-scales, and complex shapes in real industrial scenarios. Summary of the Invention

[0004] The purpose of this application is to provide a surface defect detection method, equipment and medium based on multi-scale adaptive guidance, which can improve the accuracy of surface defect detection and reduce the missed detection rate and false detection rate of industrial surface defect detection.

[0005] To achieve the above objectives, this application provides the following solutions: In a first aspect, the present application provides a surface defect detection method based on multi-scale adaptive guidance, comprising: Acquire an image of the object to be detected; The image of the object to be inspected is input into a pre-trained surface defect detection network to predict the surface defect detection results of the object to be inspected; the surface defect detection network includes a multi-scale feature extraction subnetwork, a multi-scale adaptive guidance subnetwork, a feature fusion subnetwork and a detection head subnetwork connected in sequence; the multi-scale adaptive guidance subnetwork is used to adaptively fuse multiple feature maps of different scales extracted by the multi-scale feature extraction subnetwork to obtain adaptive fused feature maps of different scales.

[0006] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned surface defect detection method based on multi-scale adaptive guidance.

[0007] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned surface defect detection method based on multi-scale adaptive guidance.

[0008] According to the specific embodiments provided in this application, this application discloses the following technical effects: The present application provides a surface defect detection method, device and medium based on multi-scale adaptive guidance, which obtains the image of the object to be detected and inputs it into a pre-trained surface defect detection network to predict the surface defect detection results of the object to be detected; wherein, a multi-scale adaptive guidance subnetwork is introduced between the multi-scale feature extraction subnetwork and the feature fusion subnetwork in the surface defect detection network, which effectively integrates multi-scale feature information and can also reduce the loss of weak texture defects and minor defects, thereby improving the accuracy of surface defect detection and reducing the missed detection rate and false detection rate of industrial surface defect detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0010] Figure 1 This is a diagram of an application environment of a surface defect detection method based on multi-scale adaptive guidance in one embodiment of the present application; Figure 2 A schematic flow chart of a surface defect detection method based on multi-scale adaptive guidance provided in one embodiment of the present application; Figure 3 A schematic diagram of an image of an object to be detected provided in one embodiment of the present application; Figure 4 A schematic diagram of surface defect detection results of an object to be detected provided in one embodiment of the present application; Figure 5 A schematic diagram of the structure of a surface defect detection network provided in one embodiment of the present application; Figure 6 A schematic diagram of the structure of a multi-scale feature extraction subnetwork provided in one embodiment of the present application; Figure 7 A schematic diagram of the structure of a multi-scale adaptive guidance sub-network provided in one embodiment of the present application; Figure 8 A schematic diagram of the structure of a multi-scale adaptive fusion module provided in one embodiment of the present application; Figure 9A schematic diagram of the structure of a full-dimensional dynamic feature extraction module provided in one embodiment of the present application; Figure 10 A diagram of the full-dimensional dynamic convolution structure provided in one embodiment of the present application; Figure 11 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0011] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0012] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0013] The surface defect detection method based on multi-scale adaptive guidance provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal communicates with the server through the network. The data storage system can store the data that the server needs to process. The data storage system can be set up separately, integrated on the server, or placed on the cloud or other servers. The terminal can send the image of the object to be detected to the server. After the server receives the image of the object to be detected, the server inputs the image of the object to be detected into the pre-trained surface defect detection network to predict the surface defect detection result of the object to be detected; the surface defect detection network includes a multi-scale feature extraction subnetwork, a multi-scale adaptive guidance subnetwork, a feature fusion subnetwork and a detection head subnetwork connected in sequence; the multi-scale adaptive guidance subnetwork is used to adaptively fuse the feature maps of multiple different scales extracted by the multi-scale feature extraction subnetwork to obtain adaptive fusion feature maps of different scales. The server can feed back the obtained surface defect detection result of the object to be detected to the terminal. In addition, in some embodiments, the surface defect detection method based on multi-scale adaptive guidance can also be implemented by a server or a terminal alone, such as the terminal can directly perform surface defect detection based on multi-scale adaptive guidance on the image of the object to be detected, or the server can obtain the image of the object to be detected from the data storage system and perform surface defect detection based on multi-scale adaptive guidance.

[0014] The terminals may be, but are not limited to, various desktop computers, laptops, smart phones, tablet computers, IoT devices, and portable wearable devices. The server may be implemented as an independent server or a server cluster consisting of multiple servers, or as a cloud server.

[0015] In an exemplary embodiment, Figure 2 As shown, a surface defect detection method based on multi-scale adaptive guidance is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The server in is used as an example for description, including the following steps 101 to 102.

[0016] Step 101: Obtain an image of the object to be detected. Figure 3 Shown is an image from a dataset of industrial steel surface defects.

[0017] Step 102: Input the image of the object to be detected into the pre-trained surface defect detection network to predict the surface defect detection result of the object to be detected, such as Figure 4 As shown. Figure 5 As shown, the surface defect detection network includes a multi-scale feature extraction subnetwork, a multi-scale adaptive guidance subnetwork, a feature fusion subnetwork and a detection head subnetwork connected in sequence; the multi-scale adaptive guidance subnetwork is used to adaptively fuse multiple feature maps of different scales extracted by the multi-scale feature extraction subnetwork to obtain adaptive fused feature maps of different scales.

[0018] Implement the above steps 101 to 102, obtain the image of the object to be detected and input it into the pre-trained surface defect detection network to predict the surface defect detection result of the object to be detected; wherein, a multi-scale adaptive guidance sub-network is introduced between the multi-scale feature extraction sub-network and the feature fusion sub-network in the surface defect detection network, which effectively integrates the multi-scale feature information and reduces the loss of weak texture defects and minor defects, thereby improving the accuracy of surface defect detection and reducing the missed detection rate and false detection rate of industrial surface defect detection.

[0019] In another exemplary embodiment of the present application, in step 102, the image of the object to be inspected is input into a pre-trained surface defect detection network to predict the surface defect detection result of the object to be inspected, which specifically includes: (1) Input the image of the object to be detected into the multi-scale feature extraction sub-network to extract multiple feature maps of different scales.

[0020] (2) A multi-scale adaptive guidance sub-network is used to adaptively fuse feature maps of multiple scales to obtain adaptive fused feature maps of different scales.

[0021] (3) The feature fusion sub-network is used to splice and fuse the adaptive fusion features of different scales to obtain a spliced ​​fusion feature map.

[0022] (4) Input the spliced ​​fusion feature map into the detection head sub-network and output the surface defect detection results of the object to be detected.

[0023] In another exemplary embodiment of the present application, Figure 6 As shown, the multi-scale feature extraction sub-network includes: N FasterNet modules connected in series (corresponding to Figure 6 The input of the first FasterNet module is the image of the object to be detected; each FasterNet module outputs feature maps of different scales. As an example, N is 4. Figure 6 Four feature maps of different scales are shown, corresponding to Figure 6 C1 to C4 in. Figure 6 In the example, an Embedding module is connected in series to the input of the first FasterNet module, and a Merging module is provided between two adjacent FasterNet modules. Figure 6 The H and W in the image represent the length and width of the image respectively.

[0024] In another exemplary embodiment of the present application, Figure 7 As shown, the multi-scale adaptive guidance sub-network includes: N-2 multi-scale adaptive fusion modules. As an example, N is 4. Figure 7 Two multi-scale adaptive fusion modules are shown in Figure 7 MAG in ).

[0025] like Figure 8 As shown, each multi-scale adaptive fusion module includes an upsampling layer (corresponding to Figure 8 2x up in), downsampling layer (corresponding to Figure 8 2x down in ), first splicing layer, second splicing layer, third splicing layer, first convolution layer, second convolution layer, activation layer (corresponding to Figure 8 Sigmiod in ), multiplication layer and the first addition layer.

[0026] The output of the upsampling layer is connected to the input of the first splicing layer, the output of the downsampling layer is connected to the input of the second splicing layer, the output of the first splicing layer is connected to the input of the third splicing layer via the first convolutional layer, the output of the second splicing layer is connected to the input of the third splicing layer via the second convolutional layer, the output of the third splicing layer is connected to the input of the multiplication layer via the activation layer, the output of the multiplication layer is connected to the input of the first addition layer, and the output of the first addition layer is connected to the input of the feature fusion sub-network.

[0027] The upsampling layer of the m-th multi-scale adaptive fusion module is connected to the output of the m+2-th FasterNet module, the downsampling layer of the m-th multi-scale adaptive fusion module is connected to the output of the m-th FasterNet module, and the inputs of the first splicing layer, the second splicing layer, the multiplication layer, and the first addition layer of the m-th multi-scale adaptive fusion module are also connected to the output of the m+1-th FasterNet module; m = 1, 2, ..., N-2; N represents the number of scale types of feature maps extracted by the multi-scale feature extraction subnetwork.

[0028] This application constructs a Figure 8 The multi-scale adaptive fusion module shown in the figure fully utilizes the shallow and deep features at different levels in the multi-scale feature extraction subnetwork. By adaptively fusing the detail information in the shallow layers with the semantic information in the deep layers, this module not only effectively integrates multi-scale feature information but also reduces the loss of weak texture defects and minor defects. The calculation formula for the mth multi-scale adaptive fusion module is shown below: ; Where, 、 They represent the shallow feature map (feature map of the mth scale), the intermediate feature map (feature map of the m+1th scale), and the deep feature map (feature map of the m+2th scale) input to the multi-scale adaptive fusion module respectively; and Respectively represent the shallow feature map and deep feature map after preliminary fusion; among them, represents the downsampling operation; Represents an upsampling operation; Indicates that the feature map is spliced ​​on the feature channel; Represents the convolution operation; Indicates that the feature maps are spliced ​​in the feature space; Represents the Sigmoid activation function; Represents the multiplication operation of the feature map and the generated weight; Represents the element-by-element addition operation of the feature map; Represents the feature map output by the multi-scale adaptive fusion module. Figure 7Taking the feature maps of the four scales shown in the figure as an example, the multi-scale adaptive guidance sub-network is used to adaptively optimize the feature maps of different levels output by the multi-scale feature extraction sub-network. The calculation formula is shown as follows: ; Where, 、 、 and Represents the feature maps of four different scales output by the multi-scale feature extraction sub-network; Indicates that the multi-scale adaptive fusion module performs weighted addition operation on the feature map; and Represents the adaptive fusion feature maps of two different scales output by the multi-scale adaptive fusion module.

[0029] In another exemplary embodiment of the present application, Figure 5 As shown, the feature fusion sub-network includes: N-2 first splicing modules connected in sequence, a first full-dimensional dynamic feature extraction module (corresponding to Figure 5 ODFE Block1 in the ODFE Block1), three second full-dimensional dynamic feature extraction modules connected in sequence (corresponding to Figure 5 ODFE Block2 in the ODFE Block) and a second splicing module provided between every two adjacent second full-dimensional dynamic feature extraction modules.

[0030] The output of the last first splicing module is connected to the input of the first second full-dimensional dynamic feature extraction module.

[0031] The inputs of the N-2 first splicing modules are connected to the outputs of the N-2 multi-scale adaptive fusion modules in a one-to-one correspondence.

[0032] The input of the first second splicing module is also connected to the output of the first first full-dimensional dynamic feature extraction module; the input of the second second splicing module and the input of the first first splicing module are also connected to the output of the last FasterNet module in the multi-scale feature extraction subnetwork.

[0033] The outputs of the three second full-dimensional dynamic feature extraction modules are connected to the inputs of the three detection heads in the detection head sub-network in a one-to-one correspondence.

[0034] The feature fusion sub-network is used to accurately locate defects and perform splicing and fusion operations on the feature maps of different resolutions output by the multi-scale adaptive guidance sub-network. Specifically, this process covers two methods: top-down splicing and fusion and bottom-up splicing and fusion. Figure 5 Taking the splicing and fusion of the outputs of two multi-scale adaptive fusion modules as an example, the corresponding calculation formulas are shown as follows: ; ; Where, Represents the Nth (N=4) scale feature map output by the multi-scale feature extraction subnetwork; and They represent feature maps of two different scales output by the multi-scale adaptive guidance sub-network; Feature maps are spliced ​​on feature channels; and Respectively represent the first full-dimensional dynamic feature extraction module and the second full-dimensional dynamic feature extraction module; and Respectively represent the feature maps of two different scales output by top-down splicing fusion, for Figure 5 The output of ODFE Block1, for Figure 5 The output of the first ODFE Block2; and They represent two feature maps of different scales output by bottom-up splicing and fusion; for Figure 5 The output of the second ODFE Block2, It is the output of the third ODFE Block2.

[0035] In another exemplary embodiment of the present application, Figure 9 As shown, Figure 9 (a) is the specific structure of the full-dimensional dynamic feature extraction module (referring to the first full-dimensional dynamic feature extraction module or the second full-dimensional dynamic feature extraction module). Taking the first full-dimensional dynamic feature extraction module as an example, the first full-dimensional dynamic feature extraction module includes: the first CBS layer, the second CBS layer, the segmentation layer (corresponding to Figure 9 Split in), the fourth splicing layer (corresponding to Figure 9 Concat in ) and multiple ODC Botleneck modules connected in series. Figure 9 (b) in the figure shows the specific structure of the ODC Bottleneck module. Each ODC Bottleneck module includes a second addition layer and multiple full-dimensional dynamic convolution layers; multiple full-dimensional dynamic convolution layers and the second addition layer are connected in sequence; the input of the second addition layer is also connected to the input of the first full-dimensional dynamic convolution layer. Figure 9 As shown in (b), the full-dimensional dynamic convolution layer uses batch normalization BN and activation function SiLU.

[0036] The input of the first CBS layer is connected to the output of the first splicing module to which the input end of the first full-dimensional dynamic feature extraction module belongs.

[0037] The output of the first CBS layer is connected to the input of the segmentation layer, and the output of the segmentation layer is connected to the input of the fourth splicing layer and the input of the first full-dimensional dynamic convolution layer of the first ODC Bottleneck module. The output of the second addition layer of each ODC Bottleneck module is also connected to the input of the fourth splicing layer, and the output of the fourth splicing layer is connected to the input of the second CBS layer. The output of the second CBS layer is the output of the first full-dimensional dynamic feature extraction module to which it belongs.

[0038] This application constructs a Figure 9 The full-dimensional dynamic feature extraction module shown (referring to the first full-dimensional dynamic feature extraction module or the second full-dimensional dynamic feature extraction module) realizes more efficient feature extraction. The calculation formula is as follows:

[0039] Where, and They represent the input feature map and output feature map of the full-dimensional dynamic feature extraction module respectively; and Indicates that the feature map is segmented on the feature channel ( ) Output feature map; and They represent the input feature map and output feature map of the preliminary extraction module (ODC Bottleneck module); () represents the full-dimensional dynamic convolution operation; ODC_Neck() represents the execution operation of the ODC Bottleneck module. Conv() in this formula refers to Figure 9 In the “CBS, S=1, K=1”, S (stride) represents the stride of the convolution and K (kernal size) represents the size of the convolution kernel.

[0040] This application combines full-dimensional dynamic convolution with a feature fusion subnetwork to define a new feature fusion subnetwork, which can be called a feature refinement fusion subnetwork. The convolution kernel sampling position of the full-dimensional dynamic convolution is no longer fixed, but is dynamically adjusted according to the characteristics of the input feature map. It can more effectively adapt to the feature extraction of complex surface defect structures, allowing the feature fusion subnetwork to obtain more accurate defect location information. Figure 10 As shown, the full-dimensional dynamic convolution formula is as follows:

[0041] In the formula, the full-dimensional dynamic convolution introduces a multi-dimensional attention mechanism with a parallel strategy. By multiplying different attention weights along the convolution position, channel, filter, and convolution kernel dimensions, the convolution operation can adapt to the differences in each dimension of the input. and Represents the input feature map and output feature map (with high and width of / aisle); Represents the convolution kernel The attention scalar, i=1, 2, ..., n; n represents the number of dynamic convolution kernels; , and Represents the convolution kernel along the spatial dimension, input channel dimension and The kernel space outputs three attention scalars in the channel dimension; Represents multiplication operations along different dimensions of the kernel space. Figure 10 In , k represents the k×k spatial position; That is the convolution kernel ; Represents a convolution filter.

[0042] Figure 10 It describes how the feature map is dot-multiplied with the attention weights generated in four different dimensions in the full-dimensional dynamic convolution. Figure 10 (a) to (d) in the figure respectively express the multiplication of feature maps with different weights of convolution at different positions, the multiplication of feature maps with different weights on different channels, the multiplication of feature maps with different weights on different filters, and the multiplication of feature maps with different weights on different convolution kernels.

[0043] In another exemplary embodiment of the present application, the surface defect detection method based on multi-scale adaptive guidance further includes: a training process of a pre-trained surface defect detection network, specifically: (1) Obtain image samples of the object to be inspected; each image sample corresponds to a real surface defect detection result.

[0044] (2) Input the image sample of the object to be detected into the initial surface defect detection network and output the surface defect detection sample prediction result.

[0045] (3) Construct a loss function and calculate the loss error based on the predicted results of the surface defect detection samples and the corresponding actual surface defect detection results; the loss function includes a bounding box regression loss term and a classification loss term; a modulation factor that changes the weight coefficient is introduced into the bounding box regression loss term.

[0046] This application defines a new bounding box regression loss function: FCIOU loss function. This loss function introduces a modulation factor that can change the weight coefficient, giving a larger loss to the regression box with a higher IOU value, thereby accelerating the convergence of the prediction box and improving the regression accuracy. The calculation formula is as follows: ; Where, represents the bounding box regression loss; represents the FCIOU loss function; represents the CIOU loss function; IOU is the intersection-union ratio between the predicted box and the real box; the modulation factor It is a hyperparameter with a value range between 0 and 1; the CIOU loss function calculation formula is as follows: ; Where, Represents the center point of the prediction box The center point of the real frame The Euclidean distance between is the diagonal distance between the minimum bounding rectangle of the predicted box and the real box; and Represents the difference between the aspect ratio of the predicted box and the real box; IOU is the intersection-union ratio between the predicted box and the real box, and its calculation formula is as follows: ; Where, and Represent the predicted box and the true box respectively; The bounding box regression loss function and classification loss function are used to calculate the bounding box regression loss and classification loss between the target prediction result obtained by the model and the target's true annotation information. The classification loss calculation formula is as follows: ; Where, is the sample size, is the number of categories, is an image sample Corresponding category The true label (0 or 1), is the surface defect detection network for image samples Predicted as class probability.

[0047] Therefore, the loss function of the surface defect detection network is expressed as: ; Where L represents the loss error; and is the loss term weight.

[0048] (4) The initial surface defect detection network is adjusted according to the back propagation of the loss error until the loss error converges or the maximum number of training iterations is reached, and a pre-trained surface defect detection network is obtained.

[0049] In this application, a multi-scale feature extraction sub-network is used to quickly extract coarse features from the input image of the object to be detected; a multi-scale adaptive guidance sub-network is constructed to adaptively optimize the extracted multi-scale feature maps; a feature fusion sub-network is constructed to first accurately locate defects on the feature maps of different resolutions obtained by the multi-scale adaptive guidance sub-network, and then perform splicing and fusion; the spliced ​​and fused feature maps are then sent to the detection head sub-network to obtain the target prediction results; a novel hybrid loss function is constructed to calculate the classification loss and bounding box regression loss for the target prediction results obtained by the detection head and the true target annotation information, and the model parameters of the surface defect detection network are adjusted through back propagation based on these two types of losses to obtain a surface defect detection network based on multi-scale adaptive guidance. This application introduces a multi-scale adaptive guidance sub-network between the multi-scale feature extraction sub-network and the feature fusion sub-network to reduce the loss of weak target information, and introduces a full-dimensional dynamic feature extraction module in the feature fusion sub-network to enhance the feature extraction capabilities of multi-scale targets and complex shape targets, and construct an improved bounding box regression loss function. This application adopts a multi-scale adaptive guidance sub-network to fully utilize the shallow features and deep features at different levels in the multi-scale feature extraction sub-network to reduce the loss of weak texture defects and tiny defects; further adopts a feature fusion sub-network to accurately adapt to the feature extraction of complex surface defect structures; finally, combined with a new hybrid loss function, the surface defect detection network's ability to locate and identify defects is improved, the false detection rate and missed detection rate of industrial surface defect detection are reduced, and the robustness of the model is improved.

[0050] The present application also provides an application scenario that applies the above-mentioned surface defect detection method based on multi-scale adaptive guidance. Specifically: The surface defect detection method based on multi-scale adaptive guidance provided in this embodiment can be applied in the industrial steel surface defect detection scenario. The scenario includes an image acquisition link and a surface defect detection link; the image acquisition link is used to acquire defect images of the industrial steel to be inspected; the surface defect detection link is used to identify and locate the surface defects of the industrial steel to be inspected based on the acquired defect images. The surface defect detection method based on multi-scale adaptive guidance provided in this embodiment belongs to the surface defect detection link.

[0051] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 11As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store surface defect detection results based on the surface defect detection network and the predicted object to be detected. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a surface defect detection method based on multi-scale adaptive guidance is implemented.

[0052] Those skilled in the art will understand that Figure 11 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above-mentioned method embodiments when executing the computer program.

[0053] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0054] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0055] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0056] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0057] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0058] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A surface defect detection method based on multi-scale adaptive guidance, characterized in that: include: Acquire an image of the object to be detected; Input the image of the object to be inspected into the pre-trained surface defect detection network to predict the surface defect detection results of the object to be inspected; The surface defect detection network includes a multi-scale feature extraction subnetwork, a multi-scale adaptive guidance subnetwork, a feature fusion subnetwork and a detection head subnetwork connected in sequence; the multi-scale adaptive guidance subnetwork is used to adaptively fuse multiple feature maps of different scales extracted by the multi-scale feature extraction subnetwork to obtain adaptive fused feature maps of different scales.

2. The surface defect detection method based on multi-scale adaptive guidance according to claim 1, characterized in that: The multi-scale feature extraction sub-network includes: N FasterNet modules connected in series; the input of the first FasterNet module is the image of the object to be detected; each FasterNet module outputs a feature map of different scales; The feature fusion subnetwork is used to splice and fuse adaptive fusion features of different scales to obtain a spliced ​​fusion feature map; the detection head subnetwork is used to output the surface defect detection result of the object to be detected based on the spliced ​​fusion feature map.

3. The surface defect detection method based on multi-scale adaptive guidance according to claim 2, characterized in that: The multi-scale adaptive guidance sub-network includes: N-2 multi-scale adaptive fusion modules; Each multi-scale adaptive fusion module includes an upsampling layer, a downsampling layer, a first splicing layer, a second splicing layer, a third splicing layer, a first convolution layer, a second convolution layer, an activation layer, a multiplication layer and a first addition layer; The output of the upsampling layer is connected to the input of the first splicing layer, the output of the downsampling layer is connected to the input of the second splicing layer, the output of the first splicing layer is connected to the input of the third splicing layer via the first convolutional layer, the output of the second splicing layer is connected to the input of the third splicing layer via the second convolutional layer, the output of the third splicing layer is connected to the input of the multiplication layer via the activation layer, the output of the multiplication layer is connected to the input of the first addition layer, and the output of the first addition layer is connected to the input of the feature fusion sub-network; The upsampling layer of the m-th multi-scale adaptive fusion module is connected to the output of the m+2-th FasterNet module, the downsampling layer of the m-th multi-scale adaptive fusion module is connected to the output of the m-th FasterNet module, and the inputs of the first splicing layer, the second splicing layer, the multiplication layer, and the first addition layer of the m-th multi-scale adaptive fusion module are also connected to the output of the m+1-th FasterNet module; m = 1, 2, ..., N-2; N represents the number of scale types of feature maps extracted by the multi-scale feature extraction subnetwork.

4. The surface defect detection method based on multi-scale adaptive guidance according to claim 3 is characterized in that: The feature fusion sub-network includes: N-2 first splicing modules connected in sequence, a first full-dimensional dynamic feature extraction module provided between every two adjacent first splicing modules, three second full-dimensional dynamic feature extraction modules connected in sequence, and a second splicing module provided between every two adjacent second full-dimensional dynamic feature extraction modules; The output of the last first splicing module is connected to the input of the first second full-dimensional dynamic feature extraction module; The inputs of the N-2 first splicing modules are connected to the outputs of the N-2 multi-scale adaptive fusion modules in a one-to-one correspondence; The input of the first second splicing module is also connected to the output of the first first full-dimensional dynamic feature extraction module; the input of the second second splicing module and the input of the first first splicing module are also connected to the output of the last FasterNet module in the multi-scale feature extraction subnetwork; The outputs of the three second full-dimensional dynamic feature extraction modules are connected to the inputs of the three detection heads in the detection head sub-network in a one-to-one correspondence.

5. The surface defect detection method based on multi-scale adaptive guidance according to claim 4 is characterized in that: The first full-dimensional dynamic feature extraction module includes: a first CBS layer, a second CBS layer, a segmentation layer, a fourth splicing layer, and multiple ODC Bottleneck modules connected in series; each ODC Bottleneck module includes a second addition layer and multiple full-dimensional dynamic convolution layers; the multiple full-dimensional dynamic convolution layers and the second addition layer are connected in sequence; the input of the second addition layer is also connected to the input of the first full-dimensional dynamic convolution layer; The input of the first CBS layer is connected to the output of the first splicing module to which the input end of the first full-dimensional dynamic feature extraction module belongs; The output of the first CBS layer is connected to the input of the segmentation layer, and the output of the segmentation layer is connected to the input of the fourth splicing layer and the input of the first full-dimensional dynamic convolution layer of the first ODC Bottleneck module. The output of the second addition layer of each ODC Bottleneck module is also connected to the input of the fourth splicing layer, and the output of the fourth splicing layer is connected to the input of the second CBS layer. The output of the second CBS layer is the output of the first full-dimensional dynamic feature extraction module to which it belongs.

6. The surface defect detection method based on multi-scale adaptive guidance according to claim 1, characterized in that: Input the image of the object to be inspected into the pre-trained surface defect detection network to predict the surface defect detection results of the object to be inspected, including: Input the image of the object to be detected into the multi-scale feature extraction sub-network to extract multiple feature maps of different scales; The multi-scale adaptive guidance sub-network is used to adaptively fuse multiple feature maps of different scales to obtain adaptive fused feature maps of different scales; The feature fusion sub-network is used to splice and fuse the adaptive fusion features of different scales to obtain a spliced ​​fusion feature map; The spliced ​​fusion feature map is input into the detection head sub-network, and the surface defect detection results of the object to be detected are output.

7. The surface defect detection method based on multi-scale adaptive guidance according to claim 1, characterized in that: The surface defect detection method based on multi-scale adaptive guidance further includes: a training process of a pre-trained surface defect detection network, specifically: Obtain image samples of the object to be inspected; each image sample corresponds to a real surface defect detection result; Input the image sample of the object to be detected into the initial surface defect detection network and output the surface defect detection sample prediction result; A loss function is constructed, and the loss error is calculated based on the predicted results of surface defect detection samples and the corresponding actual surface defect detection results. The loss function includes a bounding box regression loss term and a classification loss term. A modulation factor that changes the weight coefficient is introduced into the bounding box regression loss term. The initial surface defect detection network is adjusted according to the back propagation of the loss error until the loss error converges or the maximum number of training iterations is reached, thereby obtaining a pre-trained surface defect detection network.

8. The surface defect detection method based on multi-scale adaptive guidance according to claim 7, characterized in that: The expression of the loss function is: ; in, ; ; ; Where L represents the loss error; represents the bounding box regression loss term; Represents the classification loss term; IOU is the intersection-union ratio between the predicted box and the real box; modulation factor It is a hyperparameter with a value range between 0 and 1; Represents the center point of the prediction box The center point of the real frame The Euclidean distance between is the diagonal distance between the minimum bounding rectangle of the predicted box and the real box; Represents the difference between the aspect ratios of the predicted boxes; Represents the difference between the aspect ratios of the real boxes; M represents the number of samples; is the number of categories, is an image sample Corresponding category The true label of is the surface defect detection network for image samples Predicted as class probability.

9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the surface defect detection method based on multi-scale adaptive guidance according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the surface defect detection method based on multi-scale adaptive guidance according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Battery defect detection method and system based on FCS-YOLOv8 algorithm

    CN120219321A

  • Steel plate surface defect detection method based on improved YOLOv8

    CN120374543A

  • Surface defect small target detection method based on multi-scale feature interaction

    CN120374613A

  • Fundus image quality evaluation method and device based on multi-source and multi-scale feature fusion

    US20230274427A1

Cited By

  • Sheet metal part defect detection method and device, terminal and medium

    CN122023430A

  • Sheet metal part defect detection method and device, terminal and medium

    CN122023430B