A method and system for detecting surface defects of objects based on reconstruction network
Through the U-net network of ConvNeXt block and channel attention mechanism, combined with mask processing and multiple loss functions, the accuracy and generalization problems of unsupervised detection models in defect detection are solved, and efficient surface defect detection is achieved.
Patent Information
- Application Number
- CN202310090377.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-09
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2043-02-09
AI Technical Summary
Existing unsupervised surface defect detection models only use normal samples during training, but can still reconstruct defect areas during testing, resulting in reduced detection accuracy. In addition, traditional methods are greatly affected by the environment and lighting, and are difficult to adapt to diverse defects and high-dimensional feature spaces.
The ConvNeXt block is used to build the U-net network, and the channel attention mechanism and mask processing are introduced. Combined with the MSE, SSIM and GMS loss functions, defects are judged by gradient amplitude similarity, the network generalization ability is reduced, and the reconstruction of important channel information is enhanced.
It improves the accuracy of unsupervised detection, is applicable to a variety of industrial environments, reduces the need for manual labeling, and is suitable for surface defect detection of textiles, steel, plastic products and printed materials, improving detection results.
Smart Images

Figure CN116051523B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of quality control, and in particular relates to a method and system for detecting surface defects of an object based on a reconstruction network. Background Art
[0002] Machine vision is widely used in industrial inspection, medical applications and video surveillance technology. Using machine vision to detect surface defects of objects has greatly improved production efficiency and reduced industrial costs.
[0003] In visual inspection, traditional defect detection typically involves a four-step process: image preprocessing, image segmentation, image feature extraction, and surface defect target identification. However, in real-world inspections, the inspection system is often affected by environmental and lighting factors, significantly reducing detection accuracy. Furthermore, defects often manifest in diverse forms, and the types of defects vary significantly across different inspection objects. This results in low-efficiency defect feature extraction, making it difficult to find standard sample images as a reference, significantly complicating inspection. In factory inspections, the data volume is vast, and the dimensionality of the feature space is high, making algorithms incapable of extracting effective defect information from this massive data set. With the advent of deep learning, researchers have proposed deep learning-based defect detection solutions that can better address the challenges presented by traditional detection algorithms. These approaches primarily include supervised and unsupervised approaches. Since surface defect samples are often difficult to obtain in most real-world scenarios, and the types of surface defects often appear uncertain, supervised detection models struggle to generalize to surface defects that the network has never encountered before. To address these issues, unsupervised surface defect detection models are more adaptable to real-world applications. These models use only normal samples (i.e., samples without defects) during training, allowing them to detect surface defects during testing. There are generally two types of unsupervised detection models: those based on feature extraction networks and those based on reconstruction algorithms.
[0004] Among them, the method based on feature extraction network usually does not process the image directly, but extracts image features through CNN pre-trained on large datasets such as ImageNet, establishes the distribution of these normal features, and then obtains the surface defect score by calculating the distance between the test image features and the normal feature distribution, thereby detecting the surface defect area in the image.
[0005] Reconstruction-based surface defect detection methods typically employ autoencoder networks, which are trained by minimizing the reconstruction error of normal samples. Their core assumption is that, since the model is trained entirely on normal samples, it cannot accurately reconstruct the defective areas of the sample under test when a defect is present. The difference between the reconstructed image and the input image is then used as an indicator of the defect severity. Consequently, defect detection accuracy relies entirely on the assumption that only normal samples can be successfully reconstructed, while defective samples fail. Typically, autoencoders or generative adversarial networks (GANs) are used to encode and reconstruct the image. However, a common drawback of reconstruction-based methods is that, due to the strong generalization ability of neural networks, even if the model is trained only on normal samples, when tested with defective samples, the defective areas within the sample can still be reconstructed, thus reducing the ability to distinguish defective areas. To address this issue, some researchers have proposed memory-based autoencoders. This memory module stores image features of normal data and uses the memory of selected normal data to reconstruct the test sample, thereby improving the distinction between normal and surface defect samples. Another approach is to mask some image information and have the network predict the masked areas, thus transforming the reconstruction problem into a repair problem. Although both methods have achieved better results, the fully convolutional neural network (CNN) focuses on the features of local information, and its effect in modeling long-range contextual information will be greatly reduced, which makes it difficult to remove larger defect areas. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to address the deficiencies in the above-mentioned existing technologies and provide a method and system for detecting surface defects of objects based on a reconstruction network, which is used to solve the technical problem of surface defect detection in the surface quality control of items such as textiles, steel, plastic products, and printed products that is not applicable to the existing technology.
[0007] The present invention adopts the following technical solutions:
[0008] A method for detecting surface defects of an object based on a reconstruction network comprises the following steps:
[0009] The ConvNeXt block is used to build a network model. The channel attention mechanism module is introduced into the built network model to obtain a U-net network model based on the ConvNeXt block and the channel attention module.
[0010] Based on the obtained U-net network model, set the loss function; train the U-net network model and save the weight file generated by the training;
[0011] The weight file is loaded into the U-net network model, and then the masked image is input into the U-net network model loaded with the weight file. The anomaly is located according to the anomaly detection function GMS to complete the defect detection.
[0012] Specifically, the network module built using the ConvNeXt block includes a depth-wise separable convolutional layer, a fully connected layer, a Layer Norm, and a GELU activation function, and also introduces a channel attention mechanism module.
[0013] Specifically, the U-net network model based on the ConvNeXt block and the channel attention module is as follows:
[0014] Use 2×2 conv layer and stride 2 for spatial downsampling; use ConvNeXt block to obtain feature maps, add layer normalization LN and activation function GELU where the spatial resolution changes, and realize channel recalibration through channel attention mechanism module while skipping links.
[0015] Furthermore, the channel attention mechanism module is recalibrated and reactivated to adjust the output U to obtain the c-th feature map as follows:
[0016]
[0017] Among them, F s (u c ,s c ) is the channel multiplication between the scale factor vector and the feature map, Refers to the feature map.
[0018] Furthermore, the scaling factor vector s is calculated as follows:
[0019] s=σ(W2·δ(W1·z))
[0020] Among them, δ is the ReLU function, σ is the sigmoid activation function, W1 and W2 represent the weight matrices of two consecutive fully connected layers, and z is the result of performing global pooling on each channel of the input U.
[0021] Specifically, the gradient magnitude similarity GMS is used as the surface defect score of the image, and the loss function during model training is as follows:
[0022]
[0023] in, is the mean square error between the input image and the output result, is the structural similarity error between the input image and the output result, is the gradient amplitude similarity error, is the loss function during model training, λ S =λ G =1.
[0024] Specifically, the training epochs are set to 350, the Adam optimizer is used, and the learning rate is 0.0001 to train the U-net network model.
[0025] Specifically, the image with the mask is:
[0026] Each input image is divided into a grid of dimension N, and the N grids are randomly divided into n unequal sets to obtain n non-overlapping masks. The n masks are multiplied with the image respectively to obtain n images with different masks.
[0027] Furthermore, the total number of grids N is:
[0028]
[0029] Where W and H are the width and height of the image, and k is a hyperparameter that controls the grid size.
[0030] In a second aspect, an embodiment of the present invention provides a surface defect detection system for an object based on a reconstruction network, comprising:
[0031] Build a module and use ConvNeXt block to build a network model; introduce the channel attention mechanism module into the built network model to obtain a U-net network model based on ConvNeXt block and channel attention module;
[0032] The training module sets the loss function based on the U-net network model obtained by building the module; trains the U-net network model and saves the weight file generated by the training;
[0033] The detection module loads the weight file obtained by the training module into the U-net network model, and then inputs the image with the mask into the U-net network model loaded with the weight file. The anomaly is located according to the anomaly detection function GMS to complete the defect detection.
[0034] Compared with the prior art, the present invention has at least the following beneficial effects:
[0035] The present invention discloses a method for detecting surface defects of objects based on a reconstruction network. An autoencoding neural network based on an unsupervised detection mode is designed to solve the problems that in the actual production process, there are more good products than defective samples, and defect types are difficult to obtain, and the forms of defects are numerous and not fixed. In the training stage, the present invention only uses normal sample images as a training set, aiming to train a model that has a good fit for normal samples. In the testing stage, the sample to be tested is input into the network. Since the network has never seen defective samples, the defective area in the defective sample is difficult to reconstruct. The network will output a repaired image, and then by comparing the difference between the test image and the output image, it can be determined whether the sample has defects. The present invention is verified on the large-scale industrial surface defect detection dataset MvTec AD. The verification results show that the present invention has a higher accuracy rate in the detection process and can more effectively ensure the control of product appearance quality in industrial inspection.
[0036] Furthermore, in actual samples, since the area where the defect exists is often related to a large number of pixels around it, there are certain requirements for the long-distance modeling of the model. However, CNN is limited by the size of its receptive field. Considering the lightweight of the model, ConvNeXt Block is used to build the detection network. This can make the network have a larger receptive field, achieve better long-distance modeling effects, and have fewer parameters, making the model more lightweight.
[0037] Furthermore, for all channels in the feature map, in order to strengthen the enhancement of important channel information and suppress unimportant channel information, a channel attention mechanism is adopted. By adding this mechanism to the network, not only can important channels be selectively enhanced and less useful channels be suppressed, but it can also ensure that the network does not simply learn to reconstruct occluded areas based on linear interpolation context information, so as to achieve better recovery of texture and details.
[0038] Furthermore, the output U of the ConvNeXt block is recalibrated using the channel attention mechanism module to obtain the output X, where the c-th feature map of X is represented as The feature map is the result of multiplying the information of each channel by its corresponding channel scale factor. This is done to enhance the information of important channels while reducing the proportion of information of unimportant channels, thereby selecting more useful information.
[0039] Furthermore, the setting of the scaling factor utilizes the characteristics of the back-propagation derivation of the neural network, with the aim of solving the corresponding weight of the importance of each channel so that the information of important channels can be strengthened. At the same time, the application of the sigmoid activation function will prevent the network from reconstructing the input based on a simple linear relationship.
[0040] Furthermore, the loss function adopted is a combination of SSIM similarity, MSE and GMS. This not only uses MSE to perform regression fitting on single-point pixels, but also fits the overall structure of the image at the same time, which greatly reduces the occurrence of image degradation and blurring caused by using only MSE fitting.
[0041] Furthermore, the optimizer used in network training is Adam, which can iteratively update the neural network weights based on the training data. It is a commonly used optimizer in network training. The epochs used in training are 350 and the learning rate is 0.0001. This is the result of multiple experiments on the MvTec dataset. For other datasets, the parameters can be adjusted according to the characteristics of the corresponding dataset.
[0042] Furthermore, due to the strong generalization ability of neural networks, even though the network is trained only on normal samples, it can still generalize to surface defect areas during testing. In other words, surface defect areas can still be reconstructed, significantly reducing the network's accuracy in detecting surface defects. To avoid this, image mask generation was added, sequentially masking different areas of an image. This method enables the network to better learn the characteristic information in the image while avoiding excessive generalization.
[0043] Furthermore, the total number of grids N depends on the hyperparameter k that controls the grid size. This is because when masking, the entire image needs to be randomly masked. Dividing the image into grids can ensure the uniformity of the mask size and the randomness of the masking.
[0044] It can be understood that the beneficial effects of the second aspect mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0045] In summary, the present invention can be effectively applied to product surface defect detection. Compared with traditional template matching, it does not require high-precision image registration technology. Compared with traditional test algorithms, it can be suitable for a variety of industrial environments. Compared with supervised surface defect detection technology, it does not require a lot of manual labeling. Compared with other unsupervised detection models, the network designed by the present invention can better model images, reconstruct better pictures, and achieve better results.
[0046] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is the flow chart of the training phase;
[0048] Figure 2 This is the flow chart for the testing phase;
[0049] Figure 3 The input image and the input image when three non-intersecting random masks are added, where (a) is the input image, (b) is the input image masked by the first random mask, (c) is the input image masked by the second random mask, and (d) is the input image masked by the third random mask;
[0050] Figure 4 Schematic diagram of network composition modules, where (a) is the ConvNeXt block and (b) is the channel attention mechanism module;
[0051] Figure 5 Schematic diagram of the network structure;
[0052] Figure 6 This figure shows the surface defect detection results of the present invention on the MVTec AD dataset. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0054] In the description of the present invention, it is to be understood that the terms “include” and “comprise” indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0055] It should also be understood that the terms used in the present specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0056] It should be further understood that the term "and / or" as used in the present specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items. For example, A and / or B may represent: A alone, A and B simultaneously, or B alone. In addition, the character " / " herein generally indicates that the associated items are in an "or" relationship.
[0057] It should be understood that although the terms "first," "second," and "third" may be used to describe preset ranges in embodiments of the present invention, these preset ranges should not be limited to these terms. These terms are merely used to distinguish one preset range from another. For example, without departing from the scope of embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0058] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0059] The accompanying drawings illustrate various schematic diagrams of structures according to embodiments disclosed herein. These figures are not drawn to scale; for clarity, some details are exaggerated and some details may be omitted. The shapes of the various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are merely exemplary and may deviate in practice due to manufacturing tolerances or technical limitations. Those skilled in the art may design regions / layers with different shapes, sizes, and relative positions as needed.
[0060] The present invention provides a method for detecting surface defects of objects based on a reconstruction network, which adopts an autoencoder network based on ConvNeXtBlock. Compared with the convolution kernel used in the traditional convolutional neural network, it has a larger receptive field and fewer fitting parameters than the Transformer, making the model more lightweight, and most importantly, it can better reconstruct occluded pixels. At the same time, the present invention uses random coverage of the pixels of the input image to reduce the generalization ability of the network, and then passes the feature map obtained by ConvNeXt Block through a channel attention module. The module can learn to use global information to selectively strengthen important channels and suppress less useful channels, and can also ensure that the network does not simply learn to reconstruct occluded areas based on linear interpolation context information. The present invention is suitable for surface quality control of textiles, steel, plastic products, and printed products, and is used to realize surface defect detection for these categories of items.
[0061] See also Figure 1 and Figure 2 The present invention provides a method for detecting surface defects of an object based on a reconstruction network, which includes a training phase and a testing phase. The specific steps are as follows:
[0062] S1. Create an image with a mask;
[0063] The masked image is created as follows:
[0064] Each input image is divided into where W and H are the width and height of the image, and k is a hyperparameter that controls the size of the grid.
[0065] Randomly divide N grids into n unequal sets {i=1, 2, ...n}. Starting from i=1, set the area corresponding to the i-th set to 0, and set the areas corresponding to the remaining n-1 sets to 1.
[0066] Finally, n disjoint masks are obtained, and these n masks are multiplied with the image respectively, that is, n images with different masks are obtained. Figure 3 The result of taking n as 3
[0067] S2, use ConvNeXt block to build the network model;
[0068] In the field of computer vision, Transformer is usually used to expand the receptive field of the convolution kernel. This is because compared with CNN, Transformer can receive global information, that is, the information of each point in the feature map is no longer controlled only by the information in the convolution kernel, but is obtained by the information of each point in the feature map. However, due to its global receptive field, the training of Transformer usually takes longer. Taking this into account, the present invention still uses full convolution. The network constructed by the present invention is built by ConvNeXt block, such as Figure 4 As shown in (a), although it also consists of a pure CNN and convolution kernels, its module design uses large-scale convolution kernels to ensure a larger receptive field while maintaining information validity. Furthermore, the network's larger receptive field facilitates reconstruction of occluded areas using adjacent information. Furthermore, compared to the Transformer, the CNN has fewer parameters and is more lightweight.
[0069] A basic U-net network is built using the ConvNeXt block. In this network, there are four downsampling layers and four upsampling layers. Downsampling is achieved using a 2×2 conv layer with a stride of 2, and upsampling is achieved using a 2x bilinear interpolation. After each downsampling layer or upsampling layer, a normalization function LN and an activation function GELU are added to stabilize training. Then, the ConvNeXt block and skip links are introduced to complete the construction of the entire network.
[0070] S3, introduce the channel attention mechanism module into the network model built in step S2 to obtain a U-net network model based on ConvNeXtblock and channel attention module;
[0071] The feature maps extracted by the ConvNeXt block are processed by a channel attention module, which computes a weighted score for each channel in the feature map. This module is used to selectively enhance important information and suppress less useful channel information by using global information. This module provides a mechanism that not only allows the network to perform feature recalibration but also ensures that the network does not simply learn to reconstruct masked regions based on linearly interpolated contextual information, effectively recovering high-frequency information.
[0072] It is the feature map extracted by ConvNeXt block. Generally speaking, is obtained by performing global pooling on each channel, so that the cth element of z is calculated by the following formula:
[0073]
[0074] Then, the scale factor vector Calculated by the following formula:
[0075] s=σ(W2·δ(W1·z)) (2)
[0076] Among them, δ represents the ReLU function, σ represents the sigmoid activation function, denote the weight matrices of two consecutive fully connected (FC) layers, and r is the dimensionality reduction rate.
[0077] Finally, the c-th feature map in the output X is obtained by recalibrating the reactivated U through the channel attention module.
[0078]
[0079] in, F s (u c ,s c ) refers to the channel multiplication between the scale factor vector and the feature map, Refers to the feature map.
[0080] See also Figure 5The network structure designed for surface defect detection and localization in this paper uses a 2×2 conv layer with a stride of 2 for spatial downsampling. A ConvNeXt block replaces traditional convolutional layers and transformers to obtain feature maps. Furthermore, layer normalization (LN) and activation functions (GELU) are added where spatial resolution changes to help stabilize training. Channel recalibration is then achieved through a channel attention mechanism. Skip links are also added to utilize feature maps of varying scales to better reconstruct detailed textures during training.
[0081] S4. Based on the U-net network model obtained in step S3, set the loss functions including MSE, SSIM and GMS;
[0082] The original and reconstructed images are compared using the pixel-wise MSE loss. To reflect the perceptual differences between the reconstructed and original images, structural similarity (SSIM) and gradient magnitude similarity (GMS) are used to reflect structural differences. Therefore, the final loss function to be minimized is the weighted sum of the three aforementioned losses, as shown in Equation 5. Gradient magnitude similarity is used as the surface defect score for the image.
[0083]
[0084] Among them, λ S =λ G =1.
[0085] S5, training the U-net network model of step S4;
[0086] The training epochs were set to 350, and the Adam optimizer was used with a learning rate of 0.0001 to train the network model.
[0087] S6. Save the model trained in step S5;
[0088] When the model training in step S5 is completed, the weight file generated by the training is saved.
[0089] See also Figure 2 ,The specific steps of the testing phase are as follows:
[0090] S7. In the testing phase, for the test sample, the masked input image is obtained through step S1, the weight file saved in step S6 is loaded into the U-net network model obtained in step S3, and then the masked input image is input into the U-net network model that has been loaded with the weight file. Finally, the anomaly is located according to the anomaly detection function GMS, and the defect detection result is output.
[0091] In another embodiment of the present invention, a reconstruction network-based object surface defect detection system is provided, which can be used to implement the above-mentioned reconstruction network-based object surface defect detection method. Specifically, the reconstruction network-based object surface defect detection system includes a building module, a training module and a detection module.
[0092] Among them, the building module uses ConvNeXt block to build a network model; the channel attention mechanism module is introduced into the built network model to obtain the U-net network model based on ConvNeXt block and channel attention module;
[0093] The training module sets the loss function based on the U-net network model obtained by building the module; trains the U-net network model and saves the weight file generated by the training;
[0094] The detection module loads the weight file obtained by the training module into the U-net network model, and then inputs the image with the mask into the U-net network model loaded with the weight file. The anomaly is located according to the anomaly detection function GMS to complete the defect detection.
[0095] In another embodiment of the present invention, a terminal device is provided, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the object surface defect detection method based on the reconstruction network, including:
[0096] Use ConvNeXt block to build a network model; introduce the channel attention mechanism module into the built network model to obtain a U-net network model based on ConvNeXt block and channel attention module; set the loss function based on the obtained U-net network model; train the U-net network model and save the weight file generated by the training; load the weight file into the U-net network model, and then input the masked image into the U-net network model loaded with the weight file, realize anomaly location according to the anomaly detection function GMS, and complete defect detection.
[0097] In another embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory.
[0098] The processor may load and execute one or more instructions stored in a computer-readable storage medium to implement the corresponding steps of the object surface defect detection method based on the reconstruction network in the above embodiment; the processor may load and execute the following steps:
[0099] Use ConvNeXt block to build a network model; introduce the channel attention mechanism module into the built network model to obtain a U-net network model based on ConvNeXt block and channel attention module; set the loss function based on the obtained U-net network model; train the U-net network model and save the weight file generated by the training; load the weight file into the U-net network model, and then input the masked image into the U-net network model loaded with the weight file, realize anomaly location according to the anomaly detection function GMS, and complete defect detection.
[0100] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0101] The present invention has been verified and proven effective on the public dataset MVTec AD, which is the dataset currently used for mainstream anomaly detection model verification. The MvTec AD dataset contains a total of 15 categories, of which 5 categories are texture data; the remaining 10 categories are object data. 3629 images are used for training and verification, and 1725 images are used for testing, where the training set only contains normal samples. The test set contains 73 different surface defects. For example, surface defects (scratches, dents), structural defects (partial deformation of objects), or defects caused by the missing parts of objects. The AUROC used is used as a result metric for image-level anomaly detection and pixel-level anomaly localization. AUROC: (Area Under the Curve) is defined as the area under the ROC curve, which means the expectation that a uniformly drawn random positive sample ranks before a uniformly drawn random negative sample. AUROC is a value between 0 and 1. When the AUROC value is closer to 1, it means that the classifier can better classify positive and negative samples. The following are the results of the present invention:
[0102] Table 1. Comparison of the surface defect detection results of the proposed method and the existing optimal surface defect detection algorithm based on image completion on the MVTec AD dataset (AUROC%)
[0103]
[0104] Table 2. Comparison of the surface defect segmentation results of this method and the existing optimal surface defect detection algorithm based on image completion on the MVTec AD dataset (AUROC%)
[0105]
[0106] Table 1 describes the comparison of the surface defect detection task results of the method of the present invention and the existing optimal surface defect detection algorithm based on image completion on the dataset, that is, the comparison of the detection results at the image level. It can be found that in the detection results of the texture category, the method proposed in the present invention is 2.8% higher than the existing method, and in the object category, it is 4.9% higher.
[0107] Table 2 describes the comparison of the surface defect segmentation task results of the method of the present invention and the existing optimal surface defect detection algorithm based on image completion on the data set, that is, the comparison of pixel-level detection results. The results show that in different categories, the pixel-level detection task of the present invention has achieved better results.
[0108] See also Figure 6 , shows an example of the results of anomaly detection using the MVTecAD dataset. Rows 1, 3, and 5 are samples with defects, while rows 2, 4, and 6 are schematic diagrams of the anomaly detection results. As can be seen from the figure, for both texture and object categories, the proposed method can effectively detect surface defects. In actual production testing, there is no need to collect defective samples; only normal samples can be used to train the anomaly detection model, making it applicable to a variety of industrial testing environments.
[0109] In summary, the present invention provides a method and system for detecting surface defects of objects based on a reconstruction network, which can be effectively applied to surface defect detection of objects, can be suitable for a variety of industrial environments, does not require a large amount of manual labeling, can better model images, reconstruct better pictures, and achieve better detection results. It is suitable for surface quality control of textiles, steel, plastic products, and printed products, and is used to realize surface defect detection of these categories of items.
[0110] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0111] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0112] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0113] In the embodiments provided by the present invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical functional division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of devices or units, and can be electrical, mechanical, or other forms.
[0114] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0115] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0116] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0117] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0118] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0120] The above content is only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.
Claims
1. A method for detecting surface defects of an object based on a reconstruction network, characterized in that: The following steps are involved: The ConvNeXt block is used to build a network model, including a depthwise separable convolutional layer, a fully connected layer, a Layer Norm, and a GELU activation function. The channel attention mechanism module is also introduced. The U-net network model based on the ConvNeXt block and the channel attention module is obtained by introducing the channel attention mechanism module into the built network model. Specifically: Use 2×2 conv layers with a stride of 2 for spatial downsampling; use ConvNeXt blocks to obtain feature maps, add Layer Norm and GELU activation functions where the spatial resolution changes, and implement channel recalibration through the channel attention mechanism module, while skipping links; Based on the obtained U-net network model, a loss function is set; the U-net network model is trained, and the weight file generated by the training is saved. The gradient magnitude similarity GMS is used as the surface defect score of the image. The loss function during model training is as follows: in, is the mean square error between the input image and the output result, is the structural similarity error between the input image and the output result, is the gradient amplitude similarity error, is the loss function during model training, ; The weight file is loaded into the U-net network model, and then the masked image is input into the U-net network model loaded with the weight file. The anomaly is located according to the anomaly detection function GMS to complete the defect detection.
2. The object surface defect detection method based on reconstruction network according to claim 1 is characterized in that: The output result of the feature map U extracted by the ConvNeXt block after the channel attention mechanism module is recalibrated and reactivated The cth feature map in as follows: in, is the channel multiplication between the scale factor vector and the feature map, Refers to the feature map.
3. The object surface defect detection method based on reconstruction network according to claim 2, characterized in that: Scale factor vector The calculation is as follows: Among them, δ is the ReLU function, σ is the sigmoid activation function, , denote the weight matrices of two consecutive fully connected layers, is the result of performing global pooling on each channel of the input U.
4. The object surface defect detection method based on reconstruction network according to claim 1, characterized in that: The training epochs are set to 350, the Adam optimizer is used, and the learning rate is 0.0001 to train the U-net network model.
5. The object surface defect detection method based on reconstruction network according to claim 1, characterized in that: The image with the mask is specifically: Each input image is divided into The N grids are randomly divided into n unequal sets to obtain n non-overlapping masks, and the n masks are multiplied with the image respectively to obtain n images with different masks.
6. The object surface defect detection method based on reconstruction network according to claim 5, characterized in that: Total number of grids for: Where W and H are the width and height of the image, and k is a hyperparameter that controls the grid size.
7. A surface defect detection system based on a reconstruction network, characterized in that: include: Build a module and use ConvNeXt block to build a network model, including a depth-wise separable convolutional layer, a fully connected layer, Layer Norm, and a GELU activation function. At the same time, introduce a channel attention mechanism module. Introduce the channel attention mechanism module into the built network model to obtain a U-net network model based on ConvNeXt block and channel attention module. Specifically: Use 2×2 conv layers with a stride of 2 for spatial downsampling; use ConvNeXt blocks to obtain feature maps, add layer normalization LN and activation function GELU where the spatial resolution changes, and implement channel recalibration through the channel attention mechanism module, while skipping links; The training module sets the loss function based on the U-net network model obtained by building the module; trains the U-net network model, saves the weight file generated by the training, and uses the gradient magnitude similarity GMS as the surface defect score of the image. The loss function during model training is as follows: in, is the mean square error between the input image and the output result, is the structural similarity error between the input image and the output result, is the gradient amplitude similarity error, is the loss function during model training, ; The detection module loads the weight file obtained by the training module into the U-net network model, and then inputs the image with the mask into the U-net network model loaded with the weight file. The anomaly is located according to the anomaly detection function GMS to complete the defect detection.
8. The object surface defect detection system based on reconstruction network according to claim 7, characterized in that: The c-th feature map in the output X obtained by reactivating the adjusted U after the channel attention mechanism module is recalibrated as follows: in, is the channel multiplication between the scale factor vector and the feature map, refers to the feature map, the scale factor vector vector The calculation is as follows: Among them, δ is the ReLU function, σ is the sigmoid activation function, , denote the weight matrices of two consecutive fully connected layers, is the result of performing global pooling on each channel of the input U.
9. The object surface defect detection system based on reconstruction network according to claim 7, characterized in that: The training epochs are set to 350, the Adam optimizer is used, and the learning rate is 0.0001 to train the U-net network model.
10. The object surface defect detection system based on reconstruction network according to claim 7, characterized in that: The image with the mask is specifically: Each input image is divided into The grids are randomly divided into n unequal sets to obtain n non-overlapping masks. The n masks are multiplied with the image to obtain n images with different masks. The total number of grids is for: Where W and H are the width and height of the image, and k is a hyperparameter that controls the grid size.
Citation Information
Patent Citations
Image processing method and system for CT (Computed Tomography) image of intestinal part
CN115439471A
Product defect detection method and apparatus
WO2022121531A1